Read-level methylation detection
The method classifies DNA regions by methylation states and identifies tumor-informative regions from methylation sequencing data, addressing the lack of effective methods in existing technologies and improving our understanding of biological processes like cancer development.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-10-01
- Publication Date
- 2026-04-09
AI Technical Summary
Existing technologies lack effective methods for classifying DNA regions based on methylation states and identifying tumor-informative regions from methylation sequencing data, which is crucial for understanding complex biological processes like cancer development.
A method is provided for classifying DNA regions based on methylation states by processing signals from methylation sequencing data, identifying tumor-informative regions, and quantifying tumor signals through comparison of methylation states, followed by assigning classifications to samples.
Enables accurate classification of DNA regions and identification of tumor-informative regions, enhancing our understanding of biological processes such as cancer development.
Smart Images

Figure US2025048888_09042026_PF_FP_ABST
Abstract
Description
READ-LEVEL METHYLATION DETECTIONBACKGROUND
[0001] Aspects of the present disclosure relate generally to devices, systems, and methods providing biological or chemical analysis. Various protocols in biological or chemical research involve performing a large number of controlled reactions on local support surfaces or within predefined reaction chambers. The designated reactions may then be observed or detected, and subsequent analysis may help identify or reveal properties of chemicals involved in the reaction. For example, in some multiplex assays, an unknown analyte having an identifiable label (e.g., fluorescent label) may be exposed to thousands of known probes under controlled conditions. Each known probe may be deposited into a corresponding well of a flow cell channel. Observing any chemical reactions that occur between the known probes and the unknown analyte within the wells may help identify or reveal properties of the analyte. Other examples of such protocols include known DNA sequencing processes, such as sequencing-by-synthesis (SBS) or cyclic-array sequencing.
[0002] While a variety of devices, systems, and methods have been made and used to perform biological or chemical analysis, it is believed that no one prior to the inventor(s) has made or used the devices and techniques described herein.
[0003] With the preceding in mind, the various processes in the human body are complex in nature, having both genetic and epigenetic components. By way of example, cancer development is one such complex process, as are processes related to embryonic and fetal development. With respect to epigenetic mechanisms, one such mechanism is methylation, which often occurs via the addition of a methyl group at the 5' carbon of cytosine residues. This methyl addition to cytosine residues often occurs at CpG dinucleotides. Regions of a genome that contain CpG dinucleotides at a higher frequency than expected are often characterized as CpG islands. Such CpG islands may be present in gene promoter regions. Such methylation is generally understood or believed to serve a regulatory or control function with respect to gene expression, with methylation or increased methylation of such CpG islands often being associated with repression of expression ofthe gene in question (i.e., turning the gene “off’). Correspondingly, abnormalities in the methylation associated with epigenetic control processes may be relevant to numerous disorders and disease states, including but not limited to cancer.SUMMARY
[0004] A summary of certain embodiments disclosed herein is set forth below. It should be understood that these aspects are presented merely to provide the reader with a brief summary of these certain embodiments and that these aspects are not intended to limit the scope of this disclosure. Indeed, this disclosure may encompass a variety of aspects that may not be set forth below.
[0005] In accordance with certain embodiments, a method is provided for classifying one or more regions of a DNA sequence based on methylation state. In accordance with this method, signals are accessed or acquired that correspond to a set of methylation sequencing data comprising a plurality of reads within or overlapping with the one or more regions of the DNA sequence. The signal data is processed to classify reads of the plurality of reads based on methylation state. The one or more regions of the DNA sequence are classified based on the methylation state of reads in or overlapping with each respective region.
[0006] In accordance with further embodiments, a method is provided for classifying a sample. In accordance with this method, one or more tumor informative regions are identified from a set of target regions based on respective methylation states of the target regions. The methylation states of respective regions of a sample corresponding to the tumor informative regions are compared. A tumor signal is quantified based on the comparison between the methylation states of respective regions to the tumor informative regions. A classification is assigned to the sample based on the quantified tumor signal.
[0007] Various refinements of the features noted above may exist in relation to various aspects of the present disclosure. Further features may also be incorporated in these various aspects as well. These refinements and additional features may exist individually or in any combination. For instance, various features discussed below in relation to one or more of the illustrated embodiments may be incorporated into any of the above-described aspects of the present disclosure alone or in any combination. The brief summary presented above is intended only to familiarize the readerwith certain aspects and contexts of embodiments of the present disclosure without limitation to the claimed subject matter.BRIEF DESCRIPTION OF THE DRAWINGS
[0008] FIG. 1 depicts a schematic view of an example of a system that may be used to provide biological or chemical analysis, in accordance with aspects of the present disclosure.
[0009] FIG. 2 depicts a cross-sectional view of an example of a flow cell that may be used in the system of FIG. 1, in accordance with aspects of the present disclosure.
[0010] FIG. 3 depicts a schematic view of another example of a system that may be used to provide biological or chemical analysis, in accordance with aspects of the present disclosure.
[0011] FIG. 4 depicts a schematic view of an example of a networked system in which the system of FIG. 1 or 3 may be incorporated, in accordance with aspects of the present disclosure.
[0012] FIG. 5 depicts a schematic view of an example of a base calling arrangement that may be carried out using the system of FIG. 1, 3, or 4, in accordance with aspects of the present disclosure.
[0013] FIG. 6 depicts a schematic view of an example of a base caller training technique that may be implemented using the system of FIG. 1, 3, or 4, in accordance with aspects of the present disclosure.
[0014] FIG. 7 depicts a plot of the extent of methylation sequencing reads based on percentage methylation, in accordance with aspects of the present disclosure.
[0015] FIG. 8 depicts a plot illustrating the overlap of sequencing reads within a target region (e.g., CpG island), in accordance with aspects of the present disclosure.
[0016] FIG. 9 depicts results of a K-means clustering approach (two clusters) to analyzing methylation data to determine separation thresholds, in accordance with aspects of the present disclosure.
[0017] FIG. 10 depicts results of a K-means clustering approach (three clusters) to analyzing methylation data to determine separation thresholds, in accordance with aspects of the presentdisclosure.
[0018] FIG. 11 depicts further results of a K-means clustering approach (three clusters) to analyzing methylation data to determine separation thresholds, in accordance with aspects of the present disclosure.
[0019] FIG. 12 depicts results of a K-means clustering approach (four clusters) to analyzing methylation data to determine separation thresholds, in accordance with aspects of the present disclosure.
[0020] FIG. 13 depicts a plot illustrating the effect of data filtration based on the number of CpGs per read, in accordance with aspects of the present disclosure.
[0021] FIG. 14 depicts a plot illustrating the effect of data filtration based on percentage read overlap with a CpG island, in accordance with aspects of the present disclosure.
[0022] FIG. 15 depicts a 3D plot depicting methylated, unmethylated, and semimethylated read frequency, in accordance with aspects of the present disclosure.
[0023] FIG. 16 depicts a frequency plot of the percentage of methylated, unmethylated, and semimethylated reads in an island, in accordance with aspects of the present disclosure.
[0024] FIG. 17 depicts each CpG island visually shaded based on whether the islands are methylated or unmethylated in a single sample, in accordance with aspects of the present disclosure.
[0025] FIG. 18 depicts results of using a tumor sample to identify tumor-informative CpG islands, in accordance with aspects of the present disclosure.
[0026] FIG. 19 depicts curves generated by simulation of titration results across a range from 0 % to 100% tumor, in accordance with aspects of the present disclosure.
[0027] FIG. 20 depicts read-level methylation analysis signal intensity plotted with base-level analysis signal intensity with respect to the titrated tumor percentages, in accordance with aspects of the present disclosure.
[0028] FIG. 21 depicts a process flow of steps that may be performed as part of read-levelmethylation analysis, in accordance with aspects of the present disclosure.DET AILED DESCRIPTION
[0029] The following detailed description of certain examples will be better understood when read in conjunction with the appended drawings. To the extent that the figures illustrate diagrams of the functional blocks of various examples, the functional blocks are not necessarily indicative of the division between hardware components. Thus, for example, one or more of the functional blocks (e g., processors or memories) may be implemented in a single piece of hardware (e.g., a general -purpose signal processor or random-access memory, hard disk, or the like). Similarly, the programs may be stand-alone programs, may be incorporated as subroutines in an operating system, may be functions in an installed software package, and the like. It should be understood that the various examples are not limited to the arrangements and instrumentality shown in the drawings.
[0030] I. Overview of System for Biological or Chemical Analysis
[0031] Examples described herein may be used in various biological or chemical processes and systems for academic analysis, commercial analysis, or other analysis. More specifically, examples described herein may be used in various processes and systems where it is desired to detect an event, property, quality, or characteristic that is indicative of a designated reaction (e.g., methylation). Bioassay systems such as those described herein may be configured to perform a plurality of designated reactions that may be detected individually or collectively. For example, bioassay systems may be used to sequence a dense array of nucleic acid features through iterative cycles of enzymatic manipulation and image acquisition. In some examples, nucleic acids can be attached to a surface and amplified. Examples of such amplification are described in U.S. Pat. No. 7,741,463, entitled “Method of Preparing Libraries of Template Polynucleotides,” issued June 22, 2010, the disclosure of which is incorporated by reference herein, in its entirety; and / or U.S. Pat. No. 7,270,981, entitled “Recombinase Polymerase Amplification,” issued September 18, 2007, the disclosure of which is incorporated by reference herein, in its entirety.
[0032] Components that are used in the bioassay systems may include one or more microfluidic channels that deliver reagents or other reaction components to a reaction site. Thereaction sites may be randomly distributed across a substantially planar surface; or may be patterned across a substantially planar surface. Each of the reaction sites may be imaged to detect light from the reaction site. The signals indicating photons emitted from the reaction sites and detected by image sensors may provide illumination values. These illumination values may be combined into an image indicating photons as detected from the reaction sites. These images may be further analyzed to identify compositions, reactions, conditions, etc., at each reaction site.
[0033] II. Examples of Fluidics Devices and Fluid Flow Paths - Example of System with Higher Volume Throughput
[0034] FIG. 1 illustrates a schematic diagram of an example of a system (100) that may be used to perform an analysis on one or more samples of interest. In some implementations, the sample may include one or more clusters of nucleotides (e.g., DNA) that have been linearized to form a single stranded DNA (sstDNA). In the implementation shown, system (100) is configured to receive a flow cell cartridge assembly (102) including a flow cell assembly (103) and a sample cartridge (104). System (100) includes a flow cell receptacle (122) that receives flow cell cartridge assembly (102), a vacuum chuck (124) that supports flow cell assembly (103), and a flow cell interface (126) that is used to establish a fluidic coupling between system (100) and flow cell assembly (103). Flow cell interface (126) may include one or more manifolds. System (100) further includes a sipper manifold assembly (106), a sample loading manifold assembly (108), and a pump manifold assembly (110). System (100) also includes a drive assembly (112), a controller (114), an imaging system (116), and a waste reservoir (118). Controller (114) is electrically and / or communicatively coupled to drive assembly (112) and to imaging system (116); and is configured to cause drive assembly (112) and / or the imaging system (116) to perform various functions as disclosed herein.
[0035] In the present example, flow cell assembly (103) includes a flow cell (128) having a channel (130) and defining a plurality of first openings (132), which are fluidically coupled to the channel (130) and arranged on a first side (134) of the channel (130). Flow cell (128) further includes a plurality of second openings (136) fluidically coupled to the channel (130) and arranged on a second side (138) of the channel (130). Fluid may thus flow through flow cell (128) via channel. While the flow cell (128) is shown including one channel (130), flow cell (128) may include two or more channels (130). Flow cell assembly (103) also includes a flow cell manifoldassembly (140) coupled to flow cell (128) and having a first manifold fluidic line (142) and a second manifold fluidic line (144). Flow cell manifold assembly (140) may be in the form of a laminate including a plurality of layers as discussed in more detail below.
[0036] In the implementation shown, first manifold fluidic line (142) has a first fluidic line opening (146) and is fluidically coupled to each of the first openings (132) of flow cell (128); and second manifold fluidic line (144) has a second fluidic line opening (148) and is fluidically coupled to each of the second openings (136). As shown, flow cell assembly (103) includes gaskets (150) coupled to flow cell manifold assembly (140) and fluidically coupled to fluidic line openings (146, 148). In some implementations where flow cell (128) includes a plurality of channels (130), flow cell manifold assembly (140) may include additional fluidic lines (152) that couple first fluidic line openings (146) to a single manifold port (154). In such implementations, a single gasket (150) may be coupled to flow cell manifold assembly (140) that surrounds the manifold port (154) and is in fluidic communication with a plurality of channels (130). In operation, flow cell interface (126) engages with corresponding gaskets (150) to establish a fluidic coupling between system (100) and flow cell (128). The engagement between flow cell interface (126) and gaskets (150) reduces or eliminates fluid leakage between flow cell interface (126) and flow cell (128).
[0037] In the implementation shown, first manifold fluidic line (142) has a portion (156) that is substantially parallel to a longitudinal axis (158) of channel (130); and second manifold fluidic line (144) has a portion (160) that is substantially parallel to longitudinal axis (158) of channel (130). Additionally, first manifold fluidic line (142) is shown being at least partially adjacent a first end (162) of flow cell (128) and spaced from a second end (164) of flow cell (128); and second manifold fluidic line (144) is shown being at least partially adjacent second end (164) of flow cell (128) and spaced from first end (162). Other arrangements of manifold fluidic lines (142, 144) may prove suitable, however.
[0038] In the implementation shown, system (100) includes a sample cartridge receptacle (166) that receives sample cartridge (104) that carries one or more samples of interest (e.g., an analyte). System (100) also includes a sample cartridge interface (168) that establishes a fluidic connection with sample cartridge (104). Sample loading manifold assembly (108) includes one or more sample valves (170). Pump manifold assembly (110) includes one or more pumps (172), one or more pump valves (174), and a cache (176). Valves (170, 174) and pumps (172) may takeany suitable form. Cache (176) may include a serpentine cache and may temporarily store one or more reaction components during, for example, bypass manipulations of the system (100). While cache (176) is shown being included in pump manifold assembly (110), cache (176) may alternatively be located elsewhere (e.g., in sipper manifold assembly (106) or in another manifold downstream of a bypass fluidic line (178), etc ).
[0039] Sample loading manifold assembly (108) and pump manifold assembly (110) flow one or more samples of interest from sample cartridge (104) through a fluidic line (180) toward flow cell cartridge assembly (102). In some implementations, sample loading manifold assembly (108) may individually load or address each channel (130) of flow cell (128) with a respective sample of interest. The process of loading channel (130) with a sample of interest may occur automatically using system (100). As shown in FIG. 1, sample cartridge (104) and sample loading manifold assembly (108) are positioned downstream of flow cell cartridge assembly (102). In the implementation shown, sample loading manifold assembly (108) is coupled between flow cell cartridge assembly (102) and pump manifold assembly (110). To draw a sample of interest from sample cartridge (104) and toward pump manifold assembly (110), sample valves (170), pump valves (174), and / or pumps (172) may be selectively actuated to urge the sample of interest toward pump manifold assembly (110). Sample cartridge (104) may include a plurality of sample reservoirs that are selectively fluidically accessible via the corresponding sample valves (170). To individually flow the sample of interest toward channel (130) of flow cell (128) and away from pump manifold assembly (110), sample valves (170), pump valves (174), and / or pumps (172) may be selectively actuated to urge the sample of interest toward flow cell cartridge assembly (102) and into respective channels (130) of flow cell (128).
[0040] Drive assembly (112) interfaces with sipper manifold assembly (106) and pump manifold assembly (110) to flow one or more reagents that interact with the sample within flow cell (128). In some scenarios, a reversible terminator is attached to the reagent to allow a single nucleotide to be incorporated onto a growing DNA strand. In some such implementations, one or more of the nucleotides has a unique fluorescent label that emits a color when excited. The color (or absence thereof) is used to detect the corresponding nucleotide. In the implementation shown, imaging system (116) excites one or more of the identifiable labels (e.g., a fluorescent label) and thereafter obtains image data for the identifiable labels. The labels may be excited by incident lightand / or a laser and the image data may include one or more colors emitted by the respective labels in response to the excitation. The image data (e.g., detection data) may be analyzed by system (100). Examples of features and functionalities that may be incorporated into imaging system (116) will be described in greater detail below.
[0041] After the image data is obtained, drive assembly (112) interfaces with sipper manifold assembly (106) and pump manifold assembly (110) to flow another reaction component (e.g., a reagent) through flow cell (128) that is thereafter received by waste reservoir (118) via a primary waste fluidic line (182) and / or otherwise exhausted by system (100). Some reaction components may perform a flushing operation that chemically cleaves the fluorescent label and the reversible terminator from the sstDNA. The sstDNA may then be ready for another cycle.
[0042] The primary waste fluidic line (182) is coupled between pump manifold assembly (110) and waste reservoir (118). In some implementations, pumps (172) and / or pump valves (174) of pump manifold assembly (110) selectively flow the reaction components from flow cell cartridge assembly (102), through fluidic line (180) and sample loading manifold assembly (108) to primary waste fluidic line (182). Flow cell cartridge assembly (102) is coupled to a central valve (184) via flow cell interface (126). Central valve (184) is coupled with flow cell interface (126) via a fluidic line (185). An auxiliary waste fluidic line (186) is coupled to central valve (184) and to waste reservoir (118). In some implementations, auxiliary waste fluidic line (186) receives excess fluid of a sample of interest from flow cell cartridge assembly (102), via central valve (184), and flows the excess fluid of the sample of interest to waste reservoir (118) when back loading the sample of interest into flow cell (128), as described herein.
[0043] Sipper manifold assembly (106) includes a shared line valve (188) and a bypass valve (190). Shared line valve (188) may be referred to as a reagent selector valve. Central valve (184) and the valves (188, 190) of sipper manifold assembly (106) may be selectively actuated to control the flow of fluid through fluidic lines (192, 194, 196). Sipper manifold assembly (106) may be coupled to a corresponding number of reagent reservoirs (198) via reagent sippers (200). Reagent reservoirs (198) may contain fluid (e.g., reagent and / or another reaction component). In some implementations, sipper manifold assembly (106) includes a plurality of ports. Each port of sipper manifold assembly (106) may receive one of the reagent sippers (200). Reagent sippers (200) may be referred to as fluidic lines. Some forms of reagent sippers (200) may include an array of sippertubes extending downwardly along the z-dimension from ports in the body of sipper manifold assembly (106). Reagent reservoirs (198) may be provided in a cartridge, and the tubes of reagent sippers (200) may be configured to be inserted into corresponding reagent reservoirs (198) in the reagent cartridge so that liquid reagent may be drawn from each reagent reservoir (198) into the sipper manifold assembly (106).
[0044] Shared line valve (188) of sipper manifold assembly (106) is coupled to central valve (184) via shared reagent fluidic line (196). Different reagents may flow through shared reagent fluidic line (196) at different times. In some versions, when performing a flushing operation before changing between one reagent and another, pump manifold assembly (110) may draw wash buffer through shared reagent fluidic line (196), central valve (184), and flow cell cartridge assembly (102).
[0045] Bypass valve (190) of sipper manifold assembly (106) is coupled to central valve (184) via dedicated reagent fluidic lines (194, 196). Each of the dedicated reagent fluidic lines (194, 196) may be associated with a single reagent. The fluids that may flow through dedicated reagent fluidic lines (194, 196) may be used during sequencing operations and may include a cleave reagent, an incorporation reagent, a scan reagent, a cleave wash, and / or a wash buffer.
[0046] Bypass valve (190) is also coupled to cache (176) of pump manifold assembly (110) via bypass fluidic line (178). One or more reagent priming operations, hydration operations, mixing operations, and / or transfer operations may be performed using bypass fluidic line (178). The priming operations, the hydration operations, the mixing operations, and / or the transfer operations may be performed independent of flow cell cartridge assembly (102). Thus, the operations using bypass fluidic line (178) may occur during, for example, incubation of one or more samples of interest within flow cell cartridge assembly (102). That is, shared line valve (188) may be utilized independently of bypass valve (190) such that bypass valve (190) may utilize bypass fluidic line (178) and / or cache (176) to perform one or more operations while shared line valve (188) and / or central valve (184) simultaneously, substantially simultaneously, or offset synchronously perform other operations.
[0047] Drive assembly (112) includes a pump drive assembly (202) and a valve drive assembly (204). Pump drive assembly (202) may be adapted to interface with one or more pumps (172) topump fluid through flow cell (128) and / or to load one or more samples of interest into flow cell (128). Valve drive assembly (204) may be adapted to interface with one or more of the valves (170, 174, 184, 188, 190) to control the position of the corresponding valves (170, 174, 184, 188, 190).
[0048] Controller (114) of the present example includes a user interface (206), a communication interface (208), one or more processors (210), and a memory (212) storing instructions executable by the one or more processors (210) to perform various functions including the disclosed implementations. User interface (206), communication interface (133), and memory (212) are electrically and / or communicatively coupled to the one or more processors (210). User interface (206) may be adapted to receive input from a user and to provide information to the user associated with the operation of system (100) and / or an analysis taking place. User interface (206) may include a touch screen, a display, a keyboard, a speaker(s), a mouse, a track ball, and / or a voice recognition system.
[0049] Communication interface (208) is adapted to enable communication between system (100) and a remote system(s) (e.g., computers) via a network(s) (e.g., the Internet, an intranet, a local-area network (LAN), a wide-area network (WAN), a coaxial-cable network, a wireless network, a wired network, a satellite network, a digital subscriber line (DSL) network, a cellular network, a Bluetooth connection, a near field communication (NFC) connection, etc.). Some of the communications provided to the remote system may be associated with analysis results, imaging data, etc. generated or otherwise obtained by system (100). Some of the communications provided to system (100) may be associated with a fluidics analysis operation, patient records, and / or a protocol(s) to be executed by system (100).
[0050] The one or more processors (210) and / or system (100) may include one or more of a processor-based system(s) or a microprocessor-based system(s). In some implementations, the one or more processors (210) and / or system (100) includes one or more of a programmable processor, a programmable controller, a microprocessor, a microcontroller, a graphics processing unit (GPU), a digital signal processor (DSP), a reduced-instruction set computer (RISC), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a field programmable logic device (FPLD), a logic circuit, and / or another logic-based device executing various functions including the ones described herein.
[0051] Memory (212) may include one or more of a semiconductor memory, a magnetically readable memory, an optical memory, a hard disk drive (HDD), an optical storage drive, a solid- state storage device, a solid-state drive (SSD), a flash memory, a read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable readonly memory (EEPROM), a random-access memory (RAM), a non-volatile RAM (NVRAM) memory, a compact disc (CD), a compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a Blu-ray disk, a redundant array of independent disks (RAID) system, a cache and / or any other storage device or storage disk in which information is stored for any duration (e.g., permanently, temporarily, for extended periods of time, for buffering, for caching).
[0052] III. Examples of Flow Cell Structures
[0053] As noted above, a system (100) may execute reactions in a flow cell (128) and / or perform analysis on one or more samples of interest in a flow cell (128). The following describes examples of forms that such flow cells (128) may take, it being understood that flow cells (128) may take various other forms and have various other features in addition to or in lieu of the features described below.
[0054] Example of Single-Surface Patterned Flow Cell
[0055] FIG. 2 shows an example of a flow cell (400) that includes a patterned substrate (402), which includes depressions (404) separated by interstitial regions (406), and surface chemistry (410, 412) positioned in the depressions (404). Depressions (404) may be in the form of microwells or nanowells. Depressions (404) may be configured to contain nucleic acid strands or other oligonucleotides and thereby provide a reaction site for SBS and / or for other kinds of processes. In some versions, each depression (404) has a cylindraceous configuration, with a generally circular cross-sectional profile. In some other versions, each depression (404) has a polygonal (e.g., hexagonal, octagonal, square, rectangular, elliptical, etc.) cross-sectional profile. Alternatively, depressions (404) may have any other suitable configuration. It should also be understood that depressions (404) may be arranged in any suitable pattern, including but not limited to a grid pattern.
[0056] Surface chemistry (410, 412) of the present example includes functionalized coating layer (410) and primers (412). While not shown, it is to be understood that the depressions(404) may also have surface preparation or treatment chemistry (e g., silane or a silane derivative) positioned between the substrate (402) and the functionalized coating layer (410). This same surface preparation or treatment chemistry may also be positioned on the interstitial regions (406). In the present example, a hydrogel (440) is applied before lid (420) is bonded to substrate (402). Hydrogel (440) covers surface chemistry (410, 412) in depressions (404), and at least a portion of the patterned substrate (402) (e.g., those interstitial regions (406) that are not also bonding regions (422)). By way of example only, hydrogel (440) may comprise PAZAM, crosslinked polyacrylamide, agarose gel, etc.
[0057] Flow cell (400) of this example further includes a lid (420) bonded to bonding region(s) (422) of patterned substrate (402). In the example shown in FIG. 2, lid (420) includes a top portion (424) that is connected to several sidewalls (426), and these components (424, 426) define a portion of each of the six flow channels (430A, 430B, 430C, 430D, 430E, 430F). The respective sidewalls (426) isolate one flow channel (430A, 430B, 430C, 430D, 430E, 430F) from each adjacent flow channel (430A, 430B, 430C, 430D, 430E, 430F). Each flow channel (430A, 430B, 430C, 430D, 430E, 430F) is in selective fluid communication with a respective set of depressions (404).
[0058] Lid (420) may be bonded to bonding region (422) of substrate (402) using any suitable technique, such as laser bonding, diffusion bonding, anodic bonding, eutectic bonding, plasma activation bonding, glass frit bonding, or other methods known in the art. In some versions, a spacer layer (428) may be used to bond lid (420) to bonding region (422). Spacer layer (428) may comprise any material that will seal at least some of interstitial regions (404) (e.g., bonding region (422)) of substrate (402) and lid (420) together. While not shown, lid (420) or the patterned substrate (402) may include inlet and outlet ports that are to fluidically engage other ports (not shown), such as those of sample cartridge interface (168), for directing fluid(s) into the respective flow channels (430A, 430B, 430C, 430D, 430E, 430F) (e.g., from a reagent cartridge or other fluid storage system) and out of the flow channel (e.g., to waste reservoir (118) or another waste removal system). Flow channels (430A, 430B, 430C, 430D, 430E, 430F) may serve to, for example, selectively introduce reaction components or reactants to hydrogel (440) and the underlying surface chemistry (410, 412) in order initiate designated reactions in / at depressions (404).
[0059] While flow cell (400) includes a pattern of depressions (404) to provide an array of reaction sites, other variations may provide reaction sites on or at various other kinds of structural features, including but not limited to continuously planer surfaces and / or protruding surfaces, etc. By way of further example only, flow cell (400) may be constructed and operable in accordance with at least some of the teachings of U.S. Pat. No. 10,919,033, entitled “Flow Cells with Hydrogel Coating,” issued February 16, 2021, the disclosure of which is incorporated by reference herein, in its entirety.
[0060] IV. Examples of Imaging System Features
[0061] As noted above, system (100) includes an imaging system (116) that excites one or more identifiable labels (e.g., a fluorescent label) in samples in reaction sites provided by depressions (404, 462, 464) of a flow cell (128, 400); and thereafter obtains image data for the identifiable labels. This image data is used to identify nucleotides as part of a nucleic acid sequencing process. Alternatively, the image data may be used for various other purposes. The following description provides details on how some versions of imaging system (116) may be configured and operable.
[0062] FIG. 3 illustrates a schematic diagram of another example of a system (500) that may be used to perform an analysis on one or more samples of interest. Except as otherwise described below, system (500) of this example may be configured and operable like systems (100) described above. System (500) is configured to perform a large number of parallel reactions within a flow cell (510). Flow cell (510) may be configured and operable like flow cells (400) described above or may have any other suitable configuration. Flow cell (510) may thus include one or more flow channels that receive a solution from system (500) and direct the solution toward reaction sites of flow cell (510).
[0063] System (500) includes a system controller (520) that may communicate with the various components, assemblies, and sub-systems of the system (500). Controller (520) may be configured and operable like controllers (114) described above. An imaging assembly (522) of system (500) includes a light emitting assembly (550) that emits light that reaches reaction sites on flow cell (510). Light emitting assembly (550) may include an incoherent light emitter (e.g., emit light beams output by one or more excitation diodes), or a coherent light emitter such asemitter of light output by one or more lasers or laser diodes. In some implementations, light emitting assembly (550) may include a plurality of different light sources (not shown), each light source emitting light of a different wavelength range. Some versions of light emitting assembly (550) may also include one or more collimating lenses (not shown), a light structuring optical assembly (not shown), a projection lens (not shown) that is operable to adjust a structured beam shape and path, epifluorescence microscopy components, and / or other components. Although system (500) is illustrated as having a single light emitting assembly (550), multiple light emitting assemblies (550) may be included in some other implementations.
[0064] In the present example, the light from light emitting assembly (550) is directed by dichroic mirror assembly (546) through an objective lens assembly (542) onto a sample of a flow cell (510), which is positioned on a motion stage (570). In the case of fluorescent microscopy of a sample, a fluorescent element associated with the sample of interest fluoresces in response to the excitation light, and the resultant light is collected by objective lens assembly (542) and is directed to an image sensor of camera system (540) to detect the emitted fluorescence. In some implementations, a tube lens assembly may be positioned between the objective lens assembly (542) and the dichroic mirror assembly (546) or between the dichroic mirror (546) and the image sensor of the camera system (540). A moveable lens element may be translatable along a longitudinal axis of the tube lens assembly to account for focusing on an upper interior surface or lower interior surface of the flow cell (510) and / or spherical aberration introduced by movement of the objective lens assembly (542).
[0065] In the present example, a fdter switching assembly (544) is interposed between dichroic mirror assembly (546) and camera system (540). Filter switching assembly (544) includes one or more emission filters that may be used to pass through particular ranges of emission wavelengths and block (or reflect) other ranges of emission wavelengths. For example, emission filters may be used to direct different wavelength ranges of emitted light to different image sensors of the camera system (540) of imaging assembly (522). For instance, the emission filters may be implemented as dichroic mirrors that direct emission light of different wavelengths from flow cell (510) to different image sensors of camera system (540). In some variations, a projection lens is interposed between filter switching assembly (544) and camera system (540). Filter switching assembly (544) may be omitted in some versions.
[0066] System (500) further includes a fluid delivery assembly (590) that may direct the flow of reagents (e.g., fluorescently labeled nucleotides, buffers, enzymes, cleavage reagents, etc.) to (and through) flow cell (510) and waste valve (580). Fluid delivery assembly (590) may be configured and operable like the various fluid delivery components described herein. System (500) of the present example also includes a temperature station actuator (530) and heater / cooler (532) that may optionally regulate the temperature of conditions of the fluids within the flow cell (510). In some implementations, the heater / cooler (532) may be fixed to sample stage (570), upon which the flow cell (510) is placed, and / or may be integrated into sample stage (570).
[0067] Flow cell (510) may be removably mounted on sample stage (570), which may provide movement and alignment of flow cell (510) relative to objective lens assembly (542). Sample stage (570) may have one or more actuators to allow sample stage (570) to move in any of three dimensions. For example, actuators may be provided to allow sample stage (570) to move in the x, y, and z directions relative to objective lens assembly (542), tilt relative to objective lens assembly (542), and / or otherwise move relative to objective lens assembly (542). Movement of sample stage (570) may allow one or more sample locations on flow cell (510) to be positioned in optical alignment with objective lens assembly (542). Movement of sample stage (570) relative to objective lens assembly (542) may be achieved by moving sample stage (570) itself, by moving objective lens assembly (542), by moving some other component of imaging assembly (522), by moving some other component of system (500), or any combination of the foregoing. For instance, in some implementations, the sample stage (570) may be actuatable in the x and y directions relative to the objective lens assembly (542) while a focus component (562) or z-stage may move the objective lens assembly (542) along the z direction relative to the sample stage (570).
[0068] In some implementations, a focus component (562) may be included to control positioning of one or more elements of objective lens assembly (542) relative to the flow cell (510) in the focus direction (e.g., along the z-axis or z-dimension). Focus component (562) may include one or more actuators physically coupled to the objective lens assembly (542), the optical stage, the sample stage (570), or a combination thereof, to move flow cell (510) on sample stage (570) relative to the objective lens assembly (542) to provide proper focusing for the imaging operation. In the present example, the focus component (562) utilizes a focus tracking module (560) that is configured to detect a displacement of the objective lens assembly (542) relative to a portion ofthe flow cell (510) and output data indicative of an in-focus position to the focus component (562) or a component thereof or operable to control the focus component (562), such as controller (520), to move the objective lens assembly (542) to position the corresponding portion of the flow cell (510) in focus of the objective lens assembly (542).
[0069] In some implementations, an actuator of focus component (562) or for sample stage (570) may be physically coupled to objective lens assembly (542), the optical stage, sample stage (570), or a combination thereof, such as, for example, by mechanical, magnetic, fluidic, or other attachment or contact directly or indirectly to or with the stage or a component thereof. The actuator of focus component (562) may be configured to move objective lens assembly (542) in the z-direction while maintaining sample stage (570) in the same plane (e.g., maintaining a level or horizontal attitude, perpendicular to the optical axis). In some implementations, sample stage (570) includes an x direction actuator and a y direction actuator to form an x-y stage. Sample stage (570) may also be configured to include one or more tip or tilt actuators to tip or tilt sample stage (570) and / or a portion thereof, to account for any slope in its surfaces.
[0070] Camera system (540) may include one or more image sensors to monitor and track the imaging (e g., sequencing) of flow cell (510). Camera system (540) may be implemented, for example, as a CCD or CMOS image sensor camera, but other image sensor technologies (e.g., active pixel sensor) may be used. By way of further example only, camera system (540) may include a dual-sensor time delay integration (TDI) camera, a single-sensor camera, a camera with one or more two-dimensional image sensors, and / or other kinds of camera technologies. While camera system (540) and associated optical components are shown as being positioned above flow cell (510) in FIG. 3, one or more image sensors or other camera components may be incorporated into system (500) in numerous other ways as will be apparent to those skilled in the art in view of the teachings herein. For instance, one or more image sensors may be positioned under flow cell (510), such as within the sample stage (570) or below the sample stage (570); or may even be integrated into flow cell (510).
[0071] V. Examples of Data Processing Features
[0072] A. Example of Networked Data Processing Arrangement
[0073] As noted above, a system (100, 500) may include a controller (114, 520) that isconfigured to process data, execute algorithms, etc., as needed to perform a sequencing operation or other kind of operation. In some scenarios, system (100, 500) may be coupled with other devices via a network to perform further data processing, data storage, execution of algorithms, etc. FIG. 4 shows an example of such an arrangement. In particular, FIG. 4 shows a networked system (800) that includes a sequencing device (810), a server device (820), a client device (830), and a local device (840), with all devices (810, 820, 830, 840) being coupled together via a network (850). Network (850) may take any suitable form as will be apparent to those skilled in the art in view of the teachings herein.
[0074] As shown in FIG. 4, sequencing device (810) comprises a computing device and a sequencing device system (812) for sequencing a genomic sample or other nucleic-acid polymer. In some versions, by executing sequencing device system (812) using a processor, sequencing device (810) analyzes nucleotide fragments or oligonucleotides extracted from genomic samples to generate nucleotide reads or other data utilizing computer implemented methods and systems either directly or indirectly on sequencing device (810). More particularly, sequencing device (810) receives nucleotide- sample slides (e.g., flow cells (128, 400, 510)) comprising nucleotide fragments extracted from samples and further copies and determines the nucleobase sequence of such extracted nucleotide fragments. It should be understood that sequencing device (810) may represent a version of systems (100, 500) described above.
[0075] In some versions, the sequencing device (810) utilizes SBS to sequence nucleotide fragments into nucleotide reads and determine nucleobase calls for the nucleotide reads. By executing sequencing device system (812), sequencing device (810) may further store the nucleobase calls as part of base-call data that is formatted as a binary base call (BCL) file and send the BCL file to the local device (840) and / or the server device(s) (820). Sequencing device (810) may communicate the BCL file and / or other data to local device (840) and / or client device (830) via network (850) or directly (i.e., bypassing network (850)).
[0076] In some scenarios, local device (840) is located at or near a same physical location of sequencing device (810). For instance, local device (840) and sequencing device (810) may be integrated into a single computing device. Local device (840) may run sequencing system (814) to generate, receive, analyze, store, and transmit digital data, such as by receiving base-call data or determining variant calls based on analyzing such base-call data. By executing software in theform of sequencing system (814), local device (840) may align nucleotide reads with a structural variation graph genome (824) and determine genetic variants based on the aligned nucleotide reads. Local device (840) may also send data to client device (830), including a variant call fde (VCF) or other information indicating nucleobase calls, sequencing metrics, error data, or other metrics.
[0077] Server device(s) (820) may be located remotely from the local device (840) and sequencing device (810). Server device(s) (820) may comprise a distributed collection of servers, where server device(s) (820) include a number of server devices distributed across network (850) and located in the same or different physical locations. Similar to local device (840), server device(s) (820) may include a version of sequencing system (814). Accordingly, server device(s) (820) may generate, receive, analyze, store, and transmit digital data, such as by receiving basecall data or determining variant calls based on analyzing such base-call data. As indicated above, sequencing device (810) may send (and server device(s) (820) may receive) base-call data from sequencing device (810). Server device(s) (820) may also send data to client device (830), including VCFs or other sequencing related information.
[0078] As indicated above, as part of server device(s) (820) or local device (840), sequencing system (814) may generate or implement a structural variation graph genome with alternate contiguous sequences representing structural variant haplotypes. For instance, system (814) may identify candidate structural variants of a threshold frequency (or that otherwise satisfy another occurrence threshold) within a genomic sample database. From among the candidate structural variants, sequencing system (814) selects structural variant haplotypes based on one or both of satisfying another occurrence threshold and finding flanking variants adjacent to particular structural variant haplotypes. Sequencing system (814) may likewise select reference haplotypes of genomic regions corresponding to the selected structural variant haplotypes from a reference genome. Based on the selected haplotypes, sequencing system (814) generates a structural variation graph genome comprising both alternate contiguous sequences representing the structural variant haplotypes and reference sequences representing the reference haplotypes. Based on comparing nucleotide reads of a genomic sample with alternate contiguous sequences representing structural variant haplotypes, sequencing system (814) can determine nucleobase calls for the genomic sample.
[0079] By executing a sequencing application (832), client device (830) may generate, store, receive, and send digital data. In particular, client device (830) may receive sequencing data from local device (840) or receive call files (e.g., BCL) and sequencing metrics from sequencing device (810). Furthermore, client device (830) may communicate with local device (840) or server device(s) (820) to receive a VCF comprising nucleobase calls and / or other metrics, such as a base- call-quality metrics or pass-filter metrics. Client device (830) may accordingly present or display information pertaining to variant calls or other nucleobase calls within a graphical user interface of sequencing application (832) to a user associated with client device (830). For example, client device (830) may present structural variant calls and / or sequencing metrics for a sequenced genomic sample within a graphical user interface of sequencing application (832).
[0080] As shown in FIG. 4, sequencing application (832) is included in client device (830). Sequencing application (832) may include a web application or a native application stored and executed on client device (830) (e.g., a mobile application, desktop application). Sequencing application (832) may include instructions that (when executed) cause client device (830) to receive data from sequencing system (814) and present, for display at client device (830), basecall data or data from a VCF. Furthermore, sequencing application (832) may instruct client device (830) to display summaries for multiple sequencing runs.
[0081] As further illustrated in FIG. 4, a version of sequencing system (814) may be located and implemented (e.g., entirely or in part) on client device (830) or sequencing device (810). In some versions, sequencing system (814) is implemented by one or more other components of networked system (800), such as local device (840). In particular, sequencing system (814) may be implemented in a variety of different ways across sequencing device (810), local device (840), server device(s) (820), and client device (830). For example, sequencing system (814) may be downloaded from server device(s) (820) to sequencing system (814) and / or local device (840) where all or part of the functionality of sequencing system (814) is performed at each respective device within networked system (800).
[0082] B. Examples of Base Calling Schemes
[0083] FIG. 5 illustrates a system (900) that employs two or more base callers for base calling operations on the raw images (i.e., sensor data) output by image sensors in a sequencing machinesequencing machine (910). Sequencing machine (910) of this example includes a flow cell (912), which includes a plurality of tiles (914). Each tile (914) includes a plurality of clusters (916). Sequencing machine (910) may be understood to represent a version of systems (100, 500) or sequencing device (810) described above; while flow cell (912) may be understood to represent a version of flow cells (128, 400, 510) described above. Sequencing machine (910) thus outputs sensor data (920) comprising raw images from the tiles (914) of flow cell (912).
[0084] In the present example, system (900) comprises a first base caller (922) and a second base caller (926), though some variations may include more than two base callers (922, 926). Each base caller (922, 926) of this example outputs corresponding base call classification information. For example, first base caller (922) outputs first base call classification information (924); and second base caller (926) outputs second base call classification information (928). A base calling combining module (930) generates final base calls (932), based on one or both first base call classification information (924) and / or second base call classification information (928). In some versions, first base caller (922) is a neural-network based base-caller; while second base caller (926) is a non-neural network based base-caller. For example, first base caller (922) may include a non-linear system employing one or more neural network models for base calling. The first base caller (922) may also be referred to as a DeepRTA (Deep Real Time Analysis) base caller or Deep Neural Network base caller.
[0085] By way of further example only, second base caller (926) may include, at least in part, a linear system used for base calling. For example, some versions of second base caller (926) do not employ a neural network for base calling (or use a smaller neural network model for base calling, compared to a larger neural network model used by first base caller (922)). Second base caller (926) may also be referred to as an RTA (Real Time Analysis) base caller. An RTA base caller may use linear intensity extractors to extract features from sequencing images for base calling. In some such versions, RTA performs a template generation step to produce a template image that identifies locations of clusters (916) on a tile (914) using sequencing images from some number of initial sequencing cycles called template cycles. The template image is used as a reference for subsequent registration and intensity extraction steps. The template image is generated by detecting and merging bright spots in each sequencing image of the template cycles, which in turn involves sharpening a sequencing image (e.g., using the Laplacian convolution),determining an “on” threshold by a spatially segregated Otsu approach, and subsequent five-pixel local maximum detection with subpixel location interpolation.
[0086] In another example, locations of clusters (916) on a tile (914) are identified using fiducial markers. A solid support upon which a biological specimen is imaged may include such fiducial markers, to facilitate determination of the orientation of the specimen or the image thereof in relation to probes that are attached to the solid support. Examples of fiducials include, but are not limited to, beads (with or without fluorescent moieties or moieties such as nucleic acids to which labeled probes can be bound), fluorescent molecules attached at known or determinable features, or structures that combine morphological shapes with fluorescent moieties.
[0087] RTA then registers a current sequencing image against the template image. This is achieved by using image correlation to align the current sequencing image to the template image on a sub-region, or by using non-linear transformations (e.g., a full six-parameter linear affine transformation). RTA generates a color matrix to correct cross-talk between color channels of the sequencing images. RTA implements empirical phasing correction to compensate noise in the sequencing images caused by phase errors. After different corrections are applied to the sequencing images, RTA extracts signal intensities for each spot location in the sequencing images. For example, for a given spot location, signal intensity may be extracted by determining a weighted average of the intensity of the pixels in a spot location. For example, a weighted average of the center pixel and neighboring pixels may be performed using bilinear or bicubic interpolation. In some implementations, each spot location in the image may comprise a few pixels (e g., 1-5 pixels). RTA then spatially normalizes the extracted signal intensities to account for variation in illumination across the sampled imaged. For example, intensity values may be normalized such that a 5th and 95th percentiles have values of 0 and 1, respectively. The normalized signal intensities for the image (e.g., normalized intensities for each channel) may be used to calculate mean chastity for the plurality of spots in the image.
[0088] In some implementations, RTA uses an equalizer to maximize the signal-to-noise ratio of the extracted signal intensities. The equalizer may be trained (e.g., using least square estimation, adaptive equalization algorithm) to maximize the signal-to-noise ratio of cluster intensity data in sequencing images. In some implementations, the equalizer is a lookup table (LUT) bank with a plurality of LUTs with subpixel resolution, also referred to as “equalizer filters” or “convolutionkernels.” By way of example only, the number of LUTs in the equalizer may depend on the number of subpixels into which pixels of the sequencing images can be divided. For example, if the pixels are divisible into n-by-n subpixels (e.g., 5 x 5 subpixels), then the equalizer generates n2LUTs (e g., 25 LUTs).
[0089] In some implementations of training the equalizer, data from the sequencing images is binned by well subpixel location. It should be understood that a “well” may include depressions (404) of a flow cell (400) or any other kind of reaction site (e.g., in a flow cell or otherwise). In an example of sequencing images being binned by well subpixel location, for a 5 x 5 LUT, l / 25th of the wells have a center that is in bin (1,1) (e.g., the upper left comer of a sensor pixel), l / 25th of the wells are in bin (1,2), and so on. The equalizer coefficients for each bin may be determined using least squares estimation on the subset of data from the wells corresponding to the respective bins. This way, the resulting estimated equalizer coefficients are different for each bin. Each LUT / equalizer filter / convolution kernel has a plurality of coefficients that are learned from the training. The number of coefficients in a LUT may correspond to the number of pixels that are used for base calling a cluster. For example, if a local grid of pixels (image or pixel patch) that is used to base call a cluster is of size p xp (e.g., 9 x 9 pixel patch), then each LUT has p2coefficients (e.g., 81 coefficients). The training may produce equalizer coefficients that are configured to mix / combine intensity values of pixels that depict intensity emissions from a target cluster being base called and intensity emissions from one or more adjacent clusters in a manner that maximizes the signal-to-noise ratio. The signal maximized in the signal-to-noise ratio is the intensity emissions from the target cluster, and the noise minimized in the signal-to-noise ratio is the intensity emissions from the adjacent clusters, i.e., spatial crosstalk, plus some random noise (e.g., to account for background intensity emissions). The equalizer coefficients are used as weights and the mixing / combining includes executing element-wise multiplication between the equalizer coefficients and the intensity values of the pixels to calculate a weighted sum of the intensity values of the pixels, i.e., a convolution operation.
[0090] RTA then performs base calling by fitting a mathematical model to the optimized intensity data. Suitable mathematical models that can be used include, for example, a k-means clustering algorithm, a k-means-like clustering algorithm, expectation maximization clustering algorithm, a histogram -based method, and the like. Four Gaussian distributions may be fit to theset of two-channel intensity data such that one distribution is applied for each of the four nucleotides represented in the data set. In some implementations, an expectation maximization (EM) algorithm may be applied. As a result of the EM algorithm, for each X, Y value (referring to each of the two channel intensities respectively) a value may be generated which represents the likelihood that a certain X, Y intensity value belongs to one of four Gaussian distributions to which the data is fitted. Where four bases give four separate distributions, each X, Y intensity value will also have four associated likelihood values, one for each of the four bases. The maximum of the four likelihood values indicates the base call. For example, if a cluster is “off’ in both channels, the base call is G. If the cluster is “off’ in one channel and “on” in another channel the base call is either C or T (depending on which channel is on), and if the cluster is “on” in both channels the base call is A.
[0091] In some implementations of RTA, the base calling errors get averaged out across many training examples. In some other implementations, the ground truth may be sourced using aligned genomic data, which may provide better quality because aligned genomic data may use reference genome and truth information that incorporate the knowledge gained from multiple sequencing platforms and sequencing runs to average out the noise. The ground truth may include basespecific intensity values (or feature values) that reliably represent intensity profiles of bases A, C, G, and T, respectively. A base caller like RTA base caller (926) base calls clusters by processing the sequencing images and producing, for each base call, color-wise intensity values / outputs. The color-wise intensity values may be considered base-wise intensity values because, depending on the type of chemistry (e g., 2-color chemistry or 4-color chemistry), the colors map to each of the bases A, C, G, and T. The base with the closest matching intensity profile is called.
[0092] A trainer may train base caller (926) and generate the trained coefficients of the sharpening masks using various training techniques. FIG. 6 shows one implementation of an adaptive technique that may be used to train base caller (926), e.g., using an offline or online mode. Here, the logic is y = x.h + d, where x is the input pixel intensities, h is the sharpening mask coefficients, d is the DC offset. In some implementations, x and h are row and column vectors respectively, with length 81. This vector model is equivalent to a dot product of 9 x 9 matrices representing input pixels and coefficients. The cost is the expected value of error squared. The gradient update moves each coefficient in a direction that reduces the expected value of errorsquared. Applying this update generates a new estimate of the coefficients that moves them in a direction that (on average) reduces the mean squared error (MSE). In some implementations, Mu is a small constant used to change the adaptation rate / convergence speed. A DC term update can be calculated in a similar way. A gain term update also can be calculated in a similar way.
[0093] In some implementations, since linear interpolation is applied on the coefficient sets, the updates are applied slightly differently in the following manner:
[0094] h(q, n+1) = h(q, n) + lambda q. mu.x(n). e(n)
[0095] In the equation above, h(q, n) is weight q at cycle n, lambda q is the linear interpolation weight for a particular set of coefficients and can include four updates per output due to linear interpolation in two dimensions. The recursive least-squares technique extends the least squares technique to a recursive algorithm.
[0096] C. Examples of Structural Variation Graph Genome Generation
[0097] In some scenarios, a secondary analysis may be performed iteratively while sequence reads are generated by a sequencing system such as systems (100, 500, 814) described herein. Secondary analyses may encompass both alignment of sequence reads to a reference sequence (e.g., the human reference genome sequence) and utilization of this alignment to detect differences between a sample and the reference. Secondary analyses may enable detection of genetic differences, variant detection and genotyping, identification of single nucleotide polymorphisms (SNPs), small insertions and deletion (indels) and structural changes in the DNA, such as copy number variants (CNVs) and chromosomal rearrangements.
[0098] By performing secondary analyses while sequence reads are generated, system (100, 500, 814) may determine preliminary variant calls iteratively in real-time (or with zero or low latency). Final results of variant determinations may be available soon after (or immediately after) the end of a sequencing run. Alternatively, a sequencing run may be terminated early if variant calls are available with sufficient confidence during the run. In some scenarios, only information related to variant determinations (e.g., variant calls) is transferred off the sequencing system (100, 500, 814). This may decrease, or minimize, the data bandwidth required in comparison to performing the variant determinations in a system that is external. In addition, only variant information may be sent to a computing system (e.g., a cloud computing system) for furtherprocessing. In this example, sequencing runs may be terminated prior to completion of an entire sequencing process. For example, if the identity of a pathogen of interest is determined after a number of sequencing cycles of a sequencing run, the sequencing run may be terminated. Thus, the time to a particular answer (e.g., pathogen identification) may be decreased. In some implementations, outputs and intermediate results of system (100, 500, 814) may include histograms of duplicates, exact matches, single and double SNPs, and single and double indels.
[0099] VI. Examples of read-level somatic methyl detection
[0100] The described sequencing systems and methodologies described herein may be utilized in various contexts for the generation of methylation (e g., epigenetic) data with respect to a given genome or sample. By way of example, and as discussed herein, sequencing-based methylation data may be generated using various techniques, one of which is bisulfite sequencing (BiS). In bisulfite sequencing approaches methylated cytosines may be detected in DNA (e.g., genomic DNA) by treating a DNA sample with sodium bisulfite prior to sequencing. In response to the bisulfite treatment, unmethylated cytosines are deaminated to uracils, which upon sequencing convert to thymidines. Conversely, methylated cytosines are not deaminated and are read as cytosines. The location of the methylated cytosines is then determined by comparison of bisulfite treated sequences to a reference genome or to untreated sequences. In this manner, methylation data for a sample may be obtained using sequencing techniques.
[0101] By way of a further example, enzymatic methyl sequencing (EMS or EM-Seq) is another approach by which methylation data may be obtained using sequencing techniques. In EMS approaches, non-destructive enzymatic reactions are employed which address certain technical biases that may be introduced by BiS techniques. In EMS approaches, a two-step conversion process is employed. In the first reaction TET2 and T4-BGT convert methylated cytosine (and hydroxymethylcytosine) into products that are not deaminated by APOBEC3A (i.e., methylated cytosines are modified to be protected from deamination). In the second reaction, APOBEC3A is used to deaminate unmodified cytosines to uracils. Methylations may then be determined using sequencing techniques and systems as described herein.
[0102] With this in mind, conventional methylation analysis approaches work in a site-specific manner, identifying methylation status site-by-site, even within a larger region. However, suchconventional methylation assessment approaches may suffer from various drawbacks, such as low sensitivity to low tumor mutational burden (TMB) contexts and / or in the context of low shedding tumors. In addition, base calling inaccuracy may act to confound calls of DNA variants from single reads. Similarly, in other contexts related to read-level analyses, methyl conversion errors may confound confident single-read and base methylation estimates. Further, ambiguous methylation states in the subject may generate background noise, making it more difficult to identify useful diagnostic methylation signal.
[0103] As discussed herein, the presently described techniques and systems address such failings of such conventional methylation assessment approaches as well as issues impeding implementation of read-level approaches. For example, in contrast to conventional base-level (i.e., site-specific) analysis, the presently described techniques assess methylation at the read-level in the context of a sequencing operation (e.g., next generation sequencing (NGS) operation). In accordance with the presently-described read-level approaches, issues related to low sensitivity in the context of low TMB or low shedding tumors are mitigated by looking for abnormalities in methylation, as opposed to DNA variants, thereby increasing the potential search space. Likewise, issues related to base calling inaccuracy and / or methyl conversion errors are addressed by assessing read-level methylation state using multiple CpG sites to effectively decouple technical error from classification accuracy. Further, and as discussed herein, issues associated with classifying read-level methylation state regardless of methyl conversion accuracy may be addressed by the use of unsupervised clustering to auto-detect appropriate thresholds, even in the context single-sample data inputs. Additionally, as discussed herein, background noise attributable to ambiguous methylation states may be addressed by selection of genomic features with binary methylation stated in tumor / normal pairs to facilitate confident detection of abnormalities.
[0104] With the preceding in mind, sequencing technology systems and methods as described above may be utilized as part of assessing methylation levels of a subject. Such methylation levels may in turn be used for generating a diagnosis and, correspondingly, recommending, selecting, or implementing a treatment plan (e g., a personalized treatment plan) or protocol, such as a pharmacological, radiological, surgical, treatment plan or a suitable combination of such treatment techniques. Further, methylation assessment as discussed herein may be ongoing over time so asto allow or facilitate monitoring of treatment efficacy of an ongoing or prior treatment, adjustment to an ongoing treatment plan, and / or confirmation of a remission state for a subject.
[0105] By way of example, methylation may be a useful tool in identifying and quantifying markers for tumors. In particular, low tumor mutation burden (TMB) and / or low shedding tumors are difficult to detect solely using DNA-based techniques due to the scarcity of signal. As a result, methylation has become a popular tool, such as within the liquid biopsy field, due to the relative increase in signal relative to DNA-based techniques since the methylation signal is not dictated by TMB. Indeed, abnormal methylation patterns may be associated with non-mutated sites. Further, with respect to DNA-based techniques there are sensitivity limitations related to sequencing error rates (as well as amplification errors and library preparation errors) that are relevant when attempting to detect a limited number (e.g., a single) event in a limited number (e.g., a single) read. In particular, in DNA-based techniques it is common to only get one alteration per read, which means sequencing error may be a limiting factor. Conversely, based on methylation one may observe multiple alterations per read (e.g., multiple target bases), thereby providing greater confidence that an observed deviation is real and not an error. Further, methylation may serve as a tissue-specific marker, which may be of particular importance in cancer diagnosis and in treatment selection and parameterization.
[0106] Conventional base-level approaches for assessing methylation look at a single site of methylation (e.g., a single CpG site) in contrast to the presently described approaches which quantify methylation across a sequencing read. With this in mind, the techniques described herein may, in certain implementations, be used to identify a quantity or percentage of methylated sites across a DNA sequencing read, as opposed to conventional approaches which only identify whether a single site (e.g., a CpG site) is methylated, is not methylated, or is methylated some percentage of the time among a number of reads. As used herein a “read” may be understood to be a sequence or strand sequenced from a DNA fragment, such as during a next-generation sequencing (NGS) operation. Correspondingly, a “read length” refers to the number of base pairs (BP) sequenced from such a fragment. Read lengths may vary based on the sequencing reagents and / or sequencing instrument employed, with more chemistry cycles typically generating longer reads. In practice, reads may be generated as single continuous reads (which may be, for example, 300 bp to 400 bp in length or longer) or as paired-end reads (which may be 150 bp to 200 bp inlength or longer). For example, when sequencing cell-free DNA (cfDNA) fragments which are typically 120-150 bp in length, the entire length of the fragment can be read by a single continuous read. As another example, when sequencing a larger target that is 600-800 bp in length, paired end reads can capture up to 200 bp or more from each end of the target sequence.
[0107] As discussed herein, methylation data is acquired for such reads at the “read-level” using techniques described based on a methylation status or quantitative metric of a respective read, in contrast with position-level or site-specific data acquired site-by-site for a strand of DNA. That is, read-level data as described herein is data at the whole strand or sub-region level, not just for a nucleotide site or set of sites. Further, such read level data may be generated and presented in a clinically useful time frame (e.g., minutes, tens of minutes, or hours).
[0108] Cancer screening examples are employed herein to provide a useful example and context for the presently described technique. However, it should be understood that the epigenetic information obtainable by the present techniques may be used in a multitude of contexts where regulatory pathways may be implicated in an abnormality or disorder or may otherwise be useful in monitoring (either once or periodically over a period of time) a physiological process. By way of example, the presently described techniques may be useful in the context of reproductive health (e.g., assessing methylation changes or abnormalities during fetal development), in the context of cardiovascular disease (e.g., assessing methylation changes or abnormalities associated with cardiac or vascular function at a moment in time or over a period of time), in the context of neurodegeneration, including Alzheimer’s (e.g., assessing methylation changes or abnormalities associated with neural function, cognition, and / or memory at a moment in time or over a period of time), and / or in the context of metabolic disorders, such as diabetes (e.g., assessing methylation changes or abnormalities associated with metabolic or anabolic function at a moment in time or over a period of time). Further, the presently described techniques may be used in the context of monitoring transplant acceptance (or rejection) where methylation markers may be indicative of the subject’s adaptation to and acceptance of the transplanted tissue or organ. More generally, the presently described techniques may be used in contexts where abnormal methylation patterns are deemed to be of interest or of significance. Additionally, the presently described techniques may be of particular usefulness in situations or uses where a genetic signal (e.g., a signal based directly or indirectly on DNA sequence) is low (e.g., at or near a limit of detectability) and / or is unclear,such as due to noise, artifacts, or likely error (e.g., sequencing error). In such situations, methylation analysis may help refine the analysis and allow one to observe alterations earlier in time.
[0109] In the oncology context, the presently described techniques may be particularly useful in enabling certain clinical application, such as cancer screening and minimal residual disease (MRD) testing. In particular, such applications may be difficult due to the need to detect abnormalities from a single read. In such contexts a clinician may be looking for a small population of abnormal molecules. As discussed herein, methylation is regulated in regions, not single bases, which is one benefit of a read-based approach as described. Further, detection of multiple, co-occurring abnormalities, which is possible with the presently described techniques, greatly increases confidence in the veracity of clinically informative results.
[0110] With the preceding in mind, certain of the analyses described herein utilized a paired tumor / normal cell line set of data. Such paired tumor / normal cell line data represented tumor / normal cells from the same original donor, with tumor cells being derived from carcinoma cells and normal cells being derived from lymphocytes. Differing sequencing techniques were employed to determine if sequence technique might be a relevant variable of interest. In particular, bisulfite sequencing (BiS or BS-Seq) and enzymatic methyl sequencing (EMS or EM-Seq) were employed in evaluating the present read-based methylation analyses. As part of the studies performed in evaluating the presently described techniques, the correlation between EMS and BiS data was studied with a focus on CpG islands (i.e., regions of the genome that have higher CpG density than average), which tend to be sites of gene expression regulation. A CpG island is typically defined as genomic locus of at least 200 bp with >50% GC content and an observed-to- expected ratio of CG dinucleotides exceeding 0.6. Based on these analyses it was observed that data obtained via BiS and EMS for methylation level correlated well within the same samples.
[0111] Additionally, multiple CpG islands were analyzed to determine whether such islands exhibited opposite methylation levels for the tumor / normal cell lines to a substantial degree. Such opposite methylation states between the tumor / normal cell lines (i.e., low methylation in tumor / high methylation in normal or high methylation in tumor / low methylation in normal) would provide an indication of the utility of using methylation state of certain CpG islands in a clinical context. For chr!8 it was observed that approximately 60% of chr!8 CpG islands have similarmethylation levels (i.e., states), with approximately 238 CpGislands being hypom ethylated in both cell lines and approximately 57 CpG islands being hypermethylated in both cell lines. Conversely, it was observed that approximately 5% of chrl8 CpG islands have extremely discordant methylation patterns between the tumor / normal cell lines, with approximately 16 CpG islands being hypermethylated in lymphocyte but hypomethylated in carcinoma and approximately 8 islands being hypomethylated in lymphocyte but hypermethylated in carcinoma. Assuming such frequencies are consistent across the genome, one may expect approximately 500 to 1,500 CpG islands to exhibit opposite methylation patterns for these sample types. Based on this analysis it was determined that there is sufficient difference between the sample types (e.g., ~5%) that such islands having discordant methylation patterns could be used as a proxy for detecting differences in somatic-type cases.
[0112] Classification of methylation states at a read-level (as opposed to a site or singleposition level) was also investigated. As part of this study, various thresholding requirements were considered such as, but not limited to: the number of CpG sites needed, whether filtering is beneficial, the amount of overlap with a target region that is needed, can thresholds be autodetermined on a per sample basis or are global threshold preferred, how many thresholds are appropriate, and so forth.
[0113] Turning to FIG. 7, a graph is shown depicting a plot of the extent of methylation in sequencing reads based on percentage methylation (i.e., methylated CpGs / total CpGs). As shown in FIG. 7, the distribution of methylation percentage is bimodal, with peaks at or near 0 and 1, signifying that the reads under analysis (from CpG islands) tend to fall into methylated or unmethylated states. In particular, approximately 60% of the reads have no methylated CpG sites, approximately 9% have all methylated CpG sites, and approximately 31% of reads have a mix of methylated and unmethylated CpG sites (i.e., between 0 and 1). Turning to FIG. 8, it was also observed that most reads were contained fully within a CpG island. The number of CpG sites per read was also investigated. The majority of reads had 0 methylated CpGs while unmethylated CpGs had a bimodal distribution with peaks at 0 and 15. Most reads contained between 5 and 25 total CpGs (i.e., methylated and unmethylated CpG sites) with reads corresponding to an approximately 200 bp insert (i.e., approximately the size of a cfDNA strand).
[0114] In certain implementations of the techniques described herein reads may becharacterized and fdtered based on whether they are fully or partly contained within a target region (e.g., CpG island). For example, a read pair (e.g., a 200 base read pair) may be excluded from consideration if all 200 bases are not within a CpG island in a context where 100% overlap with the target region is specified. Conversely, in other contexts the read may be partially overlapping (e g., 50%, 75%, 80%, 85%, 90%, 95% overlapping, and so forth) with the target region to be retained. For example, in a 90% overlap threshold context, if a read pair is 200 bases, at least 180 bases would need to be within a CpG island to meet the 90% overlap threshold. Based on the overlap criteria established (e.g., fully within the target region or partially overlapping with the target region), reads not meeting the overlap criteria may be excluded (e.g., filtered out) from analysis. Alternatively, in other embodiments no overlap criterion may be specified and, thus, reads not excluded based on overlap with the target region. As one example, if a target region is large, (e.g., at least 2X the median insert length), a more stringent thresholding could be applied than if the region is smaller. Alternatively, where a region is very small, only a segment of the read under analysis might be considered.
[0115] Another question examined is whether the number of CpG sites affect the percentage of methyl cytosine measurements that are expected. That is, if many CpG sites are present, does it become less likely that the region in question is fully methylated. In a study in which measurements were thresholded based upon the number of CpG sites, it was observed that the number of CpG sites and percentage of methylation observed were not meaningfully correlated.
[0116] With the preceding in mind, in certain embodiments it may be useful to automatically determine methylation thresholds for a given sample or read as opposed to using global or predetermined thresholds. Suitability of K-means clustering was investigated for the purpose of automated threshold determination that is tuned to whatever dataset is provided as an input. By way of background, K-means clustering is a technique by which a number (n) of observations or measurements are partitioned into a specified number (k) of clusters or groupings. Each observation or measurement is assigned or grouped into the cluster having the closest mean. As may be appreciated however, other statistical approaches may also be suitable for generating one or more sample or dataset specific methylation thresholds and corresponding grouping or binning solutions. Such sample specific thresholding may be useful in view of variations that may exist between sample preparations (e.g., in terms of variations in conversion rate, accuracy, and so forthbetween sample preparations).
[0117] To determine the suitability of K-means clustering in classifying the methylation status of read-level sequencing data, frequency distributions were provided as input (such as shown in FIG. 7) to a K-means clustering routine to generate different numbers of groupings. In particular, k values of 2, 3, 4, and 5 were employed and the results reviewed, with higher k values correspondingly increasing boundary stringency. One observes that requiring a read to be within a group with the lowest or highest methylation state to be classified as unmethylated or methylated, respectively, causes higher k values and increasing boundary stringency. Turning to FIG. 9, for two clusters (i.e., k = 2), corresponding methylated and unmethylated groupings were generated with a threshold at approximately 45%. Similarly, and turning to FIG. 10, for three clusters (i.e., k = 3), unmethylated, intermediate, and highly methylated groupings were generated, with a threshold (i.e., unmethylated threshold) separating unmethylated and intermediate at approximately 23% and a threshold (i.e., methylated threshold) separating intermediate and highly methylated at approximately 68%. For k values greater than 3, additional intermediate groupings were added, but no clinical relevance was believed to be added by such additional intermediate groupings.
[0118] Additionally, the effects of up-front read filtering (i.e., limiting to reads having 1 or more or 5 or more CpG sites) and of filtering on percent target overlap (i.e., limiting to full overlap (100%) with a target region (e.g., CpG island) or partial overlap (>0%) with a target region) on such thresholds were evaluated in the k = 3 context. In general, however, up-front filtering and percent target overlap filtering did not significantly alter the unmethylated and methylated thresholds. In some implementations, however, one or both of up-front filtering and percent target overlap filtering may be employed for improved results or to address sample quality issues, errors introduced by one or both of the amplification or sequencing processes, or other data quality issues. This can be especially evident when evaluating reads with a single CpG site, because sequencing error can be a confounding factor.
[0119] With the preceding in mind, efforts were made to determine a suitable number of clusters or groupings based on a K-means thresholding methodology. As part of this analysis overlap with a CpG island and the number of CpGs were considered. Turning to FIG. 9, for k = 2, centroid 0 was found to exhibit 3.568% methylation, to have 87.9% overlap with a CpG island,and to have a 0.236 normalized number of CpGs (~18); centroid 1 was found to exhibit 86.5% methylation, to have 81.2% overlap with a CpG island, and to have a 0.194 normalized number of CpGs (—15). For comparison, and turning to FIG. 11, for k = 3, centroid 0 was found to exhibit 7.06% methylation, to have 44.8% overlap with a CpG island, and to have a 0.153 normalized number of CpGs (—12); centroid 1 was found to exhibit 87.9% methylation, to have 82.7% overlap with a CpG island, and to have a 0.197 normalized number of CpGs (—15); and centroid 2 was found to exhibit 3.48% methylation, to have 97.4% overlap with a CpG island, and to have a 0.254 normalized number of CpGs (~19). Further, and turning to FIG. 12, for k = 4, centroid 0 was found to exhibit 5.02% methylation, to have 46.2% overlap with a CpG island, and to have a 0.157 normalized number of CpGs (—12); centroid 1 was found to exhibit 3.26% methylation, to have 97.5% overlap with a CpG island, and to have a 0.254 normalized number of CpGs (—19); centroid 2 was found to exhibit 87.2% methylation, to have 44.8% overlap with a CpG island, and to have a 0.140 normalized number of CpGs (~11), and centroid 3 was found to exhibit 86.4% methylation, to have 95.9% overlap with a CpG island, and to have a 0.216 normalized number of CpGs (—16). As may be noted, as additional groups or clusters are added, better discrimination (e.g., separation) between methylated and unmethylated groups may be obtained (i.e., one can identify less distinct methylation pattern groupings in the added groups. However, in terms of tradeoffs, additional clusters also increase the likelihood of overfitting the data. By way of further characterization, for k = 3, Read %mCpG levels were tabulated for tumor / normal sample data evaluated using both BiS and EMS, as shown in Table 1.Table 1
[0120] Based on the preceding discussion and factors, a number of variables may be utilized in a read-based methylation analysis, such as approaches including automated thresholding, such as using K-means clustering our other suitable statistical techniques. Relevant factors that may be incorporated into such analyses and / or thresholding include, but are not limited to: pre-filtering sequence read data based on number of CpGs (e.g., 1+, 3+, 5+, 10+, and so forth), use of automated thresholding using statistical techniques (e.g., K-means clustering using one or more of read percent methylated CpG sites (% mCpG), CpG count, target overlap (e.g., 75% 85%, 90%, 95%, 100% CpG island overlap), and so forth), for clustering techniques, the number of clusters to be employed (e.g., 2, 3, 4, 5, or more clusters), filtering based on centroid center distances for one or more methylation patters (e.g., low and / or high methylation), filtering reads in low CpG count and / or low target overlap clusters, use of algorithmic thresholding (e.g., low methylation threshold defined as remaining low %mCpG centroid value + (high - low) / 4 and / or high methylation threshold defined as remaining high %mCpG centroid value - (high - low) / 4), and so forth.
[0121] With the preceding in mind, and by way of example, in one implementation criteria for a read-level methylation analysis may be employed that include one or more of: filtering out reads with less than 5 CpGs, filtering out reads with less than 95% overlap with a CpG island, employing K-means clustering with % methylated CpG sites (% mCpG) and k = 3, and classifying based on the k-Means labels as low, intermediate or high methylation. In such an implementation, one or both of the number and percentage of reads in each methylation category may be quantified for each CpG island. By way of example, this process was performed for different tumor / normal cell line samples and the technical replicate concordance assessed. As discussed below, tumor / normal results were compared to evaluate differences in the CpG islands when looking at methylation data at the read-level (as opposed to site- or position-level).
[0122] With respect to filtering out reads with less than x CpGs, such a step may be useful since fewer CpGs mean less information is available to distinguish methylated / unmethylated status. An example of such filtration based on number of CpGs per read is graphically illustrated as FIG. 1 where a threshold of >5 CpGs is shown. Such a threshold eliminates between 2.7% - 8.6% of reads. Similarly, with respect to filtering based on percent read overlap with a CpG island, such a step may be useful since lower overlap with a target region could lead to irrelevant signalchanges. An example of such filtration based on percentage read overlap with a CpG island is graphically illustrated as FIG. 14 where a threshold of 95% overlap is shown.
[0123] As may be appreciated, the preceding illustrates the capability to classify or characterize reads based on methylation state (e.g., methylated, unmethylated, semimethylated). Further works was done to analyze the data at the CpG island level to characterize such islands based on methylation state using read data. As part of this analysis, whether there was correlation between the different methylation states in CpG islands was studied (both based on counts and on percent methylation of the respective reads) and little to no correlation was observed. That is, based on this work it was determined that the different methylation states were largely independent of one another.
[0124] Additionally, it was observed that some of the sample had more (or less) coverage of CpG islands. By way of example, some samples had dropouts, and correspondingly less coverage. In the study, it was observed that 97% of CpG islands had coverage in all samples. Correspondingly, to simplify analysis the work set was filtered down to those islands where all of the samples contained coverage.
[0125] With this in mind, FIG. 15 depicts a 3D plot depicting methylated, unmethylated, and semimethylated read frequency for individual CpG islands. As shown, data density for the CpG sites is highest in the regions encircled by ovals 950. These high data density encircled regions correspond to: (1) CpG islands with a mix of unmethylated and some partially methylated (i.e., semimethylated) reads, and (2) CpG islands with a mix of methylated and some partially methylated reads. This graphically supports the presence of bimodal segregation on a region level of whether a CpG island state is methylated or unmethylated.
[0126] Similarly, and turning to FIG. 16, this concept is illustrated in a frequency plot of the percentage of methylated, unmethylated, and semimethylated reads in islands. In this example, the different categories of reads were extracted to allow visualization of what percentage of reads within islands the categories correspond to. As in the prior examples, the methylated and unmethylated distributions are bimodal. That is, one generally has mostly unmethylated reads or few or no unmethylated reads. Similarly, one generally has mostly methylated reads or few or no methylated reads. With respect to the semimethylated reads, these reads are skewed towardsuncommon (i.e., toward 0), so one generally has none or a small number of semi methylated reads. In particular, with respect to numeric results and with reference to FIG. 16, 38% of the CpG islands have only unmethylated reads while 9% have only methylated reads. With respect to other thresholds, 52% of the CpG islands have >90% unmethylated reads and 16% have >90% methylated reads. Similarly, 58% of the CpG islands have >75% unmethylated reads and 19% have >75% methylated reads. With this in mind, CpG island threshold may be established for classification of such CpG islands as methylated or unmethylated. For example, a methylated CpG island may be characterized as one having >75%, >85% >90% or >95% methylation and 0%, <5%, <10%, or <15% unmethylation (e.g., >90% methylation and 0% unmethylation). Conversely, an unmethylated CpG island may be characterized as one having >75%, >85% >90% or >95% unmethylation and 0%, <5%, <10%, or <15% methylation (e.g., >90% unmethylation and 0% methylation).
[0127] Of the samples studied there were 667 CpG islands where at least one sample was categorized as methylated and another as unmethylated, corresponding to 2.5% of those with coverage in all samples. Such divisions were particularly notable between the tumor / normal samples, with 584 (88%) of the CpG islands having opposite methylation patterns for at least one LP06 (normal lymphocyte)-LP07 (carcinoma) pair (i.e., a tumor / normal pair).
[0128] This is graphically illustrated in FIG. 17, in which approximately one-third of CpG islands showed consistent methylation states across replication, with original sample identity (i.e., tumor or normal) being a clear factor in what are generally binary (e.g., bimodal) methylation patterns for respective CpG islands. In particular, with respect to FIG. 17, this figure depicts each CpG island visually shaded based on whether the islands are methylated or unmethylated in a single sample. Hierarchical clustering was performed to see what the groupings were. Separation is generally discernible between the LP06 (normal) and LP07 (tumor) groups based on methylation state.
[0129] With the preceding in mind, further study was performed to evaluate the use of readlevel methylation analysis for somatic tumor detection. In accordance with these studies, a first tumor / normal sample (e.g., corresponding to a first set of sampled collected from a patient) were used to identify which CpG islands were tumor informative and what their respective tumor informative sates would be (i.e., methylated or unmethylated). Other data sets (separately preparedbut from the same original material) were used and the reads mixed together to get different tumor fractions (i.e., to titrate down the tumor signal into the normal signal). This was then used to determine at what levels tumor apparent reads could still be ascertained.
[0130] More particularly, for both the EMS and BiS methodologies, steps were performed comprising:(1) For each technical replicate ID:(A)Extract CpG islands that are methylated in one of either tumor or normal of the method-replicate sample and unmethylated in the other(B)For each other technical replicate:(i) Get a count of how many of the target CpG islands contain methylated or unmethylated reads (whichever matches the expected tumor status)• Perform initially on the pure tumor / normal samples to obtain a baseline and max signal / noise estimates(ii) Run Monte Carlo simulations to mix the reads together from pairs of LP06 and LP07• %LP07 (tumor): 10%, 1%, 0.1%, 0.01%, 0.001% (lowest point is 1 in 100,000• Even up potential coverage differences between samples(2) Generate Limit of Detection (LoD) type curves, evaluate for:• How tight are the replicates• How far above background is the signal at each “titration”• Attempt to quantify sample level identification rates.
[0131] Based on this protocol, a pure tumor sample was used to identify which CpG islands were tumor informative (e.g., if CpG islands are identified in one sample, do they consistently show up in other samples as well so as to be a reliable indicator). In such a context, an informativeisland corresponds to a cancer or tumor signature that can be expected to be seen in affected cell lines and indicative of the respective tumor type. Turning to FIG. 18, results of such an approach are depicted. In this illustration, as shown in the lower left, 5.6% of the CpG islands in the noncancer sample had the tumor signature while, as shown in the upper right, 99.6 of the CpG island in the cancer sample had the tumor signature. That is, the tumor signature in tumor samples was observed to be very consistent and rarely missed a CpG island that should have had the signature methylated or unmethylated read.
[0132] As Turning to FIG. 19, curves are depicted (BiS on the left and EMS on the right) that were generated from performing a Monte Carlo simulation to simulate titration results across a range from 0 to 1. The detection threshold, visually indicated via a horizontal dashed line, was defined as:(1) Detection Threshold = [i0%+ (1.645 x <J00O) + (1.645 x cr0 0010 / o) which corresponds to 0% signal + noise at 0% + the signal of a very low tumor fraction. Respective arrows 980 indicate for each methodology (i.e., Bis or EMS) where the median breakpoint is observed for each respective method and titration curve (i.e., where at least 50% of the samples are detecting more CpG islands as altered above the detection threshold. Comparable sensitivity was obtained for the BiS methodology (approximately 180 bp inserts) and EMS methodology (-210 bp inserts). In particular, for BiS greater than 50% of the samples are above the detection threshold at 0.2% tumor while for EMS greater than 50% of the samples are above the detection threshold at 0.3% tumor. For both methodologies 100% of the samples are above the detection threshold at 0.4% tumor.
[0133] As noted herein, a primary distinction with the presently described techniques and prior approaches is the reliance on read-level analysis as opposed to base-level analysis (i.e., looking at single CpG positions). With respect to the differing approaches, signal and detection were compared between the presently described read-level approach and the conventional base-level analysis. Turning to FIG. 20, read-level methylation analysis signal intensity is plotted with baselevel analysis signal intensity with respect to the titrated tumor percentages described above. Baseline signal correction was performed by subtracting the average signal of 0% tumor data. As may be seen, read-level analysis has higher signal across all tumor fractions and spikes earlier thanbase-level analysis, though base-level analysis has lower technical variance. This is believed to illustrate a strength of the read-level analysis technique, which takes advantage of the ability to capture a change in methylation because multiple changes per molecule are being observed. In addition, the percentage of samples identified as abnormal by both read-level analysis and baselevel analysis was assessed. Read-level analysis appears generally comparable in terms of sensitivity to base-level analysis in this assessment
[0134] Whether a coverage filter improves performance of the read-level analysis was also studied. In particular, whether requiring 5 or 10 reads as well as requiring the minimum number of reads in all data sets or only the training data sets (i.e., the samples used to identify which regions (e.g., CpG islands) were tumor informative) was evaluated. Use of a coverage filter was determined to provide improvement in terms of increased signal for all % tumor values. With this in mind, in some implementations a coverage filter (e.g., a minimum of 5 reads) may be employed, such as for the training data sets.
[0135] As a further proof-of-concept of the present approach, testing was performed using known problem data. In particular, data sets were analyzed using read-level and base-level techniques where the samples were generated using enzymatic conversion techniques. In general, such enzymatic conversion techniques are associated with false positive conversion of unmethylated bases and an erroneous lack of conversion of methylated bases at low levels. Such errors are typically detrimental to base-level methylation analysis approaches.
[0136] In terms of real-level methylation with such problem data, general data set distributions generated using BiS and EMS relative to the problem data remained generally bimodal (i.e., methylated and unmethylated), though less clean. In particular, a low peak of methylated reads was observed in the problem data (due to false positive conversion of unmethylated bases) and the peak corresponding to full methylation was reduced relative to a non-problematic data set. With respect to other parameters looked at and discussed herein, with respect to the BiS and EMS data relative to the problem data, the same number of CpGs per read were observed and the same percentage overlap with CpG islands was observed. In addition, the shape of the histogram after filtering based on the percentage of methylated CpGs was also unchanged.
[0137] With respect to the K-means clustering for auto-thresholding, this technique alsocontinued to perform well with the problem data. While more noise was present and more observations occurred in the middle (i.e., partially methylated) range, separation thresholds for separating the methylated and unmethylated reads were still observed to occur at substantially similar boundaries.
[0138] Looking at the island-level (i.e., target region) read distributions, bimodal peaks were still observed (though less prominent) for methylated and unmethylated reads when analyzing the problem data at the read-level. Consistent with other results, more noise was observed in the range between peaks for the problem data.
[0139] In addition, the BiS, EMS, and problem data were analyzed using the prior described technique extracting CpG islands that showed at least two methylation states for reads. It was again observed that there was excellent separation of the tumor and normal samples, with the main differentiator between the samples being their origin. In particular, the same methylation pattern was observed in the problem data, BiS data, and EMS data within their respective sample types. Overall, the methylation state pattern matched between all methods (BiS, EMS, and problem data), with unmethylated islands having similar signal consistency between replicates for all methods. Methylated islands had a lower classification rate for the problem data, but the same overall pattern as the other methods.
[0140] The problem data was also used in the Monte Carlo simulation and titration analyses relative to the BiS and EMS data sets. With respect to the problem data, there was higher background signal and replicates were more variable. Sample detection was lower for the problem data, likely due to fewer CpG islands being identified as tumor informative in the problem data in the first steps. In particular, approximately 30 tumor informative CpG islands were identified in the problem date compared to approximately 250 in the non-problem data sets. It may be noted, however, that the lower sample detection in the problem data was still an improvement over what was seen with base-level detection approaches. In particular, using read-level analysis sample detection was at >50% at 3% tumor and at 2% tumor for the two different problem data sets used. Conversely, for base-level analysis approaches >10% tumor was required for detection.
[0141] With the preceding in mind, FIG. 21 depicts steps related herein for read-level analysis as a process flow. As may be appreciated, certain of the steps may be omitted depending on thegoal or desired end-product of a particular processing operation. Similarly, additional steps may be added or repeated depending on the goal or desired end-product of a particular processing operation. Turning to FIG. 21, at step 1000 a browser extensible data (BED) fde and / or binary alignment map (BAM) file (such as may be generated by a next generation sequencing (NGS) operation) may be provided as an input. Paired-end (PE) sequence may be extracted from target regions (e.g., CpG islands). In one embodiment such extraction may be performed using an automated script (e.g., a Python script) configured to read through the bands and extract the paired- end read data. In one such implementation, the script finds mate pairs and generates a table (i.e., a tabular format) or other suitable data structure with the extracted paired-end data. In some embodiments, for any paired read data where there is overlap in the aligned coordinates between R1 and R2, the methylation status can be inferred from R1 for the overlapping region.
[0142] In the depicted example a filtering step 1004 is performed on the extracted paired-end data, though such a filtering step may be omitted in other embodiments. In practice, such a filtering step 1004 may be to reduce file size and focus the analysis, correspondingly increasing computational efficiency of the operation. By way of example, filtering step 1004 may require that a minimum number of reads be included (e.g., 1, 3, 5, or 10 or some other suitable numeric threshold), may or may require that read have full or partial overlap with a CpG island or other target region to be included, and / or may require another suitable coverage metric for inclusion in the data being analyzed. If employed, such a filtering step may be configurable, and thus one or more filtering criteria may be an input to such a filtering step 1004.
[0143] At step 1008, read methylation states may be classified for the filtered (or unfiltered) read data. As discussed herein, in certain embodiments this may be accomplished via K-means clustering, though other suitable statistical clustering or analytic techniques (including but not limited to DBSCAN, Gaussian mixtures, Spectral clustering, and Support Vector Machines, affinity propagation and / or mean-shift) may be employed in place of or in addition to K-means clustering. In particular, at this step an automated routine may be employed to autodetect thresholds for distinguishing methylate and unmethylated read states or methylated, unmethylated, and semimethylated read states. In this manner, determination of methylation state classification thresholds may be both automated and sample specific so that such thresholds are derived from the sample data itself. As may be appreciated, inputs to such a step may include specification ofthe threshold detection routine or technique to be employed as well as parameters (e.g., a number of clusters) to be employed if needed. Alternatively, a hard threshold determined by prior empirical work may be employed at this step if dynamic determination of read methylation state thresholds is not employed.
[0144] At step 1012, target regions (e.g., CpG islands) are classified based on methylation state based on the read-level methylation state classifications generated at step 1008. Such region level classifications may be performed using a hard threshold, a dynamic or algorithmic threshold (such as a sample specific threshold), or a combination of these options. The output of the step 1012, therefore, will be a set of target regions, such as CpG islands, classified as methylated or unmethylated or as methylated, unmethylated, or neither (e.g., semimethylated or undeterminable).
[0145] The preceding steps 1000 through 1012, as discussed herein, are performed on a single sample (i.e., single sample processing). Steps 1016 through 1024 discussed below, in contrast, may occur after such single sample processing and may be performed to assess performance and utility of the methylation assessment operation.
[0146] Turning to step 1016, the target regions classified based on methylation state from step 1012 may be used to identify tumor informative target regions (e.g., tumor informative CpG islands). Such a step may be facilitated by the use of known tumor / normal pairs used for the purpose of identifying which CpG islands have different methylation state between the tumor and normal cell lines. Alternatively or in addition, literature or database searches may be used to facilitate the identification of expected tumor informative regions for diverse cancers and population backgrounds.
[0147] At step 1020 the tumor informative regions identified as step 1016 may be used in the analysis of an unknown sample to quantify how much signal of a potential tumor is present in the unknown sample based on the observed methylation states at tumor informative target regions. By way of example, in certain implementations a target threshold may be specified or otherwise provided as an input that sets a threshold coverage and / or a threshold number of CpGs having a tumor informative state that may be indicative of a tumor. Based on such a threshold or thresholds and the methylation states of tumor informative regions within the unknown sample, detected target regions having an adverse methylation state at tumor informative regions may be quantified.
[0148] The detected signal from step 1020 (e.g., the quantified metric of target regions having an adverse methylation state) may be used at step 1024 to classify the unknown sample as indicative of a tumor or now indicative of a tumor. Such a classification may be made based on one or more thresholds such as statistical confidence, requiring a threshold number of regions (e.g., CpG islands) to be observed and to have specified signal intensity within them, and so forth. Different classifications assigned to the unknown sample may correspond to different treatment types and / or protocols and / or may be used in the ongoing assessment (i.e., monitoring) of a patient undergoing treatment or after completion of treatment (e g., longitudinal monitoring).
[0149] With the preceding discussion in mind, it may be understood that the presently described read-level methylation techniques may be utilized to generate a cancer diagnosis and treatment plan at early stages, thereby vastly improving prognosis for a patient. Likewise, in a non-diagnostic context, the presently described read-level methylation techniques may be utilized to monitor a cancer patient during treatment to determine treatment efficacy and, thereby, serve as a tool to modify a treatment plan or protocol over the course of the treatment to thereby improve patient outcomes. In particular methylation-based techniques such as those described herein may be particularly useful for assessing the extent of tumor tissue present in a patient at a given time, and thus may be useful in ongoing monitoring of treatment efficacy and / or remission status.
[0150] While cancer detection and treatment are noted specifically as a particular use case of interest, it may be appreciated that the present technique may facilitate detection, quantification, and / or treatment of other genetic or epigenetic disorders associated with identifiable methylation patters at the read-level of analysis, and in particular may facilitate the detection of such disorders while at the sub-clinical level (i.e., early stages). Further, such goals may be achieved using a relatively non-invasive sample collection protocol, such as via a liquid biopsy sample that may include, but is not limited to, a blood draw of a patient.
[0151] As noted above, results and / or outputs of the presently described read-level methylation techniques may have uses in making treatment decisions and / or implementing a treatment with respect to a patient. In this manner, personalized treatment decisions may be informed for a patient based upon criteria such as a quantified measure of methylation at tumor informative CpG regions or other target regions and / or a classification of a sample form the subject based on quantitative, read-level measures of methylation at regions of interest. In particular, to the extent a read-levelmethylation level is associated with a tumor type and / or tissue origin based upon the observed methylation patterns, therapeutic decision making for the patient may be informed so as to make appropriate treatment decisions based on read-level methylation data. For example, tissue of origin information determined based upon read-level methylation analysis in conjunction with genetic sequencing information may be utilized to inform or design specialized and / or personalized treatments for a patient, such as an anti-body based treatment.
[0152] VII. Examples of Combinations
[0153] The following examples relate to various non-exhaustive ways in which the teachings herein may be combined or applied. The following examples are not intended to restrict the coverage of any claims that may be presented at any time in this application or in subsequent fdings of this application. No disclaimer is intended. The following examples are being provided for nothing more than merely illustrative purposes. It is contemplated that the various teachings herein may be arranged and applied in numerous other ways. It is also contemplated that some variations may omit certain features referred to in the below examples. Therefore, none of the aspects or features referred to below should be deemed critical unless otherwise explicitly indicated as such at a later date by the inventors or by a successor in interest to the inventors. If any claims are presented in this application or in subsequent filings related to this application that include additional features beyond those referred to below, those additional features shall not be presumed to have been added for any reason relating to patentability.
[0154] Example 1
[0155] A method for classifying one or more regions of a DNA sequence based on methylation state, comprising: accessing or acquiring signals corresponding to a set of methylation sequencing data comprising a plurality of reads within or overlapping with the one or more regions of the DNA sequence; processing the signals to classify reads of the plurality of reads based on methylation state; and classifying the one or more regions of the DNA sequence based on the methylation state of reads in or overlapping with each respective region.
[0156] Example 2
[0157] The method of Example 1, wherein the one or more regions comprise one or more CpG islands.
[0158] Example 3
[0159] The method of Examples 1 or 2, wherein the methylation sequencing data comprises bisulfite sequencing (BiS) data or enzymatic methyl sequencing (EMS) data.
[0160] Example 4
[0161] The method of Examples 1-3, wherein the plurality of reads are paired-end reads of the one or more regions.
[0162] Example 5
[0163] The method of Examples 1 -4, wherein accessing or acquiring the signals corresponding to the set of methylation sequencing data comprises extracting paired-end sequencing data using an automated script configured to process a sequencing data file generated based on an output of a sequencing system.
[0164] Example 6
[0165] The method of Examples 1-5, further comprising: performing a filtering step on the plurality of reads prior to classifying the reads based on methylation state, wherein the filtering step comprises one or both of: filtering based on a minimum number of reads; or filtering based on percent overlap with a respective region of the one or more regions.
[0166] Example 7
[0167] The method of Examples 1-6, wherein classifying reads comprises classifying reads as methylated or unmethylated.
[0168] Example 8
[0169] The method of Examples 1-6, wherein classifying reads comprises classifying reads as methylated, unmethylated, or semimethylated.
[0170] Example 9
[0171] The method of Examples 1-8, further comprising: performing a statistical thresholding analysis to generate a sample-specific methylated threshold and a sample specific unmethylatedthreshold, wherein classifying the reads is based on the sample-specific methylated threshold and the sample-specific unmethylated threshold.
[0172] Example 10
[0173] The method of Example 9, wherein the statistical thresholding analysis comprises K- means clustering.
[0174] Example 11
[0175] The method of Examples 1-10, wherein the one or more regions are classified as methylated or unmethylated or as methylated, unmethylated, or semimethylated.
[0176] Example 12
[0177] A method for classifying a sample, comprising: identifying one or more tumor informative regions from a set of target regions based on respective methylation states of the target regions; comparing the methylation states of respective regions of a sample corresponding to the tumor informative regions; quantifying a tumor signal based on the comparison between the methylation states of respective regions to the tumor informative regions; and assigning a classification to the sample based on the quantified tumor signal.
[0178] Example 13
[0179] The method of Example 12, wherein target regions of the set of target regions comprise one or more CpG islands.
[0180] Example 14
[0181] The method of Examples 12 or 13, wherein the methylation states of the set of target regions are classified based methylation states assigned to sequencing reads.
[0182] Example 15
[0183] The method of Examples 12-14, wherein quantifying the tumor signal is based at least in part on a target threshold that defines one or both of a threshold coverage or a threshold number of CpGs having a tumor informative state that is indicative of a tumor.
[0184] Example 16
[0185] The method of Examples 12-15, wherein quantifying the tumor signal comprises generating one or more quantitative metrics based on detected target regions having an adverse methylation state at tumor informative regions.
[0186] Example 17
[0187] The method of Examples 12-16, wherein the classification is indicative of the presence or absence of a tumor.
[0188] Example 18
[0189] The method of Examples 12-17, wherein assigning the classification is based on a threshold number of tumor informative regions in the sample contributing to the tumor signal and having a threshold intensity.
[0190] Example 19
[0191] The method of Examples 12-18, further comprising: determining one or both of a personalized treatment type or protocol based on the classification assigned to the sample.
[0192] Example 20
[0193] The method of Examples 12-18, further comprising: determining an efficacy of an ongoing or prior clinical treatment based on the classification assigned to the sample.
[0194] VIII. Miscellaneous
[0195] While the foregoing examples are provided in the context of a system (100) that may be used in nucleotide sequencing processes, and in particular the sequencing of methylation data, the teachings herein may also be readily applied in other contexts, including in systems that perform other processes (i.e., other than nucleotide sequencing procedures). The teachings herein are thus not necessarily limited to systems that are used to perform nucleotide sequencing processes.
[0196] It is to be understood that the subject matter described herein is not limited in its application to the details of construction and the arrangement of components set forth in thedescription herein or illustrated in the drawings hereof. The subject matter described herein is capable of other implementations and of being practiced or of being carried out in various ways. Also, it is to be understood that the phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. As used herein, an element or step recited in the singular and proceeded with the word “a” or “an” should be understood as not excluding plural of said elements or steps, unless such exclusion is explicitly stated. Furthermore, references to “one example” are not intended to be interpreted as excluding the existence of additional examples that also incorporate the recited features. The use of “including,” “comprising,” or “having” and variations thereof herein is meant to encompass the items listed thereafter and equivalents thereof as well as additional items.
[0197] When used in the claims, the term “set” should be understood as one or more things which are grouped together. Similarly, when used in the claims “based on” should be understood as indicating that one thing is determined at least in part by what it is specified as being “based on.” Where one thing is required to be exclusively determined by another thing, then that thing will be referred to as being “exclusively based on” that which it is determined by.
[0198] Unless specified or limited otherwise, the terms “mounted,” “connected,” “supported,” and “coupled” and variations thereof are used broadly and encompass both direct and indirect mountings, connections, supports, and couplings. Further, “connected” and “coupled” are not restricted to physical or mechanical connections or couplings. Also, it is to be understood that phraseology and terminology used herein with reference to device or element orientation (such as, for example, terms like “above,” “below,” “front,” “rear,” “distal,” “proximal,” and the like) are only used to simplify description of one or more examples described herein, and do not alone indicate or imply that the device or element referred to must have a particular orientation. In addition, terms such as “outer” and “inner” are used herein for purposes of description and are not intended to indicate or imply relative importance or significance.
[0199] It is to be understood that the above description is intended to be illustrative, and not restrictive. For example, the above-described examples (and / or aspects thereof) may be used in combination with each other. In addition, many modifications may be made to adapt a particular situation or material to the teachings of the presently described subject matter without departing from its scope. While the dimensions, types of materials and coatings described herein are intendedto define the parameters of the disclosed subject matter, they are by no means limiting and instead illustrations. Many further examples will be apparent to those of skill in the art upon reviewing the above description. The scope of the disclosed subject matter should, therefore, be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. In the appended claims, the terms “including” and “in which” are used as the plain- English equivalents of the respective terms “comprising” and “wherein.” Moreover, in the following claims, the terms “first,” “second,” and “third,” etc. are used merely as labels, and are not intended to impose numerical requirements on their objects. Further, the limitations of the following claims are not written in means — plus-function format and are not intended to be interpreted based on 35 U.S.C. § 112(f) paragraph, unless and until such claim limitations expressly use the phrase “means for” followed by a statement of function void of further structure.
[0200] The following claims recite aspects of certain examples of the disclosed subject matter and are considered to be part of the above disclosure. These aspects may be combined with one another.
Claims
What is claimed is:
1. A method for classifying one or more regions of a DNA sequence based on methylation state, comprising: accessing or acquiring signals corresponding to a set of methylation sequencing data comprising a plurality of reads within or overlapping with the one or more regions of the DNA sequence; processing the signals to classify reads of the plurality of reads based on methylation state; and classifying the one or more regions of the DNA sequence based on the methylation state of reads in or overlapping with each respective region.
2. The method of claim 1, wherein the one or more regions comprise one or more CpG islands.
3. The method of claims 1, wherein the methylation sequencing data comprises bisulfite sequencing (BiS) data or enzymatic methyl sequencing (EMS) data.
4. The method of claim 1, wherein the plurality of reads are paired-end reads of the one or more regions.
5. The method of claim 1, wherein accessing or acquiring the signals corresponding to the set of methylation sequencing data comprises extracting paired-end sequencing data using an automated script configured to process a sequencing data file generated based on an output of a sequencing system.
6. The method of claim 1, further comprising: performing a filtering step on the plurality of reads prior to classifying the reads based on methylation state, wherein the filtering step comprises one or both of: filtering based on a minimum number of reads; orfiltering based on percent overlap with a respective region of the one or more regions.
7. The method of claim 1, wherein classifying reads comprises classifying reads as methylated or unmethylated.
8. The method of claim 1, wherein classifying reads comprises classifying reads as methylated, unmethylated, or semimethylated.
9. The method of claim 1, further comprising: performing a statistical thresholding analysis to generate a sample-specific methylated threshold and a sample specific unmethylated threshold; wherein classifying the reads is based on the sample-specific methylated threshold and the sample-specific unmethylated threshold.
10. The method of claim 9, wherein the statistical thresholding analysis comprises K- means clustering.
11. The method of claim 1, wherein the one or more regions are classified as methylated or unmethylated or as methylated, unmethylated, or semimethylated.
12. A method for classifying a sample, comprising: identifying one or more tumor informative regions from a set of target regions based on respective methylation states of the target regions; comparing the methylation states of respective regions of a sample corresponding to the tumor informative regions; quantifying a tumor signal based on the comparison between the methylation states of respective regions to the tumor informative regions; and assigning a classification to the sample based on the quantified tumor signal.
13. The method of claim 12, wherein target regions of the set of target regions comprise one or more CpG islands.
14. The method of claim 12, wherein the methylation states of the set of target regions are classified based methylation states assigned to sequencing reads.
15. The method of claim 12, wherein quantifying the tumor signal is based at least in part on a target threshold that defines one or both of a threshold coverage or a threshold number of CpGs having a tumor informative state that is indicative of a tumor.
16. The method of claim 12 wherein quantifying the tumor signal comprises generating one or more quantitative metrics based on detected target regions having an adverse methylation state at tumor informative regions.
17. The method of claim 12, wherein the classification is indicative of the presence or absence of a tumor.
18. The method of claim 12, wherein assigning the classification is based on a threshold number of tumor informative regions in the sample contributing to the tumor signal and having a threshold intensity.
19. The method of claim 12, further comprising: determining one or both of a personalized treatment type or protocol based on the classification assigned to the sample.
20. The method of claim 12, further comprising: determining an efficacy of an ongoing or prior clinical treatment based on the classification assigned to the sample.
Citation Information
Patent Citations
Flow cells with hydrogel coating
US10919033B2
Recombinase polymerase amplification
US7270981B2
Method of preparing libraries of template polynucleotides
US7741463B2
Cell-free DNA methylation patterns for disease and condition analysis
WO2017212428A1
Methylation markers and targeted methylation probe panel
WO2020069350A1