Precision, recall, and f-score for continuous distributions
By deriving precision, recall, and F-score metrics for continuously distributed methylation signal data, the method and system provide accurate evaluation of automated routines, addressing the limitations of binary distribution-based assessments.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- ILLUMINA INC
- Filing Date
- 2025-10-29
- Publication Date
- 2026-05-07
AI Technical Summary
Existing methods for evaluating the accuracy of processor-executable routines for methylation signal data, which have a continuous distribution, are inadequate as they rely on binary distribution-based precision, recall, and F-score metrics, leading to complex and inaccurate assessments.
A method and system for deriving precision, recall, and F-score metrics based on continuously distributed methylation signal data, allowing for accurate evaluation of automated routines operating on nucleic acid samples.
Enables precise and reliable assessment of the accuracy of automated routines in processing methylation signal data, providing a comparative measure of their performance.
Smart Images

Figure US2025053166_07052026_PF_FP_ABST
Abstract
Description
PRECISION, RECALL, AND F-SCORE FOR CONTINUOUS DISTRIBUTIONSTECHNICAL FIELD
[0001] The present disclosure relates generally to the derivation and use of precision, recall, and F-score metrics in the context of data having a continuous distribution. Such precision and recall metrics may in turn be used in the evaluation or assessment of processor-executable routines configured to characterize the data, such as making calls related to a state of an underlying physical construct (e.g., a methylation state of all or part of a nucleic acid strand). Aspects of the present disclosure relate generally to devices, systems, and methods for deriving and utilizing such precision, recall, and F-score metrics.BACKGROUND
[0002] Various protocols in biological or chemical research involve performing a large number of controlled reactions on local support surfaces or within predefined reaction chambers. The designated reactions may then be observed or detected, and subsequent analysis may help identify or reveal properties of chemicals involved in the reaction. For example, in some multiplex assays, an unknown analyte having an identifiable label (e.g., fluorescent label) may be exposed to thousands of known probes under controlled conditions. Each known probe may be deposited into a corresponding well of a flow cell channel. Observing any chemical reactions that occur between the known probes and the unknown analyte within the wells may help identify or reveal properties of the analyte. Other examples of such protocols include known DNA sequencing processes, such as sequencing-by-synthesis (SBS) or cyclic-array sequencing.
[0003] While a variety of devices, systems, and methods have been made and used to perform biological or chemical analysis, it is believed that no one prior to the inventor(s) has made or used the devices and techniques described herein.
[0004] With the preceding in mind, the various processes in the human body are complex in nature, having both genetic and epigenetic components. By way of example, cancer development is one such complex process, as are processes related to embryonic and fetal development. Withrespect to epigenetic mechanisms, one such mechanism is methylation, which often occurs via the addition of a methyl group at the 5' carbon of cytosine residues. This methyl addition to cytosine residues often occurs at CpG dinucleotides. Regions of a genome that contain CpG dinucleotides at a higher frequency than expected are often characterized as CpG islands. Such methylation is generally understood or believed to serve a regulatory or control function with respect to gene expression, with methylation or increased methylation of such CpG islands often being associated with repression of expression of the gene in question (i.e., turning the gene “off’). Correspondingly, abnormalities in the methylation associated with epigenetic control processes may be relevant to numerous disorders and disease states, including but not limited to cancer.
[0005] The techniques used to sequence such regions and derive a methylation signal that may be used to evaluate such regions typically rely on processor-executable routines or software to evaluate the methylation signals generated by the sequencing process. Such routines are useful due to the complexity and quantity of the data involved and the need for timely analysis. However, evaluating the sufficiency (e.g., accuracy or other quality metrics) of such routines in their operation may be difficult. In particular, three metrics to evaluate such metrics that are commonly employed are precision, recall, and F-score. However, the techniques for deriving these metrics are based on the underlying data having a binary distribution. Methylation signal data, however, has a continuous distribution and thus precision, recall, and F-score metrics may not be directly derived from such signals. With this in mind, assessment of the routines used to process and evaluate such data may be complex and / or inaccurate.SUMMARY
[0006] A summary of certain embodiments disclosed herein is set forth below. It should be understood that these aspects are presented merely to provide the reader with a brief summary of these certain embodiments and that these aspects are not intended to limit the scope of this disclosure. Indeed, this disclosure may encompass a variety of aspects that may not be set forth below.
[0007] In accordance with certain embodiments, a method is provided for deriving one or more of a precision, a recall, or an F-score metric. In accordance with this method, a set of data is acquired or accessed. The data is continuously or non-negative continuously distributed. Basedon an expected signal and an observed signal, one or more of a precision, a recall, or an F-score are determined for one or more automated routines configured to operate on the set of data. One or more of the precision, the recall, or the F-score provide a comparative measure of the accuracy of the one or more automated routines.
[0008] In accordance with further embodiments, a method is provided for deriving one or more of a precision, a recall, or an F-score metric based on a differentially methylated signal (DMS). In accordance with this method, a DMS for a set of nucleic acid samples is accessed or acquired. The DMS is continuously or non-negative continuously distributed, based on an expected DMS and an observed DMS, one or more of a precision, a recall, or an F-score are determined for one or more automated routines configured to operate on the DMS to determine one or more differentially methylated regions (DMRs). One or more of the precision, the recall, or the F-score provide a comparative measure of the accuracy of the one or more automated routines at determining DMRs.
[0009] In accordance with additional embodiments, a processor-based system is provided. In accordance with such embodiments, the processor-based system comprises at least one processor configured to execute stored routines and one or more tangible storage media storing processorexecutable routines. The processor-executable routines, when executed by the at least one processor, cause acts to be performed comprising: acquiring or accessing a set of data, wherein the data is continuously or non-negative continuously distributed, and based on an expected signal and an observed signal, determining one or more of a precision, a recall, or an F-score for one or more automated routines configured to operate on the set of data. One or more of the precision, the recall, or the F-score provide a comparative measure of the accuracy of the one or more automated routines.
[0010] Various refinements of the features noted above may exist in relation to various aspects of the present disclosure. Further features may also be incorporated in these various aspects as well. These refinements and additional features may exist individually or in any combination. For instance, various features discussed below in relation to one or more of the illustrated embodiments may be incorporated into any of the above-described aspects of the present disclosure alone or in any combination. The brief summary presented above is intended only to familiarize the reader with certain aspects and contexts of embodiments of the present disclosure without limitation to the claimed subject matter.BRIEF DESCRIPTION OF THE DRAWINGS
[0011] FIG. 1 depicts a schematic view of an example of a system that may be used to provide biological or chemical analysis, in accordance with aspects of the present disclosure.
[0012] FIG. 2 depicts a cross-sectional view of an example of a flow cell that may be used in the system of FIG. 1, in accordance with aspects of the present disclosure.
[0013] FIG. 3 depicts a schematic view of another example of a system that may be used to provide biological or chemical analysis, in accordance with aspects of the present disclosure.
[0014] FIG. 4 depicts a schematic view of an example of a networked system in which the system of FIG. 1 or 3 may be incorporated, in accordance with aspects of the present disclosure.
[0015] FIG. 5 depicts a schematic view of an example of a base calling arrangement that may be carried out using the system of FIG. 1, 3, or 4, in accordance with aspects of the present disclosure.
[0016] FIG. 6 depicts a schematic view of an example of a base caller training technique that may be implemented using the system of FIG. 1, 3, or 4, in accordance with aspects of the present disclosure.
[0017] FIG. 7 depicts a plot of expected and observed differentially methylated signals (DMS), in accordance with aspects of the present disclosure.
[0018] FIG. 8 depicts a plot of a conventional binary data distribution defining a possible precision and recall space, in accordance with aspects of the present disclosure.
[0019] FIG. 9 depicts a plot extending the precision and recall space to continuous distribution data in accordance with aspects of the present disclosure.
[0020] FIG. 10 depicts the plot of FIG. 9 updated to include the angular range cp as may be used in the generalization of precision and recall to data having a continuous distribution, in accordance with aspects of the present disclosure.DETAILED DESCRIPTION
[0021] The following detailed description of certain examples will be better understood when read in conjunction with the appended drawings. To the extent that the figures illustrate diagrams of the functional blocks of various examples, the functional blocks are not necessarily indicative of the division between hardware components. Thus, for example, one or more of the functional blocks (e.g., processors or memories) may be implemented in a single piece of hardware (e.g., a general-purpose signal processor or random-access memory, hard disk, or the like). Similarly, the programs may be stand-alone programs, may be incorporated as subroutines in an operating system, may be functions in an installed software package, and the like. It should be understood that the various examples are not limited to the arrangements and instrumentality shown in the drawings.
[0022] I. Overview of System for Biological or Chemical Analysis
[0023] Examples described herein may be used in various biological or chemical processes and systems for academic analysis, commercial analysis, or other analysis. More specifically, examples described herein may be used in various processes and systems where it is desired to detect an event, property, quality, or characteristic that is indicative of a designated reaction (e.g., methylation). Bioassay systems such as those described herein may be configured to perform a plurality of designated reactions that may be detected individually or collectively. For example, bioassay systems may be used to sequence a dense array of nucleic acid features through iterative cycles of enzymatic manipulation and image acquisition. In some examples, nucleic acids can be attached to a surface and amplified. Examples of such amplification are described in U. S. Pat. No.7,741,463, entitled “Method of Preparing Libraries of Template Polynucleotides,” issued June 22, 2010, the disclosure of which is incorporated by reference herein, in its entirety; and / or U. S. Pat. No. 7,270,981, entitled “Recombinase Polymerase Amplification,” issued September 18, 2007, the disclosure of which is incorporated by reference herein, in its entirety.
[0024] Components that are used in the bioassay systems may include one or more microfluidic channels that deliver reagents or other reaction components to a reaction site. The reaction sites may be randomly distributed across a substantially planar surface; or may be patterned across a substantially planar surface. Each of the reaction sites may be imaged to detectlight from the reaction site. The signals indicating photons emitted from the reaction sites and detected by image sensors may provide illumination values. These illumination values may be combined into an image indicating photons as detected from the reaction sites. These images may be further analyzed to identify compositions, reactions, conditions, etc., at each reaction site.
[0025] II. Examples of Fluidics Devices and Fluid Flow Paths - Example of System with Higher Volume Throughput
[0026] FIG. 1 illustrates a schematic diagram of an example of a system (100) that may be used to perform an analysis on one or more samples of interest. In some implementations, the sample may include one or more clusters of nucleotides (e.g., DNA) that have been linearized to form a single stranded DNA (sstDNA). In the implementation shown, system (100) is configured to receive a flow cell cartridge assembly (102) including a flow cell assembly (103) and a sample cartridge (104). System (100) includes a flow cell receptacle (122) that receives flow cell cartridge assembly (102), a vacuum chuck (124) that supports flow cell assembly (103), and a flow cell interface (126) that is used to establish a fluidic coupling between system (100) and flow cell assembly (103). Flow cell interface (126) may include one or more manifolds. System (100) further includes a sipper manifold assembly (106), a sample loading manifold assembly (108), and a pump manifold assembly (110). System (100) also includes a drive assembly (112), a controller (114), an imaging system (116), and a waste reservoir (118). Controller (114) is electrically and / or communicatively coupled to drive assembly (112) and to imaging system (116); and is configured to cause drive assembly (112) and / or the imaging system (116) to perform various functions as disclosed herein.
[0027] In the present example, flow cell assembly (103) includes a flow cell (128) having a channel (130) and defining a plurality of first openings (132), which are fluidically coupled to the channel (130) and arranged on a first side (134) of the channel (130). Flow cell (128) further includes a plurality of second openings (136) fluidically coupled to the channel (130) and arranged on a second side (138) of the channel (130). Fluid may thus flow through flow cell (128) via channel. While the flow cell (128) is shown including one channel (130), flow cell (128) may include two or more channels (130). Flow cell assembly (103) also includes a flow cell manifold assembly (140) coupled to flow cell (128) and having a first manifold fluidic line (142) and a second manifold fluidic line (144). Flow cell manifold assembly (140) may be in the form of alaminate including a plurality of layers as discussed in more detail below.
[0028] In the implementation shown, first manifold fluidic line (142) has a first fluidic line opening (146) and is fluidically coupled to each of the first openings (132) of flow cell (128); and second manifold fluidic line (144) has a second fluidic line opening (148) and is fluidically coupled to each of the second openings (136). As shown, flow cell assembly (103) includes gaskets (150) coupled to flow cell manifold assembly (140) and fluidically coupled to fluidic line openings (146, 148). In some implementations where flow cell (128) includes a plurality of channels (130), flow cell manifold assembly (140) may include additional fluidic lines (152) that couple first fluidic line openings (146) to a single manifold port (154). In such implementations, a single gasket (150) may be coupled to flow cell manifold assembly (140) that surrounds the manifold port (154) and is in fluidic communication with a plurality of channels (130). In operation, flow cell interface (126) engages with corresponding gaskets (150) to establish a fluidic coupling between system (100) and flow cell (128). The engagement between flow cell interface (126) and gaskets (150) reduces or eliminates fluid leakage between flow cell interface (126) and flow cell (128).
[0029] In the implementation shown, first manifold fluidic line (142) has a portion (156) that is substantially parallel to a longitudinal axis (158) of channel (130); and second manifold fluidic line (144) has a portion (160) that is substantially parallel to longitudinal axis (158) of channel (130). Additionally, first manifold fluidic line (142) is shown being at least partially adjacent a first end (162) of flow cell (128) and spaced from a second end (164) of flow cell (128); and second manifold fluidic line (144) is shown being at least partially adjacent second end (164) of flow cell (128) and spaced from first end (162). Other arrangements of manifold fluidic lines (142, 144) may prove suitable, however.
[0030] In the implementation shown, system (100) includes a sample cartridge receptacle (166) that receives sample cartridge (104) that carries one or more samples of interest (e.g., an analyte). System (100) also includes a sample cartridge interface (168) that establishes a fluidic connection with sample cartridge (104). Sample loading manifold assembly (108) includes one or more sample valves (170). Pump manifold assembly (110) includes one or more pumps (172), one or more pump valves (174), and a cache (176). Valves (170, 174) and pumps (172) may take any suitable form. Cache (176) may include a serpentine cache and may temporarily store one or more reaction components during, for example, bypass manipulations of the system (100). Whilecache (176) is shown being included in pump manifold assembly (110), cache (176) may alternatively be located elsewhere (e.g., in sipper manifold assembly (106) or in another manifold downstream of a bypass fluidic line (178), etc.).
[0031] Sample loading manifold assembly (108) and pump manifold assembly (110) flow one or more samples of interest from sample cartridge (104) through a fluidic line (180) toward flow cell cartridge assembly (102). In some implementations, sample loading manifold assembly (108) may individually load or address each channel (130) of flow cell (128) with a respective sample of interest. The process of loading channel (130) with a sample of interest may occur automatically using system (100). As shown in FIG. 1, sample cartridge (104) and sample loading manifold assembly (108) are positioned downstream of flow cell cartridge assembly (102). In the implementation shown, sample loading manifold assembly (108) is coupled between flow cell cartridge assembly (102) and pump manifold assembly (110). To draw a sample of interest from sample cartridge (104) and toward pump manifold assembly (110), sample valves (170), pump valves (174), and / or pumps (172) may be selectively actuated to urge the sample of interest toward pump manifold assembly (110). Sample cartridge (104) may include a plurality of sample reservoirs that are selectively fluidically accessible via the corresponding sample valves (170). To individually flow the sample of interest toward channel (130) of flow cell (128) and away from pump manifold assembly (110), sample valves (170), pump valves (174), and / or pumps (172) may be selectively actuated to urge the sample of interest toward flow cell cartridge assembly (102) and into respective channels (130) of flow cell (128).
[0032] Drive assembly (112) interfaces with sipper manifold assembly (106) and pump manifold assembly (110) to flow one or more reagents that interact with the sample within flow cell (128). In some scenarios, a reversible terminator is attached to the reagent to allow a single nucleotide to be incorporated onto a growing DNA strand. In some such implementations, one or more of the nucleotides has a unique fluorescent label that emits a color when excited. The color (or absence thereof) is used to detect the corresponding nucleotide. In the implementation shown, imaging system (116) excites one or more of the identifiable labels (e.g., a fluorescent label) and thereafter obtains image data for the identifiable labels. The labels may be excited by incident light and / or a laser and the image data may include one or more colors emitted by the respective labels in response to the excitation. The image data (e.g., detection data) may be analyzed by system(100). Examples of features and functionalities that may be incorporated into imaging system (116) will be described in greater detail below.
[0033] After the image data is obtained, drive assembly (112) interfaces with sipper manifold assembly (106) and pump manifold assembly (110) to flow another reaction component (e.g., a reagent) through flow cell (128) that is thereafter received by waste reservoir (118) via a primary waste fluidic line (182) and / or otherwise exhausted by system (100). Some reaction components may perform a flushing operation that chemically cleaves the fluorescent label and the reversible terminator from the sstDNA. The sstDNA may then be ready for another cycle.
[0034] The primary waste fluidic line (182) is coupled between pump manifold assembly (110) and waste reservoir (118). In some implementations, pumps (172) and / or pump valves (174) of pump manifold assembly (110) selectively flow the reaction components from flow cell cartridge assembly (102), through fluidic line (180) and sample loading manifold assembly (108) to primary waste fluidic line (182). Flow cell cartridge assembly (102) is coupled to a central valve (184) via flow cell interface (126). Central valve (184) is coupled with flow cell interface (126) via a fluidic line (185). An auxiliary waste fluidic line (186) is coupled to central valve (184) and to waste reservoir (118). In some implementations, auxiliary waste fluidic line (186) receives excess fluid of a sample of interest from flow cell cartridge assembly (102), via central valve (184), and flows the excess fluid of the sample of interest to waste reservoir (118) when back loading the sample of interest into flow cell (128), as described herein.
[0035] Sipper manifold assembly (106) includes a shared line valve (188) and a bypass valve (190). Shared line valve (188) may be referred to as a reagent selector valve. Central valve (184) and the valves (188, 190) of sipper manifold assembly (106) may be selectively actuated to control the flow of fluid through fluidic lines (192, 194, 196). Sipper manifold assembly (106) may be coupled to a corresponding number of reagent reservoirs (198) via reagent sippers (200). Reagent reservoirs (198) may contain fluid (e g., reagent and / or another reaction component). In some implementations, sipper manifold assembly (106) includes a plurality of ports. Each port of sipper manifold assembly (106) may receive one of the reagent sippers (200). Reagent sippers (200) may be referred to as fluidic lines. Some forms of reagent sippers (200) may include an array of sipper tubes extending downwardly along the z-dimension from ports in the body of sipper manifold assembly (106). Reagent reservoirs (198) may be provided in a cartridge, and the tubes of reagentsippers (200) may be configured to be inserted into corresponding reagent reservoirs (198) in the reagent cartridge so that liquid reagent may be drawn from each reagent reservoir (198) into the sipper manifold assembly (106).
[0036] Shared line valve (188) of sipper manifold assembly (106) is coupled to central valve (184) via shared reagent fluidic line (196). Different reagents may flow through shared reagent fluidic line (196) at different times. In some versions, when performing a flushing operation before changing between one reagent and another, pump manifold assembly (110) may draw wash buffer through shared reagent fluidic line (196), central valve (184), and flow cell cartridge assembly (102).
[0037] Bypass valve (190) of sipper manifold assembly (106) is coupled to central valve (184) via dedicated reagent fluidic lines (194, 196). Each of the dedicated reagent fluidic lines (194, 196) may be associated with a single reagent. The fluids that may flow through dedicated reagent fluidic lines (194, 196) may be used during sequencing operations and may include a cleave reagent, an incorporation reagent, a scan reagent, a cleave wash, and / or a wash buffer.
[0038] Bypass valve (190) is also coupled to cache (176) of pump manifold assembly (110) via bypass fluidic line (178). One or more reagent priming operations, hydration operations, mixing operations, and / or transfer operations may be performed using bypass fluidic line (178). The priming operations, the hydration operations, the mixing operations, and / or the transfer operations may be performed independent of flow cell cartridge assembly (102). Thus, the operations using bypass fluidic line (178) may occur during, for example, incubation of one or more samples of interest within flow cell cartridge assembly (102). That is, shared line valve (188) may be utilized independently of bypass valve (190) such that bypass valve (190) may utilize bypass fluidic line (178) and / or cache (176) to perform one or more operations while shared line valve (188) and / or central valve (184) simultaneously, substantially simultaneously, or offset synchronously perform other operations.
[0039] Drive assembly (112) includes a pump drive assembly (202) and a valve drive assembly (204). Pump drive assembly (202) may be adapted to interface with one or more pumps (172) to pump fluid through flow cell (128) and / or to load one or more samples of interest into flow cell (128). Valve drive assembly (204) may be adapted to interface with one or more of the valves(170, 174, 184, 188, 190) to control the position of the corresponding valves (170, 174, 184, 188, 190).
[0040] Controller (114) of the present example includes a user interface (206), a communication interface (208), one or more processors (210), and a tangible memory (212) storing instructions executable by the one or more processors (210) to perform various functions including the disclosed implementations. User interface (206), communication interface (133), and memory (212) are electrically and / or communicatively coupled to the one or more processors (210) that may in practice receive, process, and / or transform various signals generated by sensors and / or other signal generating components of the system into other processed or revised signals that may in turn be communicated (e.g., transmitted), processed and / or transformed as part of the techniques discussed herein. User interface (206) may be adapted to receive input (e.g., input or interface signals) from a user and to provide information to the user associated with the operation of system (100) and / or an analysis taking place. User interface (206) may include a touch screen, a display, a keyboard, a speaker(s), a mouse, a track ball, and / or a voice recognition system.
[0041] Communication interface (208) is adapted to enable communication between system (100) and a remote system(s) (e.g., computers) via a network(s) (e.g., the Internet, an intranet, a local-area network (LAN), a wide-area network (WAN), a coaxial-cable network, a wireless network, a wired network, a satellite network, a digital subscriber line (DSL) network, a cellular network, a Bluetooth connection, a near field communication (NFC) connection, etc.). Some of the communications provided to the remote system may take the form of respective signals associated with analysis results, imaging data, etc. generated or otherwise obtained by system (100). Some of the communications provided to system (100) may take the form of respective signals associated with a fluidics analysis operation, patient records, and / or a protocol(s) to be executed by system (100).
[0042] The one or more processors (210) and / or system (100) may include one or more of a processor-based system(s) or a microprocessor-based system(s). In some implementations, the one or more processors (210) and / or system (100) includes one or more of a programmable processor, a programmable controller, a microprocessor, a microcontroller, a graphics processing unit (GPU), a digital signal processor (DSP), a reduced-instruction set computer (RISC), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a field programmable logicdevice (FPLD), a logic circuit, and / or another logic-based device executing various functions including the ones described herein.
[0043] Memory (212) may include one or more of a semiconductor memory, a magnetically readable memory, an optical memory, a hard disk drive (HDD), an optical storage drive, a solid-state storage device, a solid-state drive (SSD), a flash memory, a read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable readonly memory (EEPROM), a random-access memory (RAM), a non-volatile RAM (NVRAM) memory, a compact disc (CD), a compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a Blu-ray disk, a redundant array of independent disks (RAID) system, a cache and / or any other storage device or storage disk in which information is stored for any duration (e.g., permanently, temporarily, for extended periods of time, for buffering, for caching).
[0044] III. Examples of Flow Cell Structures
[0045] As noted above, a system (100) may execute reactions in a flow cell (128) and / or perform analysis on one or more samples of interest in a flow cell (128). The following describes examples of forms that such flow cells (128) may take, it being understood that flow cells (128) may take various other forms and have various other features in addition to or in lieu of the features described below.
[0046] Example of Single-Surface Patterned Flow Cell
[0047] FIG. 2 shows an example of a flow cell (400) that includes a patterned substrate (402), which includes depressions (404) separated by interstitial regions (406), and surface chemistry (410, 412) positioned in the depressions (404). Depressions (404) may be in the form of microwells or nanowells. Depressions (404) may be configured to contain nucleic acid strands or other oligonucleotides and thereby provide a reaction site for SBS and / or for other kinds of processes. In some versions, each depression (404) has a cylindraceous configuration, with a generally circular cross-sectional profile. In some other versions, each depression (404) has a polygonal (e.g., hexagonal, octagonal, square, rectangular, elliptical, etc.) cross-sectional profile. Alternatively, depressions (404) may have any other suitable configuration. It should also be understood that depressions (404) may be arranged in any suitable pattern, including but not limited to a grid pattern.
[0048] Surface chemistry (410, 412) of the present example includes functionalized coating layer (410) and primers (412). While not shown, it is to be understood that the depressions (404) may also have surface preparation or treatment chemistry (e.g., silane or a silane derivative) positioned between the substrate (402) and the functionalized coating layer (410). This same surface preparation or treatment chemistry may also be positioned on the interstitial regions (406). In the present example, a hydrogel (440) is applied before lid (420) is bonded to substrate (402). Hydrogel (440) covers surface chemistry (410, 412) in depressions (404), and at least a portion of the patterned substrate (402) (e.g., those interstitial regions (406) that are not also bonding regions (422)). By way of example only, hydrogel (440) may comprise PAZAM, crosslinked polyacrylamide, agarose gel, etc.
[0049] Flow cell (400) of this example further includes a lid (420) bonded to bonding region(s) (422) of patterned substrate (402). In the example shown in FIG. 2, lid (420) includes a top portion (424) that is connected to several sidewalls (426), and these components (424, 426) define a portion of each of the six flow channels (430A, 430B, 430C, 430D, 430E, 430F). The respective sidewalls (426) isolate one flow channel (430A, 430B, 430C, 430D, 430E, 430F) from each adjacent flow channel (430A, 430B, 430C, 430D, 430E, 430F). Each flow channel (430A, 430B, 430C, 430D, 430E, 430F) is in selective fluid communication with a respective set of depressions (404).
[0050] Lid (420) may be bonded to bonding region (422) of substrate (402) using any suitable technique, such as laser bonding, diffusion bonding, anodic bonding, eutectic bonding, plasma activation bonding, glass frit bonding, or other methods known in the art. In some versions, a spacer layer (428) may be used to bond lid (420) to bonding region (422). Spacer layer (428) may comprise any material that will seal at least some of interstitial regions (404) (e.g., bonding region (422)) of substrate (402) and lid (420) together. While not shown, lid (420) or the patterned substrate (402) may include inlet and outlet ports that are to fluidically engage other ports (not shown), such as those of sample cartridge interface (168), for directing fluid(s) into the respective flow channels (430A, 430B, 430C, 430D, 430E, 430F) (e.g., from a reagent cartridge or other fluid storage system) and out of the flow channel (e.g., to waste reservoir (118) or another waste removal system). Flow channels (430A, 430B, 430C, 430D, 430E, 430F) may serve to, for example, selectively introduce reaction components or reactants to hydrogel (440) and theunderlying surface chemistry (410, 412) in order initiate designated reactions in / at depressions (404).
[0051] While flow cell (400) includes a pattern of depressions (404) to provide an array of reaction sites, other variations may provide reaction sites on or at various other kinds of structural features, including but not limited to continuously planer surfaces and / or protruding surfaces, etc. By way of further example only, flow cell (400) may be constructed and operable in accordance with at least some of the teachings of U. S. Pat. No. 10,919,033, entitled “Flow Cells with Hydrogel Coating,” issued February 16, 2021, the disclosure of which is incorporated by reference herein, in its entirety.
[0052] IV. Examples of Imaging System Features
[0053] As noted above, system (100) includes an imaging system (116) that excites one or more identifiable labels (e.g., a fluorescent label) in samples in reaction sites provided by depressions (404, 462, 464) of a flow cell (128, 400); and thereafter obtains image data for the identifiable labels. This image data is used to identify nucleotides as part of a nucleic acid sequencing process. Alternatively, the image data may be used for various other purposes. The following description provides details on how some versions of imaging system (116) may be configured and operable.
[0054] FIG. 3 illustrates a schematic diagram of another example of a system (500) that may be used to perform an analysis on one or more samples of interest. Except as otherwise described below, system (500) of this example may be configured and operable like systems (100) described above. System (500) is configured to perform a large number of parallel reactions within a flow cell (510). Flow cell (510) may be configured and operable like flow cells (400) described above or may have any other suitable configuration. Flow cell (510) may thus include one or more flow channels that receive a solution from system (500) and direct the solution toward reaction sites of flow cell (510).
[0055] System (500) includes a system controller (520) that may communicate via respective signals with the various components, assemblies, and sub-systems of the system (500). Controller (520) may be configured and operable like controllers (114) described above. An imaging assembly (522) of system (500) includes a light emitting assembly (550) that, responsive to controlsignals, emits light that reaches reaction sites on flow cell (510). Light emitting assembly (550) may include an incoherent light emitter (e.g., emit light beams output by one or more excitation diodes), or a coherent light emitter such as emitter of light output by one or more lasers or laser diodes. In some implementations, light emitting assembly (550) may include a plurality of different light sources (not shown), each light source emitting light of a different wavelength range. Some versions of light emitting assembly (550) may also include one or more collimating lenses (not shown), a light structuring optical assembly (not shown), a projection lens (not shown) that is operable to adjust a structured beam shape and path, epifluorescence microscopy components, and / or other components. Although system (500) is illustrated as having a single light emitting assembly (550), multiple light emitting assemblies (550) may be included in some other implementations.
[0056] In the present example, the light from light emitting assembly (550) is directed by dichroic mirror assembly (546) through an objective lens assembly (542) onto a sample of a flow cell (510), which is positioned on a motion stage (570). In the case of fluorescent microscopy of a sample, a fluorescent element associated with the sample of interest fluoresces in response to the excitation light, and the resultant light is collected by objective lens assembly (542) and is directed to an image sensor of camera system (540) to detect the emitted fluorescence and to generate or transmit a corresponding fluorescence or color signal for downstream processing in accordance with techniques discussed herein. In some implementations, a tube lens assembly may be positioned between the objective lens assembly (542) and the dichroic mirror assembly (546) or between the dichroic mirror (546) and the image sensor of the camera system (540). A moveable lens element may be translatable along a longitudinal axis of the tube lens assembly to account for focusing on an upper interior surface or lower interior surface of the flow cell (510) and / or spherical aberration introduced by movement of the objective lens assembly (542).
[0057] In the present example, a filter switching assembly (544) is interposed between dichroic mirror assembly (546) and camera system (540). Filter switching assembly (544) includes one or more emission filters that may be used to pass through particular ranges of emission wavelengths and block (or reflect) other ranges of emission wavelengths. For example, emission filters may be used or controlled to direct different wavelength ranges of emitted light to different image sensors of the camera system (540) of imaging assembly (522). For instance, the emission filters may beimplemented as dichroic mirrors that direct emission light of different wavelengths from flow cell (510) to different image sensors of camera system (540). In some variations, a projection lens is interposed between filter switching assembly (544) and camera system (540). Filter switching assembly (544) may be omitted in some versions.
[0058] System (500) further includes a fluid delivery assembly (590) that may receive control signals causing the fluid delivery assembly to direct the flow of reagents (e g., fluorescently labeled nucleotides, buffers, enzymes, cleavage reagents, etc.) to (and through) flow cell (510) and waste valve (580). Fluid delivery assembly (590) may be configured and operable like the various fluid delivery components described herein. System (500) of the present example also includes a temperature station actuator (530) and heater / cooler (532) that may optionally regulate the temperature of conditions of the fluids within the flow cell (510). In some implementations, the heater / cooler (532) may be fixed to sample stage (570), upon which the flow cell (510) is placed, and / or may be integrated into sample stage (570).
[0059] Flow cell (510) may be removably mounted on sample stage (570), which may be controlled via signals to provide movement and alignment of flow cell (510) relative to objective lens assembly (542). Sample stage (570) may have one or more actuators to allow sample stage (570) to move in any of three dimensions. For example, actuators may be provided to allow sample stage (570) to move in the x, y, and z directions relative to objective lens assembly (542), tilt relative to objective lens assembly (542), and / or otherwise move relative to objective lens assembly (542). Movement of sample stage (570) may allow one or more sample locations on flow cell (510) to be positioned in optical alignment with objective lens assembly (542). Movement of sample stage (570) relative to objective lens assembly (542) may be achieved by moving sample stage (570) itself, by moving objective lens assembly (542), by moving some other component of imaging assembly (522), by moving some other component of system (500), or any combination of the foregoing. For instance, in some implementations, the sample stage (570) may be actuatable in the x and y directions relative to the objective lens assembly (542) while a focus component (562) or z-stage may move the objective lens assembly (542) along the z direction relative to the sample stage (570).
[0060] In some implementations, a focus component (562) may be included and used to control positioning of one or more elements of objective lens assembly (542) relative to the flowcell (510) in the focus direction (e.g., along the z-axis or z-dimension). Focus component (562) may include one or more actuators physically coupled to the objective lens assembly (542), the optical stage, the sample stage (570), or a combination thereof, to move flow cell (510) on sample stage (570) relative to the objective lens assembly (542) to provide proper focusing for the imaging operation. In the present example, the focus component (562) utilizes a focus tracking module (560) that is configured to detect a displacement of the objective lens assembly (542) relative to a portion of the flow cell (510) and output data indicative of an in-focus position to the focus component (562) or a component thereof or operable to control the focus component (562), such as controller (520), to move the objective lens assembly (542) to position the corresponding portion of the flow cell (510) in focus of the objective lens assembly (542).
[0061] In some implementations, an actuator of focus component (562) or for sample stage (570) may be physically coupled to objective lens assembly (542), the optical stage, sample stage (570), or a combination thereof, such as, for example, by mechanical, magnetic, fluidic, or other attachment or contact directly or indirectly to or with the stage or a component thereof. The actuator of focus component (562) may be configured to move objective lens assembly (542) in the z-direction while maintaining sample stage (570) in the same plane (e.g., maintaining a level or horizontal attitude, perpendicular to the optical axis). In some implementations, sample stage (570) includes an x direction actuator and a y direction actuator to form an x-y stage. Sample stage (570) may also be configured to include one or more tip or tilt actuators to tip or tilt sample stage (570) and / or a portion thereof, to account for any slope in its surfaces.
[0062] Camera system (540) may include one or more image sensors to monitor and track the imaging (e.g., sequencing) of flow cell (510). Camera system (540) may be implemented, for example, as a CCD or CMOS image sensor camera configured to generate and communicate camera output signals (e.g., fluorescence signals, color signals, intensity signals, and so forth), but other image sensor technologies (e.g., active pixel sensor) may be used. By way of further example only, camera system (540) may include a dual-sensor time delay integration (TDI) camera, a single-sensor camera, a camera with one or more two-dimensional image sensors, and / or other kinds of camera technologies. While camera system (540) and associated optical components are shown as being positioned above flow cell (510) in FIG. 3, one or more image sensors or other camera components may be incorporated into system (500) in numerous other ways as will beapparent to those skilled in the art in view of the teachings herein. For instance, one or more image sensors may be positioned under flow cell (510), such as within the sample stage (570) or below the sample stage (570); or may even be integrated into flow cell (510).
[0063] V. Examples of Data Processing Features
[0064] A. Example of Networked Data Processing Arrangement
[0065] As noted above, a system (100, 500) may include a controller (114, 520) that is configured to process signals corresponding to or representative of data, execute algorithms, etc., as needed to perform a sequencing operation or other kind of operation. In some scenarios, system (100, 500) may be coupled with other devices via a network to perform further data processing, data storage, execution of algorithms, etc. FIG. 4 shows an example of such an arrangement. In particular, FIG. 4 shows a networked system (800) that includes a sequencing device (810), a server device (820), a client device (830), and a local device (840), with all devices (810, 820, 830, 840) being coupled together via a network (850). Network (850) may take any suitable form as will be apparent to those skilled in the art in view of the teachings herein.
[0066] As shown in FIG. 4, sequencing device (810) comprises a computing device and a sequencing device system (812) for sequencing a genomic sample or other nucleic-acid polymer. In some versions, by executing sequencing device system (812) using a processor, sequencing device (810) analyzes nucleotide fragments or oligonucleotides extracted from genomic samples to generate nucleotide reads (or other read signals or data) utilizing computer implemented methods and systems either directly or indirectly on sequencing device (810). More particularly, sequencing device (810) receives nucleotide-sample slides (e.g., flow cells (128, 400, 510)) comprising nucleotide fragments extracted from samples and further copies and determines the nucleobase sequence of such extracted nucleotide fragments. It should be understood that sequencing device (810) may represent a version of systems (100, 500) described above.
[0067] In some versions, the sequencing device (810) utilizes SBS to sequence nucleotide fragments into nucleotide reads and determine nucleobase calls for the nucleotide reads. By executing sequencing device system (812), sequencing device (810) may further store the nucleobase calls as part of base-call data that is formatted as a binary base call (BCL) file and send the BCL file to the local device (840) and / or the server device(s) (820). Sequencing device (810)may communicate signals corresponding to the BCL file and / or other data to local device (840) and / or client device (830) via network (850) or directly (i.e., bypassing network (850)).
[0068] In some scenarios, local device (840) is located at or near a same physical location of sequencing device (810). For instance, local device (840) and sequencing device (810) may be integrated into a single computing device. Local device (840) may run sequencing system (814) to generate, receive, analyze, store, and transmit digital data signals, such as by receiving base-call signals or determining variant calls based on analyzing such base-call signals or data. By executing software in the form of sequencing system (814), local device (840) may align nucleotide reads with a structural variation graph genome (824) and determine genetic variants based on the aligned nucleotide reads. Local device (840) may also send data to client device (830), including a variant call file (VCF) or other information indicating nucleobase calls, sequencing metrics, error data, or other metrics.
[0069] Server device(s) (820) may be located remotely from the local device (840) and sequencing device (810). Server device(s) (820) may comprise a distributed collection of servers, where server device(s) (820) include a number of server devices distributed across network (850) and located in the same or different physical locations. Similar to local device (840), server device(s) (820) may include a version of sequencing system (814). Accordingly, server device(s) (820) may generate, receive, analyze, store, and transmit digital data signals, such as by receiving base-call signals or determining variant calls based on analyzing such base-call signals or data. As indicated above, sequencing device (810) may send (and server device(s) (820) may receive) basecall signals or data from sequencing device (810). Server device(s) (820) may also send signals or data to client device (830), including VCFs or other sequencing related information.
[0070] As indicated above, as part of server device(s) (820) or local device (840), sequencing system (814) may generate or implement a structural variation graph genome with alternate contiguous sequences representing structural variant haplotypes. For instance, system (814) may identify candidate structural variants of a threshold frequency (or that otherwise satisfy another occurrence threshold) within a genomic sample database. From among the candidate structural variants, sequencing system (814) selects structural variant haplotypes based on one or both of satisfying another occurrence threshold and finding flanking variants adjacent to particular structural variant haplotypes. Sequencing system (814) may likewise select reference haplotypesof genomic regions corresponding to the selected structural variant haplotypes from a reference genome. Based on the selected haplotypes, sequencing system (814) generates a structural variation graph genome comprising both alternate contiguous sequences representing the structural variant haplotypes and reference sequences representing the reference haplotypes. Based on comparing nucleotide reads of a genomic sample with alternate contiguous sequences representing structural variant haplotypes, sequencing system (814) can determine nucleobase calls for the genomic sample.
[0071] By executing a sequencing application (832), client device (830) may generate, store, receive, and send digital signal data. In particular, client device (830) may receive sequencing data from local device (840) or receive call files (e g., BCL) and sequencing metrics from sequencing device (810). Furthermore, client device (830) may communicate with local device (840) or server device(s) (820) to receive a VCF comprising nucleobase calls and / or other metrics, such as a base-call-quality metrics or pass-filter metrics. Client device (830) may accordingly present or display information pertaining to variant calls or other nucleobase calls within a graphical user interface of sequencing application (832) to a user associated with client device (830). For example, client device (830) may present structural variant calls and / or sequencing metrics for a sequenced genomic sample within a graphical user interface of sequencing application (832).
[0072] As shown in FIG. 4, sequencing application (832) is included in client device (830). Sequencing application (832) may include a web application or a native application stored and executed on client device (830) (e g., a mobile application, desktop application). Sequencing application (832) may include instructions that (when executed) cause client device (830) to receive data signals from sequencing system (814) and present, for display at client device (830), base-call data or data from a VCF. Furthermore, sequencing application (832) may instruct client device (830) to display summaries for multiple sequencing runs.
[0073] As further illustrated in FIG. 4, a version of sequencing system (814) may be located and implemented (e.g., entirely or in part) on client device (830) or sequencing device (810). In some versions, sequencing system (814) is implemented by one or more other components of networked system (800), such as local device (840). In particular, sequencing system (814) may be implemented in a variety of different ways across sequencing device (810), local device (840),server device(s) (820), and client device (830). For example, sequencing system (814) may be downloaded from server device(s) (820) to sequencing system (814) and / or local device (840) where all or part of the functionality of sequencing system (814) is performed at each respective device within networked system (800).
[0074] B. Examples of Base Calling Schemes
[0075] FIG. 5 illustrates a system (900) that employs two or more base callers for base calling operations on the raw images (j.e., sensor data) output as image, color, or intensity signal data by image sensors in a sequencing machine (910). Sequencing machine (910) of this example includes a flow cell (912), which includes a plurality of tiles (914). Each tile (914) includes a plurality of clusters (916). Sequencing machine (910) may be understood to represent a version of systems (100, 500) or sequencing device (810) described above; while flow cell (912) may be understood to represent a version of flow cells (128, 400, 510) described above. Sequencing machine (910) thus outputs sensor data (920) comprising raw images (e.g., image, RGB, and / or intensity signals) of the tiles (914) of flow cell (912).
[0076] In the present example, system (900) comprises a first base caller (922) and a second base caller (926) which operate on the signals generated by the image sensors, though some variations may include more than two base callers (922, 926). Each base caller (922, 926) of this example outputs signal data corresponding to base call classification information. For example, first base caller (922) outputs first base call classification information (924); and second base caller (926) outputs second base call classification information (928). A base calling combining module (930) generates final base calls (932), based on the signals corresponding to one or both first base call classification information (924) and / or second base call classification information (928). In some versions, first base caller (922) is a neural -network based base-caller; while second base caller (926) is a non-neural network-based base-caller. For example, first base caller (922) may include a non-linear system employing one or more neural network models for base calling. The first base caller (922) may also be referred to as a DeepRTA (Deep Real Time Analysis) base caller or Deep Neural Network base caller.
[0077] By way of further example only, second base caller (926) may include, at least in part, a linear system used for base calling. For example, some versions of second base caller (926) donot employ a neural network for base calling (or use a smaller neural network model for base calling, compared to a larger neural network model used by first base caller (922)). Second base caller (926) may also be referred to as an RTA (Real Time Analysis) base caller. An RTA base caller may use linear intensity extractors to extract features from sequencing images for base calling. In some such versions, RTA performs a template generation step to produce a template image that identifies locations of clusters (916) on a tile (914) using sequencing images from some number of initial sequencing cycles called template cycles. The template image is used as a reference for subsequent registration and intensity extraction steps. The template image is generated by detecting and merging bright spots in each sequencing image of the template cycles, which in turn involves sharpening a sequencing image (e.g., using the Laplacian convolution), determining an “on” threshold by a spatially segregated Otsu approach, and subsequent five-pixel local maximum detection with subpixel location interpolation.
[0078] In another example, locations of clusters (916) on a tile (914) are identified using fiducial markers. A solid support upon which a biological specimen is imaged may include such fiducial markers, to facilitate determination of the orientation of the specimen or the image thereof in relation to probes that are attached to the solid support. Examples of fiducials include, but are not limited to, beads (with or without fluorescent moieties or moieties such as nucleic acids to which labeled probes can be bound), fluorescent molecules attached at known or determinable features, or structures that combine morphological shapes with fluorescent moieties.
[0079] RTA then registers a current sequencing image against the template image. This is achieved by using image correlation to align the current sequencing image to the template image on a sub-region, or by using non-linear transformations (e.g., a full six-parameter linear affine transformation). RTA generates a color matrix to correct cross-talk between color channels of the sequencing images. RTA implements empirical phasing correction to compensate for noise in the sequencing images caused by phase errors. After different corrections are applied to the sequencing images, RTA extracts signal intensities for each spot location in the sequencing images. For example, for a given spot location, signal intensity may be extracted by determining a weighted average of the intensity of the pixels in a spot location. For example, a weighted average of the center pixel and neighboring pixels may be performed using bilinear or bicubic interpolation. In some implementations, each spot location in the image may comprise a few pixels (e.g., 1-5pixels). RTA then spatially normalizes the extracted signal intensities to account for variation in illumination across the sampled imaged. For example, intensity values may be normalized such that a 5th and 95th percentiles have values of 0 and 1, respectively. The normalized signal intensities for the image (e.g., normalized intensities for each channel) may be used to calculate mean chastity for the plurality of spots in the image.
[0080] In some implementations, RTA uses an equalizer to maximize the signal-to-noise ratio of the extracted signal intensities. The equalizer may be trained (e.g., using least square estimation, adaptive equalization algorithm) to maximize the signal-to-noise ratio of cluster intensity data in sequencing images. In some implementations, the equalizer is a lookup table (LUT) bank with a plurality of LUTs with subpixel resolution, also referred to as “equalizer filters” or “convolution kernels.” By way of example only, the number of LUTs in the equalizer may depend on the number of subpixels into which pixels of the sequencing images can be divided. For example, if the pixels are divisible into / -by-n subpixels (e.g., 5 x 5 subpixels), then the equalizer generates / / 2LUTs (e.g., 25 LUTs).
[0081] In some implementations of training the equalizer, data from the sequencing images is binned by well subpixel location. It should be understood that a “well” may include depressions (404) of a flow cell (400) or any other kind of reaction site (e.g., in a flow cell or otherwise). In an example of sequencing images being binned by well subpixel location, for a 5 x 5 LUT, l / 25th of the wells have a center that is in bin (1,1) (e.g., the upper left comer of a sensor pixel), l / 25th of the wells are in bin (1,2), and so on. The equalizer coefficients for each bin may be determined using least squares estimation on the subset of data from the wells corresponding to the respective bins. This way, the resulting estimated equalizer coefficients are different for each bin. Each LUT / equalizer filter / convolution kernel has a plurality of coefficients that are learned from the training. The number of coefficients in a LUT may correspond to the number of pixels that are used for base calling a cluster. For example, if a local grid of pixels (image or pixel patch) that is used to base call a cluster is of size p xp (e.g., 9x9 pixel patch), then each LUT has p2coefficients (e.g., 81 coefficients). The training may produce equalizer coefficients that are configured to mix / combine intensity values of pixels that depict intensity emissions from a target cluster being base called and intensity emissions from one or more adjacent clusters in a manner that maximizes the signal-to-noise ratio. The signal maximized in the signal-to-noise ratio is the intensityemissions from the target cluster, and the noise minimized in the signal-to-noise ratio is the intensity emissions from the adjacent clusters, i.e., spatial crosstalk, plus some random noise (e.g., to account for background intensity emissions). The equalizer coefficients are used as weights and the mixing / combining includes executing element-wise multiplication between the equalizer coefficients and the intensity values of the pixels to calculate a weighted sum of the intensity values of the pixels, i.e., a convolution operation.
[0082] RTA then performs base calling by fitting a mathematical model to the optimized intensity signals. Suitable mathematical models that can be used include, for example, a k-means clustering algorithm, a k-means-like clustering algorithm, expectation maximization clustering algorithm, a histogram -based method, and the like. Four Gaussian distributions may be fit to the set of two-channel intensity signal data such that one distribution is applied for each of the four nucleotides represented in the data set. In some implementations, an expectation maximization (EM) algorithm may be applied. As a result of the EM algorithm, for each X, Y value (referring to each of the two channel intensities respectively) a value may be generated which represents the likelihood that a certain X, Y intensity value belongs to one of four Gaussian distributions to which the data is fitted. Where four bases give four separate distributions, each X, Y intensity value will also have four associated likelihood values, one for each of the four bases. The maximum of the four likelihood values indicates the base call. For example, if a cluster is “off’ in both channels, the base call is G. If the cluster is “off’ in one channel and “on” in another channel the base call is either C or T (depending on which channel is on), and if the cluster is “on” in both channels the base call is A.
[0083] In some implementations of RTA, the base calling errors get averaged out across many training examples. In some other implementations, the ground truth may be sourced using aligned genomic data, which may provide better quality because aligned genomic data may use reference genome and truth information that incorporate the knowledge gained from multiple sequencing platforms and sequencing runs to average out the noise. The ground truth may include basespecific intensity values (or feature values) that reliably represent intensity profiles of bases A, C, G, and T, respectively. A base caller like RTA base caller (926) base calls clusters by processing the sequencing images and producing, for each base call, color-wise intensity values / outputs. The color-wise intensity values may be considered base-wise intensity values because, depending onthe type of chemistry (e.g., 2-color chemistry or 4-color chemistry), the colors map to each of the bases A, C, G, and T. The base with the closest matching intensity profile is called.
[0084] A trainer may train base caller (926) and generate the trained coefficients of the sharpening masks using various training techniques. FIG. 6 shows one implementation of an adaptive technique that may be used to train base caller (926), e.g., using an offline or online mode. Here, the logic is y = x.h + d, where x is the input pixel intensities, h is the sharpening mask coefficients, d is the DC offset. In some implementations, x and h are row and column vectors respectively, with length 81. This vector model is equivalent to a dot product of 9 x 9 matrices representing input pixels and coefficients. The cost is the expected value of error squared. The gradient update moves each coefficient in a direction that reduces the expected value of error squared. Applying this update generates a new estimate of the coefficients that moves them in a direction that (on average) reduces the mean squared error (MSE). In some implementations, Mu is a small constant used to change the adaptation rate / convergence speed. A DC term update can be calculated in a similar way. A gain term update also can be calculated in a similar way.
[0085] In some implementations, since linear interpolation is applied on the coefficient sets, the updates are applied slightly differently in the following manner:
[0086] h(q, n+1) = h(q, n) + lambda q. mu.x(n). e(n)
[0087] In the equation above, h(q, n) is weight q at cycle n, lambda q is the linear interpolation weight for a particular set of coefficients and can include four updates per output due to linear interpolation in two dimensions. The recursive least-squares technique extends the least squares technique to a recursive algorithm.
[0088] C. Examples of Structural Variation Graph Genome Generation
[0089] In some scenarios, a secondary analysis may be performed iteratively while sequence reads (e.g., sequence read signals) are generated by a sequencing system such as systems (100, 500, 814) described herein. Secondary analyses may encompass both alignment of sequence reads to a reference sequence (e.g., the human reference genome sequence) and utilization of this alignment to detect differences between a sample and the reference. Secondary analyses may enable detection of genetic differences, variant detection and genotyping, identification of singlenucleotide polymorphisms (SNPs), small insertions and deletion (indels) and structural changes in the DNA, such as copy number variants (CNVs) and chromosomal rearrangements.
[0090] By performing secondary analyses while sequence reads are generated, system (100, 500, 814) may determine preliminary variant calls iteratively in real-time (or with zero or low latency). Final results of variant determinations may be available soon after (or immediately after) the end of a sequencing run. Alternatively, a sequencing run may be terminated early if variant calls are available with sufficient confidence during the run. In some scenarios, only information related to variant determinations (e.g., variant calls) is transferred off the sequencing system (100, 500, 814). This may decrease, or minimize, the data bandwidth required in comparison to performing the variant determinations in a system that is external. In addition, only variant information may be sent to a computing system (e.g., a cloud computing system) for further processing. In this example, sequencing runs may be terminated prior to completion of an entire sequencing process. For example, if the identity of a pathogen of interest is determined after a number of sequencing cycles of a sequencing run, the sequencing run may be terminated. Thus, the time to a particular answer (e.g., pathogen identification) may be decreased. In some implementations, outputs and intermediate results of system (100, 500, 814) may include histograms of duplicates, exact matches, single and double SNPs, and single and double indels.
[0091] VI. Examples of Precision, Recall, and F-score Determination for Continuous Distributions
[0092] The described sequencing systems and methodologies described herein may be utilized in various contexts for the generation of sequence data pertaining to genetic or epigenetic characteristics of a sample, such as may be obtained for a patient or subject. As may be appreciated, analysis of signal data, at one or more of the primary analysis (i.e., the production of nucleic acid sequence data via conversion of raw sequencing instrument signals to nucleotide and sequence reads), secondary analysis (i.e., the identification of DNA variants via read alignment to a reference genome and variant calling), or tertiary analysis (i.e., the interpretation and annotation of identified variants including filtering, prioritization, and / or the drawing of biological or clinical inferences or conclusions) stage is typically performed using software and / or other automated routines or techniques as the nature of the signal data to be processed, the extent of the data in question, and / or the time limits imposed in which the data is meaningfully useful do not lendthemselves to non-automated or manual computation.
[0093] In practice assessing the performance of such software or executable routines may be based on one or more metrics by which the comparative performance of a technique may be evaluated. By way of example, three such metrics that may be employed in the present context are precision, recall, and F-score. As used herein precision, recall, and F-score may be understood to be two evaluation metrics used to measure the performance of a classifier. In prior practice, precision, recall, and F-score have typically been used in a binary classification context, with no mechanism known to extend their use to continuous data distributions. While the presently described examples related to techniques for calculating precision, recall, and F-score may be provided in the context of bioinformatics, it should be appreciated that the present discussion may be equally applicable to other fields of computational science, such as machine learning, computer vision, and so forth.
[0094] By way of explanation, as used herein “precision” (also known as “positive predictive value”) may be understood to correspond to the accuracy or correctness of “positive” predictions (i.e., out of all instances identified or characterized as true, how many were actually true). With this in mind, “precision” may be understood to be the ratio between the true positives and all identified positives (i.e., true positives and false positives combined). Mathematically, this may be expressed as:„.. True Positives (TP)(1) Precision = - - — - - — - ' ’ True Positives (TP)+False Positives (FP)
[0095] “Recall” (also known as “sensitivity”), as used herein, may be understood to correspond to the completeness of positive predictions or of capturing all relevant instances (i.e., out of all instances that were actually true, what proportion did the system correctly identify as being true). With this in mind, “recall” is a measure of a model or routine correctly identifying true positives out of the population of all actual positives. Mathematically, this may be expressed as:„,, True Positives (TP)(2) Recall = - '7True Positives (TP)+False Negatives (FN)
[0096] Additionally, a further metric to which the present techniques may be extended is theF-score (or Fi- score), which is also a measure of predictive performance and is calculated from precision and recall. The F-score is typically understood to be the harmonic mean of precision and recall. Mathematically this may be expressed as:Use of the Fi-score allows a single value to be looked at in balancing the trade-offs between precision and recall, and thus may help balance the process by simply aiming to obtain a good Fi-score. Similarly, other metrics based on precision and / or recall may benefit from the techniques described herein. By way of example, a Precision-Recall curve (PRC), which typically represents precision along the y-axis and recall along the x-axis, may be plotted using precision and recall values calculated using the techniques described herein.
[0097] With the preceding in mind, one issue arising from current techniques for calculating precision and recall is that current techniques are only adapted for use with binary data distributions. Certain data of interest, however, may be best represented using continuous distributions. By way of example, in the context of genetics and epigenetics, differentially methylated signals (DMSs) represent an example of a continuously distributed variable that may be assessed using software or other executable routines for which the precision, recall, and F-score would be valuable in evaluating performance. In particular, in the differentially methylated region (DMR) context each DMR is not characterized by a binary value (e.g., 0 or 1 or True or False) but by a continuous variable, differentially methylated signal (DMS), which can be represented by values between -1 to 1. In practice such DMRs may range in size from a single base (e.g., a single cytosine also known as differentially methylated position (DMP)) up to thousands of bases (though it should be understood that in a generalized scenario the presently disclosed approaches do not have an upper or lower size limitation).
[0098] With respect to methylation sequencing data giving rise to the DMSs, and by way of background, sequencing-based methylation data may be generated using various techniques, one of which is bisulfite sequencing (BiS). In bisulfite sequencing approaches methylated cytosines may be detected in DNA (e.g., genomic DNA) by treating a DNA sample with sodium bisulfite prior to sequencing. In response to the bisulfite treatment, unmethylated cytosines are deaminated to uracils, which upon sequencing convert to thymidines. Conversely, methylated cytosines arenot deaminated and are read as cytosines. The location of the methylated cytosines is then determined by comparison of bisulfite treated sequences to a reference genome or to untreated sequences. In this manner, methylation data for a sample may be obtained using sequencing techniques.
[0099] By way of a further example, enzymatic methyl sequencing (EMS or EM-Seq) is another approach by which methylation data may be obtained using sequencing techniques. In EMS approaches, non-destructive enzymatic reactions are employed which address certain technical biases that may be introduced by BiS techniques. In EMS approaches, a two-step conversion process is employed. In the first reaction TET2 and T4-BGT convert methylated cytosine (and hydroxymethyl cytosine) into products that are not deaminated by APOBEC3A (i.e., methylated cytosines are modified to be protected from deamination). In the second reaction, APOBEC3A is used to deaminate unmodified cytosines to uracils. Methylations may then be determined using sequencing techniques and systems as described herein.
[0100] With the preceding in mind, sequencing technology and methods as described above may be utilized as part of assessing methylation levels of a subject. Such methylation levels may in turn be used for generating a diagnosis and, correspondingly, recommending, selecting, or implementing a treatment plan (e.g., a personalized treatment plan) or protocol, such as a pharmacological, radiological, surgical, treatment plan or a suitable combination of such treatment techniques. Further, methylation assessment as discussed herein may be ongoing over time so as to allow or facilitate monitoring of treatment efficacy of an ongoing or prior treatment, adjustment to an ongoing treatment plan, and / or confirmation of a remission state for a subject. Regardless of the manner in which a methylation signal is obtained, the presently described approaches for generating precision and recall metrics are applicable and, further, may be applied in various contexts, such as whole genome sequencing, gene panel sequencing, whole exome sequencing, and so forth.
[0101] As discussed herein, methylation sequencing data may be characterized or processed as a differential or differentially methylated signal (DMS) at different nucleotide sites (e.g., CpG sites) either alone or within a larger region. In practice, such sites or regions are automatically “called” or otherwise characterized using specialized routines or software. As discussed herein, precision, recall, and F-score metrics may be utilized in assessing the accuracy and performanceof such software. However, precision, recall, and F-score conventionally are calculated in the context of binary data distributions (e.g., only taking values 0 or 1). Conversely, DMS, as discussed herein, is a continuously distributed value, and thus can have negative as well as positive values (e.g., a range of -1 to 1). Correspondingly, precision, recall, and F-score are difficult to meaningfully calculate in the context of software or processor-implemented routines directed to making calls based on DMS or on other continuously distributed values.
[0102] With this in mind, and turning to FIG. 7, a plot is depicted of an expected (i.e., ground truth) and observed or measured DMS. Due to the continuous nature of the DMS values, conventional (i.e., binary-based) methodologies for calculating precision and recall cannot be directly applied to estimate the accuracy of the automated routines employed as such precision and recall calculations conventionally involve counting the number of true positive (TP), false positive (FP), and false negative (FN) CpG sites, selection of which is not possible using DMS without further assumptions. In such circumstances, under conventional approaches, an additional parameter is specified. This additional parameter typically comprises a threshold that is used to classify CpGs based on the strength of the differential methylation, essentially binning the sites into discrete, non-continuous characterizations suitable for precision and recall calculation using conventional techniques directed toward binary data distributions. In particular, the existing techniques for calculating precision and recall in these circumstances transform DMS values to DMRs, as shown in FIG. 7, by applying a threshold criteria and calculating precision, recall, and F-score by counting FP, TP, and FN CpGs. By way of example, in one embodiment employing a threshold, such as a 10% threshold as shown in FIG. 7:TP CpGs - expected and observed differential methylation is of the same sign and above the threshold;FP CpGs - absolute value of observed differential methylation is above the threshold, whereas expected differential methylation is of opposite sign or its absolute value is below the threshold; andFN CpGs - absolute value of expected differential methylation is above the threshold, whereas observed differential methylation is of opposite sign or its absolute value is below the threshold.
[0103] However, this conventional methodology fails in certain circumstances. For example, the conventional approaches based on thresholding to address continuously distributed data can give a 100% assessment for non-identical signals (e.g., precision = recall = F-score = 1). Conversely, the presently disclosed techniques calculate realistic precision, recall, and F-score values in such circumstances. Additionally, precision, recall, and F-score derived using conventional approaches based on thresholding strongly depend on DMR strength. Conversely, the presently disclosed techniques calculate precision, recall, and F-score values that are independent of DMR strength. Further, precision, recall, and F-score derived using conventional approaches linearly depend on DMR density. Conversely, the presently disclosed techniques calculate precision, recall, and F-score values that are independent of DMR density.
[0104] With the preceding in mind, the presently contemplated precision, recall, and F-score calculation techniques are suitable for use with data having a continuous distribution. In contrast, with prior techniques precision, recall, and F-score could only realistically be calculated for data having binary distributions for the reasons set forth above. The newly contemplated approaches, therefore, allow precision, recall, and F-score metrics to be employed in various technological fields and endeavors for which they were previously poorly suited due to the prevalence of continuously distributed data.
[0105] By way of example, and as discussed herein, in the context of methylation data conventional methods directed to binary data distributions transform differentially methylated signal (DMS) acquired as part of a sequencing operation to differentially methylated regions (DMRs) by applying threshold criteria, as discussed above. Based on the count of the different DMRs (e g., true positives (TP) CpGs, false positives (FP) CpGs, and false negatives (FN) CpGs) the precision, recall, and / or F-score are calculated.
[0106] In accordance with the presently disclosed techniques, however, the differentially methylated signal (DMS) at individual CpGs is itself directly operated on (as opposed to thresholded or binned to determine DMRs) to calculate precision, recall, and F-score for automated DMR calling routines. As discussed herein, in certain implementations this calculation is performed without the use of another parameter (e.g., threshold value or parameter). Examples of equations that may be used to calculate precision, recall, and F-score directly from the DMS are shown in equations (4) - (6) below:(4)(5)where x is the expected signal (e.g., expected DMS) and y is the observed or measured signal (e.g., observed DMS). Correspondingly, the Fi-score for a variable having a continuous distribution may be represented as:
[0107] One further aspect of the present approaches to extending precision, recall, and F-score to continuously distributed data is the ability to identify instances of ideal disagreement between precision and recall. As discussed below, identifying instances of ideal disagreement provides the ability to identify and address errors in the routines or software being evaluated. By way of example, and turning to FIG. 8, a plot is depicted of a conventional binary data distribution defining a possible precision and recall space. In this example, expected values are plotted on the x axis and observed values are plotted on the axis. Due to the underlying binary distribution of the data, values can be only 0 or 1. In such a context, precision and recall may respectively be calculated with equations (4) - (5), which are valid for binary distribution as well. In this instance, where all xt= yfand therefore p, r = 1 there is ideal agreement and the model, software or routines being evaluated based on precision and recall are fully accurate. Correspondingly, where all x;- ■ yt = 0 (i.e., eitheror yfis zero for any i) and therefore p, r = 0 there is total disagreement and the model, software or routines being evaluated based on precision and recall are fully inaccurate.
[0108] Turning to FIG. 9, extension of this example to continuous distribution data is shown via a similar plot. In this example, however, the expected values (x) and observed values (y) each extend to real numbers (though it should be understood that the presently disclosed approaches may be used with either unbounded or bounded continuously distributed data, e.g., a range from -1 to 1, or from -100% to 100%), as shown in FIG. 9 where the values -1 are highlighted to simplify further explanation. Due to this expanded range into negative values, continuous distributions introduce the possibility of a false opposite (FO) event when expected ( / ) and observed (y^)signals are of the same amplitude but of opposite sign: xt= — yz. This in turn introduces possibility of ideal disagreement in which all xt= — ytand therefore p, r = -1 (see equations (4) - (5)). In such a situation, ideal disagreement corresponds to the routines or software being fully accurate, but with an underlying positive or negative sign being reversed, such as due to a programming error. Correction of the error, however, results in ideal agreement and is indicative of the routine or software being fully accurate.
[0109] With FIG. 9 in mind and with reference to Equations (1) and (2), the following observations and definitions may be noted. Precision and recall are mean values of a function fxt>yi) with measures (yt) and / r(x(), respectively:The function (%, y) is symmetric(9) f x,y) = f(y,x)does not exceed one in absolute value(10) l / (,y)| < 1does not depend on the length of the vector r = (x,y)(11) f ax, ay) = f(x,y) for anyand takes the following values(12) FN: / (l,0) = 0(13) TP: / (1,1) = 1(14) FP: (0,l) = 0(15) FO: / (-1,1) = -1and so forth.
[0110] With the preceding in mind and turning to FIG. 10, it may be observed that:Linear measure ^(x) = x can be chosen for non-negative continuous distributions giving the following equations for precision and recall:y2xiytvX?+y?y‘(17) p = —t‘ytDue to the non-negativity constraint of the measure, linear measure cannot be used for continuous distributions of a general type (i.e., taking positive and negative values). Therefore, quadratic measure ju(x) = x2can be chosen instead which gives Equations (4)-(5) for precision and recall.With reference to the Equation (3), the following observations may be noted. F-score is mean value of the function f(xi>yi) with measure ^(xz) + - y )'.For continuous distributions, the equation (6) can be used to calculate F-score. For non-negative continuous distributions, the following formula can be also used:
[0111] It should be understood that the presently disclosed approaches are not limited to discrete sets (Xj,y;) and may be used with continuous variables as well ( ( ),y( )), where is a single or multiple continuous variables. In such case, precision, recall and F-score for continuous distribution can be calculated with equationsThe equations (7)-(8), (17)-(20) are transferred from discrete sets to continuous variables in the same way.
[0112] With the preceding in mind, the techniques for deriving precision, recall, and / or F-score may have numerous applications, including in biometrics and machine learning. By way of example, a real-world genomic application where precision, recall and F-score for non-negative continuous distributions can be used may be the detection of somatic mutations. Since somatic mutation rate can be used to detect cancer, these metrics may be of significance. In particular, somatic mutations are characterised with the variant type, variant location, and variant allele fraction (VAF), which takes non-negative values. Zero-value VAFs can be interpreted as a missing mutation. This in turn can be interpreted as a false positive (FP) event if expected VAF of a mutation is zero whereas observed VAF is above zero, or a false negative (FN) event if expected VAF of a mutation is above zero whereas observed VAF is zero. In the general case, somatic mutation calling software should not only find all mutations but should also accurately compute their VAFs. With this in mind, the equations (7), (8), and (19) as discussed herein can be used to assess performance of such somatic mutation calling software.VII. Examples of Combinations
[0113] The following examples relate to various non-exhaustive ways in which the teachings herein may be combined or applied. The following examples are not intended to restrict the coverage of any claims that may be presented at any time in this application or in subsequent filings of this application. No disclaimer is intended. The following examples are being provided for nothing more than merely illustrative purposes. It is contemplated that the various teachings herein may be arranged and applied in numerous other ways. It is also contemplated that some variations may omit certain features referred to in the below examples. Therefore, none of the aspects or features referred to below should be deemed critical unless otherwise explicitly indicated as such at a later date by the inventors or by a successor in interest to the inventors. If any claims arepresented in this application or in subsequent filings related to this application that include additional features beyond those referred to below, those additional features shall not be presumed to have been added for any reason relating to patentability.
[0114] Example 1
[0115] A method for deriving one or more of a precision, a recall, or an F-score metric, comprising: acquiring or accessing a set of data, wherein the data is continuously or non-negative continuously distributed; and based on an expected signal and an observed signal, determining one or more of a precision, a recall, or an F-score for one or more automated routines configured to operate on the set of data. One or more of the precision, the recall, or the F-score provide a comparative measure of the accuracy of the one or more automated routines.
[0116] Example 2
[0117] The method of Example 1, wherein one or more of the precision, the recall, or the F-score are determined without use of a threshold to categorize the set of data into a binary data set.
[0118] Example s
[0119] The method of Example 1, wherein the set of data comprises a differentially methylated signal.
[0120] Example 4
[0121] The method of Example 1, wherein the precision is determined in accordance with:or its mathematical equivalent, wherein x is the expected signal and is the observed signal.
[0122] Example 5
[0123] The method of Example 1, wherein the recall is determined in accordance with:or its mathematical equivalent, wherein x is the expected signal and y is the observed signal.
[0124] Example 6
[0125] The method of Example 1, wherein the F-score is determined in accordance with:or its mathematical equivalent, wherein x is the expected signal and y is the observed signal.
[0126] Example 7
[0127] The method of Example 1, wherein an ideal disagreement is identified for the precision and the recall.
[0128] Example 8
[0129] The method of Example 7, further comprising updating the one or more automated routines to address an incorrect positive or negative sign, wherein the ideal disagreement then becomes an ideal agreement between the precision and the recall.
[0130] Example 9
[0131] A method for deriving one or more of a precision, a recall, or an F-score metric based on a differentially methylated signal (DMS), comprising: acquiring or accessing a DMS for a set of nucleic acid samples, wherein the DMS is continuously or non-negative continuously distributed; and based on an expected DMS and an observed DMS, determining one or more of a precision, a recall, or an F-score for one or more automated routines configured to operate on the DMS to determine one or more differentially methylated regions (DMRs). One or more of the precision, the recall, or the F-score provide a comparative measure of the accuracy of the one or more automated routines at determining DMRs.
[0132] Example 10
[0133] The method of Example 9, wherein one or more of the precision, the recall, or the F-score are determined without use of a threshold to categorize DMS into DMRs.
[0134] Example 11
[0135] The method of Example 9, wherein the precision is determined in accordance with:or its mathematical equivalent, wherein x is the expected DMS andy is the observed DMS.[00136J Example 12
[0137] The method of Example 9, wherein the recall is determined in accordance with:or its mathematical equivalent, wherein x is the expected DMS andy is the observed DMS.
[0138] Example 13
[0139] The method of Example 9, wherein the F-score is determined in accordance with:or its mathematical equivalent, wherein x is the expected signal andy is the observed signal.
[0140] Example 14
[0141] The method of Example 9, wherein an ideal disagreement is identified for the precision and the recall.
[0142] Example 15
[0143] The method of Example 14, further comprising updating the one or more automated routines to address an incorrect positive or negative sign, wherein the ideal disagreement thenbecomes an ideal agreement between the precision and the recall.
[0144] Example 16
[0145] A processor-based system, comprising: at least one processor configured to execute stored routines; and one or more tangible storage media storing processor-executable routines. The processor-executable routines, when executed by the at least one processor, cause acts to be performed comprising: acquiring or accessing a set of data, wherein the data is continuously or non-negative continuously distributed; and based on an expected signal and an observed signal, determining one or more of a precision, a recall, or an F-score for one or more automated routines configured to operate on the set of data. One or more of the precision, the recall, or the F-score provide a comparative measure of the accuracy of the one or more automated routines.
[0146] Example 17
[0147] The processor-based system of Example 16, wherein one or more of the precision, the recall, or the F-score are determined without use of a threshold to categorize the set of data into a binary data set.
[0148] Example 18
[0149] The processor-based system of Example 16, wherein the set of data comprises a differentially methylated signal.
[0150] Example 19
[0151] The processor-based system of Example 16, wherein the precision is determined in accordance with:or its mathematical equivalent, wherein x is the expected signal and y is the observed signal.
[0152] Example 20
[0153] The processor-based system of Example 16, wherein the recall is determined inaccordance with:or its mathematical equivalent, wherein x is the expected signal and y is the observed signal.
[0154] Example 21
[0155] The processor-based system of Example 16, wherein the F-score is determined in accordance with:or its mathematical equivalent, wherein x is the expected signal and y is the observed signal.
[0156] Example 22
[0157] The processor-based system of Example 16, wherein an ideal disagreement is identified for the precision and the recall.
[0158] Example 23
[0159] The processor-based system of Example 22, further comprising updating the one or more automated routines to address an incorrect positive or negative sign, wherein the ideal disagreement then becomes an ideal agreement between the precision and the recall.
[0160] VIII. Miscellaneous
[0161] While the foregoing examples are provided in the context of a system (100) that may be used in processing data having a continuous distribution, such as differentially methylated signal (DMS) data, and the derivation of precision and recall metrics to evaluate routines and / or software used to process calls made using such data, the teachings herein may also be readily applied in other contexts, including in systems that perform other processes (i.e., other than nucleotide sequencing procedures). The teachings herein are thus not necessarily limited to systems that are used to perform nucleotide sequencing processes or methylation signal analysis.
[0162] It is to be understood that the subject matter described herein is not limited in itsapplication to the details of construction and the arrangement of components set forth in the description herein or illustrated in the drawings hereof. The subject matter described herein is capable of other implementations and of being practiced or of being carried out in various ways. Also, it is to be understood that the phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. As used herein, an element or step recited in the singular and proceeded with the word “a” or “an” should be understood as not excluding plural of said elements or steps, unless such exclusion is explicitly stated. Furthermore, references to “one example” are not intended to be interpreted as excluding the existence of additional examples that also incorporate the recited features. The use of “including,” “comprising,” or “having” and variations thereof herein is meant to encompass the items listed thereafter and equivalents thereof as well as additional items.
[0163] When used in the claims, the term “set” should be understood as one or more things which are grouped together. Similarly, when used in the claims “based on” should be understood as indicating that one thing is determined at least in part by what it is specified as being “based on.” Where one thing is required to be exclusively determined by another thing, then that thing will be referred to as being “exclusively based on” that which it is determined by.
[0164] Unless specified or limited otherwise, the terms “mounted,” “connected,” “supported,” and “coupled” and variations thereof are used broadly and encompass both direct and indirect mountings, connections, supports, and couplings. Further, “connected” and “coupled” are not restricted to physical or mechanical connections or couplings. Also, it is to be understood that phraseology and terminology used herein with reference to device or element orientation (such as, for example, terms like “above,” “below,” “front,” “rear,” “distal,” “proximal,” and the like) are only used to simplify description of one or more examples described herein, and do not alone indicate or imply that the device or element referred to must have a particular orientation. In addition, terms such as “outer” and “inner” are used herein for purposes of description and are not intended to indicate or imply relative importance or significance.
[0165] It is to be understood that the above description is intended to be illustrative, and not restrictive. For example, the above-described examples (and / or aspects thereof) may be used in combination with each other. In addition, many modifications may be made to adapt a particular situation or material to the teachings of the presently described subject matter without departingfrom its scope. While the dimensions, types of materials and coatings described herein are intended to define the parameters of the disclosed subject matter, they are by no means limiting and instead illustrations. Many further examples will be apparent to those of skill in the art upon reviewing the above description. The scope of the disclosed subject matter should, therefore, be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. In the appended claims, the terms “including” and “in which” are used as the plain-English equivalents of the respective terms “comprising” and “wherein.” Moreover, in the following claims, the terms “first,” “second,” and “third,” etc. are used merely as labels, and are not intended to impose numerical requirements on their objects. Further, the limitations of the following claims are not written in means — plus-function format and are not intended to be interpreted based on 35 U. S. C. §112(f) paragraph, unless and until such claim limitations expressly use the phrase “means for” followed by a statement of function void of further structure.
[0166] The following claims recite aspects of certain examples of the disclosed subject matter and are considered to be part of the above disclosure. These aspects may be combined with one another.
Claims
1. CLAIMS2.What is claimed is:
1. A method for deriving one or more of a precision, a recall, or an F-score metric, comprising:4.acquiring or accessing a set of data, wherein the data is continuously or non-negative continuously distributed; and5.based on an expected signal and an observed signal, determining one or more of a precision, a recall, or an F-score for one or more automated routines configured to operate on the set of data;6.wherein one or more of the precision, the recall, or the F-score provide a comparative measure of the accuracy of the one or more automated routines.
2. The method of claim 1, wherein one or more of the precision, the recall, or the F-score are determined without use of a threshold to categorize the set of data into a binary data set.
3. The method of claim 1, wherein the set of data comprises a differentially methylated signal.
4. The method of claim 1, wherein the precision is determined in accordance with:
13. 15.or its mathematical equivalent, wherein x is the expected signal and ' is the observed signal.
5. The method of claim 1, wherein the recall is determined in accordance with:
20.
21. or its mathematical equivalent, wherein x is the expected signal andy is the observed signal.
6. The method of claim 1, wherein the F-score is determined in accordance with:
24. 26.or its mathematical equivalent, wherein x is the expected signal andy is the observed signal.
7. The method of claim 1, wherein an ideal disagreement is identified for the precision and the recall.
8. The method of claim 7, further comprising updating the one or more automated routines to address an incorrect positive or negative sign, wherein the ideal disagreement then becomes an ideal agreement between the precision and the recall.
9. A method for deriving one or more of a precision, a recall, or an F-score metric based on a differentially methylated signal (DMS), comprising:30.acquiring or accessing a DMS for a set of nucleic acid samples, wherein the DMS is continuously or non-negative continuously distributed; and31.based on an expected DMS and an observed DMS, determining one or more of a precision, a recall, or an F-score for one or more automated routines configured to operate on the DMS to determine one or more differentially methylated regions (DMRs);32.wherein one or more of the precision, the recall, or the F-score provide a comparative measure of the accuracy of the one or more automated routines at determining DMRs.
10. The method of claim 9, wherein one or more of the precision, the recall, or the F-score are determined without use of a threshold to categorize DMS into DMRs.
11. The method of claim 9, wherein the precision is determined in accordance with:
36. 38.or its mathematical equivalent, wherein x is the expected DMS and is the observed DMS.
12. The method of claim 9, wherein the recall is determined in accordance with:40.[Mathematical formula rendered as image - OCR garbled representation of recall formula]41.r42.or its mathematical equivalent, wherein x is the expected DMS and is the observed DMS.
13. The method of claim 9, wherein the F-score is determined in accordance with:
45. 47.or its mathematical equivalent, wherein x is the expected signal and y is the observed signal.
14. The method of claim 9, wherein an ideal disagreement is identified for the precision and the recall.
15. The method of claim 14, further comprising updating the one or more automated routines to address an incorrect positive or negative sign, wherein the ideal disagreement then becomes an ideal agreement between the precision and the recall.
16. A processor-based system, comprising:51.at least one processor configured to execute stored routines;52.one or more tangible storage media storing processor-executable routines, wherein the processor-executable routines, when executed by the at least one processor, cause acts to be performed comprising: acquiring or accessing a set of data, wherein the data is continuously or nonnegative continuously distributed; and53.based on an expected signal and an observed signal, determining one or more of a precision, a recall, or an F-score for one or more automated routines configured to operate on the set of data;54.wherein one or more of the precision, the recall, or the F-score provide a comparative measure of the accuracy of the one or more automated routines.
17. The processor-based system of claim 16, wherein one or more of the precision, the recall, or the F-score are determined without use of a threshold to categorize the set of data into a binary data set.
18. The processor-based system of claim 16, wherein the set of data comprises a differentially methylated signal.
19. The processor-based system of claim 16, wherein the precision is determined in accordance with:
61. 63.or its mathematical equivalent, wherein x is the expected signal and y is the observed signal.
20. The processor-based system of claim 16, wherein the recall is determined in accordance with:
68. 70.or its mathematical equivalent, wherein x is the expected signal and is the observed signal.
21. The processor-based system of claim 16, wherein the F-score is determined in accordance with:
72. 74.or its mathematical equivalent, wherein x is the expected signal and y is the observed signal.
22. The processor-based system of claim 16, wherein an ideal disagreement is identified for the precision and the recall.
23. The processor-based system of claim 22, further comprising updating the one or more automated routines to address an incorrect positive or negative sign, wherein the ideal disagreement then becomes an ideal agreement between the precision and the recall.
Citation Information
Patent Citations
Flow cells with hydrogel coating
US10919033B2
Recombinase polymerase amplification
US7270981B2
Method of preparing libraries of template polynucleotides
US7741463B2