Systems and methods for minimum residual disease (MRD) detection

A genome-wide mutational integration workflow with machine learning improves MRD detection sensitivity by identifying somatic variants through WGS, addressing the limitations of targeted sequencing panels in detecting low levels of ctDNA.

WO2025181217A1PCT designated stage Publication Date: 2025-09-04F HOFFMANN LA ROCHE & CO AG +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/055311
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-10
Filing Date
2025-02-27
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Current MRD detection methods, particularly those relying on targeted sequencing panels, struggle to detect low levels of ctDNA due to variability in somatic mutations among patients, leading to false-negatives and inadequate sensitivity.

Method used

A genome-wide mutational integration workflow using low-coverage whole genome sequencing (WGS) combined with a machine learning algorithm, specifically a random forest classifier, to identify and filter germline mutations and sequencing artifacts, enabling sensitive detection of somatic variants for MRD assessment across different cancer types.

Benefits of technology

The approach enhances sensitivity in detecting low-burden cancer by capturing a broader spectrum of mutations, improving the accuracy of predicting tumor recurrence and guiding personalized treatment decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000032_0000
    Figure 00000032_0000
  • Figure 00000033_0000
    Figure 00000033_0000
  • Figure 00000034_0000
    Figure 00000034_0000
Patent Text Reader

Abstract

The disclosure is related to using machine learning algorithms to detect minimum residual disease (MRD) in cancer patients. For example, a method includes detecting somatic variants from pre-treatment samples from a patient using a variant caller. The method also includes filtering any germline mutations and sequencing artifacts from the detected somatic mutations. The method also includes classifying sequencing reads covering remaining somatic variants from the filtering using a machine learning model to identify somatic variants used for minimum residual disease (MRD) detection. The machine learning model may be trained using a random forest classifier based on a plurality of features. The method also includes analyzing the identified somatic variants to determine whether the patient is positive or negative for MRD.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR MINIMUM RESIDUAL DISEASE (MRD) DETECTION

[0001] Embodiments of the disclosure related generally to disease detection, and more specifically to using machine learning algorithms to detect minimum residual disease (MRD) in cancer patients.BACKGROUND

[0002] Minimal residual disease (MRD) is a small number of cells, e.g., cancer cells, that are left in the human body after a patient receives treatment. Unfortunately, these cells may come back and cause relapse in the patient. MRD detection may be used as an approach for predicting tumor recurrence and identifying patients who may benefit from adjuvant therapy following surgery or radiotherapy. MRD assessment using circulating tumor DNA (ctDNA) is a minimally invasive approach for MRD detection. Currently, many methods for MRD detection rely on targeted sequencing panels, but these methods have limitations in detecting low levels of ctDNA. One reason why targeted sequencing panels may not detect ctDNA in some patients is that not all patients share somatic mutations that are easily targetable, which may lead to false-negatives.SUMMARY OF THE DISCLOSURE

[0003] Embodiments of the disclosure related generally to disease detection, and more specifically to using machine learning algorithms to detect minimum residual disease (MRD).

[0004] In an aspect, a method includes detecting somatic variants from pre-treatment samples from a patient using a variant caller. The method also includes filtering any germline mutations and sequencing artifacts from the detected somatic variants. The method also includes classifyingsequencing reads covering remaining somatic variants from the filtering using a machine learning model to identify somatic variants used for minimum residual disease (MRD) detection. The machine learning model is trained using a random forest classifier based on a plurality of features. The method also includes analyzing the identified somatic variants to determine whether the patient is positive or negative for MRD.

[0005] In another aspect, the present disclosure includes a system having a memory and a processor coupled to the memory. The processor is configured to detect somatic variants from pretreatment samples from a patient using a variant caller. The processor is configured to filter any germline mutations and sequencing artifacts from the detected somatic variants. The processor is configured to classify sequencing reads covering remaining somatic variants from the filtering using a machine learning model to identify somatic variants used for minimum residual disease (MRD) detection. The machine learning model is trained using a random forest classifier based on a plurality of features. The processor is configured to analyze the identified somatic variants to determine whether the patient is positive or negative for MRD.

[0006] In another aspect, the detecting the somatic variants is based on a whole genome sequencing (WGS).

[0007] In another aspect, the plurality of features used to train the machine learning model is based on a type of sequencing used on the pre-treatment samples.

[0008] In another aspect, when the type of sequencing comprises a sequencing by synthesis process, the machine learning model is trained using a first set of features from among the plurality of features and when the type of sequencing comprises a sequencing by expansion process, the machine learning model is trained using a second set of features from among the plurality offeatures. In some aspects, a combination of features used in the first set of features is different than a combination of features used in the second set of features.

[0009] In another aspect, the plurality of features are assigned different weights.

[0010] In another aspect, the filtering any germline mutations and sequencing artifacts is based on a plurality of databases defining the germline mutations and / or sequencing artifacts.

[0011] In another aspect, the method includes generating a report based on the analyzing the identified somatic variants. In another aspect, the processor is further configured to generate a report based on the analyzing the identified somatic variants.

[0012] In another aspect, the machine learning model is trained using positive and negative training datasets. The positive training dataset includes germline mutations and the negative training dataset includes sequencing artifacts.

[0013] In another aspect, the analyzing the identified somatic variants includes comparing the identified somatic variants in a pre- or post-treatment sample with a set of healthy samples.

[0014] In another aspect, when a number of the identified somatic variants in the pre- or post-treatment sample exceeds a number of the identified somatic variants in the set of healthy samples, the analyzing the identified somatic variants renders a positive determination for MRD, and when the number of the identified somatic variants in the pre- or post-treatment sample is less than the number of identified somatic variants in the set of healthy samples, the analyzing the identified somatic variants renders a negative determination for MRD.

[0015] In another aspect, the analyzing the identified somatic variants includes comparing a normalized number of the identified somatic variants in a pre- or post-treatment sample with a normalized set of somatic variants in a set of healthy samples.

[0016] In another aspect, when the normalized number of the identified somatic in the pre- or post-treatment sample variants exceeds the normalized set of somatic variants in the set of healthy samples, the analyzing the identified somatic variants renders a positive determination for MRD, and when the normalized number of the identified somatic variants in the pre- or posttreatment sample is less than the normalized set of somatic variants in the set of healthy samples, the analyzing the identified somatic variants renders a negative determination for MRD.BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The novel features of the disclosure are set forth with particularity in the claims that follow. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings of which:

[0018] FIG. 1 is a top view of an embodiment of a nanopore sensor chip having an array of nanopore cells.

[0019] FIG. 2 is a block diagram illustrating an embodiment of a method for detecting minimum residual disease (MRD), according to aspects of the present disclosure.

[0020] FIG. 3 is a block diagram illustrating an embodiment of a computer system, according to aspects of the present disclosure.

[0021] FIG. 4 illustrates an embodiment of a nanopore cell in a nanopore sensor chip, according to aspects of the present disclosure.DETAILED DESCRIPTION

[0022] The disclosure described here is a system and method based on a machine learning algorithm for detecting minimum residual disease (MRD).

[0023] To overcome the limitations of current MRD detection techniques, the present disclosure is directed to a genome-wide mutational integration workflow for ctDNA detection based on low-coverage whole genome sequencing (WGS). The techniques described herein allow for tracking a broader spectrum of mutations with lower coverage, enabling highly sensitive detection of low-burden cancer across different cancer types without the need for a cancer-specific target panel. By using WGS to track a larger number of mutations, the techniques of the present disclosure achieve better sensitivity in detecting ctDNA, which is critical for predicting tumor recurrence and guiding personalized treatment decisions. The present disclosure provides for a highly sensitive single nucleotide variant (SNV) MRD workflow that facilitates the development of multimodal and multi-analyte based MRD assays.

[0024] According to aspects of the present disclosure, a tumor-informed SNV MRD workflow uses whole genome sequencing (WGS) data. The WGS data may be based on data from pre-treatment tumor tissue, pre-treatment normal whole blood or peripheral blood mononuclear cells (PBMCs), and pre- and post-treatment plasma samples. In some embodiments, the pretreatment tumor tissue and post-treatment plasma may be used as base requirements, PBMCs may be used to improve sensitivity and specificity of the processes described herein, and pre-treatment plasma may be used as a ctDNA baseline. The workflow includes tracking a large number of somatic mutations across the whole genome, ensuring the capturing of patient-specific mutations that cannot be detected by targeted panels. By capturing thousands to tens of thousands of somaticmutations per patient, the WGS-based SNV MRD workflow of the present disclosure can achieve higher sensitivity even with much lower sequencing depth compared to targeted sequencing.

[0025] FIG. l is a top view of an embodiment of a nanopore sensor chip 100 having an array 140 of nanopore cells 150. Each nanopore cell 150 includes a control circuit integrated on a silicon substrate of nanopore sensor chip 100. In some embodiments, side walls 136 are included in array 140 to separate groups of nanopore cells 150 so that each group can receive a different sample for characterization. Each nanopore cell can be used to sequence a nucleic acid. In some embodiments, nanopore sensor chip 100 includes a cover plate 130. In some embodiments, nanopore sensor chip 100 also includes a plurality of pins 110 for interfacing with other circuits, such as a computer processor.

[0026] In some embodiments, nanopore sensor chip 100 includes multiple chips in a same package, such as, for example, a Multi-Chip Module (MCM) or System-in-Package (SiP). The chips can include, for example, a memory, a processor, a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), data converters, a high-speed I / O interface, etc.

[0027] In some embodiments, nanopore sensor chip 100 is coupled to (e.g., docked to) a nanochip workstation 120, which can include various components for carrying out (e.g., automatically carrying out) various embodiments of the processes disclosed herein. For example, the nanochip workstation 120 may include the computing system shown with respect to FIG. 3. These processes can include, for example, analyte delivery mechanisms, such as pipettes for delivering lipid suspension or other membrane structure suspension, analyte solution, and / or other liquids, suspension or solids. The nanochip workstation components can further include robotic arms, one or more computer processors, and / or memory. A plurality of polynucleotides can bedetected on array 140 of nanopore cells 150. In some embodiments, each nanopore cell 150 is individually addressable.

[0028] Nanopore cells 150 in nanopore sensor chip 100 can be implemented in many different ways. For example, in some embodiments, tags of different sizes and / or chemical structures are attached to different nucleotides in a nucleic acid molecule to be sequenced. In some embodiments, a complementary strand to a template of the nucleic acid molecule to be sequenced may be synthesized by hybridizing differently polymer-tagged nucleotides with the template. In some implementations, the nucleic acid molecule and the attached tags both move through the nanopore, and an ion current passing through the nanopore can indicate the nucleotide that is in the nanopore because of the particular size and / or structure of the tag attached to the nucleotide. In some implementations, only the tags are moved into the nanopore. There can also be many different ways to detect the different tags in the nanopores.

[0029] In some embodiments, the nanochip workstation 120 may execute a machine learning model 155. The machine learning model 155 may analyze sequencing data from the nanopore sensor chip 100 to classify reads as circulating tumor DNA (ctDNA) reads. In some embodiments, the machine learning model 155 may be trained on a set of healthy samples to identify which reads contain germline and artifact variants. The set of healthy samples may have variable provenance and quality control (QC) stats. Additionally, the machine learning model 155 of the present disclosure may be trained to remove reads that contain errors from library preparation and sequencing and, as a result, lower the background noise. In other words, in some embodiments, the machine learning model 155 may determine if reads covering the set of somatic variants are high-quality reads or low-quality reads, remove any low-quality reads, and output the remaining high-quality reads covering the set of somatic variants. If any somatic variants do nothave enough high-quality reads supporting it, the machine learning model 155 may remove such somatic variants from the set of somatic variants to be used in MRD detection, and the machine learning model 155 may then use the smaller set of somatic variants in the MRD detection. Although the machine learning model 155 is described as operating on the nanochip workstation 120, it should be understood by those of ordinary skill that the machine learning model 155 may similarly be executed on an external computing device.

[0030] FIG. 4 illustrates an embodiment of an example nanopore cell 400 in a nanopore sensor chip, such as nanopore cell 150 in nanopore sensor chip 100 of FIG. 1, that can be used to characterize a polynucleotide or a polypeptide. Nanopore cell 400 can include a well 405 formed of dielectric layers 401 and 404; a membrane, such as a lipid bilayer 414 formed over well 405; and a sample chamber 415 on lipid bilayer 414 and separated from well 405 by lipid bilayer 414. Well 405 can contain a volume of electrolyte 406, and sample chamber 415 can hold bulk electrolyte 408 containing a nanopore, e.g., a soluble protein nanopore transmembrane molecular complexes (PNTMC), and the analyte of interest (e.g., a nucleic acid molecule to be sequenced).

[0031] Nanopore cell 400 can include a working electrode 402 at the bottom of well 405 and a counter electrode 410 disposed in sample chamber 415. A signal source 428 can apply a voltage signal between working electrode 402 and counter electrode 410. A single nanopore (e.g., a PNTMC) can be inserted into lipid bilayer 414 by an electroporation process caused by the voltage signal, thereby forming a nanopore 416 in lipid bilayer 414. The individual membranes (e.g., lipid bilayers 414 or other membrane structures) in the array can be neither chemically nor electrically connected to each other. Thus, each nanopore cell in the array can be an independent sequencing machine, producing data unique to the single polymer molecule associated with the nanopore thatoperates on the analyte of interest and modulates the ionic current through the otherwise impermeable lipid bilayer.

[0032] As shown in FIG. 4, nanopore cell 400 can be formed on a substrate 430, such as a silicon substrate. Dielectric layer 401 can be formed on substrate 430. Dielectric material used to form dielectric layer 401 can include, for example, glass, oxides, nitrides, and the like. An electric circuit 422 for controlling electrical stimulation and for processing the signal detected from nanopore cell 400 can be formed on substrate 430 and / or within dielectric layer 401. For example, a plurality of patterned metal layers (e.g., metal 1 to metal 6) can be formed in dielectric layer 401, and a plurality of active devices (e.g., transistors) can be fabricated on substrate 430. In some embodiments, signal source 428 is included as a part of electric circuit 422. Electric circuit 422 can include, for example, amplifiers, integrators, analog-to-digital converters, noise filters, feedback control logic, and / or various other components. Electric circuit 422 can be further coupled to a processor 424 that is coupled to a memory 426, where processor 424 can analyze the sequencing data as described herein.

[0033] Working electrode 402 can be formed on dielectric layer 401 , and can form at least a part of the bottom of well 405. In some embodiments, working electrode 402 is a metal electrode. For non-faradaic conduction, working electrode 402 can be made of metals or other materials that are resistant to corrosion and oxidation, such as, for example, platinum, gold, titanium nitride, and graphite. For example, working electrode 402 can be a platinum electrode with electroplated platinum. In another example, working electrode 402 can be a titanium nitride (TiN) working electrode. Working electrode 402 can be porous, thereby increasing its surface area and a resulting capacitance associated with working electrode 402. Because the working electrode of a nanoporecell can be independent from the working electrode of another nanopore cell, the working electrode can be referred to as cell electrode in this disclosure.

[0034] Dielectric layer 404 can be formed above dielectric layer 401. Dielectric layer 404 forms the walls surrounding well 405. Dielectric material used to form dielectric layer 404 can include, for example, glass, oxide, silicon mononitride (SiN), polyimide, or other suitable hydrophobic insulating material. The top surface of dielectric layer 404 can be silanized. The silanization can form a hydrophobic layer 420 above the top surface of dielectric layer 404. In some embodiments, hydrophobic layer 420 has a thickness of about 1.5 nanometer (nm).

[0035] Well 405 formed by the dielectric layer walls 404 includes volume of electrolyte 406 above working electrode 402. Volume of electrolyte 406 can be buffered and can include one or more of the following: lithium chloride (LiCl), sodium chloride (NaCl), potassium chloride (KC1), lithium glutamate, sodium glutamate, potassium glutamate, lithium acetate, sodium acetate, potassium acetate, calcium chloride (CaCh), strontium chloride (SrCh), manganese chloride (MnCh), and magnesium chloride (MgCh). In some embodiments, volume of electrolyte 406 has a thickness of about three microns (pm).

[0036] As also shown in FIG. 4, a membrane can be formed on top of dielectric layer 404 and spanning across well 405. In some embodiments, the membrane includes a lipid monolayer 418 formed on top of hydrophobic layer 420. As the membrane reaches the opening of well 405, lipid monolayer 408 can transition to lipid bilayer 414 that spans across the opening of well 405. The lipid bilayer can comprise or consist of lipids, such as a phospholipid, for example, selected from diphytanoyl-phosphatidylcholine (DPhPC), l,2-diphytanoyl-sn-glycero-3-phosphocholine, 1,2- di-O-phytanyl-sn-glycero-3 -phosphocholine (DoPhPC), palmitoyl-oleoyl-phosphatidylcholine (POPC), dioleoyl-phosphatidyl-methylester (DOPME), dipalmitoylphosphatidylcholine (DPPC),phosphatidylcholine, phosphatidylethanolamine, phosphatidylserine, phosphatidic acid, phosphatidylinositol, phosphatidylglycerol, sphingomyelin, 1 ,2-di-O-phytanyl-sn-glycerol, 1 ,2- dipalmitoyl-sn-glycero-3-phosphoethanolamine-N-[methoxy(polyethylene glycol)-350], 1 ,2- dipalmitoyl-sn-glycero-3-phosphoethanolamine-N-[methoxy(polyethylene glycol)-550], 1 ,2- dipalmitoyl-sn-glycero-3-phosphoethanolamine-N-[methoxy(polyethylene glycol)-750], 1 ,2- dipalmitoyl-sn-glycero-3-phosphoethanolamine-N-[methoxy(polyethylene glycol)-l 000], 1 ,2- dipalmitoyl-sn-glycero-3-phosphoethanolamine-N-[methoxy(polyethylene glycol)-2000], 1 ,2- dioleoyl-sn-glycero-3-phosphoethanolamine-N-lactosyl, GM1 Ganglioside,Lysophosphatidylcholine (LPC), or any combination thereof. Other phospholipid derivatives may also be used, such as phosphatidic acid derivatives (e.g., DMPA, DDPA, DSPA), phosphatidylchohne derivatives (e.g., DDPC, DLPC, DMPC, DPPC, DSPC, DOPC, POPC, DEPC), phosphatidylglycerol derivatives (e.g., DMPG, DPPG, DSPG, POPG), phosphatidylethanolamine derivatives (e.g., DMPE, DPPE, DSPE DOPE), phosphatidylserine derivatives (e.g., DOPS), PEG phospholipid derivatives (e.g, mPEG-phospholipid, polyglycerinphospholipid, funcitionalized-phospholipid, terminal activated-phospholipid), diphytanoyl phospholipids (e.g., DPhPC, DOPhPC, DPhPE, and DOPhPE), for example. In some embodiments, the bilayer can be formed using non-lipid based materials, such as amphiphilic block copolymers (e.g, poly(butadiene)-block-poly(ethylene oxide), PEG diblock copolymers, PEG triblock copolymers, PPG triblock copolymers, and poloxamers) and other amphiphilic copolymers, which may be nonionic or ionic. In some embodiments, the bilayer can be formed from a combination of lipid based materials and non-lipid based materials. In some embodiments, the bilayer materials can be delivered in a solvent phase including one or more organic solventssuch as alkanes (e.g., decane, tridecane, hexadecane, etc.), and / or one or more silicone oils (e.g.,AR-20).

[0037] As shown, lipid bilayer 414 is embedded with a single nanopore 416, e.g., formed by a single PNTMC. As described above, nanopore 416 can be formed by inserting a single PNTMC into lipid bilayer 414 by electroporation. Nanopore 416 can be large enough for passing at least a portion of the analyte of interest and / or small ions (e.g., Na+, K+, Ca2+, CI") between the two sides of lipid bilayer 414.

[0038] Sample chamber 415 is over lipid bilayer 414, and can hold a solution of the analyte of interest for characterization. The solution can be an aqueous solution containing bulk electrolyte 408 and buffered to an optimum ion concentration and maintained at an optimum pH to keep the nanopore 416 open. Nanopore 416 crosses lipid bilayer 414 and provides the only path for ionic flow from bulk electrolyte 408 to working electrode 402. In addition to nanopores (e.g., PNTMCs) and the analyte of interest, bulk electrolyte 408 can further include one or more of the following: lithium chloride (LiCl), sodium chloride (NaCl), potassium chloride (KC1), lithium glutamate, sodium glutamate, potassium glutamate, lithium acetate, sodium acetate, potassium acetate, calcium chloride (CaCh), strontium chloride (SrCh), manganese chloride (MnCh), and magnesium chloride (MgCh).

[0039] Counter electrode (CE) 410 can be an electrochemical potential sensor. In some embodiments, counter electrode 410 is shared between a plurality of nanopore cells, and can therefore be referred to as a common electrode. In some cases, the common potential and the common electrode can be common to all nanopore cells, or at least all nanopore cells within a particular grouping. The common electrode can be configured to apply a common potential to the bulk electrolyte 408 in contact with the nanopore 416. Counter electrode 410 and workingelectrode 402 can be coupled to signal source 428 for providing electrical stimulus (e.g., voltage bias) across lipid bilayer 414, and can be used for sensing electrical characteristics of lipid bilayer 414 (e.g., resistance, capacitance, and ionic current flow). In some embodiments, nanopore cell 400 can also include a reference electrode 412.

[0040] In some embodiments, various checks are made during creation of the nanopore cell as part of calibration. Once a nanopore cell is created, further calibration steps can be performed, e.g., to identify nanopore cells that are performing as desired (e.g., one nanopore in the cell). Such calibration checks can include physical checks, voltage calibration, open channel calibration, and identification of cells with a single nanopore.

[0041] Nanopore cells in nanopore sensor chip, such as nanopore cells 150 in nanopore sensor chip 100, can enable parallel sequencing using a single molecule nanopore based sequencing by synthesis technique.

[0042] FIG. 2 is a block diagram illustrating an embodiment of a method 200 for detecting minimum residual disease (MRD), according to aspects of the present disclosure. At step 210, the method 200 may include detecting, using a processor, e.g., processor 310 of FIG. 3, somatic variants from pre-treatment samples from a patient using a variant caller. In some embodiments, the pre-treatment samples may include pre-treatment tumor-normal samples from the patient. The variant caller may be, for example, a deep learning-based variant caller. As one example, the deep learning-based variant caller may be the NeuSomatic network developed by Sahraeian et al. It should be understood by those of ordinary skill in the art that this is merely one example of a variant caller and that other variant callers may be implemented in accordance with aspects of the present disclosure.

[0043] In some embodiments, the detecting the somatic variants may be based on data obtained using a whole genome sequencing (WGS) process of the nanopore sensor chip 100. For example, the WGS process may include a synthesis based sequencing process, as should be understood by those of ordinary skill in the art. As another example, the WGS process may include a sequencing by expansion process. It should be understood by those of ordinary skill in the art that the nanopore sensor chip 100 is merely one example of a sequencer that may be used to perform the WGS process, and that other sequencers may also be used in accordance with aspects of the present disclosure. A somatic mutation may be any alteration at the level of DNA in somatic tissues occurring after fertilization, as should be understood by those of ordinary skill in the art. By using the WGS process, the present disclosure provides for improved performance compared to traditional targeted sequencing assays in both in-silico admixture samples and dilution series. For example, the processes described herein may detect tumor fractions (TFs) as low as 10A-6 in in-silico admixture samples and as low as 10A-5 in in-vitro dilution series. In contrast, existing processes may detect TFs as low as 10A-4 in in-vitro samples.

[0044] At step 220, the method 200 further includes filtering any germline mutations and sequencing artifacts from the detected somatic mutations. Germline mutations may be any changes to a person’s DNA that is inherited from the egg and sperm cells during conception. Sequencing artifacts may be any variations introduced by non-biological processes during the sequencing process. In some embodiments, the germline mutations and sequencing artifacts may be defined in a plurality of databases. The plurality of databases may include, but are not limited to, public databases such as 1000 Genome, ExAC, COSMIC, and TCGA. It should be understood by those of ordinary skill in the art that these are merely example databases that define germline mutations and sequencing artifacts, and that other databases may be used in accordance with aspects of thepresent disclosure. In some embodiments, any SNVs in the publicly available databases may be annotated. For example, the annotations may be added using Nirvana, a publicly available annotator used to identify clinical-grade annotation of genomic variants. In some embodiments, filtering any germline mutations and sequencing artifacts may also be based on common biomarker SNVs stored on the nanochip workstation 120. Applying these filters improves signal-to-noise and sensitivity for detecting ctDNA in plasma.

[0045] In some embodiments, the filtering any germline mutations and sequencing artifacts may include comparing the detected somatic mutations in the sample with a threshold frequency. For example, the detected somatic mutations in the sample may be categorized as a germline mutation and / or a sequencing artifact, and filtered accordingly, when the detected somatic mutations satisfies the threshold frequency. For example, the threshold frequency may be an allele frequency that is more than 0.001 for each of a number of different geographic regions, e.g., East Asia, South Asia, Europe, Africa, and the Americas. In other words, when a populationwide allele frequency is less than 0.001, the detected somatic mutations may be categorized as a somatic mutation, whereas when the allele frequency is more than 0.001, the detected somatic mutations may be categorized as a germline mutation and / or a common sequencing artifact. As another example, the filtering any germline mutations and sequencing artifacts may also include removing any germline mutations based on a normal sample from the patient. As a result, any remaining somatic mutations may be patient specific, which may then be searched for in pretreatment and / or post- treatment samples from the patient. As a result, the remaining somatic mutations may be used to detect any MRD in the sample. It should be understood by those of ordinary skill in the art that the population- wide allele frequency is distinguishable from the allelefrequency within each ctDNA sample, which measures a fraction of mutated DNA molecules (and by proxy, mutated cells) within that sample.

[0046] At step 230, the method 200 includes classifying sequencing reads covering remaining somatic variants from the filtering using a machine learning model, e.g., machine learning model 155 of FIG. 1, to identify somatic variants used for MRD detection. For example, the classifying the sequencing reads may include classifying the sequencing reads as non-noisy sequencing reads and noisy sequencing reads. Additionally, the noisy sequencing reads may be suppressed (or removed). The non-noisy sequencing reads may then be used to identify somatic variants in the patient that may be used for MRD detection. In this way, the machine learning model 155 may reduce background noise by removing any sequencing reads that contained errors from library preparation and sequencing.

[0047] The machine learning model 155 may be trained using a random forest classifier based on a plurality of features. For example, the plurality of features may include, but are not limited to: a base quality score for a variant; a mapping quality of the read; a mean base quality for the whole read; a ternary value (0, 1 , 2) that determines a support of the variant by a first read and a second read (if the variant is in an overlap between the first and second reads, it will be 1, if only one of the reads support it, it will be 0, and if the variant is outside the overlap region, it is 2); a mean base quality for the aligned portion of the read; a number of mis-matching bases; a fraction of identical bases in gapped q-align and s-align sequences; a position of the variant in the read, corrected for clipping and orientation; an edit distance of the whole read; an “A to G” or “T to C” mutation; a “C to T” or “G to A” mutation; a “C to A” to “G to T” mutation; an “A to C” or “T to G” mutation; an “A to T” or “T to A” mutation; a “C to G” or “G to C” mutation; a number of matching bases; a binary value that determines if the read is properly mapped; a binary value thatdetermines if the read is the first read or the second read; a number of insertion events; a number of deletion events; a length of a homo-polymer at a position of the variant; a mean quality score of the homo-polymer bases at a position of the variant; a distance to the closest indels within the read; a trinucleotide context of the variant; an edit distance within 3 base pair (bp) window; an edit distance within 5bp window; an edit distance within 7bp window; and a length of the fragment.

[0048] In some embodiments, the plurality of features used to train the machine learning model 155 may be based on a type of sequencing used on the pre-treatment samples. For example, when the sequencing comprises the synthesis based sequencing process, the machine learning model 155 may trained using a first set of features from among the plurality of features, and when the sequencing comprises the sequencing by expansion process, the machine learning model 155 may be trained using a second set of features from among the plurality of features. In some embodiments, a combination of features used in the first set of features is different than a combination of features used in the second set of features. Thus, the machine learning model 155 may be uniquely trained based on a sequencing type.

[0049] In some embodiments, the plurality of features may be assigned different weights. The different weights assigned to the plurality of features may be based on which features are stronger indicators of MRD. For example, for features that are stronger indicators of MRD, the weight assigned may be greater than the weight assigned to features that are weaker indicators of MRD. In some embodiments, a total value of the weights assigned to the features may be 100%.

[0050] In some embodiments, the machine learning model 155 may be trained using positive and negative training datasets. For example, the positive training dataset may include data defining germline mutations and the negative training dataset may include data indicating sequencing artifacts. As the machine learning model 155 is trained using data unrelated to anyparticular type of cancer, the machine learning model 155 is cancer agnostic and, as such, can be used to detect any type of cancer.

[0051] At step 240, the method 200 includes analyzing the identified somatic variants to determine whether the patient is positive or negative for minimum residual disease (MRD). In some embodiments, the analyzing the identified somatic variants may include analyzing a signal- to-noise ratio based on the identified somatic variants to determine whether the patient is positive or negative for MRD. As one example, the analyzing the identified somatic variants may include comparing the identified somatic variants with a set of healthy samples. In some embodiments, the set of healthy samples may include samples from at least one third party, i.e., someone other than the patient. In some embodiments, the set of healthy samples may include samples for a plurality of third parties. When a number of the identified somatic variants in the pre- or post-treatment cfDNA sample exceeds a number of the identified somatic variants in the set of healthy samples, the analyzing the identified somatic variants may render a positive determination for MRD. In contrast, when the number of the identified somatic variants in the pre- or post-treatment cfDNA sample is not higher than the number of identified somatic variants in the set of healthy samples, the analyzing the identified somatic variants may render a negative determination for MRD.

[0052] As another example, the analyzing the identified somatic variants may include comparing a normalized number of the identified somatic variants in a pre- or post-treatment cfDNA sample with a normalized set of somatic variants in a set of healthy samples. For example, in some embodiments, the identified somatic variants in the pre- or post-treatment cfDNA sample and / or the set of healthy samples may be normalized based on a total number of sequencing reads. By normalizing the data, the identified somatic variants may be converted to a standardized format, such that the data may be more readily compared to one another. In some embodiments, when thenormalized number of the identified somatic variants in the pre- or post-treatment cfDNA sample exceeds the normalized set of somatic variants in the set of healthy samples, the analyzing the identified somatic variants in the pre- or post-treatment cfDNA sample renders a positive determination for MRD, and when the normalized number of the identified somatic variants in the pre- or post-treatment cfDNA sample is less than the normalized set identified somatic variants in the set of healthy samples, the analyzing the identified somatic variants renders a negative determination for MRD.

[0053] As another example, the analyzing the identified somatic variants may include comparing a ratio between the number of identified somatic variants in the pre- or post-treatment cfDNA sample and the number of the identified somatic variants in the set of healthy samples to a threshold level. When the ratio exceeds the threshold level, the analyzing the identified somatic variants may render a positive determination for MRD. In contrast, when the ratio does not exceed the threshold level, the analyzing the identified somatic variants may render a negative determination for minimum residual disease MRD. For example, the threshold of the ratio may be 1: 1, 3:2, or 2:1. It should be understood by those ordinary skill in the art that these are merely example thresholds and other thresholds are contemplated in accordance with aspects of the present disclosure.

[0054] In some embodiments, at step 250, the method 200 may optionally include generating a report indicating whether the determination of MRD is positive or negative. That is, the report may be used to indicate whether the patient is at risk of relapsing. Using this report, one or more healthcare providers may develop at least one treatment plan for the patient. In some embodiments, the report may include the ratio, which somatic mutations were analyzed, and so on.

[0055] FIG. 3 is a block diagram illustrating one embodiment of a computer system 300 configured to implement one or more aspects of the present disclosure. For example, the method 200 of FIG. 2 may be implemented using the computer system 300.

[0056] As shown in FIG. 3, computing system 300 can include a processor 310, a memory 320, a storage device 330, and input / output devices 340. Processor 310, memory 320, storage device 330, and input / output devices 340 can be interconnected via system bus 350. Processor 310 is capable of processing instructions for execution within the computing system 300. In some example embodiments, processor 310 can be a single-threaded processor. Alternately, processor 310 can be a multi-threaded processor. Processor 310 is capable of processing instructions stored in memory 320 and / or on the storage device 330 to display graphical information for a user interface provided via the input / output device 340.

[0057] Memory 320 is a computer readable medium such as volatile or non-volatile that stores information within computing system 300. Memory 320 can store data structures representing configuration object databases, for example. Storage device 330 is capable of providing persistent storage for computing system 300. Storage device 330 can be a floppy disk device, a hard disk device, an optical disk device, or a tape device, or other suitable persistent storage means. Input / output device 340 provides input / output operations for the computing system 300. In some example embodiments, input / output device 340 includes a keyboard and / or pointing device. In various implementations, the input / output device 340 includes a display unit for displaying graphical user interfaces.

[0058] According to some example embodiments, input / output device 340 can provide input / output operations for a network device. For example, input / output device 340 can includeEthernet ports or other networking ports to communicate with one or more wired and / or wireless networks (e.g., a local area network (LAN), a wide area network (WAN), the Internet).

[0059] In some example embodiments, computing system 300 can be used to execute various interactive computer software applications that can be used for organization, analysis and / or storage of data in various formats. Alternatively, computing system 300 can be used to execute any type of software applications. These applications can be used to perform various functionalities, e.g., planning functionalities (e.g., generating, managing, editing of spreadsheet documents, word processing documents, and / or any other objects, etc.), computing functionalities, communications functionalities, etc. The applications can include various add-in functionalities or can be standalone computing products and / or functionalities. Upon activation within the applications, the functionalities can be used to generate the user interface provided via input / output device 340. The user interface can be generated and presented to a user by computing system 300 (e.g., on a computer screen monitor, etc.).

[0060] One or more aspects or features of the subject matter described herein can be realized in digital electronic circuitry, integrated circuitry, specially designed application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) computer hardware, firmware, software, and / or combinations thereof. These various aspects or features can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. Therelationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0061] These computer programs, which can also be referred to as programs, software, software applications, applications, components, or code, include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object- oriented programming language, and / or in assembly / machine language. As used herein, the term “machine-readable medium” refers to any computer program product, apparatus and / or device, such as for example magnetic discs, optical disks, memory, and Programmable Logic Devices (PLDs), used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor. The machine-readable medium can store such machine instructions non-transitorily, such as for example as would a non-transient solid-state memory or a magnetic hard drive or any equivalent storage medium. The machine-readable medium can alternatively or additionally store such machine instructions in a transient manner, such as for example, as would a processor cache or other random access memory associated with one or more physical processor cores.

[0062] To provide for interaction with a user, one or more aspects or features of the subject matter described herein can be implemented on a computer having a display device, such as for example a cathode ray tube (CRT) or a liquid crystal display (LCD) or a light emitting diode (LED) monitor for displaying information to the user and a keyboard and a pointing device, such as for example a mouse or a trackball, by which the user may provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well. For example, feedbackprovided to the user can be any form of sensory feedback, such as for example visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including acoustic, speech, or tactile input. Other possible input devices include touch screens or other touch-sensitive devices such as single or multi-point resistive or capacitive track pads, voice recognition hardware and software, optical scanners, optical pointers, digital image capture devices and associated interpretation software, and the like.

[0063] When a feature or element is herein referred to as being “on” another feature or element, it can be directly on the other feature or element or intervening features and / or elements may also be present. In contrast, when a feature or element is referred to as being “directly on” another feature or element, there are no intervening features or elements present. It will also be understood that, when a feature or element is referred to as being “connected”, “attached” or “coupled” to another feature or element, it can be directly connected, attached or coupled to the other feature or element or intervening features or elements may be present. In contrast, when a feature or element is referred to as being “directly connected”, “directly attached” or “directly coupled” to another feature or element, there are no intervening features or elements present. Although described or shown with respect to one embodiment, the features and elements so described or shown can apply to other embodiments. It will also be appreciated by those of skill in the art that references to a structure or feature that is disposed “adjacent” another feature may have portions that overlap or underlie the adjacent feature.

[0064] Terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. For example, as used herein, the singular forms “a”, “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” and / or “comprising,”when used in this specification, specify the presence of stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof. As used herein, the term “and / or” includes any and all combinations of one or more of the associated listed items and may be abbreviated as “ / ”.

[0065] Spatially relative terms, such as “under”, “below”, “lower”, “over”, “upper” and the like, may be used herein for ease of description to describe one element or feature’s relationship to another element(s) or feature(s) as illustrated in the figures. It will be understood that the spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. For example, if a device in the figures is inverted, elements described as “under” or “beneath” other elements or features would then be oriented “over” the other elements or features. Thus, the exemplary term “under” can encompass both an orientation of over and under. The device may be otherwise oriented (rotated 90 degrees or at other orientations) and the spatially relative descriptors used herein interpreted accordingly. Similarly, the terms “upwardly”, “downwardly”, “vertical”, “horizontal” and the like are used herein for the purpose of explanation only unless specifically indicated otherwise.

[0066] Although the terms “first” and “second” may be used herein to describe various features / elements (including steps), these features / elements should not be limited by these terms, unless the context indicates otherwise. These terms may be used to distinguish one feature / element from another feature / element. Thus, a first feature / element discussed below could be termed a second feature / element, and similarly, a second feature / element discussed below could be termed a first feature / element without departing from the teachings of the present disclosure.

[0067] Throughout this specification and the claims which follow, unless the context requires otherwise, the word “comprise”, and variations such as “comprises” and “comprising” means various components can be co-jointly employed in the methods and articles (e.g., compositions and apparatuses including device and methods). For example, the term “comprising” will be understood to imply the inclusion of any stated elements or steps but not the exclusion of any other elements or steps.

[0068] As used herein in the specification and claims, including as used in the examples and unless otherwise expressly specified, all numbers may be read as if prefaced by the word “about” or “approximately,” even if the term does not expressly appear. The phrase “about” or “approximately” may be used when describing magnitude and / or position to indicate that the value and / or position described is within a reasonable expected range of values and / or positions. For example, a numeric value may have a value that is + / - 0.1% of the stated value (or range of values), + / - 1% of the stated value (or range of values), + / - 2% of the stated value (or range of values), + / - 5% of the stated value (or range of values), + / - 10% of the stated value (or range of values), etc. Any numerical values given herein should also be understood to include about or approximately that value, unless the context indicates otherwise. For example, if the value “10” is disclosed, then “about 10” is also disclosed. Any numerical range recited herein is intended to include all subranges subsumed therein. It is also understood that when a value is disclosed that “less than or equal to” the value, “greater than or equal to the value” and possible ranges between values are also disclosed, as appropriately understood by the skilled artisan. For example, if the value “X” is disclosed the “less than or equal to X” as well as “greater than or equal to X” (e.g., where X is a numerical value) is also disclosed. It is also understood that the throughout the application, data is provided in a number of different formats, and that this data, represents endpoints and startingpoints, and ranges for any combination of the data points. For example, if a particular data point “10” and a particular data point “15” are disclosed, it is understood that greater than, greater than or equal to, less than, less than or equal to, and equal to 10 and 15 are considered disclosed as well as between 10 and 15. It is also understood that each unit between two particular units are also disclosed. For example, if 10 and 15 are disclosed, then 11, 12, 13, and 14 are also disclosed.

[0069] Although various illustrative embodiments are described above, any of a number of changes may be made to various embodiments without departing from the scope of the disclosure as described by the claims. For example, the order in which various described method steps are performed may often be changed in alternative embodiments, and in other alternative embodiments one or more method steps may be skipped altogether. Optional features of various device and system embodiments may be included in some embodiments and not in others. Therefore, the foregoing description is provided primarily for exemplary purposes and should not be interpreted to limit the scope of the disclosure as it is set forth in the claims.

[0070] The examples and illustrations included herein show, by way of illustration and not of limitation, specific embodiments in which the subject matter may be practiced. As mentioned, other embodiments may be utilized and derived there from, such that structural and logical substitutions and changes may be made without departing from the scope of this disclosure. Such embodiments of the inventive subject matter may be referred to herein individually or collectively by the term “invention” merely for convenience and without intending to voluntarily limit the scope of this application to any single invention or inventive concept, if more than one is, in fact, disclosed. Thus, although specific embodiments have been illustrated and described herein, any arrangement calculated to achieve the same purpose may be substituted for the specific embodiments shown. This disclosure is intended to cover any and all adaptations or variations ofvarious embodiments. Combinations of the above embodiments, and other embodiments not specifically described herein, will be apparent to those of skill in the art upon reviewing the above description.

Claims

CLAIMSWhat is claimed is:

1. A method comprising: detecting somatic variants from pre-treatment samples from a patient using a variant caller; filtering any germline mutations and sequencing artifacts from the detected somatic mutations; classifying sequencing reads covering remaining somatic variants from the filtering using a machine learning model to identify somatic variants used for minimum residual disease (MRD) detection, wherein the machine learning model is trained using a random forest classifier based on a plurality of features; and analyzing the identified somatic variants to determine whether the patient is positive or negative for MRD.

2. The method of claim 1, wherein the detecting the somatic variants is based on a whole genome sequencing (WGS).

3. The method of claim 2, wherein the plurality of features used to train the machine learning model is based on a type of sequencing used on the pre-treatment samples.

4. The method of claim 3, wherein: when the type of sequencing comprises a sequencing by synthesis process, the machine learning model is trained using a first set of features from among the plurality of features; andwhen the type of sequencing comprises a sequencing by expansion process, the machine learning model is trained using a second set of features from among the plurality of features, wherein a combination of features used in the first set of features is different than a combination of features used in the second set of features.

5. The method of claim 1, wherein the plurality of features are assigned different weights.

6. The method of claim 1, wherein the filtering any germline mutations and sequencing artifacts is based on a plurality of databases defining the germline mutations and / or sequencing artifacts.

7. The method of claim 1, further comprising generating a report based on the analyzing the identified somatic variants.

8. The method of claim 1, wherein the machine learning model is trained using positive and negative training datasets, the positive training dataset comprising germline mutations and the negative training dataset comprising sequencing artifacts.

9. The method of claim 1 , wherein the analyzing the identified somatic variants comprises comparing the identified somatic variants in a pre- or post-treatment sample with a set of healthy samples.

10. The method of claim 9, wherein, when a number of the identified somatic variants in the pre- or post-treatment sample exceeds a number of the identified somatic variants in the set of healthy samples, the analyzing the identified somatic variants renders a positive determinationfor MRD, and when the number of the identified somatic variants in the pre- or post-treatment sample is less than the number of identified somatic variants in the set of healthy samples, the analyzing the identified somatic variants renders a negative determination for MRD.

11. The method of claim 9, wherein the analyzing the identified somatic variants comprises comparing a normalized number of the identified somatic variants in a pre- or post-treatment sample with a normalized set of somatic variants in a set of healthy samples.

12. The method of claim 11, wherein when the normalized number of the identified somatic variants in the pre- or post-treatment sample exceeds the normalized set of somatic variants in the set of healthy samples, the analyzing the identified somatic variants renders a positive determination for MRD, and when the normalized number of the identified somatic variants in the pre- or post-treatment sample is not higher than the normalized set of somatic variants in the set of healthy samples, the analyzing the identified somatic variants renders a negative determination for MRD.

13. A system comprising: a memory; and a processor coupled to the memory and configured to: detect somatic variants from pre-treatment samples from a patient using a variant caller; filter any germline mutations and sequencing artifacts from the detected somatic mutations;classify sequencing reads covering remaining somatic variants from the filtering using a machine learning model to identify somatic variants used for minimum residual disease (MRD) detection, wherein the machine learning model is trained using a random forest classifier based on a plurality of features; and analyze the identified somatic variants to determine whether the patient is positive or negative for MRD.

14. The system of claim 13, wherein the detecting the somatic variants is based on a whole genome sequencing (WGS).

15. The system of claim 13, wherein the plurality of features used to train the machine learning model is based on a type of sequencing used on the pre-treatment samples.

16. The system of claim 15, wherein: when the type of sequencing comprises a sequencing by synthesis process, the machine learning model is trained using a first set of features from among the plurality of features; and when the type of sequencing comprises a sequencing by expansion process, the machine learning model is trained using a second set of features from among the plurality of features, and wherein a combination of features used in the first set of features is different than a combination of features used in the second set of features.

Citation Information

Patent Citations

  • Systems and methods for detection of residual disease

    US20210002728A1

  • Detecting somatic single nucleotide variants from cell-free nucleic acid with application to minimal residual disease monitoring

    US20210125683A1