Systems and methods for training a machine learning model using fragmentomic features and applications thereof

A machine learning model trained with fragmentomic features effectively addresses the challenges of cfDNA analysis in medical diagnostics, achieving high sensitivity and specificity for disease detection and monitoring.

WO2025240896A1PCT designated stage Publication Date: 2025-11-20ROCHE SEQUENCING SOLUTIONS INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/029821
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-17
Filing Date
2025-05-16
Publication Date
2025-11-20

AI Technical Summary

Technical Problem

Analyzing cell-free DNA (cfDNA) for medical diagnostics is challenging due to wide ranges of concentrations and fragment sizes, which complicates disease detection, monitoring, and personalized medicine applications.

Method used

A method involving a machine learning model trained using fragmentomic features, including nucleosome depleted regions, fragment size entropy, fragment end profile, fragment end motif, and transcription factor occupancy, is recursively trained and validated to predict disease risk, therapeutic response, and disease recurrence, utilizing a dataset with known disease status and sequencing data.

Benefits of technology

The method achieves high sensitivity and specificity in detecting diseases like lung and colorectal cancer, with sensitivities between 85% and 95% and specificity of at least 90%, enabling accurate disease prediction and monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000010_0001
    Figure IMGF000010_0001
  • Figure IMGF000033_0001
    Figure IMGF000033_0001
  • Figure 00000042_0000
    Figure 00000042_0000
Patent Text Reader

Abstract

The disclosure is related to training a machine learning model using fragmentomic features. A method includes obtaining a dataset comprising data indicating a known disease status of each of a plurality of individuals and sequencing data for each of the plurality of individuals. The method includes recursively training a machine learning model to predict a status of a patient. Recursively training the machine learning model includes: randomly partitioning the dataset into a first partition and a second partition; selecting one or more fragmentomic features for training the machine learning model; using the selected fragmentomic features and the first partition to train the machine learning model; and applying the machine learning model to the second partition to determine a performance of the machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR TRAINING A MACHINE LEARNING MODEL USING FRAGMENTOMIC FEATURES AND APPLICATIONS THEREOF CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority of U.S. Provisional Application Serial No. 63 / 649,230, filed May 17, 2024, the disclosure of which is incorporated herein by reference. FIELD

[0002] Embodiments of the disclosure related generally to training a machine learning model using fragmentomic features, and more specifically to using machine learning model to predict whether the patient is at risk of developing a disease, determining how responsive a patient is to a prescribed therapy, and / or monitoring the patient for recurrence of a disease. BACKGROUND

[0003] Plasma-based diagnostics are considered to be an important area for non-invasiveassays and personalized medicine with application in a wide range of fields including oncology, infectious and cardiovascular diseases, and autoimmune disorders. For example, cell-free DNA (cfDNA) is a minimally invasive and real-time biomarker, and cfDNA based diagnostic approaches hold promise for improving disease detection, monitoring, and patient care. However, due to wide ranges of cfDNA concentrations and various fragment sizes, analyzing cfDNA still remains a challenge in medical diagnostics. SUMMARY OF THE DISCLOSURE

[0004] In one aspect, a method includes obtaining, by one or more computing devices, a dataset comprising data indicating a known disease status of each of a plurality of individuals and sequencing data for each of the plurality of individuals. The sequencing data includes1 P38835-WO-1fragmentomic signals. The method also includes recursively training, by the one or more computing devices, a machine learning model.

[0005] The recursively training the machine learning model includes randomly partitioning thedataset into a first partition and a second partition. The recursively training the machine learning model also includes selecting one or more fragmentomic features for training the machine learning model. The one or more fragmentomic features are selected based on fragmentomic features detected in the first partition. The recursively training the machine learning model also includes using the selected fragmentomic features and the first partition to train the machine learning model. The machine learning model is trained to at least one of: predict whether the patient is at risk of developing a disease; monitor response to therapeutic treatment; and monitor a risk of disease reoccurrence. The recursively training the machine learning model also includes applying the machine learning model to the second partition to determine a performance of the machine learning model. The machine learning model is recursively trained a predetermined number of iterations.

[0006] In one embodiment, the recursively training the machine learning model furtherincludes removing any redundant fragmentomic features from the selected fragmentomic features.

[0007] In some aspects, the selecting one or more fragmentomic features is based on at leastone of a targeted genomic region and a binned genomic region. The targeted genomic region can be based on a gene expression level, a blood expression level, tissue specific open chromatin, and / or a differential transcription factor binding site.

[0008] In one aspect, the recursively training the machine learning model further includes analyzing the selected fragmentomic features to determine whether the selected fragmentomic features satisfy a threshold requirement. The selected fragmentomic features are used to train the machine learning model when the threshold requirement is satisfied.2 P38835-WO-1

[0009] For example, the analyzing the selected features includes recursively analyzing the selected features by: performing a hyper-parameter tuning; recursively subsampling the selected fragmentomic features and analyzing the subsamples a number of times; determining whether any of the selected fragmentomic features in the subsamples exceed the threshold requirement; and using any of the selected fragmentomic features that exceeded the threshold requirement to train the machine learning model.

[0010] In a specific embodiment, the method further includes recursively analyzing the selected fragmentomic features until at least one of the selected fragmentomic features in the subsample exceeds the threshold requirement.

[0011] The one or more fragmentomic features comprise nucleosome depleted regions, a fragment size entropy, a fragment end profile, a fragment end motif, a fragment size index, and a transcription factor activity.

[0012] In a specific embodiment, the known disease status of each of a plurality of individualscomprises a positive or negative indication of the disease for each individual of the plurality of individuals.

[0013] In some aspects, the dataset further includes data indicating demographic informationfor each individual of the plurality of individuals.

[0014] In one embodiment, the first partition and the second partition comprise distinct datafrom one another. The first partition and the second partition can include at least one common data entry.

[0015] In some aspects, the machine learning model is trained using a logistic regression model.

[0016] In one embodiment, the method further includes determining an average performance of the machine learning model based on the number of iterations, and in a specific embodiment,3 P38835-WO-1the method further includes determining whether the average performance satisfies a performance requirement.

[0017] In some aspects, the method further includes retraining the machine learning model when the average performance of the machine learning model fails to satisfying a performance criteria, including but not limited to, an accuracy and a sensitivity at a predetermined specificity.

[0018] In some embodiments, a system includes a memory and a processor coupled to the memory. The processor is configured to obtain a dataset from the memory. The dataset includes data indicating a known disease status of each of a plurality of individuals and sequencing data for each of the plurality of individuals and the sequencing data includes fragmentomic signals. The processor is also configured to recursively train a machine learning model.

[0019] To recursively train the machine learning model, the processor is further configured torandomly partition the dataset into a first partition and a second partition. To recursively train the machine learning model, the processor is further configured to select one or more fragmentomic features for training the machine learning model. The one or more fragmentomic features are selected based on fragmentomic features detected in the first partition. To recursively train the machine learning model, the processor is further configured to use the selected fragmentomic features and the first partition to train the machine learning model. The machine learning model is trained to at least one of: predict whether the patient is at risk of developing a disease; monitor response to therapeutic treatment; and monitor a risk of disease reoccurrence. To recursively train the machine learning model, the processor is further configured to apply the machine learning model to the second partition to determine a performance of the machine learning model. The machine learning model is recursively trained a predetermined number of iterations.4 P38835-WO-1

[0020] A non-transitory computer readable medium storing instructions that, when executed by a hardware processor, causes the hardware processor to obtain a dataset from a memory. The dataset includes data indicating a known disease status of each of a plurality of individuals and sequencing data for each of the plurality of individuals. The sequencing data includes fragmentomic signals. The instructions further cause the hardware processor to recursively train a machine learning model.

[0021] To recursively train the machine learning model, the instructions further cause the processor to randomly partition the dataset into a first partition and a second partition. To recursively train the machine learning model, the instructions further cause the processor to select one or more fragmentomic features for training the machine learning model. The one or more fragmentomic features are selected based on fragmentomic features detected in the first partition. To recursively train the machine learning model, the instructions further cause the processor to use the selected fragmentomic features and the first partition to train the machine learning model. The machine learning model is trained to at least one of: predict whether the patient is at risk of developing a disease; monitor response to therapeutic treatment; and monitor a risk of disease reoccurrence. To recursively train the machine learning model, the instructions further cause the processor to apply the machine learning model to the second partition to determine a performance of the machine learning model. The machine learning model is recursively trained a predetermined number of iterations. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The features of the disclosure are set forth with particularity in the claims that follow. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings of which:5 P38835-WO-1

[0023] FIG.1 is an embodiment of a measurement system as described herein.

[0024] FIG.2. a block diagram illustrating an embodiment of a computer system, according toaspects of the present disclosure.

[0025] FIG.3 is a block diagram illustration an embodiment of a method for training a machine learning model as described herein.

[0026] FIG. 4 is a block diagram illustration an embodiment of a method for recursively training the machine learning model, according to aspects of the present disclosure.

[0027] FIG.5 is a block diagram illustrating a method for performing a clinical assessment for a patient as described herein.

[0028] FIG.6 is a schematic illustrating a concatenated classifier as described herein.

[0029] FIG.7 is a schematic illustrating stacked ensemble classifiers as described herein. DETAILED DESCRIPTION

[0030] The present disclosure is directed to a fragmentomics based diagnostic workflow for cfDNA biomarker discovery process and disease detection process via sophisticated machine learning models based on several fragmentomics feature types. As should be understood by those of ordinary skill in the art, cfDNA may be protected from DNA digestion by DNA binding proteins (e.g. nucleosomes, transcription factors, etc.), and DNA binding proteins may give rise to cfDNA fragments from specific genomic loci with unique length and end motif profiles. Additionally, unprotected “open” chromatin may be heavily digested into short cfDNA fragments (e.g., less than 50 base pair (bp)), which are underrepresented in sequencing workflows allowing for the estimation of open chromatin regions.

[0031] The cfDNA fragmentomics feature types described herein inform on chromatin structure and accessibility, enzymatic fragmentation signatures and DNA methylation, all of6 P38835-WO-1which may be linked to tissue specific molecular biology. The fragmentomics feature types described herein are based on: 1) estimating / scoring transcript expression from cfDNA by measuring chromatin accessibility in promoter regions of candidate genes; 2) estimating / scoring transcription factor activity by measuring nucleosome protection / positioning at known transcription factor binding sites; and 3) estimating tissue of origin by measuring fragment endpoint profiles in experimentally determined tissue specific open chromatin regions. The workflow of the present disclosure may be based on a combination of biology and data-driven approaches to identify regions of the genome, where differential fragmentomics features may be enriched and demonstrate how these features may be harnessed for the early detection of certain diseases. As one example, the workflow described herein can be used for the early detection of lung cancer and colorectal cancer with sensitivities of between 85% and 95% and at a specificity of at least 90%. As another example, the workflow of the present disclosure was used to evaluate the performance of our fragment size index feature to positively predict lung ctDNA at 10-3(0.001) limit of detection.

[0032] In implementation, a machine learning model may be used to evaluate cfDNA extracted from a patient sample that is sequenced using whole genome sequencing or targeted enrichment sequencing, and processed through an optimized high throughput next generation sequencing data processing and quality control pipeline.

[0033] To train the machine learning model, experimental data may be mined for diseaseand / or tissue specific signatures, which may then be targeted through fragmentomics feature extraction. In some embodiments, the extracted fragmentomics features may be reduced through a principal component analysis and the resultant principal components may be used as fragmentomics features for training the machine learning model.7 P38835-WO-1

[0034] In some embodiments, the extracted fragmentomic features may include, but are not limited to, nucleosome depleted regions (NDR), fragment size entropy (FSE), fragment endpoint profile (FEP), fragment end motif (FEM), fragment size index (FSI), and transcription factor occupancy (TFO). In some embodiments, NDR may be the measure of nucleosome occupancy at a number of different regions and used as an inference of gene expression and for the detection of chromatin accessibility. For example, in some embodiments, NDRs may be found in promoter regions, exon regions, and / or intron regions.

[0035] In some embodiments, FSE may be the measure of fragment length diversity in a number of different regions (e.g., promoter regions, exon regions, and / or intron regions), and used as an inference of gene expression and for the detection of chromatin accessibility. For example, FSE may increase at open chromatin sites, such as promoter regions.

[0036] According to aspects of the present disclosure, normalized cfDNA coverage at transcription start sites (TSS) may be used to infer nucleosomal occupancy and may be inversely correlated with gene expression, while fragment size entropy may be positively correlated with gene expression. In other words, FSE may be positively correlated with the whole blood gene expression levels, while NDR coverage may be negatively correlated with it. As such, the relationships between gene expression and NDR / FSE metrics at differentially expressed regions (e.g., gene promoters) may be used as biomarker descriptors for disease prediction, such as, but not limited to, colorectal cancer, lung cancer, ovarian cancer, liver cancer, esophageal cancer, breast cancer, bladder cancer, pancreatic cancer, and head and neck cancer.

[0037] In some embodiments, FEP may be an inference of open-chromatin using fragment end point clusters and used to infer tissue-of-origin, nucleosome occupancy, and chromatin structure. For example, FEPs may be used to resolve open chromatin length based on chromatin8 P38835-WO-1state and mono-nucleosomal fragments. In some embodiments, the present disclosure employs a dynamic implementation of FEP to enable accurate detection of variable length open chromatin regions (OCRs). This may include, for example, extracting mono-nucleosomal fragment start / end counts around a center of target region of interest (ROI), normalizing counts by a maximum count within the ROI, smoothing the fragment start / end counts using, for example, a locally weighted scatterplot smoothing technique, defining a start of OCR as a first peak on a first side of the TSS, and define the end of OCR as a first peak on a second side of the OCR. In some implementations, for specific target regions like TSS NDRs or known TFBS, a semi-dynamic method may be employed, whereby the OCR overlaps with the target ROI. The aforementioned methodologies allow for sensitive detection and resolution of subtle nucleosome repositioning in different samples.

[0038] In some embodiments, FEM may be a cfDNA profile using fragment endpoint motifs, and FEM may be used to determine nucleosome positioning, chromatin organization and / or structure, and nuclease activity. FEM features may include, for example, raw motif frequencies and / or motif entropy score (MES). According to some aspects, FEMs may include blunt-end motifs and break-point motifs. In some embodiments, blunt-end motifs may be 4 bp in length and include 256 motifs and break-point motifs may be 6 bp in length and include 46,656 motifs. It should be understood by those of ordinary skill in the art that the break-point motifs may be different lengths and the number of motifs may be changed accordingly. In some embodiments, the MES may be determined using equation (1): , where N is the total number of

[0039] In some embodiments, FSI may be a weighted fragment length ratio across the genomeand used for whole genome profiling of aberrant fragments. The FSI may include ratios of short9 P38835-WO-1length to long length fragments, short length to medium length fragments, and / or medium length to long length fragments. For example, short length fragments may be 30-135 bp, medium length fragments may be 135-200 bp, and long length fragments may be 200-600 bp.

[0040] In some embodiments, TFO may be an inference of transcription factor binding siteoccupancy and used to describe differential transcription factor binding / activity.

[0041] FIG.1 illustrates a measurement system 100 according to an embodiment of the present invention. The system as shown includes a sample holder 101, sample 105, and assay components 108. The sample 105 can be contacted with assay components 108 in the sample holder 101 to perform an assay in the measurement system 100 on the sample or a component thereof, wherein the assay generates a signal indicative of a physical characteristic 115. An example of a sample holder can be a multi-well plate, microfluidic chip, array, other suitable assay consumable device. Physical characteristic 115 (e.g., a voltage, a current, optical, or other physical characteristic), from the sample is detected by detector 102. Detector 102 can take a measurement at intervals (e.g., periodic intervals) to obtain data points that make up a data signal. In one embodiment, an analog-to-digital converter converts an analog signal from the detector into digital form at a plurality of times. In one embodiment, detector 102 may be a voltage, current, or optical measurement device. The sample holder 101 and detector 102 can be provided as an assay device in a single unit. A data signal 125 is sent from detector 102 to logic system 103. Data signal 125 may be stored in a local memory 135, an external memory 104, or a storage device 145.

[0042] Logic system 103 may be, or may include, a computer system, ASIC, microprocessor, etc. It may also include or be coupled with a display (e.g., monitor, LED display, etc.) and a user input device (e.g., mouse, keyboard, buttons, etc.). Logic system 103 and the other components may be part of a stand-alone or network connected computer system, or they may10 P38835-WO-1be directly attached to or incorporated in a device (e.g., a sequencing device) that includes detector 102 and / or sample holder 101. Logic system 103 may also include software that executes in a processor 120. Logic system 103 may include a computer readable medium storing instructions for controlling system 100 to perform any of the methods described herein. For example, logic system 103 can provide commands to a system that includes sample holder 101 such that sequencing or other physical operations are performed. Such physical operations can be performed in a particular order, e.g., with reagents being added and removed in a particular order. Such physical operations may be performed by a robotics system, e.g., including a robotic arm, as may be used to obtain a sample and perform an assay.

[0043] Any of the computer systems mentioned herein may utilize any suitable number of subsystems. In some embodiments, a computer system includes a single computer apparatus, where the subsystems can be the components of the computer apparatus. In other embodiments, a computer system can include multiple computer apparatuses, each being a subsystem, with internal components. A computer system can include desktop and laptop computers, tablets, mobile phones, other mobile devices, and cloud-based systems.

[0044] The logic system 103 includes subsystems shown in FIG.2 interconnected via a systembus 200. Additional subsystems such as a printer 204, keyboard 209, storage device(s) 210, monitor 207 (e.g., a display screen, such as an LED), which is coupled to display adapter 206, and others are shown. Peripherals and input / output (I / O) devices, which couple to I / O controller 201, can be connected to the computer system by any number of means known in the art such as input / output (I / O) port 208 (e.g., USB, Thunderbolt, Lightning). For example, I / O port 208 or external interface 211 (e.g. Ethernet, Wi-Fi, etc.) can be used to connect computer system 110 to a wide area network such as the Internet, a mouse input device, or a scanner. The interconnection via system bus 200 allows the central processor 203 to communicate with each subsystem and to control the execution of a plurality of instructions from system memory 20211 P38835-WO-1or the storage device(s) 210 (e.g., a fixed disk, such as a hard drive, or optical disk), as well as the exchange of information between subsystems. The system memory 202 and / or the storage device(s) 210 may embody a computer readable medium. Another subsystem is a data collection device 205, such as a camera, microphone, accelerometer, and the like. Any of the data mentioned herein can be output from one component to another component and can be output to the user.

[0045] A computer system can include a plurality of the same components or subsystems, e.g., connected together by external interface 211, by an internal interface, or via removable storage devices that can be connected and removed from one component to another component. In some embodiments, computer systems, subsystem, or apparatuses can communicate over a network. In such instances, one computer can be considered a client and another computer a server, where each can be part of a same computer system. A client and a server can each include multiple systems, subsystems, or components.

[0046] Aspects of embodiments can be implemented in the form of control logic usinghardware circuitry (e.g., an application specific integrated circuit or field programmable gate array) and / or using computer software with a generally programmable processor in a modular or integrated manner. As used herein, a processor can include a single-core processor, multi- core processor on a same integrated chip, or multiple processing units on a single circuit board or networked, as well as dedicated hardware. Based on the disclosure and teachings provided herein, a person of ordinary skill in the art will know and appreciate other ways and / or methods to implement embodiments of the present invention using hardware and a combination of hardware and software.

[0047] Any of the software components or functions described in this application may be implemented as software code to be executed by a processor using any suitable computer12 P38835-WO-1language such as, for example, Java, C, C++, C#, Objective-C, Swift, or scripting language such as Perl or Python using, for example, conventional or object-oriented techniques. The software code may be stored as a series of instructions or commands on a computer readable medium for storage and / or transmission. A suitable non-transitory computer readable medium can include random access memory (RAM), a read only memory (ROM), a magnetic medium such as a hard-drive or a floppy disk, or an optical medium such as a compact disk (CD) or DVD (digital versatile disk) or Blu-ray disk, flash memory, and the like. The computer readable medium may be any combination of such storage or transmission devices.

[0048] Such programs may also be encoded and transmitted using carrier signals adapted for transmission via wired, optical, and / or wireless networks conforming to a variety of protocols, including the Internet. As such, a computer readable medium may be created using a data signal encoded with such programs. Computer readable media encoded with the program code may be packaged with a compatible device or provided separately from other devices (e.g., via Internet download). Any such computer readable medium may reside on or within a single computer product (e.g., a hard drive, a CD, or an entire computer system), and may be present on or within different computer products within a system or network. A computer system may include a monitor, printer, or other suitable display for providing any of the results mentioned herein to a user.

[0049] Referring back to FIG.1, in some embodiments, the system 100 may execute a machine learning model 155. It should be understood by those of ordinary skill in the art that the measurement system 100 is merely one example of a sequencer that may be used to perform the processes described herein, and that other sequencers may also be used in accordance with aspects of the present disclosure. That is, the machine learning model 155 may be executed by any known sequencer in accordance with aspects of the present disclosure. In some embodiments, the machine learning model 155 may be trained using the measurement system13 P38835-WO-1120 or an external computing device. The machine learning model may be trained by obtaining, using processor 120 of FIG. 1, a dataset comprising data indicating a known disease status of each of a plurality of individuals. The dataset may be stored, for example, on memory 135. The dataset may include records pertaining to each individual of the plurality of individuals. Each record may include, for example, the known disease status, sequencing data from a sample obtained from a respective individual, and / or one or more pieces of demographic information of the respective individual. It should be understood by those of ordinary skill in the art that while the present disclosure is described with respect to cancer, the processes described herein may be used for any disease that sheds cfDNA.

[0050] In some embodiments, the known disease status of each of a plurality of individuals may include a positive or negative indication of the disease for each individual of the plurality of individuals. For example, the disease status may indicate whether a particular individual has cancer (e.g., a positive indication) or is cancer-free (e.g., a negative indication).

[0051] In some embodiments, the sequencing data may include hundreds to thousands offragmentomic features obtained from the sample. As discussed herein, the fragmentomic features may include, but are not limited to, NDRs, FSEs, FEPs, FEMs, FSIs, and TFOs. As each record incudes hundreds to thousands of fragmentomic features, the overall dataset consequently includes thousands of fragmentomic features, which may be analyzed and selected for training the machine learning model 155.

[0052] The demographic information may include, but is not limited to, age, gender, location, height, weight, race, ethnicity, and smoking history (e.g., never smoked, current smoker, former smoker and when they quit smoking, etc.). It should be understood by those of ordinary skill in the art that these are merely examples of different demographic data that may be14 P38835-WO-1included in the dataset and that other demographic information is further contemplated in accordance with aspects of the present disclosure.

[0053] Furthermore, the machine learning model 155 may be recursively trained using theprocessor 203. For example, to train the machine learning model 155, the processor 120 may randomly partition the dataset into a first partition and a second partition. In some embodiments, the dataset may be randomly partitioned at each iteration of training the machine learning model, such that the first partition used to train the machine learning model is different at each iteration. Additionally, by randomly partitioning the records at each iteration, the processes of the present disclosure enable the machine learning model 155 to be trained using a smaller dataset.

[0054] In some embodiments, the partitioning can be based on a predetermined ratio of volume of data. For example, the first partition may contain two-thirds of the dataset and the second partition may include one-third of the data (or vice-versa). It should be understood by those of ordinary skill in the art that this is merely one example for partitioning of the dataset and that the dataset may be partitioned in any proportion in accordance with aspects of the present disclosure. In some instances, the first partition and the second partition include distinct records from one another. In other words, none of the records in the first partition overlap with the records in the second partition, and vice-versa. In other instances, the first partition and the second partition may include at least one common record.

[0055] In some embodiments, the first partition may be used for training the machine learning model 155 and the second partition may be used for testing the performance of the machine learning model 155. In this way, the processes described herein allow for records in the first partition to be used for training purposes while assessing the performance of the machine learning model against the known disease status of the records in the second partition. In15 P38835-WO-1accordance with aspects of the present disclosure, the machine learning model 155 may be blinded to the known disease status of each record in the second partition.

[0056] Additionally, to train the machine learning model, one or more fragmentomic featuresmay be selected for training the machine learning model 155. The one or more fragmentomic features may include, but are not limited to, NDRs, FSEs, FEPs, FEMs, FSIs, and TFOs. In some embodiments, the one or more fragmentomic features may be selected by analyzing the first partition. To achieve this, the processor 120 may analyze the first partition to identify relevant genomic regions of the sequencing data for each record. For example, the selecting one or more fragmentomic features may be based on at least one of a targeted genomic region and a binned genomic region and the fragmentomic features identified within the region(s). The targeted genomic region may be based on a gene expression level, a blood expression level, tissue specific open chromatin, and / or a differential transcription factor binding site. Based on the relevant genomic regions, the processor 120 may then select the one or more fragmentomic features to use in training the machine learning model 155. By selecting the one or more fragmentomic features at each iteration, the processes of the present disclosure avoid overfitting when training the machine learning model 155.

[0057] To speed up the training process and to improve the accuracy of the machine learning model 155, the processor 120 may perform dimensionality reduction on the selected fragmentomic features. The dimensionality reduction may include removing certain fragmentomic features from the selected fragmentomic features. The dimensionality reduction can be performed using various statistical techniques including, for example, analysis of variance (ANOVA) or a principal component analysis (PCA). The dimensionality reduction of the selected fragmentomic features increases the interpretability of the selected fragmentomic features, while simultaneously minimizing information loss. In some embodiments, the dimensionality reduction may identify a plurality of principal components (PCs) and a16 P38835-WO-1predetermined number of the PCs may be selected as the one or more fragmentomic features used to train the machine learning model 155. For example, the number of PCs selected to train the machine learning model 155 may be ten (10) PCs, although it should be understood that more or less PCs may be used in accordance with aspects of the present disclosure.

[0058] In some embodiments, to perform the dimensionality reduction , the processor 120 may normalize raw data associated with the selected fragmentomic features. For example, normalizing the raw data associated with the selected fragmentomic features may include calculating a Z-score for each of the selected fragmentomic features and / or defining a boundaries for a feature range for each of the selected fragmentomic features. As a result, the dimensionality reduction can be applied to standardized format for the selected fragmentomic features, such that each of the selected fragmentomic features is evaluated equally.

[0059] In some embodiments, to further select the one or more fragmentomic features, the processor 120 may remove redundant fragmentomic features from the selected fragmentomic features. For example, redundant features may include fragmentomic features of the same type, e.g., NDRs from multiple regions, multiple FSEs, multiple FSIs, and so on. As another example, redundant features may include fragmentomic features of different types that provide similar genomic information. For example, redundant features may include FSI’s in different bins or realted NDRs and FSEs. In some embodiments, determining which redundant features to keep may be randomly performed by the processor 120. In some embodiments, the selected fragmentomic features may be considered redundant when the fragmentomic features provide a same type of information, and a fragmentomic feature with a lower relative coverage score may be removed.

[0060] In some embodiments, the processor 120 may curate the records in the first partition based on the selected fragmentomic features. For example, the processor 120 may remove any records with null sets for the selected fragmentomic features. That is, if any record does not17 P38835-WO-1include sequencing data that includes the selected fragmentomic features, such records may be removed from the first partition. As a result, these records do not introduce any sort of bias when training the machine learning model 155. As another example, the processor 120 may remove features that contain more than 10% null sets, and fill the null sets with medians of the features.

[0061] In some embodiments, to select the one or more fragmentomic features, the processor 120 may recursively analyze the selected fragmentomic features to determine whether the selected fragmentomic features satisfy a threshold requirement. For example, the threshold requirement may be 90%. In this way, the present disclosure may implement a data-driven approach for selecting fragmentomic features for training the machine learning model 155. In some embodiments, when the selected fragmentomic features satisfy the threshold requirement, such fragmentomic features may be used to train the machine learning model 155. In accordance with aspects of the present disclosure, this data-driven approach may be executed when training the machine learning model 155 on a new disease type, e.g., a new type of cancer.

[0062] Recursively analyzing the selected fragmentomic features may include performing ahyper-parameter tuning. For example, the hypermeter tuning may include modifying one or more values. For example, the one or more values may include, but are not limited to, a stop criteria, a training penalty value, and / or the number of PCs. Various techniques can be used for the hyper-parameter tuning including, for example, Bayesian optimization, gradient-based optimization, random search, etc.

[0063] In some embodiments, recursively analyzing the selected fragmentomic features may further include recursively subsampling the selected fragmentomic features and analyzing the subsamples a predetermined number of times. For example, the predetermined number of times may be, for example, at least one hundred (100) iterations. It should be understood by those of ordinary skill in the art that this is merely one example of the number of times the selected18 P38835-WO-1fragmentomic features may be analyzed and that any number of iterations is contemplated in accordance with aspects of the present disclosure. In some embodiments, analyzing the subsamples may be performed using a Lasso model (e.g., a regularized linear regression model) to select the fragmentomic features, which may be used to predict ctDNA content in samples. In further embodiments, analyzing the subsamples may include performing any parsimony- encouraging analysis. The analysis may also include performing, using a portion of data from the second partition, a hyper-parameter search in each iteration of the feature selection process to reduce a negative impact of suboptimal hyper-parameters on the selected one or more fragmentomic features.

[0064] In some embodiments, recursively analyzing the selected fragmentomic features may further include determining whether any of the selected fragmentomic features in the subsamples exceed the threshold requirement. For example, the threshold requirement may be a frequency threshold for the selected fragmentomic features.

[0065] In some embodiments, recursively analyzing the selected fragmentomic features mayfurther include using any of the selected fragmentomic features that exceeded the threshold requirement to train the machine learning model 155. The processor 120 may recursively analyze the selected fragmentomic features until at least one of the selected fragmentomic features exceeds the threshold requirement. To do so, the processor 120 may retain a predetermined percentage of the selected fragmentomic features and the process may be repeated starting with the hyper-parameter tuning. For example, between 5% and 25% of the selected fragmentomic features may be retained in accordance with aspects of the present disclosure.

[0066] The processor 120 may use the selected fragmentomic features and the first partition to train the machine learning model 155. In some embodiments, the machine learning model 15519 P38835-WO-1may be trained for multiple use cases. For example, the machine learning model 155 may be trained to predict whether the patient is at risk of developing a disease, determining how responsive a patient is to a prescribed therapy, and / or monitoring the patient for recurrence of a disease.

[0067] The machine learning model 155 may be trained using a classification algorithm. Forexample, the machine learning model 155 may be trained using a logistic regression algorithm (with and without penalty), a support vector classifier (SVC) algorithm, a LASSO regression algorithm, a naive bayes algorithm, a decision tree, or a support vector machines algorithm. In some embodiments, a penalty C may be 0.001, 0.01, 0.1, where C is an argument for adjusting penalty strength. In further embodiments, the machine learning model 155 may be trained using gradient boosting (XGBoost). In some embodiments, for a binary classification (at risk of disease development or not), the machine learning model 155 may be trained using logistic regression, XGBoost, and / or SVC, whereas for a quantitative classification, the machine learning model 155 may be trained using a LASSO regression algorithm. In further embodiments, the algorithm used to train the machine learning model 155 may be based on a training objective. For example, for testing an accuracy of the machine learning model, the algorithm used to train the machine learning model 155 may be a logistic regression algorithm.

[0068] Integration of selected features may be carried out in a number of ways. In someembodiments some or all of the selected features may be merged or concatenated into a single combined feature matrix prior to classification. FIG. 6 shows a schematic exemplifying the techniques of feature concatenation. As shown in FIG. 6, any set of features may be concatenated into a single combined feature matrix to be used to train the global classifier and produce the final prediction. In some instances, once concatenated, PCA or other dimensionality reduction techniques may be applied to the concatenated feature matrix, treating the number of principal components as a tunable hyperparameter as described above. The20 P38835-WO-1PCA-reduced concatenated features may then be used to train any suitable classifier types such as Random Forest (RF), Logistic Regression (LR) or Extreme Gradient Boosting (XGB), with model hyperparameters optimized using grid search.

[0069] In some embodiments, select features or feature groups may be classified independentlyand these independent classifiers may be stacked in an ensemble of classifiers. FIG. 7 shows a schematic exemplifying the techniques of a stacking ensemble approach. As shown in FIG. 7, each feature (or in some instances, groups of features) may be used to train separate base classifiers. The base classifiers used may include any of the described classifiers above (e.g. RF, LR, or XGB) and as above, PCA or other dimensionality reduction may be performed prior to training. Although described herein as machine learning models such as RF, LR, or XGB models, it will be understood that the base classifiers need not be machine learning models. In some embodiments, the base classifier may use any classifying method that generates a predictive number. For example, the base classifier may be a statistical metric or score for a given feature that generates a predictive probability. The predicted probabilities from the base classifiers may then be used as inputs to a second-level meta-classifier which may provide the final prediction. In some embodiments, the meta-classifier may be implemented as a logistic regression model. Alternatively, RF or XGB classifiers may be used for the meta-classifier. It will be understood that this stacking architecture allows the meta-classifier to integrate outputs from specialized models, and thus leverage the strengths of each feature-specific classifier.

[0070] In some embodiments, a hybrid approach may be employed using both the concatenation and the stacking approaches described above. For example, certain subsets or groups of features may be concatenated as described with respect to FIG.6 but then each group may be stacked with other groups or other individual features as described with respect to FIG. 7.21 P38835-WO-1

[0071] In some embodiments, in order to identify the most effective feature integration strategy, a comparative analysis may be conducted between concatenation, stacking, and hybrid models using cross-validation. The final selection between the approaches may be based on empirical performance metrics, ensuring that the workflow adopts the ensemble strategy that offers the highest predictive accuracy and clinical relevance.

[0072] In some embodiments, processor 120 may similarly curate the second partition based on the selected fragmentomic features. For example, the processor 120 may remove any records in the second partition with null sets for the selected fragmentomic features. That is, if any record does not include sequencing data that includes the selected fragmentomic features, such records are removed from the second partition prior to training the machine learning model 155. As another example, the processor 120 may remove features that contain more than 10% null sets, and fill the null sets with medians of the features. The processor 120 may also perform the dimensionality reduction on the fragmentomic features on the records in the second partition.

[0073] The processor 120 may apply the machine learning model 155 to the second partition to determine a performance of the machine learning model 155. For example, the machine learning model 155 may be applied to the second partition to predict whether the patient is at risk of developing a disease. As the disease status of each individual is known, the accuracy of the predictions from the machine learning model 155 may be assessed against the known disease status. To do so, the processor 120 may calculate an accuracy of the machine learning model 155 by comparing a prediction for each record against its known disease status.

[0074] The processor 120 may recursively train the machine learning model 155 a predetermined number of iterations. For example, the predetermined number of iterations may be ten (10), twenty (20), fifty (50), or one hundred (100) iterations. It should be understood by those of ordinary skill in the art that these are merely examples of the number of iterations that22 P38835-WO-1the machine learning model 155 may be trained, and that the machine learning model may be trained using any number of iterations.

[0075] At each iteration of recursively training the machine learning model 155, the processor120 may determine the performance of the machine learning model 155, as discussed herein. At the conclusion of the predetermined number of iterations, the processor 120 may determine an average performance of the machine learning model 155 based on the performance of the machine learning model 155 at each iteration. That is, the overall performance of the machine learning model 155 may be determined by calculating the average of the performance of the machine learning model 155 at each iteration.

[0076] The processor 120 may then determine whether the average performance of the machine learning model 155 satisfies a performance requirement. The performance requirement may be, for example, a ninety percent (90%) sensitivity at ninety percent (90%) specificity. In some embodiments, when the average performance of the machine learning model fails 155 to satisfy the performance criteria, the machine learning model 155 may be retrained using the processes described herein; otherwise, the machine learning model 155 may be finalized and an assay may be developed accordingly. In some embodiments, retraining the machine learning model 155 may include modifying one or more parameters of the machine learning model 155 or modifying one or more features used to train the machine learning model 155.

[0077] FIG. 3 is a block diagram illustration an embodiment of a method 300 for training a machine learning model. At step 310, the method 300 may include obtaining, by one or more computing devices, a dataset comprising data indicating a known disease status of each of a plurality of individuals and sequencing data for each of the plurality of individuals. In some embodiments, the sequencing data may include fragmentomic signals. The dataset may be23 P38835-WO-1stored, for example, on memory 135 of the logic system shown in FIG. 1. The dataset may include records pertaining to each individual of the plurality of individuals. Each record may include, for example, the known disease status, sequencing data including fragmentomic features, and / or one or more pieces of demographic information.

[0078] At step 320, the method 300 may include recursively training, by the one or morecomputing devices, a machine learning model, e.g., the machine learning model 155 of FIG.1. A method 400 for recursively training the machine learning model is described with respect to FIG. 4. At step 410, the method 400 includes randomly partitioning the dataset into a first partition and a second partition. In some embodiments, the first partition may be used for training the machine learning model and the second partition may be used for testing a performance of the machine learning model. In some instances, the first partition and the second partition include distinct records from one another. In other words, none of the records in the first partition overlaps with the records in the second partition, and vice-versa. In other instances, the first partition and the second partition may include at least one common record.

[0079] At step 420, the method 400 may also include selecting one or more fragmentomicfeatures for training the machine learning model. The one or more fragmentomic features may include, but are not limited to, NDRs, FSEs, FEPs, FEMs, FSIs, and TFOs. In some embodiments, the one or more fragmentomic features may be selected by analyzing the first partition. For example, the first partition may be analyzed to identify relevant genomic regions of the sequencing data for each record, and the one or more fragmentomic features may be selected based on at least one of a targeted genomic region and a binned genomic region and the fragmentomic features identified within the region(s). The targeted genomic region may be based on a gene expression level, a blood expression level, tissue specific open chromatin, and / or a differential transcription factor activity. The one or more fragmentomic features may then be selected based on the relevant genomic regions. By selecting the one or more24 P38835-WO-1fragmentomic features at each iteration, the processes of the present disclosure avoid overfitting when training the machine learning model.

[0080] Additionally, selecting the one or more fragmentomic features may include performing a dimensionality reduction on the selected fragmentomic features. The dimensionality reduction may include removing certain fragmentomic features from the selected fragmentomic features. The dimensionality reduction can be performed using various statistical techniques including, for example, analysis of variance (ANOVA) or a principal component analysis (PCA). In some embodiments, the dimensionality reduction may identify a plurality of principal components (PCs) and a predetermined number of the PCs (e.g., 10 PCs) may be selected as the one or more fragmentomic features used to train the machine learning model.

[0081] To perform the dimensionality reduction , raw data associated with the selected fragmentomic features may be normalized. For example, normalizing the raw data associated with the selected fragmentomic features may include calculating a Z-score for each of the selected fragmentomic features and / or defining a boundaries for a feature range for each of the selected fragmentomic features. As a result, the dimensionality reduction can be applied to standardized format for the selected fragmentomic features.

[0082] In some embodiments, selecting the one or more fragmentomic features may includeremoving redundant fragmentomic features from the selected fragmentomic features. For example, redundant features may include fragmentomic features of the same type, e.g., NDRs from multiple regions, multiple FSEs, multiple FSIs, and so on. As another example, redundant features may include fragmentomic features of different types that provide similar genomic information.

[0083] In some embodiments, selecting the one or more fragmentomic features may include recursively analyzing the selected fragmentomic features to determine whether the selected fragmentomic features satisfy a threshold requirement. When the selected fragmentomic25 P38835-WO-1features satisfy the threshold requirement, such fragmentomic features may be used to train the machine learning model. Recursively analyzing the selected fragmentomic features may include performing a hyper-parameter tuning. For example, the hypermeter tuning may include modifying one or more values. Various techniques can be used for the hyper-parameter tuning including, for example, Bayesian optimization, gradient-based optimization, random search, etc.

[0084] Recursively analyzing the selected fragmentomic features may further include recursively subsampling the selected fragmentomic features and analyzing the subsamples a predetermined number of times (e.g., at least one hundred (100). In some embodiments, the analysis may include applying a LASSO model (e.g.,a Lasso regularized linear regression) to select the fragmentomic features, which may be used to predict ctDNA content in samples. In further embodiments, analyzing the subsamples may include performing any parsimony- encouraging analysis. The analysis may also include performing, using a portion of data from the second partition, a hyper-parameter search in each iteration of the feature selection process to reduce a negative impact of suboptimal hyper-parameters on the selected one or more fragmentomic features.

[0085] In some embodiments, recursively analyzing the selected fragmentomic features may further include determining whether any of the selected fragmentomic features in the subsamples exceed the threshold requirement and using any of the selected fragmentomic features that exceeded the threshold requirement to train the machine learning model. Recursively analyzing the selected fragmentomic features until at least one of the selected fragmentomic features exceeds the threshold requirement. To do so, a predetermined percent of the selected fragmentomic features may be retained and the process may be repeated starting with the hyper-parameter tuning.26 P38835-WO-1

[0086] At step 430, the method 400 may include using the selected fragmentomic features and the first partition to train the machine learning model. In some embodiments, the machine learning model may be trained for multiple different use cases. For example, the machine learning model may be trained to predict whether the patient is at risk of developing a disease, determining how responsive a patient is to a prescribed therapy, and / or monitoring the patient for recurrence of a disease. The machine learning model may be trained using a classification algorithm. For example, the machine learning model may be trained using a logistic regression algorithm (with and without penalty), a support vector classifier algorithm, a LASSO regression algorithm, a naive bayes algorithm, a decision tree, or a support vector machines algorithm. In further embodiments, the machine learning model may be trained using gradient boosting (XGBoost). In some embodiments, for a binary classification (at risk of disease development or not), the machine learning model may be trained using logistic regression, XGBoost and SVC, whereas for a quantitative classification, the machine learning model may be trained using a LASSO regression algorithm.

[0087] At step 440, the method 400 may include applying the machine learning model to thesecond partition to determine a performance of the machine learning model. For example, the machine learning model may be applied to the second partition to predict whether the patient is at risk of developing a disease. As the disease status of each individual is known, the accuracy of the predictions from the machine learning model may be assessed against the known disease status. Determining a performance of the machine learning model may thus include calculating an accuracy of the machine learning model by comparing a prediction for each record against its known disease status.

[0088] At step 330, the method 300 may include determining an average performance of the machine learning model based on the number of iterations. At the conclusion of the predetermined number of iterations, determining the average performance of the machine27 P38835-WO-1learning model based on the performance of the machine learning model at each iteration. That is, the overall performance of the machine learning model may be determined by calculating the average of the performance of the machine learning model at each iteration.

[0089] At step 340, the method 400 may include determining whether the average performancesatisfies a performance requirement. The performance requirement may be, for example, a ninety percent (90%) accuracy at ninety percent (90%) specificity. In some embodiments, when the average performance of the machine learning model fails to satisfy the performance criteria, the machine learning model may be retrained using the processes described herein; otherwise, the machine learning model may be finalized and an assay may be developed accordingly.

[0090] FIG.5 is a block diagram illustrating a method 500 for performing a clinical assessment for a patient, according to aspects of the present disclosure. The clinical assessment may include early disease detection and classification, therapy selection and molecular response, and / or disease monitoring and reoccurrence risk. In some embodiments, the method 500 can be performed by, for example, using the measurement system 100 of FIG.1 or the computing system of FIG.2.

[0091] At step 510, the method 500 may include receiving, at a computing device (e.g., the computing device of FIG. 2), sequencing data of a patient. In some embodiments, the sequencing data may be generated by the device shown in FIG.1. The sequencing data may be based on a sequencing by synthesis process. As another example, the sequencing may be based on a sequencing by expansion process. The sequencing data can be derived from either double- stranded or single-stranded DNA molecule samples of the patient. In some embodiments, the sequencing data may be based on a targeted enrichment process on samples collected from the patient.28 P38835-WO-1

[0092] At step 520, the method 500 may include preprocessing, using the computing device, the sequencing data. For example, preprocessing the sequencing data may include identifying any sequencing reads with failed samples and filtering the failed samples accordingly. As another example, preprocessing the sequencing data may also include mapping any remaining sequencing reads and extracting any fragmentomic features from the mapped sequencing reads.

[0093] At step 530, the method 500 may include executing, using the computing device, a machine learning model (e.g., the machine learning model 155 of FIG. 1) to analyze the preprocessed sequencing data to perform a clinical assessment. For example, the clinical assessment may include early disease detection and classification, therapy selection and molecular response, and / or disease monitoring and reoccurrence risk. As discussed herein, the machine learning model 155 may include a trained classifier based on linear regression. In some examples, the machine learning model 155 can be trained using the processes described herein.

[0094] At step 540, the method 500 may include generating, using the computing device, anoutput of the clinical assessment. The output can include display signals, audio signals, etc., to indicate a result of the assessment. In some embodiments, the processor 120 may also control a medication administration system to administer (or not to administer) a medication to the patient based on the result of assessment. In further embodiments, the processor 120 may also generate a recommended therapeutic treatment plan for the patient based on the result of the assessment.

[0095] The methods FIGs. 3, FIG.4, and / or FIG.5 may be implemented, in part or in whole, using the computer system 200.

[0096] In some example embodiments, computing system 200 can be used to execute various interactive computer software applications that can be used for organization, analysis and / or29 P38835-WO-1storage of data in various formats. Alternatively, computing system 200 can be used to execute any type of software applications. These applications can be used to perform various functionalities, e.g., planning functionalities (e.g., generating, managing, editing of spreadsheet documents, word processing documents, and / or any other objects, etc.), computing functionalities, communications functionalities, etc. The applications can include various add- in functionalities or can be standalone computing products and / or functionalities. Upon activation within the applications, the functionalities can be used to generate the user interface provided via input / output device 208. The user interface can be generated and presented to a user by computing system 200 (e.g., on a computer screen monitor, etc.).

[0097] One or more aspects or features of the subject matter described herein can be realized in digital electronic circuitry, integrated circuitry, specially designed application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) computer hardware, firmware, software, and / or combinations thereof. These various aspects or features can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0098] These computer programs, which can also be referred to as programs, software, software applications, applications, components, or code, include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object- oriented programming language, and / or in assembly / machine language. As used herein, the30 P38835-WO-1term “machine-readable medium” refers to any computer program product, apparatus and / or device, such as for example magnetic discs, optical disks, memory, and Programmable Logic Devices (PLDs), used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor. The machine-readable medium can store such machine instructions non-transitorily, such as for example as would a non-transient solid-state memory or a magnetic hard drive or any equivalent storage medium. The machine-readable medium can alternatively or additionally store such machine instructions in a transient manner, such as for example, as would a processor cache or other random access memory associated with one or more physical processor cores.

[0099] To provide for interaction with a user, one or more aspects or features of the subject matter described herein can be implemented on a computer having a display device, such as for example a cathode ray tube (CRT) or a liquid crystal display (LCD) or a light emitting diode (LED) monitor for displaying information to the user and a keyboard and a pointing device, such as for example a mouse or a trackball, by which the user may provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well. For example, feedback provided to the user can be any form of sensory feedback, such as for example visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including acoustic, speech, or tactile input. Other possible input devices include touch screens or other touch-sensitive devices such as single or multi-point resistive or capacitive track pads, voice recognition hardware and software, optical scanners, optical pointers, digital image capture devices and associated interpretation software, and the like.31 P38835-WO-1

[0100] When a feature or element is herein referred to as being “on” another feature or element, it can be directly on the other feature or element or intervening features and / or elements may also be present. In contrast, when a feature or element is referred to as being “directly on” another feature or element, there are no intervening features or elements present. It will also be understood that, when a feature or element is referred to as being “connected," “attached” or “coupled” to another feature or element, it can be directly connected, attached or coupled to the other feature or element or intervening features or elements may be present. In contrast, when a feature or element is referred to as being “directly connected," “directly attached” or “directly coupled” to another feature or element, there are no intervening features or elements present. Although described or shown with respect to one embodiment, the features and elements so described or shown can apply to other embodiments. It will also be appreciated by those of skill in the art that references to a structure or feature that is disposed “adjacent” another feature may have portions that overlap or underlie the adjacent feature.

[0101] Terminology used herein is for the purpose of describing particular embodiments onlyand is not intended to be limiting of the disclosure. For example, as used herein, the singular forms “a," “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof. As used herein, the term “and / or” includes any and all combinations of one or more of the associated listed items and may be abbreviated as “ / ”. Spatially re

[0102] lative terms, such “under," “below," “lower," “over," “upper” and the like, may be used herein for ease of description to describe one element or feature’s relationship to another element(s) or feature(s) as illustrated in the figures. It will be understood that the32 P38835-WO-1spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. For example, if a device in the figures is inverted, elements described as “under” or “beneath” other elements or features would then be oriented “over” the other elements or features. Thus, the exemplary term “under” can encompass both an orientation of over and under. The device may be otherwise oriented (rotated 90 degrees or at other orientations) and the spatially relative descriptors used herein interpreted accordingly. Similarly, the terms “upwardly," “downwardly," “vertical," “horizontal” and the like are used herein for the purpose of explanation only unless specifically indicated otherwise.

[0103] Although the terms “first” and “second” may be used herein to describe various features / elements (including steps), these features / elements should not be limited by these terms, unless the context indicates otherwise. These terms may be used to distinguish one feature / element from another feature / element. Thus, a first feature / element discussed below could be termed a second feature / element, and similarly, a second feature / element discussed below could be termed a first feature / element without departing from the teachings of the present disclosure.

[0104] Throughout this specification and the claims which follow, unless the context requiresotherwise, the word “comprise," and variations such as “comprises” and “comprising” means various components can be co-jointly employed in the methods and articles (e.g., compositions and apparatuses including device and methods). For example, the term “comprising” will be understood to imply the inclusion of any stated elements or steps but not the exclusion of any other elements or steps.

[0105] As used herein in the specification and claims, including as used in the examples and unless otherwise expressly specified, all numbers may be read as if prefaced by the word “about” or “approximately,” even if the term does not expressly appear. The phrase “about” or33 P38835-WO-1“approximately” may be used when describing magnitude and / or position to indicate that the value and / or position described is within a reasonable expected range of values and / or positions. For example, a numeric value may have a value that is + / - 0.1% of the stated value (or range of values), + / - 1% of the stated value (or range of values), + / - 2% of the stated value (or range of values), + / - 5% of the stated value (or range of values), + / - 10% of the stated value (or range of values), etc. Any numerical values given herein should also be understood to include about or approximately that value, unless the context indicates otherwise. For example, if the value “10” is disclosed, then “about 10” is also disclosed. Any numerical range recited herein is intended to include all sub-ranges subsumed therein. It is also understood that when a value is disclosed that “less than or equal to” the value, “greater than or equal to the value” and possible ranges between values are also disclosed, as appropriately understood by the skilled artisan. For example, if the value “X” is disclosed the “less than or equal to X” as well as “greater than or equal to X” (e.g., where X is a numerical value) is also disclosed. It is also understood that the throughout the application, data is provided in a number of different formats, and that this data, represents endpoints and starting points, and ranges for any combination of the data points. For example, if a particular data point “10” and a particular data point “15” are disclosed, it is understood that greater than, greater than or equal to, less than, less than or equal to, and equal to 10 and 15 are considered disclosed as well as between 10 and 15. It is also understood that each unit between two particular units are also disclosed. For example, if 10 and 15 are disclosed, then 11, 12, 13, and 14 are also disclosed.

[0106] Although various illustrative embodiments are described above, any of a number of changes may be made to various embodiments without departing from the scope of the disclosure as described by the claims. For example, the order in which various described method steps are performed may often be changed in alternative embodiments, and in other alternative embodiments one or more method steps may be skipped altogether. Optional34 P38835-WO-1features of various device and system embodiments may be included in some embodiments and not in others. Therefore, the foregoing description is provided primarily for exemplary purposes and should not be interpreted to limit the scope of the disclosure as it is set forth in the claims.

[0107] The examples and illustrations included herein show, by way of illustration and not of limitation, specific embodiments in which the subject matter may be practiced. As mentioned, other embodiments may be utilized and derived there from, such that structural and logical substitutions and changes may be made without departing from the scope of this disclosure. Such embodiments of the inventive subject matter may be referred to herein individually or collectively by the term “invention” merely for convenience and without intending to voluntarily limit the scope of this application to any single invention or inventive concept, if more than one is, in fact, disclosed. Thus, although specific embodiments have been illustrated and described herein, any arrangement calculated to achieve the same purpose may be substituted for the specific embodiments shown. This disclosure is intended to cover any and all adaptations or variations of various embodiments. Combinations of the above embodiments, and other embodiments not specifically described herein, will be apparent to those of skill in the art upon reviewing the above description.

[0108] All patents, patent applications, publications, and descriptions mentioned herein areincorporated by reference in their entirety for all purposes. None is admitted to be prior art.35 P38835-WO-1

Claims

CLAIMS What is claimed is:

1. A method comprising: obtaining, by one or more computing devices, a dataset comprising data indicating a known disease status of each of a plurality of individuals and sequencing data for each of the plurality of individuals, the sequencing data including fragmentomic signals; and recursively training, by the one or more computing devices, a machine learning model, wherein the recursively training the machine learning model comprises: randomly partitioning the dataset into a first partition and a second partition; selecting one or more fragmentomic features for training the machine learning model, wherein the one or more fragmentomic features are selected based on fragmentomic features detected in the first partition; using the selected fragmentomic features and the first partition to train the machine learning model, wherein the machine learning model is trained to at least one of predict whether the patient is at risk of developing a disease, monitor response to therapeutic treatment, and monitor a risk of disease reoccurrence; and applying the machine learning model to the second partition to determine a performance of the machine learning model, wherein the machine learning model is recursively trained a predetermined number of iterations.

2. The method of claim 1, wherein the recursively training the machine learning model further comprises removing any redundant fragmentomic features from the selected fragmentomic features.36 P38835-WO-13. The method of any one of the preceding claims, wherein the selecting one or more fragmentomic features is based on at least one of a targeted genomic region and a binned genomic region.

4. The method of claim 3, wherein the targeted genomic region is based on: a gene expression level; a blood expression level; tissue specific open chromatin; and a differential transcription factor binding site.

5. The method of any one of the preceding claims, wherein the recursively training the machine learning model further comprises analyzing the selected fragmentomic features to determine whether the selected fragmentomic features satisfy a threshold requirement, wherein the selected fragmentomic features are used to train the machine learning model when the threshold requirement is satisfied.

6. The method of claim 5, wherein analyzing the selected features comprises recursively analyzing the selected features by: performing a hyper-parameter tuning; recursively subsampling the selected fragmentomic features and analyzing the subsamples a number of times; determining whether any of the selected fragmentomic features in the subsamples exceed the threshold requirement; and using any of the selected fragmentomic features that exceeded the threshold requirement to train the machine learning model.37 P38835-WO-17. The method of claim 6, further comprising recursively analyzing the selected fragmentomic features until at least one of the selected fragmentomic features in the subsample exceeds the threshold requirement.

8. The method of any one of the preceding claims, wherein the one or more fragmentomic features comprise nucleosome depleted regions, a fragment size entropy, a fragment end profile, a fragment end motif, a fragment size index, and transcription factor activity score.

9. The method of any one of the preceding claims, wherein the known disease status of each of a plurality of individuals comprises a positive or negative indication of the disease for each individual of the plurality of individuals.

10. The method of any one of the preceding claims, wherein the dataset further comprises data indicating demographic information for each individual of the plurality of individuals.

11. The method of any one of the preceding claims, wherein the first partition and the second partition comprise distinct data from one another.

12. The method of any one of the preceding claims, wherein the first partition and the second partition comprise at least one common data entry.

13. The method of any one of the preceding claims, wherein the machine learning model is trained using a logistic regression model.

14. The method of any one of the preceding claims, further comprising determining an average performance of the machine learning model based on the number of iterations.38 P38835-WO-115. The method of claim 14, further comprising determining whether the average performance satisfies a performance requirement.

16. The method of claim 15, further comprising retraining the machine learning model when the average performance of the machine learning model fails to satisfy a performance criteria.

17. The method of claim 14, wherein the performance criteria comprises an accuracy and a sensitivity at a predetermined specificity.

18. The method of claim 1, wherein the step of using the selected fragmentomic features and the first partition to train the machine learning model comprises concatenating each of the selected fragmentomic features into a single feature matrix and training the machine learning model using the single feature matrix.

19. The method of claim 1, wherein the machine learning model comprises a first machine learning model, and the selected one or more fragmentomic features for training the machine learning model comprise a first set of fragmentomic features, the method further comprising: selecting additional sets of fragmentomic features including one or more fragmentomic features other than the first set of fragmentomic features to train additional machine learning models; and aggregating outputs of the first machine learning model and of each additional machine learning model into a meta machine learning model to generate a final prediction regarding the patient.39 P38835-WO-120. A system comprising: a memory; a processor coupled to the memory and configured to: obtain a dataset from the memory, the dataset comprising data indicating a known disease status of each of a plurality of individuals and sequencing data for each of the plurality of individuals, wherein the sequencing data includes fragmentomic signals; and recursively train a machine learning model, wherein to recursively train the machine learning model, the processor is further configured to: randomly partition the dataset into a first partition and a second partition; select one or more fragmentomic features for training the machine learning model, wherein the one or more fragmentomic features are selected based on fragmentomic features detected in the first partition; use the selected fragmentomic features and the first partition to train the machine learning model, wherein the machine learning model is trained to at least one of predict whether the patient is at risk of developing a disease, monitor response to therapeutic treatment, and monitor a risk of disease reoccurrence; and apply the machine learning model to the second partition to determine a performance of the machine learning model, wherein the machine learning model is recursively trained a predetermined number of iterations.

21. A non-transitory computer readable medium storing instructions that, when executed by a hardware processor, causes the hardware processor to: obtain a dataset from a memory, the dataset comprising data indicating a known disease status of each of a plurality of individuals and sequencing data for each of the plurality of individuals, wherein the sequencing data includes fragmentomic signals; and40 P38835-WO-1recursively train a machine learning model, wherein to recursively train the machine learning model, wherein the instructions further cause the processor: randomly partition the dataset into a first partition and a second partition; select one or more fragmentomic features for training the machine learning model, wherein the one or more fragmentomic features are selected based on fragmentomic features detected in the first partition; use the selected fragmentomic features and the first partition to train the machine learning model, wherein the machine learning model is trained to at least one of predict whether the patient is at risk of developing a disease, monitor response to therapeutic treatment, and monitor a risk of disease reoccurrence; and apply the machine learning model to the second partition to determine a performance of the machine learning model, wherein the machine learning model is recursively trained a predetermined number of iterations.41 P38835-WO-1

Citation Information

Patent Citations

  • Methods, systems and devices comprising support vector machine for regulatory sequence features

    WO2016183348A1

  • Analysis of fragment ends in DNA

    WO2022226389A1

  • Multi-omics assessment

    WO2023240046A2

  • Detecting liver cancer using cell-free DNA fragmentation

    WO2024098073A1