SYSTEM AND METHOD FOR IMPROVING PREDICTIVE MODELING OF OUTCOMES BASED ON ctDNA PROFILE

By training models on diverse patient cohorts and transforming genomic features into biologically relevant pathways, the method addresses biases and standardization issues in ctDNA diagnostics, improving predictive accuracy and clinical relevance.

WO2026060310A1PCT designated stage Publication Date: 2026-03-19LANTERN PHARMA INC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Existing predictive models for ctDNA-based diagnostics are biased towards ctDNA-positive patients, fail to capture pathway-level interactions, and lack standardization across assays, leading to inaccurate and non-generalizable outcomes.

Method used

A method that trains models on cohorts including both ctDNA-positive and ctDNA-negative patients, transforms variant-level signals into biologically grounded pathway and network features, and harmonizes ctDNA measurements to improve model accuracy and generalizability.

Benefits of technology

The method achieves superior discrimination and calibration across patient strata, identifying key biological pathways and enhancing predictive power for clinical outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025046268_19032026_PF_FP_ABST
    Figure US2025046268_19032026_PF_FP_ABST
Patent Text Reader

Abstract

A system and method for improving predictive modeling of patient outcomes by transforming ctDNA-linked genomic data into engineered pathway and network features and training machine-learning models that explicitly incorporate ctDNA status and amount while controlling selection bias. By broadening the cohort to include both ctDNA-positive and ctDNA-negative patients, harmonizing ctDNA across assays, and iteratively reducing features using Shapley-based importance, the approach improves predictive accuracy, calibration, and clinical relevance.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Attorney Docket No. 197585-010941 PCT

[0002] System and Method for Improving Predictive Modeling of Outcomes Based on ctDNA Profile

[0003] Cross-Reference To Related Applications

[0004] [1] This application claims the benefit of U.S. Provisional Application No. 63 / 694,152, Filed September 12, 2024. the entire disclosure of which is hereby incorporated by reference in its entirety.

[0005] Technical Field

[0006] [2] This application relates to medical diagnostics, machine learning, artificial intelligence, and precision oncology; More particularly, it concerns systems and methods for improving the accuracy and generalizability of outcome predictions by training models on cohorts that include both circulating tumor DNA (ctDNA)-positive and ctDNA-negative patients and by transforming variant-level signals into biologically grounded pathway and network features.

[0007] Background

[0008] [3] Liquid-biopsy assays that detect tumor-derived ctDNA can be increasingly used for diagnosis, response monitoring, and minimal-residual-disease (MRD) assessment. Yet many published and operational predictive models have been trained only on ctDNA-positive patients or on narrowly defined variant sets tied to targeted therapies.

[0009] [4] Such practices introduce systematic error. First, such as for ctDNA methods based on ctDNA concentration level associated with response, restricting training to ctDNA-positive samples produces selection bias toward higher tumor burden or adverse biology, which in turn degrades external validity and causes calibration drift when models can be applied to ty pical clinic populations that include ctDNA-negative patients. Second, over-focusing on single drug-gene pairs underperforms when clinical benefit or resistance is driven by pathway-level interactions, compensatory’ signaling, protein-protein interaction (PPI) network effects, or fragmentomic / methylomic context, none of which can be captured by single-variant heuristics. Third, ctDNA quantitation is inconsistent across assays platforms — owing to differing units and limits of detection — so treating “ctDNA-positive” as a monolithic risk factor ignores assayspecific thresholds and latent structure within ctDNA profiles.

[0010] [5] Accordingly, there is a need for bias-robust, biologically contextual models trained on cohorts that include both ctDNA-positive and ctDNA-negative patients; that translate variant- Attorney Docket No. 197585-010941 PCT level observations into pathway and network features; and that incorporate rigorous ctDNA harmonization, feature reduction, and interpretable outputs suitable for clinical use.

[0011] Summary

[0012] [6] This application provides a method for improving predictive modeling of patient outcomes by incorporating data from both ctDNA-positive and ctDNA-negative patients together with a novel feature-engineering approach. This broadens the scope of the predictive model and enables more accurate and statistically significant comparisons of patient outcomes. Matched genomic features — i.e., altered genomic profiles stemming from individual gene variants identified as tumor mutations in ctDNA liquid-biopsy assays and / or tissue sequencing — can be processed into functionally related biological pathways and / or PPIs to generate pathway features that provide biologically relevant context and enhance predictive power.

[0013] [7] In certain embodiments, a method improves outcome prediction by: (i) constructing a cohort that includes both ctDNA-positive and ctDNA-negative patients; (ii) harmonizing clinical and genomic inputs; (lii) transforming matched genomic features (e.g.. tumor mutations detected by ctDNA liquid biopsy or tissue NGS) into pathway features (e.g., Reactome / KEGG / GO gene sets) and / or PPI-derived network features; (iv) training a model that explicitly incorporates ctDNA status / amount with the engineered features; and (v) iteratively reducing features via Shapley-based (SHAP) importance ranking to yield a compact, accurate model.

[0014] [8] In certain embodiments, automated training randomly samples model classes, hyperparameters, and feature subsets across rounds; after each round, features can be ranked by mean SHAP importance across top performers, a top quantile (e g., 10-30%) is retained, and the process can repeat until a compact feature set (e.g., <10% of sample size) achieves prespecified validation performance (e.g., >0.85 AUROC for classification or >0.70 C-index for survival). The resulting models can be used to (i) predict outcomes (e.g., response, PFS, OS), (ii) identify candidate biomarkers / pathways with prognostic or predictive value, and (iii) support trial enrollment or therapy selection.

[0015] [9] Certain embodiments include: (1) Patient selection — identify a cohort including both ctDNA-positive and ctDNA-negative individuals to better reflect the general population; (2) Data collection — collect clinical and genomic data, including ctDNA status, clinical outcomes, and treatment information; (3) Model development — develop a predictive model using the collected data and engineered features, integrating ctDNA status and quantitative profiles as Attorney Docket No. 197585-010941 PCT key variables; (4) Outcome prediction — use the model to predict outcomes and identify influential features via model-based feature-importance methods, yielding candidates for independent biomarkers; and (5) Comparison — compare performance against traditional models that only consider ctDNA status or pre-specified mutations, highlighting improvements in accuracy, statistical significance, and clinical relevance.

[0016]

[0010] Another aspect includes iterative feature reduction and model training in which original features (binary mutations and pathway features) can be used in automated model training and selection. Model algorithms, hyperparameters, and feature subsets can be randomly sampled from an available pool; after multiple models can be trained, features can be ranked by mean SHAP values and the top 20% can be retained. Iterations continue until a highly accurate model with fewer features (less than 10% of the sample size) and validation accuracy of at least 85% is achieved. Hyperparameters can be fine-tuned based on parameters that improved performance in prior rounds.

[0017] Brief Description of the Drawings

[0018]

[0011] FIG. 1 is a schematic diagram of a system for improving predictive modeling of outcomes based on ctDNA profiles according to embodiments of the present disclosure.

[0019]

[0012] FIG. 2 is a schematic of the training pipeline, from cohort assembly to iterative SHAP- guided reduction according to embodiments of the present disclosure;

[0020]

[0013] FIG. 3 shows generation of pathway and PPI network features from variant calls according to embodiments of the present disclosure;

[0021]

[0014] FIG. 4 depicts ctDNA harmonization;

[0022]

[0015] FIG. 5 illustrates nested cross-validation, calibration, and stratified evaluation according to embodiments of the present disclosure;

[0023]

[0016] FIG. 6 is an example interpretability report according to embodiments of the present disclosure;

[0024]

[0017] FIG. 7 shows clinical deployment architecture for inference and decision support according to embodiments of the present disclosure;

[0025]

[0018] FIG. 8 provides confusion-matrix and calibration plots comparing a ctDNA-only model to the disclosed model (prophetic) according to embodiments of the present disclosure;

[0026]

[0019] FIG. 9 is a flow-chart for tri al -matching using predicted outcomes and biomarker explanations according to embodiments of the present disclosure.

[0027]

[0020] FIG. 10 a schematic diagram of a machine in the form of a computer system within which a set of instructions, when executed, may cause the machine to improve predicting Attorney Docket No. 197585-010941 PCT modeling of outcomes, such as based on ctDNA profiles, according to embodiments of the present disclosure.

[0028] Definitions

[0029]

[0021] The terms “computer engine” and “engine” can identify at least one software component and / or a combination of at least one software component and at least one hardware component which are designed / programmed / configured to manage / control other software and / or hardware components (such as the libraries, software development kits (SDKs), objects, and the like).

[0030]

[0022] The term “ctDNA-positive” can mean presence of circulating tumor DNA above an assay-specific LOD or threshold and may be expressed as binary status and / or continuous metrics (e.g., tumor fraction, variant-allele-frequency (VAF) burden, fragments / mL, methylation scores).

[0031]

[0023] The term “ctDNA-negative” can mean no detectable circulating tumor DNA or levels below the assay LOD.

[0032]

[0024] The term “matched genomic features” can mean a set of altered genes / variants for a patient derived from liquid biopsy (ctDNA) and / or tissue sequencing (e.g., SNVs / indels, copynumber alterations, fusions).

[0033]

[0025] The term “pathway features” can mean numeric summaries per pathway / gene set derived from matched genomic features (e.g.. mutation load in a pathway, enrichment scores, topology-aware propagation scores), given as a fraction between 0 and 1 where the number of mutations in a pathway are divided by the total number of genes of that pathway that were measured.

[0034]

[0026] The term “PPI network features” can mean features computed on a PPI graph (e.g., node embeddings, centrality-weighted mutation scores, random-walk diffusion).

[0035]

[0027] The term “model” can mean any machine-learning or statistical model including but not limited to gradient boosting, random forests, generalized linear models, survival models (Cox, DeepSurv), neural networks, or ensembles.

[0036]

[0028] The term “outcome” can mean a clinical endpoint such as RECIST response, disease control, PFS, OS, or a time-to-event composite.

[0037]

[0029] The term “SHAP” can mean Shapley additive explanations used to quantify feature importance and support interpretability. Attorney Docket No. 197585-010941 PCT

[0038] Detailed Description

[0039]

[0030] This application discloses exemplary approaches to patient selection and data integration that addresses limitations of existing predictive models. Prior models frequently use ctDNA status as the main or sole criterion for evaluating treatment response and / or evaluate only a small set of mutations relevant to targeted therapies. These practices limit discovery' of novel mechanisms of tumor response and can misinform treatment when ctDNA levels can be elevated yet the mutational profile is favorable for a given therapy, or when response mechanisms operate in genes and pathways other than those previously implicated. Moreover, ctDNA assessments lack standardization across assays and treatment types, and many indications lack a consistent link between absolute ctDNA level and outcome.

[0040]

[0031] In some embodiments, raw ctDNA outputs — such as variant-allele frequency (VAF) values, tumor fraction, fragments per milliliter, and related liquid-biopsy metrics — can be transformed into (a) standardized continuous measures via within-assay z-scaling and (b) binary' flags at assay-specific limits of detection or thresholds. The model optionally learns an adaptive ctDNA positivity threshold via cross-validated search to account for heterogeneous platform characteristics, thereby enabling consistent treatment of ctDNA status and amount across laboratories and disease settings.

[0041]

[0032] In certain embodiments, variant calls can be normalized to HGVS conventions and collapsed to gene-level indicators denoting presence or absence of alterations, with optional quantitative descriptors including alteration-class burden (SNV. CNV. fusion), maximum VAF, clonal / subclonal status, and fragmentomic or methylation scores when available. Clinical covariates can include ECOG performance status, line of therapy, stage, age, sex, prior therapies, and laboratory' values. Missingness may be handled with model-aware imputation (e.g., iterative imputation) while preserving leakage controls during cross-validation.

[0042]

[0033] In some embodiments, matched genomic features can be mapped to curated pathway libraries (e.g., Reactome, KEGG, Gene Ontology ) and to a protein-protein interaction (PPI) graph to derive biologically contextual representations. Pathway features can include counts or fractions of altered genes per pathway, enrichment or activity scores (e.g., GSVA-like computations), and topology -aware diffusion or random-walk propagation scores on the PPI network. Network embeddings (e g., node2vec) may be aggregated over altered nodes to yield compact features that capture convergent mechanisms not visible at the single-gene level.

[0043]

[0034] In certain embodiments, training proceeds under nested cross-validation. For each outer fold, a model family and hyperparameters can be randomly sampled from a predefined search space (e.g., gradient-boosted trees with specified depth and learning rate, penalized Cox Attorney Docket No. 197585-010941 PCT models, neural networks). Models can be trained on baseline features comprising clinical variables, ctDNA status and amount, variant indicators, and pathway / PPI features. Features can be ranked by mean Shapley additive explanation (SHAP) importance across the top-performing configurations, and atop quantile (e.g., approximately 10-30%) is retained. This selection-and- retrain cycle is iterated until stopping criteria can be met, such as a plateau in validation AUROC or concordance index and a retained feature count below a target ratio (e.g., fewer than 0. 1 x the number of patients). Class imbalance may be addressed via reweighting, focal loss, or synthetic oversampling. The final model set may be an ensemble that averages or stacks top models.

[0044]

[0035] In some embodiments, probabilistic outputs can be calibrated using isotonic regression or Platt scaling, and performance is evaluated across strata defined by ctDNA status (ctDNA- positive versus ctDNA-negative), assay type, and disease subtype to verify discrimination and calibration stability. Fairness and robustness checks may include subgroup error analysis and sensitivity to site-specific distributions.

[0045]

[0036] In certain embodiments, at inference time the engine ingests new patient data — including ctDNA positivity / negativify and any continuous ctDNA measures — computes variant, pathway, and optional PPI features, and applies the trained model to produce risk scores or survival functions. The system returns explanations identifying features driving the prediction (e.g., global importance and per-patient SHAP waterfall plots) and may trigger clinical decision support (CDS) actions such as recommending continuation versus switch of therapy, ordering reflex testing, or suggesting clinical-trial screening when resistance- associated pathways predominate. Thresholds for CDS actions can be learned during validation or configured by the user.

[0046]

[0037] In certain prophetic examples, a simulated non-small cell lung cancer (NSCLC) cohort (N=l,200; 55% ctDNA-positive, 45% ctDNA-negative) with realistic panel coverage is used to train a gradient-boosting model incorporating pathway and PPI features together with ctDNA harmonization via the iterative SHAP pipeline described herein. Relative to a ctDNA-only baseline classifier, the disclosed model achieves superior cross-validated discrimination for 6- month disease control, maintains calibration across ctDNA-positive and ctDNA-negative strata, and identifies DNA-damage response and oxidative-stress pathways as leading explanatory features despite heterogeneous single-gene drivers. These results can be prophetic and serve to illustrate potential operation and advantages.

[0047]

[0038] In some embodiments, the system executes on one or more processors (e.g., x86, ARM, GPU / TPU) with memory and storage sufficient to host feature libraries, models, and audit logs. Attorney Docket No. 197585-010941 PCT

[0048] Implementations may employ compiled and / or interpreted languages (e.g., C / C++, Python, R, Java) and be deployed as client, server, or web-enabled applications with HIPAA / GDPR- compliant access controls. Model artifacts, calibration parameters, and explanation outputs may be serialized for versioning and reproducibility.

[0049]

[0039] Once trained, these models can predict patient, tumor, and drug characteristics using data from liquid biopsies or other sources. The models can be validated on both public and private datasets and can be used to inform clinical decisions, such as predicting drug efficacy or selecting patients for clinical trials.

[0050]

[0040] Another aspect include performance in predicting outcomes across a more diverse patient population, which is not the focus of current models that concentrate on specific genes or cfDNA-status alone.

[0051]

[0041] Examples of hardware elements may include processors, microprocessors, circuits, circuit elements (e.g., transistors, resistors, capacitors, inductors, and so forth), integrated circuits, application specific integrated circuits (ASIC), programmable logic devices (PLD), digital signal processors (DSP), field programmable gate array (FPGA), logic gates, registers, semiconductor device, chips, microchips, chip sets, and so forth. In some embodiments, the one or more processors may be implemented as a Complex Instruction Set Computer (CISC) or Reduced Instruction Set Computer (RISC) processors; x86 instruction set compatible processors, multi-core, or any other microprocessor or central processing unit (CPU). In various implementations, the one or more processors may be dual-core processor(s), dual-core mobile processor(s), and so forth.

[0052]

[0042] Computer-related systems, computer systems, and systems, as used herein, include any combination of hardware and software. Examples of software may include software components, programs, applications, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, procedures, software interfaces, application program interfaces (API), instruction sets, computer code, computer code segments, words, values, symbols, or any combination thereof. Determining whether an embodiment is implemented using hardware elements and / or software elements may vary’ in accordance with any number of factors, such as desired computational rate, power levels, heat tolerances, processing cycle budget, input data rates, output data rates, memory resources, data bus speeds and other design or performance constraints.

[0053]

[0043] For the purposes of this disclosure a module is a software, hardware, or firmware (or combinations thereof) system, process or functionality, or component thereof, that performs or facilitates the processes, features, and / or functions described herein (with or without human Attorney Docket No. 197585-010941 PCT interaction or augmentation). A module can include sub-modules. Software components of a module may be stored on a computer readable medium for execution by a processor. Modules may be integral to one or more servers, or be loaded and executed by one or more servers. One or more modules may be grouped into an engine or an application.

[0054]

[0044] One or more aspects of at least one embodiment may be implemented by representative instructions stored on a machine-readable medium which represents various logic within the processor, which when read by a machine causes the machine to fabricate logic to perform the techniques described herein. Such representations, known as “IP cores,” may be stored on a tangible, machine readable medium and supplied to various customers or manufacturing facilities to load into the fabrication machines that make the logic or processor. Of note, various embodiments described herein may, of course, be implemented using any appropriate hardware and / or computing software languages (e.g., C++, Objective-C, Swift, Java, JavaScript, Python, R, Perl, QT, and the like).

[0055]

[0045] For example, exemplary software specifically programmed in accordance with one or more principles of the present disclosure may be downloadable from a network, for example, a website, as a stand-alone product or as an add-in package for installation in an existing softw are application. For example, exemplary software specifically programmed in accordance with one or more principles of the present disclosure may also be available as a client-server software application, or as a web-enabled software application. For example, exemplary software specifically programmed in accordance with one or more principles of the present disclosure may also be embodied as a software package installed on a hardware device.

[0056]

[0046] For the purposes of this disclosure the term “user” or “patient” can be understood to refer to a user of an application or applications as described herein and / or a consumer of data supplied by a data provider. By w ay of example, and not limitation, the term “user” or “patient” can refer to a person who receives data provided by the data or service provider over the Internet in a browser session, or can refer to an automated softw are application which receives the data and stores or processes the data. Those skilled in the art will recognize that the methods and systems of the present disclosure may be implemented in many manners and as such cannot to be limited by the foregoing exemplary embodiments and examples. In other words, functional elements being performed by single or multiple components, in various combinations of hardware and software or firmware, and individual functions, may be distributed among software applications at either the client level or server level or both. In this regard, any number of the features of the different embodiments described herein may be Attorney Docket No. 197585-010941 PCT combined into single or multiple embodiments, and alternate embodiments having fewer than, or more than, all of the features described herein can be possible.

[0057]

[0047] Functionality may also be, in whole or in part, distributed among multiple components, in manners now known or to become known. Thus, myriad software / hardware / firmware combinations can be possible in achieving the functions, features, interfaces and preferences described herein. Moreover, the scope of the present disclosure covers conventionally known manners for carrying out the described features and functions and interfaces, as well as those variations and modifications that may be made to the hardware or software or firmware components described herein as would be understood by those skilled in the art now and hereafter.

[0058]

[0048] In certain embodiments, a computer-implemented method for training a predictive model for modeling of outcomes, such as patient outcomes, is provided. The method can include obtaining a training cohort that includes both ctDNA-positive and ctDNA-negative patients. The method can include receiving, for each patient, clinical data and matched genomic features comprising tumor alterations detected by liquid biopsy and / or tissue sequencing. The method can include transforming at least a portion of the matched genomic features into pathway features and / or protein-protein interaction (PPI) network features. The method can include harmonizing ctDNA measurements by encoding ctDNA status and at least one normalized continuous ctDNA measure. The method can include training one or more machine-learning models using the clinical data, the harmonized ctDNA measurements, the matched genomic features, and the pathway and / or PPI network features. The method can include computing feature importances using Shapley-based explanations and iteratively reducing a feature set by retaining a top quantile of features across training rounds. The method can further include outputting a trained model and an associated compact feature set usable to predict a clinical outcome.

[0059]

[0049] In certain embodiments, the outcome can include at least one of objective response, disease control, progression-free survival, overall survival, or a time-to-event endpoint. In certain embodiments, harmonizing ctDNA measurements can include at least one of: within- assay z-scaling, learning an adaptive threshold for ctDNA positivity, or jointly modeling tumor fraction and variant-allele-frequency burden. In certain embodiments, the pathway features can be computed using one or more curated gene-set libraries selected from Reactome, KEGG, and Gene Ontology. In certain embodiments, the PPI network features can include at least one of network-propagation scores, random-walk diffusion scores, centrality-weighted mutation scores, or aggregated node embeddings. In certain embodiments, the method can further Attorney Docket No. 197585-010941 PCT include iteratively reducing the feature set comprises retaining between 10% and 30% of features by mean Shapley importance across a plurality of trained models.

[0060]

[0050] In certain embodiments, the compact feature set can contain fewer than 10% as many features as patients in the training cohort. In certain embodiments, the method can include nested cross-validation and probability calibration using isotonic regression or Platt scaling. In certain embodiments, model families can be randomly sampled from a collection comprising gradient-boosted trees, random forests, generalized linear models. Cox proportional hazards models, neural networks, and stacked ensembles. In certain embodiments, generating the explanation can include outputting at least one of global feature importance, per-patient SHAP waterfall plots, pathway attribution tables, or counterfactual feature analyses. In certain embodiments, the method can further include routing the outcome prediction to a clinical decision support rule that recommends at least one action selected from continuing current therapy, switching therapy, ordering reflex testing, or screening the patient for a clinical trial.

[0061]

[0051] In certain embodiments, matched genomic features can include one or more of singlenucleotide variants, insertions / deletions, copy-number alterations, gene fusions, methylation markers, or fragmentomic metrics. In certain embodiments, the method can further include controlling selection bias using at least one of stratified sampling, inverse-probability weighting, or site / assay batch-effect correction. In certain embodiments, the outcome prediction can be evaluated and reported separately for ctDNA-positive and ctDNA-negative strata to verily calibration. In certain embodiments, the trained model can achieve at least one of: an area under the receiver-operating characteristic curve of 0.85 or greater for a binary outcome on a validation set; or a concordance index of 0.70 or greater for a survival outcome on a validation set. In certain embodiments, the method can include enforcing pathway-level sparsity’ by applying group-lasso or structured regularization. In certain embodiments, the matched genomic features can be obtained from a liquid-biopsy panel and optionally reconciled with a tissue panel using graph-based imputation of missing genes. In certain embodiments, a non-transitory computer-readable medium storing instructions that, when executed by one or more processors, can cause the processors to perform the method of claim 1.

[0062]

[0052] In certain embodiments, a system for training a predictive model of patient outcomes can include: one or more processors; and a non-transitory memory storing instructions that, when executed by the one or more processors, cause the system to perform various operations. For example, the system can be configured to obtain a training cohort that includes both ctDNA-positive and ctDNA-negative patients; receive, for each patient, clinical data and matched genomic features comprising tumor alterations detected by liquid biopsy and / or tissue Attorney Docket No. 197585-010941 PCT sequencing; generate engineered features by transforming at least a portion of the matched genomic features into pathway features and / or PPI network features; harmonize ctDNA measurements by encoding ctDNA status and at least one normalized continuous ctDNA measure; train one or more machine-learning models using the clinical data, the harmonized ctDNA measurements, the matched genomic features, and the engineered features; compute feature importances using Shapley-based explanations and iteratively reduce a feature set by retaining a top quantile of features across training rounds; and output a trained model and an associated compact feature set usable to predict a clinical outcome.

[0063]

[0053] In certain embodiments, the training can employ nested cross-validation and probability calibration using isotonic regression or Platt scaling. In certain embodiments, the system can be further configured to mitigate selection bias using at least one of stratified sampling to preserve ctDNA-positive / negative prevalence, inverse-probability or propensity weighting to account for ctDNA testing patterns, or batch-effect correction across assay panels. In certain embodiments, pathway-level sparsity can be enforced using group-lasso or structured regularization. In certain embodiments, the system can further include a validation module configured to report discrimination and calibration separately for ctDNA-positive and ctDNA- negative strata, and an audit module configured to store model versioning, datasets used for training, selected features, and calibration parameters.

[0064]

[0054] Referring now also to Figure 1, an exemplary system 100 for improving predictive modeling of outcomes based on ctDNA profiles according to embodiments of the present disclosure is illustrated. The system 100 can be utilized to perform any of the features and functionality provided in Figures 2, 3, 4, 5, 6, 7, 8, and 9. Notably, the system 100 may be configured to support, but is not limited to supporting, healthcare systems, predictive modeling of outcomes, patient intake systems, medical diagnosis systems, medical condition analysis systems, automation systems, data analytics systems and services, data collation and processing systems and services, artificial intelligence services and systems, machine learning services and systems, content deliver}7services, cloud computing services, satellite services, telephone services, voice-over-internet protocol services (VoIP), software as a service (SaaS) applications, platform as a service (PaaS) applications, mobile applications and services, and / or any other computing applications and services. In certain embodiments,, the system 100 may include a first user 101, who may utilize a first user device 102 to access data, content, and sen ices, or to perform a variety of other tasks and functions. As an example, the first user device 102 can be utilized to access an application, devices, and / or components of the system 100 that provide any or all of the operative functions of the system 100. For example, the first Attorney Docket No. 197585-010941 PCT user 101 may utilize the first user device 102 to access an application having a user interface that enables the first user 101 to submit personal data into the system 100 to register the first user 101 with the system 100 for purposes of patient intake and assessing the first user’s outcomes in response to medicine or a pharmaceutical product. In certain embodiments, the first user 101 may seek to have a medical condition assessed and / or diagnosed, particularly if the first user 101 is experiencing symptoms that may need to be treated by a provider, such as a physician, providing the pharmaceutical product. In certain embodiments, the first user 101 may be a bystander, any type of person, a robot, a humanoid, any type of user, or a combination thereof, that may be located in a particular environment.

[0065]

[0055] In certain embodiments, the first user 101 may be a person that may be experiencing a medical condition, may be seeking to having a health checkup, may be seeking a medical treatment, or a combination thereof. For example, the first user 101 may be a patient of a provider, such as a physician (e.g., the second user 110). In certain embodiments, the first user device 102 may be utilized by the first user to interact with the system 100, other users of the system 100, or a combination thereof. In certain embodiments, the first user device 102 may include a memory 103 that includes instructions, and a processor 104 that executes the instructions from the memory 103 to perform the various operations that are performed by the first user device 102. In certain embodiments, the processor 104 may be hardware, software, or a combination thereof. The first user device 102 may also include an interface 105 (e.g., screen, monitor, graphical user interface, etc.) that may enable the first user 101 to interact with various applications executing on the first user device 102 and to interact with the system 100. In certain embodiments, the first user device 102 may be and / or may include a computer, any type of sensor, a laptop, a set-top-box, a tablet device, a phablet, a server, a mobile device, a smartphone, a smart watch, and / or any other type of computing device. Illustratively, the first user device 102 is shown as a smartphone device in Figure 1. In certain embodiments, the first user device 102 may be utilized by the first user 101 to control and / or provide some or all of the operative functionality of the system 100.

[0066]

[0056] In addition to using first user device 102, the first user 101 may also utilize and / or have access to additional user devices. As with first user device 102, the first user 101 may utilize the additional user devices to transmit signals to access various online services and content. The additional user devices may include memories that include instructions, and processors that executes the instructions from the memories to perform the various operations that are performed by the additional user devices. In certain embodiments, the processors of the additional user devices may be hardware, software, or a combination thereof. The additional Attorney Docket No. 197585-010941 PCT user devices may also include interfaces that may enable the first user 101 to interact with various applications executing on the additional user devices and to interact with the system 100. In certain embodiments, the first user device 102 and / or the additional user devices may be and / or may include a computer, any type of sensor, a laptop, a set-top-box, a tablet device, a phablet, a server, a mobile device, a smartphone, a smart watch, and / or any other type of computing device, and / or any combination thereof.

[0067]

[0057] Sensors may include, but are not limited to, cameras, motion sensors, acoustic / audio sensors, pressure sensors, temperature sensors, light sensors, heart-rate sensors, blood pressure sensors, sweat detection sensors, breath-detection sensors, stress-detection sensors, any type of health sensor, humidity sensors, any type of sensors, or a combination thereof. In certain embodiments, the sensor data generated by the sensors may be utilized to facilitate an assessment of a medical condition, the patient’s response to a pharmaceutical drug, or medical diagnosis of the first user 101. For example, if atemperature sensor of the first user device 102 detects an elevated human body temperature, the system 100 may be configured to utilize the temperature sensor data to generate an assessment that the first user 101 is likely experiencing a fever, which may have resulted from a bacterial infection, viral infection, or a combination thereof. As another example, if a heart rate sensor provides sensor data indicative of atrial fibrillation, the sensor data may facilitate the system 100 in generating an assessment indicating that the first user is experience atrial fibrillation, such as in response to a particular drug.

[0068]

[0058] The first user device 102 and / or additional user devices may belong to and / or form a communications network. In certain embodiments, the communications network may be a local, mesh, or other network that enables and / or facilitates various aspects of the functionality of the system 100. In certain embodiments, the communications network may be formed between the first user device 102 and additional user devices through the use of any type of wireless or other protocol and / or technology. For example, user devices may communicate with one another in the communications network by utilizing any protocol and / or wireless technology, satellite, fiber, or any combination thereof. Notably, the communications network may be configured to communicatively link with and / or communicate with any other network of the system 100 and / or outside the system 100.

[0069]

[0059] In certain embodiments, the first user device 102 and additional user devices belonging to the communications network may share and exchange data with each other via the communications network. For example, the user devices may share information associated with a user (e.g., patient) with each other, information relating to lab results, information relating to medical or physical examinations conducted by a physician on a user and / or Attorney Docket No. 197585-010941 PCT pharmaceutical products taken by the user, information corresponding to and / or associated with modeling of outcomes generated by the system 100, information corresponding to and / or associated with sensor data generated by sensors of the system 100, information relating to the various components of the user devices, information associated with images and / or content accessed by a user of the user devices, any other information, or any combination thereof.

[0070]

[0060] In addition to the first user 101, the system 100 may also include a second user 110. The second user 110 may be a person that may conduct examinations of the first user 101. facilitate treatment of the first user 101, recommend protocols for the first user 101 to follow, refer the first user 101 to another medical personnel, or any combination thereof. For example, in certain embodiments, the second user 110 may be a physician, nurse, technician, intake professional, pharmacist, or other individual that work at a hospital, medical practice, any other location, or a combination thereof. In certain embodiments, the second user device 111 may be utilized by the second user 110 to transmit signals to request various types of content, sendees, and data provided by and / or accessible by communications network 135 or any other network in the system 100. In certain embodiments, the second user device 111 may be utilized by the second user 110 to view patient data, generate assessments associate with medical conditions and / or diagnoses of patients, conduct predictive modeling of outcomes based on ctDNA profiles, perform any operative functionality of the system 100, or a combination thereof. In further embodiments, the second user 110 may be a person, a humanoid, an animal, any type of user, or any combination thereof. The second user device 111 may include a memory 112 that includes instructions, and a processor 113 that executes the instructions from the memor\' 112 to perform the various operations that are performed by the second user device 111. In certain embodiments, the processor 113 may be hardware, software, or a combination thereof. The second user device 111 may also include an interface 114 (e.g.. screen, monitor, graphical user interface, etc.) that may enable the first user 101 to interact with various applications executing on the second user device 111 and, in certain embodiments, to interact with the system 100. In certain embodiments, the second user device 111 may be a computer, a laptop, a set-top-box, a tablet device, a phablet, a server, a mobile device, a smartphone, a smart watch, and / or any other type of computing device. Illustratively, the second user device 111 is shown as a mobile device in Figure 1. In certain embodiments, as with the first user device 102, the second user device 111 may also include sensors, such as, but are not limited to, cameras, audio sensors, motion sensors, pressure sensors, temperature sensors, light sensors, heart-rate sensors, blood pressure sensors, sweat detection sensors, breath-detection sensors, Attorney Docket No. 197585-010941 PCT stress-detection sensors, any type of health sensor, humidity sensors, any type of sensors, or a combination thereof.

[0071]

[0061] In certain embodiments, the first user device 102, the additional user devices, and / or the second user device 111 may have any number of software applications and / or application sendees stored and / or accessible thereon. For example, the first user device 102, the additional user devices, and / or the second user device 111 may include applications for controlling and / or accessing the operative features and functionality of the system 100, applications for controlling and / or accessing any device of the system 100, applications including machine learning models for providing predictive modeling of patient outcomes based on ctDNA profiles, any other type of applications, any types of application services, or a combination thereof. In certain embodiments, the software applications may support the functionality provided by the system 100 and methods described in the present disclosure. In certain embodiments, the software applications and services may include one or more graphical user interfaces so as to enable the first and / or potentially second users 101, 110 to readily interact with the software applications. The software applications and services may also be utilized by the first and / or potentially second users 101. 110 to interact with any device in the system 100. any netw ork in the system 100, or any combination thereof. In certain embodiments, the first user device 102, the additional user devices, and / or potentially the second user device 111 may include associated telephone numbers, device identities, or any other identifiers to uniquely identity’ the first user device 102. the additional user devices, and / or the second user device 1 11.

[0072]

[0062] The system 100 may also include a communications network 135. The communications network 135 may be under the control of a service provider, any designated user, a computer, another network, or a combination thereof. The communications network 135 of the system 100 may be configured to link each of the devices in the system 100 to one another. For example, the communications network 135 may be utilized by the first user device 102 to connect with other devices within or outside communications network 135. Additionally, the communications network 135 may be configured to transmit, generate, and receive any information and data traversing the system 100. In certain embodiments, the communications network 135 may include any number of servers, databases, or other componentry. The communications network 135 may also include and be connected to a mesh network, a local network, a cloud-computing network, an IMS network, a VoIP network, a security network, a VoLTE network, a wireless network, an Ethernet network, a satellite network, a broadband network, a cellular network, a private network, a cable network, the Internet, an internet Attorney Docket No. 197585-010941 PCT protocol network, MPLS network, a content distribution network, any network, or any combination thereof. Illustratively, servers 140, 145, and 150 are shown as being included within communications network 135. In certain embodiments, the communications network 135 may be part of a single autonomous system that is located in a particular geographic region or be part of multiple autonomous systems that span several geographic regions.

[0073]

[0063] Notably, the functionality of the system 100 may be supported and executed by using any combination of the servers 140. 145. 150, and 160. The servers 140. 145, and 150 may reside in communications network 135, however, in certain embodiments, the servers 140, 145, 150 may reside outside communications network 135. The servers 140, 145, and 150 may provide and serve as a server service that performs the various operations and functions provided by the system 100. In certain embodiments, the server 140 may include a memory

[0074] 141 that includes instructions, and a processor 142 that executes the instructions from the memory 141 to perform various operations that are performed by the server 140. The processor

[0075] 142 may be hardware, software, or a combination thereof. Similarly, the server 145 may include a memory 146 that includes instructions, and a processor 147 that executes the instructions from the memory 146 to perform the various operations that are performed by the server 145. Furthermore, the server 150 may include a memory 151 that includes instructions, and a processor 152 that executes the instructions from the memory 151 to perform the various operations that are performed by the server 150. In certain embodiments, the servers 140, 145, 150, and 160 may be network servers, routers, gateways, switches, media distribution hubs, signal transfer points, service control points, service switching points, firewalls, routers, edge devices, nodes, computers, mobile devices, or any other suitable computing device, or any combination thereof. In certain embodiments, the servers 140, 145, 150 may be communicatively linked to the communications network 135, any network, any device in the system 100, or any combination thereof.

[0076]

[0064] The database 155 of the system 100 may be utilized to store and relay information that traverses the sy stem 100, cache content that traverses the system 100, store data about each of the devices in the system 100 and perform any other typical functions of a database. In certain embodiments, the database 155 may be connected to or reside within the communications network 135, any other network, or a combination thereof. In certain embodiments, the database 155 may sene as a central repository for any information associated with any of the devices and information associated with the system 100. Furthermore, the database 155 may include a processor and memory or may be connected to a processor and memory to perform the various operation associated with the database 155. In certain embodiments, the database Attorney Docket No. 197585-010941 PCT

[0077] 155 may be connected to the servers 140, 145, 150, 160, the first user device 102, the second user device 111 , the additional user devices, any devices in the system 100. any process of the system 100, any program of the system 100, any other device, any network, or any combination thereof.

[0078]

[0065] The database 155 may also store information and metadata obtained from the system 100, store metadata and other information associated with the first and second users 101, 110, store artificial intelligence models utilized in the system 100. store sensor data and / or content obtained from a patient, store predictions made by the system 100 and / or artificial intelligence models, store predictions of outcomes based on ctDNA profiles, storing confidence scores relating to predictions made, store threshold values for confidence scores, responses outputted and / or facilitated by the system 100. store information associated with anything determined or detected via the system 100, store information and / or content utilized to train the artificial intelligence models, store information associated with behaviors and / or actions conducted by individuals, store user profiles associated with the first and second users 101, 110, store device profiles associated with any device in the system 100, store communications traversing the system 100, store user preferences, store information associated with any device or signal in the system 100, store information relating to patterns of usage relating to the user devices 102, 111, store any information obtained from any of the networks in the system 100, store historical data associated with the first and second users 101, 110, store device characteristics, store information relating to any devices associated with the first and second users 101, 110. store information associated with the communications network 135, store any information generated and / or processed by the system 100, store any of the information disclosed for any of the operations and functions disclosed for the system 100 herewith, store any information traversing the system 100, or any combination thereof. In certain embodiments, the database 155 may be configured to store information associated with the patient’s health status, digital records, lab results, information associated with medicine and / or pharmaceutical drugs to be administered to the patient, information relating to patient visits and medical conditions, information associated with medical complaints made by a patient or determined by the system 100, information associated with recommendations for treatments to be done for the patient, information identifying the patient and / or physician, any other information of the system 100, or a combination thereof. Furthermore, the database 155 may be configured to process queries sent to it by any device in the system 100.

[0079]

[0066] In certain embodiments, the system 100 may incorporate the use of any number of machine learning and / or artificial intelligence models that can include may comprise software, Attorney Docket No. 197585-010941 PCT hardware, or a combination thereof. In certain embodiments, the system 100 may include one or more machine learning models supporting the functionality of the system 100. In certain embodiments, an machine learning and / or artificial intelligence model can be a file, program, module, and / or process that may be trained by the system 100 (or other system) to recognize certain patterns, conducting predictive modeling of outcomes based on ctDNA profiles, performing predictions as described in the present disclosure, or a combination thereof. For example, the machine learning model(s) may be trained to perform any of the functionality illustrates and / or described in Figures 2, 3, 4, 5, 6,7, 8, and 9. In certain embodiments, the machine learning and / or artificial intelligence model may be, may include, and / or may utilize a Deep Convolutional Neural Network, a one-dimensional convolutional neural network, a two-dimensional convolutional neural network, a Long Short-Term Memory network, any type of machine learning system, any type of artificial intelligence system, or a combination thereof. Additionally, in certain embodiments, the artificial intelligence model can incorporate the use of any t pe of artificial intelligence and / or machine learning algorithms to facilitate the operation of the artificial intelligence model(s).

[0080]

[0067] The system 100 may train the artificial intelligence model(s) to reason and leam from data fed into the system 100 so that the model(s) may generate and / or facilitate the generation of predictions about new data and information that is fed into the system 100 for analysis. For example, the system 100 may train a machine learning and / or artificial intelligence model using various types of data, information, and / or content, such as, but not limited to, images, video content, audio content, text content, augmented reality content, virtual reality content, information relating to patterns, information relating to behaviors, information relating to characteristics of users, information relating to environments, sensor data, pharmaceutical drug information (e.g., expected responses to drugs, etc.), any data associated with the foregoing, any type of data, or a combination thereof. As additional data and / or content is fed into the model(s) over time, the model's ability to predict patient outcomes can improve and be more finely tuned.

[0081]

[0068] In certain embodiments, the artificial intelligence models supporting the functionality of the system 100 can be trained continuously, at periodic intervals, or at the option of a controller of the artificial intelligence models. In certain embodiments, the artificial intelligence models may be trained on the accuracy of the predictions made by the artificial intelligence models, any information generated and / or used by the system 100, or a combination thereof. In certain embodiments, the artificial intelligence models may be trained with any type of content associated with medical complaints, assessments, and the like. Attorney Docket No. 197585-010941 PCT

[0082] Operatively, the system 100 may operate and / or execute the functionality as described and illustrated in Figures as otherwise described herein.

[0083]

[0069] Notably, as shown in Figure 1, the system 100 may perform any of the operative functions disclosed herein by utilizing the processing capabilities of server 160, the storage capacity of the database 155, or any other component of the system 100 to perform the operative functions disclosed herein. The server 160 may include one or more processors 162 that may be configured to process any of the various functions of the system 100. The processors 162 may be software, hardware, or a combination of hardware and software. Additionally, the server 160 may also include a memory' 161, which stores instructions that the processors 162 may execute to perform various operations of the system 100. For example, the sen' er 160 may assist in processing loads handled by the various devices in the system 100, such as, but not limited to, and performing any other operations conducted in the system 100 or otherwise. In one embodiment, multiple servers 160 may be utilized to process the functions of the system 100. The server 160 and other devices in the system 100, may utilize the database 155 for storing data about the devices in the system 100 or any other information that is associated with the system 100. In one embodiment, multiple databases 155 may be utilized to store data in the system 100.

[0084]

[0070] Referring now also to Figure 2, Figure 2 is a schematic of a training pipeline 200, from cohort assembly to iterative SHAP-guided reduction according to embodiments of the present disclosure. In certain embodiments, the training pipeline can utilize a binary mutation matrix 202 to represent mutations (e.g. changes in genetic sequences) in a binary form (e g., Is and 0s) rather than full nucleotide or amino acid codes. The training pipeline 200 can include conducting reactome pathway mutation enrichment 204 and then conducting random groupings 206 for various gene pathways. In certain embodiments, the training pipeline 200 can then include providing the random groupings to one or more neural networks 208 for processing. For example, the neural networks 208 can conduct feature ranking using SHAP-guided reduction for model n and feature f and the most important features to utilize can be predicted and / or determined. The training pipeline 200 can then include conducting re-randomization 210. Then the final predictive models 212 can be generated. Model deconstruction can be performed and potential biomarkers 214, drug mechanisms 216, and response prediction algorithms 218 provided and / or utilized.

[0085]

[0071] Referring now also to Figure 3. Figure 3 shows generation of pathway and PPI network features from variant calls according to embodiments of the present disclosure. For example, Figure 3 illustrates and example using biological groupings to reduce sparsity of mutations. Attorney Docket No. 197585-010941 PCT

[0086] Additionally, Figure 3 illustrates graphs depicting an increase in the number of non-zero values after conducting feature engineering.

[0087]

[0072] Referring now also to Figure 4, Figure 4 illustrates a method 400 featuring ctDNA harmonization, such as to generate a standardized ctDNA dataset, according to embodiments of the present disclosure. There can be a plurality of hospitals, such as Hospital A 402 with a Platform 1. Hospital B 404 with a Platform 2. and Lab C 406 with a Platform 3. At 408, raw data, such as raw ctDNA data, can be obtained from a variety of data sources, which can include, but is not limited to. Hospital A, Hospital B and Lab C. At 410, a data quality check 410 can be conducted by the system 100 to remove inaccurate data, faulty data, fraudulent data, duplicative data, and / or other data that can reduce the quality of the data. In certain embodiments, a threshold amount and / or type of data may need to have a quality level to be acceptable. If the quality check fails, the method 400 can proceed to 412 and flag the data for further review. At 414, the method 400 can include having an individual and / or another system conduct curation of the data to adjust the quality level to meet the required quality7level. At step 416, the method 400 can include normalizing coverage. At step 418, the method 400 can include conducting standardization of VAF reporting. At step 420, the method 400 can include harmonizing gene names 420, applying quality filters at step 422, and conducting batch effect correction at step 424. At step 426, the method 400 can include generating a standardized ctDNA dataset.

[0088]

[0073] FIG. 5 illustrates nested cross-validation, calibration, and stratified evaluation method 500 according to embodiments of the present disclosure. In certain embodiments, the method 500, at step 502, can include obtaining a dataset 502 (e.g., a standardized ctDNA dataset). At step 504, the method 500 can include selecting and / or providing a holdout test data set from the dataset 502. At step 506, the method 500 can include also selecting and / or providing development dataset 506 from the dataset 502. At step 508, the method 500 can include conducting nested cross-validation on the development dataset 506. At step 510, the method 500 can include selecting the best machine learning model for processing and / or handing the validated dataset (e.g., a model having a threshold accuracy level, speed level, capability level, etc.). At step 512. the method 500 can include, such as by utilizing the best / optimal machine learning model of the system 100, a feature reduction loop based on analyzing features in the dataset that underwent the cross-validation. At step 514, the method 500 can include retraining the model, and, at step 516, the method 500 can include assessing feature importance of the features. At step 518, the method 500 can include removing features identified as low- importance / relevance. At step 520, the method 500 can, at step 520, evaluate the performance Attorney Docket No. 197585-010941 PCT of the machine learning model and, if performance is not acceptable, the method 500 can proceed back to step 512 until the performance is acceptable (e.g., athreshold level of accuracy, speed, etc.). If the performance is acceptable, the method 500 can proceed to step 522 and can including stopping and utilizing a previous feature set. At step 524, the method 500 can include holding out evaluation, and, at step 526, the method 500 can include providing clinical performance metrics.

[0089]

[0074] FIG. 6 is an example interpretability report 600 according to embodiments of the present disclosure. In certain embodiments, the interpretability report 600 can include, but is not limited to, including a patient identifier (“ID”) 602, which can also include a risk score that can be a value between 0 to 1. The report 600 can also include clinical features 604, mutation information (e.g., burden high, medium, or low), biomarkers 604 (e.g.. biomarkers, such as TMB and MIS), and pathway dysregulation (e.g., PI3K / AKT: active or inactive; P53: inactive or active). In certain embodiments, the report 600 can include a SHAP analysis 610, KRAS values, biological pathways, ctDNA levels, and age information. In certain embodiments, the report 600 can include treatment recommendations 620 to treat a particular health condition, recommended trials (e.g., NCT123456 PI3K inhibitor; NC789012 combination therapy), and indications of what toe avoid 624 (e.g., avoid EGFR-targeted therapy.

[0090]

[0075] FIG. 7 shows clinical deployment architecture 700 for inference and decision support according to embodiments of the present disclosure. In certain embodiments, the architecture 700 can include a clinical interface layer 702 that can include a digital dashboard 704, a software patient portal 706, electronic health records (“EHR”) integration 708. The architecture 700 can also include an API Gateway 710, which can include authentication 712, rate limiting 714, and request routing 716. The information from the clinical interface layer 702 can be provided to the authentication 712 software of the API gateway 710 for authentication. The architecture 700 can further include a machine learning processing layer 720, which can include a model inference sen-ice 722, a SHAP explainer 724, a risk calculator 726, and atrial matcher 728 (e.g., to match patients to clinical trials). The architecture can also include a data processing layer 730 that can include a ctDNA parser 732, a feature engineering component 736, a data harmonization component 734, and a quality control 738 component. In certain embodiments, the architecture 700 can include a storage layer 740, which can include a patient database 742, a model registry 744 for storing machine learning models, a clinical trials database 746, and audit logs 748.

[0091]

[0076] FIG. 8 provides confusion-matrix and calibration plots 820, 804 comparing a ctDNA- only model to the disclosed model (prophetic) according to embodiments of the present Attorney Docket No. 197585-010941 PCT disclosure. FIG. 9 is a flowchart illustrating a method 900 for trial-matching using predicted outcomes and biomarker explanations according to embodiments of the present disclosure. At 902, the method 900 can include providing ctDNA mutation data to a machine learning response model 904 for processing and generating predictions. The machine learning model 904 outputs can be utilized for population response stratification 906, conducting populationbased trial cohorts 908, generating individual response scores 910, individual treatment selections 912. SHAP analyses 914, identification of key biomarkers 916, determination of biomarker-drive therapies 918, and then utilizing some or all of the foregoing to generate ranked treatment recommendations 920. At 922, the method 900 can include generating clinical decision support output.

[0092]

[0077]

[0093] EXAMPLES

[0094]

[0078] Referring now also to Figure 10, at least a portion of the methodologies and techniques described with respect to the exemplary embodiments of the system 100 can incorporate a machine, such as. but not limited to. computer system 1000, or other computing device within which a set of instructions, when executed, may cause the machine to perform any one or more of the methodologies or functions discussed above. The machine may be configured to facilitate various operations conducted by the system 100 and / or other components shown in the Figures and present disclosure. For example, the machine may be configured to, but is not limited to, assist the system 100 by providing processing power to assist with processing loads experienced in the system 100, by providing storage capacity for storing instructions or data traversing the system 100, or by assisting with any other operations conducted by or within the system 100. As another example, the computer system 1000 may assist with generating models associated with generating predictions of outcomes based on ctDNA profiles, any type of predictions generated by the system 100, or a combination thereof. As another example, the computer system 1000 may assist with binary mutation matrices, reactome pathway mutation enrichment, random grouping, neural netw ork functionality, featuring ranking, determining important features, conducting re-randomization, conducting model deconstruction, identifying potential markers, identifying drug mechanisms, utilizing response prediction algorithms, performing any other functionality provided by the system 100, or a combination thereof.

[0095]

[0079] In certain embodiments, the machine may operate as a standalone device. In some embodiments, the machine may be connected (e.g., using communications network 135, another network, or a combination thereof) to and assist with operations performed by other machines and systems, such as, but not limited to, the first user device 102, the second user device Attorney Docket No. 197585-010941 PCT

[0096] Ill, the server 140, the server 145, the server 150, the database 155, the server 160, any other system, program, and / or device, or any combination thereof. The machine may be connected with any component in the system 100. In a networked deployment, the machine may operate in the capacity of a server or a client user machine in a server-client user network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine may comprise a server computer, a client user computer, a personal computer (PC), a tablet PC, a laptop computer, a desktop computer, a control system, a network router, switch or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by that machine. Further, while a single machine is illustrated, the term “machine” shall also be taken to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein.

[0097]

[0080] The computer system 1000 may include a processor 1002 (e.g., a central processing unit (CPU), a graphics processing unit (GPU, or both), a main memory 1004 and a static memory 1006, which communicate with each other via a bus 1008. The computer system 1000 may further include a video display unit 610, which may be. but is not limited to, a liquid crystal display (LCD), a flat panel, a solid-state display, or a cathode ray tube (CRT). The computer system 1000 may include an input device 1012, such as, but not limited to, a keyboard, a cursor control device 1014, such as, but not limited to, a mouse. a disk drive unit 1016, a signal generation device 1018. such as, but not limited to. a speaker or remote control, and a network interface device 1020.

[0098]

[0081] In certain embodiments, the disk drive unit 1016 may include a machine-readable medium 622 on which is stored one or more sets of instructions 1024, such as, but not limited to, software embodying any one or more of the methodologies or functions described herein, including those methods illustrated above. The instructions 1024 may also reside, completely or at least partially, within the main memory 1004, the static memory 1006, or within the processor 1002, or a combination thereof, during execution thereof by the computer system 1000. In certain embodiments, the main memory 1004 and the processor 1002 also may constitute machine-readable media.

[0099]

[0082] Dedicated hardware implementations including, but not limited to, application specific integrated circuits, programmable logic arrays and other hardware devices can likewise be constructed to implement the methods described herein. Applications that may include the apparatus and systems of various embodiments broadly include a variety of electronic and computer systems. Some embodiments implement functions in two or more specific interconnected Attorney Docket No. 197585-010941 PCT hardware modules or devices with related control and data signals communicated between and through the modules, or as portions of an application-specific integrated circuit. Thus, the example system is applicable to software, firmware, and hardware implementations.

[0100]

[0083] In accordance with various embodiments of the present disclosure, the methods described herein are intended for operation as software programs running on a computer processor. Furthermore, software implementations can include, but not limited to, distributed processing or component / object distributed processing, parallel processing, or virtual machine processing can also be constructed to implement the methods described herein.

[0101]

[0084] The present disclosure contemplates a machine-readable medium 1022 containing instructions 624 so that a device connected to the communications network 135, another network, or a combination thereof, can send or receive voice, video or data, and communicate over the communications network 135, another network, or a combination thereof, using the instructions. The instructions 1024 may further be transmitted or received over the communications network 135, another network, or a combination thereof, via the network interface device 1020.

[0102]

[0085] While the machine-readable medium 1022 is shown in an example embodiment to be a single medium, the term "machine-readable medium" should be taken to include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) that store the one or more sets of instructions. The term "machine-readable medium" shall also be taken to include any medium that is capable of storing, encoding or carrying a set of instructions for execution by the machine and that causes the machine to perform any one or more of the methodologies of the present disclosure.

[0103]

[0086] The terms "machine-readable medium," "machine-readable device," or "computer- readable device" shall accordingly be taken to include, but not be limited to: memory devices, solid-state memories such as a memory card or other package that houses one or more readonly (non-volatile) memories, random access memories, or other re-writable (volatile) memories; magneto-optical or optical medium such as a disk or tape; or other self-contained information archive or set of archives is considered a distribution medium equivalent to a tangible storage medium. In certain embodiments, the "machine-readable medium," "machine-readable device." or "computer-readable device" may be non-transitory. and, in certain embodiments, may not include a wave or signal per se. Accordingly, the disclosure is considered to include any one or more of a machine-readable medium or a distribution medium, as listed herein and including art-recognized equivalents and successor media, in which the software implementations herein are stored.

[0104]

[0087] The illustrations of arrangements described herein are intended to provide a general Attorney Docket No. 197585-010941 PCT understanding of the structure of various embodiments, and they are not intended to serve as a complete description of all the elements and features of apparatus and systems that might make use of the structures described herein. Other arrangements may be utilized and derived therefrom, such that structural and logical substitutions and changes may be made without departing from the scope of this disclosure. Figures are also merely representational and may not be drawn to scale. Certain proportions thereof may be exaggerated, while others may be minimized. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense.

[0105]

[0088] Thus, although specific arrangements have been illustrated and described herein, it should be appreciated that any arrangement calculated to achieve the same purpose may be substituted for the specific arrangement shown. This disclosure is intended to cover any and all adaptations or variations of various embodiments and arrangements of the invention. Combinations of the above arrangements, and other arrangements not specifically described herein, will be apparent to those of skill in the art upon reviewing the above description. Therefore, it is intended that the disclosure is not limited to the particular arrangement(s) disclosed as the best mode contemplated for carrying out this invention, but that the invention will include all embodiments and arrangements falling within the scope of the appended claims.

[0106]

[0089] The foregoing is provided for purposes of illustrating, explaining, and describing embodiments of this invention. Modifications and adaptations to these embodiments will be apparent to those skilled in the art and may be made without departing from the scope or spirit of this invention. Upon reviewing the aforementioned embodiments, it would be evident to an artisan with ordinary’ skill in the art that said embodiments can be modified, reduced, or enhanced without departing from the scope and spirit of the claims described below.

[0107]

[0090] While various embodiments have been described for purposes of this disclosure, such embodiments should not be deemed to limit the teaching of this disclosure to those embodiments. Various changes and modifications may be made to the elements and operations described above to obtain a result that remains within the scope of the systems and processes described in this disclosure.

Claims

Attorney Docket No. 197585-010941 PCTClaims1. A computer-implemented method of training a predictive model of patient outcomes, comprising:(a) obtaining a training cohort that includes both ctDNA-positive and ctDNA-negative patients;(b) receiving, for each patient, clinical data and matched genomic features comprising tumor alterations detected by liquid biopsy and / or tissue sequencing;(c) transforming at least a portion of the matched genomic features into pathway features and / or protein-protein interaction (PPI) network features;(d) harmonizing ctDNA measurements by encoding ctDNA status and at least one normalized continuous ctDNA measure;(e) training one or more machine-learning models using the clinical data, the harmonized ctDNA measurements, the matched genomic features, and the pathway and / or PPI network features;(I) computing feature importances using Shapley-based explanations and iteratively reducing a feature set by retaining a top quantile of features across training rounds; and(g) outputting a trained model and an associated compact feature set usable to predict a clinical outcome.

2. The method of claim 1 , wherein the outcome comprises at least one of obj ective response, disease control, progression-free survival, overall survival, or a time-to-event endpoint.

3. The method of claim 1, wherein harmonizing ctDNA measurements comprises at least one of within-assay z-scahng, learning an adaptive threshold for ctDNA positivity, or jointly modeling tumor fraction and variant-allele-frequency burden.

4. The method of claim 1, wherein the pathway features are computed using one or more curated gene-set libraries selected from Reactome, KEGG, and Gene Ontology’.

5. The method of claim 1. wherein the PPI network features comprise at least one of network-propagation scores, random-walk diffusion scores, centrality -weighted mutation scores, or aggregated node embeddings.

6. The method of claim 1, wherein iteratively reducing the feature set comprises retaining between 10% and 30% of features by mean Shapley importance across a plurality of trained models.Attorney Docket No. 197585-010941 PCT7. The method of claim 1, wherein the compact feature set contains fewer than 10% as many features as patients in the training cohort.

8. The method of claim 1, further comprising nested cross-validation and probability calibration using isotonic regression or Platt scaling.

9. The method of claim 1, wherein model families are randomly sampled from a collection comprising gradient-boosted trees, random forests, generalized linear models, Cox proportional hazards models, neural networks, and stacked ensembles.

10. The method of claim 10, wherein generating the explanation comprises outputting at least one of global feature importance, per-patient SHAP waterfall plots, pathway attribution tables, or counterfactual feature analyses.

11. The method of claim 10. further comprising routing the outcome prediction to a clinical decision support rule that recommends at least one action selected from continuing current therapy, switching therapy, ordering reflex testing, or screening the patient for a clinical trial.

12. The method of claim 1, wherein matched genomic features comprise one or more of single-nucleotide variants, insertions / deletions. copy -number alterations, gene fusions, methylation markers, or fragmentomic metrics.

13. The method of claim 1, further comprising controlling selection bias using at least one of stratified sampling, inverse-probability weighting, or site / assay batch-effect correction.

14. The method of claim 1, wherein the outcome prediction is evaluated and reported separately for ctDNA-positive and ctDNA-negative strata to verify calibration.

15. The method of claim 1, wherein the trained model achieves at least one of: an area under the receiver-operating characteristic curve of 0.85 or greater for a binary outcome on a validation set; or a concordance index of 0.70 or greater for a survival outcome on a validation set.

16. The method of claim 1, further comprising enforcing pathway-level sparsity by applying group-lasso or structured regularization.

17. The method of claim 10. wherein the matched genomic features are obtained from a liquid-biopsy panel and optionally reconciled with a tissue panel using graph-based imputation of missing genes.

18. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the processors to perform the method of claim 1.Attorney Docket No. 197585-010941 PCT19. A system for training a predictive model of patient outcomes, comprising: one or more processors; and a non-transitory memory storing instructions that, when executed by the one or more processors, cause the system to:(a) obtain a training cohort that includes both ctDNA-positive and ctDNA- negative patients;(b) receive, for each patient, clinical data and matched genomic features comprising tumor alterations detected by liquid biopsy and / or tissue sequencing;(c) generate engineered features by transforming at least a portion of the matched genomic features into pathway features and / or PPI network features;(d) harmonize ctDNA measurements by encoding ctDNA status and at least one normalized continuous ctDNA measure;(e) train one or more machine-learning models using the clinical data, the harmonized ctDNA measurements, the matched genomic features, and the engineered features;(!) compute feature importances using Shapley-based explanations and iteratively reduce a feature set by retaining a top quantile of features across training rounds; and(g) output a trained model and an associated compact feature set usable to predict a clinical outcome.

20. The system of claim 19, wherein training employs nested cross-validation and probability calibration using isotonic regression or Platt scaling.

21. The system of claim 19, further configured to mitigate selection bias using at least one of stratified sampling to preserve ctDNA-positive / negative prevalence, inverseprobability or propensity weighting to account for ctDNA testing patterns, or batch-effect correction across assay panels.

22. The system of claim 19, wherein pathway-level sparsity is enforced using group-lasso or structured regularization.

23. The system of claim 19, further comprising a validation module configured to report discrimination and calibration separately for ctDNA-positive and ctDNA-negative strata, and an audit module configured to store model versioning, datasets used for training, selected features, and calibration parameters.

Citation Information

Patent Citations

  • Pathway recognition algorithm using data integration on genomic models (paradigm)

    US20190341123A1

  • Novel Markers for Detecting Microsatellite Instability in Cancer and Determining Synthetic Lethality with Inhibition of the DNA Base Excision Repair Pathway

    US20200255913A1

  • Generating machine learning models using genetic data

    US20230222311A1

  • Techniques for designing patient-specific panels and methods of use thereof for detecting minimal residual disease

    WO2024129844A1