Methods and systems for stratifying patient cancer risk using computational oncology and molecular data

By employing machine learning models trained with molecular data and incorporating advanced data filtering and correction techniques, the method effectively addresses the limitations of conventional cancer risk stratification, enabling personalized and accurate treatment strategies.

JP2025081280APending Publication Date: 2025-05-27TEMPUS AI INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024199305
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-15
Filing Date
2024-11-14
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Conventional cancer diagnosis and treatment models struggle with accurately stratifying cancer risk due to the requirement for homogeneous input data, leading to inaccurate predictions and challenges in deciding between invasive and less invasive treatment options for patients with similar clinical profiles.

Method used

A computer-implemented method using machine learning models trained with molecular data to determine a patient's cancer risk profile, incorporating univariate gene selection, RNA bias correction, and multivariate gene selection to filter and correct training data, and including a survival model to generate a matched treatment strategy based on the patient's molecular data risk.

Benefits of technology

This approach enables more accurate and nuanced cancer risk stratification, allowing for personalized treatment strategies that can avoid overtreatment or undertreatment by leveraging molecular data to refine clinical risk assessments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025081280000001_ABST
    Figure 2025081280000001_ABST
Patent Text Reader

Abstract

To provide an implemented method, computing system, and computer-readable medium for stratifying a patient cancer risk using molecular data.SOLUTION: An implemented method includes: receiving molecular data; processing the molecular data using a machine learning model; and generating a matched treatment strategy for a patient on the basis of the patient's molecular data risk. A computer-implemented method, computing system, and computer-readable medium for training a machine learning model to stratify a patient cancer risk using molecular data include: receiving a patient training dataset and a reference training dataset; selecting a cohort of the patient; selecting a small subset of genes using univariate selection; generating a corrected reference training dataset; selecting a smaller subset of genes using multivariate selection; training a survival model; and (g) selecting a decision threshold to identify a patient population.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims priority to U.S. Patent Application No. 63 / 599,471, filed on November 15, 2023, entitled "METHODS AND SYSTEMS FOR STRATIFYING PATIENT CANCER RISK USING COMPUTATIONAL ONCOLOGY AND MOLECULAR DATA", which is hereby incorporated by reference in its entirety.

[0002] The present disclosure is directed to methods and systems for stratifying a patient's cancer risk using computational oncology and molecular data, and more particularly, to techniques for training and operating one or more machine - learning models to process a patient's molecular data to predict the prognosis of the patient's cancer risk profile.

Background Art

[0003] The description of the background art provided herein is for the purpose of generally presenting the context of the present disclosure. The research of the inventors named at present, to the extent described in this background art section, and aspects of the description that may not be considered prior art at the time of filing in another way are not admitted to be prior art to the present disclosure, either expressly or implicitly.

[0004] Conventional cancer diagnosis and treatment models (e.g., simple classifiers) have several drawbacks. For example, such models require strictly homogeneous input data (e.g., censored patient data often cannot be used for training), and any deviation in the availability of training data can make conventional modeling inaccurate.

[0005] Regarding the decision-making for at least one type of cancer and the corresponding treatment, clinicians are currently struggling to understand a clinically homogeneous group of patients who do not present sufficient symptoms or diseases to enable an appropriate diagnosis of the patient's risk profile based on clinical criteria. Specifically, for a particular patient, it is not always clear whether the clinician should select a more invasive treatment for the patient (e.g., the patient has a worse prognosis and a higher risk) or a less invasive treatment for the patient (e.g., the patient has a better prognosis and a lower risk).

[0006] As an illustrative example of cancer type and the corresponding treatment decision-making, the decision-making for adjuvant therapy of endometrioid endometrial cancer currently relies on risk stratification using clinical features such as histology, grade, stage, and lymphovascular space invasion (LVSI). Recently, a molecular classification system derived from TCGA evaluated in GOG-210 and PORTEC-3 defined four prognostic subtypes based on POLE, MSI-H / MMR-D, and p53 alterations. This molecular approach, although valuable, still has significant limitations, such as its applicability to the majority of EEC patients classified as having a non-specific molecular profile (NSMP), and the potential need to address pathogenic and prognostic heterogeneity within the MMR-D and TP53 subtypes.

[0007] The high- to intermediate-risk group may include a significant percentage (e.g., 9%) of patients who have distant recurrence of cancer and would benefit from early treatment. Some of the patients in this group may potentially benefit from early intervention treatment, but considering the aforementioned homogeneity, the decision of whether to treat these patients early is opaque.

[0008] Many of these patients undergo resection and receive adjuvant therapy while recovering at home. However, many early-stage cases without metastasis appear essentially the same from a clinical perspective, and clinicians do not know whether to escalate, de-escalate, or maintain the patient's treatment strategy. This can lead directly to overtreatment or undertreatment when the patient is actually at low risk (a certain percentage of these cases will ultimately develop distant recurrence). Furthermore, generally, conventional techniques do not properly systematize cohort selection.

[0009] Therefore, there is an opportunity for improved platforms and technologies to stratify a patient's cancer risk using computational oncology and molecular data by enhancing data availability, systematizing cohort selection criteria, and actively determining a patient's prognosis. SUMMARY OF THE INVENTION

[0010] In one aspect, a computer-implemented method for stratifying a patient's cancer risk using molecular data includes: (a) receiving, via one or more processors, molecular data corresponding to a patient; and (b) processing the molecular data using a machine learning model via one or more processors to determine the patient's molecular data risk, wherein the machine learning model is trained using a patient training data set and a reference training data set, the machine learning model uses univariate gene selection, RNA bias correction, and multivariate gene selection to filter and correct its training data, and the machine learning model includes a survival model; and (c) generating a matched treatment strategy corresponding to the patient based on the patient's molecular data risk.

[0011] In another aspect, a computer-implemented method for training a machine learning model to stratify a patient's cancer risk using molecular data includes: (a) receiving, via one or more processors, (i) a patient training data set including molecular data for each of a plurality of patients and (ii) a reference training data set including molecular data for each of the plurality of patients; (b) selecting, via one or more processors, a patient cohort from the patient training data set; (c) selecting, via one or more processors, a small subset of genes from the patient training data set using univariate gene selection; (d) generating, by processing the reference training data set to correct biases in the molecular data of the plurality of patients, a corrected reference training data set; (e) selecting, via one or more processors, an even smaller subset of genes from the small subset of genes using multivariate gene selection; (f) training, via one or more processors, a survival model, the training including determining a set of hyperparameters; and (g) selecting, via one or more processors, a decision threshold for identifying a population of patients having an RNA risk profile.

[0012] In yet another aspect, a computing system includes one or more processors and one or more memories storing computer-executable instructions that, when executed by the one or more processors, cause the computing system to: (a) receive molecular data corresponding to a patient; (b) process the molecular data using a machine learning model to determine a molecular data risk for the patient, the machine learning model being trained using a patient training data set and a reference training data set, the machine learning model using univariate gene selection, RNA bias correction, and multivariate gene selection to filter and correct its training data, the machine learning model including a survival model; and (c) generate a matched treatment strategy corresponding to the patient based on the molecular data risk for the patient.

[0013] In yet another aspect, a computing system includes one or more processors and one or more memories storing computer-executable instructions, which when executed by the one or more processors cause the computing system to: (a) receive (i) a patient training data set including molecular data of each of a plurality of patients and (ii) a reference training data set including molecular data of each of the plurality of patients; (b) select a patient cohort from the patient training data set; (c) select a small subset of genes from the patient training data set using univariate gene selection; (d) process the reference training data set to generate a corrected reference training data set by correcting biases in the molecular data of the plurality of patients; (e) select a smaller subset of genes from the small subset of genes using multivariate gene selection; (f) train a survival model, the training including determining a set of hyperparameters; and (g) select a decision threshold for identifying a population of patients having an RNA risk profile.

[0014] In another aspect, a computer-readable medium includes computer-executable instructions that, when executed by one or more processors, cause a computer to: (a) receive molecular data corresponding to a patient; (b) process the molecular data using a machine learning model to determine a molecular data risk for the patient, the machine learning model being trained using a patient training data set and a reference training data set, the machine learning model using univariate gene selection, RNA bias correction, and multivariate gene selection to filter and correct its training data, and the machine learning model including a survival model; and (c) generate a matched treatment strategy corresponding to the patient based on the molecular data risk of the patient.

[0015] In yet another aspect, a computer-readable medium includes computer-executable instructions that, when executed by one or more processors, cause the computer to: (a) (i) receive a patient training data set that includes the molecular data of each of a plurality of patients, and (ii) receive a reference training data set that includes the molecular data of each of the plurality of patients; (b) select a patient cohort from the patient training data set; (c) use univariate gene selection to select a small subset of genes from the patient training data set; (d) process the reference training data set to correct the bias of the molecular data of the plurality of patients, thereby generating a corrected reference training data set; (e) use multivariate gene selection to select a smaller subset of genes from the small subset of genes; (f) train a survival model, where training includes determining a set of hyperparameters; and (g) select a decision threshold for identifying a population of patients having an RNA risk profile.

Brief Description of the Drawings

[0016] The patent or application file contains at least one drawing created in color. Copies of this patent or patent application publication, including the color drawing, will be provided by the Patent Office upon request and payment of the necessary fees.

[0017] The drawings described below depict various aspects of the systems and methods disclosed herein. It should be understood that each drawing depicts an example of an aspect of the system and method.

[0018]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5A

Figure 5B

Figure 5C

Figure 6A

Figure 6B

Figure 6C

Figure 6D

Figure 6E

Figure 6F

Figure 7

DETAILED DESCRIPTION

[0019] Overview The present technique is directed to methods and systems for stratifying a patient's cancer risk using molecular data, and more specifically, techniques for training and operating one or more machine learning models to process a patient's molecular data to prognostically predict the patient's cancer risk profile.

[0020] Conventional diagnosis and treatment of cancer patients rely on classical histopathological and immunohistochemical techniques. Molecular characterization of tumor samples provides additional insights into cancer biology beyond classical clinical factors such as stage, grade, and histology. The present technique can identify patient tumors with a negative prognosis from within patient groups known to be equivalent in terms of clinical factors regarding the risk of disease progression. The present technique can further generate recommendations regarding more (or less) invasive treatment for these patients.

[0021] In some examples, the predicted prognosis or risk profile of a patient, determined by processing molecular data using a machine learning model to check whether the patient aligns with the inclusion or exclusion criteria of a particular trial, may be utilized. Alternatively, the predicted prognosis or risk profile of a large number of patients within a database may be employed to estimate the number of patients who may be suitable candidates for a particular trial. The machine learning model employed for these purposes may be trained using a patient training dataset and a reference training dataset. This model can employ univariate gene selection, RNA bias correction, and multivariate gene selection to filter and correct its training data. Additionally, the model may incorporate a survival model. Based on the molecular data risk of a patient, a matched treatment strategy may be generated, which may include, but is not limited to, systemic therapy, external beam radiation therapy, brachytherapy, or observation. This approach enables a more nuanced understanding of a patient's risk profile based on molecular data, which may be particularly useful for stratifying patients within clinical risk groups such as low clinical risk, low-to-intermediate clinical risk, high-to-intermediate clinical risk, or high clinical risk. For example, patients within the high-to-intermediate clinical risk group determined to have a high molecular risk may be matched with more invasive treatment strategies such as systemic therapy or external beam radiation therapy. Conversely, patients within the same clinical risk group but with a low molecular risk may instead be matched with observation, potentially avoiding unnecessary treatments that may have adverse effects. The methods and systems described herein leverage the power of machine learning and molecular data to make more informed decisions regarding a patient's treatment strategy and trial inclusion, providing significant advancements in the fields of patient care and clinical trial selection.

[0022] That is, the molecular profiling technique of the present invention may also (or alternatively) be used in the context of computational oncology to identify patients with a low risk of disease progression from within a group of patients having a similar clinical profile. For these patients identified by molecular data as having a good prognosis, treatment escalation may be recommended.

[0023] For example, data generated by assays such as the Target Oncology DNA Sequencing Panel (also referred to as Tempus xT Solid Tumor+Normal Match or xT) and / or the whole transcriptome RNA seq panel (also referred to as xR or Tempus xR) can be used to train and operate the prognostic models contemplated herein. In particular, RNA seq data from xR can be used to predict the risk of disease progression. Some related events that characterize the risk of disease progression include progression-free survival (PFS), disease-free survival (DFS), recurrence-free survival (RFS), and overall survival (OS).

[0024] Constructing a prognostic model on an RNA expression (xR) dataset is not straightforward due to technical challenges such as patient selection (for training), biomarker identification, survival modeling, and selection of the operating point (i.e., the decision threshold that separates patients with good prognosis from those with poor prognosis, and the threshold is applied to the output of the survival model). Furthermore, for a given cancer type and the metric of interest, it may not be possible to train a prognostic model on in-house patient data or to properly validate such a trained model because a dataset that well represents the true patient population seen in a clinic may not be available (e.g., due to insufficient holdout data to form an independent validation dataset for the validation of such a prognostic model).

[0025] To address this problem, the technique can include machine learning techniques for constructing a facility-specific RNA-based prognostic model trained on an external publicly available RNA expression dataset with clinical outcomes. Since these models are trained on external datasets, the associated Tempus data can be used for held-out independent validation. Even when the associated clinical outcomes are not available in Tempus, or when the Tempus patient population does not represent the true clinical population, the machine learning techniques can be used to train an RNA-based prognostic model on an external RNA dataset, such that these models can be evaluated against samples from an xR assay at the time of evaluation. In particular, when training early-stage cancer progression models, externally available datasets tend to be more heterogeneous from the perspective of the patient population, resulting in their being a good representation of the patient population, which leads to the training of robust and reproducible RNA-based prognostic models.

[0026] The technique can be used to identify a matching patient population between a cohort (e.g., the The Cancer Genome Atlas (TCGA) data cohort) and a Tempus cohort. The technique can be used to perform patient selection based on molecular information, as opposed to general conventional methods of selecting a patient population according to clinical criteria. The limitation of the clinical patient selection paradigm is that it does not account for the molecular diversity of the underlying selected population. Understanding and controlling molecular diversity is important for successfully creating molecular algorithms that are robust and can be validated across multiple cohorts.

[0027] An example of molecular cohort selection considered herein as a motivating example is the example of endometrioid endometrial cancer. Specifically, the machine learning technique can be used to train an RNA-based prognostic model for risk stratification of high- to intermediate-risk endometrioid endometrial cancer patients.

[0028] When examining molecular predictors of response to therapy, it was empirically observed in TCGA that a different set of genes than those found in the proprietary database was predictive. Using a greedy algorithm, the clinical criteria of the proprietary database funnel were programmatically updated, and ultimately, this technique identified molecular matching between the proprietary database and TCGA (mainly with the exclusion of uterine sarcomas). This enabled prognostic predictors to be trained on TCGA data and then validated on proprietary database data.

[0029] Specifically, this technique may include training of a prognostic gene signature in a cross-validation experimental setting in the endometrial cancer TCGA cohort. This may include modeling the log-rank p-value of 0.078, median survival difference of 12 months, and hazard ratio of 2.82 for stage I+II patients in the proprietary database. In this example, the proprietary database PFS data may not be used for model training. Instead, a larger TCGA cohort may be used for model training. The prognostic model may be validated in (i) the corrected proprietary database v1 cohort and (ii) the proprietary database v2 cohort, neither of which was used in training. The training cohort (TCGA) may be corrected to the proprietary database RNA v2, and thus, the prognostic model may be native to RNA v2. Since the prognostic model is native to RNA v2, risk prediction can be performed on future RNA v2 samples without modification.

[0030] The training described above is highly advantageous and improves the prognostic prediction method by enabling more datasets to be used. Furthermore, as described above, by training on an external data source, internally generated data can be used for validation. This is a particularly important result since training may require considerably more data than validation (e.g., the data for training is three times or more the data for validation). In short, this advancement has made algorithms that were previously infeasible to train feasible.

[0031] On the one hand, in conventional research, the training cohort of the institution's own cancer prognosis model had to be carefully curated. However, with this technique, the institution's prognosis model can be trained using publicly available data. This enables the effective use of the large amount of computing resources that would otherwise be required for the sequencing and curation of the training cohort. The institution's own database RNA dataset can be utilized for independent holdout validation. Furthermore, the training strategy leads to a robust and reproducible RNA-based prognosis model because this technique can (i) train with larger publicly available datasets and (ii) select biomarkers (e.g., genes) in a multi-stage cascade model setup.

[0032] Exemplary Computing Environment FIG. 1 illustrates an exemplary computing environment 100 for implementing this technique according to some aspects. Environment 100 includes computing resources for processing RNA data using machine learning to train a model to stratify a patient's cancer risk, and more specifically, to prognostically classify a patient's molecular risk as high or low, and / or for further processing / computation based on such classification, such as recommendations and / or reports regarding treatment strategies.

[0033] Computing environment 100 can include a molecular risk state prediction computing device 102, a client computing device 104, an electronic network 106, a sequencer system 108, and an electronic database 110. Molecular risk state prediction computing device 102 can include an application programming interface 112 that enables programmatic access to molecular risk state prediction computing device 102. The components of computing environment 100 can be communicatively connected to each other via electronic network 106 in some aspects. Each will be described in more detail herein.

[0034] The molecular risk state prediction computing device 102 can implement, among other things, the training and operation of a machine learning model for predicting the molecular cancer risk of one or more patients, patient identification, and report generation. In some aspects, the molecular risk state prediction computing device 102 can be implemented as one or more computing devices (e.g., one or more servers, one or more laptops, one or more mobile computing devices, one or more tablets, one or more wearable devices, one or more cloud computing virtual instances, etc.). The molecular risk state prediction computing device 102 can include one or more processors 120, one or more network interface controllers 122, one or more memories 124, an input device 126, and an output device 128.

[0035] In some aspects, the one or more processors 120 can include one or more central processing units, one or more graphics processing units, one or more field programmable gate arrays, one or more application specific integrated circuits, one or more tensor processing units, one or more digital signal processors, one or more neural processing units, one or more RISC-V processors, one or more coprocessors, one or more special processors / accelerators for artificial intelligence or machine learning specific applications, one or more microcontrollers, etc.

[0036] The molecular risk state prediction computing device 102 can include one or more network interface controllers 122, such as an Ethernet network interface controller, a wireless network interface controller, etc. The network interface controller 122 can include advanced functions, such as hardware acceleration, special networking protocols, etc., in some aspects.

[0037] The memory 124 of the molecular risk state prediction computing device 102 may include a volatile and / or non-volatile storage medium. For example, the memory 124 may include one or more random access memories, one or more read-only memories, one or more cache memories, one or more hard disk drives, one or more solid state drives, one or more non-volatile memory express, one or more optical drives, one or more universal serial bus flash drives, one or more external hard drives, one or more network-connected storage devices, one or more cloud storage instances, one or more tape drives, and the like.

[0038] The memory 124 may store one or more modules 130, for example, as one or more sets of computer-executable instructions. In some embodiments, the module 130 may include additional storage devices such as one or more operating systems (e.g., Microsoft Windows®, GNU / Linux®, Mac OSX®, etc.). The operating system may be configured to execute the module 130 during the operation of the molecular risk state prediction computing device 102. For example, the module 130 may include additional modules and / or services for receiving and processing quantitative data. The module 130 may be implemented using any suitable computer programming language (e.g., Python®, JavaScript®, C, C++, Rust, C#, Swift, Java®, Go, LISP, Ruby, Fortran, etc.). The memory may be non-transitory memory.

[0039] Module 130 may include a machine learning model training module 152 that includes a plurality of sub - modules. Specifically, the sub - modules may include a cohort selection module 154, a bias correction module 156, a clinical risk module 158, and a survival modeling module 160. In some aspects, more or fewer modules 130 may be included. Module 130 may be configured to communicate with each other (e.g., via inter - process communication, via a bus, message queue, socket, etc.).

[0040] The machine learning model training module 152 may include a set of computer - executable instructions for training one or more machine learning models based on training data. The machine learning model training module 152 may take input data, for example, in the form of a dataset, and use it to train a machine learning model. The machine learning model training module 152 may prepare the input data by performing data cleaning, feature engineering, data splitting (into training and validation sets), and handling missing or outlier values. The machine learning model training module 152 may select a machine learning algorithm or model architecture to use for the task at hand. Specifically, the machine learning model training module 152 may include a set of computer - executable instructions for implementing a machine learning training architecture such as the cascade model architecture 200 shown in FIG. 2. In some aspects, the machine learning model training module 152 may delegate training steps to one or more of the sub - modules 154 - 160 to perform one or more stages of the training process.

[0041] The machine learning model training module 152 may include instructions for performing hyperparameter tuning (e.g., settings or configurations of the model that need to be specified before training rather than learned from data). The machine learning model training module 152 may use grid search or other techniques to specify hyperparameters. For example, the machine learning model training module 152 may include instructions for selecting the following hyperparameters: RNA correction, univariate gene selection, penalty, L1 ratio, Hcoef, Lcoef, and alpha.

[0042] The machine learning model training module 152 may include instructions for training a selected machine learning model on training data. The training process may include optimizing the parameters of the model for making predictions. In some aspects, one or more free / open-source software libraries may be used to implement one or more training strategies. Examples of such libraries include Scikit-learn, Python®, and Lifelines. For example, in some aspects, the survival modeling module 160 may implement the Cox proportional hazards algorithm using CoxPHFitter from the Lifelines library (https: / / lifelines.readthedocs.io / en / latest / fitters / regression / CoxPHFitter.html). After training, the machine learning model training module 152 may evaluate the performance of the trained model using validation data. The validation technique may include cross-validation, as will be discussed in more detail below.

[0043] The machine learning model training module 152 may include instructions for serializing and deserializing the stored model. This enables the machine learning model training module 152 to store the trained model as data, reload the model without retraining, and use the trained model for predictions.

[0044] The cohort selection module 154 may include instructions for identifying one or more patients, as will be discussed in more detail below. Specifically, the cohort selection module 154 may include instructions for performing univariate gene selection on a patient training dataset, which may include data from public sources (e.g., TCGA) and / or proprietary sources.

[0045] The bias correction module 156 may include instructions for removing bias between datasets, as will be discussed in more detail below.

[0046] The clinical risk module 158 may include instructions for determining a patient's clinical risk, as discussed herein. Clinical risk generally relates to risks that can be quantified by a clinician by referring to available information without molecular analysis. The clinical risk module 158 may include instructions for determining clinical data from patient records (e.g., via public or proprietary datasets). The clinical risk module 158 may include instructions for retrieving clinical data (e.g., electronic health records) from the database 110, via the API 112, or via another source.

[0047] The survival modeling module 160 includes computer-executable instructions for generating a survival curve. The instructions may also generate data corresponding to the prediction of time to an event using an algorithm (e.g., Cox proportional hazards) that can use censored data, and these algorithms can simultaneously evaluate the impact of multiple variables on survival time. The instructions may include additional or different algorithms, such as the Kaplan-Meier estimator, which can be used to determine the impact of different covariates on survival probability and survival time.

[0048] The model operation module 170 may include computer-executable instructions for operating one or more trained machine learning models. For example, in some aspects, the model operation module 170 may access next-generation sequencing data via the sequencer 108 and / or the database 110. The model operation module 170 may load one or more models trained by the model training module 152. The model operation module 170 may receive raw next-generation sequencing data, preprocess it, apply one or more trained machine learning models, and generate one or more predictions corresponding to the patient's next-generation sequencing data (e.g., molecular risk prediction, matched therapy, etc.).

[0049] The model operation module 170 may receive next-generation sequencing data including DNA sequencing data (e.g., whole-genome sequencing, exome sequencing), RNA sequencing data (RNAseq), ChIP sequencing data (ChIP-seq), etc. The model operation module 170 may include instructions for receiving and processing data encoded in a plurality of different formats (FASTQ, BAM, VCF, etc.).

[0050] The model operation module 170 may perform quality control, read alignment, variant calling, and data normalization.

[0051] The model operation module 170 may perform feature engineering to convert raw sequencing data into features that can be used by one or more machine learning models trained by the machine learning training module 170.

[0052] The model operation module 170 can preprocess the received sequencing data by processing it using one or more trained machine models (e.g., deep learning models, neural networks, survival models, support vector machines, etc.). The model operation module 170 can generate model performance statistics such as accuracy, precision, recall, F1 score, or AUC-ROC for classification techniques.

[0053] In some aspects, the report generation module 158 can generate a report that includes predictions regarding the reliability of the classification of patient data. This data can be plotted on a heatmap, for example, as shown in FIG. 5A. Generally, such a heatmap can be a graphical representation of data where the individual values contained in a matrix are colored and the intensity is shown on an axis, representing higher expression values of gene covariants. The squares within the plot can correspond to the values from the matrix and can represent the magnitude of the values therein. Instructions can include generating digital visual artifacts to aid understanding. For example, in FIG. 5, red represents higher values (higher expression).

[0054] For example, the digital artifacts generated by the report generation module 158 can be presented to clinicians, patients, etc., and can take forms such as text documents, digital presentations, word processing documents, etc. For example, the report generation module 158 can generate one or more reports that include a patient's RNA risk score and / or one or more matching therapies. The report can include stratified clinical risk and / or stratified molecular (e.g., RNA) risk scores, as shown in FIGS. 4, 5B, and 5C. In some aspects, the report generation module 158 can include instructions for generating visual reports for comparison / validation, such as the graphs of FIGS. 6A - 6F.

[0055] The report generation module 158 may include computer-executable instructions for generating machine-readable results. For example, the client computing device 104 may be accessed by a user to view the results of what was generated by the prediction computing device 102. For example, the user may access a mobile device, a laptop device, a thin client, etc., embodied as the client computing device 104 to view simulation and reliability scoring results and / or reports regarding samples whose values were processed by the prediction computing device 102. Information from the prediction computing device 102 may be transmitted via the network 106 (e.g., for display via a viewer application 180).

[0056] The electronic network 106 may communicatively couple the elements of the environment 100. The network 106 may include a public network such as the Internet, a private network such as a research institution's or a company's private network, and / or any combination thereof. The network 106 may include a local area network (LAN), a wide area network (WAN), a cellular network, a satellite network, and / or other network infrastructure, whether wireless or wired.

[0057] In some embodiments, network 106 may be communicatively coupled to and / or be part of a cloud-based platform (e.g., cloud computing infrastructure). Network 106 may utilize communication protocols including packet-based and / or datagram-based protocols such as Internet Protocol, Transmission Control Protocol, User Datagram Protocol, and / or other types of protocols. Network 106 may include one or more devices that facilitate network communication and / or form a hardware infrastructure for the network, such as one or more switches, one or more routers, one or more gateways, one or more access points (such as wireless access points), one or more firewalls, one or more base stations, one or more repeaters, one or more backbone devices, and the like.

[0058] Sequencer system 108 may include a next-generation sequencer such as a Tempus xT RNA sequencing whole exome capture transcriptome assay.

[0059] Electronic database 110 may include one or more suitable electronic databases for storing and retrieving data, such as relational databases (e.g., MySQL® database, Oracle database, Microsoft SQL Server database, PostgreSQL database, etc.). Electronic database 110 may be a NoSQL database such as a key-value store, a graph database, a document store, etc. Electronic database 110 may be an object-oriented database, a hierarchical database, a spatial database, a time-series database, an in-memory database, etc. In some embodiments, some or all of electronic database 110 may be distributed.

[0060] During operation, one or more sequencer executions can be performed using sequencer 108 by the company that operates / controls environment 100 or by another party. The results of the sequencer can be received as sequencer data by the molecular risk state prediction computing device 102. The molecular risk state prediction computing device 102 can preprocess the sequencer data and optionally store some or all of it in the electronic database 110 of FIG. 1. The molecular risk state prediction computing device 102 can load one or more trained models, provide the sequencer data as input to the one or more trained models, and receive predictions regarding molecular risk from the one or more trained models. The molecular risk state prediction computing device 102 can provide the sequencer data (e.g., RNA Seq data) to a machine learning model trained as a molecular signature. The molecular risk state prediction computing device 102 can generate predictions regarding samples corresponding to RNA Seq. For example, the predictions can correspond to the molecular risk profile (e.g., RNA risk) of a patient. The molecular risk state prediction computing device 102 can communicate the results of the modeling to the patient and / or clinician. The molecular risk state prediction computing device 102 can access the patient's current clinical risk profile to determine whether the model-predicted RNA risk is the same. The molecular risk state prediction computing device 102 can generate one or more matches of treatment strategies to the patient's current molecular and / or clinical risk profile (e.g., as discussed below with respect to FIG. 4).

[0061] In some aspects, database 110 may include additional clinical and molecular patient data. For example, database 110 may include qualitative insights regarding a patient's history, symptoms, and clinical inferences underlying treatment decisions, providing a narrative context to the quantitative data also stored within the database. Database 110 may include lab results and test results, from basic blood tests, imaging, or analysis of images, to more complex genetic screening. Database 110 may store detailed diagnoses, as well as DNA and RNA sequencing data. In some aspects, database 110 may include methylation assay results, or other information related to epigenetic modifications or disease etiology. In some aspects, database 110 may include treatment response data or other follow-up information related to how a patient responds to various treatments. Database 110 may also incorporate results from the methods described in the application, such as risk profiles generated through the application's trained multi-stage machine learning architecture. By integrating these results, the database enhances the utility of molecular data in clinical decision-making, enabling healthcare providers to identify patients who may benefit from targeted therapies based on the patient's risk profile. This ability represents a significant advancement in the field of oncology, where risk profiles are an important factor in determining the most appropriate treatment for cancer patients.

[0062] By the time molecular risk state prediction computing device 102 uses the model thus trained, the trained model has already been trained using a training data set as described herein, and one or more submodels that are part of the model architecture (e.g., model architecture 200) are trained individually (e.g., using their own and / or public training data sets) to perform tasks such as univariate gene selection, RNA bias correction, multivariate gene selection, survival model training, and threshold selection.

[0063] The prediction of the molecular risk state prediction computing device 102 can be stored, for example, in the memory 124 or the database 110. These results can be provided directly to other elements of the environment 100 via, for example, the network 100. These results can also be further processed, for example, to identify / notify the patient / clinician and / or to generate a digital report by the reporting generation module 158.

[0064] Exemplary computer-implemented machine learning model FIG. 2 depicts an exemplary computer-implemented machine learning model training method 200 according to some aspects. The method 400 can be implemented using the environment 100 of FIG. 1. In some aspects, the method 200 includes identifying a patient training data set by performing cohort selection (block 202a), as discussed below. The patient training data set can be received from a database such as the database 110 of FIG. 1 and can include public data (e.g., data from The Cancer Genome Atlas (TCGA)) and / or proprietary data sets.

[0065] FIG. 3 depicts a method 300 that can be used in block 202a for the method to perform cohort selection in the case of endometrioid endometrial cancer, as discussed in more detail below. The method 200 can include performing univariate gene selection / biomarker identification (block 202b).

[0066] Method 200 may include performing dataset correction for technical RNA bias (e.g., domain adaptation). With respect to a reference training dataset (e.g., a matching Tempus RNA dataset) (block 202d) (block 202c). For example, method 200 may include correcting the technical RNA bias between TCGA and the Tempus assay using a reference cohort of Tempus primary patients (e.g., n = 111), and selecting relevant bias correction hyperparameters in a three-way cross-validation experiment. The corrected cohort may be referred to as the Tempus Adapted TCGA Cohort (TATC). In some embodiments, the SpinAdapt technology may be used to perform correction of RNA data across laboratories. This may enable external datasets (e.g., TCGA) to appear as if they were internally sequenced and be extensible to other data modalities such as DNA and electronic medical records. Using SpinAdapt may remove the unnecessary bioinformatics burden and enable the application of existing molecular predictors to external data without retraining. SpinAdapt is a privacy-protecting technology, which is advantageous. SpinAdapt is the only known algorithm that enables application to new expected data by sharing summary statistics without the need to share actual patient data. For example, the techniques described in any of the following publications may be used for RNA correction in this technique. U.S. Patent Application No. 17 / 405,025, titled "Systems and Methods for Homogenization of Disparate Datasets," filed on August 8, 2021, U.S. Patent Application No. 17 / 548,118, titled "Systems and Methods for Homogenization of Disparate Datasets," filed on December 10, 2021, and U.S. Patent Application No. 17 / 548,084, titled "Systems and Methods for Homogenization of Disparate Datasets," filed on December 10, 2021.Each of the foregoing publications is hereby incorporated by reference in its entirety for all purposes.

[0067] Method 200 may include performing multivariate gene selection / gene signature (block 202e). The model training in blocks 202b - 202e is further described below with respect to FIGS. 5A, 5B, and 5C. Method 200 may include using a multivariate cox proportional hazards (CoxPH) model with an elastic net penalty to select a gene signature on the TATC cohort. The associated hyperparameters may be selected in a 3 - fold cross - validation experiment. For gene signature characterization, the selection procedure may be repeated for a number of TATC bootstraps (e.g., 1000), and method 200 may select genes enriched in a characteristic gene set including estrogen, proliferation, and invasion gene groups.

[0068] Method 200 may include training a survival model (e.g., a cox proportional hazards (CoxPH) model) (block 202f). Method 200 may include using hyperparameter optimization to select the hyperparameters of blocks 202b - 202e as discussed herein (block 202g). Method 200 may include training the survival model in block 202f by performing blocks 202b - 202e using the patient training dataset selected in block 202a, the gene selection hyperparameters and survival model hyperparameters determined in block 202g. Method 200 may include selecting a decision threshold regarding the output of the trained survival model (e.g., the 70th percentile of the log partial - hazard scores evaluated on the corrected training dataset) to identify a molecular high - risk patient population (i.e., a population with negative or poor prognosis).

[0069] Specifically, the binary decision threshold can be set to the 60th, 65th, 70th, 75th, or 80th percentile of the log partial hazard scores on the training data. Method 200 can select a cutoff percentile based on the model performance in the training cohort. In some embodiments, this threshold model can classify patients as having high molecular risk (molecular risk score > 0.255) or low molecular risk (MR score ≤ 0.255).

[0070] Method 200 can train a survival model (e.g., CoxPH) on the selected genes to predict molecular risk. The molecular risk can be characterized by the log partial hazard score, and thus the molecular risk is linearly proportional to the expression of the selected genes. The selected gene signature can include: ALPL, APOBEC3G, BCL9, BNIPL, CARD10, CDKN2A, CENPF, EDN1, FAM83D, GGH, GNLY, HSPA1A, ISM1, ITGAL, KCNH3, KIF2C, LAMA3, LRRC23, MICA, MSLN, NKG7, PALM3, TDRKH, TPX2. In some embodiments, the selected gene signature my can include one or more of the foregoing genes, two or more of the foregoing genes, three or more of the foregoing genes, four or more of the foregoing genes, five or more of the foregoing genes, or all of the foregoing genes.

[0071] To evaluate the stability of the gene signature, method 200 can repeat the training on multiple bootstrap samples of the training dataset using preselected hyperparameters. The idea of repeating through bootstrapping is to analyze various bootstraps of the training population, compile a list of genes consistently selected across bootstraps, and ensure that the gene signature is reproducible across bootstraps. In some embodiments, method 200 can perform the bootstrap experiment 250 times using the selected hyperparameters, and method 200 can add to the list the top 100 genes most frequently selected across 250 bootstraps.

[0072] Method 200 can analyze which of the genes within a gene signature were among the top 100 bootstrap genes (i.e., signature stability). In empirical tests, 80% of the genes from a signature (e.g., 24 genes) were among the top approximately 35 genes from the bootstrap list, and nearly all of the signature genes were among the top 100 bootstrap genes.

[0073] Method 200 can also perform gene set enrichment analysis. For example, method 200 can perform gene set enrichment analysis on the list of the top 100 bootstrap genes to characterize the types of genes selected by method 200 with a pre - selected set of hyperparameters. The following characteristic gene sets have a statistically significant overlap with the gene list (adjusted p - value < 0.05): G2 - M checkpoint, early estrogen response, late estrogen response, E2F target, epithelial - mesenchymal transition, angiogenesis, MEL18 DN.V1 UP, P53 DN.V1 DN.

[0074] Method 200 can train a survival model with a selected gene signature (e.g., 24 genes). A patient's molecular / RNA risk score (low - molecular - risk or high - molecular - risk) can be defined as the log - partial - hazard output of the survival model. The partial - hazard function can be a time - invariant scaling coefficient in the hazard function that captures the contribution of a covariate (e.g., a gene) to the hazard rate. The logarithm of the log - partial - hazard score e (k) For each unit change, the hazard rate can be scaled by a factor of k. For example, if the log - partial - hazard molecular risk score increases by e (2) units, the hazard rate can double.

[0075] Aspects of an exemplary computer - implemented cohort selection method Figure 3 depicts an exemplary optional cohort selection method 300 before the pipeline. In some examples, there is no universally excellent machine learning model that functions well for all target populations. Instead, each machine learning model generally functions best for a specific target population. The present technique can define the target population. Accordingly, the present technique may include a cohort selection module that cuts or stratifies patient population training data so that the model learns best with respect to the target population. Thereby, the present technique becomes capable of training with a large dataset aligned with the target population. For example, the cohort selection module may remove histology based on a particular attribute (e.g., cancer tumor, sarcoma, mixed cells, adenocarcinoma, etc.) from the training data through the exploration process (e.g., via a loop). In some aspects, the cohort selection module may use a greedy algorithm strategy to successively use one histology attribute to exclude data and determine whether doing so affects matching the training dataset to the target population.

[0076] For some tasks, the greedy algorithm strategy can result in a statistically significant improvement in matching with the target population. For example, an empirical test using the greedy algorithm demonstrated that removing sarcoma from the training dataset for training a cancer prognosis predictor (e.g., an endometrial cancer prognosis predictor) resulted in much better alignment of the training data with the target population. Effectively, the greedy algorithm excludes patients with diseases (e.g., sarcoma) not seen in the target population from the training dataset.

[0077] As contemplated, the technique can include training an RNA seq-based gene expression profiling machine learning model to prognostically label patients as high RNA risk or low RNA risk, where high RNA risk patients are more likely to have an event (e.g., a progression event) and low RNA risk patients are less likely to have an event. This gene expression profiling machine learning model can, in some embodiments, be validated using Tempus data and trained using public data (e.g., TCGA data). The technique can include, in particular, advantageous improvements over conventional techniques by enabling the use of censored patient data for training. For example, the survival modeling module 158 can include instructions for training a survival model (e.g., a Cox proportional hazards (CoxPH) model) to model the survival of censored and uncensored patients and to enable a larger number of trainings.

[0078] Specifically, when attempting to generate a prediction to stratify patients as high RNA risk or low RNA risk, the event of interest can be a distant recurrence (cancer that recurs in a different part of the body than where it was first detected). For example, the survival modeling module 156 can process multiple patient data at a 5-year follow-up time point to determine whether the associated patients had a distant recurrence in the previous 5 years. Of course, this time period can be set to different values or determined dynamically during the study. In any case, this type of modeling can be difficult to achieve using a classifier due to several complex factors.

[0079] The first complex factor is the lack of follow-up observations. For example, some patients have not received a full five-year follow-up, and their records may be associated with partial data (e.g., six months, two years, three years, or less). Such patients' records may be incomplete, but the data can still reflect the absence (or presence) of distant recurrence. Traditional classification models cannot use the training data of such patients because they classify only patients who had distant recurrence or did not have distant recurrence within a certain period common to all patients.

[0080] On the other hand, this technique can capture such partial data of the truncated patients and use it for model training (e.g., using a survival model). Therefore, this technique advantageously enables the use of more training data including the data of truncated patients and non-truncated patients, enables training on more patient data, and thus improves the robustness of the model.

[0081] The clinical risk module 154 may include a set of computer-executable instructions for determining a patient's clinical risk group. Here, the clinical risk is different from the RNA risk. Specifically, the clinical risk refers to the risk group assigned to the patient by the clinician, as shown in FIG. 4. In contrast, this technique can also generate RNA risk annotations, classifications, and / or scores for a patient, representing more refined risks, based on, for example, processing the patient's RNA information using a trained model. The RNA risk annotation is different from the clinical risk category in FIG. 4. For example, a patient may be in the high clinical risk group 402d but have a calculated low-risk RNA risk annotation. Therefore, in many cases, the clinical risk and RNA risk annotations do not match, and it is intended that they may not match.

[0082] The clinical risk module 154 can determine the clinical risk group of a patient by referring to clinical factors (e.g., age, stage, fascial invasion status, lymphatic space invasion (LVSI) status, and / or other clinical factors). The clinical risk module 154 can classify patients according to predetermined criteria. For example, the clinical risk module 154 may include instructions for identifying patients with high to intermediate risk of a given cancer (e.g., endometrioid endometrial cancer). For example, the clinical risk module 154 may include a set of rules for determining whether a patient is a member of a high to intermediate risk population based on the patient's values for one or more clinical factors. It should also be understood that the clinical risk module 154 can classify each patient into other clinical groups, cohorts, and populations based on the patient's clinical factors. For example, another patient cohort (regardless of risk) is patients at an early stage.

[0083] Exemplary computer-implemented aspects of high to intermediate risk cohort selection As contemplated, in some aspects, the technique may attempt to stratify a patient's cancer risk by targeting the high to intermediate clinical risk cohort of early stage endometrioid endometrial cancer patients. However, this cohort may have multiple complex definitions. For example, during the cohort and model development process, the technique may include instructions for implementing a patient cohort selection algorithm.

[0084] For example, in some aspects, the algorithm implemented by the clinical risk module 154 may require that the patient have endometrioid histology and be at a specific stage (e.g., stage I or II). Stage II patients may or may not be high to intermediate risk patients, and thus, in some cases, the algorithm may apply published criteria.

[0085] For example, the clinical risk module 154 is incorporated herein by reference in its entirety for all purposes by Henry M Keys et al., “A phase III trial of surgery with or without adjunctive external pelvic radiation therapy in intermediate risk endometrioid endometrial adenocarcinoma: a Gynecologic Oncology Group study”, Gynecologic Oncology, Volume 92, Issue 3, 2004, Pages 744 - 751, ISSN 0090 - 8258, https: / / doi.org / 10.1016 / j.ygyno.2003.11.048 (https: / / www.sciencedirect.com / science / article / pii / S0090825803008631) (hereinafter referred to as “GOG”), or is incorporated herein by reference in its entirety for all purposes by B.G. Wortman et al., “Ten-year results of the PORTEC-2 trial for high-intermediate risk endometrial carcinoma: improving patient selection for adjuvant therapy” Nature, British Journal of Cancer (2018) 119:1067 - 1074, https: / / doi.org / 10.1038 / s41416-18-0310-8 (hereinafter referred to as “PORTEC-2”), or is incorporated herein by reference in its entirety for all purposes by Scholten et al., the instructions may include rules and criteria codified from those described in "Postoperative radiotherapy for Stage 1 endometrial carcinoma: Long - term outcome of the randomized PORTEC trial with central pathology review", International Journal of Radiation Oncology*Biology*Physics, Volume 63, Issue 3, 2005, Pages 834 - 838, ISSN 0360 - 3016, https: / / doi.org / 10.1016 / j.ijrobp.2005.03.007 (https: / / www.sciencedirect.com / science / article / pii / S0360301605004190) (hereinafter referred to as "PORTEC").

[0086] In some aspects, the patient cohort selection algorithm implemented by the clinical risk module 153 may further require that the patient be stage II or grade 3 with deep invasion.

[0087] Table 1 below summarizes the definitions that the clinical risk module 154 can implement in code and apply to patient data (e.g., the patient's electronic medical record).

Table 1 - 1

Table 1 - 2

[0088] Figure 3 depicts a method 300 for performing cohort selection in endometrial cancer according to some embodiments. Method 300 may include identifying a patient (baseline) having a uterine subtype and a primary site in the endometrium or uterus (block 301). Method 300 may include excluding any patient who is not RNA V1 (block 302). Method 300 may include identifying (and including) only progression-free period eligible patients (block 303). Method 300 may include identifying (and excluding) patients having data related to sarcoma cancer (block 304). Method 300 may include identifying (and excluding) patients having data related to serous tissue cancer and squamous tissue cancer (block 305). Method 300 may include identifying (and excluding) patients having data related to sarcoma cancer, serous tissue cancer, and squamous tissue cancer.

[0089] The patient cohort selection algorithm represents an advantageous improvement over conventional techniques that do not define whether a patient is high to intermediate, as discussed above and as confirmed by clinician interviews and clinical trial reviews.

[0090] Exemplary aspects of computer-implemented risk stratification model training Using RNA, patients within this group are stratified into high-risk RNA patients and low-risk RNA patients. A clinician (e.g., oncologist) can then use this information to treat patients in the high to intermediate clinical risk group with chemotherapy even if they are early-stage patients.

[0091] Stepwise feature selection, signature selection is used. This is a lower-cost computational method for performing coarser gene selection.

[0092] The 20,000 genes are curated into 1,000 genes. The smaller candidate genes are set to perform more complex, computationally intensive, and more data-hungry methods for more refined feature selection.

[0093] A method implemented for coarser selections that are less accurate, less precise, and faster, followed by more complex models on a refined set of candidates.

[0094] This makes the problem computationally feasible and also enables more data-hungry methods to be run on a smaller set of candidates.

[0095] RNA bias correction can be performed to make two or more datasets with different distributions more similar so that they can be used together as training data. This bias correction can be performed between two steps. By doing this in the middle, learning of bias correction between biased RNA datasets can be done using sufficient data to learn a good mapping between the two biased RNA datasets without the model having so much data (i.e., a very large number of genes) that it is overwhelmed or without so many genes that the RNA correction mapping fails to fit.

[0096] An ordered combination of univariate selection, followed by bias correction, followed by multivariate selection enables learning of a good mapping between biased RNA datasets without having so many genes that it becomes difficult to learn the bias.

[0097] This technique may include a model trained using molecular data (e.g., RNA data), process a patient's molecular data, and generate a high RNA risk annotation indicating to a clinician or other reviewer that among high to intermediate clinical risk groups, the patient is more likely to have a distant recurrence and should be escalated to chemotherapy earlier. Then, for example, instead of sending high to intermediate clinical patients home after resection, the patient can be scheduled immediately and aggressively for chemotherapy treatment (for example), which is likely to lead to a favorable outcome for the patient. This technique can identify additional subgroups beyond the high RNA risk.

[0098] For example, FIG. 4 depicts a table 400 of a plurality of clinical risk groups, each of which can be stratified for RNA risk using the present molecular modeling technique. As discussed herein, the modeling approach can be deployed to further stratify risk within each of the clinical risk groups 402. Table 400 includes a low clinical risk group 402a, a low-to-medium clinical risk group 402b, a high-to-medium clinical risk group 402c, and a high clinical risk group 402d. Table 400 also includes several clinical treatment strategies 404a-404d, including radiotherapy and chemotherapy. The question presented to the clinician is which of the four treatment strategies should be pursued for a given patient. In some embodiments, the technique can be used to model the RNA-based risk of patients in the high-to-medium clinical risk group 402c. For example, when a trained model operated by the model manipulation module 160 annotates a patient in the high-to-medium clinical risk group 402c as having a high RNA risk, the patient RNA-based risk stratification computing device 102 can generate an indication that the patient's care should be escalated to clinical treatment strategy 404a or clinical treatment strategy 404b. When the technique annotates a patient in the high-to-medium clinical risk group 402c as having a low RNA risk, the patient can be de-escalated to the observation clinical treatment strategy 404d, thereby avoiding the implementation of a potentially unnecessary clinical treatment strategy 404c (brachytherapy).

[0099] Brachytherapy is known to cause several common side effects, including fatigue, skin irritation, sexual dysfunction, bowel problems, and nausea. Accordingly, the technique includes improvements that are more advantageous than the current state of the art. Specifically, by processing a patient's RNA using a trained machine learning model, the technique can stratify an opaque / uniform clinical risk group into more differentiated RNA risk annotations, thereby avoiding unnecessary treatments that are likely to cause adverse events or side effects and / or that may negatively impact a patient's outcome.

[0100] This advantage is further amplified when considering the high clinical risk group 402d of FIG. 4. This group consists of patients who have been assessed as having a high clinical risk. Thus, purely on a clinical basis, treatment strategies include treatment strategy 404a (in some instances, systemic therapy, i.e., chemotherapy) and external beam radiation therapy (EBRT). Both of these treatments pose common and significant risks and side effects to patients, including hair loss, fatigue, easy bruising and bleeding, infection, anemia, nausea, sleep disturbances, diarrhea, and pain. The present technique can potentially stratify some of the patients in the high clinical risk group 402d with a low RNA risk annotation, which may enable a clinician to recommend avoiding a disruptive medical treatment that can significantly impact a patient's outcome and be associated with adverse events or side effects. Thus, the present technique represents a significant advancement and improvement over conventional cancer prognostic modeling techniques (e.g., techniques based solely on clinical data without considering molecular data).

[0101] This RNA risk modeling technique is applicable to endometrioid endometrial cancer, acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), adolescent cancer, adrenocortical carcinoma, AIDS-related cancer, Kaposi sarcoma (soft tissue sarcoma), AIDS-related lymphoma (lymphoma), primary CNS lymphoma (lymphoma), anal cancer, appendiceal cancer, astrocytoma, pediatric (brain cancer), atypical teratoid / rhabdoid tumor, pediatric, central nervous system (brain cancer), basal cell carcinoma of the skin, bile duct cancer, bladder cancer, bone cancer (including Ewing sarcoma, osteosarcoma, and malignant fibrous histiocytoma), brain tumor, breast cancer, bronchial tumor (lung cancer), Burkitt lymphoma, carcinoid tumor (gastrointestinal), cancer of unknown primary, cardiac (heart) tumor, pediatric, central nervous system, atypical teratoid / rhabdoid tumor, pediatric (brain cancer), medulloblastoma and other CNS fetal tumors, pediatric (brain cancer), germ cell tumor, pediatric (brain cancer), primary CNS lymphoma, cervical cancer, pediatric cancer, childhood cancer, rare bile duct cancer, chordoma, pediatric (bone cancer), chronic lymphocytic leukemia (CLL), chronic myelogenous leukemia (CML), chronic myeloproliferative neoplasm, colorectal cancer, craniopharyngioma, pediatric (brain cancer), cutaneous T-cell lymphoma, ductal carcinoma in situ (DCIS), pediatric (brain cancer), endometrial cancer (uterine cancer), epithelioma, pediatric (brain cancer), esophageal cancer, nasal neuroblastoma (head and neck cancer), Ewing sarcoma (bone cancer), extracranial germ cell tumor, pediatric, extragonadal germ cell tumor, eye cancer, intraocular melanoma, retinoblastoma, fallopian tube cancer, fibrous histiocytoma of bone, malignant, and osteosarcoma, gallbladder cancer, gastric (stomach) cancer, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST) (soft tissue sarcoma), germ cell tumor, pediatric central nervous system germ cell tumor (brain cancer), pediatric extracranial germ cell tumor, extragonadal germ cell tumor, ovarian germ cell tumor, testicular cancer, gestational trophoblastic disease, hairy cell leukemia, head and neck cancer, heart tumor, pediatric, hepatocellular (liver) cancer, histiocytosis, Langerhans cell, Hodgkin lymphoma, hypopharyngeal cancer (head and neck cancer), intraocular melanoma, islet cell tumor, pancreatic neuroendocrine tumor, Kaposi sarcoma (soft tissue sarcoma), kidney (renal cell) cancer, Langerhans cell histiocytosis, laryngeal cancer (head and neck cancer), leukemia, lip and oral cavity cancer (head and neck cancer), liver cancer, lung cancer (non-small cell, small cell, pleuropulmonary blastoma, and tracheobronchial tumor), lymphoma, male breast cancer,Prognosis prediction applicable to multiple different types of cancers, including malignant fibrous histiocytoma and osteosarcoma of bone, melanoma, melanoma, intraocular (eye), Merkel cell carcinoma (skin cancer), mesothelioma, malignant, metastatic cancer state, metastatic squamous cell carcinoma of unknown primary origin of the neck (head and neck cancer), midline duct carcinoma with NUT gene alteration, oral cancer (head and neck cancer), multiple endocrine neoplasia syndrome, multiple myeloma / plasma cell neoplasm, mycosis fungoides (lymphoma), myelodysplastic syndrome, myelodysplasia / myeloproliferative neoplasm, chronic myelogenous leukemia (CML), acute myelogenous leukemia (AML), chronic myeloproliferative neoplasm, cancer of the nasal cavity and paranasal sinuses (head and neck cancer), nasopharyngeal cancer (head and neck cancer), neuroblastoma, non-Hodgkin lymphoma, non-small cell lung cancer, oral cancer, cancer of the lips and oral cavity and oropharyngeal cancer (head and neck cancer), osteosarcoma and malignant fibrous histiocytoma of bone, ovarian cancer, pancreatic cancer, pancreatic neuroendocrine tumor (islet cell tumor), papillomatosis (pediatric larynx), paraganglioma, cancer of the paranasal sinuses and nasal cavity (head and neck cancer), parathyroid cancer, penile cancer, pharyngeal cancer (head and neck cancer), pheochromocytoma, pituitary tumor, plasma cell neoplasm / multiple myeloma, pleuropulmonary blastoma (lung cancer), breast cancer during pregnancy, primary central nervous system (CNS) lymphoma, primary peritoneal cancer, prostate cancer, rectal cancer, recurrent cancer, renal cell (kidney) cancer, retinoblastoma, rhabdomyosarcoma, pediatric (soft tissue sarcoma), salivary gland cancer (head and neck cancer), pediatric rhabdomyosarcoma (soft tissue sarcoma), pediatric hemangioma (soft tissue sarcoma), Ewing sarcoma (bone cancer), Kaposi sarcoma (soft tissue sarcoma), osteosarcoma (bone cancer), soft tissue sarcoma, uterine sarcoma, Sézary syndrome (lymphoma), skin cancer, small cell lung cancer, small intestine cancer, soft tissue sarcoma, squamous cell carcinoma of the skin, squamous cell carcinoma of the neck of unknown primary origin, metastatic (head and neck cancer), gastric cancer (Stomach (Gastric) Cancer), T cell lymphoma, lymphoma (mycosis fungoides and Sézary syndrome), testicular cancer, laryngeal cancer (head and neck cancer), superior pharyngeal cancer, oral pharyngeal cancer, hypopharyngeal cancer, thymoma and thymic carcinoma, thyroid cancer, tracheobronchial tumor (lung cancer), transitional cell carcinoma of the renal pelvis and ureter (kidney (renal cell) cancer), ureter and renal pelvis, transitional cell carcinoma (kidney (renal cell) cancer, urethral cancer, cervical cancer, endometrium, uterine sarcoma, vaginal cancer, hemangioma (soft tissue sarcoma), or vulvar cancer, etc.

[0102] For example, breast cancer prognostic models can be trained to determine molecular risk, which can be used as contemplated herein to determine which patients to escalate. In particular, such models can escalate newly diagnosed HR+ / HER2- patients to endocrine therapy and chemotherapy. Another cancer diagnostic model can be trained and used to predict molecular risk, which can be used to escalate newly diagnosed low-stage prostate cancer for some patients. Generally, during the treatment of many early-stage solid organ cancers, there is a question of which systemic therapy, before or after surgery, is best for the patient to reduce the likelihood of recurrence and death. An advantageous improvement of the present technique is to base such a determination of whether to give systemic therapy to early-stage patients on prognostic diagnosis (risk stratification) using clinical, pathological, and / or molecular biomarkers.

[0103] In some embodiments, the diagnosis can include non-glioma brain (epithelial tumor, hemangioblastoma, medulloblastoma, meningioma), breast (ductal, lobular), colon, endometrium (endometrial, serous endometrial, endometrial stromal sarcoma), gastroesophageal (esophageal adenocarcinoma, stomach), gastrointestinal stromal tumor, glioma (glioblastoma, anaplastic glioma), head and neck adenocarcinoma, blood (acute lymphoblastic leukemia, acute myeloid leukemia, b-cell lymphoma, chronic lymphocytic leukemia, chronic myeloid leukemia, Rosai-Dorfman, t-cell lymphoma), hepatobiliary (cholangiocarcinoma, gallbladder, liver), lung adenocarcinoma, melanoma, mesothelioma, neuroendocrine (gastrointestinal neuroendocrine, high-grade neuroendocrine lung, low-grade neuroendocrine lung, pancreatic neuroendocrine, cutaneous neuroendocrine), ovary (ovarian clear cell, ovarian granulosa, ovarian serous), pancreas, prostate, kidney (chromophobe kidney, renal clear cell, renal papillary), sarcoma (chondrosarcoma, chordoma, Ewing sarcoma, fibrosarcoma, leiomyosarcoma, liposarcoma, osteosarcoma, rhabdomyosarcoma, synovial sarcoma, angiosarcoma), squamous epithelium (cervical, esophageal squamous, head and neck squamous, lung squamous, cutaneous / squamous basal), thymus, thyroid, or urothelium.

[0104] In some embodiments, the diagnosis may include one or more entries of ICD-10-CM, or International Classification of Disease. The ICD provides a way to classify diseases, injuries, and causes of death. The World Health Organization (WHO) has published the ICD to standardize the way of recording and tracking cases of diagnosed diseases, including cancer. For example, classification from any chapter of cancer from ICD or Chapter 2, C and D codes. C codes include neoplasms of the lip, oral cavity, and pharynx (C00-C14), neoplasms of the digestive organs (C15-C26), neoplasms of the respiratory system and intrathoracic organs (C30-C39), neoplasms of the mesothelium and soft tissues (C45), neoplasms of the bone, joints, and articular cartilage (C40-C41), neoplasms of the skin (melanoma, Merkel cell, and other skin histologies) (C43, C44, C4a), Kaposi sarcoma (46), neoplasms of the peripheral nerves and autonomic nervous system, retroperitoneum, peritoneum, and soft tissues (C47, C48, C49), neoplasms of the breast and female genital organs (C50-C58), neoplasms of the male genital organs (C60-C63), neoplasms of the urinary tract (C64-C68), neoplasms of the eye, brain, and other parts of the central nervous system (C69-C72), neoplasms of the thyroid, other endocrine glands, and unspecified local sites (C73-C76), malignant neuroendocrine tumors (C7a._), secondary neuroendocrine tumors (C7B), neoplasms of other unspecified local sites (C76-80), secondary and unspecified malignant neoplasms of lymph nodes (C77), secondary cancers of the respiratory and digestive systems, other unspecified sites (C78-80), malignant neoplasms without site designation (C80), lymphomas, or malignant neoplasms of hematopoietic and related tissues (C81-C96).

[0105] Exemplary Machine Learning Model Training Returning to FIG. 2, method 200 includes performing univariate gene selection / biomarker identification (block 202b). Method 200 may include performing dataset correction for technical RNA bias (e.g., domain adaptation) (block 202c) with respect to a reference dataset (e.g., a matching Tempus RNA dataset) (block 202d). Method 200 may include performing multivariate gene selection / gene signature (block 202e). For example, the training data in block 202 may include a dataset having RNASeq data and related clinical outcomes. For example, the following training datasets may be used: 1. TCGA endometrial (N = 400, clinical assessment item = PFS) a. TCGA stage 1 and 2 (N = 300) b. TCGA HIR (N = 170) 2. Tempus endometrial primary (N = 110, clinical assessment item = PFS) 3. Tempus endometrial HIR (N = 60, clinical assessment item = PFS) 4. In-house (N = 181, clinical assessment item = distant recurrence) a. In-house - non-HIR (N = 70) b. In-house HIR (N = 100) i. GOG (N = 50) ii. PORTEC (N = 80) iii. GOG or PORTEC (N = 100)

[0106] As described above, method 200 may include hyperparameter optimization in block 202g. For example, during hyperparameter search / optimization, performance may be evaluated on in-house - non-HIR for hyperparameter selection. - Hyperparameters: - RNA correction method: [z-score, spinadapt, mean correction] - Reference for RNA correction: [Tempus primary] - Univariate gene selection: [300, 400, 500] - Penalty: [1e-3, 1e-2, 1e-1, 1e0, 1e1] - L1 ratio: [1e-2, 1e-1, 0.2, 0.3, 0.5, 1] - Hcoef: [0.05, 0.1, 0.25] - Lcoef: [0.05, 0.1, 0.25] - Alpha: [0.01, 0.05, 0.1, 0.15, 0.2, 0.25, 0.3, 0.5]

[0107] Hyperparameter Search: Performance Metric Train on TCGA stage 1 and 2 patients with Tempus RNA V2 endometrial cancer - corrected RNASeq. Optimize hyperparameters for the in - institution non - HIR dataset. The performance metrics are shown in Table 2 below. [Table 2]

[0108] Once the hyperparameters are selected, Method 200 can be trained on all TCGA stage 1 - 2 patients (n = 323) using Method 200.

[0109] In block 202a, Method 200 can perform univariate gene selection to reduce the number of acceptable genes. Method 200 can regress the expression of each gene on survival data across all patients (e.g., using a univariate cox model). Method 200 can set a tunable parameter by selecting the K most important genes (e.g., K is about 500). In some aspects, the expression of each gene can be regressed against survival data across all TCGA patients using a univariate cox model to select a set of candidate genes (e.g., 1000 genes).

[0110] For example, in block 202f, Method 200 can perform multivariate gene selection using a multivariate (e.g., CoxPH) survival model with an Elastic Net penalty. Tunable parameters can include a regularization term, L1 and L2 penalty weights.

Number

[0111] These tunable parameters determine the number of genes, and they can be selected (e.g., in a 3-fold CV experiment). Method 200 may, in block 202f, select the inventors' genes and initialize model training using the selected hyperparameters to train a CoxPH model. Method 200 may perform threshold selection in block 202g by selecting a threshold that is the 70th percentile of the log partial hazard scores on the TCGA stage 12 dataset. FIG. 5A depicts a heatmap 510 showing the clustering of the final gene set as determined by method 200 of FIG. 2.

[0112] The method may evaluate a survival model (e.g., CoxPH) trained on the TCGA stage 12 training dataset via calibration plots for the low-risk and high-risk groups. The two groups can be selected based on a single threshold determined based on the log partial hazard scores of the TCGA stage 12 patients. FIG. 5B depicts a plot 510 of the Kaplan-Meier (KM)-adjusted event risk rate, where the x-axis corresponds to the percentile of the TCGA stage 12 log partial hazard score of the decision threshold.

[0113] Any point on the high-risk curve gives the 4-year KM-adjusted risk of patients with a log partial hazard score above the threshold reported on the x-axis, and the percentile is calculated at TCGA stage 12. Any point on the low-risk curve gives the 4-year KM-adjusted risk of patients with a log partial hazard score below the threshold reported on the x-axis. Of course, method 200 can calculate risk curves for any suitable time scale as long as sufficient data exists.

[0114] Figure 5B demonstrates a practical application of the present technique. First, method 200 shows that patients can be further stratified based on their risk levels. Patients with higher log partial hazard scores are more likely to be at higher risk of an event (e.g., cancer recurrence) within a time period. Figure 5B also demonstrates that different thresholds of the log partial hazard score can be used to predict outcomes. By observing where the risk changes significantly, it helps to determine the most appropriate cut-off point for clinical decision-making. Figure 5B also enables clinicians to compare risks across cancer stages, make better treatment decisions, and identify higher-risk patients.

[0115] Figure 5C depicts a survival plot 520 of patients with TCGA stage 12, and the decision threshold is set at the 70th percentile of the log partial hazard score on the TCGA stage 12 dataset.

[0116] Using the present invention, publicly available datasets can be utilized to train RNA-based prognostic models. As a result, the models not only become less expensive to train but also more reproducible across patient cohorts from heterogeneous sources. These models can be used to predict (i) the risk of distant recurrence, (ii) the risk of local recurrence, and (iii) the risk of cancer progression.

[0117] Furthermore, these models can be easily evaluated on in-house sequenced datasets after standard normalization procedures (VST normalization or log-transformed transcripts per million).

[0118] As described above, there are unmet needs in understanding the risks for cancer patients, particularly high- to intermediate-risk patients. This molecular classification and gene profiling technique directly addresses these unmet needs and can be used to predict the risk of distant recurrence in early-stage endometrioid endometrial cancer, focusing on high- to intermediate-risk patients. The RNA-seq-based gene expression profiler of this technique can be trained using TCGA data, resulting in a gene signature (e.g., a 24-gene signature) that classifies (e.g., labels, profiles, or annotates) the respective molecular risk (e.g., RNA-based risk prediction) of endometrioid endometrial cancer patients as either high or low. This profiling can be used to further stratify opaque / homogeneous patient groups.

[0119] Next, this gene expression profiler machine learning technique can be tested in an unselected cohort of endometrioid endometrial cancer patients (e.g., N = ~1000) from Tempus to examine its association with known pathologic or molecular prognostic features.

[0120] Empirical tests showed that this gene expression profiler machine learning technique demonstrated a significant enrichment of high molecular risk in G3 compared to G1 / 2 histology (p-value < 5e-8). A high correlation was found between the molecular risk score and the copy number change score (t-test p-value < 1e-5). Next, a clinical evaluation was performed in an early-stage endometrioid endometrial cancer case-control cohort of patients (N = 109) in whom 4-year recurrence from within the institution was documented or there were no recurrence events. Across the entire cohort, the high molecular risk group had a significantly higher distant recurrence rate compared to the low molecular risk group (HR = 4.8, N = 109). Next, subgroup analysis was performed in the clinically important high- to intermediate-clinical risk groups of patients. In this subgroup, the high molecular risk group showed a significantly higher distant recurrence rate compared to the low molecular risk group (HR = 8.0, N = 56). Finally, considering the significance of genomic biomarkers in the evolution of FIGO staging of endometrioid endometrial cancer, outcomes were stratified by the TCGA subtypes established as the reference standard, and subgroup analysis was performed in patients classified as not having a specific molecular profile (NSMP). Among patients who were NSMP, the high molecular risk molecular group showed a significantly higher distant recurrence rate compared to the low molecular risk group (HR = 7.92, N = 67). These evaluation studies demonstrated the performance of the gene expression profiler molecular risk machine learning model for distant recurrence risk stratification in early-stage endometrioid endometrial cancer, particularly in high- to intermediate-clinical risk patients, and can be used clinically to inform adjuvant clinical management.

[0121] Figure 7 depicts an exemplary digital report of a patient's molecular risk assessment and disease recurrence from pathological analysis, according to some embodiments. The report includes the patient's diagnosis of endometrial cancer and several sections that detail the results and their potential clinical significance. At the top of the report, the patient's name (redacted), diagnosis of endometrial cancer, accession number (redacted), and the following placeholders for visual details are displayed: - Date of birth (not shown) - Gender (female) - Physician (example name: Thomas) - Institution (not shown) - Test Data Institution (collected on January 4, 2021; received on May 4, 2023) - Tumor Specimen (endometrium) - Inspection Data Management Pathology Lab Information (Endometrial Algo) - Tumor Percentage (50%)

[0122] Also, the patient currently states that they are not eligible to participate in clinical trials in the database.

[0123] In the "Molecular Risk" section, the patient's risk group is highlighted as "MR-HIGH", the risk score is 62, and the threshold for the risk group on a 0 - 100 graphical scale is shown as 25. The text states, "This patient is predicted to have a high risk of distant recurrence if treated with brachytherapy or observation alone."

[0124] The "4-Year Distant Recurrence" section indicates that this patient is predicted to have a 31% risk of distant recurrence over 4 years if treated with brachytherapy or observation alone.

[0125] The "TCGA Subtype" section shows the subtype presented in filled bubble form as "NSMP".

[0126] At the bottom of the report, it includes the CLIA number, signature / report date (January 14, 2023), laboratory medical director, Tempus ID#, and pipeline version (3.10.0), as well as an electronic signature by the laboratory address (Tempus Labs, Inc. ● 600 West Chicago Avenue, Ste 510 ● Chicago, IL ● 60654 ● tempus.com ● support@tempus.com).

[0127] The report depicted in Figure 7, which provides a detailed assessment of a patient's molecular risk and a prediction of disease recurrence in patients diagnosed with endometrial cancer, can be generated through a series of computational and analytical processes outlined in the disclosed methods and systems for stratifying a patient's cancer risk using computational oncology and molecular data. For example, first, molecular data corresponding to a patient, which may include RNA sequencing data, can be received by a molecular risk state prediction computing device 102. This device is equipped with one or more processors that store computer-executable instructions for processing the molecular data, and a memory. The processing can include the use of a machine learning model trained using a patient training dataset and a reference training dataset. This model can employ techniques such as univariate gene selection, RNA bias correction, and multivariate gene selection to filter and correct its training data. Further, the model can incorporate a survival model, which can be, for example, a Cox proportional hazards model. The training of the machine learning model can include several steps, including the selection of a cohort of patients from the patient training dataset, the selection of genes using univariate and multivariate gene selection methods, and the correction of biases in the patient's molecular data. This process can also include optimizing hyperparameters to improve the accuracy of the model in predicting molecular risk. The training data used for this purpose can be from various sources, including public datasets such as The Cancer Genome Atlas (TCGA), and proprietary datasets from institutions (e.g., Tempus AI, Inc.).

[0128] Once the model is trained, the model processes the received molecular data to determine the risk of the patient's molecular data. This can include analyzing the patient's molecular data against the trained model to classify the patient's risk as high or low. In the case of the report in Figure 7, the patient is classified as having a high molecular risk ("MR-HIGH") with a specific risk score shown on a graphical scale.

[0129] Based on the patient's molecular data risk, a matched treatment strategy is generated. This strategy takes into account the patient's molecular risk profile and proposes appropriate treatment options. In the case of the patient in Figure 7, the report indicates that the risk of distant recurrence is high when treated with brachytherapy or observation only, suggesting that alternative treatment strategies may be more appropriate.

[0130] The report also includes additional sections such as "4-year distant recurrence," which provides a quantitative risk assessment of distant recurrence over a 4-year period, and "TCGA subtype," which classifies the patient's cancer subtype based on molecular data.

[0131] The generation of this report can be facilitated by a report generation module 158, which compiles the analysis results, predictions, and recommendations into an understandable format. This module may also include instructions for generating visual reports for comparison / validation that enhance the interpretability of the data for clinicians and patients.

[0132] By combining the various embodiments described above, further embodiments can be provided. All U.S. patents, U.S. patent application publications, U.S. patent applications, foreign patents, foreign patent applications, and non-patent publications mentioned herein and / or listed in the application data sheet are hereby incorporated by reference in their entirety for all purposes. Aspects of the embodiments can be modified as necessary to adopt concepts from the various patents, applications, and publications to provide further embodiments.

[0133] In light of the above detailed description, these and other modifications can be made to the embodiments. Generally, in the following claims, the terms used should not be construed as limiting the claims to the specific embodiments disclosed herein and in the claims, but rather the claims should be construed to include all possible embodiments together with the full scope of equivalents to which such claims are entitled. Accordingly, the claims are not limited by the present disclosure.

[0134] Aspects of the techniques described in this disclosure may include any of the following aspects, alone or in combination.

[0135] 1. A computer-implemented method for stratifying a patient's cancer risk using molecular data, comprising: receiving, via one or more processors, molecular data corresponding to a patient; processing, via one or more processors, the molecular data using a machine learning model to determine a molecular data risk for the patient, wherein the machine learning model is trained using a patient training data set and a reference training data set, and the machine learning model uses univariate gene selection, RNA bias correction, and multivariate gene selection to filter and correct the training data, and the machine learning model includes a survival model; and generating, based on the molecular data risk for the patient, a matched treatment strategy corresponding to the patient.

[0136] 2. The computer-implemented method of aspect 1, wherein the survival model is a Cox proportional hazards model.

[0137] 3. The computer-implemented method of aspect 1 or 2, wherein the molecular data corresponding to the patient includes RNA seq data.

[0138] 4. The computer-implemented method of any of aspects 1-3, wherein the cancer is endometrial cancer and the machine learning model is trained on a cohort of patient data selected using a greedy algorithm.

[0139] 5. The computer-implemented method of aspect 4, wherein the greedy algorithm includes identifying patients having a uterine subtype and a primary site of endometrium or uterus, identifying progression-free survival eligible patients, identifying patients having sarcoma cancer, identifying patients having serous tissue cancer and squamous tissue cancer, and identifying patients having sarcoma cancer, serous tissue cancer, and squamous tissue cancer.

[0140] 6. The computer-implemented method according to any one of aspects 1 to 5, wherein the patient has at least one prior clinical risk group assignment of low clinical risk, low to intermediate clinical risk, high to intermediate clinical risk, or high clinical risk.

[0141] 7. Generating a treatment strategy for a patient based on the molecular data risk of the patient includes generating a treatment strategy based on both (i) a prior clinical risk group and (ii) the molecular risk of the patient, the computer-implemented method according to aspect 6.

[0142] 8. The computer-implemented method according to aspect 7, wherein the prior clinical risk group is high to intermediate, the molecular risk of the patient is high, and the treatment strategy is at least one of systemic therapy or external beam radiation therapy.

[0143] 9. The computer-implemented method according to aspect 7 or 8, wherein the prior clinical risk group is high to intermediate, the molecular risk of the patient is low, and the matched treatment strategy is observation.

[0144] 10. The computer-implemented method according to any one of aspects 7 to 9, wherein the prior clinical risk group is low to intermediate, the molecular risk of the patient is high, and the matched treatment strategy is at least one of brachytherapy, external beam radiation therapy, or systemic therapy.

[0145] 11. The computer-implemented method according to any one of aspects 7 to 10, wherein the prior clinical risk group is high, the molecular risk of the patient is low, and the matched treatment strategy is observation.

[0146] 12. The computer-implemented method according to any one of aspects 7 to 11, wherein the prior clinical risk group is high, the molecular risk of the patient is low, and the matched treatment strategy is observation.

[0147] 13. The computer-implemented method according to any one of aspects 7 to 12, wherein the prior clinical risk group is high, the molecular risk of the patient is high, and the matched treatment strategy is at least one of systemic therapy or external beam radiation therapy.

[0148] 14. The computer-implemented method according to any one of aspects 1 to 13, wherein the matched treatment strategy includes at least one of systemic therapy, external beam radiation therapy, brachytherapy, or observation.

[0149] 15. A computer-implemented method for training a machine learning model to stratify a patient's cancer risk using molecular data, the method comprising, via one or more processors: (i) receiving a patient training dataset including the molecular data of each of a plurality of patients, and (ii) a reference training dataset including the molecular data of each of a plurality of patients; selecting a patient cohort from the patient training dataset via one or more processors; selecting a small subset of genes from the patient training dataset using univariate gene selection via one or more processors; generating a corrected reference training dataset by processing the reference training dataset to correct the bias of the molecular data of the plurality of patients; selecting an even smaller gene from the small subset of genes using multivariate gene selection via one or more processors; training a survival model via one or more processors, wherein training includes determining a set of hyperparameters; and selecting a decision threshold for identifying a population of patients having an RNA risk profile via one or more processors.

[0150] 16. The cancer risk of the patient is the risk of endometrial cancer, and selecting a patient cohort from the patient's training dataset via one or more processors includes applying a greedy algorithm via one or more processors, the greedy algorithm having uterine subtypes and identifying patients with a primary site in the endometrium or uterus, identifying progression-free survival eligible patients, identifying patients with sarcoma cancer, identifying patients with serous tissue cancer and squamous tissue cancer, and identifying patients with sarcoma cancer, serous tissue cancer, and squamous tissue cancer, thereby excluding patient data, the computer-implemented method according to aspect 15.

[0151] 17. The computer-implemented method according to aspect 15 or 16, wherein the molecular data includes at least some transcriptome data.

[0152] 18. The computer-implemented method according to any one of aspects 15 to 17, wherein the molecular data includes at least some data generated via RNA seq.

[0153] 19. Receiving a reference training dataset including the respective molecular data of a plurality of patients via one or more processors includes receiving the reference training dataset from a next-generation sequencing platform, the computer-implemented method according to any one of aspects 15 to 18.

[0154] 20. A computing system comprising one or more processors and one or more memories storing computer-executable instructions, the computer-executable instructions, when executed by the one or more processors, causing the computing system to receive molecular data corresponding to a patient and process the molecular data using a machine learning model to determine a molecular data risk for the patient, the machine learning model being trained using a patient training data set and a reference training data set, the machine learning model using univariate gene selection, RNA bias correction, and multivariate gene selection to filter and correct its training data,

[0155] the machine learning model including a survival model, causing the computing system to perform processing and generate a matched treatment strategy corresponding to the patient based on the patient's molecular data risk.

[0156] 21. The computing system according to aspect 20, wherein one or more memories store instructions that, when executed, cause the computing system to perform the functions described in any of aspects 2-14.

[0157] 22. A computing system comprising one or more processors and one or more memories storing computer-executable instructions, the computer-executable instructions, when executed by the one or more processors, causing the computing system to: (i) receive a patient training data set including the molecular data of each of a plurality of patients and a reference training data set including the molecular data of each of the plurality of patients; (ii) select a patient cohort from the patient training data set; (iii) select a small subset of genes from the patient training data set using univariate gene selection; (iv) process the reference training data set to generate a corrected reference training data set by correcting the bias of the molecular data of the plurality of patients; (v) select a smaller subset of genes from the small subset of genes using multivariate gene selection; (vi) train a survival model, the training including determining a set of hyperparameters; and (vii) select a decision threshold for identifying a population of patients having an RNA risk profile.

[0158] 23. The computing system according to aspect 22, wherein one or more memories store instructions that, when executed, cause the computing system to perform the functions described in any of aspects 16-19.

[0159] 22. A computer-readable medium storing computer-executable instructions that, when executed by one or more processors, cause the computer to receive molecular data corresponding to a patient and process the molecular data using a machine learning model to determine a molecular data risk for the patient, wherein the machine learning model is trained using a patient training data set and a reference training data set, the machine learning model uses univariate gene selection, RNA bias correction, and multivariate gene selection to filter and correct its training data, the machine learning model includes a survival model, and cause the computer to generate a matched treatment strategy corresponding to the patient based on the molecular data risk of the patient.

[0160] 23. The computer-readable medium according to claim 22, storing instructions that, when executed, cause the computer to perform the functions described in any of aspects 2 to 14.

[0161] 24. A computer-readable medium storing computer-executable instructions that, when executed by one or more processors, cause the computer to receive (i) a patient training data set including molecular data of each of a plurality of patients and (ii) a reference training data set including molecular data of each of a plurality of patients, select a patient cohort from the patient training data set, select a small subset of genes from the patient training data set using univariate gene selection, process the reference training data set to correct the bias of the molecular data of the plurality of patients to generate a corrected reference training data set, select an even smaller subset of genes from the small subset of genes using multivariate gene selection, train a survival model, wherein training includes determining a set of hyperparameters, and select a decision threshold for identifying a patient population having an RNA risk profile.

[0162] 25. A computer-readable medium storing instructions that, when executed, cause a computer to perform the functions according to any one of aspects 16 to 19, as recited in claim 22.

[0163] Additional Considerations The computer-readable medium may include executable computer-readable code stored on a computer for programming the computer with the techniques of this specification (e.g., including a processor and a GPU). Examples of such computer-readable storage media include hard disks, CD-ROMs, digital versatile disks (DVDs), optical storage devices, magnetic storage devices, ROM (read-only memory), PROM (programmable read-only memory), EPROM (erasable programmable read-only memory), EEPROM (electrically erasable programmable read-only memory), and flash memory. More generally, the processing unit of computing device 1300 may represent a CPU-type processing unit, a GPU-type processing unit, a TPU-type processing unit, a field programmable gate array (FPGA), another class of digital signal processor (DSP), or other hardware logic components drivable by a CPU.

[0164] A system for executing the methods described herein may include a computing device, and more specifically, may be implemented on one or more processing units, such as a central processing unit (CPU), and / or one or more graphics processing units (GPUs) including a cluster of CPUs and / or GPUs. The features and functions described may be stored on one or more non-transitory computer-readable media of the computing device and then implemented. The computer-readable media may include, for example, an operating system and software modules, or “engines,” that implement the methods described herein. These engines may be stored as a set of non-transitory computer-executable instructions. The computing device may be a distributed computing system such as Amazon Web Services, Google Cloud Platform, Microsoft Azure, or other public, private, and / or hybrid cloud computing solutions.

[0165] The computing device includes a network interface communicatively coupled to a network for communicating to and / or from a portable personal computer, smartphone, electronic document, tablet, and / or desktop personal computer, or other computing device. The computing device further includes an I / O interface connected to devices such as a digital display, user input device, and the like.

[0166] The functions of the engine can be implemented in distributed computing devices interconnected with each other via communication links. In other embodiments, the functions of the system can be distributed among any number of devices, including the portable personal computer, smartphone, electronic document, tablet, and desktop personal computer device shown. The computing devices can be communicatively coupled to a network and another network. The network can be a public network such as the Internet, a private network such as a research institution or corporate network, or any combination thereof. The network can include local area network (LAN), wide area network (WAN), cellular, satellite, or other network infrastructure, whether wireless or wired. The network can utilize communication protocols including packet-based and / or datagram-based protocols such as Internet Protocol (IP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), or other types of protocols. Further, the network can include a plurality of devices that facilitate network communication and / or form the hardware infrastructure of the network, such as switches, routers, gateways, access points (such as wireless access points as shown), firewalls, base stations, repeaters, backbone devices, etc.

[0167] A computer-readable medium may include executable computer-readable code stored on a computer for programming a computer with the techniques of this specification (e.g., including a processor and a GPU). Examples of such computer-readable storage media include hard disks, CD-ROMs, digital versatile disks (DVDs), optical storage devices, magnetic storage devices, ROM (read-only memory), PROM (programmable read-only memory), EPROM (erasable programmable read-only memory), EEPROM (electrically erasable programmable read-only memory), and flash memory. More generally, the processing unit of a computing device may represent a CPU-type processing unit, a GPU-type processing unit, a field programmable gate array (FPGA), another class of digital signal processor (DSP), or other hardware logic components that can be driven by a CPU.

[0168] Throughout this specification, multiple instances may implement components, operations, or structures described as a single instance. Although individual operations of one or more methods are illustrated and described as separate operations, one or more of the individual operations may not be performed simultaneously, and the operations need not be performed in the illustrated order. Structures and functions presented as separate components within an exemplary configuration may be implemented as a combined structure or component. Similarly, structures and functions presented as a single component may be implemented as separate components or multiple components.

[0169] Additionally, certain aspects are described herein as including logic or some routines, subroutines, applications, or instructions. These can constitute either software (e.g., code embodied on a machine-readable medium or in a transmission signal) or hardware. In hardware, routines and the like are tangible units capable of performing specific operations and can be configured or arranged in a specific manner. In an exemplary aspect, one or more computer systems (e.g., stand-alone, client, or server computer systems), or one or more hardware modules of a computer system (e.g., a processor or group of processors), can be configured as a hardware module that operates to perform the specific operations described herein by software (e.g., an application or a part of an application).

[0170] In various aspects, the hardware module can be implemented mechanically or electronically. For example, the hardware module can include dedicated circuitry or logic (e.g., a special-purpose processor such as a microcontroller, a field programmable gate array (FPGA), or an application specific integrated circuit (ASIC)) that is permanently configured to perform a specific operation. The hardware module can also include programmable logic or circuitry (e.g., that included within a general-purpose processor or other programmable processor) that is temporarily configured by software to perform a specific operation. It will be appreciated that whether to implement the hardware module mechanically, with dedicated and permanently configured circuitry, or with temporarily configured circuitry (e.g., configured by software) can be determined considering cost and time.

[0171] Accordingly, the term "hardware module" should be understood to encompass a tangible entity, one that is physically constructed or permanently configured (e.g., embedded in hardware) or temporarily configured (e.g., programmed) to operate in a particular manner or to perform certain operations described herein. Considering the case where a hardware module is temporarily configured (e.g., programmed), each of the hardware modules need not be configured or instantiated at any given instance. For example, if a hardware module includes a general-purpose processor configured using software, the general-purpose processor can be configured as different hardware modules at different times. Thus, software can configure a processor, for example, to constitute a particular hardware module at one time and a different hardware module at another time.

[0172] A hardware module can provide information to and receive information from other hardware modules. Accordingly, the described hardware modules can be regarded as communicatively coupled. If multiple such hardware modules exist simultaneously, communication can be achieved via signal transmission connecting the hardware modules (e.g., via appropriate circuitry and buses). In the case where multiple hardware modules are configured or instantiated at different times, communication between such hardware modules can be achieved, for example, via storage and retrieval of information in a memory structure accessed by the multiple hardware modules. For example, a hardware module can perform an operation and store the output of the operation in a memory device to which the hardware module is communicatively coupled. Subsequently, a further hardware module can later access the memory device to retrieve and process the stored output. A hardware module can also initiate communication with an input or output device and operate on a resource (e.g., collect information).

[0173] The various operations of the exemplary methods described herein can be performed, at least in part, by one or more processors temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors can constitute processor-implemented modules that operate to perform one or more operations or functions. In some exemplary aspects, the modules referred to herein can include processor-implemented modules.

[0174] Similarly, the methods or routines described herein can be at least in part processor-implemented. For example, at least some of the operations of a method can be performed by one or more processors or processor-implemented hardware modules. Certain performance of the operations can exist not only within a single machine but also be distributed among one or more processors deployed across several machines. In some exemplary aspects, one or more processors can be located in a single location (e.g., within a home environment, within a workplace environment, or as a server farm), but in other aspects, the processors can be distributed across multiple locations.

[0175] Certain performance of the operations can exist not only within a single machine but also be distributed among one or more processors deployed across several machines. In some exemplary aspects, one or more processors or processor-implemented modules can be located in a single location (e.g., within a home environment, within a workplace environment, or as a server farm). In other exemplary aspects, one or more processors or processor-implemented modules can be distributed across several locations.

[0176] Unless otherwise specified, the discussions in this specification using terms such as "processing", "computing", "calculating", "determining", "presenting", "displaying", etc. may refer to operations or processes of a machine (e.g., a computer) that manipulates or transforms data represented as a physical (e.g., electronic, magnetic, or optical) quantity within one or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other mechanical components that receive, store, transmit, or display information.

[0177] As used herein, any reference to "one aspect" or "aspect" means that a particular element, feature, structure, or characteristic described in conjunction with that aspect is included in at least one aspect. The appearances of the phrase "in one aspect" in various places in this specification do not necessarily all refer to the same aspect.

[0178] Some aspects may be described using the expressions "coupled" and "connected" along with their derivatives. For example, some aspects may be described using the term "coupled" to indicate that two or more elements are in direct physical or electrical contact. However, the term "coupled" may also mean that two or more elements are not in direct contact with each other but still cooperate or interact with each other. The aspects are not limited to this context.

[0179] As used herein, the terms "comprises," "comprising," "includes," "including," "has," "having," or any other variation thereof are intended to cover a non-exclusive inclusion. For example, a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to only those elements, but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Further, unless expressly stated to the contrary, "or" as used herein means "and / or." For example, the condition A or B is satisfied by any one of the following: A is true (or present) and B is false (or absent), A is false (or absent) and B is true (or present), and both A and B are true (or present).

[0180] In addition, the use of "a" or "an" is employed to describe elements and components of aspects of this specification. This is done merely for convenience and to give a general sense of the description. This description should be read to include one or at least one, and the singular also includes the plural unless it is obvious that the contrary is meant.

[0181] This detailed description is to be construed as merely illustrative, and not as limiting, since it is not possible, and even if it were not impracticable, to describe all possible embodiments. Many alternative embodiments can be implemented using any of the techniques developed after the filing date of this technology or this patent application.

Claims

1. 1. A computer-implemented method for stratifying a patient's cancer risk using molecular data, comprising: receiving, via one or more processors, molecular data corresponding to the patient; processing the molecular data using a machine learning model, via one or more processors, to determine a molecular data risk for the patient; the machine learning model is trained using a patient training dataset and / or a reference training dataset; the machine learning model uses univariate gene selection, RNA bias correction, and multivariate gene selection to filter and correct its training data; processing the machine learning model, wherein the machine learning model comprises a survival model; and generating a matched treatment strategy corresponding to the patient based on the molecular data risk of the patient.

2. The computer-implemented method of claim 1 , wherein the survival model is a Cox proportional hazards model.

3. The computer-implemented method of claim 1 , wherein the molecular data corresponding to the patient comprises RNA seq data.

4. 2. The computer-implemented method of claim 1, wherein the cancer is endometrial cancer and the machine learning model was trained on a cohort of patient data selected using a greedy algorithm.

5. The greedy algorithm, identifying a patient having a uterine subtype and a primary site being endometrial or uterine; Identifying progression-free eligible patients; and Identifying a patient having a sarcoma cancer; Identifying patients with serous and squamous cell carcinoma; and identifying patients having sarcoma cancer, serous cancer, and squamous cancer.

6. 10. The computer-implemented method of claim 1, wherein the patient has a pre-existing clinical risk group assignment of at least one of: low clinical risk, low to intermediate clinical risk, high to intermediate clinical risk, or high clinical risk.

7. 7. The computer-implemented method of claim 6, wherein generating the treatment strategy for the patient based on the patient's molecular data risk comprises generating the treatment strategy based on both (i) the pre-existing clinical risk group, and (ii) the patient's molecular risk.

8. 8. The computer-implemented method of claim 7, wherein the pre-existing clinical risk group is high to intermediate, the patient's molecular risk is high, and the treatment strategy is at least one of systemic therapy or external beam radiation therapy.

9. 8. The computer-implemented method of claim 7, wherein the pre-existing clinical risk group is high to intermediate, the patient's molecular risk is low, and the matched treatment strategy is observation.

10. 8. The computer-implemented method of claim 7, wherein the pre-existing clinical risk group is low to intermediate, the patient's molecular risk is high, and the matched treatment strategy is at least one of brachytherapy, external beam radiation therapy, or systemic therapy.

11. 8. The computer-implemented method of claim 7, wherein the pre-existing clinical risk group is high, the patient's molecular risk is low, and the matched treatment strategy is observation.

12. 8. The computer-implemented method of claim 7, wherein the pre-existing clinical risk group is high, the patient's molecular risk is low, and the matched treatment strategy is observation.

13. 8. The computer-implemented method of claim 7, wherein the pre-existing clinical risk group is high, the patient's molecular risk is high, and the matched treatment strategy is at least one of systemic therapy or external beam radiation therapy.

14. The computer-implemented method of claim 1 , wherein the matched treatment strategy comprises at least one of systemic therapy, external beam radiation therapy, brachytherapy, or observation.

15. 1. A computing system comprising: one or more processors; and one or more memories storing computer-executable instructions, the computer-executable instructions, when executed by the one or more processors, causing the computing system to: receiving molecular data corresponding to the patient; processing the molecular data using a machine learning model to determine a molecular data risk for the patient; The machine learning model is trained using a patient training dataset and a reference training dataset; the machine learning model uses univariate gene selection, RNA bias correction, and multivariate gene selection to filter and correct its training data; processing the machine learning model, wherein the machine learning model comprises a survival model; and generating a matched treatment strategy corresponding to the patient based on the molecular data risk of the patient.

16. 16. The computing system of claim 15, wherein the cancer is endometrial cancer and the machine learning model was trained on a cohort of patient data selected using a greedy algorithm.

17. The memory stores instructions that, when executed, cause the computing system to: identifying a patient having a uterine subtype and a primary site being endometrial or uterine; Identifying progression-free eligible patients; and Identifying a patient having a sarcoma cancer; Identifying patients with serous and squamous cell carcinoma; and identifying patients having sarcoma cancer, serous cancer, and squamous cancer.

18. 16. The computing system of claim 15, wherein the patient has a pre-existing clinical risk group assignment of at least one of: low clinical risk, low to intermediate clinical risk, high to intermediate clinical risk, or high clinical risk.

19. A computer-readable medium having stored thereon computer-executable instructions that, when executed by one or more processors, cause a computer to: receiving molecular data corresponding to the patient; processing the molecular data using a machine learning model to determine a molecular data risk for the patient; The machine learning model is trained using a patient training dataset and a reference training dataset; the machine learning model uses univariate gene selection, RNA bias correction, and multivariate gene selection to filter and correct its training data; processing the machine learning model, wherein the machine learning model comprises a survival model; and generating a matched treatment strategy corresponding to the patient based on the molecular data risk of the patient.

20. 20. The computer readable medium of claim 19, wherein the patient has a pre-existing clinical risk group assignment of at least one of: low clinical risk, low to intermediate clinical risk, high to intermediate clinical risk, or high clinical risk.