Methods and processes for predicting and analyzing response, progression, and survival of patient cohorts

The system addresses the challenge of analyzing vast patient data by using predictive modeling and outlier identification to improve treatment outcomes and patient survival rates.

JP7689494B2Active Publication Date: 2025-06-06テンパスエーアイインコーポレイテッド
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2021538761
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-12-31
Filing Date
2019-12-31
Publication Date
2025-06-06
Estimated Expiration
2039-12-31

AI Technical Summary

Technical Problem

Current medical technologies lack effective methods to quickly and comprehensively analyze vast amounts of patient data, including demographic, clinical, genomic, and treatment information, to predict patient response, progression, and survival.

Method used

A system and user interface that utilizes existing datasets to define patient cohorts and identify significant inflection points in patient attribute distributions, enabling predictive modeling and outlier identification to improve treatment outcomes.

Benefits of technology

Facilitates the discovery of therapeutic insights through automated analysis of patient data, enabling more informed treatment decisions and potentially improving patient survival rates and treatment responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007689494000005
    Figure 0007689494000005
  • Figure 0007689494000006
    Figure 0007689494000006
  • Figure 0007689494000007
    Figure 0007689494000007
Patent Text Reader

Abstract

Systems and methods are provided for analyzing data stores of de-identified patient data to generate one or more dynamic user interfaces that can be used to predict the likely response of a particular patient population or cohort when subjected to a particular treatment. Automated analysis of patterns emerging in patient clinical, molecular, phenotypic, and response data is facilitated by a variety of user interfaces, providing an efficient and intuitive way for clinicians to evaluate large data sets and aid in the potential discovery of insights of therapeutic significance.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to methods and processes for predicting and analysing response, progression and survival of patient cohorts. [Background technology]

[0002] In some medical fields, such as cancer research and treatment, vast amounts of data for each patient may be generated and collected. This data may include demographic information such as the patient's age, sex, height, weight, smoking history, geographic location, and other non-medical information. The data may also include clinical components such as tumor type, location, size, and stage of the tumor, as well as treatment data including medication, dosage, treatment, mortality, and other outcome / response data. In addition, more advanced analyses may also include genomic information about the patient and / or tumor, including genetic markers, mutations, and other information from fields including the proteome, transcriptome, epigenome, metabolome, microbiome, and other multi-omics fields.

[0003] Despite this abundance of data, there is a lack of meaningful ways to compile and analyze it quickly, efficiently, and comprehensively.

[0004] What is needed therefore are user interfaces, systems, and methods that overcome one or more of these challenges. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] U.S. Provisional Patent Application No. 62 / 746,997 [Patent Document 2] U.S. Patent Application Serial No. 16 / 289,027 [Patent Document 3] U.S. Patent No. 10,395,772 [Patent Document 4] PCT International Application No. PCT / US19 / 56713 [Patent Document 5] U.S. Patent Application Serial No. 16 / 679,054 Summary of the Invention [Means for solving the problem]

[0006] In one embodiment, a system and user interface is provided for predicting the expected response of a particular patient population or cohort when administered a particular treatment. To accomplish these predictions, the system uses an existing dataset to define a sample patient population, or "cohort," and identifies one or more significant inflection points in the distribution of patients exhibiting each attribute of interest in the cohort relative to the distribution of the general patient population, thereby targeting the expected survival and / or response predictions to the particular patient population.

[0007] The system described herein facilitates the discovery of insights of therapeutic significance through automated analysis of patterns emerging within patient clinical, molecular, phenotypic, and response data, and by enabling further investigation via a fully integrated, responsive user interface.

[0008] In one embodiment, the invention provides a method for identifying an outlier group of patients, comprising: 1) selecting a patient cohort comprising a plurality of patients; 2) calculating a mean survival rate for the patient cohort; 3) selecting a plurality of clinical or molecular features associated with the patient cohort; and 4) for each feature of the plurality of features, a) identifying a plurality of data values ​​associated with the feature; and b) for each data value of the plurality of data values ​​associated with the feature, i) dividing the patient cohort into a first subgroup and a second subgroup of the plurality of patients based on whether each patient of the plurality of patients survived during an outlier time period; ii) determining a difference between the number of patients in the first subgroup and the number of patients in the second subgroup; and iii) determining a difference between the number of patients in the first subgroup and the number of patients in the second subgroup. 5) creating a new node of the tree structure based on the data value that results in a maximum difference between the number of patients in the first subgroup and the number of patients in the second subgroup; 6) creating a first branch from the new node based on the first subgroup; 7) creating a second branch from the new node based on the second subgroup; 8) for each of the first and second branches, repeating steps 4) b) i-iii) and 5) based on the patients in the first and second subgroups, respectively, until a maximum number of nodes or branches have been created or a node includes less than a minimum number of patients; and 9) identifying at least one node including an outlier group of patients.

[0009] In yet another embodiment, the present invention provides a method for implementing a predictive model, comprising: receiving a plurality of data for a plurality of patients for a period of time; identifying a plurality of patient timepoints within the period of time for each of the plurality of patients; for each patient of the plurality of patients and for each patient timepoint of the plurality of patient timepoints, calculating an outcome target for an outcome event within a horizontal time window based on the plurality of data for the plurality of patients, identifying a plurality of prior features, and determining a state of each of the plurality of prior features at the patient timepoint; for each patient timepoint of the plurality of timepoints having a valid outcome target, and for each combination of the horizontal time window and outcome event, identifying a plurality of forward features; and generating a plurality of sets of predictions for the plurality of patients based on the plurality of prior features and the plurality of forward features.

[0010] In yet another embodiment, the invention provides a method, comprising receiving patient information for one or more patients; identifying one or more interactions for each of the one or more patients based at least in part on the received patient information; generating one or more timeline metrics for one or more targets in each of the one or more interactions that identify whether each of the one or more targets appears within a time period of occurrence of the interaction; identifying, for each timeline metric of the one or more timeline metrics, whether the patient may be subjected to one or more status characteristics within the time period; training a target prediction model for each of the one or more targets based at least in part on the one or more status characteristics; and associating a prediction for each patient from the target prediction model for each of the one or more targets with a respective one or more timeline metrics of the one or more timeline metrics.

[0011] In some embodiments, the method includes: 1) selecting a patient cohort including a patient group of a plurality of patients; 2) identifying a common anchor time point from a set of anchor points associated with each of the patient groups, the common anchor point being shared by each of the patient groups in the cohort; 3) for each patient of the patient group, aligning a timeline associated with each patient of the patient group to the common anchor point; 4) identifying an outcome target; 5) retrieving a generated plurality of sets of predictions each including a predicted target value for each patient of the patient group and for each of a plurality of forward features and a plurality of prior features; and 6) generating a plurality of decision trees, wherein for each decision tree of the plurality of decision trees: a) for each feature of the plurality of forward features and a plurality of prior features, i) subgrouping the group of patients into a first subgroup based on a difference between a predicted target value and an actual target value; ii) determining a difference between the number of patients in the first subgroup and the number of patients in the second subgroup; iii) selecting the feature that results in a difference that is the largest difference between the number of patients in the first subgroup and the number of patients in the second subgroup; 7) creating a new node of the tree structure based on the feature that results in the largest difference between the number of patients in the first subgroup and the number of patients in the second subgroup; 8) creating a first branch from the new node based on the first subgroup; 9) creating a second branch from the new node based on the second subgroup; and 10) for each of the first branch and the second branch, repeating steps 6) a) i-iii) and 7) based on the patients in the first subgroup and the second subgroup, respectively, until a maximum number of nodes or branches have been created or a node contains less than a minimum number of patients.

[0012] In other embodiments, the method may further include receiving a plurality of predictions, an outcome target, a subset of a plurality of forward features corresponding to the outcome target, and a cohort of patients comprising a subset of the plurality of patients; receiving anchor points; and for each patient in the cohort with the anchor points, providing a predictive model having a selected subset of the plurality of forward features and a difference between each of the plurality of predictions and the outcome target; and generating a decision tree based on determining, for each feature of the selected subset of the plurality of forward features, a maximum difference between each of the plurality of predictions and the outcome target, wherein the decision tree comprises a plurality of leaf nodes and one or more branch nodes, each of the one or more branch nodes comprising a pair of branches each of which comprises a leaf node or a branch node, and each of the plurality of leaf nodes of the decision tree comprising a number of patients from the cohort of patients.

[0013] The foregoing and other aspects and advantages of the present invention will become apparent from the following description. Reference is now made to the accompanying drawings, which form a part hereof, in which preferred embodiments of the invention are shown. However, such embodiments do not necessarily represent the full scope of the invention, and reference is therefore made to the claims herein for interpreting the scope of the invention.

[0014] Further objects, features, and advantages of the present disclosure will become apparent from the following detailed description taken in conjunction with the accompanying drawings illustrating exemplary embodiments of the present disclosure. [Brief description of the drawings]

[0015] [Figure 1] FIG. 1 is an exemplary system diagram of back-end and front-end components for predicting and analyzing response, progression, and survival of a patient cohort. [Diagram 2] FIG. 13 illustrates an example of a patient cohort selection filtering interface. [Diagram 3] FIG. 1 illustrates an example of a cohort funnel and population analysis user interface. [Figure 4]FIG. 1 illustrates another example of a cohort funnel and population analysis user interface. [Diagram 5] FIG. 1 illustrates another example of a cohort funnel and population analysis user interface. [Figure 6] FIG. 1 illustrates another example of a cohort funnel and population analysis user interface. [Figure 7] FIG. 1 illustrates another example of a cohort funnel and population analysis user interface. [Figure 8] FIG. 1 illustrates another example of a cohort funnel and population analysis user interface. [Figure 9] FIG. 1 illustrates another example of a cohort funnel and population analysis user interface. [Figure 10] FIG. 13 illustrates an example of a Data Summary window in a Patient Timeline Analysis user interface. [Figure 11] FIG. 13 illustrates another example of a data summary window in a patient timeline analysis user interface. [Figure 12] FIG. 13 illustrates another example of a data summary window in a patient timeline analysis user interface. [Figure 13] FIG. 13 illustrates another example of a data summary window in a patient timeline analysis user interface. [Figure 14] FIG. 13 illustrates another example of a data summary window in a patient timeline analysis user interface. [Figure 15] FIG. 1 illustrates an example of a patient survival analysis user interface. [Figure 16] FIG. 13 illustrates another example of a patient survival analysis user interface. [Figure 17] FIG. 13 illustrates another example of a patient survival analysis user interface. [Figure 18] FIG. 13 illustrates another example of a patient survival analysis user interface. [Figure 19] FIG. 13 illustrates another example of a patient survival analysis user interface. [Figure 20] FIG. 13 illustrates another example of a patient survival analysis user interface. [Figure 21] FIG. 13 illustrates an example of a patient event likelihood analysis user interface. [Figure 22] FIG. 13 illustrates another example of a patient event likelihood analysis user interface. [Diagram 23] FIG. 13 illustrates another example of a patient event likelihood analysis user interface. [Figure 24] FIG. 13 illustrates another example of a patient event likelihood analysis user interface. [Figure 25A] FIG. 13 illustrates an example of a binary decision tree for determining outliers usable with respect to a patient event likelihood analysis user interface. [Figure 25B] FIG. 13 illustrates an example of a binary decision tree for determining outliers usable with respect to a patient event likelihood analysis user interface. [Figure 26] FIG. 13 illustrates an example sample timeline of anchor events with associated exacerbation windows. [Figure 27A] FIG. 1 illustrates an example of adaptive feature ranking according to an embodiment of the SAFE algorithm. [Figure 27B] FIG. 1 illustrates an example of adaptive feature ranking according to an embodiment of the SAFE algorithm. [Figure 27C] FIG. 1 illustrates an example of handling correlated features by an embodiment of the SAFE algorithm. [Figure 27D] FIG. 1 illustrates an example of sample level importance assignment according to an embodiment of the SAFE algorithm. [Figure 27E] FIG. 1 illustrates an example of sample level importance assignment according to an embodiment of the SAFE algorithm. [Figure 28] FIG. 13 shows an example of using patient folds for cross-validation. [Figure 29]FIG. 1 illustrates an example of a user interface of an interactive analytics portal for generating analytics via one or more notebooks according to some embodiments. [Diagram 30] FIG. 1 illustrates an example workbook creation interface of the interactive analytics portal for creating a new workbook according to one embodiment. [Diagram 31] FIG. 13 illustrates an example of opening a pre-configured template from a custom workbook widget in the notebook user interface. [Diagram 32] 11A-11C are diagrams illustrating responses from the notebook user interface when a user drags a workbook into the display window. [Diagram 33] 13 illustrates an example cell edit view of a custom workbook after a user loads the workbook into the workbook editor and selects edit from the cell UIE. [Diagram 34] FIG. 1 is a block diagram illustration of one implementation of a computer system in which some implementations of the present disclosure may operate. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0016] With reference to the accompanying figures, and in particular with reference to Figure 1, a system 10 for predicting and analyzing response, progression, and survival of a patient cohort may include a back-end tier 12 including a patient data store 14 accessible by a patient cohort selector module 16 in communication with a patient cohort timeline data storage 18. The patient cohort selector module 16 interacts with a front-end tier 20 including an interactive analytic portal 22, which may be implemented, by way of example, via a web browser, to enable on-demand filtering and analysis of the data store 14.

[0017] The interactive analytics portal 22 may comprise a number of user interfaces, including an interactive cohort selection filtering interface 24 that allows a user to query and filter elements of the data store 14, as described in more detail below. As described in more detail below, the portal 22 may also include a cohort funnel and population analysis interface 26, a patient timeline analysis user interface 28, a patient survival analysis user interface 30, and a patient event likelihood analysis user interface 32. The portal 22 may further comprise a patient-based analysis user interface 34 and one or more patient future analysis user interfaces 36.

[0018] Referring again to FIG. 1, the backend tier 12 may also include a distributed computing and modeling tier 38 that receives data from the patient cohort timeline data storage 18 and provides input to a number of modules, including a time-to-event modeling module 40 that drives the patient survival analysis user interface 30, an event likelihood module 42 that calculates the likelihood of one or more events received in the patient event likelihood analysis user interface 32 for subsequent display in that user interface, a next event modeling module 44 that generates one or more models of next events for subsequent display in the patient next event analysis user interface 34, and one or more future modeling modules 46 that generate one or more future models for subsequent display in one or more patient future analysis user interfaces 36.

[0019] The patient data store 14 may be a pre-existing data set that includes patient medical history, such as demographics, comorbidities, diagnoses and recurrences, medications, surgeries, and other treatments, along with details of their responses and side effects. The patient data store may also contain details of patient genetic / molecular sequencing and gene mutations related to the patient, as well as the results of organoid modeling. In one embodiment, these data sets may be generated from one or more sources. For example, the institutions implementing the system may be able to pull from all of their records, and all records from all physicians and / or patients involved with the institutions may be available to agents, physicians, researchers, or other authorized members of the institutions. Similarly, physicians may pull from all of their records, for example, the records of all of their patients. Alternatively, some system users may be able to purchase or license aspects of a dataset when they do not have immediate access to a sufficiently robust dataset, when they are looking for even more records, and / or when they are looking for a particular type of data, such as data reflecting patients with several primary cancers, metastases by site of origin and / or site of diagnosis, recurrences by site of origin, metastases, or site of diagnosis, etc.

[0020] Features and Feature Modules The patient data store may include one or more feature modules that may comprise a collection of features available for all patients in the system 10. These features may be used to generate and model artificial intelligence classifiers in the system 10. While the feature range across all patients is informationally dense, a patient's Feature Set may be sparsely populated throughout the collective feature range of all features across all patients. For example, the feature range across all patients may extend to tens of thousands of features, but a patient's unique Feature Set may include only a subset of hundreds or thousands of features of the collective feature range based on the records available for that patient.

[0021] The feature collection may include a diverse set of fields available within a patient's health record. Clinical information may be based on fields entered into an electronic medical record (EMR) or electronic health record (EHR) by a physician, nurse, or other medical professional or representative. Other clinical information, such as molecular fields from gene sequencing reports, may be curated from other sources. Sequencing may include next generation sequencing (NGS) and may be long read, short read, or other forms of sequencing the patient's somatic cells and / or normal genome. A comprehensive collection of features within an additional feature module may combine together various features across various medical disciplines that may include diagnosis, response to treatment plans, genetic profiles, clinical and phenotypic features, and / or other medical, geographic, demographic, clinical, molecular, or genetic features. For example, a subset of features may include molecular data features, such as features derived from RNA feature module or DNA feature module sequencing.

[0022] Another subset of features, imaging features from the imaging feature module, may include features identified through review of specimens by pathologists, such as review of stained H&E or IHC slides. As another example, the subset of features may include derived features obtained from analysis of the individual and combined results of such feature sets. Features derived from DNA and RNA sequencing may include genetic variants from the variant science module present in the sequenced tissue. Further analysis of genetic variants may include additional steps such as identifying single or multi-nucleotide polymorphisms, identifying whether the mutation is an insertion or deletion event, identifying loss or gain of function, identifying fusions, calculating copy number changes, calculating microsatellite instability, calculating tumor gene mutation burden, or other structural changes within DNA and RNA. Analysis of slides for H&E or IHC staining may reveal features such as tumor infiltration, programmed cell death ligand 1 (PD-L1) status, human leukocyte antigen (HLA) status, or other immunological features.

[0023] Features derived from structured, curated, or electronic medical or health records may include clinical features such as diagnosis, symptoms, treatment, outcome; patient name, date of birth, sex, ethnicity, date of death, address, smoking status, patient demographics such as date of diagnosis of cancer, disease, illness, diabetes, depression, other physical or mental illness; medical history, family medical history, clinical diagnosis such as date of first consultation, date of metastasis diagnosis; cancer stage, tumor characterization, tissue of origin, course of therapy, treatment group, clinical trial, medications prescribed or taken, surgery, radiation therapy, imaging, side effects, associated outcomes; genetic testing and laboratory information such as performance score, laboratory tests, pathology results, prognostic indicators; genetic testing date, testing provider used, testing method used such as gene sequencing method or gene panel, genes included, variants, genetic results such as expression levels / status, or dates corresponding to any of the above.

[0024] Features may be derived from information from additional medical or research-based omics fields, including proteome, transcriptome, epigenome, metabolome, microbiome, and other multi-omics fields. Features derived from organoid modeling labs may include DNA and RNA sequencing information closely related to each organoid, and results from treatments applied to those organoids. Features derived from imaging data may further include reports associated with stained slides, tumor size, tumor size difference over time, including treatment during change, and even machine learning approaches to classify PDL1 status, HLA status, or other characteristics from imaging data. Other features may include additional derived feature sets from other machine learning approaches based at least in part on any new features and / or combinations of the above features. For example, imaging results may need to be combined with MSI calculations derived on RNA expression to determine additional further imaging features. In another example, the machine learning model may generate the likelihood that a patient's cancer will metastasize to a particular organ, or the patient's future probability of metastasis to yet another organ in the body. Other features that can be extracted from medical information may also be used. There are thousands of features, and the above listing of feature types is merely representative and should not be construed as an exhaustive listing of features.

[0025] A variation module may be one or more microservices, servers, scripts, or other executable algorithms that generate variation features associated with de-identified patient features from a feature collection. A variation module may provide variation that takes input from a feature collection and stores it. An exemplary variation module may include one or more of the following variations as a collection of variation modules: A SNP (single nucleotide polymorphism) module may identify single base substitutions that occur at specific positions in the genome, with each variation occurring with some degree of prominence (e.g., >1%) in the population. For example, at a specific base position, or locus, in the human genome, a C-nucleotide may occur in most people, but an A-nucleotide may occur in a minority of people. This means that there is a SNP at this specific position, and the two possible nucleotide changes, C or A, are said to be allelic for this position. SNPs reveal differences in our vulnerability to a wide range of diseases (e.g., sickle cell anemia, β-thalassemia, cystic fibrosis are caused by SNPs). The severity of disease and the way the body responds to treatment are also manifestations of genetic variation. For example, single base mutations in the APOE (apolipoprotein E) gene are associated with a lower risk of Alzheimer's disease. Single nucleotide variants (SNVs) are single nucleotide mutations with no frequency restriction that can occur somatically. Somatic single nucleotide mutations (e.g., caused by cancer) are sometimes called single nucleotide changes. The MNP (multiple nucleotide polymorphism) module can identify the substitution of consecutive nucleotides at specific positions in the genome. The InDels module can identify the insertion or deletion of bases in the genome of an organism, which are classified among small genetic variations. Microindels are defined as indels that result in a net change of 1 to 50 nucleotides, although they usually measure 1 to 10,000 base pairs in length. Indels can be contrasted with SNPs or point mutations. Indels insert and delete nucleotides from a sequence, whereas point mutations are a form of substitution that replaces one of the nucleotides without changing the total number in the DNA.Indels, either insertions or deletions, can be used as genetic markers of natural populations, especially in phylogenetic studies. Indel frequencies tend to be significantly lower than those of single nucleotide polymorphisms (SNPs), except in close highly repetitive regions, including homopolymers and microsatellites. The MSI (microsatellite instability) module can identify genetic hypermutability resulting from DNA mismatch repair gene (MMR) defects. The presence of MSI provides phenotypic evidence that MMR is not functioning properly. MMR corrects errors that occur naturally during DNA replication, such as single-base mismatches or short insertions and deletions. Proteins involved in MMR correct polymerase errors by forming a complex that binds to mismatched sections of DNA, deleting the error, and inserting the correct sequence in its place. Cells with dysfunctional MMR are unable to correct errors that occur during DNA replication and, as a result, accumulate errors. This causes the creation of new microsatellite fragments. Polymerase chain reaction-based assays can reveal these new microsatellites and provide evidence of the presence of MSI. Microsatellites are repetitive sequences of DNA. These sequences can be made up of repeating units of one to six base pairs in length. Although the length of these microsatellites varies greatly from person to person, contributing to an individual's DNA "fingerprinting," each individual has a set length of microsatellites. The most common microsatellites in humans are dinucleotide repeats consisting of C- and A-nucleotides that occur tens of thousands of times in the genome. Microsatellites are also called simple sequence repeats (SSRs). The tumor mutation burden (TMB) module is a predictive biomarker that may identify a measure of mutations carried by tumor cells and is being investigated to assess its association with response to cancer immunotherapy (IO) therapy. Tumor cells with high TMB may harbor more neoantigens and thus increase anticancer T cells within the tumor microenvironment and periphery. These neoantigens may be recognized by T cells and trigger an antitumor response.In recent years, TMB has emerged as a quantitative marker that can help predict potential immunotherapy response among various cancers, including melanoma, lung cancer, and bladder cancer. TMB is defined as the total number of mutations per coding region of the tumor genome. Importantly, TMB is consistent and reproducible. This provides a quantitative measure that can be used to appropriately inform treatment decisions, such as the selection of targeted or immunotherapy, and enrollment in clinical trials. CNV (copy number change) modules may identify deviations from the normal genome, and their subsequent effects, from analyzing genes, variants, alleles, or nucleotide sequences. CNV is a phenomenon in which structural variations can occur in sections of nucleotides or base pairs that contain repeats, deletions, or inversions. Fusion modules may identify hybrid genes that form from two previously separate genes. This can occur as a result of translocations, interstitial deletions, or chromosomal inversions. Gene fusions play an important role in tumorigenesis. Fusion genes can contribute to tumorigenesis because they can produce abnormal proteins that are significantly more active than non-fused genes. Fusion genes are often cancer-causing oncogenes, including BCR-ABL, TEL-AML1 (all with t(12;21)), AML1-ETO (M2 AML with t(8;21)), and TMPRSS2-ERG with an interstitial deletion on chromosome 21, which often appears in prostate cancer. In the case of TMPRSS2-ERG, the fusion product regulates prostate cancer by inhibiting androgen receptor (AR) signaling and suppressing AR expression by oncogenic ETS transcription factors. Most fusion genes are found in hematological cancers, sarcomas, and prostate cancer. BCAM-AKT2 is a unique fusion gene specific to high-grade serous ovarian cancer. Oncogenic fusion genes may generate gene products with new or different functions from the two fusion partners. Alternatively, proto-oncogenes are fused to strong promoters, thereby setting the oncogenic function to function by upregulation caused by the strong promoter of the upstream fusion partner. The latter is common in lymphomas, where an oncogene is juxtaposed to the promoter of an immunoglobulin gene.Oncogenic fusion transcripts can also be caused by trans-splicing or read-through events. Because chromosome translocations play such an important role in neoplasia, a specialized database of chromosome aberrations and gene fusions in cancer has been created. This database is called the Mitelman Database of Chromosome Aberrations and Gene Fusions in Cancer. The IHC (Immunohistochemistry) module can identify antigens (proteins) in cells of tissue sections by utilizing the principle of antibodies specifically binding to antigens in living tissue. IHC staining is widely used in the diagnosis of abnormal cells such as those found in cancerous tumors. Certain molecular markers are characteristic of certain cellular events such as proliferation and cell death (apoptosis). IHC is also widely used in basic research to understand the distribution and localization of biomarkers and proteins that are differentially expressed in different parts of living tissue. Visualizing antibody-antigen interactions can be accomplished in a number of ways. In the most common case, antibodies are conjugated to enzymes such as peroxidase, which can catalyze a color-producing reaction during immunoperoxidase staining. Alternatively, antibodies may also be tagged with fluorophores such as fluorescein or rhodamine in immunofluorescence. Approximations from RNA expression data, H&E slide imaging data, or other data may be generated. The therapy module may identify differences in the cancer cells (or other cells near the cancer cells) that may aid in growth and proliferation, and drugs that "target" these differences. Treatment with these drugs is called targeted therapy. For example, many targeted drugs aim at the internal "programming" of cancer cells that makes them different from normal healthy cells while leaving most healthy cells untouched. Targeted drugs may block or turn off chemical signals that tell cancer cells to grow and divide, change proteins in cancer cells so that they die, stop making new blood vessels that feed the cancer cells, trigger the immune system to kill the cancer cells, or deliver toxins to cancer cells to kill them but not to normal cells. Some targeted drugs are more "targeted" than others.Some target only a single change in cancer cells, others may affect several different changes. Others boost the way the body fights cancer cells. This may affect where these drugs act and what side effects they cause. Matching targeted therapies may involve identifying the patient's therapy target and meeting other inclusion or exclusion criteria. The VUS (Variant of Unknown Clinical Significance) module may identify variants that are called but cannot be classified as pathogenic or benign at the time of the call. VUS may be cataloged from publications on VUS to identify whether they can be classified as benign or pathogenic. The testing module may identify and test hypotheses for treating cancers with specific characteristics by matching patient characteristics with clinical trials. These trials have inclusion and exclusion criteria that must be matched with entries that can be populated and structured from publications, test reports, or other documents. The amplification module may identify genes that are disproportionately inflated in counts relative to other genes. Amplification can cause high-count genes to become dormant, overactive, or behave in another unexpected way. Amplification can be detected at the gene level, variant level, RNA transcription or expression level, or even protein level. Detection can be performed at all different detection mechanisms or levels and validated against each other. Isoform modules can identify alternative splicing (AS), a biological process in which multiple mRNAs (isoforms) are generated from the same gene transcript through different combinations of exons and introns. Large-scale genomics studies estimate that 30-60% of mammalian genes are alternatively spliced. The possible patterns of alternative splicing for a gene can be highly complex, and the complexity increases exponentially as the number of introns in the gene increases.Alternative splicing prediction in silico can identify genomic loci through searching mRNA sequences against genomic sequences, extract sequences against genomic loci to extend sequences on both ends up to 20kb, search genomic sequences (repeated sequences are masked), extract splicing pairs (two boundaries of an alignment gap that have a GT-AG consensus or have more than two expressed sequence tags aligned on both ends of the gap), assemble splicing pairs according to their coordinates, determine gene boundaries (splice pair predictions are generated to this point), generate predicted gene structures by aligning mRNA sequences to genomic templates, and find alternative splice isoforms by comparing splice pair predictions with gene structure predictions to find large insertions or deletions within a set of mRNAs that share most aligned sequences. Pathway modules can identify defects in DNA repair pathways that allow cancer cells to accumulate genomic alterations that contribute to their malignant phenotype. Cancerous tumors rely on a residual DNA repair capacity to survive damage induced by genotoxic stress, which leads to the isolation and inactivation of DNA repair pathways in cancer cells. DNA repair pathways are generally considered to be mutually exclusive units of machinery that handle different types of damage at different cell cycle stages. However, recent preclinical studies have provided strong evidence that multifunctional DNA repair hubs, which involve multiple traditional DNA repair pathways, are frequently altered in cancer. Identifying pathways that may be affected may lead to important considerations regarding patient treatment. The Raw Count module may identify counts of variants detected from sequencing data. For DNA, this may be the number of reads from sequencing that correspond to a particular variant in a gene. For RNA, this may be gene expression counts or transcriptome counts from sequencing.

[0026] Structural variant classification may include evaluating features from the feature collection, changes from the change module, and other classifications from among itself from one or more classification modules. In structural variant classification, a classification may be provided to the stored classification storage. An exemplary classification module may include a classification of CNVs, where "reportable" means that the CNV has been identified in one or more reference databases as affecting tumor cancer characterization, disease state, or pharmacogenomics, "non-reportable" means that the CNV has not been identified as such, and "conflicting evidence" may mean that the CNV has both evidence suggesting "reportable" and "non-reportable". Furthermore, the classification of therapeutic relevance is similarly confirmed from the reference dataset reference of therapies that may be affected by the detection (or non-detection) of the CNV. Other classifications may include the application of machine learning algorithms, neural networks, regression methods, graph techniques, inductive reasoning approaches, or other artificial intelligence evaluations in the module. A classifier for clinical trials may include evaluating variants identified from alteration modules that have been identified as significant or reportable, evaluating all available clinical trials to identify inclusion and exclusion criteria, mapping the patient's variants and other information to the inclusion and exclusion criteria, and classifying clinical trials as applicable or inapplicable to the patient. Similar classifications may be performed for therapeutic, loss-of-function, gain-of-function, diagnostic, microsatellite instability, tumor mutation burden, indels, SNPs, MNPs, fusions, and other alterations that may be classified based on the results of alteration modules.

[0027] Each of the feature collections, variation modules, structural variants, and feature store may be communicatively coupled to a data bus to transfer data between each module for processing and / or storage. In another embodiment, each of the feature collections, variation modules, structural variants, and feature store may be communicatively coupled to one another for independent communication without sharing a data bus.

[0028] In addition to the above features and listed modules, the feature modules may further include one or more of the following modules within their respective modules, either as sub-modules or as stand-alone modules:

[0029] The germline / somatic DNA feature module may include feature collections associated with information derived from the DNA of the patient or the patient's tumor. These features may include raw sequencing results, genes, mutations, variant calls, and variant characterizations, as stored in FASTQ, BAM, VCF, or other sequencing file types known in the art. Genomic information from the patient's normal sample may be stored as germline, and genomic information from the patient's tumor sample may be stored as somatic.

[0030] The RNA feature module may include feature collections that are associated with information derived from the patient's DNA, such as transcriptome information. These features may include raw sequencing results, transcriptome expression, genes, mutations, variant calls, and variant characterizations.

[0031] The metadata module may include feature collections related to the human genome, protein structures, and effects such as changes in energy stability based on protein structure.

[0032] The clinical module may include feature collections associated with information derived from the patient's clinical records and records from the patient's family. These may be extracted from unstructured clinical documents, EMRs, EHRs, or other sources of patient history. Information may include the patient's symptoms, diagnoses, treatments, medications, therapies, hospice, response to treatments, laboratory test results, medical history, respective geographic location, demographics, or other characteristics of the patient that may be found in the patient's medical record. Information regarding treatments, medications, therapies, and the like may be captured as recommendations or prescriptions and / or confirmation that such treatments, medications, therapies, and the like have been administered or taken.

[0033] The imaging module may include feature collections associated with information derived from a patient's imaging record. The imaging record may include H&E slides, IHC slides, radiology images, and other medical images that may be ordered by a physician in the course of diagnosing and treating various illnesses and diseases. These features may include TMB, ploidy, purity, nuclear-cytoplasmic ratio, macronuclei, cell state changes, biological pathway activation, hormone receptor changes, immune cell infiltration, immune biomarkers of MMR, MSI, PDL1, CD3, FOXP3, HRD, PTEN, PIK3CA, collagen or stromal composition, appearance, density, or characteristics, tumor budding, size, grade, metastasis, immune status, chromatin morphology, and other characteristics of cells, tissues, or tumors for prognostic purposes.

[0034] Epigenomic modules, such as epigenomic modules from omics, may contain feature collections associated with information derived from modifications of DNA that regulate the expression of genes, rather than changes in DNA sequence. These modifications are often the result of environmental factors based on what the patient breathes, eats, or drinks. These features may include DNA methylation, histone modifications, or other factors that inactivate genes or cause changes to gene function without changing the sequence of nucleotides within the gene.

[0035] A microbiome module, such as a microbiome module from omics, may include a collection of features associated with information derived from a patient's viruses and bacteria. These features may include viral infections that may affect the treatment and diagnosis of some diseases, as well as bacteria present in the patient's gastrointestinal tract that may affect the effectiveness of medicines taken by the patient.

[0036] A proteome module, such as a proteome module from omics, may include a collection of features associated with information derived from proteins produced in a patient. These features may include the composition, structure, and activity of a protein, when and where the protein is expressed, the protein's production rate, degradation rate, and steady-state abundance, how the protein is modified, for example, post-translational modifications such as phosphorylation, trafficking of the protein between subcellular compartments, participation of the protein in metabolic pathways, protein-protein interactions, or modifications of the protein after it is translated from RNA, such as phosphorylation, ubiquitination, methylation, acetylation, glycosylation, oxidation, or nitrosylation.

[0037] Additional omics modules are also often included in omics, such as feature collections associated with all the different fields of omics, which are collections of features that include the study of changes in cognitive processes associated with genetic profiles; cognitive genomics, which are collections of features that include the study of genomic structure and function relationships between different biological species or strains; comparative genomics, which are collections of features that include the study of gene and protein function and interactions, including transcriptomics; functional genomics, which are collections of features that include studies concerned with the large-scale analysis of gene-gene, protein-protein, or protein-ligand interactions; interactomics, which are collections of features that include the study of metagenomics, such as genetic material recovered directly from environmental samples; metagenomics, which are collections of features that include the study of genetic influences on nervous system development and function; neurogenomics, which are collections of features that include the study of the entire collection of gene families found within a given species; and pangenomics, which are collections of features that include the study of an individual's genotype after the genotype is known and made public. Genomics is a collection of features that involves the study of an individual's genome, including the sequencing and analysis of the individual's genome; personal genomics is a collection of features that involves the study of supporting genome structure, including protein and RNA binders, alternative DNA structures, and chemical modifications on DNA; epigenomics is a collection of features that involves the study of the suite of genomic components that form the cell nucleus as a complex dynamic biological system; nucleomics is a collection of features that involves the study of cellular lipids, including modifications made to specific lipid groups produced by the patient; lipidomics is a collection of features that involves the study of proteins, including modifications made to specific proteins produced by the patient; proteomics is a collection of features that involves the study of large sets of proteins involved in the immune response; immunoproteomics is a collection of features that involves the study of large sets of proteins involved in the immune response;Nutriproteomics, a collection of features that includes research into identifying molecular targets of nutritional and non-nutritional components of the diet, including the use of proteomic mass spectrometry data for protein expression studies; Nutriproteomics, a collection of features that includes research into biological studies at the intersection of proteomics and genomics, including data that identifies gene annotations; Proteogenomics, a collection of features that includes the study of the three-dimensional structure of all proteins encoded by a given genome using a combination of modeling approaches; Structural genomics, a collection of features that includes the study of sugars and carbohydrates and their effects in patients; Glycomics, a collection of features that includes the study of the intersection between the food and nutrition domains through the application and integration of technologies to improve consumer well-being, health, and knowledge; Foodomics, a collection of features that includes the study of RNA molecules, including mRNA, rRNA, tRNA, and other non-coding RNA produced within cells; Transcriptomics, a collection of features that includes the study of chemical processes involving metabolic products or the solids that certain cellular processes leave behind. metabolomics, a collection of features that involves the study of quantitative measurements of the dynamic multi-parametric metabolic response of cells to pathophysiological stimuli or genetic modifications; metabonomics, a collection of features that involves the study of genetic variations in the interaction between diet and health with associations with susceptible subgroups; nutrigenetics, a collection of features that involves the study of changes in cognitive processes associated with genetic profiles; cognitive genomics, a collection of features that involves the study of the effect of the sum of variations in the human genome on drugs; pharmacogenomics, a collection of features that involves the study of the effect of variations in the human microbiome on drugs; pharmacomicrobiomics, a collection of features that involves the study of the activity of genes and proteins in specific cells or tissues of an organism in response to toxic substances; toxicogenomics, a collection of features that involves the study of the processes used by mitochondrial proteins to interact; mitointeractome,Psychogenomics, a collection of features that includes the study of processes that apply the powerful tools of genomics and proteomics to better understand the biological substrates of normal behavior and diseases of the brain that manifest as behavioral disorders; Psychogenomics, a collection of features that includes the application of psychogenomics to the study of drug addiction to develop more effective treatments for these diseases as well as objective diagnostic tools, preventative measures, and cures; Psychogenomics, a collection of features that includes the study of stem cell biology that establishes stem cells as a model system for understanding human biology and disease states; Stem Cell Genomics, a collection of features that includes the study of neural connections in the brain; Connectomics, a collection of features that includes the study of the genomes of the microbial population that live in the digestive tract. microbiomics, a collection of features including the study of quantitative cell analysis and the use of bioimaging methods and bioinformatics; cellomics, a collection of features including the study of tomography and omics methods to understand tissue or cellular biochemistry at high spatial resolution from imaging mass spectrometry data; tomomics, a collection of features including the study of high-throughput machine measurements of patient behavior; ethomics, a collection of features including the study of video analysis paradigms inspired by genomics principles, including videoomics, where a continuous image sequence, i.e., a video, can be interpreted as the capture of a single image unfolding over a time course of mutations that reveals insights into the patient.

[0038] Feature sets for DNA-related (molecular) features may include a unique calculation of the maximum effect a gene may have from the sequencing results of the gene, and these genes may include ABCB1-somatic, ACTA2-germline, ACTC1-germline, ALK-fluorescence_in_situ_hybridization_(fish), ALK-immunohistochemistry_(ihc), ALK-md_dictated, ALK-somatic, AMER1-somatic, APC-gene_mutation_analysis, APC-germline, APC-somatic, APOB-germline, APOB-somatic, AR-somatic, ARHGAP35-somatic, ARID1A-somatic, ARID1B-somatic, ARID2-somatic, ASXL1-somatic, ATM-gene_mutation_analysis, ATM-germline, ATM-somatic, ATP7B-germline, ATR-somatic, ATRX-somatic, AXI N2-germline, BACH1-germline, BCL11B-somatic, BCLAF1-somatic, BCOR-somatic, BCORL1-somatic, BCR-somatic, BMPR1A-germline, BRAF-gene_mu tation_analysis, BRAF-md_dictated, BRAF-somatic, BRCA1-germline, BRCA1-somatic, BRCA2-germline, BRCA2-somatic, BRD4-somatic, BRIP1-ge rmline, CACNA1S-germline, CARD11-somatic, CASR-somatic, CD274-immunohistochemistry_(ihc), CD274-md_dictated, CDH1-germline, CDH1-som atic, CDK12-germline, CDKN2A-immunohistochemistry_(ihc), CDKN2A-germline, CDKN2A-somatic, CEBPA-germline, CEBPA-somatic, CFTR-somatic、CHD2-somatic、CHD4-somatic、CHEK2-germline、CIC-somatic、COL3A1-germline、CREBBP-somatic、CTNNB1-somatic、CUX1-somatic、DICER1-somatic、DOT1L-somatic、DPYD-somatic、DSC2-germline、DSG2-germline、DSP-germline、DYNC2H1-somatic、EGFR-gene_mutation_analysis、EGFR-immunohistochemistry_(ihc)、EGFR-md_dictated、EGFR-germline、EGFR-somatic、EP300-somatic、EPCAM-germline、EPHA2-somatic、EPHA7-somatic、EPHB1-somatic、ERBB2-fluorescence_in_situ_hybridization_(fish)、ERBB2-immunohistochemistry_(ihc)、ERBB2-md_dictated、ERBB2-somatic、ERBB3-somatic、ERBB4-somatic、ESR1-immunohistochemistry_(ihc)、ESR1-somatic、ETV6-germline、FANCA-germline、FANCA-somatic、FANCD2-germline、FANCI-germline、FANCL-germline、FANCM-somatic、FAT1-somatic、FBN1-germline、FBXW7-somatic、FGFR3-somatic、FH-germline、FLCN-germline、FLG-somatic、FLT1-somatic、FLT4-somatic、GATA2-germline、GATA3-somatic、GATA4-somatic、GATA6-somatic、GLA-germline、GNAS-somatic、GRIN2A-somatic、GRM3-somatic、HDAC4-somatic、HGF-somatic、IDH1-somatic、IKZF1-somatic、IRS2-somatic、JAK3-somatic、KCNH2-germline、KCNQ1-germline、KDM5A-somatic、KDM5C-somatic、KDM6A-somatic、KDR-somatic、KEAP1-somatic、KEL-somatic、KIF1B-somatic、KMT2A-fluorescence_in_situ_hybridization_(fish)、KMT2A-somatic、KMT2B-somatic、KMT2C-somatic、KMT2D-somatic、KRAS-gene_mutation_analysis、KRAS-md_dictated、KRAS-somatic、LDLR-germline、LMNA-germline、LRP1B-somatic、MAP3K1-somatic、MED12-somatic、MEN1-germline、MET-fluorescence_in_situ_hybridization_(fish)、MET-somatic、MKI67-immunohistochemistry_(ihc)、MKI67-somatic、MLH1-germline、MSH2-germline、MSH3-germline、MSH6-germline、MSH6-somatic、MTOR-somatic、MUTYH-germline、MYBPC3-germline、MYCN-somatic、MYH11-germline、MYH11-somatic、MYH7-germline、MYL2-germline、MYL3-germline、NBN-germline、NCOR1-somatic、NCOR2-somatic、NF1-somatic、NF2-germline、NOTCH1-somatic、NOTCH2-somatic、NOTCH3-somatic、NRG1-somatic、NSD1-somatic、NTRK1-somatic、NTRK3-somatic、NUP98-somatic、OTC-germline、PALB2-germline、PALLD-somatic、PBRM1-somatic、PCSK9-germline、PDGFRA-somatic、PDGFRB-somatic、PGR-immunohistochemistry_(ihc)、PIK3C2B-somatic、PIK3CA-somatic、PIK3CG-somatic、PIK3R1-somatic、PIK3R2-somatic、PKP2-germline、PLCG2-somatic、PML-somatic、PMS2-germline、POLD1-germline、POLD1-somatic、POLE-germline、POLE-somatic、PREX2-somatic、PRKAG2-germline、PTCH1-somatic、PTEN-fluorescence_in_situ_hybridization_(fish)、PTEN-gene_mutation_analysis、PTEN-germline、PTEN-somatic、PTPN13-somatic、PTPRD-somatic、RAD51B-germline、RAD51C-germline、RAD51D-germline、RAD52-germline、RAD54L-germline、RANBP2-somatic、RB1-germline、RB1-somatic、RBM10-somatic、RECQL4-somatic、RET-fluorescence_in_situ_hybridization_(fish)、RET-germline、RET-somatic、RICTOR-somatic、RNF43-somatic、ROS1-fluorescence_in_situ_hybridization_(fish)、ROS1-md_dictated、ROS1-somatic、RPTOR-somatic、RUNX1-germline、RUNX1T1-somatic、RYR1-germline、RYR2-germline、SCN5A-germline、SDHAF2-germline、SDHB-germline、SDHC-germline、SDHD-germline、SETBP1-somatic、SETD2-somatic、SH2B3-somatic、SLIT2-somatic、SLX4-somatic、SMAD3-germline、SMAD4-germline、SMAD4-somatic、SMARCA4-somatic、SOX9-somatic、SPEN-somatic、STAG2-somatic、STK11-gene_mutation_analysis、STK11-germline, STK11-somatic, TAF1-somatic, TBX3-somatic, TCF7L2-somatic, TERT-somatic, TET2-somatic, TGFBR1-germline, TGFBR2-germline, TGFBR2-somatic, TMEM43-germline, TNNI3-germline, TNNT2-germline, TP53-gene_mutation_analysis, TP53-immunohistochemistry_(ihc), TP53-md_dictated, TP53-germline, TP53-somatic, TPM1-germline, TSC1-germline, TSC1 -somatic, TSC2-germline, TSC2-somatic, VHL-germline, WT1-germline, WT1-somatic, XRCC3-germline, and ZFHX3-somatic. ,

[0039] A sufficiently robust collection of features may include all of the features disclosed above, however, models and predictions based on available features may include models optimized and trained from a selection of features that is significantly more restricted than the exhaustive feature set. Such constrained feature sets may include tens to hundreds of features. For example, the constrained feature set of the model may include the genomic results of sequencing the patient's tumor, derived features based on the genomic results, the origin of the patient's tumor, the patient's age at diagnosis, the patient's gender and race, and symptoms the patient presented to the physician during a routine checkup.

[0040] The feature store may augment a patient's feature set by applying machine learning and analytics by selecting from any features, alterations, or computed outputs derived from the patient's features or alterations in those features. Such feature stores may generate new features from the original features found in the feature module, or may identify and store key insights or analyses based on the features. The selection of features may be based on the alterations or calculations to be generated and may include genomic single or multi-nucleotide polymorphism insertions or deletions, tumor mutation burden, microsatellite instability, copy number alterations, fusions, or other such calculations. Exemplary outputs of alterations or calculations generated that may inform future alterations or calculations include findings of hypertrophic cardiomyopathy (HCM) and variants in MYH7. Previously classified variants may be identified in the patient's genome that may inform the classification of novel variants or indicate further risk of disease. An exemplary approach may include enriching the variants and their respective classifications to identify regions in MYH7 that are associated with HCM. Any novel variants detected from the patient's sequencing that are localized to this region would increase the patient's risk for HCM. Features that can be exploited for such change detection include the structure of MYH7 and the classification of variants therein. Models focused on enrichment can separate such variants.

[0041] Artificial Intelligence Model The artificial intelligence models referred to herein may be gradient boosting models, random forest models, neural networks (NNs), regression models, naive Bayes models, or machine learning algorithms (MLAs). The MLAs or NNs may be trained from a training dataset. In an exemplary predictive profile, the training dataset may include patient imaging, pathology, clinical, and / or molecular reports and details, such as those curated from EHRs or gene sequencing reports. MLA includes supervised algorithms (e.g., algorithms where features / classifications in the dataset are annotated) using linear regression, logistic regression, decision trees, classification and regression trees, naive Bayes, nearest neighbor clustering, unsupervised algorithms (e.g., algorithms where features / classifications in the dataset are not annotated) using a priori, means clustering, principal component analysis, random forests, adaptive boosting, semi-supervised algorithms (e.g., algorithms where features / classifications in the dataset are not annotated) using generative approaches (e.g., mixtures of Gaussians, mixtures of multinomial distributions, hidden Markov models), sparse separation, graph-based approaches (e.g., mincut, harmonic functions, manifold regularization), heuristic approaches, or support vector machines. NN includes conditional random fields, convolutional neural networks, attention-based neural networks, deep learning, long short-term memory networks, or other neural models, and the training dataset includes pathology reports covering multiple tumor samples, RNA expression data for each sample, and imaging data for each sample. Although MLA and neural network identify different approaches of machine learning, these terms may be used interchangeably herein.Thus, unless otherwise specified, a reference to MLA may include a corresponding NN, or a reference to NN may include a corresponding MLA.Training may include providing an optimized data set, labeling these features as they appear in patient records, and training the MLA to predict or classify based on new inputs.Artificial NNs are efficient computational models and have shown strength in solving difficult problems in artificial intelligence. They have also been shown to be universal approximators (capable of expressing a wide range of functions when given the appropriate parameters). Some MLAs identify important features and identify coefficients, or weights, for them. The coefficients are multiplied with the frequency of occurrence of the feature to generate a score, and some classifications may be predicted by the MLA when the score of one or more features exceeds a threshold. The coefficient schema may be combined with a rule-based schema to generate more complex predictions, such as predictions based on multiple features. For example, 10 important features may be identified in different classifications. A list of coefficients may exist for the important features and a rule set may exist for the classification. The rule set may be based on the number of occurrences of the feature, the scaled weights of the feature, or other qualitative and quantitative evaluation of the features coded in logic known to those skilled in the art. In other MLAs, the features may be organized in a binary tree structure. For example, a key feature that distinguishes most classifications may exist as the root of a binary tree and as each subsequent branch in the tree until a classification can be given based on reaching a terminal node of the tree. For example, a binary tree may have a root node that tests a first feature. The occurrence or non-occurrence of this feature must be present (a binary decision) and the logic may traverse the branch that is true for the item being classified. Additional rules may be based on thresholds, ranges, or other qualitative and quantitative tests. Supervised methods are useful when the training data set has many known values ​​or annotations, but due to the nature of EMR / EHR documents, many annotations may not be given. When exploring large amounts of unlabeled data, unsupervised methods are beneficial for binning / bucketing instances in the data set. A single instance of the above model, or two or more such instances combined, may constitute a model for the purposes of the model, artificial intelligence, neural network, or machine learning algorithm herein.

[0042] A series of conversion steps may be performed to convert data from the patient data store into a format suitable for analysis. A variety of modern machine learning algorithms may be utilized to train models directed to predicting expected survival and / or response for a particular patient population. Exemplary data store 14 is described in further detail in U.S. Provisional Patent Application No. 62 / 746,997, filed October 17, 2018, entitled "Data Based Cancer Research and Treatment Systems and Methods," U.S. Patent Application No. 16 / 289,027, filed February 28, 2019, and issued August 27, 2019 as U.S. Patent No. 10,395,772, and PCT International Application No. PCT / US19 / 56713, filed October 17, 2019, entitled "Data Based Cancer Research and Treatment Systems and Methods," each of which is incorporated herein by reference in its entirety.

[0043] The system may include a data distribution pipeline for the bulk transmission of de-identified clinical and molecular records, and may include separate storage for de-identified and identified data to maintain compliance with data privacy and applicable laws or guidelines, such as the Health Insurance Portability and Accountability Act.

[0044] The raw input data and / or the transformed, normalized, and / or predicted data may be stored in one or more relational databases for further access by the system to perform one or more comparison or analysis functions, as described in more detail herein. The data model used to build the relational database may be used to store, organize, display, and / or interpret a significant amount of different data, e.g., dozens of tables containing hundreds of different columns. Unlike standard data models such as OMOP or QDM, this data model may generate inherent links within and between tables to directly relate various clinical attributes, thereby facilitating the incorporation, interpretation, and analysis of complex clinical attributes.

[0045] After the relevant data has been received, transformed, and manipulated as described above, the system may comprise a number of modules such that the desired dynamic user interface can be generated, as described above with respect to the system diagram of FIG. 1.

[0046] Patient Cohort Filtering User Interface 2, a first embodiment of the patient cohort selection filtering interface 24 may be provided as a side pane 200 provided along the height (or alternatively the length) of a display screen through which attribute criteria 202 (clinical, molecular, demographic, etc.) can be specified by a user to define a patient population of interest for further analysis. The side pane 200 may be hidden or expanded by selecting it, dragging it, double-clicking it, etc.

[0047] Additionally or alternatively, the system may recognize one or more attributes defined for the tumor data stored by the system, which may be, for example, genotypic, phenotypic, ancestry, or demographic. The various selectable attribute criteria may reflect patient-related metadata stored in the patient data store 14, and example metadata may include, for example, project name (which may reflect a database that stores a list of patients) 204, gender 206, race 208, cancer, cancer site 210, cancer name 212, metastasis, cancer name 214, tumor site 216 (which may reflect where the tumor was identified), stage 218 (such as I, II, III, IV, and unknown), M stage 220 (such as m0, m1, m2, m3, and unknown), medication (such as name 222 or component 224), sequencing 226 (such as gene name or variant), MSI (microsatellite instability) status 228, TMB (tumor mutation burden) status (not shown), procedure 230 (such as by name), or death (such as by event name 232 or cause of death 234).

[0048] The system may also allow the user to filter the patient data according to any of the criteria listed herein, including those listed under the heading "Features and Feature Modules," and may include additional criteria, i.e., one or more of facility, demographics, molecular data, evaluation, site of diagnosis, tumor characterization, treatment, or one or more internal criteria. The facility option may allow the user to filter based on a particular facility. The demographic option may allow the user to sort, for example, by one or more of gender, death status, age at first presentation, or race. The molecular data option may allow the user to filter according to variant calls (e.g., when there is molecular data available for the patient, what the specific gene name, mutation, mutation effect, and / or sample type is), abstracted variants (e.g., including gene name and / or sequencing method), MSI status (e.g., stable, low, or high), or TMB status (e.g., selectable within or outside a user-defined range). The evaluations may allow the user to filter according to various system-defined criteria, such as smoking status and / or menopausal status. The diagnosis site may allow the user to filter according to primary and / or metastatic site. The tumor characterization may allow the user to filter according to one or more tumor-related criteria, such as grade, histology, stage, TNM classification (TNM), and / or each respective T, N, and / or M value. The treatment may allow the user to select from among various treatment-related options including, for example, ingredients, regimens, treatment types, etc.

[0049] Some criteria may allow the user to select from multiple sub-criteria that may be prompted after an initial criterion is selected. Other criteria may present the user with a binary option, for example, deceased or not. Still other criteria may present the user with a slider or range type option, for example, age at initial diagnosis may be presented as a slider with lower and upper values ​​that the user can select. Still further, for any of these options, the system may present the user with a radio button or slider that alternately determines whether the system should include or exclude the patient based on the selected criterion. It should be understood that the examples described herein do not limit the scope of the type of information that may be used as a criterion. Any type of medical information that can be stored in a structured format may be used as a criterion.

[0050] In another embodiment, the user interface may include a natural language search style bar to facilitate filter criteria definition for the cohort, for example, in the "Ask Gene" tab 236 of the user interface or via text entry in the filtering interface. In one aspect, a query may be specified via keyboard typed input or via machine interpreted dictation to define one or more of the subsequent layers of the cohort funnel (described in more detail in the next section). Thus, for example, when employing conventional natural language processing software or techniques, given the input "breast cancer patients", the system recognizes the filter "cancer_site==breast cancer" and adds this as the next layer of filtering. Similarly, the system recognizes the input "pancreatic patients with rejection to gemcitabine" and converts it into multiple successive layers of filtering, for example, "cancer_site==pancreatic cancer AND drug==gemcitabine AND rejection==not null".

[0051] In a second aspect, natural language processing may allow a user to use the system to directly query for general insights, thereby causing the system to refine the patient cohort through one or more funnel levels and display an appropriate summary panel in the user interface. Thus, if the system were to receive the query "What is the 5-year progression-free survival rate for stage III colorectal cancer patients after radiation therapy?", it would translate this into a set of filters such as "cancer_site==colorectal" AND "stage==III" AND "treatment==radiotherapy" and then display the 5-year progression-free survival rate using, for example, the patient survival analysis user interface 30. Similarly, a query such as "What percentage of female lung cancer patients are postmenopausal at the time of diagnosis" would be translated into a set of patients such as "gender==female", "cancer_site==lung", "temporal==time of diagnosis", and determine how many of the resulting patients had data reflecting a postmenopausal status, determine the associated percentages, and display the results, for example, via one or more statistical summary charts.

[0052] Cohort Funnel and Population Analysis User Interface 3-9, cohort funnel and population analysis user interface 26 may be configured to enable a user to perform analysis of a cohort with a goal of identifying important inflection points in the distribution of patients exhibiting each attribute of interest with respect to the distribution in the general patient population or in a patient population whose data is stored in patient data store 14. In one embodiment, the filtering and selection of additional patient-related criteria described above with respect to FIG. 2 may be used in connection with cohort funnel and population analysis user interface 26.

[0053] In another embodiment, the system may include a selectable button or icon that opens a dialog box 238 showing multiple selectable tabs, each tab representing the same or similar filtering criteria described above (demographics, molecular data, evaluation, site of diagnosis, tumor characterization, and treatment). Selecting each tab may present the user with the same or similar options for each respective filter as described above (e.g., selecting "Demographics" presents the user with additional options related to gender, circumstances of death, age at initial diagnosis, or race). The user may then select one or more options, select "Next," and then select whether it is an inclusion or exclusion filter, and the corresponding selection is added to a funnel (described in more detail below) and the icon moves down the next successively narrower portion of the funnel.

[0054] Additionally or alternatively, looking at a cohort, or set of patients, in the database, the system allows filtering by multiple clinical and molecular factors via menus 240. For example, with respect to clinical factors, the system may include filters based on patient demographics 242, cancer site 244, tumor characterization 246, or molecular data 248, which may further include their own subsets of filterable options 242, such as tissue 250, stage 252, and / or grade-based options 254 (see FIG. 4) for tumor characterization. With respect to molecular factors, the system may allow filtering according to variant calls 256, abstracted variants 258, MSI 260, and / or TMB 262.

[0055] Although the examples described herein provide analysis for various cancer types, it should be understood that in other embodiments the system may be used to show filtered views of other medical conditions, and in those situations the selections will be different to specifically focus on conditions related to other diseases.

[0056] The cohort funnel and population analysis user interface 26 may visually display the number of patients in a data set either all at once or incrementally after receiving a user selection of multiple filtering criteria. In one embodiment, a display of patient frequency by filtering attribute may be provided using an interactive funnel chart 264. As can be seen in Figures 3-9, with each selection, the user interface 26 is updated to illustrate a reduction in results matching the filtering criteria, e.g., as more filtering criteria are added, there are fewer patients matching all of the selected criteria after receiving each of the user's filtering factors.

[0057] The above filtering may be performed after receiving each user selection of filter criteria, with the funnel 264 updating to indicate the extent of the data set refinement after each filter selection. In that situation, the filtering menu 240 as described above may remain visible within each tab when switched to, or may collapse to the side, or may be represented as a summary 266 of the selected filtering options to inform the user of the reduced data set / size.

[0058] For each filtering method described above, the factor combination may be based on a Boolean combination. An exemplary Boolean combination may include allowing a user to select whether to search for patients with "A AND B," "A OR B," "A AND NOT B," "B AND NOT A," etc., for filtering factors A and B.

[0059] The final filtered cohort of interest may form the basis for further detailed analysis in the modules described below or in other user interfaces. The population of interest is referred to as a "cohort." The user interface may provide fixed-function attribute selectors that are appropriately pre-populated based on available data attributes in the patient data store.

[0060] The display may further show geographic location clustering plots of patients and / or comparison of demographic distribution with publicly available statistics and / or privately curated statistics.

[0061] Patient Timeline Analysis Module In addition, the system may include a patient timeline analysis module 28 that allows the user to review the sequence of events in each patient's bedside life, it being understood that this data may be anonymized as described above to protect the confidentiality of the patient data.

[0062] After the user has provided all of the desired filter criteria, for example, via the Cohort Funnel & Population Analysis user interface 26, the system allows the user to analyze the filtered subset of patients. With respect to the user interface depicted in the figure, this procedure may be accomplished by selecting the "Analyze Cohort" option 268 presented in the upper right corner of the interface 26.

[0063] 10, after requesting an analysis of the filtered subset of patients, the user interface may generate a data summary window in the patient timeline analysis user interface 28, providing information about the selected subset of patients, such as a number of other distributions in clinical and molecular features, in one or more regions 300. In one embodiment, a first region 300a may include demographic information, such as a mean patient age 302 and / or a plot of patient ages 304. A second region 300b may include additional demographic information, such as gender information 306, for the subset of patients. A third region 300c may include a summary of certain clinical data, including, for example, an analysis of medications 308 taken by each of the patients in the subset. Similarly, a fourth region 300d may include molecular data for each of the patients, such as an analysis of each genomic variant or alteration 310 carried by the patients in the subset.

[0064] The user interface 28 also allows the user to perform queries on the data summary information presented in the data summary window or area 300 and further sort the data using, for example, the control panel 312. For example, as can be seen in Figs. 11-14, the system may be configured to sort patient data based on one or more factors including, for example, gender 314, histology 316, menopausal status 318, response 320, smoking status 322, disease stage 324, and surgical procedure 326. Selecting one or more of these options may not reduce the patient sample size, as was the case above when describing the filtering summarized in the data summary window. Instead, the sorting function may subdivide the summarized information into one or more subcategories. For example, Figs. 11 and 12 show medication information 308 sorted by additional response data 328 overlaid in the data summary window 300c, along with a legend 330 explaining the overlaid response data.

[0065] 13-14, the subset of patients selected by the user may also be compared, e.g., via drop-down menu 332, to a second subset (or "cohort") of patients, thereby facilitating side-by-side analysis of the groups. Doing so may allow the user to quickly and easily see any similarities between the subsets, as well as any notable differences.

[0066] In one embodiment, a high-level overview event timeline Gantt-style chart is provided combined with a tabular details panel. This display may also allow visualization and comparison of multiple patients simultaneously on a normalized timeline with the goal of identifying both areas of overlap and potential discontinuities between patient subsets.

[0067] Patient "Survival Rate" Analysis Module The system may further perform survival analysis of subsets of patients using a patient survival analysis user interface 30, as seen in Figures 15-20. This modeling and visualization component may allow a user to interactively explore time to event (and probability in time) curves and their confidence intervals for subgroups of a filtered cohort of interest. The start of the time series and the target event, along with attributes that cluster groups of patients within a selected population, may be selected and dynamically modified by the user, all while the curve visualizer reactively adapts to the provided parameters.

[0068] To provide the user with flexibility in defining the scope of its analysis, the system may allow the user to select one or both of the beginning and ending events on which the analysis is based. Exemplary beginning events include initial primary disease diagnosis, progression, metastasis, regression, identification of first primary cancer, initial prescription of medication, etc. Conversely, exemplary ending events include progression, metastasis, recurrence, death, time period, and treatment start / end dates. Selecting a beginning event sets an anchor point for all patients from which the curve begins, and selecting an ending event sets a horizontal line along which the curve will be projected.

[0069] As can be seen in FIG. 15 , the analysis may be presented to the user in the form of a plot 300 of end events 302, e.g., progression-free survival or overall survival, versus time 304. Progression for these purposes may reflect the appearance of one or more progression events, e.g., metastatic events, recurrences, specific measures of progression to or independent of a drug, specific tumor size or change in tumor size, or enriched measurements (e.g., measurements indirectly extracted from an underlying clinical dataset). Exemplary enriched measurements may include detection of stage change (e.g., by detecting that a categorization of stage 2 has changed to stage 3), regression, or via inference (e.g., both stage 3 and metastasis are inferred from detection of stages 2 and 4, but no detection of stage 3).

[0070] In addition, the system may be configured to allow the user to focus or zoom in on a particular time span within the plot, as can be seen in FIG. 16. In particular, the user may be able to zoom in on only the x-axis, only the y-axis, or both the x-axis and y-axis simultaneously. This feature may be particularly useful depending on the type of disease being analyzed, as some aggressive diseases benefit from analyzing smaller time windows compared to other diseases. For example, survival rates for patients with pancreatic cancer tend to be significantly lower than other types of cancer, and therefore, when analyzing pancreatic cancer, it may be useful for the user to zoom in on a shorter period of time, for example, from a window of about 5 years to a window of about 1 year.

[0071] 17-20, the user interface 30 may also be configured to modify its display by receiving user input corresponding to additional grouping or sorting criteria to present survival information for smaller groups within the subset. Those criteria may be clinical or molecular factors, and the user interface 30 may include selectors such as one or more drop-down menus that allow the user to select, for example, either the start event 306 or the end event 308, as well as gender 310, gene 312, tissue 314, regimen 316, smoking status 318, disease stage 320, surgical procedure 322, etc.

[0072] Then, as shown in FIG. 18, selecting one of the criteria may present the user with multiple options related to that criterion. For example, selecting "Regimens" may cause the system to use one or more value sets to fill in selectable fields generated in the user interface and prompt the user to select one or more of the specific drug regimens 324 that one or more of the patients in the subset will receive. Thus, as shown in FIG. 19, selecting the "Gemcitabine+Paclitaxel" option 326 followed by the "FOLFIRINOX" option 328 results in the system analyzing the patient subset data to determine which patient records contain data corresponding to either of the selected regimens, recalculating survival statistics for those separate groups of patients, and updating the user interface to include separate survival plots 330, 332 for each regimen. Adding a group / adding two or more selections may result in the system plotting them on the same chart for side-by-side viewing, and the user interface may generate a legend 334 with names, colors, and sample sizes to distinguish each group.

[0073] As can be seen in FIG. 20, the system may allow for a higher level of analysis by calculating and overlaying statistical ranges for the survival analysis. In particular, the system may calculate confidence intervals for each data set requested by the user and display those confidence intervals 336, 338 for the survival plots 330, 332. In one case, the desired confidence intervals may be user set. In another case, the confidence intervals may be pre-set by the system, for example, 68% (1 standard deviation) intervals, 95% (2 standard deviations) intervals, or 99.7% (3 standard deviations) intervals. The confidence intervals may be calculated as Kaplan Meier confidence intervals or using another type of statistical analysis, as will be appreciated by those skilled in the art.

[0074] As will be appreciated from the preceding discussion, the utility of the system is based on its ability to highlight high importance features and interaction pathways that drive these predictions, and to further pinpoint cohorts of patients that exhibit levels of response that deviate significantly from the expected norm. In this context, high importance may be considered based on the importance of the feature to the outcome of the prediction. In particular, features that give the greatest weight to the prediction may be designated as high importance features. The system and user interface provide an intuitive and efficient method for patient selection and cohort definition given specific inclusion and / or exclusion criteria. The system also provides a robust user interface that facilitates internal research and analysis, including research and analysis of the impact of specific clinical and / or molecular attributes, as well as drug dosages, combinations, and / or other treatment protocols, on treatment outcomes and patient survival rates for potentially large and otherwise intractable patient sample sizes.

[0075] The modeling and visualization framework described herein allows users to interactively explore automatically detected patterns in the clinical and genomic data of filtered patient cohorts and analyze their relationship to treatment response and / or survival chances. This analysis can guide users to more informed treatment decisions for patients earlier in the cycle than without the use of the system and user interface. This analysis is also useful in the context of clinical trials, providing robust, data-backed inclusion and / or exclusion analysis of clinical trials. The system is supported by an extensive library of clinical and molecular data and integrates and applies various algorithms and concepts related to clinical analysis and machine learning to generate a fully integrated, interactive user interface.

[0076] Outlier Analysis Module 21-24, in another embodiment, the system may include additional user interfaces, such as a patient event likelihood analysis user interface 32, to quickly and effectively determine the presence of one or more outliers in a group of patients being analyzed. For example, the interface of FIG. 21 allows a user to visually determine how one or more groups of patients naturally separate in data based on progression-free survival. The user interface includes a first region 400 that includes a number of indicators 402 representing a number of patient groups, where each patient in a given group has commonality with other patients in the group. For example, the commonality may be based on one or more of the attributes described above, additional system-defined tumor relationship criteria used for filtering, and other medical information that may be stored in a structured format that may be identified by the system. In addition, groups may be formed from the absence of any attribute. For example, commonality may be found by groups that have never taken a drug, have never received a treatment, or otherwise share the absence of one or more attributes. This region may be similar to a radar plot 406 in that the indicators are plotted radially away from and circumferentially around a central indicator 408, with the radial distance from the central indicator 408 reflecting the similarity between the patients represented by the central radially spaced indicators and the circumferential distance between the radially spaced indicators reflecting the similarity between the patients represented by those indicators. In this instance, the similarity with respect to radial distance may be based exclusively or solely on the criterion / criteria governing the outlier analysis.For example, when analyzing a patient group for progression-free survival ("PFS"), the center point or indicator 408 may be based on a particular proportion or percentage (e.g., 10%, 25%, 50%, 75%, or other percentage) of the PFS of the entire cohort over the time period being evaluated, and the radial distance from the center point or indicator 408 may be indicative of the progression-free survival of the group of patients reflected by the respective indicator 402 such that groups of patients with better than a particular percentage PFS are plotted above the center point or indicator 408 and groups of patients with worse than a particular percentage PFS are plotted below the center point or indicator 408, and the distance from the center point on the X-axis may be derived based on the size of the population, the difference between observed and expected PFS, or a similar metric.

[0077] Additionally, the user interface may include a second region 410 including a control panel 412 for filtering, selecting, or otherwise highlighting a subset of patients as outliers in the first region. Setting a value or range in the control panel may generate an overlay 414 on the radar plot (see FIG. 22), which may be in the form of a circle centered on the central indicator 408, with the radius of the circle related to the value or range received from the user in the second region 410. In this aspect, the user may select a value that applies equally in both directions with respect to the reference patient. For example, the user may select "25%", which may be reflected as a range of -25% to +25% such that the overlay may be a uniform circle surrounding the central point or indicator 408. Alternatively, the system may receive multiple values ​​from the user, some representing positive ranges and some representing negative ranges, such as "-20% to +25%". Values ​​may be received via text input, dropdowns, or selected by clicking on the respective location on the graph. In that case, the overlay may take the form of two separate hemispheres with different radii, which reflect the values ​​received from the user. As can be seen in Figures 21 and 22, the values ​​may indicate a percentage of deviation from whatever value is relative to the center point or indicator 408. For example, Figures 21 and 22 display the progression free survival (PFS) percentages of various clusters of patients centered around a patient with a PFS value of 0%. Figure 21 includes an overlay 414 with a range of ±10%, while Figure 22 shows how the overlay adjusts when the range is modified to ±30%. It will be appreciated that the center point or indicator 408 may be associated with a non-zero value, for example, a patient with a PFS of 20%. In that case, the ±10% range encapsulates a cluster of patients in the range of 10-30% PFS, and the ±30% range encapsulates a cluster of patients in the range of -10-50%.In either case, after the system receives the user input, the indicators covered by the overlay may change visual appearance, for example, to be displayed in a lighter gray or in some other less noticeable form, as shown in Figure 22, with values ​​416 (shown in histogram form in the upper right corner of Figure 22) that are outside the outlier threshold 414 displayed in a darker color (e.g., blue or shading) and values ​​418 within the outlier threshold 414 displayed in a lighter color (e.g., light gray or no shading). That is, indicators that are outside the overlay may remain highlighted or otherwise more easily visually distinguishable, thereby identifying them as representing outliers.

[0078] 23-24, the first region 400 of the user interface may include a plot 420 of multiple patient groups that differs from the radar-type plot just described. In this embodiment, the x-axis 422 may represent the number of patients in a given group, represented by an indicator, and the y-axis 434 may represent the degree of deviation from the norm / norms being considered. As a result of these display parameters, the user interface 32 presents the largest patient group 436 furthest from the y-axis and the largest outlier group 438 furthest from the x-axis 422. (It should be understood that for both this user interface and the previously described user interfaces, the origin may not reflect a value of 0 for either the y-axis or the radial dimension, respectively. Instead, the origin may reflect a basal level of the criterion / criteria being analyzed. For example, in the case of progression free survival, the basal group may have a 2-year survival rate of 15%. A deviation may then be determined with respect to that 15% value to assess the presence of an outlier. Such deviations may be additive, ±20% being 0% to 35% (0% instead of -5%, since there can be no negative survival rates), or multiplicative, ±20% being 12% to 18%.)

[0079] Similar to the previously described user interfaces, the interface of Figs. 23-24 may include a second region 410 that includes a control panel 412 for modifying the presentation of the identifiers in the first panel 400. Again, similar to that interface, the control panel may allow the user to make uniform or independent selections for the positive and negative sides of the scale. In particular, as can be seen in Fig. 24, the control panel 412 in this instance allows the user to independently select the positive and negative ranges in the search for outliers. After each selection is made, the user interface 32 may dynamically adjust to cover, obscure, de-highlight, remove, or otherwise differentiate indicators that fall within the zone selected by the user from outlier indicators that fall outside of that zone. As described above, depending on the configuration of the x-axis and y-axis, the user interface 32 may be configured to allow the user to quickly identify which outlier group is the furthest removed from the representative patient / group because that outlier group is the furthest removed from the x-axis in the positive direction, the negative direction, or both directions. Similarly, the user interface 32 may be configured to facilitate a user to quickly visually determine which patient group has the greatest number of patients because that group is furthest away from the y-axis in the positive direction, the negative direction, or both. Additionally, the axis combinations may enable a user to quickly visually determine which indicators warrant further investigation, for example, by the user visually determining which indicators provide an ideal tradeoff between degree of deviation / outlier and patient size.

[0080] With respect to any of the outlier user interfaces described above, the interface may further comprise a third region 440 that provides information specific to the selected node when the system receives user input corresponding to a given indicator, for example, by clicking on that indicator 436 in the first region of the interface, as seen in FIG. 24. In one embodiment, the additional information may include a comparison of the criterion / criteria being evaluated as compared to the value of the entire population used to generate the interface in the first region. The information in this region may also include the total number of patients in the record set, the number of patients for which the record set has been filtered based on one or more different criteria, and an identification of the population size of the selected node as part of an inline plot, where these size comparisons may help inform the user about the potential significance of the outlier group.

[0081] In addition, with respect to any of the outlier user interfaces described above, the algorithm for determining the presence of outliers may be based on a binary tree 500 as shown in FIG. 25A and FIG. 25B. To generate such a tree, the system may separate each feature into its respective category. For each category, the system may determine which subset of the cohort has the largest spread of progression-free survival versus non-survival, treating the split feature that produced the largest spread as the edge between the nodes and the feature itself as a node. The system may continue this analysis until it encounters a leaf. For example, the mutation column may be split into "with mutation" and "without mutation," and the age option may be set by the user to "over 50" and "under 50." The system may then determine what the maximum cutoff age for survival is and use that as the binary decision point. Of all of these categories, each with a binary choice splitting into two groups, the system may determine which has better survival and which has worse survival, and compare those decisions across all columns to find the group with the largest difference. The category with the greatest difference is the first node split in the tree that continues to split with additional nodes, forming multiple branches with the category criteria for groups being the edges between each node. Each branch ends in a leaf, which is simply a split of all the features that came before to identify the group of people with the highest PFS in the cohort according to the split above it. In one embodiment, the system may treat each leaf as an outlier. Alternatively, an outlier may be some particularly divergent feature. For example, an outlier leaf may be one that deviates from a user-entered value or expected value by some threshold, for example, one standard deviation or more away from the expected threshold.

[0082] In some cases, data in the branches may be lost when the system extrapolates all the way to the leaves. In such cases, the system may scan for features the current patient has in common with the outlier patients and suggest clinical process changes that may place them in a new bucket (leaf / node) of patients with higher outliers. For example, if a branch has a high PFS in a node, but loses distinction by the time the branch resolves to a leaf, the system may identify the node with the highest PFS as the leaf.

[0083] To generate the expected survival rate for the population, the system may rely on a predictive algorithm built on the survival rates of patients in dataset 14. Alternatively, the system may use an external source for the PFS prediction, such as FDA published PFS for some cancers or treatments. The system may then compare the expected survival rate to the observed PFS rate for the population to determine outliers.

[0084] In one particular embodiment, a method is provided for identifying one or more outlier groups of patients. The method includes selecting a cohort of patients, the cohort including a plurality of patients. The selection of the cohort may be based on identifying a group of patients having a particular condition, such as a particular disease. In one particular embodiment, the cohort may include a group (e.g., tens, hundreds, thousands, or more) of patients having non-small cell lung cancer or breast cancer. Other groupings based on other criteria are also possible.

[0085] In various embodiments, the next step of the method may include calculating the average survival rate for the cohort of patients. For example, it may be determined that, on average, these patients survive for a certain amount of time (e.g., a number of months, such as 63 months) based on the available data.

[0086] In some embodiments, another step of the method may include selecting a plurality of clinical or molecular characteristics associated with the cohort of patients. The clinical or molecular characteristics associated with the cohort of patients may include one or more of a genetic marker, a procedure performed on the patient, a drug treatment given to the patient, the age at which the patient was diagnosed, the age at which the patient was treated, or a lifestyle indicator. In certain embodiments, the clinical or molecular characteristics of the patient may include the smoking status of the patient (e.g., yes, no, unknown), a DNA mutation associated with the patient (e.g., KRAS, BRAF, EGFR, etc.), the age of the patient at diagnosis or treatment (e.g., one or more integers within a particular age range, such as 18-115 years), or one or more therapies or medications received by the patient.

[0087] In some embodiments, information about a cohort of patients may be used to generate a tree structure, and nodes of the tree structure may include one or more patients that are outliers, i.e., patients that exhibit significantly different survival rates (shorter or longer) for a given set of conditions. Thus, to generate the tree structure, for each characteristic of the plurality of characteristics, the method may include identifying a plurality of data values ​​associated with the characteristic. For each data value of the plurality of data values ​​associated with the characteristic, the method may include dividing the cohort of patients into a first subgroup and a second subgroup of the plurality of patients based on a criterion such as whether each patient of the plurality of patients survived in an outlier time period, determining a difference between the number of patients in the first subgroup and the number of patients in the second subgroup, and selecting the data value that results in a difference that is the largest difference between the number of patients in the first subgroup and the number of patients in the second subgroup.

[0088] This procedure may be repeated for each data value of each characteristic. For example, for embodiments in which the characteristic relates to age, the data values ​​include a range of ages, starting with a low age range such as ages 18, 19, 20, 21, ..., and going to an upper age range such as age 115 (or another suitable value). Then, in one particular example, if age=20 and the time period is x years (e.g., 5 years), then a first cohort of patients may be those who died x years after diagnosis at age 20, and a second cohort of patients may be those who did not die within x years of diagnosis at age 20.

[0089] To determine the difference, the number of patients who did not survive within a particular time is considered a first subgroup of patients, and the number of patients who survived within a particular time is considered a second subgroup of patients. Then, for each data value associated with each characteristic, the difference between the number of patients in the first subgroup and the number of patients in the second subgroup is determined. This difference may be divided by the total number of patients in the first and second subgroups and expressed as a decimal value between 0 and 1 (e.g., if 400 patients died x years after diagnosis at age 20 and 100 patients did not die x years after diagnosis at age 20, the difference is 400-100=300, which is divided by 500, the total number in the two groups, to obtain a difference of 0.6). The particular data value with the largest such difference may be retained while the procedure is performed to determine the nodes of the tree structure (e.g., the largest difference may be a difference of 0.7 at age=44).

[0090] The method may further include creating a new node of the tree structure based on the data value that results in the largest difference between the number of patients in the first subgroup and the number of patients in the second subgroup (e.g., one node may be created for age=44). After a particular data value is identified as having the largest difference, the method may then include creating branches from the node, including creating a first branch from the new node based on the first subgroup and creating a second branch from the new node based on the second subgroup. Some examples of potential nodes may include: smoking=yes, difference=0.8; DNA mutation=KRAS, difference=0.78; age=82 years, difference=0.9; sex=male, difference=0.6. Based on this information, the "age" characteristic has the largest difference and may be selected and branches based on ages 82 years and older and ages younger than 82 years may be created.

[0091] The tree structure may continue to be built by repeating the steps above, including splitting the cohort into subgroups for each characteristic and each data value of each characteristic. The starting cohort for each subsequent iteration is the group of patients in a particular node at the starting point. This procedure is repeated at each node based on the patients in the first subgroup and the second subgroup, respectively. This procedure continues until one or both of the following conditions are met: (1) a maximum number of nodes or branches have been created, or (2) the node contains fewer than a minimum number of patients. When the procedure is completed, the method may include identifying at least one node from the tree structure that contains an outlier group of patients.

[0092] Smart Cohorts In various embodiments, a predictive model can be developed that facilitates the identification of one or more cohorts of patients whose disease progression and / or survival probability is substantially different from expected, for example, significantly longer or shorter than expected.The information from these cohorts can then be examined to identify one or more key factors that may potentially contribute to the survival profile of the cohort.The identification of smart cohorts can be used to provide precision medicine results for specific patients, to help identify potential areas of interest for drug research, and / or to identify unexpected potential for expanding drug patient coverage.

[0093] Given a set of patient timelines, in various embodiments, the Smart Cohort module has three objectives and attempts to answer one or more of the following questions:

[0094] 1. What is the likelihood (i.e., "survival rate") of each patient to survive beyond Y years (or live at least Y years progression-free), measured at each event point in the patient's timeline?

[0095] 2.What are the major factors that most influence expected survival outcomes?

[0096] 3. Which subset of patients exhibit a combination of these factors that makes them stand out as an outlier cohort with respect to survival profile versus expectation at a user-specified anchor timeline event (e.g., time of diagnosis at stage IV), and what are the characteristics of these patients?

[0097] This problem can be approached from a time series modeling perspective, with time snapshots of feature states, and a binary classification objective. In some embodiments, a tree-based supervised clustering approach can be used to help identify patient groups of interest, although other embodiments include other analysis and visualization methods.

[0098] The inherent temporal nature of the problem is complicated by the fact that the target survival rate at anchor point T may depend as much on what happens to the patient after point T as it does on what happened before point T. As such, expected future survival rates cannot be simply modeled using event history alone; future events cannot be included in the model without invalidating the model as a recommender or inadvertently introducing information leakage into the features, which may result in overfitting.

[0099] In some embodiments, a hybrid two-model approach may be employed, in one part of which a history-only model is trained to derive the "expected values" at each time point, and in another part of the approach, a forward clustering model is developed to isolate the discrepancy between expected and observed survival along with associated features.

[0100] Thus, in some embodiments, a hybrid approach may include:

[0101] 1. We build a dataset that utilizes only backward-looking features derived at each event point on the timeline.

[0102] 2. Train a model on such a dataset to derive predictions of expected future survival at each time point.

[0103] 3. Tag these expected survival predictions at each time point to act as best estimate prior probabilities using all historical information content.

[0104] 4. Construct a "forward-looking" feature set for each time point, ensuring that no implicit survival duration information is incorporated into the features (in some cases historical priors may be included as features in this set as well).

[0105] 5. Train a “summary / clustering” model using the forward feature set.

[0106] At this point, following the "training" step, a decision can be made as to whether to limit how prospective the features for this portion can be. For example, it may not make sense to include features observed over the next 2 years if you are trying to predict 1-year survival. In addition, you may want to give less importance to features that occur far from the anchor event. Finally, you may want to exclude event points that are observed after the outcome event of interest, even if the event occurred within the X-year boundary. For example, if the first progression event was observed within 6 months and you are predicting 2-year PFS, then for that patient, you should exclude all events between 6 months and 2 years.

[0107] 6. For each of the forward clusters, the expected survival predictions based on the forward model are compared to the actual survival, and clusters with large deviations from the expected survival predictions are identified, along with their constituent forward feature sets.

[0108] Thus, the model is directed to determining how future events may affect the expected survival predicted by prior events, regardless of whether the expected survival prediction for a particular subcluster is higher than the expected survival prediction for a different cluster (but also looking at the root causes of deviations in expected survival predictions). That is, it is of interest to know whether the next action will affect the patient's survival, or whether the patient's survival is determined solely by events already experienced.

[0109] A predictive model may be implemented based on data from multiple patients, using information on patient history and treatment along with information on patient survival. To perform temporal alignment of data from multiple patients, one or more anchor points (also referred to as "patient time points") may be identified within the data (FIG. 26). Anchor points identify time points that are common to all or at least many of the patients and may help standardize the time course of data for events such as disease progression. Anchor points may include events such as time of first diagnosis, time of first metastasis, or time of first treatment, although other anchor point events are possible. FIG. 26 shows a prediction model for patient P based on a common anchor event. 1 , P 2 , P 3 , ..., P n This shows the alignment of the timelines.

[0110] There may be some imprecision with respect to the time of some anchor point events, for example, the date of first diagnosis may occur several weeks earlier or later (e.g., with respect to when the illness began) for a given patient due to the time the patient first notices symptoms or sees a doctor to receive a diagnosis, allowing for a lack of precision. Thus, in some embodiments, the anchor points may include a tolerance window before and / or after the anchor point date, which may add flexibility to the modeling procedure. In various embodiments, the tolerance window may be ±1 day, 3 days, 1 week, 2 weeks, 1 month, 2 months, 3 months, or other suitable period. FIG. 26 shows an anchor event (set on January 1st) and a 12-month progression window thereafter. The anchor event may have a ±15-day tolerance window associated with it. In addition, the progression window may have a 3-month tolerance window, so the progression reference point window may extend back in time from 3 months before January 1st to October 1st.

[0111] With respect to predictive models, in various embodiments, multiple data are obtained or received for multiple patients over a period of time (e.g., a time span spanning each of the patient's medical history from the time of the patient's diagnosis to the present or time of death, although the medical history may begin prior to diagnosis).

[0112] The data is processed to identify a number of patient time points (anchor points) occurring within the time period covered by each patient's data. As described above, anchor points or patient time points may include time points associated with any patient interaction with the healthcare system, including any interaction with an individual or facility that provides healthcare or obtains healthcare information, such as a healthcare provider, a genetic sequencing facility, a hospital outpatient or inpatient facility, etc. Patient time points may be identified by a date that is attached to or associated with each data in the received set of patient data.

[0113] In general, both temporal and static features may be derived from patient data, but the analysis at this stage is purely retrospective to avoid future information leakage. Different categories or classes of features include "time since last / first XXX", "number of XXX", or "demographics". Extracting features may include multiple lookback horizons, for example features may be bounded to the past 12 months or based on continuous historical analysis.

[0114] In one particular example, for hypothetical patient A, four time points may be identified: date of biopsy collection, July 1, 2018 (KRAS PL1S147GLU mutation with identified high SNP effect), initiation of anastrozar and lotinib, August 1, 2018, administration of radiation therapy, November 1, 2018, reported treatment outcome: disease progression from stage 1 to stage 2, January 1, 2019, imaging performed, July 1, 2018 and November 1, 2018. The other patients B, C, D... each have their own set of time points that may correspond to some of the same events (e.g., diagnosis, medication initiation, imaging, etc.), or to different events, or to a combination of some of the same events and some different events.

[0115] Based on the data for each of the patients and for each patient time point, an outcome target for the outcome event may be calculated within a horizon time window, a number of prior features may be identified, and the status of each of the number of prior features at the patient time point may be determined. The outcome event may include a patient and / or disease state, such as progression or death, and the outcome target may be described with a target label, such as "yes" or "no," indicating whether the outcome occurs within a particular horizontal time window from the patient time point / anchor point, along with an end date. The horizon time window may include any suitable period, such as 3 months, 6 months, 9 months, 12 months, 24 months, 36 months, 48 ​​months, or 60 months.

[0116] For the hypothetical patient A, the analysis of progression events occurring within 6 months of the time point is as follows:

[0117] Patient A: July 1, 2018 -- Progression within 12 months -- Yes, January 1, 2019

[0118] Patient A: August 1, 2018 -- Progression within 12 months -- Yes, January 1, 2019

[0119] Patient A: November 1, 2018 -- Progression within 12 months -- Yes, January 1, 2019

[0120] Patient A: January 1, 2019 - Progression within 12 months - Ineffective

[0121] The data for Patient A included information on the report of progression from Stage 1 to Stage 2 on January 1, 2019, so there is a valid outcome target of "Yes" for "Progression within 12 months" for each of the first three time points. However, the analysis for the final time point is displayed as "null" because there is no patient information available to inform the model after this date. Although progression was reported on this date, there is no further information available for Patient A after this date.

[0122] The prior features may include various features related to the patient's condition and / or treatment. In various embodiments, the prior features may include time / time-based events or features, structural or biological features, or molecular / genetic features, among other categories. In certain embodiments, the prior features may include one or more of the following: time since starting a particular drug therapy, time since taking a particular drug, time since last progressive treatment outcome (e.g., patient response to a drug), time since metastasis, maximum tumor size to date / last recorded tumor size, most severe effect of identified SNPs (e.g., low effect, high effect), or RNA features (e.g., expression levels per gene / transcript). In some embodiments, the data may require additional processing, such as using an autoencoder to reduce the dimensionality of the feature space.

[0123] The status of each pre-feature can be determined at each of the patient timepoints. For a hypothetical patient A, the status of three features (time since starting medication A, time since last imaging, and highest SNP effect identified by lab A) for each of the four patient timepoints are shown below (note that the value of "time since taking medication A" at the first patient timepoint is "null" since patient A did not take medication A until the next timepoint):

[0124] Patient A: July 1, 2018

[0125] Time since starting medication A: null

[0126] Time since last imaging: 0 days

[0127] Highest SNP effect as identified by Lab A: Germline: KRAS: High (5)

[0128] Patient A: August 1, 2018

[0129] Time since starting medication A: 0 days

[0130] Time since last imaging: 1 month

[0131] Highest SNP effect as identified by Lab A: Germline: KRAS: High (5)

[0132] Patient A: November 1, 2018

[0133] Time since starting medication A: 3 months

[0134] Time since last imaging: 0 days

[0135] Highest SNP effect as identified by Lab A: Germline: KRAS: High (5)

[0136] Patient A: January 1, 2019

[0137] Time since starting medication A: 5 months

[0138] Time since last imaging: 2 months

[0139] Highest SNP effect as identified by Lab A: Germline: KRAS: High (5)

[0140] Next, for each patient time point among the multiple time points having a valid outcome target, and for each combination of horizon time window and outcome event, multiple forward features can be identified. The combination of horizon time window and outcome event can include "progression within 6 months", "progression within 12 months", "progression within 24 months", "progression within 60 months", "death within 6 months", "death within 12 months", "death within 24 months", "death within 60 months", etc.

[0141] For patient A, using the horizon time window / outcome event combination of "progression within 12 months," prospective characteristics may include:

[0142] Patient A: From July 1, 2018

[0143] Will the patient take drug A from the time point onwards until the end date (yes)?

[0144] Did the patient take drug A before the time point (no)

[0145] Highest SNP effect as identified by Lab A: Germline: KRAS: High (5)

[0146] Patient A: August 1, 2018~

[0147] Will the patient take drug A from the time point onwards until the end date? (No)

[0148] Highest SNP effect as identified by Lab A: Germline: KRAS: High (5)

[0149] Did the patient take drug A before the time point (yes)

[0150] Patient A: From November 1, 2018

[0151] Will the patient take drug A from the time point onwards until the end date? (No)

[0152] Highest SNP effect as identified by Lab A: Germline: KRAS: High (5)

[0153] Did the patient take drug A before the time point (yes)

[0154] At this point, a set of predictions for a number of patients may be generated based on the prior features and the forward features, and a predictive model may be generated based on the set of predictions using machine learning. In some embodiments, the predictive model may be generated using gradient boosting.

[0155] The plurality of sets of predictions may be divided into a plurality of folds, where each fold contains data corresponding to a subset or subgroup of a plurality of patients such that data for each patient is kept within the same fold (Figure 28). Thus, a machine learning procedure such as gradient boosting can be trained using a subset of the folds. For example, if there are eight folds, the gradient boosting algorithm may be run on seven of the eight folds. The remaining fold not used for training is run through the model for prediction purposes, and the difference between the predicted and actual results can be used to adjust the model before subsequent rounds of training are run. This can be repeated with different folds that are omitted from the training step and used for prediction and / or adjustment of the model. More generally, if there are N folds, training can be run on X < N folds and prediction can be run using N - X folds. When generating a prediction model, various parameters can be adjusted, including the learning rate, maximum depth of the tree, minimum leaf size, etc. (depending on the type of model). The goal is a model that learns the relationship between the pre - characteristics of all patients leading to the target outcome. Predictions are received from the model at each patient time point and associated or linked with the corresponding outcome target. In some embodiments, eight folds are cross - validated and an additional two folds can be a full hold - out for another testing purpose. The folds may be stratified by a combination of multiple characteristics such as target, gender, cancer, patient event count, etc.

[0156] After generating multiple predictions, this information can be used to identify one or more "smart cohorts", i.e., one or more cohorts of patients whose disease progression and / or survival chances are substantially different from expected, e.g., significantly longer or shorter than expected. In general, a decision tree is constructed using the prediction information, thereby identifying various potential smart cohorts, which can ultimately be grouped into various leaf nodes of the decision tree. Two approaches for constructing the decision tree are disclosed herein, referred to as offline smart cohorts and online smart cohorts.

[0157] Offline Smart Cohorts In some embodiments, a method for identifying a cohort of patients may be developed. The method may include selecting a cohort of patients that includes a plurality of patients, for example, a cohort of 500 breast cancer patients. Generally, the cohort may be selected based on the patients having a particular condition in common, for example, a particular disease.

[0158] The method may also include identifying a common anchor time point from the set of anchor points associated with each of the patient groups, the common anchor point being shared by each of the groups of patients in the cohort. Selecting a common point between all patients facilitates visualization of the data and also allows for preventing the same patient from appearing multiple times in the model at each of the patient's available anchors. Possible anchor points include time of diagnosis, time of treatment, time of metastasis, and other times. In one particular embodiment, the time of diagnosis may be selected as the anchor point.

[0159] For each patient in the group of patients, the timelines associated with each of the group of patients may be aligned to a common anchor point. An outcome target may then be identified, such as disease progression within 12 months. A set of previously generated predictions, each including a predicted target value, may then be retrieved for each patient in the group of patients and for each of the multiple antecedent features and multiple prior features. The predictions may include information such as that shown in Table 1.

[0160] [Table 1]

[0161] More generally, the "target prediction" may take the form of a "probability of survival (PFS) in X months," "death in X months," "likelihood of taking a drug in X months," "likelihood of other targets in X months," etc., and may take the form of a decimal value between 0 and 1. The "target actual" value is essentially a binary yes / no value, denoted by 1 or 0, representing the occurrence or non-occurrence of an event within X months. In various embodiments, the feature set may include pre- and / or forward features, e.g., any of the features disclosed herein, including those described under the heading "Features and Feature Models." The pre-features may include one or more of age, sex, treatment (e.g., drug, procedure, therapy, etc.), sequencing / lab / imaging results. The forward features, further described below, may include future occurring events, treatments, etc., between the anchor point and the observed target.

[0162] In various embodiments, hundreds or thousands (or other larger numbers) of decision trees may be generated using this information, e.g., using a procedure similar to that described above for the outlier procedure. For each of the decision trees constructed, the following procedure may be performed for each feature in the plurality of leading features and the plurality of prior features.

[0163] The group of patients may be divided into a first subgroup and a second subgroup based on the difference between the predicted target value and the actual target value.

[0164] - the difference between the number of patients in the first subgroup and the number of patients in the second subgroup may be determined.

[0165] The feature that results in a difference that is the largest difference between the number of patients in the first subgroup and the number of patients in the second subgroup may be selected.

[0166] A new node of the tree structure may be created based on the feature that results in the largest difference between the number of patients in the first subgroup and the number of patients in the second subgroup. A first branch may be created from the new node based on the first subgroup, and a second branch may be created from the new node based on the second subgroup. The step of building a decision tree may then be repeated for each of the first and second branches based on the patients in the first and second subgroups. This may continue until completion as defined by either a maximum number of nodes or branches have been created or a particular node contains fewer patients than the minimum number for all nodes and branches.

[0167] The goal of constructing the decision tree is to predict for each patient the difference between the predicted and actual outcome for the target based on the features in the feature set by clustering the patients based on which feature most accurately predicts the difference between the predicted and actual outcome.

[0168] In some embodiments, the method may include determining a similarity metric by determining how often a given patient ends up in the same leaf node of the tree as other patients across hundreds or thousands of decision trees. Thus, for each patient in the group of patients, the method may include identifying the co-occurrences of the given patient appearing in each of a plurality of leaf nodes with each of the other patients of the plurality of patients across hundreds or thousands of decision trees. The similarity metric may be determined for the given patient based on the sum of the co-occurrences divided by the total number of nodes in which the given patient falls across all of the hundreds or thousands of decision trees constructed and analyzed. In some embodiments, a database of patient-to-patient similarity metrics may be generated based on determining the similarity metric for each of the plurality of patients. In other embodiments, the similarity metric may be displayed, for example, as a cohort trader plot. Additionally, the data may be displayed in association with one or more of the steps outlined above for identifying at least one of the plurality of features.

[0169] The method may further include determining a similarity metric for a new patient, i.e., a patient different from the initial group of patients. The new patient may be matched with a subgroup of patients corresponding to a particular leaf node of the plurality of leaf nodes based on determining the similarity metric. A treatment may then be identified for the new patient based on matching the new patient with the subgroup of patients. Furthermore, the database of patient-to-patient similarity metrics may be processed using a dimensionality reduction algorithm to identify a particular cohort of patients that have shared features, such as shared prior features or shared forward features. In general, dimensionality reduction identifies several subgroupings (e.g., K subgroups), where each of the subgroups 1-k has some characteristic in common across the groupings identified from the entire patient cohort (standard population grouping).

[0170] Online Smart Cohorts In addition to the multiple predictions, the system may receive an outcome target, a subset of multiple antecedent features corresponding to the outcome target, and a patient cohort that includes a subset of the multiple patients. The cohort may be a group that shares a condition or trait of interest, for example, the cohort may be a group of 20,000 breast cancer patients. This group is then subdivided using a decision tree to find one or more specific subgroups of interest for further investigation.

[0171] Table 2 shows an example of the type of forecast data that may be received.

[0172] [Table 2]

[0173] The forward features may include various future actions or states related to the patient and, in some embodiments, may be used to advise patients with certain medical conditions. Some of the forward features may be "actionable", i.e., may include things that a given patient can do to change the prognosis or outcome. For example, a doctor or other clinician may be able to take some steps or actions to improve the patient's prognosis (e.g., prescribe a drug or combination of drugs, prescribe a particular treatment such as surgery, chemotherapy, or radiation, send a tumor sample for sequencing and receive molecular information such as testing for DNA markers). Certain molecular features may or may not be considered actionable based on whether the molecular information obtained is relevant to a subsequent action or step. In various embodiments, features such as lab results, imaging results, tumor characterization (e.g., histology, grade, TNM stage, etc.) may not be included as forward features to avoid suggesting to a patient to take actions that are not within their control, such as "lower N stage", "higher hemoglobin concentration", etc.

[0174] In various embodiments, this information can also be used to counsel specific patient groups, for example, for stage N patients with X mutations, treatments A and B taken together will improve the probability of survival (PFS) within 12 months. For example, stage 4 breast cancer patients with KRAS mutations are predicted to progress based on placement in the cohort (90% predicted progression), and should receive anastrozar in combination with lotinib as an intervention to improve PFS within 12 months based on predictions beyond the selected anchor point at the time of first metastasis (60% predicted progression). Other specific courses of action can also be determined based on the data.

[0175] Examples of predictions include predictions of survival probability within 12 months for patients A and B and time points T1 (January 1, 2018) and T2 (May 1, 2018), expressed as probability values ​​between 0 and 1, as shown in Table 3.

[0176] [Table 3]

[0177] The outcome target may be the probability of survival within 12 months, given as 0 or 1, as shown in Table 4.

[0178] [Table 4]

[0179] Below is an example of a subset of multiple forward features (FD1, FD2, FD3, each shown below) corresponding to an outcome target that includes forward data corresponding to probability of survival within 12 months.

[0180] January 1, 2018:

[0181] FD1 (patient takes anastrozar and lotinib): (Yes)

[0182] FD2 (Patient receives radiotherapy):....

[0183] FD3 (Patient undergoes surgery): ....

[0184] May 1, 2018:

[0185] FD1 (patient takes anastrozar and lotinib): (Yes)

[0186] FD2 (Patient receives radiotherapy):....

[0187] FD3 (Patient undergoes surgery): ....

[0188] The system may also receive anchor points or patient time points, such as time of first diagnosis, time of first metastasis, time of first treatment, and the like.

[0189] A subset of the multiple forward features may be selected. These features may include drugs (future and past) and even sequencing (somatic sequencing (future or past), germline sequencing, etc.). For each patient in the cohort with the anchor point, a predictive model may comprise the selected subset of the multiple forward features, and a difference between each of the multiple predictions and the outcome target may be determined.

[0190] For example, a model may receive data such as the following:

[0191] Patient A: [.95~1], [drug and sequencing dataset]

[0192] Patient B: [.92~1], [drug and sequencing dataset]

[0193] Patient C: [.63~0], [drug and sequencing dataset]

[0194] The data may include information such as "drugs and sequencing data sets at anchor points," which may include an N x M table of patients and their respective features. Each feature may include information such as:

[0195] Patient A: July 1, 2018 (anchor point date) -

[0196] Column 1: Does the patient take Drug A after the time point and before the endpoint date (Yes)

[0197] Column 2: Did the patient take drug A before the time point (No)

[0198] Column 3: Highest SNP effect as identified by Lab A: Germline: KRAS: High (5)

[0199] A decision tree may then be generated based on determining, for each feature of the selected subset of the plurality of forward features, a maximum difference between each of the plurality of predictions and the outcome target. The decision tree may include a plurality of leaf nodes and one or more branch nodes, each of which may include a pair of branches that include the leaf node or the branch node, and the branches are formed based on the features selected from the subset of the plurality of forward features.

[0200] Each of the plurality of leaf nodes of the decision tree may include a number of patients from the cohort of patients. In some embodiments, the decision tree may continue splitting based on the difference between each of the plurality of predictions and the outcome target until the number of patients in a particular leaf node of the plurality of leaf nodes is less than the minimum number of patients. In other embodiments, the decision tree may continue splitting based on the difference between each of the plurality of predictions and the outcome target until the number of levels of the decision tree reaches a certain number, i.e., equal to the maximum number of levels. In one particular example, each patient's status with respect to the feature "KRAS somatic: medical history>3" may be used to split the branch node into two branches based on whether each patient's medical history importance value for this marker is greater than 3 (high importance).

[0201] Leaf nodes of decision trees provide information that can be used to identify cohorts of interest. In some cases, leaf nodes may have high values ​​for the prediction target because the predictions are, on average, significantly higher than the target values. For patient C in the above example, the predictions indicated that patient C was likely to progress, but this did not happen. In other cases, leaf nodes may also generate low negative values ​​for the "prediction-target" difference, e.g., the prediction-target value may be [.05-1]=-.95, indicating that the patient is unlikely to progress, but may still progress in some cases. However, in some cases, leaf nodes may have values ​​close to zero, indicating that the model made an accurate prediction. The smart cohort procedure focuses on cases where the patient's actual outcome deviated significantly from the expected outcome, because these groups of patients can provide information about what can be done to change the disease progression trajectory, while the cohorts with the prediction-target difference closest to zero inform the model what features are most important for reliable prediction.

[0202] In some embodiments, analytics may be run on one or more of the leaf nodes of the decision tree, where the analytics parses the leaf branches to make them meaningful. Only a subset of the features sent to the model is considered to create the splits. In one embodiment where the subset of features includes "drug" and "molecule", a particular leaf may indicate "Variant effect on KRAS (somatic) protein (post-anchor):>1" (molecular feature) and "Do not take drug: pembrolizumab" (medical feature). Thus, analytics may be run on the data to improve the overall quality and accuracy of the splits and the resulting leaf nodes. In certain cases (although not relevant when drug and molecular features are used for splits), analytics may be used to parse branching information and make otherwise ambiguous information meaningful, i.e., information indicating "gender is not male" may be set to "gender is female".

[0203] In another case where segmentation involves models based on drugs and molecular features, analytics can be used to map the data into specific categories and / or ranges to make the data meaningful. For example, ranges can be presented as follows:

[0204] Variant effect on KRAS (somatic) protein (post-anchor): =>1

[0205] This can be mapped to:

[0206] Variant effect on KRAS (somatic) protein (post-anchor): = 1 ("negative")

[0207] Here, the term "negative" indicates "tested and confirmed to be not mutated" (as opposed to an unknown state).

[0208] In some embodiments, the analysis that leads to generating branches from nodes requires that all of the patients in the resulting leaf nodes meet certain requirements, i.e., the procedure may require 100% cohort participation to form a branch. However, in some cases, this requirement of 100% cohort participation may cause features derived from the tree to miss statistically relevant cohort features. Thus, in some embodiments, a subset-aware feature effect (SAFE) algorithm may be implemented to allow a particular leaf to include features that are shared by fewer than all patients in the leaf cohort (e.g., shared by 95%), but not all patients in the entire cohort (e.g., 95%).

[0209] In various embodiments, the smart cohort algorithm can be run in observational mode (using no predictions, only targets, e.g., 0 or 1) or algorithmic mode (using predictions, e.g., prediction-target[.95-1]).

[0210] The SAFE algorithm has been developed to return actionable feature importance ranks based on selected subpopulations of patients without the need for retraining the underlying model. Given predictions from a pre-trained global multi-cancer model for a patient population, the SAFE algorithm can interactively and quickly derive approximate high-level importance ranks. In addition, feature importance ranks can be intelligently and dynamically adjusted to be relevant given selected subset cohorts of the population without the need to retrain the global model. To optimize interpretability, in some embodiments, the SAFE feature importance algorithm can be made to explicitly handle assigning appropriate importance to correlated features, agnostic to the underlying machine learning model used. The SAFE algorithm can also provide the ability to explore feature importance in "feature + prediction" datasets where targets are not necessarily defined. Finally, for more continuous features, the SAFE algorithm can allow for deeper exploration of changes in feature importance with changes in feature values.

[0211] In one embodiment, the SAFE algorithm may include calculating a population mean prediction. The algorithm may then include encoding categorical feature levels as the difference between the predicted value and the population mean prediction, with rarely occurring levels grouped together. The algorithm may further include clustering or bucketing the continuous features and processing these features as in the previous step. Next, the algorithm may include aggregating the average value per category level (pE(p)) for each feature. Finally, the algorithm may include assigning an overall feature importance for each feature as a frequency-weighted sum of the absolute values ​​of all values.

[0212] As can be seen using the approach described above, the algorithm does not explicitly rely on the presence of a target variable to derive importance rankings, but instead requires only features and predictions. As such, it can be effectively applied to predictions made on unlabeled datasets, and further generalized to predictions obtained from different types of machine learning (ML) algorithms.

[0213] 27A and 27B show examples of adaptive feature ranking according to an embodiment of the SAFE algorithm. FIG. 27A shows a list of the top 10 features from an overall model based exclusively on breast cancer patients. FIG. 27B shows a list of the top 10 features from the dataset from FIG. 27A after creating a subset for colorectal stage 4 patients. As can be seen in FIG. 27B, some features that are likely to be associated with colorectal patients (e.g., "historical-took_medication:irinotecan" and "historical-took_medication:bevacizumab") have higher rankings and higher values ​​in the subset for colorectal stage 4 patients. On the other hand, features that are not relevant to colorectal stage 4 patients (e.g., "cancer:lung_cancer" and "cancer:pancreatic_cancer") do not appear in the list in FIG. 27B. FIG. 27C continues the example of FIG. 27A and FIG. 27B and shows an example of handling correlated features. Continuing with the colorectal example from FIG. 27B, FIG. 27C shows that after adding duplicate dummy columns based on two features, namely, “historical-took_medication:irinotecan” and “historical-took_medication:capecitabine”, these duplicate columns are properly sorted with the other values ​​associated with colorectal stage 4, as expected.

[0214] 27D and 27E show an example of sample-level importance assignment according to an embodiment of the SAFE algorithm. Given the derivation of the SAFE algorithm, one advantage is that each instance of each feature value is assigned an "impact" value that represents its co-occurrence with the observed deviation from the expected mean, which then allows for examining the variation in impact per change in feature value. FIG. 27D shows a box plot grouped according to the feature "historical-took_medication:irinotecan". FIG. 27E shows a box plot grouped according to final disease stage. FIG. 27D shows that features that co-occur with "historical-took_medication:irinotecan" with a value of 1 have a greater impact than features associated with a value of 0, as would be expected for a subset of colorectal stage 4. FIG. 27E shows the greater impact associated with later disease stages.

[0215] Although the SAFE algorithm does not directly take feature interactions into account, these values ​​can be derived from manually constructed composite features. In addition, the SAFE algorithm aims to convey how each feature impacts the prediction from the underlying model, which is used as an indirect proxy of feature importance for predicting the target, but this assumes the validity of the model.

[0216] Notebooks In various embodiments, one or more statistical models and analyses may be combined to address specific purposes and used to solve many problems through variations of the initial analysis. Such combinations of statistical models and analyses may be stored as notebooks in the interactive analysis portal 22. Notebooks are a feature of the interactive analysis portal 22 and provide an easily accessible framework for building statistical models and analyses. Once the statistical models and analyses are developed, they may then be shared with different users to analyze and find answers to scientific and business questions other than the one originally posed.

[0217] 1) The Interactive Analysis Portal 22 allows customization of inputs with a simple and intuitive point-and-click / drag-and-drop interface to narrow down cohorts for analysis. Cohorts selected through either the Interactive Analysis Portal 22, outliers, smart cohorts, or other portals of the Interactive Analysis Portal 22 can be sent to the notebook for processing.

[0218] 2) A custom application interface (API) with a library of function calls that interface with the interactive analytics portal 22, the underlying authorized databases, and any supported statistical models, visualizations, mathematical models, and other provided operations may be provided to users to integrate their notebooks or workbooks with the data, function calls, and other resources of the interactive analytics portal 22. Exemplary function calls may include listing authorized data sources, selecting a data source, filtering data sources, listing clinical events for patients in the current filtered cohort, identifying fusions from RNA or DNA, identifying genes from RNA or DNA, identifying matching clinical trials, identifying DNA variants, immunohistochemistry (IHC), identifying RNA expression, identifying therapies in a cohort, identifying potential therapies applicable to treat patients in a cohort, and other cohort or data set processing.

[0219] 3) The interactive analytics portal 22 allows the notebook generation to run one or more statistical models, analyses, and visualizations or reports of results on the refined cohort without having the user code anything within the notebook, as the selected model, analysis, visualization, or report in the notebook itself is configured to accept the cohort from the interactive analytics portal 22 and provide the analysis on the cohort as is, without user intervention at the code level. Some models may have hyperparameters or tuning parameters that may be selected, or the model itself may identify optimal parameters to be applied based on the cohort and / or other models, analyses, visualizations, or reports at run time.

[0220] 4) The interactive analysis portal 22 displays the prepared results to the user based on the selected notebook.

[0221] 5) The associated user may then select the notebook that has been generated prior to applying the selected analysis to the refined cohort without having the user code or re-code anything within the notebook, as the notebook itself is configured to accept the cohort from the interactive analysis portal 22 and provide notebook results without user intervention.

[0222] 6) Users may track the computational resources used by their notebooks to understand the cost of cloud computing or hardware resources on the network, and may track the popularity of their notebooks to determine the effectiveness of the statistical analysis they provide through their notebooks.

[0223] In some embodiments, the notebook benefits users by allowing the interactive analytics portal 22 to provide custom templates to the user's selected data and leverage pre-built healthcare statistical models to provide results to users who are not familiar with programming. In-house teams may analyze the curated data to support new healthcare insights that can help both improve patient care and improve life science research. Similarly, external users can easily access this proprietary real-world data for analysis and access the proprietary statistical models.

[0224] The billing model for users may be provided on a subscription or on-demand basis. For example, a user may subscribe to one or more data sets for a period of time, such as a monthly or yearly subscription, or the user may pay per access for the use of data and notebooks, such as paying a fee to load notebooks corresponding to specific cohorts and generate instant results for consumption. Users may also desire a benchmarking and optimization portal that can be used to review and optimize the use of storage and computing resources.

[0225] Generating the notebook may be performed in a GUI for editing notebooks. The user may configure a report page for the notebook. The report page may include text, images, and graphs that are selected and written by the user. Pre-configured elements may be selected from a list, such as a drop-down list or a drag-and-drop menu. Pre-configured elements include statistical analysis modules and machine learning models. For example, a user may want to run a linear regression on the data for a particular feature. When the user selects linear regression, a menu with checkboxes may be displayed with the features of the dataset that should be fed into the linear regression model. Once filled out, a template for reporting the results of the linear regression for the selected feature may be added to the report page at a location identified by the drop location of the active cursor or drag-and-drop element. If the user wants to use a machine learning model to solve the problem, the model may be added to the sheet. A header may be written that identifies the model, the hypertuning parameters, and the reported results. Then, in some cases, a previously trained model may be applied to the current cohort. In other cases, a model may be trained on the fly, for example, by selecting an outcome that is associated with the annotated features on which the model should be trained. In unsupervised machine learning models, the models may not require selection of annotated features because features are identified during training. In some embodiments, if the selected statistical model requires results from a trained model that is not computed in the template, the template may automatically add the trained model to generate the required results before inserting the selected statistical model into the notebook.

[0226] Statistical analysis models may be pre-designed to calculate the arithmetic mean of a cohort for a selected feature, the standard deviation / distribution of a cohort for a selected feature, regression relationships between variables for a selected feature, a sample size determination model to subset a cohort into optimal subpopulations for analysis, or a t-test module to identify statistically significant features and correlations in a cohort. Other pre-computed statistical analysis modules may perform cohort analysis to identify significant correlations and / or features within a cohort, data mining to identify meaningful patterns, or data dredging to match statistical models to data and report which models are applicable and add those models to the notebook.

[0227] The machine learning model may apply linear regression algorithms, nonlinear regression, logistic regression algorithms, classification models, bootstrap resampling models, subset selection models, dimensionality reduction models, tree-based models (such as bagging, boosting, random forests), and other supervised or unsupervised models. Once each model is selected, a target output may be requested from the user that specifies which features the model should identify, classify, and / or report. For example, the user may select a model that identifies which features are most closely correlated to patient survival in the cohort or which features are most closely correlated to positive treatment outcomes in the cohort. The user may also select which of the model's classification labels the user wants the model to classify. In an example where the model may classify a cohort according to five labels, the user may specify one or more labels as a binary classification (patient has label, patient does not have label), such as whether a patient with a tumor of unknown origin originates from the breast, lung, or brain. The user may select breast only to identify for any tumor of unknown origin whether the tumor may be classified as originating from the breast or not.

[0228] FIG. 29 illustrates an example user interface of the interactive analytics portal 22 for generating analytics via one or more notebooks according to one embodiment.

[0229] The notebook user interface 2900 may be accessed by selecting Notebook from the interactive analysis portal 22, such as via a sidebar menu 2910, either before or after filtering the patient database to a desired cohort of patients via the interactive cohort selection filtering process 24.

[0230] Notebooks, or workbooks, may be curated internally with company labels by team members knowledgeable in data science, machine learning, or other fields that routinely perform analytics on patient data, and presented to the user via the custom workbook widget 2920. The custom workbook widget may be presented as a searchable list, a searchable icon, a scrolling window that may scroll horizontally or vertically to display additional workbooks, or an expandable window that expands to provide access to all workbooks the user is authorized to access. Workbooks may be represented by icons and associated text, as illustrated in workbook 2960. Users may also create personalized workbooks that may be accessed via the my workbooks widget 2930. A workbook display window 2950 may be provided to display workbooks selected from the widgets 2920 or 2930. A new workbook may be created by the user by selecting blank workbook 2940. After selecting blank workbook 2940, a workbook creation interface may open.

[0231] FIG. 30 illustrates an example workbook creation interface of the interactive analytics portal 22 for creating a new workbook according to one embodiment.

[0232] A workbook creation interface 3000 may be provided to a user after selecting a blank workbook from the notebook user interface. A text entry user interface element (UIE) 3010 may be provided to name the workbook for identification, retrieval, and indexing after creation. A series of buttons and drop-down menu UIE 3020 may be provided to separate grouped elements of the user interface. The UIE 3020 may assist a user in building and structuring a presentation of a workbook. A cell UIE may provide selections related to a currently selected cell in a window 3040 having a block of code, such as a command to run the currently selected cell, a command to terminate the currently selected cell, a command to add a cell, a command to delete a cell, a command to run all cells, a command to run all cells above, a command to run all cells below, or a command to terminate all cells. The kernel UIE may provide selections related to one or more programming languages ​​and / or languages ​​available to the user, such as Python, Structured Query Language (SQL), R, Spark, Haskell, Ruby, Typescript, Javascript, Perl, Lua, C, C++, Matlab, Java, Emu86, and other kernels. Selecting a kernel from the kernel UIE reloads the workbook so that the cell executes commands from the respective language. The widget UIE may provide selections related to one or more code snippets supported for the active kernel. The code snippets may include code for creating visualizations such as graphs or plots, code for simple arithmetic operations such as calculating the mean or standard deviation, or code for more complex operations such as calculating a distribution and displaying the respective curves. A set of icon UIE 3030 may be provided, each icon representing a popular command executed from the UIE 3020.Exemplary popular commands may include saving the document, adding a new cell, cutting or pasting code or a cell, rearranging a cell by moving it up or down the page relative to any other cell, or executing / terminating code in the active cell.

[0233] One or more cells may be present in the window 3040 into which the user may insert one or more lines of code for the active kernel. The user may enter code or commands into the cells that may operate on the active database or patient cohort. Executing the cell executes the entered code or command. Any output, such as standard output, error messages, print statements, etc., is displayed directly below the cell after execution. Additionally, text widgets may be inserted that provide formatting and associated text based on the code from one or more cells. Such text widgets may provide a simple, easy-to-read format for the results of executing the code. In one embodiment, the text widgets may be presented as markdown cells that support HTML, indented lists, text formatting, TeX / LaTeX formulas, and inline tables.

[0234] In one example, a code block may perform arithmetic operations on a matrix of values. The associated output, such as printing a matrix, results in a series of brackets, parentheses, and commas that are difficult to understand. A visualization widget accepts a variable containing a matrix and provides an image where the matrix values ​​are visible in a displayable table format that represents the matrix instead of a potentially confusing text output. Cells accept all commands associated with each supported kernel and programming language. Cells may import modules or libraries from other sources (such as dask, fastparaquet, pandas, or other libraries), support data structures, support conditional statements and logical loops, and even establish and invoke functions. Cell output is generated asynchronously with the execution of the code so that the user can see instantaneous output from active code. If the output exceeds a preconfigured limit on the number of visible rows, the output becomes scrollable text and may auto-scroll with new entries or scroll after user input.

[0235] One or more templates may be provided in the template window 3050 for the convenience of the user. A template may include one or more cells that are pre-configured to operate on input data, such as a filtered patient cohort, execute one or more cells of code to generate logical results, and execute one or more cells of text or visualization to report the results of the executed logic on the input data in an easy-to-use manner. There may be templates for charts, graphs, regression, dimensionality reduction, classification, RNA or DNA normalization, and other commonly used features on the templates available to the user. Templates may be provided with the dataset or custom created by the user to be shared with other users.

[0236] FIG. 31 illustrates the action of opening a pre-configured template from a custom workbook widget in the notebook user interface.

[0237] Returning to the notebook user interface 2900, a user may populate the workbook display window 2950 with a custom workbook widget 2920 by clicking and dragging the desired workbook from the widget to the display window. In one example, a user may select a workbook 2960 with a mouse cursor and drag the workbook to the display window 2950, ​​as illustrated at 3120. Other intuitive mouse, keyboard, or gesture commands may be implemented instead of or in addition to clicking and dragging.

[0238] FIG. 32 illustrates the response from the notebook user interface when a user drags a workbook into the display window.

[0239] The notebook editor 3200 may auto-populate the title 3210 and one or more cells 3240A-D based on the workbook selected by the user. The user may further use the text entry UIE 3220 to rename the workbook using Edit Workbook. The user may change the configuration of the workbook via a series of buttons, and a drop-down menu UIE 3220 may be provided to separate grouped elements of the user interface. The UIE 3220 may assist the user in building and structuring the presentation of the workbook. The cell UIE may provide selections related to the currently selected cell 3240A-D with a block of code, such as a command to run the currently selected cell, a command to terminate the currently selected cell, a command to add a cell, a command to remove a cell, a command to run all cells, a command to run all cells above, a command to run all cells below, or a command to terminate all cells. The kernel UIE may provide selections related to one or more programming languages ​​and / or languages ​​available to the user, such as Python, Structured Query Language (SQL), R, Spark, Haskell, Ruby, Typescript, Javascript, Perl, Lua, C, C++, Matlab, Java, Emu86, and other kernels. Selecting a kernel from the kernel UIE reloads the workbook so that cells execute commands from the respective language. The widget UIE may provide selections related to one or more code snippets supported for the active kernel. The code snippets may include code for creating visualizations such as graphs or plots, code for simple arithmetic operations such as calculating the mean or standard deviation, or code for more complex operations such as calculating a distribution and displaying the respective curves. The user may further modify the configuration of the workbook via a set of icons UIE 3230 that may be provided, each icon representing a popular command executed from the UIE 3220.Exemplary popular commands may include saving the document, adding a new cell, cutting or pasting code or a cell, rearranging a cell by moving it up or down the page relative to any other cell, or executing / terminating code in the active cell.

[0240] A user may edit the source code of each of cells 3240A-D by selecting the cell and selecting the cell UIE option for editing or pressing the associated keyboard shortcut.

[0241] FIG. 33 illustrates a cell edit view of a custom workbook after a user loads the workbook into the workbook editor 3300 and selects edit from the cell UIE.

[0242] After entering the cell edit view of the workbook with cells 3240A-D, cells 3310A and 3310B become visible (3310C-D not shown). Cell 3310A displays code to generate a survival curve 3240A based on the trend difference between the control and treated cohorts of patients. Cell 3310B displays code to generate a scatter plot 3240B (not shown) based on normalized RNA expression for two selected RNA transcriptomes in the filtered cohort of patients. Similar cells 3310C-D (not shown) can be generated for the scatter plot and box plot 3240C-D (not shown), respectively.

[0243] Users can edit the code to modify the workbook for their own purposes, and even add or remove cells to create new customized workbooks.

[0244] In the cell editing view, the user may see that one or more templates may be provided in the template window 3050 for the user's convenience. The template may include one or more cells that are pre-configured to operate on input data, such as a filtered patient cohort, execute one or more cells of code to generate logical results, and execute one or more cells of text or visualization to report the results of the executed logic on the input data in an easy-to-use manner. There may be templates for charts, graphs, regression, dimensionality reduction, classification, RNA or DNA normalization, and other commonly used features on the templates available to the user. The templates may be provided with the dataset or may be custom created by the user to be shared with other users.

[0245] A user may drag any template into a cell to populate the cell with code for generating the template's associated visualization or mathematical operation.

[0246] A user may access a user interface of patient databases provided to the user by association with an institution or medical facility that has a subscription to each patient database. Custom workbooks may also be provided for each database where a workbook is selected for applicability to patients in each database. Accessing the user interface may generate resources within the cloud computing environment that can access authorized databases and / or workbooks. The user's resource usage within the cloud computing environment may be monitored and tracked to facilitate accurate billing for resources consumed by the user. The user may request and purchase other databases of patients. Patient databases may be purchased based on characteristics of the patients therein. For example, a user may desire a database of patients diagnosed with breast cancer. A look-up table (LUT) or cancer ontology is consulted to obtain alternative matches for breast cancer, such as ductal carcinoma, breast cancer, breast cancer, breast epithelial carcinoma, or other related terms. Patients that meet the requested diagnosis and any of the alternative terms from the LUT or cancer ontology may be compiled into a database and delivered to the user. The user may then perform statistical analysis and research on the data in accordance with the disclosure herein.

[0247] Other web interfaces may be incorporated into the interactive analytics portal 22 similar to the outlier portal, smart cohort portal, and notebook portal described above. One such other web interface may include using propensity scoring to identify the effect of a therapy, procedure, clinical trial, or other medical event on a patient's medical condition. Propensity scoring and related web interfaces are described in further detail in U.S. Patent Application No. 16 / 679,054, filed November 8, 2019, entitled "Evaluating Effect of Event on Condition Using Propensity Scoring," which is incorporated herein by reference in its entirety.

[0248] 34 illustrates an example machine of a computer system 3400 upon which a set of instructions may be executed to cause the machine to perform one or more of the methods described herein. In alternative implementations, the machine may be connected (e.g., networked) to other machines in a LAN, an intranet, an extranet, and / or the Internet.

[0249] The machine may operate in the capacity of a server or a client machine in a client-server network environment, or as a peer machine in a peer-to-peer (or distributed) network environment, or as a server or client machine in a cloud computing infrastructure or environment. The machine may be a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a mobile phone, a web appliance, a server, a network router, a switch or bridge, or any machine capable of executing a set of instructions (sequentially or otherwise) that specify actions to be performed by the machine. Moreover, although a single machine is illustrated, the term "machine" shall be taken to include any collection of machines that individually or in concert execute a set of instructions (or multiple sets of instructions) to perform any one or more of the methodologies described herein.

[0250] The exemplary computer system 3400 includes a processing device 3402, a main memory 3404 (read-only memory (ROM), flash memory, dynamic random access memory (DRAM), such as synchronous DRAM (SDRAM) or DRAM, etc.), a static memory 3406 (flash memory, static random access memory (SRAM), etc.), and a data storage device 3418, which communicate with each other via a bus 3430.

[0251] The processing device 3402 represents one or more general-purpose processing devices, such as a microprocessor, a central processing unit, or the like. More specifically, the processing device may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, or a processor implementing other instruction sets or a processor implementing a combination of instruction sets. The processing device 3402 may also be one or more special-purpose processing devices, such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a network processor, or the like. The processing device 3402 is configured to execute instructions 3422 to perform the operations and steps described herein.

[0252] The computer system 3400 may further include a network interface device 3408 for connecting to a LAN, an intranet, the Internet, and / or an extranet. The computer system 3400 may also include a video display unit 3410 (such as a liquid crystal display (LCD) or a cathode ray tube (CRT)), an alphanumeric input device 3412 (such as a keyboard), a cursor control device 3414 (such as a mouse), a signal generation device 3416 (such as a speaker), and a graphic processing unit 3424 (such as a graphics card).

[0253] The data storage device 3418 may be a machine-readable storage medium (also known as a computer-readable medium) on which is stored one or more sets of instructions or software 3422 that employs any one or more of the methods or functions described herein. The instructions 3422 may also reside, completely or at least partially, within the main memory 3404 and / or the processing device 3402 during execution by the computer system 3400, with the main memory 3404 and the processing device 3402 also constituting machine-readable media.

[0254] In one embodiment, the instructions 3422 include a software library including instructions for an interactive analytics portal (such as interactive analytics portal 22 of FIG. 1) and / or methods for functioning as an interactive analytics portal. The instructions 3422 may further include instructions for a patient filtering module 3426 (such as interactive cohort selection filtering interface 24 of FIG. 1) and a patient analytics module 3428 (such as cohort funnel and population analysis interface 26, patient timeline analysis user interface 28, patient survival analysis user interface 30, and / or patient event likelihood analysis user interface 32 of FIG. 1). Although the data storage device 3148 / machine readable storage medium is shown in the exemplary implementation as being a single medium, the term "machine readable storage medium" should be interpreted as including a single medium or multiple media (such as a centralized or distributed database, and / or associated caches and servers) that store one or more sets of instructions. The term "machine-readable storage medium" should also be construed as including any medium capable of storing or encoding a set of instructions executed by a machine and causing a machine to perform any one or more of the methods of the present disclosure. The term "machine-readable storage medium" should be appropriately construed as including, but not limited to, solid-state memory, optical media, and magnetic media. The term "machine-readable storage medium" should appropriately exclude temporary storage media such as signals, unless otherwise specified by identifying the machine-readable storage medium as a temporary storage medium or a temporary machine-readable storage medium.

[0255] In another implementation, the virtual machine 3440 may include modules for executing instructions for the patient filtering module 3426 (such as the interactive cohort selection filtering interface 24 of FIG. 1 ) and the patient analytics module 3428 (such as the cohort funnel and population analysis interface 26, the patient timeline analysis user interface 28, the patient survival analysis user interface 30, and / or the patient event likelihood analysis user interface 32 of FIG. 1 ). In computing, a virtual machine (VM) is an emulation of a computer system. Virtual machines are based on computer architectures and provide the functionality of a physical computer. Their implementation may involve dedicated hardware, software, or a combination of hardware and software.

[0256] Some portions of the preceding detailed description are presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by data processing engineers to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, contemplated to be a self-consistent sequence of operations leading to a desired result. These operations are those requiring physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical and magnetic signals that can be stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient at times, principally for reasons of common usage, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like.

[0257] It should be remembered, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. As is apparent from the above description, unless otherwise noted, throughout the description, descriptions utilizing terms such as "identifying" or "providing" or "calculating" or "determining" or similar terms will be understood to refer to the actions and processes of a computer system, or similar electronic computing device, that manipulates and converts data represented as physical (electronic) quantities in the computer system's registers and memory into other data that are similarly represented as physical quantities in the computer system's memory or registers or other such information storage devices.

[0258] The present disclosure also relates to an apparatus for performing the operations herein. This apparatus may be specially constructed for the intended purposes or may comprise a general-purpose computer selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a computer-readable storage medium, such as any type of disk, including, but not limited to, a floppy disk, an optical disk, a CD-ROM, and a magneto-optical disk, a read-only memory (ROM), a random access memory (RAM), an EPROM, an EEPROM, a magnetic or optical card, or any type of medium suitable for storing electronic instructions, each of which is coupled to a computer bus.

[0259] The algorithms and displays presented herein are not inherently related to a particular computer or other apparatus. Various general-purpose systems may be used with programs in accordance with the teachings herein, or it may prove convenient to construct a more specialized apparatus to perform the methods. A structure for a variety of these systems may appear as set forth in the description that follows. In addition, the present disclosure has not been described with reference to any particular programming language. It will be understood that a variety of programming languages ​​may be used to implement the teachings of the present disclosure as described herein.

[0260] The present disclosure may be provided as a computer program product or software, which may include a machine-readable medium having instructions stored thereon that may be used to program a computer system (or other electronic device) to perform a process according to the present disclosure. A machine-readable storage medium comprises a mechanism for storing information in a form readable by a machine (such as a computer). For example, machine-readable (such as a computer-readable) medium includes read-only memory ("ROM"), random access memory ("RAM"), magnetic disk storage media, optical storage media, flash memory devices, and the like.

[0261] In the foregoing specification, implementations of the present disclosure have been described with reference to certain exemplary implementations. It will be apparent that various modifications may be made thereto without departing from the broader spirit and scope of the implementations of the present disclosure as defined in the following claims. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense.

[0262] It will be apparent to those skilled in the art that numerous changes and modifications can be made in the specific embodiments of the invention described above without departing from the scope of the invention, and the whole of the foregoing description is therefore to be interpreted in an illustrative rather than a restrictive sense. [Explanation of symbols]

[0263] 10. System 12 Backend Layer 14 Patient Data Store 16 Patient Cohort Selector Module 18 Patient Cohort Timeline Data Storage 20 Front-end layer 22 Interactive Analysis Portal 24 Interactive Cohort Selection Filtering Interface 26 Cohort Funnel and Population Analysis Interface, Cohort Funnel and Population Analysis User Interface 28 Patient Timeline Analysis User Interface 30 Patient Survival Analysis User Interface 32 Patient Event Likelihood Analysis User Interface 34 Patient-specific Analysis User Interface 36 Patient Future Analysis User Interface 38 Distributed Computing and Modeling Layer 40 Time to Event Modeling Module 42 Event Likelihood Module 204 Project name (may reflect a database that stores a list of patients) 206 Gender 208 Race 210 Cancer site 212, 214 Cancer name 216 Tumor site 218 Stage 220 M stage 222 Name 224 ingredients 226 Sequencing 228 MSI (microsatellite instability) status 230 Procedure 232 Event Name 234 Cause of death 236 "Ask Gene" tab 238 Dialog Box 242 Patient Demographics, Filterable Options 244 Cancer site 246 Tumor Characterization 248 Molecular Data 250 organizations 252 Stage 254 Degree-Based Options 256 variant calls 258 Abstracted Variants 260 MSI 262 TMB 264 Interactive Funnel Chart 266 Summary 268 "Analyze Cohort" option 300 areas 300a First Area 300b Second Area 300c Third Area, Data Summary Window 300d The Fourth Region 302 Mean patient age, terminated events 304 Plot of patient age, time 306 Gender information, start event 308 Drugs, Drug Information, End Events 310 Genomic variants or alterations, gender 312 Control Panel, Gene 314 Gender, organization 316 Organization, Prescription Plan 318 Menopausal status, smoking status 320 Response, Stage 322 Smoking status, surgical procedures 324 Disease stage, specific drug regimen 326 Surgical Procedures 328 Response Data 330 Legend, Survival Plot 332 Drop-down menu, survival plot 334 Legend 336, 338 Confidence interval 400 1st area, 1st panel 402 Indicator 406 Radar Plot 408 Center Indicator, Center Point or Indicator 410 Second Area 412 Control Panel 414 Overlay, Outlier Threshold 416, 418 Value 420 Plot 422 x-axis 434 y-axis 436 Largest patient group, indicator 438 Largest Outlier Group 440 The Third Region 500 binary tree 2900 Notebook User Interface 2910 Sidebar Menu 2920 Custom Workbook Widget 2930 My Workbooks Widget 2940 Blank Workbook 2950 Workbook Display Window 2960 Workbook 3000 Workbook Generation Interface 3010 Text Input User Interface Elements (UIE) 3020 Button and Dropdown Menu UIE 3030 Icon UIE 3040 Window 3050 Template Window 3200 Notebook Editor 3210 Titles 3220 Text input UIE, drop-down menu UIE 3240AD Cell, Survival Curve 3240B Cell, Scatter 3240C Cell, Scatter and Box Plots 3240D Cell, Scatter and Box Plots 3300 Workbook Editor 3310A, 3310B, 3310C, 3310D Cells 3400 Computer Systems 3402 Processing Device 3404 Main Memory 3406 Static Memory 3408 Network Interface Device 3410 Video Display Unit 3412 Alphanumeric Input Device 3414 Cursor Control Device 3416 Signal Generating Device 3418 Data storage devices 3422 Instructions, one or more instruction sets or software 3424 Graphic Processing Unit 3426 Patient Filtering Module 3428 Patient Analytics Module 3430 Bus 3440 Virtual Machines

Claims

1. A method implemented by a computer system for implementing a predictive model for predicting response, progression, or survival for a plurality of subjects, comprising: receiving a plurality of data for a plurality of subjects, the plurality of data spanning a period of time, the plurality of subjects being subjects having the plurality of data for training the predictive model, and a plurality of first features, a plurality of second features and their respective states for generating a set of predictions; aligning timelines associated with the plurality of data with respect to a common anchor point shared by the plurality of subjects; identifying, for each of the plurality of subjects, a plurality of time points of interest within the time period, the plurality of time points of interest including at least one time point after the common anchor point, each time point of interest being associated with a respective datum of the plurality of data; For each subject of the plurality of subjects, and for each subject time point of the plurality of subject time points, based on the plurality of data for the plurality of subjects, determining whether an outcome target for an outcome event occurs within a horizontal time window measured from the subject time point, the outcome event relating to a subject and / or disease state, and the outcome target indicating whether the outcome event will occur; generating a first feature set comprising a plurality of first features, the plurality of first features comprising temporal features determined relative to the time points of interest; determining a state of each of the plurality of first features at the time of interest; For each time point of interest among the plurality of time points of interest having a valid outcome target, and for each combination of a horizontal time window and an outcome event, identifying a plurality of second features, the plurality of second features including temporal features determined relative to the time point of interest; determining a state of each of the plurality of second features for each combination of horizontal time window and outcome event at each time point of interest having a valid outcome target; and generating a plurality of sets of predictions for the plurality of objects using the predictive model generated using machine learning based on the plurality of data including the plurality of first features, the plurality of second features, and their respective states.

2. 1) selecting a cohort of subjects comprising a group of subjects of said plurality of subjects; 2) identifying a common anchor time point from a set of anchor points associated with each of the groups of subjects, the common anchor point being shared by each of the groups of subjects within the cohort; 3) for each object of the group of objects, aligning a timeline associated with each object of the group of objects to the common anchor point; 4) identifying an outcome target; 5) retrieving, for each subject of the group of subjects and for each of the plurality of second features and the plurality of first features, the generated plurality of sets of predictions each including a predicted target value; 6) generating a classification model based on the plurality of sets of predictions, each set including a predicted target value.

3. The classification model includes a plurality of decision trees, and the method further comprises: 7) generating the plurality of decision trees, for each of the plurality of decision trees, a) for each of the plurality of second features and the plurality of first features, i) dividing the group of subjects into a first subgroup and a second subgroup based on a difference between the predicted target value and an actual target value; ii) determining the difference between the number of subjects in the first subgroup and the number of subjects in the second subgroup; 3. The method of claim 2, further comprising the step of: b) selecting the feature that results in a difference that is the largest difference between the number of subjects in the first subgroup and the number of subjects in the second subgroup.

4. 8) creating a new node of a tree structure based on the feature that results in the largest difference between the number of objects in the first subgroup and the number of objects in the second subgroup; 9) creating a first branch from the new node based on the first subgroup; 10) creating a second branch from the new node based on the second subgroup; 11) for each of the first and second branches, repeating steps 7) a) i-ii) or 7) b) and 8) - 10) based on the subjects in the first subgroup and the second subgroup, respectively: The maximum number of nodes or branches have already been created, or and repeating until either the node contains less than the minimum number of objects.

5. the classification model includes a plurality of models, each of which includes a plurality of terminal nodes; The method comprises: For each subject in said group of subjects, identifying co-occurrences of the object occurring in each of the plurality of terminal nodes with each other of the group of objects; determining a similarity metric for the object based on the sum of co-occurrences divided by the number of models; and generating a database of subject-to-subject similarity metrics based on determining the similarity metric for each of the plurality of subjects.

6. The method of claim 5 , further comprising displaying the similarity metric for each subject as a cluster plot.

7. 4. The method of claim 3, further comprising displaying data associated with one or more of steps 7) a) i-ii) or 7) b) to identify at least one of the plurality of first features and the plurality of second features.

8. determining a similarity metric for a new object that is different from said group of objects; matching the new object with a subgroup of objects corresponding to a particular one of the plurality of terminal nodes based on the determined similarity metric; and identifying a treatment for the new subject based on matching the new subject to the subgroup of subjects.

9. 6. The method of claim 5, further comprising processing the database of subject-subject similarity metrics using a dimensionality reduction algorithm to identify specific cohorts of subjects that have shared characteristics.

10. The method of claim 9 , wherein the shared characteristic comprises at least one of a first shared characteristic or a second shared characteristic.

11. The step of selecting a cohort of subjects comprises:

3. The method of claim 2, further comprising the step of selecting a cohort of subjects having a particular condition.

12. The method of claim 11 , wherein the particular condition comprises a disease.

13. 3. The method of claim 2, wherein the common anchor point comprises a time associated with at least one of a diagnosis, treatment, or metastasis event.

14. 1. A method for implementing a predictive model for predicting response, progression, or survival for a plurality of subjects, the method comprising: receiving a plurality of data for a plurality of subjects over a period of time, the subjects being subjects having the plurality of data for training the predictive model, and a plurality of first features, a plurality of second features and their respective states for generating a set of a plurality of predictions; aligning a timeline associated with the plurality of data to a common anchor point shared by the plurality of subjects; identifying, for each of the plurality of subjects, a plurality of time points of interest within the time period, the plurality of time points of interest including at least one time point after the common anchor point, each time point of interest being associated with a respective datum of the plurality of data; For each subject of the plurality of subjects, and for each subject time point of the plurality of subject time points, based on the plurality of data for the plurality of subjects, determining whether an outcome target for an outcome event occurs within a horizon time window measured from the subject time point, the outcome event relating to a subject and / or disease state, and the outcome target indicating whether the outcome event will occur; identifying a plurality of first features, the first features including temporal features determined relative to the time points of interest; determining a state of each of the plurality of first features at the time of interest; For each time point of interest among the plurality of time points of interest having a valid outcome target, and for each combination of a horizontal time window and an outcome event, identifying a plurality of second features, the plurality of second features including temporal features determined relative to the time point of interest; determining a state of each of the plurality of second features for each combination of horizontal time window and outcome event at each time point of interest having a valid outcome target; generating a plurality of sets of predictions for the plurality of objects using the predictive model generated using machine learning based on the plurality of data including the plurality of first features, the plurality of second features, and their respective states.

15. generating the predictive model using machine learning based on the plurality of data, 15. The method of claim 1 or 14, further comprising generating the predictive model based on the plurality of data using gradient boosting.

16. the sets of predictions for the plurality of subjects are divided into a plurality of folds, each fold including data corresponding to a subset of the plurality of subjects; generating the predictive model using gradient boosting based on the plurality of data, 16. The method of claim 15, further comprising generating the predictive model using gradient boosting based on a subset of the plurality of folds.

17. receiving a plurality of forecasts from the forecasting model corresponding to each of the plurality of time points of interest; and associating each of the plurality of predictions with a corresponding outcome target.

18. receiving the plurality of predictions, an outcome target, a subset of the plurality of second features corresponding to the outcome target, and a cohort of subjects comprising a subset of the plurality of subjects; receiving an anchor point; providing, for each subject in the cohort having the anchor points, the predictive model having a selected subset of the plurality of second features and a difference between each of the plurality of predictions and the outcome target; 20. The method of claim 17, further comprising: generating a classification model based on determining, for each feature of the selected subset of the plurality of second features, a maximum difference between each of the plurality of predictions and the outcome target.

19. the classification model comprises a decision tree; The method further includes generating the decision tree based on determining the maximum difference between each of the plurality of predictions and the outcome target; the decision tree includes a plurality of leaf nodes and one or more branch nodes; each of the one or more branch nodes includes a pair of branches each including a leaf node or a branch node; 20. The method of claim 18, wherein each of the plurality of leaf nodes of the decision tree includes a number of subjects from the cohort of subjects.

20. generating the decision tree based on determining the maximum difference between each of the plurality of predictions and the outcome target, 20. The method of claim 19, further comprising generating the decision tree until a number of objects in a leaf node of the plurality of leaf nodes is less than a minimum number of objects.

21. generating the decision tree based on determining the maximum difference between each of the plurality of predictions and the outcome target, 20. The method of claim 19, further comprising generating the decision tree until the number of levels in the decision tree is equal to a maximum number of levels.

22. each of the one or more branch nodes includes a pair of branches each including a leaf node or a branch node; The method of claim 19 , wherein the branches are formed based on features selected from the subset of the plurality of second features.

Citation Information

Patent Citations

  • PCT/US19/56713

  • US10,395,772

  • Mobile supplementation, extraction, and analysis of health records

    US10395772B1

  • Evaluating effect of event on condition using propensity scoring

    US11715565B2

  • System and method for predicting long-term patient outcome

    US20120041772A1