Techniques for generating predictive outcomes for oncological treatment lines using artificial intelligence
AI-based methods predict cancer treatment outcomes and side effects by analyzing genetic profiles and validating treatment rationale, addressing the challenges of complex mutations and guideline compliance in oncology.
Patent Information
- Application Number
- JP2023533883
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-12-07
- Filing Date
- 2021-10-06
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2041-10-06
AI Technical Summary
Current oncological treatment selection for cancer is challenging due to complex genetic mutations, individualized toxicity levels, and the need to ensure treatment lines comply with guidelines, making it difficult to predict treatment outcomes and side effects effectively.
A computer-implemented method using AI to predict treatment outcomes and side effects by analyzing genetic mutation profiles and comparing them with trained similarity models, and validating treatment rationale against oncology guidelines.
Enhances personalized treatment selection and side effect prediction, ensuring compliance with established medical guidelines, thereby improving treatment efficacy and safety.
Smart Images

Figure 0007739428000003 
Figure 0007739428000004 
Figure 0007739428000005
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of and priority to European Patent Application No. 20212280.0, filed December 7, 2020, which is incorporated herein by reference in its entirety for all purposes.
[0002] Field The methods and systems disclosed herein generally relate to techniques for using artificial intelligence (AI) to facilitate the selection of a treatment line for a subject diagnosed with cancer. More specifically, the methods and systems disclosed herein relate to techniques for using AI to (1) predict a subject's treatment outcome and cancer progression based on the mutational profiles of other subjects across cancer types, (2) predict subject-specific side effects of candidate lines of therapy for treating cancer, and / or (3) automatically verify whether the reasons contributing to the selection of a particular treatment line (e.g., as represented by specific features in the subject record) comply with oncology treatment guidelines. [Background technology]
[0003] background Cancer is one of the leading causes of death worldwide. It can develop anywhere in the human body. However, there are several common locations where cancer can develop. For example, major cancer types include breast, lung, colon, and blood cancer. Regardless of the type, cancer involves the unrestrained division of some of the body's cells and can spread to other tissues around the body. In healthy individuals, cell division, which creates new cells, is generally balanced by the death of older or damaged cells. However, in individuals diagnosed with cancer, this balance is disrupted. Cancer causes the uncontrolled growth of abnormal cells in the body, even when new cells are not needed. The uncontrolled growth of abnormal cells can form tumors in the body's tissues. In some cases, abnormal cells can detach from tumors, travel through the body's bloodstream, and attach to tissues in new areas of the body, potentially forming new tumors.
[0004] The uncontrolled growth of these abnormal cells is caused by genetic mutations in cellular deoxyribonucleic acid (DNA). Genetic mutations are often caused by inherited genes. However, mutations can also be induced by environmental factors. For example, toxic exposures (e.g., exposure to carcinogens, radiation, and tobacco), lifestyle-related factors (e.g., obesity, diet, and alcohol intake), age, drugs, hormones, random chance, and certain infections (e.g., hepatitis, human papillomavirus (HPV), and Epstein-Barr virus) can cause cancer-associated genetic mutations in otherwise healthy individuals.
[0005] Oncology, the study and treatment of cancerous cells, presents several unique and significant challenges. First, a particular cancer can be caused by a complex combination of multiple mutations across different genes. Modern cancer research suggests that the progression of cancer pathways in a subject involves complex dependencies and interactions between multiple genetic mutations. Cancer often develops when a protein produced by one mutation interacts with a protein produced by another mutation. For example, in certain blood cancers, subjects fared much worse if the primary mutation JAK2 V617F (the driving mutation) was activated before the secondary mutation identified as TET2. Conversely, subjects with an activating TET2 mutation before the JAK2 V617F driving mutation had significantly better clinical outcomes. Furthermore, advances in genetic testing have enabled the identification of subject-specific molecular subsets, which can be assessed to select specific treatments based on their molecular characteristics. However, these advances have also presented many challenges, such as obtaining correct genotyping of tumor samples. Thus, identifying therapeutic lines for treating cancer is uniquely difficult compared to other diseases, since targeting a primary mutation, for example by gene replacement therapy, may activate or exacerbate the effects of a secondary mutation, which may exacerbate the cancer. Thus, isolating the cause of cancer can be particularly difficult.
[0006] Second, oncological treatment lines often involve levels of toxicity that can be harmful to the subject. For example, depending on the subject's unique risk factors, certain chemotherapy and immunosuppressant drugs can cause life-threatening side effects. Therefore, cancer treatment selection depends heavily on the individual's unique progression-free survival. Furthermore, there are a wide variety of side effects depending on the treatment line. Furthermore, treatment selection varies depending on the subject's subjective risk tolerance. For example, if a group of subjects with the same cancer at the same stage have a 15% three-year survival probability, subjects in the group may be willing to accept different aggressive treatments, with some of the group being willing to accept aggressive treatments such as high-dose radiation therapy, while a different portion of the group may be willing to accept only less aggressive treatments such as combination therapy. Therefore, treatment selection and side effect assessment are inherently difficult in oncological settings.
[0007] Third, certain lines of treatment require approval before they can be implemented. For example, a physician intending to administer gene replacement therapy to a subject may require prior approval if the therapy targets a mutation different from those commonly targeted by other therapies. Organizations such as the National Comprehensive Cancer Network (NCCN) and the American Society of Clinical Oncology (ASCO) have established guidelines for treating cancer. Identifying whether the reasons underlying the selection of a line of treatment for a subject conform to existing guidelines is challenging because it is difficult to identify the characteristics that contributed to treatment selection. In some cases, a literature review may be necessary. Because treatments are often selected using the treating physician's knowledge base, it is difficult to objectively identify the characteristics that contributed to treatment selection.
[0008] U.S. Patent Application Publication No. 2020 / 0370124 discloses a system and method for predicting the effectiveness of cancer treatment in a subject. The disclosed system and method are based on determining that the number, percentage, or ratio of specific types of single nucleotide variants (SNVs) in the nucleic acid of subjects with cancer who respond to treatment differs from that of subjects who do not respond to treatment. The SNVs identified in the nucleic acid molecules can be used to determine multiple metrics that form a profile, such that subjects who are likely to respond to cancer treatment typically have a different profile from subjects who are unlikely to respond to cancer treatment. The multiple metrics are then applied to a computational model, which is selected based on specific subject attributes. The computational model determines a therapeutic index, such as a numerical percentage, based on the multiple metrics, and the therapeutic index indicates predicted responsiveness to the cancer treatment.
[0009] Therefore, to improve the effectiveness of treatment for individual subjects diagnosed with cancer, there is a need for improved individual selection of lines of therapy for subjects diagnosed with cancer, individualized assessment of side effects, and verification that lines of therapy comply with existing guidelines. Summary of the Invention
[0010] overview In some embodiments, a computer-implemented method for predicting a subject-specific outcome of a line of oncological treatment is provided. The method can include identifying a particular subject diagnosed with a type of cancer and obtaining a genetic dataset corresponding to the particular subject. A line of treatment can be proposed to be administered to the particular subject. The genetic dataset can include a mutation profile, which can include molecular features of the subject's tumor, such as a molecular pattern, a mutation order (e.g., representing a series of multiple genetic mutations mutated at different times), etc. The computer-implemented method can also include identifying a set of other subjects diagnosed with the same type of cancer as the subject. Each of the other subjects can be receiving a line of treatment and can be associated with a treatment outcome. The computer-implemented method can also include obtaining another genetic dataset for each other subject in the set of other subjects. The other genetic dataset can include another mutation profile. The computer-implemented method can include, for each other subject in the set of other subjects, inputting the mutation profile of the particular subject and another mutation profile of the other subject into a trained similarity model. The trained similarity model can be trained to generate a similarity weight representing a predicted degree to which the mutation profile of the particular subject is similar to another mutation profile of the other subject. The computer-implemented method can include determining a predicted treatment outcome for administering a line of therapy to the particular subject based on the similarity weights output by the trained similarity model. If it is determined that at least one of the similarity weights output by the similarity model is within a threshold, the computer-implemented method can include identifying one of the other subjects based on the determination and assigning the treatment outcome of the identified other subject as the predicted treatment outcome for the particular subject.If it is determined that none of the similarity weights output by the similarity model are within the threshold, the computer-implemented method can include identifying another set of subjects diagnosed with a different type of cancer than the particular subject to search for mutation profiles similar to the mutation profile of the particular subject.
[0011] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more of the methods disclosed herein.
[0012] In some embodiments, a computer program product is provided that is tangibly embodied in a non-transitory machine-readable storage medium and includes instructions configured to cause one or more data processors to perform some or all of one or more of the methods disclosed herein.
[0013] Some embodiments of the present disclosure include a system including one or more processors. In some embodiments, the system includes a non-transitory computer-readable storage medium including instructions that, when executed on the one or more processors, cause the one or more processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium including instructions configured to cause one or more processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein.
[0014] The terms and expressions which have been employed are used as terms of description rather than of limitation, and there is no intention in the use of such terms and expressions to exclude any equivalents of the features shown and described, or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, although the claimed invention has been specifically disclosed by embodiments and optional features, it will be understood that modifications and variations of the concepts disclosed herein may be resorted to by those skilled in the art, and that such modifications and variations are deemed to be within the scope of the invention as defined by the appended claims. [Brief explanation of the drawings]
[0015] The present disclosure is described in conjunction with the accompanying drawings, in which: [Figure 1] 1 illustrates a network environment in which cloud-based applications are hosted, in accordance with some aspects of the present disclosure. [Figure 2] 1 is a flowchart illustrating an example of a process performed by a cloud-based application to deliver an abbreviated subject record to a user device in connection with a consultation broadcast requesting assistance with a subject's treatment, according to some aspects of the present disclosure. [Figure 3] 10 is a flowchart illustrating an example process for monitoring a user integration of a treatment plan definition (e.g., a decision tree or treatment workflow) and automatically updating the treatment plan definition based on the results of the monitoring, according to some aspects of the present disclosure. [Figure 4] 1 is a flowchart illustrating an example of a process for recommending a treatment for a subject, according to some aspects of the present disclosure. [Figure 5] 1 is a flowchart illustrating an example of a process for obfuscating query results to comply with data privacy regulations in accordance with certain aspects of the present disclosure. [Figure 6] 1 is a flowchart illustrating an example of a process for communicating with a user using a bot script, such as a chatbot, in accordance with some aspects of the present disclosure. [Figure 7] FIG. 1 is a block diagram illustrating an example of a network environment for deploying an AI model trained to facilitate subject-specific identification of treatments and treatment schedules for subjects diagnosed with cancer, according to some aspects of the present disclosure. [Figure 8] FIG. 1 is a block diagram illustrating an example of a network environment for deploying an AI model trained to predict treatment outcomes and cancer progression in subjects diagnosed with cancer, according to some embodiments of the present disclosure. [Figure 9] FIG. 1 is a block diagram illustrating an example of a network environment for deploying an AI model trained to predict subject-specific side effects of an oncology treatment line, according to some embodiments of the present disclosure. [Figure 10] FIG. 1 is a block diagram illustrating an example of a network environment for deploying an AI model trained to identify factors that contribute to the selection of a given line of treatment, according to some aspects of the present disclosure. [Figure 11] 1 is a flowchart illustrating an example of a process for predicting treatment outcome and cancer progression in a subject diagnosed with cancer, according to some embodiments of the present disclosure. [Figure 12] 1 is a flowchart illustrating an example of a process for predicting subject-specific side effects of a mutation-targeted treatment, according to some embodiments of the present disclosure. [Figure 13] 1 is a flowchart illustrating an example process for developing an AI model to identify factors that contribute to the selection of a given treatment, according to some aspects of the present disclosure.
[0016] In the accompanying drawings, similar components and / or features may have the same reference label. Furthermore, various components of the same type may be distinguished by a dash following the reference label and a second label that distinguishes between the similar components. When only a first reference label is used in this specification, the description is applicable to any of the similar components having the same first reference label, regardless of the second reference label. DETAILED DESCRIPTION OF THE INVENTION
[0017] Detailed Description I. Overview Cancer is a highly complex disease. It can develop anywhere in the human body. In some cases, cancer is hereditary, while in other cases, cancer can develop in response to environmental factors. Regardless of the origin of cancer development, there is often a complex combination of genetic mutations along the progression of the cancer pathway. For example, tumors consist of billions of cells, and each cell may individually harbor different mutations. Therefore, monitoring and responding to cancer progression is an extremely challenging task, as cancer cells can evolve or adapt to treatment lines.
[0018] In oncology settings, understanding the mechanisms underlying cancer typically involves frequently obtaining genetic data of cancerous cells to detect changes in the cells. Modern oncology practice uses genetic data to identify specific gene mutations and gene mutation sequences that contribute to cancerous cell growth. Mutational profiles can include molecular characteristics of tumors, such as the order in which individual gene mutations are activated (e.g., mutation order). In certain cases, cancer can develop after a specific group of gene mutations is activated according to the pattern indicated by the mutation profile. Therefore, using genetic data to facilitate mutation identification is beneficial. However, identifying the appropriate line of therapy for treating cancerous cells presents another complex consideration. Furthermore, identifying an oncological treatment line is particularly challenging due to the wide range of side effects and uncertainty of treatment outcomes exhibited across subjects diagnosed with cancer.
[0019] Certain aspects of the present disclosure relate to deploying AI models trained to perform tasks that solve complex cancer-specific problems. AI techniques can produce predictive results from dense or seemingly unrelated datasets to assist physicians in making clinical decisions when treating subjects diagnosed with cancer. Certain aspects of the present disclosure provide cloud-based oncology applications configured with AI systems capable of performing predictive functions. AI-based techniques can be used to learn patterns and correlations across complex datasets of various data types (e.g., structured datasets, unstructured datasets, streaming data) from different sources. While oncological diseases are characterized by complexity and uncertainty, certain aspects of the present disclosure relate to the execution of specialized AI models to facilitate the selection of a line of treatment in a manner that is contextual to the genetic profile of each individual subject.
[0020] Certain aspects of the present disclosure relate to AI systems configured to perform specific predictive functions, such as predicting treatment outcomes and subsequent cancer progression for individual subjects (e.g., patients) based on the subject's mutational profile across cancer types, predicting side effects specific to subjects who have responded to a line of treatment, and automatically validating the rationale for selecting a line of treatment (e.g., a particular targeted therapy for treating breast cancer) for an individual subject to comply with oncology guidelines.
[0021] Certain aspects of the present disclosure relate to a cloud-based oncology application configured to generate predictions of treatment outcomes for proposed lines of treatment administered to an individual subject. The predictions can be based on mutation profiles of subjects with the same cancer or different cancer types as the individual subject. For example, the mutation profile represents, among other molecular characteristics, the order in which genes mutate over time (e.g., mutation order or mutation pattern). The mutation profile can influence clinical decision-making regarding diagnosis and selection of treatment lines. Certain aspects of the present disclosure relate to executing a specialized similarity-based AI model trained to automatically identify when the mutation profile of a subject with, for example, breast cancer is similar to the mutation profile of another subject with lung cancer. For example, targeted therapy administered to a subject with lung cancer can be beneficial regarding the effectiveness of a particular line of treatment for the subject with breast cancer. The specialized similarity-based AI model can be trained based on a training dataset of pairs or mutation profiles of subjects with the same or different types of cancer (one mutation profile representing one subject and the other mutation profile representing the other subject). Each pair can be labeled as similar or dissimilar. A learning algorithm can be implemented to automatically learn which patterns exhibited by the mutation profiles are similar to one another. Once trained, the specialized similarity-based AI model can output a similarity weight, which is a value that represents the degree to which one mutation profile of a subject is similar to another mutation profile of another subject.
[0022] Certain aspects of the present disclosure also relate to a cloud-based oncology application configured to generate predictions of side effects for lines of therapy based on the context of characteristics of a particular subject. The oncology application can be used to build a graphical mapping between lines of therapy and various side effects associated with the lines of therapy. In some examples, the graphical mapping can represent an ontology describing types of lines of therapy, characteristics of each line of therapy (e.g., side effects, progression-free survival), and relationships between the lines of therapy and the characteristics. The graphical mapping can be stored as a knowledge graph that is accessed each time a user requests a subject-specific prediction of side effects for a line of therapy. When a user interacts with the oncology application to request a prediction of subject-specific side effects for a line of therapy, the oncology application can query the knowledge graph using subject characteristics for the particular subject. An inference engine can perform logical inference tasks to identify which treatments and / or side effects in the knowledge graph are logically related to the subject characteristics of the particular subject. The output of the inference engine represents the subject-specific side effects for the line of therapy. It will be understood that the present disclosure is not limited to mapping lines of therapy to their corresponding side effects. Progression free survival versus line of treatment or any other variable can be graphically mapped and stored in the knowledge graph as an ontology.
[0023] Certain aspects of the present disclosure also relate to a cloud-based oncology application configured to use AI-based algorithms to evaluate subject data for cancer subjects with specific cancer types and the treatments administered to those cancer subjects to automatically learn the reasons for each individual cancer subject's treatment. For example, the oncology application can automatically predict that the reason a particular lung cancer subject will be treated with a specific targeted therapy treatment is that the lung cancer subject has a driver mutation in the HER2 gene. The oncology application can then compare the predicted reasons for various treatments with a set of guidelines or rules established by reputable medical organizations such as NCCN and ASCO. If guidelines do not exist, the oncology application can also identify potential new guidelines based on treatments administered to target specific mutations, the corresponding treatment outcomes of those treatments, and the subject's progression-free survival after the treatment.
[0024] The applications (e.g., running locally on the device and / or using at least in part the results of computations performed on one or more remote and / or cloud servers) can be used (for example) by a subject with cancer and / or a caregiver caring for the subject with cancer. The present application may perform one or more operations disclosed herein. In some examples, one or more applications can facilitate communication between a subject with cancer and a caregiver. An oncology application is related to a tumor-specific treatment workflow, while in some implementations, applications can be related to other specific cancer types, such as a cloud-based breast cancer application, a cloud-based lung cancer application, a cloud-based colon cancer application, a cloud-based blood cancer application, etc. Each cancer-type-specific application can be distinguished from other applications based, for example, on variables the application makes available. Such communication can facilitate (for example) alerting a caregiver to abnormal symptoms and / or can facilitate telemedicine (which may be particularly beneficial, for example, when a subject or part of the community has a contagious disease, when a subject has mobility impairments, and / or when a subject is physically distant from a caregiver's facility).
[0025] II. Overview of Cancer Subtypes, Diagnostic Protocols, Related Medical Tests, Progression Assessment, and Available Treatments II.A. Causes of Cancer According to the World Health Organization, approximately one in six deaths can be attributed to cancer, making it the second leading cause of death worldwide. Cancer is a group of diseases characterized by the uncontrolled growth of abnormal cells in the body. This uncontrolled growth is caused by genetic changes, such as mutations in cellular DNA. While these mutations are often caused by inherited genetic predispositions or traits, other factors, including environmental / toxic exposures (e.g., exposure to carcinogens, radiation, and tobacco), lifestyle-related factors (e.g., obesity, diet, and alcohol consumption), age, drugs, hormones, chance and infectious diseases (e.g., hepatitis, HPV, and Epstein-Barr virus), can cause cancer-associated genetic changes in individuals. Despite advances in screening, diagnosis, and treatment, cancer rates are increasing as more people live longer and engage in causative lifestyle behaviors.
[0026] II.B. Cancer Type There are over 100 types of cancer, including cancers that form solid tumors, such as breast, skin, lung, colon, and prostate cancer, to name a few. According to the American Institute for Cancer Research, there were an estimated 18 million cases of cancer worldwide in 2018. Of these, 9.5 million cases were in men and 8.5 million cases were in women. Lung and breast cancer are the most common cancers worldwide, each contributing approximately 12.3% of the total number of new cases in 2018. While lung cancer was the most common cancer in men, breast cancer was the most common cancer in women worldwide. Colorectal cancer was the third most common cancer, with 1.8 million new cases in 2018, followed by prostate cancer as the fourth most common cancer, with over 1,275,000 new cases in 2018.
[0027] Cancer also includes blood or blood cancers that affect the production and function of blood cells. Examples include leukemia (e.g., acute leukemia, acute lymphocytic leukemia, acute myeloid leukemia, and chronic lymphocytic leukemia (CLL)), lymphoma (e.g., Hodgkin's disease or non-Hodgkin's disease lymphoma (e.g., diffuse anaplastic lymphoma kinase (ALK)-negative, large B-cell lymphoma (DLBCL); follicular lymphoma (FL); diffuse ALK-positive DLBCL; ALK-positive, ALK+ anaplastic large cell lymphoma (ALCL); acute myeloid lymphoma (AML)); and multiple myeloma.
[0028] II.B.1. Breast cancer Breast cancer is the most common invasive cancer in women, but it can also occur in men. Breast cancer often develops in cells from the lining of the milk ducts and the lobules that supply these ducts with milk. Cancers that develop in the ducts are known as ductal carcinoma, and cancers that develop in the lobules are known as lobular carcinoma. Although rare, inflammatory breast cancer is another type of breast cancer, accounting for approximately 1-5% of all breast cancers. These cancers can be broadly divided into subgroups according to specific biomarkers that have been established to predict response to treatment: (1) hormone receptor (ER+ and / or PR+)-positive and Her2-negative (Her2- breast cancer), (2) hormone receptor-positive (ER+ and / or PR+) and Her2-positive (Her2+) breast cancer, (3) hormone receptor-negative (ER-) and Her2-positive (Her2+) breast cancer, and (4) hormone receptor-negative (ER-) and Her2-negative (Her2-) (triple-negative) breast cancer.
[0029] II.B.1.i.Clinical symptoms Symptoms of breast cancer include a breast lump, bloody discharge from the nipple, thickening or swelling of the breast, breast pain, irritation or dimpling of the breast skin, redness or flaky skin on the nipple or breast, nipple pain, itching, change in breast color, or a breast rash.
[0030] II.B.1.ii. Diagnosis Although many clinical symptoms are associated with breast cancer, it is often identified by routine mammography screening. Breast cancer can be diagnosed by several tests, including mammograms, ultrasound, magnetic resonance imaging (MRI), and biopsies.
[0031] Genetic testing for mutations associated with increased risk of breast cancer (e.g., BRCA1 and BRCA2 mutations) can also be performed after breast cancer is diagnosed to determine the best course of treatment. Other diagnostic assays (e.g., the VENTANA Her2Dual ISH test (Roche, Basel, Switzerland)) can be used to identify HER2-positive breast cancers for targeted therapy with trastuzumab (Herceptin, Roche, Basel, Switzerland).
[0032] Breast cancer generally has four stages characterized by the medical community as follows:
[0033] Stage 0 is the earliest stage of breast cancer. At this stage, abnormal cells are present, but the cancer has not spread to other parts of the breast. This stage is often called carcinoma in situ or carcinoma in situ.
[0034] Stage 1 is the earliest stage of invasive breast cancer, meaning the cancer has grown or spread into nearby or surrounding breast tissue. The tumor is usually about 2 centimeters or less in size. At this stage, the cancer may or may not have spread to the lymph nodes.
[0035] Stage 2 also indicates invasive breast cancer, in which the tumor may have grown to about 5 centimeters, sometimes larger. The cancer may or may not have spread to the lymph nodes.
[0036] Stage 3 is the stage of invasive breast cancer where the cancer has usually spread to the lymph nodes. Inflammatory breast cancer begins at stage 3 because it involves the skin.
[0037] Stage 4 is often called "metastatic," meaning the cancer has spread beyond the breast and nearby lymph nodes to other parts of the body. II.B.1.iii. Subtyping
[0038] Once breast cancer is diagnosed, it is often subtyped based on the hormone receptors expressed by the tumor cells to determine the course of treatment. The four major female breast cancer subtypes, in order of prevalence, are as follows:
[0039] (1) hormone receptor (ER+ and / or PR+) positive and Her2 negative (Her2- breast cancer (luminal A breast cancer)), (2) hormone receptor negative (ER-) and Her2 negative (Her2-) (triple negative) breast cancer, (3) hormone receptor positive (ER+ and / or PR+) and Her2 positive (Her2+) breast cancer (luminal B breast cancer), and (4) hormone receptor negative (ER-) and Her2 positive (Her2+) breast cancer (HER2-enriched breast cancer).
[0040] II.B.1.iv. Treatment The standard of care for breast cancer is a multidisciplinary approach incorporating surgery, radiation therapy, and drug treatments. The standard of care for breast cancer is determined by both the disease (e.g., tumor, stage, and pace of disease) and patient characteristics (e.g., age, biomarker expression, and intrinsic phenotype). General guidance regarding treatment options is provided in NCCN guidelines (e.g., NCCN Clinical Practice Guidelines in Oncology, Breast Cancer, version 2.2016, National Comprehensive Cancer Network, 2016, pp. 1-202) and ESMO guidelines (e.g., Senkus, E., et al., Primary Breast Cancer: ESMO Clinical Practice Guidelines for diagnosis, treatment and follow-up. Annals of Oncology 2015; 26(Suppl. 5):v8-v30; and Cardoso F., et al., Locally recurrent or metastatic breast cancer: ESMO Clinical Practice Guidelines for diagnosis, treatment and follow-up. Annals of Oncology 2012; 23(Suppl. 7):vii11-vii19.).
[0041] II.B.1.iv.a. Early stage or nonmetastatic breast cancer The standard treatment for early stage or non-metastatic breast cancer is typically mastectomy or breast-conserving surgery followed by radiation therapy or systemic therapy.
[0042] If the subject is hormone receptor (ER+ and / or PR+) positive and Her2 negative (Her2-), endocrine therapy (e.g., tamoxifen, GnRH agonist, aromatase inhibitor) with or without chemotherapy can be administered. If chemotherapy is administered, its type and dosage are selected according to tumor burden and / or biomarker expression. Neoadjuvant therapy to reduce tumor burden before surgery can also be used. Exemplary neoadjuvant therapies include tamoxifen or aromatase inhibitor with or without chemotherapy.
[0043] If the subject is hormone receptor (ER+ and / or PR+) positive and Her2 positive (Her2+), hormone therapy and anti-Her2 therapy can be administered with or without chemotherapy.Exemplary treatments include trastuzumab (Herceptin® (Roche, Basel, Switzerland)), chemotherapy and tamoxifen, or aromatase inhibitor administration.Neoadjuvant therapy (for example, administration of trastuzumab or pertuzumab with chemotherapy) can also be used.
[0044] If the subject is hormone receptor negative (ER-) and Her2 positive (Her2+), anti-Her2 therapy and chemotherapy can be administered. Neoadjuvant therapy (e.g., administration of trastuzumab or pertuzumab with chemotherapy) can also be used.
[0045] If the subject is hormone receptor negative (ER-) and Her2 negative (Her2-), chemotherapy can be administered. Chemotherapy can also be administered as neoadjuvant therapy.
[0046] Numerous chemotherapeutic agents are available for the treatment of early-stage or non-metastatic breast cancer, including, but not limited to, cyclophosphamide (Cytoxan), docetaxel (Taxotere), paclitaxel (Taxol), doxorubicin (Adriamycin), epirubicin (Ellence), and methotrexate (Maxtrex), which can be administered as monotherapy or combination therapy. For example, for the treatment of Her2+ breast cancer, docetaxel, carboplatin, and trastuzumab can be administered in combination. Other examples include the administration of trastuzumab and paclitaxel, or the administration of doxorubicin and cyclophosphamide followed by the administration of paclitaxel and trastuzumab.
[0047] II.B.1.iv.b. Advanced or metastatic breast cancer The standard treatment for advanced or metastatic breast cancer is often surgery. In some cases, chemotherapy is administered before or after surgery. Radiation therapy and / or hormone therapy (for ER+ positive tumors) can be administered after surgery.
[0048] If the subject is hormone receptor (ER+ and / or PR+) positive and postmenopausal, hormone therapy can include tamoxifen, aromatase inhibitors (anstrozole, letrozole, or exemestane), cyclin-dependent kinase inhibitors (palbociclib), or fulvestrant (anti-estrogen therapy).
[0049] If the subject is hormone receptor (ER+ and / or PR+) positive and premenopausal, hormone therapy can include tamoxifen or LHRH agonist.Targeted therapy such as trastuzumab (Herceptin (Roche, Basel, Switzerland)), bevacizumab (Avastin® (Roche, Basel, Switzerland)), lapatinib, pertuzumab, mTOR inhibitor, T-DM1 (trastuzumab emtansine), or palbociclib and letrozole can also be administered.In some cases, if the subject is Her2+, (1) pertuzumab only, (2) trastuzumab and pertuzumab, (3) trastuzumab and chemotherapy, or (4) lapatinib and chemotherapy are administered to the subject as first-line therapy.In some cases, Avastin® is administered in combination with paclitaxel to treat HER2-negative breast cancer in patients who have not yet received chemotherapy for metastatic breast cancer.
[0050] Numerous chemotherapeutic agents are available for the treatment of advanced or metastatic breast cancer, including, but not limited to, capecitabine (Xeloda® (Roche, Basel, Switzerland)), gemcitabine (Cynzar), carboplatin (Paraplatin), cisplatin (Platinol), cyclophosphamide (C) (Cytoxen), docetaxel (T) (Taxotere), paclitaxel, (T) (Taxol), doxorubicin (A) (Adriamycin), epirubicin (E) (Ellence), eribulin (Halaven), 5-fluorouracil (5-FU, Adrsil), ixabepilone (Ixempra), liposomal doxorubicin (doxil), methotrexate (M) (Maxtrex), albumin-bound paclitaxel (Abraxane), and vinorelbine (Navelbine).
[0051] II.B.1.iv.c. Early stage or nonmetastatic breast cancer The standard of care for triple-negative breast cancer (TNBC) is determined by both disease (stage, pace of disease, etc.) and patient (age, comorbidities, symptoms, etc.) characteristics.
[0052] Patients with early-stage, potentially resectable locally advanced TNBC (i.e., no distant metastatic disease) are managed with locoregional therapy (surgical resection with or without radiation therapy) with or without systemic chemotherapy.
[0053] Surgical treatment can be breast-conserving (i.e., lumpectomy, which focuses on removing the primary tumor at the margins) or more extensive (i.e., mastectomy, which aims to completely remove all breast tissue). Radiation therapy is typically administered postoperatively to the breast / chest wall and / or regional lymph nodes to kill any microscopic cancer cells remaining after surgery. In breast-conserving surgery, radiation is directed to the remaining breast tissue and sometimes to regional lymph nodes (including axillary nodes). Even in cases of mastectomy, radiation may be administered if factors predict a high risk of local recurrence.
[0054] Depending on tumor and patient characteristics, chemotherapy can be administered in adjuvant (postoperative) or neoadjuvant (preoperative) settings. Further guidance for treating early and locally advanced TNBC is provided in Solin LJ., Clin Br Cancer.2009,9:96-100; Freedman GM, et al., Cancer.2009,115:946-951; Heemskerk-Gerritsen BAM, et al., Ann Surg Oncol.2007,14:3335-3344; and Kell MR, et al., MBJ.2007,334:437-438.
[0055] Systemic chemotherapy is the standard treatment for patients with metastatic TNBC, but no standard regimen or sequence exists. Cytotoxic chemotherapy options are similar to those for other subtypes. Single-agent cytotoxic chemotherapy agents, such as anthracyclines (e.g., doxorubicin, epirubicin), taxanes (e.g., paclitaxel, docetaxel), antimetabolites (e.g., capecitabine, gemcitabine), non-taxane microtubule inhibitors (e.g., vinorelbine, eribulin, exabepilone), platinum (e.g., cisplatin, carboplatin), and alkylating agents (e.g., cyclophosphamide), are generally considered primary options for patients with metastatic TNBC. However, combination chemotherapy regimens may be used in cases of aggressive disease and visceral involvement. Treatment may involve the sequential use of different single-agent treatments. Palliative surgery and radiation may be utilized as needed to manage local complications.
[0056] II.B.2. Colorectal cancer Colorectal cancer, also known as bowel cancer or colon cancer, is any cancer that affects the colon and / or rectum. Colorectal cancer begins in the large intestine (colon). Colon cancer typically affects older people, but can occur at any age. It usually begins as small, noncancerous clumps of cells called polyps that form on the lining of the colon. Over time, some of these polyps can become colon cancer.
[0057] II.B.2.i.Clinical symptoms Symptoms of colon cancer include rectal bleeding or bloody stools, cramps, gas, abdominal pain, persistent changes in bowel habits including diarrhea or constipation, weakness or fatigue, and unexplained weight loss. Many people with colon cancer do not experience symptoms in the early stages of the disease. If symptoms do occur, they are likely to vary depending on the size and location of the cancer in the large intestine.
[0058] II.B.2.ii. Diagnosis Physicians recommend screening tests for healthy subjects without signs or symptoms of colon cancer to look for signs of colon cancer or noncancerous colon polyps. Physicians generally recommend that individuals at average risk for colon cancer begin screening around age 50. Detecting colon cancer at its earliest stages offers the greatest opportunity for successful treatment.
[0059] In addition to a physical exam, one or more of the following tests may be used to diagnose colorectal cancer: colonoscopy, biopsy, molecular testing of the tumor, blood tests, computed tomography (CT or CAT) scan, MRI, proctoscopy, ultrasound, and X-ray. Often, if suspected colorectal cancer is found by any screening or diagnostic test, it is biopsied during colonoscopy.
[0060] If the biopsy indicates the presence of colon cancer, additional genetic testing may be performed to further classify the colon cancer. For example, alterations in any of the mismatch repair genes (MLH1, MSH2, MSH6, and PMS2) can be detected to identify subjects with Lynch syndrome, a genetic disorder that increases a person's risk of developing colon cancer.
[0061] The stages of colon cancer are characterized by the medical community as follows:
[0062] Stage 0 is the earliest stage of colon cancer. This stage is also known as carcinoma in situ or intramucosal (Tis). At this stage, the cancer has not grown beyond the inner lining (mucosa) of the colon or rectum.
[0063] Stage I is characterized by growth of the cancer through the muscularis mucosae into the submucosa and sometimes into the muscularis propria. It has not spread to nearby lymph nodes or distant sites.
[0064] Stage IIA is characterized by cancer growth into the outermost layers of the colon or rectum, but not through them. At this stage, the cancer has not spread to nearby lymph nodes or distant sites. Stage II colon cancer can be subdivided into three stages: Stage IIA - The cancer has spread to the serosa or outer colon wall but has not crossed that outer barrier. Stage IIB - The cancer has spread beyond the serosa but has not affected nearby organs. Stage IIC - The cancer has affected the serosa and nearby organs.
[0065] Stage III is characterized by cancer growth beyond the lining of the colon, affecting the lymph nodes. At this stage, despite lymph node involvement, the cancer has not yet spread to other organs in the body. This stage is further divided into three categories: IIIA-IIIC. Where the cancer falls into these categories depends on a complex combination of which layers of the colon wall are affected and how many lymph nodes are affected.
[0066] Stage IV is characterized by metastatic growth that has spread via the blood and lymph nodes to other organs in the body.
[0067] II.B.2.iii. Treatment The standard treatment for colon cancer depends on the stage of the colon cancer: Stages 0-III colon cancer are typically treated with surgery.
[0068] Treatment for stage 0 colon cancer is a polypectomy, which is usually performed during a colonoscopy. During this procedure, the doctor can remove all of the malignant cells. If the cells have affected a larger area, a resection can be performed during the colonoscopy.
[0069] For patients with stage I colon cancer, a partial colectomy is performed to remove the diseased area. This surgical procedure may involve rejoining the parts of the colon that are still healthy.
[0070] Stage II cancer is treated with surgery to remove the affected area. In some cases, chemotherapy may also be recommended. High-grade or abnormal cancer cells or tumors that have caused colon blockage or perforation may require further treatment. If the surgeon is unable to remove all of the cancer cells, radiation may also be recommended to kill any remaining cancer cells and reduce the risk of recurrence.
[0071] All stages of stage III colon cancer require surgery to remove the affected area. In some cases, chemotherapy and / or radiation therapy may be administered. In some cases, radiation therapy may also be recommended for patients who are not healthy enough for surgery or who may still have cancer cells in their body after surgery.
[0072] Patients with stage IV colon cancer may undergo surgery to remove small areas or metastases of the affected organ. However, in many cases, the areas are too large to be removed. Therefore, targeted therapy, usually combined with chemotherapy, is used to treat stage IV / metastatic cancer (mCRC).
[0073] While there is no single standard treatment for mCRC, common first-line treatment regimens include administration of a fluoropyrimidine (e.g., fluorouracil (5-FU) or capecitabine) in various combinations and schedules with irinotecan and / or oxaliplatin. Bevacizumab (Avastin®), cetuximab, or panitumumab may be combined with any of the first-line chemotherapy treatments, such as Xeloda. In some cases, maintenance therapy is administered. Administration of maintenance therapy depends on the choice of first-line chemotherapy, but is often a combination of a fluoropyrimidine and bevacizumab.
[0074] Second-line therapy can also be used. In addition to the treatments listed above, aflibercept or ramucirumab can be used in combination with FOLFIRI (fluorouracil + leucovorin + irinotecan), depending on the first-line therapy chosen.
[0075] Third-line therapy can also be used. For example, if the cancer is RAS wild-type and has not been previously treated with an EGFR antibody, cetuximab or panitumumab can be administered, optionally in combination with chemotherapy. Regorafenib or a combination of trifluridine and tipiracil can also be used as third-line therapy. In some cases, colorectal patients who are unlikely to respond to anti-EGFR monoclonal antibody therapy can be identified using the cobas® KRAS Mutation Test or the cobas® KRAS Mutation Test v2 (Roche, Basel, Switzerland), which detects mutations in codons 12, 13, and 61 of the KRAS gene in formalin-fixed, paraffin-embedded tissues from colorectal cancer patients.
[0076] II.B.3. Lung cancer Lung cancer typically begins in cells lining the bronchi and parts of the lungs (such as the bronchioles or alveoli). Approximately 80-85% of lung cancers are non-small cell lung cancer (NSCLC), which can be divided into the following subtypes: adenocarcinoma, squamous cell carcinoma, and large cell carcinoma. These subtypes are often grouped together as NSCLC because their treatment and prognosis are often similar. Approximately 10-15% of all lung cancers are small cell lung cancer (SCLC), which tends to grow and spread more quickly than NSCLC.
[0077] II.B.3.i.Clinical symptoms Symptoms of lung cancer include persistent cough, hemoptysis, chest pain, hoarseness, loss of appetite, unexplained weight loss, shortness of breath, fatigue, non-healing infections, and wheezing.
[0078] II.B.3.ii. Diagnosis Lung cancer can be detected using imaging tests (e.g., X-rays, CT scans, or MRIs), sputum cytology, and / or tissue biopsies. Biopsies can be performed using bronchoscopy, mediastinoscopy, or needle biopsies. Biopsy samples can also be obtained from lymph nodes or tissues where cancer may have spread, for example, from the liver.
[0079] Once a diagnosis of lung cancer is made, the type and stage of the lung cancer are determined. Staging tests can include imaging procedures that allow a doctor to determine whether the cancer has spread beyond the lungs. These tests include CT, MRI, positron emission tomography (PET), and bone scans.
[0080] Several diagnostic assays for stratifying and classifying lung cancer are available. For example, the VENTANA ROS1 (SP384) rabbit monoclonal primary antibody assay (Roche, Basel, Switzerland) can be used to identify ROS1-positive cancers, an aggressive form of cancer that occurs in approximately 1-2% of NSCLC patients. The VENTANA ALK (D5F3) CDx assay (Roche, Basel, Switzerland) can be used to help identify NSCLC patients eligible for treatment with XALKORI® (crizotinib), ZYKADIA® (ceritinib), or ALECENSA® (alectinib). p40 (BC28) mouse monoclonal primary antibody assay (Roche, Basel, Switzerland), TTF-1 (SP141) rabbit monoclonal primary antibody assay (Roche, Basel, Switzerland), Cytokeratin 5 / 6 (D5 / 16B4) mouse monoclonal primary antibody assay (Roche, Basel, Switzerland), and Napsin A (MRQ-60) mouse monoclonal primary antibody assay (Roche, Basel, Switzerland) can also be used to stratify lung cancer.
[0081] II.B.3.ii.a. NSCLC The stages of NSCLC are as follows:
[0082] Stage 0 is also known as carcinoma in situ. At this stage, the cancer is small in size and has not spread to deeper lung tissue or outside the lung.
[0083] Stage I is characterized by cancer in a single lung that may be present in the underlying lung tissue but has not spread to lymph nodes. This stage is divided into Stage Ia and Stage Ib. In Stage Ia, the tumor is 3 centimeters or smaller. In Stage Ib, the tumor is 3-4 centimeters in size or the tumor is 4 centimeters or smaller and one or more of the following is present: (1) cancer has spread to the main bronchus but not to the carina; (2) cancer has spread to the innermost layer of the membrane that covers the lung; and / or (3) partial or complete collapse of the lung or pneumonitis has developed.
[0084] Stage II includes possible spread to nearby lymph nodes and the chest wall. This stage is divided into stage IIa and stage IIb. Stage IIa cancer represents a tumor larger than 4 centimeters but not larger than 5 centimeters that has not spread to nearby lymph nodes. Stage IIb lung cancer represents a tumor not larger than 5 centimeters that has spread to lymph nodes. Stage IIb cancer can also be a tumor larger than 5 centimeters wide that has not spread to lymph nodes.
[0085] Stage III involves continued spread from the lungs to the lymph nodes. If the cancer has spread only to lymph nodes on the same side of the chest where it started, it's called Stage IIIa. If the cancer has spread to lymph nodes on the other side of the chest or above the collarbone, it's called Stage IIIb.
[0086] Stage IV is the most advanced metastatic stage of the disease. At this stage, the cancer has spread beyond the lungs to other areas of the body. Approximately 40% of NSCLC patients are diagnosed at stage IV, with a five-year survival rate of less than 10%.
[0087] II.B.3.ii.b.SCLC The stages of SCLC are characterized by the medical community as follows:
[0088] Limited-stage or stage 1 SCLC is lung cancer that occurs on only one side of the chest and involves a single area of the lung, lymph nodes, or both.
[0089] Extensive stage, or stage 2, SCLC is lung cancer that has spread to the other side of the chest, outside the chest, or to other parts of the body.
[0090] II.B.3.iii. Treatment II.B.3.iii.a. NSCLC Surgery is often recommended for patients with stage I or II NSCLC and can offer the best chance for a cure. Based on risk factors, surgery with or without adjuvant chemotherapy (or radiation if the patient is not a surgical candidate) is generally appropriate for stages Ib and II.
[0091] The standard treatment for stage I and II NSCLC is surgery with adjuvant chemotherapy. For example, platinum chemotherapy agents such as cisplatin or carboplatin can be administered in combination with vinorelbine, etoposide, vinblastine, gemcitabine, docetaxel, pemetrexed, or paclitaxel.
[0092] The standard treatment for locally advanced disease (stage IIIa or IIIb) is chemoradiotherapy. Treatment recommendations include the use of concurrent chemotherapy and radiation, or sequential chemotherapy and radiation. Selected patients (mainly stage IIIa patients) can be candidates for surgery, and these patients can receive chemotherapy alone or with radiation before surgical resection. Stage IIIa and IIIb disease are typically treated with a combination of chemotherapy and radiation if the patient is not a surgical candidate.
[0093] Chemotherapy and radiation therapy are preferably given concurrently, although in patients with poor performance status, these therapies may be given sequentially. The decision to treat a patient with concurrent chemoradiation therapy rather than surgery, radiation, or chemotherapy should be made by a multidisciplinary team including a medical oncologist, radiation therapist, and thoracic surgeon.
[0094] Patients with metastatic disease (stage IV) or recurrent disease after primary therapy (e.g., surgery and / or radiation) should consider primary chemotherapy to improve quality of life, relieve symptoms, and improve overall survival. For example, platinum chemotherapy agents such as cisplatin or carboplatin can be administered in combination with vinorelbine, etoposide, vinblastine, gemcitabine, docetaxel, pemetrexed, or paclitaxel.
[0095] For example, monotherapy with paclitaxel, docetaxel, gemcitabine, vinorelbine, or pemetrexeb is a reasonable first-line option in patents or elderly patients with good performance status.
[0096] Second-line chemotherapy can be administered for metastatic or recurrent disease after disease progression after first-line therapy. Exemplary second-line regimens are: nivolumab; pembrolizumab in PD-L1-positive tumors (patients with EGFR or ALK gene tumor abnormalities must have disease progression before receiving pembrolizumab); docetaxel and ramucirumab; nintedanib and docetaxel; erlotinib (Tarceva® (Roche, Basel, Switzerland)); and afatinib. Erlotinib alone remains the standard treatment in the second-line setting.
[0097] After disease progression after first- and second-line therapy, third-line chemotherapy is used for advanced or recurrent NSCLC. Options include erlotinib, ramucirumab, and nivolumab.
[0098] For patients with advanced (stage IV) disease who have a disease response or stable disease after completing first-line chemotherapy, maintenance chemotherapy for metastatic or recurrent disease in the form of switch maintenance chemotherapy or continued maintenance therapy can be considered.
[0099] Switch maintenance chemotherapy involves administering chemotherapy using different agents than those used in first-line therapy. Continuation maintenance therapy involves administering chemotherapy containing agents that were part of first-line therapy after completing four to six cycles of first-line therapy.
[0100] II.B.3.iii.b.SCLC SCLC of any stage typically responds initially to treatment, but responses are usually short-lived. Chemotherapy, with or without radiation therapy, is administered depending on the stage of the disease. In many patients, chemotherapy prolongs survival and improves quality of life sufficiently to warrant its use. Surgery generally plays no role in the treatment of SCLC, but it can be curative in rare patients with small, localized tumors (such as solitary pulmonary nodules) who undergo surgical resection before the tumor is identified as SCLC.
[0101] Limited-stage SCLC is commonly treated with a combination of chemotherapy agents. For example, platinum chemotherapy agents such as cisplatin or carboplatin can be administered in combination with vinorelbine, etoposide, vinblastine, gemcitabine, docetaxel, pemetrexed, or paclitaxel.
[0102] In persistent SCLC, chemotherapy alone, either as monotherapy or combination therapy, is often used. Examples of such chemotherapy agents include irinotecan, topotecan, vinca alkaloids (e.g., vinblastine, vincristine, vinorelbine), alkylating agents (e.g., cyclophosphamide, ifosfamide), doxorubicin, taxanes (e.g., docetaxel, paclitaxel), and gemcitabine. Some combinations include platinum chemotherapy agents such as cisplatin or carboplatin combined with etoposide, irinotecan, topotecan, and gemcitabine. In some cases, cyclophosphamide, doxorubicin, and vincristine are administered as first-line chemotherapy.
[0103] Patients with disease that recurs more than six months after completion of first-line chemotherapy can be treated again with the original first-line regimen (typically a platinum-based combination).
[0104] II.B.4. Blood cancer Most hematological, or blood, cancers begin in the bone marrow, where blood cells are made. Blood cancers occur when abnormal blood cells grow out of control and interfere with the function of normal blood cells. There are three main types of blood cancer:
[0105] Leukemia occurs when the body produces too many abnormal white blood cells, interfering with the bone marrow's ability to produce red blood cells and platelets.
[0106] Lymphoma is a blood cancer that affects the lymphatic system. In lymphoma, abnormal mutated lymphocytes grow and produce more abnormal lymphocytes. Over time, these abnormal lymphocytes become lymphoma cells, which damage the immune system.
[0107] Myeloma is a cancer of plasma cells, which are white blood cells that produce antibodies to fight disease and infection. Myeloma cells interfere with the normal production of antibodies, thus weakening the body's immune system and making the body more susceptible to infection.
[0108] II.B.4.i.Clinical symptoms Symptoms of blood cancer include anemia, poor blood clotting, abnormal bruising, bleeding gums, rash, heavy periods, black or red streaked bowel movements, fever, night sweats, a lump in the neck or armpit, unexplained weight loss, and bone pain.
[0109] II.B.4.ii. Diagnosis To diagnose leukemia, a physical exam and complete blood count (CBC) test are performed, which can identify abnormal levels of white blood cells compared to red blood cells and platelets. In some cases, a bone marrow biopsy is performed to diagnose and / or identify the type of leukemia. Once a diagnosis is made, leukemia can also be staged. For example, the stages for CLL, the most common type of leukemia in adults over the age of 19, are as follows:
[0110] Stage 0 is when there are too many white blood cells (lymphocytes) in the blood, but other blood cell counts are near normal. There are usually no other symptoms of leukemia. The cancer is growing slowly, and this stage is low risk.
[0111] Stage I is the medium-risk stage, when the blood has too many lymphocytes. At this stage, the lymph nodes are larger than normal, but other organs are normal in size. Typically, the red blood cell and platelet counts are also near normal.
[0112] Stage II is a moderate-risk stage in which there are too many lymphocytes in the blood and the spleen is enlarged. The lymph nodes may also be larger than normal. The red blood cell and platelet counts are near normal.
[0113] Stage III is a high-risk stage in which the blood has too many lymphocytes and the patient is anemic (i.e., too few red blood cells). In addition, the lymph nodes, liver, or spleen may be larger than normal. The platelet count is near normal.
[0114] Stage IV is a high-risk stage when the blood has too many lymphocytes and too few platelets. At this stage, the lymph nodes, liver, or spleen may be larger than normal, and the patient may be anemic.
[0115] Diagnosis of lymphoma usually involves a lymph node biopsy. In some cases, X-rays, blood tests, CT scans, and / or PET scans can be used to detect enlarged lymph nodes. Once a diagnosis is made, lymphoma can also be staged. The stages of lymphoma are as follows:
[0116] Stage 1 involves only one area or site, such as a lymph node or lymphatic structure.
[0117] Stage 2 involves two or more lymph node regions or two or more lymph node structures. In this stage, the involved areas are on the same side of the body.
[0118] Stage 3 involves the lymph node areas and structures on both sides of the body.
[0119] Stage 4 involves other organs besides the lymph nodes, and lymphatic structures are involved throughout the body. These organs can include the bone marrow, liver, or lungs.
[0120] To diagnose myeloma, one or more of the following may be used: CBC test, blood tests, urine tests, bone marrow biopsy, X-ray, MRI, PET and CT scan to confirm the presence and extent of myeloma.
[0121] II.B.4.iii. Treatment Treatment for blood cancer depends on the type and stage of the cancer, as well as the extent of the disease and other basic health parameters. Treatment options include radiation therapy, chemotherapy, immunotherapy, and stem cell transplantation.
[0122] II.B.4.iii.a B-cell lymphoma B-cell lymphomas make up the majority (approximately 85%) of non-Hodgkin's lymphomas (NHL) in the U.S. DLBCL, FL, and CLL are the most common types of B-cell lymphoma.
[0123] II.B.4.iii.b. Diffuse large B-cell lymphoma (DLBCL) Treatment of DLBCL will vary depending on the stage and sub-indication of DLBCL, but the standard of care for most patients is R-CHOP (rituximab (Mabthera / Rituxan (Roche, Basel, Switzerland)), cyclophosphamide, hydroxydaunorubicin, vincristine, and prednisolone) chemotherapy.
[0124] Treatment for first relapse of DLBCL is typically based on whether or not the patient intends to proceed to autologous stem cell transplantation. For patients intending transplantation, typical regimens are R-ICE (rituximab, ifosfamide, carboplatin, and etoposide) and R-DHAP (rituximab, dexamethasone, high-dose cytarabine, and cisplatin) or, less commonly, R-ESHAP (rituximab, etoposide, solumedrone, high-dose cytarabine, and cisplatin). Other regimens, such as R-Benda (rituximab and bendamustine) and R-Borte (rituximab and bortezomib), are typically reserved for patients who are not eligible for transplant due to factors such as age and the presence of comorbidities. In some cases, adult patients with relapsed or refractory DLBCL who are not candidates for stem cell transplantation are given polazutuzumab vedotin (Polivy®, (Roche, Basel, Switzerland)) in combination with bendamustine and rituximab. If there is a second relapse of DLBCL, R-ICE, R-ESHAP, BR, R-Benda, R-DHAP, or R-Hyper-CVAD (rituximab, hyperfractionated cyclophosphamide, doxorubicin, vincristine, and dexamethasone) can be administered.
[0125] II.B.4.iii.c. Follicular lymphoma (FL) Treatment for FL varies depending on the subindication of FL and the standard of care, but first-line chemotherapy treatments include rituximab (R), R-CHOP (rituximab, cyclophosphamide, hydroxydaunorubicin, vincristine, and prednisolone) chemotherapy, R-Benda, and R-CVP (rituximab, cyclophosphamide, vincristine, and prednisolone). First-line maintenance therapy for FL is usually rituximab.
[0126] If a first relapse of FL occurs, patients typically receive a different regimen from the first-line therapy, such as R-CHOP, R-CVP, R-Benda, or R-DHP. If a second relapse occurs, patients can be administered R-Benda, R-ICE, or idelalisib.
[0127] In some cases, tazemetostat can be administered to patients with recurrent or refractory FL whose tumors are positive for enhancer of zeste homolog 2 (EZH2) gene mutations and who have received at least two prior systemic therapies. FDA-approved tests for the detection of EZH2 mutations are available; for example, the cobas® EZH2 Mutation Test (Roche, Basel, Switzerland) can be used to identify mutations in DNA extracted from formalin-fixed, paraffin-embedded human FL tumor tissue.
[0128] II.B.4.iii.d. Chronic lymphocytic leukemia (CLL) CLL is commonly diagnosed in older adults, with a median age at diagnosis of 72 years. For this reason, the 2013 International CLL Workshop proposed that CLL patient fitness be a better determinant for patient selection and treatment goal identification. This fitness classification is necessary because it can: (1) accurately classify patient life expectancy unrelated to CLL (i.e., other health issues); (2) determine a patient's ability to tolerate aggressive chemotherapy, including predicting treatment modifications and discontinuations; and (3) enable more consistent patient stratification and selection across clinical trials. Researchers now recognize extensive disease heterogeneity due to underlying tumor biology (e.g., 17p and 11q deletions). (Fit vs. Frail Assessment Strategies in CLL, New Evidence Oncology Issue—October 2015). In CLL, patients are treated according to their fitness status (fit or unfit), whether they have specific mutations, and whether they are treated for initial disease onset or recurrence.
[0129] Treatments for CLL vary, but often the patient's condition is monitored without treatment until signs or symptoms appear or change. Once the decision to administer treatment is made, options include radiation therapy, chemotherapy, and targeted therapy.
[0130] Depending on the subindication of CLL, FCR (fludarabine, cyclophosphamide, and rituximab) is often used as the standard first-line chemotherapy regimen for eligible patients. Benda-R can be used for patients with a history of previous infection. An alternative first-line option for less compatible patients is the combination of chlorambucil with an anti-CD20 antibody (e.g., rituximab, ofatumumab, or obinutuzumab). For patients with TP53 or del(17p) mutations, a BCR receptor antagonist with or without rituxamib can be administered. Alternatively, hematopoietic stem cell transplantation can be considered for patients in remission.
[0131] If a patient has relapsed or refractory CLL, the patient can be administered a BCL2 antagonist with or without rituximab. Alternatively, R-Benda or FCR can be administered to the patient. Other regimens for relapsed CLL include ibrutinib, idelalisib and rituximab, or allogeneic hematopoietic stem cell transplantation. If a patient has relapsed CLL and has a TP53 mutation or del(17p) mutation, the patient can be administered a BCL2 antagonist with or without rituximab. Alternatively, other regimens include ibrutinib, idelalisib and rituximab, or allogeneic hematopoietic stem cell transplantation.
[0132] Supportive care regimens can also be administered to patients being treated or who have been treated for cancer. These include medications for chemotherapy- and / or radiation-induced nausea and vomiting (e.g., Kytril® (Roche, Basel, Switzerland)); anti-anemic drugs (e.g., NeoRecorman (Roche, Basel, Switzerland)); medications to treat or prevent bone metastases (e.g., Bondronat® (Roche, Basel, Switzerland)); and treatment for neutropenia (e.g., Neupogen® (Roche, Basel, Switzerland)), to name a few.
[0133] III. Overview of Cloud-Based Network Architecture for Deploying Intelligent Functions The technology relates to configuring a server to execute code that enables a user of an entity (e.g., a physician) to run machine learning or AI techniques using subject records. The subject records contain complex combinations of data elements that characterize the subject. As an illustrative example, a subject record may contain a combination of thousands of data fields. Some data fields may contain fixed, non-numeric values (e.g., the subject's ethnicity), other data fields may contain unstructured text data (e.g., notes written by a physician), other data fields may contain a time-varying series of collected measurements (e.g., glycosylated hemoglobin measurements taken two to four times a year), and other data fields may contain images (e.g., an MRI of the subject's brain). Because machine learning and AI models are often configured to process data in numeric or vector format, the complexity and distribution of data types and formats in subject records makes processing subject records technically difficult, if not impossible. In light of this intended technical problem, certain aspects and features of the present disclosure relate to the transformation of subject records into a transformed representation, such as a vector representation, that characterizes the various data elements of the subject record.
[0134] The technology relates to converting non-numeric values contained in subject records into numerical representations (e.g., feature vectors) that can be input into machine learning or AI models to generate predicted outputs. A server executing the code provides a technical effect of solving a desired technical problem by converting subject records into a converted representation consumable by a machine learning or AI model. "Consumable" may refer to data in a format or form that a machine learning or AI model is configured to process to generate predicted outputs. Machine learning or AI models are not configured to process subject records (as they reside stored in a data registry) due to the complex combination of data elements in multiple data formats and data types contained in each individual subject record. By way of example, for a given subject record, data elements may include a longitudinal sequence of events (e.g., immunization records), other data elements may include measurements obtained from the subject (e.g., vitals), still other data elements may include text entered by a user (e.g., notes recorded by a physician), and still other data elements may be images (e.g., X-rays). Limited or simple analysis can be performed on subject records (before any transformations), such as grouping subjects based on the values of data elements (e.g., age groups). However, as the complexity and size of subject records reach big data scale, limited or simple analysis becomes problematic or infeasible. To process and extract analytical assessments from subject records at big data scale, machine learning or AI techniques can be used to data mine the subject records. However, the machine learning or AI model is configured to receive numeric inputs or vector inputs. For example, a clustering operation such as k-means clustering is configured to receive vectors as inputs.Thus, to perform clustering operations on subject records, the present disclosure provides a technical effect of solving a desired technical problem by converting the subject records into a transformed representation consumable by a machine learning or AI model, such as a numeric vector representation. Intelligent analysis can be performed on the subject records in the transformed representation. Non-limiting examples of intelligent analysis (performed when the server executes the code) include automatically detecting subject groups using clustering techniques, generating output that predicts a particular outcome based on values of data elements in the subject records, and identifying existing subject records that are similar to a given or new subject record.
[0135] By way of illustration and non-limiting example only, a subject's subject record may include four data elements: a first data element includes a unique code representing a diagnosis of a condition; a second data element includes an MRI of the subject's brain; a third data element includes a time-varying series of measurements, such as blood pressure measurements over a one-year period; and a fourth data element includes unstructured notes, e.g., notes of symptoms detected by examining or performing one or more tests. According to certain implementations, each of the first, second, third, and fourth data elements may be converted to a transformation representation (e.g., a vector). The technique used to convert the values contained within the four data elements may depend on the type of data contained in the data elements. For the first data element, for example, the unique code representing a diagnosis may be represented as a fixed-length vector, the size of the vector being determined by the size of a vocabulary of codes, and each code in the vocabulary being represented by a vector element of the fixed-length vector. One or more unique codes contained within the first data element may be compared to the vocabulary of codes. If the unique code matches a code in the vocabulary, a "1" may be assigned to the vector element at the position of the vector corresponding to the unique code, and all remaining vector elements of the vector may be assigned a "0." In light of the above, a first vector may be generated to represent the value of the first data element. As another example, for the second data element, a latent space representation of an image may be generated using a trained autoencoder neural network. The latent space representation of the input image may be a dimensionally reduced version of the input image. The trained autoencoder neural network may include two models: an encoder model and a decoder model. The encoder model may be trained to extract a subset of salient features from a set of features detected in the image. The salient features (e.g., keypoints) may be regions of high intensity in the image (e.g., object edges). The output of the encoder model may be a latent space representation of the input image.The latent space representation may be output by a hidden layer of a trained autoencoder model, and therefore, the latent space representation may be interpretable only by the server. A decoder model can be trained to reconstruct the original input image from a subset of the extracted salient features. The output of the encoder model can be used as a feature vector representing the pixel values of the image included in the second data element. In light of the above, a second vector (e.g., a latent space representation) can be generated to represent the image included in the second data element. As another example, for a third data element, a time-varying sequence of measurements can be numerically represented. In some implementations, the time-varying sequence can be represented by a sum of instances where measurements are taken from a subject. In other implementations, the time-varying sequence can be numerically represented using the arithmetic mean, average, or median of values of measurements taken over measurement instances occurring over a period of time (e.g., one year). In other implementations, the frequency of measurements can be calculated and used to numerically represent the time-varying sequence of measurements. In light of the above, a third vector can be generated to represent the time-varying sequence of values included in the third data element. As yet another example, the notes entered by the user for the fourth data element can be processed and vectorized using any number of natural language processing (NLP) text vectorization techniques. In some implementations, a word-to-vector machine learning model, such as a Word2Vec model, can be executed to convert the notes included in the fourth data element into a single vector representation. In other implementations, a convolutional neural network can be trained to detect words or numbers in the text indicating symptoms, treatments, or diagnoses from the notes included in the fourth data element. In light of the above, a fourth vector can be generated to represent the text of the notes included in the fourth data element as a vector representation. Thus, the final feature vector representing the entire subject record can be a vector of vectors including a concatenation of the first vector, the second vector, the third vector, and the fourth vector.In another example, the average of the first vector, the second vector, the third vector, and the fourth vector can be used to numerically represent the entire subject record. Other combinations of the first vector, the second vector, the third vector, and the fourth vector can be used to generate a final feature vector that numerically represents the entire subject record.
[0136] In some implementations, instead of generating a vector that numerically represents each data element in the subject record, a technique can be implemented to reduce the dimensionality of the subject record by identifying and selecting a subset of data elements from the set of data elements. The subset of data elements can represent “important” data elements, and the “importance” of the data elements can be determined based on predictions using feature extraction techniques such as singular value decomposition (SVD). For example, converting the subject record into a transformed representation consumable by machine learning and AI models can include performing one or more feature extraction techniques on non-numeric values contained in the data elements of the subject record to generate a feature vector that numerically represents the decomposed version of the non-numeric values. In some implementations, the feature extraction technique can include, for example, reducing the dimensionality of a set of data elements (e.g., each data element representing a characteristic or dimension of the subject) in the subject record to an optimal subset of features that can be used, for example, to predict an outcome or event. Reducing the dimensionality of the set of data elements can include reducing N data elements to a subset of M elements, where M is less than N. In these implementations, each element of the subset of M elements can be converted to a numerical value. In some implementations, a feature vector can be generated to represent N data elements of a subject record. The feature vector can include a vector for each data element of a set of data elements. For example, the feature vector can be a numerical representation of a complex combination of the data elements of a subject record. Each non-numeric value in the data elements of a subject record can be vectorized to generate a representative vector. Vectors representing a set of data elements in a subject record can be concatenated or combined (e.g., as an average or weighted average) to generate a feature vector that numerically characterizes the entire set of data elements of the subject record. The feature vector can be consumed by a trained machine learning or AI model. Once a feature vector for a subject record is generated, the subject record can be evaluated individually or in groups of other subject records using machine learning models and AI techniques.After the feature vectors representing each subject record are generated and stored, the feature vectors for the subject records stored in the central data store can be input into a machine learning or AI model, or other advanced analysis can be performed on the numerical representation of the subject records. For example, two different subject records can be compared along one or more dimensions. A dimension can represent a feature or data element of a subject record along which a comparison between two or more subject records is made. By way of example, a data element in a first subject record includes text entered by a first user (e.g., a physician) describing the first subject's symptoms. The text (e.g., the value of the data element in the first subject record) can be vectorized using the text vectorization techniques described above (e.g., Word2Vec) to generate a first vector that numerically represents the text associated with the data element. The text vectorization technique can generate an N-dimensional word vector for each word contained in the text. A matching data element in a second subject record (e.g., a data element in another patient record that also includes text entered by a physician describing the other patient's symptoms) can include text entered by a second user describing the second subject's symptoms. The text (e.g., the value of a data element in the second subject record) can be vectorized using the text vectorization techniques described above to generate a second vector (e.g., an N-dimensional word vector) representing the text associated with the data element. The server can compare the first vector to the second vector in Euclidean or cosine space to quantify the similarity or dissimilarity between the first patient record and the second patient record, at least with respect to the dimension of the patient's presentation of symptoms. If the first vector and the second vector are close to each other in Euclidean space (i.e., if the Euclidean distance between the first vector and the second vector is small), then the symptoms experienced by the first subject (as described in the text of the data element) are likely to be similar to the symptoms experienced by the second subject (as described in the text of the data element).However, if the Euclidean distance between the first vector and the second vector is large or exceeds a threshold distance (e.g., or if the Euclidean distance is above a threshold), the symptoms experienced by the first subject can be predicted to be different from the symptoms experienced by the second subject.
[0137] In some implementations, the server can be configured to run an application that enables users of an entity to build a data registry that serves to store subject records for subsequent processing. The data in a subject record can include unstructured data, such as electronic copies of physician notes and / or responses to open-ended questions. The unstructured data can be captured in the data registry by mapping portions of the unstructured data to fixed portions (e.g., data elements) of a structured data record. The structure of the structured data record can be defined using specifications from a module that corresponds (for example) to a particular use case (e.g., a particular disease, a particular test). For example, each word (i.e., text) in the unstructured note data can be converted to a numeric representation, and the various numeric representations associated with the unstructured note data can be decomposed (e.g., using SVD) to find words that describe a particular set of symptoms exhibited by the subject. The decomposition of the numeric representation of the unstructured note data can remove non-informative words such as "and," "the," and "or." The remaining words represent a particular set of symptoms. Some portions of the note data may be irrelevant with respect to a data element in the structured data and / or may be more or less specific than the data contained in the data element. In some examples, various mapping (e.g., mapping a "poor balance" symptom to a "neurological" symptom), NLP, or interface-based techniques (e.g., prompting a user for new information) may be used to obtain a structured data record. The interface may also be used to receive input identifying new information about a new or existing subject, and the interface may include input components and options that map to the structure of the data record.
[0138] The technology further relates to configuring a cloud-based application to convert non-numeric values contained in data elements of a subject record into numeric representations, so that the cloud-based application can perform intelligent analytical functions using the numeric representations (e.g., converted representations) of the subject record stored in the data registry. The conversion of non-numeric values of data elements of a subject record into numeric representations can depend on the type of data contained in the data element. For example, for data elements containing text, such as notes taken by a user, the text can be converted into a numeric representation of the text using an NLP technique such as Word2Vec or other text vectorization techniques. As another example, for data elements containing image frames of an image (e.g., an MRI) or video (e.g., an ultrasound video), each image or image frame can be converted into a numeric representation (e.g., a vector) using a trained autoencoder neural network trained to generate a latent space representation of the input image. The condensed representation (e.g., latent space representation) of the input image can serve as a vector that numerically represents the input image. As yet another example, for data elements containing a time-varying information sequence (e.g., events occurring over a period of time), the time-varying information can be represented as a numeric representation using several exemplary transformations. In some examples, a count of an event may be used as a vector representing the time-varying information. In other examples, the frequency or rate of an event occurring (e.g., once a week, once a month, once a year) can be used as a vector representing the time-varying information. In still other examples, an average or combination of measurements associated with each event in the time-varying information can be used as a vector representing the time-varying information. The present disclosure is not limited to these examples, and other numerical representations of the time-varying information can be used as vectors representing the numerical representation. Intelligent analytical functions can be performed by running machine learning or AI models trained using the data records. Model outputs can be used to indicate specific analyses extracted from the data records.
[0139] In some examples, transmission of data from subject records can be provided to create treatment plans for individual subjects. For example, subject record information (e.g., complying with data privacy restrictions by selecting data omission and / or obfuscation) can be broadcast and / or transmitted to a selected group of user devices. For example, a broadcast may be sent to user devices associated with similar data records in response to input from the user corresponding to a request to initiate a consultation with a user associated with a similar subject. If the user receiving the broadcast accepts the consultation request (via providing corresponding input), a secure data channel can be established between the users, and potentially more subject records can be shared (e.g., while complying with data privacy restrictions applicable to the two users). Subject records similar to a given subject can be identified by performing nearest neighbor techniques using vector representations of two or more subject records. The nearest neighbor technique can be performed by comparing vectors of individual data elements across multiple subject records (e.g., nearest neighbors can be determined relative to a dimension or feature of the subject record). Alternatively, the nearest neighbor technique can be performed by comparing an overall vector characterizing an entire subject record with an overall vector characterizing another entire subject record. The overall vector may be a concatenation of the individual vectors representing the values of the data elements, or may be an average or combination of the individual vectors representing the values of the data elements.
[0140] As another example, one or more processed data records may be returned in response to a query for subject records that match certain constraints. In some examples, a first user may submit a query identifying a first subject record. The query may correspond to a request to identify other subject records similar to the first subject record. The server may convert the first subject record into a converted representation using certain conversion techniques described above. Alternatively, the converted representation of the first subject record may be previously generated and stored in the database. Regardless of whether the converted representation of the first subject record is generated before or after receiving the query, converting the first subject record into the converted representation of the first subject record may include generating a vectorization of one or more non-numeric values of data elements of the first subject record. Vectorizing one or more non-numeric values contained within the first subject record may include generating a numeric vector representation for each value contained in each data element of the first subject record (e.g., for non-numeric text, such as notes). The various vector representations may be concatenated or otherwise combined (e.g., an average may be calculated) to generate a feature vector representing the entire first subject record. A vector representation that numerically represents a first subject record can be compared to vector representations of other subject records in a domain space (e.g., Euclidean space or cosine space). For example, if the Euclidean distance between the two vector representations is within a threshold distance, the two subject records associated with the two vector representations can be interpreted (e.g., by a server) as similar with respect to at least one or more dimensions.
[0141] For each data element in a subject record, the technique used to generate a vector representation of the values associated with the data element may depend on the type of data associated with the data element. In some examples, a data element in a subject record may be associated with one or more images, such as an X-ray of the subject. A feature extraction technique may be performed to generate a vector representation of each image associated with the data element. For example, the server may be configured to run a trained autoencoder neural network to generate a dimensionality-reduced version of the image. The trained autoencoder neural network may include two models: an encoder model and a decoder model. The encoder model may be trained to extract a subset of salient features from a set of features detected in the image. The salient features (e.g., keypoints) may be regions of high intensity in the image (e.g., object edges). The output of the encoder model may be a latent space representation of the input image. The latent space representation may be output by a hidden layer of the trained autoencoder model; therefore, the latent space representation may be interpretable only by the server. The subset of salient features of the latent space representation characterizing the subject record may be compared to the subset of salient features of the latent space representation characterizing other subject records to gain specific analytical insights. The decoder model may be trained to reconstruct the original input image from the extracted subset of salient features. The output of the encoder model may be a vector representation of the data elements associated with the image comprising the subject record. In another example, a keypoint matching technique may be performed to match keypoints of an image contained in the data elements of a first subject record to keypoints of another image contained in the data elements of a second subject record. The vector representation (e.g., latent space representation) of the input image may be consumable by a machine learning or AI model, such that two different subject records (each containing an image) can be compared to each other to determine the similarity or dissimilarity between the two different subject records.
[0142] To illustrate, and by way of non-limiting example only, a magnetic resonance image (MRI) of a subject's brain is captured. The MRI is stored in a subject record associated with the subject. The server is configured to generate a transformed representation, such as a vector representation, of the MRI included in the subject record using feature extraction techniques, such as keypoint detection, auto-encoding to a latent space representation, SVD, and other suitable computer vision techniques. The vector representation of the data element including the MRI is concatenated or combined (e.g., averaged) with the vector representations of each remaining data element in the set of data elements to generate a feature vector characterizing the entire subject record. A user can access the application to query a database of other subject records to search for a subset of other subject records that include MRIs similar to the MRI of the subject's brain. Identifying other subject records that are similar to the subject record (at least in terms of similarity between the MRIs) can include computing k-nearest neighbors of the subject record. For example, the transformed representation may be plotted (visually or internally by the computing system) on a domain space, such as Euclidean space or cosine space. The transformed representation for each other subject record may be plotted (visually or internally by the computing system). Nearest neighbor techniques can be performed to compare the vector representation of the subject record with vector representations of other subject records to identify the k nearest neighbors to the subject vector. The identified k nearest neighbors can be predicted to have MRIs similar to the MRI of the subject's brain. Each other subject record identified as a nearest neighbor can be identified and retrieved for further evaluation or processing using an application.
[0143] In some implementations, the computing system can perform data processing techniques (e.g., nearest neighbor techniques) to identify similar subject records. In this search, various data elements can be differentially weighted (e.g., according to predefined data element weightings, the importance of matching various data elements, and / or user input indicating the prevalence of particular data element values across the subject record set). When searching a set of records for potential matches, some records may be missing values for various data elements. In these cases, when evaluating potential matches, it can be determined (for example) that the data element values do not match and / or that the data elements are unweighted. Handling of missing values can depend on the values of the data elements across the set of records and / or the distribution of the values of the data elements in the query.
[0144] Additionally, some techniques involve defining and using a set of rules used to identify potential treatment regimens for a subject given a set of symptoms identified in the subject record. For example, a target subject record may represent a target subject who recently experienced three symptoms: upper respiratory tract infection, fever, and sore throat. The three symptoms may be written as text within a data element of the target subject record (e.g., with separation between words marked by tags such as semicolons). A server, such as cloud server 135, may input the text "upper respiratory tract infection," "fever," and "sore throat" individually into a trained Word2Vec model or other text-to-vector model, such as vocabulary mapping. The Word2Vec model may be trained to generate a vector representation for each word representing a symptom. The vector representations of the three symptoms may be averaged to generate a single vector representation for the "symptoms" data element of the target subject record. The single vector representation for the "symptoms" data element of the target subject record may be processed to identify other subject records containing similar words in the "symptoms" data element. Each subject record stored in a database may be associated with an existing "symptoms" data element, converted into a numerical representation, such as a vector. The vector of the "symptoms" data element can be plotted and compared to the vector of the "symptoms" data element of the target subject record. The server can identify the vector closest to the vector characterizing the "symptoms" data element. The vector of the "symptoms" data element closest to the vector of the target subject record can be predicted to be similar to the subject. The subject record associated with the vector closest to the vector of the target subject record may be identified and further evaluated to determine a treatment regimen to be provided to that subject. The treatment provided to the subject associated with the vector closest to the vector for the target subject record can be used as a potential treatment regimen for treating the target subject. Furthermore, each potential treatment regimen can be weighted by the responsiveness experienced by other subjects. The potential treatment regimens can be classified according to the responsiveness experienced by other subjects.
[0145] The set of rules can be defined based on user interaction with a user interface, which can include specification of particular criteria and associated particular medical procedures and / or selection of one or more previously defined rules (specifying criteria and procedures). For example, one or more existing rules can be presented via the interface, and a user can select a rule to incorporate into a rule base associated with an account associated with the user. The one or more rules may be selected from among a set of rules defined by multiple users (e.g., associated with one or more institutions) and / or may be generated based on rules created by multiple users. When a user selects a rule for incorporation into the rule base, the application can generate a feedback signal to the cloud server 135. The feedback signal can include metadata associated with the user's selection. The metadata can indicate whether the rule was incorporated into the rule base without changes or with changes. If the rule base was modified, the metadata indicates what changes were made to the rule. The metadata can also indicate whether the rule was rejected, deleted, or otherwise determined to be not useful to the user. To illustrate, by way of non-limiting example, a computing system may detect that rules associating one or more particular types of symptoms and / or test results with a given treatment are relatively frequently defined and / or selected by a user, and the computing system may then generate general rules relating to the particular types of symptoms and / or test results and treatments. General rules may be defined to have (for example) the most restrictive, the most inclusive, or median criteria. In some examples, a user's rule base may be processed to detect overlaps of any criteria between rules. If overlaps are identified, a warning may be presented that an overlap has been identified. The rules in the rule base may be used to evaluate and classify subject records to define populations associated with the subject records.Evaluating a subject record using a rule may be implemented as a decision tree, for example, in that the first criterion of the rule is compared to an attribute contained in the subject record. If the first criterion is met, the next criterion is compared to an attribute contained in the subject record. If the next criterion is met, the comparison continues for each criterion contained in the rule. The comparison may continue even if the next criterion is not met. In this case, the fact that the criterion (and any others contained in the rule) was not met is remembered and presented to the user device along with the criteria that were met.
[0146] Additionally, embodiments of the present disclosure provide a cloud-based application configured to exchange subject information with external entities without violating data privacy regulations. The cloud-based application is configured to automatically evaluate data privacy regulations involved in sharing subject information across various jurisdictions. The cloud-based application is configured to execute protocols that obfuscate or otherwise modify the subject information, thereby algorithmically ensuring compliance with data privacy regulations.
[0147] IV. Network Environment for Hosting Cloud-Based Applications Configured with Intelligent Capabilities 1 illustrates a network environment 100 in which an embodiment of a cloud-based application is hosted. Network environment 100 may include a cloud network 130 that includes a cloud server 135, a data registry 140, and an AI system 145. Cloud server 135 may execute source code underlying the cloud-based application. Data registry 140 may store data records captured from or identified using one or more user devices, such as computers 105, laptops 110, and mobile devices 115.
[0148] The data records stored in the data registry 140 may be structured according to a skeletal structure of fixed parts (e.g., data elements). The computer 105, the laptop 110, and the mobile device 115 may each be operated by various users. For example, the computer 105 may be operated by a doctor, the laptop 110 may be operated by an administrator of an entity, and the mobile device 115 may be operated by a subject. The mobile device 115 may connect to the cloud network 130 using the gateway 120 and the network 125. In some examples, the computer 105, the laptop 110, and the mobile device 115 are associated with the same entity (e.g., the same hospital). In other examples, the computer 105, the laptop 110, and the mobile device 115 are associated with different entities (e.g., different hospitals). The user devices of the computer 105, the laptop 110, and the mobile device 115 are examples for illustrative purposes, and therefore the present disclosure is not limited thereto. The network environment 100 can include any number or configuration of user devices of any device type.
[0149] In some embodiments, cloud server 135 can obtain data (e.g., subject records) for storage in data registry 140 by interacting with either computer 105, laptop 110, or mobile device 115. For example, computer 105 interacts with cloud server 135 by using an interface to select subject records or other data records stored locally (e.g., stored on a network local to computer 105) for inclusion in data registry 140. As another example, computer 105 interacts with an interface to provide cloud server 135 with the address (e.g., network location) of a database that stores subject records or other data records. Cloud server 135 then retrieves the data records from the database and incorporates the data records into data registry 140.
[0150] In some embodiments, the computer 105, the laptop 110, and the mobile device 115 are associated with different entities (e.g., a medical center). Data records that the cloud server 135 retrieves from the computer 105, the laptop 110, and the mobile device 115 may be stored in different data registries. Data records from each of the computer 105, the laptop 110, and the mobile device 115 may be stored within the cloud network 130, but the data records are not intermingled. For example, the computer 105 cannot access data records retrieved from the laptop 110 due to constraints imposed by data privacy regulations. However, the cloud server 135 may be configured to automatically obfuscate, obscure, or mask portions of the data records when the data records are queried by different entities. Thus, data records ingested from an entity can be disclosed to different entities in an obfuscated, obscured, or masked form to comply with data privacy regulations.
[0151] Once the data records are collected from the computers 105, laptops 110, and mobile devices 115, the data records can be used as training data for training machine learning or AI models to provide the intelligent analytics functionality described herein. If a user device associated with an entity queries the data registry 140 and the query results include data records originating from different entities, the data records may be available for querying by any entity, given that these data records may be provided or exposed to the user device in an obfuscated format that complies with data privacy regulations.
[0152] The cloud server 135 can be configured in a specialized manner to execute code that, when executed, causes the cloud server 135 to perform an intelligent function using the transformed representation of the subject record (e.g., a vector that numerically represents information stored in the subject record). For example, the intelligent function can be performed by executing code using the cloud server 135. The executed code can represent a trained neural network model. The neural network model may be trained to perform intelligent functions such as predicting a subject's responsiveness to a treatment regimen, identifying similar patients, generating treatment regimen recommendations for the patient, and other intelligent functions. The neural network model can be trained using a training dataset that includes subject records of subjects who have previously been treated for a condition and experienced an outcome (e.g., overcoming the condition, increasing the severity of the condition, decreasing the severity of the condition, etc.). Furthermore, the executed code can be configured to cause the cloud server 135 to convert non-numeric values in the existing subject record into a numerical representation (e.g., the transformed representation) that can be processed by the trained neural network model. For example, code executed by cloud server 135 can be configured to receive as input each subject record of a set of subject records, and for each subject record, the code, when executed, can cause cloud server 135 to perform the operations described herein to convert each data element of each subject record into a transformed representation, such as a vector representation. Performing the intelligent function can include inputting at least a portion of the data records stored in data registry 140 into a trained machine learning or AI model to generate output for further analysis. In some embodiments, the output can be used to extract patterns within the data records or to predict values or outcomes associated with data fields of the data records. Various embodiments of intelligent functions performed by cloud server 135 are described below.
[0153] In some embodiments, the cloud server 135 is configured to enable a user device (e.g., operated by a physician) to access a cloud-based application to send a consultation broadcast to a set of destination devices. The consultation broadcast can be a request for support or assistance regarding the treatment of a subject associated with the subject record. The destination devices can be user devices operated by other users associated with other entities (e.g., physicians at other medical centers). If the destination devices accept the request for assistance associated with the consultation broadcast, the cloud-based application can generate a condensed representation of the subject record that omits or hides certain data fields of the subject record. The condensed representation can comply with data privacy regulations, such that the condensed representation of the subject record cannot be used to uniquely identify the subject associated by the subject record. The cloud-based application can send the condensed representation of the subject record to the destination devices that accepted the request for assistance. The users operating the destination devices can evaluate the condensed representation and communicate with the user devices using a communication channel to discuss options for treating the subject. For example, the communication channel can be configured as a secure chat room that allows a user device (e.g., operated by a physician requesting a consultation) to securely communicate with a destination device (e.g., operated by another physician providing a consultation).
[0154] In some embodiments, the cloud server 135 is configured to provide a treatment plan definition interface to the user device. The treatment plan definition interface allows the user device to define a treatment plan for the condition. For example, the treatment plan can be a workflow for treating a subject according to the condition. The workflow can include one or more criteria for defining a population of subjects as having the condition. The workflow can also include a specific type of treatment for the condition. The cloud server 135 receives and stores a treatment plan definition for a particular condition from each user device in a collection of user devices. A cloud-based application can distribute a treatment plan for a given condition to the set of user devices. Two or more user devices in the set of user devices may be associated with different entities. Each of the two or more user devices can offer the option to integrate any portion or the entire treatment plan into a customer rule set. The cloud server 135 can monitor whether a user device integrates the entire shared treatment plan or a portion of the treatment plan. The interaction between the user device and the shared treatment plan can be used to determine whether to update the treatment plan or rules created based on the treatment plan.
[0155] In some embodiments, the cloud server 135 allows a user operating a user device to access a cloud-based application to determine a suggested treatment for a subject with a certain condition. The user device loads an interface associated with the cloud-based application. The interface allows the user operating the user device to select a subject record associated with a subject being treated by the user. The cloud-based application can evaluate other subject records to identify previously treated subjects similar to the subject being treated by the user. Similarity between subjects can be determined, for example, using an array representation of the subject record. The array representation (e.g., a transformed representation such as a vector, an N-dimensional matrix, or any numerical representation of a non-numeric value) can be any numerical and / or categorical representation of the values of the data fields of the subject record. For example, the array representation of the subject record can be a vector representation of the subject record in a domain space, such as in Euclidean space. In some examples, the cloud server 135 can be configured to convert the entire subject record into a numerical representation, such as a vector. For a given subject record, the cloud server 135 can evaluate each data element to determine the type of data contained or included in that data element. The type of data can inform the cloud server 135 about which process or technique to perform to convert the numeric or non-numeric values of that data element into a numeric representation. As an illustrative example, the cloud server 135 can convert non-numeric values of data elements in a subject record (e.g., text in a doctor's note) into a numeric representation (e.g., a vector). The conversion can include using an NLP technique, such as Word2Vec or other text vectorization technique, to generate a numeric value representing each word of the text. The generated numeric values can serve as vectors that can be input into a trained neural network to perform intelligent analysis.As another illustrative example, for data elements including images (e.g., MRI data) or image frames (e.g., ultrasound video data) from a video, each image or image frame can be converted to a numerical representation (e.g., a vector) using a trained autoencoder neural network trained to generate a latent space representation of the input image. The condensed representation (e.g., latent space representation) of the input image can serve as a numerical representation of the input image. This numerical representation can be input into a neural network or other machine learning model to perform intelligent analysis of the associated subject record. As yet another example, for data elements including a time-varying information sequence (e.g., events occurring or measurements taken from a subject over a period of time), the time-varying information can be represented as a numerical representation using several illustrative transformations. In some examples, a count of events may be used as a vector representing the time-varying information. For example, if four measurements were taken on a subject over a year, the numerical representation could be "4." In another example, the frequency or rate at which an event occurs (e.g., once a week, once a month, once a year) can be used as a vector representing the time-varying information. In yet another example, an average or combination of measurements associated with each event in the time-varying information can be used as a vector representing the time-varying information. The disclosure is not limited to these examples, and other numerical representations of the time-varying information can be used as a vector representing the numerical representation.
[0156] The AI system 145 can be configured to collect datasets on a large scale, convert the collected datasets into curated training data, run learning algorithms using the curated training data, and store detected patterns, correlations, and / or relationships in the training data in one or more trained AI models. In some implementations, the AI system 145 can be configured to perform specific predictive functions, such as predicting treatment outcomes and cancer progression in a particular subject based on the subject's mutation profile across cancer types, predicting treatment survival likelihood for a particular subject using enriched subject-specific datasets, and automatically verifying whether features contributing to treatment selection comply with oncology guidelines. In some implementations, the output of the AI system 145 can predict treatment outcomes and / or cancer progression in a particular subject, as described in more detail with respect to FIGS. 8 and 11 . In other implementations, the output of the AI system 145 can predict treatment survival likelihood for a particular subject, as described in more detail with respect to FIGS. 9 and 12 . In other implementations, as described in more detail with respect to Figures 10 and 13, the output of the AI system 145 can classify whether the subject characteristics that contributed to the treatment selection are in accordance with existing tumor guidelines.
[0157] In some examples, multiple values in an array correspond to a single field. For example, the value of a data element may be represented by multiple binary values generated via one-hot encoding. As another example, each value of multiple values in a single data element of a subject record may be individually converted to a numeric representation, as described above. Numeric representations representing each value of the multiple values may be combined into a single numeric representation corresponding to the data element. Combining multiple numeric representations may be performed using any vector combination technique, such as averaging vector magnitudes, adding vectors, or concatenating multiple vectors into a single vector. In some examples, the cloud-based application may generate an array representation of each subject record in a group of subject records. Similarity between two subject records may be represented by comparing the two array representations to determine the distance between them. Instead of comparing the numeric representation of an entire subject record with other numeric representations of other subject records, subject records may also be compared along a dimension (e.g., data element). For example, comparing two subject records along a dimension may include comparing the numeric representation of a data element in a subject record with other numeric representations of matching data elements in other subject records. Furthermore, the cloud-based application may be configured to identify subjects that are nearest to a subject record selected by a user device using the interface. The nearest neighbors may be determined by comparing the numerical representations of the various subject records with the numerical representation of the target subject record. The cloud-based application can identify treatments previously performed on the subjects that are nearest neighbors. The cloud-based application can make the treatments previously performed on the nearest neighbors available on an interface.
[0158] In some embodiments, the cloud server 135 is configured to create queries that search a database of previously treated subjects. The cloud server 135 can execute the queries and retrieve subject records that satisfy the query constraints. However, when presenting the query results, the cloud-based application can present complete subject records only for subjects who have been or are being treated by the user creating the query. The cloud-based application masks or otherwise obfuscates portions of the subject records for subjects not treated by the user creating the query. Masking or obfuscating portions of the subject records included in the query results allows users to comply with data privacy regulations. In some embodiments, the query results (whether or not the query results are obfuscated) can be automatically evaluated for patterns or common attributes within the subject records.
[0159] In some embodiments, the cloud server 135 embeds a chatbot in the cloud-based application. The chatbot is configured to automatically communicate with a user device. The chatbot can communicate with the user device in a communication session in which messages are exchanged between the user device and the chatbot. The chatbot may be configured to select an answer to a question received from the user device. The chatbot can select an answer from a knowledge base accessible to the cloud-based application. When a user device sends a question to the chatbot and the chatbot does not have an existing answer stored in the knowledge base, a different representation of the question for which there is an existing answer stored in the knowledge base is presented. A user communicating with the chatbot can prompt whether the answer provided by the chatbot is accurate or useful.
[0160] It will be understood that any machine learning or AI algorithm can be implemented to generate any of the trained machine learning models described herein. Various types and techniques of AI-based machine learning models can be trained and then implemented to generate one or more outputs that predict user outcomes for performing a protocol or function. Non-limiting examples of models include naive Bayes models, random forest or gradient boosting models, logistic regression models, deep learning neural networks, ensemble models, supervised learning models, unsupervised learning models, collaborative filtering models, and any other suitable machine learning or AI models.
[0161] It will be understood that the cloud-based application can be configured to perform intelligent functions with respect to consulting external physicians, determining diagnoses, and suggesting treatments for any disease, condition, area of study, or disorder, including, but not limited to, COVID-19, oncology, including the following cancers: cancers of the lung, breast, colorectal, prostate, stomach, liver, cervix (uterine cervix), esophagus, bladder, kidney, pancreas, endometrium, oral cavity, thyroid, brain, ovary, skin, and gallbladder; solid tumors such as sarcomas and carcinomas; cancers of the immune system, including lymphomas (such as Hodgkin's lymphoma or non-Hodgkin's lymphoma); and cancers of the bone marrow, such as blood (hematological cancers) and leukemias (such as acute lymphocytic leukemia (ALL) and acute myeloid leukemia (AML)), lymphomas, and myelomas. Further disorders include blood disorders such as anemia, bleeding disorders such as hemophilia, blood clots, ophthalmologic disorders including diabetic retinopathy, glaucoma, and macular degeneration, neurological disorders including multiple sclerosis, Parkinson's disease, spinal muscular atrophy, Huntington's disease, amyotrophic lateral sclerosis (ALS), and Alzheimer's disease, and autoimmune disorders including multiple sclerosis, diabetes, systemic lupus erythematosus, myasthenia gravis, inflammatory bowel disease (IBD), psoriasis, Guillain-Barré syndrome, chronic inflammatory demyelinating polyneuropathy (CIDP), Graves' disease, Hashimoto's thyroiditis, eczema, vasculitis, allergies, and asthma.
[0162] Other diseases and disorders include, but are not limited to, kidney disease, liver disease, heart disease, stroke, gastrointestinal disorders such as celiac disease, Crohn's disease, diverticular disease, irritable bowel syndrome (IBS), gastroesophageal reflux disease (GERD) and peptic ulcers, arthritis, sexually transmitted diseases, high blood pressure, bacterial and viral infections, parasitic infections, connective tissue diseases, celiac disease, osteoporosis, diabetes, lupus, diseases of the central and peripheral nervous system such as attention deficit / hyperactivity disorder (ADHD), catalepsy, encephalitis, epilepsy and seizures, peripheral neuropathy, meningitis, migraine, myelopathy, autism, bipolar disorder, and depression.
[0163] IV.A. The cloud-based application allows user devices to broadcast consultation requests to other user devices and automatically condense subject records to comply with data privacy regulations. 2 is a flowchart illustrating a process 200 performed by a cloud-based application to deliver a condensed subject record to a user device in connection with a consultation broadcast requesting assistance regarding the subject's treatment. Process 200 can be performed by cloud server 135 to enable user devices associated with different entities (e.g., hospitals) to collaborate or consult regarding the treatment for a subject while complying with data privacy regulations.
[0164] Process 200 begins at block 210, where cloud server 135 receives a set of attributes from a user device. Each attribute in the set of attributes can represent any characteristic of the subject (e.g., a patient). The set of attributes can be identified by a user using an interface provided by cloud server 135. For example, the set of attributes identifies the subject's demographic information and recent symptoms experienced by the subject. Non-limiting examples of demographic information include age, gender, ethnicity, state or city of residence, income range, education level, or any other suitable information. Non-limiting examples of recent symptoms include subjects who currently or recently (e.g., at last clinic visit, intake, within 24 hours, within one week) had a particular symptom (e.g., shortness of breath, fever above a threshold temperature, blood pressure above a threshold blood pressure).
[0165] In block 220, the cloud server 135 generates a record for the subject. The record can be a data element including one or more data fields. The record indicates each of a set of attributes associated with the subject. The record can be stored in a central data store, such as the data registry 140 or any other cloud-based database. In block 230, the cloud server 135 receives a request submitted by a user using the interface. The request can be to initiate a consultation broadcast. For example, a user associated with an entity is a physician at a medical center treating the subject. The user can operate a user device to access a cloud-based application and broadcast a request for assistance in treating the subject. The broadcast can be sent to a set of other user devices associated with a different entity.
[0166] In block 240, the cloud server 135 queries the central data store using one or more recent symptoms included in the set of attributes associated with the subject. The query result includes a set of other records, each record in the set of other records associated with a different subject. In some examples, the cloud server 135 can query the central data store to identify other subject records similar to the subject record. Similarity can be determined by comparing the transformed representation of the entire subject record with the transformed representation of each of the other subject records. The comparison of the transformed representations can result in a distance (e.g., Euclidean distance) representing the similarity between the two subject records. In other examples, similarity can be determined based on values included in data elements. For example, the target subject record can include a subject data element including text representing symptoms experienced by the subject. Each other subject record stored in the central data store can also include a data element including text representing the associated subject's symptoms. The cloud server 135 can convert the text included in the subject data element into a numerical representation using techniques described above (e.g., trained convolutional neural networks, text vectorization techniques such as Word2Vec, etc.). The numeric representation of the text included in the subject data element can be compared to the numeric representation of the text included in the matching data element for each other subject record. The result of the comparison (e.g., in a domain space such as Euclidean space) between the two numeric representations can indicate the degree to which the text included in the subject data element is similar to the text included in the data element for the other subject record. In block 250, the cloud server 135 identifies a set of destination addresses (e.g., other user devices associated with different entities). Each destination address in the set of destination addresses is associated with a caregiver of another subject associated with one or more other records in the set of other records identified in block 240. In block 260, the cloud server 135 generates a condensed representation of the record for the subject.The condensed representation of a record omits, obscures, or obfuscates at least a portion of the record. The condensed representation of a record cannot be used to uniquely identify the subject associated with the record and can therefore be exchanged between external systems without violating data privacy regulations. The cloud server 135 can perform any masking or obfuscation technique to generate the condensed representation of the record.
[0167] At block 270, the cloud server 135 utilizes a contracted representation of the record with a connection input component (e.g., a selectable link, such as a hyperlink, that establishes a communication channel) for each destination address in the set of destination addresses. The connection input component may be a selectable element presented at each destination address. Non-limiting examples of a connection input component include a button, a link, an input element, and other suitable selectable elements. At block 280, the cloud server 135 receives a communication from a destination device associated with the destination address. The communication includes an indication that a user operating the destination device selected the connection input component associated with the contracted representation of the record. At block 290, the cloud server 135 establishes a communication channel between the user device for which the connection input component was selected and the destination device. The communication channel enables a user operating the user device (e.g., a physician treating the subject) to exchange messages or other data (e.g., a video feed) with a destination device (e.g., a physician at another hospital who has agreed to assist in the patient's treatment) associated with the destination address for which the connection input component was selected.
[0168] In some embodiments, the cloud server 135 is configured to automatically determine the location of the user device and the location of the destination device for which the connection input component was selected. The cloud server 135 can also compare the locations to determine whether to generate a condensed representation of the record. For example, in block 260, the cloud server 135 can generate a condensed representation of the record because the cloud server 135 determines that each destination address in the set of destination addresses is not co-located with the user device that initiated the consultation broadcast. In this case, the cloud server 135 can automatically determine to generate a condensed representation of the record to comply with data privacy regulations. As another example, if the set of destination addresses are associated with the same entity as the user device that initiated the consultation broadcast, the cloud server 135 can transmit the record in its entirety (e.g., without obfuscating any portion of the record) to the destination device associated with the destination address while complying with data privacy regulations.
[0169] In some embodiments, the cloud server 135 generates a plurality of other condensed record representations. Each of the plurality of other condensed record representations is associated with a different subject. The cloud server 135 transmits the plurality of other condensed record representations to a user device and receives a communication from the user device identifying a selection of a subset of the plurality of other condensed record representations. Each of a set of destination addresses is represented by one of the condensed record representations. For example, generating the condensed record representation includes determining the jurisdiction of the other subject associated with the condensed record representation, determining data privacy rules governing the exchange of subject records within the jurisdiction, and generating the condensed record representation to comply with the data privacy rules. A first other condensed record representation of the plurality of other condensed record representations may include a particular type of data. A second other condensed record representation of the plurality of other condensed record representations may omit or obscure the particular type of data. For example, the particular type of data may be contact information, identifying information such as name and social security number, and other appropriate information that can be used to uniquely identify the other subject.
[0170] In some implementations, the communication can be received at a central data store. The communication can be sent by a user device operated by a user and can include an identifier of a target subject record for the target subject. Once received at the central data store, the communication can cause the central data store to query a set of stored subject records to identify an incomplete subset of the set of subject records. Each subject record of the incomplete subset can be identified and included in the incomplete subset because the subject record is determined to be similar to the target subject record along at least one dimension. The similarity between two subject records along a dimension can represent similarity with respect to data elements of the subject records, such as similarity with respect to symptoms, diagnoses, treatments, or any other suitable data elements. The one or more dimensions along which similarity or dissimilarity is determined can be automatically defined or user-defined. Determining the similarity or dissimilarity between the target subject record and each subject record of the set of subject records stored in the central data store may include at least the following operations: retrieving the target subject record based on the identifier included in the communication; generating a transformed representation of the target subject record (or retrieving an existing transformed representation of the target subject record); and performing a clustering operation using the transformed representation of the target subject record and the transformed representation of each subject record in the set of subject records. The clustering operation may be performed on one or more dimensions (e.g., one or more features of the subject records). For example, the clustering operation may cluster the set of subject records stored in the central data store based on a data element that includes a value representing a symptom of the subject. The transformed representation of the target subject record may include a vector representation of a data element that includes a value representing a symptom of the subject. The vector representation of this data element of the target subject record may be compared to the vector representation of a corresponding data element in each subject record in the set of subject records to define clusters of subject records.Each cluster of subject records can define a group of one or more subject records that share a common characteristic associated with a data element selected as a dimension of similarity. For each cluster of subject records, a Euclidean distance between the transformed representation of the target subject record and other transformed representations of the set of subject records can be calculated. A subject record can be determined to be similar to a target subject record, for example, if the Euclidean distance between the transformed representation of the subject record and the transformed representation of the target subject record is within a threshold.
[0171] IV.B. Update of Shareable Treatment Plan Definitions Based on Aggregate User Integration 3 is a flowchart illustrating a process 300 for monitoring a user's integration of a treatment plan definition (e.g., a decision tree or processing workflow) and automatically updating the treatment plan definition based on the monitoring results. Process 300 can be executed by cloud server 135 to enable a user device to define a treatment plan for treating a population of symptomatic subjects. The user device may distribute the treatment plan definition to user devices connected to an internal or external network. A user device that receives the treatment plan definition can decide whether to integrate the treatment plan definition into a custom rule base. Integration into the custom rule base can be monitored and used to automatically modify the treatment plan definition.
[0172] In block 310, the cloud server 135 stores interface data that causes a treatment plan definition interface to be displayed when a user device loads the interface data. The treatment plan definition interface is provided to each user device in the set of user devices when the user device accesses the cloud server 135 to navigate to the treatment plan definition interface. In some embodiments, the treatment plan definition interface allows a user to define a treatment plan for treating a population of subjects with a condition (e.g., lymphoma).
[0173] In block 320, the cloud server 135 receives a set of communications. Each communication in the set of communications is received from a user device in the set of user devices and generated in response to an interaction between the user device and the treatment plan definition interface. In some embodiments, the communication includes, for example, one or more criteria for defining a population of subject records. Each criterion may be represented by a variable type. For example, the variable type may be a value or variable used as a condition for the criterion. The variable type of a rule's criterion may also be any value of a condition that constrains a population of subjects into an incomplete subgroup. For example, the variable type of a rule defining a population of pregnant women is "IF 'subject is pregnant'." The criteria may be a filter condition for filtering a pool of subject records. For example, criteria for defining a population of subject records associated with subjects who may develop lymphoma may include the filter conditions "abnormality in ALK" and "age 60 or older." The communication may also include a specific type of treatment for the condition. A particular type of treatment can be associated with taking a particular action (e.g., undergoing surgery) or refraining from a particular action (e.g., reducing salt intake) suggested to treat symptoms associated with the subjects represented by the population of subject records.
[0174] In block 330, the cloud server 135 stores the set of rules in a central data store, such as the data registry 140 or any other centralized server in the cloud network 130. Each rule in the set of rules includes one or more criteria and a specific treatment type to be included in a communication from a user device. As an illustrative example, the rule represents a treatment workflow for treating lymphoma in a subject. The rule includes the following criteria (e.g., a condition following an "IF" statement) and a next action (e.g., a specific treatment type defined or selected by the user following a "THEN" statement): "IF 'LYMPH NODE BIOPSY INDICATES THE PRESENCE OF LYMPHOMA CELLS' AND 'BLOOD TEST INDICATES THE PRESENCE OF LYMPHOMA CELLS', THEN 'TREAT WITH CHEMOTHERAPY' AND 'ACTIVE SURVEILLANCE'." Additionally, each rule in the set of rules is stored in association with an identifier corresponding to the user device from which the communication was received.
[0175] In block 340, the cloud server 135 identifies a subset of the set of rules that are available among entities via the treatment plan definition interface. The subset of rules can include a subset of the set of rules that are associated with symptoms and distributed to external systems, such as other medical centers, for evaluation. For example, rules can be selected for inclusion in the subset of rules by evaluating characteristics of the rules or identifiers associated with the rules. The characteristics of the rules can include a code or flag stored or attached to the stored rules. The code or flag indicates that the rule is publicly available to external systems (e.g., utilized by the entities).
[0176] In block 350, for each rule in the subset of rules identified in block 340, the cloud server 135 monitors interactions with the rule. The interactions may include an external entity (e.g., external to the entity associated with the user who defined the treatment plan associated with the rule) integrating the rule into a custom rule base. For example, a user device associated with the external entity (e.g., another hospital) evaluates rules available from the external entity. The evaluation includes determining whether the rule is suitable for integration into a rule set defined by the external entity. A rule may be suitable if a user device associated with the external entity indicates that a treatment workflow defined using the rule is suitable for treating the condition corresponding to the rule. Continuing with the illustrative example above, a rule for treating lymphoma may be made available by an external medical center. A user associated with the external medical center determines that the rule for treating lymphoma is suitable for integration into a rule set defined by the external medical center. Thus, after the rule is integrated into the custom rule base defined by the external medical center, other users associated with the external medical center can execute the integrated rule by selecting the integrated rule from the custom rule base. Additionally, the cloud server 135 monitors the integration of available rules by detecting a signal that is generated or caused to be generated when the treatment plan definition interface receives input from a user device associated with an external entity corresponding to the integration of a rule into the custom rule base.
[0177] As another illustrative example, a user device associated with an external entity uses a treatment plan definition to integrate a modified version of a rule specified by an interaction into a custom rule base. The modified version of a rule specified by an interaction is a portion of the rule selected for integration into the custom rule base. Selecting a portion of a rule for integration includes selecting fewer criteria than all criteria included in the rule for integration into the custom rule base. Continuing with the illustrative example above, a user device associated with an external entity selects the criterion "IF 'lymph node biopsy indicates the presence of lymphoma cells'" for integration into the custom rule base, but the user device does not select the criterion "blood test reveals the presence of lymphoma cells" for integration into the custom rule base. Thus, the interaction-specific modified version of the rule integrated into the custom rule base is "IF 'lymph node biopsy indicates the presence of lymphoma cells', THEN 'treat with chemotherapy' and 'active surveillance'." The criterion "blood test reveals the presence of lymphoma cells" is removed from the rule to create a modified version of the rule specified by the interaction and incorporated into the custom rule base.
[0178] In block 360, the cloud server 135 may detect that a modified version of the rule specified by the interaction has been integrated into the custom rule base defined by the external entity. Upon detection, the cloud server 135 may update the rule stored in the central data store of the cloud network 130. The rule may be updated based on the monitored interaction. The term "based on" in this example corresponds to "after evaluating" or "using the results of the evaluation" of the monitored interaction. For example, the cloud server 135 detects that a user device associated with the external entity has integrated a modified version of the rule specified by the interaction. In response to detecting the modified version of the rule specified by the interaction, the cloud server 135 may update the rule stored in the central data store from the existing rule to the modified version of the rule specified by the interaction.
[0179] In some embodiments, cloud server 135 updates the rules by generating an updated version that is utilized across external entities. Another, original version can remain unupdated and available to a user associated with the user device as the source of one or more received communications that identified the criteria and a particular type of action. For example, cloud server 135 updates a rule stored in a central data store, but cloud server 135 does not update another rule in the set of rules stored in the central data store.
[0180] In some embodiments, the cloud server 135 can update the rules when an update condition is met. The update condition may be a threshold. For example, the threshold can be the number or percentage of external entities that have integrated a modified version of the rule into their custom rule base. As another example, the update condition may be determined using the output of a trained machine learning model. Illustratively, the cloud server 135 can input detection signals received from external entities into a multi-armed bandit model that automatically determines whether and / or when to employ a rule and / or whether and when to employ an updated version of the rule. By way of non-limiting example only, the rules may be defined as executable code such that, when executed, the rules automatically query a central data store to identify a subset of the set of subject records for further analysis. Furthermore, the rules may include one or more treatment protocols for treating subjects associated with the identified subset of subject records. The rules can define a subset of the set of subject records and be defined as a workflow for processing the subset associated with the subset of subject records. For example, a rule may include one or more criteria for filtering subject records from a set of subject records and executing a particular treatment protocol on subjects associated with the remaining subject records (e.g., the subject records remaining after filtering has been performed on the set of subject records). Although a rule is defined by a user of a first entity, the rule can be accepted (e.g., integrated into the rule base of the second entity), modified, or completely rejected by an external user (e.g., a doctor working at another hospital) of a second entity (e.g., the first entity and the second entity are two different medical facilities). In some examples, a feedback signal can be sent to cloud server 135 each time the external user of the second entity accepts a rule and thus fully integrates it into its rule base.In another example, a feedback signal may be sent to the cloud server 135 each time a user of the second entity changes a rule. In another example, a feedback signal may be sent to the cloud server 135 each time a user of the second entity completely rejects a rule. In each of the above examples, the feedback signal may include the rule (e.g., a rule identifier) and data indicating whether the rule was accepted, modified, or rejected. The multi-armed bandit model (executable by the cloud server 135) may be configured to intelligently select one of the original rule, the modified rule, or an entirely different rule to broadcast to external users of other entities. The selection of the original rule, the modified rule, or the different rule may be based at least in part on the configuration of the multi-armed bandit. In some examples, the multi-armed bandit may be configured by an ε-greedy search technique. In the ε-greedy search technique, the multi-armed bandit model may select the original rule to broadcast to external users of other entities with a probability of 1-ε, where ε represents the probability of searching for a new or modified rule. Thus, the multi-armed bandit model can select a modified version of the original rule or an entirely new rule with a defined probability of ε. The multi-armed bandit model can modify ε based on feedback signals received from other entities. For example, if the feedback signals indicate that a rule has been modified in a particular way by different external users a threshold number of times, the multi-armed bandit model can learn to select the modified rule in a particular way to broadcast to external users instead of broadcasting the original rule.
[0181] In some embodiments, the cloud server 135 identifies multiple rules in the set of rules that include criteria corresponding to the same variable type and identify the same or similar types of treatments. The variable type may be a value or variable used as a condition of the criteria. The variable type of the criteria of a rule may also be any value of a condition that constrains a population of subjects into subgroups. For example, the variable type of a rule that defines a population of pregnant women is "IF 'subject is pregnant'." The cloud server 135 determines the new rule, which is a contracted representation of the multiple rules, when the new rule is sent to a server typically operated by another entity.
[0182] In some embodiments, the cloud server 135 provides another interface configured to receive a set of attributes from a subject, e.g., a user who operates a user device to access the other interface and select a subject record that includes the set of attributes using the other interface. The selection of the subject record can cause the cloud server 135 to receive the set of attributes for the subject. The cloud server 135 identifies (e.g., determines) a particular rule whose criteria are met based on the set of attributes for the subject. For example, the cloud server 135 evaluates the set of attributes for the subject record against the criteria of a rule stored in a central data store. Illustratively, if the set of attributes includes a data field with a value of "pregnant" and the rule includes a single criterion of "IF 'subject is pregnant'," the cloud server 135 identifies this rule. The cloud server 135 updates the other interface to present the particular rule and each particular type of action associated with the particular rule.
[0183] In some embodiments, the rule criteria are variable types related to specific demographic variables and / or specific symptom-type variables. Non-limiting examples of demographic variables include any information that characterizes the subject's demographics, such as age, sex, ethnicity, race, income level, education level, location, and other suitable items of demographic information. Non-limiting examples of symptom-type variables indicate whether the subject currently or recently (e.g., at the last visit, at intake, within 24 hours, within one week) has experienced a specific symptom (e.g., shortness of breath, loss of consciousness, fever above a threshold temperature, blood pressure above a threshold blood pressure).
[0184] In some embodiments, cloud server 135 monitors data in a registry of subject records, such as subject records stored in data registry 140. Cloud server 135 monitors data in the registry of subject records for each rule in the subset of rules (identified in block 340). Cloud server 135 identifies a set of subjects for which the criteria of the rule are met and a particular treatment has been previously prescribed for the subject. For each of the set of subjects, cloud server 135 identifies the subject's reported status as indicated from or using an assessment or test. For example, the reported status is any information that characterizes the subject's status in aspects such as whether the subject has been discharged from the hospital, whether the subject is alive, the subject's blood pressure measurement, the number of times the subject has awakened during a sleep stage, and other suitable status. Cloud server 135 determines an estimated responsiveness metric for the set of subjects to a particular treatment based on the reported status. For example, if a particular treatment of a rule is to prescribe medication, the estimated responsiveness metric is a representation of the extent to which the medication addressed the symptoms or conditions experienced by the subject. As a non-limiting example, the estimated responsiveness metric for a set of subjects can be the average, weighted average, or any sum of the scores assigned to each subject in the set of subjects. The score can represent or measure the effectiveness of the subject's responsiveness to the treatment. In some examples, the cloud server 135 can generate a score representing the effectiveness of the subject's responsiveness to the treatment by using clustering techniques. Illustratively, as a non-limiting example, the set of subject records can represent subjects who previously underwent a particular treatment protocol to treat a condition. Each subject record in the set of subject records can be labeled (e.g., by a user) as having one of a positive responsiveness to the particular treatment protocol, a neutral responsiveness to the particular treatment protocol, or a negative responsiveness to the particular treatment protocol. The set of subject records can then be divided into three subsets (e.g., clusters).A first subset of subject records can correspond to subjects who had a positive responsiveness to a particular treatment protocol, a second subset of subject records can correspond to subjects who had a neutral responsiveness to a particular treatment protocol, and a third subset of subject records can correspond to subjects who had a neutral responsiveness to a particular treatment protocol. The cloud server 135 can convert each subject record in the first subset of subject records to a transformed representation according to the implementations described above. The cloud server 135 can also convert each subject record in the second subset of subject records to a transformed representation using the techniques described above. Finally, the cloud server 135 can convert each subject record for a third subject in the subject records to a transformed representation using the techniques described above. In some implementations, determining the predicted responsiveness of a new subject to a particular treatment protocol can include converting the new subject record of the new subject to a new transformed representation. The new transformed representation can be compared to the transformed representation of each cluster or subset of subject records in a domain space (e.g., Euclidean space). If the new transformed representation is closest to the centroid of the transformed representations associated with the first subset, the new subject is predicted to have a positive responsiveness to the particular treatment. If the new transformed representation is closest to the centroid of the transformed representations of the second subset, the new subject is predicted to have a neutral responsiveness to the particular treatment. Finally, if the new transformed representation is closest to the centroid of the transformed representations of the third subset, the new subject is predicted to have a negative responsiveness to the particular treatment protocol. The centroid may be a multidimensional average of the transformed representations associated with the subsets. The cloud server 135 may cause the estimated responsiveness metrics of the subset of the set of rules and the set of subjects to be displayed or presented in the treatment plan definition interface.
[0185] IV.C. Presentation of Treatment Recommendations with Associated Efficacy Using Treatments Prescribed to Similar Subjects 4 is a flowchart illustrating a process 400 for recommending a treatment for a subject. Process 400 can be performed by cloud server 135 to display the recommended treatments for the subject and the effectiveness of each recommended treatment on a user device associated with a medical entity. The recommended treatments can be identified using results of evaluating the effectiveness of treatments previously prescribed for similar subjects.
[0186] In block 410, the cloud server 135 receives input corresponding to a subject record characterizing an aspect of the subject. The input is received from a user device associated with the entity. Additionally, the input is received in response to the user device selecting or identifying the subject record using an interface associated with an instance of a platform configured to manage a registry of subject records. The user device can access the interface by loading interface data stored on a web server (not shown) connected within the cloud network 130. The web server may be included in the cloud server 135 or may run on the cloud server.
[0187] In block 420, the cloud server 135 extracts a set of subject attributes from the subject record received in block 410. The subject attributes characterize aspects of the subject. Non-limiting examples of subject attributes include any information found in an electronic health record, any demographic information, age, sex, ethnicity, recent or past symptoms, symptoms, symptom severity, and any other suitable information that characterizes the subject.
[0188] In block 430, the cloud server 135 generates an array representation of the subject record using the set of subject attributes. For example, the array representation is a vector representation of the values contained in the subject record. The vector representation may be a vector in a domain space, such as Euclidean space. However, the array representation may be any numerical representation of the values of the data fields of the subject record. In some embodiments, the cloud server 135 may perform a feature decomposition technique, such as SVD, to generate values that represent the set of subject attributes of the array representation of the subject record.
[0189] In block 440, the cloud server 135 accesses a set of other sequence representations characterizing a plurality of other subjects. The sequence representations in the set of other sequence representations may be vector representations of subject records characterizing another subject (e.g., one of the plurality of other subjects).
[0190] In block 450, the cloud server 135 determines a similarity score representing the similarity between the sequence representation representing the subject and each of the sequence representations of other subjects. For example, the similarity score is calculated using a function of the distance (in domain space) between the sequence representation representing the subject and the sequence representations representing the other subjects. As an illustrative and non-limiting example, the similarity score may be calculated using a range from "0" to "1," with "0" representing a distance above a defined threshold and "1" representing the sequence representations having no distance between them. By way of illustrative and non-limiting example only, the similarity score may be based on the Euclidean distance between two sequence representations (e.g., vectors).
[0191] At block 460, the cloud server 135 identifies a first subset of the plurality of other subjects. A subject may be included in the first subset if the similarity score associated with the subject falls within a predetermined absolute or relative range. Similarly, at block 470, the cloud server 135 identifies a second subset of the plurality of other subjects. However, a subject may be included in the second subset if the subject's similarity score falls within another predetermined range.
[0192] In block 480, the cloud server 135 obtains record data for each subject in the first and second subsets of the plurality of other subjects. The record data includes attributes included in the subject record that characterize the subject. For example, the subject record data identifies a treatment the subject received and the subject's responsiveness to the treatment. Responsiveness to the treatment can be represented by text (e.g., "The subject responded positively to the treatment") or a score indicating the degree to which the subject responded positively or negatively to the treatment (e.g., a score from "0" to "1," with "0" indicating a negative responsiveness and "1" indicating a positive responsiveness). In some examples, treatment responsiveness can indicate the degree to which the subject responded positively to a treatment previously administered to the subject. For example, treatment responsiveness may be numeric (e.g., a score from "0" to "10") or non-numeric (e.g., words assigned to represent responsiveness, such as "positive," "neutral," or "negative"). In some examples, treatment responsiveness for previously treated subjects may be user-defined. In other examples, treatment responsiveness may be automatically determined based on test or measurement results from a user. For example, treatment responsiveness may be determined automatically based on values contained in a blood test performed on the subject.
[0193] At block 490, the cloud server 135 generates an output that is presented to an interface on the user device. The output may, for example, indicate one or more treatment recommendations for the subject. The one or more treatment recommendations may be determined based, for example, on treatments received by other subjects in the first and second subsets, treatment responsiveness of the subjects in the first and second subsets, and differences between subject attributes of the subjects in the second subset and subject attributes of the subjects.
[0194] In some embodiments, the cloud server 135 determines that the subject and one of the subjects from the first or second subset are being or were treated by the same medical entity. The cloud server 135 determines that the subject from the first or second subset and another subject are being or were treated by different medical entities. The cloud server 135 can utilize different obfuscated versions of the subject's record via an interface. The cloud-based application can automatically provide different obfuscated versions of the record to an entity based on various constraints on data sharing imposed by data privacy regulations in different jurisdictions. In some embodiments, the cloud server 135 identifies the first and second subsets of subject records by performing a clustering operation on a transformed representation of the set of subject records.
[0195] IV.D. Automatic Obfuscation of Query Results from External Entities 5 is a flowchart illustrating a process 500 for obfuscating query results to comply with data privacy regulations. Process 500 may be performed by cloud server 135 as an execution rule to ensure that data sharing of subject records with external entities complies with data privacy regulations. A cloud-based application may enable a user device to query data registry 140 for subject records that meet query constraints. However, the query results may include data records originating from external entities. Thus, process 500 enables cloud server 135 to provide additional information about treatment from external entities to a user device while complying with data privacy regulations.
[0196] In block 510, the cloud server 135 receives a query from a user device associated with a first entity. For example, the first entity may be a medical center associated with a first set of subject records. The query may include a set of symptoms associated with a medical condition or any other information that constrains the query search of the data registry 140.
[0197] At block 520, the cloud server 135 queries the database using the query received from the user device. At block 530, the cloud server 135 generates a dataset of query results corresponding to a set of symptoms and associated with a medical condition. For example, the user device submits a query for subject records of subjects diagnosed with lymphoma. The query results include at least one subject record from a first set of subject records (originating from or created at a first entity) and at least one subject record from a second set of subject records associated with a second entity (e.g., a different medical center than the first entity). Each of the subject records from the first set of subject records and the second set of subject records can include a set of subject attributes. The subject attributes can characterize any aspect of the subject.
[0198] In block 540, the cloud server 135 presents (e.g., utilizes or otherwise makes available) the entire subject attribute set to the user device for the subject records included in the first subject record set because these records originate from the first entity. Fully presenting the subject records includes making the set of attributes included in the subject record available by the user device for evaluation or interaction using an interface. In block 550, the cloud server 135 also, or alternatively, utilizes an incomplete subset of the subject attribute set for each subject record included in the second subject record set to the user device. Providing the incomplete subset of the set of subject attributes provides anonymity to the subject because the incomplete subset of the subject attributes cannot be used to uniquely identify the subject. For example, providing the incomplete subset may include available four of ten subject attributes to anonymize the subject associated with the ten subject attributes. In some embodiments, in block 550, the cloud server 135 utilizes an obfuscated set of subject attributes for each subject record included in the second subject. Obfuscating the set of attributes includes reducing the granularity of the information provided. For example, instead of utilizing the subject attribute of the subject's address, the obfuscated attribute may be the zip code or the state the subject lives in. Regardless of whether an incomplete subject or obfuscated subset is available, the cloud server 135 anonymizes the subject associated with the subject record.
[0199] IV.E. Chatbot Integration with Self-Learning Knowledge Bases 6 is a flowchart illustrating a process 600 for communicating with a user using a bot script, such as a chatbot. Process 600 can be executed by the cloud server 135 to automatically link new questions provided by a user to existing questions in a knowledge base to provide answers to the new questions. The chatbot can be configured to provide answers to questions associated with symptoms.
[0200] In block 605, the cloud server 135 defines a knowledge base including a set of answers. The knowledge base may be a data structure stored in memory. The data structure stores text representing a set of answers to defined questions. Each answer may be selectable by the chatbot in response to a question received from the user device during a communication session. The knowledge base may be automatically defined (e.g., by retrieving text from a data source and parsing the text using NLP techniques) or user-defined (e.g., by a researcher or physician).
[0201] In block 610, the cloud server 135 receives a communication from a particular user device. The communication corresponds to a request to initiate a communication session with a particular chatbot. For example, a doctor or a subject may operate the user device to communicate with the chatbot in a chat session. The cloud server 135 (or a module stored within the cloud server 135) may manage or establish the communication session between the user device and the chatbot. In block 615, the cloud server 135 receives a particular question from the particular user device during the communication session. The question may be a string of characters that is processed using NLP techniques.
[0202] In block 620, the cloud server 135 queries the knowledge base using at least some words extracted from the specific question. The words may be extracted from the string representing the specific question using NLP techniques. In block 625, the cloud server 135 determines that the knowledge base does not contain a representation of the specific question. In this case, the received question may be newly submitted to the chatbot. In block 630, the cloud server 135 identifies other question representations from the knowledge base. The cloud server 135 can identify other question representations by comparing the question received from the user device with other question representations stored in the knowledge base. For example, if similarity is determined based on analyzing the question representation using NLP techniques, the cloud server 135 identifies other question representations.
[0203] At block 635, the cloud server 135 retrieves an answer from a set of answers associated with other question representations in the knowledge base. At block 640, the answer retrieved at block 635 is sent to the particular user device as an answer to the received question, even if the knowledge base did not contain the received question representation. At block 645, the cloud server 135 receives an indication from the particular user device. For example, the indication may be received in response to the user device indicating that the answer provided by the chatbot responded to the particular question.
[0204] In block 650, the cloud server 135 updates the knowledge base to include the particular question representation or a different representation of the particular question. For example, storing the question representation may include storing keywords included in the question in a data structure. The cloud server 135 may also associate the same or different representation of the particular question with more suitable answers sent to the particular user device.
[0205] In some embodiments, the cloud server 135 accesses a subject record associated with a particular user device. The cloud server 135 determines multiple answers to a particular question. The cloud server 135 then selects an answer from the set of answers. However, the selection of an answer is based at least in part on one or more values included in the subject record associated with the particular user device. For example, the values included in the subject record may represent symptoms recently experienced by the subject. The chatbot may be configured to select an answer according to the symptoms recently experienced by the subject. In some examples, the cloud server 135 may access a learning-to-rank machine learning model trained to predict the order of each answer in the set of answers. The learning-to-rank machine learning model may be trained using a training set of answers. Each answer in the training set of answers may be labeled with one or more symptoms and a relevance score for the symptoms. The relevance score may represent the relevance of the associated answer to a given symptom of the one or more symptoms. The relevance score may be user-defined or automatically determined based on certain factors, such as the frequency of words (e.g., words related to symptoms) in the training answers. The training set of answers may differ from the set of answers used when the chatbot is operating in a production environment. The learn-to-rank machine learning model can learn how to order a set of answers (to be used in a production environment) in terms of their relevance to symptoms (detected from the subject profile) based on patterns learned by the learn-to-rank model (e.g., patterns between the relevance scores associated with a training set of labeled answers for each symptom of one or more symptoms). The chatbot can select an answer from the set of answers to be used in a production environment based on the predicted answer ordering. In some examples, each answer in the set of answers may be associated with a tag or code that indicates one or more symptoms associated with the answer. The cloud server 135 can compare a value representing a symptom recently experienced by the subject with the tag or code associated with each answer.
[0206] V. A network environment configured to provide an oncology application that facilitates intelligent clinical decision-making for subjects diagnosed with cancer FIG. 7 is a block diagram illustrating an example of a network environment for deploying a trained AI model to facilitate subject-specific identification of a treatment and treatment schedule for a subject diagnosed with cancer, according to some embodiments of the present disclosure. The network environment 700 may include a user device 110 and an AI system 702. The user device 110 can interact with the AI system 702 using a network 736 (e.g., any public or private network), thereby facilitating the exchange of communications between the user device 110 and the AI system 702. The AI system 702 may be another implementation of the AI system 145 described with respect to FIG. 1. The user device 110 can be operated by a user, such as a physician or other medical professional treating a subject diagnosed with cancer. The user device 110 can send requests to the AI system 702 using an application programming interface (API) 704 to trigger specific functions (e.g., cloud-based services).
[0207] In some implementations, a physician treating a particular patient can operate user device 110 to access available oncology applications (e.g., modules) using a cloud-based network, such as cloud network 130. The oncology applications can be configured to perform certain predictive functions executed using AI system 702. Non-limiting examples of predictive functions include predicting an individual patient's treatment outcome and subsequent cancer progression based on the patient's mutation sequence across cancer types, creating enriched patient data and predicting progression-free survival associated with a candidate line of treatment, or automatically verifying whether a particular treatment for a subject was selected in accordance with a medical institution's guidelines and potentially suggesting new guidelines for cancer treatment based on the verified treatment. While FIG. 7 shows a single user device 110, it will be understood that any number of user devices or other computing devices, such as a cloud-based server, can interact with AI system 702.
[0208] The AI system 702 can perform predictive functions using, for example, a query resolver 706, an AI model training system 708, and an AI model execution system 710. The query resolver 706, when executed using one or more cloud-based servers of the AI system 702, can include executable code that causes a workflow to be performed, including receiving a query from the user device 110, processing the query by relaying the query to other components of the AI system 702, and resolving the query to complete performance of the predictive function by sending a query response to the user device 110. Several data structures (e.g., databases) for storing data can facilitate the predictive functions that the AI system 702 can perform. In some implementations, the data structures can store training data 716, validation data 718, test data 720, subject records from a data registry 722, an AI model 724, treatments 726, treatment schedules 728, clinical tests 730, and subject group identifiers 732. The various components of the AI system 702 can communicate with each other using a communications network 734.
[0209] The AI model training system 708 can facilitate training of the AI model using the training data 716. For example, the AI model training system 708 can execute code (e.g., executed by a processor, such as a physical or virtual central processing unit (CPU) of a cloud-based server) that inputs the training data 716 into a learning algorithm. The learning algorithm can be executed to detect patterns or correlations between data points included in the training data 716. The detected patterns or correlations can be stored as an AI model that is trained to generate outputs that predict outcomes based on the stored patterns or correlations in response to receiving input (e.g., new, previously unseen input data, such as subject records for subjects not included in the training data 716).
[0210] In some implementations, as described in more detail with respect to Figures 8 and 11, the AI model training system 708 can facilitate training of an unsupervised learning model used to cluster treatment outcomes for particular treatments. In other implementations, as described in more detail with respect to Figures 9 and 12, the AI model training system 708 can facilitate training of a knowledge graph (or knowledge model) used to predict progression-free survival of a particular treatment for a particular subject having a particular cancer type. In other implementations, as described in more detail with respect to Figures 10 and 13, the AI model training system 708 can facilitate training of a neural network model that automatically classifies the reasons that contributed to the selection of a proposed or predicted treatment as guideline-compliant or guideline-noncompliant.
[0211] The learning algorithms executed by the AI system 702 can include any supervised, unsupervised, semi-supervised, reinforcement, and / or ensemble learning algorithms. Non-limiting examples of learning algorithms that can be executed by the AI system 702 are included in Table 1 below. The selection of a learning algorithm by the AI system 702 to train an AI model can be based, for example, on the type and size of at least a portion of the training data 716 and the target predictive results aimed at the predictive functions that the AI system 702 is capable of performing. The various learning algorithms provided in Table 1 can be used as learning algorithms to train any of the AI-based models described herein. [Table 1]
[0212] Furthermore, during the process of training various AI models, the AI model training system 708 can interact with training data 716, validation data 718, and testing data 720. The training data 716 is a data set that is input to a learning algorithm. The learning algorithm detects patterns, correlations, or relationships between data points in the training data 716. However, the patterns, correlations, or relationships (e.g., parameters) detected by the learning algorithm may cause the training data 716 to be overfitted. Overfitting occurs when the analysis performed by the learning algorithm (e.g., that generated the pattern, correlation, or relationship) corresponds exactly or substantially exactly to the training data 716. In this case, the analysis performed by the learning algorithm may not accurately serve as a basis for predicting new, previously unseen input data. Therefore, the validation data 718 is a different data set from the training data 716 and is used to correct the patterns, correlations, or relationships to prevent overfitting of the training data 716. When multiple learning algorithms are run on the training data 716, the validation data 718 can be used to identify the learning algorithm with the best performance on new input data (e.g., input data not included in the training data 716). The validation data 718 can be used to generate an error function that can be evaluated to determine the performance of each learning algorithm on the new input data. For example, patterns, correlations, or relationships detected in the training data 716 by each of the various learning algorithms can be stored in various AI models. The error function of each AI model on the new input data can be evaluated using the validation data 718. The AI model with the smallest error function can be selected. Finally, the test data 720 is a separate data set independent of each of the training data 716 and the validation data 718. The test data 720 can be input into a selected AI model to test the overall performance of the selected AI model.
[0213] In some implementations, the training data 716, the validation data 718, and the testing data 720 can be segmented across a single larger dataset. For example, the dataset can be segmented into three data subsets. The training data 716 can be one of the three data subsets, the validation data 718 can be another of the three data subsets, and the testing data 720 can be the last of the three data subsets. In some implementations, the dataset segmented into three or more subsets can include any data or data type. Non-limiting examples of data or data types that can be included in the datasets from which the training data 716, validation data 718, and / or testing data 720 are generated include radiology image data, MRI data, genetic profile data, clinical data (e.g., measurements, treatment, treatment response, diagnosis, severity, medical history), subject-generated data (e.g., notes entered by a subject with breast cancer), physician or medical professional-generated data (e.g., physician's notes), audio data representing telephone calls between a patient and a physician or other medical professional, administrative data, claims data, health surveys (e.g., Health Risk Assessment (HRS) surveys), third-party or vendor information (e.g., out-of-network lab results), public databases related to the subject (e.g., medical journals related to the subject's condition), subject demographics, immunizations, radiology reports, pathology reports, utilization information, metadata representing biological samples, social data (e.g., education level, employment status), community specifications, etc. In some examples, at least a portion of the subject record can be initially identified via communications from a device operated by the subject (e.g., received at a caregiver device and / or a remote server). In some implementations, at least some features of the subject record include or are based on one or more photographs (e.g., collected on the subject's device or collected by a medical professional operating an imaging device).In some examples, at least a portion of the subject-specific data was initially identified via and / or received from an electronic medical record corresponding to the subject.
[0214] The AI model execution system 710 may be implemented using executable code that, when executed by a processor (e.g., a physical or virtual CPU in a cloud-based network such as cloud network 130), executes an instance of a particular trained AI model and generates output. The output may predict a particular clinical decision related to oncology or other particular cancers, such as breast cancer, lung cancer, colon cancer, and blood cancer.
[0215] To illustrate, and by way of non-limiting example only, the AI model execution system 710 receives a request from the query resolver 706 (e.g., the request originates from a user device 110 operated by a user, such as a physician evaluating different options for treatment lines for a particular subject). The request from the user device 110 is for the AI system 702 to predict the treatment outcome of administering alpelisib (a chemotherapy drug) to a particular subject with breast cancer that has a PIK3CA mutation. PIK3CA mutations are involved in many types of cancer, including breast cancer, lung cancer, colon cancer, ovarian cancer, brain cancer, and gastric cancer. PIK3CA mutations produce an altered p110α subunit, allowing PI3K to signal without stopping. However, unconstrained signaling can cause cells to divide in an uncontrolled manner, potentially resulting in cancer. Alpelisib chemotherapy treatment inhibits PI3K, which reduces the likelihood of tumor growth by imposing constraints on PI3K signaling. However, alpelisib can have side effects that vary in severity. The query resolver 706 processes the request and identifies which trained AI model to select to perform the prediction. In response to receiving the request, the AI system 702 generates a prediction of the treatment outcome of administering alpelisib to the particular subject using the selected AI model and a subject record characterizing the particular subject's characteristics. The selected trained AI model generates an output predicting that alpelisib will have low efficacy due to the particular subject's characteristics, such as high insulin resistance, also detected in the particular subject. The prediction functionality described in this example is further described with respect to FIGS. 8 and 11.
[0216] As another illustrative, and non-limiting example only, a physician evaluates whether a tumor necrosis factor (TNF)-related apoptosis-inducing ligand (TRAIL) targeted therapy treatment should be administered to a particular user. While there is a wide range of possible side effects of varying severity, TRAIL treatment is generally intended to reduce tumor growth. The AI system 702 is configured to generate a predictive output to assist the physician in determining the likely side effects of administering TRAIL treatment to a particular subject. Accordingly, the user device 110 operated by the physician sends a request to the AI system 702 to generate a prediction of the side effects that the particular subject is likely to experience in response to receiving the TRAIL treatment. The AI system 702 retrieves or accesses a knowledge graph, which is a graph of nodes representing various relationships between treatments and the side effects of those treatments. The knowledge graph includes a set of triplet sentences: treatment, relationship to side effect, and side effect. Each triplet sentence represents an association between the treatment and the side effect. A learning algorithm can be run on the entire set of triplet sentences in the knowledge graph to learn various relationships between treatments, subject characteristics (e.g., gene mutations), and side effects. The TRAIL treatment and subject record for a particular subject are input into an AI model trained using the knowledge graph. The output is that the side effects of administering the TRAIL treatment to a particular subject are predicted to be rare negative side effects of conditions that promote tumor growth. The prediction function described in this example is further explained with respect to Figures 9 and 12.
[0217] As yet another illustrative, non-limiting example only, the user device 110 sends a request to the AI system 702 to predict whether a physician's reasons for performing a procedure on a particular subject comply with oncology guidelines. For example, the guidelines include the NCCN Guidelines for Clinical Practice in Oncology. Before performing the procedure, the physician can receive an automated assessment of whether the physician's reasons for selecting a particular procedure comply with existing treatment guidelines. The AI system 702 can select a neural network trained to classify the list of reasons and whether the proposed procedure complies with existing oncology guidelines. The prediction functionality described in this example is further described with respect to FIGS. 10 and 13.
[0218] Certain AI models may present the technical challenge of memorizing portions of the training data 716 during the training process. Memorizing portions of the training data 716 can occur when a trained AI model simply outputs data elements included in the training data 716 in response to receiving input data. Data leakage refers to an AI model that simply outputs data elements from the training data in response to the input of new, previously unseen data. In some cases, an AI model memorizes training data when the AI model overfits to the training data. An overfitted AI model memorizes noise included in the training data (e.g., memorizes data elements from the training data that are not relevant to the learning task). Thus, an AI model does not generalize predictions about new, previously unseen input data when the AI model exhibits data leakage.
[0219] If the training data contains confidential or personal data about a subject, a data leak could violate privacy regulations. As an illustrative and non-limiting example, the training data 716 includes a subject record containing a value representing that the subject (characterized by the subject record) has a genetic mutation associated with early onset of Alzheimer's disease. The value representing the presence of an Alzheimer's disease genetic mutation is susceptibility data or personal data. Accordingly, various privacy laws and regulations prohibit the unauthorized disclosure of a subject's confidential or personal data (e.g., the Health Insurance Portability and Accountability Act (HIPAA)). However, if a trained AI model overfits the training data 716, a technical challenge arises: the trained AI model could leak (e.g., unintentionally disclose externally or to an unauthorized user) the value representing that the subject has an Alzheimer's disease genetic mutation. In some scenarios, a privacy violation could occur if an adversarial user device (e.g., operated by a user intentionally attempting to extract sensitive information from the AI model) is able to send inputs to the trained AI model and receive the corresponding outputs generated by the AI model. For example, if an adversarial user device accesses a trained AI model using the public API, the adversarial user device can send inputs to the trained AI model and receive outputs generated by the trained AI model. The adversarial user device can then evaluate various outputs received from the trained AI model to infer confidential or personal data about the training data used to train the AI model. Non-limiting examples of confidential or personal data that can be inferred include values indicative of the presence of a particular genetic mutation in a particular subject; the presence or absence of a subject record in the training data; the presence or absence of a particular subject in a particular clinical trial; a correlation between a phenotype exhibited by a particular subject and the particular subject's genetic predisposition to developing a particular disease, such as breast cancer; characteristics of a particular subject's genetic profile; and any other confidential or personal data.
[0220] To address the technical challenges related to data leakage, certain aspects and features of the present disclosure relate to configuring a data leakage detector 712 to detect and prevent data leakage when the AI model execution system 710 executes any of the trained AI models stored in the AI model data store 724. In some implementations, the data leakage detector 712 can execute specific data leakage prevention protocols on the training data 716, the validation data 718, the testing data 720, and / or the AI model 724. Executing the data leakage prevention protocols on the training data 716, the validation data 718, the testing data 720, and / or the AI model 724 can suppress or prevent leakage of sensitive data by the trained AI model. Non-limiting examples of data leakage prevention protocols executed on the data include encryption of sensitive or personal data included in subject records, data sanitization, data normalization, robust statistics, adversarial training, differential privacy, federated learning, homomorphic encryption, and other suitable techniques for suppressing or preventing leakage of sensitive data characterizing the subject.
[0221] Referring again to FIG. 7 , a subject record may include data elements that characterize subject features using numerous dimensions (e.g., hundreds or thousands of feature dimensions). While certain feature dimensions within a subject record may be useful for a target task, other feature dimensions within the subject record may represent noisy data (e.g., features that are not useful for the subject task). The high dimensionality of subject records creates technical challenges related to inputting the subject record (or a numerical representation thereof) as part of the predictive functions provided by various AI models associated with the AI system 702. Certain aspects and features of the present disclosure relate to a noisy feature detector 714 that provides a solution to the above-mentioned technical challenges. In some implementations, the noisy feature detector 714 can be configured to convert a high-dimensional subject record into a low-dimensional subject record by classifying a subset of subject features of a set of subject features included in the subject record as noise. For example, the noisy feature detector 714 can implement a two-class classification model trained to classify subject features as either predictive of a subject task or noise. It will be appreciated that the noisy feature detector 714 can also be a multi-class classification model capable of classifying subject features of subject records into one or more of multiple classes (e.g., noisy data, useful but not predictive of subject task, and useful and predictive of subject task). Reducing the dimensionality of subject records improves the computational efficiency of the AI system 702 by reducing the number of feature dimensions of subject records that the AI model execution system 710 processes when providing predictive functions. Non-limiting examples of techniques for reducing the dimensionality of subject records include reducing features based on criteria, reducing features based on feature categories, feature selection techniques, eliminating features classified as noise by a trained classifier model, and other suitable techniques.
[0222] VI. A network environment configured to provide an oncology application that uses artificial intelligence techniques to predict treatment outcomes and cancer progression Cancerous primary mutations can be preferentially associated with secondary or tertiary mutations that further contribute to the development of cancer in a subject. For example, a particular genetic mutation often associated with cancer may not cause cancer by itself; rather, a mixture of several preferentially associated mutations, when present together and activated in a specific order, may cause cancerous cell proliferation. In certain cancers, for example, tumors can only develop if a primary mutation is activated followed by a secondary mutation. Therefore, selecting a targeted therapy treatment can be difficult because targeting (e.g., inhibiting) one genetic mutation may activate a secondary or tertiary genetic mutation, further complicating the subject's cancer. Identifying the effects of specific targeted therapy treatments on a given genetic mutation across various cancer types can benefit physicians.
[0223] 8 is a block diagram illustrating an example of a network environment for deploying an AI model trained to predict treatment outcomes and cancer progression for subjects diagnosed with cancer, according to some embodiments of the present disclosure. The network environment 800 may include a user device 110 and an AI system 802. The AI system 802 may be similar to the AI system 702 shown in FIG. 7. However, the components of the AI system 802 may differ from the components of the AI system 702.
[0224] The AI system 802 can be configured to identify subjects similar to a particular subject in terms of mutation order. The AI system 802 can be configured to filter, cluster, and generate similarities using an AI model and subject records. In some implementations, the AI system 802 can be configured to train a neural network to learn how to detect similar subjects across cancer types, such that the similarity is based on patterns detected in the subjects' mutation profiles. The mutation profile, such as the mutation order indicated by the mutation profile, does not need to be exactly the same between two subjects for the subjects to be considered similar. In other implementations, the AI system 802 can be configured to train a dynamic neural network to learn aspects of similarity between two or more subject records, such as, for example, that the similarity is based on mutation order or other molecular characteristics indicated by the mutation profile. As a non-limiting example, the dynamic neural network is configured with input-dependent neurons that allow the dynamic neural network to adaptively modify to accommodate various inputs. In some implementations, the AI system 802 can be configured to learn similarities between two or more subject records using meta-learning techniques. For example, meta-learning can include learning to update certain parameters of the meta-learning model. The meta-learning model can be based on any similarity learning technique, such as initialization-based techniques, hallucination-based techniques, and metric learning-based techniques.
[0225] In some implementations, training the neural network of the AI system 802 to learn how to detect similar subject records based on mutation order can include creating a dataset of pairs of subject records. The pairs of subject records need not have the same mutation order. However, the mutation order between two subject records may differ slightly in some cases and may differ significantly in other cases. In some examples, pairs of slightly different subject records can be labeled as similar subject records, while pairs of subject records with significantly different mutation orders can be labeled as dissimilar subject records. The neural network can execute a learning algorithm to learn combinations and sequences of mutation orders that exist when two mutation orders are different but similar. Similarly, the neural network can execute a learning algorithm to learn combinations and sequences of mutation orders that exist when two mutation orders are different but dissimilar.
[0226] To illustrate by way of non-limiting example, a particular subject has breast cancer. The user device 110 can operate a cloud-based oncology application, causing the application to access a subject record 804 that characterizes the particular subject. For example, the particular subject has an ID# of 4123; a PTEN, TP53, BRCA1, and PIK3CA mutation sequence; and a cancer classification of stage I breast cancer. Subject record 806 has an ID# of 5316; a TP53, BCL2, and BRCA2 mutation sequence; and a cancer classification of stage II breast cancer. Subject record 808 has an ID# of 3142; a TP53, KRAS, and EGFR mutation sequence; and a cancer classification of stage IIIA lung cancer. Subject record 810 has an ID# of 2551; a TP53, BRCA1, KRAS, and PIK3CA mutation sequence; and a cancer classification of stage 0 colon cancer. Finally, subject record 812 has an ID# of 5456; a mutation sequence of PTEN, TP53, BCL10, GSTT1; and a cancer classification of stage IV hematological cancer. The mutation sequences for each of subject records 804-812 are summarized in Table 2 below. [Table 2]
[0227] A treating physician is evaluating potential treatments for a particular subject. The physician can operate the user device 110 to cause the user device 110 to generate a request (using a cloud-based oncology application) to identify subjects across different cancer types with similar gene mutation sequences. Querying or filtering the subject records may not identify all similar subject records due to slight differences in mutation sequence, such as intervening mutations in the mutation chain. The AI system 802 can output a prediction that subject record 804 and subject record 810 are similar in terms of mutation sequence. Both subject record 804 and subject record 810 share a mutation order sequence for TP53, BRCA1, and PIK3CA, but subject record 810 has an intervening mutation in KRAS.
[0228] The AI system 802 can send a response to the request received from the user device 110. The response can indicate that the (anonymized) subject record 810 closely (though not exactly) matches the mutation order of the subject record 804. Once similar subjects have been identified based on the mutation order (and potentially other factors), a physician can evaluate the treatments given to the similar subjects to determine the predicted effectiveness of those treatments for the particular subject.
[0229] Advantageously, AI system 802 can identify subject records similar to a given subject record, even if the similar subject records are associated with different cancer types. As illustrated in FIG. 8, the subject associated with subject record 810 was treated with alpelisib to target a PIK3CA gene mutation, and the treatment outcome was effective. Thus, a physician can select alpelisib to treat the subject associated with subject record 810 because the subject also has a PIK3CA mutation in a similar mutation order as subject record 804.
[0230] Furthermore, the cancer progression of a subject associated with subject record 810 may be useful in predicting the cancer progression of a subject associated with subject record 804, even if the subjects have different types of cancer. The fact that two subjects have similar mutation sequences indicates that the two subjects are likely to experience similar cancer progression, despite having different types of cancer.
[0231] By way of further illustration, and by way of non-limiting example only, the cloud-based oncology application can identify primary mutations, secondary mutations, tertiary mutations, etc. detected from a particular subject's genetic profile. The cloud-based oncology application can be configured to detect other breast cancer subjects with the same mutation sequence. If another breast cancer subject has the same mutation sequence, a physician can evaluate the breast cancer-specific treatment administered to the other subject. However, it may be possible that other subjects within the same cancer type do not have the same mutation sequence as the subject associated with the subject record 804. In this case, certain implementations of the present disclosure include continuing to search for subject records with similar mutation sequences but across different cancer types.
[0232] The cloud-based oncology application can also evaluate the clinical outcomes of a given targeted therapy treatment administered to other breast cancer patients with the same mutation sequence to predict the treatment outcome for a particular patient and the likely evolution of breast cancer mutations for that particular patient after the targeted therapy treatment is administered. If the oncology application cannot find other breast cancer patients with the same mutation sequence as a particular patient, the oncology application can look at patients with other cancer types, such as lung cancer. For example, the oncology application can identify a group of lung cancer patients with the same mutation sequence as a particular patient, or a group of lung cancer patients with at least the same secondary or tertiary mutations as a particular breast cancer patient. The oncology application can then evaluate the clinical outcomes of a given targeted therapy treatment administered to the identified group of lung cancer patients to predict the treatment outcome for the particular breast cancer patient.
[0233] VII. A network environment configured to predict specific side effects of oncological lines of treatment using artificial intelligence techniques FIG. 9 is a block diagram illustrating an example of a network environment for deploying an AI model trained to predict subject-specific side effects of oncological treatment, according to some embodiments of the present disclosure. The network environment 900 may include an AI system 902 and data stores 910-922 for storing various contextual information about a subject, such as a subject being treated at a medical facility. While FIG. 9 illustrates seven data stores (e.g., data stores 910-922), it will be understood that FIG. 9 is exemplary and that any number of data stores may be included in the network environment 900. The AI system 902 may be similar to the AI system 702 illustrated in FIG. 7. However, the components of the AI system 902 may differ from the components of the AI system 702. The components of the AI system 902 illustrated in FIG. 9 may be in addition to, instead of, or part of any of the components of the AI system 702 illustrated in FIG. 7.
[0234] In some implementations, the AI system 902 can be configured to automatically predict the specific side effects that a particular subject is likely to experience in response to receiving an oncological treatment, such as a targeted therapy. The AI system 902 can include a knowledge graph 904, an enriched subject record generator 906, and an enriched subject record data store 908.
[0235] In some implementations, the knowledge graph 904 can include a graphical representation of nodes and edges that map treatments to associated side effects, integrating the mapping into an ontology. For example, the knowledge graph 904 can be trained using a large set of triplet sentences. The first word or phrase of a given triplet is a treatment, such as alpelisib. The second word or phrase of a given triplet is a relationship between the treatment and the side effect, such as "less than 30% show this side effect." The third word or phrase of a given triplet is a side effect. As an illustrative example, a triplet could include [alpelisib, 10%-30% of subjects, low blood count]. Triplets can be created by individually connecting a treatment to each of its side effects. In some implementations, the knowledge graph 904 can be trained based on a treatment side effect ontology 922. The ontology can be a set of nodes connecting treatments to their side effects. An edge connecting two nodes represents the relationship between a treatment and a side effect (e.g., the proportion of subjects who experience the side effect or the characteristics of subjects who typically experience the side effect). The treatment side effect ontology 922 can be created using any medical journal or drug specification.
[0236] Additionally, knowledge graph 904 includes an inference engine trained to generate output based on the relationships between treatments and side effects captured in knowledge graph 904. In some implementations, the inference engine may be trained to output logical inferences based on knowledge graph 904 and input data (e.g., a proposed treatment to be administered to a subject). The inference engine infers which information to extract from knowledge graph 904 based on the interference generated by the inference module. The inference can be used to evaluate input, recommend actions, or update rules, for example, if the proposed treatment is a targeted treatment for alpelidib and if knowledge graph 904 includes a connection between a first node representing alpelidib and a second node representing a lung problem. In this example, if the subject has asthma, the inference engine can automatically render a logical inference that a particular subject is likely to experience lung problems.
[0237] The enriched subject record generator 906 can extract contextual information about a particular subject from the data stores 910-920. For example, the enriched subject record generator 906 can query each data store 910-920 using a unique subject identifier to obtain contextual information about the subject. The contextual information obtained for a given subject can be appended together into an enriched subject record and stored in the enriched subject record data store 908. For example, the enriched subject record for a given subject can include a subject-specific data set that is more robust than the original subject record (e.g., electronic health record). The genetic profile data store 910 can store various genetic profiles of a subject. The radiology images 912 can store, for example, various images captured by or in association with a hospital radiology department. The medical research data store 914 can include medical journals or publications that include data points related to symptoms associated with the subject. For example, if the original subject record includes a data element indicating that the subject was diagnosed with breast cancer, the enriched subject record generator 906 can retrieve information about breast cancer stage from the medical research data store 914 for inclusion in the enriched subject record associated with the subject. The clinical information data store 916 can store clinical information characterizing the patient, such as third-party lab work, emergency room visits, and measurements from the patient. The claims data 918 can include historical health insurance information about the insured, such as explanation of benefits, costs covered by the insurance carrier, and costs covered by the insured. Finally, the subject-provided input data store 920 stores data received directly from interactions with the subject. For example, a subject may maintain a journal of side effects after undergoing chemotherapy. Subject notes are stored in the subject-provided input 920.
[0238] VIII. The cloud-based application is configured to detect the reasons underlying treatment selection and automatically classify the detected reasons as guideline compliant or not. 10 is a block diagram illustrating an example of a network environment for deploying a reinforcement learner trained to select an action, according to some embodiments of the present disclosure. The network environment 1000 can include an AI system 1002. The AI system 1002 can be similar to the AI system 702 shown in FIG. 7. However, the components of the AI system 1002 can be different from the components of the AI system 702. The components of the AI system 1002 shown in FIG. 10 can be in addition to, instead of, or part of any of the components of the AI system 702 shown in FIG. 7.
[0239] There are several clinical practice guidelines in the field of oncology. The guidelines are defined by medical organizations such as NCCN and ASCO. For example, NCCN publishes guidelines for treating various cancer types. The reasons underlying treatment selection often depend heavily on the experience and expertise of the treating physician. Therefore, determining whether the reasons for selecting or proposing a treatment comply with oncology treatment guidelines is a difficult and manual process. Certain implementations of the present disclosure relate to an automated AI-based technique for verifying whether the reasons for predicting treatment for a particular subject with cancer comply with existing guidelines.
[0240] In some implementations, the AI system 1002 can be configured to include an AI model execution system 1004 and a treatment guideline validation system 1006. Further, for example, the AI system 1002 can be configured to generate predictive outputs, such as predicting the treatment outcome of a given targeted therapy (similar to FIGS. 8 and 11 ) and predicting the specific side effects that a particular subject is likely to experience in response to a given treatment (similar to FIGS. 9 and 12 ). The AI model execution system 1004 can be similar to the AI model execution system 710 in that the AI model execution system 1004 can execute any AI model stored in the AI model data store 724.
[0241] In some implementations, the AI model execution system 1004 can be configured to detect feature importance in each instance the AI model is executed and a prediction is generated. Feature importance refers to a category of algorithms that assign scores to input features of a predictive AI model. The score assigned to an input feature represents the importance or degree of contribution the input feature made to the output of the AI model. Using the score, the AI model execution system 1004 can also generate a second output (e.g., secondary to the predicted output, such as a prediction of treatment selection). The second output represents one or more input features that contributed to generating the predicted output. The input features that contributed to generating the output can represent the reason for the treatment being proposed or predicted for selection by the AI model.
[0242] As an illustrative example, a subject has a TP53 mutation and breast cancer. Inputting the subject's subject record 1008 into a predictive AI model predicts a treatment 1010 of "proposed targeted therapy = reintroduction of p53 using replication-deficient adenovirus (Ad-p53)." The predictive AI model predicted a treatment 1010 indicating the subject's proposed or predicted treatment, but the reason for the proposed treatment is unknown. Thus, according to certain implementations described herein, the AI model execution system 1004 can be configured to perform a feature importance technique to generate a second output representing one or more input features that serve as reasons for the proposed treatment. Continuing with the illustrative example, the feature importance technique is performed to detect that the Ad-p53 treatment was proposed because the particular subject had a TP53 mutation. The Ad-p53 treatment acts as a TP53 inhibitor that can improve the subject's progression-free survival. Non-limiting examples of feature importance techniques include linear regression feature importance, logistic regression feature importance, decision tree feature importance, random forest feature importance, XGBoost feature importance, permutation feature importance, feature selection with importance, and any other suitable feature importance technique.
[0243] In some implementations, the input of the subject record 1008 is also input into a treatment guideline validation system 1006. Additionally, a treatment 1010 indicating a proposed treatment of Ad-p53 to inhibit TP53 mutations or replace wild-type p53 protein can be input into the treatment guideline validation system 1006. Finally, features identified as contributing to the output of the predictive AI model are also input into the treatment guideline validation system 1006. The output of the treatment guideline validation system 1006 may be a classification of the reason the predicted treatment was selected into one of several categories, referred to as a compliance class. To illustrate, and by way of non-limiting example only, the compliance class may include "guideline compliant," "guideline non-compliant," or "recommended for creation of a new guideline for treatment." In the above example, the reason for proposing Ad-p53 treatment (e.g., detection of a TP53 mutation in the subject's genetic profile) can be input into the treatment guideline validation system 1006, which then outputs a guideline classification 1012 of "meets guideline."
[0244] In some implementations, the treatment guideline validation system 1006 can be a neural network classifier model trained to classify subject records, predicted treatments, and features that contributed to the predicted treatments, for example, as “guideline-compliant,” “guideline-non-compliant,” or “create a new guideline.” The training dataset can include a labeled dataset of data records. Each record can include one or more features of the subject, the disease the subject was diagnosed with, the treatment administered to the subject, and the features that led to the treating physician's decision to administer the treatment. Furthermore, each record can be labeled as “guideline-compliant,” “guideline-non-compliant,” or “create a new guideline.” A supervised machine learning algorithm can be run on the training dataset to learn correlations within the training data. In some implementations, the treatment guideline validation system 1012 can be an inference engine that generates inferences about whether the input “reasons” for selecting a cancer treatment logically reflect existing guidelines. Furthermore, in some examples, the “create a new guideline” compliance class is invoked to classify the proposed treatment selection based on the reasons for selecting the treatment, the treatment itself, and if the guideline results in an indeterminate output.
[0245] IX. Cloud-based applications can use artificial intelligence techniques to predict treatment outcomes for specific subjects. 11 is a flowchart illustrating an example of a process for predicting treatment outcomes and cancer progression for a subject diagnosed with cancer according to some embodiments of the present disclosure. Process 1100 can be performed by any of the components shown in FIGS. 1 and 7-10. For example, process 1100 can be performed by AI system 802. Furthermore, process 1100 can be performed to execute an AI model that generates an output predicting the treatment outcome of a particular treatment proposed to be performed on a particular subject.
[0246] Process 1100 begins at block 1105, where the AI system 802 accesses or retrieves a subject record corresponding to, for example, a particular subject (e.g., a subject being treated at a hospital). The subject record (e.g., an electronic medical record or electronic health record) can include any number of features (e.g., data elements including values such as vaccinations, medication history, age, demographics, etc.) collected from or on behalf of the subject. The subject record can include a set of features that characterize an aspect of the subject. For example, the subject record can include a feature indicating that the subject has been diagnosed with stage I breast cancer, among many other features.
[0247] In some examples, the genetic profile is associated with a subject record. For example, a subject associated with the subject record may have undergone genetic testing for various purposes, such as to confirm a disease diagnosis or to identify the effectiveness of a particular treatment. The genetic profile of a particular subject can provide the results of the genetic test. For example, the genetic profile of a particular subject may include information about particular genes (e.g., any detected gene mutations, levels of gene expression). The genetic profile may be useful for various purposes, such as diagnosing a disease, selecting a treatment to administer to a subject, or evaluating a proposed treatment, e.g., the side effects of a particular drug. In some implementations, the AI system 802 retrieves the genetic profile associated with the subject record accessed in block 1105. Further, the AI system 802 can extract the subject's mutation order from the genetic profile. The AI system 802 can also use the genetic profile or the subject record to identify the type of cancer the subject has been diagnosed with and the proposed or predicted treatment. For example, as shown in Figure 8, the mutation order shown in a subject's genetic profile may be [Mutation #1=PTEN], [Mutation #2=TP53], [Mutation #3=BRCA1], [Mutation #4=PIK3CA].
[0248] Non-limiting examples of features that may be included in a subject record include radiology imaging data, MRI data, genetic profile data, clinical data (e.g., measurements, treatment, treatment response, diagnosis, severity, medical history), subject-generated data (e.g., notes entered by a subject receiving chemotherapy), physician or healthcare professional-generated data (e.g., physician's notes), voice data representing telephone calls between a patient and a physician or other healthcare professional, administrative data, billing data, health surveys (e.g., HRS surveys), third-party or vendor information (e.g., out-of-network lab results), public databases related to the subject (e.g., medical journals related to the subject's symptoms), subject demographics, immunizations, radiology reports, pathology reports, utilization information, metadata representing biological samples, social data (e.g., education level, employment status), community specifications, etc.
[0249] In block 1110, the AI system 802 can identify groups of other subject records (e.g., other de-identified subject records associated with the medical facility). The AI system 802 can also filter the groups of subject records by the same cancer type (e.g., to form a smaller subgroup of only subject records associated with a breast cancer diagnosis). The subgroup of subject records can also be further filtered by proposed treatment (e.g., combination therapy treatment).
[0250] In block 1115, the AI system 802 may also perform a clustering operation on the vectorized subject records included in the subgroups based on the treatment outcomes of the proposed treatments. For example, the clustering operation may be any density-based, hierarchical, partitioning, or grid-based technique for clustering data points. The clustering operation may cluster the vectorized subject records of the subgroups by treatment outcome. Non-limiting examples of proposed or predicted treatments may include chemotherapy in general, specific chemotherapy drugs, radiation therapy, combination therapy, surgery, and other suitable treatments for treating cancer. Furthermore, non-limiting examples of treatment outcomes may be any outcome after administration of a treatment that causes a change in the subject's state (e.g., a change in psychological state, a change in physical state, a change in social state) that has a positive or negative impact on the subject's health. In some implementations, the treatment outcomes may be segmented into categories, thresholds, or ranges, such as a percentage range of increase or decrease in gene expression values after administration of a targeted therapy treatment. The clustering operation at block 1120 results in one or more clusters of subject records for the subjects in the subgroup. The subject records included in each cluster can be associated with the same or similar treatment and treatment outcome.
[0251] In block 1120, the AI system 802 can perform a mutation order similarity determination between a particular subject record and each of the other records in each cluster. For example, the AI system 802 can include a neural network trained to learn how to detect similar subject records based on mutation order. The training data can include a dataset of subject record pairs. The subject record pairs need not have the same mutation order. However, the mutation order between two subject records may differ slightly in some cases and may differ significantly in other cases. In some examples, a pair of slightly different subject records can be labeled as similar subject records, while a pair of subject records with significantly different mutation orders can be labeled as dissimilar subject records. The neural network can execute a learning algorithm to learn the combinations and sequences of mutation orders that exist when two mutation orders are different but similar. Similarly, the neural network can execute a learning algorithm to learn the combinations and sequences of mutation orders that exist when two mutation orders are different but dissimilar.
[0252] In block 1125, the AI system 802 may generate a similarity measure between the vector representation of the subject record characterizing the particular subject and the vector representation of each other subject record determined to be similar to the particular subject record in block 1120. Non-limiting examples of techniques for generating similarity measures include Euclidean distance, Manhattan distance, Minkowski distance, cosine similarity, Jaccard similarity, and other suitable techniques.
[0253] At decision block 1130, the AI system 802 can determine whether any of the similarities generated at block 1125 fall within a distance range associated with a cluster. For example, if the similarity between the vector representation of a particular subject's subject record and the vector representation of another subject record is within a threshold distance of the cluster, the similarity can fall within the range of that cluster. If the output of decision block 1130 is "Yes," process 1100 proceeds to block 1135, where the AI system 802 uses the treatment outcomes associated with the cluster (identified or selected at decision block 1130) to generate a prediction of a treatment outcome for the particular subject.
[0254] If the output of decision block 1130 is "NO," process 1100 proceeds to block 1140. In block 1140, AI system 802 can re-filter the group of other subject records by the same mutation order, rather than by cancer type. Thus, unlike the filtered subgroups formed in block 1120, the new filtered subgroups formed in block 1140 include subject records with the same mutation order as the particular subject, but with various cancer types that may differ from the cancer type associated with the particular subject. AI system 802 can also re-run the clustering operation on the new filtered subgroups for each treatment outcome. Finally, AI system 802 can regenerate the similarity between the vectorized subject record of the particular subject and each of the other subject records.
[0255] At decision block 1145, the AI system 802 can determine whether any of the similarities generated at block 1140 fall within a distance range (e.g., Euclidean distance) associated with a cluster. For example, if the similarity between the vector representations of a particular subject's subject record is within a threshold distance of the cluster, the similarity can fall within that cluster's range. If the output of decision block 1145 is "Yes," the process 1100 proceeds to block 1150, where the AI system 802 uses the treatment outcomes associated with the cluster (identified or selected at decision block 1145) to generate a prediction of the treatment outcome for the particular subject. If the output of decision block 1145 is "No," the process 1100 returns to block 1140 and re-filters the other subject records by different cancer types.
[0256] X. A cloud-based application can automatically predict the outcome of a mutation-targeted treatment for a particular subject. 12 is a flowchart illustrating an example of a process for predicting subject-specific treatment outcomes of mutation-targeted treatments according to some embodiments of the present disclosure. Process 1200 can be performed by any of the components shown in FIGS. 1 and 7-10. For example, process 1200 can be performed by AI system 902. Furthermore, process 1200 can be performed to run an AI model that generates an output predicting the survival advantage of a proposed treatment for a subject diagnosed with cancer.
[0257] Process 1200 begins at block 1210, where AI system 902 retrieves a subject record identifying a particular subject and characterizing the particular subject. For example, the subject record can be retrieved from a data registry, such as data registry 722. The subject record can be accessed automatically at regular or irregular time intervals, or in response to a user input that triggers a predictive function, as described in more detail herein. As an illustrative example, AI system 902 can identify a particular subject based on input received from a user device (e.g., user device 110). AI system 902 can detect a unique subject identifier (e.g., patient code) that uniquely identifies the particular subject from the input received from the user device. AI system 902 can then query the data registry using the unique subject identifier.
[0258] In block 1220, the AI system 902 (e.g., via the enriched subject record generator 906) can also query other databases for contextual information characterizing a particular subject. Non-limiting examples of other databases that the AI system 902 can query include a genetic profile data store 910, a radiology image data store 912, a medical research data store 914, a clinical data store 916, a billing data store 918, and a subject-provided input data store 920. In some examples, the AI system 902 can query the genetic profile data store 910 using a unique subject identifier for the results of genetic tests performed on a particular subject. Illustratively, a gene panel may have been sequenced for a particular subject, and the results of the gene sequencing may be stored in a genetic profile in the genetic profile data store 910. In some examples, the AI system 902 can query the billing data store 918 to obtain health insurance claims submitted by or on behalf of a particular subject.
[0259] At block 1230, the AI system 902 (e.g., via the enriched subject record generator 906) can generate an enriched subject record for the particular subject. The enriched subject record for the particular subject can include the original subject record (obtained at block 1210) characterizing the particular subject and contextual information (obtained at block 1220) characterizing the particular subject. For example, all or a portion of the contextual information for the particular subject can be appended to the original subject record obtained at block 1210. In some implementations, the enriched subject profile can include at least a portion of the particular subject's genetic profile. For example, the subject profile can include known genetic mutations detected from a gene panel performed on the particular subject. The particular subject's genetic profile is often stored separately or independently from the subject record characterizing the particular subject. Thus, as a technical advantage, the enriched subject record generator 906 can store or append at least a portion of the particular subject's genetic profile to the subject record. The enriched subject record can then be processed using the AI system 902 to perform certain predictive functions.
[0260] In block 1240, the AI system 902 can convert the enriched subject record into a query of a knowledge model (e.g., knowledge graph 904). In some implementations, converting the enriched subject record into a query can include converting each data element of the enriched subject record into a numerical representation (e.g., a vector) and then combining (e.g., using addition, averaging, or concatenation) the numerical representations of each data element into a single numerical representation that represents the entire enriched subject record. In some implementations, converting the enriched subject record into a query can include generating an array of vectors, such that each element represents a value of a data element in the enriched subject record. In some implementations, converting the enriched subject model into a query can include extracting values from the enriched subject model and forming an input graph of the extracted values. The input graph can serve as an input to the knowledge model. For example, the AI system 902 can extract mutations detected from the genetic profile and proposed treatments included in the enriched subject record. The AI system 902 can convert the extracted mutations and proposed treatments into an input graph, where the detected mutations are nodes connected to another node representing the subject's disease or health state, which is then connected to yet another node representing the proposed treatment. The input graph can be used to query the knowledge model to predict the inherent survival advantage of the proposed treatment for a particular subject.
[0261] Further, at block 1240, the input graph may or may not include a proposed treatment for treating the subject. If the input graph includes a unique proposed treatment for a particular subject, process 1200 may proceed to block 1250. At block 1250, the AI system 902 may query the knowledge model using an input graph that includes a node representing the unique proposed treatment for the subject. In response to the query, the knowledge model may generate an output that specifically represents the contextual survival advantage of the proposed treatment for the particular subject. However, at block 1240, the knowledge model may also receive as input an input graph without a node representing the proposed treatment. In this situation, process 1200 proceeds to block 1270, where the knowledge model is queried using the input graph (e.g., not including the proposed treatment). For example, at block 1270, the knowledge model may be queried to identify available candidate treatments, taking into account the contextual information included in the enriched subject record. Furthermore, the knowledge model may also store several potential survival advantages for each candidate treatment. Then, in block 1280, the knowledge model can also output a subject-specific survival advantage for each candidate treatment.
[0262] XI. A cloud-based application can automatically predict the subject characteristics that contributed to treatment prediction and determine whether the predicted subject characteristics comply with hospital guidelines. 13 is a flowchart illustrating an example process for developing an AI model to identify factors (e.g., subject-related features) that contributed to a given treatment prediction output by an AI system, according to some aspects of the present disclosure. Process 1300 can be performed by any of the components illustrated in FIGS. 1 and 7-10. For example, process 1300 can be performed by AI system 1002. Furthermore, process 1300 can be performed to automatically verify whether the subject features that contributed to the treatment prediction by the AI system comply with existing guidelines (e.g., guidelines established by a medical institution).
[0263] Process 1300 begins at block 1310, where AI system 1002 accesses or retrieves a subject record stored in a data registry, such as data registry 722. The subject record may characterize a particular subject diagnosed with cancer, such as breast cancer. At block 1320, the subject record accessed or retrieved at block 1310 may be converted to a numerical representation (e.g., a vector representation) using various implementations described herein (e.g., as described with reference to FIGS. 1-6). The subject record may be converted or vectorized to a numerical representation in advance or in real time or substantially real time by execution of block 1310.
[0264] At block 1330, the numerical representation can be input to a trained AI model for processing, e.g., using the AI model execution system 710. While block 1330 can be performed using any AI model, such as the AI model described with respect to FIG. 7 for illustrative purposes, the trained AI model can output a prediction of a treatment to administer to the subject. It will be understood that the trained AI model performed at block 1330 can also be any of the AI models described with respect to FIGS. 12 and 13. Regardless of which AI model is performed at block 1330, the AI model can be trained to generate two outputs. For example, at block 1340, the AI model outputs a prediction of a treatment to administer to a particular subject, and at block 1350, the AI model also outputs features (e.g., data elements of a particular subject record that drove or contributed to the prediction of the selected treatment). As an illustrative example, a subject has stage I breast cancer. The subject's genetic profile indicates that the subject has a PIK3CA mutation in addition to PTEN, TP53, and BRCA1. PIK3CA mutations can lead to overactivation of PI3Kα, a key upstream component of the PI3K pathway. The trained AI model learned from the training data that there is a high correlation between subjects with breast cancer who have PIK3CA mutations and those who are treated with alpelisib. Alpelisib treatment inhibits both the PI3K and ER pathways. Therefore, if the AI model detects that a particular subject has a PIK3CA mutation and has been diagnosed with breast cancer, the AI model generates an output selecting alpelisib as the optimal treatment for the particular subject. The trained AI model also detects that the characteristics of the PIK3CA mutation and the characteristics of the breast cancer diagnosis contributed to the prediction of alpelisib as the optimal treatment for the particular subject.
[0265] In block 1360, a treatment guideline validation system can receive as input the treatment prediction (generated in block 1340) and the features predicted to have contributed to the treatment prediction (generated in block 1350). In some implementations, the treatment guideline validation system can be a neural network classifier model trained to classify the subject record, the predicted treatment, and the features that contributed to the predicted treatment as, for example, "guideline-compliant," "non-guideline-compliant," or "create a new guideline." The training dataset can include a labeled dataset of data records. Each record can include one or more features of the subject, the disease for which the subject was diagnosed, the treatment administered to the subject, and the features that led to the treating physician's decision to administer the treatment. Further, each record can be labeled as "guideline-compliant," "non-guideline-compliant," or "create a new guideline." A supervised machine learning algorithm can be run on the training dataset to learn correlations within the training data. Once trained, in block 1370, the treatment guideline validation system can classify the proposed treatment and the reasons for selecting the proposed treatment as "guideline compliant" (block 1372), "non-guideline compliant treatment" (block 1374), or "create a new guideline for treatment" (block 1376).
[0266] XII. Further Considerations Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium containing instructions configured to cause one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein.
[0267] The terms and expressions which have been employed are used as terms of description rather than of limitation, and there is no intention in the use of such terms and expressions to exclude any equivalents of the features shown and described, or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, although the claimed invention has been specifically disclosed by embodiments and optional features, it will be understood that modifications and variations of the concepts disclosed herein may be resorted to by those skilled in the art, and that such modifications and variations are deemed to be within the scope of the invention as defined by the appended claims.
[0268] The following description provides only preferred exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description of preferred exemplary embodiments will provide those skilled in the art with an enabling description for implementing various embodiments. It will be understood that various changes can be made in the function and arrangement of elements without departing from the spirit and scope of the appended claims.
[0269] Specific details are given in the following description to provide a thorough understanding of the embodiments. However, it will be understood that the embodiments may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form so as not to obscure the embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail so as to avoid obscuring the embodiments.
[0270] XIII. Further Examples As used below, any reference to a series of examples should be understood disjunctively as a reference to each of those examples (e.g., "Examples 1-4" should be understood as "Examples 1, 2, 3, or 4").
[0271] Example 1 is a computer-implemented method for predicting a subject-specific outcome of an oncological line of treatment, comprising: identifying a particular subject diagnosed with a type of cancer, wherein the line of treatment is proposed to be administered to the particular subject; obtaining a genetic dataset corresponding to the particular subject, wherein the genetic dataset includes a mutation order, the mutation order including a series of multiple genetic mutations that were mutated at different times; identifying a set of other subjects diagnosed with the same type of cancer as the subject, and each other subject who has received the line of treatment and is associated with a treatment outcome; obtaining another genetic dataset for each other subject in the set of other subjects, wherein the other genetic dataset includes a different mutation order; and for each other subject in the set of other subjects, inputting the mutation order of the particular subject and the other mutation orders of the other subjects into a trained similarity model. a trained similarity model to generate similarity weights representing a predicted degree to which a mutation sequence of the particular subject is similar to another mutation sequence of another subject; and determining a predicted treatment outcome for administering a line of therapy to the particular subject based on the similarity weights output by the trained similarity model; if it is determined that at least one of the similarity weights output by the similarity model is within a threshold, identifying one of the other subjects based on the determination and assigning the treatment outcome of the identified other subject as the predicted treatment outcome for the particular subject; and / or if it is determined that none of the similarity weights output by the similarity model are within the threshold, identifying another set of subjects diagnosed with a different type of cancer from the particular subject to search for mutation sequences similar to the mutation sequence of the particular subject.
[0272] Example 2 is a computer-implemented method for predicting a subject-specific outcome of an oncological treatment line described in Example 1, further including: obtaining further mutation orders for each other subject of another set of other subjects, wherein each other subject of the other set has a different type of cancer than the particular subject; obtaining further mutation orders for each other subject of the other set of other subjects; inputting the mutation order of the particular subject and the other mutation orders of the other subjects of the other set into a trained similarity model, and determining, based on the similarity weights output by the trained similarity model, that at least one of the similarity weights output by the similarity model is within a threshold; and, based on the determination, identifying one of the other subjects of the other set and assigning the treatment outcome of the identified other subject of the other set as a predicted treatment outcome for the particular subject.
[0273] Example 3 is a computer-implemented method for predicting a subject-specific outcome of an oncology line of treatment described in Example 1 or 2, further comprising performing a clustering operation on a set of other subject records, the clustering operation being based on one or more outcomes of the line of treatment and forming one or more clusters.
[0274] Example 4 is a computer-implemented method for predicting subject-specific outcomes for the oncological treatment lines described in Examples 1-3, in which a similarity model is trained using a training dataset, the training dataset including pairs of mutation sequences labeled as similar or dissimilar.
[0275] Example 5 is a computer-implemented method for predicting subject-specific outcomes for the oncology treatment lines described in Examples 1-4, wherein the predicted treatment outcomes include one or more subject-specific side effects or progression-free survival that are specific to a particular subject characteristic.
[0276] Example 6 is a computer-implemented method for predicting a subject-specific outcome of an oncology treatment line described in Examples 1-5, wherein the contextual information associated with a particular subject includes a genetic profile associated with the subject.
[0277] Example 7 is a computer-implemented method for predicting a subject-specific outcome for an oncology treatment line described in Examples 1-6, further comprising generating contextual information associated with the particular subject by querying a genetic profile data store for a genetic profile associated with the particular subject, querying a radiology image data store for one or more radiology images associated with the particular subject, querying a medical research data store for content data regarding at least one feature attributable to the particular subject, querying a clinical information data store for clinical information associated with the particular subject, querying a claims data store for one or more health insurance claims submitted by or on behalf of the particular subject, and / or querying a subject-provided input data store for subject data provided by the particular subject, wherein the subject data is in one or more data formats.
[0278] Example 8 is a computer-implemented method for predicting a subject-specific outcome for an oncology treatment line described in Examples 1-7, wherein the treatment outcome includes one or more subject-specific side effects that are output at the subject's computing device using a chatbot.
[0279] Example 9 is a computer-implemented method for predicting a subject-specific outcome of an oncology treatment line described in Examples 1-8, wherein the subject record includes data identified in an electronic medical record corresponding to the subject.
[0280] Example 10 is a computer-implemented method for predicting a subject-specific outcome of an oncological treatment line described in Examples 1-9, wherein the type of cancer diagnosed in the subject includes at least one of breast cancer, lung cancer, colon cancer, or hematological cancer.
[0281] Example 11 is a computer-implemented method for predicting subject-specific outcomes for the oncology treatment lines described in Examples 1-10, wherein the knowledge graph is accessible using a cloud-based oncology application configured to provide predictive functionality for clinical decision-making.
[0282] Example 12 is a computer-implemented method for predicting a subject-specific outcome for an oncology treatment line described in Examples 1-11, further comprising: detecting a data leak associated with the inference module, where the data leak exposes a feature of the set of features included in the subject record or exposes an item of contextual information associated with the subject; and, in response to detecting a data leak associated with the inference module, executing a data leak prevention protocol to prevent or block disclosure of the feature of the set of features included in the subject record.
[0283] Example 13 is a computer-implemented method for predicting a subject-specific outcome for an oncology treatment line described in Examples 1-12, further comprising using a feature selection model to generate a dimension-reduced subject record characterizing the subject, wherein the dimension-reduced subject record removes one or more features from a set of features included in the subject record, generating a dimension-reduced subject record in which the one or more features are characterized as noise.
[0284] Example 14 is a system comprising one or more processors and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more processors, cause the one or more processors to perform some or all of one or more computer-implemented methods disclosed herein.
[0285] Example 15 is a computer program product tangibly embodied in a non-transitory machine-readable storage medium that includes instructions configured to cause one or more data processors to perform some or all of one or more computer-implemented methods disclosed herein.
[0286] Example 16 is a computer-implemented method for predicting subject-specific side effects for an oncological line of therapy, the computer-implemented method including: accessing a knowledge graph representing an ontology for mapping side effects to lines of therapy for treating cancer; obtaining a subject record associated with the subject, the subject record including a set of features characterizing the subject, the subject having been diagnosed with a type of cancer, and the subject record including a candidate line of therapy for the subject; querying one or more data stores for contextual information uniquely characterizing the subject; generating an enriched subject record by appending the contextual information to the subject record; converting the enriched subject record into input data for the knowledge graph; inputting the input data into the knowledge graph; and generating one or more subject-specific side effect predictions for the candidate line of therapy based on an output of the knowledge graph, wherein the one or more subject-specific side effects are identified based on the mapping of side effects to lines of therapy.
[0287] Example 17 is a computer-implemented method for predicting subject-specific side effects of an oncology treatment line described in Example 16, wherein a knowledge graph is defined based on a set of triplet sentences, each triplet sentence of the set of triplet sentences includes three data elements, the three data elements include a treatment line for treating the cancer, a side effect of the treatment line, and a relationship between the treatment line and the side effect, and the mapping of the side effect to the treatment line is based on the set of triplet sentences.
[0288] Example 18 is a computer-implemented method for predicting subject-specific side effects of an oncological treatment line described in Example 16 or 17, wherein the knowledge graph further includes an inference module configured to generate logical inferences based on mapping of side effects to candidate lines of treatment included in the input data and lines of treatment defined by the knowledge graph.
[0289] Example 19 is a computer-implemented method for predicting subject-specific side effects of an oncology treatment line described in Examples 16-18, wherein the logical inferences generated by the inference module identify an incomplete subset of side effects from the set of side effects included in the knowledge graph, and identify the incomplete subset of side effects that correspond to one or more subject-specific side effects predicted to occur after the candidate line of treatment is administered to the subject.
[0290] Example 20 is a computer-implemented method for predicting subject-specific side effects of an oncology treatment line described in Examples 16-19, wherein the set of triplet sentences defining the knowledge graph is based on medical research and / or the one or more subject-specific side effects include progression-free survival specific to a subject characteristic.
[0291] Example 21 is a computer-implemented method for predicting subject-specific side effects of the oncology treatment lines described in Examples 16-20, wherein the contextual information includes a genetic profile associated with the subject.
[0292] Example 22 is a computer-implemented method for predicting subject-specific side effects of an oncology treatment line described in Examples 16-21, wherein querying one or more data stores further includes querying a genetic profile data store for a genetic profile associated with the subject, querying a radiological image data store for one or more radiological images associated with the subject, querying a medical research data store for content data regarding at least one feature attributable to the patient, querying a clinical information data store for clinical information associated with the subject, querying a claims data store for one or more health insurance claims submitted by or on behalf of the subject, and / or querying a subject-provided input data store for subject data provided by the subject, wherein the subject data is in one or more data formats.
[0293] Example 23 is a computer-implemented method for predicting subject-specific side effects of an oncology treatment line described in Examples 16-22, wherein one or more subject-specific side effects are output at the subject's computing device using a chatbot.
[0294] Example 24 is a computer-implemented method for predicting subject-specific side effects of an oncology treatment line described in Examples 16-23, wherein the subject record includes data identified in an electronic medical record corresponding to the subject.
[0295] Example 25 is a computer-implemented method for predicting subject-specific side effects of the oncology treatment line described in Examples 16-24, wherein the type of cancer diagnosed in the subject includes at least one of breast cancer, lung cancer, colon cancer, or hematological cancer.
[0296] Example 26 is a computer-implemented method for predicting subject-specific side effects of an oncology treatment line described in Examples 16-25, wherein the knowledge graph is accessible using a cloud-based oncology application configured to provide predictive functionality for clinical decision-making.
[0297] Example 27 is a computer-implemented method for predicting a subject-specific side effect of an oncology treatment line described in Examples 16 to 26, further comprising: detecting a data leak associated with the inference module, wherein the data leak exposes a feature of the set of features included in the subject record or exposes an item of contextual information associated with the subject; and, in response to detecting a data leak associated with the inference module, executing a data leak prevention protocol to prevent or block disclosure of the feature of the set of features included in the subject record.
[0298] Example 28 is a computer-implemented method for predicting subject-specific side effects of an oncological treatment line described in Examples 16-27, further comprising using a feature selection model to generate a dimension-reduced subject record characterizing the subject, wherein the dimension-reduced subject record removes one or more features from a set of features included in the subject record to generate a dimension-reduced subject record in which the one or more features are characterized as noise.
[0299] Example 29 is a system comprising one or more processors and a non-transitory computer-readable storage medium comprising instructions that, when executed on the one or more processors, cause the one or more processors to perform some or all of one or more computer-implemented methods disclosed herein.
[0300] Example 30 is a computer program product tangibly embodied in a non-transitory machine-readable storage medium that includes instructions configured to cause one or more data processors to perform some or all of one or more computer-implemented methods disclosed herein.
Claims
1. 1. A computer-implemented method for predicting subject-specific outcomes of an oncological treatment line, comprising: obtaining a genetic dataset corresponding to a particular subject identified using the input subject identifier, the genetic dataset including a mutation profile indicative of one or more molecular characteristics of the particular subject; extracting a cancer type and line of treatment associated with a particular subject from a subject record that associates a subject identifier, a type of cancer the subject was diagnosed with, a line of treatment for the subject, and a treatment outcome, and extracting a first set of other subjects from the subject record that are associated with the extracted cancer type and line of treatment; obtaining another genetic dataset for each other subject of the first set, the another genetic dataset including another first mutation profile; For each other subject in the first set, inputting the mutation profile of the particular subject and the first mutation profile of the other subjects in the first set into a trained similarity model, the trained similarity model being trained to generate a similarity weight representing a predicted degree to which the mutation profile of the particular subject is similar to the first mutation profile of the other subjects in the first set; determining a predicted treatment outcome for administering the line of therapy to the particular subject based on the similarity weights output by the trained similarity model; and Including, determining a predicted treatment outcome for administering the line of treatment to the particular subject includes, when at least one of the similarity weights output by the similarity model is determined to be within a threshold, assigning as the predicted treatment outcome for the particular subject a treatment outcome associated in the subject record with another subject of the first set whose similarity weight is determined to be within the threshold generated by the trained similarity model; and / or When it is determined that none of the similarity weights output by the similarity model are within the threshold, a computer-implemented method includes extracting from the subject record a second set of other subjects, different from the first set, who are associated in the subject record with a different type of cancer than the particular subject, in order to search for mutation profiles similar to the mutation profile of the particular subject.
2. obtaining a second mutation profile for each other subject of the second set, the second mutation profile being different from the first mutation profile; for each other subject of the second set, inputting the mutation profile of the particular subject and the second mutation profiles of the other subjects of the second set into the trained similarity model; determining, based on the similarity weights output by the trained similarity model, that at least one of the similarity weights output by the similarity model is within the threshold; assigning as the predicted treatment outcome for the particular subject a treatment outcome associated in the subject record with other subjects of the second set whose similarity weights are determined to be within the threshold generated by the trained similarity model; and / or 2. The computer-implemented method for predicting a subject-specific outcome of an oncology treatment line of claim 1, wherein the mutation profile comprises a mutation order associated with the particular subject, the mutation order representing a series of multiple genetic mutations mutated at different time points.
3. 3. The computer-implemented method for predicting a subject-specific outcome of an oncology line of treatment of claim 1 or 2, further comprising: performing a clustering operation on a set of other subject records, the clustering operation being based on one or more outcomes of the line of treatment and forming one or more clusters.
4. 4. The computer-implemented method for predicting a subject-specific outcome of an oncological treatment line of any one of claims 1 to 3, wherein the similarity model is trained using a training dataset, the training dataset comprising pairs of mutation profiles labeled as similar or dissimilar.
5. 5. The computer-implemented method for predicting subject-specific outcomes of an oncology treatment line of any one of claims 1 to 4, wherein the predicted treatment outcomes include one or more subject-specific side effects or progression-free survival specific to the particular subject characteristics.
6. 4. The computer-implemented method for predicting a subject-specific outcome of an oncology treatment line of claim 3, wherein the contextual information associated with the particular subject comprises a genetic profile associated with the subject.
7. querying a genetic profile data store for the genetic profile associated with the particular subject; querying a radiological image data store for one or more radiological images associated with the particular subject; querying a medical research data store for content data relating to at least one characteristic attributable to the particular subject; querying a clinical information data store for clinical information associated with the particular subject; Querying a claims data store for one or more health insurance claims submitted by or on behalf of said particular subject; and / or querying a subject-provided input data store for subject data provided by the particular subject, the subject data being in one or more data formats; 7. The computer-implemented method for predicting a subject-specific outcome of an oncology line of treatment of claim 6, further comprising generating the contextual information associated with the particular subject by:
8. 8. The computer-implemented method for predicting a subject-specific outcome of an oncology line of treatment according to any one of claims 1 to 7, wherein the treatment outcome comprises one or more subject-specific side effects output at the subject's computing device using a chatbot.
9. 9. A computer-implemented method for predicting subject-specific outcomes of an oncology line of treatment according to any one of claims 1 to 8, wherein the subject record comprises data identified in an electronic medical record corresponding to the subject.
10. 10. A computer-implemented method for predicting a subject-specific outcome of an oncology treatment line according to any one of claims 1 to 9, wherein the type of cancer the subject is diagnosed with comprises at least one or more of breast cancer, lung cancer, colon cancer, or hematological cancer.
11. 11. A computer-implemented method for predicting subject-specific outcomes of an oncology line of treatment according to any one of claims 1 to 10, wherein the knowledge graph is accessible using a cloud-based oncology application configured to provide predictive capabilities for clinical decision making.
12. Detecting a data breach associated with an inference module, the data breach exposing a feature of a set of features included in the subject record or exposing an item of the contextual information associated with the subject; In response to detecting a data leak associated with the inference module, executing a data leak prevention protocol to prevent or block the feature of the set of features included in the subject record from being exposed.
10. The computer-implemented method for predicting subject-specific outcomes of an oncological treatment line of claim 6, further comprising:
13. 13. The computer-implemented method for predicting a subject-specific outcome of an oncology line of treatment of claim 12, further comprising: using a feature selection model to generate a dimensionality-reduced subject record characterizing the subject, wherein the dimensionality-reduced subject record removes one or more features from the set of features included in the subject record, generating a dimensionality-reduced subject record in which the one or more features are characterized as noise.
14. 1. A system comprising: one or more processors; a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more processors, cause the one or more processors to perform part or all of the computer-implemented method of any one of claims 1 to 13.
15. A computer program comprising instructions arranged to cause one or more data processors to carry out part or all of the computer-implemented method of any one of claims 1 to 13.
Citation Information
Patent Citations
Systems and methods for predicting efficacy of cancer treatments
JP2021503149A
Systems and methods for predicting the efficacy of cancer therapy
WO2019095017A1