Individual and Cohort Pharmacological Phenotype Prediction Platform
Through a machine learning-based system, combining patient omics, sociological and environmental data to predict pharmacological phenotypes, the problem of existing systems failing to effectively predict pharmacological phenotypes is solved, and more accurate drug selection and higher therapeutic effects are achieved.
Patent Information
- Application Number
- CN201880046200.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-02-21
- Filing Date
- 2018-05-11
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2038-05-11
AI Technical Summary
Existing systems have failed to effectively use chromatin status, genomic regulatory elements, epigenomics, proteomics, metabolomics, or transcriptomics to predict the pharmacological phenotype of patients and have not taken into account the effects of environmental and sociological characteristics.
Develop a system based on machine learning technology to generate statistical models to predict pharmacological phenotypes by analyzing patients’ omics, sociological, and environmental data. The system can acquire social and environmental data from patients at multiple time points, combining the omics data, and train models to predict drug responses, disease risk, and other pharmacological phenotypes.
Accurate prediction of patients' pharmacological phenotypes is achieved, helping healthcare providers choose the best medication, reduce the possibility of adverse drug reactions, and improve treatment effectiveness.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit and priority of (1) U.S. Provisional Application Serial No. 62 / 505,422, filed May 12, 2017, entitled "Individual and Cohort Pharmacological Phenotype Prediction Platform", and (2) U.S. Provisional Application Serial No. 62 / 633,355, filed February 21, 2018, entitled "Individual and Cohort Pharmacological Phenotype Prediction Platform", the entire disclosures of each of which are hereby expressly incorporated herein by reference. Field of the Invention
[0003] This application relates to pharmacological patient phenotypes, and more particularly, to a method and system for predicting drug response phenotypes of patients and stratified cohorts of patients based on the biological, ancestry, demographic, clinical, sociological, and environmental characteristics of the patients using machine learning and statistical techniques. Background of the Invention
[0004] Today, it is possible to predict the drug responses of some patients based on their coding genomes. Specific genetic traits can be mapped to specific responses to drugs, and drugs can be selected for patients based on their predicted responses.
[0005] However, non-coding genomic variants account for the vast majority of genetic trait differences, such as patient drug responses, adverse drug reactions, and disease risks. The integration of epigenomic regulation studies with genome-wide association studies (GWAS) has also shown that epigenomic alterations can indicate disease risks, drug responses, and adverse drug reactions in humans and animals in a wide range of medical specialties and drug research settings. In addition, phenotype variations associated with diseases may be determined by differences in chromatin states previously attributed to genetic differences.
[0006] Current systems do not utilize chromatin states, genomics regulatory elements, epigenomics, proteomics, metabolomics, or transcriptomics to predict patient pharmacological phenotypes. Current systems also do not consider environmental and sociological characteristics that may alter genetic traits to determine pharmacological phenotypes. Additionally, such systems do not utilize machine learning techniques to train the system to adapt to changes over time in biological characteristics and / or pharmacological phenotypes corresponding to biological characteristics.
[0007] Accordingly, there is a need for a system that can accurately predict pharmacological phenotypes (including pharmacological responses, disease risks, drug abuse, or other pharmacological phenotypes) based on omics features (including genomics, epigenomics, chromatin state, proteomics, metabolomics, transcriptomics, etc.) and near real-time sociological and environmental characteristics of patients. Summary of the Invention
[0008] To predict the pharmacological phenotypes of patients, various machine learning techniques can be used to train a pharmacological phenotype prediction system. More specifically, the pharmacological phenotype prediction system can be trained to analyze the omics, sociological, and environmental data of patients to predict the responses of the patients to various drugs, the likelihood of drug abuse by the patients, the risks of various diseases, or any other pharmacological phenotypes of the patients. The pharmacological phenotype prediction system can be trained by obtaining the omics, sociological, and environmental data (also referred to herein as "training data") of a group of patients (also referred to herein as "training patients").
[0009] In some embodiments, the sociological and environmental data of patients can be obtained at multiple time points to detail the experiences of the patients. For each training patient, the pharmacological phenotype prediction system can obtain the pharmacological phenotypes of the patient as training data, such as whether the patient has a drug abuse problem, the chronic diseases of the patient, the responses of the patient to various drugs prescribed to the patient, etc. The various machine learning techniques can be used to analyze the training data to generate a statistical model, which can be used to predict the responses of the patient to various drugs, the likelihood of drug abuse by the patient, the risks of various diseases, or any other pharmacological phenotypes of the patient. For example, the statistical model can be a neural network generated based on a combination of network analysis of gene regulatory networks and environmental effects on gene expression.
[0010] After the training period, the pharmacological phenotype prediction system can receive the omics, sociological, and environmental data of patients whose pharmacological phenotypes are unknown at several time points (e.g., lithium has not been prescribed to a bipolar disorder patient yet, so it is not yet clear how the patient responds to lithium). The omics, sociological, and environmental data can be applied to the statistical model to predict the pharmacological phenotypes of the patients, and these phenotypes can be displayed on the client device of a healthcare provider.
[0011] For example, for a specific drug, the pharmacogenetic phenotype prediction system can determine the likelihood that the patient will experience an adverse drug reaction. Additionally, the pharmacogenetic phenotype prediction system can generate an indication of the predicted efficacy or appropriate dosage of the patient's drug. In some embodiments, the likelihood that the patient will experience an adverse drug reaction can be compared to a threshold likelihood, and the predicted efficacy can be compared to a threshold efficacy. When the likelihood exceeds the threshold likelihood, the predicted efficacy is less than the threshold efficacy, and / or when the combination of the likelihood of an adverse drug reaction and the predicted efficacy exceeds a threshold, an indication of the likelihood and / or the efficacy of the drug can be provided to the healthcare provider. Accordingly, the healthcare provider can change the dosage, not prescribe the drug to the patient, or recommend an alternative drug with higher efficacy to the patient.
[0012] In this way, the pharmacogenetic phenotype prediction system can identify the best drug for a patient suffering from a specific disease. For example, for a specific disease, the pharmacogenetic phenotype prediction system can select one drug out of several drugs designed to treat the disease that has the greatest predicted efficacy for the patient and the lowest likelihood and / or severity of an adverse drug reaction. This embodiment advantageously allows the healthcare provider to accurately and effectively identify the best drug to recommend and prescribe to the patient. Additionally, by incorporating omics, sociology, and environmental data to generate the statistical model, this embodiment advantageously includes a comprehensive bioinformatics analysis of the patient's biological characteristics, which may change over time. This comprehensive bioinformatics analysis provides a more accurate prediction system that can not only predict pharmacogenetic phenotypes based on the patient's inherent characteristics but also incorporate sociological and environmental traits that change over time and may alter the expression of genetic traits.
[0013] Furthermore, by generating a statistical model that accurately predicts the risk of disease and the likelihood of an adverse drug reaction, the healthcare provider can proactively address these issues before the patient exhibits disease symptoms or begins to develop a drug abuse problem or other disease symptoms.
[0014] In one embodiment, a computer-implemented method is provided for identifying pharmacologic phenotypes using statistical modeling and machine learning techniques. The method includes obtaining a set of training data that includes, for each of a plurality of first patients, the following data: omics data that indicates a biological characteristic of the first patient, social omics and environmental data that indicates the experiences of the first patient collected over time, and phenomics data that indicates at least one of the following: a response to one or more drugs, whether the first patient has experienced an adverse drug reaction or drug abuse, or one or more chronic diseases of the first patient. The method further includes: generating, based on the set of training data, a statistical model for determining a pharmacologic phenotype; receiving a set of omics data and social omics and environmental data for a second patient collected over a period of time; applying the omics data and the social omics and environmental data for the second patient to the statistical model to determine one or more pharmacologic phenotypes for the second patient; and providing the one or more pharmacologic phenotypes for the second patient for presentation to a healthcare provider, where the healthcare provider recommends a treatment plan for the second patient based on the one or more pharmacologic phenotypes.
[0015] In another embodiment, a computing device is provided for identifying pharmacologic phenotypes using statistical modeling and machine learning techniques. The computing device includes a communication network, one or more processors, and a non-transitory computer-readable memory coupled to the one or more processors and storing instructions thereon. The instructions, when executed by the one or more processors, cause the system to obtain a set of training data that includes, for each of a plurality of first patients, the following data: omics data that indicates a biological characteristic of the first patient, social omics and environmental data that indicates the experiences of the first patient collected over time, and phenomics data that indicates at least one of the following: a response to one or more drugs, whether the first patient has experienced an adverse drug reaction or drug abuse, or one or more chronic diseases of the first patient. The instructions further cause the system to: generate, based on the set of training data, a statistical model for determining a pharmacologic phenotype; receive a set of omics data and social omics and environmental data for a second patient collected over a period of time; apply the omics data and the social omics and environmental data for the second patient to the statistical model to determine one or more pharmacologic phenotypes for the second patient; and provide, via the communication network, the one or more pharmacologic phenotypes for the second patient for presentation to a healthcare provider, where the healthcare provider recommends a treatment plan for the second patient based on the pharmacologic phenotype. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1AA block diagram of a computer network and system on which an exemplary pharmacologic phenotype prediction system according to the presently described embodiments may operate is shown;
[0017] Figure 1B is a block diagram of an exemplary pharmacologic phenotype assessment server that may operate in a Figure 1A system according to the presently described embodiments;
[0018] Figure 1C is a block diagram of an exemplary client device that may operate in a Figure 1A system according to the presently described embodiments;
[0019] Figure 2 depicts example omics, sociology, and environmental data that may be provided to a pharmacologic phenotype prediction system according to the presently described embodiments;
[0020] Figure 3 depicts a detailed view of a process performed by a pharmacologic phenotype prediction system according to the presently described embodiments;
[0021] Figure 4A depicts an exemplary representation of a bioinformatics analysis of permissive candidate variants associated with a particular pharmacologic phenotype according to the presently described embodiments, and a schematic diagram representing an exemplary transcriptional spatial hierarchy in the human genome;
[0022] Figure 4B is a block diagram representing an exemplary method for identifying omics data corresponding to a particular pharmacologic phenotype using machine learning techniques according to the presently described embodiments;
[0023] Figure 4C depicts an exemplary gene regulatory network for a patient according to the presently described embodiments;
[0024] Figure 4D is a block diagram representing another exemplary method for identifying omics data corresponding to a particular pharmacologic phenotype using machine learning techniques according to the presently described embodiments;
[0025] Figure 4E is a block diagram of single nucleotide polymorphisms (SNPs) identified at each stage of a method described in Figure 4D when identifying omics data corresponding to a warfarin phenotype;
[0026] Figure 4F depicts an exemplary warfarin response pathway according to embodiments described in the present invention;
[0027] Figure 4G depicts an exemplary lithium response pathway according to embodiments described in the present invention;
[0028] Figure 5 is a block diagram showing an exemplary process for generating omics data from a patient's biological sample;
[0029] Figure 6 depicts an example timeline of a patient, which includes example omics, phenomics, socialomics, physiomics, and environmental data collected over time as determined by a pharmacologic phenotype prediction system according to the presently described embodiments, as well as the patient's pharmacologic phenotype; and
[0030] Figure 7 shows a flowchart representing an exemplary method for identifying a pharmacologic phenotype using machine learning techniques according to the presently described embodiments. Detailed Description
[0031] Although the following text sets forth a detailed description of many different embodiments, it should be understood that the legal scope of this description is defined by the words of the claims set forth at the end of this disclosure. The detailed description should be interpreted as being merely exemplary and not describing every possible embodiment, as it would be impractical, if not impossible, to describe every possible embodiment. Many alternative embodiments can be implemented using current technology or technology developed after the filing date of this patent application, and such embodiments will still fall within the scope of the claims.
[0032] It should also be understood that unless a term is explicitly defined in this patent using a sentence such as "As used herein, the term '______' is defined herein to mean..." or a similar sentence, there is no intent to limit the meaning of that term, whether expressly or by implication, beyond its ordinary or common meaning, and such term should not be construed as being limited in scope by any statement made in any section of this patent (other than the language of the claims). To the extent that any term recited in the claims at the end of this patent is referred to in this patent in a manner consistent with a single meaning, this is done only for clarity so as not to confuse the reader and is not intended to limit such claim term by implication or otherwise to that single meaning. Finally, unless a claim element is defined by reference to the word "means" and the recitation of a function without any structure, the scope of any claim element is not intended to be interpreted in accordance with the provisions of 35 U.S.C. § 112, paragraph 6.
[0033] Thus, as used herein, the term "healthcare provider" can refer to any provider of medical or health services. For example, a healthcare worker can be a doctor, clinician, nurse, physician assistant, insurer, pharmacist, hospital, clinical institution, pharmacy technician, pharmaceutical company, research scientist, other medical organization, or medical professional authorized to prescribe medical products and medications for a patient, among others.
[0034] As used herein, the term "patient" can refer to any human or other organism or combination thereof whose health, lifespan, or other medical outcome is the object of clinical or research interest, study, or endeavor.
[0035] Additionally, as used herein, the term "omics" can refer to a series of molecular biology techniques related to the biological functions within cells and the interactions of other functions in the human body. For example, omics can include genomics, epigenomics, chromatin state, transcriptomics, proteomics, metabolomics, biological networks, and system models, etc. Omics data may be specific to individual time points and specific cell tissues and cell lines. Therefore, omics data collection is related to these characteristics, and omics data can also be collected and used for multiple tissues, lineages, and time points related to the patient phenotypes of interest. The omics of a patient may be related to biomarkers of multiple phenotypes, such as the pharmacological response to drugs, disease risk, complications, drug abuse problems, etc. Omics data can be generated and collected for a specific set of medical decision-making purposes at discrete time points, and omics data can also be collected from the total record of omics data collected for an individual patient at various points in the past.
[0036] As used herein, the term "pharmacological phenotype" can refer to any distinguishable phenotype that may affect drug therapy, patient lifespan and outcomes, quality of life, etc. in clinical care, the management and finance of clinical care, and pharmaceutical and other medical and biomedical research on humans and other organisms. Such phenotypes can include pharmacokinetic (PK) and pharmacodynamic (PD) phenotypes, including all phenotypes of the rates and characteristics of drug absorption, distribution, metabolism, and excretion (ADME), as well as drug responses related to drug efficacy, drug treatment dose, half-life, plasma level, clearance rate, etc., and adverse drug events, adverse drug reactions, and the corresponding severity of adverse drug events or adverse drug reactions, organ damage, drug abuse and dependence and their likelihood, as well as body weight and its changes, mood and behavior changes and disturbances. Such phenotypes can also include favorable and adverse responses to drug combinations, drug-gene interactions, social and environmental factors, dietary factors, etc. The phenotypes may also include compliance with pharmacological or non-pharmacological treatment regimens. The phenotypes may also include medical phenotypes, such as the tendency of a patient to contract a certain disease or complication, the outcome and prognosis of the disease, whether a patient will exhibit specific disease symptoms, and the patient's outcomes (such as lifespan, clinical scores and parameters, test results, healthcare expenditures), and other phenotypes.
[0037] In addition, as used herein, the term "drug phenomics" may refer to an individual patient's pharmacologic phenotype based on the integration of genomics, epigenomics, omics, drug metabolomics, socialomics, electronic health records (EHRs), and other patient data, and matched to stratified patient cohorts and population datasets enabled by machine learning.
[0038] As used herein, "precision patient phenotype" may refer to a comprehensive analysis of drug phenomics data to provide a patient treatment profile for precise and accurate clinical decision-making, which may be updated periodically to incorporate changing patient phenomics data.
[0039] As used herein, the term "phenotypic shift" may refer to periodic changes in a clinical patient phenotype that recur or occur intermittently over time according to disease progression, sociological and environmental factors, and / or the results of initial, ongoing, or changing pharmacologic and non-pharmacologic treatments, which is essentially a longitudinal record of a patient's clinical progression.
[0040] In addition, as used herein, the term "disease susceptibility" may refer to risk factors that are directly genetically inherited or modified epigenetically through transgenes.
[0041] As used herein, "socialomics risk factors" may refer to sociological and cultural clinical risk factors related to: behaviors harmful to oneself or others; adverse cultural environments, economic and community living conditions; neglect and abuse during childhood and / or adolescence, which are referred to as adverse childhood experiences (ACEs); adult traumas related to sexual, physical, and psychological abuse; other acute or chronic trauma events (e.g., military conflict, crime, breakdown, disease, death in the family); exacerbated or chronic stress caused by adverse conditions; age-related health conditions, isolation, or cognitive conditions.
[0042] As used herein, "disease diagnosis" may refer to a possible or definitive diagnosis that leads to treatment decision-making. As used herein, "treatment option" may refer to one or more pharmacologic and / or non-pharmacologic treatments that mitigate, neutralize, or improve a patient's condition.
[0043] Similarly, as used herein, the term "initial treatment response" may refer to disease stabilization, lack of response, improvement in clinical response, or adverse events (AEs) due to drug treatment within the first few weeks to months; may involve dose adjustment or adjunctive medications. The time period is typically six months to one year.
[0044] As used herein, the term "relapse response" may refer to cyclic changes in a patient's response to treatment due to pharmacological adverse reactions, drug-drug interactions, drug dose changes, new or relapsing complications, trauma, stress, and other social omics factors, which changes are measured by biological samples (such as but not limited to blood, urine, sweat (e.g., cortisol), odor) or by remote sensing, transmitters, or other active or passive data collection methods.
[0045] As used herein, the term "environment" shall refer to any object, substance, emanation, condition, experience, communication, or information external to or originating from a human or other animal or organism, which occur in the present or past (including biological generations prior to such human or other organism), occur at one or more discrete time points or over a period of time, and which may affect or alter the physical, biological, chemical, physiological, medical, psychological, or psychiatric characteristics of such human or other organism as shown in a measurable, identifiable, or otherwise significant manner. These conditions may include the type, quantity, quality, presence / absence, timing, or other characteristics of food, nutritional supplements, minerals, water and other liquids, clothing, sanitation, and other goods and services to which a human or other organism has been exposed, as well as exposure to chemicals, the atmosphere, and organisms currently or in the past, whether through the skin or by ingestion, inhalation, intubation, speculation, or other means. These conditions may include temperature, noise, light, electromagnetic and / or particle radiation, vibration, mechanical shock or stress, drugs, medical procedures, and implants. These conditions may also include occupational attributes, job duties, and recreational substances. These conditions may also include medical adverse events, such as exposure to toxins, poisons, microorganisms, viruses, and other agents, as well as physical effects, lacerations, contusions, punctures, and concussions.
[0046] These conditions may also include social factors, such as adverse childhood experiences (ACE) as well as stress, trauma, abuse, poverty and other economic conditions, food insecurity and hunger, imprisonment, interpersonal conflict, violence, and other experiences. These conditions may also include the presence or absence of parents, children, siblings, and other family members and acquaintances, including the type, quality, and duration of such relationships. These conditions may also include educational and professional experiences and achievements, religious services and guidance, and social interactions and engagements. These conditions may also include social omics risk factors as well as body modifications, including tattoos, implants, piercings, and pins.
[0047] As used herein, the term "concurrent pharmacogenomic substance exposure" can refer to a subtype of environmental elements that can be unidirectional interactions, concurrent interactions, pharmacokinetic or pharmacodynamic interactions, drug-environment interactions. For recent or ongoing exposures at the time of analysis, for example, if there is a documented interaction indicating that an environmental factor alone induces or inhibits the activity of a specific enzyme related to drug metabolism by ≥20%, or alters drug action by ≥20%, then such an interaction can be considered to have clinical significance. Such interactions can include exposure forms ranging from food to herbal / vitamin supplements to voluntary and involuntary toxic exposures. The likelihood of such an interaction can be measured numerically.
[0048] For simplicity, throughout the discussion, a patient having data that has been used as training data to generate a statistical model can be referred to herein as a "training patient", and a patient having data that is applied to the statistical model to predict a pharmacologic phenotype can be referred to as a "current patient". However, this is for ease of discussion. Data from a "current patient" can be added to the training data, and the training data can be updated continuously or periodically to keep the statistical model up-to-date. Additionally, a training patient can also have data that is applied to the statistical model to predict a pharmacologic phenotype.
[0049] Additionally, throughout the discussion, a current patient can be described as a patient whose pharmacologic phenotype is unknown, while a training patient can be described as a patient whose pharmacologic phenotype is known. More specifically, the pharmacologic phenotype of the current patient is unknown, and prediction is made using the relationship between the omics and socialomics, physiomics, and environmental data of the training patient and the previously or currently determined pharmacologic phenotype of the training patient. Thus, the training patient has a known, previously or currently determined pharmacologic phenotype. The current patient has an unknown pharmacologic phenotype. However, in some embodiments, a training patient may have other unknown pharmacologic phenotypes while having some known pharmacologic phenotypes for training a pharmacologic phenotype prediction system. Additionally, a current patient may have some known, previously or currently determined pharmacologic phenotypes while having unknown pharmacologic phenotypes to be predicted by the pharmacologic phenotype prediction system.
[0050] In general, techniques for identifying pharmacological phenotypes based on omics, social omics, physiomics, and environmental characteristics can be implemented in one or more client devices, one or more network servers, or a system comprising a combination of these devices. However, for clarity, the examples below mainly focus on one embodiment in which a pharmacological phenotype assessment server obtains a set of training data. In some embodiments, the training data can be obtained from client devices. For example, a healthcare provider can obtain a biological sample (e.g., from saliva, cheek swab, sweat, skin sample, biopsy, blood sample, urine, feces, sweat, lymph fluid, bone, bone marrow, hair, odor, etc.) for measuring the omics of a patient and provide the laboratory results obtained by analyzing the biological sample to the pharmacological phenotype assessment server.
[0051] An example process 500 for generating omics data from a patient's biological sample is shown in Figure 5 . The process can be performed by an analysis laboratory or other suitable institution. At block 502, a healthcare provider obtains a patient's biological sample and sends it to an analysis laboratory for analysis. The biological sample can include the patient's saliva, sweat, skin, blood, urine, feces, sweat, lymph fluid, bone marrow, hair, cheek cells, odor, etc. Then at block 504, cells are extracted from the biological sample and reprogrammed into stem cells, such as induced pluripotent stem cells (iPSCs), at block 506. Then at block 508, the iPSCs are differentiated into various tissues, such as neurons, cardiomyocytes, etc., and analyzed at block 510 to obtain omics data. The omics data can include genomic data, epigenomic data, transcriptomic data, proteomic data, karyomic data, metabolomic data, and / or biological networks. As described in more detail below with reference to Figures 4A - 4C , SNPs, genes, and genomic regions can be identified as being related to a specific pharmacological phenotype. When analyzing a patient's omics, social omics, physiomics, and environmental data for a specific pharmacological phenotype or a set of pharmacological phenotypes (e.g., a pharmacological phenotype indicating a response to valproic acid), the iPSCs can be analyzed for the identified SNPs, genes, and genomic regions related to the specific pharmacological phenotype. More generally, the omics data to be analyzed can be selected based on the omics data identified as being related to the set of pharmacological phenotypes being examined for the patient.
[0052] More specifically, cells are reprogrammed into iPSCs by introducing transcription factors or "reprogramming factors" or other reagents into a given cell type. For example, cells can be reprogrammed into iPSCs using the Yamanaka factors (comprising the transcription factors Oct4, Sox2, cMyc, and Klf4). The iPSCs can then be differentiated into a variety of tissues such as neurons, adipocytes, cardiomyocytes, pancreatic islet beta cells, and the like. After differentiating the iPSCs, various analytical techniques (such as DNA methylation analysis, DNase footprinting analysis, filter binding analysis, etc.) can be used to analyze the differentiated iPSCs to identify epigenomic information. In fact, the pharmacologic phenotype prediction system performs a virtual biopsy, and the differentiated iPSCs have at least to some extent the phenotypic and epigenomic characteristics of their corresponding tissues.
[0053] In the above embodiments, cells are extracted from a patient's biological sample, reprogrammed into stem cells, differentiated into various tissues, and analyzed to obtain omics data (differentiated, reprogrammed cell assays). Alternatively, in certain embodiments, the patient's biological sample is assayed without extracting cells (cell-free assays). In other embodiments, cells are extracted from a patient's biological sample and analyzed without reprogramming or differentiating the cells (primary cell assays). In other embodiments, cells are reprogrammed into iPSCs and analyzed without differentiating the cells (reprogrammed stem cell assays). For example, iPSCs can be analyzed without differentiation to obtain stem cell omics. Although these are just some example processes for generating omics data from a patient's biological sample, the analysis can be performed at any suitable stage in the process, and the omics data can be generated in any suitable manner.
[0054] Healthcare providers can also obtain physiological metrics including vital signs, sleep cycles, circadian rhythms, and the like. In addition, healthcare providers can obtain data related to pharmacometabolomics, including metabolites (such as acetate, lactate, etc.) that are products of metabolism and pharmacometabolomic metabolites of drugs. Metabolites can be identified, for example, by spectroscopy or spectrometry performed on the patient's biological sample in a laboratory, and the results can be provided to the healthcare provider as the patient's metabolic profile. The metabolic profile can then be used to identify metabolic disease signatures, identify compounds that can alter drug response, identify metabolite variables, and map the metabolite variables to known metabolic and biological pathways, and the like.
[0055] In some embodiments, a pharmacologic phenotype prediction system can utilize pharmacometabolomics data that includes a systematic assessment of the presence or absence and / or quantitative levels of multiple drugs and drug metabolites. Such information can be collected from whole blood, citrated blood, blood spots, other tissues, and body fluids, among others. The pharmacologic phenotype prediction system can utilize one or more pre-existing instances of pharmacometabolomics data in an EHR system or other database, and / or data for a current treatment or pharmacologic phenotype prediction query. Data for prescription drugs, over-the-counter drugs, non-prescription drugs, illicit drugs, etc. can be collected simultaneously. The concentrations of drugs and metabolites can be measured by techniques that include mass spectrometry and other forms of spectroscopy and spectrometry and / or nuclear magnetic resonance, antibodies, and affinity testing, among others. Such information can be used in embodiments to detect drug abuse or off-label use, measure compliance with prescription drugs, detect other prescription or non-prescription drugs used by a patient or prescribed at other clinics to evaluate the metabolite status of the patient and other purposes, etc., and to make treatment recommendations that include prescribing, discontinuing, and substituting drugs, as well as dose and regimen changes, mode of administration, monitoring, testing and diagnosis, expert referral, additional diagnosis, other treatment methods, etc.
[0056] In other embodiments, physiological metrics can be obtained from a patient's client computing device, fitness tracker, or quantified self-report / passive reporting methods. In another example, a healthcare provider can obtain a patient questionnaire (including questions related to the patient's demographics, medical history, socioeconomic status, law enforcement history, sleep cycle, circadian rhythm, etc.), and can provide the results of the patient questionnaire to a pharmacologic phenotype assessment server. Training data can be obtained from an electronic medical record (EMR) located on an EMR server and / or from multi-pharmacy data located on a multi-pharmacy server that aggregates pharmacy data for patients from multiple pharmacies. In some embodiments, training data can be obtained from a combination of sources that includes several servers (e.g., an EMR server, a multi-pharmacy server, etc.) as well as client devices of healthcare providers and patients. For example, training data for a specific patient can be obtained by cross-referencing the patient's personal history data (e.g., the patient's occupation, place of residence, etc.) with broader longitudinal data for these characteristics (e.g., data in the Human Exposome Project).
[0057] In addition to providing training data to a pharmacologic phenotype assessment server that includes omics data for patients with known pharmacologic phenotypes, the pharmacologic phenotype assessment server also obtains baseline omics levels, omics distributions, or consortium omics data that can be used to train the pharmacologic phenotype assessment server with any other suitable omics data.
[0058] In any case, a subset of the training data can be associated with the training patients corresponding to the subset of the training data. Additionally, for example, a pharmacologic phenotype assessment server can assign subsets of training patients and corresponding training data to cohorts based on demographics. Then, the training data can be used to train the pharmacologic phenotype assessment server to generate a statistical model for predicting the pharmacologic phenotype of a patient. Various machine learning techniques can be used to train the pharmacologic phenotype assessment server.
[0059] After the pharmacologic phenotype assessment server is trained, omics data, socialomics data, physiomics data, and environmental data of a current patient, whose pharmacologic phenotype may be collected at multiple time points, can be received. In some embodiments, the pharmacologic phenotype assessment server can obtain an indication of the disease or disorder that the current patient has, to identify the best drug for treating each disease. This can include stress-related diseases such as post-traumatic stress disorder (PTSD), depression, suicidal tendency, circadian rhythm disorder, substance use disorder, phobia, stress ulcer, acute stress disorder, stress-related diseases included in the Oxford Handbook of Psychiatry, etc. The disease or disorder that the current patient has can also include bipolar disorder, schizophrenia, autism spectrum disorder, and attention deficit hyperactivity disorder (ADHD). Additionally, this can include generalized anxiety disorder and anxious depression, as well as non-psychiatric complications such as irritable bowel syndrome (IBS), inflammatory bowel disease (IBD), Crohn's disease, gastritis, gastric and duodenal ulcers, and gastroesophageal reflux disease (GERD). Further, the disease or disorder that the current patient has can include heart disease, fibromyalgia, chronic fatigue syndrome, etc. The pharmacologic phenotypes of these diseases or disorders can include pharmacologic phenotypes associated with any current and future drugs and / or other methods for treating the corresponding diseases or disorders.
[0060] Then, for example, various machine learning techniques can be used to analyze the omics, socialomics, physiomics, and environmental data to predict one or more pharmacologic phenotypes of the patient. An indication of the pharmacologic phenotype can be transmitted to a client device of a healthcare provider for the healthcare provider to examine and determine an appropriate treatment process based on the pharmacologic phenotype. Pharmacologic phenotypes can be predicted in a clinical setting as well as in a research setting for drug development and insurance applications. In a research setting, the pharmacologic phenotypes of potential patient cohorts related to an experimental drug may be predicted in a research protocol. Patients can be selected for experimental treatment based on their predicted pharmacologic phenotypes related to the experimental drug.
[0061] Refer to Figure 1A, the exemplary pharmacogenetic phenotype prediction system 100 uses various machine learning techniques to predict a patient's pharmacogenetic phenotype (precise patient phenotype) based on the patient's omics, socialomics, physiomics, and environmental data. The pharmacogenetic phenotype prediction system 100 can obtain training data for a training patient cohort and can analyze this data to identify relationships between the omics, socialomics, physiomics, and environmental data and the pharmacogenetic phenotypes included in the training data. The pharmacogenetic phenotype prediction system 100 can then generate a statistical model for predicting pharmacogenetic phenotypes based on the analysis. When a patient's pharmacogenetic phenotype is unknown (e.g., lithium has not been prescribed to a bipolar disorder patient, so it is not yet clear how the patient will respond to lithium), the pharmacogenetic phenotype prediction system 100 can obtain the patient's omics, socialomics, physiomics, and environmental data and apply the omics, socialomics, physiomics, and environmental data to the statistical model to predict the patient's pharmacogenetic phenotype. For example, the pharmacogenetic phenotype prediction system 100 can predict the likelihood that a patient will have an adverse reaction to a particular drug, can predict the efficacy or appropriate dose of a drug, etc. The pharmacogenetic phenotype prediction system 100 can perform clinical decision support (CDSS) in a clinical setting to predict the patient's precise patient phenotype. Additionally, the pharmacogenetic phenotype prediction system 100 can conduct drug research to develop companion diagnostic tests to identify patients who will have a good or adverse reaction to a developed or approved drug and who will experience fewer side effects or no side effects. Further, the pharmacogenetic phenotype prediction system 100 can be used in the context of experimental treatment to recommend to researchers experimental drugs and / or doses to be prescribed to current patients in the context of a clinical study.
[0062] The pharmacogenetic phenotype prediction system 100 includes a pharmacogenetic phenotype assessment server 102 and a plurality of client devices 106 - 116 that can be communicatively connected via a network 130, as described below. In one embodiment, the pharmacogenetic phenotype assessment server 102 and the client devices 106 - 116 can communicate over the communication network 130 via wireless signals 120, which can be any suitable local area network or wide area network, including WiFi networks, Bluetooth networks, cellular networks (such as 3G, 4G, Long Term Evolution (LTE), 5G), the Internet, etc. In some cases, the client devices 106 - 116 can communicate with the communication network 130 via an intervening wireless or wired device 118, which can be a wireless router, a wireless repeater, a base transceiver station of a mobile phone provider, etc. By way of example, the client devices 106 - 116 can include a tablet computer 106, a smartwatch 107, a network-enabled cellular phone 108, a wearable computing device (such as Google Glass TM or 109), a personal digital assistant (PDA) 110, a mobile device smart phone 112 (also referred to herein as a "mobile device"), a laptop computer 114, a desktop computer 116, a wearable biosensor, a portable media player (not shown), a phablet, any device configured to perform wired or wireless RF (radio frequency) communication, etc. In addition, any other suitable client device that records a patient's omics data, clinical data, demographic data, polypharmacy data, social omics data, physiomics data, or other environmental data may also communicate with the pharmacogenetic phenotype assessment server 102.
[0063] In some embodiments, a patient may input data into the desktop computer 116, such as answers in response to a patient questionnaire that includes questions related to the patient's demographics, medical history, socioeconomic status, law enforcement history, sleep cycle, circadian rhythm, etc. In other embodiments, a healthcare provider may input data.
[0064] Each of the client devices 106 - 116 may interact with the pharmacogenetic phenotype assessment server 102 to send a patient's omics data, clinical data, demographic data, polypharmacy data, social omics data, physiomics data, or other environmental data. In some embodiments, social omics, physiomics, and environmental data may be collected periodically (e.g., monthly, quarterly, semi - annually, etc.) to identify changes in the patient's social and environmental conditions over time (e.g., from unemployed to employed, single to married, etc.). Similarly, in some embodiments, at least some of the patient's social omics, physiomics, and environmental data may be recorded by a healthcare provider via the healthcare provider's client devices 106 - 116, or may be self - reported via the patient's client devices 106 - 116.
[0065] Each client device 106 - 116 may also interact with the pharmacogenetic phenotype assessment server 102 to receive one or more indications of the predicted pharmacogenetic phenotype of the current patient. The indications may include recommendations for drugs to be prescribed to the current patient, for which the current patient has the highest expected response (e.g., the highest combination of efficacy and minimal adverse drug reactions and severity of the reaction). The indications may also include the risks of various diseases for the current patient, such as the likelihood of getting the disease, risk categories (e.g., low, medium, or high risk), etc. In addition, the indications may include the likelihood of drug abuse, such as a numerical likelihood or likelihood category (e.g., low, medium, or high likelihood).
[0066] In an example implementation, the Pharmacological Phenotype Assessment Server 102 can be a cloud-based server, an application server, a web server, etc., and includes a memory 150, one or more processors (CPUs) 142 (such as a microprocessor coupled to the memory 150), a network interface unit 144, and an I / O module 148, which can be, for example, a keyboard or a touch screen.
[0067] The Pharmacological Phenotype Assessment Server 102 can also be communicatively connected to the Consortium Omics / Environmental / Physiological / Demographic / Pharmacy Information Database 154. The Consortium Omics / Environmental / Physiological / Demographic / Pharmaceutical Information Database 154 can store training data and statistical models for determining pharmacological phenotypes, where the training data includes omics data of training patients, genome-based ethnicity data, clinical data, demographic data, multi-pharmacy data, social omics data, physiological omics data, or other environmental data. The Consortium Omics / Environmental / Physiological / Demographic / Pharmacy Information Database 154 can also include a consortium omics database, an academic omics database, and a pharmacy database, including (for example) RxNorm, drug-drug interactions (such as FDA black box labels), drug-gene interactions, and others. In some embodiments, to determine a pharmacological phenotype, the Pharmacological Phenotype Assessment Server 102 can retrieve patient information of each training patient from the Consortium Omics / Environmental / Physiological / Demographic / Pharmacy Information Database 154.
[0068] The memory 150 can be a tangible non-transitory memory and can include any type of suitable memory module, including random access memory (RAM), read-only memory (ROM), flash memory, other types of persistent memory, etc. The memory 150 can store, for example, instructions for an operating system (OS) 152 that can be executed on the processor 142, and the operating system can be any type of suitable operating system, such as a modern smartphone operating system. The memory 150 can also store, for example, instructions for a machine learning engine 146 that can be executed on the processor 142, and the machine learning engine can include a training module 160 and a phenotype assessment module 162. The Pharmacological Phenotype Assessment Server 102 is described in more detail below with reference to Figure 1B In some embodiments, the machine learning engine 146 can be part of one or more of the client devices 106-116, the Pharmacological Phenotype Assessment Server 102, or a combination of the Pharmacological Phenotype Assessment Server 102 and the client devices 106-116.
[0069] In any case, the machine learning engine 146 can receive electronic data from the client devices 106 - 116. For example, the machine learning engine 146 can obtain a set of training data by receiving omics data, clinical data, demographic data, polypharmacy data, social omics data, physiological omics data, or other environmental data, etc. Additionally, the machine learning engine 146 can obtain a set of training data by receiving phenomics data related to the pharmacological phenotypes of the training patients (such as chronic diseases suffered by the training patients, responses to drugs previously prescribed to the training patients, whether each of the training patients has a drug abuse problem, etc.).
[0070] Thus, the training module 160 can classify omics data, social omics data, physiomics data, and environmental data into specific pharmacological phenotypes, such as drug abuse, specific types of chronic diseases, adverse drug reactions to specific drugs, efficacy levels of specific drugs, etc. Then, the training module 160 can analyze the classified omics data, social omics data, physiomics data, and environmental data to generate statistical models for each pharmacological phenotype. For example, a first statistical model can be generated to determine the likelihood that a current patient will experience a drug abuse problem, a second statistical model can be generated to determine the risk of having one disease, a third statistical model can be generated to determine the risk of having another disease, a fourth statistical model can be generated to determine the likelihood of having a negative reaction to a specific drug, etc. In some embodiments, each statistical model can be combined in any suitable manner to generate an overall statistical model for predicting each of the pharmacological phenotypes.In any case, various machine learning techniques can be used to analyze the training data set, including but not limited to regression algorithms (e.g., ordinary least squares regression, linear regression, logistic regression, stepwise regression, multivariate adaptive regression splines, locally estimated scatterplot smoothing, etc.), instance-based algorithms (e.g., k-nearest neighbor, learning vector quantization, self-organizing map, locally weighted learning, etc.), regularization algorithms (e.g., ridge regression, least absolute shrinkage and selection operator, elastic net, least angle regression, etc.), decision tree algorithms (e.g., classification and regression tree, C4.5, C5, chi-squared automatic interaction detection, decision stump, M5, conditional decision tree, etc.), clustering algorithms (e.g., k-means, k-medians, expectation maximization, hierarchical clustering, spectral clustering, mean shift, density-based spatial clustering of applications with noise, ordering points to identify the clustering structure, etc.), association rule learning algorithms (e.g., Apriori algorithm, Eclat algorithm, etc.), Bayesian algorithms (e.g., naive Bayes, Gaussian naive Bayes, multinomial naive Bayes, averaged one-dependence estimators, Bayesian belief network, Bayesian network, etc.), artificial neural networks (e.g., perceptron, Hopfield network, radial basis function network, etc.), deep learning algorithms (e.g., multi-layer perceptron, deep Boltzmann machine, deep belief network, convolutional neural network, stacked autoencoder, generative adversarial network, etc.), dimensionality reduction algorithms (e.g., principal component analysis, principal component regression, partial least squares regression, Sammon mapping, multidimensional scaling, projection pursuit, linear discriminant analysis, mixture discriminant analysis, quadratic discriminant analysis, flexible discriminant analysis, factor analysis, independent component analysis, non-negative matrix factorization, t-distributed stochastic neighbor embedding, etc.), ensemble algorithms (e.g., boosting, bagging, AdaBoost, stacking generalization, gradient boosting machine, gradient boosting regression tree, random decision forest, etc.), reinforcement learning (e.g., temporal difference learning, Q-learning, learning automata, state-action-reward-state-action, etc.), support vector machines, mixture models, evolutionary algorithms, probabilistic graphical models, etc.
[0071] In the testing phase, the training module 160 can compare the test omics data, social omics data, physiological omics data, and environmental omics data of the test patient with the statistical model to determine the likelihood that the test patient has a specific pharmacological phenotype.
[0072] If the training module 160 makes correct judgments more frequently than a predetermined threshold amount, the statistical model can be provided to the phenotype evaluation module 162. On the other hand, if the training module 160 does not make correct judgments more frequently than the predetermined threshold amount, the training module 160 can continue to obtain training data for further training.
[0073] The phenotypic evaluation module 162 can obtain a statistical model and a set of omics data, social omics, physiomics, and environmental data of the current patient, and the data can be collected over a period of time (e.g., one month, three months, six months, one year, etc.). For example, biological samples of the current patient (e.g., blood samples, saliva, biopsies, bone marrow, hair, etc.) can be analyzed in a laboratory to obtain genomics data, epigenomics data, transcriptomics data, proteomics data, karyomics data, and / or metabolomics data of the current patient. Then the omics data can be provided to the phenotypic evaluation module 162. Additionally, clinical data of the patient can be provided from the EMR server or client devices 106-116 of the healthcare provider. Multi-pharmacy data can be provided from a multi-pharmacy server or from several pharmacy servers, and demographic data, social omics data, physiomics data, and other environmental data can be provided from the client devices 106-116 of the healthcare provider or the client devices 106-116 of the current patient.
[0074] Then, the omics, social omics, physiomics, and environmental data can be applied to the statistical model generated by the training module 160. Based on the analysis, the phenotypic evaluation module 162 can determine the likelihood indicating that the current patient has certain pharmacological phenotypes or other semi-quantitative and quantitative metrics, such as the likelihood of drug abuse, the likelihood of various diseases, the overall rating of the predicted response to various drugs, etc. The phenotypic evaluation module 162 can cause the likelihood to be displayed on the user interface for inspection by the healthcare provider. Each likelihood can be expressed as a probability (e.g., 0.6), a percentage (e.g., 80%), a category in a set of categories (e.g., "high", "medium", or "low"), and / or in any other suitable way.
[0075] The pharmacological phenotype evaluation server 102 can communicate with the client devices 106-116 via the network 130. The digital network 130 can be a private network, a secure public internet, a virtual private network, and / or some other type of network, such as a dedicated access line, an ordinary conventional telephone line, a satellite link, a combination of these, etc. In the case where the digital network 130 includes the Internet, data communication can be carried out on the digital network 130 via the Internet communication protocol.
[0076] Now turn to Figure 1B, the pharmacology phenotype assessment server 102 may include a controller 224. The controller 224 may include a program memory 226, a microcontroller or microprocessor (MP) 228, a random access memory (RAM) 230, and / or input / output (I / O) circuitry 234, all of which may be interconnected via an address / data bus 232. In some embodiments, the controller 224 may also include a database 239, or otherwise be communicatively coupled to the database or other data storage mechanisms (e.g., one or more hard disk drives, optical storage drives, solid state storage devices, etc.). The database 239 may include data such as patient information, training data, risk analysis templates, web page templates, and / or web pages, as well as other data necessary for interacting with users via the network 130. The database 239 may include data similar to the consortium omics / environmental / physiomics / demographics / pharmacy information database 154 described above with reference to Figure 1A and / or the data sources 325a-d described below with reference to Figure 3 (e.g., the biomedical training set 325a, the pharmacology database 325b, the environmental data 325c, and the granulated data 325d).
[0077] It should be understood that although Figure 1B only one microprocessor 228 is depicted, the controller 224 may include multiple microprocessors 228. Similarly, the memory of the controller 224 may include multiple RAMs 230 and / or multiple program memories 226. Although Figure 1B the I / O circuitry 234 is described as a single block, the I / O circuitry 234 may include many different types of I / O circuitry. The controller 224 may implement one or more RAMs 230 and / or program memories 226 as, for example, semiconductor memories, magnetically readable memories, and / or optically readable memories.
[0078] As Figure 1B shown, the program memory 226 and / or the RAM 230 may store various applications for execution by the microprocessor 228. For example, a user interface application 236 may provide a user interface 102 to the pharmacology phenotype assessment server, which may, for example, allow a system administrator to configure, troubleshoot, or test various aspects of server operation. A server application 238 may operate to receive a set of omics data, social omics data, physiomics data, and environmental data of a current patient, determine an indication of the likelihood that the current patient has a pharmacology phenotype or other semi-quantitative and quantitative metrics, and send an indication of the likelihood to the client devices 106-116 of healthcare providers. The server application 238 may be a single module 238 or multiple modules 238A, 238B, such as a training module 160 and a phenotype assessment module 162.
[0079] Although in Figure 1B the server application 238 is depicted as including two modules 238A and 238B, the server application 238 can include any number of modules that perform tasks related to the implementation of the pharmacological phenotype assessment server 102. It should be understood that although only one pharmacological phenotype assessment server 102 is depicted in Figure 1B , multiple pharmacological phenotype assessment servers 102 can be provided for distributing server loads, serving different web pages, etc. These multiple pharmacological phenotype assessment servers 102 can include web servers, entity-specific servers (e.g., servers, etc.), servers located in retail or private networks, etc.
[0080] Now referring to Figure 1C , the laptop computer 114 (or any one of the client devices 106 - 116) can include a display 240, a communication unit 258, a user input device (not shown), and a controller 242 similar to the pharmacological phenotype assessment server 102. Similar to the controller 224, the controller 242 can include a program memory 246, a microcontroller or microprocessor (MP) 248, a random access memory (RAM) 250, and / or input / output (I / O) circuitry 254, all of which can be interconnected via an address / data bus 252. The program memory 246 can include an operating system 260, a data storage device 262, multiple software applications 264, and / or multiple software routines 268. For example, the operating system 260 can include Microsoft OS and so on. The data storage device 262 can include data such as patient information, application data of multiple applications 264, routine data of multiple routines 268, etc., and / or other data necessary for interacting with the pharmacological phenotype assessment server 102 via the digital network 130. In some embodiments, the controller 242 can also include other data storage mechanisms (e.g., one or more hard disk drives, optical storage drives, solid-state storage devices, etc.) resident within the laptop computer 114, or otherwise communicatively connected to the other data storage mechanisms.
[0081] The communication unit 258 can communicate with the pharmacogenetic phenotype assessment server 102 via any suitable wireless communication protocol network such as a wireless telephone network (e.g., GSM, CDMA, LTE, etc.), a Wi-Fi network (802.11 standard), a WiMAX network, a Bluetooth network, etc. The user input device (not shown) can include a "soft" keyboard displayed on the display 240 of the laptop computer 114, an external hardware keyboard (e.g., a Bluetooth keyboard) communicating via a wired or wireless connection, an external mouse, a microphone for receiving voice input, or any other suitable user input device. As discussed with reference to the reference controller 224, it should be understood that although Figure 1C only one microprocessor 248 is depicted, the controller 242 can include multiple microprocessors 248. Similarly, the memory of the controller 242 can include multiple RAMs 250 and / or multiple program memories 246. Although Figure 1C the I / O circuit 254 is described as a single block, the I / O circuit 254 can include many different types of I / O circuits. The controller 242 can implement one or more RAMs 250 and / or program memories 246 as, for example, semiconductor memories, magnetically readable memories, and / or optically readable memories.
[0082] In addition to other software applications, one or more of the processors 248 can be adapted and configured to execute any one or more of the multiple software applications 264 residing in the program memory 246 and / or any one or more of the multiple software routines 268. One of the multiple applications 264 can be a client application 266, which can be implemented as a series of machine-readable instructions for performing various tasks associated with receiving information at the laptop computer 114, displaying information on the laptop computer, and / or sending information from the laptop computer.
[0083] One of the multiple applications 264 can be a native application and / or a web browser 270 (such as Apple’s Google Chrome TM , Microsoft Internet and Mozilla ), which can be implemented as a series of machine-readable instructions for receiving, interpreting, and / or displaying web page information from the pharmacogenetic phenotype assessment server 102 while also receiving input from a user such as a healthcare provider. Another of the multiple applications can include an embedded web browser 276, which can be implemented as a series of machine-readable instructions for receiving, interpreting, and / or displaying web page information from the pharmacogenetic phenotype assessment server 102.
[0084] One of the multiple routines may include a risk analysis display routine 272 that obtains the likelihood that the current patient has certain pharmacologic phenotypes and displays the likelihood and / or an indication of a recommendation for treating the current patient on the display 240. Another of the multiple routines may include a data input routine 274 that obtains the socialomics, physiomics, and environmentalomics data of the current patient from a healthcare provider and sends the received socialomics, physiomics, and environmentalomics data, along with the previously stored socialomics, physiomics, and environmentalomics data of the current patient (e.g., environmentalomics data collected during a previous visit), to the pharmacologic phenotype assessment server 102.
[0085] Preferably, a user may initiate the client application 266 from a client device (such as one of the client devices 106 - 116) to communicate with the pharmacologic phenotype assessment server 102 to implement the pharmacologic phenotype prediction system 100. Additionally, the user may also initiate or instantiate any other suitable user interface application (e.g., a native application or a web browser 270, or any other application among the multiple software applications 264) to access the pharmacologic phenotype assessment server 102 to implement the pharmacologic phenotype prediction system 100.
[0086] As described above, Figure 1A The illustrated pharmacologic phenotype assessment server 102 may include a memory 150 that stores instructions for the machine learning engine 146 that can be executed on the processor 142. The machine learning engine 146 may include a training module 160 and a phenotype assessment module 162.
[0087] Figure 2shows omics, socionomics, physiomics, and environmental data that can be provided to a pharmacologic phenotype prediction system 100, which in turn predicts pharmacologic phenotypes in a clinical or research setting. The omics, socionomics, physiomics, and environmental data are divided into four categories: individual / cohort and population omics and drug metabolomics 302; exposome 304; socionomics demographics and stress / trauma 306; and medical physiomics, structured or unstructured electronic health records (EHRs), laboratory values, stress and abuse factors and trauma, and medical outcome data 308. However, this is for illustrative purposes only. The exposome 304, socionomics demographics and stress / trauma 306, and medical physiomics, structured or unstructured EHRs, laboratory values, stress and abuse factors and trauma, and medical outcome data 308 can be included as part of the socionomics, physiomics, and environmental data, while the individual / cohort and population omics and drug metabolomics 302 can be included as part of the omics data. Additionally, the individual / cohort and population omics and drug metabolomics 302, exposome 304, socionomics demographics and stress / trauma 306, and medical physiomics, structured or unstructured EHRs, laboratory values, stress and abuse factors and trauma, and medical outcome data 308 can be classified and / or organized in any other suitable manner.
[0088] In any case, the individual / cohort and population omics and drug metabolomics 302 can include genomics, epigenomics, chromatin state, transcriptomics, proteomics, metabolomics, biological networks, and system models, among others, each of which can be extracted from or at least related to the genome. The individual / cohort and population omics and drug metabolomics 302 can also include chemical mapping of discrete molecular entities within tissues to various pharmacologic phenotypes. The discrete molecular entities can be metabolites (such as acetic acid, lactic acid, etc.) that are products of metabolism and drug metabolome metabolites of drugs.
[0089] The exposome 304 can include information indicating the patient's environment, such as the location of the patient's residence, the type of residence, the size of the residence, the quality of the residence, the patient's work environment (including the location of the patient's workplace), the distance from the patient's residence to the workplace, how the patient is treated at the workplace and / or residence, etc. The exposome 304 can also include any other environmental exposures experienced by the patient, including climate factors, lifestyle factors (e.g., tobacco, alcohol), diet, physical activity, pollutants, radiation, infections, educational level, etc.
[0090] Sociomics, demographics, and stress / trauma 306 may include demographic data such as gender, ethnicity, age, income, marital status, education level, language, etc. Sociomics, demographics, and stress / trauma 306 may also include other family data, cultural conditions, circadian rhythm data, age-related health conditions, isolation or cognitive conditions, economic and community living conditions, etc. In addition, sociomics, demographics, and stress / trauma 306 may include trauma, domestic violence, law enforcement history, or any other stress or abuse factors. In some embodiments, stress and abuse factors during childhood may be quantified by an Adverse Childhood Experiences (ACE) score, which assesses different types of abuse, neglect, and other experiences during a difficult childhood. This may include physical, emotional, and sexual abuse, physical and emotional neglect, mental illness within the family, domestic violence within the family, divorce, substance abuse within the family, incarcerated relatives, etc.
[0091] In addition, medical physiomics, structured or unstructured EHRs, laboratory values, stress and abuse factors, trauma, and medical outcome data 308 may include trauma, domestic violence, law enforcement history, or any other stress or abuse factors. In some embodiments, stress and abuse factors during childhood may be quantified by an Adverse Childhood Experiences (ACE) score, which assesses different types of abuse, neglect, and other experiences during a difficult childhood. This may include physical, emotional, and sexual abuse, physical and emotional neglect, mental illness within the family, domestic violence within the family, divorce, substance abuse within the family, incarcerated relatives, etc. Medical physiomics, structured or unstructured EHRs, laboratory values, stress and abuse factors, trauma, and medical outcome data 308 may also include clinical data, polypharmacy data, and physiological characteristics such as human functions related to genes and proteins. In addition, medical outcome data may include the pharmacologic phenotypes of specific patients or patient cohorts. Additionally, medical outcome data may include information indicating drug or treatment efficacy, adverse drug events or adverse drug reactions, disease stabilization, lack of response, improvement in clinical response, etc.
[0092] Personal / cohort and population omics and pharmacometabolomics 302, exposome 304, social omics demographics and stress / trauma 306, and medical physiomics, structured or unstructured EHR, laboratory values, stress and abuse factors and trauma, and medical outcome data 308 of a training patient cohort can be provided as training data to a pharmacologic phenotype prediction system 100 to generate a statistical model for predicting a pharmacologic phenotype. Additionally, individual omics and pharmacometabolomics 302, exposome 304, social omics demographics and stress / trauma 306, and medical physiomics, structured or unstructured EHR, laboratory values, stress and abuse factors and trauma, and medical outcome data 308 or some portions thereof can be obtained from a current patient to apply to the statistical model to predict the pharmacologic phenotype or precise patient phenotype of the current patient.
[0093] Figure 3 A detailed view 320 of a process performed by the pharmacologic phenotype prediction system 100 is shown. As Figure 2 shown, the pharmacologic phenotype prediction system 100 obtains personal / cohort and population omics and pharmacometabolomics 302, exposome 304, social omics demographics and stress / trauma 306, and medical physiomics, structured or unstructured EHR, laboratory values, stress and abuse factors and trauma, and medical outcome data 308 from a training patient cohort as training data to train a machine learning engine 146. In some instances, personal / cohort and population omics and pharmacometabolomics 302 and their respective correlations with a pharmacologic phenotype can be obtained from GWAS, candidate gene association studies, and / or other machine learning methods, as described in more detail below. Training data can also be obtained from several data sources 325a-d including a biomedical training set 325a, a pharmacology database 325b, environmental data 325c, and data segmented by granularity 325d.
[0094] The biomedical training set 325a includes omics, pharmacometabolomics, medical physiomics, EHR, laboratory values, medical outcomes, and stress and abuse factors and trauma similar to the omics and pharmacometabolomics 302 and medical physiomics, structured or unstructured EHR, laboratory values, stress and abuse factors and trauma, and medical outcome data 308 described above with reference to Figure 2 The pharmacology database 325b includes pharmacy records, drug databases, drug-drug interactions, drug-gene interactions, etc. Additionally, the social omics and environmental data 325c includes social omics demographics and stress / trauma similar to the social omics demographics and stress / trauma 306 described above with reference to Figure 2The described exposome 304, social omics demographics, and stress / trauma 306 are similar social omics, demographics, and exposomes. Granularity-divided data 325d can identify any data in the biomedical training set 325a, pharmacological database 325b, and environmental data 325c corresponding to an individual patient, patient cohort, or group of patients, respectively.
[0095] Then, the machine learning engine 146 can utilize omics, social omics, physiomics, environmental, and phenomics data of a cohort or group of training patients to generate a statistical model for predicting pharmacological phenotypes using machine learning techniques. In some embodiments, the machine learning engine 146 can analyze the relationship between omics data and pharmacological phenotypes to identify single nucleotide polymorphisms (SNPs), genes, and genomic regions highly correlated with a specific pharmacological phenotype. This is discussed in more detail below with reference to Figure 4B and 4D which.
[0096] In addition, the machine learning engine 146 can classify a cohort or group of training patients having at least some of the identified SNPs, genes, and genomic regions or any suitable combination thereof according to the phenomics data of each training patient in the cohort or group as having a specific pharmacological phenotype or not having a specific pharmacological phenotype. The machine learning engine 146 can further analyze the social omics, physiomics, and environmental data of a cohort or group of training patients corresponding to each category to generate a statistical model. For example, the machine learning engine 146 can perform statistical measurements on the social omics, physiomics, and environmental data of each category to distinguish the social omics, physiomics, and environmental data of a subset of training patients having a specific pharmacological phenotype and a subset of training patients not having a specific pharmacological phenotype. Supervised learning algorithms (such as classification and regression) can be used to train the machine learning engine 146. Unsupervised learning algorithms (such as dimensionality reduction and clustering) can also be used to train the machine learning engine 146.
[0097] In any case, the machine learning engine 146 can receive an input 330 from the current patient or current patient group without knowing whether the current patient or current patient group has a specific pharmacologic phenotype, the input comprising omics, social omics, physiologic omics, and environmental data. The input can comprise any one of the above individual omics and pharmacometabolomics 302, exposome 304, social omics demographics and stress / trauma 306, and medical physiologic omics, structured or unstructured EHR, laboratory values, stress and abuse factors and trauma, and medical outcome data 308. For example, for a single current patient, the input can comprise personal data, omics laboratory tests, personal physiologic omics, EHR data, medication history, and environmental data. For a group of current patients, the input can comprise personal data, cohort omics, physiologic omics, EHR data, medication history, and environmental data.
[0098] The omics, social omics, physiologic omics, and environmental data of the current patient or current patient group can be applied to a statistical model included in the machine learning engine 146 to predict the pharmacologic phenotype of the current patient or current patient group. For example, the machine learning engine 146 can predict the likelihood of a negative reaction to warfarin. Additionally or alternatively, the machine learning engine 146 can generate a response score indicative of the efficacy of warfarin in treating thrombosis in the current patient, the efficacy being discounted due to the adverse effects of warfarin on the current patient.
[0099] The likelihood, response score, or other semi - quantitative or quantitative metric can be analyzed by the pharmacophenotypic clinical decision support engine 335 within the pharmacophenotypic prediction system 100 to recommend to a healthcare provider a drug and / or dose to be prescribed to the current patient. In some embodiments, the response scores for each of the drugs available for a particular medical indication can be ranked, and the pharmacophenotypic clinical decision support engine 335 can recommend to the healthcare provider the highest - ranked drug to be prescribed to the current patient. Doses of a particular drug can also be ranked. In another example, when the likelihood of a negative reaction to a drug selection is higher than a threshold score for the current patient, the pharmacogenomic clinical decision support engine 335 can recommend a different drug for the particular medical indication. Additionally, the pharmacophenotypic clinical decision support engine 335 can compare the recommended drug with the current patient's polypharmacy data included in the environmental data and / or medical record. If the current patient is taking a drug that is incompatible with the recommended drug, the pharmacophenotypic clinical decision support engine 335 may recommend the drug with the next - highest response score or another drug with a likelihood of negative reaction that is not higher than the threshold likelihood. In other embodiments, when the pharmacophenotype is the likelihood of drug abuse, the pharmacogenomic clinical decision support engine 335 can recommend early intervention, or when the pharmacophenotype is disease risk, the pharmacophenotypic clinical decision support engine 335 can recommend screening and / or treatment options to proactively address the issue. In other embodiments where the pharmacophenotype is different from those described above, the pharmacophenotypic clinical decision support engine 335 can recommend other drug therapies, informative analyses, or other courses of action.
[0100] The likelihood, response score, or other semi - quantitative or quantitative metric of a pharmacophenotype can also be used in drug research 340 in various capacities. For example, researchers developing a drug can use this method to develop a companion diagnostic test to identify patients who will have a good or adverse response to the developed or approved drug and will have fewer or no side effects. Additionally, researchers screening or comparing multiple molecular entities as putative drugs and having comparative data on molecular experiments conducted with these drugs can use these methods to prospectively evaluate the likely effects and adverse events on a population to determine the entities or priorities to be developed during the development process. Further, these methods can be used in the context of experimental treatment to recommend to researchers an experimental drug and / or dose to be prescribed to the current patient in the context of a clinical study. Finally, these methods can be used for the explicit construction or model generation of pharmacogenomic tests that will be conducted outside of an integrated CDSS environment.
[0101] The predicted pharmacologic phenotypes and / or recommendations provided by the pharmacologic phenotype clinical decision support engine 335 or the drug research tool 340 can be provided to the data sources 325a-d in a feedback loop. Then, the omics, social omics, physiomics, environmental, and phenomics data of the current patient are used as training data to further train the machine learning engine 146 for use by other current patients. In this way, the machine learning engine 146 can continuously update the statistical model to reflect at least an almost real-time representation of the social omics, physiomics, environmental, and omics data.
[0102] Figure 4A An exemplary representation depicting the biological context and interactions of multiple different omics modalities within the 4D nucleome, and examples of bioinformatics analysis of omics data in this context are shown. The illustration 450 depicts chromosomes located in regions bound to chromatin in the cell nucleus. Euchromatin is characterized by a specific combination of DNase 1 hypersensitivity and histone marks that define active genomics regulatory elements such as promoter H3K4me3 and H3K27ac, and enhancer H3K4me1 and H3K27ac. Enhancers can increase or decrease transcription in their target genes, which can be located proximally in sequence and / or spatially (e.g., Hi-C or ChIA-PET data, or genome architecture mapping or combinatorial chromatin capture) and / or functionally linked to the enhancer individually or in combination (e.g., through molecular QTL linkage). Heterochromatin is located in the interior of chromosomal regions and at the periphery of the nucleus, near the nuclear lamina and nucleoli, and is characterized by its own repressive chromatin marks and DNA-binding proteins, as well as spatial compaction and linker histones. Recent studies have shown that in the brain, the DNA sequence CAC is a common site of methylation, which is contrary to other tissues where CpG is most often methylated. Additionally, in the brain, a unique reactive species carrying epigenomics information, 5-hydroxymethylcytosine (5hmC), is relatively common. In contrast, in the periphery, methylcytosine (hmC) is common.
[0103] Figure 4A A schematic diagram representing an exemplary spatial hierarchy 460 of transcriptional organization first determined by chromatin conformation capture methods is also shown. The spatial hierarchy 460 includes a Hi-C map 462 showing the multi-scale hierarchy of transcriptional regulation. In this illustration, the normalized frequency of spatial interactions between genomic portions (on the X-axis) and other genomic portions (on the Y-axis) is represented by a color gradient to generate a two-dimensional map of chromatin organization.
[0104] This figure can be generated using "bins" of a fixed length representing DNA sequences, or bins representing increments of cleavage sites or sets thereof, or functional elements such as genes, chromatin state segments, loop domains, chromatin domains, TADs, etc. Thresholds can be used to identify contacts in various normalization modes of distance, overall contact propensity, and other elements. For example, in bins with variable sequence lengths, and in cases where the square genomic regions described by a pair of bins may have variable sizes and shapes, normalization methods can be designed to replace traditional methods that rely on fixed bins. The contact density as a function of distance can be fit to an integrable function, which can be integrated over the rectangular region of bin pairs to produce an expected value of contacts mapped to the square genomic region. Statistical tests such as Poisson distribution p-values, for which the Benjamini false discovery rate can be applied, can be used to compare the expected value with the raw or normalized read counts mapped to the square genomic region to generate, in an adjusted manner, sets of enriched and depleted chromatin contacts at a certain distance, either locally or genome-wide. This can be done for various analytical purposes, including detecting target genes of genomic variants and performing genome-wide analyses of contacts.
[0105] The spatial hierarchy 460 also includes a visualization 464 of nuclear and subnuclear transcriptional topologies as shown in the Hi-C map 462. As shown in the visualization 464, chromosomes fill most of the available volume of the nucleoplasm as regions (CTs), and contain outer A and B compartments composed of euchromatin and heterochromatin, respectively. Active genes tend to localize at the periphery of CTs, and interchromosomal loops between CTs provide the basis for trans enhancer-promoter and promoter-promoter spatial interactions. The A and B chromatin compartments of CTs contain topologically associated domains (TADs), the average length of the linear sequences of which is within approximately 1 Mb. TADs can first be characterized using chromatin conformation capture methods such as Hi-C, in which the initial scaling is consistent with the fractal globule model, while high-resolution studies of enhancer-promoter loops within TADs, the organization of TAD boundary proteins including CCCTC-binding factor (CTCF) and cohesin (RAD21), and direct imaging support the loop extrusion model of TAD organization. Features of transcriptional units, including frequently interacting regulatory elements (FIREs), include an example located within an intron of the GRIN2A gene on chromosome 16.
[0106] Machine learning methods can use Figure 4A the epigenomic profiling and / or bioinformatics analysis described in Figure 4CThe gene regulatory network shown. As described above, the training module 160 can generate a statistical model for each pharmacological phenotype. Figure 4B is a block diagram showing an exemplary method 400 for using machine learning techniques to identify omics data corresponding to a specific pharmacological phenotype. The method 400 can be executed on the pharmacological phenotype assessment server 102. In some embodiments, the method 400 can be implemented in a set of instructions stored on a non-transitory computer-readable memory and executable on one or more processors on the pharmacological phenotype assessment server 102. For example, the method 400 can be executed by Figure 1A the training module 160 within the machine learning engine 146.
[0107] At block 402, a statistical test (e.g., via GWAS or candidate gene association study) is performed for each of several non-coding or coding SNPs, genes, and genomic regions in the genome to determine the relationship between the SNP and a specific pharmacological phenotype, which can be drug response, adverse drug reaction, adverse drug event, dose, disease risk, etc. (e.g., the response of a patient with depression to ketamine). When the statistical test shows a significant relationship between the SNP and the specific pharmacological phenotype (e.g., a p-value less than a threshold probability using the null hypothesis), then the SNP is determined to be associated with the specific pharmacological phenotype. In some embodiments, SNPs can be identified based on Figure 4A the bioinformatics analysis shown.
[0108] Then, at block 404, a linkage disequilibrium analysis is performed on the SNPs associated with the specific pharmacological phenotype to identify which SNPs are independent of each other. For example, when a group of SNPs are all associated with the same pharmacological phenotype and are in tight linkage disequilibrium (e.g., LD > 0.9), the group of SNPs may be linked, resulting in unclear which SNPs in the group are the SNPs that produce the association with the pharmacological phenotype. A linkage disequilibrium analysis can be performed to identify each SNP (effect SNP) that may produce the association with the pharmacological phenotype. More specifically, the linkage disequilibrium analysis can be performed by comparing the SNP (original SNP) with a database of SNPs (e.g., from the 1000 Genomes Project) to find the SNPs that are linked to the original SNP. In some embodiments, the ethnic group of the GWAS or candidate gene association study can be identified, and the SNPs and linkage disequilibrium coefficients can be retrieved from a database of SNPs corresponding to the identified ethnic group. Then, the pharmacological phenotype assessment server 102 can generate a set of permissive candidate variants of all SNPs that are in tight linkage disequilibrium with the original SNP (block 406), where the original SNP is associated with a specific pharmacological phenotype from a GWAS or other candidate association study.
[0109] In addition, the set of permissive candidate variants (box 406) can include somatic SNPs of genes (box 420) with known or suspected relevance to the pharmacologic phenotype under study, molecular QTLs targeting the gene body, and SNPs residing in genomic regions or networks with known or suspected relevance to the pharmacologic phenotype under study (box 422). In any case, the permissive candidate variants (box 406) can be subjected to bioinformatics analysis to filter the permissive candidate variants into a subset of intermediate candidate variants (box 410), which can then be ranked (e.g., by a scoring system) (box 412). Then, the SNPs, genes, and genomic regions with the highest rank in the subset of intermediate candidate variants that have known or suspected relevance to the pharmacologic phenotype under study (e.g., a rank higher than a threshold rank or a score higher than a threshold score) can be identified as SNPs, genes, and genomic regions causally related to a specific pharmacologic phenotype (box 414). For example, the SNPs, genes, and genomic regions may be related to drug response, adverse drug reactions, adverse drug events, disease risk, dose, complications, drug abuse, drug-gene interactions, drug-drug interactions, polypharmacy interactions, etc.
[0110] More specifically, to filter out a portion of the permissive candidate variants to produce a subset of intermediate candidate variants (box 410), the regulatory function of the genomic region surrounding the permissive candidate variant is evaluated (box 408a) to determine whether its sequence context (e.g., allele) affects the regulatory function (variant-dependence) (box 408b) and to identify its target gene (box 408c).
[0111] To evaluate whether a permissive candidate variant is functional, bioinformatics analysis can be used to determine whether the permissive candidate variant is located in open chromatin, as indicated by DNase I hypersensitivity. Figure 4A An exemplary illustration 450 of the bioinformatics analysis is depicted in.
[0112] Various machine learning techniques such as support vector machines (SVMs) can be used to determine variant-dependence. For example, an SVM can be used with permissive candidate variants to create a hyperplane for classifying k-mers, gap k-mers, or other local sequence features in a DNA sequence. The tendency of a specific allele of an SNP to produce a state change in a portion of the genome nearby can be measured using an SVM. This can indicate the level of importance of the SNP for a specific epigenomic profile (omics modality) in the tissue or cell line used to train the SVM. Additionally or alternatively, variant-dependence can be determined by identifying SNPs that alter transcription factor binding using a position weight matrix (PWM) or other algorithms for this purpose.
[0113] Various bioinformatics and machine learning techniques can also be used to identify target genes for permissive candidate variants. Quantitative trait locus (QTL) mapping can be used to identify target genes, thereby identifying associations between permissive candidate variants and the omics status of gene expression and / or genetic loci. Biological methods and datasets, as well as software analysis mapping systems for cis-eQTL, trans-eQTL, dsQTL, esQTL, hQTL, haQTL, eQTL, meQTL, pQTL, rQTL, etc., can be utilized. Permissive candidate variants can regulate gene expression on the same chromosome (cis-regulatory elements) or can regulate gene expression on another chromosome (trans-regulatory elements). However, the mapping system may have redundant sampling and incorrect relationships. Therefore, machine learning techniques are used to perform additive correction to fill in sparse data.
[0114] To determine whether a functional permissive candidate variant maintains regulatory control over nearby genes, bioinformatics analysis can determine whether the permissive candidate variant is hypomethylated, whether the permissive candidate variant is associated with histone marks indicating transcription start sites, and / or whether the permissive candidate variant enhances RNA, promoter RNA, or other RNA.
[0115] Methods for determining long-range interactions between permissive candidate variants and the genes they regulate can include Hi-C chromatin conformation capture, ChIA-PET, chromatin immunoprecipitation sequencing (ChIP-seq), and QTL analysis. Such methods can also be used to identify target genes of permissive candidate variants. Information can be combined with QTL data, and information density can be increased by using matrix densification methods or other various machine learning techniques to detect or simulate other contacts.
[0116] In any case, each permissive candidate variant can be scored and / or ranked based on its regulatory function (box 408a), variant dependence (box 408b), and target gene (box 408c) for a specific pharmacological phenotype. Permissive candidate variants with scores higher than a threshold score and / or ranks higher than a threshold rank or other score or rank criteria can be included in a subset of intermediate candidate variants (box 410).
[0117] Then, machine learning techniques are used to score and / or rank a subset of the intermediate candidate variants (box 410) relative to each other. For example, a bipartite graph analysis can be performed on the subset of intermediate candidate variants, where the intermediate candidate variants are represented by nodes in the graph and the relationships between two intermediate candidate variants are represented by edges. The intermediate candidate variants can be partitioned into disjoint sets, where any member within a disjoint set has no relationship with any other member within that set. In some embodiments, a particular weight can be assigned to the relative strength of a particular relationship between two intermediate candidate variants. Then, each intermediate candidate variant can be scored based on the number of relationships the intermediate candidate variant has with other intermediate candidate variants from other disjoint sets. In some embodiments, each intermediate candidate variant is scored based on the total weight assigned to each relationship between the intermediate candidate variant and other intermediate candidate variants.
[0118] In any case, then the top-ranked SNPs, genes, and genomic regions in the subset of intermediate candidate variants (e.g., ranked higher than a threshold rank or scored higher than a threshold score) can be identified as the SNPs, genes, and genomic regions associated with a particular pharmacologic phenotype (box 414). Then, when predicting whether the current patient has a particular pharmacologic phenotype, the identified SNPs, genes, and genomic regions for the particular pharmacologic phenotype can be analyzed for the current patient.
[0119] When performing methods 400 and 800 using server 102, and in other embodiments, sensitive, proprietary, or valuable data can be protected by using encryption and / or secure execution techniques and / or a remotely located computing device that is subject to additional security safeguards. Such data can include patient data that complies with HIPAA or other confidentiality and regulatory requirements, data restricted by patient or client privilege, proprietary data of a commercial entity, or other such data. Such data can be encrypted for transmission and decrypted for analysis, and such data can be analyzed in an anonymized or encrypted form by using mathematical transforms such as hash tables, elliptic curves, or other metrics. In such analysis, trusted execution techniques, trusted platform modules, and other similar techniques can be used. A data representation form that omits or obscures personal health information (PHI) can be used, particularly when preparing reports and diagnostic information and distributing reports and diagnostic information to healthcare practitioners.
[0120] Figure 4C An example gene regulatory network 470 or genomic region that includes genes and SNPs indicative of a pharmacologic phenotype is shown. The example gene regulatory network 470 can be identified using method 400 described above with reference to Figure 4B or can be identified based on a GWAS or candidate gene association study. The gene regulatory network 470 can be located within the central nervous system or within any other suitable system within the human body.
[0121] In any case, the gene regulatory network 470 includes genes BCDEF (reference number 472), DEFGH (reference number 474), ABCF (reference number 476), IJKLM (reference number 478), MNOP (reference number 480), LMNOP (reference number 482), PQRS (reference number 484), HIJKLM (reference number 486), XYZ (reference number 488), CDEFG (reference number 490), and ABCDEF (reference number 492).
[0122] The gene regulatory network 470 includes several non-coding SNPs located within introns, promoters, and intergenic regions, including non-coding SNPs related to transcription, which are significantly associated with the response of a specific patient cohort to drug X. For example, SNP2 found in the BCDEF gene 472 on chromosome 1 indicates drug X response and disease risk. In another example, SNP3, which has strong linkage disequilibrium (e.g., LD > 0.8) with SNP2, is also found in the BCDEF gene 472 on chromosome 1, and the SNP3 indicates an adverse drug reaction related to drug X. The gene regulatory network 470 also includes interchromosomal interactions, which provide a basis for subsets of trans enhancer-promoter and promoter-promoter spatial interactions. For example, SNP15 found in the enhancer region of the gene PQRS (reference number 484) on chromosome 1 that interacts with the gene HIJKLM (reference number 486) within chromosome 6 indicates an adverse drug reaction related to drug X. In some cases, one or more variants, genes, or enhancers in the gene regulatory network 470 may be located on sex chromosomes.
[0123] The interconnected genes (e.g., gene IJKLM (reference number 478) and gene LMNOP (reference number 482)) within the gene regulatory network 470 are depicted using single or double arrows Figure 4C Each connection may include numerical or categorical coefficients (e.g., P, C, V, T) that can be further described in legend 494. In some embodiments, the numerical or categorical coefficients indicate the relationship between the interconnected genes (e.g., activation, translocation, expression, inhibition, etc.).
[0124] The exemplary gene regulatory network 470 is merely one example of the omics data that can be obtained from GWAS, candidate gene association studies, and / or used to train patients to train the pharmacology phenotype prediction system 100. Additional gene regulatory networks can be obtained in conjunction with additional or alternative genomics data, epigenomics data, transcriptomics data, proteomics data, chromosomics data, or metabolomics data.
[0125] In addition to referring to Figure 4BBeyond the method 400 described, Figure 4D Another exemplary method 800 for using machine learning techniques to identify omics data corresponding to a specific pharmacological phenotype is shown. Method 800 can be executed on the pharmacological phenotype assessment server 102. In some embodiments, method 800 can be implemented in a set of instructions stored on a non-transitory computer-readable memory and executable on one or more processors on the pharmacological phenotype assessment server 102. For example, method 800 can be executed by Figure 1A the training module 160 within the machine learning engine 146 of
[0126] In method 800, permissive candidate variants are identified (block 810) in a manner similar to that in the Figure 4B method 400 described above. More specifically, at blocks 802 and 804, a statistical test (e.g., via GWAS or candidate gene association study) is performed for each of several non-coding or coding SNPs, genes, and genomic regions in the genome to determine the relationship between the SNP and a specific pharmacological phenotype, which can be drug response, adverse drug reaction, adverse drug event, dose, disease risk, etc. (e.g., response to warfarin). When the statistical test shows a significant relationship between the SNP and the specific pharmacological phenotype (e.g., p-value less than the threshold probability using the null hypothesis), the SNP is determined to be associated with the specific pharmacological phenotype. Then, at block 806, a linkage disequilibrium analysis is performed on the SNPs associated with the specific pharmacological phenotype and other SNPs to identify which SNPs are in linkage disequilibrium with the associated SNPs. Linkage disequilibrium analysis can be performed by comparing the SNP (the original SNP) with a database of SNPs (e.g., from the 1000 Genomes Project) to find the SNPs linked to the original SNP. In some embodiments, the ethnic group of the population used in the GWAS or candidate gene association study can be identified, and data from a matched population can be used to find SNPs with a significant linkage disequilibrium coefficient in the database of SNPs of the identified ethnic group. The physical SNPs of genes known or suspected to be relevant to the pharmacological phenotype under study can also be identified (block 808).
[0127] Then, a bioinformatics analysis can be performed on the permissive candidate variants (block 810) to filter the permissive candidate variants into a subset of intermediate candidate variants (block 814). Figure 4BMethod 400 filters permissive candidate variants based on their status as putative expression regulatory variants (e.g., according to regulatory function, the dependence of the regulatory function on variant alleles, and the presence of an identifiable target gene relationship). In method 800, permissive candidate variants are filtered based on expression regulatory variants (boxes 812a - 812c) or coding variants (box 812d) to produce a subset of intermediate candidate variants (814). To filter based on coding variants, method 800 determines whether a permissive candidate variant is a non - synonymous coding variant with a significant minor allele frequency (e.g., a minor allele frequency of at least 0.01).
[0128] More specifically, each permissive candidate variant can be scored and / or ranked based on its expression regulatory variants (e.g., according to the regulatory function for a specific pharmacological phenotype (box 812a), variant dependence (box 812b), and target gene (box 812c)). Permissive candidate variants with a score higher than a threshold expression regulatory variant score and / or a rank higher than a threshold expression regulatory variant rank or other score or rank criteria can be included in the subset of intermediate candidate variants (box 814). Additionally, each permissive candidate variant can be scored and / or ranked based on its coding variant (e.g., according to whether it is a non - synonymous coding variant with a significant minor allele frequency for a specific pharmacological phenotype). Permissive candidate variants with a score higher than a threshold coding variant score and / or a rank higher than a threshold coding variant rank or other score or rank criteria can also be included in the subset of intermediate candidate variants (box 814).
[0129] Then, at box 816, the intermediate candidate variants are associated with target genes, and pathway analysis is performed on genes expressed in relevant tissues (e.g., based on genotype - tissue expression (GTEx) data), such as using pathway analysis for pathway mapping and gene set enrichment. Gene sets associated with important and relevant pathways are identified, and regulatory variants and coding variants that affect the gene sets are identified as candidate variants (box 818).
[0130] Refer to Figure 4E describes Figure 4D an example application of the method in to a set of warfarin phenotypes. Warfarin is an anticoagulant used for the prevention and treatment of venous thromboembolism in heart disease and other situations where control of blood clotting is required. The dose requirements vary by up to 10 - fold between patients, and warfarin remains a commonly prescribed drug despite the availability of other anticoagulants recently. Thus, the above - described method can be used to predict a patient's response to warfarin and to determine whether to administer warfarin or other anticoagulants to a patient and the dosage.
[0131] Figure 4E shows a representation inFigure 4D A block diagram of single nucleotide polymorphisms (SNPs) identified in each stage of method 800 for identifying omics data corresponding to warfarin phenotypes.
[0132] To identify associations and candidate genes for a set of warfarin phenotypes, 23 genome-wide association studies (GWAS) were used in healthy patients for warfarin response and other pharmacologic phenotypes of warfarin, venous thromboembolism risk, and baseline anticoagulant protein levels. Input data from populations around the world were used, including European, East Asian, South Asian, African, and U.S. cohorts. In this example, warfarin phenotypes include several phenotype categories such as warfarin response, ADE, and disease / background. The warfarin response category includes the warfarin phenotypes: warfarin maintenance dose. The ADE category includes the warfarin phenotypes: hemostatic factors and hematologic phenotypes, hemorrhagic endpoint coagulation, and thrombin generation potential phenotypes. The disease / background category includes the warfarin phenotypes: venous thromboembolism, thromboembolism, thrombus, thrombosis, coagulation, bleeding, C4b-binding protein level, activated partial thromboplastin time, anticoagulant level, factor XI, prothrombin time, platelet thrombosis. Based on these 23 GWAS and 23 additional variants, a total of 204 SNPs were identified as association and candidate gene inputs (block 852).
[0133] Then, linkage disequilibrium analysis was performed on the 204 SNPs, and for a total of 4492 SNPs identified as permissive candidate variants, the physical SNPs of the 204 SNPs were also identified (block 854). Then, the expression regulatory variant workflow was applied to the 4492 SNPs, resulting in a total of 186 SNPs in 57 genes (block 856). As Figure 4D shown, the gene expression test of block 814 was applied to the 186 SNPs, resulting in a total of 66 SNPs in 30 genes. In addition, the coding variant workflow was also applied to the 4492 SNPs, resulting in a total of 37 SNPs with a minor allele frequency of at least 0.01 (block 858). As Figure 4D shown, the gene expression test of block 814 was also applied to the 37 SNPs, resulting in a total of 22 SNPs in 17 genes. Thus, the combined output of the expression regulatory variant workflow and the coding variant workflow is 87 SNPs in 41 genes (block 860). Finally, pathway analysis was performed on the 87 SNPs to identify a single pathway with 74 SNPs in 31 genes (block 862).
[0134] The pathway can be referred to as the warfarin response pathway and includes genes expressed in the liver, small intestine, and vasculature. Figure 4F An example warfarin response pathway 870 including genes and SNPs indicative of warfarin phenotypes is shown. It can be used as described above with reference to Figure 4DThe described method 800 is used to identify the exemplary warfarin response pathway 870. In any case, the warfarin response pathway includes the following genes: Aldo-keto reductase family 1 member C3 (AKR1C3), Cytochrome P450 family 2 subfamily C member 19 (CYP2C19), Cytochrome P450 family 2 subfamily C member 8 (CYP2C8), Cytochrome P450 family 2 subfamily C member 9 (CYP2C9), Cytochrome P450 family 4 subfamily F member 2 (CYP4F2), Coagulation factor V (F5), Coagulation factor VII (F7), Coagulation factor X (F10), Coagulation factor XI (F11), Fibrinogen gamma chain (FGG), Orosomucoid 1 (ORM1), Serine protease 53 (PRSS53), Vitamin K epoxide reductase complex subunit 1 (VKORC1), Syntaxin 4 (STX4), Coagulation factor XIII A chain (F13A1), Protein C receptor (PROCR), Von Willebrand factor (VWF), Complement factor H-related 5 (CFHR5), Fibrinogen alpha chain (FGA), Flavin-containing monooxygenase 5 (FMO5), Histidine-rich glycoprotein (HRG), Kininogen 1 (KNG1), Surf-4 (SURF4), Alpha-1,3-N-acetylgalactosaminyltransferase and alpha-1,3-galactosyltransferase (ABO), Lysozyme (LYZ), Polycomb group ring finger 3 (PCGF3), Serine protease 8 (PRSS8), Transient receptor potential cation channel subfamily C member 4 associated pattern (TRPC4AP), Solute carrier family 44 member 2 (SLC44A2), Sphingosine kinase 1 (SPHK1), and Ubiquitin-specific peptidase 7 (USP7).
[0135] The 74 SNPs (not shown) included in 31 genes are: rs12775913 (regulatory SNP), rs346803 (regulatory SNP), rs346797 (regulatory SNP), rs762635 (regulatory SNP), and rs76896860 (regulatory SNP) included in the AKR1C3 gene (expressed in the liver); rs3758581 (coding SNP) included in the CYP2C19 gene (expressed in the liver); rs10509681 (coding SNP) and rs11572080 (coding SNP) included in the CYP2C8 gene (expressed in the liver); rs1057910 (coding SNP), rs1799853 (coding SNP), and rs7900194 (coding SNP) included in the CYP2C9 gene (expressed in the liver); rs2108622 (coding SNP) included in the CYP4F2 gene (expressed in the liver); rs6009 (regulatory SNP), rs11441998 (regulatory SNP), rs2026045 (regulatory SNP), rs34580812 (regulatory SNP), rs749767 (regulatory SNP), rs9378928 (regulatory SNP), and rs7937890 (regulatory SNP) included in the F5 gene (expressed in the liver); rs7552487 (regulatory SNP), rs6681619 (regulatory SNP), rs8102532 (regulatory SNP), rs491098 (coding SNP), and rs6046 (coding SNP) included in the F7 gene (expressed in the liver); rs11150596 (regulatory SNP) and rs11150596 (regulatory SNP) included in the F10 gene (expressed in the liver); rs2165743 (regulatory SNP) and rs11252944 (regulatory SNP) included in the F11 gene (expressed in the liver); rs8050894 (regulatory SNP) included in the FGG gene (expressed in the liver); rs10982156 (regulatory SNP) included in the ORM1 gene; rs7199949 (coding SNP) included in the PRSS53 gene (expressed in the liver); rs2884737 (regulatory SNP), rs9934438 (regulatory SNP), rs897984 (regulatory SNP), and rs17708472 (regulatory SNP) included in the VKORC1 gene (expressed in the liver); rs35675346 (regulatory SNP) and rs33988698 (regulatory SNP) included in the STX4 gene (expressed in the small intestine); rs5985 (coding SNP) included in the F13A1 gene (expressed in the vasculature); rs867186 (coding SNP) included in the PROCR gene (expressed in the vasculature);rs75648520 (regulatory SNP), rs55734215 (regulatory SNP), rs12244584 (regulatory SNP), and rs1063856 (coding SNP) contained in the VWF gene (expressed in the vasculature); rs674302 (regulatory SNP) contained in the CFHR5 gene (expressed in the liver); rs12928852 (regulatory SNP) and rs6050 (coding SNP) contained in the FGA gene (expressed in the liver); rs8060857 (regulatory SNP) and rs7475662 (regulatory SNP) contained in the FMO5 gene (expressed in the liver); rs9898 (coding SNP) contained in the HRG gene (expressed in the liver); rs710446 (coding SNP) contained in the KNG1 gene (expressed in the liver); rs11577661 (regulatory SNP) contained in the SURF4 gene (expressed in the liver); rs11427024 (regulatory SNP), rs6684766 (regulatory SNP), rs2303222 (regulatory SNP), rs1088838 (regulatory SNP), rs13130318 (regulatory SNP), and rs12951513 (regulatory SNP) contained in the ABO gene (expressed in the small intestine); rs8118005 (regulatory SNP) contained in the LYZ gene (expressed in the small intestine); rs76649221 (regulatory SNP), rs9332511 (regulatory SNP), and rs6588133 (regulatory SNP) contained in the PCGF3 gene (expressed in the small intestine); rs11281612 (regulatory SNP) contained in the PRSS8 gene (expressed in the small intestine); rs11589005 (regulatory SNP), rs8062719 (regulatory SNP), rs889555 (regulatory SNP), rs36101491 (regulatory SNP), rs7426380 (regulatory SNP), rs6579208 (regulatory SNP), rs77420750 (regulatory SNP), and rs73905041 (coding SNP) contained in the TRPC4AP gene (expressed in the small intestine); rs3211770 (regulatory SNP), rs3211770 (regulatory SNP), rs3087969 (coding SNP), and rs2288904 (coding SNP) contained in the SLC44A2 gene (expressed in the vasculature); rs683790 (regulatory SNP) and rs346803 (coding SNP) contained in the SPHK1 gene (expressed in the vasculature); and rs201033241 (coding SNP) contained in the USP7 gene (expressed in the vasculature).;
[0136] Except for warfarin, Figure 4B and4D The method described can also be applied to a set of lithium phenotypes as well as any other pharmacological phenotypes. By Figure 4B and 4DThe method described in [reference] can be applied to the lithium phenotype to identify the lithium response pathways of 78 SNPs in 12 genes. The lithium response pathways include the following genes: ankyrin 3 (ANK3), aryl hydrocarbon nuclear translocator-like (ARNTL), voltage-gated calcium channel auxiliary subunit gamma 2 (CACNG2), voltage-gated calcium channel auxiliary subunit alpha 1 C (CACNA1C), cyclin-dependent kinase inhibitor 1A (CDKN1A), cAMP response element-binding protein 1 (CREB1), AMPA-type glutamate receptor subunit 1 (GRIA2), glycogen synthase kinase 3 beta (GSK3B), nuclear receptor subfamily 1, group D, member 1 (NR1D1), solute carrier family 1 member 2 (SLC1A2), serotonin receptor 1A (HTR1A), and TRAF2 and NCK interacting kinase (TNIK).The 78 SNPs contained in 12 genes are as follows: rs2185502, rs10821792, rs1938540, rs3808943, rs61847646, rs75314561, rs61846516, rs10994397, rs10994318, rs61847579, rs12412727, rs10994308, rs4948418, rs4948412, rs4948413, rs4948416, rs10821745, rs10994336, rs10994360, rs9633532, rs1938526, rs10994322 and rs10994321 contained in the ANK3 gene; rs10766075, rs7938308, rs10832017, rs4603287, rs7934154, rs12361893, rs4414197, rs4757140, rs4757141, rs61882122, rs11022755, rs11022754, rs1481892, rs1481891, rs4353253, rs4756764, rs2403662, rs4237700, rs10832018, rs12290622, rs7928655, rs34148132, rs4146388, rs4146387, rs7949336, rs4757139, rs7107287 and rs1351525 contained in the ARNTL gene; rs2284017 and rs2284016 contained in the CACNG2 gene; rs2007044 and rs1016388 contained in the CACNA1C gene; rs3176336, rs3176333, rs3176334, rs3176320, rs4135240, rs2395655 and rs733590 contained in the CDKN1A gene; rs10932201 contained in the CREB1 gene; rs78957301 contained in the GRIA2 gene; rs334558 contained in the GSK3B gene; rs2314339 contained in the NR1D1 gene; rs3794088, rs3794087, rs4354668, rs12418812, rs1923294, rs5791047, rs111885243, rs752949 and rs16927292 contained in the SLC1A2 gene; rs6449693 and rs878567 contained in the HTR1A gene; and rs7372276 contained in the TNIK gene.
[0137] InFigure 4G The lithium reaction pathway 890 is depicted in FIG. 890. Lithium is a psychotherapeutic drug used to treat mental illness / disorders. The above method can be used to predict a patient's response to lithium and determine whether to administer lithium or other psychotherapeutic drugs to the patient and the dosage to be administered. In any case, the method can be used to predict a patient's response to lithium or other psychotherapeutic drugs. Figure 4B and 4D Methods 400, 800 described in are used to identify an example lithium response pathway 890. Each gene in the lithium response pathway 890 is expressed in a portion of the brain including the frontal lobe, insula, temporal cortex, cingulate cortex, amygdala, hippocampus, anterior caudate nucleus, thalamus, motor cortex, fusiform cortex, substantia nigra, cerebellum, and hypothalamus.
[0138] In some embodiments, the pharmacological phenotype prediction system 100 can test the presence of the identified SNPs, genes, and genomic regions in the current patient to determine whether the current patient has a specific pharmacological phenotype. For example, the identified SNPs, genes, and genomic regions may indicate a negative response to valproic acid used to treat TBI. When the current patient suffers from TBI, a biological sample of the current patient can be provided and, for example, the above description of the pharmacological phenotype can be used to predict the pharmacological phenotype. Figure 5 The process 500 for generating omics data is used to analyze the biological sample to detect whether there are identified SNPs, genes and genomic regions. When the current patient has at least some identified SNPs, genes and genomic regions indicating a negative reaction to valproic acid, valproic acid is not administered to the current patient. In other embodiments, the identified SNPs, genes and genomic regions are scored, combined and / or weighted in any suitable manner to determine which combinations indicate a negative reaction to valproic acid. The scoring or weighting system is then applied to the SNPs, genes and genomic regions in the current patient's biological sample to determine whether the current patient has a combination indicating a negative reaction to valproic acid.
[0139] In any case, the pharmacologic phenotype prediction system 100 can provide identified SNPs, genes, and genomic regions indicative of a particular pharmacologic phenotype as omics data for the pharmacologic phenotype. The machine learning engine 146 can obtain omics data with socio-omics, physio-omics, and environmental data for patients with at least some of the identified SNPs, genes, and genomic regions, as well as the phenomics data of the training patients. In this way, the machine learning engine 146 can classify training patients with identified SNPs, genes, and genomic regions indicative of a particular pharmacologic phenotype (e.g., a negative response to ketamine for treating depression) as having or not having the particular pharmacologic phenotype. Then, the socio-omics, physio-omics, and environmental data can be used to distinguish training patients with the identified SNPs, genes, and genomic regions who do have the particular pharmacologic phenotype from those who do not have the particular pharmacologic phenotype.
[0140] For example, when the machine learning technique is a decision tree, a decision tree containing multiple nodes can be generated, with each node representing a test on the data of the current patient. The nodes can be connected by branches, with each branch representing the result of a test or other measurement or an observable / recordable state (e.g., "yes" branch and "no" branch), where the branches can be weighted and the leaf nodes can indicate whether the current patient has the pharmacologic phenotype. In other embodiments, the leaf nodes indicate, for example, the likelihood of the pharmacologic phenotype determined by summing or combining the weighted branches, or the leaf nodes can indicate a score that can be compared with a threshold to determine whether the current patient has the pharmacologic phenotype. In any case, a decision tree can be generated, with the nodes near the top of the decision tree representing tests on the omics data of the current patient, e.g., as indicated by the identified SNPs, genes, and genomic regions. When the current patient has a suitable combination of identified SNPs, genes, and genomic regions indicative of a particular pharmacologic phenotype (e.g., a negative response to ketamine for treating depression), the decision tree branches to several nodes that represent tests on the socio-omics, physio-omics, and environmental data of the current patient.
[0141] In another example, when the machine learning technique is SVM, for each training patient, the pharmacologic phenotype assessment server 102 obtains the social omics, physiologic omics, and environmental data, the identified SNPs, genes, and genomic regions of the training patient that indicate a specific pharmacologic phenotype, and an indication that indicates whether the training patient has a specific pharmacologic phenotype (e.g., a negative response to ketamine used to treat depression) as a training vector. The SVM obtains each training vector and creates a statistical model for determining whether a current patient has a specific pharmacologic phenotype by generating a hyperplane that separates a first subset of training vectors corresponding to training patients with the pharmacologic phenotype from a second subset of training vectors corresponding to training patients without the pharmacologic phenotype.
[0142] Environmental, physiologic omics, and social omics data can be obtained for a training patient or a cohort of training patients having identified SNPs, genes, and genomic regions that indicate a specific pharmacologic phenotype. More specifically, the training module 160 can obtain, for example, a set of training data from the client devices 106-116 and / or one or several servers (e.g., an EMR server, a polypharmacy server, etc.), the training data can include the omics data of several training patients as well as social omics, physiologic omics, and environmental data, where the pharmacologic phenotype of the training patients is known (e.g., previously determined or currently determined) and is also provided in the training data. The environmental, physiologic omics, and social omics data can include clinical data, demographic data, polypharmacy data, socioeconomic data, educational data, drug abuse data, diet and exercise data, law enforcement data, circadian rhythm data, family data, or any other suitable data that indicates the social status or environmental conditions of the patient.
[0143] In an exemplary scenario, omics, phenomics, socialomics, physiomics, and environmental data of training patients are collected over a first time period (e.g., one year). Although the patients can be training patients, the results of the pharmacologic phenotype prediction system 100 can also be determined based on the omics, phenomics, socialomics, physiomics, and environmental data of the patients. As described above, a training patient having training data for training the pharmacologic phenotype prediction system 100 can also become a current patient for predicting an unknown pharmacologic phenotype of the training patient. In this example, omics, phenomics, socialomics, physiomics, and environmental data can be collected during January to December of the first year. Although the omics, phenomics, socialomics, physiomics, and environmental data can be for a single training patient, omics, phenomics, socialomics, physiomics, and environmental data of a cohort of training patients can also be collected. For example, omics, phenomics, socialomics, physiomics, and environmental data of a cohort of training patients can be collected, where each patient has identified SNPs, genes, and genomic regions indicative of a negative response to drug X used to treat disease Y.
[0144] The omics, phenomics, socialomics, physiomics, and environmental data can include Figure 2 the individual / cohort and population omics and pharmacometabolomics 302, exposome 304, socialomics demographics and stress / trauma 306, and medical physiomics, EHR, laboratory values, stress and abuse factors and trauma, and medical outcome data 308 shown in. However, these are just a few examples of the omics, phenomics, socialomics, physiomics, and environmental data that can be obtained from training patients to train the pharmacologic phenotype assessment server 102. Additional socialomics, physiomics, and environmental data can also be included, such as circadian rhythm data indicating the sleep and other recurrent lifestyle temporal patterns of the training patients.
[0145] The omics and pharmacometabolomics data can be the result of pharmacogenomic analysis of a biological sample of a training patient. The exposome data for the first year can include the employment status and place of residence of the training patient in August of the first year.
[0146] The medical physiomics, EHR, laboratory values, stress and abuse factors and trauma, and medical outcome data can indicate the law enforcement experience of the training patient from January to December of the first year.
[0147] Medical physiomics, EHR, laboratory values, stress and abuse factors and trauma, and medical outcome data can also include polypharmacy data, which includes the drugs prescribed to a training patient in the form of prescriptions and the times at which the drugs were prescribed. For example, in March of Year 1, Drug A was prescribed to the training patient to treat Disease 1, Drug B to treat Disease 2, and Drug C to treat Disease 3, and in August of Year 1, Drug D was prescribed to treat Disease 4. In August of Year 1, the training patient also tapered off Drug C. Additionally, medical physiomics, EHR, laboratory values, stress and abuse factors and trauma, and medical outcome data can also include phenomics data indicating the pharmacological phenotype of the training patient, such as the diseases the training patient was diagnosed with. For example, in January of Year 1, the training patient was diagnosed with Diseases 1-3.
[0148] In addition, the data can include other phenomics data, such as information describing drug efficacy and / or side effects. For example, the training patient may have experienced side effects of Drug C, and thus, after experiencing these side effects, the training patient may have tapered off Drug C. The training patient may also have had a positive response to taking Drug C, which can indicate the efficacy of the drug on the training patient.
[0149] Sociomics and demographic data can indicate the family status of the training patient. For example, sociomics and demographic data can indicate the marital status and number of children of the training patient in January of Year 1. Sociomics and demographic data can also indicate the amount of income of the training patient from January to December of Year 1. Sociomics and demographic data can also indicate available information about family members who have or have not been affected by relevant diseases and complications, as well as their treatment responses and other pharmacological phenotypes.
[0150] In this exemplary scenario, omics, phenomics, sociomics, physiomics, and environmental data can also include data collected from the training patient during January to December of Year 2. In Year 2, the omics and drug metabolomics data include the proteomics and transcriptomics of the training patient obtained in March and the results of pharmacogenomics analysis obtained in August.
[0151] The omics and drug metabolomics data for Year 2 can include the results of an inpatient metabolic examination of the training patient in March of Year 2, which indicates the presence of a toxic metabolite of Drug B. The omics and drug metabolomics data can also include the results of another inpatient metabolic examination of the training patient in August of Year 2, which indicates normal blood levels of Drug E.
[0152] Medical physiomics, EHR, laboratory values, stress and abuse factors and trauma, and medical outcome data for the second year indicate the results of various mental health, substance abuse, and stress and trauma questionnaires administered to the training patient.
[0153] Medical physiomics, EHR, laboratory values, stress and abuse factors and trauma, and medical outcome data for the second year also indicate that the training patient had a negative reaction to Drug C in February and subsequently stopped taking Drug C. Conversely, Drug F was prescribed to the patient in March of the second year and Drug E was prescribed to the patient in August. The training patient then had a positive reaction to taking Drugs E and F, which can indicate the efficacy of said drugs on the training patient.
[0154] In addition, in response to the detection of the presence of a toxic metabolite of Drug B (as indicated by the omics and pharmacometabolomics data of the training patient), the training patient stopped taking Drug B in March of the second year. Medical physiomics, EHR, laboratory values, stress and abuse factors and trauma, and medical outcome data further indicate that the diagnosis of Disease 1 in the training patient was controlled in April of the second year and the dosage of Drug A was reduced. The social omics and demographic data for the second year also indicate the income amount of the training patient from January to December of the second year.
[0155] The information contained in the omics, phenomics, social omics, physiomics, and environmental data for the first and second years can be used to train the pharmacology phenotype assessment server 102 to generate a statistical model. Similar information can be collected for some training patients (e.g., dozens, hundreds, thousands) comprising a training patient cohort or population, and the omics data can be combined to generate a statistical model. In the generation of such a statistical model, important features can be identified and a model using such a restricted information set can be generated to allow prediction of the phenotype of a current patient using the limited information set.
[0156] For example, as described above, omics, phenomics, socialomics, physiomics, and environmental data for a cohort of training patients can be obtained, where each training patient has identified SNPs, genes, and genomic regions indicative of a negative response to drug X for treating disease Y. The training module 160 can classify a first subset of the omics, phenomics, socialomics, physiomics, and environmental data (e.g., the identified SNPs, genes, and genomic regions of a cohort of training patients indicative of a negative response to drug X for treating disease Y) as corresponding to training patients with a negative response to drug X, and can classify a second subset of the omics, phenomics, socialomics, physiomics, and environmental data as corresponding to training patients without a negative response to drug X. In some embodiments, the training module 160 can perform statistical measurements on the subsets of the omics, phenomics, socialomics, physiomics, and environmental data for each classification (training patients with a negative response to drug X and training patients without a negative response to drug X). For example, the training module 160 can determine the average income, average ACE score, etc., of the training patients corresponding to each classification.
[0157] Then, the training module 160 can generate a statistical model (e.g., a decision tree, neural network, hyperplane, linear or non-linear regression coefficients, etc.) to predict whether a current patient has a negative response to drug X for treating disease Y based on the statistical measurements for each classification. For example, the statistical model can be a decision tree with several nodes connected by branches, where each node represents a test on the omics data related to the identified SNPs, genes, and genomic regions indicative of a negative response to drug X for treating disease Y. The branches can contain weights or scores for different SNPs, genes, and genomic regions, and when the combined weight or score of the current patient exceeds a threshold, this can indicate that the current patient has a combination of SNPs, genes, and genomic regions indicative of a negative response to drug X.
[0158] A decision tree can further include several nodes connected by branches, where each node represents a test on social omics, physiomics, and environmental data. The first node can test whether the annual income of the current patient is greater than $20,000, which is connected by a "yes" branch to a second node that tests whether the ACE score of the current patient is greater than 5, which is connected by a "yes" branch to a third node that tests whether the current patient has experienced domestic violence, which is connected by a "no" branch to a leaf node that predicts whether the current patient will have a negative reaction to drug X. The leaf branch can indicate the likelihood that the current patient will have a negative reaction to drug X, which can be determined based on the results of each test at the respective nodes and / or the weights assigned to the corresponding branches. In some embodiments, the leaf branch can indicate a reaction score to drug X, which can also be determined based on the results of each test at the respective nodes and / or the weights assigned to the corresponding branches. The reaction score can indicate the efficacy of the drug in treating the corresponding disease, which is discounted due to the adverse effects of the drug on the patient.
[0159] As described in this exemplary scenario, the omics, social omics, physiomics, and environmental data of each of several training patients can be combined with an indication of the pharmacological phenotype of the training patient (e.g., the disease the training patient was diagnosed with, the reaction to the drug, drug abuse problems) to generate a statistical model for predicting the pharmacological phenotype of the current patient. In some embodiments, the omics, social omics, physiomics, and environmental data of the training patients can be combined with omics and / or genomics data from previous studies (e.g., GWAS) to identify relationships between genes, transcription factors, proteins, metabolites, chromatin states, the environment, and the pharmacological phenotype of the patient.
[0160] In any case, the omics, social omics, physiomics, and environmental data of each training patient are combined to generate a statistical model. In some embodiments, the omics, social omics, physiomics, and environmental data of the training patients can be classified as corresponding to training patients diagnosed with a specific disease (or not diagnosed with a specific disease), classified as corresponding to training patients with a substance abuse problem (or without a substance abuse problem), classified as corresponding to training patients with a special response to a drug, or classified in any other suitable manner. In some cases, the omics, social omics, physiomics, and environmental data of the same training patient can be subdivided into multiple pharmacologic phenotypes at different time periods. For example, as the social omics, physiomics, and environmental data of the training patient change, the omics data of the training patient may change over time, and at a first time period, the training patient may have one mental illness, while at a second time period, the training patient may have another mental illness.
[0161] Similarly, in some embodiments, the omics, social omics, physiomics, and environmental data of the training patients can be classified according to demographics. For example, the omics, social omics, physiomics, and environmental data of training patients of European descent can be assigned to one cohort, while the omics, social omics, physiomics, and environmental data of training patients of Chinese descent can be assigned to another cohort. In another example, the omics, social omics, physiomics, and environmental data of training patients between 25 and 35 years old can be assigned to one cohort, while the omics, social omics, physiomics, and environmental data of training patients between 35 and 45 years old can be assigned to another cohort.
[0162] In such an embodiment, the training module 160 can generate different statistical models for each pharmacologic phenotype and / or for each cohort (e.g., a cohort separated based on demographics). For example, a first statistical model can be generated to determine the likelihood that a current patient will experience a substance abuse problem, a second statistical model can be generated to determine the risk of one disease, a third statistical model can be generated to determine the risk of another disease, a fourth statistical model can be generated to determine the likelihood of a negative response to a specific drug, etc. In other embodiments, the training module 160 can generate a single statistical model to determine the likelihood that a current patient has any pharmacologic phenotype, or can generate any number of statistical models to determine the likelihood that a current patient has any number of pharmacologic phenotypes.
[0163] Once the omics, socialomics, physiomics, and environmental data are classified into subsets corresponding to respective cohorts and / or pharmacological phenotypes, the omics, socialomics, physiomics, and environmental data for a particular pharmacological phenotype can be analyzed to generate a statistical model. A neural network, deep learning, decision tree, support vector machine, or any of the above machine learning methods can be used to generate the statistical model. For example, the omics, socialomics, physiomics, and environmental data of training patients of European descent can be analyzed to determine that there is a high correlation between training patients with SNP 2 found in the XYZ gene, being unemployed and having a record of violent crime, and developing schizophrenia. Thus, European descent patients with SNP 2 found in the XYZ gene, being unemployed and having a record of violent crime may be at high risk of developing schizophrenia.
[0164] When the machine learning technique is a neural network or deep learning, the training module 160 can generate a graph having input nodes, intermediate or "hidden" nodes, edges, and output nodes. The nodes can represent tests or functions performed on the omics, socialomics, physiomics, and environmental data, and the edges can represent the connections between the nodes. In some embodiments, the output node can contain an indication of the pharmacological phenotype or pharmacological phenotype likelihood. In some embodiments, the edges can be weighted based on the strength of the tests or functions of the previous nodes when determining the pharmacological phenotype.
[0165] Thus, for pharmacological phenotype determination, the types of omics, socialomics, physiomics, and environmental data of previous nodes with higher weights may be more important than the types of omics, socialomics, physiomics, and environmental data of previous nodes with lower weights. By identifying the most important omics, socialomics, physiomics, and environmental data, the training module 160 can eliminate the least important and potentially misleading and / or random noise omics, socialomics, physiomics, and environmental data from the statistical model. Additionally, the most important omics data for the pharmacological phenotype (determined by ranking, weighting, or scoring omics data above a threshold) can be used to select the types of omics data to analyze a patient's biological sample.
[0166] For example, a neural network can include four input nodes representing omics, social omics, physiomics, and environmental data, each of which is connected to several hidden nodes. The hidden nodes are then connected to an output node that indicates the likelihood that the current patient has bipolar disorder. The connections can have assigned weights, and the hidden nodes can include tests or functions performed on the omics, social omics, physiomics, and environmental data. In some embodiments, the test or function can be a distribution determined by training data or prior studies such as GWAS. For example, a patient with a specific SNP and in the 98th percentile of the income distribution is less likely to contract disease Y than a patient with the same SNP and in the 10th percentile of the income distribution.
[0167] In some embodiments, the hidden nodes can be connected to several output nodes, each of which indicates the likelihood that the current patient will develop a different disease, the likelihood that the current patient will have a substance abuse problem, and / or the likelihood or response score of the current patient to a specific drug. In this example, the four input nodes can include the patient's ancestry, the patient's current income, and / or the change in the patient's income in the previous year, the presence or absence of SNP 13 in the patient's LMNOP gene, and the patient's poor sleep pattern.
[0168] In some embodiments, each of the four input nodes can assign numerical values to the patient's omics, social omics, physiomics, and environmental data, and a test or function can be applied to the numerical values at the hidden nodes. Then, for example, the results of the test or function can be weighted and / or aggregated to determine a response score to lithium. The response score can indicate the efficacy of the drug in treating the corresponding disease, which is discounted due to the adverse effects of the drug on the patient. In this example, the response score for lithium may be high (e.g., 80 out of 100), indicating that the patient will respond positively when taking lithium to treat bipolar disorder. Therefore, a healthcare provider can prescribe lithium to treat the patient's manic-depressive illness. In some embodiments, the response score for lithium used to treat manic-depressive illness can be compared with the response scores of other prescription drugs used to treat manic-depressive illness. Then, the drugs can be ranked according to their respective response scores, and the drug with the highest rank can be recommended to the healthcare provider for prescription to the patient. However, this is just one example of the input and result output of a statistical model for determining phenotypes. In other examples, any number of input nodes can include several types of omics data, social omics data, and environmental data of the patient. Additionally, any number of output nodes can determine the likelihood or risk of developing different diseases, the likelihood of substance abuse problems, the likelihood of complications, etc.
[0169] Since additional training data are collected, weights, nodes, and / or connections can be adjusted. In this way, the statistical model is updated continuously or periodically to reflect at least an almost real-time representation of the social omics, physiomics, environmental, and omics data.
[0170] In some embodiments, machine learning techniques can be used to identify cohort or demographic markers to classify patients as having or not having a specific pharmacologic phenotype. As shown in the above examples, tests or functions included in hidden nodes can be developed by performing statistical measurements on the social omics, physiomics, and environmental data of a first subset and a second subset of training patients having and not having a specific pharmacologic phenotype, respectively. Statistical measurements can be used to determine the most important variables included in the social omics, physiomics, and environmental data of the training patients to distinguish the training patients having the pharmacologic phenotype from the training patients not having the pharmacologic phenotype. In this way, the tests or functions included in the hidden nodes are not necessarily a priori.
[0171] After generating a statistical model using machine learning techniques as described above (e.g., neural networks, deep learning, decision trees, support vector machines, etc.), the training module 160 can test the statistical model using test omics, social omics, physiomics, and environmental data from a test patient and the test patient's pharmacologic phenotype. The test patient can be a patient whose pharmacologic phenotype is known. However, for testing purposes, the training module 160 can determine the likelihood that the test patient has various pharmacologic phenotypes by comparing the test omics, social omics, physiomics, and environmental data of the test patient with the statistical model generated using machine learning techniques.
[0172] For example, the training module 160 can traverse the nodes from a neural network using the test omics, social omics, physiomics, and environmental data of the test patient. When the training module 160 reaches a result node indicating the likelihood or response score of a specific pharmacologic phenotype, the likelihood or response score can be compared with the known pharmacologic phenotype of the test patient.
[0173] In some embodiments, if the likelihood that the test patient has a specific pharmacologic phenotype (e.g., disease Y) is greater than 0.5 and the known pharmacologic phenotype of the test patient is that they do in fact have disease Y, then the determination can be considered correct. In another example, if the response score of a drug that elicits a strong response in the test patient without harmful side effects is the highest, then the determination can be considered correct. In other embodiments, when the known pharmacologic phenotype of the test patient is that they have a specific pharmacologic phenotype, the likelihood may have to be higher than 0.7, or some other predetermined threshold of the determined likelihood must be considered correct.
[0174] In addition, in some embodiments, when the accuracy rate of the training module 160 exceeds a predetermined threshold amount of time, the statistical model can be presented to the phenotypic evaluation module 162. On the other hand, if the accuracy rate of the training module 160 does not exceed the threshold amount, the training module 160 can continue to obtain a training data set for further training.
[0175] Once the statistical model has been sufficiently tested to verify its accuracy, the phenotypic evaluation module 162 can obtain the statistical model. Based on the statistical model, the phenotypic evaluation module 162 can determine the likelihood that the current patient has various pharmacological phenotypes. The phenotypic evaluation module 162 can obtain omics, socialomics, physiomics, and environmental data of the current patient without knowing whether the current patient has certain pharmacological phenotypes. Socialomics, physiomics, and environmental data can be collected at several time points, and the data can be similar to the socialomics and environmental data collected for the training patients described in the above exemplary scenarios.
[0176] More specifically, socialomics, physiomics, and environmental data can include clinical data, such as the medical history of the current patient, including the diseases that the current patient has been diagnosed with, the results of laboratory tests, and the procedures performed on the patient, the family history of the current patient, etc. Socialomics, physiomics, and environmental data can also include polypharmacy data, such as each drug prescribed to the current patient in the form of a prescription within a specific time period, the duration of each prescription, the number of refills, whether the current patient has refilled each drug on time, etc. In addition, socialomics, physiomics, and environmental data can include demographic data, such as the ethnicity or race of the current patient, the age of the current patient, the weight of the current patient, the gender of the current patient, the place of residence of the current patient, etc. Furthermore, socialomics, physiomics, and environmental data can include the socioeconomic data of the current patient, such as the amount and / or source of income of the current patient, educational data (e.g., high school diploma, GED, college graduate, master's degree, etc.), and diet and exercise data indicating the frequency of exercise of the current patient, the eating habits of the current patient, the amount of weight gain or weight loss within a specific time period, etc. In addition, socialomics, physiomics, and environmental data can include family data, such as the marital status of the current patient, the number of children and family members currently living with the current patient, law enforcement data indicating the criminal record of the current patient and whether the current patient has ever been a victim of abuse or other crimes, drug abuse data indicating whether the current patient has or has ever had a drug abuse problem, and circadian rhythm data indicating the sleep pattern of the current patient.
[0177] In addition to socialomics, physiomics, and environmental data, the phenotypic evaluation module 162 can also obtain omics data of the current patient. The omics data can be similar to Figure 4CThe omics data shown. More specifically, the omics data may include genomics data indicating gene traits, epigenomics data indicating gene expression, transcriptomics data indicating DNA transcription, proteomics data indicating proteins expressed by the genome, chromosomics data indicating the chromatin state in the genome, and / or metabolomics data indicating metabolites in the genome.
[0178] Then, the phenotype assessment module 162 may apply the omics, social omics, physiomics, and environmental data of the current patient to a statistical model to determine the likelihood that the current patient has various pharmacological phenotypes. When several statistical models are generated, the phenotype assessment module 162 may apply the omics, social omics, physiomics, and environmental data of the current patient to each of the statistical models to determine, for example, the likelihood that the current patient has disease Y, has a drug abuse problem, and develops complications.
[0179] In some embodiments, a healthcare provider may provide a request for a specific type of pharmacological phenotype (such as the predicted response of the current patient to a specific drug) or a request for the best drug for treating a specific disease. Accordingly, the phenotype assessment module 162 may apply a statistical model or a portion of a statistical model generated in response to the request of the healthcare provider. Then, the healthcare provider may receive an indication of the best drug and dosage for treating a specific disease from the pharmacological phenotype assessment server 102.
[0180] In other embodiments, the pharmacological phenotype assessment server 102 may evaluate the omics, social omics, physiomics, and environmental data by applying the omics, social omics, physiomics, and environmental data of the current patient to a statistical model to determine the likelihood or response score of each of several pharmacological phenotypes. Then, the pharmacological phenotype assessment server 102 may generate a risk analysis display for review by the healthcare provider of the current patient.
[0181] Risk analysis may display indicators that can include patient biographical information, such as the patient's name, date of birth, address, etc. Risk analysis may also display an indication for each of the likelihoods of various pharmacologic phenotypes or other semi - quantitative and quantitative metrics, which likelihoods or other semi - quantitative and quantitative metrics may be expressed as a probability (e.g., 0.6), a percentage (e.g., 80%), a category in a set of categories (e.g., "high risk", "medium risk", or "low risk"), and / or in any other suitable manner. Additionally, risk analysis may display an indication of a response score for a drug, which response score may be numerical (e.g., 75 out of 100), categorical (e.g., "strong response", "adverse reaction", etc.), or in any other suitable manner. Further, a description of each likelihood or response score and the corresponding pharmacologic phenotype may be displayed (e.g., "high risk for disease Y"). These risk factors and levels may be stored and processed in a quantitative or semi - quantitative form within the internal workings of a statistical model, but may be translated into qualitative terms for output to care providers and patients.
[0182] In some embodiments, the pharmacologic phenotype assessment server 102 may compare the likelihood of a pharmacologic phenotype to a likelihood threshold (e.g., 0.5) and may include the pharmacologic phenotype in the risk analysis display when the likelihood of the pharmacologic phenotype exceeds the likelihood threshold. For example, only diseases that are at high risk for the current patient may be included in the risk analysis display. In another example, the risk analysis display may include an indication that the current patient may have a substance abuse problem when the likelihood of the current patient's substance abuse problem exceeds the likelihood threshold. In this way, healthcare providers may recommend or provide early intervention to the current patient. Similarly, in some embodiments, the response scores for each drug corresponding to a particular disease may be ranked (e.g., from highest to lowest). The drugs and the corresponding response scores may be provided in ranked order on the risk analysis display. In other embodiments, only the drug with the highest rank for a particular disease may be included in the risk analysis display.
[0183] In addition to displaying one or more top-ranked drugs (and / or other therapies) for a particular disease as a recommendation to healthcare providers for prescribing to a patient, the risk analysis shows that the recommended dose of the drug for the current patient can also be included. The risk analysis shows that any social omics, physiomics, environmental, or demographic information that may cause a change in the response of the current patient to the drug (e.g., changes in diet, exercise, exposure, etc.) can also be included. In addition, the risk analysis shows that recommendations for changing the current patient's existing therapy can be included by changing the dose, changing the drug, or eliminating the treatment regimen or other methods that the current patient is taking based on their polypharmacy data. For example, when the recommended drug may render one or several drugs in the current patient's polypharmacy data redundant, the risk analysis shows that recommendations for discontinuing these drugs can be included.
[0184] In some embodiments, if the patient is taking a drug that is incompatible with the top-ranked drug, the pharmacologic phenotype assessment server 102 can recommend a drug with a lower response score (or other drug-specific property) but higher compatibility with the existing therapy (or other polypharmacy properties). For example, the pharmacologic phenotype assessment server 102 can obtain the polypharmacy data of the current patient and compare the drugs within the polypharmacy data with the recommended drug to check for contraindications. The risk analysis shows that the top-ranked drug with no contraindications to any of the drugs prescribed to the current patient can be included.
[0185] In addition to the clinical setting, pharmacologic phenotypes can be predicted in a research setting for drug development and insurance applications. In the research setting, the pharmacologic phenotypes of potential patient cohorts related to novel drugs, experimental drugs, or repurposed drugs may be predicted in a research protocol. Patients can be selected for experimental treatment based on their predicted pharmacologic phenotypes related to the experimental drug.
[0186] In addition, when the pharmacologic phenotype of the current patient becomes known (e.g., the current patient has certain pharmacologic phenotypes after a threshold amount of time, such as one year), the omics, social omics, physiomics, environmental, and phenomics data of the current patient can be added to the training data, and the statistical model can be updated accordingly.
[0187] Figure 6Illustrated is an example timeline 600 of a current patient, in which omics, social omics, physiomics, environmental, and phenomics data of the current patient are collected over time. Then, the pharmacologic phenotype prediction system 100 can analyze the collected data of the current patient according to a statistical model to predict the pharmacologic phenotype of the current patient. More specifically, in the example timeline 600, the diagnosis, treatment methods, and outcomes 602 of the current patient are collected. Also collected are the medical physiomics, EHR, laboratory values, stress, abuse factors, and trauma 604 of the current patient, as well as the social omics and demographics 608, omics and drug metabolomics 610, and exposome data 612 of the current patient. In addition, the phenomics data 606 of the current patient (which can indicate the medical outcomes 602 of the current patient) are collected.
[0188] As described above, in a clinical treatment and / or pharmacology or other biomedical research setting, the pharmacologic phenotype prediction system 100 can include a clinical decision support module 614 for clinicians and a clinical decision support module 616 for researchers.
[0189] In the clinical decision support module 614 for clinicians, a statistical model is generated in a manner similar to the above based on training data from individual training patients or a training patient cohort / group. The omics, social omics, physiomics, environmental, and phenomics data of the current patient (e.g., diagnosis, treatment methods, and outcomes 602; medical physiomics, EHR, laboratory values, stress, abuse factors, and trauma 604; phenomics data 606; social omics and demographics 608; omics and drug metabolomics 610; and exposome data 612) are applied to the statistical model to predict the pharmacologic phenotype of the current patient. The pharmacologic phenotype can include disease risk or conditions, drug recommendations, drug adverse reaction scores, overall drug response scores, etc. However, these are merely a few examples of pharmacologic phenotypes. Additional or alternative pharmacologic phenotypes are described throughout the text.
[0190] In the clinical decision support module 616 for researchers, a statistical model is generated in a manner similar to the above based on training data from individual patients or a patient cohort / group. The omics, social omics, physiomics, environmental, and phenomics data of the patient are applied to the statistical model to identify GWAS analysis results that describe the relationship between the training patient cohort and specific pharmacologic phenotypes, pharmacology / drug metabolomics results, precise phenomics analysis results, biomarkers, etc.
[0191] At a first time point within timeline 600, omics, socialomics, physiomics, and environmental data are collected from the current patient. The phenomics state of the patient is negative 622 at this time. Then, the current patient begins to experience disease symptoms and is subsequently hospitalized, resulting in a further negative phenomics state 624 at a second time point. All of this information (including the patient's response to treatment due to hospitalization 626) is provided to the clinical decision support module 614 of the clinician. Then, the clinical decision support module 614 of the clinician identifies the pharmacological phenotype of the current patient based on the omics, socialomics, physiomics, and environmental data and provides treatment options 628 that, for example, have the highest predicted response for the current patient. Accordingly, the phenomics state of the current patient changes from negative 622, 624 to positive 630 and remains positive 632 at a subsequent time point.
[0192] Figure 7 FIG. depicts a flowchart of an exemplary method 700 for identifying a pharmacological phenotype using machine learning techniques. Method 700 may be executed on a pharmacological phenotype assessment server 102. In some embodiments, method 700 may be implemented in a set of instructions stored on a non-transitory computer-readable memory and executable on one or more processors of pharmacological phenotype assessment server 102. For example, method 700 may be executed by Figure 1A the training module 160 and the phenotype assessment module 162 within the machine learning engine 146 of.
[0193] At block 702, the training module 160 may obtain a set of training data that includes omics, socialomics, physiomics, and environmental data of training patients, where it is known whether the training patients have a pharmacological phenotype (e.g., have a current or previously determined pharmacological phenotype). Environmental, physiomics, and socialomics data may include clinical data, demographic data, polypharmacy data, socioeconomic data, educational data, drug abuse data, diet and exercise data, law enforcement data, circadian rhythm data, family data, or any other suitable data indicating the patient's environment. Omics data may include genomics data indicating gene traits, epigenomics data indicating gene expression, transcriptomics data indicating DNA transcription, proteomics data indicating proteins expressed by the genome, chromosomics data indicating the chromatin state in the genome, and / or metabolomics data indicating metabolites in the genome. As described above, the omics, socialomics, physiomics, and environmental data of the training patients may be obtained at several time points (e.g., over a three-year time span).
[0194] Social omics, physiomics, and environmental data can be obtained from electronic medical records (EMRs) located on an EMR server and / or from multi-pharmacy data located on a multi-pharmacy server that aggregates pharmacy data of patients from multiple pharmacies. Additionally, social omics, physiomics, and environmental data can be obtained from a trained patient's healthcare provider or from the trained patient's self-report. In some embodiments, training data can be obtained from a combination of sources including several servers (e.g., an EMR server, a multi-pharmacy server, etc.) and client devices 106-116 of healthcare providers and patients.
[0195] Omics data can be obtained from the client devices 106-116 of healthcare providers. For example, a healthcare provider can obtain a biological sample (e.g., from saliva, biopsy, blood sample, bone marrow, hair, sweat, odor, etc.) for measuring the omics of a patient and provide the laboratory results obtained by analyzing the biological sample to a pharmacologic phenotype assessment server 102. In other embodiments, omics data can be directly obtained from a laboratory that analyzes biological samples. In other embodiments, omics data can be obtained from a GWAS or a candidate gene association study that describes the relationship between a trained patient cohort and a specific pharmacologic phenotype.
[0196] The training module 160 can also obtain phenomics data related to the pharmacologic phenotype of a trained patient, such as chronic diseases that the trained patient has, the pharmacologic response to drugs previously prescribed to the trained patient, whether each of the trained patients has a drug abuse problem, etc.
[0197] Then, the training module 160 can classify the omics, social omics, physiomics, and environmental data according to the pharmacologic phenotype of the trained patient associated with the omics, social omics, physiomics, and environmental data (block 704). The pharmacologic phenotype can include at least some diseases, drug abuse problems, pharmacologic responses to various drugs, complications, etc. that the trained patients are diagnosed with. In some cases, the omics, social omics, physiomics, and environmental data of the same trained patient can be subdivided into multiple pharmacologic phenotypes at different time periods. For example, as the environmental data of a trained patient changes, the omics data of the trained patient may change over time, and at a first time period, the trained patient may have a mental illness, while at a second time period, the trained patient may have another mental illness or complication.
[0198] Similarly, in some embodiments, the omics, social omics, physiomics, and environmental data of the training patients can be classified according to demographics. For example, the omics, social omics, physiomics, and environmental data of training patients of European descent can be assigned to one cohort, while the omics, social omics, physiomics, and environmental data of training patients of Chinese descent can be assigned to another cohort. In another example, the omics, social omics, physiomics, and environmental data of training patients between the ages of 25 and 35 can be assigned to one cohort, while the omics, social omics, physiomics, and environmental data of training patients between the ages of 35 and 45 can be assigned to another cohort.
[0199] Then, various machine learning techniques can be used to analyze the omics, social omics, physiomics, and environmental data of the training patients and their respective pharmacologic phenotypes to generate a statistical model (block 706) for determining the likelihood or other semi - quantitative and quantitative metrics indicating that the current patient has various pharmacologic phenotypes. The statistical model can also be used to determine response scores, doses, or any other suitable indication of the predicted response of the current patient to various drugs.
[0200] For example, as referred to above Figure 4B it was described that various machine learning techniques were used to analyze statistical tests from GWAS or candidate gene association studies that indicated the relationship between a training patient cohort and a specific pharmacologic phenotype, thereby scoring and / or ranking the variants identified in the study and / or variants that are in tight linkage with the identified variants. The top - ranked variants can be identified as SNPs, genes, and genomic regions associated or strongly associated with a specific pharmacologic phenotype.
[0201] For example, for the warfarin phenotype, the warfarin response pathway can be identified (such as Figure 4FAs shown, the warfarin response pathway contains 74 SNPs in 31 genes expressed in the liver, small intestine, and vasculature. The warfarin response pathway contains the following genes: AKR1C3 (expressed in the liver), CYP2C19 (expressed in the liver), CYP2C8 (expressed in the liver), CYP2C9 (expressed in the liver), CYP4F2 (expressed in the liver), F5 (expressed in the liver), F7 (expressed in the liver), F10 (expressed in the liver), F11 (expressed in the liver), FGG (expressed in the liver), ORM1 (expressed in the liver), PRSS53 (expressed in the liver), VKORC1 (expressed in the liver), STX4 (expressed in the small intestine), F13A1 (expressed in the vasculature), PROCR (expressed in the vasculature), VWF (expressed in the vasculature), CFHR5 (expressed in the liver), FGA (expressed in the liver), FMO5 (expressed in the liver), HRG (expressed in the liver), KNG1 (expressed in the liver), SURF4 (expressed in the liver), ABO (expressed in the small intestine), LYZ (expressed in the small intestine), PCGF3 (expressed in the small intestine), PRSS8 (expressed in the small intestine), TRPC4AP (expressed in the small intestine), SLC44A2 (expressed in the vasculature), SPHK1 (expressed in the vasculature), and USP7 (expressed in the vasculature). The 74 SNPs contained in the 31 genes are: rs12775913 (regulatory SNP), rs346803 (regulatory SNP), rs346797 (regulatory SNP), rs762635 (regulatory SNP), and rs76896860 (regulatory SNP) contained in the AKR1C3 gene; rs3758581 (coding SNP) contained in the CYP2C19 gene; rs10509681 (coding SNP) and rs11572080 (coding SNP) contained in the CYP2C8 gene; rs1057910 (coding SNP), rs1799853 (coding SNP), and rs7900194 (coding SNP) contained in the CYP2C9 gene; rs2108622 (coding SNP) contained in the CYP4F2 gene; rs6009 (regulatory SNP), rs11441998 (regulatory SNP), rs2026045 (regulatory SNP), rs34580812 (regulatory SNP), rs749767 (regulatory SNP), rs9378928 (regulatory SNP), and rs7937890 (regulatory SNP) contained in the F5 gene; rs7552487 (regulatory SNP), rs6681619 (regulatory SNP), rs8102532 (regulatory SNP), rs491098 (coding SNP), and rs6046 (coding SNP) contained in the F7 gene;rs11150596 (regulatory SNP) and rs11150596 (regulatory SNP) contained in the F10 gene; rs2165743 (regulatory SNP) and rs11252944 (regulatory SNP) contained in the F11 gene; rs8050894 (regulatory SNP) contained in the FGG gene; rs10982156 (regulatory SNP) contained in the ORM1 gene; rs7199949 (coding SNP) contained in the PRSS53 gene; rs2884737 (regulatory SNP), rs9934438 (regulatory SNP), rs897984 (regulatory SNP) and rs17708472 (regulatory SNP) contained in the VKORC1 gene; rs35675346 (regulatory SNP) and rs33988698 (regulatory SNP) contained in the STX4 gene; rs5985 (coding SNP) contained in the F13A1 gene; rs867186 (coding SNP) contained in the PROCR gene; rs75648520 (regulatory SNP), rs55734215 (regulatory SNP), rs12244584 (regulatory SNP) and rs1063856 (coding SNP) contained in the VWF gene; rs674302 (regulatory SNP) contained in the CFHR5 gene; rs12928852 (regulatory SNP) and rs6050 (coding SNP) contained in the FGA gene; rs8060857 (regulatory SNP) and rs7475662 (regulatory SNP) contained in the FMO5 gene; rs9898 (coding SNP) contained in the HRG gene; rs710446 (coding SNP) contained in the KNG1 gene; rs11577661 (regulatory SNP) contained in the SURF4 gene; rs11427024 (regulatory SNP), rs6684766 (regulatory SNP), rs2303222 (regulatory SNP), rs1088838 (regulatory SNP), rs13130318 (regulatory SNP) and rs12951513 (regulatory SNP) contained in the ABO gene; rs8118005 (regulatory SNP) contained in the LYZ gene; rs76649221 (regulatory SNP), rs9332511 (regulatory SNP) and rs6588133 (regulatory SNP) contained in the PCGF3 gene; rs11281612 (regulatory SNP) contained in the PRSS8 gene;rs11589005 (regulatory SNP), rs8062719 (regulatory SNP), rs889555 (regulatory SNP), rs36101491 (regulatory SNP), rs7426380 (regulatory SNP), rs6579208 (regulatory SNP), rs77420750 (regulatory SNP), and rs73905041 (coding SNP) contained in the TRPC4AP gene; rs3211770 (regulatory SNP), rs3211770 (regulatory SNP), rs3087969 (coding SNP), and rs2288904 (coding SNP) contained in the SLC44A2 gene; rs683790 (regulatory SNP) and rs346803 (coding SNP) contained in the SPHK1 gene; and rs201033241 (coding SNP) contained in the USP7 gene.;
[0202] In another example, for the lithium phenotype, a lithium response pathway can be identified, the lithium response pathway comprising 78 SNPs in 12 genes expressed in the brain.
[0203] Environmental data, socialomics data, physiomics data, and phenomics data of training patients with a suitable combination of SNPs, genes, and genomic regions can be obtained to distinguish training patients with the identified SNPs, genes, and genomic regions who do have a specific pharmacological phenotype from training patients with the identified SNPs, genes, and genomic regions who do not have a specific pharmacological phenotype.
[0204] Statistical models for predicting pharmacological phenotypes can be generated using machine learning techniques, including but not limited to regression algorithms (e.g., ordinary least squares regression, linear regression, logistic regression, stepwise regression, multivariate adaptive regression splines, locally estimated scatterplot smoothing, etc.), instance-based algorithms (e.g., k-nearest neighbor, learning vector quantization, self-organizing maps, locally weighted learning, etc.), regularization algorithms (e.g., ridge regression, least absolute shrinkage and selection operator, elastic net, least angle regression, etc.), decision tree algorithms (e.g., classification and regression trees, C4.5, C5, chi-squared automatic interaction detection, decision stumps, M5, conditional decision trees, etc.), clustering algorithms (e.g., k-means, k-medians, expectation maximization, hierarchical clustering, spectral clustering, mean shift, density-based spatial clustering of applications with noise, ordering points to identify the clustering structure, etc.), association rule learning algorithms (e.g., Apriori algorithm, Eclat algorithm, etc.), Bayesian algorithms (e.g., naive Bayes, Gaussian naive Bayes, multinomial naive Bayes, averaged one-dependence estimators, Bayesian belief networks, Bayesian networks, etc.), artificial neural networks (e.g., perceptrons, Hopfield networks, radial basis function networks, etc.), deep learning algorithms (e.g., multi-layer perceptrons, deep Boltzmann machines, deep belief networks, convolutional neural networks, stacked autoencoders, generative adversarial networks, etc.), dimensionality reduction algorithms (e.g., principal component analysis, principal component regression, partial least squares regression, Sammon mapping, multidimensional scaling, projection pursuit, linear discriminant analysis, mixture discriminant analysis, quadratic discriminant analysis, flexible discriminant analysis, factor analysis, independent component analysis, non-negative matrix factorization, t-distributed stochastic neighbor embedding, etc.), ensemble algorithms (e.g., boosting, bagging, AdaBoost, stacking generalization, gradient boosting machines, gradient boosting regression trees, random decision forests, etc.), reinforcement learning (e.g., temporal difference learning, Q-learning, learning automata, state-action-reward-state-action, etc.), support vector machines, mixture models, evolutionary algorithms, probabilistic graphical models, etc. For example, in addition to statistical measurements of social omics, physiomics, and environmental data (e.g., average income of training patients corresponding to each classification or social omics data such as average ACE scores, etc.), statistical models can also be generated based on identified SNPs, genes, and genomic regions.
[0205] In addition, the training module 160 can generate several statistical models for several pharmacological phenotypes. For example, a first statistical model can be generated to determine the likelihood that the current patient will experience a substance abuse problem, a second statistical model can be generated to determine the risk of having a disease, a third statistical model can be generated to determine the risk of having another disease, a fourth statistical model can be generated to determine the likelihood of having a negative reaction to a specific drug, and so on. In any case, each statistical model can be a graphical model, a decision tree, a probability distribution, or any other suitable model for determining the likelihood that the current patient has a certain pharmacological phenotype or a drug response score based on the training data.
[0206] At block 708, omics, social omics, physiomics, and environmental data of the current patient can be obtained. A process similar to process 500 described above with reference to Figure 5 can be used to obtain the omics data. For example, a healthcare provider can obtain a biological sample of the patient and send it to an analysis laboratory for analysis. Then, cells are extracted from the biological sample and reprogrammed into stem cells, such as iPSCs. Then, the iPSCs are differentiated into various tissues, such as neurons, cardiomyocytes, etc., and analyzed to obtain the omics data. The omics data can include genomic data, epigenomic data, transcriptomic data, proteomic data, karyomic data, metabolomic data, and / or biological networks. Specifically, the omics data can include a quantitative assessment of the current drugs of the patient by performing metabolomic measurements on the patient samples.
[0207] The social omics, physiomics, and environmental data, such as in the exposome, social omics and demographics, and medical physiomics, EHR, laboratory values, stress and abuse factors and trauma, and medical outcome data as described above, can be collected at several time points. For example, the social omics, physiomics, and environmental data of the current patient can indicate that the current patient is single, then gets married in year 1, and then gets divorced in year 2. The social omics, physiomics, and environmental data can also indicate that the current patient is employed in year 1 and then loses the job in year 2. In addition, the social omics, physiomics, and environmental data can indicate that the current patient is a victim of domestic abuse in year 1. The longitudinal data can be compared with the similar experiences of the patients trained during a similar time period as shown in the statistical model.
[0208] Then, at block 710, the omics, socialomics, physiomics, and environmental data of the current patient can be applied to a statistical model to determine the pharmacological phenotype of the current patient. The pharmacological phenotype can include the likelihood that the current patient has various diseases or a response score indicating the expected response of the current patient to various drugs and the recommended dosage of the drugs. For example, if the statistical model is a neural network, the phenotype assessment module 162 can use the omics, socialomics, physiomics, and environmental data of the current patient to traverse the nodes of the neural network to reach the respective output nodes, so as to determine the likelihood or response score. If several statistical models are generated, the phenotype assessment module 162 can apply the omics, socialomics, physiomics, and environmental data to each of the statistical models to determine, for example, the likelihood or risk of having bipolar disorder, the likelihood or risk of having schizophrenia, the likelihood of having a substance abuse problem, the response score of taking lithium to treat bipolar disorder, etc.
[0209] For example, the omics data of the current patient can be analyzed to identify SNPs and genes in the omics data of the current patient that are identical to any of the 74 SNPs or 31 genes associated with the warfarin phenotype in the warfarin response pathway, so as to determine whether the current patient has any of the warfarin phenotypes. Additionally, the socialomics, physiomics, and environmental data of the current patient can also be applied to a warfarin statistical model to determine the warfarin phenotype of the current patient. In addition to the statistical measurements performed on the socialomics, physiomics, and environmental data, a warfarin statistical model can also be generated based on the identified 74 SNPs, 31 genes, and the warfarin response pathway.
[0210] In another example, the omics data of the current patient can be analyzed to identify SNPs and genes in the omics data of the current patient that are identical to any of the 78 SNPs or 12 genes associated with the lithium phenotype in the lithium response pathway, so as to determine whether the current patient has any of the lithium phenotypes. Additionally, the socialomics, physiomics, and environmental data of the current patient can also be applied to a lithium statistical model to determine the lithium phenotype of the current patient. In addition to the statistical measurements performed on the socialomics, physiomics, and environmental data, a lithium statistical model can also be generated based on the identified 78 SNPs, 12 genes, and the lithium response pathway.
[0211] At block 712, the phenotypic assessment module 162 can display on the user interface of the healthcare provider's client device one or more indications of the pharmacological phenotype of the current patient. For example, the phenotypic assessment module 162 can generate a risk analysis display that includes an indication of each of the likelihoods of various pharmacological phenotypes or other semi - quantitative and quantitative metrics, which can be expressed as a probability (e.g., 0.6), a percentage (e.g., 80%), a category in a set of categories (e.g., "high risk", "medium risk", or "low risk"), and / or in any other suitable manner. Additionally, the risk analysis display can include an indication of a response score for a drug, which can be numerical (e.g., 75 out of 100), categorical (e.g., "strong response", "adverse reaction", etc.), or in any other suitable manner. Each likelihood or response score and a description of the corresponding pharmacological phenotype (e.g., "high risk for disease Y") can be displayed. In this way, the healthcare provider of the current patient can view the indications of the pharmacological phenotype of the current patient and develop an appropriate treatment plan or treatment regimen. For example, the healthcare provider can prescribe a drug for treating a particular disease that has the highest response score for drugs treating that particular disease.
[0212] In some embodiments, the pharmacological phenotype assessment server 102 can compare the likelihood of a pharmacological phenotype to a likelihood threshold (e.g., 0.5) and can include the pharmacological phenotype in the risk analysis display when the likelihood of the pharmacological phenotype exceeds the likelihood threshold. For example, only diseases that are at high risk for the current patient can be included in the risk analysis display. In another example, the risk analysis display can include an indication that the current patient may have a drug abuse problem when the likelihood of the current patient's drug abuse problem exceeds the likelihood threshold. In this way, the healthcare provider can recommend or provide early intervention to the current patient. Also, in some embodiments, the response scores for each drug corresponding to a particular disease can be ranked (e.g., from highest to lowest). The drugs and the corresponding response scores can be provided in ranked order on the risk analysis display. In other embodiments, only the drug with the highest rank for a particular disease can be included in the risk analysis display.
[0213] In addition to displaying the top-ranked drugs for a particular disease as a recommendation for a healthcare provider to prescribe to a patient, the risk analysis indicates that the recommended dosage of the drugs being taken by the current patient can also be included. The risk analysis indicates that any social omics, physiological omics, environmental, or demographic information that may cause a change in the response of the current patient to the drugs (e.g., changes in diet, exercise, exposure, etc.) can also be included. Additionally, the risk analysis indicates that recommendations for increasing or decreasing the amount of drugs the current patient is taking based on the current patient's polypharmacy data can be included. For example, when the recommended drug may render one or more of the drugs in the current patient's polypharmacy data redundant, the risk analysis indicates that recommendations for discontinuing these drugs can be included. Such recommendations can relate to one or more drugs, drug combinations, or other treatment measures.
[0214] In some embodiments, if the patient is taking a drug that is incompatible with the top-ranked drug, the pharmacogenetic phenotype assessment server 102 can recommend the drug with the next highest response score. For example, the pharmacogenetic phenotype assessment server 102 can obtain the current patient's polypharmacy data and compare the drugs within the polypharmacy data with the recommended drug to check for contraindications. The risk analysis can include the top-ranked drug that has no contraindications with any of the drugs prescribed to the current patient.
[0215] As in the example regarding warfarin above, the omics data of the current patient can be compared with 74 SNPs or 31 genes related to the warfarin phenotype in the warfarin response pathway to determine whether the current patient has any of the warfarin phenotypes. Then, the pharmacogenetic phenotype assessment server 102 or the healthcare provider can determine based on the comparison whether warfarin or another anticoagulant should be administered to the current patient. The recommended dosage of warfarin can also be determined. For example, the current patient may have an SNP or gene in the warfarin response pathway related to a negative response to warfarin. Therefore, the pharmacogenetic phenotype assessment server 102 can recommend another anticoagulant to administer to the current patient. In another example, the current patient may have an SNP or gene in the warfarin response pathway related to the warfarin dosage phenotype. Therefore, the pharmacogenetic phenotype assessment server 102 can provide a recommended dosage based on the warfarin dosage phenotype to administer warfarin to the current patient. In yet another example, the current patient may have an SNP or gene in the warfarin response pathway related to disease risk, where warfarin can actively prevent blood clotting, coagulation, or thrombus formation. In any case, the healthcare provider can administer warfarin to the current patient at the recommended dosage or can administer another anticoagulant.
[0216] As in the example of lithium above, the omics data of the current patient can be compared with 78 SNPs or 12 genes related to the lithium phenotype in the lithium response pathway to determine whether the current patient has any of the lithium phenotypes. Then, the pharmacologic phenotype assessment server 102 or the healthcare provider can determine whether lithium or another psychiatric therapeutic drug should be administered to the current patient based on the comparison. A recommended dose of lithium can also be determined. For example, the current patient may have an SNP or gene in the lithium response pathway related to a negative response to lithium. Thus, the pharmacologic phenotype assessment server 102 can recommend another psychiatric therapeutic drug to administer to the current patient. In another example, the current patient may have an SNP or gene in the lithium response pathway related to the lithium dose phenotype. Thus, the pharmacologic phenotype assessment server 102 can provide a recommended dose based on the lithium dose phenotype to administer lithium to the current patient. In any case, the healthcare provider can administer lithium to the current patient at the recommended dose or can administer another psychiatric therapeutic drug.
[0217] In addition to the clinical setting, pharmacologic phenotypes can be predicted in a research setting for drug development and insurance applications. In the research setting, the pharmacologic phenotypes of a potential patient cohort related to an experimental drug may be predicted in a research protocol. Patients can be selected for experimental treatment based on their predicted pharmacologic phenotypes related to the experimental drug.
[0218] Furthermore, when the pharmacologic phenotype of the current patient becomes known (e.g., the current patient has a pharmacologic phenotype after a threshold amount of time, such as one year), then the omics, social omics, physiologic omics, and environmental and phenomic data of the current patient can be added to the training data (block 714), and the statistical model can be updated accordingly. In some embodiments, the omics, social omics, physiologic omics, environmental, and phenomic data are stored in several data sources 716, such as Figure 3 the data sources 325a-d described in. Then, the training module 160 can retrieve data from the data sources 716 to further train the model.
[0219] Throughout the specification, multiple instances may implement components, operations, or structures described as a single instance. Although the individual operations of one or more methods are shown and described as separate operations, one or more of the individual operations may be performed simultaneously and the operations need not be performed in the order shown. Structures and functions presented as separate components in an example configuration may be implemented as a combined structure or component. Similarly, structures and functions presented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements fall within the scope of the subject matter herein.
[0220] Additionally, this document describes certain embodiments as including logic or multiple routines, subroutines, applications, or instructions. These can constitute software (e.g., code embodied on a machine-readable medium or in a transmitted signal) or hardware. In hardware, a routine, etc. is a tangible unit capable of performing certain operations and can be configured or arranged in a certain manner. In an example embodiment, one or more computer systems (e.g., standalone client or server computer systems) or one or more hardware modules of a computer system (e.g., a processor or a group of processors) can be configured by software (e.g., an application or a portion of an application) to operate as a hardware module that performs certain operations as described herein.
[0221] In various embodiments, a hardware module can be implemented mechanically or electronically. For example, a hardware module can include dedicated circuitry or logic that is permanently configured to perform certain operations (e.g., a dedicated processor, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC)). A hardware module can also include programmable logic or circuitry that is temporarily configured by software to perform certain operations (e.g., as included in a general purpose processor or other programmable processor). It should be appreciated that the decision to implement a hardware module mechanically in dedicated and permanently configured circuitry or in temporarily configured circuitry (e.g., configured by software) may be driven by cost and time considerations.
[0222] Accordingly, the term "hardware module" should be understood to encompass a tangible entity that is physically constructed, permanently configured (e.g., hardwired) or temporarily configured (e.g., programmed) to operate in a certain manner or to perform any of the operations described herein. Considering embodiments in which a hardware module is temporarily configured (e.g., programmed), it is not necessary to configure or instantiate every hardware module at any given time. For example, in cases where a hardware module includes a general purpose processor that is configured by software, the general purpose processor can be configured to corresponding different hardware modules at different times. Thus, software can configure a processor, for example, to constitute a particular hardware module at one time and a different hardware module at a different time.
[0223] Hardware modules can provide information to and receive information from other hardware modules. Thus, the described hardware modules can be considered to be communicatively coupled. In cases where multiple such hardware modules are present simultaneously, communication can be achieved through signal transmission that connects the hardware modules (e.g., via appropriate circuitry and buses). In embodiments where multiple hardware modules are configured or instantiated at different times, such communication between the hardware modules can be achieved, for example, by storing and retrieving information in a memory structure accessible to the multiple hardware modules. For example, one hardware module can perform an operation and store the output of the operation in a memory device to which it is communicatively coupled. Then, another hardware module can access the memory device at a later time to retrieve and process the stored output. Hardware modules can also initiate communication with input or output devices and can operate on resources (e.g., collections of information).
[0224] The various operations of the example methods described herein can be performed, at least in part, by one or more processors temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors can constitute processor-implemented modules that operate to perform one or more operations or functions. In some example embodiments, the modules referred to herein can include processor-implemented modules.
[0225] Similarly, the methods or routines described herein can be implemented, at least in part, by processors. For example, at least some of the operations of the method can be performed by one or more processors or processor-implemented hardware modules. The execution of some of the operations can be distributed among one or more processors that reside not only within a single machine but are also deployed across multiple machines. In some example embodiments, one or more processors can be located in a single location (e.g., in a home environment, in an office environment, or as a server farm), while in other embodiments, the processors can be distributed across multiple locations.
[0226] The execution of some of the operations can be distributed among one or more processors that reside not only within a single machine but are also deployed across multiple machines. In some example embodiments, one or more processors or processor-implemented modules can be located in a single geographical location (e.g., in a home environment, an office environment, or a server farm). In other example embodiments, one or more processors or processor-implemented modules can be distributed across multiple geographical locations.
[0227] Unless otherwise specifically stated, discussions herein using terms such as "processing", "operation", "computation", "determination", "presentation", "display", etc. can refer to actions or processes of a machine (e.g., a computer) that manipulates or transforms data, where the data is represented as physical (e.g., electrical, magnetic, or optical) quantities in one or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other machine components that receive, store, transmit, or display information.
[0228] As used herein, any reference to "an embodiment" or "embodiments" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. The phrase "in an embodiment" appearing in various places in the specification does not necessarily refer to the same embodiment.
[0229] Some embodiments may be described using the expressions "coupled" and "connected" and their derivatives. For example, the term "coupled" may be used to describe some embodiments to indicate that two or more elements are in direct physical or electrical contact. However, the term "coupled" may also mean that two or more elements are not in direct contact with each other but still cooperate or interact. Embodiments are not limited to these scopes.
[0230] As used herein, the terms "comprises / comprising", "includes / including", "has / having", or any other variation thereof are intended to cover non-exclusive inclusion. For example, a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to those elements, but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Further, unless expressly stated to the contrary, "or" refers to the inclusive "or" rather than the exclusive "or". For example, any of the following satisfies the condition A or B: A is true (or present) and B is false (or absent), A is false (or absent) and B is true (or present), and both A and B are true (or present).
[0231] Additionally, "a / an" is used to describe elements and components of embodiments herein. This is merely for convenience and to provide a general description. This description and the appended claims should be understood to include one or at least one, and the singular also includes the plural unless clearly stated otherwise.
[0232] This detailed description should be construed as merely providing examples and not describing every possible embodiment, as it would be impractical, if not impossible, to describe every possible embodiment. Many alternative embodiments can be implemented using current technology or technology developed after the filing date of this application.
Claims
1. A computing device for identifying a lithium phenotype using statistical modeling and machine learning techniques, the computing device comprising: a communication network; one or more processors; and a non-transitory computer-readable memory coupled to the one or more processors and storing instructions thereon, the instructions when executed by the one or more processors cause the computing device to: obtain a set of training data, the training data comprising the following data for each of a plurality of first patients: omics data, which indicates the biological characteristics of the first patient, including one or more of the following: genomics data, transcriptomics data, proteomics data, or metabolomics data, socio-omics and environmental data, which refers to the experiences of the first patient collected at multiple instances over a period of time, including the exposures encountered by the first patient, and phenomics data, which indicates: the response to lithium; generate a statistical model for determining a lithium phenotype based on the set of training data; receive a set of omics data and socio-omics and environmental data of a second patient collected over a period of time; apply the omics data and the socio-omics and environmental data of the second patient to the statistical model to determine one or more lithium phenotypes of the second patient; and provide, via the communication network, the one or more lithium phenotypes of the second patient for display to a healthcare provider, wherein the healthcare provider recommends a treatment plan to the second patient based on the lithium phenotype, wherein the socio-omics and environmental data indicates the experiences of the first patient collected over time, which includes at least one of the following: clinical data, which indicates the medical history of the first patient; demographic data of the first patient, the socioeconomic data of the first patient; or polypharmacy data, which indicates each drug prescribed to the first patient in prescription form collected from at least two time instances, wherein the one or more lithium phenotypes of the second patient includes at least one of the following: the predicted response of the second patient to lithium, an adverse drug event or adverse drug reaction to lithium, wherein the statistical model is generated using one or more machine learning techniques.
2. The computing device according to claim 1, wherein the instructions further cause the computing device to: determine the efficacy of lithium for the second patient based on the omics data and the socio-omics and environmental data of the second patient; determine one or more expected adverse drug reactions regarding the second patient based on the omics data and the socio-omics and environmental data of the second patient; determine the overall value of lithium for the second patient by combining the efficacy and the one or more expected adverse drug reactions.
3. The computing device according to any one of claims 1 to 2, wherein in order to provide the lithium phenotype of the second patient for display to a healthcare provider, the instructions cause the computing device to: generate a risk analysis display of the second patient, which includes: the predicted response of the second patient to lithium.
4. The computing device according to any one of claims 1 to 2, wherein the instructions further cause the computing device to: After a threshold time period, receive the phenomics data of the second patient, the phenomics data indicating: response to lithium; Store the omics data, the socialomics and environmental data, and the phenomics data of the second patient in a knowledge base; and Update the set of training data to include data from the knowledge base.
5. The computing device according to any one of claims 1 to 2, wherein the omics data comprises at least one of genomics data, epigenomics data, transcriptomics data, proteomics data, chromosomics data or metabolomics data.
Citation Information
Patent Citations
System and methods for the production of personalized drug products
CN103250176A
Learning health systems and methods
US20150324527A1