Pharmacological phenotype prediction platform for individuals and queues
By integrating machine learning with omics, sociology, and environmental data, the pharmacological phenotype prediction system solves the problem of inaccurate drug response prediction in existing technologies and realizes personalized drug treatment recommendations and disease risk management.
Patent Information
- Application Number
- CN202510686581.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-02-21
- Filing Date
- 2018-05-11
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies fail to effectively utilize chromatin states, genomic regulatory elements, epigenomics, proteomics, metabolomics, and transcriptomics to predict patients' pharmacological phenotypes, and do not consider changes in biological and sociological characteristics, resulting in inaccurate predictions of drug responses.
The pharmacological phenotype prediction system is trained through machine learning technology, integrating omics, sociological and environmental data, generating statistical models, and analyzing patients' biological, sociological and environmental characteristics to predict drug responses and disease risks.
It provides a more accurate drug response and disease risk prediction system, which can provide personalized treatment recommendations based on the patient's near real-time changes, reduce adverse reactions and improve treatment effects.
Smart Images

Figure CN120636518A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application is a divisional application of application number: 201880046200.7, entitled “Individual and Cohort Pharmacological Phenotype Prediction Platform.” This application claims priority to and the benefit of (1) provisional U.S. application serial number 62 / 505,422, filed on May 12, 2017, entitled “Individual and Cohort Pharmacological Phenotype Prediction Platform,” and (2) provisional U.S. application serial number 62 / 633,355, filed on February 21, 2018, entitled “Individual and Cohort Pharmacological Phenotype Prediction Platform,” the entire disclosure of each of which is hereby expressly incorporated herein by reference. Technical Field
[0003] The present application relates to pharmacological patient phenotyping, and more particularly, to a method and system for predicting drug response phenotypes of patients and stratified cohorts of patients based on their biological, ancestral, demographic, clinical, sociological, and environmental characteristics using machine learning and statistical techniques. Background Art
[0004] Today, drug responses can be predicted for some patients based on their coding genome. Specific genetic traits can be mapped to specific responses to drugs, and drugs can be selected for patients based on their predicted responses.
[0005] However, noncoding genomic variants account for the vast majority of differences in genetic traits, such as patient drug response, adverse drug reactions, and disease risk. The integration of epigenomic regulation research with genome-wide association studies (GWAS) has also shown that epigenomic alterations can indicate disease risk, drug response, and adverse drug reactions in humans and animals across a wide range of medical specialties and drug research settings. Furthermore, phenotypic variation associated with disease may be determined by differences in chromatin states that were previously attributed to genetic differences.
[0006] Current systems do not utilize chromatin state, genomic regulatory elements, epigenomics, proteomics, metabolomics, or transcriptomics to predict a patient's pharmacological phenotype. Current systems also do not consider environmental and social characteristics that may alter genetic traits to determine pharmacological phenotype. Furthermore, such systems do not utilize machine learning techniques to train the system to adapt to changes in biological traits and / or pharmacological phenotypes corresponding to biological traits over time.
[0007] Therefore, a system is needed to accurately predict pharmacological phenotypes (including pharmacological response, disease risk, drug abuse or other pharmacological phenotypes) based on omics features (including genomics, epigenomics, chromatin state, proteomics, metabolomics, transcriptomics, etc.) and patients' near real-time social and environmental characteristics. Summary of the Invention
[0008] In order to predict the pharmacological phenotype of a patient, various machine learning techniques can be used to train a pharmacological phenotype prediction system. More specifically, the pharmacological phenotype prediction system can be trained to analyze the patient's omics, sociology, and environmental data to predict the patient's response to various drugs, the patient's likelihood of drug abuse, the risk of various diseases, or any other pharmacological phenotype of the patient. The pharmacological phenotype prediction system can be trained by obtaining the omics, sociology, and environmental data (also referred to herein as "training data") of a group of patients (also referred to herein as "training patients").
[0009] In some embodiments, the patient's sociological and environmental data can be obtained at multiple time points to record the patient's experience in detail. For each training patient, the pharmacological phenotype prediction system can obtain the patient's pharmacological phenotype as training data, such as whether the patient has a drug abuse problem, the patient's chronic disease, the patient's response to various drugs prescribed to the patient, etc. The various machine learning techniques can be used to analyze the training data to generate a statistical model, which can be used to predict the patient's response to various drugs, the patient's likelihood of drug abuse, the risk of various diseases, or any other pharmacological phenotype of the patient. For example, the statistical model can be a neural network generated based on a combination of network analysis of gene regulatory networks and environmental influences on gene expression.
[0010] After the training period, the pharmacological phenotype prediction system can receive omics, sociology, and environmental data collected at several time points for patients whose pharmacological phenotypes are unknown (e.g., lithium has not yet been prescribed to a bipolar patient, and therefore the patient's response to lithium is unknown). The omics, sociology, and environmental data can be applied to the statistical model to predict the patient's pharmacological phenotype, and these phenotypes can be displayed on a healthcare provider's client device.
[0011] For example, for a particular drug, the pharmacological phenotype prediction system can determine the likelihood of the patient experiencing an adverse drug reaction. In addition, the pharmacological phenotype prediction system can generate an indication of the predicted efficacy or appropriate dosage of the drug for the patient. In some embodiments, the likelihood of the patient experiencing an adverse drug reaction can be compared with a threshold likelihood, and the predicted efficacy can be compared with a threshold efficacy. When the likelihood exceeds the threshold likelihood, the predicted efficacy is less than the threshold efficacy, and / or when the combination of the likelihood of an adverse drug reaction and the predicted efficacy exceeds a threshold, an indication of the likelihood and / or efficacy of the drug can be provided to the healthcare provider. Therefore, the healthcare provider can change the dosage, not prescribe the drug to the patient, or recommend an alternative drug with a higher efficacy to the patient.
[0012] In this way, the pharmacological phenotype prediction system can identify the best drug for a patient suffering from a particular disease. For example, for a particular disease, the pharmacological phenotype prediction system can select one of several drugs designed to treat the disease that has the greatest predicted efficacy for the patient and the least likelihood and / or severity of adverse drug reactions. This embodiment advantageously allows healthcare providers to accurately and efficiently identify the best drug to recommend and prescribe to the patient. In addition, by combining omics, sociological, and environmental data to generate the statistical model, this embodiment advantageously includes a comprehensive bioinformatics analysis of the patient's biological characteristics, which may change over time. This comprehensive bioinformatics analysis provides a more accurate prediction system that can not only predict the pharmacological phenotype based on the patient's inherent characteristics, but can also incorporate sociological and environmental traits that change over time and may alter the expression of genetic traits.
[0013] Furthermore, by generating statistical models that accurately predict disease risk and likelihood of adverse drug reactions, the healthcare provider can proactively address these disease symptoms before the patient exhibits disease symptoms or begins to develop substance abuse problems or other disorders.
[0014] In one embodiment, a computer-implemented method for identifying a pharmacological phenotype using statistical modeling and machine learning techniques is provided. The method comprises obtaining a set of training data, the training data comprising the following data for each of a plurality of first patients: omics data indicating a biological characteristic of the first patient, sociomic and environmental data indicating the experiences of the first patient collected over time, and phenomic data indicating at least one of: response to one or more drugs, whether the first patient experiences adverse drug reactions or drug abuse, or one or more chronic diseases of the first patient. The method further comprises: generating a statistical model for determining a pharmacological phenotype based on the set of training data; receiving a set of omics data and sociomic and environmental data for a second patient collected over a period of time; applying the omics data and sociomic and environmental data for the second patient to the statistical model to determine one or more pharmacological phenotypes for the second patient; and providing the one or more pharmacological phenotypes of the second patient for presentation to a healthcare provider, wherein the healthcare provider recommends a treatment regimen for the second patient based on the one or more pharmacological phenotypes.
[0015] In another embodiment, a computing device for identifying pharmacological phenotypes using statistical modeling and machine learning techniques is provided. The computing device includes a communication network, one or more processors, and a non-transitory computer-readable memory coupled to the one or more processors and storing instructions thereon. The instructions, when executed by the one or more processors, cause the system to acquire a set of training data, the training data including the following data for each of a plurality of first patients: omics data indicating a biological characteristic of the first patient, sociomic and environmental data indicating the experiences of the first patient collected over time, and phenomic data indicating at least one of the following: response to one or more drugs, whether the first patient experiences adverse drug reactions or drug abuse, or one or more chronic diseases of the first patient. The instructions further cause the system to: generate a statistical model for determining a pharmacological phenotype based on the set of training data; receive a set of omics data and sociomic and environmental data of a second patient collected over a period of time; apply the omics data and the sociomic and environmental data of the second patient to the statistical model to determine one or more pharmacological phenotypes of the second patient; and provide the one or more pharmacological phenotypes of the second patient for presentation to a healthcare provider via the communication network, wherein the healthcare provider recommends a treatment regimen to the second patient based on the pharmacological phenotype. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1Ashows a block diagram of a computer network and system on which an exemplary pharmacological phenotype prediction system according to the presently described embodiments may operate;
[0017] Figure 1B is possible according to the presently described embodiment Figure 1A A block diagram of an exemplary pharmacological phenotyping server operating in a system of
[0018] Figure 1C is possible according to the presently described embodiment Figure 1A A block diagram of an exemplary client device operating in a system;
[0019] Figure 2 depicts example omics, sociology, and environmental data that may be provided to a pharmacological phenotype prediction system according to the presently described embodiments;
[0020] Figure 3 depicts a detailed view of the process performed by a pharmacological phenotype prediction system according to the presently described embodiments;
[0021] Figure 4A depicts an exemplary representation of a bioinformatics analysis of permissive candidate variants associated with a particular pharmacological phenotype according to the presently described embodiments, and a schematic diagram representing an exemplary transcriptional spatial hierarchy in the human genome;
[0022] Figure 4B is a block diagram representing an exemplary method for using machine learning techniques to identify omics data corresponding to a specific pharmacological phenotype according to the presently described embodiments;
[0023] Figure 4C depicts an exemplary gene regulatory network for a patient according to the presently described embodiments;
[0024] Figure 4D is a block diagram representing another exemplary method for using machine learning techniques to identify omics data corresponding to a specific pharmacological phenotype, according to the presently described embodiments;
[0025] Figure 4E is when identifying omics data corresponding to the warfarin phenotype Figure 4D Box diagram of the single nucleotide polymorphisms (SNPs) identified in each stage of the method described in;
[0026] Figure 4F An exemplary warfarin reaction pathway according to embodiments described herein is depicted;
[0027] Figure 4G depicts an exemplary lithium reaction pathway according to an embodiment described herein;
[0028] Figure 5 is a block diagram representing an exemplary process for generating omics data from a patient's biological sample;
[0029] Figure 6 depicts an example timeline of a patient including example omics, phenomics, sociomics, physiomics, and environmental data collected over time and the patient's pharmacological phenotype as determined by a pharmacological phenotype prediction system according to the presently described embodiments; and
[0030] Figure 7 A flow chart representing an exemplary method for identifying pharmacological phenotypes using machine learning techniques according to the presently described embodiments is shown. DETAILED DESCRIPTION
[0031] Although the following text sets forth a detailed description of many different embodiments, it should be understood that the legal scope of this description is defined by the words of the claims set forth at the end of this disclosure. The detailed description should be construed as exemplary only and does not describe every possible embodiment, as describing every possible embodiment would be impractical, if not impossible. Numerous alternative embodiments can be implemented using current technology or technology developed after the filing date of this patent, which would still fall within the scope of the claims.
[0032] It should also be understood that unless a term is expressly defined in this patent using the sentence "As used herein, the term '______' is defined herein to mean ..." or similar sentences, no limitation of the meaning of said term, whether express or by implication, beyond its plain or ordinary meaning is intended, and such term should not be construed as limited in scope based on any statement made in any section of this patent (except in the language of the claims). To the extent that any term recited in the claims at the end of this patent is referenced in this patent in a manner consistent with a single meaning, this is done solely for clarity so as not to confuse the reader, and is not intended to limit such claim term, by implication or otherwise, to that single meaning. Finally, unless a claim element is defined by reference to the word "means" and the function of any recitation of structure, it is not intended to construe the scope of any claim element under the application of 35 U.S.C. § 112 (sixth paragraph).
[0033] Therefore, as used herein, the term "healthcare provider" can refer to any provider of medical or health services. For example, a health care provider can be a physician, clinician, nurse, physician's assistant, insurance provider, pharmacist, hospital, clinical setting, pharmacy technician, pharmaceutical company, research scientist, other medical organization, or medical professional licensed to prescribe medical products and medications to patients.
[0034] As used herein, the term "patient" may refer to any human or other organism, or combination thereof, whose health, lifespan, or other medical outcome is the target of clinical or research interest, study, or effort.
[0035] In addition, as used herein, the term "omics" can refer to a series of molecular biology techniques relevant to the interaction of other functions in the biological function of the cell and the human body. For example, omics can include genomics, epigenomics, chromatin state, transcriptomics, proteomics, metabolomics, biological networks and system models, etc. Omics data may be specific to each time point and specific cell tissues and cell lines. Therefore, omics data collection is relevant to these features, and omics data can also be collected and used for multiple tissues, pedigrees and time points relevant to the patient's phenotype being paid attention to. Patient's omics may be relevant to the biomarkers of multiple phenotypes, such as pharmacological responses to drugs, disease risks, complications, drug abuse problems, etc. Omics data can be generated and collected for a group of specific medical decision-making purposes at discrete time points, and omics data can also be collected from the total record of omics data collected for a single patient at various points in the past.
[0036] As used herein, the term "pharmacological phenotype" may refer to any discernible phenotype that may affect drug therapy, patient lifespan and results, quality of life, etc. in terms of clinical care, management and finance of clinical care, and pharmaceutical and other medical and biomedical research for humans and other organisms. Such phenotypes may include pharmacokinetic (PK) and pharmacodynamic (PD) phenotypes, all phenotypes of the rate and characteristics of the absorption, distribution, metabolism and excretion (ADME) of the drug, and drug reactions related to drug efficacy, drug treatment dose, half-life, plasma levels, clearance, etc., as well as adverse drug events, adverse drug reactions and adverse drug events or adverse drug reactions, organ damage, drug abuse and dependence, and their likelihood, as well as body weight and its change, mood and behavior changes and interference. Such phenotypes may also include favorable and unfavorable reactions to drug combinations, drug and gene interactions, social and environmental factors, dietary factors, etc. The phenotype may also include compliance with pharmacological or non-pharmacological treatment regimens. The phenotype may also include medical phenotypes, such as a patient's predisposition to contracting a certain disease or complication, the outcome and prognosis of a disease, whether a patient will experience specific disease symptoms, and patient outcomes (e.g., lifespan, clinical scores and parameters, test results, healthcare expenditures), and other phenotypes.
[0037] Furthermore, as used herein, the term "pharmacophenotyping" may refer to the pharmacological phenotype of an individual patient based on the integration of genomics, epigenomics, omics, pharmaco-metabolomics, sociomics, electronic health records (EHRs), and other patient data, matched to stratified patient cohorts and population datasets enabled by machine learning.
[0038] As used herein, "precise patient phenotyping" may refer to the comprehensive analysis of drug phenotypic data to provide a precise and accurate patient treatment profile for clinical decisions, which may be regularly updated to incorporate changing patient phenotypic data.
[0039] As used herein, the term "phenotypic transition" can refer to cyclical changes in a clinical patient's phenotype that recur or occur intermittently over time based on disease progression, sociological and environmental factors, and / or the results of initial, ongoing, or changing pharmacological and non-pharmacological treatments, essentially a longitudinal record of the patient's clinical progression.
[0040] Furthermore, as used herein, the term "disease susceptibility" may refer to risk factors associated with direct genetic inheritance or epigenetic modification through transgenes.
[0041] As used herein, "sociomic risk factors" can refer to sociological and cultural clinical risk factors associated with: behaviors that are harmful to oneself or others; adverse cultural, economic, and community living conditions; neglect and abuse during childhood and / or adolescence, which are referred to as adverse childhood experiences (ACEs); adult trauma associated with sexual, physical, and psychological abuse; other acute or chronic traumatic events (e.g., military conflict, crime, breakdown, illness, death in the family); exacerbated or prolonged stress caused by adverse conditions; and age-related health conditions, isolation, or cognitive conditions.
[0042] As used herein, "disease diagnosis" may refer to a possible or confirmed diagnosis that leads to a treatment decision. As used herein, "treatment options" may refer to one or more pharmacological and / or non-pharmacological treatments that alleviate, neutralize, or improve a patient's condition.
[0043] Likewise, as used herein, the term "initial therapeutic response" can refer to stabilization of disease, lack of response, improvement in clinical response, or adverse events (AEs) resulting from drug treatment during the first few weeks to months; this may involve dose adjustments or adjunctive medications. This time period typically ranges from six months to one year.
[0044] As used herein, the term "relapse response" may refer to cyclical changes in a patient's response to treatment due to adverse pharmacological effects, drug-drug interactions, changes in drug dosage, new or recurring complications, trauma, stress, and other socio-omic factors, as measured by biological samples (such as, but not limited to, blood, urine, sweat (e.g., Cortisol), odor) or by remote sensing, transmitters, or other active or passive data collection methods.
[0045] As used herein, the term "environment" shall mean any object, substance, emanation, condition, experience, communication, or information external to or originating from a human or other animal or organism, occurring now or in the past (including prior biological generations of such human or other organism), at one or more discrete points in time or over a period of time, that may affect or alter the physical, biological, chemical, physiological, medical, psychological, or psychiatric characteristics of such human or other organism in a measurable, identifiable, or other significant way. Such conditions may include the type, quantity, quality, presence / absence, timing, or other characteristics of food, nutritional supplements, minerals, water and other liquids, clothing, sanitary facilities, and other goods and services to which a human or other organism has come into contact, as well as current or past exposure to chemicals, atmosphere, and organisms, whether through the skin or by ingestion, inhalation, intubation, speculative, or other means. Such conditions may include temperature, noise, light, electromagnetic and / or particle radiation, vibration, mechanical shock or stress, medications, medical procedures, and implants. Such conditions may also include occupational attributes, job responsibilities, and recreational substances. These conditions may also include medical adverse events such as exposure to toxins, poisons, microorganisms, viruses, and other agents, as well as physical impacts, lacerations, contusions, punctures, and concussions.
[0046] These conditions may also include social factors such as adverse childhood experiences (ACEs) and stress, trauma, abuse, poverty and other economic conditions, food insecurity and hunger, incarceration, interpersonal conflict, violence, and other experiences. These conditions may also include the presence or absence of parents, children, siblings, and other family members and acquaintances, including the type, quality, and duration of such relationships. These conditions may also include educational and professional experience and achievements, religious services and guidance, and social contact and interaction. These conditions may also include sociogenomic risk factors and body modifications, including tattoos, implants, piercings, and plugs.
[0047] As used herein, the term "simultaneous pharmacogenomic exposure" can refer to a subset of environmental factors, including unidirectional interactions, simultaneous interactions, pharmacokinetic or pharmacodynamic interactions, and drug-environment interactions. For example, if there is a documented interaction demonstrating that an environmental factor alone induces or inhibits the activity of a specific enzyme involved in drug metabolism by 20%, or alters the drug's action by 20%, then the interaction is considered clinically important. Such interactions can encompass exposures ranging from foods to herbal / vitamin supplements to voluntary and involuntary toxic exposures. The likelihood of such an interaction can be numerically measured.
[0048] For simplicity, throughout this discussion, a patient whose data is used as training data to generate a statistical model may be referred to herein as a "training patient," and a patient whose data is applied to the statistical model to predict a pharmacological phenotype may be referred to as a "current patient." However, this is for ease of discussion. Data from the "current patient" may be added to the training data, and the training data may be continuously or periodically updated to keep the statistical model up to date. Additionally, a training patient may also have data applied to the statistical model to predict a pharmacological phenotype.
[0049] In addition, throughout this discussion, the current patient may be described as a patient for whom it is unknown whether or not they have certain pharmacological phenotypes, while the training patient may be described as a patient for whom the pharmacological phenotype is known. More specifically, the pharmacological phenotype of the current patient is unknown, and predictions are made using the relationship between the omics and sociomics, physiomics, and environmental data of the training patient and the previously or currently determined pharmacological phenotypes of the training patient. Therefore, the training patient has a known, previously or currently determined pharmacological phenotype. The current patient has an unknown pharmacological phenotype. However, in some embodiments, the training patient may have other unknown pharmacological phenotypes while having some known pharmacological phenotypes for training the pharmacological phenotype prediction system. In addition, the current patient may have some known, previously or currently determined pharmacological phenotypes while having an unknown pharmacological phenotype that will be predicted by the pharmacological phenotype prediction system.
[0050] In general, the technology for identifying pharmacological phenotypes based on omics, sociomics, physiological omics and environmental characteristics can be implemented in one or more client devices, one or more network servers or a system comprising a combination of these devices. However, for clarity, the examples below focus on one embodiment, in which a pharmacological phenotype assessment server obtains a set of training data. In some embodiments, training data can be obtained from a client device. For example, a healthcare provider can obtain a biological sample for measuring the omics of a patient (e.g., from saliva, cheek swabs, sweat, skin samples, biopsies, blood samples, urine, feces, sweat, lymph, bones, bone marrow, hair, odor, etc.), and provide the laboratory results obtained by analyzing the biological sample to the pharmacological phenotype assessment server.
[0051] exist Figure 5 An example process 500 for generating omics data from a patient's biological sample is shown in FIG. The process can be performed by an analytical laboratory or other suitable institution. At box 502, a healthcare provider obtains a biological sample from a patient and sends it to an analytical laboratory for analysis. The biological sample can include the patient's saliva, sweat, skin, blood, urine, feces, sweat, lymph, bone marrow, hair, cheek cells, odor, etc. At box 504, cells are then extracted from the biological sample and reprogrammed into stem cells, such as induced pluripotent stem cells (iPSCs), at box 506. The iPSCs are then differentiated into various tissues, such as neurons, cardiomyocytes, etc., at box 508, and analyzed at box 510 to obtain omics data. The omics data can include genomic data, epigenomic data, transcriptomic data, proteomic data, chromosomic data, metabolomic data, and / or biological networks. As described below with reference to Figures 4A-4C In more detail, SNP, gene and genomics region can be identified as being relevant to a specific pharmacological phenotype. When analyzing the patient's omics, socio-omics, physiological omics and environmental data for a specific pharmacological phenotype or a group of pharmacological phenotypes (e.g., indicating the pharmacological phenotype of the reaction to valproic acid), iPSC can be analyzed for the SNP, gene and genomics region identified that are relevant to the specific pharmacological phenotype. More generally, the omics data to be analyzed can be selected based on the omics data that are identified as being relevant to the pharmacological phenotype set being examined with the patient.
[0052] More specifically, by introducing transcription factors or " reprogramming factors " or other reagents into a given cell type, cells are reprogrammed into iPSC. For example, cells can be reprogrammed into iPSC using Yamanaka factors (comprising transcription factors Oct4, Sox2, cMyc and Klf4). iPSC can then be differentiated into various tissues, such as neurons, adipocytes, cardiomyocytes, pancreatic beta cells, etc. After differentiation of iPSC, various analytical techniques (such as DNA methylation analysis, DNAse footprint analysis, filter binding analysis, etc.) can be used to analyze the differentiated iPSC to identify epigenomic information. In fact, the pharmacological phenotype prediction system performs a virtual biopsy, and the differentiated iPSC has the phenotype and epigenomic characteristics of its corresponding tissue at least to a certain extent.
[0053] In the above-described embodiment, cells are extracted from the patient's biological sample, reprogrammed into stem cells, differentiated into various tissues, and analyzed to obtain group data (differentiation, reprogrammed cell analysis method). Alternatively, in certain embodiments, the patient's biological sample is measured without extracting cells (cell-free analysis method). In other embodiments, cells are extracted from the patient's biological sample, and analyzed (primary cell analysis method) when the cell is not reprogrammed or differentiated. In other embodiments, cells are reprogrammed into iPSC, and analyzed (reprogrammed stem cell analysis method) when the cell is not differentiated. For example, iPSC can be analyzed to obtain stem cell group studies without differentiation. Although these are just some example processes for generating group data from the patient's biological sample, analysis can be performed at any suitable stage in the process, and group data can be generated in any suitable manner.
[0054] Healthcare providers can also obtain physiological measurements including vital signs, sleep cycles, circadian rhythms, etc. In addition, healthcare providers can obtain data related to pharmaco-metabolomics, including metabolites that are products of metabolism (such as acetic acid, lactic acid, etc.) and pharmaco-metabolomics metabolites of drugs. Metabolites can be identified by spectroscopy or spectroscopy performed on biological samples of patients in the laboratory, for example, and the results can be provided to healthcare providers as a metabolic profile of the patient. The metabolic profile can then be used to identify metabolic disease characteristics, identify compounds that can alter drug responses, identify metabolite variables and map the metabolite variables to known metabolic and biological pathways, etc.
[0055] In certain embodiments, the pharmacological phenotype prediction system can utilize drug metabolomics data, and the drug metabolomics data comprises a systematic assessment of the presence or absence and / or quantitative levels of multiple drugs and drug metabolites. Such information can be collected from whole blood, citrated blood, blood spots, other tissues and body fluids, etc. The pharmacological phenotype prediction system can utilize one or more drug metabolomics data instances pre-existing in the EHR system or other databases, and / or data for current treatment or pharmacological phenotype prediction inquiries. Data of prescription drugs, over-the-counter drugs, non-prescription drugs, illegal drugs, etc. can be collected simultaneously. The concentration of drugs and metabolites can be measured by techniques including mass spectrometry and other forms of spectroscopy and spectroscopy and / or nuclear magnetic resonance, antibodies and affinity testing, etc. Such information can be used in embodiments to detect drug abuse or off-label use, measure compliance with prescribed medications, detect other prescription or over-the-counter medications used by the patient or prescribed at other clinics, to assess the patient's metabolite status and other purposes, and to make treatment recommendations, including prescribing, discontinuing and substituting medications, as well as changes in dosage and regimen, mode of administration, monitoring, testing and diagnosis, specialist referral, additional diagnostics, other treatment methods, etc.
[0056] In other embodiments, physiological measurements can be obtained from the patient's client computing device, fitness tracker, or quantified self-report / passive reporting method. In another example, the healthcare provider can obtain a patient questionnaire (including questions about the patient's demographics, medical history, socioeconomic status, law enforcement history, sleep cycle, circadian rhythm, etc.), and the results of the patient questionnaire can be provided to the pharmacological phenotyping assessment server. Training data can be obtained from the electronic medical record (EMR) located on the EMR server and / or from the multi-pharmacy data located on the multi-pharmacy server, which aggregates the pharmacy data of patients from multiple pharmacies. In some embodiments, training data can be obtained from a combination of sources including several servers (e.g., EMR servers, multi-pharmacy servers, etc.) and the client devices of the healthcare provider and the patient. For example, training data for a specific patient can be obtained by cross-referencing the patient's personal historical data (e.g., the patient's occupation, place of residence, etc.) with more extensive longitudinal data on these characteristics (e.g., data in the Human Exposome Project).
[0057] In addition to providing training data to the pharmacological phenotype assessment server including omics data for training patients with known pharmacological phenotypes, the pharmacological phenotype assessment server also obtains consortium omics data of baseline omics levels, omics distributions, or any other suitable omics data that can be used to train the pharmacological phenotype assessment server.
[0058] In any case, the subset of training data can be associated with the training patient corresponding to the subset of training data. Furthermore, for example, the pharmacological phenotype assessment server can assign the subset of training patients and the corresponding training data to cohorts based on demographics. The training data can then be used to train the pharmacological phenotype assessment server to generate a statistical model for predicting the patient's pharmacological phenotype. Various machine learning techniques can be used to train the pharmacological phenotype assessment server.
[0059] After the pharmacological phenotype assessment server is trained, the omics data, sociomics data, physiological omics data and environmental data of the current patient whose pharmacological phenotype is unknown, which may be collected at multiple time points, can be received. In some embodiments, the pharmacological phenotype assessment server can obtain an indication of the disease or condition suffered by the current patient to identify the best drug for treating each disease. This can include stress-related diseases such as post-traumatic stress disorder (PTSD), depression, suicidal tendencies, circadian rhythm disorders, substance abuse disorders, phobias, stress ulcers, acute stress disorders, and stress-related diseases included in the Oxford Handbook of Psychiatry. The disease or condition suffered by the current patient can also include bipolar disorder, schizophrenia, autism spectrum disorder and attention deficit hyperactivity disorder (ADHD). In addition, this may include generalized anxiety disorder and anxiety-depressive disorders and non-psychiatric complications such as irritable bowel syndrome (IBS), inflammatory bowel disease (IBD), Crohn's disease, gastritis, gastric and duodenal ulcers, and gastroesophageal reflux disease (GERD). In addition, the disease or condition currently suffered by the patient may include heart disease, fibromyalgia, chronic fatigue syndrome, etc. The pharmacological phenotype of these diseases or conditions may include the pharmacological phenotype associated with any current and future drugs and / or other methods for treating the corresponding disease or condition.
[0060] For example, various machine learning techniques can then be used to analyze omics, sociomics, physiomic, and environmental data to predict one or more pharmacological phenotypes for the patient. An indication of the pharmacological phenotype can be transmitted to a healthcare provider's client device for review by the healthcare provider and determination of an appropriate course of treatment based on the pharmacological phenotype. The pharmacological phenotype can be predicted in clinical settings as well as in research settings for drug development and insurance applications. In a research setting, the pharmacological phenotype of a potential patient cohort associated with an experimental drug in a research program might be predicted. Patients can be selected for the experimental treatment based on their predicted pharmacological phenotype associated with the experimental drug.
[0061] Reference Figure 1A, an example pharmacological phenotype prediction system 100 uses various machine learning techniques to predict a patient's pharmacological phenotype (precise patient phenotype) based on the patient's omics, sociomics, physiological omics, and environmental data. The pharmacological phenotype prediction system 100 can obtain training data for a training patient cohort, and these data can be analyzed to identify the relationship between the omics, sociomics, physiological omics, and environmental data and the pharmacological phenotype contained in the training data. The pharmacological phenotype prediction system 100 can then generate a statistical model for predicting the pharmacological phenotype based on the analysis. When the patient's pharmacological phenotype is unknown (for example, lithium has not yet been prescribed to a bipolar disorder patient, so it is unclear how the patient responds to lithium), the pharmacological phenotype prediction system 100 can obtain the patient's omics, sociomics, physiological omics, and environmental data and apply the omics, sociomics, physiological omics, and environmental data to the statistical model to predict the patient's pharmacological phenotype. For example, the pharmacological phenotype prediction system 100 can predict the likelihood that a patient will have an adverse reaction to a particular drug, can predict the efficacy or appropriate dosage of the drug, etc. The pharmacological phenotype prediction system 100 can be used to perform clinical decision support (CDSS) in a clinical setting to predict the patient's precise patient phenotype. In addition, the pharmacological phenotype prediction system 100 can be used to conduct drug research to develop supporting diagnostic tests to identify patients who will have a good or bad response to the developed or approved drug and will have fewer or no side effects. In addition, the pharmacological phenotype prediction system 100 can be used in the context of experimental treatment to recommend to researchers experimental drugs and / or dosages to be prescribed to the current patient in the context of clinical research.
[0062] The pharmacological phenotype prediction system 100 includes a pharmacological phenotype evaluation server 102 and a plurality of client devices 106-116 that can be communicatively connected via a network 130, as described below. In one embodiment, the pharmacological phenotype evaluation server 102 and the client devices 106-116 can communicate via wireless signals 120 over a communication network 130, which can be any suitable local area network or wide area network, including a WiFi network, a Bluetooth network, a cellular network (such as 3G, 4G, long term evolution (LTE), 5G), the Internet, etc. In some cases, the client devices 106-116 can communicate with the communication network 130 via an intervening wireless or wired device 118, which can be a wireless router, a wireless repeater, a base station transceiver of a mobile phone provider, etc. For example, the client devices 106-116 can include a tablet computer 106, a smart watch 107, a network-enabled cellular phone 108, a wearable computing device (such as Google Glass TM or 109 ), a personal digital assistant (PDA) 110 , a mobile device smartphone 112 (also referred to herein as a “mobile device”), a laptop computer 114 , a desktop computer 116 , a wearable biosensor, a portable media player (not shown), a tablet phone, any device configured for wired or wireless RF (radio frequency) communication, etc. In addition, any other suitable client device that records a patient's omics data, clinical data, demographic data, multi-pharmacy data, socio-omics data, physiological omics data, or other environmental data can also communicate with the pharmacological phenotype assessment server 102 .
[0063] In some embodiments, the patient may enter data into the desktop computer 116, such as answers to a patient questionnaire containing questions regarding the patient's demographics, medical history, socioeconomic status, law enforcement history, sleep cycles, circadian rhythms, etc. In other embodiments, a healthcare provider may enter the data.
[0064] Each of the client devices 106-116 can interact with the pharmacological phenotyping server 102 to send the patient's omics data, clinical data, demographic data, multi-pharmacy data, socio-omics data, physio-omics data, or other environmental data. In some embodiments, socio-omics, physio-omics, and environmental data can be collected periodically (e.g., monthly, every three months, every six months, etc.) to identify changes in the patient's sociological status and environment over time (e.g., from unemployed to employed, single to married, etc.). Similarly, in some embodiments, at least some of the patient's socio-omics, physio-omics, and environmental data can be recorded by a healthcare provider via the healthcare provider's client device 106-116, or can be self-reported via the patient's client device 106-116.
[0065] Each client device 106-116 may also interact with the pharmacological phenotype assessment server 102 to receive one or more indications of a predicted pharmacological phenotype for the current patient. The indication may include a recommendation for a drug to be prescribed to the current patient for which the current patient has the highest expected response (e.g., the highest combination of efficacy and minimal adverse drug reactions and severity of reactions). The indication may also include the current patient's risk for various diseases, such as the likelihood of developing the disease, a risk category (e.g., low, medium, or high risk), etc. In addition, the indication may include the likelihood of drug abuse, such as a numerical likelihood or a likelihood category (e.g., low, medium, or high likelihood).
[0066] In an example embodiment, the pharmacological phenotype assessment server 102 can be a cloud-based server, an application server, a web server, etc., and includes a memory 150, one or more processors (CPUs) 142 (such as microprocessors coupled to the memory 150), a network interface unit 144, and an I / O module 148, which can be, for example, a keyboard or a touch screen.
[0067] The pharmacological phenotype assessment server 102 can also be communicatively connected to a consortium omics / environmental / physiological omics / demographics / pharmacy information database 154. The consortium omics / environmental / physiological omics / demographics / pharmacy information database 154 can store training data and statistical models for determining pharmacological phenotypes, wherein the training data includes omics data of training patients, whole genome-based ethnic data, clinical data, demographic data, multiple pharmacy data, socio-omics data, physiological omics data, or other environmental data. The consortium omics / environmental / physiological omics / demographics / pharmacy information database 154 can also include a consortium omics database and an academic omics database and a pharmacy database, including (for example) RxNorm, drug-drug interactions (such as FDA black box labeling), drug-gene interactions, and others. In some embodiments, to determine the pharmacological phenotype, the pharmacological phenotype assessment server 102 can retrieve patient information for each training patient from the consortium omics / environmental / physiological omics / demographics / pharmacy information database 154.
[0068] The memory 150 may be a tangible, non-transitory memory and may include any type of suitable memory module, including random access memory (RAM), read-only memory (ROM), flash memory, other types of persistent memory, etc. The memory 150 may store, for example, instructions for an operating system (OS) 152 that can be executed on the processor 142, which may be any type of suitable operating system, such as a modern smartphone operating system. The memory 150 may also store, for example, instructions for a machine learning engine 146 that can be executed on the processor 142, which may include a training module 160 and a phenotype assessment module 162. Figure 1B The pharmacological phenotype assessment server 102 is described in greater detail. In some embodiments, the machine learning engine 146 can be part of one or more of the client devices 106-116, the pharmacological phenotype assessment server 102, or a combination of the pharmacological phenotype assessment server 102 and the client devices 106-116.
[0069] In any case, the machine learning engine 146 can receive electronic data from the client devices 106-116. For example, the machine learning engine 146 can obtain a set of training data by receiving omics data, clinical data, demographic data, multi-pharmacy data, socio-omics data, physio-omics data, or other environmental data. In addition, the machine learning engine 146 can obtain a set of training data by receiving phenotypic data related to the pharmacological phenotype of the training patients (e.g., chronic diseases suffered by the training patients, responses to medications previously prescribed to the training patients, whether each of the training patients has a substance abuse problem, etc.).
[0070] Thus, the training module 160 can classify the omics data, sociomic data, physiological data, and environmental data into specific pharmacological phenotypes, such as drug abuse, specific types of chronic diseases, adverse drug reactions to specific drugs, the level of efficacy of specific drugs, etc. The training module 160 can then analyze the classified omics data, sociomic data, physiological data, and environmental data to generate a statistical model for each pharmacological phenotype. For example, a first statistical model can be generated to determine the likelihood that a current patient will experience a drug abuse problem, a second statistical model can be generated to determine the risk of developing a disease, a third statistical model can be generated to determine the risk of developing another disease, a fourth statistical model can be generated to determine the likelihood of having a negative reaction to a specific drug, and so on. In some embodiments, each statistical model can be combined in any suitable manner to generate an overall statistical model for predicting each of the pharmacological phenotypes.In any case, the training data set can be analyzed using a variety of machine learning techniques, including but not limited to regression algorithms (e.g., ordinary least squares regression, linear regression, logistic regression, stepwise regression, multivariate adaptive regression splines, local estimated scatter plot smoothing, etc.), instance-based algorithms (e.g., k-nearest neighbors, learning vector quantization, self-organizing maps, local weighted learning, etc.), regularization algorithms (e.g., ridge regression, least absolute shrinkage and selection operators, elastic net, least angle regression, etc.), decision tree algorithms (e.g., classification and regression trees, iterative bisection3, C4.5, C5, chi-square automatic interaction detection, decision stumps, M5, conditional decision trees, etc.), clustering algorithms (e.g., k-means, k-medians, expectation maximization, hierarchical clustering, etc.), , spectral clustering, mean shift, density-based spatial clustering of applications with noise, sorting points for identifying cluster structures, etc.), association rule learning algorithms (e.g., prior algorithms, Eclat algorithms, etc.), Bayesian algorithms (e.g., naive Bayes, Gaussian naive Bayes, multinomial naive Bayes, averaged single dependency estimator, Bayesian belief network, Bayesian network, etc.), artificial neural networks (e.g., perceptron, Hopfield network, radial basis function network, etc.), deep learning algorithms (e.g., multilayer perceptron, deep Boltzmann machine, deep belief network, convolutional neural network, stacked autoencoder, generative adversarial network, etc.), dimensionality reduction algorithms (e.g., principal component analysis, principal component regression, partial least squares regression, Sammon map, etc.), mapping), multidimensional scaling, projection pursuit, linear discriminant analysis, hybrid discriminant analysis, quadratic discriminant analysis, flexible discriminant analysis, factor analysis, independent component entity analysis, non-negative matrix factorization, t-distributed random neighbor embedding, etc.), ensemble algorithms (e.g., boosting, bootstrap aggregation, AdaBoost, stacked generalization, gradient boosting machines, gradient boosted regression trees, random decision forests, etc.), reinforcement learning (e.g., temporal difference learning, Q-learning, learning automata, state-action-reward-state-action, etc.), support vector machines, hybrid models, evolutionary algorithms, probabilistic graphical models, etc.
[0071] During the testing phase, the training module 160 can compare the test patient's test-omics data, sociomics data, physiomic data, and environmental data with the statistical model to determine the likelihood that the test patient has a specific pharmacological phenotype.
[0072] If the training module 160 makes the correct judgment more often than the predetermined threshold amount, the statistical model can be provided to the phenotype assessment module 162. On the other hand, if the training module 160 does not make the correct judgment more often than the predetermined threshold amount, the training module 160 can continue to acquire training data for further training.
[0073] Phenotype assessment module 162 can obtain statistical models and a set of omic data, sociomics, physiological omics and environmental data of the current patient, and the data can be collected over a period of time (e.g., one month, three months, six months, one year, etc.). For example, biological samples (e.g., blood samples, saliva, biopsies, bone marrow, hair, etc.) of the current patient can be analyzed in a laboratory to obtain the genomic data, epigenomic data, transcriptomic data, proteomic data, genomic data and / or metabolomic data of the current patient. The omic data can then be provided to phenotype assessment module 162. In addition, the clinical data of the patient can be provided from the EMR server or client device 106-116 of the healthcare provider. Multi-pharmacy data can be provided from multiple pharmacy servers or from several pharmacy servers, and demographic data, sociomics data, physiological omics data and other environmental data can be provided from the client device 106-116 of the healthcare provider or the client device 106-116 of the current patient.
[0074] The omics, sociomics, physiomic, and environmental data can then be applied to the statistical model generated by the training module 160. Based on the analysis, the phenotype assessment module 162 can determine a probability or other semi-quantitative and quantitative measure indicating that the current patient has certain pharmacological phenotypes, such as the likelihood of drug abuse, the likelihood of various diseases, an overall rating of predicted response to various drugs, etc. The phenotype assessment module 162 can cause the likelihood to be displayed on a user interface for review by a healthcare provider. Each likelihood can be expressed as a probability (e.g., 0.6), a percentage (e.g., 80%), one of a set of categories (e.g., "high," "medium," or "low"), and / or in any other suitable manner.
[0075] The pharmacological phenotype assessment server 102 can communicate with the client devices 106-116 via a network 130. The digital network 130 can be a private network, a secure public internet, a virtual private network, and / or some other type of network, such as a dedicated access line, an ordinary conventional telephone line, a satellite link, a combination of these, etc. In the case where the digital network 130 includes the Internet, data communication can be performed over the digital network 130 using Internet communication protocols.
[0076] Now go to Figure 1B, the pharmacological phenotype assessment server 102 may include a controller 224. The controller 224 may include a program memory 226, a microcontroller or microprocessor (MP) 228, a random access memory (RAM) 230, and / or input / output (I / O) circuits 234, all of which may be interconnected via an address / data bus 232. In some embodiments, the controller 224 may also include a database 239, or be otherwise communicatively connected to the database or other data storage mechanism (e.g., one or more hard disk drives, optical storage drives, solid-state storage devices, etc.). The database 239 may include data such as patient information, training data, risk analysis templates, web page templates and / or web pages, and other data necessary for interaction with a user via the network 130. The database 239 may include the same data as described above with reference to Figure 1A The combined omics / environmental / physiomic / demographic / pharmacy information database 154 described and / or referenced below Figure 3 The data sources 325a-d described (eg, biomedical training set 325a, pharmacology database 325b, environmental data 325c, and granularly segmented data 325d) are similar data.
[0077] It should be understood that although Figure 1B Only one microprocessor 228 is depicted, but the controller 224 may include multiple microprocessors 228. Similarly, the memory of the controller 224 may include multiple RAMs 230 and / or multiple program memories 226. Although Figure 1B I / O circuitry 234 is depicted as a single block, but I / O circuitry 234 may include many different types of I / O circuitry. Controller 224 may implement one or more RAMs 230 and / or program memories 226 as, for example, semiconductor memory, magnetically readable memory, and / or optically readable memory.
[0078] like Figure 1B As shown, the program memory 226 and / or RAM 230 can store various applications for execution by the microprocessor 228. For example, the user interface application 236 can provide a user interface 102 to the pharmacological phenotype assessment server, which can, for example, allow a system administrator to configure, troubleshoot, or test various aspects of the server's operation. The server application 238 can be operable to receive a set of omics data, socio-omics data, physio-omics data, and environmental data for a current patient, determine a likelihood or other semi-quantitative and quantitative metrics indicating that the current patient has a pharmacological phenotype, and send an indication of the likelihood to the healthcare provider's client device 106-116. The server application 238 can be a single module 238 or multiple modules 238A, 238B, such as the training module 160 and the phenotype assessment module 162.
[0079] Despite Figure 1B 238B, the server application 238 may include any number of modules that perform tasks associated with the implementation of the pharmacological phenotype assessment server 102. Figure 1B Only one pharmacological phenotype evaluation server 102 is depicted in FIG, but multiple pharmacological phenotype evaluation servers 102 may be provided for distributing server load, serving different web pages, etc. These multiple pharmacological phenotype evaluation servers 102 may include web servers, entity-specific servers (e.g., servers, etc.), servers located in retail or private networks, etc.
[0080] Now refer to Figure 1C , the laptop computer 114 (or any of the client devices 106-116) may include a display 240, a communication unit 258, a user input device (not shown), and, like the pharmacological phenotype assessment server 102, a controller 242. Similar to the controller 224, the controller 242 may include a program memory 246, a microcontroller or microprocessor (MP) 248, a random access memory (RAM) 250, and / or input / output (I / O) circuits 254, all of which may be interconnected via an address / data bus 252. The program memory 246 may include an operating system 260, a data storage device 262, a plurality of software applications 264, and / or a plurality of software routines 268. For example, the operating system 260 may include a Microsoft OS The data storage device 262 may contain data such as patient information, application data for a plurality of applications 264, routine data for a plurality of routines 268, and / or other data necessary for interaction with the pharmacological phenotype assessment server 102 via the digital network 130. In some embodiments, the controller 242 may also include other data storage mechanisms (e.g., one or more hard disk drives, optical storage drives, solid-state storage devices, etc.) residing within the laptop computer 114 or otherwise be communicatively connected thereto.
[0081] The communication unit 258 can communicate with the pharmacological phenotype assessment server 102 via any suitable wireless communication protocol network, such as a wireless telephone network (e.g., GSM, CDMA, LTE, etc.), a Wi-Fi network (802.11 standard), a WiMAX network, a Bluetooth network, etc. The user input device (not shown) can include a "soft" keyboard displayed on the display 240 of the laptop computer 114, an external hardware keyboard (e.g., a Bluetooth keyboard) that communicates via a wired or wireless connection, an external mouse, a microphone for receiving voice input, or any other suitable user input device. As discussed with reference to the controller 224, it should be understood that although Figure 1C Only one microprocessor 248 is depicted, but the controller 242 may include multiple microprocessors 248. Similarly, the memory of the controller 242 may include multiple RAMs 250 and / or multiple program memories 246. Although Figure 1C I / O circuitry 254 is depicted as a single block, but I / O circuitry 254 may include many different types of I / O circuitry. Controller 242 may implement one or more RAMs 250 and / or program memories 246 as, for example, semiconductor memory, magnetically readable memory, and / or optically readable memory.
[0082] The one or more processors 248 may be adapted and configured to execute, among other software applications, any one or more of a plurality of software applications 264 and / or any one or more of a plurality of software routines 268 resident in the program memory 246. One of the plurality of applications 264 may be a client application 266, which may be implemented as a series of machine-readable instructions for performing various tasks associated with receiving information at, displaying information on, and / or sending information from the laptop computer 114.
[0083] An application in the plurality of applications 264 may be a native application and / or a web browser 270 (eg, Apple's Google Chrome TM 、Microsoft Internet and Mozilla ), the local application and / or web browser may be implemented as a series of machine-readable instructions for receiving, interpreting, and / or displaying web page information from the pharmacological phenotype assessment server 102, while also receiving input from a user, such as a healthcare provider. Another application among the plurality of applications may include an embedded web browser 276, which may be implemented as a series of machine-readable instructions for receiving, interpreting, and / or displaying web page information from the pharmacological phenotype assessment server 102.
[0084] One of the plurality of routines may include a risk analysis display routine 272 that obtains a likelihood that the current patient has a certain pharmacological phenotype and displays the likelihood and / or an indication of a recommendation for treating the current patient on the display 240. Another of the plurality of routines may include a data input routine 274 that obtains socio-omic, physio-omic, and environmental data for the current patient from a healthcare provider and sends the received socio-omic, physio-omic, and environmental data to the pharmacological phenotype assessment server 102 along with previously stored socio-omic, physio-omic, and environmental data for the current patient (e.g., environmental data collected during a previous visit).
[0085] Preferably, the user can launch the client application 266 from a client device (e.g., one of the client devices 106-116) to communicate with the pharmacological phenotype assessment server 102, thereby implementing the pharmacological phenotype prediction system 100. Alternatively, the user can launch or instantiate any other suitable user interface application (e.g., a native application or web browser 270, or any other application in the plurality of software applications 264) to access the pharmacological phenotype assessment server 102, thereby implementing the pharmacological phenotype prediction system 100.
[0086] As mentioned above, Figure 1A The illustrated pharmacological phenotype assessment server 102 can include a memory 150 that can store instructions for a machine learning engine 146 that can be executed on a processor 142. The machine learning engine 146 can include a training module 160 and a phenotype assessment module 162.
[0087] Figure 2The omics, sociomics, physiomic, and environmental data that can be provided to the pharmacological phenotype prediction system 100 are shown, which in turn predicts the pharmacological phenotype in a clinical or research setting. The omics, sociomics, physiomic, and environmental data are divided into four categories: individual / cohort and population omics and drug metabolomics 302; exposure group 304; sociomics demographics and stress / trauma 306; and medical physiomic, structured or unstructured electronic health records (EHR), laboratory values, stress and abuse factors and trauma and medical outcome data 308. However, this is for ease of illustration only. Exposure group 304, social omics demographics and stress / trauma 306, and medical physiology group, structured or unstructured EHR, laboratory values, stress and abuse factors and trauma and medical outcome data 308 can be included as part of the social omics, physiological omics and environmental data, while individual / cohort and population omics and drug metabolomics 302 can be included as part of the omics data. In addition, individual / cohort and population omics and drug metabolomics 302, exposure group 304, social omics demographics and stress / trauma 306, and medical physiology group, structured or unstructured EHR, laboratory values, stress and abuse factors and trauma and medical outcome data 308 can be categorized and / or organized in any other suitable manner.
[0088] In any case, individual / cohort and population omics and drug metabolomics 302 can include genomics, epigenomics, chromatin states, transcriptomics, proteomics, metabolomics, biological networks and system models, etc., each of which can be extracted from or at least related to the genome. Individual / cohort and population omics and drug metabolomics 302 can also include chemical mapping of discrete molecular entities within the tissue to various pharmacological phenotypes. Discrete molecular entities can be metabolites (such as acetic acid, lactic acid, etc.) as products of metabolism and drug metabolome metabolites of drugs.
[0089] The exposure group 304 may include information indicating the patient's environment, such as the location of the patient's residence, the type of residence, the size of the residence, the quality of the residence, the patient's work environment (including the location of the patient's workplace), the distance from the patient's residence to the workplace, how the patient is treated at the workplace and / or residence, etc. The exposure group 304 may also include any other environmental exposures experienced by the patient, including climatic factors, lifestyle factors (e.g., tobacco, alcohol), diet, physical activity, pollutants, radiation, infection, education level, etc.
[0090] Socio-omics, demographics and stress / trauma 306 can include demographic data such as gender, ancestry, age, income, marital status, education level, language, etc. Socio-omics, demographics and stress / trauma 306 can also include other family data, cultural conditions, circadian rhythm data, age-related health conditions, isolation or cognitive conditions, economic and community living conditions, etc. In addition, socio-omics, demographics and stress / trauma 306 can include trauma, domestic violence, law enforcement history, or any other stress or abuse factors. In some embodiments, stress and abuse factors in childhood can be quantified by adverse childhood experiences (ACE) scores, which assess different types of abuse, neglect, and other measures during difficult childhoods. This can include physical, emotional and sexual abuse, physical and emotional neglect, mental illness within the family, domestic violence within the family, divorce, substance abuse within the family, imprisoned relatives, etc.
[0091] In addition, the medical physiology, structured or unstructured EHR, laboratory values, stress and abuse factors and trauma, and medical outcome data 308 can include trauma, domestic violence, law enforcement history, or any other stress or abuse factors. In some embodiments, childhood stress and abuse factors can be quantified using the Adverse Childhood Experiences (ACE) score, which assesses different types of abuse, neglect, and other measures of difficult childhood. This may include physical, emotional, and sexual abuse, physical and emotional neglect, mental illness within the family, domestic violence within the family, divorce, substance abuse within the family, incarcerated relatives, etc. The medical physiology, structured or unstructured EHR, laboratory values, stress and abuse factors and trauma, and medical outcome data 308 can also include clinical data, multi-pharmacy data, and physiological characteristics, such as human body functions related to genes and proteins. In addition, the medical outcome data can include the pharmacological phenotype of a specific patient or patient cohort. In addition, the medical outcome data can include information indicating drug or treatment efficacy, adverse drug events or adverse drug reactions, stable disease, lack of response, improvement in clinical response, etc.
[0092] The individual / cohort and population omics and pharmaco-metabolomics 302, exposure group 304, socio-omics demographics and stress / trauma 306, and medical physiology omics, structured or unstructured EHR, laboratory values, stress and abuse factors and trauma, and medical outcome data 308 of the training patient cohort can be provided as training data to the pharmacological phenotype prediction system 100 to generate a statistical model for predicting the pharmacological phenotype. In addition, the individual omics and pharmaco-metabolomics 302, exposure group 304, socio-omics demographics and stress / trauma 306, and medical physiology omics, structured or unstructured EHR, laboratory values, stress and abuse factors and trauma, and medical outcome data 308, or some portion thereof, can be obtained from the current patient to be applied to the statistical model to predict the pharmacological phenotype of the current patient or the precise patient phenotype.
[0093] Figure 3 A detailed view 320 of the process performed by the pharmacological phenotype prediction system 100 is shown. Figure 2 As shown, the pharmacological phenotype prediction system 100 obtains individual / cohort and population omics and drug metabolomics 302, exposure group 304, social group demographics and stress / trauma 306, and medical physiological group, structured or unstructured EHR, laboratory values, stress and abuse factors and trauma and medical outcome data 308 from the training patient cohort as training data to train the machine learning engine 146. In some examples, individual / cohort and population omics and drug metabolomics 302 and their respective correlations with pharmacological phenotypes can be obtained from GWAS, candidate gene association studies and / or other machine learning methods, as described in more detail below. Training data can also be obtained from several data sources 325a-d including biomedical training sets 325a, pharmacological databases 325b, environmental data 325c, and data segmented by granularity 325d.
[0094] Biomedical training set 325a includes the same Figure 2 The omics and drug metabolomics 302 described above and the medical physiology, structured or unstructured EHR, laboratory values, stress and abuse factors and trauma and medical outcome data 308 are similar to the omics, drug metabolomics, medical physiology, EHR, laboratory values, medical outcomes and stress and abuse factors and trauma. The pharmacology database 325b contains pharmacy records, drug databases, drug-drug interactions, drug-gene interactions, etc. In addition, the socio-omics and environmental data 325c contains the same as those described above. Figure 2Similar socio-omics, demographics, and exposomes are described for exposomes 304, socio-omics demographics, and stress / trauma 306. Granular data 325d can identify any data from the biomedical training set 325a, pharmacology database 325b, and environmental data 325c corresponding to a single patient, a cohort of patients, or a group of patients, respectively.
[0095] The machine learning engine 146 can then utilize the omics, sociomics, physiological and environmental, and phenotypic data of a group of training patients or a cohort of training patients to generate a statistical model for predicting pharmacological phenotypes using machine learning techniques. In some embodiments, the machine learning engine 146 can analyze the relationship between omics data and pharmacological phenotypes to identify single nucleotide polymorphisms (SNPs), genes, and genomic regions that are highly correlated with specific pharmacological phenotypes. Figure 4B and 4D This is discussed in more detail.
[0096] In addition, the machine learning engine 146 can classify a team of training patients or a group of training patients with at least some identified SNPs, genes and genomic regions or any suitable combination thereof as having a specific pharmacological phenotype or not having a specific pharmacological phenotype based on the phenotypic data of each training patient in the cohort or population. The machine learning engine 146 can further analyze the sociogenomics, physiognomics and environmental data of a team of training patients or a group of training patients corresponding to each category to generate a statistical model. For example, the machine learning engine 146 can perform statistical measurements on the sociogenomics, physiognomics and environmental data of each category to distinguish between the sociogenomics, physiognomics and environmental data of a subset of training patients with a specific pharmacological phenotype and a subset of training patients without a specific pharmacological phenotype. Supervised learning algorithms (such as classification and regression) can be used to train the machine learning engine 146. Unsupervised learning algorithms (such as dimensionality reduction and clustering) can also be used to train the machine learning engine 146.
[0097] In any case, the machine learning engine 146 can receive input 330 from the current patient or current patient group without knowing whether the current patient or current patient group has a specific pharmacological phenotype. The input can include any of the above-mentioned individual omics and drug metabolomics 302, exposure group 304, social omics demographics and stress / trauma 306, and medical physiology, structured or unstructured EHR, laboratory values, stress and abuse factors and trauma, and medical outcome data 308. For example, for a single current patient, the input can include personal data, omics laboratory tests, personal physiology, EHR data, medication history, and environmental data. For a group of current patients, the input can include personal data, cohort omics, physiology, EHR data, medication history, and environmental data.
[0098] The omics, sociomics, physiomic, and environmental data for the current patient or the current group of patients can be applied to the statistical model contained in the machine learning engine 146 to predict the pharmacological phenotype of the current patient or the current group of patients. For example, the machine learning engine 146 can predict the likelihood of a negative response to warfarin. Additionally or alternatively, the machine learning engine 146 can generate a response score that indicates the efficacy of warfarin in treating thrombosis in the current patient, where the efficacy is discounted by the adverse effects of warfarin on the current patient.
[0099] The pharmacological phenotype clinical decision support engine 335 within the pharmacological phenotype prediction system 100 can analyze the likelihood, response score, or other semi-quantitative or quantitative metrics to recommend to the healthcare provider the medication and / or dosage to be prescribed to the current patient. In some embodiments, the response score of each medication available for a specific medical indication can be ranked, and the pharmacological phenotype clinical decision support engine 335 can recommend the highest-ranked medication to the healthcare provider to be prescribed to the current patient. The dosage of a specific medication can also be ranked. In another example, when the likelihood of a negative reaction to a medication selection is higher than a threshold score for the current patient, the pharmacogenomics clinical decision support engine 335 can recommend a different medication for the specific medical indication. In addition, the pharmacological phenotype clinical decision support engine 335 can compare the recommended medication with the current patient's multi-pharmacy data contained in the environmental data and / or medical records. If the patient is currently taking a medication that is incompatible with the recommended medication, the pharmacological phenotype clinical decision support engine 335 may recommend a medication with the next highest response score or another medication whose likelihood of a negative reaction is no greater than a threshold likelihood. In other embodiments, when the pharmacological phenotype is a likelihood of drug abuse, the pharmacogenomics clinical decision support engine 335 may recommend early intervention, or when the pharmacological phenotype is a disease risk, the pharmacological phenotype clinical decision support engine 335 may recommend screening and / or treatment options to proactively address the issue. In other embodiments where the pharmacological phenotype differs from that described above, the pharmacological phenotype clinical decision support engine 335 may recommend alternative medications, informative analyses, or other courses of action.
[0100] The likelihood of a pharmacological phenotype, response scores, or other semi-quantitative or quantitative measures can also be used in various capacities in drug research. 340 For example, researchers developing drugs can use this approach to develop companion diagnostic tests to identify patients who will respond favorably or poorly to a drug being developed or approved and who will experience fewer or no side effects. In addition, researchers screening or comparing multiple molecular entities for use as putative drugs and having comparative data from molecular experiments using these drugs can use these methods to prospectively assess the likely effects and adverse events of a population to determine which entities to develop or prioritize during development. Additionally, these methods can be used in the context of experimental treatments to recommend experimental drugs and / or doses to researchers for prescription to current patients in the context of a clinical study. Finally, these methods can be used for explicit construction or model generation of pharmacogenomic tests that will be performed outside of an integrated CDSS environment.
[0101] The predicted pharmacological phenotype and / or recommendation provided by the pharmacological phenotype clinical decision support engine 335 or the drug research tool 340 can be provided to the data source 325a-d in a feedback loop. The omics, sociomics, physiomics, and environmental and phenotypic data of the current patient are then used as training data to further train the machine learning engine 146 for subsequent use with other current patients. In this way, the machine learning engine 146 can continuously update the statistical model to reflect at least near real-time representations of the sociomics, physiomics, environmental, and omics data.
[0102] Figure 4A Describe the biological background of a variety of different omics patterns in the 4D nuclear group and the exemplary representation of the interaction, as well as the example of the bioinformatics analysis of omics data carried out in this context. Diagram 450 depicts the chromosome positioned in the region bound to chromatin in the cell nucleus. Euchromatin is characterized by a specific combination of DNase 1 hypersensitivity and histone marks, which define active genomics regulatory elements, such as promoters H3K4me3 and H3K27ac and enhancers H3K4me1 and H3K27ac. Enhancers can increase or decrease the transcription in their target genes, which can be located proximal to the sequence and / or spatially (e.g., Hi-C or ChIA-PET data, or genome architecture mapping or combined chromatin capture) and / or functionally connected to enhancers individually or in combination (connected by, for example, molecular QTL). Heterochromatin is located in the interior of the chromosome region and the periphery of the nucleus, close to the nuclear lamina and nucleolus. Heterochromatin is characterized by its own repressive chromatin marks and DNA binding proteins, as well as spatial compaction and linker histones. Recent research indicates that in the brain, the DNA sequence CAC is a common site of methylation, in contrast to other tissues where CpG is most commonly methylated. Furthermore, 5-hydroxymethylcytosine (5hmC), a reactive species that carries unique epigenomic information, is relatively common in the brain. In contrast, methylcytosine (hmC) is prevalent in the periphery.
[0103] Figure 4A Also depicted is a schematic diagram of an exemplary spatial hierarchy 460 representing transcriptional organization first determined by a chromatin conformation capture method. The spatial hierarchy 460 includes a Hi-C map 462 showing a multi-scale hierarchy of transcriptional regulation. In this illustration, the normalized frequency of spatial interactions between parts of the genome (on the X-axis) and other parts of the genome (on the Y-axis) is represented by a color gradient to generate a two-dimensional map of chromatin organization.
[0104] This graph can be generated using "bins" representing fixed lengths of DNA sequences, or bins representing cut site increments or their collections, or functional elements (such as genes, chromatin state segments, loop domains, chromatin domains, TADs, etc.). Contacts can be identified using thresholds in various normalized modes for distance, overall contact propensity, and other elements. For example, in bins where sequence length is not fixed, and where a square genomic region described by a pair of bins may have a variable size and shape, a normalization method can be designed to replace the traditional method that relies on a fixed bin. The contact density as a function of distance can be fitted to an integrable function that can be integrated over a rectangular area of the bin pair to produce an expected value of the contact mapped to the square genomic region. The expected value can be compared to the original or standardized read count mapped to the square genomic region using, for example, a statistical test to which the Benjamini false discovery rate can be applied (such as a Poisson distribution p-value) to generate a collection of enriched and depleted chromatin contacts in an adjusted manner at a certain distance within a local or genome-wide range. This can be performed for a variety of analytical purposes, including detecting target genes for genomic variants and performing genome-wide analysis of exposures.
[0105] The spatial hierarchy 460 also includes a visualization 464 of the nuclear and subnuclear transcriptional topology shown in the Hi-C map 462. As shown in the visualization 464, chromosomes fill most of the available volume of the nucleoplasm as regions (CTs) and contain circumscribed A and B compartments composed of euchromatin and heterochromatin, respectively. Active genes tend to be located at the periphery of the CTs, and interchromosomal loops between the CTs provide the basis for spatial interactions between trans-enhancer-promoter and promoter-promoter interactions. The A and B chromatin compartments of the CTs contain topologically associated domains (TADs), the average length of the linear sequence of which is within approximately 1 Mb. TADs can be characterized initially using chromatin conformation capture methods such as Hi-C, in which initial scaling is consistent with a fractal sphere model, while high-resolution studies of enhancer-promoter loops within TADs, the organization of TAD boundary proteins (including CCCTC-binding protein (CTCF) and cohesin (RAD21)), and direct imaging support the loop extrusion model of TAD organization. Characterization of transcriptional units containing frequently interacting regulatory elements (FIREs) includes an example located within an intron of the GRIN2A gene on chromosome 16.
[0106] Machine learning methods can be used Figure 4A Epigenomic tracking and / or bioinformatics analysis as described in the present invention can be used to identify associations between genes, SNPs, and genomic regions and pharmacological phenotypes. For example, epigenomic tracking and / or bioinformatics analysis can be used to generate Figure 4CAs described above, the training module 160 can generate a statistical model for each pharmacological phenotype. Figure 4B is a block diagram illustrating an exemplary method 400 for identifying omics data corresponding to a specific pharmacological phenotype using machine learning techniques. The method 400 can be executed on the pharmacological phenotype assessment server 102. In some embodiments, the method 400 can be implemented in a set of instructions stored on a non-transitory computer readable memory and executable on one or more processors on the pharmacological phenotype assessment server 102. For example, the method 400 can be executed by Figure 1A The training module 160 within the machine learning engine 146 is executed.
[0107] At block 402, a statistical test is performed (e.g., by GWAS or candidate gene association studies) on each of several non-coding or coding SNPs, genes, and genomic regions in the genome to determine a relationship between the SNP and a particular pharmacological phenotype, which can be a drug response, an adverse drug reaction, an adverse drug event, a dosage, a disease risk, etc. (e.g., a patient's response to ketamine for depression). When the statistical test shows a significant relationship between the SNP and the particular pharmacological phenotype (e.g., the p-value is less than a threshold probability using the null hypothesis), then the SNP is determined to be associated with the particular pharmacological phenotype. In some embodiments, the SNP can be associated with the particular pharmacological phenotype based on the p-value of the SNP. Figure 4A Bioinformatics analysis was performed to identify SNPs.
[0108] Then, at box 404, the SNP relevant to specific pharmacological phenotype is carried out linkage disequilibrium analysis, to identify which SNPs are independent of each other. For example, when a group of SNPs is all relevant to identical pharmacological phenotype and is in tight linkage disequilibrium state (for example, LD>0.9), this group of SNPs may be linked, causing unclear which SNPs in this group are the SNPs that produce the correlation with the pharmacological phenotype. Linkage disequilibrium analysis can be carried out to identify each SNP (effect SNP) that may produce the correlation with the pharmacological phenotype. More specifically, linkage disequilibrium analysis can be carried out by comparing SNP (original SNP) with the database of SNP (for example, from 1000 genome projects) to find the SNP linked to the original SNP. In certain embodiments, the ethnic group of GWAS or candidate gene association study can be identified, and SNP and linkage disequilibrium coefficient can be retrieved from the database of the SNP corresponding to the ethnic group identified. The pharmacological phenotype assessment server 102 may then generate a set of allowable candidate variants for all SNPs in close linkage disequilibrium with the original SNP associated with the particular pharmacological phenotype from a GWAS or other candidate association study (block 406 ).
[0109] In addition, the set of permissive candidate variants (box 406) can include body SNPs (box 420) of genes with known or suspected correlations with the pharmacological phenotype being studied, molecular QTLs targeting the gene bodies, and SNPs (box 422) residing in genomic regions or networks with known or suspected correlations with the pharmacological phenotype being studied. In any case, the permissive candidate variants (box 406) can be subjected to bioinformatics analysis to filter the permissive candidate variants into a subset of intermediate candidate variants (box 410), which can then be ranked (e.g., by a scoring system) (box 412). The highest-ranked SNPs, genes, and genomic regions (e.g., ranked above a threshold rank or having a score above a threshold score) in the subset of intermediate candidate variants that have known or suspected correlations with the pharmacological phenotype being studied can then be identified as SNPs, genes, and genomic regions (box 414) that are causally related to the specific pharmacological phenotype. For example, SNPs, genes, and genomic regions may be associated with drug response, adverse drug reaction, adverse drug event, disease risk, dosage, complications, drug abuse, drug-gene interaction, drug-drug interaction, polypharmacy interaction, etc.
[0110] More specifically, to filter out a portion of permissive candidate variants to generate a subset of intermediate candidate variants (box 410), the regulatory functions of the genomic regions surrounding the permissive candidate variants are evaluated (box 408a) to determine whether their sequence context (e.g., allele) affects the regulatory function (variant dependency) (box 408b) and to determine their target genes (box 408c).
[0111] To assess whether a permissive candidate variant is functional, bioinformatics analysis can be used to determine whether the permissive candidate variant is localized in open chromatin, as indicated by DNase I hypersensitivity. Figure 4A An exemplary diagram 450 of the bioinformatics analysis is depicted in FIG.
[0112] Can use as various machine learning techniques such as support vector machine (SVM) to determine variant dependency.For example, SVM can be used together with permissive candidate variant to create the hyperplane that is used to classify the k-aggregate, room k-aggregate or other local sequence features in DNA sequence.The specific allele of SNP can be used to measure the tendency of the state change of genome near part.This can indicate the importance level of SNP to the specific epigenome tracking (omics modality) in the tissue or cell line for training SVM.In addition or alternatively, can by using position weight matrix (PWM) or for this purpose other algorithm identification change the SNP that transcription factor combines and determine variant dependency.
[0113] Can also use various bioinformatics techniques and machine learning techniques to determine the target gene of permissibility candidate variant.Can use quantitative trait locus (QTL) mapping to identify target gene, thereby identify the association between the group state of permissibility candidate variant and gene expression and / or genetic locus.Can utilize biological method and data set and cis-eQTL, the software analysis mapping system of trans-eQTL, dsQTL, esQTL, hQTL, haQTL, eQTL, meQTL, pQTL, rQTL etc. permissibility candidate variant can regulate the gene expression (cis-regulatory element) on the same chromosome or can regulate the gene expression (trans-regulatory element) on another chromosome.But mapping system may have unnecessary sampling and wrong relationship.Therefore, use machine learning techniques to perform addition correction to fill sparse data.
[0114] To determine whether a functional permissive candidate variant maintains regulatory control over nearby genes, bioinformatics analysis can determine whether the permissive candidate variant is hypomethylated, whether the permissive candidate variant is associated with histone marks that indicate transcription start sites, and / or whether the permissive candidate variant enhances RNA, promoter RNA, or other RNA.
[0115] Methods for determining long-range interactions between permissive candidate variants and the genes they regulate can include Hi-C chromatin conformation capture, ChIA-PET, chromatin immunoprecipitation sequencing (ChIP-seq), and QTL analysis. Such methods can also be used to determine the target genes of permissive candidate variants. This information can be merged with QTL data, and other contacts can be detected or simulated by increasing information density using matrix densification methods or various other machine learning techniques.
[0116] In any case, each permissive candidate variant can be scored and / or ranked based on regulatory function for a particular pharmacological phenotype (block 408a), variant dependency (block 408b), and target gene (block 408c). Permissive candidate variants that score above a threshold score and / or rank above a threshold rank or other score or ranking criteria can be included in a subset of intermediate candidate variants (block 410).
[0117] Then, the subsets of the intermediate candidate variants (frame 410) are scored and / or sorted relative to each other using machine learning techniques. For example, a bipartite graph analysis can be performed on the subsets of the intermediate candidate variants, wherein the intermediate candidate variants are represented by the nodes in the graph, and the relationship between the two intermediate candidate variants is represented by the edges. The intermediate candidate variants can be divided into disjoint sets, wherein any member in the disjoint sets has no relationship with each other. In some embodiments, the relative strength of the specific relationship between the two intermediate candidate variants can be assigned a specific weight. Each intermediate candidate variant can then be scored based on the number of relationships that the intermediate candidate variant has with other intermediate candidate variants from other disjoint sets. In some embodiments, each intermediate candidate variant is scored based on the total weight of each relationship assigned to the intermediate candidate variant and other intermediate candidate variants.
[0118] In any case, the highest-ranked SNPs, genes, and genomic regions (e.g., ranked above a threshold rank or having a score above a threshold score) in the subset of intermediate candidate variants can then be identified as SNPs, genes, and genomic regions associated with a specific pharmacological phenotype (block 414). The identified SNPs, genes, and genomic regions for the specific pharmacological phenotype can then be analyzed for the current patient when predicting whether the current patient has the specific pharmacological phenotype.
[0119] When methods 400 and 800 are performed using server 102, and in other embodiments, sensitive, proprietary, or valuable data can be protected by using encryption and / or secure execution techniques and / or remote computing devices that are subject to additional security protections. Such data can include patient data that complies with HIPAA or other confidentiality and regulatory requirements, data subject to patient or client privilege restrictions, proprietary data of a business entity, or other such data. Such data can be encrypted for transmission and decrypted for analysis, and such data can be analyzed in an anonymous or encrypted form using mathematical transformations such as hash tables, elliptic curves, or other metrics. In such analysis, trusted execution techniques, trusted platform modules, and other similar techniques can be used. Data representations that omit or obfuscate personal health information (PHI) can be used, particularly when preparing reports and diagnostic information and distributing them to healthcare practitioners.
[0120] Figure 4C An example gene regulatory network 470 or genomic region containing genes and SNPs indicative of a pharmacological phenotype is shown. Figure 4B Method 400 is described to identify an example gene regulatory network 470 and / or can be identified based on GWAS or candidate gene association studies.Gene regulatory network 470 can be located within the central nervous system or within any other suitable system in the human body.
[0121] In any case, the gene regulatory network 470 includes the genes BCDEF (reference number 472), DEFGH (reference number 474), ABCF (reference number 476), IJKLM (reference number 478), MNOP (reference number 480), LMNOP (reference number 482), PQRS (reference number 484), HIJKLM (reference number 486), XYZ (reference number 488), CDEFG (reference number 490) and ABCDEF (reference number 492).
[0122] The gene regulatory network 470 includes several non-coding SNPs located in introns, promoters, and intergenic regions, including non-coding SNPs associated with transcription, which are significantly associated with the response of a specific patient cohort to drug X. For example, SNP2 found in the BCDEF gene 472 on chromosome 1 indicates drug X response and disease risk. In another example, SNP3 with tight linkage disequilibrium (e.g., LD>0.8) with SNP2 is also found in the BCDEF gene 472 on chromosome 1, and the SNP3 indicates an adverse drug reaction associated with drug X. The gene regulatory network 470 also includes interchromosomal interactions, which provide a basis for a subset of trans-enhancer-promoter and promoter-promoter spatial interactions. For example, SNP15 found in the enhancer region of the gene PQRS (reference number 484) on chromosome 1 that interacts with the gene HIJKLM (reference number 486) on chromosome 6 indicates an adverse drug reaction associated with drug X. In some cases, one or more variants, genes, or enhancers in gene regulatory network 470 may be located on a sex chromosome.
[0123] Use single or double arrows to Figure 4C 478 and LMNOP (reference number 482) within a gene regulatory network 470. Each connection can include a numerical value or classification coefficient (e.g., P, C, V, T), which can be further described in legend 494. In some embodiments, the numerical value or classification coefficient indicates the relationship between the interconnected genes (e.g., activation, translocation, expression, repression, etc.).
[0124] The example gene regulatory network 470 is only one example of omics data that may be obtained from GWAS, candidate gene association studies, and / or training patients to train the pharmacological phenotype prediction system 100. Additional gene regulatory networks may be obtained with additional or alternative genomic data, epigenomic data, transcriptomic data, proteomic data, chromomic data, or metabolomic data.
[0125] In addition to reference Figure 4BIn addition to the method 400 described, Figure 4D Another exemplary method 800 for identifying omics data corresponding to a specific pharmacological phenotype phase using machine learning techniques is shown. The method 800 can be executed on the pharmacological phenotype assessment server 102. In some embodiments, the method 800 can be implemented in a set of instructions stored on a non-transitory computer readable memory and executable on one or more processors on the pharmacological phenotype assessment server 102. For example, the method 800 can be executed by Figure 1A The training module 160 within the machine learning engine 146 is executed.
[0126] In method 800, the Figure 4B In the method 400 of , similar mode is identified permissibility candidate variant (frame 810).More specifically, at frame 802 and 804 places, for several non-coding or coding SNPs in genome, each in gene and genomics region performs statistical test (for example, by GWAS or candidate gene association study), to determine the relationship between SNP and specific pharmacology phenotype, described pharmacology phenotype can be drug reaction, adverse drug reaction, adverse drug events, dosage, disease risk etc. (for example, to the reaction of warfarin).When statistical test shows that there is significant relationship between SNP and specific pharmacology phenotype (for example, p value is less than the threshold probability using null hypothesis), then determine that SNP is relevant to specific pharmacology phenotype.Then, at frame 806 places, SNP and other SNP relevant to specific pharmacology phenotype are carried out linkage disequilibrium analysis, to identify which SNP and related SNP are in linkage disequilibrium.Can be by comparing SNP (original SNP) with the database of SNP (for example, from 1000 genome projects) to find the SNP linked with original SNP and carry out linkage disequilibrium analysis. In some embodiments, the ethnic groups of the population used in the GWAS or candidate gene association study can be identified, and the data of the matched population can be used to find SNPs with significant linkage disequilibrium coefficients in the database of SNPs for the identified ethnic groups. Somatic SNPs of genes with known or suspected associations with the pharmacological phenotype being studied can also be identified (block 808).
[0127] The permissive candidate variants (block 810 ) may then be subjected to a bioinformatics analysis to filter the permissive candidate variants into a subset of intermediate candidate variants (block 814 ). Figure 4BMethod 400 filters permissive candidate variants based on the status of permissive candidate variants as putative expression regulation variants (e.g., based on regulatory function, dependency of the regulatory function on variant alleles, and the presence of identifiable target gene relationships). In method 800, permissive candidate variants are filtered based on expression regulation variants (frames 812a-812c) or coding variants (frame 812d) to generate a subset of intermediate candidate variants (814). In order to filter based on coding variants, method 800 determines whether the permissive candidate variant is a non-synonymous coding variant with a significant minor allele frequency (e.g., a minor allele frequency of at least 0.01).
[0128] More specifically, each permissive candidate variant can be scored and / or ranked based on its expression regulation variant (e.g., according to regulatory function (box 812a), variant dependency (box 812b), and target gene (box 812c) for a particular pharmacological phenotype). Permissive candidate variants that score above a threshold expression regulation variant score and / or rank above a threshold expression regulation variant rank or other score or ranking criteria can be included in a subset of intermediate candidate variants (box 814). Additionally, each permissive candidate variant can be scored and / or ranked based on its coding variant (e.g., according to whether it is a non-synonymous coding variant with a significant minor allele frequency for a particular pharmacological phenotype). Permissive candidate variants that score above a threshold coding variant score and / or rank above a threshold coding variant rank or other score or ranking criteria can also be included in a subset of intermediate candidate variants (box 814).
[0129] Then, at block 816, the intermediate candidate variants are associated with the target gene and pathway analysis is performed on the gene expressed in the relevant tissue (e.g., based on genotype-tissue expression (GTEx) data), such as using Pathway analysis performs pathway mapping and gene set enrichment. Gene sets associated with important and relevant pathways are identified, and regulatory variants and coding variants that affect the gene sets are identified as candidate variants (block 818).
[0130] Reference Figure 4E Described Figure 4D An example application of the methods described above is to a panel of warfarin phenotypes. Warfarin is an anticoagulant used to prevent and treat venous thromboembolism in heart disease and other conditions requiring controlled coagulation. Dosage requirements vary up to 10-fold between patients, and despite the recent availability of other anticoagulants, warfarin remains a commonly prescribed medication. Therefore, the methods described above can be used to predict a patient's response to warfarin and determine whether to administer warfarin or another anticoagulant, as well as the dosage.
[0131] Figure 4E Shown is the representation in Figure 4D Block diagram 850 of the single nucleotide polymorphisms (SNPs) identified in each stage of the method 800 described in for identifying omics data corresponding to a warfarin phenotype.
[0132] To identify associations and candidate genes for a set of warfarin phenotypes, 23 GWAS were performed in healthy patients on warfarin response and other pharmacological phenotypes of warfarin, venous thromboembolism risk, and baseline anticoagulant protein levels. Input data from populations around the world were used, including European, East Asian, South Asian, African, and American cohorts. In this example, the warfarin phenotype included several phenotypic categories, such as warfarin response, ADE, and disease / background. The warfarin response category included warfarin phenotype: warfarin maintenance dose. The ADE category included warfarin phenotype: hemostatic factors and hematological phenotype, hemorrhagic terminal coagulation, and thrombin generation potential phenotype. The disease / background category included warfarin phenotype: venous thromboembolism, thrombus, thrombosis, coagulation, bleeding, C4b binding protein level, activated partial thromboplastin time, anticoagulation level, factor XI, prothrombin time, platelet thrombosis. Based on these 23 GWAS and 23 additional variants, a total of 204 SNPs were identified as associations and candidate gene inputs (box 852).
[0133] Then, linkage disequilibrium analysis was performed on the 204 SNPs, and for a total of 4492 SNPs identified as permissive candidate variants, somatic SNPs were also identified for the 204 SNPs (block 854). The expression regulation variant workflow was then applied to the 4492 SNPs, resulting in a total of 186 SNPs in 57 genes (block 856). Figure 4D As shown, the gene expression test of block 814 was applied to 186 SNPs, resulting in a total of 66 SNPs in 30 genes. In addition, the coding variant workflow was also applied to 4492 SNPs, resulting in a total of 37 SNPs with a minor allele frequency of at least 0.01 (block 858). Figure 4D As shown, the gene expression test of block 814 is also applied to 37 SNPs, resulting in a total of 22 SNPs in 17 genes. Thus, the combined output of the expression regulation variant workflow and the coding variant workflow is 87 SNPs in 41 genes (block 860). Finally, pathway analysis is performed on the 87 SNPs to identify a single pathway with 74 SNPs in 31 genes (block 862).
[0134] The pathway can be referred to as the warfarin response pathway and includes genes expressed in the liver, small intestine, and vasculature. Figure 4F An example warfarin response pathway 870 is shown, comprising genes and SNPs indicative of a warfarin phenotype. Figure 4DThe method 800 is described to identify the example warfarin response pathway 870. In any case, the warfarin response pathway includes the following genes: aldo-keto reductase family 1 member C3 (AKR1C3), cytochrome P450 family 2 subfamily C member 19 (CYP2C19), cytochrome P450 family 2 subfamily C member 8 member (CYP2C8), cytochrome P450 family 2 subfamily C member 9 (CYP2C9), cytochrome P450 family 4 subfamily F member 2 (CYP4F2), coagulation factor V (F5), coagulation factor VII (F7), coagulation factor X (F10), coagulation factor XI (F11), fibrinogen gamma chain (FGG), oromucoid 1 (ORM1), serine protease 53 (PRSS53), vitamin K epoxide reductase complex subunit 1 (VKORC1), synthin 4 (STX4), coagulation factor XIII A chain (F13A1), protein C receptor (PROCR), von Willebrand factor (VWF), complement factor H-related 5 (CFHR5), fibrinogen alpha chain (FGA), flavin-containing monooxygenase 5 (FMO5), histidine-rich glycoprotein (HRG), kininogen 1 (KNG1), surplus 4 (SURF4), α1-3-N-acetylgalactosaminyltransferase and α1-3-galactosyltransferase (ABO), lysozyme (LYZ), polycomb family RING finger 3 (PCGF3), serine protease 8 (PRSS8), transient receptor potential cation channel subfamily C member 4-related pattern (TRPC4AP), solute carrier family 44 member 2 (SLC44A2), sphingosine kinase 1 (SPHK1), and ubiquitin-specific peptidase 7 (USP7).
[0135] The 74 SNPs (not shown) contained in 31 genes were: rs12775913 (regulatory SNP), rs346803 (regulatory SNP), rs346797 (regulatory SNP), rs762635 (regulatory SNP), and rs76896860 (regulatory SNP) contained in the AKR1C3 gene (expressed in the liver); rs3758581 (coding SNP) contained in the CYP2C19 gene (expressed in the liver); rs10509681 (coding SNP) and rs11572080 (coding SNP) contained in the CYP2C8 gene (expressed in the liver); and rs12775913 (regulatory SNP), rs346803 (regulatory SNP), rs346797 (regulatory SNP), rs762635 (regulatory SNP), and rs76896860 (regulatory SNP) contained in the AKR1C3 gene (expressed in the liver). 057910 (coding SNP), rs1799853 (coding SNP), and rs7900194 (coding SNP) contained in the CYP4F2 gene (expressed in the liver); rs2108622 (coding SNP) contained in the CYP4F2 gene (expressed in the liver); rs6009 (regulatory SNP), rs11441998 (regulatory SNP), rs2026045 (regulatory SNP), rs34580812 (regulatory SNP), rs749767 (regulatory SNP), rs9378928 (regulatory SNP), and rs7937890 (regulatory SNP) contained in the F5 gene (expressed in the liver); rs7552487 (regulatory SNP) contained in the F7 gene (expressed in the liver); The rs107750773 (SNP) is a 100-fold increase in the number of SNPs in the 15- to 20-year-old female fetal stem cell lineage (FSL) and is expressed in the liver. The rs107750773 (SNP) is a 100-fold increase in the number of SNPs in the 15- to 20-year-old female fetal stem cell lineage (FSL) and is expressed in the liver. The rs107750773 (SNP) is a 100-fold increase in the number of SNPs in the 15- to 20-year-old female fetal stem cell lineage (FSL) and is expressed in the liver. rs7199949 (coding SNP) in the 53 gene (expressed in the liver); rs2884737 (regulatory SNP), rs9934438 (regulatory SNP), rs897984 (regulatory SNP), and rs17708472 (regulatory SNP) in the VKORC1 gene (expressed in the liver); rs35675346 (regulatory SNP) and rs33988698 (regulatory SNP) in the STX4 gene (expressed in the small intestine); rs5985 (coding SNP) in the F13A1 gene (expressed in the vasculature); and rs867186 (coding SNP) in the PROCR gene (expressed in the vasculature).The rs75648520 (regulatory SNP), rs55734215 (regulatory SNP), rs12244584 (regulatory SNP), and rs1063856 (coding SNP) are contained in the VWF gene (expressed in the vasculature); rs674302 (regulatory SNP) is contained in the CFHR5 gene (expressed in the liver); rs12928852 (regulatory SNP) and rs6050 (coding SNP) are contained in the FGA gene (expressed in the liver); rs8060857 (regulatory SNP) and rs7475662 (regulatory SNP) are contained in the FMO5 gene (expressed in the liver); and rs12928852 (regulatory SNP) and rs6050 (coding SNP) are contained in the HRG gene. rs9898 (coding SNP) in the KNG1 gene (expressed in the liver); rs710446 (coding SNP) in the KNG1 gene (expressed in the liver); rs11577661 (regulatory SNP) in the SURF4 gene (expressed in the liver); rs11427024 (regulatory SNP), rs6684766 (regulatory SNP), rs2303222 (regulatory SNP), rs1088838 (regulatory SNP), rs13130318 (regulatory SNP), and rs12951513 (regulatory SNP) in the ABO gene (expressed in the small intestine); and rs11427024 (regulatory SNP), rs6684766 (regulatory SNP), rs2303222 (regulatory SNP), rs1088838 (regulatory SNP), rs13130318 (regulatory SNP), and rs12951513 (regulatory SNP) in the LYZ gene (expressed in the small intestine). rs8118005 (regulatory SNP); rs76649221 (regulatory SNP), rs9332511 (regulatory SNP), and rs6588133 (regulatory SNP) contained in the PCGF3 gene (expressed in the small intestine); rs11281612 (regulatory SNP) contained in the PRSS8 gene (expressed in the small intestine); rs11589005 (regulatory SNP), rs8062719 (regulatory SNP), rs889555 (regulatory SNP), rs36101491 (regulatory SNP), rs7426380 (regulatory SNP), and rs6 rs579208 (regulatory SNP), rs77420750 (regulatory SNP), and rs73905041 (coding SNP) in the SLC44A2 gene (expressed in the vasculature); rs3211770 (regulatory SNP), rs3211770 (regulatory SNP), rs3087969 (coding SNP), and rs2288904 (coding SNP) in the SLC44A2 gene (expressed in the vasculature); rs683790 (regulatory SNP) and rs346803 (coding SNP) in the SPHK1 gene (expressed in the vasculature); and rs201033241 (coding SNP) in the USP7 gene (expressed in the vasculature).
[0136] In addition to warfarin, Figure 4B and4D The methods described in can also be applied to a panel of lithium phenotypes as well as any other pharmacological phenotypes. Figure 4B and 4DThe method described in
[15] was applied to the lithium phenotype and identified 78 SNPs in 12 genes involved in the lithium response pathway. The lithium response pathway includes the following genes: ankyrin 3 (ANK3), aryl hydrocarbon nuclear transporter receptor-like (ARNTL), voltage-gated calcium channel auxiliary subunit gamma 2 (CACNG2), voltage-gated calcium channel auxiliary subunit alpha 1C (CACNA1C), cyclin-dependent kinase inhibitor 1A (CDKN1A), cAMP response element binding protein 1 (CREB1), AMPA-type glutamate positive receptor subunit 1 (GRIA2), glycogen synthase kinase 3 beta (GSK3B), nuclear receptor subfamily 1, group D, member 1 (NR1D1), solute carrier family 1, member 2 (SLC1A2), serotonin receptor 1A (HTR1A), and TRAF2 and NCK-interacting kinase (TNIK).The 78 SNPs included in 12 genes are: rs2185502, rs10821792, rs1938540, rs3808943, rs61847646, rs75314561, rs61846516, rs10994397, rs10994318, rs61847579, rs12412727, rs10994308, rs4948418, rs4948412, rs4948413, rs4948416, rs10821745, rs10994336, rs10994360, rs9633532, and rs 1938526, rs10994322, and rs10994321; rs10766075, rs7938308, rs10832017, rs4603287, rs7934154, rs12361893, rs4414197, rs4757140, rs4757141, rs61882122, rs11022755, rs11022754, rs1481892, rs1481891, rs4353253, rs4756764, rs2403662, rs4237700, rs10832018, and rs1236122 rs290622, rs7928655, rs34148132, rs4146388, rs4146387, rs7949336, rs4757139, rs7107287, and rs1351525 in the CACNG2 gene; rs2284017 and rs2284016 in the CACNA1C gene; rs3176336, rs3176333, rs3176334, rs3176320, rs4135240, rs2395655, and rs733590 in the CDKN1A gene; and rs3176336, rs3176333, rs3176334, rs3176320, rs4135240, rs2395655, and rs733590 in the CRE gene. rs10932201 in the B1 gene; rs78957301 in the GRIA2 gene; rs334558 in the GSK3B gene; rs2314339 in the NR1D1 gene; rs3794088, rs3794087, rs4354668, rs12418812, rs1923294, rs5791047, rs111885243, rs752949, and rs16927292 in the SLC1A2 gene; rs6449693 and rs878567 in the HTR1A gene; and rs7372276 in the TNIK gene.
[0137] exist Figure 4G The lithium reaction pathway 890 is depicted in FIG. Lithium is a psychotropic drug used to treat mental illness / disorders. The above method can be used to predict a patient's response to lithium and determine whether to administer lithium or other psychotropic drugs to the patient and the dosage to be administered. In any case, the method can be used to predict a patient's response to lithium or other psychotropic drugs. Figure 4B and 4D Methods 400, 800 described in
[00155] were used to identify an example lithium response pathway 890. Each gene in the lithium response pathway 890 is expressed in a portion of the brain including the frontal lobe, insula, temporal cortex, cingulate cortex, amygdala, hippocampus, caudate nucleus, thalamus, motor cortex, fusiform cortex, substantia nigra, cerebellum, and hypothalamus.
[0138] In some embodiments, the pharmacological phenotype prediction system 100 can test whether the identified SNPs, genes, and genomic regions are present in the current patient to determine whether the current patient has a specific pharmacological phenotype. For example, the identified SNPs, genes, and genomic regions may indicate a negative response to valproic acid used to treat TBI. When the current patient suffers from TBI, a biological sample of the current patient can be provided and, for example, the pharmacological phenotype prediction system 100 can be used to predict the presence of the identified SNPs, genes, and genomic regions in the current patient to determine whether the current patient has a specific pharmacological phenotype. Figure 5 The process 500 for generating omics data is used to analyze the biological sample to detect whether the identified SNPs, genes, and genomic regions are present. When the current patient has at least some of the identified SNPs, genes, and genomic regions that indicate a negative reaction to valproic acid, valproic acid is not administered to the current patient. In other embodiments, the identified SNPs, genes, and genomic regions are scored, combined, and / or weighted in any suitable manner to determine which combinations indicate a negative reaction to valproic acid. The scoring or weighting system is then applied to the SNPs, genes, and genomic regions in the current patient's biological sample to determine whether the current patient has a combination that indicates a negative reaction to valproic acid.
[0139] In any case, the pharmacological phenotype prediction system 100 can provide identified SNPs, genes, and genomic regions that indicate a specific pharmacological phenotype as the omics data of the pharmacological phenotype. The machine learning engine 146 can obtain omics data with sociomics, physiological omics, and environmental data for training patients with at least some of the identified SNPs, genes, and genomic regions, as well as the phenotypic omics data of the training patients. In this way, the machine learning engine 146 can classify training patients with identified SNPs, genes, and genomic regions that indicate a specific pharmacological phenotype (e.g., a negative response to ketamine for treating depression) as having a specific pharmacological phenotype or not having a specific pharmacological phenotype. The sociomics, physiological omics, and environmental data can then be used to distinguish training patients with identified SNPs, genes, and genomic regions that do have a specific pharmacological phenotype from training patients with identified SNPs, genes, and genomic regions that do not have a specific pharmacological phenotype.
[0140] For example, when the machine learning technique is a decision tree, a decision tree comprising multiple nodes can be generated, each node representing a test of the current patient's data. The nodes can be connected by branches, each branch representing the result of a test or other measurement or an observable / recordable state (e.g., a "yes" branch and a "no" branch), wherein the branches can be weighted and the leaf nodes can indicate whether the current patient has a pharmacological phenotype. In other embodiments, the leaf nodes indicate the likelihood of the pharmacological phenotype, for example, determined by aggregating or combining weighted branches, or the leaf nodes can indicate a score that can be compared to a threshold to determine whether the current patient has a pharmacological phenotype. In any case, a decision tree can be generated, with the nodes near the top of the decision tree representing tests of the current patient's omics data, for example, as shown by identified SNPs, genes, and genomic regions. When the current patient has an appropriate combination of identified SNPs, genes, and genomic regions that indicate a specific pharmacological phenotype (e.g., a negative response to ketamine for treating depression), the decision tree branches to several nodes representing tests of the current patient's sociogenomic, physiological, and environmental data.
[0141] In another example, when the machine learning technique is an SVM, for each training patient, the pharmacological phenotype assessment server 102 obtains sociomic, physiomic, and environmental data, identified SNPs, genes, and genomic regions of the training patient that indicate a specific pharmacological phenotype, and an indication of whether the training patient has the specific pharmacological phenotype (e.g., a negative response to ketamine for treating depression) as training vectors. The SVM obtains each training vector and creates a statistical model for determining whether the current patient has the specific pharmacological phenotype by generating a hyperplane that separates a first subset of training vectors corresponding to training patients with the pharmacological phenotype from a second subset of training vectors corresponding to training patients without the pharmacological phenotype.
[0142] Environmental, physiological and sociomic data can be obtained for a training patient or a cohort of training patients having identified SNPs, genes and genomic regions that indicate a specific pharmacological phenotype. More specifically, the training module 160 can obtain a set of training data, for example, from the client devices 106-116 and / or one or more servers (e.g., an EMR server, a multi-pharmacy server, etc.), wherein the training data can include omics data for several training patients and sociomics, physiological and environmental data, wherein the pharmacological phenotype of the training patient is known (e.g., previously determined or currently determined) and is also provided in the training data. Environmental, physiological and sociomic data can include clinical data, demographic data, multi-pharmacy data, socioeconomic data, educational data, drug abuse data, diet and exercise data, law enforcement data, circadian rhythm data, family data, or any other suitable data indicating the patient's social status or environmental conditions.
[0143] In an exemplary case, the omics, phenomics, sociomics, physiology and environmental data of a training patient are collected in a first time period (e.g., one year). Although the patient can be a training patient, the result of the pharmacological phenotype prediction system 100 can also be determined based on the patient's omics, phenomics, sociomics, physiology and environmental data. As described above, the training patient with the training data for training the pharmacological phenotype prediction system 100 can also become the current patient for predicting the unknown pharmacological phenotype of the training patient. In this example, omics, phenomics, sociomics, physiology and environmental data can be collected during January to December of the first year. Although omics, phenomics, sociomics, physiology and environmental data can be for a single training patient, the omics, phenomics, sociomics, physiology and environmental data of a team of training patients can also be collected. For example, omics, phenomics, sociomics, physiomic, and environmental data can be collected for a cohort of training patients, each with identified SNPs, genes, and genomic regions that indicate a negative response to drug X for the treatment of disease Y.
[0144] Omics, phenomics, sociomics, physiomic and environmental data can include Figure 2 302, exposomes 304, sociomics demographics and stress / trauma 306, and medical physiomic, EHR, laboratory values, stress and abuse factors and trauma and medical outcome data 308. However, these are just a few examples of the omic, phenomic, sociomic, physiomic and environmental data that can be acquired from training patients to train the pharmacological phenotyping server 102. Additional sociomic, physiomic and environmental data may also be included, such as circadian rhythm data indicating sleep and other recurring lifestyle temporal patterns of the training patients.
[0145] The omics and pharmaco-metabolomics data may be the results of pharmacogenomic analysis of biological samples from the training patients.The exposome data for year 1 may include the employment status and residence of the training patients in August of year 1.
[0146] Medical physiology, EHR, laboratory values, stress and abuse factors and trauma, and medical outcome data can indicate the training patients' law enforcement experiences during January to December of Year 1.
[0147] The medical physiology, EHR, laboratory values, stress and abuse factors and trauma, and medical outcome data may also include multi-pharmacy data, which includes medications prescribed to the training patient and the time when the medications were prescribed. For example, in March of Year 1, the training patient was prescribed medication A to treat disease 1, medication B to treat disease 2, and medication C to treat disease 3, and in August of Year 1, medication D to treat disease 4. In August of Year 1, the training patient also gradually stopped taking medication C. In addition, the medical physiology, EHR, laboratory values, stress and abuse factors and trauma, and medical outcome data may also include phenotypic data indicating the pharmacological phenotype of the training patient (such as the disease with which the training patient was diagnosed). For example, in January of Year 1, the training patient was diagnosed with diseases 1-3.
[0148] In addition, the data may include other phenotypic data, such as information describing the efficacy and / or side effects of a drug. For example, a training patient may experience side effects of drug C, and thus, after experiencing these side effects, the training patient may gradually stop taking drug C. A training patient may also have a positive response to taking drug C, which may indicate the efficacy of the drug in the training patient.
[0149] Sociomic and demographic data can indicate the family status of the training patient. For example, sociomic and demographic data can indicate the marital status and number of children of the training patient in January of Year 1. Sociomic and demographic data can also indicate the amount of income of the training patient from January to December of Year 1. Sociomic and demographic data can also indicate available information about family members who are or are not affected by the relevant disease and complications, as well as their treatment response and other pharmacological phenotypes.
[0150] In this exemplary scenario, the omics, phenomics, sociomics, physiomic, and environmental data may also include data collected from training patients between January and December of Year 2. In Year 2, the omics and pharmaco-metabolomics data include proteomic and transcriptomic analysis results for training patients obtained in March and pharmacogenomic analysis results obtained in August.
[0151] The omics and pharmaco-metabolomics data for year 2 may include the results of an inpatient metabolic exam for the training patient in March of year 2, which indicates the presence of toxic metabolites of drug B. The omics and pharmaco-metabolomics data may also include the results of another inpatient metabolic exam for the training patient in August of year 2, which indicates normal blood levels of drug E.
[0152] Year 2 medical physiomic, EHR, laboratory values, stress and abuse factors and trauma, and medical outcome data indicate the results of various mental health, substance abuse, and stress and trauma questionnaires administered to training patients.
[0153] The medical physiology, EHR, laboratory values, stress and abuse factors and trauma, and medical outcome data for Year 2 also indicated that the training patient had a negative reaction to Drug C in February and subsequently stopped taking Drug C. In contrast, the patient was prescribed Drug F in March and Drug E in August of Year 2. The training patient then had a positive reaction to taking Drugs E and F, which may indicate the efficacy of the drugs for the training patient.
[0154] In addition, in response to detecting the presence of toxic metabolites of Drug B (as indicated by the training patient's omics and pharmaco-metabolomics data), the training patient discontinued Drug B in March of Year 2. Medical physio-omics, EHR, laboratory values, stress and abuse factors and trauma, and medical outcome data further indicated that the training patient's Disease 1 diagnosis was controlled in April of Year 2, and the dose of Drug A was reduced. The Year 2 socio-omics and demographic data also indicated the amount of income the training patient had from January to December of Year 2.
[0155] The pharmacological phenotype assessment server 102 can be trained using information contained in the omics, phenotypic, sociomic, physiological, and environmental data from Years 1 and 2 to generate a statistical model. Similar information can be collected for a number of training patients (e.g., tens, hundreds, or thousands) comprising a training patient cohort or population, and the omics data can be combined to generate a statistical model. In generating such a statistical model, important features can be identified, and a model using this limited information set can be generated to allow the phenotype of the current patient to be predicted using the limited information set.
[0156] For example, as described above, omics, phenomics, sociomics, physiomics, and environmental data can be obtained for a cohort of training patients, each of which has identified SNPs, genes, and genomic regions that indicate a negative response to drug X for treating disease Y. The training module 160 can classify a first subset of the omics, phenomics, sociomics, physiomics, and environmental data (e.g., the identified SNPs, genes, and genomic regions for the cohort of training patients that indicate a negative response to drug X for treating disease Y) as corresponding to training patients who had a negative response to drug X, and can classify a second subset of the omics, phenomics, sociomics, physiomics, and environmental data as corresponding to training patients who did not have a negative response to drug X. In some embodiments, the training module 160 can perform statistical measures on each classified subset of the omics, phenomics, sociomics, physiomics, and environmental data (training patients who had a negative response to drug X and training patients who did not have a negative response to drug X). For example, the training module 160 may determine the average income, average ACE score, etc. of the training patients corresponding to each classification.
[0157] The training module 160 can then generate a statistical model (e.g., a decision tree, a neural network, a hyperplane, linear or nonlinear regression coefficients, etc.) to predict whether the current patient has a negative response to drug X for treating disease Y based on the statistical measures for each classification. For example, the statistical model can be a decision tree having several nodes connected by branches, where each node represents a test of omics data related to identified SNPs, genes, and genomic regions that indicate a negative response to drug X for treating disease Y. The branches can contain weights or scores for different SNPs, genes, and genomic regions, and when the combined weight or score for the current patient exceeds a threshold, this can indicate that the current patient has a combination of SNPs, genes, and genomic regions that indicate a negative response to drug X.
[0158] The decision tree can further include several nodes connected by branches, each representing a test of sociomic, physiomic, and environmental data. A first node can test whether the current patient's annual income is greater than $20,000. This node is connected to a second node via a "yes" branch. This second node tests whether the current patient's ACE score is greater than 5. This node is connected to a third node via a "yes" branch. This third node tests whether the current patient has experienced domestic violence. This node is connected to a leaf node via a "no" branch. This leaf node predicts whether the current patient will have a negative reaction to drug X. A leaf branch can indicate the likelihood that the current patient will have a negative reaction to drug X, which can be determined based on the results of each test at each node and / or the weights assigned to the corresponding branches. In some embodiments, a leaf branch can indicate a response score for drug X, which can also be determined based on the results of each test at each node and / or the weights assigned to the corresponding branches. The response score can indicate the drug's efficacy in treating the corresponding disease, discounted by the drug's adverse effects on the patient.
[0159] As described in this exemplary scenario, the omic, sociomic, physiomic, and environmental data for each of several training patients can be combined with an indication of the training patient's pharmacological phenotype (e.g., the disease the training patient was diagnosed with, the response to a drug, a substance abuse problem) to generate a statistical model for predicting the current patient's pharmacological phenotype. In some embodiments, the omic, sociomic, physiomic, and environmental data for the training patient can be combined with omic and / or genomic data from prior studies (e.g., GWAS) to identify relationships between genes, transcription factors, proteins, metabolites, chromatin states, the environment, and the patient's pharmacological phenotype.
[0160] In any case, the omics, sociomics, physiomic and environmental data of each training patient are combined to generate a statistical model. In some embodiments, the omics, sociomics, physiomic and environmental data of the training patient can be classified as corresponding to a training patient diagnosed with a specific disease (or not diagnosed as suffering from a specific disease), classified as corresponding to a training patient suffering from a drug abuse problem (or not suffering from a drug abuse problem), classified as corresponding to a training patient with a special reaction to a drug, or classified in any other suitable manner. In some cases, the omics, sociomics, physiomic and environmental data of the same training patient can be subdivided into multiple pharmacological phenotypes at different time periods. For example, as the sociomics, physiomic and environmental data of the training patient change, the omics data of the training patient may change over time, and in a first time period, the training patient may suffer from one mental illness, while in a second time period, the training patient may suffer from another mental illness.
[0161] Similarly, in some embodiments, the omics, sociomics, physiomic, and environmental data of training patients can be categorized based on demographics. For example, the omics, sociomics, physiomic, and environmental data of training patients of European descent can be assigned to one cohort, while the omics, sociomics, physiomic, and environmental data of training patients of Chinese descent can be assigned to another cohort. In another example, the omics, sociomics, physiomic, and environmental data of training patients aged between 25 and 35 years old can be assigned to one cohort, while the omics, sociomics, physiomic, and environmental data of training patients aged between 35 and 45 years old can be assigned to another cohort.
[0162] In such embodiments, the training module 160 can generate a different statistical model for each pharmacological phenotype and / or for each cohort (e.g., a cohort separated based on demographics). For example, a first statistical model can be generated to determine the likelihood that the current patient will experience a substance abuse problem, a second statistical model can be generated to determine the risk of one disease, a third statistical model can be generated to determine the risk of another disease, a fourth statistical model can be generated to determine the likelihood of having a negative reaction to a particular drug, and so on. In other embodiments, the training module 160 can generate a single statistical model to determine the likelihood that the current patient has any pharmacological phenotype, or can generate any number of statistical models to determine the likelihood that the current patient has any number of pharmacological phenotypes.
[0163] Once the omics, sociomics, physiological omics, and environmental data are classified into subsets corresponding to individual cohorts and / or pharmacological phenotypes, the omics, sociomics, physiological omics, and environmental data of a specific pharmacological phenotype can be analyzed to generate a statistical model. A statistical model can be generated using a neural network, deep learning, a decision tree, a support vector machine, or any of the above machine learning methods. For example, the omics, sociomics, physiological omics, and environmental data of training patients of European descent can be analyzed to determine that there is a high correlation between training patients with SNP 2 found in the XYZ gene, being unemployed, and having a record of violent crime and developing schizophrenia. Therefore, patients of European descent with SNP 2 found in the XYZ gene, being unemployed, and having a record of violent crime may be at high risk for developing schizophrenia.
[0164] When the machine learning technique is a neural network or deep learning, the training module 160 can generate a graph having input nodes, intermediate or "hidden" nodes, edges, and output nodes. The nodes can represent tests or functions performed on omics, sociomics, physiological omics, and environmental data, and the edges can represent connections between nodes. In some embodiments, the output nodes can include an indication of a pharmacological phenotype or the likelihood of a pharmacological phenotype. In some embodiments, the edges can be weighted according to the strength of the test or function of the previous node when determining the pharmacological phenotype.
[0165] Therefore, for the determination of pharmacological phenotype, the type of omics, sociomics, physiology and environmental data of the previous node with higher weight may be more important than the type of omics, sociomics, physiology and environmental data of the previous node with lower weight. By identifying the most important omics, sociomics, physiology and environmental data, the training module 160 can eliminate the least important and possibly misleading and / or random noise omics, sociomics, physiology and environmental data from the statistical model. In addition, the most important omics data of the pharmacological phenotype (determined by ranking, weighting or scoring the omics data above the threshold) can be used to select the type of omics data to analyze the patient's biological sample.
[0166] In some embodiments, the present invention provides a method for the detection of bipolar disorder in patients with bipolar disorder. For example, a neural network can include four input nodes representing omics, sociomics, physiological omics and environmental data, and the input nodes are connected to several hidden nodes. The hidden nodes are then connected to output nodes, and the output nodes indicate the possibility that the current patient suffers from bipolar disorder. The connection can have the weight assigned, and the hidden nodes can include tests or functions performed on omics, sociomics, physiological omics and environmental data. In certain embodiments, the test or function can be the distribution determined by training data or previous studies (such as GWAS). For example, the possibility that the patient with specific SNP and in the 98th percentile of income distribution is lower than that with the same SNP and in the 10th percentile of income distribution.
[0167] In some embodiments, a hidden node can be connected to several output nodes, each of which indicates the likelihood that the current patient will have different diseases, the likelihood that the current patient will have a substance abuse problem, and / or the likelihood or response score of the current patient to a specific drug. In this example, the four input nodes can include the patient's ancestry, the patient's current income and / or the patient's change in income over the past year, the presence of SNP 13 in the patient's LMNOP gene, and the patient's poor sleep patterns.
[0168] In some embodiments, each of the four input nodes can assign a numerical value to the patient's omics, sociomics, physiomic, and environmental data, and tests or functions can be applied to the numerical values at the hidden nodes. The results of these tests or functions can then be weighted and / or aggregated, for example, to determine a response score for lithium. The response score can indicate the efficacy of a drug in treating the corresponding condition, discounted by the drug's adverse effects on the patient. In this example, lithium's response score may be high (e.g., 80 out of 100), indicating that the patient will respond positively to lithium when taking it to treat bipolar disorder. Therefore, a healthcare provider can prescribe lithium to treat the patient's bipolar disorder. In some embodiments, the response score for lithium used to treat bipolar disorder can be compared with the response scores of other prescription medications used to treat bipolar disorder. The medications can then be ranked based on their respective response scores, and the highest-ranked medication can be recommended to the healthcare provider for prescription to the patient. However, this is merely one example of the inputs and outputs of a statistical model used to determine a phenotype. In other examples, any number of input nodes can contain several types of omics, sociomics, and environmental data for the patient. Additionally, any number of output nodes can determine the likelihood or risk of developing different diseases, the likelihood of substance abuse problems, the likelihood of complications, and so on.
[0169] As additional training data is collected, weights, nodes, and / or connections can be adjusted. In this way, the statistical model is continuously or periodically updated to reflect at least a near real-time representation of the sociomics, physiomics, environmental, and omics data.
[0170] In some embodiments, machine learning techniques can be used to identify cohorts or demographic markers to classify patients as having or not having a specific pharmacological phenotype. As shown in the above examples, statistical measurements can be taken of the sociomics, physiologic and environmental data of a first and second subset of training patients with and without a specific pharmacological phenotype, to develop a test or function contained in a hidden node. Statistical measurements can be used to determine the most important variables contained in the sociomics, physiologic and environmental data of training patients to distinguish between training patients with and without a pharmacological phenotype. In this way, the test or function contained in the hidden node is not necessarily a priori.
[0171] After generating a statistical model using machine learning techniques as described above (e.g., neural networks, deep learning, decision trees, support vector machines, etc.), the training module 160 can test the statistical model using test omics, sociomics, physiological omics, and environmental data from the test patient and the test patient's pharmacological phenotype. The test patient can be a patient with a known pharmacological phenotype. However, for testing purposes, the training module 160 can determine the likelihood that the test patient has various pharmacological phenotypes by comparing the test omics, sociomics, physiological omics, and environmental data of the test patient with the statistical model generated using machine learning techniques.
[0172] For example, the training module 160 can traverse nodes from a neural network using the test-omics, sociomics, physiomic, and environmental data of a test patient. When the training module 160 reaches a result node indicating the likelihood or response score of a particular pharmacological phenotype, the likelihood or response score can be compared with known pharmacological phenotypes of the test patient.
[0173] In some embodiments, the determination may be considered correct if the probability that the test patient has a particular pharmacological phenotype (e.g., disease Y) is greater than 0.5 and the known pharmacological phenotype of the test patient is that they do have disease Y. In another example, the determination may be considered correct if the drug that elicits a strong response without deleterious side effects in the test patient has the highest response score. In other embodiments, when the known pharmacological phenotype of the test patient is that they have a particular pharmacological phenotype, the probability may have to be higher than 0.7, or some other predetermined threshold value of the determined probability may have to be considered correct.
[0174] Furthermore, in some embodiments, when the accuracy of the training module 160 exceeds a predetermined threshold amount of time, the statistical model may be presented to the phenotype assessment module 162. On the other hand, if the accuracy of the training module 160 does not exceed the threshold amount, the training module 160 may continue to acquire training data sets for further training.
[0175] Once the statistical model has been fully tested to verify its accuracy, the phenotype assessment module 162 can obtain the statistical model. Based on the statistical model, the phenotype assessment module 162 can determine the likelihood that the current patient has various pharmacological phenotypes. The phenotype assessment module 162 can obtain the omics, sociomics, physiomic, and environmental data for the current patient without knowing whether the current patient has certain pharmacological phenotypes. The sociomics, physiomic, and environmental data can be collected at several time points and can be similar to the sociomics and environmental data collected for the training patient described in the exemplary scenario above.
[0176] More specifically, the socio-omics, physio-omics, and environmental data may include clinical data, such as the current patient's medical history, including diseases the current patient has been diagnosed with, results of laboratory tests and procedures performed on the patient, the current patient's family medical history, etc. The socio-omics, physio-omics, and environmental data may also include multi-pharmacy data, such as each medication prescribed to the current patient within a specific time period, the duration of each prescription, the number of refills, whether the current patient has refilled each medication on time, etc. In addition, the socio-omics, physio-omics, and environmental data may include demographic data, such as the current patient's ancestry or race, the current patient's age, the current patient's weight, the current patient's gender, the current patient's place of residence, etc. Furthermore, the socio-omics, physio-omics, and environmental data may include socioeconomic data of the current patient, such as the current patient's income amount and / or source of income, education data (e.g., high school diploma, GED, college graduate, master's degree, etc.), and diet and exercise data indicating how often the current patient exercises, the current patient's eating habits, the amount of weight gained or lost within a specific time period, etc. Additionally, the sociomic, physiomic, and environmental data may include family data such as the current patient's marital status, the number of children and family members currently living with the current patient, law enforcement data indicating the current patient's criminal record and whether the current patient has been a victim of abuse or other crime, substance abuse data indicating whether the current patient has or has had a substance abuse problem, and circadian rhythm data indicating the current patient's sleep patterns.
[0177] In addition to sociological, physiological, and environmental data, the phenotype assessment module 162 can also obtain omics data of the current patient. The omics data can be similar to Figure 4CMore specifically, the omics data may include genomic data indicating gene traits, epigenomic data indicating gene expression, transcriptomic data indicating DNA transcription, proteomic data indicating proteins expressed by a genome, chromomic data indicating chromatin states in a genome, and / or metabolomic data indicating metabolites in a genome.
[0178] The phenotype assessment module 162 can then apply the current patient's omics, sociomics, physiomic, and environmental data to the statistical model to determine the likelihood that the current patient has various pharmacological phenotypes. When several statistical models are generated, the phenotype assessment module 162 can apply the current patient's omics, sociomics, physiomic, and environmental data to each of the statistical models to determine, for example, the likelihood that the current patient has disease Y, has a substance abuse problem, and has complications.
[0179] In some embodiments, a healthcare provider may provide a request for a specific type of pharmacological phenotype (e.g., a patient's current predicted response to a specific drug), or a request for the best drug to treat a specific disease. Accordingly, the phenotype evaluation module 162 may apply a statistical model or portion of a statistical model generated in response to the healthcare provider's request. The healthcare provider may then receive an indication of the best drug and dosage for treating the specific disease from the pharmacological phenotype evaluation server 102.
[0180] In other embodiments, the pharmacological phenotype assessment server 102 can assess the pharmacological, sociological, physiological, and environmental data of the current patient by applying the data to a statistical model to determine a likelihood or response score for each of several pharmacological phenotypes. The pharmacological phenotype assessment server 102 can then generate a risk analysis display for review by the current patient's healthcare provider.
[0181] The risk analysis display may include an indication of patient biographical information, such as the patient's name, date of birth, address, etc. The risk analysis display may also include an indication of each of the likelihoods or other semi-quantitative and quantitative measures of various pharmacological phenotypes, which may be expressed as a probability (e.g., 0.6), a percentage (e.g., 80%), one of a set of categories (e.g., "high risk," "moderate risk," or "low risk"), and / or in any other suitable manner. Additionally, the risk analysis display may include an indication of the response score to the drug, which may be numerical (e.g., 75 out of 100), categorical (e.g., "strong response," "adverse response," etc.), or in any other suitable manner. Additionally, each likelihood or response score may be displayed along with a description of the corresponding pharmacological phenotype (e.g., "high risk for disease Y"). These risk factors and levels may be stored and processed in quantitative or semi-quantitative form within the internal working environment of the statistical model, but may be converted into qualitative terms for output to care providers and patients.
[0182] In some embodiments, the pharmacological phenotype assessment server 102 can compare the likelihood of the pharmacological phenotype with a likelihood threshold (e.g., 0.5) and can include the pharmacological phenotype in the risk analysis display when the likelihood of the pharmacological phenotype exceeds the likelihood threshold. For example, only diseases that are high risk for the current patient can be included in the risk analysis display. In another example, the risk analysis display can include an indication that the current patient may have a drug abuse problem when the likelihood of the current patient having a drug abuse problem exceeds the likelihood threshold. In this way, a healthcare provider can recommend or provide early intervention to the current patient. Similarly, in some embodiments, the response score of each drug corresponding to a specific disease can be ranked (e.g., from highest to lowest). The drugs and corresponding response scores can be provided in the order of ranking on the risk analysis display. In other embodiments, only the highest-ranked drugs for a specific disease can be included in the risk analysis display.
[0183] In addition to displaying the highest-ranked drug or drugs (and / or other therapies) for a particular disease as recommendations to the healthcare provider for prescribing to the patient, the risk analysis display may also include recommended dosages of the drugs for the current patient. The risk analysis display may also include any socio-omics, physio-omics, environmental, or demographic information that may cause the current patient's response to the drug to change (e.g., changes in diet, exercise, exposure, etc.). In addition, the risk analysis display may include recommendations to change the current patient's existing therapy by changing the dosage, changing the drug, or eliminating the current patient's treatment plan based on their polypharmacy data, or other methods. For example, when the recommended drug may make one or more drugs in the current patient's polypharmacy data redundant, the risk analysis display may include recommendations to stop taking these drugs.
[0184] In some embodiments, if a patient is currently taking a medication that is incompatible with the highest-ranked medication, the pharmacological phenotyping server 102 may recommend a medication that has a lower response score (or other drug-specific attribute) but is more compatible with existing therapies (or other polypharmacy attributes). For example, the pharmacological phenotyping server 102 may obtain polypharmacy data for the current patient and compare the medications in the polypharmacy data with the recommended medications to check for contraindications. The risk analysis display may include the highest-ranked medication that has no contraindications with any medication prescribed to the current patient.
[0185] Beyond clinical settings, pharmacologic phenotypes can be predicted in research settings for drug development and insurance applications. In research settings, the pharmacologic phenotypes of potential patient cohorts associated with novel, experimental, or repurposed drugs might be predicted in a study plan. Patients can be selected for experimental treatment based on their predicted pharmacologic phenotypes in relation to the experimental drug.
[0186] Furthermore, as the pharmacological phenotype of the current patient becomes known (e.g., the current patient has a certain pharmacological phenotype after a threshold amount of time, such as one year), the current patient's omics, sociomics, physiomics, and environmental, as well as phenotypic data can be added to the training data, and the statistical model can be updated accordingly.
[0187] Figure 6An example timeline 600 of a current patient is shown, wherein the current patient's omics, sociomics, physiomic and environmental, and phenotypic data are collected over time. The pharmacological phenotype prediction system 100 can then analyze the collected data of the current patient based on a statistical model to predict the pharmacological phenotype of the current patient. More specifically, in the example timeline 600, the current patient's diagnosis, treatment methods, and outcomes 602 are collected. The current patient's medical physiology, EHR, laboratory values, stress and abuse factors, and trauma 604 are also collected, as well as the current patient's sociomics and demographics 608, omics and drug metabolomics 610, and exposure group data 612. In addition, the current patient's phenotypic data 606 (which can indicate the current patient's medical outcomes 602) is also collected.
[0188] As described above, in a clinical treatment and / or pharmacology or other biomedical research environment, the pharmacology phenotype prediction system 100 may include a clinical decision support module 614 for clinicians and a clinical decision support module 616 for researchers.
[0189] In the clinician's clinical decision support module 614, based on the training data from individual training patients or training patient cohorts / colonies, a statistical model is generated in a manner similar to the above. The current patient's omics, sociomics, physiologic, environmental, and phenotypic data (e.g., diagnosis, treatment methods, and results 602; medical physiology, EHR, laboratory values, stress and abuse factors, and trauma 604; phenotypic data 606; sociomics and demographics 608; omics and drug metabolomics 610, and exposure group data 612) are applied to a statistical model to predict the current patient's pharmacological phenotype. Pharmacological phenotypes can include disease risk or conditions, drug recommendations, adverse drug reaction scores, overall drug response scores, etc. However, these are just a few examples of pharmacological phenotypes. Additional or alternative pharmacological phenotypes are described in the full text.
[0190] In the researcher's clinical decision support module 616, statistical models are generated in a manner similar to that described above based on training data from individual patients or patient cohorts / populations. The patient's omics, sociomics, physiomic, environmental, and phenomics data are applied to the statistical model to identify GWAS analysis results that describe the relationship between the training patient cohort and specific pharmacological phenotypes, pharmacological / pharmacometabolomics results, precise phenomics analysis results, biomarkers, etc.
[0191] At the first time point in the timeline 600, omics, sociomics, physiological omics and environmental data are collected from the current patient. The patient's phenotypic omics state is now negative 622. Then, the current patient begins to experience disease symptoms and is subsequently hospitalized, resulting in a further negative phenotypic omics state 624 at the second time point. All of this information (including the patient's response to treatment due to hospitalization 626) is provided to the clinician's clinical decision support module 614. The clinician's clinical decision support module 614 then identifies the pharmacological phenotype of the current patient based on omics, sociomics, physiological omics and environmental data, and provides, for example, a treatment option 628 with the highest predicted response to the current patient. Therefore, the current patient's phenotypic omics state changes from negative 622, 624 to positive 630 and remains positive 632 at a subsequent time point.
[0192] Figure 7 A flow chart representing an exemplary method 700 for identifying a pharmacological phenotype using machine learning techniques is depicted. The method 700 may be executed on the pharmacological phenotype assessment server 102. In some embodiments, the method 700 may be implemented in a set of instructions stored on a non-transitory computer readable memory and executable on one or more processors on the pharmacological phenotype assessment server 102. For example, the method 700 may be executed by Figure 1A The training module 160 and the phenotype assessment module 162 are executed within the machine learning engine 146.
[0193] At frame 702, training module 160 can obtain one group of training data, and described training data comprises the group of science, social group of science, physiological group of science and environmental data of training patient, wherein known training patient has pharmacological phenotype (for example, with current or previously determined pharmacological phenotype).Environmental science, physiological group of science and social group of science data can comprise any other suitable data of clinical data, demographic data, polypharmacy data, socioeconomic data, education data, drug abuse data, diet and exercise data, law enforcement data, circadian rhythm data, family data or the environment of indication patient.Omics data can comprise the genomics data of indication gene traits, the epigenomics data of indication gene expression, the transcriptomics data of indication DNA transcription, the proteomics data of indication by genomic expression protein, the chromosome group of science data of chromatin state in indication genome and / or the metabolomics data of metabolite in indication genome.As mentioned above, can obtain the group of science, social group of science, physiological group of science and environmental data of training patient at several time points (for example, on the time span of 3 years).
[0194] The socio-omic, physio-omic, and environmental data can be obtained from electronic medical records (EMRs) located on an EMR server and / or from multi-pharmacy data located on a multi-pharmacy server that aggregates pharmacy data for patients from multiple pharmacies. Additionally, the socio-omic, physio-omic, and environmental data can be obtained from the training patient's healthcare provider or from self-reports of the training patient. In some embodiments, the training data can be obtained from a combination of sources including several servers (e.g., an EMR server, a multi-pharmacy server, etc.) and client devices 106-116 of healthcare providers and patients.
[0195] Omics data can be obtained from the client device 106-116 of the healthcare provider. For example, the healthcare provider can obtain a biological sample (e.g., from saliva, biopsy, blood sample, bone marrow, hair, sweat, odor, etc.) for measuring the patient's omics and provide the laboratory results obtained by analyzing the biological sample to the pharmacological phenotype assessment server 102. In other embodiments, omics data can be obtained directly from the laboratory analyzing the biological sample. In other embodiments, omics data can be obtained from GWAS or candidate gene association studies that describe the relationship between a training patient cohort and a specific pharmacological phenotype.
[0196] The training module 160 can also obtain phenotypic data related to the pharmacological phenotype of the training patients, such as the chronic diseases suffered by the training patients, the pharmacological responses to drugs previously prescribed to the training patients, whether each of the training patients has a drug abuse problem, etc.
[0197] The training module 160 can then classify the omic, sociomic, physiological, and environmental data according to the pharmacological phenotype of the training patient associated with the omic, sociomic, physiological, and environmental data (block 704). The pharmacological phenotype can include at least some of the diseases that the training patient has been diagnosed with, substance abuse problems, pharmacological responses to various drugs, complications, etc. In some cases, the omic, sociomic, physiological, and environmental data of the same training patient can be subdivided into multiple pharmacological phenotypes at different time periods. For example, as the environmental data of the training patient changes, the omic data of the training patient may change over time, and in a first time period, the training patient may have one mental illness, while in a second time period, the training patient may have another mental illness or complication.
[0198] Similarly, in some embodiments, the omics, sociomics, physiomic, and environmental data of training patients can be categorized based on demographics. For example, the omics, sociomics, physiomic, and environmental data of training patients of European descent can be assigned to one cohort, while the omics, sociomics, physiomic, and environmental data of training patients of Chinese descent can be assigned to another cohort. In another example, the omics, sociomics, physiomic, and environmental data of training patients aged between 25 and 35 years old can be assigned to one cohort, while the omics, sociomics, physiomic, and environmental data of training patients aged between 35 and 45 years old can be assigned to another cohort.
[0199] Various machine learning techniques can then be used to analyze the training patient's omics, sociomics, physiomic, and environmental data and their respective pharmacological phenotypes to generate statistical models that indicate the likelihood or other semi-quantitative and quantitative metrics that the current patient has various pharmacological phenotypes (block 706). The statistical models can also be used to determine response scores, dosages, or any other suitable indication of the current patient's predicted response to various drugs.
[0200] For example, as mentioned above Figure 4B Described, use various machine learning techniques to analyze statistical tests from GWAS or candidate gene association studies, the relationship between the study indication training patient cohort and specific pharmacological phenotype, thereby to the variants identified in the study and / or variants with close association with the identified variants are scored and / or ranked. The highest ranked variants can be identified as SNPs, genes and genomic regions that are associated or strongly associated with the specific pharmacological phenotype.
[0201] For example, for the warfarin phenotype, warfarin response pathways (e.g. Figure 4FThe warfarin response pathway includes 74 SNPs in 31 genes expressed in the liver, small intestine, and vasculature. The warfarin response pathway includes the following genes: AKR1C3 (expressed in the liver), CYP2C19 (expressed in the liver), CYP2C8 (expressed in the liver), CYP2C9 (expressed in the liver), CYP4F2 (expressed in the liver), F5 (expressed in the liver), F7 (expressed in the liver), F10 (expressed in the liver), F11 (expressed in the liver), FGG (expressed in the liver), ORM1 (expressed in the liver), PRSS53 (expressed in the liver), VKORC1 (expressed in the liver), STX4 (expressed in the small intestine), F13A1 (expressed in the vasculature), PR OCR (expressed in the vasculature), VWF (expressed in the vasculature), CFHR5 (expressed in the liver), FGA (expressed in the liver), FMO5 (expressed in the liver), HRG (expressed in the liver), KNG1 (expressed in the liver), SURF4 (expressed in the liver), ABO (expressed in the small intestine), LYZ (expressed in the small intestine), PCGF3 (expressed in the small intestine), PRSS8 (expressed in the small intestine), TRPC4AP (expressed in the small intestine), SLC44A2 (expressed in the vasculature), SPHK1 (expressed in the vasculature), and USP7 (expressed in the vasculature). The 74 SNPs included in 31 genes were: rs12775913 (regulatory SNP), rs346803 (regulatory SNP), rs346797 (regulatory SNP), rs762635 (regulatory SNP), and rs76896860 (regulatory SNP) in the AKR1C3 gene; rs3758581 (coding SNP) in the CYP2C19 gene; rs10509681 (coding SNP) and rs11572080 (coding SNP) in the CYP2C8 gene; rs1057910 (coding SNP), rs1799853 (coding SNP), and rs7900194 (coding SNP) in the CYP2C9 gene. coding SNP); rs2108622 (coding SNP) contained in the CYP4F2 gene; rs6009 (regulatory SNP), rs11441998 (regulatory SNP), rs2026045 (regulatory SNP), rs34580812 (regulatory SNP), rs749767 (regulatory SNP), rs9378928 (regulatory SNP), and rs7937890 (regulatory SNP) contained in the F5 gene; rs7552487 (regulatory SNP), rs6681619 (regulatory SNP), rs8102532 (regulatory SNP), rs491098 (coding SNP), and rs6046 (coding SNP) contained in the F7 gene;rs11150596 (regulatory SNP) and rs11150596 (regulatory SNP) contained in the F10 gene; rs2165743 (regulatory SNP) and rs11252944 (regulatory SNP) contained in the F11 gene; rs8050894 (regulatory SNP) contained in the FGG gene; rs10982156 (regulatory SNP) contained in the ORM1 gene; rs7199949 (coding SNP) contained in the PRSS53 gene; rs2884737 (regulatory SNP), rs9934438 ( rs897984 (regulatory SNP), and rs17708472 (regulatory SNP); rs35675346 (regulatory SNP) and rs33988698 (regulatory SNP) contained in the STX4 gene; rs5985 (coding SNP) contained in the F13A1 gene; rs867186 (coding SNP) contained in the PROCR gene; rs75648520 (regulatory SNP), rs55734215 (regulatory SNP), rs12244584 (regulatory SNP), and rs1063856 (coding SNP) contained in the VWF gene. SNP); rs674302 (regulatory SNP) in the CFHR5 gene; rs12928852 (regulatory SNP) and rs6050 (coding SNP) in the FGA gene; rs8060857 (regulatory SNP) and rs7475662 (regulatory SNP) in the FMO5 gene; rs9898 (coding SNP) in the HRG gene; rs710446 (coding SNP) in the KNG1 gene; rs11577661 (regulatory SNP) in the SURF4 gene; rs11427 (regulatory SNP) in the ABO gene. 024 (regulatory SNP), rs6684766 (regulatory SNP), rs2303222 (regulatory SNP), rs1088838 (regulatory SNP), rs13130318 (regulatory SNP), and rs12951513 (regulatory SNP); rs8118005 (regulatory SNP) contained in the LYZ gene; rs76649221 (regulatory SNP), rs9332511 (regulatory SNP), and rs6588133 (regulatory SNP) contained in the PCGF3 gene; rs11281612 (regulatory SNP) contained in the PRSS8 gene;rs11589005 (regulatory SNP), rs8062719 (regulatory SNP), rs889555 (regulatory SNP), rs36101491 (regulatory SNP), rs7426380 (regulatory SNP), rs6579208 (regulatory SNP), rs77420750 (regulatory SNP), and rs73905041 (coding SNP) in the TRPC4AP gene; rs3211770 (regulatory SNP), rs3211770 (regulatory SNP), rs3087969 (coding SNP), and rs2288904 (coding SNP) in the SLC44A2 gene; rs683790 (regulatory SNP) and rs346803 (coding SNP) in the SPHK1 gene; and rs201033241 (coding SNP) in the USP7 gene.
[0202] In another example, for the lithium phenotype, a lithium response pathway can be identified that includes 78 SNPs in 12 genes expressed in the brain.
[0203] Environmental, sociomic, physiological, and phenotypic data can be obtained for training patients with appropriate combinations of SNPs, genes, and genomic regions to distinguish between training patients who have identified SNPs, genes, and genomic regions and do have a specific pharmacological phenotype and training patients who have identified SNPs, genes, and genomic regions but do not have a specific pharmacological phenotype.
[0204] Statistical models for predicting pharmacological phenotypes can be generated using machine learning techniques, including but not limited to regression algorithms (e.g., ordinary least squares regression, linear regression, logistic regression, stepwise regression, multivariate adaptive regression splines, local estimated scatter plot smoothing, etc.), instance-based algorithms (e.g., k-nearest neighbors, learning vector quantization, self-organizing maps, local weighted learning, etc.), regularization algorithms (e.g., ridge regression, least absolute shrinkage and selection operators, elastic net, least angle regression, etc.), decision tree algorithms (e.g., classification and regression trees, iterative bisection3, C4.5, C5, chi-square automatic interaction detection, decision stumps, M5, conditional decision trees, etc.), clustering algorithms (e.g., k-means, k-medians, expectation maximization, hierarchical clustering, etc.), , spectral clustering, mean shift, density-based spatial clustering of applications with noise, sorting points for identifying cluster structures, etc.), association rule learning algorithms (e.g., prior algorithms, Eclat algorithms, etc.), Bayesian algorithms (e.g., naive Bayes, Gaussian naive Bayes, multinomial naive Bayes, averaged single dependency estimator, Bayesian belief network, Bayesian network, etc.), artificial neural networks (e.g., perceptron, Hopfield network, radial basis function network, etc.), deep learning algorithms (e.g., multilayer perceptron, deep Boltzmann machine, deep belief network, convolutional neural network, stacked autoencoder, generative adversarial network, etc.), dimensionality reduction algorithms (e.g., principal component analysis, principal component regression, partial least squares regression, Sammon map, etc.), Mapping), multidimensional scaling, projection pursuit, linear discriminant analysis, mixed discriminant analysis, quadratic discriminant analysis, flexible discriminant analysis, factor analysis, independent component entity analysis, non-negative matrix factorization, t-distributed stochastic neighbor embedding, etc.), ensemble algorithms (e.g., boosting, bootstrap aggregation, AdaBoost, stacked generalization, gradient boosting machines, gradient boosted regression trees, random decision forests, etc.), reinforcement learning (e.g., temporal difference learning, Q-learning, learning automata, state-action-reward-state-action, etc.), support vector machines, mixture models, evolutionary algorithms, probabilistic graphical models, etc. For example, in addition to statistical measures of sociomic, physiological, and environmental data (e.g., the average income of training patients corresponding to each classification or sociomic data such as the average ACE score, etc.), statistical models can also be generated based on identified SNPs, genes, and genomic regions.
[0205] Furthermore, the training module 160 can generate several statistical models for several pharmacological phenotypes. For example, a first statistical model can be generated to determine the likelihood that the current patient will experience a substance abuse problem, a second statistical model can be generated to determine the risk of developing a disease, a third statistical model can be generated to determine the risk of developing another disease, a fourth statistical model can be generated to determine the likelihood of having a negative reaction to a particular drug, and so on. In any case, each statistical model can be a graphical model, a decision tree, a probability distribution, or any other suitable model for determining the likelihood that the current patient has a certain pharmacological phenotype or a drug response score based on the training data.
[0206] At block 708, the omics, sociomics, physiomic, and environmental data for the current patient may be obtained. Figure 5 The described process 500 is similar to the process of obtaining omics data. For example, a healthcare provider can obtain a biological sample of a patient and send it to an analytical laboratory for analysis. Cells are then extracted from the biological sample and reprogrammed into stem cells, such as iPSCs. The iPSCs are then differentiated into various tissues, such as neurons, cardiomyocytes, etc., and analyzed to obtain omics data. The omics data can include genomic data, epigenomic data, transcriptomic data, proteomic data, genomic data, metabolomic data and / or biological networks. Specifically, the omics data can include quantitatively evaluating the patient's current medication by performing metabolomic measurements on the patient sample.
[0207] Socio-, physio-, and environmental data, for example, in the exposure groups, socio- and demographics as described above, as well as medical physio-, EHR, laboratory values, stress and abuse factors and trauma, and medical outcome data, can be collected at several time points. For example, the socio-, physio-, and environmental data for the current patient may indicate that the current patient was single, then married in year 1, and then divorced in year 2. The socio-, physio-, and environmental data may also indicate that the current patient was employed in year 1, then lost his job in year 2. Furthermore, the socio-, physio-, and environmental data may indicate that the current patient was a victim of domestic abuse in year 1. The longitudinal data can be compared to similar experiences of training patients over similar time periods as shown in the statistical model.
[0208] Then, at block 710, the omics, sociomics, physiomic, and environmental data of the current patient may be applied to the statistical model to determine the pharmacological phenotype of the current patient. The pharmacological phenotype may include the likelihood that the current patient suffers from various diseases or a response score indicating the current patient's expected response to various drugs and the recommended dosage of the drug. For example, if the statistical model is a neural network, the phenotype assessment module 162 may use the omics, sociomics, physiomic, and environmental data of the current patient to traverse the nodes of the neural network to reach each output node to determine the likelihood or response score. If several statistical models are generated, the phenotype assessment module 162 may apply the omics, sociomics, physiomic, and environmental data to each of the statistical models to determine, for example, the likelihood or risk of suffering from bipolar disorder, the likelihood or risk of suffering from schizophrenia, the likelihood of having a substance abuse problem, the response score for taking lithium to treat bipolar disorder, and the like.
[0209] For example, the omics data of the current patient can be analyzed to identify SNPs and genes in the omics data of the current patient that are identical to any of the 74 SNPs or 31 genes in the warfarin response pathway that are associated with the warfarin phenotype, thereby determining whether the current patient has any of the warfarin phenotypes. In addition, the sociomics, physiomic, and environmental data of the current patient can be applied to a warfarin statistical model to determine the warfarin phenotype of the current patient. In addition to the statistical measurements performed on the sociomics, physiomic, and environmental data, a warfarin statistical model can also be generated based on the identified 74 SNPs, 31 genes, and warfarin response pathway.
[0210] In another example, the omics data of the current patient can be analyzed to identify SNPs and genes in the omics data of the current patient that are identical to any of the 78 SNPs or 12 genes in the lithium response pathway that are associated with the lithium phenotype, thereby determining whether the current patient has any of the lithium phenotypes. In addition, the sociomics, physiomic, and environmental data of the current patient can be applied to a lithium statistical model to determine the lithium phenotype of the current patient. In addition to the statistical measurements performed on the sociomics, physiomic, and environmental data, a lithium statistical model can be generated based on the identified 78 SNPs, 12 genes, and lithium response pathways.
[0211] At block 712, the phenotype assessment module 162 may display one or more indications of the pharmacological phenotype of the current patient on a user interface of the healthcare provider's client device. For example, the phenotype assessment module 162 may generate a risk analysis display that includes an indication of each of the likelihood or other semi-quantitative and quantitative measures of various pharmacological phenotypes, which may be expressed as a probability (e.g., 0.6), a percentage (e.g., 80%), one of a set of categories (e.g., "high risk," "moderate risk," or "low risk"), and / or in any other suitable manner. Additionally, the risk analysis display may include an indication of a response score to the drug, which may be numerical (e.g., 75 out of 100), categorical (e.g., "strong response," "adverse reaction," etc.), or in any other suitable manner. Each likelihood or response score may be displayed along with a description of the corresponding pharmacological phenotype (e.g., "high risk for disease Y"). In this manner, the healthcare provider of the current patient can review the indication of the pharmacological phenotype of the current patient and develop an appropriate treatment plan or regimen. For example, a healthcare provider may prescribe a drug to treat a specific disease that has the highest response score among drugs that treat the specific disease.
[0212] In some embodiments, the pharmacological phenotype assessment server 102 can compare the likelihood of the pharmacological phenotype with a likelihood threshold (e.g., 0.5) and can include the pharmacological phenotype in the risk analysis display when the likelihood of the pharmacological phenotype exceeds the likelihood threshold. For example, only diseases that are high risk for the current patient can be included in the risk analysis display. In another example, the risk analysis display can include an indication that the current patient may have a drug abuse problem when the likelihood of the current patient having a drug abuse problem exceeds the likelihood threshold. In this way, a healthcare provider can recommend or provide early intervention to the current patient. Similarly, in some embodiments, the response score of each drug corresponding to a specific disease can be ranked (e.g., from highest to lowest). The drugs and corresponding response scores can be provided in the order of ranking on the risk analysis display. In other embodiments, only the highest-ranked drugs for a specific disease can be included in the risk analysis display.
[0213] In addition to displaying the highest-ranked medications for a particular disease as recommendations to healthcare providers for prescribing to a patient, the risk analysis display may also include recommended dosages of medications for the current patient. The risk analysis display may also include any socio-omics, physio-omics, environmental, or demographic information that may cause the current patient's response to the medication to change (e.g., changes in diet, exercise, exposure, etc.). In addition, the risk analysis display may include recommendations to increase or decrease the amount of medication the current patient is taking based on their polypharmacy data. For example, when the recommended medication may make one or more medications in the current patient's polypharmacy data redundant, the risk analysis display may include recommendations to stop taking these medications. Such recommendations may involve one or more medications, combinations of medications, or other therapeutic measures.
[0214] In some embodiments, if a patient is currently taking a medication that is incompatible with the highest-ranked medication, the pharmacological phenotyping server 102 may recommend a medication with the next highest response score. For example, the pharmacological phenotyping server 102 may access polypharmacy data for the current patient and compare the medications in the polypharmacy data with the recommended medications to check for contraindications. The risk analysis display may include the highest-ranked medication that has no contraindications with any medication prescribed to the current patient.
[0215] As in the example above about warfarin, the omics data of the current patient can be compared with 74 SNPs or 31 genes in the warfarin reaction pathway that are associated with the warfarin phenotype to determine whether the current patient has any of the warfarin phenotypes. Then, the pharmacological phenotype assessment server 102 or a healthcare provider can determine whether warfarin or another anticoagulant should be administered to the current patient based on the comparison. The recommended dose of warfarin can also be determined. For example, the current patient may have a SNP or gene in the warfarin reaction pathway that is associated with a negative reaction to warfarin. Therefore, the pharmacological phenotype assessment server 102 can recommend another anticoagulant to administer to the current patient. In another example, the current patient may have a SNP or gene in the warfarin reaction pathway that is associated with a warfarin dosage phenotype. Therefore, the pharmacological phenotype assessment server 102 can provide a recommended dose based on the warfarin dosage phenotype to administer warfarin to the current patient. In yet another example, the current patient may have a SNP or gene in a warfarin response pathway that is associated with disease risk, wherein warfarin can actively prevent coagulation, clotting, or thrombosis. In any case, the healthcare provider can administer warfarin to the current patient at the recommended dose, or can administer another anticoagulant.
[0216] As in the example above regarding lithium, the omics data of the current patient can be compared with 78 SNPs or 12 genes in the lithium reaction pathway that are associated with the lithium phenotype to determine whether the current patient has any of the lithium phenotypes. The pharmacological phenotype assessment server 102 or the healthcare provider can then determine whether lithium or another psychotropic drug should be administered to the current patient based on the comparison. The recommended dose of lithium can also be determined. For example, the current patient may have a SNP or gene in the lithium reaction pathway that is associated with a negative reaction to lithium. Therefore, the pharmacological phenotype assessment server 102 can recommend another psychotropic drug to be administered to the current patient. In another example, the current patient may have a SNP or gene in the lithium reaction pathway that is associated with a lithium dosage phenotype. Therefore, the pharmacological phenotype assessment server 102 can provide a recommended dose to administer lithium to the current patient based on the lithium dosage phenotype. In any case, the healthcare provider can administer lithium to the current patient at the recommended dose, or can administer another psychotropic drug.
[0217] In addition to clinical settings, pharmacologic phenotypes can be predicted in research settings for drug development and insurance applications. In a research setting, the pharmacologic phenotype of a potential patient cohort associated with an experimental drug might be predicted during a study plan. Patients can be selected for experimental treatment based on their predicted pharmacologic phenotype associated with the experimental drug.
[0218] Additionally, when the pharmacological phenotype of the current patient becomes known (e.g., the current patient has a pharmacological phenotype after a threshold amount of time, such as one year), the current patient's omics, sociomics, physiomics, and environmental and phenomics data can be added to the training data (block 714), and the statistical model can be updated accordingly. In some embodiments, the omics, sociomics, physiomics, environmental, and phenomics data are stored in several data sources 716, such as Figure 3 The training module 160 can then retrieve data from the data source 716 to further train the model.
[0219] Throughout the entire specification, multiple instances can be implemented as components, operations or structures described as single instances. Although the individual operations of one or more methods are shown and described as separate operations, one or more of the individual operations can be performed simultaneously, and the operations do not need to be performed in the order shown. The structure and function presented as a separate component in the example configuration can be implemented as a combined structure or component. Similarly, the structure and function presented as a single component can be implemented as a separate component. These and other variations, modifications, additions and improvements all fall within the scope of this paper theme.
[0220] In addition, some embodiments are described herein as comprising logic or multiple routines, subroutines, applications or instructions. These can constitute software (e.g., code embodied on a machine-readable medium or in a transmission signal) or hardware. In hardware, routines, etc. are tangible units that can perform certain operations and can be configured or arranged in a certain way. In example embodiments, one or more computer systems (e.g., independent client or server computer systems) or one or more hardware modules (e.g., processors or processor groups) of a computer system can be configured by software (e.g., applications or application portions) to operate as hardware modules that perform certain operations as described herein.
[0221] In various embodiments, hardware modules may be implemented mechanically or electronically. For example, a hardware module may include dedicated circuitry or logic that is permanently configured to perform certain operations (e.g., a dedicated processor such as a field programmable gate array (FPGA) or an application-specific integrated circuit (ASIC)). A hardware module may also include programmable logic or circuitry (e.g., as contained in a dedicated processor or other programmable processor) that is temporarily configured to perform certain operations via software. It will be appreciated that the decision to implement a hardware module mechanically, in a dedicated and permanently configured circuit, or in a temporarily configured circuit (e.g., configured via software), may be driven by cost and time considerations.
[0222] Accordingly, the term "hardware module" should be understood to encompass a tangible entity that is physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate in a certain manner or perform any of the operations described herein. Contemplating embodiments in which hardware modules are temporarily configured (e.g., programmed), it is not necessary to configure or instantiate each hardware module at any one time. For example, where the hardware modules include a general-purpose processor configured using software, the general-purpose processor can be configured into corresponding different hardware modules at different times. Thus, software can configure a processor, for example, to constitute a particular hardware module at one time and to constitute a different hardware module at a different time.
[0223] A hardware module can provide information to other hardware modules and receive information from other hardware modules. Thus, the described hardware modules can be considered to be communicatively connected. When multiple such hardware modules exist simultaneously, communication can be achieved by signal transmission (e.g., through appropriate circuits and buses) connecting the hardware modules. In an embodiment where multiple hardware modules are configured or instantiated at different times, communication between such hardware modules can be achieved, for example, by storing and retrieving information in a memory structure accessible to multiple hardware modules. For example, a hardware module can perform an operation and store the output of the operation in a memory device to which it is communicatively connected. Then, another hardware module can access the memory device at a later time to retrieve and process the stored output. The hardware module can also initiate communication with an input or output device and can operate on a resource (e.g., an information collection).
[0224] The various operations of the example methods described herein may be performed, at least in part, by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors may constitute processor-implemented modules that operate to perform one or more operations or functions. In some example embodiments, the modules mentioned herein may include processor-implemented modules.
[0225] Similarly, the methods or routines described herein can be implemented at least in part by a processor. For example, at least some of the operations in the method's operation can be performed by one or more processors or a hardware module implemented by the processor. Some of the execution in the operation can be distributed between one or more processors that not only reside in a single machine but also across multiple machine deployments. In some example embodiments, one or more processors can be located in a single location (e.g., in a home environment, in an office environment, or as a server farm), and in other embodiments, the processor can be distributed across multiple locations.
[0226] The execution of some of the operations may be distributed between one or more processors that reside not only within a single machine but also across multiple machines. In some example embodiments, one or more processors or processor-implemented modules may be located in a single geographic location (e.g., in a home environment, an office environment, or a server farm). In other example embodiments, one or more processors or processor-implemented modules may be distributed across multiple geographic locations.
[0227] Unless specifically stated otherwise, discussions herein using terms such as "process," "calculate," "calculate," "determine," "present," "display," etc. may refer to the action or process of a machine (e.g., a computer) that manipulates or transforms data represented as physical (e.g., electrical, magnetic, or optical) quantities in one or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other components of the machine that receive, store, send, or display information.
[0228] As used herein, any reference to "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. The appearances of the phrase "in one embodiment" in various places in the specification are not necessarily all referring to the same embodiment.
[0229] The expressions "coupled" and "connected" and their derivatives may be used to describe some embodiments. For example, the term "coupled" may be used to describe some embodiments to indicate that two or more elements are in direct physical or electrical contact. However, the term "coupled" may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other. The embodiments are not limited in scope to these terms.
[0230] As used herein, the terms "comprises / comprising," "includes / including," "has / having," or any other variations thereof, are intended to cover a non-exclusive inclusion. For example, a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to only those elements but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Furthermore, unless expressly stated to the contrary, "or" refers to an inclusive "or" and not to an exclusive "or." For example, any of the following satisfies condition A or B: A is true (or exists) and B is false (or does not exist), A is false (or does not exist) and B is true (or exists), and both A and B are true (or exist).
[0231] In addition, "a" and "an" are used to describe elements and components of the embodiments herein. This is merely for convenience and to provide a general description. This description and the appended claims should be understood to include one or at least one, and unless otherwise clearly indicated, the singular also includes the plural.
[0232] This detailed description should be construed as providing examples only and does not describe every possible embodiment because describing every possible embodiment would be impractical, if not impossible. Many alternative embodiments can be implemented using current technology or technology developed after the filing date of this application.
Claims
1. A computing device for identifying a pharmacological phenotype, the computing device comprising: one or more processors; as well as a non-transitory computer-readable memory coupled to the one or more processors and storing thereon instructions that, when executed by the one or more processors, cause the computing device to: Identify multiple single nucleotide polymorphisms (SNPs) associated with a specific set of pharmacological phenotypes; comparing the plurality of SNPs to a database of SNPs to identify additional SNPs linked to the plurality of SNPs, wherein the plurality of SNPs and the additional SNPs are included in a set of permissive candidate variants; performing a bioinformatics analysis to filter the set of permissive candidate variants into a subset of intermediate candidate variants based on at least one of: regulatory function, variant dependency, presence of a target gene relationship of the permissive candidate variant, or whether the permissive candidate variant is a non-synonymous coding variant with a minor allele frequency; Network analysis is performed to filter a subset of intermediate candidate variants to candidate variants of SNPs, genes associated with SNPs, and / or networks associated with SNPs that are causally related to the specific set of pharmacological phenotypes.
2. The computing device according to claim 1, wherein: To perform a bioinformatics analysis to filter the set of permissive candidate variants into a subset of intermediate candidate variants, the instructions cause the computing device to: The regulatory function of the genomic regions surrounding the set of permissive candidate variants is assessed.
3. A computing device according to claim 2, wherein the regulatory function of the genomic region surrounding the set of permissive candidate variants is evaluated to achieve at least one of the following: (1) determine whether the sequence context of the set of permissive candidate variants affects the regulatory function, or (2) determine the target gene of the set of permissive candidate variants.
4. The computing device of claim 3, wherein the target genes of the set of permissive candidate variants are determined using machine learning techniques.
5. The computing device according to claim 1, wherein: To perform a bioinformatics analysis to filter the set of permissive candidate variants into a subset of intermediate candidate variants, the instructions cause the computing device to: Each permissive candidate variant is scored based on at least one of the following: regulatory function, variant dependency, the presence of a permissive candidate variant-target gene relationship, or whether the permissive candidate variant is a nonsynonymous coding variant with a minor allele frequency; as well as One or more permissive candidate variants with scores above a threshold are assigned to a subset of intermediate candidate variants. The computing device according to claim 1 , wherein: The specific set of pharmacological phenotypes comprises a set of warfarin phenotypes, and the plurality of SNPs identified as causally associated with the set of warfarin phenotypes comprises a warfarin response pathway.
7. The computing device of claim 1 , wherein the specific set of pharmacological phenotypes comprises at least one of: a predicted response to one or more drugs, a risk of one or more diseases, an adverse drug event or adverse drug reaction to the one or more drugs, or a potential for drug abuse.
8. The computing device of claim 1 , wherein the instructions further cause the computing device to: obtaining biological samples from patients; comparing the biological sample to SNPs causally associated with the specific set of pharmacological phenotypes; and Based on the comparison, an indication of one or more pharmacological phenotypes of the patient is provided for display to a healthcare provider.
9. The computing device of claim 8, wherein to provide one or more pharmacological phenotypes of a patient to a healthcare provider for display, the instructions cause the computing device to: A risk profile is generated for the patient comprising at least one of the patient's predicted response to each of one or more drugs, risk of one or more diseases, or potential for drug abuse.