System and method for predicting effective and safe drug therapy
A computer-implemented system with a neural network language model efficiently classifies feature vectors for drug response predictions, addressing the complexity of pharmacogenetic guideline development and enhancing personalized medicine by reducing adverse drug reactions.
Patent Information
- Application Number
- PCT/US2025/024165
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-14
- Filing Date
- 2025-04-10
- Publication Date
- 2025-10-16
AI Technical Summary
The current pace of developing pharmacogenetic guidelines is lengthy and complex, making personalized medicine benefits inaccessible to a wide range of patients.
A system and method using a computer-implemented approach with a structured pharmacogenomic dataset and a neural network language model to classify feature vectors as serviceable or non-serviceable for drug response predictions, generating personalized drug therapy recommendations through machine learning.
Accelerates the development of pharmacogenetic guidelines, providing high-accuracy, personalized drug therapy recommendations that enhance efficacy and reduce adverse drug reactions.
Smart Images

Figure US2025024165_16102025_PF_FP_ABST
Abstract
Description
SYSTEM AND METHOD FOR PREDICTING EFFECTIVE AND SAFE DRUG THERAPYCROSS REFERENCE
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 633,469, filed on April 12, 2024, and U.S. Provisional Application No. 63 / 758,970, filed on February 14, 2025, each of which is entirely incorporated herein by reference.FIELD
[0002] In some aspects, the present disclosure relates to personalized medicine and more particularly, to a system and method for predicting effective drug therapy.BACKGROUND
[0003] Modern medical practice increasingly relies on pharmacogenetics as a key component of personalized medicine. Pharmacogenetic studies contribute to the development of guidelines that consider individual genetic characteristics of patients when prescribing drugs. Currently, according to the Pharmacogenomics Knowledgebase (PharmGKB) database, which specializes in pharmacogenomic knowledge, there are 202 pharmacogenetic guidelines, each of which undergoes a rigorous review and approval process.
[0004] The development of new guidelines is a lengthy and complex process that requires not only the collection and analysis of a large volume of scientific data but also multiple expert reviews and iterations to ensure the accuracy and reliability of the recommendations provided. The problem is that at the current pace of development, many potential benefits of pharmacogenetic testing remain inaccessible to a wide range of patients.SUMMARY
[0005] In the present disclosure, various methods and systems are provided for efficiently providing pharmacogenetic guidelines. A method can include inputting a drug prescription and a patient genotype and analyzing the drug prescription and the patient genotype to provide a drug therapy recommendation. The analysis can include comparing the patient genotype and the drug name to a drug profile database, a molecular biomarker database, hospital drug reaction records, and omics databases, and calculating the patient response to the drug. The system can include a computer input system for entering a drug prescription and a patient genotype and a computer processor for analyzing the drug prescription and patient genotype by comparing them to databases, including a drug profile database, a molecular biomarker database, hospital drug reactions records, and omics database.
[0006] In some aspects, the present disclosure provides a computer-implemented system comprising: a database comprising a structured pharmacogenomic dataset; a machine learning model comprising a neural network language model; and computer-executable instructions configured to: train the machine learning model, using the structured pharmacogenomic dataset, to: classify a feature vector as (1) a serviceable feature vector or (2) a non-serviceable feature vector for a drug response prediction; and generate for the serviceable feature vector, using the neural network language model, the drug response prediction in natural language; update the structured pharmacogenomic dataset with a new pharmacogenomic dataset to generate an updated structured pharmacogenomic dataset; and train a new machine learning model, using the updated structured pharmacogenomic dataset, such that the new machine learning model has more serviceable feature vectors than the machine learning model.
[0007] In some aspects, the present disclosure provides a computer-implemented system comprising: a database comprising a structured pharmacogenomic dataset; a machine learning model comprising a neural network language model; and computer-executable instructions configured to: train the machine learning model, using the structured pharmacogenomic dataset, to: classify a feature vector as (1) a serviceable feature vector or (2) a non-serviceable feature vector for a drug response prediction; and generate for the serviceable feature vector, using the neural network language model, the drug response prediction in natural language; update the structured pharmacogenomic dataset with a new pharmacogenomic dataset to generate an updated structured pharmacogenomic dataset; and retrain the machine learning model, using the updated structured pharmacogenomic dataset, to add new serviceable feature vectors to the machine learning model.
[0008] In some embodiments, the structured pharmacogenomic dataset comprises structural representations of pharmaceutical substances. In some embodiments, the structural representations comprise: molecular structures, fingerprints, latent vectors, SMILES (Simplified Molecular Input Line Entry System), InChi (International Chemical Identifier), SDF (Structure Data File), topological descriptors, 3D conformations, graph adjacency matrices, node and edge feature vectors for each atom and bond, autoencoder-derived embeddings, transformer-based embeddings for small molecules, graph neural network latent vectors, quantum chemical descriptors (partial charges, H0M0-LUM0 levels), protein-ligand interaction features (docking scores, binding affinity), pharmacophore or toxicophore features, or any combination thereof. In some embodiments, the structured pharmacogenomic dataset comprises physical or chemical properties of pharmaceutical substances. In some embodiments, the physical or chemical properties comprise: pharmacokinetic properties, pharmacodynamic properties, molecular weights, logP, logD, pKa, solubility, melting point, boiling point, hydrogen bond donor count,hydrogen bond acceptor count, topological polar surface area (tPSA), rotatable bond count, refractivity, or any combination thereof. In some embodiments, the structured pharmacogenomic dataset comprises genetic profiles of subjects. In some embodiments, the genetic profiles comprise: genomic sequences, mutations, SNPs, copy number variations (CNVs), insertions, deletions, haplotypes, splicing variants, epigenetic modifications, gene expression levels, genomic rearrangements, or any combination thereof. In some embodiments, the feature vector comprises a structural representation of a pharmaceutical substance. In some embodiments, the feature vector comprises a physical or chemical property of a pharmaceutical substance.
[0009] In some embodiments, (i) the serviceable feature vector, (ii) the non-serviceable feature vector, or (iii) both is input to the neural network language model as a structured prompt. In some embodiments, the structured prompt comprises: a feature definition, a feature label, a feature value, a prefix, a suffix, a question, a chain of thought, or any combination thereof. In some embodiments, the feature definition comprises context for a feature in natural language. In some embodiments, the feature label comprises a database identifier for a feature. In some embodiments, the feature value comprises a quantitative value, a qualitative value, In some embodiments, the structured prompt comprises a plurality of features. In some embodiments, the plurality of features comprises: a molecular structure of a pharmaceutical substance, a molecular fingerprint of a pharmaceutical substance, a physicochemical characteristic of a pharmaceutical substance, a pharmacodynamic genetic characteristic of a pharmaceutical substance, a pharmacokinetic genetic characteristic of a pharmaceutical substance, pharmacoepigenomics of a pharmaceutical substance, pharmacotranscriptomics of a pharmaceutical substance, pharmacoproteomics of a pharmaceutical substance, pharmacometabolomics of a pharmaceutical substance, a clinical outcome of a pharmaceutical substance, allele variants, or any combination thereof.
[0010] In some embodiments, the serviceable feature vector indicates that a clinically effective recommendation of the drug response prediction can be generated for the feature vector. In some embodiments, the non-serviceable feature vector indicates that a clinically effective recommendation of the drug response prediction cannot be generated for the feature vector. In some embodiments, the serviceable feature vector defines a genomic feature of a subject that is not found in the structured pharmacogenomic dataset. In some embodiments, the serviceable feature vector defines a combination of genomic features of the subject that is not found in the structured pharmacogenomic dataset. In some embodiments, the serviceable feature vector defines a pharmacological substance that is not found in the structured pharmacogenomic dataset. In some embodiments, the serviceable feature vector defines a combination of pharmacological substances that is not found in the structured pharmacogenomic dataset. Insome embodiments, the serviceable feature vector defines a combination of (i) a pharmacological substance and (ii) genomic features of the subject that is not found in the structured pharmacogenomic dataset. In some embodiments, the serviceable feature vector defines a combination of (i) pharmacological substances and (ii) genomic features of the subject that is not found in the structured pharmacogenomic dataset.
[0011] In some embodiments, the feature vector is classified as the serviceable feature vector or the non-serviceable feature vector based on a confidence score. In some embodiments, the confidence score is determined by processing the feature vector to quantify a similarity between (i) a pharmaceutical substance defined by the feature vector and (ii) a different pharmaceutical substance defined by the structured pharmacogenomic dataset. In some embodiments, the confidence score is determined by processing the feature vector to quantify (i) a pharmacokinetic property defined by the structured pharmacogenomic dataset, (ii) a pharmacodynamic property defined by the structured pharmacogenomic dataset, (iii) a predicted pharmacokinetic property, (iv) a predicted pharmacodynamic property, or any combination thereof. In some embodiments, the confidence score is determined by an association between a pharmaceutical substance and genomic features of the subject. In some embodiments, the confidence score is determined by processing an association between a pharmaceutical substance and (i) epigenetic features, (ii) proteomic features, (iii) transcriptomic features, (iv) metabolomic features of the subject, or any combination thereof. In some embodiments, the confidence score is determined by processing an availability and relevance of clinical outcome data for a pharmacological substance defined by the feature vector. In some embodiments, the feature vector is classified as the serviceable feature vector when the confidence score exceeds a predetermined threshold. In some embodiments, the predetermined threshold is 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, 99.91%, 99.92%, 99.93%, 99.94%, 99.95%, 99.96%, 99.97%, 99.98%, 99.99%, 99.991%, 99.992%, 99.993%, 99.994%, 99.995%, 99.996%, 99.997%, 99.998%, 99.999% confidence that the feature vector is a serviceable feature vector. In some embodiments, the predetermined threshold is 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, 99.91%, 99.92%, 99.93%, 99.94%, 99.95%, 99.96%, 99.97%, 99.98%, 99.99%, 99.991%, 99.992%, 99.993%, 99.994%, 99.995%, 99.996%, 99.997%, 99.998%, 99.999% confidence that the feature vector is a non-serviceable feature vector.
[0012] In some embodiments, the new pharmacogenomic dataset comprises: raw genetic data, clinical outcome data, scientific papers, guidelines, electronic health records, comorbidities, environmental factors, concomitant medications, adverse event reports, real-world evidence from wearable devices, retrospective and prospective cohort data, structured references from curateddatabases (e.g., ClinVar, PharmGKB), laboratory test results, phenotype data, demographic profiles, patient compliance data, toxicology studies, and multi-center consortia datasets, or any combination thereof. In some embodiments, the computer-executable instructions are configured to extract features useful for drug response prediction in the new pharmacogenomic dataset. In some embodiments, the computer-executable instructions are configured to filter mutations in the new pharmacogenomic dataset. In some embodiments, the computer-executable instructions are configured to isolate clinically significant markers, from the new pharmacogenomic dataset, that are relevant to adverse drug reactions or efficacy. In some embodiments, the computerexecutable instructions are configured to filter pharmaceutical substances in the new pharmacogenomic dataset. In some embodiments, the computer-executable instructions are configured to, in the new pharmacogenomic dataset, parse unstructured data (e.g., scientific papers, guidelines) into structured formats. In some embodiments, the computer-executable instructions are configured to, in the new pharmacogenomic dataset, annotate data with standardized ontologies (e.g., gene and drug nomenclatures). In some embodiments, the computer-executable instructions are configured to, in the new pharmacogenomic dataset, remove duplicates and resolve inconsistent entries. In some embodiments, the computerexecutable instructions are configured to apply dimensionality reduction or feature engineering techniques. In some embodiments, the computer-executable instructions are configured to integrate the new pharmacogenomic dataset with the structured pharmacogenomic dataset. In some embodiments, the computer-executable instructions are configured to, using the updated structured pharmacogenomic dataset, generate updated feature vectors for training the new machine learning model.
[0013] In some embodiments, the machine learning model comprises a binary classifier. In some embodiments, the binary classifier comprises a random forest model. In some embodiments, the binary classifier comprises a boosted tree algorithm. In some embodiments, the neural network language model comprises an autoregressive model. In some embodiments, the neural network language model comprises a transformer. In some embodiments, the neural network language model comprises a large language model.
[0014] In some embodiments, the drug response prediction comprises a personalized regimen for administering a pharmaceutical substance for a subject. In some embodiments, the drug response prediction comprises a personalized dosing guideline for a subject. In some embodiments, the drug response prediction comprises a personalized drug selection for a subject. In some embodiments, the drug response prediction comprises a prediction of an adverse drug reaction in a subject. In some embodiments, a label of the personalized regimen, the personalized dosing guideline, the personalized drug selection, or any combination thereof can be printed. In someembodiments, the label can be applied to a packaging for a pharmaceutical substance. In some embodiments, multiple drug response predictions are generated to statistically filter hallucinated drug response predictions. In some embodiments, the natural language is English, Spanish, German, French, Russian, Mandarin Chinese, Cantonese Chinese, French, Portuguese, Hindi, Korean, or Japanese.
[0015] In some aspects, the present disclosure provides a computer-implemented system comprising: a database comprising a structured dataset, wherein the structured dataset comprises a pharmacogenomic dataset relating at least 1,200 pharmacological substances and at least 350 genes; a machine learning model comprising a neural network language model; and computerexecutable instructions configured to train the machine learning model to generate a drug response prediction in natural language, wherein the training is based on the structured dataset.
[0016] In some aspects, the present disclosure provides a computer-implemented system comprising: a machine learning model comprising a neural network language model, wherein the machine learning model is trained using a structured dataset comprising a pharmacogenomic dataset relating at least 1,200 pharmacological substances and at least 350 genes; and computerexecutable instructions configured to process a feature vector, using the machine learning model, to generate a drug response prediction in natural language.
[0017] In some embodiments, the pharmacogenomic dataset comprises at least 1,700, 2,500, 3,500, 5,000, 10,000, or 21,000 pharmacological substances. In some embodiments, the pharmacogenomic dataset comprises at most 2,500, 3,500, 5,000, 10,000, 21,000, 30,000 pharmacological substances. In some embodiments, the pharmacogenomic dataset comprises at least 700, 1,300, 2,400, 4,500, 10,000, or 20,000 genes. In some embodiments, the pharmacogenomic dataset comprises at most 1,300, 2,400, 4,500, 10,000, 20,000, 25,000 genes.
[0018] In some aspects, the present disclosure provides a computer-implemented system comprising: a database comprising a structured dataset, wherein the structured dataset comprises a pharmacogenomic dataset and at least one of: (i) an epigenomic dataset, (ii) a transcriptomic dataset, (iii) a proteomic dataset, (iv) a metabolomic dataset, (v) a lipidomic dataset, and (vi) a secretomic dataset; a machine learning model comprising a neural network language model; and computer-executable instructions configured to: train the machine learning model, using the structured dataset, to: (i) classify a feature vector as (1) a serviceable feature vector or (2) a non- serviceable feature vector for a drug response prediction; and (ii) generate for the serviceable feature vector, using the neural network language model, the drug response prediction in natural language.
[0019] In some embodiments, the epigenomic dataset is a pharmacoepigenomic dataset. In some embodiments, the transcriptomic dataset is a pharmacotranscriptomic dataset. In someembodiments, the proteomic dataset is a pharmacoproteomic dataset. In some embodiments, the metabolomic dataset is a pharmacometabolomic dataset. In some embodiments, the lipidomic dataset is a pharmacolipidomic dataset. In some embodiments, the secretomic dataset is a pharmacosecretomic dataset.
[0020] In some embodiments, the drug response prediction is a prediction for using a pharmaceutical substance to treat a disease. In some embodiments, the disease is a cancer. In some embodiments, the pharmaceutical substance is a chemotherapeutic. In some embodiments, the chemotherapeutic is an EGFR inhibitor for treating lung cancer, or a PARP inhibitor for treating BRCA-mutated breast cancer. In some embodiments, the drug response prediction comprises a response prediction of an immune checkpoint inhibitor based on tumor mutational burden. In some embodiments, the drug response prediction is based tumor mutations, pharmacogenomic biomarkers, immunotherapy response predictors, or any combination thereof. In some embodiments, the disease is an autoimmune disease. In some embodiments, the autoimmune disease is rheumatoid arthritis, lupus, or multiple sclerosis. In some embodiments, the drug response prediction is based on genetic susceptibility to drug-induced side effects. In some embodiments, the drug response prediction comprises a providing a therapy selection among an array of therapies. In some embodiments, the array of therapies comprises an array of biologies. In some embodiments, the array of biologies comprises TNF inhibitors and IL-6 inhibitors. In some embodiments, the pharmaceutical substance comprises methotrexate, TNF inhibitors, or IL-6 inhibitors. In some embodiments, the disease is a metabolic disorder. In some embodiments, the metabolic disorder is obesity. In some embodiments, the pharmaceutical substance comprises a GLP-1 agonist, semaglutide, tirzepatide, or any combination thereof. In some embodiments, the disease is a psychiatric or a neurological disease. In some embodiments, the disease is epilepsy or a neurodegenerative disease. In some embodiments, the pharmaceutical substance comprises an antidepressant or an antipsychotic. In some embodiments, the pharmaceutical substance comprises an SSRI or a tricyclic antidepressant. In some embodiments, the drug response prediction provides a recommending dosing regimen of the pharmaceutical substance based on predicted metabolism of the pharmaceutical substance by a subject. In some embodiments, the disease is a cardiovascular disease. In some embodiments, the pharmaceutical substance comprises an antiplatelet or warfarin. In some embodiments, the drug response prediction is based on clopidogrel response based on CYP2C19 status. In some embodiments, the drug response prediction comprises VKORCl / CYP2C9-guided adjustments.
[0021] In some embodiments, the drug response prediction indicates a set of candidate subjects that are predicted to benefit from administration of a pharmaceutical substance. In some embodiments, the drug response prediction indicates a set of candidate subjects that are predictednot to experience adverse effects from administration of a pharmaceutical substance. In some embodiments, the drug response prediction indicates a set of candidate subjects that are predicted to experience adverse effects from administration of a pharmaceutical substance. In some embodiments, the drug response prediction identifies drug-gene interactions for a use of the pharmaceutical substance to treat a rare disease. In some embodiments, the drug response prediction comprises a companion diagnostics. In some embodiments, the system generates list of subjects having a biomarkers that match a biomarker-based clinical trial enrollment criteria. In some embodiments, the system is in operable communication with a personal device of a user to provide telemedicine services to the user. In some embodiments, the system generates policy pricing based on the drug response prediction. In some embodiments, the system generates a preventive healthcare program based on the drug response prediction. In some embodiments, the system is implemented within a wearable device. In some embodiments, the system is in operable communication with a wearable device configured to provide the system with user health data. In some embodiments, the drug response prediction is based on real-time updates of a subject’s or a plurality of subjects’ epigenomic, transcriptomic, proteomic, or metabolomic state.
[0022] In some aspects, the present disclosure provides a computer-implemented system comprising: a database comprising an internal pharmacogenomic dataset encrypted based on an internal cipher; a neural network language model; and computer-executable instructions configured to: create a secure connection with a client to receive an external pharmacogenomic dataset encrypted based on an external cipher; receive the external pharmacogenomic dataset through the secure connection; decrypt the external pharmacogenomic dataset based on the external cipher; add the external pharmacogenomic dataset to the internal pharmacogenomic dataset; decrypt the internal pharmacogenomic dataset based on the internal cipher; and train the neural network language model, using the internal pharmacogenomic dataset, to generate a drug response prediction in natural language.
[0023] In some embodiments, the computer-executable instructions are configured to provide access privileges to a plurality of clients, wherein the access privileges are based on a client profile. In some embodiments, the access privileges permit access to the neural network language model and not the database. In some embodiments, the access privileges permit access to the database and not the neural network language model. In some embodiments, the access privileges permit access to the neural network language model and the database. In some embodiments, the client is an organization. In some embodiments, the client is an individual user account. In some embodiments, the access privileges comprise (i) read access, (ii) write access, (iii) execute access, (iv) delete access, (v) full privileged access, or any combination thereof.
[0024] In some embodiments, the client profile comprises jurisdiction. In some embodiments, the jurisdiction comprises a jurisdiction of registration. In some embodiments, the jurisdiction comprises a jurisdiction of access. In some embodiments, the jurisdiction comprises US jurisdiction, European jurisdiction, Chinese jurisdiction, South Korean jurisdiction, Australian jurisdiction, African jurisdiction, Russian jurisdiction, Canadian jurisdiction, Japanese jurisdiction, Indian jurisdiction, Brazilian jurisdiction.
[0025] In some embodiments, the computer-executable instructions are configured to log traffic to and from the computer-implemented system. In some embodiments, the computer-executable instructions are configured to log user access to the computer-implemented system. In some embodiments, the computer-executable instructions are configured to log changes to the computer-implemented system.
[0026] In some aspects, the present disclosure provides a computer-implemented system comprising: a neural network language model; a plurality of computers, each computer comprising: a database comprising a pharmacogenomic dataset, wherein the pharmacogenomic dataset of each computer in the plurality of computers is different from one another; a containerized computer-executable instructions configured to train the neural network language model using the computer’s computational resources and the computer’s pharmacogenomic dataset, without sharing the computer’s pharmacogenomic dataset with another computer in the plurality of computers.
[0027] In some embodiments, the system further comprises a central computer, wherein the central computer is configured to aggregate the trained neural network language model from the plurality of computers. In some embodiments, the containerized computer computer-executable instructions are configured to run using different architectures of different computers. In some embodiments, the containerized computer computer-executable instructions are configured to load balance.
[0028] In some embodiments, the system further comprises a second containerized computerexecutable instructions configured to use the neural network language model to generate a drug response prediction in natural language for a user of the computer, without sharing the user’s input to the neural network language model with another computer in the plurality of computers.
[0029] In some embodiments, the first containerized computer computer-executable instructions are configured to run in parallel at different computers. In some embodiments, each computer further comprises a third containerized computer-executable instructions for transferring the pharmacogenomic dataset to another computer. In some embodiments, each computer further comprises a fourth containerized computer-executable instructions for generating a report of computer activities performed at the computer. In some embodiments, the containerizedcomputer computer-executable instructions are configured to partition the pharmacogenomic dataset and cache the pharmacogenomic dataset across multiple servers.
[0030] In some aspects, the present disclosure provides a method for predicting a patient response to a drug, the method comprising: extracting data from at least one database of correspondence between genetic alleles and drug responses; integrating the data using ML / Al algorithms to provide a set of drug response predictions; iteratively adding and integrating new data to the set of drug response predictions, wherein the new data comprises at least one of a new drug; a new patient; and a new allele; and at least one of calculating a score for the patient response to the drug using the set of drug response predictions, wherein the score corresponds to a prediction to use: a standard dose of the drug; an adjusted dose of the drug; or an alternative drug; calculating a specific therapy with a specific drug for the patient; providing a recommendation to a physician based on a relation between the genetic alleles and drug therapy outcomes; and calculating a new drug therapy by iteratively adding and integrating new data to the set of drug response predictions.
[0031] In some aspects, the present disclosure provides a computer-implemented method for providing pharmacogenetic guidelines, the method comprising: inputting a drug prescription and a patient genotype; and analyzing the drug prescription and the patient genotype to provide a drug therapy recommendation; wherein the analyzing comprises: comparing the patient genotype and drug name to a drug profile database and a molecular biomarker database and calculating a score for a patient response to the drug, wherein the score corresponds to a prediction to use: a standard dose of the drug; an adjusted dose of the drug; or an alternative drug.
[0032] In some embodiments, the analyzing comprises: comparing the patient genotype and the drug name to a drug profile database and a molecular biomarker database iteratively using artificial intelligence and generating a specific drug therapy recommendation for a specific patient. In some embodiments, the analyzing comprises: comparing the patient genotype and the drug name to a drug profile database and a molecular biomarker database iteratively using artificial intelligence and generating a drug recommendation for the drug to a hospital. In some embodiments, the analyzing comprises: comparing the patient genotype and the drug name to a drug profile database, a molecular biomarker database and a hospital database of reactions to the drug iteratively using artificial intelligence and generating a drug recommendation for the drug to the hospital. In some embodiments, the analyzing comprises: comparing the patient genotype and the drug name to a drug profile database, a molecular biomarker database, a hospital database of reactions to the drug and at least one of an omics database iteratively using artificial intelligence and generating a drug recommendation for the drug to a hospital using artificial intelligence.
[0033] In some aspects, the present disclosure provides a system for providing pharmacogenetic guidelines, the system comprising: a computer input system for entering a drug prescription and a patient genotype; and a computer processor for analyzing the drug prescription and the patient genotype to provide a drug therapy recommendation; wherein the analyzing comprises: comparing the patient genotype and a drug name to a drug profile database and a molecular biomarker database and calculating a score for a patient response to the drug, wherein the score corresponds to a prediction to use: a standard dose of the drug; an adjusted dose of the drug; or an alternative drug.
[0034] In some embodiments, the analyzing comprises: comparing the patient genotype and the drug name to a drug profile database and a molecular biomarker database iteratively using artificial intelligence and generating a specific drug therapy recommendation for a specific patient. In some embodiments, the analyzing comprises: comparing the patient genotype and the drug name to a drug profile database and a molecular biomarker database iteratively using artificial intelligence and generating a drug recommendation for the drug to a hospital. In some embodiments, the analyzing comprises: comparing the patient genotype and the drug name to a drug profile database, a molecular biomarker database and a hospital database of reactions to the drug iteratively using artificial intelligence and generating a drug recommendation for the drug to the hospital. In some embodiments, the analyzing comprises: comparing the patient genotype and the drug name to a drug profile database, a molecular biomarker database, a hospital database of reactions to the drug and at least one of an omics database iteratively using artificial intelligence and generating a drug recommendation for the drug to a hospital using artificial intelligence.
[0035] In some aspects, the present disclosure provides a computer program product comprising a computer-readable medium having computer-executable code encoded therein, the computerexecutable code adapted to be executed to implement any one of the computer-implemented methods or computer-implemented systems disclosed herein.
[0036] In some aspects, the present disclosure provides a computer-implemented method comprising: training a machine learning model comprising a neural network language model, using a structured pharmacogenomic dataset, to: classify a feature vector as (1) a serviceable feature vector or (2) a non-serviceable feature vector for a drug response prediction; and generate for the serviceable feature vector, using the neural network language model, the drug response prediction in natural language; updating the structured pharmacogenomic dataset with a new pharmacogenomic dataset to generate an updated structured pharmacogenomic dataset; and retraining the machine learning model, using the updated structured pharmacogenomic dataset, to add new serviceable feature vectors to the machine learning model.
[0037] In some aspects, the present disclosure provides a computer-implemented method, comprising: using a machine learning model comprising a neural network language model to process a feature vector to generate a drug response prediction in natural language, wherein the machine learning model is trained based on a structured dataset comprising a pharmacogenomic dataset relating at least 1,200 pharmacological substances and at least 350 genes.
[0038] In some aspects, the present disclosure provides a computer-implemented method, comprising: training a machine learning model comprising a neural network language model to process a feature vector to generate a drug response prediction in natural language, wherein the training is based on a structured dataset comprising a pharmacogenomic dataset relating at least 1,200 pharmacological substances and at least 350 genes.
[0039] In some aspects, the present disclosure provides a computer-implemented method, comprising: using a machine learning model comprising a neural network language model to: classify a feature vector as (1) a serviceable feature vector or (2) a non-serviceable feature vector for a drug response prediction; and generate a drug response prediction in natural language; wherein the machine learning model is trained based on a structured dataset comprising a pharmacogenomic dataset and at least one of: (i) an epigenomic dataset, (ii) a transcriptomic dataset, (iii) a proteomic dataset, (iv) a metabolomic dataset, (v) a lipidomic dataset, and (vi) a secretomic dataset.
[0040] In some aspects, the present disclosure provides a computer-implemented method, comprising: training a machine learning model comprising a neural network language model to: classify a feature vector as (1) a serviceable feature vector or (2) a non-serviceable feature vector for a drug response prediction; and generate a drug response prediction in natural language; wherein the training is based on a structured dataset a pharmacogenomic dataset and at least one of: (i) an epigenomic dataset, (ii) a transcriptomic dataset, (iii) a proteomic dataset, (iv) a metabolomic dataset, (v) a lipidomic dataset, and (vi) a secretomic dataset.
[0041] In some aspects, the present disclosure provides a computer-implemented method, comprising: creating a secure connection with a client to receive an external pharmacogenomic dataset encrypted based on an external cipher; receiving the external pharmacogenomic dataset through the secure connection; decrypting the external pharmacogenomic dataset based on the external cipher; adding the external pharmacogenomic dataset to the internal pharmacogenomic dataset; decrypting the internal pharmacogenomic dataset based on an internal cipher; and training the neural network language model, using the internal pharmacogenomic dataset, to generate a drug response prediction in natural language.
[0042] In some aspects, the present disclosure provides a computer-implemented method, comprising: training a neural network language model using a plurality of computers, whereineach computer comprises: a database comprising a pharmacogenomic dataset, wherein the pharmacogenomic dataset of each computer in the plurality of computers is different from one another; and a containerized computer-executable instructions configured to train the neural network language model using the computer’s computational resources and the computer’s pharmacogenomic dataset, without sharing the computer’s pharmacogenomic dataset with another computer in the plurality of computers.
[0043] In some aspects, the present disclosure provides a non-transitory computer-readable storage media encoded with a computer program including instructions executable by one or more processors to implement any one of the computer-implemented methods or computer- implemented systems disclosed herein.
[0044] In some aspects, the present disclosure provides a computer-implemented system comprising: (a) a digital processing device comprising: (b) at least one processor, an operating system configured to perform executable instructions, a memory, and a computer program including instructions executable by the digital processing device to perform any one of the computer-implemented methods or computer-implemented systems disclosed herein.INCORPORATION BY REFERENCE
[0045] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The novel features of the disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings of which:
[0047] FIG. 1 is a flow chart of an embodiment of the data entry and analysis of the present disclosure.
[0048] FIG. 2 is a confusion matrix visualization for pharmacogenetic classification performance.
[0049] FIG. 3 shows composite ROC curves demonstrating multiclass classification performance.
[0050] FIG. 4 is a Precision-Recall curve for the developed model.
[0051] FIG. 5 is a visualization of the importance of each variable in predicting the target variable, wherein variables with higher values are more important and contribute more to the model’s decision-making process.
[0052] FIG. 6 shows an example architecture of a system, in accordance with some embodiments.
[0053] FIG. 7 shows performance comparison of gradient boosting models - Fl -Score, Precision, and Recall Metrics.
[0054] FIG. 8 shows a frequency distribution of recommendation lengths (in Words).
[0055] FIG. 9 shows a frequency distribution of recommendation lengths (in Words) within the Recommendation Class (serviceable class).
[0056] FIG. 10 shows a feature importance plot of the CatBoost Model.
[0057] FIG. 11 shows receiver operating characteristic (ROC) Curve of the CatBoost Model
[0058] FIG. 12 shows evaluating model performance - classification and regression metrics.
[0059] FIG. 13 shows evaluation of BLEU scores across different models.
[0060] FIG. 14 shows evaluation of ROUGE-1 scores across different models.
[0061] FIG. 15 shows a detailed comparison of models based on ROUGE scores.
[0062] FIG. 16 shows a workflow for analyzing a sample with the Al-powered model.
[0063] FIG. 17 shows a workflow for coordinating various data streams to train a neural network language model.
[0064] FIG. 18 shows a computer system, in accordance with some embodiments.
[0065] FIG. 19 shows integration of epigenomic, transcriptomic, proteomic, and metabolomics data into a large language model.
[0066] FIG. 20 shows a data pipeline for training a pharmacogenomic recommendation model and for fine-tuning of the model. Training Data Processing and Model Architecture. The diagram illustrates the data pipeline for training a pharmacogenomic recommendation model. Data is collected from sources like PharmGKB and PubChem, processed into structured data vectors, and fed into a machine learning framework. The model comprises a fine-tuned LLM (leveraging LoRa and a modified tokenizer) and a gradient boosting model. These components work together to generate pharmacogenomic recommendations based on both textual and structured data.
[0067] FIG. 21 shows a data pipeline an inference pipeline for generating pharmacogenomic recommendations. Preprocessed data vectors, collected from sources like PharmGKB andPubChem, can be fed into a machine learning model. The model can comprise a fine-tuned LLM (utilizing LoRa and a modified tokenizer) and a gradient boosting model. These components can process the input data and generate a final pharmacogenomic recommendation based on learned patterns and structured knowledge.
[0068] FIG. 22 shows a representative diagram showing iterative recommendation refinement. The diagram illustrates how a model can evolve a base recommendation iteratively based on clinical outcomes. At each step, previous recommendations and / or outcomes can be fed back into the system, refining the next recommendation. This transformation chain can ensure that pharmacogenomic recommendations dynamically adapt to real-world clinical data, improving accuracy over multiple iterations. The model can be further fine-tuned on a diverse range of medical data, including clinical guidelines and clinical outcomes. During each iteration, the model can be provided with the initial recommendation, the clinical outcome, and supplementary information regarding patients with similar clinical profiles. This supplementary information can include modifications made to the patients’ treatment plans along with the corresponding treatment results. The approach can allow the model to generate the most accurate and effective treatment modification tailored to each individual case. By incorporating data from analogous clinical scenarios and their outcomes, the model can continuously refine its recommendations, ensuring that the proposed treatment adjustments are both contextually relevant and optimized for efficacy. This iterative process can enhance the model's ability to adapt to complex clinical situations, ultimately improving patient outcomes through personalized and evidence-based treatment strategies.
[0069] FIG. 23 shows a representative workflow incorporating a two-recommendation architecture. The diagram illustrates a workflow where a first model generates the initial (base) recommendation, while a second model focuses on modifying and refining recommendations based on clinical outcomes. The process can start with user-provided data vectors, which are processed by the first model to generate a preliminary recommendation. Clinical outcomes can then be incorporated into the second model, which adapts the recommendation dynamically, and improve accuracy and relevance.
[0070] FIG. 24 shows an illustrative data flow incorporating an upload interface (UI) a pharmacogenomic recommendation model. This diagram illustrates the sequential processing of data across different modules in the system. The UI can interact with the UI server, which can manage communication between the front-end and back-end. Data can be passed to the Data Loader, which can process and forward the data to the ML Analyzer. After generating a recommendation, the system can process the data further through an additional ML Analyzer,refining the output before sending the final answer back to the UI. This structured flow can improve efficiency of data handling and optimized recommendation generation.
[0071] FIG. 25 shows an example of a graphical user interface, where a pharmacogenomic recommendation is generated based on a selected drug (atorvastatin) and genotype (SLCO1B1 *1 / *1).
[0072] FIG. 26 shows another GUI element. The system can generate pharmacogenomic recommendations based on input genetic data. The interface can accept user-defined molecular and genetic information as input and produce a pharmacogenomic recommendation as output.
[0073] FIG. 27 shows a process for pharmacogenomic data selection and transformation. The user can select a drug and a gene in the upload interface (UI), which is sent to the UI server. The database can retrieve relevant molecular and genetic data, which can then be transformed and forwarded to the ML server for further analysis. The model can be fine-tuned to process input in the form of numerical vectors. This fine-tuning can significantly simplify the data ingestion process for end-users, particularly in scenarios where clinical data, which is inherently complex and challenging to translate into textual descriptions, is involved. By adapting the model to handle data in this format, the model can be enhanced for the specific tasks it is designed to address, improving efficiency and accuracy in processing and analyzing structured numerical data. This approach can streamline the user experience by eliminating the need for cumbersome data transformation steps as well as optimize the model's performance by aligning its architecture with the inherent characteristics of the data it is intended to process.
[0074] FIG. 28 shows a multi-omics data upload interface. A UI for uploading and selecting various types of biological data within the pharmacogenomic analysis system is illustrated. The interface can include a dropdown menu with multiple data input options, which can permit users to provide genetic and other biological information for model-driven pharmacogenomic recommendations. The lower section of the UI displays input fields for selecting specific genes and alleles, enabling customized data entry for analysis. A download button is also present, which can permit users to retrieve processed results or reports. This interface is designed to facilitate seamless integration of multi-omics data into the pharmacogenomic model pipeline, leading to comprehensive and accurate drug response predictions based on diverse biological inputs.
[0075] FIG. 29 shows a low level data upload interface. A low-level interface can allow the editing of any properties of a substance and the genes associated with it, thereby unlocking opportunities for the prediction of new pharmaceuticals or the advancement of research endeavors. This interface can provide a granular level of control, allowing users to manipulate molecular structures, biochemical properties, and genetic sequences with precision. Byfacilitating direct interaction with the fundamental attributes of substances, the interface can serve as a powerful tool for computational drug discovery, enabling researchers to model and predict the efficacy, toxicity, and interaction mechanisms of novel compounds.
[0076] FIG. 30 shows a confusion matrix comparing LLM Score predictions with expert evaluations.DETAILED DESCRIPTION
[0077] The following description and drawings are illustrative of the disclosure and are not to be construed as limiting the disclosure. Numerous specific details are described to provide a thorough understanding of various embodiments of the present disclosure. However, in certain instances, well-known or conventional details are not described in order to provide a concise discussion of embodiments of the present disclosure.
[0078] Integrating pharmacogenetics (PGx) into clinical practice can lead to significant benefits in personalized medicine. Personalized drug therapy plans can be provided to patients in a way that enhances the efficacy of the drug while minimizing adverse drug reactions (ADRs). ADRs, in particular, are a global health concern as hospital systems shoulder billions in annual expenses due to medication errors, patient readmissions arising from ADRs, and extended inpatient stays resulting from inappropriate or suboptimal prescriptions. With the advent of artificial intelligence based drug discovery, many new drugs are expected to hit the market more quickly than before. Costs of adverse drug events (ADEs) are expected to rise as a result.
[0079] However, providing personalized medicine plans and reducing ADRs is a complex problem. Increasingly diverse and comorbid patients make one-size-fits-all prescription approaches ineffective and inaccurate, which leads to complications and extended hospital stays. But personalized medicine is demanded to take into account the enormous biological diversity, even only at the genomic level, to tailor treatments for patients. The complex interplay of genetic factors influencing drug response and the labor-intensive process of developing pharmacogenetic guidelines hinders the timely application of personalized medicine.
[0080] In some aspects, the present application provides an artificial intelligence (Al) system for pharmacogenomics, and more broadly, pharmaco-omics. The system can use real-time machine learning and generative Al to match, for example, patient genetic profiles with the safest, most effective medications. The system can be a cloud-based platform that can integrate with existing clinical record-keeping and data management systems. The integration can allow the system to provide recommendations to at point-of-care. The integrated system can harness large-scale multi-omics data and real -world evidence (e.g., clinical evidence) to provide improvedrecommendations to future subjects. Over time, the system can improve the efficacy and efficiency of recommendations, which can reduce ADEs and readmissions.
[0081] In some aspects, the present disclosure provides systems and methods to accelerate the generation of PGx recommendations. Acceleration can be achieved by using a machine learning pipeline that integrates drug molecular structures, physical and chemical parameters of drugs, PKPD genetic profiles, allele combinations, and any combination thereof. Utilizing advanced artificial intelligence (Al) models, the predictive accuracy can be enhanced while providing high-quality, contextually relevant recommendations, which can facilitate the integration of pharmacogenetics into clinical practice.
[0082] The systems and methods of the present disclosure can be used to effectively accelerate the development of pharmacogenetic recommendations, while providing high accuracy and performance. Additional data layers such as clinical outcomes and multi-omics information can be incorporated, Recommendation updates can be automatically integrated. By facilitating the integration of personalized medicine into clinical practice, this approach can be used to reduce ADRs, improve therapeutic efficacy, and enhance patient outcomes.
[0083] In some aspects, the present disclosure provides a computer-implemented system comprising a database comprising a structured pharmacogenomic dataset. The computer- implemented system can comprise a machine learning model comprising a neural network language model. The computer-implemented system can comprise computer-executable instructions. The computer-executable instructions can be configured to train the machine learning model classify a feature vector. The training can be using the structured pharmacogenomic dataset. The feature vector can be classified as a serviceable feature vector or a non-serviceable feature vector for a prediction. The computer-executable instructions can be configured to generate for the serviceable feature vector, using the neural network language model, the prediction in natural language. The prediction can be a drug response prediction. The computer-executable instructions can be configured to update the structured pharmacogenomic dataset with a new pharmacogenomic dataset to generate an updated structured pharmacogenomic dataset. The computer-executable instructions can be configured to train a new machine learning model, using the updated structured pharmacogenomic dataset, such that the new machine learning model has more serviceable feature vectors than the machine learning model. The computer-executable instructions can be configured to retrain the machine learning model, using the updated structured pharmacogenomic dataset, to add new serviceable feature vectors to the machine learning model.
[0084] In some aspects, the present disclosure provides a computer-implemented system comprising a machine learning model comprising a neural network language model, wherein themachine learning model is trained using a structured dataset comprising a pharmacogenomic dataset relating at least 1,200 pharmacological substances and at least 350 genes. The computerexecutable instructions can be configured to process a feature vector, using the machine learning model, to generate a drug response prediction in natural language.
[0085] In some aspects, the present disclosure provides a computer-implemented system comprising a large language model (LLM), wherein the LLMs is trained using a structured dataset comprising a pharmacogenomic dataset relating at least 1,200 pharmacological substances and at least 350 genes. The computer-executable instructions can be configured to process a feature vector, using the large learning model, to generate a drug response prediction in natural language. In some embodiments, the large language model generates or is configured to generate a natural langue explanation of drug response predictions. In some embodiments, the large language model can be integrated into a system, wherein the system collects, updates, and optimizes clinical information. In some embodiments, the large language model can be in operable communication with a system, wherein the system collects, updates, and optimizes clinical information. In some embodiments, the large language model can be in electronic communication with a system, wherein the system collects, updates, and optimizes clinical information. In some embodiments, the integration with, operable communication with, or electronic communication with the system results in continuous updating and refining of pharmacogenomic and genetic data, thereby resulting in a drug response prediction with high clinical relevance and precision.
[0086] In some aspects, the present disclosure provides a pharmacogenomics system or a model comprising a two-stage recommendation architecture, configured to generate a drug response prediction. In some embodiments, the two-stage recommendation architecture comprises a first model that generates a pharmacogenomic recommendation and a second model that refines and adapts the pharmacogenomic recommendation based on factors including clinical outcomes, omics data, and real-world patient data. In some embodiments, the pharmacogenomics system, or model incorporates adaptive learning mechanisms, wherein the incorporation results in continuous updating and refinement based on factors including new clinical outcomes, such as hospital-specific adjustments.
[0087] In some aspects, the present disclosure provides a pharmacogenomics system that seamlessly integrates with hospital workflows, leading to personalized medication decisions. The system can harness patient genetic data and real-world outcomes to optimize therapy choices, significantly reducing adverse drug events and lowering healthcare costs.
[0088] The system can be used to provide personalized treatment, for example, by identifying high-risk medication-gene interactions before prescriptions are made. The system can be used toreduce readmissions of subjects, for example, by providing timely intervention and predicting preventable complications. The system can be used to improve patient safety & satisfaction, for example, by minimizing ADEs to elevate overall patient experience and outcomes.
[0089] The system can be used lower operational costs. For example, the system can be used optimize spending on drugs by preventing inefficient medication usage. The system can be used to reduce duration of hospital stays, for example, by reducing complications that translate to faster discharges, which increases bed turnover and revenue. The system can be used reduce legal and compliance risk, for example, by mitigating costly malpractice claims and penalty fees related to avoidable errors. The system can be used to reduce adverse drug events, for example, by proactively flagging high-risk medications prescribed to a subject given the subject’s genetics. The system can proper medication selection and dosing of a drug for a subject.
[0090] The system can be used streamline clinical workflows. The system can be used be integrated into electronic health record systems. The system can be used deliver actionable recommendations at the point of care, decreasing physician workload. The system can be used provide a scalable Al system, for example, by rolling out updates incrementally within hospital departments or across entire networks with minimal disruption.
[0091] In some embodiments, the system uses a machine learning system which recognizes when to and when not to generate a recommendation (e.g., serviceable versus non-serviceable). In some embodiments, the machine learning algorithm can determine when a plausible and safe recommendation can be generated, prior to generating the recommendation. When a machine learning system determines that an input is “serviceable”, the machine learning system can be said to have determined that it is “actionable” to generate a recommendation. In some cases, “serviceable” and “actionable” may be used interchangeably herein. Therefore, the system provides an advantage that it can provide useful outputs even while the system’s machine learning system is scaling over time to larger datasets (e.g., more drugs, more genes). This can make it easier for users to organize and contribute to dataset, because it does not require pooling a heroic amount of resources to obtain a usable model. The model can accumulate data organically over time as doctors, hospitals, and clinics use the model, even at earlier stages of model development where datasets are smaller.
[0092] In some aspects, the present disclosure provides a continuously learning pharmacogenomics Al system. The machine learning models can evolve with new data, ensuring that recommendations and predictions generated by the machine learning models stay current with emerging research, guidelines, and real world clinical outcomes. The system can incorporate large-scale multi -omics datasets and real clinical outcomes to provide robust insights to users. The system can be adapted to comply with HIPAA, GDPR, and other regulatoryguidelines which are different between jurisdictions and which change over time. The system can be used to pilot the Al system in a single ward, or provide a system-wide rollout across multiple facilities. The system can use Al to process unstructured data to curate datasets for training the Al to make better predictions.Drug Response Prediction
[0093] Various aspects of the present disclosure may be used to generate predictions about a subject’s response to an administration of a pharmaceutical substance. In some embodiments, the prediction is a drug response prediction. In some embodiments, the drug response prediction comprises a personalized regimen for administering a pharmaceutical substance for a subject. In some embodiments, the drug response prediction comprises a personalized dosing guideline for a subject. In some embodiments, the drug response prediction comprises a personalized drug selection for a subject. In some embodiments, the drug response prediction comprises a prediction of an adverse drug reaction in a subject. In some embodiments, a label of the personalized regimen, the personalized dosing guideline, the personalized drug selection, or any combination thereof can be printed. In some embodiments, the label can be applied to a packaging for a pharmaceutical substance. In some embodiments, multiple drug response predictions are generated to statistically filter hallucinated drug response predictions. In some embodiments, the natural language is English, Spanish, German, French, Russian, Mandarin Chinese, Cantonese Chinese, French, Portuguese, Hindi, Korean, or Japanese.
[0094] The pharmaceutical substance can be for treating an infection, e.g., viral infection, e.g., HIV infection, an immune disorder, e.g., autoimmune disorder, an inflammatory disease, e.g., a chronic inflammatory disorder, an infectious disease, a hematological disease, a proliferative disease, e.g., a cancer, a solid tumor, or a liquid tumor, a degenerative disease, e.g., a neurodegenerative disease, a pulmonary disease, a renal disease, a gastrointestinal disease, an endocrine disease, a dermatological disease, a genetic disease, or any combination thereof.
[0095] The pharmaceutical substance can be a small molecule, an antibody, a vaccine, a protein, a biologic, a biosimilar, a pharmaceutical composition, a pharmaceutical formulation, a diet, or any combination thereof. The pharmaceutical substance can be a biological, pharmaceutical, or chemical compound. Non-limiting examples of a pharmaceutical substance include a simple or complex organic or inorganic molecule, a peptide, a protein, an oligonucleotide, an epigenetic modulator, hormones (steroidal or peptide), fusion molecules, an antibody, an antibody derivative, antibody fragment, a vitamin derivative, a carbohydrate, a toxin, a vaccine, e.g., cancer vaccine, a chemotherapeutic compound, radiotherapies (y-rays, X-rays, and / or the directed delivery of radioisotopes, microwaves, and UV radiation), gene therapies (e.g.,antisense, retroviral therapy) and other immunotherapies. Additional non-limiting examples of a pharmaceutical substance include small molecule inhibitors, monoclonal antibodies (mAbs), sdAbs, chimeric antigen receptors (CARs), CAR T-cell therapy, and antibody-drug conjugates (ADCs), and bispecific antibodies. In some embodiments, a pharmaceutical substance is a biologic. Non-limiting examples of biologies include vaccines, blood, and blood components, allergenics, somatic cells, gene therapy, tissues, and recombinant therapeutic proteins.
[0096] The drug response prediction can be a prediction for treating an infection, e.g., viral infection, e.g., HIV infection, an immune disorder, e.g., autoimmune disorder, an inflammatory disease, e.g., a chronic inflammatory disorder, an infectious disease, a hematological disease, a proliferative disease, e.g., a cancer, a solid tumor, or a liquid tumor, a degenerative disease, e.g., a neurodeg enerative disease, a pulmonary disease, a renal disease, a gastrointestinal disease, an endocrine disease, a dermatological disease, a genetic disease, or any combination thereof.
[0097] In some embodiments, a drug response prediction may be reprocessed when a data gap or data inconsistency is detected. This can increase the validity and accuracy of final outputs. In some embodiments, a drug response prediction may be flagged and / or withheld for human review. For example, a high-risk drug response prediction, such as an extreme side-effect likelihood that may cause severe harm to a subject, can be reviewed by a human before final report generation. In some embodiments, a drug response prediction may be based on omics data and / or clinical data. In some embodiments, a drug response prediction may be based on omics data and / or clinical data when omic data and / or clinical data is detected to be available.Generated drug response predictions can be stored in a database.
[0098] Non-limiting of applications for a drug response prediction as disclosed herein include supporting clinical decisions, guiding and assisting drug development, and assisting in preventative healthcare applications. A drug response prediction provided herein can be automated to provide automated decisions. A drug response prediction provided herein can be integrated into electronic health records. A drug response prediction can be in operable communication with electronic health records. A drug response prediction can be in electronic communication with electronic health records. A drug response prediction provided herein can provide clinicians with immediate pharmacogenomic recommendations at the point of care. A drug response prediction provided herein can provide personalized treatment decisions. Nonlimiting applications for which a drug response prediction provided herein can provide personalized treatment decisions include: oncology, autoimmune disease, obesity and metabolic disorders, psychiatry, neurology, cariology, endocrinology, anesthesiology, dermatology, pediatrics, immunology, urology, obstetrics and gynecology, nephrology, pulmonology, hematology, infectious disease, hepatology, or any combination thereof.
[0099] A drug response prediction provided herein can leverage multiple data sources, including but not limited to pharmacogenomic and transcriptomic data, to guide biotechnological research and drug development. A drug response prediction provided herein can identity novel drug-gene interactions. Identification of novel-gene drug interactions can allow for repurposing of medications in rare disease applications. A drug response model provided herein can integrate biomarkers from data sources, including but not limited to clinical trial enrollment criteria, to develop companion diagnostics for applications including the matching of patients to precision medicine trials.
[0100] A drug response prediction provided herein can be integrated into electronic health records. A drug response prediction provided herein can be in operable communication with electronic health records. A drug response prediction provided herein can be in electronic communication with electronic health records. A drug response prediction provided herein can be utilized in telemedicine platforms. A drug response prediction provided herein can be integrated into a wearable-device for continuous health monitoring. A drug response model provided here can be integrated into health insurance models for risk stratification and predictive modeling in applications including but not limited to policy pricing and preventive healthcare programs.
[0101] In some embodiments, a drug response prediction can be deployed in various configurations to balance performance, privacy, and scalability. Non-limiting examples of drug response prediction deployment include: on-premise deployment for institutions requiring strict data control and compliance with regulatory standards; cloud-based deployment, leveraging distributed processing for high-performance inference; edge computing implementation, enabling real-time processing on local devices without reliance on external serves; and adaptive inference, allowing for real-time adjustments in computational complexity based on system constraints.
[0102] By providing personalized dosing guidelines and drug selections based on a subject’s genetic profile, drug response prediction can reduce the incidence of ADRs, improve therapeutic efficacy, and contribute to better patient outcomes.Machine Learning
[0103] In some embodiments, the structured pharmacogenomic dataset comprises structural representations of pharmaceutical substances. In some embodiments, the structural representations comprise: molecular structures, fingerprints, latent vectors, SMILES (Simplified Molecular Input Line Entry System), InChi (International Chemical Identifier), SDF (Structure Data File), topological descriptors, 3D conformations, graph adjacency matrices, node and edgefeature vectors for each atom and bond, autoencoder-derived embeddings, transformer-based embeddings for small molecules, graph neural network latent vectors, quantum chemical descriptors (partial charges, HOMO-LUMO levels), protein-ligand interaction features (docking scores, binding affinity), pharmacophore or toxicophore features, or any combination thereof.
[0104] In some embodiments, the structured pharmacogenomic dataset comprises physical or chemical properties of pharmaceutical substances. In some embodiments, the physical or chemical properties comprise: pharmacokinetic properties, pharmacodynamic properties, molecular weights, logP, logD, pKa, solubility, melting point, boiling point, hydrogen bond donor count, hydrogen bond acceptor count, topological polar surface area (tPSA), rotatable bond count, refractivity, or any combination thereof.
[0105] In some embodiments, the structured pharmacogenomic dataset comprises genetic profiles of subjects. In some embodiments, the genetic profiles comprise: genomic sequences, mutations, SNPs, copy number variations (CNVs), insertions, deletions, haplotypes, splicing variants, epigenetic modifications, gene expression levels, genomic rearrangements, or any combination thereof. In some embodiments, the feature vector comprises a structural representation of a pharmaceutical substance. In some embodiments, the feature vector comprises a physical or chemical property of a pharmaceutical substance.
[0106] For example, a feature vector may comprise the following prompt structure:System context and feature labels :Molecular structure of the drug : ### Fingerprint {med_f ingerprint } Physicochemical characteristics of the drug : ### Chem { chem_inf o } Pharmacodynamic genetic characteristic of the drug : ### Gen_PD { gen_info_pd } Pharmacokinetic genetic characteristic of the drug : ### Gen_PK { gen_inf o_pk } Allele variants : ### Allele { allels } Pharmacoepigenomics of the drug : ### Methylation {meth } Pharmacotranscriptomics : ### RNA { rna } Pharmacoproteomics : ### Proteins { protein } Pharmacometabolomics : ### Metabolites {met } Clinical Outcomes : ### Outcomes { clin_weights }
[0107] In some embodiments, (i) the serviceable feature vector, (ii) the non-serviceable feature vector, or (iii) both is input to the neural network language model as a structured prompt. In some embodiments, the structured prompt comprises: a feature definition, a feature label, a feature value, a prefix, a suffix, a question, a chain of thought, or any combination thereof. In some embodiments, the feature definition comprises context for a feature in natural language. In some embodiments, the feature label comprises a database identifier for a feature. In some embodiments, the feature value comprises a quantitative value, a qualitative value, In some embodiments, the structured prompt comprises a plurality of features. In some embodiments, the plurality of features comprises: a molecular structure of a pharmaceutical substance, a molecular fingerprint of a pharmaceutical substance, a physicochemical characteristic of a pharmaceutical substance, a pharmacodynamic genetic characteristic of a pharmaceutical substance, a pharmacokinetic genetic characteristic of a pharmaceutical substance, pharmacoepigenomics of a pharmaceutical substance, pharmacotranscriptomics of a pharmaceutical substance, pharmacoproteomics of a pharmaceutical substance, pharmacometabolomics of a pharmaceutical substance, a clinical outcome of a pharmaceutical substance, allele variants, or any combination thereof.
[0108] In some embodiments, the prompt can be engineered to optimize a response. Nonlimiting examples of engineering techniques to optimize a prompt include: dynamically generated system prompts that are tailored to specific pharmacogenomic queries that are based on user data and domain-specific constraints; incorporation of relevant reference materials, patient history, and retrieval of real-time knowledge for augmentation with domain-specific context; and adaptive prompt formatting that adjusts to the model in use to optimize response quality and computational efficiency.
[0109] In some embodiments, the serviceable feature vector indicates that a clinically effective recommendation of the drug response prediction can be generated for the feature vector. In some embodiments, the non-serviceable feature vector indicates that a clinically effective recommendation of the drug response prediction cannot be generated for the feature vector. In some embodiments, the serviceable feature vector defines a genomic feature of a subject that is not found in the structured pharmacogenomic dataset. In some embodiments, the serviceable feature vector defines a combination of genomic features of the subject that is not found in the structured pharmacogenomic dataset. In some embodiments, the serviceable feature vector defines a pharmacological substance that is not found in the structured pharmacogenomic dataset. In some embodiments, the serviceable feature vector defines a combination of pharmacological substances that is not found in the structured pharmacogenomic dataset. In some embodiments, the serviceable feature vector defines a combination of (i) a pharmacologicalsubstance and (ii) genomic features of the subject that is not found in the structured pharmacogenomic dataset. In some embodiments, the serviceable feature vector defines a combination of (i) pharmacological substances and (ii) genomic features of the subject that is not found in the structured pharmacogenomic dataset.
[0110] In some embodiments, the feature vector is classified as the serviceable feature vector or the non-serviceable feature vector based on a confidence score. In some embodiments, the confidence score is determined by processing the feature vector to quantify a similarity between (i) a pharmaceutical substance defined by the feature vector and (ii) a different pharmaceutical substance defined by the structured pharmacogenomic dataset. In some embodiments, the confidence score is determined by processing the feature vector to quantify (i) a pharmacokinetic property defined by the structured pharmacogenomic dataset, (ii) a pharmacodynamic property defined by the structured pharmacogenomic dataset, (iii) a predicted pharmacokinetic property, (iv) a predicted pharmacodynamic property, or any combination thereof. In some embodiments, the confidence score is determined by an association between a pharmaceutical substance and genomic features of the subject. In some embodiments, the confidence score is determined by processing an association between a pharmaceutical substance and (i) epigenetic features, (ii) proteomic features, (iii) transcriptomic features, (iv) metabolomic features of the subject, or any combination thereof.[oni] In some embodiments, the confidence score is determined by processing an availability and relevance of clinical outcome data for a pharmacological substance defined by the feature vector. In some embodiments, the feature vector is classified as the serviceable feature vector when the confidence score exceeds a predetermined threshold. In some embodiments, the predetermined threshold is 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, 99.91%, 99.92%, 99.93%, 99.94%, 99.95%, 99.96%, 99.97%, 99.98%, 99.99%, 99.991%, 99.992%, 99.993%, 99.994%, 99.995%, 99.996%, 99.997%, 99.998%, 99.999% confidence that the feature vector is a serviceable feature vector. In some embodiments, the predetermined threshold is 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, 99.91%, 99.92%, 99.93%, 99.94%, 99.95%, 99.96%, 99.97%, 99.98%, 99.99%, 99.991%, 99.992%, 99.993%, 99.994%, 99.995%, 99.996%, 99.997%, 99.998%, 99.999% confidence that the feature vector is a non-serviceable feature vector.
[0112] Various inputs may be provided to the machine learning model in the feature vector. The feature vector can synthesize or combine a variety of information about the pharmaceutical substance, the genetics of a subject, etc. An input representing or defining a pharmaceutical substance may be provided. The pharmaceutical substance can be a protein. The pharmaceuticalsubstance can be a small molecule, configured to bind to a protein as a ligand. The input of the protein can comprise: a primary structure of the protein, a secondary structure of the protein, a tertiary structure of the protein, a quaternary structure of the protein, a physical property of the protein, a binding site of the protein, a relative accessible surface area of the protein, or any combination thereof. The input can be embedding in a vector, e.g., using a tokenization scheme or a neural network. An input of a ligand can comprise a chemical structure of the ligand. The chemical structure can be a human-readable representation of the ligand, e.g., SMILES. The chemical structure can be an atomistic structure or a coarse-grained structure of the ligand.
[0113] In some cases, a “representation”, a “descriptor”, a “latent vector”, and the like, may refer to a description of a molecular entity. For example, commonly used descriptors include SMILES, primary sequences, or nucleic acid sequences. In some cases, a representation may be used as features or feature values to a machine learning model. In some cases, a representation be features or feature values from a machine learning model.
[0114] In some cases, the features or feature values may comprise an identifier of an atom, a functional group, an amino acid, a secondary structure, a tertiary structure, or a motif. In some cases, the features or feature values may comprise an electronic configuration, a charge, a size, a bond angle, a dihedral angle, or any combination thereof. In some cases, the motif may comprise a collection of atoms. In some cases, the features or feature values may comprise a geometric descriptor of an atom, an amino acid, a secondary structure, a tertiary structure, or a motif. In some cases, the features or feature values may comprise a physical property.
[0115] Some examples of secondary structures include alpha helices and beta sheets of proteins, where each can be formed when hydrogen bonding between residues of a protein is stabilized. Some examples of a tertiary structure can include larger geometrical features within a protein such as pockets, hairpins, concatenations, loop regions, globular regions, etc. The secondary structure of proteins may include the hydrogen bonding networks of subsections in a primary sequence. Alpha helices and beta sheets can be discerned, for example, within a Ramachandran plot of a protein due to the arrangement of the backbone amide groups hydrogen bonding. Tertiary structure may be seen as more interaction types contribute to the conformational landscape of a protein. Tertiary structures may depend on non-polar / hydrophobic van der Waal interactions, electrostatic interactions, and other various interactions described herein.
[0116] In some cases, a representation is in M0L2, PDB, MOL, PDBQ / PDBQT, SDF, CIF, CML, XML, ASN1, PARM, CRD, or TRI. In some cases, a representation comprise SMILES, SELFIES, or InChi.
[0117] In some cases, a representation comprises one or more electron configurations. In some cases, an electron configuration may comprise one or more atomic orbitals, one or moremolecular orbitals, or both. In some cases, an electron configuration may comprise valence electrons of an atom. In some cases, an electron configuration may comprise a character of an electron (e.g., s, p, d, f, and any mixtures thereof). In some cases, an electron configuration may comprise an electron spin. In some cases, an electron configuration may comprise electron density. In some cases, an electron configuration may be represented in various basis functions, including but not limited to, atomic orbitals, molecular orbitals, or plane waves.
[0118] Within various chemoinformatic formats can be differently encoded information. In some cases, an atomistic representation may comprise the relative cartesian coordinates of atoms to each other. In some cases, an atomistic representation may comprise the relative cartesian coordinates of atoms to an arbitrary point. In some cases, an atomistic representation may comprise thermodynamic estimations of values such as solvation energy, potential energy of bond lengths, bond angles, dihedral angles, 1-4 intramolecular interaction energies, intramolecular energies among adjacent bond angles, hydrogen bonding energies, and nonbonded interaction energies. In some cases, an atomistic representation may comprise atom type definitions and generalizations. In some cases, an atomistic representation may comprise polarizability parameters. In some cases, an atomistic representation may comprise Lennard- Jones van der Waal parameters. In some cases, an atomistic representation may comprise electrostatic charge parameters. In some cases, an atomistic representation may comprise bond length, bond angle, and dihedral force constants. In some cases, an atomistic representation may comprise bond length, bond angle, and dihedral equilibrium values. In some cases, an atomistic representation may comprise dihedral phase and periodicity force constants.
[0119] A graph, graph model, and graphical model can refer to a method of conceptualizing or organizing information into a graphical representation comprising nodes and edges. In some cases, a graph can refer to the principle of conceptualizing or organizing data, wherein the data may be stored in a various and alternative forms such as linked lists, dictionaries, spreadsheets, arrays, in permanent storage, in transient storage, and so on, and is not limited to specific cases disclosed herein. In some cases, the machine learning model can comprise a graph model.
[0120] In some embodiments, the new pharmacogenomic dataset comprises: raw genetic data, clinical outcome data, scientific papers, guidelines, electronic health records, comorbidities, environmental factors, concomitant medications, adverse event reports, real-world evidence from wearable devices, retrospective and prospective cohort data, structured references from curated databases (e.g., ClinVar, PharmGKB), laboratory test results, phenotype data, demographic profiles, patient compliance data, toxicology studies, and multi-center consortia datasets, or any combination thereof. In some embodiments, the computer-executable instructions are configured to extract features useful for drug response prediction in the new pharmacogenomic dataset. Insome embodiments, the computer-executable instructions are configured to filter mutations in the new pharmacogenomic dataset. In some embodiments, the computer-executable instructions are configured to isolate clinically significant markers, from the new pharmacogenomic dataset, which are relevant to adverse drug reactions or efficacy. In some embodiments, the computerexecutable instructions are configured to filter pharmaceutical substances in the new pharmacogenomic dataset. In some embodiments, the computer-executable instructions are configured to, in the new pharmacogenomic dataset, parse unstructured data (e.g., scientific papers, guidelines) into structured formats. In some embodiments, the computer-executable instructions are configured to, in the new pharmacogenomic dataset, annotate data with standardized ontologies (e.g., gene and drug nomenclatures). In some embodiments, the computer-executable instructions are configured to, in the new pharmacogenomic dataset, remove duplicates and resolve inconsistent entries. In some embodiments, the computerexecutable instructions are configured to apply dimensionality reduction or feature engineering techniques. In some embodiments, the computer-executable instructions are configured to integrate the new pharmacogenomic dataset with the structured pharmacogenomic dataset. In some embodiments, the computer-executable instructions are configured to, using the updated structured pharmacogenomic dataset, generate updated feature vectors for training the new machine learning model.
[0121] A feature vector can comprise information from a pharmacogenomic dataset. The feature vector can be structured information from the pharmacogenomic dataset. In some embodiments, a pharmacogenomic dataset can indicate an annotation for an interaction between a gene and a pharmaceutical substance. For example, the annotation may state: “The FDA label for abrocitinib gives dosing modifications for CYP2C19 poor metabolizers. It does not describe which CYP2C19 variants or tests can be used to define poor metabolizers.” The annotation can indicate a suggested prescription for a pharmaceutical substance. For example, the annotation may state: “Recommended Dosage in CYP2C19 Poor Metabolizers In patients who are known or suspected to be CYP2C19 poor metabolizers, the recommended dosage of CIBINQO is 50 mg once daily ... If an adequate response is not achieved with CIBINQO 50 mg orally daily after 12 weeks, consider increasing dosage to 100 mg orally once daily. Discontinue therapy if inadequate response is seen after dosage increase to 100 mg once daily.” The annotation can indicate a scientific reason, e.g., a mechanism, which supports the prescription. For example, the annotation may state: “CYP2C19 Poor Metabolizers In patients who are CYP2C19 poor metabolizers, the AUC of abrocitinib is increased compared to CYP2C19 normal metabolizers due to reduced metabolic clearance. Dosage reduction of CIBINQO is recommended in patients who are known or suspected to be CYP2C19 poor metabolizers based on genotype or previoushistory / experience with other CYP2C19 substrates.” The annotation may indicate a lack of interaction, or a lack of knowledge about an interaction between a gene and a pharmaceutical substance. For example, the annotation may state: “The EMA product information for abrocitinib states that CYP2C19 or CYP2C9 genotype does not have a clinically significant effect on abrocitinib exposure. There is clinical dosing information about CYP2C19 drug-drug interactions.” The annotation may indicate a type of subject that a particular prescription may be recommended for. For example, the annotation may state: “KRAZATI is indicated for the treatment of adult patients with KRAS G12C-mutated locally advanced or metastatic non-small cell lung cancer (NSCLC), as determined by an FDA-approved test, who have received at least one prior systemic therapy.” The annotation may indicate factors that may be considered by a physician in generating a prescription for a subject. For example, the annotation may state: “Amphetamine is known to inhibit monoamine oxidase, whereas the ability of amphetamine and its metabolites to inhibit various P450 isozymes and other enzymes has not been adequately elucidated. In vitro experiments with human microsomes indicate minor inhibition of CYP2D6 by amphetamine and minor inhibition of CYP1 A2, 2D6, and 3 A4 by one or more metabolites. However, due to the probability of autoinhibition and the lack of information on the concentration of these metabolites relative to in vivo concentrations, no predications regarding the potential for amphetamine or its metabolites to inhibit the metabolism of other drugs by CYP isozymes in vivo can be made.” The annotation may indicate that a pharmaceutical substance is known to trigger severe adverse drug response in some subjects. The annotation may indicate caution about prescribing a pharmaceutical substance. The annotation may indicate all pharmaceutical substances that are known to interact with a gene. The annotation may indicate all genes that are known to interact with a pharmaceutical substance. The genetic information in the annotation may indicate a particular mutation or regulatory elements that are pharmacogenomically relevant. The pharmacogenomic dataset can comprise a knowledge graph.
[0122] In some embodiments, the machine learning model comprises a binary classifier. In some embodiments, the binary classifier comprises a random forest model. In some embodiments, the binary classifier comprises a boosted tree algorithm. In some embodiments, the neural network language model comprises an autoregressive model. In some embodiments, the neural network language model comprises a transformer. In some embodiments, the neural network language model comprises a large language model.
[0123] A neural network language model can be used to create an electronic health record based on a transcription of an interaction between a subject and a medical professional. A neural network language model can be used to summarize, analyze, or extract salient information from medical literature, including textbooks, databases, scientific articles, electronic health records, orany combination thereof. A neural network language model can be used to analyze medical images and provide diagnostic or interpretive commentary in natural language. A neural network language model can be used to stratify or classify a disease of a subject into more subcategories, wherein the subcategories can be treated with different pharmaceutical substances and / or administration regimens. A neural network language model can be trained with federated learning.
[0124] The neural network language model can have various number of parameters. The neural network language model can have at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, or 900 million parameters. The neural network language model can have at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, or 900 million parameters. The neural network language model can have at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, or 900 billion parameters. The neural network language model can have at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, or 900 billion parameters.
[0125] Predictions by the neural network language model can have various BLEU scores. A BLEU score of the neural network language model can be at least 0.8, 0.85, 0.9, 0.95, 0.96, 0.97, 0.98, or 0.99. A BLEU score of the neural network language model can be at most 0.8, 0.85, 0.9, 0.95, 0.96, 0.97, 0.98, or 0.99. Predictions by the neural network language model can have various ROUGE- 1 scores. A ROUGE- 1 score of the neural network language model can be at least 0.8, 0.85, 0.9, 0.95, 0.96, 0.97, 0.98, or 0.99. A ROUGE-1 score of the neural network language model can be at most 0.8, 0.85, 0.9, 0.95, 0.96, 0.97, 0.98, or 0.99.
[0126] The machine learning model can have various Fl -scores for distinguishing serviceable feature vectors versus non-serviceable feature vectors. The machine learning model can have a Fl-score of at least 0.8, 0.85, 0.9, 0.95, 0.96, 0.97, 0.98, or 0.99. The machine learning model can have a Fl-score of at most 0.8, 0.85, 0.9, 0.95, 0.96, 0.97, 0.98, or 0.99. The machine learning model can have various precision scores for distinguishing serviceable feature vectors versus non-serviceable feature vectors. The machine learning model can have a precision score of at least 0.8, 0.85, 0.9, 0.95, 0.96, 0.97, 0.98, or 0.99. The machine learning model can have a precision score of at most 0.8, 0.85, 0.9, 0.95, 0.96, 0.97, 0.98, 0.99, or 1.0000.
[0127] Various machine learning models may be used. In some cases, the machine learning model comprises a language model, a graph model, a flow model, a generative adversarial network, a variational autoencoder, an autoregressive model, an autoencoder, a diffusion model, or any combination thereof. The machine learning model can comprise a neural network comprising various architectures, loss functions, optimization algorithms, priors, and variousother neural network design choices. In some cases, the machine learning model can comprise a neural network. In some cases, the machine learning model can comprise an autoencoder. In some cases, the machine learning model can comprise a generative model. In some cases, the machine learning model can comprise a variational autoencoder. In some cases, the machine learning model can comprise a generative adversarial network. In some cases, the machine learning model can comprise a flow model. In some cases, the machine learning model can comprise an autoregressive model. In some cases, the machine learning model can comprise a diffusion model. In some cases, the machine learning model can comprise a neural network with one or more layers. In some cases, the machine learning model can comprise a neural network with one or more fully connected layers. In some cases, the machine learning model can comprise a neural network with one or more convolutional layers. In some cases, the machine learning model can comprise a neural network with one or more message-passing layers. In some cases, the machine learning model can comprise a neural network with a bottleneck layer. In some cases, a layer may comprise an attention mechanism, a generalized message-passing graph neural network, or both. In some cases, a generalized message-passing graph neural network comprises a graph convolutional neural network.
[0128] In some cases, the machine learning model can comprise a neural network with residual blocks. In some cases, the machine learning model can comprise a neural network with attention. In some cases, the machine learning model can comprise a neural network with one or more nonlinearities. In some cases, the machine learning model can comprise a neural network with one or more dropout layers. In some cases, the machine learning model can comprise a neural network with one or more batch normalization layers. In some cases, the machine learning model can comprise a regression loss function. In some cases, the machine learning model can comprise a logistic loss function. In some cases, the machine learning model can comprise a variational loss. In some cases, the machine learning model can comprise a prior. In some cases, the machine learning model can comprise a Gaussian prior. In some cases, the machine learning model can comprise a non-Gaussian prior. In some cases, the machine learning model can comprise an adversarial loss. In some cases, the machine learning model can comprise a reconstruction loss. In some cases, the machine learning model is trained with the Adam optimizer. In some cases, the machine learning model is trained with the stochastic gradient descent optimizer. In some cases, the model learning model hyperparameters are optimized with Gaussian Processes. In some cases, the machine learning model is trained with train / validation / test data splits. In some cases, the machine learning model is trained with k-fold data splits, with any positive integer for k.
[0129] The neural network language model can be a light-weight model. The neural network language model can be trained using knowledge distillation. The neural network language model can be subjected to model pruning. The neural network language model can be 2-bit, or 3 -bit quantized. The neural network language model can be fine-tuned.
[0130] The machine learning model can comprise a variety of manifold learning algorithms. In some cases, the machine learning model can comprise a manifold learning algorithm. In some cases, the manifold learning algorithm comprises principal component analysis. In some cases, the manifold learning algorithm comprises a uniform manifold approximation algorithm. In some cases, the manifold learning algorithm comprises an isomap algorithm. In some cases, the manifold learning algorithm comprises a locally linear embedding algorithm. In some cases, the manifold learning algorithm comprises a modified locally linear embedding algorithm. In some cases, the manifold learning algorithm comprises a Hessian eigen mapping algorithm. In some cases, the manifold learning algorithm comprises a spectral embedding algorithm. In some cases, the manifold learning algorithm comprises a local tangent space alignment algorithm. In some cases, the manifold learning algorithm comprises a multi-dimensional scaling algorithm. In some cases, the manifold learning algorithm comprises a t-distributed stochastic neighbor embedding algorithm (t-SNE). In some cases, the manifold learning algorithm comprises a Barnes-Hut t- SNE algorithm.
[0131] In some cases, the methods of the disclosure further comprise reducing one or more representations using a machine learning model. The terms “reducing”, “dimensionality reduction”, “projection”, “component analysis”, “feature space reduction”, “latent space engineering”, “feature space engineering”, “representation engineering”, or “latent space embedding”, as used herein, generally refer to a method of transforming a given input data with an initial number of dimensions to another form of data that has fewer dimensions than the initial number of dimensions. In some cases, the terms can refer to the principle of reducing a set of input dimensions to a smaller set of output dimensions. In some cases, the terms can refer to the principle of reducing a set of input dimensions to a set of output dimensions of a same or larger size.
[0132] The term “normalizing”, as used herein, generally refers to a collection of methods for adjusting a dataset to align the dataset to a common scale. In some cases, a normalizing method can comprise multiplying a portion or the entirety of a dataset by a factor. In some cases, a normalizing method can comprise adding or subtracting a constant from a portion or the entirety of a dataset. In some cases, a normalizing method can comprise adjusting a portion or the entirety of a dataset to a known statistical distribution. In some cases, a normalizing method can comprise adjusting a portion or the entirety of a dataset to a normal distribution. In some cases, anormalizing method can comprise adjusting the dataset so that the signal strength of a portion or the entirety of a dataset is about the same.
[0133] Converting can comprise one or more steps of various conversions of data. In some cases, converting can comprise normalizing data. In some cases, converting can comprise performing a mathematical operation that computes a score based on a distance between 2 points in the data. In some cases, the distance can comprise a distance between two edges in a graph. In some cases, the distance can comprise a distance between two nodes in a graph. In some cases, the distance can comprise a distance between a node and an edge in a graph. In some cases, the distance can comprise a Euclidean distance. In some cases, the distance can comprise a non- Euclidean distance. In some cases, the distance can be computed in a frequency space. In some cases, the distance can be computed in Fourier space. In some cases, the distance can be computed in Laplacian space. In some cases, the distance can be computed in spectral space. In some cases, the mathematical operation can be a monotonic function based on the distance. In some cases, the mathematical operation can be a non-monotonic function based on the distance. In some cases, the mathematical operation can be an exponential decay function. In some cases, the mathematical operation can be a learned function.
[0134] In some cases, converting can comprise transforming data in one representation to another representation. In some cases, converting can comprise transforming data into another form of data with less dimensions. In some cases, converting can comprise linearizing one or more curved paths in the data. In some cases, converting can be performed on data comprising data in Euclidean space. In some cases, converting can be performed on data comprising data in graph space. In some cases, converting can be performed on data in a discrete space. In some cases, converting can be performed on data comprising data in frequency space. In some cases, converting can transform data in discrete space to continuous space, continuous space to discrete space, graph space to continuous space, continuous space to graph space, graph space to discrete space, discrete space to graph space, or any combination thereof. In some cases, converting can comprise transforming data in discrete space into a frequency domain. In some cases, converting can comprise transforming data in continuous space into a frequency domain. In some cases, converting can comprise transforming data in graph space into a frequency domain.
[0135] In some cases, reducing can comprise transforming a given input data with any initial number of dimensions to another form of data that has any number of dimensions fewer than the initial number of dimensions. In some cases, reducing can comprise transforming input data into another form of data with fewer dimensions. In some cases, reducing can comprise linearizing one or more curved paths in the input data to the output data. In some cases, reducing can be performed on data comprising data in Euclidean space. In some cases, reducing can beperformed on data comprising data in graph space. In some cases, reducing can be performed on data in a discrete space. In some cases, reducing can transform data in discrete space to continuous space, continuous space to discrete space, graph space to continuous space, continuous space to graph space, graph space to discrete space, discrete space to graph space, or any combination thereof.
[0136] In some embodiments, the drug response prediction can be filtered or post-processed following generation of the recommendation. Non-limiting filtering and post-processing techniques include: rule-based validation layers to verify that the recommendation complies with established guidelines; uncertainty estimation mechanisms, wherein cases where model confidence is low or additional information is required are flagged; dynamic ranking of recommendations, wherein responses based on empirical evidence, known genetic associations, and best-practice guidelines are prioritized; integration of feedback loops allowing for iterative refinement of recommendations based on clinical input and new data; and any combination thereof.
[0137] The terms “clustering”, “cluster analysis”, or “generating modules”, as used herein, generally refer to a method of grouping samples in a dataset by some measure of similarity. Samples can be grouped in a set space, for example, element ‘a’ is in set ‘A’. Samples can be grouped in a continuous space, for example, element ‘a’ is a point in Euclidean space with distance T away from the centroid of elements comprising cluster ‘A’. Samples can be grouped in a graph space, for example, element ‘a’ is highly connected to elements comprising cluster ‘A’. These terms can refer to the principle of organizing a plurality of elements into groups in some mathematical space based on some measure of similarity.Computing System
[0138] In some aspects, the present disclosure describes a computer-implemented system comprising: a digital processing device comprising: at least one processor, an operating system configured to perform executable instructions, a memory, and a computer program including instructions executable by the digital processing device to train a machine learning model, use a machine learning model to classify a feature vector, use a machine learning model to generate a drug response prediction, update a dataset in a database, encrypt a dataset in a database, decrypt a dataset in a database, make a secure connection with a client, provide access privileges to a computer, a database, or a machine learning model, or any combination thereof. In some aspects, the present disclosure describes a computer-implemented method, implementing any one of the methods disclosed herein in a computer system. Referring to FIG. 18, a block diagram is shown depicting an exemplary machine that includes a computer system 1800 (e.g., a processing orcomputing system) within which a set of instructions can execute for causing a device to perform or execute any one or more of the aspects and / or methodologies for training a machine learning model, using a machine learning model to classify a feature vector, using a machine learning model to generate a drug response prediction, updating a dataset in a database, encrypting a dataset in a database, decrypting a dataset in a database, making a secure connection with a client, providing access privileges to a computer, a database, or a machine learning model, or any combination thereof. The components in FIG. 18 are examples only and do not limit the scope of use or functionality of any hardware, software, embedded logic component, or a combination of two or more such components implementing particular embodiments.
[0139] Computer system 1800 may include one or more processors 1801, a memory 1803, and a storage 1808 that communicate with each other, and with other components, via a bus 1840. The bus 1840 may also link a display 1832, one or more input devices 1833 (which may, for example, include a keypad, a keyboard, a mouse, a stylus, etc.), one or more output devices 1834, one or more storage devices 1835, and various tangible storage media 1836. All of these elements may interface directly or via one or more interfaces or adaptors to the bus 1840. For instance, the various tangible storage media 1836 can interface with the bus 1840 via storage medium interface 1826. Computer system 1800 may have any suitable physical form, including but not limited to one or more integrated circuits (ICs), printed circuit boards (PCBs), mobile handheld devices (such as mobile telephones or PDAs), laptop or notebook computers, distributed computer systems, computing grids, or servers.
[0140] Computer system 1800 includes one or more processor(s) 1801 (e.g., central processing units (CPUs), general purpose graphics processing units (GPGPUs), or quantum processing units (QPUs)) that carry out functions. Computer system 1800 may be one of various high performance computing platforms. For instance, the one or more processor(s) 1801 may form a high performance computing cluster. In some embodiments, the one or more processors 1801 may form a distributed computing system connected by wired and / or wireless networks. In some embodiments, arrays of CPUs, GPUs, QPUs, or any combination thereof may be operably linked to implement any one of the methods disclosed herein. Processor(s) 1801 optionally contains a cache memory unit 1802 for temporary local storage of instructions, data, or computer addresses. Processor(s) 1801 are configured to assist in execution of computer readable instructions. Computer system 1800 may provide functionality for the components depicted in FIG. 18 as a result of the processor(s) 1801 executing non-transitory, processor-executable instructions embodied in one or more tangible computer-readable storage media, such as memory 1803, storage 1808, storage devices 1835, and / or storage medium 1836. The computer-readable media may store software that implements particular embodiments, and processor(s) 1801 may executethe software. Memory 1803 may read the software from one or more other computer-readable media (such as mass storage device(s) 1835, 1836) or from one or more other sources through a suitable interface, such as network interface 1820. The software may cause processor(s) 1801 to carry out one or more processes or one or more steps of one or more processes described or illustrated herein. Carrying out such processes or steps may include defining data structures stored in memory 1803 and modifying the data structures as directed by the software.
[0141] The memory 1803 may include various components (e.g., machine readable media) including, but not limited to, a random access memory component (e.g., RAM 1804) (e.g., static RAM (SRAM), dynamic RAM (DRAM), ferroelectric random access memory (FRAM), phasechange random access memory (PRAM), etc.), a read-only memory component (e.g., ROM 1805), and any combinations thereof. ROM 1805 may act to communicate data and instructions unidirectionally to processor(s) 1801, and RAM 1804 may act to communicate data and instructions bidirectionally with processor(s) 1801. ROM 1805 and RAM 1804 may include any suitable tangible computer-readable media described below. In one example, a basic input / output system 1806 (BIOS), including basic routines that help to transfer information between elements within computer system 1800, such as during start-up, may be stored in the memory 1803.
[0142] Fixed storage 1808 is connected bidirectionally to processor(s) 1801, optionally through storage control unit 1807. Fixed storage 1808 provides additional data storage capacity and may also include any suitable tangible computer-readable media described herein. Storage 1808 may be used to store operating system 1809, executable(s) 1810, data 1811, applications 1812 (application programs), and the like. Storage 1808 can also include an optical disk drive, a solid- state memory device (e.g., flash-based systems), or a combination of any of the above. Information in storage 1808 may, in appropriate cases, be incorporated as virtual memory in memory 1803.
[0143] In one example, storage device(s) 1835 may be removably interfaced with computer system 1800 (e.g., via an external port connector (not shown)) via a storage device interface 1825. Particularly, storage device(s) 1835 and an associated machine-readable medium may provide non-volatile and / or volatile storage of machine-readable instructions, data structures, program modules, and / or other data for the computer system 1800. In one example, software may reside, completely or partially, within a machine-readable medium on storage device(s) 1835. In another example, software may reside, completely or partially, within processor(s) 1801
[0144] Bus 1840 connects a wide variety of subsystems. Herein, reference to a bus may encompass one or more digital signal lines serving a common function, where appropriate. Bus 1840 may be any of several types of bus structures including, but not limited to, a memory bus, amemory controller, a peripheral bus, a local bus, and any combinations thereof, using any of a variety of bus architectures. As an example, and not by way of limitation, such architectures include an Industry Standard Architecture (ISA) bus, an Enhanced ISA (EISA) bus, a Micro Channel Architecture (MCA) bus, a Video Electronics Standards Association local bus (VLB), a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, an Accelerated Graphics Port (AGP) bus, HyperTransport (HTX) bus, serial advanced technology attachment (SATA) bus, and any combinations thereof.
[0145] Computer system 1800 may also include an input device 1833. In one example, a user of computer system 1800 may enter commands and / or other information into computer system 1800 via input device(s) 1833. Examples of an input device(s) 1833 include, but are not limited to, an alpha-numeric input device (e.g., a keyboard), a pointing device (e.g., a mouse or touchpad), a touchpad, a touch screen, a multi-touch screen, a joystick, a stylus, a gamepad, an audio input device (e.g., a microphone, a voice response system, etc.), an optical scanner, a video or still image capture device (e.g., a camera), and any combinations thereof. In some embodiments, the input device is a Kinect, Leap Motion, or the like. Input device(s) 1833 may be interfaced to bus 1840 via any of a variety of input interfaces 1823 (e.g., input interface 1823) including, but not limited to, serial, parallel, game port, USB, FIREWIRE, THUNDERBOLT, or any combination of the above. In some embodiments, an input device 1833 may be used to receive a command to train a machine learning model, use a machine learning model to classify a feature vector, use a machine learning model to generate a drug response prediction, update a dataset in a database, encrypt a dataset in a database, decrypt a dataset in a database, make a secure connection with a client, provide access privileges to a computer, a database, or a machine learning model, or any combination thereof. In some embodiments, training a machine learning model, using a machine learning model to classify a feature vector, using a machine learning model to generate a drug response prediction, updating a dataset in a database, encrypting a dataset in a database, decrypting a dataset in a database, making a secure connection with a client, providing access privileges to a computer, a database, or a machine learning model, or any combination thereof is triggered using human inputs through an input device 1833.
[0146] In particular embodiments, when computer system 1800 is connected to network 1830, computer system 1800 may communicate with other devices, specifically mobile devices and enterprise systems, distributed computing systems, cloud storage systems, cloud computing systems, and the like, connected to network 1830. Communications to and from computer system 1800 may be sent through network interface 1820. For example, network interface 1820 may receive incoming communications (such as requests or responses from other devices) in the form of one or more packets (such as Internet Protocol (IP) packets) from network 1830, andcomputer system 1800 may store the incoming communications in memory 1803 for processing. Computer system 1800 may similarly store outgoing communications (such as requests or responses to other devices) in the form of one or more packets in memory 1803 and communicated to network 1830 from network interface 1820. Processor(s) 1801 may access these communication packets stored in memory 1803 for processing.
[0147] Examples of the network interface 1820 include, but are not limited to, a network interface card, a modem, and any combination thereof. Examples of a network 1830 or network segment 1830 include, but are not limited to, a distributed computing system, a cloud computing system, a wide area network (WAN) (e.g., the Internet, an enterprise network), a local area network (LAN) (e.g., a network associated with an office, a building, a campus or other relatively small geographic space), a telephone network, a direct connection between two computing devices, a peer-to-peer network, and any combinations thereof. A network, such as network 1830, may employ a wired and / or a wireless mode of communication. In general, any network topology may be used.
[0148] Information and data can be displayed through a display 1832. Examples of a display 1832 include, but are not limited to, a cathode ray tube (CRT), a liquid crystal display (LCD), a thin film transistor liquid crystal display (TFT-LCD), an organic liquid crystal display (OLED) such as a passive-matrix OLED (PMOLED) or active-matrix OLED (AMOLED) display, a plasma display, and any combinations thereof. The display 1832 can interface to the processor(s) 1801, memory 1803, and fixed storage 1808, as well as other devices, such as input device(s) 1833, via the bus 1840. The display 1832 is linked to the bus 1840 via a video interface 1822, and transport of data between the display 1832 and the bus 1840 can be controlled via the graphics control 1821. In some embodiments, the display is a video projector. In some embodiments, the display is a head-mounted display (HMD) such as a VR headset. In further embodiments, suitable VR headsets include, by way of non-limiting examples, HTC Vive, Oculus Rift, Samsung Gear VR, Microsoft HoloLens, Razer OSVR, FOVE VR, Zeiss VR One, Avegant Glyph, Freefly VR headset, and the like. In still further embodiments, the display is a combination of devices such as those disclosed herein.
[0149] In addition to a display 1832, computer system 1800 may include one or more other peripheral output devices 1834 including, but not limited to, an audio speaker, a printer, a storage device, and any combinations thereof. Such peripheral output devices may be connected to the bus 1840 via an output interface 1824. Examples of an output interface 1824 include, but are not limited to, a serial port, a parallel connection, a USB port, a FIREWIRE port, a THUNDERBOLT port, and any combinations thereof.
[0150] In addition, or as an alternative, computer system 1800 may provide functionality as a result of logic hardwired or otherwise embodied in a circuit, which may operate in place of or together with software to execute one or more processes or one or more steps of one or more processes described or illustrated herein. Reference to software in this disclosure may encompass logic, and reference to logic may encompass software. Moreover, reference to a computer- readable medium may encompass a circuit (such as an IC) storing software for execution, a circuit embodying logic for execution, or both, where appropriate. The present disclosure encompasses any suitable combination of hardware, software, or both.
[0151] Those of skill in the art will appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality.
[0152] The various illustrative logical blocks, modules, and circuits described in connection with the embodiments disclosed herein may be implemented or performed with a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0153] The steps of a method or algorithm described in connection with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by one or more processor(s), or in a combination of the two. A software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a user terminal.
[0154] In accordance with the description herein, suitable computing devices include, by way of non-limiting examples, server computers, desktop computers, laptop computers, notebook computers, sub-notebook computers, netbook computers, netpad computers, set-top computers, media streaming devices, handheld computers, Internet appliances, mobile smartphones, and tablet computers.
[0155] In some embodiments, the computing device includes an operating system configured to perform executable instructions. The operating system is, for example, software, including programs and data, which manages the device’s hardware and provides services for execution of applications. Those of skill in the art will recognize that suitable server operating systems include, by way of non -limiting examples, FreeBSD, OpenBSD, NetBSD®, Linux, Apple® Mac OS X Server®, Oracle® Solaris®, Windows Server®, and Novell® NetWare®. Those of skill in the art will recognize that suitable personal computer operating systems include, by way of nonlimiting examples, Microsoft® Windows®, Apple® Mac OS X®, UNIX®, and UNIX-like operating systems such as GNU / Linux®. In some embodiments, the operating system is provided by cloud computing. Those of skill in the art will also recognize that suitable mobile smartphone operating systems include, by way of non-limiting examples, Nokia® Symbian® OS, Apple® los®, Research In Motion® BlackBerry OS®, Google® Android®, Microsoft® Windows Phone® OS, Microsoft® Windows Mobile® OS, Linux®, and Palm® WebOS®.
[0156] In some embodiments, a computer system 1800 may be accessible through a user terminal to receive user commands. The user commands may include line commands, scripts, programs, etc., and various instructions executable by the computer system 1800. A computer system 1800 may receive instructions to train a machine learning model, use a machine learning model to classify a feature vector, use a machine learning model to generate a drug response prediction, update a dataset in a database, encrypt a dataset in a database, decrypt a dataset in a database, make a secure connection with a client, provide access privileges to a computer, a database, or a machine learning model, or any combination thereof, or schedule a computing job for the computer system 1800 to carry out any instructions.Non-Transitory Computer Readable Storage Medium
[0157] In some aspects, the present disclosure describes a non-transitory computer-readable storage media encoded with a computer program including instructions executable by one or more processors to train a machine learning model, use a machine learning model to classify a feature vector, use a machine learning model to generate a drug response prediction, update a dataset in a database, encrypt a dataset in a database, decrypt a dataset in a database, make a secure connection with a client, provide access privileges to a computer, a database, or amachine learning model, or any combination thereof using any one of the methods disclosed herein. In some embodiments, a non-transitory computer-readable storage media may comprise instructions for training a machine learning model, using a machine learning model to classify a feature vector, using a machine learning model to generate a drug response prediction, updating a dataset in a database, encrypting a dataset in a database, decrypting a dataset in a database, making a secure connection with a client, providing access privileges to a computer, a database, or a machine learning model, or any combination thereof. In some embodiments, the platforms, systems, media, and methods disclosed herein include one or more non-transitory computer readable storage media encoded with a program including instructions executable by the operating system of an optionally networked computing device.
[0158] In further embodiments, a computer readable storage medium is a tangible component of a computing device. In still further embodiments, a computer readable storage medium is optionally removable from a computing device. In some embodiments, a computer readable storage medium includes, by way of non-limiting examples, flash memory devices, solid state memory, magnetic disk drives, magnetic tape drives, optical disk drives, distributed computing systems including cloud computing systems and services, and the like. In some embodiments, the program and instructions are permanently, substantially permanently, semi-permanently, or non-transitorily encoded on the media.Computer Program
[0159] In some aspects, the present disclosure describes a computer program product comprising a computer-readable medium having computer-executable code encoded therein, the computerexecutable code adapted to be executed to implement any one of the methods disclosed herein. In some embodiments, the platforms, systems, media, and methods disclosed herein include at least one computer program, or use of the same.
[0160] A computer program includes a sequence of instructions, executable by one or more processor(s) of the computing device’s CPU, written to perform a specified task. Computer readable instructions may be implemented as program modules, such as functions, objects, Application Programming Interfaces (APIs), computing data structures, and the like, that perform particular tasks or implement particular abstract data types. In light of the disclosure provided herein, those of skill in the art will recognize that a computer program may be written in various versions of various languages. In some embodiments, APIs may comprise various languages, for example, languages in various releases of TensorFlow, Theano, Keras, PyTorch, or any combination thereof which may be implemented in various releases of Python, Python3, C, C#, C++, MatLab, R, Java, or any combination thereof.
[0161] The functionality of the computer readable instructions may be combined or distributed as desired in various environments. In some embodiments, a computer program comprises one sequence of instructions. In some embodiments, a computer program comprises a plurality of sequences of instructions. In some embodiments, a computer program is provided from one location. In other embodiments, a computer program is provided from a plurality of locations. In various embodiments, a computer program includes one or more software modules. In various embodiments, a computer program includes, in part or in whole, one or more web applications, one or more standalone applications, one or more web browser plug-ins, extensions, add-ins, or add-ons, or combinations thereof.Web Application
[0162] In some embodiments, a computer program includes a web application. In some embodiments, a user may enter a query for training a machine learning model, using a machine learning model to classify a feature vector, using a machine learning model to generate a drug response prediction, updating a dataset in a database, encrypting a dataset in a database, decrypting a dataset in a database, making a secure connection with a client, providing access privileges to a computer, a database, or a machine learning model, or any combination thereof through a web application. In some embodiments, a user may command a computer to train a machine learning model, use a machine learning model to classify a feature vector, use a machine learning model to generate a drug response prediction, update a dataset in a database, encrypt a dataset in a database, decrypt a dataset in a database, make a secure connection with a client, provide access privileges to a computer, a database, or a machine learning model, or any combination thereof through a web application. In light of the disclosure provided herein, those of skill in the art will recognize that a web application, in various embodiments, utilizes one or more software frameworks and one or more database systems. In some embodiments, a web application is created upon a software framework such as Microsoft® .NET or Ruby on Rails (RoR). In some embodiments, a web application utilizes one or more database systems including, by way of non -limiting examples, relational, non-relational, object oriented, associative, XML, and document oriented database systems. In further embodiments, suitable relational database systems include, by way of non-limiting examples, Microsoft® SQL Server, mySQL™, and Oracle®. Those of skill in the art will also recognize that a web application, in various embodiments, is written in one or more versions of one or more languages. A web application may be written in one or more markup languages, presentation definition languages, client-side scripting languages, server-side coding languages, database query languages, or combinations thereof. In some embodiments, a web application is written to some extent in amarkup language such as Hypertext Markup Language (HTML), Extensible Hypertext Markup Language (XHTML), or extensible Markup Language (XML). In some embodiments, a web application is written to some extent in a presentation definition language such as Cascading Style Sheets (CSS). In some embodiments, a web application is written to some extent in a client-side scripting language such as Asynchronous JavaScript and XML (AJAX), Flash® ActionScript, JavaScript, or Silverlight®. In some embodiments, a web application is written to some extent in a server-side coding language such as Active Server Pages (ASP), ColdFusion®, Perl, Java™, JavaServer Pages (JSP), Hypertext Preprocessor (PHP), Python™, Ruby, Tel, Smalltalk, WebDNA®, or Groovy. In some embodiments, a web application is written to some extent in a database query language such as Structured Query Language (SQL). In some embodiments, a web application integrates enterprise server products such as IBM® Lotus Domino®.Mobile application
[0163] In some embodiments, a computer program includes a mobile application provided to a mobile computing device. In some embodiments, the mobile application is provided to a mobile computing device at the time it is manufactured. In other embodiments, the mobile application is provided to a mobile computing device via the computer network described herein.
[0164] In view of the disclosure provided herein, a mobile application is created by techniques known to those of skill in the art using hardware, languages, and development environments known to the art. Those of skill in the art will recognize that mobile applications are written in several languages. Suitable programming languages include, by way of non-limiting examples, C, C++, C#, Objective-C, Java™, JavaScript, Pascal, Object Pascal, Python™, Ruby, VB.NET, WML, and XHTML / HTML with or without CSS, or combinations thereof.
[0165] Suitable mobile application development environments are available from several sources. Commercially available development environments include, by way of non-limiting examples, AirplaySDK, alcheMo, Appcelerator®, Celsius, Bedrock, Flash Lite, .NET Compact Framework, Rhomobile, and WorkLight Mobile Platform. Other development environments are available without cost including, by way of non-limiting examples, Lazarus, MobiFlex, MoSync, and Phonegap. Also, mobile device manufacturers distribute software developer kits including, by way of non-limiting examples, iPhone and iPad (los) SDK, Android™ SDK, BlackBerry® SDK, BREW SDK, Palm® OS SDK, Symbian SDK, webOS SDK, and Windows® Mobile SDK.Standalone application
[0166] In some embodiments, a computer program includes a standalone application, which is a program that is run as an independent computer process, not an add-on to an existing process, e.g., not a plug-in. Those of skill in the art will recognize that standalone applications are often compiled. A compiler is a computer program(s) that transforms source code written in a programming language into binary object code such as assembly language or machine code. Suitable compiled programming languages include, by way of non-limiting examples, C, C++, Objective-C, COBOL, Delphi, Eiffel, Java™, Lisp, Python™, Visual Basic, and VB .NET, or combinations thereof. Compilation is often performed, at least in part, to create an executable program. In some embodiments, a computer program includes one or more executable complied applications.Software Modules
[0167] In some embodiments, the platforms, systems, media, and methods disclosed herein include software, server, and / or database modules, or use of the same. In view of the disclosure provided herein, software modules are created by techniques known to those of skill in the art using machines, software, and languages known to the art. The software modules disclosed herein are implemented in a multitude of ways. In various embodiments, a software module comprises a file, a section of code, a programming object, a programming structure, a distributed computing resource, a cloud computing resource, or combinations thereof. In further various embodiments, a software module comprises a plurality of files, a plurality of sections of code, a plurality of programming objects, a plurality of programming structures, a plurality of distributed computing resources, a plurality of cloud computing resources, or combinations thereof. In various embodiments, the one or more software modules comprise, by way of non-limiting examples, a web application, a mobile application, a standalone application, and a distributed or cloud computing application. In some embodiments, software modules are in one computer program or application. In other embodiments, software modules are in more than one computer program or application. In some embodiments, software modules are hosted on one machine. In other embodiments, software modules are hosted on more than one machine. In further embodiments, software modules are hosted on a distributed computing platform such as a cloud computing platform. In some embodiments, software modules are hosted on one or more machines in one location. In other embodiments, software modules are hosted on one or more machines in more than one location.Data Security
[0168] In some aspects, the present disclosure provides a computer-implemented system. In some embodiments, the computer-implemented system comprises a database comprising an internal pharmacogenomic dataset encrypted based on an internal cipher. In some embodiments, the computer-implemented system comprises a neural network language model. In some embodiments, the computer-implemented system comprises computer-executable instructions. In some embodiments, the computer-executable instructions are configured to create a secure connection with a client to receive an external pharmacogenomic dataset encrypted based on an external cipher. In some embodiments, the computer-executable instructions are configured to receive the external pharmacogenomic dataset through the secure connection. In some embodiments, the computer-executable instructions are configured to decrypt the external pharmacogenomic dataset based on the external cipher. In some embodiments, the computerexecutable instructions are configured to add the external pharmacogenomic dataset to the internal pharmacogenomic dataset. In some embodiments, the computer-executable instructions are configured to decrypt the internal pharmacogenomic dataset based on the internal cipher. In some embodiments, the computer-executable instructions are configured to train the neural network language model, using the internal pharmacogenomic dataset, to generate a drug response prediction in natural language.
[0169] Various encryption standards can be used to encrypt data. In some embodiments, an encryption standard comprises transport layer security (TLS). The TLS can be, e.g., TLS 1.2 or TLS 1.3. In some embodiments, an encryption standard comprises Advanced Encryption Standard (AES). The AES can be, e.g., AES-256. In some embodiments, an encryption standard comprises Rivest-Shamir-Adleman (RSA), OpenPGP, a hash function, message-digest algorithm, SHA-1, SHA-2, hash-based message authentication code (HMAC), Elliptic Curve Digital Signature Algorithm (ECDSA), Digital Signature Algorithm (DSA), X.509, ed25519, secure shell protocol (SSH), IEE Pl 363, Kerberos, RADIUS, CRYPTREC, or any combination thereof.
[0170] In some embodiments, the computer-executable instructions are configured to provide access privileges to a plurality of clients, wherein the access privileges are based on a client profile. In some embodiments, the access privileges permit access to the neural network language model and not the database. In some embodiments, the access privileges permit access to the database and not the neural network language model. In some embodiments, the access privileges permit access to the neural network language model and the database. In some embodiments, the client is an organization. In some embodiments, the client is an individual useraccount. In some embodiments, the access privileges comprise (i) read access, (ii) write access, (iii) execute access, (iv) delete access, (v) full privileged access, or any combination thereof.
[0171] In some embodiments, the client profile comprises jurisdiction. In some embodiments, the jurisdiction comprises a jurisdiction of registration. In some embodiments, the jurisdiction comprises a jurisdiction of access. In some embodiments, the jurisdiction comprises US jurisdiction, European jurisdiction, Chinese jurisdiction, South Korean jurisdiction, Australian jurisdiction, African jurisdiction, Russian jurisdiction, Canadian jurisdiction, Japanese jurisdiction, Indian jurisdiction, Brazilian jurisdiction.
[0172] In some embodiments, the computer-executable instructions are configured to log traffic to and from the computer-implemented system. In some embodiments, the computer-executable instructions are configured to log user access to the computer-implemented system. In some embodiments, the computer-executable instructions are configured to log changes to the computer-implemented system. The log of the computer-implemented system can be monitored. Monitoring can detect and / or prevent intrusions or suspicious computer activity. Monitoring can provide real-time alerts.
[0173] In some aspects, the present disclosure provides a computer-implemented system. In some embodiments, the computer-implemented system comprises a neural network language model. In some embodiments, the computer-implemented system comprises a plurality of computers. In some embodiments, each computer comprises a database comprising a pharmacogenomic dataset. In some embodiments, the pharmacogenomic dataset of each computer in the plurality of computers is different from one another. In some embodiments, each computer comprises a containerized computer-executable instructions. In some embodiments, containerized computer-executable instructions are configured to train the neural network language model using the computer’s computational resources and the computer’s pharmacogenomic dataset. In some embodiments, the training is performed without sharing the computer’s pharmacogenomic dataset with another computer in the plurality of computers.
[0174] In some embodiments, the system further comprises a central computer. In some embodiments, the central computer is configured to aggregate the trained neural network language model from the plurality of computers. In some embodiments, the containerized computer computer-executable instructions are configured to run using different architectures of different computers. In some embodiments, the containerized computer computer-executable instructions are configured to load balance. The load balancing can be across a plurality of graphical processing units (GPUs).
[0175] In some embodiments, the system further comprises a second containerized computerexecutable instructions configured to use the neural network language model to generate a drugresponse prediction in natural language for a user of the computer, without sharing the user’s input to the neural network language model with another computer in the plurality of computers.
[0176] In some embodiments, the first containerized computer computer-executable instructions are configured to run in parallel at different computers. In some embodiments, each computer further comprises a third containerized computer-executable instructions for transferring the pharmacogenomic dataset to another computer. In some embodiments, each computer further comprises a fourth containerized computer-executable instructions for generating a report of computer activities performed at the computer. In some embodiments, the containerized computer computer-executable instructions are configured to partition the pharmacogenomic dataset and cache the pharmacogenomic dataset across multiple servers.
[0177] The computer-implemented system can be configured to receive data deletion requests. The computer-implement system can be configured to delete data in response to data deletion requests.
[0178] The computer-implemented system can comprise containerized services. The containerized services can comprise a containerized Al inference service, a containerized report generation service, or both. The computer-implemented system can be updated by updating a subset of its containers.
[0179] The computer-implemented system can provide access upon authorization. Authorization can comprise multi-factor authentication. Network segmentation and strict authentication protocols can be used to ensure that each service authenticates itself before any data exchange.
[0180] The computer-implemented system can be hosted on a secure environment. The computer-implemented system can be hosted on a cloud computing system. The secure environment can comprise compliance certifications (e.g., HIPAA, GDPR, SOC2) and 24 / 7 monitoring. The secure environment can be, e.g., InterSystems IRIS, GCP, or Azure.
[0181] The computer-implemented system can be updated with a Continuous Integration and Continuous Delivery (CI / CD) pipeline. For example, GitHub Actions can be used to automate build, test, and deploy steps, with strict checks that prevent unverified code from entering production. The CI / CD pipeline can comprise automated static code analysis (SAST) to detect input validation or logic flaws.Databases
[0182] In some embodiments, the platforms, systems, media, and methods disclosed herein include one or more databases, or use of the same. In view of the disclosure provided herein, those of skill in the art will recognize that many databases are suitable for storage and retrieval of information about training a machine learning model, using a machine learning model to classifya feature vector, using a machine learning model to generate a drug response prediction, updating a dataset in a database, encrypting a dataset in a database, decrypting a dataset in a database, making a secure connection with a client, providing access privileges to a computer, a database, or a machine learning model, or any combination thereof. In various embodiments, suitable databases include, by way of non-limiting examples, relational databases, non-relational databases, object oriented databases, object databases, entity -relationship model databases, associative databases, XML databases, document oriented databases, and graph databases. Further non-limiting examples include SQL, PostgreSQL, MySQL, Oracle, DB2, Sybase, and MongoDB. In some embodiments, a database is Internet-based. In further embodiments, a database is web-based. In still further embodiments, a database is cloud computing-based. In a particular embodiment, a database is a distributed database. In other embodiments, a database is based on one or more local computer storage devices.
[0183] In some aspects, the present disclosure provides a computer-implemented system comprising a database comprising a structured dataset. In some embodiments, the structured dataset comprises a pharmacogenomic dataset relating at least 1,200 pharmacological substances and at least 350 genes. In some embodiments, the computer-implemented system comprises a machine learning model comprising a neural network language model. In some embodiments, the computer-implemented system comprises computer-executable instructions configured to train the machine learning model to generate a drug response prediction in natural language, wherein the training is based on the structured dataset.
[0184] In some aspects, the present disclosure provides a computer-implemented system comprising a benchmark dataset. In some embodiments, the benchmark data set comprises a pharmacogenomic dataset relating a drug, a gene known to influence the drug response, a genotype representing a specific allele combination, and a recommendation derived from established clinical guidelines. In some embodiments, the benchmark dataset comprises data related to cases with actionable recommendations, wherein a genotype impacts drug selection or dosage. In some embodiments, the benchmark dataset comprises data related to cases without actionable recommendations, wherein no genetic-based adjustment is required. In some embodiments, the benchmark data set comprises at least 220,000 examples.
[0185] In some aspects, the present disclosure provides a computer-implemented system comprising a machine learning model comprising a neural network language model. In some embodiments, the machine learning model is trained using a structured dataset comprising a pharmacogenomic dataset relating at least 1,200 pharmacological substances and at least 350 genes. In some embodiments, the computer-implemented system comprises computer-executableinstructions configured to process a feature vector, using the machine learning model, to generate a drug response prediction in natural language.
[0186] In some embodiments, the pharmacogenomic dataset comprises at least 1,700, 2,500, 3,500, 5,000, 10,000, or 21,000 pharmacological substances. In some embodiments, the pharmacogenomic dataset comprises at most 2,500, 3,500, 5,000, 10,000, 21,000, 30,000 pharmacological substances. Table 1 provides a non-limiting list of pharmacological substances that can be incorporated into a pharmacogenomic dataset. In some embodiments, the pharmacological substances are one or more pharmacological substances in Table 1. In some embodiments, the pharmacogenomic dataset comprises at least 700, 1,300, 2,400, 4,500, 10,000, or 20,000 genes. In some embodiments, the pharmacogenomic dataset comprises at most 1,300, 2,400, 4,500, 10,000, 20,000, 25,000 genes. Table 2 provides a non-limiting list of genes that can be incorporated into a pharmacogenomic dataset. In some embodiments, the genes are one or more genes in Table 2. Without being bound to a particular theory, it is found that relying on a dataset comprised of at least 1,200 pharmacological substances and at least 350 genes provides significant advantages in generating drug response predictions for combinations of drugs and genes that are unknown in the dataset.
[0187] Table 1. Pharmacological substances:
[0188] Table 2. Pharmacogenomic genes
[0189] In some aspects, the present disclosure provides a computer-implemented system comprising a database comprising a structured dataset. In some embodiments, the structured dataset comprises a pharmacogenomic dataset and at least one of: (i) an epigenomic dataset, (ii) a transcriptomic dataset, (iii) a proteomic dataset, (iv) a metabolomic dataset, (v) a lipidomic dataset, and (vi) a secretomic dataset. In some embodiments, the computer-implemented system comprises a machine learning model comprising a neural network language model. In some embodiments, the computer-implemented system comprises computer-executable instructions configured to train the machine learning model, using the structured dataset. In some embodiments, the machine learning model can be trained to classify a feature vector as (1) a serviceable feature vector or (2) a non-serviceable feature vector for a drug response prediction. In some embodiments, the machine learning model can be trained to generate for the serviceable feature vector, using the neural network language model, the drug response prediction in natural language.
[0190] In some embodiments, the epigenomic dataset is a pharmacoepigenomic dataset. In some embodiments, the transcriptomic dataset is a pharmacotranscriptomic dataset. In some embodiments, the proteomic dataset is a pharmacoproteomic dataset. In some embodiments, the metabolomic dataset is a pharmacometabolomic dataset. In some embodiments, the lipidomic dataset is a pharmacolipidomic dataset. In some embodiments, the secretomic dataset is a pharmacosecretomic dataset.
[0191] The computer-implemented system can be configured to receive, organize, store, encrypt, and / or decrypt data from an external data source. The external data source can be a computer of a hospital, contract research organization, a laboratory, a physician, a public database etc. The incoming data can comprise raw genetic data (e.g., FASTQ, VCF, Micro-Array, PCR). The genetic data can comprise subject-specific variant data. The incoming data can comprise data from a public database (e.g., PharmGKB, PubChem, PharmaVar, FDA, EMA). The computer- implemented system can be configured to cross-reference publicly available drug-gene interaction guidelines for initial filtering. The incoming data can comprise clinical outcome data. The incoming data can comprise omic data (e.g., transcriptomic, proteomic, epidemiomic,epigenomic, metabolomic, lipidomic, secretomic). Hospitals and research partners can provide omics data to create and / or update a multi-omics dataset and multi-omics model training.
[0192] The database can be implemented cloud-computing. The database can also be implemented using local storage. For example, the database can be managed using an API Buffer (e.g., Flask) to receives, validates, and queues new (genomic) datasets or parameter files. Data can then be storage securely, e.g., using InterSystems IRIS Server.
[0193] The data can be preprocessed to structure the data. In some embodiments, preprocessing comprises filtering single-nucleotide-polymorphisms (SNPs). The data, which can be initially provided in array, matrix, XML, dictionary, struct, or another type of data form, can be parsed. Then, the parsed data can be analyzed to isolate clinically significant markers relevant to adverse drug reactions or efficacy. In some embodiments, preprocessing comprises filtering for specific pharmaceutical substances. In some embodiments, filtering can comprise cross-checking patient variants with a curated drug database, linking to public sources (PharmGKB, etc.) for evidencebased guidelines, or both. In some embodiments, the data can be unstructured data comprising natural language (e.g., scientific articles, guidelines, electronic health record comprising a narrative). In some embodiments, a neural network language model can be used to extract relevant features from the unstructured data.
[0194] The database can comprise a structured dataset. The structured dataset can comprise various features for training a machine learning model. The structured dataset can comprise molecular structures of pharmaceutical substances, phy si cal / chemi cal characteristics of pharmaceutical substances, combined genotype-phenotype predictions, or any combination thereof.
[0195] The computer-implemented system can comprise computer-executable instructions that are configured to checks the queue every few minutes for tasks, gather relevant genetic markers and drug data, launches parallel processes (multiprocessing) to streamline annotation and report compilation, or any combination thereof. As many clients (or users) may initiate data transfers to the computer-implemented system, or prompt drug response predictions from the computer- implemented system, the computer-executable instructions can be used to efficiently direct a queue of tasks to various processors of the computer-implemented system. In some embodiments, a computer-implemented system can automatically restart or scale containers if performance thresholds are exceeded. This can provide high availability for processing bursts (such as bulk lab uploads). In some embodiments, computer-implemented system can manage data configurations for drug-genome matching, dose adjustments, and alert thresholds for high- risk interactions.
[0196] The database can be continuously checked for quality. Various subsets of datasets in the database can be reviewed for consistency with new datasets as well as public datasets. Checking for quality can be performed, e.g., on a daily basis or weekly basis. User feedback can be used to refine the database, which can eventually be used to update the machine learning model. As another example, new genetic panels or drug expansions can be incorporated with containerized updates.
[0197] The database can be continuously updated to improve recommendation accuracy. Nonlimiting examples of updates to the databased include proprietary datasets containing pharmacogenomic knowledge, clinical guidelines, and real-world patient data. The proprietary datasets can be used to fine-tune the database. In some embodiments, the database can be augmented with external data sources. Non-limiting examples of external data sources include: medical literature, structured knowledge graphs, and proprietary databases, including but not limited to the proprietary dataset as described above. In some embodiments, the augmentation results in enriched contextual understanding of the database.
[0198] In some embodiments, the database comprising user-provided genetic data and related pharmacogenomic inputs can be preprocessed. Non-limiting examples of preprocessing include: standardization of variant call data (e.g., VCF, JSON, XML, or other formats) to ensure compatibility with a drug response prediction input structure; feature extraction and transformation, including the encoding of genetic marks, clinical metadata, and user-specific parameters; and integration with knowledge bases, including internal knowledge bases, allowing for retrieval and inclusion of additional context in a drug response prediction prompt.
[0199] In some embodiments, the database can be managed with multiple redundancies to safeguard data. For example, data can be mirrored across multiple zones or regions to mitigate potential data loss. In some embodiments, snapshots of the data can be made into backups.Multiomics
[0200] In some embodiments, a machine learning model can integrate omic data. In some embodiments, the omic data can be in the form of an omic dataset. In some embodiments, a machine learning model may generate a prediction based on omic dataset. In some embodiments a machine learning model can integrate or incorporate data from an omic dataset to enhance the accuracy of drug response predictions. In some embodiments, an omic dataset is a single-cell omic dataset, spatial omic profiling dataset, or any combination thereof. An omic dataset can be generated using RNAseq, DNAseq, ATACseq, in-situ hybridization based approaches, methylation assays, liquid chromatography -mass spectrometry, gas chromatography-mass spectrometry, and / or proteomics, in bulk, single cell, and / or spatial variants. In someembodiments, an omic dataset comprises a hybridization and / or imaging based profiling technique. In some embodiments, an omic dataset can be generated using, DNA-seq, ATAC-seq, Ribo-seq, single-cell RNA-seq, single-cell targeted DNA sequencing, single-cell ATAC-seq, spatially-resolved single-cell sequencing, spatially-resolved transcriptomic profiling, spatially- resolved DNA-seq, or any combination thereof. In some embodiments, the genetic sequencing comprises sequencing genetic material associated with expressed proteins by the one or more cells.
[0201] In some embodiments, the proteomics comprises iTRAQ, dissociated proteomic profiling, spatially-resolved proteomics, or any combination thereof. In some embodiments, the proteomics comprises, ELISA, gel-based proteomics, chromatography based proteomics, mass spectrometry, fractionation, or any combination thereof.
[0202] In some embodiments, the omic dataset is preprocessed to normalize and align heterogenous data for incorporation or integration in the machine learning model. In some embodiments, the preprocessing occurs via a unified pipeline.
[0203] The omic dataset can be received from a plurality of computers. Each computer in the plurality of computers can comprise a database comprising at least a portion of the omic dataset. The portions of the omic dataset, across the different computers, can be different from one another. For example, the portions may have omic datasets from different subjects, drugs, genes, trascriptome, proteome, epigenome, and / or metabolome. The omic dataset can be received from the plurality of computers, where each computer comprises a containerized computer-executable instructions configured to transmit the omic dataset to another computer. The another computer can be a central computer. Another containerized computer-executable instructions can be configured to train a neural network language model, or use a neural network, to generate a drug response prediction in natural language for a user of the computer.
[0204] In some embodiments, a method of the disclosure comprises obtaining epigenomic data of a biological sample. Non-limiting examples of epigenomic data include DNA methylation and histone modifications. In some embodiments, a method of the disclosure comprises obtaining transcriptomic data. Non-limiting examples of transcriptomic data include expression levels of drug-metabolizing enzymes, transports, and target genes. In some embodiments, a method of the disclosure comprises obtaining proteomic data. Non-limiting examples of proteomic data include biomarkers correlated or indicate with metabolic capacity drug clearance efficiency.
[0205] In some embodiments, a method of the disclosure comprises obtaining a genetic sequence of a biological sample. A genetic sequence can comprise any number of genes on any number of cells. In some embodiments, a gene refers to a locatable region of a genomic sequence, corresponding to a unit of inheritance, which is associated with regulatory regions, transcribedregions, or other functional sequence regions. In some embodiments, a gene can refer to an encoding, a representation, a distinction, or a name of a gene.
[0206] In some embodiments, an omic dataset can comprise data about gene expression of genes encoding proteins involving a metabolic pathway or a metabolic process or a metabolic reaction. In some embodiments, an omic dataset can comprise data about gene expression of healthy cells. In some embodiments, an omic dataset can comprise data about gene expression of harmful cells. In some embodiments, an omic dataset can comprise data about gene expression of cancer cells. In some embodiments, an omic dataset can comprise data about gene expression of tumor cells. In some embodiments, an omic dataset can comprise data about gene expression of a healthy population of a species. In some embodiments, an omic dataset comprises data about gene expression of a population of a species with a particular disease.
[0207] In some embodiments, an omic dataset comprises a proteomic profile of a biological sample. A proteomic profile can comprise any number of proteins with any number of cells. In some embodiments, a protein can refer to any molecule comprising at least two amino acids. In some embodiments, an omic dataset comprises a proteomic profile of a biological sample from a proteomic database. In some embodiments, a proteomic database of comprises ontology terms.
[0208] In some embodiments, a proteomic database comprises a database of biological functions assigned to one or more proteins. In some embodiments, a proteomic database comprises a database assigning one or more biological functions to a protein. In some embodiments, a proteomic database comprises a database assigning one or more proteins to a biological function. In some embodiments, a proteomic database comprises a plurality of ontology terms and each gene ontology term can be assigned to a plurality of proteins. In some embodiments, a proteomic database comprises at least 100, 1000, 5000, 10000, or 20000 proteins. In some embodiments, a proteomic database comprises at most 100, 1000, 5000, 10000, or 20000 proteins.
[0209] In some embodiments, the omic dataset can be integrated with, in operable communication with, or in electronic communication with a large language model. In some embodiments, the integration, operable communication, or electronic communication results in multi-modal learning, incorporating diverse biological signals into pharmacogenomic recommendation generation. In some embodiments, the integration, operable communication, or electronic communication results in context-aware adjustment of treatment recommendations based on real-time omic data, including but not limited to transcriptomic and metabolic state.
[0210] In some embodiments, a drug response prediction incorporating the omic dataset can be benchmarked against real-world clinical outcomes.
[0211] “A,” “an,” and “the” refer to “one or more” when used in this application, including the claims. Thus, for example, reference to “a subject” includes a plurality of subjects, unless the context clearly is to the contrary (e.g., a plurality of subjects), and so forth.
[0212] While preferred embodiments of the present disclosure have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the disclosure. It should be understood that various alternatives to the embodiments of the present disclosure may be employed in practicing the present disclosure. It is intended that the following claims define the scope of the present disclosure and that methods and structures within the scope of these claims and their equivalents be covered thereby.EMBODIMENTS
[0213] Embodiment 1 comprises a computer-implemented system comprising: a database comprising a structured pharmacogenomic dataset; a machine learning model comprising a neural network language model; and computer-executable instructions configured to:(a) train the machine learning model, using the structured pharmacogenomic dataset, to:(i) classify a feature vector as (1) a serviceable feature vector or (2) a non- serviceable feature vector for a drug response prediction; and(ii) generate for the serviceable feature vector, using the neural network language model, the drug response prediction in natural language;(b) update the structured pharmacogenomic dataset with a new pharmacogenomic dataset to generate an updated structured pharmacogenomic dataset; and(c) train a new machine learning model, using the updated structured pharmacogenomic dataset, such that the new machine learning model has more serviceable feature vectors than the machine learning model.
[0214] Embodiment 2 comprises a computer-implemented system comprising: a database comprising a structured pharmacogenomic dataset; a machine learning model comprising a neural network language model; and computer-executable instructions configured to:(a) train the machine learning model, using the structured pharmacogenomic dataset, to:(i) classify a feature vector as (1) a serviceable feature vector or (2) a non- serviceable feature vector for a drug response prediction; and(ii) generate for the serviceable feature vector, using the neural network language model, the drug response prediction in natural language;(b) update the structured pharmacogenomic dataset with a new pharmacogenomic dataset to generate an updated structured pharmacogenomic dataset; and(c) retrain the machine learning model, using the updated structured pharmacogenomic dataset, to add new serviceable feature vectors to the machine learning model.
[0215] Embodiment 3 comprises the computer-implemented system of embodiment 1 or 2, wherein the structured pharmacogenomic dataset comprises structural representations of pharmaceutical substances.
[0216] Embodiment 4 comprises the computer-implemented system of embodiment 2, wherein the structural representations comprise: molecular structures, fingerprints, latent vectors, SMILES (Simplified Molecular Input Line Entry System), InChi (International Chemical Identifier), SDF (Structure Data File), topological descriptors, 3D conformations, graph adjacency matrices, node and edge feature vectors for each atom and bond, autoencoder-derived embeddings, transformer-based embeddings for small molecules, graph neural network latent vectors, quantum chemical descriptors (partial charges, H0M0-LUM0 levels), protein-ligand interaction features (docking scores, binding affinity), pharmacophore or toxicophore features, or any combination thereof.
[0217] Embodiment 5 comprises the computer-implemented system of any one of embodiments 1-4, wherein the structured pharmacogenomic dataset comprises physical or chemical properties of pharmaceutical substances.
[0218] Embodiment 6 comprises the computer-implemented system of embodiment 5, wherein the physical or chemical properties comprise: pharmacokinetic properties, pharmacodynamic properties, molecular weights, logP, logD, pKa, solubility, melting point, boiling point, hydrogen bond donor count, hydrogen bond acceptor count, topological polar surface area (tPSA), rotatable bond count, refractivity, or any combination thereof.
[0219] Embodiment 7 comprises the computer-implemented system of any one of embodiments 1-6, wherein the structured pharmacogenomic dataset comprises genetic profiles of subjects.
[0220] Embodiment 8 comprises the computer-implemented system of embodiment 7, wherein the genetic profiles comprise: genomic sequences, mutations, SNPs, copy number variations (CNVs), insertions, deletions, haplotypes, splicing variants, epigenetic modifications, gene expression levels, genomic rearrangements, or any combination thereof.
[0221] Embodiment 9 comprises the computer-implemented system of any one of embodiments 1-8, wherein the feature vector comprises a structural representation of a pharmaceutical substance.
[0222] Embodiment 10 comprises the computer-implemented system of any one of embodiments 1-9, wherein the feature vector comprises a physical or chemical property of a pharmaceutical substance.
[0223] Embodiment 11 comprises the computer-implemented system of any one of embodiments 1-10, wherein (i) the serviceable feature vector, (ii) the non-serviceable feature vector, or (iii) both is input to the neural network language model as a structured prompt.
[0224] Embodiment 12 comprises the computer-implemented system of embodiment 11, wherein the structured prompt comprises: a feature definition, a feature label, a feature value, a prefix, a suffix, a question, a chain of thought, or any combination thereof.
[0225] Embodiment 13 comprises the computer-implemented system of 12, wherein the feature definition comprises context for a feature in natural language.
[0226] Embodiment 14 comprises the computer-implemented system of 12 or 13, wherein the feature label comprises a database identifier for a feature.
[0227] Embodiment 15 comprises the computer-implemented system of any one of embodiments 12-14, wherein the feature value comprises a quantitative value, a qualitative value, or both.
[0228] Embodiment 16 comprises the computer-implemented system of any one of embodiments 11-15, wherein the structured prompt comprises a plurality of features.
[0229] Embodiment 17 comprises the computer-implemented system of 16, wherein the plurality of features comprises: a molecular structure of a pharmaceutical substance, a molecular fingerprint of a pharmaceutical substance, a physicochemical characteristic of a pharmaceutical substance, a pharmacodynamic genetic characteristic of a pharmaceutical substance, a pharmacokinetic genetic characteristic of a pharmaceutical substance, pharmacoepigenomics of a pharmaceutical substance, pharmacotranscriptomics of a pharmaceutical substance, pharmacoproteomics of a pharmaceutical substance, pharmacometabolomics of a pharmaceutical substance, a clinical outcome of a pharmaceutical substance, allele variants, a list of genes that interact with a pharmaceutical substance, or any combination thereof.
[0230] Embodiment 18 comprises the computer-implemented system of any one of embodiments 1-17, wherein the serviceable feature vector indicates that a clinically effective recommendation of the drug response prediction can be generated for the feature vector.
[0231] Embodiment 19 comprises the computer-implemented system of any one of embodiments 1-18, wherein the non-serviceable feature vector indicates that a clinicallyeffective recommendation of the drug response prediction cannot be generated for the feature vector.
[0232] Embodiment 20 comprises the computer-implemented system of any one of embodiments 1-19, wherein the serviceable feature vector defines a genomic feature of a subject that is not found in the structured pharmacogenomic dataset.
[0233] Embodiment 21 comprises the computer-implemented system of embodiment 20, wherein the serviceable feature vector defines a combination of genomic features of the subject that is not found in the structured pharmacogenomic dataset.
[0234] Embodiment 22 comprises the computer-implemented system of any one of embodiments 1-21, wherein the serviceable feature vector defines a pharmacological substance that is not found in the structured pharmacogenomic dataset.
[0235] Embodiment 23 comprises the computer-implemented system of embodiment 22, wherein the serviceable feature vector defines a combination of pharmacological substances that is not found in the structured pharmacogenomic dataset.
[0236] Embodiment 24 comprises the computer-implemented system of any one of embodiments 1-23, wherein the serviceable feature vector defines a combination of (i) a pharmacological substance and (ii) genomic features of the subject that is not found in the structured pharmacogenomic dataset.
[0237] Embodiment 25 comprises the computer-implemented system of any one of embodiments 1-24, wherein the serviceable feature vector defines a combination of (i) pharmacological substances and (ii) genomic features of the subject that is not found in the structured pharmacogenomic dataset.
[0238] Embodiment 26 comprises the computer-implemented system of any one of embodiments 1-25, wherein the feature vector is classified as the serviceable feature vector or the non-serviceable feature vector based on a confidence score.
[0239] Embodiment 27 comprises the computer-implemented system of embodiment 26, wherein the confidence score is determined by processing the feature vector to quantify a similarity between (i) a pharmaceutical substance defined by the feature vector and (ii) a different pharmaceutical substance defined by the structured pharmacogenomic dataset.
[0240] Embodiment 28 comprises the computer-implemented system of embodiment 26 or 27, wherein the confidence score is determined by processing the feature vector to quantify (i) a pharmacokinetic property defined by the structured pharmacogenomic dataset, (ii) a pharmacodynamic property defined by the structured pharmacogenomic dataset, (iii) a predicted pharmacokinetic property, (iv) a predicted pharmacodynamic property, or any combination thereof.
[0241] Embodiment 29 comprises the computer-implemented system of any one of embodiments 26-28, wherein the confidence score is determined by an association between a pharmaceutical substance and genomic features of the subject.
[0242] Embodiment 30 comprises the computer-implemented system of any one of embodiments 26-29, wherein the confidence score is determined by processing an association between a pharmaceutical substance and (i) epigenetic features, (ii) proteomic features, (iii) transcriptomic features, (iv) metabolomic features of the subject, or any combination thereof.
[0243] Embodiment 31 comprises the computer-implemented system of any one of embodiments 26-30, wherein the confidence score is determined by processing an availability and relevance of clinical outcome data for a pharmacological substance defined by the feature vector.
[0244] Embodiment 32 comprises the computer-implemented system of any one of embodiments 26-31, wherein the feature vector is classified as the serviceable feature vector when the confidence score exceeds a predetermined threshold.
[0245] Embodiment 33 comprises the computer-implemented system of embodiment 32, wherein the predetermined threshold is 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, 99.91%, 99.92%, 99.93%, 99.94%, 99.95%, 99.96%, 99.97%, 99.98%, 99.99%, 99.991%, 99.992%, 99.993%, 99.994%, 99.995%, 99.996%, 99.997%, 99.998%, 99.999% confidence that the feature vector is a serviceable feature vector.
[0246] Embodiment 34 comprises the computer-implemented system of embodiment 33, wherein the predetermined threshold is 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, 99.91%, 99.92%, 99.93%, 99.94%, 99.95%, 99.96%, 99.97%, 99.98%, 99.99%, 99.991%, 99.992%, 99.993%, 99.994%, 99.995%, 99.996%, 99.997%, 99.998%, 99.999% confidence that the feature vector is a non-serviceable feature vector.
[0247] Embodiment 35 comprises the computer-implemented system of any one of embodiments 1-34, wherein the new pharmacogenomic dataset comprises: raw genetic data, clinical outcome data, scientific papers, guidelines, electronic health records, comorbidities, environmental factors, concomitant medications, adverse event reports, real-world evidence from wearable devices, retrospective and prospective cohort data, structured references from curated databases (e.g., ClinVar, PharmGKB), laboratory test results, phenotype data, demographic profiles, patient compliance data, toxicology studies, and multi-center consortia datasets, or any combination thereof.
[0248] Embodiment 36 comprises the computer-implemented system of any one of embodiments 1-35, wherein the computer-executable instructions are configured to extract features useful for drug response prediction in the new pharmacogenomic dataset.
[0249] Embodiment 37 comprises the computer-implemented system of any one of embodiments 1-36, wherein the computer-executable instructions are configured to filter mutations in the new pharmacogenomic dataset.
[0250] Embodiment 38 comprises the computer-implemented system of any one of embodiments 1-37, wherein the computer-executable instructions are configured to isolate clinically significant markers, from the new pharmacogenomic dataset, that are relevant to adverse drug reactions or efficacy.
[0251] Embodiment 39 comprises the computer-implemented system of any one of embodiments 1-38, wherein the computer-executable instructions are configured to filter pharmaceutical substances in the new pharmacogenomic dataset.
[0252] Embodiment 40 comprises the computer-implemented system of any one of embodiments 1-39, wherein the computer-executable instructions are configured to, in the new pharmacogenomic dataset, parse unstructured data (e.g., scientific papers, guidelines) into structured formats.
[0253] Embodiment 41 comprises the computer-implemented system of any one of embodiments 1-40, wherein the computer-executable instructions are configured to, in the new pharmacogenomic dataset, annotate data with standardized ontologies (e.g., gene and drug nomenclatures).
[0254] Embodiment 42 comprises the computer-implemented system of any one of embodiments 1-41, wherein the computer-executable instructions are configured to, in the new pharmacogenomic dataset, remove duplicates and resolve inconsistent entries.
[0255] Embodiment 43 comprises the computer-implemented system of any one of embodiments 1-42, wherein the computer-executable instructions are configured to apply dimensionality reduction or feature engineering techniques.
[0256] Embodiment 44 comprises the computer-implemented system of any one of embodiments 1-43, wherein the computer-executable instructions are configured to integrate the new pharmacogenomic dataset with the structured pharmacogenomic dataset.
[0257] Embodiment 45 comprises the e computer-implemented system of any one of embodiments 1-44, wherein the computer-executable instructions are configured to, using the updated structured pharmacogenomic dataset, generate updated feature vectors for training the new machine learning model.
[0258] Embodiment 46 comprises the computer-implemented system of any one of embodiments 1-45, wherein the machine learning model comprises a binary classifier.
[0259] Embodiment 47 comprises the computer-implemented system of embodiment 46, wherein the binary classifier comprises a random forest model.
[0260] Embodiment 48 comprises the computer-implemented system of embodiment 47, wherein the binary classifier comprises a boosted tree algorithm.
[0261] Embodiment 49 comprises the computer-implemented system of any one of embodiments 1-48, wherein the neural network language model comprises an autoregressive model.
[0262] Embodiment 50 comprises the computer-implemented system of any one of embodiments 1-49, wherein the neural network language model comprises a transformer.
[0263] Embodiment 51 comprises the computer-implemented system of any one of embodiments 1-50, wherein the neural network language model comprises a large language model.
[0264] Embodiment 52 comprises the computer-implemented system of any one of embodiments 1-51, wherein the drug response prediction comprises a personalized regimen for administering a pharmaceutical substance for a subject.
[0265] Embodiment 53 comprises the computer-implemented system of any one of embodiments 1-52, wherein the drug response prediction comprises a personalized dosing guideline for a subject.
[0266] Embodiment 54 comprises the computer-implemented system of any one of embodiments 1-53, wherein the drug response prediction comprises a personalized drug selection for a subject.
[0267] Embodiment 55 comprises the computer-implemented system of any one of embodiments 1-54, wherein the drug response prediction comprises a prediction of an adverse drug reaction in a subject.
[0268] Embodiment 56 comprises the computer-implemented system of any one of embodiments 52-55, further comprising computer-executable instructions configured to print a label of the personalized regimen, the personalized dosing guideline, the personalized drug selection, or any combination thereof.
[0269] Embodiment 57 comprises the computer-implemented system of embodiment 56, further comprising computer-executable instructions configured to apply the label to a packaging for a pharmaceutical substance.
[0270] Embodiment 58 comprises the computer-implemented system of any one of embodiments 1-57, wherein multiple drug response predictions are generated to statistically filter hallucinated drug response predictions.
[0271] Embodiment 59 comprises the computer-implemented system of any one of embodiments 1-58, wherein the natural language is English, Spanish, German, French, Russian, Mandarin Chinese, Cantonese Chinese, French, Hindi, Korean, or Japanese.
[0272] Embodiment 60 comprises a computer-implemented system comprising: a database comprising a structured dataset, wherein the structured dataset comprises a pharmacogenomic dataset relating at least 1,200 pharmacological substances and at least 350 genes; a machine learning model comprising a neural network language model; and computer-executable instructions configured to train the machine learning model to generate a drug response prediction in natural language, wherein the training is based on the structured dataset.
[0273] Embodiment 61 comprises a computer-implemented system comprising: a machine learning model comprising a neural network language model, wherein the machine learning model is trained using a structured dataset comprising a pharmacogenomic dataset relating at least 1,200 pharmacological substances and at least 350 genes; and computer-executable instructions configured to process a feature vector, using the machine learning model, to generate a drug response prediction in natural language.
[0274] Embodiment 62 comprises the computer-implemented system of embodiment 60 or 61, wherein the pharmacogenomic dataset comprises at least 1,700, 2,500, 3,500, 5,000, 10,000, or 21,000 pharmacological substances.
[0275] Embodiment 63 comprises the computer-implemented system of any one of embodiments 60-62, wherein the pharmacogenomic dataset comprises at most 2,500, 3,500, 5,000, 10,000, 21,000, 30,000 pharmacological substances.
[0276] Embodiment 64 comprises the computer-implemented system of any one of embodiments 60-63, wherein the pharmacogenomic dataset comprises at least 700, 1,300, 2,400, 4,500, 10,000, or 20,000 genes.
[0277] Embodiment 65 comprises the computer-implemented system of any one of embodiments 60-64, wherein the pharmacogenomic dataset comprises at most 1,300, 2,400, 4,500, 10,000, 20,000, 25,000 genes.
[0278] Embodiment 66 comprises a computer-implemented system comprising:a database comprising a structured dataset, wherein the structured dataset comprises a pharmacogenomic dataset and at least one of: (i) an epigenomic dataset, (ii) a transcriptomic dataset, (iii) a proteomic dataset, (iv) a metabolomic dataset, (v) a lipidomic dataset, and (vi) a secretomic dataset; a machine learning model comprising a neural network language model; and computer-executable instructions configured to: train the machine learning model, using the structured dataset, to:(i) classify a feature vector as (1) a serviceable feature vector or (2) a non- serviceable feature vector for a drug response prediction; and(ii) generate for the serviceable feature vector, using the neural network language model, the drug response prediction in natural language.
[0279] Embodiment 67 comprises the computer-implemented system of embodiment 66, wherein the epigenomic dataset is a pharmacoepigenomic dataset.
[0280] Embodiment 68 comprises the computer-implemented system of embodiment 66 or 67, wherein the transcriptomic dataset is a pharmacotranscriptomic dataset.
[0281] Embodiment 69 comprises the computer-implemented system of any one of embodiments 66-68, wherein the proteomic dataset is a pharmacoproteomic dataset.
[0282] Embodiment 70 comprises the computer-implemented system of any one of embodiments 66-69, wherein the metabolomic dataset is a pharmacometabolomic dataset.
[0283] Embodiment 71 comprises the computer-implemented system of any one of embodiments 66-70, wherein the lipidomic dataset is a pharmacolipidomic dataset.
[0284] Embodiment 72 comprises the computer-implemented system of any one of embodiments 66-71, wherein the secretomic dataset is a pharmacosecretomic dataset.
[0285] Embodiment 73 comprises the computer-implemented system of any one of embodiments 1-72, wherein the drug response prediction is a prediction for using a pharmaceutical substance to treat a disease.
[0286] Embodiment 74 comprises the computer-implemented system of embodiment 73, wherein the disease is a cancer.
[0287] Embodiment 75 comprises the computer-implemented system of embodiment 74, wherein the pharmaceutical substance is a chemotherapeutic.
[0288] Embodiment 76 comprises the computer-implemented system of embodiment 75, wherein the chemotherapeutic is an EGFR inhibitor for treating lung cancer, or a PARP inhibitor for treating BRCA-mutated breast cancer.
[0289] Embodiment 77 comprises the computer-implemented system of any one of embodiments 73-76, wherein the drug response prediction comprises a response prediction of an immune checkpoint inhibitor based on tumor mutational burden.
[0290] Embodiment 78 comprises the computer-implemented system of any one of embodiments 73-77, wherein the drug response prediction is based tumor mutations, pharmacogenomic biomarkers, immunotherapy response predictors, or any combination thereof.
[0291] Embodiment 79 comprises the computer-implemented system of any one of embodiments 73-78, wherein the disease is an autoimmune disease.
[0292] Embodiment 80 comprises the computer-implemented system of embodiment 79, wherein the autoimmune disease is rheumatoid arthritis, lupus, or multiple sclerosis.
[0293] Embodiment 81 comprises the computer-implemented system of any one of embodiments 73-80, wherein the drug response prediction is based on genetic susceptibility to drug-induced side effects.
[0294] Embodiment 82 comprises the computer-implemented system of any one of embodiments 73-81, wherein the drug response prediction comprises a providing a therapy selection among an array of therapies.
[0295] Embodiment 83 comprises the computer-implemented system of embodiment 82, wherein the array of therapies comprises an array of biologies.
[0296] Embodiment 84 comprises the computer-implemented system of embodiment 83, wherein the array of biologies comprises TNF inhibitors and IL-6 inhibitors.
[0297] Embodiment 85 comprises the computer-implemented system of any one of embodiments 73-84, wherein the pharmaceutical substance comprises methotrexate, TNF inhibitors, or IL-6 inhibitors.
[0298] Embodiment 86 comprises the computer-implemented system of any one of embodiments 73-85, wherein the disease is a metabolic disorder.
[0299] Embodiment 87 comprises the computer-implemented system of embodiment 86, wherein the metabolic disorder is obesity.
[0300] Embodiment 88 comprises the computer-implemented system of embodiment 86 or 87, wherein the pharmaceutical substance comprises a GLP-1 agonist, semaglutide, tirzepatide, or any combination thereof.
[0301] Embodiment 89 comprises the computer-implemented system of any one of embodiments 73-88, wherein the disease is a psychiatric or a neurological disease.
[0302] Embodiment 90 comprises the computer-implemented system of embodiment 89, wherein the disease is epilepsy or a neurodegenerative disease.
[0303] Embodiment 91 comprises the computer-implemented system of embodiment 89 or 90, wherein the pharmaceutical substance comprises an antidepressant or an antipsychotic.
[0304] Embodiment 92 comprises the computer-implemented system of embodiment 91, wherein the pharmaceutical substance comprises an SSRI or a tricyclic antidepressant.
[0305] Embodiment 93 comprises the computer-implemented system of any one of embodiments 73-92, wherein the drug response prediction provides a recommending dosing regimen of the pharmaceutical substance based on predicted metabolism of the pharmaceutical substance by a subject.
[0306] Embodiment 94 comprises the computer-implemented system of any one of embodiments 73-93, wherein the disease is a cardiovascular disease.
[0307] Embodiment 95 comprises the computer-implemented system of embodiment 94, wherein the pharmaceutical substance comprises an antiplatelet or warfarin.
[0308] Embodiment 96 comprises the computer-implemented system of any one of embodiments 73-95, wherein the drug response prediction is based on clopidogrel response based on CYP2C19 status.
[0309] Embodiment 97 comprises the computer-implemented system of any one of embodiments 73-96, wherein the drug response prediction comprises VKORCl / CYP2C9-guided adjustments.
[0310] Embodiment 98 comprises the computer-implemented system of any one of embodiments 1-97, wherein the drug response prediction indicates a set of candidate subjects that are predicted to benefit from administration of a pharmaceutical substance.
[0311] Embodiment 99 comprises the computer-implemented system of any one of embodiments 1-98, wherein the drug response prediction indicates a set of candidate subjects that are predicted not to experience adverse effects from administration of a pharmaceutical substance.
[0312] Embodiment 100 comprises the computer-implemented system of any one of embodiments 1-99, wherein the drug response prediction indicates a set of candidate subjects that are predicted to experience adverse effects from administration of a pharmaceutical substance.
[0313] Embodiment 101 comprises the computer-implemented system of any one of embodiments 1-100, wherein the drug response prediction identifies drug-gene interactions for a use of the pharmaceutical substance to treat a rare disease.
[0314] Embodiment 102 comprises the computer-implemented system of any one of embodiments 1-101, wherein the drug response prediction comprises a companion diagnostics.
[0315] Embodiment 103 comprises the computer-implemented system of any one of embodiments 1-102, wherein the system generates list of subjects having a biomarkers that match a biomarker-based clinical trial enrollment criteria.
[0316] Embodiment 104 comprises the computer-implemented system of any one of embodiments 1-103, wherein the system is in operable communication with a personal device of a user to provide telemedicine services to the user.
[0317] Embodiment 105 comprises the computer-implemented system of any one of embodiments 1-104, wherein the system generates policy pricing based on the drug response prediction.
[0318] Embodiment 106 comprises the computer-implemented system of any one of embodiments 1-105, wherein the system generates a preventive healthcare program based on the drug response prediction.
[0319] Embodiment 107 comprises the computer-implemented system of any one of embodiments 1-106, wherein the system is implemented within a wearable device.
[0320] Embodiment 108 comprises the computer-implemented system of any one of embodiments 1-107, wherein the system is in operable communication with a wearable device configured to provide the system with user health data.
[0321] Embodiment 109 comprises the computer-implemented system of any one of embodiments 1-108, wherein the drug response prediction is based on real-time updates of a subject’s or a plurality of subjects’ epigenomic, transcriptomic, proteomic, or metabolomic state.
[0322] Embodiment 110 comprises a computer-implemented system comprising: a database comprising an internal pharmacogenomic dataset encrypted based on an internal cipher; a neural network language model; and computer-executable instructions configured to:(i) create a secure connection with a client to receive an external pharmacogenomic dataset encrypted based on an external cipher;(ii) receive the external pharmacogenomic dataset through the secure connection;(iii) decrypt the external pharmacogenomic dataset based on the external cipher;(iv) add the external pharmacogenomic dataset to the internal pharmacogenomic dataset;(v) decrypt the internal pharmacogenomic dataset based on the internal cipher; and(vi) train the neural network language model, using the internal pharmacogenomic dataset, to generate a drug response prediction in natural language.
[0323] Embodiment 111 comprises the computer-implemented system of embodiment 110, wherein the computer-executable instructions are configured to provide access privileges to a plurality of clients, wherein the access privileges are based on a client profile.
[0324] Embodiment 112 comprises the computer-implemented system of embodiment 111, wherein the access privileges permit access to the neural network language model and not the database.
[0325] Embodiment 113 comprises the computer-implemented system of embodiment 111 or 112, wherein the access privileges permit access to the database and not the neural network language model.
[0326] Embodiment 114 comprises the computer-implemented system of any one of embodiments 111-113, wherein the access privileges permit access to the neural network language model and the database.
[0327] Embodiment 115 comprises the computer-implemented system of any one of embodiments 111-114, wherein the client is an organization.
[0328] Embodiment 116 comprises the computer-implemented system of any one of embodiments 111-115, wherein the client is an individual user account.
[0329] Embodiment 117 comprises the computer-implemented system of any one of embodiments 111-116, wherein the client profile comprises jurisdiction.
[0330] Embodiment 118 comprises the computer-implemented system of embodiment 117, wherein the jurisdiction comprises a jurisdiction of registration.
[0331] Embodiment 119 comprises the computer-implemented system of embodiment 117 or 118, wherein the jurisdiction comprises a jurisdiction of access.
[0332] Embodiment 120 comprises the computer-implemented system of any one of embodiments 117-119, wherein the jurisdiction comprises US jurisdiction, European jurisdiction, Chinese jurisdiction, South Korean jurisdiction, Australian jurisdiction, African jurisdiction, Russian jurisdiction, Canadian jurisdiction, Japanese jurisdiction, Indian jurisdiction, Brazilian jurisdiction.
[0333] Embodiment 121 comprises the computer-implemented system of any one of embodiments 111-120, wherein the access privileges comprise (i) read access, (ii) write access, (iii) execute access, (iv) delete access, (v) full privileged access, or any combination thereof.
[0334] Embodiment 122 comprises the computer-implemented system of any one of embodiments 110-121, wherein the computer-executable instructions are configured to log traffic to and from the computer-implemented system.
[0335] Embodiment 123 comprises the computer-implemented system of any one of embodiments 110-122, wherein the computer-executable instructions are configured to log user access to the computer-implemented system.
[0336] Embodiment 124 comprises the computer-implemented system of any one of embodiments 110-123, wherein the computer-executable instructions are configured to log changes to the computer-implemented system.
[0337] Embodiment 125 comprises a computer-implemented system comprising:(a) a neural network language model;(b) a plurality of computers, each computer comprising:(i) a database comprising a pharmacogenomic dataset, wherein the pharmacogenomic dataset of each computer in the plurality of computers is different from one another;(ii) a containerized computer-executable instructions configured to train the neural network language model using the computer’s computational resources and the computer’s pharmacogenomic dataset, without sharing the computer’s pharmacogenomic dataset with another computer in the plurality of computers.
[0338] Embodiment 126 comprises the computer-implemented system of embodiment 125, further comprising a second containerized computer-executable instructions configured to use the neural network language model to generate a drug response prediction in natural language for a user of the computer, without sharing the user’s input to the neural network language model with another computer in the plurality of computers.
[0339] Embodiment 127 comprises the computer-implemented system of embodiment 125 or 126, wherein the containerized computer computer-executable instructions are configured to run using different architectures of different computers.
[0340] Embodiment 128 comprises the computer-implemented system of any one of embodiments 125-127, wherein the containerized computer computer-executable instructions are configured to load balance.
[0341] Embodiment 129 comprises the computer-implemented system of any one of embodiments 125-128, further comprising a central computer, wherein the central computer is configured to aggregate the trained neural network language model from the plurality of computers.
[0342] Embodiment 130 comprises the computer-implemented system of any one of embodiments 125-129, wherein the first containerized computer computer-executable instructions are configured to run in parallel at different computers.
[0343] Embodiment 131 comprises the computer-implemented system of any one of embodiments 125-130, wherein each computer further comprises a third containerized computerexecutable instructions for transferring the pharmacogenomic dataset to another computer.
[0344] Embodiment 132 comprises the computer-implemented system of any one of embodiments 125-131, wherein each computer further comprises a fourth containerized computer-executable instructions for generating a report of computer activities performed at the computer.
[0345] Embodiment 133 comprises the computer-implemented system of any one of embodiments 125-132, wherein the containerized computer computer-executable instructions are configured to partition the pharmacogenomic dataset and cache the pharmacogenomic dataset across multiple servers.
[0346] Embodiment 134 comprises a method for predicting a patient response to a drug, the method comprising: extracting data from at least one database of correspondence between genetic alleles and drug responses; integrating the data using ML / Al algorithms to provide a set of drug response predictions; iteratively adding and integrating new data to the set of drug response predictions, wherein the new data comprises at least one of: a new drug; a new patient; and a new allele; and at least one of: calculating a score for the patient response to the drug using the set of drug response predictions, wherein the score corresponds to a prediction to use: a standard dose of the drug; an adjusted dose of the drug; or an alternative drug; calculating a specific therapy with a specific drug for the patient; providing a recommendation to a physician based on a relation between the genetic alleles and drug therapy outcomes; and calculating a new drug therapy by iteratively adding and integrating new data to the set of drug response predictions.
[0347] Embodiment 135 comprises a computer-implemented method for providing pharmacogenetic guidelines, the method comprising: inputting a drug prescription and a patient genotype; andanalyzing the drug prescription and the patient genotype to provide a drug therapy recommendation; wherein the analyzing comprises: comparing the patient genotype and drug name to a drug profile database and a molecular biomarker database and calculating a score for a patient response to the drug, wherein the score corresponds to a prediction to use:1. a standard dose of the drug;2. an adjusted dose of the drug; or3. an alternative drug.
[0348] Embodiment 136 comprises the method as in embodiment 134 or 135, wherein the analyzing comprises: comparing the patient genotype and the drug name to a drug profile database and a molecular biomarker database iteratively using artificial intelligence and generating a specific drug therapy recommendation for a specific patient.
[0349] Embodiment 137 comprises the method as in any one of embodiments 134-136, wherein the analyzing comprises: comparing the patient genotype and the drug name to a drug profile database and a molecular biomarker database iteratively using artificial intelligence and generating a drug recommendation for the drug to a hospital.
[0350] Embodiment 138 comprises the method as in any one of embodiments 134-137, wherein the analyzing comprises: comparing the patient genotype and the drug name to a drug profile database, a molecular biomarker database and a hospital database of reactions to the drug iteratively using artificial intelligence and generating a drug recommendation for the drug to the hospital.
[0351] Embodiment 139 comprises the method as in any one of embodiments 134-138, wherein the analyzing comprises: comparing the patient genotype and the drug name to a drug profile database, a molecular biomarker database, a hospital database of reactions to the drug and at least one of an omics database iteratively using artificial intelligence and generating a drug recommendation for the drug to a hospital using artificial intelligence.
[0352] Embodiment 140 comprises a system for providing pharmacogenetic guidelines, the system comprising:a computer input system for entering a drug prescription and a patient genotype; and a computer processor for analyzing the drug prescription and the patient genotype to provide a drug therapy recommendation; wherein the analyzing comprises: comparing the patient genotype and a drug name to a drug profile database and a molecular biomarker database and calculating a score for a patient response to the drug, wherein the score corresponds to a prediction to use:1. a standard dose of the drug;2. an adjusted dose of the drug; or3. an alternative drug.
[0353] Embodiment 141 comprises the system as in embodiment 140, wherein the analyzing comprises: comparing the patient genotype and the drug name to a drug profile database and a molecular biomarker database iteratively using artificial intelligence and generating a specific drug therapy recommendation for a specific patient.
[0354] Embodiment 142 comprises the system as in embodiment 140 or 141, wherein the analyzing comprises: comparing the patient genotype and the drug name to a drug profile database and a molecular biomarker database iteratively using artificial intelligence and generating a drug recommendation for the drug to a hospital.
[0355] Embodiment 143 comprises the system as in any one of embodiments 140-142, wherein the analyzing comprises: comparing the patient genotype and the drug name to a drug profile database, a molecular biomarker database and a hospital database of reactions to the drug iteratively using artificial intelligence and generating a drug recommendation for the drug to the hospital.
[0356] Embodiment 144 comprises the system as in any one of embodiments 140-143, wherein the analyzing comprises: comparing the patient genotype and the drug name to a drug profile database, a molecular biomarker database, a hospital database of reactions to the drug and at least one of an omics database iteratively using artificial intelligence and generating a drug recommendation for the drug to a hospital using artificial intelligence.
[0357] Embodiment 145 comprises a computer-implemented method comprising:(a) training a machine learning model comprising a neural network language model, using a structured pharmacogenomic dataset, to:(i) classify a feature vector as (1) a serviceable feature vector or (2) a non- serviceable feature vector for a drug response prediction; and(ii) generate for the serviceable feature vector, using the neural network language model, the drug response prediction in natural language;(b) updating the structured pharmacogenomic dataset with a new pharmacogenomic dataset to generate an updated structured pharmacogenomic dataset; and(c) training a new machine learning model, using the updated structured pharmacogenomic dataset, such that the new machine learning model has more serviceable feature vectors than the machine learning model.
[0358] Embodiment 146 comprises a computer-implemented method comprising:(a) training a machine learning model comprising a neural network language model, using a structured pharmacogenomic dataset, to:(i) classify a feature vector as (1) a serviceable feature vector or (2) a non- serviceable feature vector for a drug response prediction; and(ii) generate for the serviceable feature vector, using the neural network language model, the drug response prediction in natural language;(b) updating the structured pharmacogenomic dataset with a new pharmacogenomic dataset to generate an updated structured pharmacogenomic dataset; and(c) retraining the machine learning model, using the updated structured pharmacogenomic dataset, to add new serviceable feature vectors to the machine learning model.
[0359] Embodiment 147 comprises a computer-implemented method, comprising: using a machine learning model comprising a neural network language model to process a feature vector to generate a drug response prediction in natural language, wherein the machine learning model is trained based on a structured dataset comprising a pharmacogenomic dataset relating at least 1,200 pharmacological substances and at least 350 genes.
[0360] Embodiment 148 comprises a computer-implemented method, comprising: training a machine learning model comprising a neural network language model to process a feature vector to generate a drug response prediction in natural language, wherein the training is based on astructured dataset comprising a pharmacogenomic dataset relating at least 1,200 pharmacological substances and at least 350 genes.
[0361] Embodiment 149 comprises a computer-implemented method, comprising: using a machine learning model comprising a neural network language model to:(a) classify a feature vector as (1) a serviceable feature vector or (2) a non- serviceable feature vector for a drug response prediction; and(b) generate a drug response prediction in natural language; wherein the machine learning model is trained based on a structured dataset comprising a pharmacogenomic dataset and at least one of: (i) an epigenomic dataset, (ii) a transcriptomic dataset, (iii) a proteomic dataset, (iv) a metabolomic dataset, (v) a lipidomic dataset, and (vi) a secretomic dataset.
[0362] Embodiment 150 comprises a computer-implemented method, comprising: training a machine learning model comprising a neural network language model to:(a) classify a feature vector as (1) a serviceable feature vector or (2) a non- serviceable feature vector for a drug response prediction; and(b) generate a drug response prediction in natural language; wherein the training is based on a structured dataset a pharmacogenomic dataset and at least one of: (i) an epigenomic dataset, (ii) a transcriptomic dataset, (iii) a proteomic dataset, (iv) a metabolomic dataset, (v) a lipidomic dataset, and (vi) a secretomic dataset.
[0363] Embodiment 151 comprises a computer-implemented method, comprising:(a)creating a secure connection with a client to receive an external pharmacogenomic dataset encrypted based on an external cipher;(b) receiving the external pharmacogenomic dataset through the secure connection;(c) decrypting the external pharmacogenomic dataset based on the external cipher;(d) adding the external pharmacogenomic dataset to the internal pharmacogenomic dataset;(e) decrypting the internal pharmacogenomic dataset based on an internal cipher; and(f) training the neural network language model, using the internal pharmacogenomic dataset, to generate a drug response prediction in natural language.
[0364] Embodiment 152 comprises a computer-implemented method, comprising:training a neural network language model using a plurality of computers, wherein each computer comprises:(i) a database comprising a pharmacogenomic dataset, wherein the pharmacogenomic dataset of each computer in the plurality of computers is different from one another; and(ii) a containerized computer-executable instructions configured to train the neural network language model using the computer’s computational resources and the computer’s pharmacogenomic dataset, without sharing the computer’s pharmacogenomic dataset with another computer in the plurality of computers.EXAMPLES
[0365] The following examples are provided to further illustrate some embodiments of the present disclosure, but are not intended to limit the scope of the disclosure; it will be understood by their exemplary nature that other procedures, methodologies, or techniques known to those skilled in the art may alternatively be used.Example 1: System and Method for Predicting a Patient Response to a Drug
[0366] This example provides a novel approach to pharmacogenetic guideline development for improved speed and accuracy of predicted recommendations. This is achieved using a machine learning / artificial intelligence (ML / Al) system based on accumulated pharmacogenetic data and further training it by integrating the ability to generate feedback using therapy outcome results (effective therapy, no effect, adverse reaction). This significantly accelerates the development and update of pharmacogenetic guidelines, which improves efficacy and safety of patient therapy.
[0367] By ensuring accuracy and reducing the time spent on the development and review of guidelines, machine learning opens the door for a broader and timelier implementation of pharmacogenetic testing in clinical practice, making personalized medicine accessible to a greater number of patients. This approach also implies an improvement in clinical outcomes and the optimization of healthcare processes, thereby acting as a catalyst for future innovations in medicine.
[0368] Other studies using ML / Al have involved training based on articles, whereas the present disclosure utilizes layers related to drugs (chemical and physical properties), molecular biomarkers, clinical biomarkers and solely on recommendations. This represents afundamentally different approach that allows for more than just interacting with PubMed posttraining, the approach also allows for making predictions based on chemical and physical properties to suggest new recommendations. In another study, the use of multi-omics biomarkers to enhance the efficacy and safety of NSAIDS, without any AI / ML involvement is discussed, and a further study is on general pharmacogenetics (without AI / ML).
[0369] Referring to FIG. 1, in an embodiment of the a system for providing drug guidelines using pharmacogenetics is provided. Data input 100 includes a prescription 105 of a drug 110 and a patient’s genotyping results 115 from a genetic laboratory 120.
[0370] Data input 100 occurs when a doctor prescribes a drug for a patient and enters the drug prescription into the hospital Electronic Health Record (EHR) 105. The drug name 110 is the international nonproprietary name (INN) of the drug intended to be prescribed to the patient.
[0371] Raw genotyping data from a genetic laboratory is uploaded into the EHR 120. The raw genotyping data may include results from NGS (FASTQ / VCF), micro-array, PCR Real-time, etc. When uploaded from the EHR using IRIS for Health (InterSystems), the raw genotyping data is transformed into a format convenient for the analysis 200 (i.e., a database with two columns).
[0372] The patient genotyping results 115 correspond to a specific single nucleotide polymorphism (SNP) with its genotype. Essentially, the file is a database with two columns, where the first column contains the Reference SNP (rs) name, and the second the SNP state: “wild type,” “heterozygote,” “mutant homozygote.” Additionally, the number of gene copies, phenocopying might be considered.
[0373] If patient genotyping data 115 are available, a request 125 is made to the analysis system 200 with the drug name 110 and the patient genotyping data 115. The input data 100 is analyzed by the analysis system 200 to produce an output 600. Output 600 provides a recommendation with information on how to proceed in the given clinical situation.
[0374] The analysis system 200 includes databases on which the generative Al model is trained, as well as the sources from which they are filled and updated. The analysis system 200 includes a chemical structure, genomic and protein biomarkers of drugs subsystem 300, a clinical biomarkers subsystem 400 and an omics biomarkers subsystem 500.
[0375] The chemical structure, genomic and protein biomarkers of drugs subsystem 300 includes databases used primarily for the development of ML classifiers and GenAI. Chemicals. tsv (PharmGKB) 310 is a database available for download at PharmGKB and provides the nomenclature of drugs for which genetic data are available, along with the Simplified Molecular Input Line Entry System (SMILES) representation of the drugs. In Drugs Fingerprints 315 data are generated based on the transformation of SMILES of drugs obtained from chemicals. tsv (PharmGKB) 310. Morgan Fingerprints are used for encoding drugs for Al training.
[0376] PubChem 320 is used for collecting data concerning the physical and chemical properties of drugs (such as Molecular Weight, XLogP3, Hydrogen Bond Donor Count, Hydrogen Bond Acceptor Count, Rotatable Bond Count, etc.) and proteins associated with the pharmacodynamics and pharmacokinetics of the drug.
[0377] A drugs chemical characteristics database 325 includes data related to the physical and chemical characteristics of drugs (such as Molecular Weight, XLogP3, Hydrogen Bond Donor Count, Hydrogen Bond Acceptor Count, Rotatable Bond Count, etc.), which are parsed from PubChem.
[0378] A drugs protein profile database 330 includes proteins associated with the pharmacodynamics and pharmacokinetics of drugs.
[0379] A pharmacogenetic recommendations database 335 includes existing pharmacogenetic recommendations from the CPIC and DPWG consortia for specific drugs. Recommendations are loaded to the pharmacogenetic recommendations database 335 from the annotation.zip resource of PharmGKB 340. The annotation.zip database 340 is available for downloading at PharmGKB, and facilitates the import of pharmacogenetic recommendations tailored to each drug in accordance with the relevant phenotype.
[0380] An allele combinations database 345 contains all possible combinations of alleles of clinically significant genes. Each allele combination corresponds to a specific recommendation for a specific drug. Allele combinations are loaded from the genes.zip resource of PharmGKB 350. The genes.zip database 350 is available for downloading at PharmGKB, and can be used to import various allele combinations corresponding to the phenotype under consideration.
[0381] A drugs genetic profile database 355 contains all genes for which a statistically significant impact has been demonstrated in studies of any level (in silico, in vitro, in comparative studies, etc.). Data is loaded from the automated annotations.tsv resource of PharmGKB 360. The automated annotations.tsv database 360 is available for downloading at PharmGKB, and supplies specific gene-related information pertinent to each drug.
[0382] The clinical biomarkers subsystem 400 describes clinical biomarkers that are represented by therapy outcomes, including both adverse drug reactions and the absence of therapy effect. Utilizing these outcomes improves the developed generative model, as it takes clinical outcomes into account.
[0383] The clinical biomarkers subsystem 400 includes information in the EHR about the outcomes of therapy 410. The information can be recorded in specialized sections of the EHR. Using LLM models like IRIS for Health (or equivalents if the hospital does not use InterSystems products), the chart and all medical records for a specific patient are analyzed, thus forming avalue of a variable relating to the therapy’s effectiveness, which can later be used to optimize the Gen Al algorithm.
[0384] The EHR therapy outcomes records 410 provide a number and severity of Adverse drug reactions (ADRs) database 420 that contains all records of developed adverse reactions and their severity upon the use of the drug for the corresponding patient with a unique combination of allelic variants. The EHR therapy outcomes records 410 also provide an evaluation of the efficacy of therapy database 430 that contains all records of drug therapy outcomes for the corresponding patient with a unique combination of allelic variants.
[0385] The Omics Biomarkers subsystem 500 provides the next level of improvement over the current development of pharmacogenetics guidelines. The omics biomarkers subsystem 500 includes other molecular biomarkers for personalization. These data are obtained from both public databases and proprietary data acquired from the hospital at earlier stages.
[0386] The analysis system 200 provides several levels of output 600.
[0387] One level of output is a Level 1 ML-Classifier 610. Level 1 610 allows the classification of a patient into one of the three most important clinical categories: Standard Dose, Adjusted Dose, Alternative Drug 615.
[0388] Level 2 Generative Al for guidelines 620 generates implications 625 (forms textual conclusions about deviations in drug metabolism and affinity) and recommendations 630 for optimizing therapy with a specific drug for a specific patient. At this stage, based on these recommendations, new guidelines are created for new drugs, and existing guidelines are updated.
[0389] Level 3 Generative Al for laboratories and hospitals 635 is implemented in two phases. The first phase is based only on molecular biomarkers. The implications 640 and recommendations 645 formed will be directed to the physician at the moment the physician makes a request 110 in the form of a drug name 120. The second phase is further trained on clinical data, including therapy outcomes from pharmacogenetic testing.
[0390] Level 4 Multi-omics Al for Labs and Hospitals 640 includes the remaining molecular biomarkers, which will be added iteratively to the latest version of the model. This iterative process allows for further improvements in the implications 645 and recommendations 650 for increased effectiveness and safety of drug therapy.
[0391] Description of dataset
[0392] The dataset is extracted from the publicly accessible PharmGKB database, encompassing several distinct files:
[0393] • chemicals, tsv provides the nomenclature of drugs for which genetic data were available, alongside the Simplified Molecular Input Line Entry System (SMILES) representation of the drugs.
[0394] • automated annotations, tsv supplies specific gene-related information pertinent to each drug.
[0395] • genes.zip can be used as import of various allele combinations corresponding to the phenotype under consideration.
[0396] • annotation.zip facilitates the import of pharmacogenetic recommendations tailored to each drug in accordance with the relevant phenotype.
[0397] Data preprocessing was conducted to reduce the degrees of freedom within the class labels to three discrete categories: Standard Dose, Adjusted Dose, and Use Alternative Drug.
[0398] The final structural composition of the database encompassed 7,771 samples with 1,799 features. Each sample constituted a test Morgan Fingerprint of the drug, all genes linked with any statistically significant data concerning the efficacy and safety of the drug therapy, as well as combinations of alleles for which a recommendation had been formulated.
[0399] Dividing data into training and test sets
[0400] During the course of the investigation, the original dataset was partitioned into training and testing subsets to ensure reproducibility and stability of experimental outcomes. A fixed seed (set.seed(42)) was employed to maintain consistency across random number generation.
[0401] The methodology for data partitioning included:
[0402] • Utilization of a random selection function to index the original dataset, allocating 80% of the entries to the training subset.
[0403] • The residual indices were automatically designated to the testing subset.
[0404] • This bifurcation was executed without stratification, which could lead to disparities in class distribution between the training and testing subsets. Nevertheless, it was anticipated that the voluminous nature of the data would ensure adequate representation of each class.
[0405] Model Training
[0406] Subsequent to data partitioning, the training of the models was carried out via the gradient boosting method (XGBoost), with the objective of training the model to predict pharmacogenetic recommendations accurately.
[0407] Model Quality Assessment
[0408] The evaluation of the models’ quality employed multiple methods, including:
[0409] • A confusion matrix to assess the classification accuracy, sensitivity, and specificity of the model.
[0410] • Receiver Operating Characteristic (ROC) and Precision-Recall curves to appraise the model’s capacity to distinguish between classes and to balance precision with recall.
[0411] • Feature importance visualization to identify the salient factors influencing classification outcomes.
[0412] Working with missing values
[0413] To ensure the reliability and quality of the predictive model, it was decided to exclude 55 rows with missing values from the analysis (all missing values were in the value and Genotype vector), since they could distort the training and prediction results. Thus, all missing values were excluded from the dataset before starting the model training process, which allowed for higher forecasting quality.
[0414] Software and Tools
[0415] The study was conducted using R software v 4.2.2 (2022-10-31) with the IDE R Studio v 2023.06.2. The following packages and libraries were used to analyze data and train models: ROCR (1.0-11), PRROC (1.3.1), pROC (1.18.5), xgboost (1.7.7.1), randomForest (4.7-1.1), jsonlite (1.8.8), rvest (1.0.3), tensorflow (2.14.0), scales (1.3.0), viridis (0.6.4), viridisLite (0.4.2), caret (6.0-94), lattice (0.22-5), openxlsx (4.2.5.2), lubridate (1.9.3), forcats (1.0.0), stringr (1.5.1), purrr (1.0.2), readr (2.1.4), tidyr (1.3.0), tibble (3.2.1), tidyverse (2.0.0), reshape2 (1.4.4), ggrepel (0.9.4), ggplot2 (3.4.4), keras (2.13.0), writexl (1.4.2), dplyr (1.1.4), and readxl (1.4.3).
[0416] All calculations were performed on a computer running the MacOS operating system. Ventura 13.2.1 (22D68), Ml Pro processor and 32 GB RAM.
[0417] Data preparation
[0418] To ensure reproducibility of the study results, a starting point for random number generation was set (set. seed (42)). The source data was divided into training (80% of the total) and test (20% of the total) sets using random sampling.
[0419] The distribution of the labels
[0420] To gain a more comprehensive understanding of the outcomes and the context of the accuracy metrics obtained, it is crucial to analyze the distribution of labels within our dataset. After preprocessing the raw data, we have classified the pharmacogenetic recommendations into 3 classes: “Standard Dose”, “Adjusted Dose” and “Use Alternative Drug”. This distribution is key to providing a backdrop for a more informed comprehension and assessment of the model’s accuracy.
[0421] XGBoost Model Architecture
[0422] An XGBoost model was applied for the analysis. The key features of the XGBoost methodology used in the study are as follows:
[0423] • Objective set to ‘ multi :softprob’ for multi-class classification which outputs a probability matrix for all classes.
[0424] • Number of boosting rounds was set to 100.
[0425] • Max depth of trees was limited to 6 to control over-fitting.
[0426] • Learning rate (eta) was set to 0.3 to provide a balance between speed and accuracy.
[0427] • Subsample ratio of the training instances was set to 0.8 to prevent overfitting.
[0428] • Colsample bytree, which is the subsample ratio of columns when constructing each tree, was set to 0.8 to manage feature sampling.
[0429] Model Training
[0430] The XGBoost model was trained using the aforementioned parameters. The training process involved:
[0431] • Using a tree-based boosting ensemble method, which is robust to overfitting and commonly provides high performance.
[0432] • Applying the early stopping method with a patience of 10 rounds to prevent unnecessary computations and overfitting.
[0433] • Implementing a stratified k-fold cross-validation with k=5 to ensure the model’s robustness and generalizability.
[0434] • Monitoring the evaluation metric ‘mlogloss’, which is suitable for multi-class classification problems, to guide the training process.
[0435] This XGBoost model, with its ensemble learning approach, is well-suited for the highdimensional and complex nature of pharmacogenetic data, potentially resulting in a highly accurate model.
[0436] Model Evaluation
[0437] Upon the completion of model training, the performance of our XGBoost model was assessed using a separate test dataset. The evaluation focused on generating class predictions to interpret the model’s ability to generalize to new data.
[0438] Model Performance Metrics
[0439] Confusion Matrix
[0440] A confusion matrix was generated to understand the model’s prediction accuracy for each class against the actual labels. This matrix provides a clear visualization of the performance and potential misclassifications made by the model.
[0441] ROC Curve Analysis
[0442] To evaluate the discriminatory capacity of the model for our multi-class classification task, ROC curves were constructed using the one-vs-rest (OvR) methodology. This approach involves treating each class as the positive class against all other classes combined as thenegative class. By doing so, a set of ROC curves were produced, one for each class, providing insights into the model’s ability to distinguish each class from the rest.
[0443] Feature Importance
[0444] The XGBoost algorithm provides a built-in method to evaluate the importance of features. This analysis was conducted to identify and rank the features based on their contribution to the model’s performance, offering insights into which variables had the most significant impact on predictions.
[0445] Precision-Recall (PR) Curve Analysis
[0446] Given the potential imbalance in class distribution, Precision-Recall curves were utilized for a more informative performance evaluation. The one-vs-rest scheme was also applied here, facilitating the assessment of the model’s precision and recall for each class individually.
[0447] Cross-validation
[0448] To ensure the robustness of the model’s performance metrics, a 5-fold cross-validation was conducted. This method provided a more reliable estimate of the model’s accuracy by evaluating it across different subsets of the dataset.
[0449] The metrics derived from these evaluations collectively offer a comprehensive assessment of the model’s predictive performance, ensuring that there is a robust and reliable pharmacogenetic recommendation system.
[0450] The following sections illustrate the implementation and results of these evaluation strategies using R and the XGBoost model.
[0451] Results
[0452] In the pharmacogenetic study, an XGBoost system designed to classify pharmacogenetic recommendation data was developed. The core objective of this system was to discern the complex relationships between genetic markers (RS values) and their implications for drug effectiveness and safety, thereby facilitating personalized treatment recommendations.
[0453] Overall Model Evaluation
[0454] Upon evaluation, our XGBoost model achieved an accuracy of 89.15%, indicating its robust ability to make the correct recommendations in nearly 9 out of 10 instances. This level of accuracy reflects the model’s precision and underscores its potential utility in clinical settings.
[0455] Sample Distribution Across Classes
[0456] The dataset used for training and testing the model comprised a total of 7,771 samples, distributed among three classes as follows:
[0457] • Class 1 : 2,738 samples
[0458] • Class 2: 2,979 samples
[0459] • Class 3: 2,054 samples
[0460] This distribution highlights the model’s need to handle a balanced dataset, where each class is represented with a substantial number of samples, ensuring that the model’s predictive performance is not biased toward a particular class.
[0461] Detailed Performance Analysis Using the Confusion Matrix
[0462] A confusion matrix was employed to gain deeper insight into the model’s classification accuracy across the individual classes. The matrix shed light on the sensitivity, specificity, positive predictive value (precision), and negative predictive value for each class:
[0463] • Class 0 demonstrated a sensitivity of approximately 89.44% and a specificity of around 90.94%, signifying a strong ability to correctly identify true positives and true negatives.
[0464] • Class 1 showed a sensitivity of about 80.13%, indicating a slightly lower yet still significant ability to detect true positives.
[0465] • Class 2 exhibited a sensitivity of 86.83%, reinforcing the model’s effective detection capability across various classes.
[0466] The Fl scores for each class were 86.83%, 80.54%, and 89.79% for classes 0, 1, and 2, respectively, indicating a balanced harmonic mean of precision and recall across the classes (FIG. 2)
[0467] In sum, the XGBoost model demonstrates strong performance and reliability in the task of classifying pharmacogenetic data. The model’s effectiveness is underscored by its high classification accuracy and the balanced distribution of samples across the classes. The predictive capabilities and the detailed performance metrics illustrate that the model is a promising tool for advancing personalized medicine.
[0468] ROC-analysis
[0469] To evaluate the model’s predictive quality across the different classes in the dataset, ROC curve analysis was performed, a technique widely used in binary classification that has been adapted for multi-class scenario through a one-vs-rest (OvR) approach. Each class was independently considered as the “positive” class against all others, allowing us to generate separate ROC curves and calculate the Area Under the ROC Curve (AUC) for each.
[0470] The individual AUC values for the classes were exceptionally high, with Class 1 achieving an AUC of 0.99, Class 2 with an AUC of 0.96, and Class 3 also reaching an AUC of 0.99. The mean AUC across all classes was 0.98, and the standard deviation was a minimal 0.02, indicating not only excellent model performance but also consistency across classes.
[0471] This high mean AUC, along with the low standard deviation, emphasizes the model’s superior diagnostic accuracy and its ability to differentiate between classes effectively. Such performance is crucial in the context of pharmacogenetics, where the correct classification of genetic data can significantly impact clinical decisions.
[0472] FIG. 3 showcases the ROC curve for one of the classes, reflecting an AUC of 0.99, which visually confirms the model’s near-perfect classification capability for this class. This remarkable level of efficacy in classification demonstrates the potential of the XGBoost model for further refinement and adoption in clinical settings.
[0473] Referring to FIG. 4, the precision-recall curve analysis has provided us with insightful metrics for each class, reflecting the model’s ability to maintain high precision over varying recall levels, which is especially beneficial when dealing with the often imbalanced datasets found in pharmacogenetic studies.
[0474] Actual metrics for each class are summarized below.
[0475] • For “Standard Dose” (Class 1), the model achieved a Precision-Recall AUC of 0.99 and an Average Precision score of 0.985.
[0476] • In the case of “Adjusted Dose” (Class 2), the Precision-Recall AUC was 0.96, with an Average Precision score of 0.963.
[0477] • For “Use Alternative Drug” (Class 3), the model recorded a Precision-Recall AUC of 0.99 and an Average Precision score of 0.994.
[0478] These outstanding results demonstrate the model’s exceptional ability to classify each class accurately. The Average Precision scores, being close to 1, highlight the model’s consistent performance across different thresholds, showcasing its precision in predicting true positives against the total predicted positives. The high AUC values across all classes further attest to the robustness of the model’s predictive quality. This level of performance suggests that the model is an excellent tool for pharmacogenetic classification tasks and holds significant promise for practical applications in personalized medicine.
[0479] In the current investigation, an XGBoost model was applied to classify pharmacogenetic data with the aim of optimizing treatment recommendations based on genetic markers. The feature importance analysis, a pivotal aspect of the model’s interpretability, offers insights into which factors most significantly influence the model’s predictions.
[0480] The XGBoost model was trained with a focus on maximizing predictive accuracy while minimizing overfitting, a common pitfail in machine learning applications. The model’s feature importance was assessed, revealing a hierarchy of genetic markers and their contributions to the decision-making process.
[0481] The variable importance analysis disclosed that the top features influencing the model’s decisions included ‘MF_128’, ‘MF_302’, and ‘MF_671’, with respective Gain scores of 0.176, 0.152, and 0.102. These features stood out among the rest, indicating their pivotal role in the classification task. Additionally, the gene ‘CYP3A4’, frequently implicated in drug metabolism, emerged as a significant predictor with a Gain score of 0.095.
[0482] Subsequent features of relevance were ‘MF 23’ and the ‘CYP2C9’ alleles, such as ‘CYP2C9 1’, ‘CYP2C9_9’, and several other variants within this gene family, which are known to affect drug metabolism and response. The Gain scores for these features ranged from 0.037 for ‘CYP2C9 1’ to 0.022 for ‘CYP2C9_45’, highlighting their substantial yet varying impacts on the model’s predictions.
[0483] It is noteworthy that the model’s accuracy and its interpretability through feature importance analysis were robust, even when faced with a diverse set of classes and a balanced dataset. The analysis shed light on the importance of certain genetic variants over others, paving the way for a more nuanced understanding of their influence on drug efficacy and safety.
[0484] FIG. 5 provides a visual representation of the top 20 features according to their importance, as determined by the Gain metric, which quantifies each feature’s contribution to the model’s performance. The visualization underscores the dominance of certain features in the model, reinforcing the need for a nuanced approach to pharmacogenetic data analysis.
[0485] In summary, the feature importance plot not only validates the XGBoost model’s performance but also enhances understanding of the genetic factors that are most influential in determining pharmacogenetic recommendations, which could be invaluable in the context of personalized medicine.
[0486] In this study, the XGBoost gradient boosting algorithm was applied to the classification of pharmacogenetic recommendations, representing a significant step towards the realization of personalized medicine. The system achieved an accuracy of 89.15%, substantially exceeding random prediction, indicating its clinical value.
[0487] The results of the feature importance analysis identified key genetic markers that may interact, affecting drug dosing recommendations. Integrating this interaction data can contribute to creating more accurate pharmacogenetic profiles and improve the personalization of treatment.
[0488] The ROC and Precision-Recall curve analysis underscores the model’s high diagnostic accuracy and its ability to effectively differentiate classes. It is important to discuss how these metrics can be used to improve clinical decision-making, especially in the selection of drug dosages.
[0489] The presented feature importance analysis enhances understanding of the impact of genetic markers on pharmacotherapy and can contribute to improving clinical decisions.
[0490] Machine learning may not only automate the analysis of large datasets, identify patterns, and predict outcomes with a high degree of accuracy for pharmacogenetic development, it can also aid in the discovery of new potential biomarkers that have not yet been manually identified, thereby expanding the boundaries of existing knowledge in pharmacogenetics.Example 2: Harnessing Advanced ML and Al Models to Accelerate Precision Medicine Recommendations Development
[0491] This example provides a system and method for generating precision medicine recommendation for patients.
[0492] We constructed a comprehensive dataset by integrating knowledge bases from PharmGKB, PubChem, and PharmVAR, encompassing drug fingerprints, chemical properties, and extensive pharmacogenetic information. Addressing data imbalance and complexity, we employed gradient boosting classifiers — CatBoost, LightGBM, and XGBoost — with hyperparameter optimization via Optuna. For recommendation generation, we fine-tuned LLaMA 3.1 models with 8-billion and 70-billion parameters, using structured prompts, Rank- Stabilized Low-Rank Adaptation (LoRA), and statistical generation filtering. Model performance was assessed using BLEU and ROUGE metrics.
[0493] The optimized CatBoost classifier achieved an Fl-score of 0.9838, precision of 1.0000, recall of 0.9681, and an ROC-AUC of 0.9991, outperforming previous models in distinguishing cases requiring pharmacogenetic recommendations. The fine-tuned LLaMA models generated high-quality recommendations, with the 70-billion parameter model attaining an average BLEU score of 0.8405 and ROUGE-1 score of 0.8695, indicating strong alignment with expert guidelines. These results demonstrate significant improvements over prior studies in both predictive accuracy and recommendation generation.
[0494] Adverse drug reactions (ADRs) significantly impact global health, ranking among the leading causes of mortality and annually leading to substantial healthcare expenditures. The widespread nature of ADRs not only increases healthcare costs due to a higher number of hospitalizations, but also exacerbates patient morbidity and mortality, underscoring the critical need to improve drug prescription and monitoring strategies. Pharmacogenetics, which studies how genes affect an individual’s response to medications, offers a promising path to address this issue by providing a personalized approach to minimizing the risk of ADRs. This individualized approach can assist in identifying the right drug at the right dose for a specific patient, potentially reducing the prevalence and severity of ADRs. The use of pharmacogenomic (PGx) biomolecular markers for individualizing the therapeutic approach has been shown to reduce the risk of developing ADRs by 30% and increase effectiveness by 1.5 to 2.5 times.
[0495] Research thus emphasizes the significant role of genetic variability in the absorption, distribution, metabolism, and excretion of drugs (ADME), all of which can predispose individuals to ADRs. For example, variations in genes encoding drug-metabolizing enzymes, drug transporters, and drug targets are associated with variable drug responses and side effects. PharmGKB, a leading resource, currently contains more than 200 thoroughly curatedpharmacogenetic guidelines. However, the development of new guidelines is a slow and complex process that requires extensive analysis, expert reviews, and iterations to ensure accuracy, which limits access to the benefits of pharmacogenetic testing.
[0496] Our goal is to use machine learning and Al to accelerate this process. We present a new dataset and pipeline for pharmacogenetic recommendations across various medical fields. We combine drug fingerprints, genetic profiles, and allele combinations of pharmacokinetic and pharmacodynamic proteins into a single dataset with recommendation classes: standard dose, dose adjustment, and alternative drug.
[0497] Methods
[0498] Preprocessing the Database
[0499] To develop our machine learning model, we constructed a comprehensive dataset by integrating knowledge bases from PharmGKB, PubChem, and PharmVAR. Specifically, we extracted the names of 2,152 substances with available pharmacogenetic information from the chemicals. tsv file of PharmGKB. Utilizing data from PubChem, we obtained the SMILES (Simplified Molecular Input Line Entry System) representations of these substances and subsequently generated their Morgan fingerprints. Additionally, we acquired detailed chemical and physical characteristics of the drugs from PubChem. Information on genes and allelic variants associated with these drugs was sourced from PharmGKB. Furthermore, recommendation texts from the Clinical Pharmacogenetics Implementation Consortium (CPIC) and the Dutch Pharmacogenetics Working Group (DPWG) were retrieved via PharmGKB databases.
[0500] When preparing data for training the neural network, we encountered technical challenges caused by the large volume of data. Initially, our dataset included 46,267 entries, of which 44,550 were genetic allele variants. Each allele variant could have two possible states: 1 (affects the recommendation) and 0 (does not affect the recommendation). This meant that each of the 229,825 recommendations depended on 44,550 binary values recorded in separate columns for each allele variant. A server with 1024 GB of RAM was used to process this data volume, leading to the execution of computationally intensive tasks and efficient handling of large datasets with high performance.
[0501] To address this issue without losing critical information, we transformed the data by assigning a unique numerical identifier to each allele variant. For a gene with 17,000 allele variants, each variant was assigned a number from 1 to 17,001, while 0 was reserved to indicate that no variant affected the recommendation. This approach reduced the number of columns from 44,550 to just 17, one for each gene. Instead of numerous columns filled with Is and 0s, we used a single column for each gene, where the unique number of the influencing allele variantwas recorded. For example, if the allele variant CYP2D6 *2 / *36 was assigned the number “20,” we recorded “20” in the CYP2D6 column to indicate its effect on the recommendation. If no variant influenced the recommendation, we recorded “0.”
[0502] This optimization allowed us to maintain the logical structure of the data while overcoming the technical limitations associated with data redundancy. The transformation significantly reduced resource requirements, leading to efficient neural network training that would have otherwise been impossible or highly problematic.
[0503] These features include the drugs’ Morgan fingerprints — a 1,024-dimensional vector representing the molecular structure — as well as all genes linked to statistically significant data concerning the efficacy and safety of drug therapy, and combinations of alleles for which pharmacogenetic recommendations have been formulated. The dataset was systematically divided into five components: chemical-physicochemical information, pharmacodynamic gene information (gen_information_pd), pharmacokinetic gene information (gen_information_pk), fingerprints, and allele information.
[0504] As illustrated in FIG. 6, the class distribution has 61.72% “No Recommendation” samples and 38.28% “Recommendation” samples. For model evaluation, we reserved 20% of the dataset as a test set. The remaining 80% was further partitioned into an 80 / 20 split for training and validation purposes. During model development, models were trained on the training set, with hyperparameter tuning conducted using the validation set. The final evaluation of the model’s performance was carried out on the held-out test set to assess its generalization capability.
[0505] Model Development
[0506] To develop our machine learning model, we deployed a JupyterHub server on the DigitalOcean platform, equipped with an NVIDIA Hl 00 GPU featuring 80 GB of memory. This configuration was chosen to ensure high computational performance and operational efficiency within a collaborative, multi-user environment.
[0507] Computational Environment
[0508] Python 3.12.6 was utilized as the primary programming language due to its extensive libraries and tools suitable for data analysis and machine learning tasks. JupyterHub (version 5.2.1) managed multi-user instances of Jupyter Notebooks, facilitating real-time collaborative work among researchers and providing access to computational resources and analytical tools. Essential components of the Jupyter environment included IPython (version 8.29.0) for interactive coding and data visualization, ipykemel (6.29.5) for Python support in Jupyter Notebooks, and core packages such as jupyter client (8.6.3), jupyter_core (5.7.2), and Traitlets (5.14.3). These packages ensured seamless operation and modularity of the Jupyter system.
[0509] The NVIDIA Hl 00 GPU was critical for executing computationally intensive machine learning and deep learning tasks. Its high memory bandwidth and numerous tensor cores led to efficient processing of large datasets and accelerated model training. Deploying the server on the DigitalOcean cloud platform provided the flexibility to scale computational resources as needed, an essential feature when working with large datasets and complex models (FIG. 6).
[0510] Addressing Class Imbalance
[0511] During data preparation for training the machine learning model, we encountered a significant issue of class imbalance — a common and critical challenge in data analysis and natural language processing. Specifically, 61.72% of the observations belonged to a single, low- informative class labeled “No recommendation.” This imbalance could lead to substantial bias in the model’s predictions, adversely affecting its overall performance and generalization capabilities.
[0512] To address this problem and minimize the risk of overfitting, we partitioned the dataset into two primary classes based on several key considerations:
[0513] Eliminating Imbalance: Class imbalance can cause the model to overfit the dominant class and fail to learn to recognize less-represented classes. By reclassifying the data into two more balanced classes, we facilitated more accurate model training.
[0514] Improving Generalization Ability: A model trained on imbalanced data may demonstrate high accuracy for the dominant class but perform poorly on new, unseen data. Dividing the dataset into two classes mitigated this effect, enhancing the model’s generalization capabilities and increasing its resilience to overfitting.
[0515] Leveraging Boosting Methods: In the initial phase of our experiments, we employed boosting algorithms known for their effectiveness in handling class imbalance. Boosting combines multiple weak learners into a strong predictive model, contributing to improved prediction accuracy and reduced likelihood of overfitting.
[0516] Additionally, the “No recommendation” class was inherently less informative, encompassing instances where no recommendation was necessary, unlike the “Recommendation” class, which contained actionable advice of varying lengths. This lack of informativeness hindered the model’s ability to generalize conclusions drawn from the “No recommendation” class.
[0517] Therefore, we incorporated a primary binary classifier into the model architecture to distinguish between cases requiring recommendations and those not requiring them. This delineation allowed the language model to focus exclusively on generating recommendations without processing instances where the correct action would be to provide none.
[0518] It is imperative to recognize the fundamental differences between the classes, which extend beyond the challenge of class imbalance. The “No recommendation” class is comprised exclusively of samples where the recommendation value is designated as “No recommendation,” leading to a literal interpretation of the text. This characteristic significantly diminishes the informative content of this class for the model, as evidenced by a mean and median sample length of only 2 words, with an interquartile range (IQR) of 0. In contrast, the “Recommendation” class offers specific guidance and detailed descriptions of actions, exhibiting an average length of 49 words, a median length of 34 words, and an IQR of 43. Consequently, the combination of low informativeness associated with the “No recommendation” class and its predominance within the dataset exacerbates the risk of overfitting, thereby hindering the model’s capacity to generalize effectively. Addressing these disparities is critical for enhancing overall model performance and optimizing learning efficacy.
[0519] Gradient Boosting
[0520] In the development of our machine learning model for the binary classification task of recommendation versus no recommendation, we employed gradient boosting techniques. Recognizing the diverse strengths of different implementations, we initially trained classifiers using three leading gradient boosting libraries: CatBoost (version 1.2.7), LightGBM (version 4.5.0), and XGBoost (version 2.1.3). Each library offers distinct advantages and has demonstrated robust performance in various applications; however, it was not possible to determine a priori which would perform optimally on our specific dataset.
[0521] CatBoost is specifically designed to handle categorical features efficiently and is known for its effectiveness in processing such data. LightGBM is optimized for speed and efficiency with large datasets, utilizing techniques such as histogram -based algorithms to accelerate training. XGBoost is renowned for its scalability and performance, especially in handling sparse data and implementing regularization techniques to prevent overfitting.
[0522] To systematically evaluate the performance of each library on our dataset, we conducted extensive training sessions with parallel hyperparameter optimization facilitated by the Optuna library (version 4.1.0). Optuna is an automated hyperparameter optimization framework that employs advanced search algorithms to effectively identify optimal parameter values.
[0523] For each gradient boosting library, we trained a total of 50 models with varying hyperparameter configurations. Based on predefined performance metrics, we selected the bestperforming model from each library for further analysis. Subsequently, an additional 50 experiments were conducted to fine-tune the hyperparameters of these selected models, leading to the identification of the optimal model within each class (CatBoost, XGBoost, and LightGBM). The final comparative results obtained from these models are presented in FIG. 7.
[0524] Generative Model Development
[0525] In developing the generative model for pharmacogenetic recommendations, we selected the LLaMA 3.1 architecture — a state-of-the-art open-source solution renowned for its advanced capabilities in natural language processing and generation. Both the 8-billion and 70-billion parameter versions of LLaMA 3.1 were fine-tuned and tested to evaluate performance across different model scales. Fine-tuning was conducted using Rank-Stabilized Low-Rank Adaptation (LoRA), which facilitates efficient adaptation of large models while minimizing computational resource usage.
[0526] To enhance the model’s performance, we designed a specialized structured prompt that decomposes tabular data from the dataset into smaller subcategories using custom tokens. This approach was used to lead the model to better understand and generate contextually relevant recommendations based on structured input data. Additionally, we implemented a statistical generation filter to stabilize the model’s final outputs. This filter selects the most consistent response from multiple generated outputs, providing it as the final answer. The mechanism not only mitigates hallucinations in the generated text but also offers insights into the model’s confidence levels during experimentation. It helps identify instances when the model is uncertain (multiple response options), confident (all responses match), or erroneous (no correct responses among generated options). While a detailed analysis of the generation filter has not yet been conducted, future experiments will explore its functionality in greater depth.
[0527] Implementation and Technologies
[0528] The implementation of the generative model involved several key technologies and libraries that facilitated various aspects of training and optimization:
[0529] Accelerate (version 0.34.2): Enhanced training across different computational devices, improving efficiency in distributed environments.
[0530] Bitsandbytes (version 0.44.1): Optimized memory usage in deep learning tasks, allowing for the handling of large models on available hardware.
[0531] Datasets (version 3.0.2): Provided essential tools for loading and processing datasets, leading to efficient data handling.
[0532] Huggingface-hub (version 0.26.2): Utilized for storing and distributing machine learning models, facilitating collaboration and model sharing.
[0533] Matplotlib (version 3.9.2) and Seaborn (version 0.13.2): Employed for visualizing results through graphs and charts, aiding in data analysis and interpretation.
[0534] NLTK (version 3.9.1): Supported natural language processing tasks, including tokenization and text preprocessing.
[0535] NumPy (version 1.26.4): Facilitated scientific computing tasks, such as numerical operations and array manipulations.
[0536] Optuna (version 4.1.0): Managed hyperparameter optimization, offering methods for identifying optimal parameter values to improve model performance.
[0537] Pandas (version 2.2.1): Enhanced data analysis capabilities, providing data structures and functions for manipulating structured data.
[0538] PEFT (version 0.13.2): Allowed for effective fine-tuning of models with parameterefficient techniques, reducing the computational burden.
[0539] Safetensors (version 0.4.5): Handled tensor serialization safely during training processes, ensuring data integrity.
[0540] Scikit-learn (version 1.5.2): Provided tools for machine learning algorithms, such as evaluation metrics and model selection utilities.
[0541] SciPy (version 1.14.1): Supported scientific calculations, offering modules for optimization, integration, and statistical analysis.
[0542] Tokenizers (version 0.20.3): Managed text tokenization tasks, crucial for preparing data for model input.
[0543] Torch (version 2.5.1): Served as the primary deep learning framework for building neural networks and performing computations on GPUs.
[0544] Weights & Biases (wandb, version 0.18.6): Facilitated progress tracking during experiments, leading to effective management of model development processes through result visualization.
[0545] Model Evaluation
[0546] The evaluation of the generative model focused on the alignment of generated recommendations with reference recommendations. We employed natural language processing metrics, specifically BLEU and ROUGE, to quantitatively assess the model’s performance:
[0547] BLEU (Bilingual Evaluation Understudy) : Measures the precision of n-grams in the generated text compared to the reference, indicating similarity in word choice and order.
[0548] ROUGE (Recall-Oriented Understudy for Gisting Evaluation) : Assesses the overlap of n- grams between the generated and reference texts, emphasizing content recall.
[0549] These metrics provided quantitative measures of the model’s ability to generate relevant and contextually appropriate recommendations, contributing to an understanding of its effectiveness in replicating expert advice.
[0550] Our comprehensive approach to developing a generative model for pharmacogenetic recommendations underscores the integration of advanced methodologies and state-of-the-art libraries. By leveraging the capabilities of the LLaMA 3.1 architecture and employingspecialized techniques such as structured prompting and statistical generation filtering, we achieved improved performance and reliability in generating contextually relevant recommendations based on input data. Future work will focus on a detailed analysis of the generation filter and further refinement of the model to enhance its applicability in clinical settings.
[0551] Results
[0552] Data Analysis
[0553] In this study, we aimed to enhance the performance of our gradient boosting classifier by re-optimizing hyperparameters specifically within the CatBoost framework. CatBoost was chosen due to its demonstrated effectiveness in handling categorical variables and its superior accuracy compared to other gradient boosting implementations.
[0554] Hyperparameter optimization is a critical component in machine learning model development, as it directly influences prediction quality and the model’s ability to generalize from training data while minimizing overfitting. We meticulously adjusted key hyperparameters, including the number of iterations, learning rate, and tree depth. These parameters significantly affect the model’s performance: for example, increasing tree depth can lead the model to capture more complex relationships within the data but may also increase the risk of overfitting if not properly controlled.
[0555] To further evaluate model performance, we conducted comparative tests across all finetuned LLaMA models, including both the 8-billion and 70-billion parameter versions, to identify which configuration yielded optimal results in text generation tasks. The LLaMA architecture, a powerful transformer-based model, is well-suited for natural language processing applications. We assessed model performance using BLEU and ROUGE metrics, which are standard measures for evaluating the quality of machine-generated text.
[0556] The BLEU (Bilingual Evaluation Understudy) score quantifies the similarity between generated text and reference text based on n-gram overlaps, providing insights into the fluency and accuracy of the generated content. A high BLEU score indicates that the model produces text closely resembling the reference material. ROUGE (Recall-Oriented Understudy for Gisting Evaluation) emphasizes recall by measuring the overlap of n-grams, word sequences, and word pairs between the generated and reference texts. This metric is particularly relevant in scenarios where capturing essential information is critical.
[0557] Our optimized CatBoost model achieved impressive performance metrics: an Fl -score of 0.9838, precision of 1.0000, recall of 0.9681, and an area under the receiver operating characteristic curve (ROC-AUC) of 0.9991. The mean absolute error (MAE) and mean squared error (MSE) were both 0.0130, indicating a high degree of accuracy in the model’s predictions.These metrics demonstrate high efficacy in classification tasks, with the Fl -score serving as a harmonic mean between precision and recall — particularly useful in contexts characterized by class imbalance.
[0558] Overall Frequency Distribution
[0559] The data distribution reveals a significant dominance of the “No Recommendation” class, which constitutes 61.72% of all records, while the “Recommendation” class accounts for 38.28%. This confirms the presence of a strong class imbalance in the initial dataset, which could potentially affect the performance of machine learning models.
[0560] The density plot of the overall distribution demonstrates that most features within the “No Recommendation” class are clustered with low variability. This low variance limits the model’s ability to learn meaningful patterns, as the data may lack sufficient informative value for effective training. Such a concentration of data can lead to a model that is biased toward the majority class, potentially reducing its generalization capabilities.
[0561] Data Frequency Distribution
[0562] To gain deeper insights into the data, we constructed a distribution plot excluding the “No Recommendation” class. This approach allowed us to identify key characteristics within the more informative “Recommendation” class. By removing the dominant class, the data exhibited greater variability, which facilitates the model’s ability to learn complex relationships between features and the target variable. The increased variance in the “Recommendation” class aids in capturing nuanced patterns that could improve predictive performance (FIG. 8 and FIG. 9).
[0563] These observations underscore the importance of addressing class imbalance in our dataset. By analyzing the data both with and without the “No Recommendation” class, we demonstrated that the imbalance could potentially hinder the model’s learning process. The exclusion of the majority class in the analysis provided valuable insights into the distribution of the more informative class, leading to refine our modeling approach to improve performance.
[0564] Gradient Boosting Model
[0565] Feature Importance
[0566] The gradient boosting classifier, specifically optimized using CatBoost, highlighted the relative importance of features contributing to prediction accuracy. Table 3 summarizes the top 20 features ranked by their importance scores.
[0567] Table 3. Top Features Influencing Model Predictions
[0568] The feature CYP2D6 Combinations exhibited a notably high importance score, indicating its significant impact on the model’s predictive capabilities. This aligns with the well- established role of the CYP2D6 gene in drug metabolism and pharmacogenetics. Other features such as CDCA3 , USP5 , and various molecular fingerprints (e.g., MF 923 , MF_677 ) also contributed meaningfully to the model’s predictions, highlighting the multifactorial nature of pharmacogenetic interactions (FIG. 10).
[0569] Model Performance Metrics
[0570] The optimized CatBoost model demonstrated excellent classification performance. Key evaluation metrics are as follows:
[0571] Fl-Score: 0.9838
[0572] Precision: 1.0000
[0573] Recall: 0.9681
[0574] Area Under the ROC Curve (ROC-AUC): 0.9991
[0575] Mean Absolute Error (MAE): 0.0130
[0576] Mean Squared Error (MSE): 0.0130 The high Fl-score reflects a strong balance between precision and recall, indicating that the model effectively identifies true positive cases while minimizing false positives and false negatives. A precision of 1.0000 signifies that all instances predicted as positive are indeed positive, showcasing the model’s accuracy in identifying cases requiring recommendations. The recall of 0.9681 demonstrates the model’s ability to capture a high proportion of actual positive cases.
[0577] The ROC-AUC score of 0.9991 illustrates the model’s exceptional ability to distinguish between the “Recommendation” and “No Recommendation” classes across all threshold settings. The low MAE and MSE values further confirm the model’s accuracy in predictions, indicating minimal deviation from the true values (FIG. 11 and FIG. 12)
[0578] Generative Model Performance
[0579] We evaluated the text generation quality of the fine-tuned LLaMA models using BLEU and ROUGE metrics across different configurations to assess the impact of model size and training data volume on performance.
[0580] Overall, our results demonstrate that careful hyperparameter tuning and model selection can lead to significant improvements in both classification and text generation tasks within pharmacogenetics. The optimized CatBoost classifier effectively distinguishes between cases requiring pharmacogenetic recommendations and those that do not. Concurrently, the fine-tuned LLaMA models generate high-quality, contextually relevant recommendations, as evidenced by strong BLEU and ROUGE scores. These advancements underscore the potential of integrating advanced machine learning techniques in the development of personalized medicine tools (FIGS. 13-15)
[0581] These findings highlight the most influential features driving the predictions of our gradient boosting model. Specifically, CYP2D6 Combinations emerged as the most significant feature, suggesting a strong impact on the target variable. This result aligns with the well- established role of the CYP2D6 gene in drug metabolism and its importance in pharmacogenetics.
[0582] Furthermore, we evaluated the text generation quality across various configurations of the LLaMA models to determine the effects of model size and dataset volume on performance. The results are summarized in Table 4.
[0583] Table 4. Performance Metrics for Different LLaMA Model Configurations
[0584] These results illustrate a clear progression in text generation quality as both data volume and model complexity increase. The 70-billion parameter model trained on 100,000 samples achieved the highest performance, with an average BLEU score of 0.8405 and an average ROUGE-1 score of 0.8695. In contrast, the 8-billion parameter model trained on 2,000 samples over five epochs showed the lowest performance metrics, indicating that both larger model capacity and more extensive training data contribute significantly to better text generation quality.
[0585] In summary, this research successfully enhanced both the gradient boosting classifier using CatBoost and the generative models based on LLaMA through meticulous hyperparameter tuning and rigorous evaluation. By employing established metrics such as BLEU and ROUGE for assessing text generation quality, along with classification performance metrics for evaluating predictive accuracy, we validated the effectiveness of our models. These advancements contribute significantly to improving prediction accuracy and generating high- quality text outputs, which are essential for further research and practical applications in pharmacogenetics and machine learning within this domain.
[0586] Discussion
[0587] The primary objective of this study was to harness machine learning and generative artificial intelligence (Al) to accelerate the development of pharmacogenetic recommendations, thereby addressing the critical need for personalized medicine in reducing adverse drug reactions (ADRs). Our approach involved creating a comprehensive dataset that integrated drug molecular structures, genetic profiles, and allele combinations, coupled with implementing advanced machine learning models for classification and text generation tasks. The results demonstrate significant advancements in both accurately predicting when pharmacogenetic recommendations are necessary and generating contextually relevant recommendations, highlighting the potential of Al in transforming pharmacogenetics (FIG. 16).
[0588] Interpretation of Results
[0589] Our optimized CatBoost classifier achieved exceptional performance metrics, with an Flscore of 0.9838, precision of 1.0000, recall of 0.9681, and an ROC-AUC of 0.9991. These results indicate that the model is highly effective in distinguishing between cases that require pharmacogenetic recommendations and those that do not. The high precision suggests that the classifier is extremely accurate when it predicts a recommendation is needed, while the high recall indicates it successfully identifies the vast majority of such cases. This balance between precision and recall is crucial in clinical settings where both false positives and false negatives can have significant implications for patient safety.
[0590] Feature importance analysis revealed that the CYP2D6_Combinations feature was the most influential in the model’s predictions. This is consistent with existing literature, as CYP2D6 is a well-known enzyme involved in the metabolism of many drugs and is a common focus in pharmacogenetic studies. Other significant features included genes like GNB3 and ADRA2A, which have been implicated in variable drug responses. The inclusion of molecular fingerprints among the top features underscores the importance of drug-specific properties in influencing pharmacogenetic interactions.
[0591] In the generative modeling component, the fine-tuned LLaMA models demonstrated a clear correlation between model size, training data volume, and performance. The 70-billion parameter model trained on 100,000 samples achieved the highest BLEU score of 0.8405 and ROUGE- 1 score of 0.8695, indicating a high degree of similarity between the generated recommendations and the reference texts. This suggests that larger models with more training data can capture the complex relationships between genetic variants and drug responses more effectively, resulting in more accurate and clinically relevant recommendations.
[0592] Improved Handling of Data Imbalance: By effectively addressing class imbalance, our model maintained high performance without the degradation commonly seen in prior studies, such as those by Tao et al. (2020).
[0593] Advanced Generative Capabilities: The utilization of the LLaMA models allowed us to generate high-quality pharmacogenetic recommendations, a feat not extensively explored in earlier research.
[0594] Comprehensive Integration of Data: Our dataset incorporated drug molecular structures, genetic profiles, and allele combinations, providing a more holistic approach compared to studies that focused on limited genetic markers or single-drug responses.
[0595] These advancements suggest that our model not only surpasses the predictive capabilities of previous models but also offers a more robust and scalable solution for the generation of pharmacogenetic recommendations. The integration of advanced machine learning techniques and the handling of complex, high-dimensional data set a new benchmark for future research in pharmacogenetics.
[0596] Significance in the Broader Context
[0597] The integration of Al into pharmacogenetics addresses a critical bottleneck in personalized medicine — the slow and resource-intensive process of developing pharmacogenetic guidelines. Traditional methods require extensive manual curation and expert consensus, limiting the speed at which new recommendations can be formulated and disseminated. By automating the classification and recommendation generation processes, our approach has thepotential to rapidly expand the coverage of pharmacogenetic guidelines, making personalized therapy more accessible.
[0598] Furthermore, the use of advanced Al models like LLaMA represents a significant step forward in natural language processing applications within healthcare. The ability of the generative model to produce high-quality, contextually appropriate recommendations can facilitate the integration of pharmacogenetic insights into clinical decision support systems, electronic health records, and other tools that clinicians use at the point of care.
[0599] Our study demonstrates the feasibility and effectiveness of using advanced machine learning and generative Al models to accelerate the development of pharmacogenetic recommendations. By integrating drug molecular structures, genetic profiles, and allele combinations into a comprehensive dataset — and addressing challenges such as data imbalance — we developed a high-performance CatBoost classifier and fine-tuned LLaMA models. The CatBoost classifier exhibited near-perfect classification metrics, while the 70- billion parameter LLaMA model generated high-quality, contextually appropriate recommendations, outperforming previous studies in predictive accuracy and recommendation quality.
[0600] These findings suggest that our approach can significantly enhance the scalability and accessibility of pharmacogenetic guidelines, facilitating the integration of personalized medicine into clinical practice. By providing clinicians with accurate, ALgenerated recommendations, we can reduce the incidence of adverse drug reactions and improve therapeutic efficacy. Future work should focus on expanding the dataset to include a broader range of genetic variants and drugs, optimizing model efficiency for clinical deployment, and collaborating with healthcare professionals to validate and refine the models. This will ensure the reliability, safety, and acceptance of Al-driven pharmacogenetic tools, ultimately enhancing patient care and outcomes.Example 3: Reducing Readmissions Through Personalized Dosing
[0601] This example provides a use case of reducing readmissions through personalized dosing.
[0602] A 65-year-old patient with hypertension and Type 2 diabetes is prescribed a standard beta-blocker dose post-discharge. However, the patient’s genetic polymorphism indicates a risk of poor drug metabolism, causing significant side effects and eventual readmission.
[0603] The patient’s genomic data is accessed in real time. An ML-driven risk assessment is performed to identify potential adverse reactions. The ML proposes an adjusted beta-blocker (or alternative medication) dose, significantly lowering the risk of side effects and reducing readmission likelihood.
[0604] 20-30% fewer readmissions related to medication issues are expected based on preliminary pilot data.Example 4: Streamlining Polypharmacy Management in Older Adults
[0605] This example provides a use case of streamlining polypharmacy management in older adults.
[0606] Elderly patients often take multiple medications (5+), creating high risk for drug-drug and drug-gene interactions. Clinicians have limited bandwidth to assess every possible conflict.
[0607] The patient’s full medication list is aggregated, and a multi-omics check is performed to evaluate potential interactions. The system flags problematic combinations in the EHR and suggests safer alternatives or dose adjustments.
[0608] This use case is expected to reduce ADEs, shorten inpatient stays, and improve efficiency of pharmacist workflows.Example 5: Accelerating Oncology Treatment Personalization
[0609] This example provides a use case of accelerating oncology treatment personalization.
[0610] Oncology patients frequently undergo highly toxic treatments, and oncologists seek biomarkers or genetic variants to fine-tune therapy.
[0611] Next-generation sequencing (NGS) results are integrated with real-world evidence from similar patients.
[0612] The system predicts which chemotherapy regimen or immunotherapy is most likely to yield positive outcomes, while minimizing severe side effects. The system is expected to improve treatment success rates and a more precise approach to complex oncology cases.Example 6: Multi-Omic Scaling Platform for Drug Response Prediction
[0613] This example provides a system with multi-omics capabilities, scaling real-world evidence (RWE) integration, federated learning, and broadening clinical adoption.
[0614] Data Strategy
[0615] Advanced LLM-based systems (Llama-2, GPT-like models) is integrated for real-time data extraction from new research, clinical guidelines, and literature. “Explainable Al” features are implemented to clarify model recommendations and build clinician trust. Datasets are expanded to cover 1,700+ substances, 700+ genes, and multi-omic data. Global RWE integration and multi-omics analysis will be used to expand the data.
[0616] Federated Learning is used to permit Al models to be trained at different locations or devices without exchanging data between the devices. Federated learning allows differentinstitutions to leverage their own data in combination with a central Al system. For example, hospitals can train Al models on local data without sharing private health information externally.
[0617] Proteomics, transcriptomics, and metabolomics data pipelines are integrated into the Al system to enrich patient profiles. Advanced filtering tools are leveraged to improve Al-driven biomarker discovery. Patient stratification and trial optimization recommendations are generated using the Al system to reduce ADRs, and improve efficiency of clinical trials.
[0618] Scalability Plan
[0619] Overview
[0620] The Al system has containerized microservices. Each key function (data ingestion, Al inference, report generation) runs as an independent container (Docker), leading to on-demand scaling based on load.
[0621] The Al system has cloud-agnostic deployment. The system supports both GCP and Azure, offering built-in auto-scaling capabilities and geographic redundancy to minimize latency for global users.
[0622] The Al system has InterSy stems IRIS Clusters. IRIS for Health’s clustering features are leveraged to handle spikes in data ingestion and maintain high availability, even under heavy patient data volumes.
[0623] The Al system has parallel training & inference. Python multiprocessing, GPU acceleration, and container-based deployments allow simultaneous training of new models while serving real-time clinical requests.
[0624] The Al system has federated learning rollout. By shifting model training to local hospital nodes, central server load is reduced while respecting data privacy regulations. Over time, this approach leads to near-linear scaling of Al development efforts across multiple sites.
[0625] The Al system has The LLM-based components that can run on dedicated GPU clusters, ensuring that text / data extraction does not bottleneck the rest of the platform.
[0626] The Al system has deploys advanced load-balancing solutions for Al inference (e.g., GPU scheduling).
[0627] The Al system has operational resilience. Distributed container architecture minimizes single points of failure and allows swift recovery from hardware or network outages.
[0628] The Al system can grow flexibly. A cloud-native approach can be used to rapidly adjust capacity for short-term spikes (e.g., large batch genetic data processing) or long-term user increases.
[0629] The Al system combines containerized microservices, federated learning, global cloud architecture, and proactive resource monitoring. This approach allows the Al system to adapt torising demand — from small clinics to large hospital systems — while ensuring high performance, robust security, and compliance with healthcare regulations.Example 7: Classification of Serviceable and Non-Serviceable Feature Vectors
[0630] This provides an example for a machine learning model for classifying serviceable and non-serviceable feature vectors.
[0631] Each feature vector undergoes a multi-step evaluation to determine whether the system can confidently yield a clinically effective recommendation. Specifically, the model computes a confidence threshold based on multiple factors, including (i) structural or molecular similarity to known compounds, (ii) known or inferred pharmacokinetic and pharmacodynamic data, (iii) genetic or multi-omic correlates (e.g., genomic features, epigenetic markers, transcriptomic signatures), and (iv) the availability and relevance of clinical outcome data.
[0632] If the model’s confidence score surpasses a predetermined threshold - indicating that sufficient reliable information exists to generate a robust drug response prediction - the feature vector is categorized as “serviceable.” This designation applies even if the pharmacological substance or genomic profile in question does not explicitly appear in the existing structured pharmacogenomic dataset, provided the model identifies enough analogous data (e.g., similar chemical structures or comparable genomic variations in other subjects).
[0633] Conversely, if the aggregated evidence or confidence score falls below the threshold - due to high uncertainty, paucity of relevant data, or irreconcilable inconsistencies - the feature vector is deemed “non-serviceable,” and no recommendation is produced. This ensures that the system only provides actionable clinical guidance when it possesses a justifiable degree of certainty, thereby distinguishing between scenarios where meaningful predictions can be made and those where insufficient data precludes a reliable recommendation.Example 8: Evaluation of Model Performance and Comparison with Comparators
[0634] This example provides an evaluation of the model performance in generating recommendation clinical outcomes and comparison against comparator performance.
[0635] The trained model was used to provide recommendations for 15 test clinical test cases which were evaluated against two comparator models, o4 and DeepSeek-Researcher, and expert recommendations provided by physicians from electronic health records in similar cases.Comparisons between the models were evaluated based on metrics including accuracy of match, completeness of recommendations, and adherence to clinical protocols, allowing for objective assessment of model outputs and alignment with decisions made by experienced specialists. The performance of the model on reflecting the semantic and contextual similarity to realrecommendations was assessed and scored by clinical experts according to a 0-2 scale according to the criteria below:
[0636] • 0 - dangerous recommendation
[0637] • 1 - ineffective recommendation
[0638] • 2 - clinically appropriate
[0639] The final score for each model was calculated as the average expert rating across all cases and normalized by dividing by two. Table 5 summarizes the clinical recommendation and expert evaluation for o4. Table 6 summarizes the clinical recommendation and expert evaluation for Deep Seek-Researcher. Table 7 summarizes the clinical recommendation and expert evaluation for a PGx model of the disclosure. Table 8 lists the normalized score for each of the evaluated models.
[0640] Table 5. o4 Recommendations and Evaluation
[0641] Table 6. DeepSeek-Research Recommendations and Evaluation
[0642] Table 7. Model Recommendations and Evaluation
[0643] Table 8. Model Performance
[0644] The model as disclosed herein significantly outperformed the two comparator models across the evaluated conditions. Experts assigned the model a mean score of 0.93, indicating strong alignment with clinical validity, clarity, and adherence to medical guidelines. This suggests the model of the present disclosure can provide the most reliable and actionable pharmacogenomic guidance among the evaluated models.
[0645] Benchmark Dataset Construction
[0646] To systematically evaluate the performance of our model in generating guideline-based pharmacogenomic recommendations, we first constructed a benchmark dataset having 229,826 examples.
[0647] Each example in the dataset included: a drug; a gene known to influence the drug response; a genotype representing a specific allelic combination; and a recommendation derived from established clinical guidelines.
[0648] The dataset was designed to cover all possible combinations of drugs, genes, and genotypes that have a pharmacogenomic impact, ensuring a comprehensive evaluation of model performance. The recommendations were sourced primarily from PharmGKB, along with CPIC guidelines and other authoritative pharmacogenomic databases.
[0649] To ensure robustness, the dataset included: cases with actionable recommendations, where genotype impacts drug selection or dosage; and cases without recommendations, where no genetic-based adjustment is required.
[0650] This benchmark serves as a gold standard for model evaluation, allowing us to measure how well different models align with established pharmacogenomic guidelines in real-world clinical decision-making.
[0651] Evaluation Metric for Deep Semantic Alignment Assessment
[0652] LLM Score is an evaluation metric developed to assess the deep semantic alignment between original and generated recommendations, particularly in cases where the syntactic structure significantly differs while the core meaning is preserved. The primary objective of this metric, in conjunction with a State-of-the-Art (SOTA) Large Language Model (LLM), is to evaluate the semantic similarity between the original and generated responses.
[0653] This approach is designed to prioritize minor yet critical modifications (e.g., changes in dosage) over more substantial alterations that do not degrade the overall content (e.g., additional recommendations for tests, supplementary explanatory sentences, etc.). By leveraging the capabilities of the model, the metric is able to account for interesting variations of the aforementioned issues, ensuring that the evaluation focuses on semantic similarity rather than superficial syntactic differences.
[0654] The LLM Score operates by utilizing a specialized evaluation prompt that guides the LLM to analyze and compare the semantic content of the original and generated recommendations. This prompt is structured to emphasize the importance of preserving the core meaning, even when the phrasing or structure of the generated text diverges significantly from the original. As a result, the metric is able to identify and prioritize subtle but meaningful changes, such as adjustments in medical dosages or specific procedural instructions, while de-emphasizing less critical additions or rephrasing that do not alter the fundamental intent of the recommendation.
[0655] Furthermore, the use of the model ensures that the metric can handle complex and nuanced variations in language, including idiomatic expressions, contextual dependencies, and domain-specific terminology. This allows the LLM Score to provide a robust and accurate assessment of semantic similarity, even in cases where traditional evaluation methods might fail to capture the true alignment of meaning.
[0656] An example evaluation prompt for assessing the quality of generated recommendations against guideline-based references is given below:Role:You are a deterministic metric evaluating semantic correspondence between medical recommendations. Your output must be repeatable and stable. Do not explain your score.Task:Calculate a numerical score (1-10) reflecting how accurately the Generated Recommendation reflects the Original. Use strict criteria:If recommendations are identical, the score will be 10.If the Original Recommendation states "No recommendation", but the Generated Recommendation includes specific treatment modifications, score 0.If the Original Recommendation states "No recommendation", but the Generated Recommendation assumes a standard regimen of drug use, score 9.Key Elements Matching (4 points)List medications, dosages, actions, and specifics from both texts.Award 1 point per fully matched element (direct or synonym).Deduct 1 point per missing / additional element.Numerical Accuracy (3 points)Compare all numerical values (e.g., dosages).Full points if identical, 0 if conflicting.Context Compliance (2 points)Deduct 1 point for non-standard medical advice.Deduct 1 point for irrelevant additions.Structural Faithfulness ( 1 point)Award if no reordering / merging alters meaning.Rules for Repeatability:Treat synonyms as matches if meaning is identical.Ignore explanatory text unless it introduces new facts.Numerical discrepancies nullify Numerical Accuracy entirely.Output Format:Rating (1-10): [number]Input Data:Original Recommendation:{orig_rec}Generated Recommendation:{generated_rec}
[0657] To validate the effectiveness of this metric, the prompt was applied to 30 randomly selected pharmacogenomic recommendations from the benchmark dataset. The LLM Score was compared against expert human ratings to assess correlation and reliability. Multiple SOTA LLMs and selected Deep Seek-Chat were tested due to their low computational cost, fast inference speed, and strong alignment with expert evaluations. The Mean Absolute Error (MAE) between expert evaluations and LLM Score was 1.13, compared to GPT-4o with MAE=2.13 and ol with MAE=2.97. This made DeepSeek-Chat an optimal choice for automated large-scale evaluation of pharmacogenomic recommendations.
[0658] FIG. 30 shows a confusion matrix comparing LLM Score predictions with expert evaluations. The results indicate strong agreement for high-confidence cases, with minor discrepancies in mid-range scores (4-6), where expert interpretations varied due to contextual nuances. The results demonstrate that LLM Score is a reliable and cost-effective alternative to expert evaluation, providing a scalable alternative to manual annotation without compromising clinical accuracy.
[0659] Comparison of Model 1 Performance
[0660] To evaluate the performance of the models in generating guideline-based pharmacogenomic recommendations, an experiment was conducted to assess the ability of different models to predict pharmacogenomic recommendations in alignment with clinical guidelines. Each model was tasked with generating a recommendation based on a given drug, gene, and genotype, simulating real-world clinical decision-making.
[0661] For evaluating the models, 100 random cases were selected from the benchmark dataset. The generated recommendations were compared against expert-validated guideline outputs. The models were assessed using two key metrics: (1) LLM Score measures semantic alignment between generated recommendations and expert-validated responses; and (2) BLEU Score assesses the textual similarity between model outputs and reference recommendations.
[0662] Inference Speed (measured in seconds per request) was also evaluated to measure the computational efficiency of each model, to evaluate the feasibility of the model for large-scale pharmacogenomic applications.
[0663] Results indicate that the fine-tuned model, deployed on internal servers, outperformed other reference models, achieving an LLM Score of 0.97 and a BLEU Score of 0.99, while maintaining a fast inference speed of 2.3 seconds per request. Unlike cloud-based SOTA models that require API calls with higher latency, the model operated efficiently within a controlled computing environment, reducing response times and ensuring data security.
[0664] Table 9. Performance comparison between the fine-tuned model and other SOTA models using a special prompt
[0665] A prompt such as the following was used to enhance SOTA models capabilities for recommendation generation:You are a pharmacogenomic recommendation system that provides drug dosing and safety recommendations based on genetic information.User input follows the format: [drug] [gene] [diplotype]. The [diplotype] field represents genetic variation, including star alleles (e.g., 1 / 2) or specific SNPs (e.g., c.61OT / Reference).Respond with a clear pharmacogenomic recommendation. If no specific recommendation exists, respond with "No recommendation."Recommendations should be concise but provide additional guidance if clinically relevant (e.g., dose adjustments, alternative drugs, monitoring considerations).#ExamplesUser input: tramadol CYP2D6 1 / 2Recommendation: Avoid tramadol use because of potential for toxicity. If opioid use is warranted, consider a non-codeine opioid.User input: capecitabine DPYD c.61OT / ReferenceRecommendation: Reduce starting dose by 50% followed by titration of dose based on toxicity or therapeutic drug monitoring (if available). Patients with the c.2846A>T / c.2846A>T genotype may require >50% reduction in starting dose.User input: simvastatin SLCO1B1 1 / 1Recommendation: Prescribe desired starting dose and adjust doses based on disease-specific guidelines. Other considerations: Evaluate the potential for drug-drug interactions and adjust based on renal and hepatic function.User input: ivacaftor CFTR ivacaftor-non-responsive Recommendation: Ivacaftor is not recommended.User input: amitriptyline CYP2D6 152 / 37 Recommendation: No recommendation. #Input dataUser input: {metabolizer}Recommendation:Example 9: Personalized Medicine and Pharmacogenomics Use Cases
[0666] This example provides various use cases of the model for data analysis and personalized medicine applications, allowing a user to make personalized treatment decisions based on a patient’s genetic, transcriptomic, and other multi-omics data. Clinical decision support is enables real-time integrations into electronic health records, providing clinicians with immediate recommendations at the point of care.
[0667] Oncology
[0668] The model analyzes tumor mutations, pharmacogenmoc biomarkers, and immunotherapy response predictors to recommend targeted therapies. Non-limiting examples of targeted therapies include: recommendation of EGFR inhibitors for lung cancer, PARP inhibitors for BRCA-mutated breast cancer, or immune checkpoint inhibitors based on tumor mutational burden.
[0669] Autoimmune diseases
[0670] The model assesses genetic susceptibility to drug-induced side effects (e.g., methotrexate toxicity) and can guide biological therapy selection (e.g., TNF inhibitors vs. IL-6 inhibitors) based on factors including but not limited to genetic predisposition for conditions including rheumatoid arthritis, lupus, and multiple sclerosis.
[0671] Obesity and Metabolic disorders
[0672] The model can predict patient response to drugs used in obesity and type 2 diabetes treatment, including but not limited to GLP-1 receptor agonists (e.g., semaglutide, tirzepatide). Patients who may benefit most or may experience adverse effects are identified.
[0673] Psychiatry and Neurology
[0674] Antidepressant and antipsychotic dosing is adjusted based on metabolism of known drugresponse modulating genes, including CYP2D6 and CYP2C19 for drugs including but not limited to selective serotonin-reuptake inhibitors (SSRIs) and tricyclic antidepressants.Treatments for epilepsy and neurodegenerative diseases are personalized based on the patient’s genetic, transcriptomic, and other multi-omics data.
[0675] Cardiology
[0676] Cardiovascular outcomes are predicted and improved based of pharmacogenomic recommendations for therapies including but not limited to antiplatelet therapy (e.g., clopidogrel response based on CYP2C19 status) or dosing with warfarin is adjusted (e.g., based on VKORC1 / CYP2C9 status).Example 10: Data Processing and Model Training
[0677] Data relating a drug and a pharmacogenomic response including the structural information of the drug, physiochemical description of the drug, and list of genes with proven influence on the drug were collected for a large panel of drugs from data sources including PharmGKB, PubChem, PharmVar, and CIPC guidelines. The data were cleaned by removing empty samples, such as those with missing or empty fields in the data. Following cleaning, the data were normalized with the physiochemical descriptions of the drugs normalized; normalizing binary vectors relative to other data was deemed inappropriate.
[0678] Statistical analysis of the data were conducted. Classes within the data were identified based on the drug and type of recommendation, class imbalance was detected. Class imbalance was addressed during training by oversampling rare classes combined with random undersampling of larger classes.
[0679] Recommendations were initially classified into categories including those containing actionable recommendations and those that do not necessitate changes in the treatment, classified as “no recommendation” according to PharmGKB database.
[0680] Following classification, a gradient boosting machine learning model was employed and a baseline LLM model established. To fine-tune the baseline LLM, the model was fine-tuned with the following layers modified using low-rank adaptation (LoRa), illustrated in FIG. 20 and FIG. 21. In the attention layer, non-limiting examples of modification included query projection,key projection, value projection, and output projection. In the feed-forward layer, non-limiting examples of modification included gate projection, up projection, and down projection.
[0681] The models’ tokenizer was modified prior to training. Non-limiting examples of special tokens added to the tokenizer include: “### Chem”, “### GEN_PD”, “### GEN_PK”, “### Fingerprint”, “### Allele”, with the special tokens used to segment the input data into structural groups to prevent incorrect comparison of numerical value between different groups. The modified model was fine-tuned on the balanced dataset described above.
[0682] The model data chain and recommendation can be iteratively refined. Referring to FIG. 21, an illustrative data chain and recommendation diagram reflecting evolving of a recommendation based on clinical outcomes. The model data chain can incorporate two recommendation architectures. Referring to FIG. 22, an illustrative diagram incorporating two recommendation architectures is shown. The first model generates the initial recommendation while the second model modifies and refines the recommendation based on factors including clinical outcomes and relevant extra information.
[0683] The model as described above can be used to provide a drug response prediction. The results are delivered to a user by sending the prediction results from the model server to the service server, which then displays the results on, for example, a user interface. Referring to FIG. 22, an illustrative data flow is shown.Example 11: Graphical User Interface
[0684] The system comprises a user visualization module designed to dynamically display model responses within the user interface. This module facilitates real-time interaction between the user and the Al model, ensuring that the outputs generated by the model are presented in an intuitive and accessible manner. The dynamic visualization capability allows for adaptive rendering of data, enabling users to observe changes, trends, or results as they are computed, thereby enhancing the decision-making process and user experience.
[0685] FIG. 25 shows an example of the graphical user interface, where a pharmacogenomic recommendation is generated based on a selected drug (atorvastatin) and genotype (SLCO1B1 *I / *I).
[0686] FIG. 26 shows another GUI element. A sample workflow demonstrating the system’s ability to generate pharmacogenomic recommendations based on genetic input data. The interface can accept user-defined molecular and genetic information as input and produces a pharmacogenomic recommendation as output.Example 12: Pharmacogenomic System Integrated with Multi-Omics
[0687] This example describes a prophetic workflow for integrating multi-omics with a pharmacogenomic system.
[0688] FIG. 19 shows integration of epigenomic, transcriptomic, proteomic, and metabolomics data into a large language model. The integration of new data into the system can be achieved through a combination of diverse approaches, including Retrieval-Augmented Generation (RAG) for incorporating additional clinical information, intermediate models for processing complex patterns within the data, data transformation pipelines for preprocessing and structuring data, and generative models for formulating final recommendations, among others.
[0689] The RAG framework can enhance the system’s ability to retrieve and integrate relevant clinical data from external sources, ensuring that the recommendations are informed by the most up-to-date and comprehensive information available. Intermediate models, which may include various machine learning and deep learning architectures, can be employed to analyze and interpret intricate relationships and patterns within the data. These models can provide handling for complexity and variability in clinical datasets, enabling the system to make more accurate and nuanced predictions.
[0690] Data transformation pipelines can play a crucial role in preparing the data for analysis. These pipelines can be configured to clean, normalize, and transform raw data into a format that is suitable for further processing by the system’s models. This step can ensure data integrity and that the subsequent analysis is accurate and reliable.
[0691] Generative models, such as those based on advanced neural network architectures, can be utilized to synthesize the processed information and generate actionable recommendations.These models can be configured to produce coherent and contextually appropriate outputs, which is important for providing clinicians with practical and reliable guidance.
[0692] By leveraging these complementary approaches, the system can provide a robust and flexible framework for data integration, capable of adapting to the evolving demands of clinical practice and delivering high-quality, evidence-based recommendations.
Claims
CLAIMSWhat is claimed is:
1. A computer-implemented system comprising: a database comprising a structured pharmacogenomic dataset; a machine learning model comprising a neural network language model; and computer-executable instructions configured to:(a) train the machine learning model, using the structured pharmacogenomic dataset, to:(i) classify a feature vector as (1) a serviceable feature vector or (2) a non- serviceable feature vector for a drug response prediction; and(ii) generate for the serviceable feature vector, using the neural network language model, the drug response prediction in natural language;(b) update the structured pharmacogenomic dataset with a new pharmacogenomic dataset to generate an updated structured pharmacogenomic dataset; and(c) train a new machine learning model, using the updated structured pharmacogenomic dataset, such that the new machine learning model has more serviceable feature vectors than the machine learning model.
2. The computer-implemented system of claim 1, wherein the structured pharmacogenomic dataset comprises structural representations of pharmaceutical substances.
3. The computer-implemented system of claim 1, wherein the structured pharmacogenomic dataset comprises physical or chemical properties of pharmaceutical substances.
4. The computer-implemented system of claim 1, wherein the structured pharmacogenomic dataset comprises genetic profiles of subjects.
5. The computer-implemented system of claim 1, wherein (i) the serviceable feature vector, (ii) the non-serviceable feature vector, or (iii) both is input to the neural network language model as a structured prompt.
6. The computer-implemented system of claim 1, wherein the serviceable feature vector indicates that a clinically effective recommendation of the drug response prediction can be generated for the feature vector.
7. The computer-implemented system of claim 1, wherein the non-serviceable feature vector indicates that a clinically effective recommendation of the drug response prediction cannot be generated for the feature vector.
8. The computer-implemented system of claim 1, wherein the serviceable feature vector defines a genomic feature of a subject that is not found in the structured pharmacogenomic dataset.
9. The computer-implemented system of claim 1, wherein the serviceable feature vector defines a pharmacological substance that is not found in the structured pharmacogenomic dataset.
10. The computer-implemented system of claim 1, wherein the feature vector is classified as the serviceable feature vector or the non-serviceable feature vector based on a confidence score.
11. The computer-implemented system of claim 1, wherein the computer-executable instructions are configured to extract features useful for drug response prediction in the new pharmacogenomic dataset.
12. The computer-implemented system of claim 1, wherein the computer-executable instructions are configured to filter mutations in the new pharmacogenomic dataset.
13. The computer-implemented system of claim 1, wherein the computer-executable instructions are configured to isolate clinically significant markers, from the new pharmacogenomic dataset, that are relevant to adverse drug reactions or efficacy.
14. The computer-implemented system of claim 1, wherein the computer-executable instructions are configured to filter pharmaceutical substances in the new pharmacogenomic dataset.
15. The computer-implemented system of claim 1, wherein the computer-executable instructions are configured to apply dimensionality reduction or feature engineering techniques.
16. The computer-implemented system of claim 1, wherein the computer-executable instructions are configured to integrate the new pharmacogenomic dataset with the structured pharmacogenomic dataset.
17. The computer-implemented system of claim 1, wherein the computer-executable instructions are configured to, using the updated structured pharmacogenomic dataset, generate updated feature vectors for training the new machine learning model.
18. The computer-implemented system of claim 1, wherein the drug response prediction comprises a personalized regimen for administering a pharmaceutical substance for a subject.
19. The computer-implemented system of claim 1, wherein the drug response prediction comprises a personalized dosing guideline for a subject.
20. The computer-implemented system of claim 1, wherein the drug response prediction comprises a personalized drug selection for a subject.
21. The computer-implemented system of claim 1, wherein the drug response prediction comprises a prediction of an adverse drug reaction in a subject.
22. A computer-implemented system comprising: a database comprising a structured dataset, wherein the structured dataset comprises a pharmacogenomic dataset relating at least 1,200 pharmacological substances and at least 350 genes; a machine learning model comprising a neural network language model; and computer-executable instructions configured to train the machine learning model to generate a drug response prediction in natural language, wherein the training is based on the structured dataset.
23. A computer-implemented system comprising: a machine learning model comprising a neural network language model, wherein the machine learning model is trained using a structured dataset comprising a pharmacogenomic dataset relating at least 1,200 pharmacological substances and at least 350 genes; and computer-executable instructions configured to process a feature vector, using the machine learning model, to generate a drug response prediction in natural language.
24. A computer-implemented system comprising: a database comprising a structured dataset, wherein the structured dataset comprises a pharmacogenomic dataset and at least one of: (i) an epigenomic dataset, (ii) a transcriptomic dataset, (iii) a proteomic dataset, (iv) a metabolomic dataset, (v) a lipidomic dataset, and (vi) a secretomic dataset; a machine learning model comprising a neural network language model; and computer-executable instructions configured to: train the machine learning model, using the structured dataset, to:(i) classify a feature vector as (1) a serviceable feature vector or (2) a non- serviceable feature vector for a drug response prediction; and(ii) generate for the serviceable feature vector, using the neural network language model, the drug response prediction in natural language.
25. A computer-implemented system comprising: a database comprising an internal pharmacogenomic dataset encrypted based on an internal cipher; a neural network language model; and computer-executable instructions configured to:(i) create a secure connection with a client to receive an external pharmacogenomic dataset encrypted based on an external cipher;(ii) receive the external pharmacogenomic dataset through the secure connection;(iii) decrypt the external pharmacogenomic dataset based on the external cipher;(iv) add the external pharmacogenomic dataset to the internal pharmacogenomic dataset;(v) decrypt the internal pharmacogenomic dataset based on the internal cipher; and(vi) train the neural network language model, using the internal pharmacogenomic dataset, to generate a drug response prediction in natural language.
26. A computer-implemented system comprising:(a) a neural network language model;(b) a plurality of computers, each computer comprising:(i) a database comprising a pharmacogenomic dataset, wherein the pharmacogenomic dataset of each computer in the plurality of computers is different from one another;(ii) a containerized computer-executable instructions configured to train the neural network language model using the computer’s computational resources and the computer’s pharmacogenomic dataset, without sharing the computer’s pharmacogenomic dataset with another computer in the plurality of computers.
27. A method for predicting a patient response to a drug, the method comprising: extracting data from at least one database of correspondence between genetic alleles and drug responses; integrating the data using ML / Al algorithms to provide a set of drug response predictions; iteratively adding and integrating new data to the set of drug response predictions, wherein the new data comprises at least one of: a new drug; a new patient; and a new allele; and at least one of: calculating a score for the patient response to the drug using the set of drug response predictions, wherein the score corresponds to a prediction to use: a standard dose of the drug; an adjusted dose of the drug; or an alternative drug; calculating a specific therapy with a specific drug for the patient;providing a recommendation to a physician based on a relation between the genetic alleles and drug therapy outcomes; and calculating a new drug therapy by iteratively adding and integrating new data to the set of drug response predictions.
28. A computer-implemented method for providing pharmacogenetic guidelines, the method comprising: inputting a drug prescription and a patient genotype; and analyzing the drug prescription and the patient genotype to provide a drug therapy recommendation; wherein the analyzing comprises: comparing the patient genotype and drug name to a drug profile database and a molecular biomarker database and calculating a score for a patient response to the drug, wherein the score corresponds to a prediction to use:
1. a standard dose of the drug;2. an adjusted dose of the drug; or3. an alternative drug.
29. A system for providing pharmacogenetic guidelines, the system comprising: a computer input system for entering a drug prescription and a patient genotype; and a computer processor for analyzing the drug prescription and the patient genotype to provide a drug therapy recommendation; wherein the analyzing comprises: comparing the patient genotype and a drug name to a drug profile database and a molecular biomarker database and calculating a score for a patient response to the drug, wherein the score corresponds to a prediction to use:
1. a standard dose of the drug;2. an adjusted dose of the drug; or3. an alternative drug.
30. A computer-implemented method comprising:(a) training a machine learning model comprising a neural network language model, using a structured pharmacogenomic dataset, to:(i) classify a feature vector as (1) a serviceable feature vector or (2) a non- serviceable feature vector for a drug response prediction; and(ii) generate for the serviceable feature vector, using the neural network language model, the drug response prediction in natural language;(b) updating the structured pharmacogenomic dataset with a new pharmacogenomic dataset to generate an updated structured pharmacogenomic dataset; and(c) training a new machine learning model, using the updated structured pharmacogenomic dataset, such that the new machine learning model has more serviceable feature vectors than the machine learning model.
31. A computer-implemented method comprising:(a) training a machine learning model comprising a neural network language model, using a structured pharmacogenomic dataset, to:(i) classify a feature vector as (1) a serviceable feature vector or (2) a non- serviceable feature vector for a drug response prediction; and(ii) generate for the serviceable feature vector, using the neural network language model, the drug response prediction in natural language;(b) updating the structured pharmacogenomic dataset with a new pharmacogenomic dataset to generate an updated structured pharmacogenomic dataset; and(c) retraining the machine learning model, using the updated structured pharmacogenomic dataset, to add new serviceable feature vectors to the machine learning model.
32. A computer-implemented method, comprising: using a machine learning model comprising a neural network language model to process a feature vector to generate a drug response prediction in natural language, wherein the machine learning model is trained based on a structured dataset comprising a pharmacogenomic dataset relating at least 1,200 pharmacological substances and at least 350 genes.
33. A computer-implemented method, comprising: training a machine learning model comprising a neural network language model to process a feature vector to generate a drug response prediction in natural language, wherein the training is based on a structured dataset comprising a pharmacogenomic dataset relating at least 1,200 pharmacological substances and at least 350 genes.
34. A computer-implemented method, comprising: using a machine learning model comprising a neural network language model to:(a) classify a feature vector as (1) a serviceable feature vector or (2) a non- serviceable feature vector for a drug response prediction; and(b) generate a drug response prediction in natural language;wherein the machine learning model is trained based on a structured dataset comprising a pharmacogenomic dataset and at least one of: (i) an epigenomic dataset, (ii) a transcriptomic dataset, (iii) a proteomic dataset, (iv) a metabolomic dataset, (v) a lipidomic dataset, and (vi) a secretomic dataset.
35. A computer-implemented method, comprising: training a machine learning model comprising a neural network language model to:(a) classify a feature vector as (1) a serviceable feature vector or (2) a non- serviceable feature vector for a drug response prediction; and(b) generate a drug response prediction in natural language; wherein the training is based on a structured dataset a pharmacogenomic dataset and at least one of: (i) an epigenomic dataset, (ii) a transcriptomic dataset, (iii) a proteomic dataset, (iv) a metabolomic dataset, (v) a lipidomic dataset, and (vi) a secretomic dataset.
36. A computer-implemented method, comprising:(a) creating a secure connection with a client to receive an external pharmacogenomic dataset encrypted based on an external cipher;(b) receiving the external pharmacogenomic dataset through the secure connection;(c) decrypting the external pharmacogenomic dataset based on the external cipher;(d) adding the external pharmacogenomic dataset to the internal pharmacogenomic dataset;(e) decrypting the internal pharmacogenomic dataset based on an internal cipher; and(f) training the neural network language model, using the internal pharmacogenomic dataset, to generate a drug response prediction in natural language.
37. A computer-implemented method, comprising: training a neural network language model using a plurality of computers, wherein each computer comprises:(i) a database comprising a pharmacogenomic dataset, wherein the pharmacogenomic dataset of each computer in the plurality of computers is different from one another; and(ii) a containerized computer-executable instructions configured to train the neural network language model using the computer’s computational resources and the computer’s pharmacogenomic dataset, without sharing the computer’s pharmacogenomic dataset with another computer in the plurality of computers.
Citation Information
Patent Citations
Collaborative anti-tumor multi-drug combination effect prediction method based on deep learning
CN111223577A
Pharmacogenomic Decision Support for Modulators of the NMDA, Glycine, and AMPA Receptors
US20200234810A1
Artificial intelligence assisted precision medicine enhancements to standardized laboratory diagnostic testing
US20210118559A1
Prediction of adverse drug reaction based on machine-learned models using protein function scores and clinical factors
US20210327553A1
System and method for optimizing general purpose biological network for drug response prediction using meta-reinforcement learning agent
US20230094323A1
Cited By
Boundary perception drug recommendation method based on deep model and large language model collaboration
CN121905583A