Predicting responses to preventive medications with machine learning models based on patient and headache features
Machine learning models predict migraine preventive medication responses using patient health data, addressing the trial-and-error issue in current treatments and improving treatment efficacy through personalized selection.
Patent Information
- Application Number
- PCT/US2025/033049
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-11
- Filing Date
- 2025-06-10
- Publication Date
- 2025-12-26
AI Technical Summary
Current approaches to selecting migraine preventive medications are largely trial-and-error, lacking individualization, leading to sub-optimal patient outcomes and prolonged suffering due to a 30% to 50% likelihood of effective treatment.
A method using machine learning models trained on subject health data, including headache questionnaire data and clinical features, to predict the likelihood of a positive response to specific migraine preventive medications, facilitating personalized treatment selection.
This approach potentially shortens the time to find an effective headache therapy by providing personalized treatment recommendations based on individual patient data, reducing ineffective medication trials.
Smart Images

Figure US2025033049_26122025_PF_FP_ABST
Abstract
Description
PREDICTING RESPONSES TO PREVENTIVE MEDICATIONS WITH MACHINE LEARNING MODELS BASED ON PATIENT AND HEADACHE FEATURESCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 658,784, filed on June 11, 2024, and entitled “PREDICTING RESPONSES TO PREVENTIVE MEDICATIONS WITH MACHINE LEARNING MODELS BASED ON PATIENT AND HEADACHE FEATURES,” which is herein incorporated by reference in its entirety.BACKGROUND
[0002] Many medication classes, including antidepressants, antihypertensives, antiepileptics, onabotulinumtoxinA, and calcitonin gene-related peptide (CGRP)-targeting therapies are used for migraine prevention. However, any migraine preventive medication typically offers only a 30% to 50% likelihood of yielding satisfactory outcomes, commonly defined as a 50% reduction in the number of headache or migraine days per month. The current approach to finding an effective migraine preventive treatment for an individual patient is typically a trial-and-error process with treatment recommendations mainly based on patient comorbidities, preferences, and insurance coverage. As a result, patients typically spend months, if not years, in search of a medication that works effectively for them. The lack of individualization in recommending and choosing preventive medications has resulted in sub- optimal patient outcomes and prolonged suffering. There is an unmet need to tailor migraine prevention treatments with a precision and individualized approach.SUMMARY
[0003] It is an aspect of the present disclosure to provide a method for predicting treatment response for a subject to one or more migraine preventive medications. The method includes accessing subject health data with a computer system, where the subject health data may include at least one of headache questionnaire data received from the subject or subject symptom data associated with the subject. A machine learning model is also accessed with the computer system, where the machine learning model has been trained on training data to predict a treatment response to a particular migraine preventive medication based on features in subject health data. The subject health data are input to the machine learning model using the computersystem, generating classified feature data as an output, where the classified feature data indicate a likelihood of the subject having a positive response to the particular migraine preventive medication associated with the machine learning model. The classified feature data are then output using the computer system. Other embodiments of this aspect include corresponding systems (e.g., computer systems), programs, algorithms, and / or modules, each configured to perform the operations of the methods.BRIEF DESCRIPTION OF THE DRAWINGS
[0004] FIG. 1 is a flowchart of an example method for inputting subject health data to a suitably trained machine learning model to generate classified feature data that are indicative of a prediction treatment response to one or more different medications or other treatment regimens for migraine prevention.
[0005] FIG. 2 is a flowchart of an example method for training a machine learning model to generate classified feature data that are indicative of the likelihood that a subject will positively respond to a particular preventive migraine treatment.
[0006] FIG. 3 is an example workflow for training and implementing a machine learning model to generate classified feature data that are indicative of a prediction treatment response to one or more different medications or other treatment regimens for migraine prevention.
[0007] FIG. 4 is a SHAP plot showing the top 20 important variables influencing model prediction for CGRP mAb response.
[0008] FIG. 5A is a SHAP summary plot for beta blockers.
[0009] FIG. 5B is a SHAP summary plot for tricyclic antidepressants.
[0010] FIG. 5C is a SHAP summary plot for topiramate.
[0011] FIG. 5D is a SHAP summary' plot for verapamil.
[0012] FIG. 5E is a SHAP summary plot for gabapentin.
[0013] FIG. 5F is a SHAP summary plot for onabotulinumtoxinA.
[0014] FIG. 5G is a SHAP summary plot for verapamil with additional medical comorbidity7information.
[0015] FIG. 5H is a SHAP summary' plot for tricyclic antidepressants with additional medical comorbidity information.
[0016] FIG. 6 illustrates examples of positive predictors on treatment response for a group of potential migraine preventive medications.
[0017] FIG. 7 is a block diagram of an example system for predicting treatment response to one or more particular migraine preventive medications.
[0018] FIG. 8 is a block diagram of example components that can implement the system of FIG. 7.DETAILED DESCRIPTION
[0019] Described here are systems and methods for predicting a subject’s response to different types of preventive medications for primary headache disorders (e.g., migraine) and / or secondary headache disorders (e.g.. idiopathic intracranial hypertension, post-stroke headache, post-traumatic headache, or other headache disorders), and to identify predictors for treatment responses. Advantageously, the disclosed systems and methods can facilitate precision treatment and personalized treatment approaches that can potentially shorten the time needed for patients to find an effective headache therapy. For instance, the disclosed systems and methods implement machine learning models that process and identify subject health data and headache features that are capable of predicting treatment responses to one or more headache preventive medications.
[0020] The disclosed systems and methods can advantageously be built on clinical features collected at the time of the headache consultation and before initiating a headache preventive medication. Furthermore, medication-specific prediction models for each of a plurality of different preventive medications can be constructed and implemented.
[0021] Referring now to FIG. 1 , a flowchart is illustrated as setting forth the operations of an example method for generating classified feature data using a suitably trained machine learning model. As will be described, the machine learning model takes subject health data as input data and generates classified feature data as output data. As an example, the classified feature data can be indicative of a prediction treatment response to one or more different medications or other treatment regimens.
[0022] The method includes accessing subject health data with a computer system, as indicated at block 102. Accessing the subject health data may include retrieving such data from a memory or other suitable data storage device or medium. Additionally, or alternatively, accessing the subject health data can include acquiring subject health data (e.g., directly or indirectly) from a subject and recording the subject health data with the computer system, or otherwise transferring the subject health data to the computer system. As one non-limitingexample, subj ect health data may be retrieved in whole or in part from one or more data sources, which may include electronic health records for the subject.
[0023] The subject health data can include clinical test data, subject symptom data, and subject demographic data, amongst other types of subject health data or other relevant data. In some implementations, the subject health data may include only a subset of the data used to train the machine learning model. For instance, the machine learning model may be finetuned, retrained, or otherwise updated based on subject health data and classified feature data generated by the model. As the machine learning model is finetuned or otherwise retrained on these data the updated machine learning model may be able to generate accurate predictions on fewer inputs. Additionally, other machine learning techniques may also be employed to improve the model performance, including unsupervised machine learning techniques to identify subject phenotypic clusters, reinforcement learning techniques designed to improve the accuracy of the prediction for a given subject, employing other machine learning or deep learning techniques, and employing other generative artificial intelligence (Al) or agentic Al techniques.
[0024] As a non-limiting example, the subject health data may include detailed baseline questionnaire data (e.g., subject answers to headache intake questionnaires, which may include multiple choice questions, administered during initial and follow up visits), structured answer format, demographic information extracted from the subject’s EHR data, and so on. As a nonlimiting example, 145 variables from the subject health data can be used as input, including headache characteristics and demographics information.
[0025] In some embodiments, the subject health data may include subject symptom data, which can include self-reported symptoms and / or symptoms recorded in, or derived from, the subject’s EMR data. EHR data, questionnaire data, conversational data, or the like. For example, subject symptom data may include questionnaire response data including symptoms reported by the subject when answering one or more questionnaires. Additionally or alternatively, subject symptom data may include symptom features extracted from EHR data. For instance, a large language model (LLM) or other natural language processing (NLP) based techniques can be used to extract symptom features from free text clinical notes stored in EHR data, or directly from a conversation between a patient and an Al system or chatbot.
[0026] As a non-limiting example, the subject symptom data can include baseline headache characteristics reported by the subject, including headache frequency, headache duration, headache location, headache quality, headache intensity, associated symptoms, aura.number of years with migraine, current preventive and acute therapies, and the like. Headache frequency can, for example, include a qualitative or quantitative measurement of how frequently the subject suffered headaches symptoms over an interval of time. As an example, headache frequency can be computed as the number of days with a headache captured over the trailing 28 days and compared to baseline.
[0027] In some instances, the subject symptom data may also include symptoms reported by the subject in a headache symptom diary, headache intake forms when visiting a clinician, a headache questionnaire, free text or narrative description of headache related symptoms, conversational data captured by an Al agent or a clinical provider, or the like, which may collect an individual's headache related experiences, including symptoms experienced, headache frequency and intensity, and treatment outcomes. Examples of additional subject symptom data may include the duration of migraine attacks, the number of years having headaches, headache onset ages, menstruation history or status, sleep related information (e.g., sleep state, quality of sleep, bed time, wake time, hours of sleep), detailed description of the head pain (throbbing, pressure, jabbing / stabbing), migraine attack triggers, family history, cranial autonomic symptoms, prior medication treatment trials, medication overuse, prodrome symptoms, postdrome symptoms, past medical histories, other medical conditions or medications taken, and other headache related experiences. Once a patient receives migraine preventive medications, the follow up questionnaires are given, or patients are inquired about the treatment outcomes of those medications. In some instances, the information can be obtained multiple times and for multiple medications during different visits or follow-up inquiries.
[0028] The demographic data can include subject age, subject gender, subject sex, subject race, subject ethnicity, body mass index (BMI). martial status, subject handedness, and / or other social determinants.
[0029] Additionally or alternatively, the subject health data may include other types of subject health data, including data stored in, retrieved from, extracted from, or otherwise derived from the subject’s EMR and / or EHR. The subject health data can include unstructured text obtained during or outside of a doctor’s visit, subject symptom data (e.g., physical symptoms, mental health symptoms), questionnaire response data, clinical test data, clinical laboratory data, histopathology' data, demographic data, medical comorbidities, problem lists, medications currently on or used, social determinants of health, social history’, patient preference, insurance coverage, genetic sequencing, multi-omic data, other bodily fluid-basedbiomarkers, medical imaging, data from wearable devices (e.g., physiological measurements or other data recorded with a wearable device), and other such clinical data types. Physiological measurements that may be recorded with a wearable device include heart rate, temperature, or other physical parameters. Other data that may be recorded with a wearable device may include measurements of stress, measurements of sleep state and / or sleep quality, treatment responses, medication adverse events, and the like. Additionally, the subject health data may include location data and / or environmental factors.
[0030] As one non-limiting example, subject health data can include clinical features associated with information derived from clinical records of a subject, which in some instances can also include records from family members of the subject. These clinical features and data may be abstracted from unstructured clinical documents. EMR, EHR. or other sources of subject history. The clinical features could also be directly obtained from patients through an Al system or an Al chatbot. Such data may include subject symptoms, diagnosis, treatments, medications, therapies, responses to treatments, laboratory' testing results, medical history', geographic locations, environmental factors, demographics, or other features of the subject which may be found in the subject’s EMR and / or EHR.
[0031] Additionally or alternatively, subject health data can include clinical features derived from structured, curated, EMR and / or EHR data, such as diagnoses; symptoms; therapies; outcomes; subject demographics, such as subject name, date of birth, gender, and / or ethnicity: diagnosis dates for illness, disease, or other physical or mental conditions; personal medical history'; family medical history; clinical diagnoses, such as date of initial diagnosis, ; and the like. Additionally, the subject health data may also include features such as treatments and outcomes, such as line of therapy, therapy groups, clinical trials, medications prescribed or taken, non-pharmacological treatments, behavioral interventions, neuromodulations, surgeries, imaging, adverse effects, and associated outcomes.
[0032] Examples of clinical laboratory' data and / or histopathology data can include genetic testing and laboratory information, such as performance scores, lab tests, pathology results, prognostic indicators, date of genetic testing, testing method used, and so on.
[0033] In some embodiments, the subject health data can include a collection of data and / or features including all of the datatypes disclosed above. Alternatively, the subject health data may include a selection of fewer data and / or features.
[0034] Data from other modalities, including comorbidities, genomics, non-genomic body fluid-based biomarkers, and imaging can also be incorporated to enhance the performance of the predictive models described in the present disclosure.
[0035] A trained machine learning model is then accessed with the computer system, as indicated at block 104. In general, the machine learning model is trained, or has been trained, on training data in order to predict the likelihood that a subject will respond to one or more particular medications. The machine learning model may include a gradient boosting machine (GBM) model, a distributed random forest (DRF) model, a generalized linear model (GLM) model, an XGBoost model, a stacked ensemble model, a convolutional neural network, a residual neural network, or the like. Alternatively, the machine learning model could implement other suitable machine learning or artificial intelligence algorithms, such as those based on supervised learning, unsupervised learning, deep learning, ensemble learning, dimensionality reduction, other generative Al techniques, agentic Al systems, combinations thereof, and so on. In some instances, large language models or other foundation models (e.g., models pre-trained on a large amount of medical or non-medical data) may also be incorporated to analyze direct text inputs from subjects.
[0036] Accessing the trained machine learning model may include accessing model parameters that have been optimized or otherwise estimated by training the machine learning model on training data. In some instances, retrieving the machine learning model can also include retrieving, constructing, or otherwise accessing the particular machine learning model architecture to be implemented. For instance, data pertaining to the layers in a neural network architecture (e.g., number of layers, type of layers, ordering of layers, connections between layers, hyperparameters for layers) may be retrieved, selected, constructed, or otherwise accessed.
[0037] The subject health data are then input to the one or more trained machine learning models, generating output as classified feature data, as indicated at block 106. For example, the classified feature data may indicate the probability for a particular classification (i.e., the probability that the subject health data include patterns, features, or characteristics indicative of detecting, differentiating, and / or determining the likelihood of the patient being a responder to one or more candidate drugs). The machine learning model(s) may output a probability of whether the subject will be a responder to each of a plurality of different medications. A responder classification can be defined based on a percentage reduction in headache frequency, intensity, and / or level of disability when using a particular medication.As a non-limiting example, the responder classification can be defined as a subject having a 30% reduction in headache frequency compared to baseline in response to a medication. Using such a pre-defined cutoff value, the machine learning model(s) may output a binary response (i.e., classification) that the subject will likely be a responder or not.
[0038] In some aspects, only a subset of the subject health data are input to the machine learning model(s). For instance, a subset of features in the subject health data that have been identified as relevant for predicting treatment response to a particular medication can be selected and used as an input to the relevant machine learning model. In some embodiments, the subset of subject health data can be selected based on variable importance analysis or SHAP analysis to identify features in the subject health data that are relevant for predicting treatment response to a particular medication. In some implementations, unsupervised pre-training with a deep learning architecture, such as TabNet, a deep neural network architecture, or other deep learning models or foundation models, can be used for a model to leam the representation of the overall data structure, then the embeddings derived from such unsupervised pre-training can be used as an additional or alternative input to the machine learning model(s).
[0039] The classified feature data generated by inputting the subject health data to the trained machine learning model(s) can then be displayed to a user, stored for later use or further processing, or both, as indicated at block 108. Based on these classified feature data, a clinician or other healthcare provider can make an informed selection of a particular preventive medication to administer to a subject before starting treatment for that subject, rather than administering a medication to the subject that may be ineffective.
[0040] In some instances, the updated machine learning model may require fewer inputs than the initial machine learning model. In these cases, the most relevant features of the subject health data may be identified for the updated model, such as by using a SHAP analysis, variable importance analysis, other variable selection processes, or the like. Based on one or more of these analyses, one or more relevant features of the subject health data are identified. For instance, relevant subject health data features can be selected as those with SHAP and / or variable importance values at or above a threshold. SHAP and / or variable importance values at or above a percentile rank of all subject health data features, or the like. Then, in future implementations the subject may be prompted to collect only the more limited subset of subject health data features.
[0041] In some cases, based on the classified feature data generated by the machine learning model(s). a healthcare provider may administer an effective dose of a particularmigraine preventive medication to the subject when the classified feature data indicates a likelihood that the subject will have a positive response to that medication. The classified feature data may provide probability scores or confidence levels for different medications, allowing the healthcare provider to select the medication with the highest predicted likelihood of success for the individual subject. The effective dose may be determined based on established clinical guidelines, the subject’s medical history, body weight, age, and other relevant factors. In some implementations, the dose may be initiated at a standard starting dose and then titrated up or down based on the subject’s response and tolerance. The healthcare provider may consider factors such as the subject’s previous medication trials, current medications, comorbid conditions, and potential drug interactions when determining the appropriate dosage. In some cases, the classified feature data may include an indication of a potential effective dose of the particular headache preventive medication for the subject.
[0042] For example, if the classified feature data indicates a high probability of positive response to CGRP monoclonal antibodies, the healthcare provider may administer an effective dose of erenumab. fremanezumab, galcanezumab, or eptinezumab according to established dosing protocols. Similarly, if the model predicts a favorable response to topiramate, the healthcare provider may initiate treatment with an appropriate starting dose and gradually increase the dose as tolerated to achieve therapeutic benefit.
[0043] The administration of the medication may be accompanied by monitoring for treatment response and adverse effects. Follow-up assessments may be conducted to evaluate changes in headache frequency, intensity, and associated disability. The subject’s response to the administered medication may be documented and used to further refine the predictive models through the retraining process described herein.
[0044] In some cases, if the classified feature data indicates a low probability of response to a particular medication, the healthcare provider may choose to avoid that medication and instead select an alternative treatment option with a higher predicted likelihood of success. This approach may help reduce the trial-and-error process typically associated with migraine preventive treatment selection and may lead to more rapid achievement of therapeutic benefit for the subject.
[0045] Referring now to FIG. 2, a flow chart is illustrated as setting forth the operations of an example method for training one or more machine learning models on training data, such that the one or more machine learning models are trained to receive subject health data as input data in order to generate classified feature data as output data, where the classified feature dataindicate the likelihood that a subject will positively respond to a particular migraine preventive medication.
[0046] In general, the machine learning model(s) can implement any number of different machine learning model architectures. For instance, the machine learning model(s) could implement gradient boosting machine (GBM) model, a distributed random forest (DRF) model or other random forest model architecture, a generalized linear model (GLM) model, an XGBoost model, a stacked ensemble model, or any suitable deep learning-based model, including a convolutional neural network, a residual neural network, or the like. Additionally or alternatively, models based on a large language model (LLM), other transformer-based models, other generative Al models, other agentic Al frameworks, or other foundation models may be used to directly process subject health data (e.g., text data stored in the subject health data, other data stored in the subject health data) input by the subject. Alternatively, the machine learning model could implement other suitable machine learning or artificial intelligence algorithms, such as those based on supervised learning, unsupervised learning, deep learning, ensemble learning, dimensionality reduction, combinations thereof, and so on.
[0047] The method includes accessing training data with a computer system, as indicated at block 202. Accessing the training data may include retrieving such data from a memory7or other suitable data storage device or medium.
[0048] In general, the training data can include subject health data obtained from a group or population of subjects, and may additionally or alternatively include detailed migraine features and treatment response information gathered from the group or population of subjects. The method can include assembling training data from subject health data or other relevant data (e g., migraine features, demographic or other clinical information, treatment response data) using a computer system. This operation may include assembling the subject health data or other relevant data into an appropriate data structure on which the machine learning model can be trained. Assembling the training data may include assembling subject health data and other relevant data. For instance, assembling the training data may include generating labeled data and including the labeled data in the training data. Labeled data may include subj ect health or other relevant data that have been labeled as belonging to, or otherwise being associated with, one or more different classifications or categories. For instance, labeled data may include subject health data that have been labeled as being associated with one or more predictive outcomes (e.g.. treatment responses).
[0049] One or more machine learning models are trained on the training data, as indicated at block 204. In general, the machine learning model can be trained by optimizing model parameters based on minimizing a loss function. As one non-limiting example, the loss function may be a mean squared error loss function. The machine learning model may be any suitable machine learning model, including those based on supervised learning, unsupervised learning, ensemble learning, or other learning techniques. As a non-limiting example, the machine learning model may be a GBM model, a DRF model, a GLM model, an XGBoost model, or a stacked ensemble model.
[0050] In some aspects, the machine learning model may be trained using a two stage process. In the first stage, unsupervised representation learning is used to pre-train a deep neural network architecture using TabNet. then the TabNet-embedded data can be used to represent each subject to train one or more supervised machine learning models. An example of this model training and implementation workflow is illustrated in FIG. 3.
[0051] For example, using 145 variables among all patients included in the analysis, an unsupervised representation learning may first be conducted by pre-training a deep neural network architecture using TabNet. TabNet is a deep learning interpretable structure that utilizes sequential attention for reasoning and feature selection, designed to allow more efficient learning and representation of tabular data. Because the training data may include different numbers of subjects for whom treatment response data may be available for each medication, embeddings produced by TabNet can be used to construct prediction models for each type / class of medication to predict binary outcomes (responder vs. non-responder).
[0052] The supervised machine learning process can be conducted using a pipeline that includes automatic cross-validation, hyperparameter tuning, and model selection. For each medication, subjects who were treated with the medication with available treatment response information were selected to construct a model. Subjects were randomly divided into training (85%) and held-out test (15%) sets. The training process can include five-fold cross validation within the training set among various models (e g., 30 different models), including but not limited to gradient boosting machine (GBM), distributed random forest (DRF), generalized linear model (GLM), XGBoost, and Stacked Ensemble models. The cross-validation matrices can then be used to select a top-performing model, and the performance of the top model for each medication was then evaluated in the held-out test set, referring to data not used during the model training, and served as a final check to see how well the model generalizes to unseen data.
[0053] In some implementations, the machine learning model(s) may be retrained or updated on new data to improve their performance and accuracy over time. The retraining process may be performed at regular intervals, such as daily, weekly, monthly, quarterly, or at any other suitable frequency depending on the availability of new training data and computational resources. The frequency of retraining may be adjusted based on various factors, including the rate at which new data becomes available, changes in clinical practice patterns, the introduction of new medications, and observed changes in model performance over time. In some implementations, automated systems may be employed to monitor model performance and trigger retraining when performance metrics fall below predetermined thresholds.
[0054] The retraining process may involve accessing additional subject health data, migraine features, demographic information, other clinical information and / or treatment response information that has been collected since the initial model training. This new data may include updated questionnaire responses, follow-up visit information, treatment outcomes, and other relevant clinical data from subjects who have received migraine preventive medications. In some cases, the new data may also include information from subjects who were not included in the original training dataset.
[0055] During the retraining process, the new data may be preprocessed and formatted in a manner consistent with the original training data. Missing values may be imputed using appropriate techniques, and categorical variables may be encoded as needed. The updated training dataset, which may include both the original training data and the newly collected data, may then be used to retrain one or more of the machine learning models. The retraining may involve updating model parameters through additional training iterations while maintaining the same model architecture, or it may involve selecting new model architectures that may be better suited for the expanded dataset. In some implementations, the retraining process may include hyperparameter tuning and cross-validation to optimize model performance on the updated dataset.
[0056] In some aspects, the retraining process may allow the models to be customized for specific patient populations, clinical settings, or geographic regions. For example, if new data reveals that certain patient characteristics are more prevalent in a particular population, the retrained models may be better adapted to make accurate predictions for that population.
[0057] The one or more trained or retrained machine learning models are then stored for later use, as indicated at block 206. Storing the machine learning model(s) may include storing model parameters, which have been computed or otherwise estimated by training themachine learning model(s) on the training data. Storing the trained machine learning model(s) may also include storing the particular machine learning model architecture to be implemented. For instance, data pertaining to the layers in a neural network architecture (e g., number of layers, type of layers, ordering of layers, connections between layers, hyperparameters for layers) may be stored.
[0058] In an example study, the disclosed systems and methods were tested to predict whether subjects would respond to one or more different medications for treating migraine. This example study was a cohort analysis of data prospectively collected from clinical practice at a tertiary' headache center. All data were extracted from a previously collected headache database and were preprocessed. Individual prediction models were constructed for each medication, the importance of each variable in the input data was analyzed, and the models were evaluated for potential bias.
[0059] A database of detailed headache characteristics gathered through in-depth headache intake questionnaires that have been prospectively collected for patients was used as a data source for model training and evaluation. All patients with migraine completed an electronic questionnaire prior to their initial headache consultation, either online through a patient portal before the visit, or in the waiting room on an electronic tablet prior to meeting headache specialists. The questionnaires included 53 multiple choice questions including demographics, headache description, migraine-associated symptoms, and prior medication trials. When patients returned for follow-up visits, they were given a brief questionnaire inquiring about the preventive medications they had taken during the past three months, and their current monthly headache days. All answers were incorporated into the EHR and automatically populate into a headache consultation note and follow-up visit notes. During the encounter a clinician can revise the answers in a note-writing user interface in the EHR, which verifies the answers to be reflected in the consultation notes and stored in the database. The database platform includes a variety of commonly used headache preventive medications, including antihypertensives, anti-seizure medications, antidepressants, onabotulinumtoxinA injections, calcitonin gene-related peptide (CGRP)-targeting therapies, and non-medication- based therapies. The answers to the questionnaires were stored in an electronic headache application, and data can be retrieved based on specific filter criteria. Additionally, an EHR data warehouse was queried to obtain information including race, ethnicity, marital status, and missing height / weight information.
[0060] Based on extracted data, treatment response information on 530 patients treated with CGRP mAbs and the 30% responder rate was 144 / 530 (33.4%). For multivariate analysis with 0.30 medium effect size and 145 predictor variables, it was estimated that a minimum of 500 patients treated with CGRP mAbs would be necessary to identify predictive variables and build prediction models that achieve 85% power (0.05 alpha). It was expected that the least follow-up data would be available for CGRP mAbs as they were recently added to the database. Medications for which treatment response information was available on more than 500 patients w ere used to construct individual medication prediction models
[0061] The initial consultation questionnaire included 183 variables derived from the 53 multiple choice questions in the questionnaire. A question can expand to multiple variables. For example, the question "At times during my headaches and on the same side as the head pain, I experience: Select all that apply.” expands to ten variables of different cranial autonomic symptoms like “A blocked or stuffy nose “A sense of restlessness or agitation, “None of these After extensive data preprocessing, we included 145 variables which had less than 15% missing values, comprising 132 categorical variables, such as sex, race, ethnicity, marital status, handedness, aura, headache description, location, migraine-associated symptoms, detailed cranial autonomic symptoms, hormonal association, family history, common migraine attack triggers, detailed previous acute and preventive medication treatment trials; and 13 numerical variables, including age, monthly headache days, migraine attack duration, headache intensity, BMI. height, and weight. Monthly headache days, a relevant measure in treatment response for this analysis, was obtained from the question "Over the past 4 weeks, how many days have you had a headache? Enter a number from 0-28”. In preparation for machine learning analysis, missing values were imputed with mean values for numerical variables and “unknown’7for categorical variables. All categorical variables were encoded and then all variables were scaled between 0 to 1.
[0062] Follow-up visits questionnaire data were processed based on the different preventive medications and the reported monthly headache days at the time of each visit. For each patient, the initial visit data and follow-up visit data were compared to compute the longitudinal changes of monthly headache days. To analyze a patient’s treatment response to a preventive medication (e.g., topiramate as anon-limiting example), the follow-up visit when a patient first reported having taken the medication at the time of the visit was identified and defined as the first “post-treatment visit.” The pre-treatment (baseline) visit was then identified as the visit earlier than, but closest to. the first post-treatment visit that the patient was not onthe particular medication. Visits were included if monthly headache days were available for both the pre- and first post- treatment visits, and if they were between 28 and 730 days apart to accurately reflect treatment response from a medication.
[0063] “Calculated treatment response1’ to a medication was computed using the following equation: [(pre-treatment monthly headache days - post-treatment monthly headache days) / pre-treatment monthly headache days] x 100%. A treatment responder was defined as having at least a 30% calculated treatment response. The above treatment response analysis was computed for each preventive medication, and all patients with calculated treatment responses (regardless of responder or non-responder) to at least one type / classes of the following migraine preventive medications in our analysis were identified: topiramate, betablockers (propranolol, metoprolol, atenolol, nadolol, timolol), tricyclic antidepressants (amitriptyline, nortriptyline), verapamil, gabapentin, divalproex sodium, tizanidine, venlafaxine, onabotulinumtoxinA, and CGRP monoclonal antibodies (mAbs) (erenumab, fremanezumab, galcanezumab, eptinezumab). Afterwards, individual prediction models were constructed to predict treatment responses to each type / class of medication.
[0064] Using 145 variables among all patients included in the analysis, unsupervised representation learning was conducted by pre-training a deep neural network architecture using TabNet. As there were different numbers of patients whom had treatment response data available for each medication, embeddings produced by TabNet were employed to construct prediction models for each type / class of medication to predict binary outcomes (responder versus non-responder).
[0065] The superv ised machine learning process was conducted using the H2OAutoML (3.44.0.3) pipeline, an open-source platform that allowed automatic cross- validation, hyperparameter tuning and model selection. For each medication, patients who have been treated with the medication with available treatment response information were selected to construct the model. Patients were randomly divided into training (85%) and held-out test (15%) sets. The training process included five-fold cross validation within the training set among 30 various models, including GBM, DRF, GLM, XGBoost, and Stacked Ensemble models. The cross-validation matrices were used to select the top-performing model, and the performance of the top model for each medication was then evaluated in the held-out test set, referring to data not used during the model training, and served as a final check to see how well the model generalizes to unseen data.
[0066] Area under receiver operating characteristics curve (AUROC or AUC) were obtained in the held-out test set and the prediction threshold with the maximal F 1 score was used to evaluate the precision (positive predictive value), recall (sensitivity), accuracy, and Fl score for each model. The AUC offers an aggregated assessment of model performance over various prediction thresholds and indicates the model’s ability in distinguishing between positive and negative instances, in this case, responder versus non-responder. An AUC greater than 0.5 indicates better performance than a random guess. Fl score is the harmonic mean of precision and recall.
[0067] To enhance model transparency, SHapley Additive exPlanations (SHAP) analysis for the top performing model for each medication was conducted. SHAP values quantify the importance of a variable and the direction it contributes for the model to make predictions. The analysis was summarized in a SHAP summary plot for the top performing model for each medication. As variable importance analysis is currently not available for Stacked Ensemble models, if the top model was a Stacked Ensemble model, the top ‘‘nonstacked” model was used to conduct SHAP analysis.
[0068] The performance of each model was evaluated separately for male only or female only patients to see if the model performance was comparable when predicting treatment responses.
[0069] A total of 4260 patients for whom both the initial headache consultation data and calculated treatment response information available to at least one type / classes of the following migraine preventive medications were included in the analysis: topiramate, betablockers (propranolol, metoprolol, atenolol, nadolol, timolol), tricyclic antidepressants (amitriptyline, nortriptyline), verapamil, gabapentin, divalproex sodium, tizanidine, venlafaxine, onabotulinumtoxinA, and CGRP monoclonal antibodies (mAbs) (erenumab, fremanezumab, galcanezumab, eptinezumab). The demographic information, detailed migraine features, associated symptoms and prior treatment trials among all patients are presented in Table 1. At the time of the initial headache consultation, the average age was 42.8 (14.5), the mean monthly headache days was 20.5 (8.3). the mean BMI was 28.7 (7.5), and 3356 / 4260 (78.8%) were females.Table 1: The demographic information, detailed migraine features, associated symptoms, and prior treatment trials among the 4260 patients included in our final analysis. Each of the listed answers to individual questions were variables included in prediction modeling. “Y” is yes, "N” is no.| Age| Age at the time of initial consultation: 42.8(14.4)
[0070] Among the 4260 patients, 1731 patients had been treated with only one of the ten included medications, and other patients had been treated with at least two of the ten medications. To construct individual models to predict treatment responses to each type / class of medication, 1348, 1394, 1358, 1400, 1031, 1655, and 524 patients treated with beta blockers, topiramate, tricyclic antidepressants, gabapentin, verapamil, onabotulinumtoxinA, and CGRP mAbs, respectively, were used. The demographics of individuals included in each medication model are shown in Table 2. The number of patients who achieved at least a 30% reduction in headache frequency (i.e., “responders”), ranged from 28.7% to 34.9% for each medication, as shown in Table 2. The number of patients with available treatment response information for divalproex sodium (136), tizanidine (89) and venlafaxine (139) were too small to develop individual prediction models.Table 2 The number of patients used to construct each prediction model, along with their demographics, baseline monthly headache days, and the number and percentage of 30% responders. Beta blockers include propranolol, metoprolol, atenolol, nadolol, timolol. Tricyclic antidepressants include amitriptyline and nortriptyline. CGRP mAbs include erenumab, fremanezumab, galcanezumab, eptinezumab.
[0071] In the held-out test set, the CGRP mAh prediction model achieved a high AUC of 0.825 (95% CI 0.726. 0.920) with an accuracy of 0.80 (95% CI 0.70, 0.88). The precision, recall and Fl score were 0.65 (95% CI 0.43, 0.87), 0.57 (95% CI 0.36, 0.78), and 0.60 (95% CI 0.41, 0.77), respectively. The held-out test set AUCs of prediction models for beta-blockers, tricyclic antidepressants, topiramate, verapamil, gabapentin, and onabotulinumtoxinA were 0.664 (95% CI 0.579, 0.745), 0.611 (95% CI 0.562, 0.682), 0.605 (95% CI 0.520, 0.688), 0.673 (095% CI 0.569. 0.724). 0.628 (0.533. 0.661). and 0.581 (95% CI 0.550. 0.632). respectively. The AUCs based on the train and test sets for each model, along with the evaluation metrices including precision, recall, accuracy and Fl score for each model in the held-out test set are summarized in Table 3.Table 3 Model type, performance matrices, and bias analysis of the individual prediction models predicting treatment responses for each medication. The model predicts a binary outcome- responder versus non-responder. A treatment responder is defined as having at least 30% reduction in headache frequency. GBM: Gradient Boosting Machine
[0072] The SHAP summary' plots outlining the top 20 most important variables contributing to making predictions for CGRP mAbs are presented in FIG. 4. The SHAP summary plots for other medications are illustrated in FIGS. 5A-5F.
[0073] In FIG. 4, the red color signifies higher values of the variable, while positive SHAP values (distributed on the right side of the central line) indicate a higher likelihood of being a responder. FIG. 4 suggests that having the following factors positively predictstreatment response: Fewer baseline monthly headache days, more years of having their “typical headaches’; shorter duration of severe migraine attacks, higher headache pain intensity (0-10); less severe headache-related disability, lower weight, lower BMI, older age at the time of the visit, younger age at the time first having any headache, not having used barbiturate or supplements for their headache; considering previous topiramate treatment to be effective, lower number of prior preventive medications and triptans trials, reporting lack of sleep as a migraine trigger, experiences stuffy nose during headache attacks, headache starting in the back of the head, has a family history of headache (any family member), being ambidextrous or lefthanded.
[0074] FIG. 5 A suggests that having the following factors positively predicts treatment response to beta blockers: Fewer baseline monthly headache days; more years of having their “typical headaches”; lower headache pain intensity (0-10); lower weight; lower BMI; older age at the time of the initial visit; older age at the time first having any headaches; not having tried topiramate or antiepileptics; considering previous tricyclic antidepressant treatment to be effective; low frequency or not using opioid as their acute treatment; higher number of triptans tried; lower total number of previous acute and preventive medication trials; high frequency of combination analgesic use; reporting missing meals, menstrual flow, and too little sleep as migraine attack triggers; headache ever accompanied with vomiting, headache starting as unilateral.
[0075] FIG. 5B suggests that having the following factors positively predicts treatment response to tricyclic antidepressants: fewer baseline monthly headache days, longer duration (minutes) of severe migraine attacks, less severe headache related disability, more years of having their “ty pical headaches”, lower weight, lower height, older age at the time of the initial visit, considering previous topiramate, tricyclic antidepressant and onabolulinumtoxinA treatment to be effective, lower number of total previous preventive medication trials, not having tried antiepileptics or onabolulinumtoxinA for their headache, reporting menstrual flow as a migraine attack trigger, denying flickering light as a migraine attack trigger, headache starting as unilateral, experience stuffy’ nose during headache, headache not worsened by exertion, being divorced / separated / widowed.
[0076] FIG. 5C suggests that having the following factors positively predicts treatment response to topiramate: number of baseline monthly headache days is not extremely high or low, higher BMI, lower height, younger age at the time of the initial visit, younger age at the time first having any headache, not having tried tricyclic antidepressants, topiramate or beta-blocker for their headache, lower number of total previous preventive medication trials, higher number of previous acute medication trials, using combination analgesics, medication overuse, reporting weather changes and stress as migraine attack triggers, headache ever accompanied with vomiting.
[0077] FIG. 5D suggests that having the following factors positively predicts treatment response to verapamil: Fewer baseline monthly headache days; longer duration (years) of having their '’typical headaches”; shorter duration (minutes) of severe migraine attacks, higher weight, higher height, younger age at the time first having headache, older age at the time of the initial visit, low er number of total previous preventive medication trials, not having tried tricyclic antidepressants or antiepileptics for their headache, considering previous topiramate treatment to be effective, low frequency or not using combination analgesics, denying missed meal as a migraine attack trigger, headache started as unilateral (as opposed to bilateral), headache ever accompanied with vomiting; experience stuffy nose during headache, describing their headache as “throbbing” or “j abbing”, mother not having headache.
[0078] FIG. 5E suggests that having the following factors positively predicts treatment response to gabapentin: fewer baseline monthly headache days, lower headache pain intensity (0-10), higher BMI, higher weight, lower height, older age at the time of the initial visit, older age at the time first having any headache, lower number of total previous preventive medication trials, not having tried antiepileptics or tricyclic antidepressants for their headache, low frequency or not using triptans, high frequency of acetaminophen use, headache starting in the back of the head, experience stuffy nose during headache, not having aura, brother(s) with headache, headache starting as unilateral.
[0079] FIG. 5F suggests that having the following factors positively predicts treatment response to onabotulinumtoxinA: Fewer baseline monthly headache days, longer duration of severe migraine attacks, fewer years of having their “typical headaches”, higher BMI, lower height, younger age at the time of the initial visit, low er number of total previous preventive medication trials, considering previous topiramate treatment to be effective, not having tried tricyclic antidepressants or onabotulinumtoxinA for their headache, high frequency of combination analgesics use, higher number of total triptans trials, reporting too much sleep as a migraine trigger, denying lack of sleep or flickering light as a migraine attack trigger, headache starting as unilateral, experience restlessness during headache, female.
[0080] Additionally, the top features that contribute to model performance for each medication are summarized in Table 4. The variables were summarized into the followingcategories: headache frequency, duration, intensity: BMI, height, weight; age factors; previous treatment trials; migraine attack triggers; headache description and associated symptoms; family history, and others.Table 4: Results from the SHAP analysis identifying the top 20 most important variables for each prediction model. Patients having the following features are more likely to be predicted as a 30% responder.
[0081] For CGRP mAbs. the top 5 most important features for a patient to be predicted as a 30% responder were lower baseline monthly headache days, lower weight, longer duration (years) of having their “typical headaches”, lower BMI, and shorter duration (minutes) of severe migraine attacks. Notably, baseline monthly headache days was the most important variable in all medication prediction models, except for onabotulinumtoxinA for which it was the second most important variable. Headaches typically starting unilaterally, as opposed to bilaterally, or sometimes unilaterally and sometimes bilaterally, was the most important positive predictor of being a treatment responder for onabotulinumtoxinA.
[0082] Several variables were among the top 20 across all medication prediction models, but sometimes they had opposite contributions. For example, lower BMI and lowerweight were predictive of positive treatment responses to tricyclic antidepressants, beta blockers and CGRP mAbs, while higher BMI and higher weight were predictive of positive treatment responses to onabotulinumtoxinA, topiramate, verapamil and gabapentin. Older age at the time of the initial visit predicted positive treatment responses to beta blockers, tricyclic antidepressants, gabapentin, and CGRP mAbs, while younger age predicted positive treatment response to topiramate, verapamil, and onabotulinumtoxinA. Reporting lack of sleep as a migraine attack trigger was a positive predictor for CGRP mAbs but a negative predictor for onabotulinumtoxinA treatment response. Lower number of total previous preventive medication trials was predictive of positive treatment responses to all medications. Several migraine features, such as headache starting unilaterally, headache ever accompanied with vomiting and experiencing stuffy nose on the side of the pain during headache, were predictive of positive treatment responses to many medications, as illustrated in Table 4. Aura was only a negative predictor for gabapentin, and sex (female) was only a positive predictor for onabotulinumtoxinA, with neither being in the top 20 features contributing to model prediction for most medications, as illustrated in Table 4.
[0083] Overall positive predictors for each medication are illustrated in FIG. 6.
[0084] The performance of each model, evaluated in male only and female only patients in the test sets, are presented in Table 3. The model AUCs for female only patients in the test set were higher for beta blockers and tricyclic antidepressants, while the AUC was higher for male patients for topiramate, verapamil, gabapentin, onabotulinumtoxinA, and CGRP mAbs.
[0085] From the EHR database, other medical conditions for the patients were extracted. The medical conditions were identified through ICD codes. The conditions extracted included anxiety, depression, insomnia, cancer, hypertension, seizure disorders, coronary artery disease, cognitive disorders, obesity, traumatic brain injury / concussion, hypothyroidism, hyperthyroidism, or other thyroid disorders. With this additional information, the previously stated machine learning methodology was repeated on the same cohort of patients, including employing unsupervised pre-training followed by supervised machine learning model development. The model performance for the seven different types of medications, as evaluated in the held-out test set, is listed in Table 5.Table 5: Model performance evaluated in the held-out test set for the same cohort of patients but with 13 additional medical comorbidities included as variables
[0086] Overall, it was observed that adding comorbidities information improved model performance for some medications, including beta-blockers, tricyclic antidepressants, verapamil, and gabapentin. Additionally, when examining the variable importance through variable importance analysis and Shapley analysis, it was observed that, overall, the important variables were consistent with previously identified features. Several comorbid conditions, such as depression, anxiety, and obesity, were observed in this example study to be among the most important variables when the model made predictions. For example, having anxiety, depression, and obesity positively predicted treatment response to verapamil. Not having depression predicted treatment response to tricyclic antidepressants. The SHAPley plots withadded comorbid conditions for verapamil and tricyclic antidepressants are shown in FIGS. 5G and 5H.
[0087] The example study demonstrated that a robust machine learning model can accurately predict treatment response to CGRP mAbs (AUC 0.825). The prediction models developed for six other medications had modest prediction ability (AUC ranged from 0.581 to 0.673). Other modalities, including medical comorbidities, genomic, or imaging characteristics can be incorporated into prediction modeling to enhance the model performance for migraine preventive medications.
[0088] The study demonstrated the feasibility of incorporating machine learning prediction models as a clinical decision support tool that can be used when choosing among various migraine preventive medications, in addition to considering patient comorbidities, patient preferences, and insurance coverage plans, factors that are currently used in practice. These migraine preventive treatment prediction models have the potential to facilitate a more personalized approach to migraine treatment that would shorten the trial-and-error process of finding an effective medication.
[0089] In addition to model performance, relevant clinical features that were predictive of treatment responses to each of the medications were identified. Furthermore, the summarized results of patient answers to a headache questionnaire were reported in Table 1, which contains comprehensive migraine symptoms.
[0090] Factors, including lower number of baseline monthly headache days and lower number of previous preventive medication trials, were identified to be positive predictors for treatment response to all seven preventive medications evaluated in this study. Furthermore, several factors, including weight, BMI, age, and certain migraine attack triggers were observed to be relevant variables for model prediction, but had positive or negative contributions depending on the particular preventive medication. Additionally, as weight or BMI were among the top predictors for treatment responses to all seven medications evaluated, future studies could explore the association between BMI and migraine preventive medications.
[0091] Among the medications included in predictive modeling, amitriptyline, metoprolol, propranolol, atenolol, timolol, nadolol, topiramate, onabotulinumtoxinA, and the CGRP mAbs were classified as having established efficacy or probably effective according to the American Headache Society consensus statement. Although there is no strong evidence to support the use of gabapentin and verapamil for migraine prevention, they are commonly used in the real-world setting, especially for patients who had previously tried multiple medicationsor who have other comorbid conditions. Having prediction models and identifying predictors for treatment responses to those medications could still inform future clinical practice and research.
[0092] The disclosed prediction models were based on advanced machine learning, including automating supervised machine learning tasks with AutoML pipelines that allowed comparing different model frameworks and selecting the top performing model for each medication. Deep learning was also incorporated for more efficient data representation. Employing embeddings derived from unsupervised pre-training with a deep neural network architecture, TabNet, enhanced the performance of the top CGRP mAbs prediction model from an AUC of 0.775 without TabNet embedding, to 0.823 in the held-out test set.
[0093] Results from this study demonstrate the feasibility of precision migraine treatment, i.e., predicting response to individual types of migraine preventive medications, using pre-treatment information.
[0094] FIG. 7 illustrates an example of a system 700 for predicting subject response to various migraine preventive medications in accordance with some embodiments described in the present disclosure. As shown in FIG. 7, a computing device 750 can receive one or more types of data (e.g., subject health data) from data source 702. In some embodiments, computing device 750 can execute at least a portion of a migraine medication treatment response prediction system 704 to predict a subject’s treatment response to one or more migraine preventive medications using data received from the data source 702.
[0095] Additionally or alternatively, in some embodiments, the computing device 750 can communicate information about data received from the data source 702 to a server 752 over a communication network 754, which can execute at least a portion of the migraine medication treatment response prediction system 704. In such embodiments, the server 752 can return information to the computing device 750 (and / or any other suitable computing device) indicative of an output of the migraine medication treatment response prediction system 704.
[0096] In some embodiments, computing device 750 and / or server 752 can be any suitable computing device or combination of devices, such as a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable computer, a server computer, a virtual machine being executed by a physical computing device, and so on.
[0097] In some embodiments, data source 702 can be any suitable source of data (e.g., subject health data, processed subject health data, other data extracted or derived from subject health data), another computing device (e.g.. a server storing subject health data, processedsubject health data, other data extracted or derived from subject health data), and so on. In some embodiments, data source 702 can be local to computing device 750. For example, data source 702 can be incorporated with computing device 750 (e.g., computing device 750 can be configured as part of a device for measuring, recording, estimating, acquiring, or otherwise collecting or storing data). As another example, data source 702 can be connected to computing device 750 by a cable, a direct wireless link, and so on. Additionally or alternatively, in some embodiments, data source 702 can be located locally and / or remotely from computing device 750, and can communicate data to computing device 750 (and / or server 752) via a communication network (e.g., communication network 754).
[0098] In some embodiments, communication network 754 can be any suitable communication network or combination of communication networks. For example, communication network 754 can include a Wi-Fi network (which can include one or more wireless routers, one or more switches, etc.), a peer-to-peer network (e.g., a Bluetooth network), a cellular network (e g., a 3G network, a 4G network, etc., complying with any suitable standard, such as CDMA. GSM, LTE, LTE Advanced. WiMAX, etc.), other types of wireless network, a wired network, and so on. In some embodiments, communication netw ork 754 can be a local area network, a wide area network, a public network (e.g., the Internet), a private or semi -private network (e.g., a corporate or university intranet), any other suitable ty pe of netw ork, or any suitable combination of networks. Communications links shown in FIG. 7 can each be any suitable communications link or combination of communications links, such as wired links, fiber optic links, Wi-Fi links, Bluetooth links, cellular links, and so on.
[0099] Referring now to FIG. 8, an example of hardware 800 that can be used to implement data source 702, computing device 750, and server 752 in accordance with some embodiments of the systems and methods described in the present disclosure is shown.
[0100] As shown in FIG. 8, in some embodiments, computing device 750 can include a processor 802, a display 804, one or more inputs 806, one or more communication systems 808, and / or memory 810. In some embodiments, processor 802 can be any suitable hardware processor or combination of processors, such as a central processing unit (“CPU”), a graphics processing unit (“GPU”), and so on. In some embodiments, display 804 can include any suitable display devices, such as a liquid crystal display (“LCD”) screen, a light-emitting diode (“LED”) display, an organic LED (“OLED”) display, an electrophoretic display (e.g., an “e- ink” display), a computer monitor, a touchscreen, a television, and so on. In someembodiments, inputs 806 can include any suitable input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, and so on.
[0101] In some embodiments, communications systems 808 can include any suitable hardware, firmware, and / or software for communicating information over communication network 754 and / or any other suitable communication networks. For example, communications systems 808 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 808 can include hardware, firmware, and / or software that can be used to establish a Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.
[0102] In some embodiments, memory’ 810 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for example, by processor 802 to present content using display 804, to communicate with server 752 via communications system(s) 808, and so on. Memory 810 can include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 810 can include random-access memory ('’RAM"), read-only memory (‘■ROM”), electrically programmable ROM (“EPROM”), electrically erasable ROM (“EEPROM”), other forms of volatile memory’, other forms of non-volatile memory, one or more forms of semi-volatile memory’, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some embodiments, memory 810 can have encoded thereon, or otherwise stored therein, a computer program for controlling operation of computing device 750. In such embodiments, processor 802 can execute at least a portion of the computer program to present content (e.g., images, user interfaces, graphics, tables), receive content from server 752, transmit information to server 752, and so on. For example, the processor 802 and the memory 810 can be configured to perform the methods described herein (e.g., the method of FIG. 1, the method of FIG. 2).
[0103] In some embodiments, server 752 can include a processor 812, a display 814, one or more inputs 816, one or more communications systems 818. and / or memory’ 820. In some embodiments, processor 812 can be any suitable hardware processor or combination of processors, such as a CPU, a GPU, and so on. In some embodiments, display 814 can include any suitable display devices, such as an LCD screen, LED display, OLED display, electrophoretic display, a computer monitor, a touchscreen, a television, and so on. In some embodiments, inputs 816 can include any suitable input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, and so on.
[0104] In some embodiments, communications systems 818 can include any suitable hardware, firmware, and / or software for communicating information over communication network 754 and / or any other suitable communication networks. For example, communications systems 818 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 818 can include hardware, firmware, and / or software that can be used to establish a Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.
[0105] In some embodiments, memory 820 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for example, by processor 812 to present content using display 814, to communicate with one or more computing devices 750, and so on. Memory 820 can include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 820 can include RAM, ROM, EPROM, EEPROM, other types of volatile memory, other types of non-volatile memory, one or more types of semi-volatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some embodiments, memory 820 can have encoded thereon a server program for controlling operation of server 752. In such embodiments, processor 812 can execute at least a portion of the server program to transmit information and / or content (e.g., data, images, a user interface) to one or more computing devices 750, receive information and / or content from one or more computing devices 750. receive instructions from one or more devices (e.g., a personal computer, a laptop computer, a tablet computer, a smartphone), and so on.
[0106] In some embodiments, the server 752 is configured to perform the methods described in the present disclosure. For example, the processor 812 and memory 820 can be configured to perform the methods described herein (e.g.. the method of FIG. 1, the method of FIG. 2).
[0107] In some embodiments, data source 702 can include a processor 822, one or more data acquisition systems 824, one or more communications systems 826, and / or memory 828. In some embodiments, processor 822 can be any suitable hardware processor or combination of processors, such as a CPU, a GPU, and so on. In some embodiments, the one or more data acquisition systems 824 are generally configured to acquire data, images, or both, are generally configured to acquire or otherwise receive subject health data, and can include databases; smartphones, tablet computers, or other mobile devices; computer systems; smartwatches or other wearable devices; and the like. Additionally or alternatively, in some embodiments, theone or more data acquisition systems 824 can include any suitable hardware, firmware, and / or software for coupling to and / or controlling operations of databases; smartphones, tablet computers, or other mobile devices; computer systems; smartwatches or other wearable devices; and the like. In some embodiments, one or more portions of the data acquisition system(s) 824 can be removable and / or replaceable.
[0108] Note that, although not shown, data source 702 can include any suitable inputs and / or outputs. For example, data source 702 can include input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, a trackpad, a trackball, and so on. As another example, data source 702 can include any suitable display devices, such as an LCD screen, an LED display, an OLED display, an electrophoretic display, a computer monitor, a touchscreen, a television, etc., one or more speakers, and so on.
[0109] In some embodiments, communications systems 826 can include any suitable hardware, firmware, and / or software for communicating information to computing device 750 (and, in some embodiments, over communication network 754 and / or any other suitable communication networks). For example, communications systems 826 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 826 can include hardware, firmware, and / or software that can be used to establish a wired connection using any suitable port and / or communication standard (e.g.. VGA, DVI video, USB, RS-232, etc.), Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.
[0110] In some embodiments, memory 828 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for example, by processor 822 to control the one or more data acquisition systems 824, and / or receive data from the one or more data acquisition systems 824; to generate images from data; present content (e g., data, images, a user interface) using a display; communicate with one or more computing devices 750; and so on. Memory 828 can include any suitable volatile memory', non-volatile memory', storage, or any suitable combination thereof. For example, memory 828 can include RAM. ROM, EPROM, EEPROM, other ty pes of volatile memory, other ty pes of non-volatile memory, one or more types of semi-volatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some embodiments, memory 828 can have encoded thereon, or otherwise stored therein, a program for controlling operation of data source 702. In such embodiments, processor 822 can execute at least a portion of the program to generate images, transmitinformation and / or content (e.g., data, images, a user interface) to one or more computing devices 750. receive information and / or content from one or more computing devices 750, receive instructions from one or more devices (e.g., a personal computer, a laptop computer, a tablet computer, a smartphone, etc.), and so on.
[0111] In some embodiments, any suitable computer-readable media can be used for storing instructions for performing the functions and / or processes described herein. For example, in some embodiments, computer-readable media can be transitory or non-transitory. For example, non-transitory computer-readable media can include media such as magnetic media (e.g., hard disks, floppy disks), optical media (e.g., compact discs, digital video discs, Blu-ray discs), semiconductor media (e.g., RAM, flash memory’, EPROM, EEPROM), any suitable media that is not fleeting or devoid of any semblance of permanence during transmission, and / or any suitable tangible media. As another example, transitory computer- readable media can include signals on networks, in wires, conductors, optical fibers, circuits, or any suitable media that is fleeting and devoid of any semblance of permanence during transmission, and / or any suitable intangible media.
[0112] As used herein in the context of computer implementation, unless otherwise specified or limited, the terms “component,” “system,” “module,” “framework,” and the like are intended to encompass part or all of computer-related systems that include hardware, software, a combination of hardware and software, or software in execution. For example, a component may be, but is not limited to being, a processor device, a process being executed (or executable) by a processor device, an object, an executable, a thread of execution, a computer program, or a computer. By way of illustration, both an application running on a computer and the computer can be a component. One or more components (or system, module, and so on) may reside within a process or thread of execution, may be localized on one computer, may be distributed between two or more computers or other processor devices, or may be included within another component (or system, module, and so on).
[0113] In some implementations, devices or sy stems disclosed herein can be utilized or installed using methods embodying aspects of the disclosure. Correspondingly, description herein of particular features, capabilities, or intended purposes of a device or system is generally’ intended to inherently include disclosure of a method of using such features for the intended purposes, a method of implementing such capabilities, and a method of installing disclosed (or otherwise known) components to support these purposes or capabilities. Similarly, unless otherwise indicated or limited, discussion herein of any method ofmanufacturing or using a particular device or system, including installing the device or system, is intended to inherently include disclosure, as embodiments of the disclosure, of the utilized features and implemented capabilities of such device or system.
[0114] The present disclosure has described one or more preferred embodiments, and it should be appreciated that many equivalents, alternatives, variations, and modifications, aside from those expressly stated, are possible and within the scope of the disclosure.
Claims
CLAIMS1. A method for predicting treatment response for a subj ect to one or more headache preventive medications, the method comprising:(a) accessing subject health data with a computer system, wherein the subject health data comprise at least one of headache questionnaire data received from the subject or subject symptom data associated with the subject;(b) accessing a machine learning model with the computer system, wherein the machine learning model has been trained on training data to predict a treatment response to a particular headache preventive medication based on features in subject health data;(c) inputting the subject health data to the machine learning model using the computer system, generating classified feature data as an output, wherein the classified feature data indicate a likelihood of the subject having a positive response to the particular headache preventive medication associated with the machine learning model; and(d) outputting the classified feature data using the computer system.
2. The method of claim 1, wherein the machine learning model is trained in a two stage process comprising an unsupervised representation learning stage and a supervised learning based on embeddings generated by the unsupervised learning stage.
3. The method of claim 2, wherein the unsupen ised representation learning stage includes pretraining a deep learning neural network.
4. The method of claim 3, wherein the deep learning neural netw ork comprises a TabNet neural network.
5. The method of claim 1. wherein the machine learning model is one of a gradient boosting machine (GBM) model, a distributed random forest (DRF) model, a generalized linear model (GLM) model, an XGBoost model, or a stacked ensemble model.
6. The method of claim 1. wherein the subject health data comprise both headache questionnaire data and subject symptom data.
7. The method of claim 6, wherein the subject health data further comprise subject demographic data.
8. The method of claim 7, wherein the subject demographic data comprise at least one of subject age, subject gender, subject race, subject ethnicity, or body mass index (BMI).
9. The method of claim 1. wherein the subject symptom data comprise at least one of headache frequency, headache duration, headache location, headache quality, headache intensity, associated symptoms, aura, prior medication treatment trials, or number of years with migraine.
10. The method of claim 1, wherein accessing the machine learning model with the computer system comprises accessing a plurality of machine learning model, wherein each of the plurality7of machine learning models has been trained on training data to predict the treatment response to a different one of a plurality of particular headache preventive medications based on features in subject health data; and wherein the subject health data are input to each of the plurality of machine learning models to generate a plurality of classified feature data sets as an output, wherein each of the plurality of classified feature data sets indicate the likelihood of the subject having the positive response to a different one of the plurality of particular headache preventive medications.
11. The method of claim 1 , wherein the training data comprise headache questionnaire data, subject symptom data, and treatment response data collected from a group of test subjects.
12. The method of claim 1, wherein the particular headache preventive medication comprises at least one of topiramate, a beta-blocker, a tricyclic antidepressant, verapamil,gabapentin, divalproex sodium, tizanidine, venlafaxine, onabotulinumtoxinA, or a CGRP monoclonal antibody (mAb).
13. The method of claim 12, wherein the particular headache preventive medication is the beta-blocker, wherein the beta-blocker comprises one of propranolol, metoprolol, atenolol, nadolol, or timolol.
14. The method of claim 12, wherein the particular headache preventive medication is the tricyclic antidepressant, wherein the tricyclic antidepressant comprises one of amitripty line or nortriptyline.
15. The method of claim 12, wherein the particular headache preventive medication is the CGRP mAb, wherein the CGRP mAb comprises one of erenumab, fremanezumab, galcanezumab, or eptinezumab.
16. The method of claim 1, wherein the subject symptom data are received from the subject as an input.
17. The method of claim 1. wherein the subject symptom data are extracted from electronic health record data for the subject accessed by the computer system.
18. The method of claim 1, wherein the headache preventive medication comprises a medication for treating primary headache.
19. The method of claim 1 , wherein the headache preventive medication comprises a medication for treating secondary' headache.
20. The method of claim 1. further comprising administering an effective dose of the particular migraine preventive medication to the subject based on the classified feature data indicating the likelihood of the subject having the positive response to the particular headache preventive medication.
21. The method of claim 20, further comprising monitoring the subject for treatment response following administration of the effective dose.
22. The method of claim 21, further comprising retraining the machine learning model using the treatment response to the administered effective dose of the particular preventive headache medication.
Citation Information
Patent Citations
Method and system for predicting optimal epilepsy treatment regimes
US20180211012A1
Medication recommendation system and method for treating migraine
US20210295998A1
Using Electronic Health Records and Machine Learning to Predict and Mitigate Postpartum Depression
US20210375468A1
System, device and method for safeguarding the wellbeing of patients for fluid injection
US20240115143A1
Cited By
Knowledge and data fusion driven interpretable primary headache auxiliary identification method
CN122135924A