Systems and methods for designing and conducting clinical trials and biomarker validation studies
By using predictive models to analyze, identify, and stratify patient data, the high cost and low efficiency of recruiting suitable participants in clinical trials and biomarker validation studies have been addressed, enabling more accurate and efficient patient selection and study implementation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FLAGSHIP PIONEER INNOVATION VII LLC
- Filing Date
- 2024-09-19
- Publication Date
- 2026-05-29
AI Technical Summary
In existing clinical trials and biomarker validation studies, recruiting suitable participants is costly and inefficient, and it is difficult to find and reach candidates who meet specific criteria.
Using predictive models (such as machine learning models) to receive patient data as input, providing predictions of disease status, identifying participants who meet the requirements of clinical trials or biomarker validation studies, stratifying candidate individuals by a stratification axis, selecting a cohort of individuals that conforms to the target distribution, and facilitating the implementation of related studies.
This enables more accurate and efficient identification and selection of patients who meet research requirements, improving the efficiency and success rate of clinical trials and biomarker validation studies.
Smart Images

Figure CN122122665A_ABST
Abstract
Description
Cross-references to related applications
[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 584,033, filed September 20, 2023, the entire contents of which are incorporated herein by reference. Background Technology
[0002] Clinical trials are crucial for improving healthcare outcomes and advancing medical knowledge. However, a major challenge in conducting clinical trials is recruiting suitable participants. Recruiting suitable participants is essential to ensuring the success and effectiveness of clinical trials. While recruitment methods are important, current methods often face several challenges, including high costs, inefficiency, and difficulty in finding and reaching qualified participants.
[0003] Similarly, biomarker validation studies are crucial in the development of personalized medicine, therapeutic interventions, and diagnostic tools. However, as with clinical trials, a key challenge in conducting biomarker validation studies is recruiting suitable candidates. Biomarker validation may require participants who meet specific criteria, such as a particular genetic profile, disease stage, or pre-existing condition.
[0004] Improved systems and methods are needed to recruit candidates for clinical trials, biomarker validation studies, or other longitudinal or cross-sectional studies. Summary of the Invention
[0005] In various examples, this disclosure relates to systems and methods for designing and conducting clinical trials, biomarker validation studies, and other research. Predictive models (e.g., machine learning models) are trained to receive patient data as input and provide predictions of the patient's health status or disease state as output. The health status or disease state can be or includes, for example, predictions along a health-disease axis or a disease risk axis, and / or can be expressed as a probability distribution. Similar predictions can be made for other patients, and these predictions can be used to identify a group of participants for a clinical trial (or other study). For example, a clinical trial may have a target stage or target probability distribution for a disease of interest, and model predictions can be used to select patients who meet the target stage or target probability distribution. In some cases, for example, patients can be selected such that the sum of the predicted probability distributions of the patients reaches or approximates the target probability distribution of the clinical trial. Other patient characteristics, such as age, sex, or residential address, may also be considered when selecting patients. The selected patients can then be included as participants in the clinical trial. Advantageously, in some examples, the systems and methods described herein allow for more accurate and efficient identification and selection of patients who meet the requirements of a clinical trial or other study.
[0006] In one aspect, the subject matter of this disclosure relates to a method for conducting clinical trials or biomarker validation studies. The method includes: obtaining access to a computer-implemented model trained using training data comprising multiple records of multiple individuals, each record including (i) values for multiple characteristics of a corresponding individual from the plurality of individuals, and (ii) a label providing an indication of a disease state for that corresponding individual; providing the trained computer-implemented model with multiple sets of values for the multiple characteristics, each set of values corresponding to a candidate individual from a set of candidate individuals; receiving a prediction of a disease state for each set of values from the trained computer-implemented model; identifying a group of participants from the set of candidate individuals based on the prediction of the disease state; and facilitating at least one of a clinical trial or biomarker validation study involving that group of participants.
[0007] In some examples, the values of these multiple features describe at least one of the following: demographic characteristics, medical visits, physical examinations, diagnoses, medications, medical procedures, vital sign measurements, vaccinations, laboratory results, serum, urine samples, biological samples, gene expression levels, medical images, clinical notes, radiological reports, genetic testing, biomarkers, pathology reports, health information, social determinants of health, financial data, consumer data, or any combination thereof. The model may include at least one of the following: regression models, classifiers, linear models, nonlinear models, random forests, kernel methods, Bayesian models, decision trees, or neural networks. Predictions of disease states may include probability distributions. Predictions of disease states may be on or associated with a primary hierarchical axis, which is or includes a health-disease axis or a disease risk axis.
[0008] In some implementations, identifying the group of participants includes stratifying each candidate individual from the plurality of candidate individuals according to one or more primary stratification axes. Identifying the group of participants may further include stratifying each candidate individual from the plurality of candidate individuals according to one or more secondary stratification axes. The one or more secondary stratification axes may represent one or more attributes of the plurality of individuals, and these attributes may include at least one of physiological comorbidities, demographic characteristics, or socioeconomic variables. The one or more primary stratification axes or at least one of the one or more secondary stratification axes may represent at least one of continuous or ordered dimensions. The location along the one or more primary stratification axes or at least one of the one or more secondary stratification axes may be represented by points, point estimates, or probability distributions.
[0009] In various contexts, the clinical trial or biomarker validation study may include studies of at least one of the following: progress in health transformation, efficacy of the biomarker, efficacy of a biomarker group, efficacy of a targeted therapy, efficacy of a therapeutic intervention, efficacy of a digital intervention, efficacy of a behavioral intervention, or any combination thereof. The participants in the candidate set may be enriched relative to one or more attributes, including reduced physiological heterogeneity, increased probability of an outcome event, increased propensity to respond to an intervention, increased likelihood of observing specific physiological or pathological signs, or any combination thereof.
[0010] In some examples, identifying the group of participants includes: determining a target distribution for the candidate set of individuals, which defines a probability distribution, based on the objectives of the clinical trial or the biomarker validation study; and selecting the group of participants from the candidate set to achieve a collective distribution similar to the target distribution. The target distribution may be determined based on at least one parameter to be evaluated in the clinical trial or the biomarker validation study, including at least one of a biological attribute, a biomarker, a biomarker group, a therapeutic target, or a pharmacological intervention. The target distribution may be constant or uniform with respect to disease state (e.g., along a health-disease axis). The target distribution may include a risk probability distribution. The target distribution may be regionally enriched with respect to a stratified axis associated with the disease state. Selecting the group of participants may include minimizing the number of individuals in the group. Selecting the group of participants may include: identifying multiple groups of individuals, where each group has a similar probability distribution with respect to the disease state; and determining the number of individuals to be included in the group of participants from each group (and / or individuals not in any group) to satisfy the target distribution.
[0011] In various cases, the method includes facilitating the clinical trial, which includes evaluating the efficacy of the treatment. Facilitating the clinical trial may include enrolling members from the group of participants in the clinical trial, enabling the administration of a drug, digital intervention, or behavioral intervention to these members. The method may include facilitating the biomarker validation study, which may include validating a biomarker or a biomarker set.
[0012] In various implementations, identifying the group of participants includes achieving a desired distribution of at least one covariate, which includes demographic characteristics, comorbidities, medication use, medical history, socioeconomic attributes, biological attributes, or any combination thereof. The method may include using the training data to train the computer-implemented model.
[0013] On the other hand, the subject of this disclosure relates to a system for conducting clinical trials or biomarker validation studies. The system includes one or more computer processors programmed to perform operations including: accessing or gaining access to a computer-implemented model trained using training data comprising multiple records of multiple individuals, each record including (i) values for multiple characteristics of a corresponding individual from the multiple individuals, and (ii) a label providing an indication of the disease state of the corresponding individual; providing the trained computer-implemented model with multiple sets of values for the multiple characteristics, each set of values corresponding to a candidate individual from a set of candidate individuals; receiving a prediction of the disease state for each set of values from the trained computer-implemented model; identifying a group of participants from the set of candidate individuals based on the prediction of the disease state; and facilitating at least one of a clinical trial or biomarker validation study involving that group of participants.
[0014] In some examples, the prediction of disease status may be on or associated with a primary stratification axis, which may include a health-disease axis or a disease risk axis. Identifying this group of participants may include stratifying each candidate individual from the plurality of candidate individuals according to one or more primary stratification axes and / or one or more secondary stratification axes.
[0015] Identifying the group of participants may include: determining a target distribution for the candidate individual set based on the objectives of the clinical trial or the biomarker validation study, the target distribution defining a probability distribution; and selecting the group of participants from the candidate individual set to achieve a collective distribution of the group of participants that is similar to or approximates the target distribution. The target distribution may include a risk probability distribution. Selecting the group of participants may include: identifying multiple groups of individuals, wherein individuals in each of the multiple groups have similar probability distributions regarding the disease state; and determining the number of individuals to be included in the group of participants from each group (and / or individuals not in any group) to satisfy the target distribution. These operations may include using the training data to train the computer-implemented model.
[0016] These and other objects, as well as the advantages and features of the embodiments of the invention disclosed herein, will become more apparent from the following description, drawings, and claims. Furthermore, it should be understood that the features of the various embodiments described herein are not mutually exclusive and can exist in various combinations and arrangements. Attached Figure Description
[0017] The foregoing will become clear from the following more specific description of exemplary embodiments, as illustrated in the accompanying drawings, in which the same reference numerals refer to the same parts in different views. The drawings are not necessarily to scale, but rather focus on illustrating the embodiments.
[0018] Figure 1A and Figure 1B It is a graph of the number of individuals along the health-disease axis relative to the disease state for a given disease, according to certain embodiments.
[0019] Figure 2 This is a schematic diagram of a method for identifying individual cohorts for clinical trials or other studies, according to certain embodiments.
[0020] Figure 3 This is a schematic diagram of a method for training a machine learning model according to certain embodiments.
[0021] Figure 4 This is a schematic diagram of a method for identifying a cohort of individuals to participate in a study, based on certain embodiments.
[0022] Figure 5 This is a block diagram of an example computer system according to certain embodiments. Detailed Implementation
[0023] The following is a description of exemplary embodiments. It is contemplated that the apparatuses, systems, methods, and processes of the claimed invention encompass variations and adaptations developed using information from the embodiments described herein. Adaptations and / or modifications to the apparatuses, systems, methods, and processes described herein can be performed by those skilled in the art.
[0024] It should be understood that the order of steps or the order in which certain actions are performed is irrelevant as long as the invention remains operable. Furthermore, two or more steps or actions can be performed simultaneously.
[0025] In various examples, an "organism" (which may be referred to in this text as an "individual," "subject," or "patient") can refer to an animal, plant, fungus, or any living thing. An organism can be, for example, a mammal, such as a human or a mouse.
[0026] In various examples, "disease" can refer to any functional or structural disorder in an organism. A disease can be any type of illness, including, for example: cancer (e.g., colorectal cancer, lung cancer, skin cancer, etc.), hereditary diseases (e.g., spinal muscular atrophy (SMA), Huntington's disease (HD), or Duchenne muscular dystrophy (DMD)), liver diseases (e.g., metabolic dysfunction-associated steatosis (MASLD), metabolic dysfunction-associated steatohepatitis (MASH), or cirrhosis), kidney diseases (e.g., chronic kidney disease (CKD) or polycystic kidney disease), diabetes (e.g., type II diabetes), neurodegenerative diseases (e.g., Alzheimer's disease...). AD), amyotrophic lateral sclerosis (ALS) or Parkinson's disease (PD), autoimmune diseases (e.g., multiple sclerosis (MS), rheumatoid arthritis (RA) or inflammatory bowel disease (IBD)), lung diseases (e.g., chronic obstructive pulmonary disease (COPD) or pulmonary fibrosis (e.g., idiopathic pulmonary fibrosis (IPF)), long-term COVID-19, colon polyps, age-related macular degeneration (AMD), or any other type of disease. In some examples, the disease may be, include, and / or manifest as inflammation, cellular stress response, or other types of functional impairment.
[0027] In various examples, "disease of interest" can refer to a disease that is characterized or analyzed using the predictive modeling techniques described in this paper. For example, a machine learning model can be trained to make predictions related to the disease of interest.
[0028] In various examples, a “reference domain” (which may be referred to herein as a “feature space”) can be or include a collection of features or variables (e.g., biological data, latent variables, or phenotype-specific variables) that can be used as training data for training machine learning models and / or as inputs to machine learning models. A reference domain may include, for example, gene expression levels and / or other biological features related to the disease of interest.
[0029] In various examples, "biomarker" can refer to a measurable indicator of a biological condition or state. Biomarkers can indicate the presence or progression of a disease or the effectiveness of treatment. Biomarkers can be biomolecules and can be found in tissues, blood, or body fluids. Biomarkers can be used for diagnostic purposes, such as to confirm or detect the presence of a symptom or disease of concern, or to identify individuals with a subtype of that disease. Biomarkers can be repeatedly measured to monitor individuals, for example, to assess the state of a medical condition or disease, or to look for evidence of the effects of environmental factors or medical products (or exposure to environmental factors or medical products).
[0030] In some examples, one or more predictive models (e.g., machine learning models or computer-implemented models) are used to identify individuals for inclusion in clinical trials, biomarker validation studies, cross-sectional studies, or longitudinal studies (referred to as “trials” or “studies” throughout this document for simplicity). These one or more predictive models can be configured to receive an individual’s health data as input and provide that individual’s predicted disease state, health status, and / or disease risk as output. Such predictions can be made for a set of individuals. Based on the predictions, a cohort of individuals meeting the target disease state, target health status, and / or target disease risk can be selected for the clinical trial or other study.
[0031] Figure 1A This is a graph 100 showing the number of individuals N along the health-disease axis relative to disease status for a given disease, based on certain examples. Graph 100 includes population distributions for two cohorts. A first population distribution 110 represents a cohort natively selected from the general population and indicates that the majority of individuals are healthy. A second population distribution 112 represents a cohort corresponding to a target stage 114 of the disease in a clinical trial. For example, the cohort in the second population distribution 112 might be of interest in prescribing and evaluating treatment in a clinical trial. In various examples, the systems and methods described herein can be used to identify the cohort in the second population distribution 112. The target stage 114 may correspond to a range or value of disease severity or risk. For example, a target stage may correspond to a range or value of eGFR in chronic kidney disease, visual acuity in age-related macular degeneration, time of diagnosis in Parkinson's disease, or risk of diabetes (e.g., the risk of developing diabetes).
[0032] Figure 1B This is based on a similar curve 120 for some examples, showing the number of individuals N relative to disease status along the health-disease axis for a given disease. Curve 120 includes a first population distribution 110 corresponding to a cohort selected from the place of origin. Curve 120 also includes a third population distribution 122, which has an equal or uniform representation on the health-disease axis. In various examples, the third population distribution 122 is desirable for biomarker validation studies where biomarkers can be evaluated for reliability and accuracy.
[0033] In some examples, the systems and methods described herein can be used to identify cohorts from a second population distribution 112 and a third population distribution 122. For example, a machine learning model (or other predictive model) can be configured to receive an individual's health information as input and determine the individual's disease state or health status as output. To create a cohort from the second population distribution 112, a machine learning model can be used to identify individuals with disease states corresponding to the target stage 114 targeted by the clinical trial. Such individuals can be included in the cohort, while other individuals with disease states falling outside the target stage 114 can be excluded from the cohort.
[0034] Similarly, to create a cohort of the third population distribution 122, machine learning models can be used to identify and select individuals with various disease states ranging from healthy to unhealthy on the health-disease axis. In some examples, a uniform representation of the third population distribution 122 can be obtained by selecting combinations of individuals who are healthy (e.g., low risk of developing disease), moderately healthy (e.g., moderate risk of developing disease), unhealthy (e.g., high risk of developing disease), and / or in other intermediate disease stages. The combined disease states of the selected individuals can achieve or approximate a uniform representation of the third population distribution 122.
[0035] In various examples, individuals can be selected for a cohort by considering the risk probability distribution for each individual. For example, each individual can have a risk probability distribution along a health-disease axis, and the cohort can be selected such that the sum of the risk probability distributions of the individuals in the cohort reaches a desired population distribution (or an approximation thereof), such as a third population distribution 122. In some examples, the risk probability distribution provides the probability associated with an individual's risk of having or developing a disease. Additionally or alternatively, the risk probability distribution can be associated with or derived from an individual's predicted health status along the health-disease axis. For example, an individual predicted to be healthy can have a risk probability distribution indicating a low risk of having or developing a disease. Similarly, an individual predicted to be unhealthy can have a risk probability distribution indicating a high risk of having or developing a disease. For a disease, there can be multiple risk levels, such as low, low-medium, high-medium, or high, and the risk probability distribution can provide a probability value for each risk level. The prediction model described herein can be used to calculate the risk probabilities and risk probability distributions.
[0036] Figure 2This is a schematic diagram of a method 200 for identifying a desired cohort based on certain examples. Risk probability distributions 210, 212, and 214 for individual patients are calculated using predictive models. In the depicted example, patients fall into three groups: a low-risk group (risk distribution shifted downwards) represented by risk probability distribution 210, an intermediate-risk group (risk distribution peaks in the central region) represented by risk probability distribution 212, and a high-risk group (risk distribution shifted upwards) represented by risk probability distribution 214. A certain number of patients can be selected or drawn from each group to form the final cohort. Sampling: In the depicted example, 10 patients are sampled from the low-risk group, 8 from the intermediate-risk group, and 12 from the high-risk group. Additionally or alternatively, if desired, one or more patients with unique risk probability distributions (e.g., different from risk probability distributions 210, 212, and 214) can be selected. The resulting cohort in the depicted example consists of low-risk, intermediate-risk, and high-risk subjects. Histogram 216 of the cohort shows the estimated number of patients in each of the following risk categories: low, low-medium, high-medium, and high. Threshold line 218 corresponds to the target distribution of studies involving this cohort, for example, as specified by the study sponsor.
[0037] In some examples, an "over-recruitment" of subjects can be undertaken to ensure that the overall density or number of patients exceeds a threshold line 218 for all risk categories. For example, a minimum number of patients that satisfies or approximates this target distribution and has a threshold number of patients within each risk category can be selected. Advantageously, this method enables the selection of recruitment cohorts that satisfy the target parameters based on the risk probability distribution, even if the model cannot precisely place any individual subject along the health-disease axis. While the target distribution (represented by threshold line 218) in this example is constant or uniform, it should be understood that the target distribution can be non-uniform. For example, the target distribution can resemble risk probability distributions 210, 212, or 214.
[0038] In some embodiments involving biomarker validation, it may be desirable to obtain a uniform distribution of individuals along a health-disease axis, a disease risk axis, or other dimensions (e.g., represented by a third population distribution 122 or a threshold line 218). Advantageously, method 200 can utilize predictive modeling and probability distribution-based sampling to achieve the desired distribution. In some examples, one or more models (referred to herein as "models" for simplicity) can generate a probability distribution representing the health status or risk probability of each individual, and the model can aggregate individuals to produce a broader population distribution of interest. This method can provide the expected subject count or the projected size of the recruitment cohort. The subject count can be used to evaluate the quality of method 200 or other methods used for cohort identification. Generally, stronger methods are able to meet trial requirements with smaller subject counts.
[0039] In some embodiments, one or more stratification axes may be selected for a group of individuals being considered for a study (referred to as “candidate individuals”), and the model may stratify each individual in the group according to the one or more stratification axes. Stratification axes may define, relate to, and / or provide measures of one or more inclusion criteria for the study. Such criteria may include one or more primary stratification axes, and optionally, one or more secondary stratification axes. The one or more primary stratification axes may relate to or provide measures of health or disease, such as disease state or stage (e.g., along a health-disease axis), health status, and / or disease risk (e.g., the risk of having or developing a disease along a disease risk axis). The one or more secondary stratification axes may include or relate to other characteristics, such as sex, age, race, demographic characteristics, treatment history, medication use, physical comorbidities, socioeconomic variables, or any combination thereof. For example, in addition to uniform recruitment across the health-disease axis (e.g., the primary stratification axis), the trial designers (e.g., the trial sponsor) may want an equal number of male and female participants (e.g., a secondary stratification axis). In specific examples involving chronic kidney disease, it may be desirable to select participants with a specific distribution of eGFR values (e.g., on the primary stratification axis of eGFR) and specific distributions in age, sex, and / or distance to the clinical trial site (three secondary stratification axes).
[0040] In various examples, the model and / or related techniques can be used to identify patients to be recruited for a study based on their location or probability distribution along a selected axis. The model can be used to select and / or identify subgroups of candidate individuals whose collective distribution of location or probability along one or more stratified axes is projected to resemble or approximate a target distribution for the trial (e.g., a second population distribution 112 or a third population distribution 122). This prediction can be used to recruit selected individuals for the study, for example, via invitations, emails, or notifications automatically generated to healthcare providers through an electronic healthcare record system. Alternatively or additionally, selected individuals can be recruited via primary care providers, specialist healthcare providers, or other settings. In some examples, information related to the stratified axis can be obtained from a database of candidate individuals' records.
[0041] Additionally or alternatively, risk probabilities or risk probability distributions can be used to make treatment decisions for patients. For example, patient data can be fed into the predictive model described herein, and the predictive model can determine and / or output a probability distribution representing the patient's health (e.g., a risk probability distribution). Probability distributions can be used by patients, physicians, and / or other healthcare professionals to make healthcare and / or treatment decisions for patients. In some examples, predictive models can use probability distributions to determine and / or output one or more treatment recommendations for a patient. In some examples, probability distributions can be used to derive indicators that are more familiar and interpretable to patients or physicians, such as disease risk or percentage odds.
[0042] Figure 3 This is a schematic diagram of a method 300 for training a machine learning model 310 according to certain examples. In some embodiments, model 310 may be or include a regression model, classifier, linear model, nonlinear model, random forest, kernel method, Bayesian model, decision tree (e.g., XGBoost), neural network, generative model, or other type of machine learning model. In some examples, model 310 may be or include, for example, a multilayer fully connected neural network.
[0043] In various examples, model 310 can be trained to receive individual data as input and provide predicted disease states, health conditions, and / or disease risks for that individual as output. Input data used to train model 310 may include, for example, patient data 312 (e.g., medications, diagnoses, laboratory tests, vital signs, etc.) and / or survey data 314 (if available). Model inputs may form a reference domain or feature space for model 310. In some embodiments, patient data 312 may include real-world data existing in medical records or may be created using laboratory tests, imaging examinations, other tests, consumer data, social determinants of health, genetic data, or any other relevant source. This data may be directly measured or imputed (e.g., using a regression model). Survey data 314 may be derived from questions and corresponding answers, and / or may represent additional data captured prospectively for each subject that may not be readily accessible from medical records or other data sources. Survey data 314 and / or patient data 312 may include, for example, waist-to-hip ratio, waist circumference, hip circumference, body fat percentage, blood pressure, employment status, diet, exercise, etc. In some cases, model input data may include information related to demographic characteristics (e.g., gender, race, income, etc.).
[0044] In some examples, the training data and / or model output may include a continuous representation 316 of health, which may be or include predicted values or scores along a continuous dimension of health (e.g., a health-disease axis). In some examples, the continuous representation 316 of health may form the basis of current staging paradigms, which are often abstracted into high-level categories such as “Stage II” or “Stage IV”. The continuous representation 316 of health may be used as a label in the training data or as ground truth data. The continuous representation 316 of health may be or include, for example, image-based fat scores for non-alcoholic fatty liver disease (NAFLD), estimated glomerular filtration rate (eGFR) for chronic kidney disease (CKD), and / or the Unified Parkinson's Disease Rating Scale (UPDRS) for Parkinson's disease (PD).
[0045] Other features can be used as model inputs and / or model outputs (e.g., in the training data). For example, model inputs and / or model outputs may include transition propensity, transition time, biological stage, physiological stage, and / or any dimension or variable that is consistent with the transition from health to disease. For example, for PD, model inputs may include data related to brain imaging, functional assessments by experts, diagnostic trajectories, social determinants of health, or any combination thereof. For example, for diabetes, model inputs may include data related to dietary recall, lifestyle factors, medical history, blood laboratory levels, or any combination thereof. Model outputs may include estimated time to diagnosis (e.g., for PD) and / or risk of developing a disease (e.g., diabetes). The risk of developing a disease may be stratified into high-risk, intermediate-risk, and low-risk cohorts.
[0046] With sufficient real-world data, models can be trained to analyze individuals who have or do not have a disease or condition. This enables the interpretability and / or transferability of information learned across healthcare systems and / or time, and allows models to make predictions about individuals who are not native to the data (e.g., people who have not yet been analyzed by the system or used as a training data source).
[0047] In some examples, Model 310 can be trained to selectively enrich a desired population. This can involve isolating a cohort of individuals along a relevant dimension (e.g., health status), precisely defining the cohort at or near specific points along that dimension, and / or identifying patterns that characterize the cohort. As described herein, this cohort can be selected for inclusion in clinical trials or other studies.
[0048] In some cases, the data used to generate training data or model input (e.g., laboratory results, imaging results, and / or other patient data) may lack consistency in format, language, and / or terminology. Mapping such data to a uniform or consistent format, language, and / or terminology can be beneficial for training and using the model.
[0049] In various examples, models can be used to determine when different therapeutic targets or interventions are relevant and / or effective. For instance, in NAFLD, some interventions that inhibit lipid accumulation can be administered very early, while interventions that inhibit factors that recruit immune cells or activate fibrotic cells can be applied in later stages of progression. One or more predictive biomarkers can be used (e.g., through models) to identify individuals at a particular stage of the disease and / or those more likely than others to experience beneficial or adverse effects from therapeutic targets or interventions.
[0050] In some embodiments, the systems and methods described herein can be used to generate and employ continuous time scales. In some embodiments, labels from subgroups can be used to train the model to provide labels for new or different subgroups. Training data can be generated by extracting PDFF (e.g., body fat fraction) from individuals with MRI data and / or by generating eGFR from individuals with standard laboratory tests. These generated or extracted labels or values can be used to train the model to learn how to predict these characteristics in other individuals based on more traditional data that is readily available or passively collected (e.g., data derived from electronic health records and / or surveys).
[0051] In some embodiments, two approaches can be used to identify and validate biomarkers. In the first approach, a biological study is employed where snRNA-seq data is generated from tissues collected across a spectrum from health to disease (e.g., multiple molecular stages), and then secreted proteins that may be detectable in the blood are identified. In the second approach, blood is collected from individuals recruited at different stages of a health-to-disease continuum. The blood can then be studied to identify or determine changes as the disease progresses. These studies can be biased, unbiased, proteomic, metabolomic, lipidomic, or any combination thereof.
[0052] Figure 4This is a schematic diagram of a method 400 for identifying a cohort of individuals for participation in a clinical trial or study, based on certain examples of using a predictive model (e.g., model 310). Health data, demographic data, and / or survey data (e.g., patient data 312 and survey data 314) of individual 412 are provided as input to predictive model 310. This data may be entered by individual 412, a physician, healthcare professional, clinical trial designer, or others via a user interface, or it may be entered via an electronic health record (EHR) system through automated data extraction. Model 310 may provide predictions of an individual's health status (e.g., a continuous representation of health 316), disease state, and / or disease risk as output. For example, model 310 may predict values or statuses of molecular markers (e.g., gene expression levels) across a spectrum from health to disease for individual 412. Similar predictions may be made for other individuals. In some examples, the spectrum from health to disease may capture the temporality of functional impairment without requiring longitudinal tracking of individuals.
[0053] In the depicted example, individual 412's condition 414 falls within a target state 416, which lies between a healthy state and a diseased state. Target state 416 may correspond to the intended target population for a clinical trial. For example, target state 416 may be selected to capture individuals most likely to respond to treatments associated with the clinical trial. In many chronic diseases, such as diabetes, IBD, or age-related macular degeneration (AMD), different therapies may be used for patients depending on disease severity (e.g., early, moderate, or severe). Once target state 416 is identified, model predictions are used to identify a cohort of individuals whose health condition (e.g., condition 414) falls within target state 416. Alternatively or additionally, as described herein, a probability distribution satisfying target state 416 or approximating a target distribution (e.g., risk probability distributions 210, 212, and / or 214) may be used to identify the cohort. As described herein, such individuals may be recruited for a clinical trial. Once individuals have been recruited, the clinical trial can be run by providing some or all of the recruited individuals with an intervention, therapy, or treatment 418. For example, treatment 418 could be administered to the first group of individuals in the test group, and a placebo could be administered to the second group of individuals in the control or placebo group. The results of the trial could be used to determine whether the treatment had a positive outcome for the individuals who received it.
[0054] In some examples, model 310 can provide predictions of disease risk for individual 412 and other candidate individuals considered for clinical trials. Disease risk can correspond to a transition from a healthy state to a disease state. Transitions can include, for example, early biokinetic transitions, intermediate biokinetic transitions, or late biokinetic transitions. Such biokinetic transitions can be a cascade with different characteristics in the early, intermediate, and late stages of disease. A cohort of individuals identified for the study can be used to determine a relationship 420 between molecular marker levels and biokinetic transitions between health and disease. Relationship 420 can define trends across all sampled individuals. In the depicted example, the target state 416 corresponds to an intermediate biokinetic transition. Because such individuals may be more likely to progress toward disease than healthy individuals, treatment effects may be more easily observed during the trial if they are present.
[0055] In embodiments, the systems and methods described herein (including the predictive model) can be implemented via a network (e.g., the Internet). For example, the predictive model can be accessed by a clinical study designer via a network, whereby the predictive model identifies or provides a cohort of individuals through a HIPAA-compliant and secure connection. Furthermore, the cohort can be provided in a double-blind format, enabling participants to be contacted for enrollment in the study while their identities remain confidential to the scientists running the study.
[0056] In various examples, the systems and methods described herein can be used to facilitate clinical trials, biomarker validation studies, or other research. Facilitating clinical trials (or other research) can include a variety of actions related to organizing or conducting a trial. Such actions can include, for example: recruiting (e.g., identified using the systems and methods described herein) a group of participants for a trial or study; enrolling or registering a group of participants in a trial or study; collecting biological specimens, imaging data, or other data as part of a trial or study; treating one or more trial or study participants with an intervention; applying a diagnosis to one or more participants; advising one or more participants on treatment decisions; advising one or more participants on diagnostic or screening tests; or any combination thereof.
[0057] Machine learning and neural networks
[0058] Machine learning is a method of teaching computers to learn and make decisions on their own without being explicitly programmed to perform a specific task. It involves feeding large amounts of data into a computer program, which then uses statistical analysis to identify patterns and relationships within the data. The goal is to enable the program to make predictions or decisions based on these patterns and relationships without being explicitly told how to do so.
[0059] Neural networks are machine learning algorithms inspired by the structure and function of the human brain. They consist of multiple interconnected layers of "neurons" (sometimes called nodes) that process and transmit information. Each neuron receives input from other neurons, processes the input, and passes it on to other neurons in the next layer.
[0060] In a neural network, a layer refers to a group of interconnected neurons. A neural network typically contains multiple layers, where the input layer receives raw data and the output layer produces the final prediction or decision. Between the input and output layers are one or more hidden layers that process the data and pass it to the next layer.
[0061] By training a neural network on a large dataset, the connections between neurons (called "weights") can be adjusted to improve the network's ability to make predictions or decisions. To train a neural network, data is fed through it, and the output is compared to the expected result. If the output is inaccurate, the weights are adjusted to reduce the error. This process is repeated many times, with the network continuously adjusting the weights to improve its accuracy. Once the network is fully trained, it can be used to make predictions or decisions on new data based on the patterns and relationships it has learned from the training data.
[0062] In various examples, "machine learning" can refer to a computer system applying certain techniques (e.g., pattern recognition and / or statistical inference techniques) to perform a specific task. Machine learning techniques (automated or otherwise) can be used to build data analysis models based on sample data (e.g., "training data") and to validate the models using validation data (e.g., "test data"). Sample and validation data can be organized as collections of records (e.g., "observations" or "data samples"), where each record indicates a value for a specified data field (e.g., "independent variable," "input," "feature," or "predictor") and a corresponding value for other data fields (e.g., "dependent variable," "output," or "target"). Machine learning techniques can be used to train a model to infer output values based on input values. When presented together with other data similar to or related to the sample data (e.g., "inference data"), such a model can accurately infer unknown values of the target for the inference dataset. Such a model may be referred to herein as a "machine learning model," a "predictive model," or a "computer-implemented model."
[0063] Features of a data sample can be measurable properties of an entity (e.g., cell, biological sample, person, thing, event, activity, etc.) represented or associated with the data sample. For example, a feature can be a characteristic of a cell in an organism. As another example, a feature can be the gene expression level associated with a cell. In some cases, features of a data sample are descriptions (or other information about) of an entity represented or associated with the data sample. The value of a feature can be a measurement of a corresponding attribute of the entity or an instance of information about the entity. For example, in the example above where the feature of a cell is gene expression level, the value of the feature 'expression of gene G' could be 215,000, which is the number of messenger RNA fragments in the cell that map to the region of gene G on the human reference genome. In some cases, the value of a feature can indicate a missing value (e.g., no value). For example, in the example above where the feature is gene expression level, the value of the feature could be 'NULL', indicating that the gene level cannot be measured by a given technique.
[0064] Features can also have data types. For example, features can have image data types, numeric data types, text data types (e.g., structured text data types or unstructured (“free” text data types), categorical data types, or any other suitable data type. In the example above, the feature of shape extracted from an image of a cell could be an image data type. Typically, the data type of a feature is categorical if the set of values that can be assigned to a feature is finite.
[0065] As used in this article, "developing" a machine learning model can refer to the construction of that model. Machine learning models can be constructed by a computer using a training dataset. Therefore, "developing" a machine learning model can include training it using a training dataset. In some cases (often referred to as "supervised learning"), the training dataset used to train the machine learning model may include known outcomes (e.g., labels or target values) for the individual data samples in the training dataset. For example, when training a supervised computer vision model to detect images of cats, the target value of a data sample in the training dataset could indicate whether the data sample contains an image of a cat. In other cases (often referred to as "unsupervised learning"), the training dataset does not include known outcomes for the individual data samples in the training dataset.
[0066] After development, the machine learning model can be used to generate inferences about a dataset of “inference” data. For example, after development, a computer vision model can be configured to distinguish between data samples containing images of cats and data samples that do not contain images of cats. As used in this paper, “deploying” a machine learning model can refer to using the developed machine learning model to generate inferences about data other than the training data.
[0067] Computer implementation
[0068] In some examples, some or all of the processes described above can be executed on a personal computing device, on one or more centralized computing devices, or by one or more servers via cloud-based processing. Some types of processing can occur on one device, while others can occur on another. Some or all of the data described above can be stored on a personal computing device, in a data storage device hosted on one or more centralized computing devices, and / or via cloud-based storage. Some data can be stored in one location, while others can be stored in another. In some examples, quantum computing and / or functional programming languages can be used. Electrical memory, such as flash memory, can be used.
[0069] Figure 5 This is a block diagram of an example computer system 500 that can be used to implement the techniques described in this document. General-purpose computers, network devices, mobile devices, or other electronic systems may also include at least a portion of system 500. System 500 includes a processor 510, memory 520, storage device 530, and input / output device 540. Each of components 510, 520, 530, and 540 may be interconnected, for example, using a system bus 550. Processor 510 is capable of processing instructions for execution within system 500. In some embodiments, processor 510 is a single-threaded processor. In some embodiments, processor 510 is a multi-threaded processor. Processor 510 is capable of processing instructions stored in memory 520 or on storage device 530.
[0070] The memory 520 stores information within the system 500. In some embodiments, the memory 520 is a non-transitory computer-readable medium. In some embodiments, the memory 520 is a volatile memory cell. In some embodiments, the memory 520 is a non-volatile memory cell.
[0071] Storage device 530 provides mass storage for system 500. In some embodiments, storage device 530 is a non-transitory computer-readable medium. In various embodiments, storage device 530 may include, for example, a hard disk drive, an optical disk drive, a solid-state drive, a flash drive, or some other mass storage device. For example, the storage device may store long-term data (e.g., database data, file system data, etc.). Input / output device 540 provides input / output operations for system 500. In some embodiments, input / output device 540 may include one or more of network interface devices (e.g., Ethernet cards), serial communication devices (e.g., RS-232 ports), and / or wireless interface devices (e.g., 802.11 cards, wireless modems (e.g., 3G, 4G, or 5G)). In some embodiments, input / output devices may include driver devices configured to receive input data and send output data to other input / output devices (e.g., keyboards, printers, and display devices 560). In some examples, mobile computing devices, mobile communication devices, and other devices may be used.
[0072] In some implementations, at least a portion of the methods described above can be implemented by instructions that, when executed, cause one or more processing devices to perform the processes and functions described above. Such instructions may include, for example, interpreted instructions (such as script instructions), executable code, or other instructions stored in a non-transitory computer-readable medium. Storage device 530 can be implemented in a distributed manner via a network, for example, as a server cluster or a widely distributed group of servers, or it may be implemented in a single computing device.
[0073] Although already Figure 5 Example processing systems are described herein, but embodiments of the subjects, functional operations, and processes described herein may be implemented in other types of digital electronic circuit systems, in tangibly embodied computer software or firmware, in computer hardware (including the structures disclosed herein and their structural equivalents), or in combinations thereof. Embodiments of the subjects described herein may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-volatile program carrier for execution by a data processing apparatus or for controlling the operation of a data processing apparatus. Alternatively or additionally, program instructions may be encoded on artificially generated propagated signals (e.g., machine-generated electrical, optical, or electromagnetic signals) generated for encoding information to be transmitted to a suitable receiver device for execution by the data processing apparatus. Computer storage media may be machine-readable storage devices, machine-readable storage substrates, random or serial access memory devices, or combinations thereof.
[0074] The term "system" can encompass all kinds of devices, apparatuses, and machines used for processing data, including, for example, programmable processors, computers, or multiple processors or computers. A processing system can include dedicated logic circuit systems, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, a processing system can also include code that creates the execution environment for the computer program in question, such as code that constitutes processor firmware, protocol stacks, database management systems, operating systems, or combinations thereof.
[0075] Computer programs (which may also be referred to or described as programs, software, software applications, engines, pipelines, modules, software modules, scripts, or code) can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program may, but does not necessarily, correspond to a file in a file system. A program may be stored as a portion of a file containing other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple co-located files (e.g., a file storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on a single computer or on multiple computers located at a single site or distributed across multiple sites and interconnected via a communication network.
[0076] The processes and logic flows described in this specification can be executed by one or more programmable computers, which execute one or more computer programs to perform functions by manipulating input data and generating output. These processes and logic flows can also be executed by special-purpose logic circuit systems (e.g., FPGAs (Field Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits)), and the apparatus can also be implemented as a special-purpose logic circuit system.
[0077] For example, a computer suitable for executing computer programs may include a general-purpose microprocessor or a special-purpose microprocessor or both, or any other type of central processing unit. Typically, the central processing unit receives instructions and data from read-only memory or random access memory or both. A computer typically includes a central processing unit for executing or carrying out instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., disks, magneto-optical disks, or optical disks) for storing data, or operatively coupled to receive data from or transfer data to said mass storage device, or both. However, a computer does not need to have such a device. Furthermore, a computer may be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name just a few.
[0078] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example: semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. Processors and memory may be supplemented or incorporated into dedicated logic circuitry systems.
[0079] To provide interaction with the user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. Additionally, the computer can interact with the user by sending and receiving documents from a device used by the user; for example, by sending a web page to a web browser in response to a request received from a web browser on the user's device.
[0080] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes back-end components (e.g., as a data server), or middleware components (e.g., an application server), or front-end components (e.g., a client computer having a graphical user interface or web browser through which a user can interact with embodiments of the subject matter described in this specification), or any combination of one or more such back-end, middleware, or front-end components. Components of the system can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include local area networks (“LANs”) and wide area networks (“WANs”) such as the Internet.
[0081] A computing system may include clients and servers. Clients and servers are typically geographically separated and usually interact through a communication network. The client-server relationship is established by computer programs running on the respective computers and having a client-server relationship with each other.
[0082] While this specification contains numerous specific implementation details, these details should not be construed as limiting the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed in this way, in some cases one or more features from a claimed combination may be removed from that combination, and the claimed combination may involve sub-combinations or variations thereof.
[0083] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in the specific order shown or in sequence, or requiring all shown operations to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0084] Specific embodiments of this subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims may be performed in a different order and still achieve the desired result. As an example, the process depicted in the drawings does not necessarily require the specific order or sequential sequence shown to achieve the desired result. In some embodiments, multitasking and parallel processing may be advantageous. Additional steps or stages may be provided, or steps or stages may be eliminated from the described process. Therefore, other embodiments are within the scope of the following claims.
[0085] Some embodiments
[0086] In some embodiments, a method for conducting a cross-sectional or longitudinal human study or trial includes: separating data from a database of records of multiple individuals into input data and output data, wherein the input data are static or temporal characteristics of the respective individuals collected passively or through active participation, and wherein the output data is a primary hierarchical axis representing the transition of the respective individual from relative health to relative dysfunction or disease. The method further includes: selecting one or more secondary hierarchical axes (e.g., sex, age, demographics, etc.) derived from data elements collected from the individuals' database of records. The method further includes: training a model using data from the individuals, wherein the input data and the output data (e.g., real output information) exist on these individuals or can be interpolated for each individual along the one or more hierarchical axes. The method further includes: deploying the model to generate a localized or probability distribution embedding for each individual along the one or more hierarchical axes using input data associated with each individual's health or information related to their current or future health status. The method further includes: stratifying each individual who can be recruited to the cross-sectional or longitudinal human study or trial based on the relationship between the respective collected data and the selected one or more axes. The method further includes: determining a target distribution of the patients to be recruited based on their position or probability distribution along a selected axis, wherein the position or probability of the patients along the selected axis is separate from the population distribution of the patients. The method further includes: selecting and outputting a subpopulation of the plurality of individuals that can be recruited to the cross-sectional or longitudinal human study or trial, the collective distribution of the subpopulation's position or probability along one or more selected axes expected to resemble the target distribution.
[0087] In some embodiments, the method further includes: training a model using data from individuals, wherein the input data and the actual output information exist on these individuals or can be interpolated along one or more hierarchical axes for each individual.
[0088] In some embodiments, the method further includes: deploying a model to generate a localized or probability distribution embedding for the individual along one or more hierarchical axes using input data associated with each individual's health or information related to their current or future health status.
[0089] In some embodiments, the method includes: collecting data from multiple individuals.
[0090] In some embodiments, the data of the plurality of individuals includes: demographic characteristics, medical visits, physical examinations, diagnoses, medications, procedures, vital sign measurements, vaccinations, representation of (multiple) laboratory results, collection and / or analysis of serum, urine or other biological samples, imaging, clinical notes, radiological reports, genetic testing or reports, biomarkers, pathology reports, other health-related record information, social determinants of health, financial status or behavior, or consumer status or behavior.
[0091] In some embodiments, the selected one or more hierarchical axes may be one or more primary hierarchical axes and / or may span from a relative health state to a relative disease state. In some embodiments, the one or more hierarchical axes may include one or more secondary hierarchical axes, and / or may be selected based on factors such as a known or hypothetical relationship between the one or more axes and the disease state of interest, or the availability of data elements corresponding to the one or more axes.
[0092] In some embodiments, the method further includes selecting one or more secondary hierarchical axes representing one or more attributes (including, for example, physical comorbidities, demographic characteristics, or socioeconomic variables) about the plurality of individuals.
[0093] In some embodiments, the human study or trial aims to investigate progress in health transition, validate biomarkers, validate biomarker groups, validate targets, or evaluate the efficacy or safety of pharmacological, digital, or other forms of intervention, action, or behavior modification in a population or subgroup. In some embodiments, the study of progress in health transition may be performed by a fully automated system, a semi-automated system, or by humans after generating automated indicators representing that progress.
[0094] In some embodiments, the selected subgroup is enriched on one or more attributes, including: reduced physiological heterogeneity, increased probability of outcome events, reduced heterogeneity of progression rate, increased propensity to respond to intervention, increased likelihood of observing specific physiological or pathological signs, or more broadly, increased homogeneity on the expected endpoint.
[0095] In some embodiments, the hierarchical axis represents a continuous or ordered dimension.
[0096] In some embodiments, the one or more hierarchical axes include a single data element or a combination of multiple data elements collected from multiple individuals.
[0097] In some embodiments, the position of an individual along one or more hierarchical axes can be represented by points, point estimates, or probability distributions.
[0098] In some embodiments, the method further includes: using interpolation techniques to estimate the location of individuals missing one or more relevant data elements, wherein the interpolation techniques may include multiple interpolation or regression models.
[0099] In some embodiments, the one or more hierarchical axes include learning parameters(s) derived from a model trained on data elements collected from individuals.
[0100] In some embodiments, the determined target distribution is selected based on attributes relating to the biological characteristics, biomarkers, biomarker groups, therapeutic targets, or pharmacological interventions to be evaluated.
[0101] In some embodiments, the target distribution is uniform about one or more hierarchical axes. In some embodiments, the method further includes: discovering or validating a biomarker or group of biomarkers. In some embodiments, the target distribution is enriched about a region or subregion of one or more hierarchical axes. In some embodiments, the method further includes: assessing the efficacy of an intervention. In some embodiments, the selected target distribution may differ significantly from the background distribution of the individual patients to be recruited, such that an unguided recruitment protocol intended to recruit patients constituting the target distribution would require a much larger number of recruits.
[0102] In some embodiments, the output subgroups include individuals whose distribution along one or more hierarchical axes is statistically similar to the target distribution.
[0103] In some embodiments, the selected subgroup enables rapid recruitment of individuals across dimensions that might otherwise require longitudinal studies. In some embodiments, the selected subgroup enables efficient recruitment of individuals within a desired window (e.g., a portion or subset of the dimension) within that dimension.
[0104] In some embodiments, the selected individual subgroups can be chosen to minimize or maximize certain attributes, such as minimizing the total cost or number of patients required to maintain minimum coverage in a continuous dimension.
[0105] In some embodiments, the selected individual subgroups may be chosen to maintain certain distributions or balances of covariates, including demographic characteristics, covariates, comorbidities, medication use, medical history, or other socioeconomic or biological attributes.
[0106] In some embodiments, patients are selected using optimization methods, including linear programming, gradient descent, dynamic programming, or other schemes for approximating the optimal subset of recruits.
[0107] In some embodiments, the output subgroup may report an empty output subgroup, and no subgroup exists that satisfies these target distributions and selection criteria. In some embodiments, in the case of an empty output subgroup being reported, the experiment designer is advised to refine or relax (multiple) certain target distributions or selection criteria that will return a non-empty output subgroup.
[0108] In some embodiments, the method further includes: outputting a list(s) of patients having comparable properties relative to the hierarchical axis, and the number of patients to be recruited from each list to satisfy a target distribution.
[0109] In some embodiments, anonymized tags are used to perform subgroup output of an individual.
[0110] In some embodiments, certain diseases have defining biomarkers that indicate the disease or its stage. In other diseases (such as liver disease), images can reveal how much fat or fibrosis is in the liver, while in Parkinson's disease, current diagnostic criteria are subjective measurements of an individual's tremors or sleep disturbances. Therefore, stratification can quantify what would otherwise be a subjectively defined symptom. This stratification can be a percentage group or a range of continuous scores. It can also be a discrete group.
[0111] In some embodiments, the stratification can be a risk stratification. In one embodiment, the method can be used to provide risk stratification for life insurance actuaries or other risk-based applications.
[0112] In some embodiments, the method can provide clinical decision support by offering physicians, nurses, registered nurses, physician assistants, or patients risk categories derived from a probability distribution output by the model. In some embodiments, physicians, nurses, registered nurses, physician assistants, or patients can access the model via the internet to determine a diagnosis or staging. In some embodiments, the method can be provided to physicians, nurses, registered nurses, physician assistants, or patients as a Software as a Service (SaaS) implementation. In some embodiments, the same probability distribution output by the model is used to simultaneously provide clinical decision support and clinical trial enrichment.
[0113] In some embodiments, the model is a machine learning model or a neural network. The method may further employ supervised machine learning techniques trained on ordered labels. In some embodiments, the predictive model is learned via ordered regression, neural networks, generative models, random forests, nearest neighbors, support vector machines, and Gaussian processes. Those skilled in the art will understand that other statistical, machine learning methods, and neural networks can be employed.
[0114] In some embodiments, changes in health status are represented by continuous or discrete dimensions.
[0115] In some embodiments, the method further includes: identifying or validating a biomarker or set of biomarkers that indicates a shift in health status, biological function, physiological function, disease, or multiple states. In some embodiments, the biomarker may be a paired biomarker that defines an intervention pairing of the biomarker. The method may determine that two biomarkers may be paired in indicating a symptom, or that a biomarker may be paired with an intervention. In some embodiments, the method further includes: using a model to assess whether a relationship exists between the biomarker or set of biomarkers and a shift in health status or disease. The model is a machine learning model or a neural network. In some embodiments, the method further includes: characterizing the relationship between the biomarker or set of biomarkers and a shift in health status or disease. The relationship includes the probability that the biomarker or set of biomarkers indicates a shift in health status, a confidence interval for the probability, and a time range of progression based on a range of values for the biomarker. In some embodiments, the biomarker or set of biomarkers is validated in human studies in which the participants are members of a subset identified by the model.
[0116] In some embodiments, the method further includes determining whether an arbitrary intervention is effective on a subgroup of model selection in a human study or trial. In some embodiments, the arbitrary intervention may include a pharmacological intervention, a behavioral intervention, a digital intervention, data collection, or other intentions or actions. In some embodiments, a digital intervention may be a mobile or computer application (e.g., an "app") that provides behavioral modification, information about compliance, or engagement with an application (e.g., entering health metrics, recording food or exercise data). In some embodiments, the human study or trial includes determining which stages of health status, biological function, physiological function, or disease a particular intervention is effective. In embodiments, the arbitrary intervention may be a method that determines the relevance of the intervention to changes in an individual's existing medical records and biomarkers.
[0117] In some embodiments, the method further includes enriching human studies or trials by selecting subgroups at specific stages of health, biological function, physiological function, or disease, so that the effects of the tested intervention can be determined more precisely.
[0118] In some embodiments, the method further includes enriching human studies or trials by selecting subgroups based on predicted progress, thereby enabling the determination of the effects of the tested intervention.
[0119] In some embodiments, the predicted progression is a known or existing medical consensus staging.
[0120] In some embodiments, the method further includes: conducting human studies or trials to evaluate the efficacy and safety of a protocol for using and applying a model-guided intervention targeting a subgroup.
[0121] In some embodiments, the model is configured to identify individuals in a subgroup who should receive intervention, when the identified individuals should receive intervention, or what intervention dose or regimen should be used.
[0122] In some embodiments, the output individual subgroups are anonymized.
[0123] In some embodiments, the model is a machine learning model or a neural network.
[0124] In some embodiments, the method further includes: enrolling the subgroup into clinical trials based on the stratification. In some embodiments, the method further includes: retraining the model based on the results of the clinical trials.
[0125] In some embodiments, the method further includes: treating the subgroup with a therapy based on the stratification. In some embodiments, the method further includes: retraining the model based on the results of clinical trials.
[0126] In some embodiments, a method for training a machine learning model for running cross-sectional or longitudinal human studies or trials includes: separating data from a database of records of multiple individuals into input data and output data. The input data consists of static or temporal characteristics of the respective individuals, collected passively or through active participation. The output data represents data representing the transition of the respective individual from health to dysfunction or disease. The method includes: defining one or more hierarchical axes derived from one or more data elements collected from the database of records of the multiple individuals. The method further includes: training a model using data from individuals, the input data and the output data (e.g., real output information) existing on these individuals or interpolated for each individual along the one or more hierarchical axes, the trained model being configured to be deployed to: (i) generate a localized or probability distribution embedding for each individual along the one or more hierarchical axes using input data associated with each individual's health or information related to their current or future health status; (ii) hierarchically stratify each of the plurality of individuals that can be recruited to the cross-sectional or longitudinal human study or trial based on the relationship between the corresponding collected data and the selected one or more axes; (iii) determine a target distribution of these subjects based on the position or probability distribution of the subjects to be recruited along the selected axes as defined by the goals of the study or trial, the position or probability of these subjects along the selected axes being separate from the population distribution of patients; and (iv) selecting and outputting a subgroup of the plurality of individuals that can be recruited for the cross-sectional or longitudinal human study or trial, the collective distribution of the position or probability of the subgroup along the selected one or more axes being expected to resemble the target distribution.
[0127] In some embodiments, a method of running a model trained to select a population for a cross-sectional or longitudinal human study or trial includes: deploying the model to generate a localized or probability distribution embedding for each individual along one or more hierarchical axes using input data associated with each individual's health or information related to their current or future health status. The model is trained on a database of records of individuals, which is separated into input and output data. The input data are static or temporal features of the respective individuals, collected passively or through active participation. The output data is data representing the transition of the respective individual from health to dysfunction or disease. The hierarchical axes may be defined based on one or more data elements collected from the database of records of the plurality of individuals. The model is trained using data from the individuals, the input data and the output data (e.g., real output information) existing on these individuals or interpolated for each individual along one or more hierarchical axes. Each individual among the plurality of individuals who may be recruited to the cross-sectional or longitudinal human study or trial is hierarchically stratified based on the relationship between the corresponding collected data and the selected one or more axes. The method further includes: determining a target distribution of subjects to be recruited based on the position or probability distribution of subjects along a selected axis as defined by the goals of the study or trial, the position or probability of these subjects along the selected axis being separate from the population distribution of patients. The method further includes: selecting and outputting a subgroup of the plurality of individuals that can be recruited to the cross-sectional or longitudinal human study or trial, the collective distribution of the subgroup's position or probability along one or more selected axes expected to resemble the target distribution.
[0128] the term
[0129] The wording and terminology used in this article are for descriptive purposes and should not be considered restrictive.
[0130] As used in the specification and claims, the term "approximately," the phrase "approximately equal to," and other similar phrases (e.g., "X has a value approximately equal to Y" or "X is approximately equal to Y") should be understood to mean that one value (X) is within a predetermined range of another value (Y). The predetermined range may be positive or negative 20%, 10%, 5%, 3%, 1%, 0.1%, or less than 0.1%, unless otherwise stated.
[0131] Measurements, sizes, quantities, etc., may be presented in range format herein. The range format is for convenience and brevity only and should not be construed as an immutable limitation on the scope of the invention. Therefore, a description of a range should be considered as having specifically disclosed all possible subranges and the individual values within that range. For example, a description of a range such as 10 to 20 inches should be considered as having specifically disclosed subranges, such as 10 to 11 inches, 10 to 12 inches, 10 to 13 inches, 10 to 14 inches, 11 to 12 inches, 11 to 13 inches, etc.
[0132] Unless explicitly indicated to the contrary, the indefinite article “a / an” as used in the specification and claims shall be understood to mean “at least one / an”. The phrase “and / or” as used in the specification and claims shall be understood to mean “any one or both” of the elements so linked (i.e., elements that exist together in some cases and separately in others). Multiple elements listed with “and / or” shall be interpreted in the same way, i.e., “one or more” of the elements so linked. In addition to the elements specifically indicated by the “and / or” clause, other elements may optionally be present, whether related to or unrelated to those specifically indicated. Thus, as a non-limiting example, when used in conjunction with open-ended language such as “comprising”, a reference to “A and / or B” may, in one embodiment, refer only to A (optionally including elements other than B); in another embodiment, refer only to B (optionally including elements other than A); in yet another embodiment, refer to both A and B (optionally including other elements); and so on.
[0133] As used in the specification and claims, "or" should be understood to have the same meaning as "and / or" as defined above. For example, when separating items in a list, "or" or "and / or" should be interpreted as inclusive, i.e., including multiple elements or at least one of the elements in the list, but also including more than one element, as well as optional unlisted items. Only terms that explicitly indicate the opposite meaning (such as "only one" or "exactly one," or "consisting of..." when used in a claim) will refer to including multiple elements or exactly one of the elements in the list. Generally, when preceded by an exclusive term (such as "any one," "one of," "only one," or "exactly one"), the term "or" should be interpreted only to indicate an exclusive alternative (i.e., "one or the other but not both"). When used in a claim, "consisting substantially of..." should have its usual meaning as used in the field of patent law.
[0134] As used in the specification and claims, when referring to a list of one or more elements, the phrase "at least one" should be understood to mean at least one element selected from any one or more elements in the list, but not necessarily including at least one of each element specifically listed in the list, and does not exclude any combination of elements in the list. This definition also allows for the optional presence of elements other than those specifically indicated in the list of elements referred to by the phrase "at least one," whether related to or unrelated to those specifically indicated elements. Thus, as a non-limiting example, "at least one of A and B" (or equivalently, "at least one of A or B," or equivalently, "at least one of A and / or B") in one embodiment may refer to at least one, optionally including more than one A, without B (and optionally including elements other than B); in another embodiment, it may refer to at least one, optionally including more than one B, without A (and optionally including elements other than A); in yet another embodiment, it may refer to at least one, optionally including more than one A, and at least one, optionally including more than one B (and optionally including other elements); and so on.
[0135] The use of “including,” “comprising,” “having,” “containing,” “involving,” and their variations is intended to cover the items listed thereafter and any additional items.
[0136] The use of ordinal terms such as "first," "second," and "third" to modify claim elements in claims does not imply any priority, order, or sequence of one claim element relative to another, nor does it indicate the chronological order of the actions of the method. Ordinal terms are merely labels used to distinguish one claim element with a certain name from another element with the same name (but using an ordinal term), thereby differentiating these claim elements.
[0137] While exemplary embodiments have been specifically shown and described, those skilled in the art will understand that various changes in form and detail may be made therein without departing from the scope of the embodiments covered by the appended claims.
[0138] Claims:
Claims
1. A method for conducting clinical trials or biomarker validation studies, the method comprising: Gain access to a computer-implemented model trained using training data comprising multiple records of multiple individuals, each record comprising (i) values of multiple features for a corresponding individual from the multiple individuals, and (ii) a label providing an indication of the disease state of that corresponding individual; The trained computer-implemented model is provided with multiple sets of values for the multiple features, each set of values corresponding to a candidate individual from a set of candidate individuals; The model, which has been trained by a computer, receives predictions of the disease state for each set of values. A group of participants was identified from the candidate individual set based on the prediction of the disease state; as well as Facilitate at least one of the clinical trials or biomarker validation studies involving this group of participants.
2. The method as described in claim 1, wherein, The values of these multiple features describe at least one of the following: demographic characteristics, medical visits, physical examinations, diagnoses, medications, medical procedures, vital sign measurements, vaccinations, laboratory results, serum, urine samples, biological samples, gene expression levels, medical images, clinical notes, radiological reports, genetic testing, biomarkers, pathology reports, health information, social determinants of health, financial data, consumer data, or any combination thereof.
3. The method as described in claim 1, wherein, The model includes at least one of the following: regression model, classifier, linear model, nonlinear model, random forest, kernel method, Bayesian model, decision tree, or neural network.
4. The method of claim 1, wherein, The prediction of this disease state includes a probability distribution.
5. The method of claim 1, wherein, The prediction of this disease state is based on a primary stratification axis, which includes a health-disease axis or a disease risk axis.
6. The method of claim 1, wherein, Identifying the participants in this group involves stratifying each candidate from the plurality of candidates according to one or more primary stratification axes.
7. The method of claim 6, wherein, Identifying the group of participants further includes stratifying each candidate individual from the plurality of candidate individuals according to one or more secondary stratification axes, wherein the one or more secondary stratification axes represent one or more attributes of the plurality of individuals, including at least one of physiological comorbidities, demographic characteristics, or socioeconomic variables.
8. The method of claim 7, wherein, The one or more primary hierarchical axes or at least one of the one or more secondary hierarchical axes represent at least one of continuous dimensions or ordered dimensions.
9. The method of claim 7, wherein, The location along one or more primary hierarchical axes or at least one of the one or more secondary hierarchical axes is represented by a point, a point estimate, or a probability distribution.
10. The method of claim 1, wherein, The clinical trial or biomarker validation study includes studying at least one of the following: progress in health transformation, effectiveness of the biomarker, effectiveness of a biomarker group, effectiveness of targeted therapy, efficacy of therapeutic intervention, efficacy of digital intervention, efficacy of behavioral intervention, or any combination thereof.
11. The method of claim 1, wherein, The participants in this group are enriched on one or more attributes relative to the candidate set of individuals, including at least one of the following: reduced physiological heterogeneity, increased probability of an outcome event, increased propensity to respond to an intervention, increased likelihood of observing a specific physiological or pathological sign, or any combination thereof.
12. The method of claim 1, wherein, The participants in this group were identified as including: The target distribution of the candidate individual set is determined based on the objectives of the clinical trial or the biomarker validation study, and this target distribution defines a probability distribution; and Select the group of participants from the candidate set to achieve a collective distribution of the group of participants similar to the target distribution.
13. The method of claim 12, wherein, The target distribution is determined based on at least one parameter to be evaluated in the clinical trial or the biomarker validation study, the at least one parameter including at least one of biological properties, biomarkers, biomarker groups, therapeutic targets or pharmacological interventions.
14. The method of claim 12, wherein, The distribution of the target is uniform with respect to the disease state.
15. The method of claim 12, wherein, The target distribution includes a risk probability distribution.
16. The method of claim 12, wherein, The target distribution is regionally enriched with respect to the hierarchical axis associated with the disease state.
17. The method of claim 12, wherein, Selecting this group of participants involves minimizing the number of individuals in the group.
18. The method of claim 12, wherein, The participants selected for this group include: Identify multiple groups of individuals, where each group has a similar probability distribution regarding the disease state; and Determine the number of individuals in each group that should be included in the group to satisfy the target distribution.
19. The method of claim 1, wherein, The method includes facilitating the clinical trial, wherein the clinical trial includes evaluating the efficacy of the treatment.
20. The method of claim 1, wherein, The method includes facilitating the clinical trial, wherein facilitating the clinical trial includes enrolling members from the group of participants in the clinical trial, enabling the administration of drugs, digital interventions, or behavioral interventions to these members.
21. The method of claim 1, wherein, The method includes facilitating biomarker validation studies, wherein the biomarker validation studies include validating a biomarker or a group of biomarkers.
22. The method of claim 1, wherein, Identifying the participants involves achieving the desired distribution of at least one covariate, which includes at least one of demographic characteristics, comorbidities, medication use, medical history, socioeconomic attributes, or biological attributes.
23. The method of claim 1, further comprising using the training data to train the computer-implemented model.
24. A system for performing clinical trials or biomarker validation studies, the system comprising one or more computer processors programmed to perform operations including: Gain access to a computer-implemented model trained using training data comprising multiple records of multiple individuals, each record comprising (i) values of multiple features for a corresponding individual from the multiple individuals, and (ii) a label providing an indication of the disease state of that corresponding individual; The trained computer-implemented model is provided with multiple sets of values for the multiple features, each set of values corresponding to a candidate individual from a set of candidate individuals; The model, which has been trained by a computer, receives predictions of the disease state for each set of values. A group of participants was identified from the candidate individual set based on the prediction of the disease state; as well as Facilitate at least one of the clinical trials or biomarker validation studies involving this group of participants.
25. The system of claim 24, wherein, The prediction of this disease state is based on a primary stratification axis, which includes a health-disease axis or a disease risk axis.
26. The system of claim 24, wherein, Identifying the participants in this group involves stratifying each candidate from the plurality of candidates according to one or more primary stratification axes and one or more secondary stratification axes.
27. The system of claim 24, wherein, The participants in this group were identified as including: The target distribution of the candidate individual set is determined based on the objectives of the clinical trial or the biomarker validation study, and this target distribution defines a probability distribution; and Select the group of participants from the candidate set to achieve a collective distribution of the group of participants similar to the target distribution.
28. The system of claim 27, wherein, The target distribution includes a risk probability distribution.
29. The system of claim 27, wherein, The participants selected for this group include: Identify multiple groups of individuals, where each group has a similar probability distribution regarding the disease state; and Determine the number of individuals in each group that should be included in the group to satisfy the target distribution.
30. The system of claim 24, wherein these operations further include using the training data to train the computer-implemented model.