Prediction of the progression of cognitive impairment

Predictive models using clinical trial data address the limitations of invasive cerebral amyloid detection, enabling accurate forecasting of cognitive impairment and amyloid status for improved patient management and clinical trial efficiency.

JP2026524129APending Publication Date: 2026-07-17EISAI R&D MANAGEMENT CO LTD

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
EISAI R&D MANAGEMENT CO LTD
Filing Date
2023-11-29
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Conventional methods for detecting cerebral amyloid status in patients require the administration of radioactive tracers, limiting widespread screening and increasing the difficulty of clinical trials by requiring large control and treatment groups, which can lead to patient exclusion.

Method used

A system and method for training predictive models using clinical trial data to predict cognitive impairment progression or cerebral amyloid status, utilizing baseline data including demographic, cognitive, genomic, image, and biomarker data to develop machine learning models that can forecast cognitive decline and amyloid status.

Benefits of technology

Enables accurate prediction of cognitive impairment progression and cerebral amyloid status without invasive procedures, facilitating patient selection and improving clinical trial efficiency by reducing the need for large control groups and enhancing treatment planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026524129000001_ABST
    Figure 2026524129000001_ABST
Patent Text Reader

Abstract

A platform is disclosed for training and deploying statistical or machine learning models to predict the progression of cognitive impairment or cerebral amyloid status. The statistical or machine learning models can be trained to predict the progression of cognitive impairment or cerebral amyloid status using baseline image data, biomarker data, genomic data, demographic data, cognitive data, etc. The platform can support acquiring training data, training statistical or machine learning models, and responding to prediction requests using the trained statistical or machine learning models.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-reference of related applications This application claims the benefits of U.S. Provisional Patent Application No. 63 / 513,799, filed on 14 July 2023, and U.S. Provisional Patent Application No. 63 / 593,433, filed on 26 October 2023. The provisional applications identified above are incorporated herein by reference in their entirety.

[0002] This disclosure relates to training machine learning or statistical models to predict the progression of cognitive impairment or brain amyloid status in individual subjects. [Background technology]

[0003] Patients with neurological disorders, functional impairments, or injuries may exhibit significant differences in the progression of cognitive impairment. These differences may depend on their baseline clinical and biological characteristics. These differences in the progression of cognitive impairment can limit the ability of physicians, caregivers, and patients to make appropriate decisions and plans regarding treatment and long-term care. Furthermore, such differences may increase the number of subjects required in control and treatment groups in clinical trials for treatments for such neurological disorders, functional impairments, or injuries. Apart from increasing the difficulty of such trials, the increased requirement for control groups may lead to rejection by patients who later prove to be eligible for treatment benefits. [Overview of the Initiative] [Problems that the invention aims to solve]

[0004] Conventional methods for detecting cerebral amyloid status in patients may require the administration of a radioactive tracer to the target and subsequent collection of imaging data (e.g., for PET scans). These significant requirements can prevent widespread screening for a patient's cerebral amyloid status. [Means for solving the problem]

[0005] A system and method for training a predictive model to predict the progression of cognitive impairment or cerebral amyloid status in a subject are disclosed. The predictive model can be trained using control data from clinical trials. Predictions of the progression of cognitive impairment or cerebral amyloid status in a subject can be used to manage the care of the subject, for patient selection and enrichment, or as a prognostic covariate in future clinical trials.

[0006] The disclosed embodiments include a system comprising at least one computer-readable non-temporary medium comprising at least one processor and instructions. Instructions can cause the system to perform an operation when executed by at least one processor. An operation may include acquiring training data for a first subject that meets a cognitive impairment condition. The training data may include baseline data and cognitive impairment progression data for each first subject. The baseline data for a first subject may include baseline cognitive data and image data. The image data may include one or more measurements for one or more brain regions identified as hubs or one or more composite values ​​for one or more clusters of brain regions identified as structural brain network modules, the hubs or modules can be identified using network analysis or multilevel clustering. The cognitive impairment progression data for a first subject may include repeated measurements acquired over time, which are repeated measurements acquired following the baseline cognitive data. The operation may further include training a machine learning model using the training data to predict cognitive impairment progression data for a first subject using the baseline data for the first subject. The operation may further include acquiring baseline data for a second subject that meets a cognitive impairment condition. The operation may further include predicting cognitive impairment progression data for a second subject by inputting baseline data for the second subject into a trained predictive model.

[0007] The disclosed embodiments include a system comprising at least one computer-readable non-temporary medium comprising at least one processor and instructions. Instructions can cause the system to perform an operation when executed by at least one processor. An operation may include acquiring training data for a first subject that meets a cognitive impairment condition. The training data may include baseline data and cognitive impairment progression data for each first subject. The baseline data for the first subject may include plasma biomarker data. The operation may further include training a predictive model using the training data to predict cognitive impairment progression data for the first subject using the baseline data for the first subject. The operation may further include acquiring baseline data for a second subject that meets a cognitive impairment condition. The operation may further include predicting cognitive impairment progression data for the second subject by inputting the baseline data for the second subject into the trained predictive model.

[0008] The disclosed embodiments include a system comprising at least one computer-readable non-transient medium comprising at least one processor and instructions. Instructions can cause the system to perform an operation when executed by at least one processor. An operation may include acquiring training data for a first subject that meets a cognitive impairment condition. The training data may include baseline data and brain amyloid data for each first subject. The baseline data may include plasma biomarker data. The operation may further include training a predictive model using the training data to predict the brain amyloid state for the first subject using baseline data for the first subject. The operation may further include acquiring baseline data for a second subject that meets a cognitive impairment condition. The operation may further include predicting the brain amyloid state for the second subject by inputting baseline data for the second subject into the trained predictive model.

[0009] The disclosed embodiments further include a computer-readable non-temporary medium containing instructions for designing a system to perform the operations listed above, and a method for corresponding to the operations listed above.

[0010] The general explanation above and the detailed explanation below are illustrative and explanatory only and do not limit the scope of the claims.

[0011] The accompanying drawings, incorporated into and constituting part thereof, are useful in illustrating and illustrating the principles of various exemplary embodiments, together with the description. [Brief explanation of the drawing]

[0012] [Figure 1A] Figure 1A shows the variability in the progression of individual cases of Alzheimer's disease. [Figure 1B] Figure 1B shows the variability in the progression of individual Alzheimer's disease cases. [Figure 1C] Figure 1C shows the variability in the progression of individual Alzheimer's disease cases. [Figure 1D] Figure 1D shows the variability in the progression of individual Alzheimer's disease cases. [Figure 2] Figure 2 shows an exemplary platform for developing, validating, and deploying predictive models for predicting the progression of cognitive impairment or cerebral amyloid status, consistent with the disclosed embodiments. [Figure 3] Figure 3 shows an exemplary process for predicting the progression of cognitive decline in a subject, consistent with the disclosed embodiments. [Figure 4] Figure 4 illustrates an exemplary process for identifying hubs and modules in the brain based on regional measurements, consistent with the disclosed embodiments. [Figure 5] Figure 4 shows an exemplary process for predicting the brain amyloid state of a subject, consistent with the disclosed embodiments. [Figure 6A] Figure 6A shows the results of a study on predicting the progression of cognitive impairment and disease in amyloid-positive subjects with mild cognitive impairment. [Figure 6B] Figure 6B relates to an investigation regarding the prediction of cognitive decline and disease progression in amyloid-positive subjects with mild cognitive impairment. [Figure 6C] Figure 6C relates to an investigation regarding the prediction of cognitive decline and disease progression in amyloid-positive subjects with mild cognitive impairment. [Figure 6D] Figure 6D relates to an investigation regarding the prediction of cognitive decline and disease progression in amyloid-positive subjects with mild cognitive impairment. [Figure 6E] Figure 6E relates to an investigation regarding the prediction of cognitive decline and disease progression in amyloid-positive subjects with mild cognitive impairment. [Figure 6F] Figure 6F relates to an investigation regarding the prediction of cognitive decline and disease progression in amyloid-positive subjects with mild cognitive impairment. [Figure 6G] Figure 6G relates to an investigation regarding the prediction of cognitive decline and disease progression in amyloid-positive subjects with mild cognitive impairment. [Figure 6H] Figure 6H relates to an investigation regarding the prediction of cognitive decline and disease progression in amyloid-positive subjects with mild cognitive impairment. [Figure 7A] Figure 7A relates to an investigation regarding the prediction of brain amyloid-β status using a blood-based test. [Figure 7B] Figure 7B relates to an investigation regarding the prediction of brain amyloid-β status using a blood-based test. [Figure 7C] Figure 7C relates to an investigation regarding the prediction of brain amyloid-β status using a blood-based test. [Figure 7D] Figure 7D relates to an investigation regarding the prediction of brain amyloid-β status using a blood-based test. [Figure 7E] Figure 7E relates to an investigation regarding the prediction of brain amyloid-β status using a blood-based test. [Figure 7F] Figure 7F relates to an investigation regarding the prediction of brain amyloid-β status using a blood-based test. [Figure 7G]Figure 7G shows a study on predicting brain amyloid-beta status using blood-based tests. [Figure 7H] Figure 7H shows a study on predicting brain amyloid-beta status using blood-based tests. [Figure 7I] Figure 7I shows a study on predicting brain amyloid-beta status using blood-based tests. [Figure 8A] Figure 8A relates to a study on predicting the progression of cognitive impairment. [Figure 8B] Figure 8B shows the results of a study on predicting the progression of cognitive impairment. [Figure 8C] Figure 8C relates to a study on predicting the progression of cognitive impairment. [Figure 8D] Figure 8D shows the results of a study on predicting the progression of cognitive impairment. [Figure 8E] Figure 8E relates to a study on predicting the progression of cognitive impairment. [Figure 8F] Figure 8F relates to a study on predicting the progression of cognitive impairment. [Figure 8G] Figure 8G relates to a study on predicting the progression of cognitive impairment. [Figure 8H] Figure 8H shows the results of a study on predicting the progression of cognitive impairment. [Figure 8I] Figure 8I shows the results of a study on predicting the progression of cognitive impairment. [Figure 8J] Figure 8J shows the results of a study on predicting the progression of cognitive impairment. [Figure 8K] Figure 8K shows the results of a study on predicting the progression of cognitive impairment. [Figure 8L] Figure 8L relates to a study on predicting the progression of cognitive impairment. [Figure 8M] Figure 8M shows the results of a study on predicting the progression of cognitive impairment. [Figure 8N] Figure 8N shows the results of a study on predicting the progression of cognitive impairment. [Figure 8O] Figure 80 shows the results of a study on predicting the progression of cognitive impairment. [Figure 8P]Figure 8P shows the results of a study on predicting the progression of cognitive impairment. [Figure 8Q] Figure 8Q relates to a survey on predicting the progression of cognitive impairment. [Figure 9A] Figure 9A shows a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 9B] Figure 9B shows the results of a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 9C] Figure 9C shows a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 9D] Figure 9D shows the results of a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 9E] Figure 9E shows a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 9F] Figure 9F shows the results of a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 9G] Figure 9G shows the results of a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 9H] Figure 9H shows the results of a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 9I] Figure 9I shows the results of a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 9J] Figure 9J shows the results of a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 9K] Figure 9K shows a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 9L] Figure 9L shows the results of a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 9M] Figure 9M shows the results of a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 10A]Figure 10A shows a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 10B] Figure 10B shows a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 10C] Figure 10C shows a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 10D] Figure 10D shows a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 10E] Figure 10E shows a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 10F] Figure 10F shows a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 10G] Figure 10G shows a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 10H] Figure 10H shows a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 10I] Figure 10I shows a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 10J] Figure 10J shows a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 10K] Figure 10K shows a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 10L] Figure 10L relates to a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 10M] Figure 10M shows a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 10N] Figure 10N shows a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 10O] Figure 100 shows a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 10P] Figure 10P shows a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 10Q] Figure 10Q relates to a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 10R] Figure 10R shows the results of a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 10S] Figure 10S shows the results of a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 10T] Figure 10T shows the results of a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 10U] Figure 10U relates to a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 10V] Figure 10V shows a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 10W] Figure 10W relates to a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 10X] Figure 10X shows a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 10Y] Figure 10Y shows the results of a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 10Z] Figure 10Z shows a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Figure 10AA] Figure 10AA shows a study on predicting the progression of cognitive impairment in early Alzheimer's disease. [Modes for carrying out the invention]

[0013] The following detailed description refers to the accompanying drawings. Where possible, the same reference numerals are used in the drawings and the following description to refer to the same or similar parts. While several exemplary embodiments are described herein, modifications, adaptations, and other implementations are possible. For example, substitutions, additions, or modifications may be made to components illustrated in the drawings, and the exemplary methods described herein may be modified by steps of substitution, rearrangement, removal, or addition to the methods disclosed herein. Therefore, the following detailed description is not limited to the disclosed embodiments and examples. Instead, the appropriate scope is defined by the appended claims.

[0014] Individual disease progression trajectories can vary significantly among subjects with neurological disorders, functional impairments, or injuries, depending on their baseline clinical and biological characteristics. Figures 1A and 1B show the distribution of baseline cognitive measures for individual subjects in the first and second clinical trials, respectively, while Figures 1C and 1D show repeated cognitive measures for individual subjects in the placebo arm of the first and second clinical trials over time, respectively. As is evident from these figures for subjects with early Alzheimer's disease (AD), the progression of cognitive impairment varies considerably among subjects. This variability in cognitive impairment can limit the ability of physicians, caregivers, and subjects to make appropriate decisions and plans regarding treatment and long-term care.

[0015] The disclosed embodiments include predictive models suitable for predicting the progression of cognitive impairment in AD subjects. The predictive models can be designed to take demographic data, cognitive data (e.g., one or more cognitive measures), genomic data, image data, biomarker data, or other suitable baseline data as input. The predictive models can provide, as outputs, predicted cognitive assessment scores (e.g., at a predetermined post-baseline interval, a post-baseline interval specified in the request, or any other suitable post-baseline period). A trained predictive model consistent with the disclosed embodiments may enable physicians, caregivers, and subjects to make appropriate decisions and plans regarding treatment and long-term care. Furthermore, the predictive models can be used to select appropriate patients for clinical trials or to create prognostic covariate data suitable for use in clinical trials.

[0016] The disclosed embodiments further include predictive models suitable for predicting the probability of detecting brain Aβ based on blood biomarkers. Such predictive models can be used as a screening tool for detecting brain amyloid load (for example, as part of screening for clinical trials or patient care management).

[0017] The disclosed predictive model can be trained using historical placebo-controlled data from clinical trials. Clinical trials provide a particularly strong training dataset, as subjects were evaluated at multiple point in time and screened before trial enrollment, as understood by the inventors. Screening limited subjects to those likely to have cognitive impairment associated with early-stage AD. The predictive power of the predictive model can be improved by limiting the training data to subjects with similar etiologies and stages of AD progression. Furthermore, brain imaging and biomarker data are available for a substantial proportion of these subjects and provide another source of input data for predicting cognitive impairment progression. In some embodiments, the disclosed predictive model can be trained using other datasets (e.g., research datasets).

[0018] Neurological disorders, impairments, or injuries may include any condition affecting the central nervous system that results in motor, cognitive, or behavioral impairment. Neurological disorders, impairments, or injuries may include diseases of the central nervous system (e.g., Alzheimer's disease, dementia, etc.), neurological disorders (e.g., mild cognitive impairment (MCI), etc.), or injuries (e.g., stroke, traumatic brain injury, etc.).

[0019] Predictive models may include statistical and machine learning models suitable for identifying relationships between input data and output results. Such models may include regression models (e.g., logistic regression models; ridge, lasso, or elastic network regression models; time series regression models, etc.), support vector machines, Bayesian classifiers, neural networks, decision trees, random forests, ensemble models, or other suitable statistical and machine learning models.

[0020] In some embodiments, preferred prediction models may include regularized logistic regression models and ensemble tree-based models. Regularized logistic regression models may include Bayesian elastic network models. Such models can prevent overfitting and increase model robustness by using a mixture of mixed double exponential functions before reducing model complexity. Ensemble tree-based models may include stochastic gradient boosting models that combine predictions from multiple decision trees to create a final prediction. The nodes in each of the multiple decision trees can be trained using a different random subset of input features. Therefore, individual decision trees may be different and potentially capture different signals from the data.

[0021] Predictive models can be trained using training cohorts of subjects with neurological disorders, functional impairments, or injuries. Neurological disorders may include central nervous system disorders, such as Alzheimer's disease or dementia. Neurological disorders may include mild cognitive impairment (MCI), etc. Neurological injuries may include stroke, traumatic brain injury, etc.

[0022] The training cohort may include subjects that meet the screening criteria. The screening criteria may be intended to limit the training cohort to subjects with the same neurological disorder, impairment, or injury (or a preferred combination of disorder, impairment, or injury). For example, as described herein, subjects can be screened for mild cognitive impairment or cerebral amyloid load. Limiting the training cohort to similarly positioned subjects may improve the performance of the predictive model.

[0023] A predictive model can be trained to predict the progression of cognitive impairment for an individual subject using input data for that individual subject. Input data can be acquired on or by the baseline date. If different parts of the input data are acquired on different days, one of these days may be selected as the baseline date (e.g., the day immediately preceding the baseline date, the day corresponding to the most time-sensitive or variable component of the input data, the date of acquisition of cognitive scales, the date of acquisition of image data, the date of acquisition of biomarker data, or several supplementary days based on the dates when two or more components of the input data were acquired, e.g., the imaging and cognitive day, or another preferred baseline selection day). Demographic information may be acquired on or before the baseline date. The predictive model can output predicted scores for assessment of cognitive and functional abilities (e.g., at a given post-baseline interval, a post-baseline interval specified in the request, or any other preferred post-baseline period), e.g., Clinical Dementia Severity Decision Scale (CDR-SB) scores, or other assessments described herein.

[0024] The predicted score can be absolute or relative. For example, the predicted score could be the predicted CDR-SB for the subject, or the predicted change in the CDR-SB score from the baseline CDR-SB value.

[0025] Predictive models can be trained to predict the progression of cognitive impairment for individual patients. In some embodiments, the predictive model can be designed to output an array of predicted scores. For example, assuming baseline input data, the predictive model can be designed to output predicted scores at months 3, 6, 9, 12, 15, 18, 21, 24, 27, 30, 33, 36, 48, or 60. In some embodiments, the input data for the predictive model may include duration or elapsed time. The predictive model can be designed to output predicted scores over that duration or elapsed time. As can be understood, such a model may output different predicted scores over different durations or elapsed times. For example, assuming the same baseline input data, the predictive model may predict a greater cognitive decline at month 18 than at month 3.

[0026] A predictive model can be trained and designed to predict the progression of cognitive impairment in a subject using baseline data on that subject. Baseline data may include baseline cognitive data. Baseline cognitive data can be obtained from the subject, the clinician observing or treating the subject, or another person close to the subject.

[0027] Baseline cognitive data may include one or more cognitive measures obtained using one or more cognitive assessments, such as the CDR-SB assessment, Cogstate Brief Battery assessment, International Shopping List Test assessment, Alzheimer's Disease Assessment Scale (ADAS) assessment, Alzheimer's Disease Composite Score (ADCOMS) assessment, Mini-Mental State Examination, Functional Assessment Questionnaire (FAQ), or other preferred cognitive assessments.

[0028] To be understood, cognitive assessments can include multiple components. The Cogstate Brief Battery assessment can include four components that measure different aspects of cognitive function, such as detection, identification, one-card learning, and one-back. The CDR-SB assessment can include six domains, such as memory (CDR0101), orientation (CDR0102), judgment and problem-solving (CDR0103), community activities (CDR0104), household chores and hobbies (CDR0105), and personal care (CDR0106). The ADAS-13 assessment may include components such as word recall (ADCRL), command (ADCCMD), constructive practice (ADCCPS), delayed word recall (ADCDRL), naming (ADCOF), conceptual practice (ADCIP), orientation (ADCOR), word recognition (ADCRG), recall of test instructions (ADCRI), spoken language comprehension (ADCCMP), word finder difficulty (ADCDIF), spoken language ability (ADCSL), and digit elimination (ADCNC). The ADCOMS assessment may include components such as memory, language, orientation, executive function, processing speed, visuospatial ability, and general functioning. The functional assessment questionnaire may include 10 questions about activities of daily living: paying bills (FAQ01), collecting records (FAQ02), shopping alone (FAQ03), playing games (FAQ04), heating water and turning off the stove (FAQ05), preparing a balanced meal (FAQ06), tracking facts (FAQ07), paying attention (FAQ08), remembering appointments (FAQ09), and traveling (FAQ10). As can be understood, each of these components may be associated with a score. The overall score for the assessment may be a composite of these component-level scores. In some embodiments, baseline cognitive data may include a composite score for the assessment or scores for one or more components of the assessment.

[0029] The disclosed embodiments are not limited to any particular version of the cognitive assessment described above. For example, baseline cognitive data may include ADAS-13 scores or ADAS-14 scores. The ADAS-14 questionnaire may include additional items addressing executive functions not specifically included in the ADAS-13 questionnaire. Similarly, baseline cognitive data may include FAQ IV scores or scores from earlier versions of the FAQ questionnaire.

[0030] In some embodiments, baseline data may include demographic data for the subjects. Demographic data may include age, sex, weight, body mass index (BMI), or other suitable demographic data.

[0031] In some embodiments, baseline data may include genomic data for the subject. Genomic data may include variant information for gene variants associated with neurological disorders, dysfunctions, or injuries (e.g., presence or absence of variants; variant characteristics such as deletion size, insertion size, frameshift information, copy number, single nucleotide polymorphism information; or similar). For example, genomic data may include apolipoprotein E (APOE) variant data, amyloid protein precursor (APP) variant data, presenilin-1 (PSEN1) variant data, presenilin-2 (PSEN2) variant data, clusterin (CLU) variant data, trigger receptor 2 (TREM2) expressed on myeloid cells, etc. Genomic data may further include allele count information for such variants in the subject.

[0032] In some embodiments, baseline data may include image data of the subject. Image data may be, include, or depend on features extracted from images of the subject brain (or a portion thereof). Brain images can be acquired using magnetic resonance imaging (MRI), computed tomography (CT) imaging, positron emission tomography (PET) imaging, or another preferred modality. In some embodiments, image data may include measurements such as brain region, whole brain volume, and hippocampal volume. In some embodiments, measurements may include volume, surface area, or cortical thickness measurements. Volume, surface area, or thickness measurements can be normalized (e.g., using whole brain volume or another preferred normalization factor). Brain region measurements can be extracted from images acquired using MRI, CT, PET, or another preferred imaging modality. In some embodiments, image data may include amyloid or tau PET data (e.g., whether detected amyloid or tau meets threshold criteria in the subject brain or a portion thereof). In some embodiments, the image data may include PET data of a diagnostic tracer (such as fluorodeoxyglucose, florbetaben, florbetapyr, or flutemetamol). In some embodiments, the PET data of amyloid, tau, or tracer may be represented in terms of a score (e.g., normalized or percentile score).

[0033] As described herein, baseline data may include image data specific to hubs and modules identified in the brain. Hubs and modules can be identified using data analysis or feature extraction techniques, such as network analysis or multilevel clustering. Using network analysis, suitable brain regions can be identified based on the connectivity between these brain regions. Such network analysis can identify brain regions as important based on the centrality of the brain region within a network of brain regions (or subnetworks within the entire network), the extent to which the brain region bridges different subnetworks within the network, the extent to which the brain region is connected to other important brain regions, or based on other suitable criteria.

[0034] Multilevel clustering may involve iteratively clustering brain regions together based on similarities in measurements for those brain regions. In each iteration, more (or potentially more) clusters may be created using smaller subsets of brain regions. In some embodiments, hierarchical clustering can be performed. A first clustering operation can be performed on the brain regions to create a first set of clusters. Several further clustering operations can be performed on the brain regions within each cluster in the first set of clusters.

[0035] In some embodiments, network analysis and multilevel clustering can be combined. For example, hubs and modules can be identified using multiscale embedded gene co-expression network analysis (MEGENA), recursive feature removal, etc. Image data specific to identified hubs and modules in the brain may include measurements for the hubs (e.g., volume, surface area, cortical thickness, etc.) and composite values ​​for the modules (e.g., composite values ​​based on measurements for brain regions containing the modules).

[0036] In some embodiments, baseline data may include biomarker data for the subject. Biomarkers can be measurable substances or characteristics that represent a biological process or state, such as a disease state or a response to a treatment. Biomarker data may include amyloid-beta biomarker data, tau biomarker data, neurofilamentous light peptide (NfL) biomarker data, glial fibrillary acidic protein (GFAP) biomarker data, etc. For example, biomarker data may include total tau levels, microtubule-binding region (MBTR) tau levels, phosphorylated tau levels (e.g., tau phosphorylated at 181 (p-Tau181) levels, tau phosphorylated at 217 (p-Tau217) levels, tau phosphorylated at 231 (p-Tau231) levels, etc.), neurogranin levels, Aβ1-42 levels, Aβ1-40 levels, etc. Such biomarker data can be measured in the subject's blood (e.g., plasma, serum, etc.), cerebrospinal fluid, or other suitable body fluids. For example, baseline data may include serum or plasma Aβ1-42 levels, cerebrospinal fluid Aβ1-42 levels, or Aβ1-42 levels measured in another suitable body fluid.

[0037] Biomarker data may include indicators of the presence or absence of biomarkers in a sample obtained from a subject, or the amount or concentration of biomarkers in the sample. Biomarker data can be expressed as a measured quantity or converted into a score (e.g., normalized quantity, percentile). In some embodiments, biomarker data may include the functions of multiple biomarkers. For example, biomarker data may include combinations of Aβ1-42 levels and Aβ1-40 levels (or scores), a ratio of Aβ1-42 levels to Aβ1-40 levels (or scores), a ratio of p-Tau181 (or p-Tau217 or p-Tau231) levels to Aβ1-42 (or Aβ1-40) levels (or scores), etc.

[0038] Figure 2 shows an exemplary platform 200 for developing, validating, and deploying a predictive model for predicting the progression of cognitive impairment or cerebral amyloid status, consistent with the disclosed embodiments. Platform 200 can be designed to obtain input data from other systems (e.g., image processing systems or medical testing systems not shown in Figure 2) or from records 201. Consistent with the disclosed embodiments, such data may include subject image data, biomarker data, genomic data, demographic data, cognitive data, etc. Platform 200 can be designed to create datasets suitable for training predictive models using components such as an extract-transformation-load (ETL) engine 210 and a dataset creation engine 215. Platform 200 can be designed to train models using a training engine 220. The trained models can be used in the prediction phase by the prediction engine 230. A user can control and design Platform 200 by interacting with a user device 299. The user can also provide subject data to Platform 200 and receive predictions from Platform 200 by interacting with the user device 299.

[0039] To be understood, the specific arrangement of components shown in Figure 2 is not intended to be limiting. Platform 200 may include additional components (e.g., additional databases, data sources, processing systems, etc.) or fewer components (e.g., by combining databases or processing systems). The functionality of existing components can be combined or distributed within the additional system without departing from the envisioned embodiment.

[0040] The components in Figure 2 can run using one or more computing systems (e.g., laptops, desktops, workstations, computing clusters, on-premises or off-premises cloud computing platforms). For example, a computing cluster or workstation can run the ETL engine 210, the dataset creation engine 215, or the training engine 220. As an additional example, a desktop or laptop (e.g., user device 299 or another device) can run the prediction engine 230. As an additional example, components of platform 200 (e.g., another than user device 299) can run using containerized services on a cloud computing platform. As should be understood, these examples are not intended to be limiting.

[0041] In accordance with the disclosed embodiments, Record 201 may include one or more storage locations for data available to the platform 200 for predicting the progression of cognitive impairment or cerebral amyloid status. In some embodiments, such data may include raw or processed image data of the subject. Image data may include MRI images, PET images, CT images, etc. In various embodiments, such data may include medical record information for the subject. Such medical record information may include medical records, case notes, clinical trial records, request information (e.g., association with biomarker tests), or clinical test results. In some embodiments, the medical record information may include cognitive data available to construct a training dataset (e.g., cognitive measures taken at 3, 6, 9, 12, 15, and 18-month assessments, or other intervals consistent with the disclosed embodiments).

[0042] In accordance with the disclosed embodiments, the ETL engine 210 can be designed to acquire data in various formats from one or more sources (e.g., record 201). The disclosed embodiments are not limited to any particular format of the acquired data or the method for acquiring this data. For example, the acquired data may be or may include structured data or unstructured data. The ETL engine 210 can interact with various data sources to receive or retrieve data.

[0043] The ETL engine 210 can convert the data into a suitable format and load the converted data into a target component or database of the platform 200. In some embodiments, the data conversion may include performing quality control processing on the obtained data. Such quality control processing may include verifying that the data is usable (e.g., that the subject meets the inclusion criteria for the model to be trained, or that the requested input data for the subject is complete). In some embodiments, the data conversion may include processing the data into a standard format or structure. As can be understood, the input data obtained from record 201 may not be in a format suitable for training a predictive model. Similarly, the input data obtained from different data in record 201 may have different formats. Therefore, the ETL engine 210 can decode the obtained input data so that it has a consistent format even if the input data originates from various different sources.

[0044] In some embodiments, the ETL engine 210 can enrich the image data or medical record information by creating additional data using the image data or medical record information. For example, the ETL engine 210 can convert biomarker levels into scores (e.g., using population distribution information, clinical range, etc.), normalize the image data (e.g., normalize area and volume by intracranial volume), or similar operations. In some embodiments, the ETL engine 210 can remove unnecessary or unwanted variables or data from the input dataset. For example, if a medical record contains information irrelevant to predicting the progression of cognitive impairment, the ETL engine 210 can create a version of the medical record that contains only the information relevant to predicting cognitive impairment.

[0045] In accordance with the disclosed embodiments, the ETL engine 210 can load the transformed data into another component of the platform 200, such as a dataset creation engine 215 (or a suitable data storage device from which the dataset creation engine 215 can retrieve data).

[0046] In some embodiments, the dataset creation engine 215 can be designed to create training samples or inference samples from data received from the ETL engine 210. The dataset creation engine 215 can be designed to extract any required input data features from the transformed data received from the ETL engine 210. In some embodiments, the dataset creation engine 215 can create features based on combinations of biomarkers (e.g., the ratio of Aβ1-42 to Aβ1-40 scores), determine correlations between input data (e.g., between thickness, area, or volume measurements for different brain regions), identify brain regions as modules or hubs, or perform other feature extractions, as described herein.

[0047] In some embodiments, the dataset creation engine 215 can be designed to acquire label information provided by the user through a user device 299. For example, the dataset creation engine 215 can be designed to provide data (or metadata associated with the data) received from the ETL engine 210 to a user device 299 for display. In response, the dataset creation engine 215 can receive label information (e.g., identification of a patient as having brain amyloid, cognitive measurements, etc., obtained during the evaluation of the subject).

[0048] In some embodiments, the dataset creation engine 215 can be designed to associate labels with training samples. For example, when predicting the progression of cognitive impairment, the data may include repeated cognitive measures taken over time. These repeated cognitive measures may be taken after baseline cognitive data has been acquired. The dataset creation engine 215 can associate this cognitive impairment progression data with baseline input data. The training samples may then include baseline data and associated cognitive impairment progression data for the subject. As an additional example, when predicting cerebral amyloid status, findings regarding the presence of cerebral amyloid may be described in the subject's medical record (e.g., based on visual readings of radioactive tracers in PET image data). The training samples may then include baseline data and indicators of findings regarding the presence of cerebral amyloid for the subject. The dataset creation engine 215 can be designed to store the training samples in the data storage device 205.

[0049] In accordance with the disclosed embodiments, the model storage device 203 may be a storage location for models available to components of the platform 200 (e.g., the training engine 220 or the prediction engine 230). The disclosed embodiments are not limited to any particular execution of the model storage device 203. In accordance with the disclosed embodiments, the model storage device 203 may be executed using one or more relational databases, object-oriented or document-oriented databases, tabular data stores, graph databases, distributed file systems, or other suitable data storage options.

[0050] In accordance with the disclosed embodiments, the data storage device 205 may be a storage location for created datasets available to the training engine 220 or the prediction engine 230. The disclosed embodiments are not limited to any particular execution of the data storage device 205. In accordance with the disclosed embodiments, the data storage device 205 may be executed using one or more relational databases, object-oriented or document-oriented databases, tabular data stores, graph databases, distributed file systems, or other suitable data storage options.

[0051] In accordance with the disclosed embodiments, the training engine 220 can be designed to train a model or to create and train a model. The training engine 220 can be designed to create a model (for example, in response to a command to create a particular type of trained model using an input dataset) or to obtain an existing model from the model storage device 203. The training engine 220 can be designed to create or train a model using a training dataset obtained from the data storage device 205. In some embodiments, the training engine 220 can be designed to store the trained model in the model storage device 203.

[0052] In accordance with the disclosed embodiments, the training engine 220 may include a model training component and a model evaluation component. The training engine 220 may be designed to train a model using the model training component and then determine performance metrics for the model using the model evaluation component.

[0053] In some embodiments, the training engine 220 can provide the model evaluation component with a cross-validation or holdout portion of the model and training dataset. In some embodiments, the training engine 220 can identify one or more performance metrics. Additionally or alternatively, the model evaluation component can be designed using a predetermined or default set of performance metrics. In some embodiments, the performance metrics may include confusion matrix, mean squared error, mean absolute error, sensitivity or selectivity, receiver operating characteristic curve or area under such curve, precision and recall, F-measurement, or any other suitable performance metrics. In some embodiments, the performance metric values ​​may be displayed to the user through the user device 299. The user can then interact with the training engine 220 through the user device 299 to update the model.

[0054] In some embodiments, the training engine 220 can automatically update the model being trained based on performance metrics. In various embodiments, the training engine 220 can update the model being trained in response to user input provided through the user device 299. Model updates may include one or more of the following: performing additional training (e.g., using an existing training dataset or another training dataset), modifying the model (e.g., changing the input features used by the model, changing the structure of the model, etc.), or changing the training environment (e.g., changing the training of hyperparameters, changing the training, cross-validation, and partitioning of the training dataset into holdout portions, etc.).

[0055] In accordance with the disclosed embodiments, the prediction engine 230 can be designed to use a prediction model to predict the progression of cognitive impairment or the brain amyloid load for a subject. In some embodiments, the prediction engine 230 can retrieve a trained classification model from the model storage device 203. In some embodiments, the prediction engine 230 can retrieve input data for a subject from the data storage device 205. In some embodiments, the prediction engine 230 can retrieve subject data from an alternative data storage location. This alternative data storage location may be associated with a different entity or user. For example, the prediction engine 230 can receive or retrieve subject data from a healthcare system controlled by an entity different from the entity controlling the prediction engine 230.

[0056] In accordance with the disclosed embodiments, the target data may include baseline data. Baseline data may include demographic data, cognitive data, genomic data, image data, biomarker data, or other suitable baseline data. The prediction engine 230 can apply the baseline data to a trained prediction model to provide, as output, predictions of cognitive impairment progression data or brain amyloid load for the target. The output may be provided by the prediction engine 230 to the user device 299. In some embodiments, the output may be stored in a computing device associated with the platform 200 or provided to another system.

[0057] In accordance with the disclosed embodiments, the user device 299 may provide a user interface for interacting with other components of the platform 200. The user interface may be a graphical user interface. The user interface may enable the user to design the ETL engine 210 to extract, transform, and load data according to user specifications. The user interface may enable the user to identify how the transformed data received by the dataset creation engine 215 is transformed into labeled training data (or suitable patient data). In some embodiments, the user interface may enable the user to interact with the dataset creation engine 215 to manually or semi-manually label or annotate the training data. In some embodiments, the user interface may enable the user to provide data or models to the training engine 220 for training, or to the prediction engine 230 for identification and classification.

[0058] In some embodiments, the user interface may enable the user to interact with the training engine 220 to create or select a model for training, create or select a dataset to use for training the model, or select training parameters or hyperparameters. In some embodiments, the user interface may enable the user to interact with the training engine 220 to display information about the training of the model (e.g., performance metric values, changes in missing feature values ​​during training, or other training information). In some embodiments, the user interface may enable the user to interact with the prediction engine 230 to select a training model and patient data (e.g., baseline images). In some embodiments, the user interface may enable the user to interact with the prediction engine 230 to display any indicators of identified biological structures in the patient data, store the indicators in a computing device, or transmit the indicators to another system.

[0059] The components of platform 200 can run using one or more computing devices. Such computing devices may include tablets, laptops, desktops, workstations, computing clusters, or cloud computing platforms. In some embodiments, the components of platform 200 can run using a cloud computing platform. For example, one or more of the ETL engine 210, dataset creation engine 215, training engine 220, and prediction engine 230 can run on a cloud computing platform. In some embodiments, the components of platform 200 can run using on-premises systems. For example, the record 201 or user device 299 may be or can be hosted on an on-premises system. As an additional example, the model storage device 203 or data storage device 205 may be or can be hosted on an on-premises system.

[0060] The components of platform 200 can be communicated using any preferred method. In some embodiments, two or more components of platform 200 can run as microservices or web services. Such components can be communicated using messages transmitted over a computer network. Messages can be communicated using SOAP, XML, HTTP, JSON, RCP, or any other preferred format. In some embodiments, two or more components of platform 200 can run as software, hardware, or a combined software / hardware module. Such components can be communicated using data or instructions written to or read from memory (e.g., shared memory), function calls, or any other preferred communication method.

[0061] To the best of our understanding, the specific structure of platform 200 is not intended to be limiting. In accordance with the disclosed embodiments, any two or more of the record 201, model storage device 203, or data storage device 205 can be combined or hosted on the same computing device. In accordance with the disclosed embodiments, the ETL engine 210 and the dataset creation engine 215 can be omitted from platform 200. In such embodiments, datasets formatted and designed for use by the training engine 220 or prediction engine 230 can be stored in data storage device 205 by another system or by other means. In accordance with the disclosed embodiments, the ETL engine 210 and the dataset creation engine 215 can be combined. In such embodiments, data extraction, transformation, and loading can be combined with feature extraction, labeling, and sample creation. In accordance with the disclosed embodiments, the training engine 220 and the prediction engine 230 can be combined.

[0062] Although represented by a single user device 299, platform 200 may have multiple user devices. Different user devices may be associated with different entities or different users with different roles. For example, user device 299 may be associated with a software engineer or data scientist developing a predictive model, while another user device may be associated with a clinician using the predictive model.

[0063] The user device 299 can be combined with one or more other components of the platform 200. In some embodiments, the user device 299 and at least one of the ETL engine 210, dataset creation engine 215, training engine 220, or prediction engine 230 can be run on the same computing device. In various embodiments, the user device 299 and at least one of the model storage device 203 or data storage device 205 can be run on the same computing device.

[0064] To make it understandable, platform 200 can be incorporated into a method for treating a subject or a method for conducting a clinical trial. The prediction engine 230 can use a trained predictive model and input data for a subject to predict brain amyloid load or cognitive impairment progression for the subject. The predicted brain amyloid load or cognitive impairment progression can be used to determine a patient treatment plan for the patient. In some embodiments, cognitive impairment progression for a subject can be used as a prognostic covariate in determining the therapeutic effect in a clinical trial.

[0065] Figure 3 shows an exemplary process 300 for predicting the progression of cognitive decline in a subject, consistent with the disclosed embodiments. For convenience of explanation, process 300 is described as being performed using platform 200. However, process 300 may also be performed, at least in part, using a different computing system. Process 300 may include a dataset creation phase, a training phase, and a prediction phase. In the dataset creation phase, components of platform 200 (e.g., ETL engine 210, dataset creation engine 215, etc.) can acquire training data from a database (e.g., record 201, etc.). In the training phase, components of platform 200 (e.g., training engine 220, etc.) can create or refine a predictive model for predicting cognitive decline progression data from baseline data. In the prediction phase, components of platform 200 (e.g., prediction engine 230, etc.) can use the predictive model created in the training phase to predict cognitive decline progression data from baseline data for the subject. The predicted cognitive decline progression data can be used to manage treatment for the subject or can be used as a prognostic covariate in clinical trials.

[0066] In step 310 of process 300, components of platform 200 (e.g., ETL engine 210, dataset creation engine 215, etc.) can acquire training data in accordance with the disclosed embodiments. The training data may relate to subjects that meet the criteria for cognitive impairment. A cognitive impairment can be identified as a subject having a diagnosis of neurological disorder, dysfunction, or injury (e.g., a diagnosis of AD, a diagnosis of MCI, a diagnosis of dementia, etc.) and having specific signs (e.g., amyloid positivity on PET scans; biomarker scores, e.g., p-Tau181, Aβ1-42 score, or Aβ1-40 score in plasma, serum, or cerebrospinal fluid, etc.). In some cases, cognitive impairment may be an inclusion criterion for clinical trials.

[0067] In some embodiments, components of platform 200 can acquire at least a portion of the training data from a database (e.g., record 201) or another system. In some embodiments, components of platform 200 can create at least a portion of the training data. For example, the dataset creation engine 215 can identify brain regions as hubs and clusters of brain regions as modules. The dataset creation engine 215 can identify such brain regions using network analysis or multilevel clustering. For example, as described herein with respect to Figure 4, the dataset creation engine 215 can identify such brain regions using MEGENA. The dataset creation engine 215 can create composite values ​​for modules based on measurements of brain regions containing modules.

[0068] In some embodiments, training data may include baseline data for subjects, consistent with the disclosed embodiments. In some embodiments, baseline data may include cognitive data and image data as described herein. For example, image data may include measurements for one or more brain regions identified as hubs and composite values ​​for one or more clusters of brain regions identified as modules. In some embodiments, baseline data may include demographic data as described herein. In some embodiments, baseline data may include genomic data as described herein. For example, baseline data may include ApoE4 allele counts. In some embodiments, baseline data may include biomarker data. In some embodiments, biomarker data may be plasma, serum, or cerebrospinal fluid biomarker data.

[0069] In some embodiments, training data may include cognitive impairment progression data for a subject, consistent with the disclosed embodiments. Cognitive impairment progression data may include repeated measures taken over time. In some embodiments, repeated measures may be or depend on cognitive measures as described herein. In some embodiments, repeated measures may be expressed as a function of a baseline cognitive measure and a subsequent cognitive measure (e.g., the difference between a baseline cognitive assessment and a subsequent cognitive assessment). For example, repeated measures may be or include changes in CDR-SB measures, ADCOMS measures, ADAS measures, etc., for a subject. In some embodiments, each repeated measure may include or depend on a Clinical Dementia Severity Scale (CDR-SB) measure, an Alzheimer's Disease Composite Score (ADCOMS) measure, or an Alzheimer's Disease Assessment Scale (ADAS) measure.

[0070] As can be understood, repeated measures for a subject can be obtained after baseline cognitive data for the subject has been acquired. In some embodiments, repeated measures can be obtained during repeated assessments performed after baseline cognitive data has been acquired. Such assessments can be performed for each subject at time intervals of 3 to 12 months, and repeated measures can be obtained. In some embodiments, the elapsed time between the acquisition of baseline cognitive data for a first subject and the acquisition of the final data for repeated measures may be 12 to 36 months, or 18 to 24 months, etc.

[0071] In some embodiments, the elapsed time associated with repeated measures may be implicit. For example, repeated measures can be represented as a vector (or a matrix in the case of repeated measures that qualify as a vector), and the elapsed time may be implicit at the position of the repeated measures in the vector (or the column of the repeated measures in the matrix). In some embodiments, the elapsed time associated with repeated measures may be implicit. For example, repeated measures can be represented as a tuple, each tuple containing an index of the repeated measures and the elapsed time since the baseline assessment.

[0072] In step 320 of process 300, components of platform 200 (e.g., training engine 220) can train a predictive model to predict the progression of a cognitive impairment in question, in accordance with the disclosed embodiments. In some embodiments, the training engine 220 can create a predictive model and then store the predictive model in model storage 203. In some embodiments, the training engine 220 can retrieve a predictive model from model storage 203 or another database or system and then refine the model.

[0073] In some embodiments, the training engine 220 can acquire hyperparameters for training a predictive model. The specific hyperparameters acquired may depend on the type of predictive model, and the disclosed embodiments are not limited to any particular set of hyperparameters. For example, a neural network model may have hyperparameters governing the arrangement and staging of layers, batch size, dropout, etc. As an additional example, a gradient boosting model may have hyperparameters governing the learning speed, number of trees, bagging portion, tree depth, etc.

[0074] In some embodiments, a user can interact with a user device 299 to provide hyperparameters to the training engine 220. In some embodiments, the training engine 220 can receive or retrieve hyperparameters from another component of the platform 200. In some embodiments, the training engine 220 can create suitable hyperparameters. For example, the training engine 220 can be designed to perform iteration or adaptive search of a given hyperparameter space (e.g., through a trained predictive model that evaluates the performance of the model and updates the selected hyperparameters based on the model's performance).

[0075] In some embodiments, the training engine 220 can train a predictive model using the hyperparameters and training data obtained in step 310. The disclosed embodiments are not limited to any specific code or instructions for training the model. For example, if the R statistics package is used in the training engine 220 and the predictive model is a gradient boosted model, the model can be trained using the following code: Library (GBM) x = TrainingData[, xvar] y = TrainingData[, yvar] train.data=cbind(y,x) set.seed(263) gbm.fit<-gbm(y~., data=train.data, verbose=FALSE,distribution=“gaussian”, shrinkage=0.01, interaction.depth=3, n.minobsinnode=100, n.trees=1000, cv.folds=5, bag.fraction=0.5, n.core=1 )

[0076] The training engine 220 can execute this code to determine a predictive gradient boost model using the training data and given hyperparameter values.

[0077] In some embodiments, the training engine 220 can be designed to evaluate the performance of multiple model designs using the same training dataset. The performance of a model design can be determined using k-fold cross-validation. In some embodiments, the training engine 220 can be designed to evaluate the performance of the best-performing model design by splitting the training dataset into a training subset and a validation subset. The training engine 220 can evaluate the model design by performing k-fold cross-validation using the training subset. The training engine 220 can select a model design and evaluate its performance using the validation subset.

[0078] In step 330 of process 300, components of platform 200 (e.g., ETL engine 210, dataset creation engine 215, prediction engine 230, etc.) can acquire baseline data for individual subjects. In some embodiments, individual subjects may satisfy cognitive impairment conditions. For example, individual subjects and subjects from which training data has been obtained may satisfy the same cognitive impairment condition. In some embodiments, individual subjects may satisfy similar or equivalent cognitive impairment conditions (e.g., a training subject may have a diagnosis of AD, while an individual subject may have clinical findings suggestive of AD). In some embodiments, components of platform 200 can acquire at least a portion of individual subject data from a database (e.g., record 201, etc.) or another system. For example, platform 200 can be designed to acquire prediction requests from other systems. In some embodiments, components of platform 200 can create at least a portion of individual subject data.

[0079] In some embodiments, baseline data for individual subjects (e.g., predictive baseline data) may be the same as baseline data included in the training data (e.g., training baseline data). For example, if the training baseline data includes a specific combination of demographic data, biomarker data, and cognitive measures, the predictive baseline data may include the same demographic data, biomarker data, and cognitive measures. As can be understood, obtaining predictive baseline data may involve rearranging or arranging the predictive baseline data to match the format or arrangement of the training baseline data. Similarly, obtaining predictive baseline data may involve handling missing or error values ​​in the training baseline data. Furthermore, if obtaining training baseline data involves creating specific values ​​(e.g., creating composite values ​​for modules), obtaining predictive baseline data may similarly involve creating these values.

[0080] In some embodiments, the prediction data may include elapsed time. In some embodiments, the user can interact with platform 200 to request a prediction of cognitive impairment over this elapsed time.

[0081] In step 340 of process 300, a component of platform 200 (e.g., prediction engine 230) can predict the progression of cognitive impairment for a subject, in accordance with the disclosed embodiments. In some embodiments, prediction engine 230 can input predicted baseline data into a trained prediction model. The output of the trained prediction model may be an array of predicted cognitive measures (e.g., predicted cognitive impairment progression data). The predicted cognitive impairment progression data may be implicitly or explicitly related to elapsed time since the baseline. For example, the output may be a vector (or matrix) of values, where each position in the vector (or column of the matrix) may implicitly be related to elapsed time. As an additional example, the predicted cognitive impairment progression data may be a set of tuples, each tuple containing an elapsed time and a set of predicted cognitive measures for that elapsed time.

[0082] In accordance with the disclosed embodiments, platform 200 can provide predicted cognitive impairment progression data. Platform 200 can provide the predicted cognitive impairment progression data to users of platform 200 (for example, by providing the predicted cognitive measurements to a user device 299 for display), store the predicted cognitive impairment progression data within a component of platform 200, and provide the predicted cognitive impairment progression data to another system (for example, the system that provided the prediction request).

[0083] As can be understood, predicted cognitive impairment progression data can be manually, semi-manually, or automatically assessed as an indicator of the progression of neurological disease, functional impairment, or injury. In some cases, for example, a subject may be diagnosed with mild cognitive impairment. The subject's fulfillment of cognitive impairment status may depend on such a diagnosis (or, for example, clinically equivalent findings). Predicted cognitive impairment progression data can provide an indicator of the subject's progression to Alzheimer's disease. Platform 200 can be designed to automatically assess predicted cognitive impairment progression data (e.g., using baseline data, predicted cognitive measures, and diagnostic thresholds, etc.) and provide an indicator of such predicted progression.

[0084] To make it understandable, trained predictive models can be used to screen or select patients for inclusion in a clinical trial. Cognitive impairment progression data can be predicted for candidate patients using baseline data obtained for those candidates. Candidate patients can be inducted into the trial if the predicted cognitive impairment progression data meets the inclusion criteria. In various embodiments, the inclusion criteria may depend on the final cognitive measure, the change in cognitive measure between the baseline and final cognitive measures, cognitive measures at a specific point in time after baseline, the value or coefficient of the function fit to the predicted cognitive impairment progression data, or another preferred measure. For example, a patient may be inducted into a clinical trial if the final predicted cognitive measure in the predicted cognitive impairment progression data is above a threshold or within a specific range. Such patients may have a greater need for treatment (particularly because they are more likely to experience greater cognitive decline).

[0085] As can be understood, trained predictive models can be used in trial design. As described herein, clinical trial populations can be enriched with patients who are likely to benefit from treatment. In particular, trained predictive models can be used to screen or select patients for inclusion in a clinical trial. Selected patients may be those who are likely to exhibit at least a threshold amount of cognitive decline. As can be understood, treatment effect, trial size, and trial power can be related. By selecting patients who are likely to exhibit substantial cognitive decline, it may be possible to enroll fewer patients, increase trial power, decrease the detectable treatment effect size, or some combination of the above.

[0086] Furthermore, the therapeutic benefit for such patients may be more readily apparent than in patients whose suffering is reduced. Because the therapeutic effect is more apparent (e.g., greater therapeutic effect), screening or selecting patients using trained predictive models can enable improved clinical trial designs: the number of enrolled patients can be reduced, the minimum detectable therapeutic effect can be increased, the power of the trial can be increased, or some combination of the above.

[0087] To make it understandable, predicted cognitive impairment progression data can be used to evaluate the effectiveness of treatment for neurological disorders, functional impairments, or injuries in clinical trials. Clinical trials may include multiple participants. Participants can be screened for compliance with cognitive impairment status, and baseline data can be obtained for each participant. Participants can be assigned to either the control group or the treatment group of the trial.

[0088] Using a trained predictive model, cognitive impairment progression data can be predicted for one or more participants in the treatment group of the trial and used as a prognostic covariate in the analysis of the trial's results.

[0089] For example, clinical trials may be related to the treatment of Alzheimer's disease. Using a trained predictive model, cognitive impairment progression data can be predicted for at least a subset of patients assigned to the treatment group in the clinical trial. The predicted cognitive impairment progression data can then be used as a prognostic covariate in evaluating the effectiveness of Alzheimer's disease treatment.

[0090] Figure 4 shows a process 400 for identifying hubs and modules in the brain based on brain region measurements, consistent with the disclosed embodiments. Brain region measurements may include volume, area, cortical thickness, etc. Process 400 can be used as part of process 300 described herein or as part of another process. As described herein, training baseline data may include image data, e.g., brain region measurements related to the brain region of interest. Using process 400, specific combinations of brain regions (e.g., hubs) or brain regions (e.g., modules) can be identified that are particularly representative of the brain region measurements for the interest. In this form, process 400 can reduce the dimensionality of the training baseline data and improve the robustness of the trained predictive model.

[0091] For convenience of disclosure, process 400 is described herein as being carried out by the dataset creation engine 215 of platform 200. However, this description is not intended to limit. Without departing from the assumed embodiments, process 400 can be carried out by another component of platform 200 or by another system. Similarly, this process is described as being carried out using MRI brain region measurements. However, brain region measurements obtained using any preferred imaging modality can be used without departing from the assumed embodiments.

[0092] In step 401, process 400 can be started. The dataset creation engine 215 can obtain MRI region scales for training subjects. The dataset creation engine 215 can receive or retrieve MRI region scales from another component of platform 200 (e.g., data storage device 205) or another system. The dataset creation engine 215 can create MRI region scales from MRI image data of the training subjects (e.g., using FREESURFER or another suitable tool for the analysis and visualization of neuroimaging data), and the dataset creation engine 215 can then receive or retrieve these from another component of platform 200 (e.g., data storage device 205) or another system. The MRI region scales may correspond to regions identified in neuroanatomical atlases (e.g., Desikan-Killiany atlas, Harvard-Oxford atlas, automated anatomical labeling atlas, Brainnetome atlas, etc.). In some embodiments, the dataset creation engine 215 can normalize volume and area scales by intracranial volume to reduce inter-subject variability and account for variations attributable to head size.

[0093] In step 410 of process 400, the dataset creation engine 215 can construct a planar filtered network graph using MRI region scales. The planar filtered network graph may include nodes corresponding to brain regions and edges corresponding to relationships between brain regions. For example, an edge between two nodes may correspond to the correlation between brain region measurements for a region of the training subject. The dataset creation engine 215 can determine the correlation of MRI scales across all pairs of brain regions. Pairs of regions can be ranked by correlation and then filtered based on a false detection threshold. Filtered pairs can be iteratively tested for planarity (e.g., using the Boyer-Myrvold algorithm). If a pair passes the planarity test, the network can be updated to include links in the network corresponding to the pair. The embedding process can be iterated until the termination condition is met. The termination criteria may depend on the number of pairs included in the network (for example, whether the number of included pairs is the maximum number of edges that can be embedded in a topological sphere such that no edges intersect each other), whether any pairs remain untested, or the number of pairs rejected for each pair retrieved. In this way, the dataset creation engine 215 can create a planar filtered network graph that supports the inclusion of more highly correlated brain regions.

[0094] In step 420 of process 400, the dataset creation engine 215 can perform a multilevel clustering analysis using the planar filtered network graph. The clustering analysis can attempt to optimize compactness within clusters, local clustering structure, and overall modularity. The clustering analysis can be performed iteratively. In the initial iteration, the clustering analysis may involve dividing the network graph into clusters of nodes. In each subsequent iteration, the embedded network can undergo nesting on each cluster of nodes.

[0095] In some embodiments, nested partitioning can be performed by using k-medoid clustering (or another preferred clustering into a predetermined number of clusters) with k selected through an iterative process, according to the shortest path distance (and optionally, with refined cluster boundaries using local path metrics). In each iteration, the value of k may be different, and the resulting clusterings can be evaluated using a measure of clusterability on the network (e.g., Newman's modularity). Different values ​​of K can be tried until a threshold condition is met. The threshold condition may depend on the number of k values ​​considered since the final k value that yields the best achieved compactness measure, a timing or duration condition, or another preferred condition. Candidate partitions for clusters may be partitions having the best achieved compactness measure.

[0096] In some embodiments, candidate partitions for a cluster may be acquired or rejected based on a compactness metric determined for each subcluster within the partition. The compactness metric for a subcluster may depend on path distances within the subcluster (e.g., normalized mean shortest path distance) and scaling parameters. In some embodiments, a candidate partition for a cluster may be rejected if the compactness metric for all subclusters within the cluster is greater than the compactness metric for the cluster. Otherwise, the candidate partition may be acquired.

[0097] In some embodiments, statistical significance may be related to subclusters. This statistical significance may depend on the value of the scaling parameter required to obtain the subclusters. Assuming the value of the scaling parameter required to obtain the subclusters, statistical significance may be the possibility of randomly creating subclusters that have at least a compactness measure value for the subclusters.

[0098] In some embodiments, nesting can be performed until a termination condition is met. In some embodiments, the termination condition may be met if it is not possible to identify a subcluster of a parent cluster that is more compact than the parent cluster for any value of the scaling parameter. In some embodiments, the termination condition may be met if the subcluster does not exhibit a statistical significance greater than a threshold (e.g., 0.05, or another preferred significance value). As can be understood, the disclosed embodiments are not limited to these specific termination conditions. Other preferred termination conditions (e.g., time or resource-based termination conditions) may also be used. In some embodiments, the identified clusters may be modules.

[0099] In step 430 of process 400, the dataset creation engine 215 can perform a multiscale hub analysis of the embedded network to identify relevant brain regions at each scale defined by the scaling parameters above, and across all scales. In the first step, the scaling parameter values ​​associated with significant clusters (from step 420) can be sequentially clustered (e.g., using k-medoid clustering or another clustering method) based on the connectivity of nodes within clusters at different scaling parameters. In the second step, node significance can be identified for each scale using the connectivity of nodes within clusters at that scale and the connectivity of nodes within clusters in randomly created subnetworks for that scale. In the third step, hubs can be identified by combining the significance scores of individual brain regions across all different scales.

[0100] In step 499, process 400 may terminate. The dataset creation engine 215 can create composite values ​​for the identified modules using MRI region scales for these modules. In some embodiments, a first principal component can be calculated for measurement types (e.g., volume, surface area, cortical thickness, etc.) across all brain regions included in the module. This first principal component may then be associated with the module. As can be understood, principal component values ​​can be created for multiple types of measurements (e.g., volume, surface area, cortical thickness, etc.), and the results for these types of measurements may be associated with the model. As can be understood, the disclosed embodiments are not limited to the use of principal component analysis to create values ​​for modules.

[0101] If process 400 is performed as part of process 300, the region values ​​for the identified hubs and modules can be used as input data for training a predictive model, in accordance with the disclosed embodiments.

[0102] Figure 5 shows a process 500 for predicting the brain amyloid state of a subject, consistent with the disclosed embodiments. For convenience of explanation, the process 500 is described as being carried out using platform 200. However, the process 500 may also be carried out using a different computing system, at least in part. The process 500 may include a dataset creation phase, a training phase, and a prediction phase. In the dataset creation phase, components of platform 200 (e.g., ETL engine 210, dataset creation engine 215, etc.) can acquire training data from a database (e.g., record 201, etc.). In the training phase, components of platform 200 (e.g., training engine 220, etc.) can create or refine a predictive model for predicting brain amyloid state from baseline data. In the prediction phase, components of platform 200 (e.g., prediction engine 230, etc.) can use the predictive model created in the training phase to predict brain amyloid state data for a subject from baseline data. Predicted brain amyloid status data can be used to manage treatment for subjects or as a prognostic covariate in clinical trials.

[0103] In step 510 of process 500, components of platform 200 (e.g., ETL engine 210, dataset creation engine 215, etc.) can acquire training data in accordance with the disclosed embodiments. The training data may relate to subjects that meet the criteria for cognitive impairment. Cognitive impairment can be identified as subjects having a diagnosis of neurological disorder, dysfunction, or injury (e.g., a diagnosis of AD, a diagnosis of MCI, a diagnosis of dementia, etc.) and having specific signs (e.g., positive amyloid on a PET scan; biomarker scores, e.g., p-Tau181, Aβ1-42 score, or Aβ1-40 score in plasma, serum, or cerebrospinal fluid, etc.). In some cases, cognitive impairment may be an inclusion criterion for clinical trials.

[0104] In some embodiments, components of platform 200 can retrieve at least a portion of the training data from a database (e.g., record 201) or another system. In some embodiments, components of platform 200 can create at least a portion of the training data.

[0105] In some embodiments, training data may include baseline data for subjects, consistent with the disclosed embodiments. Baseline data may include cognitive data and image data as described herein. In some embodiments, baseline data may include demographic data as described herein. In some embodiments, baseline data may include genomic data as described herein. For example, baseline data may include ApoE4 allele counts. In some embodiments, baseline data may include biomarker data. In some embodiments, biomarker data may be plasma, serum, or cerebrospinal fluid biomarker data.

[0106] In some embodiments, training data may include brain amyloid data for a subject, consistent with the disclosed embodiments. Brain amyloid data may include assessments based on image data of a subject with brain amyloid loading. As described herein, image data may be or include PET image data showing tracer uptake, or image data obtained using another preferred imaging modality. In some embodiments, brain amyloid data may be classification data (e.g., binary classification representing satisfaction of diagnostic criteria, multiclass classification indicating stage or class of amyloid plaque deposition). In some embodiments, brain amyloid data may be continuously valued data, e.g., intensity data, detection amount data (e.g., the number of pixels that satisfy the detection criteria), or other continuously valued data extracted from image data of a subject.

[0107] In step 520 of process 500, components of platform 200 (e.g., training engine 220) can train a predictive model to predict brain amyloid status for a subject, in accordance with the disclosed embodiments. In some embodiments, the training engine 220 can create a predictive model and then store it in the model memory 203. In some embodiments, the training engine 220 can retrieve a predictive model from the model memory 203 or another database or system and then refine the model.

[0108] In some embodiments, the training engine 220 can acquire hyperparameters for training a predictive model. The specific hyperparameters acquired may depend on the type of predictive model, and the disclosed embodiments are not limited to any particular set of hyperparameters. For example, a neural network model may have hyperparameters governing layer arrangement and 3D configuration, batch size, dropout, etc. As an additional example, a gradient boosted model may have hyperparameters governing learning speed, number of trees, bagging portion, tree depth, etc.

[0109] In some embodiments, a user can interact with a user device 299 to provide hyperparameters to the training engine 220. In some embodiments, the training engine 220 can receive or retrieve hyperparameters from another component of the platform 200. In some embodiments, the training engine 220 can create suitable hyperparameters. For example, the training engine 220 can be designed to perform iteration or adaptive search of a given hyperparameter space (e.g., through a trained predictive model that evaluates the performance of the model and updates the selected hyperparameters based on the model's performance).

[0110] In some embodiments, the training engine 220 can train a predictive model using hyperparameters and the training data obtained in step 510. The disclosed embodiments are not limited to any specific code or instructions for training the model. In some embodiments, the training engine 220 can be designed to evaluate the performance of multiple model designs (e.g., Monte-Carlo Logistic Lasso model, Bayesian Logistic Elastic Net model, regularized random forest model, stochastic gradient boosting machine model, etc.) using the same training dataset. The performance of the model designs can be determined using k-fold cross-validation. In some embodiments, the training engine 220 can be designed to evaluate the performance of the best-performing model design by splitting the training dataset into training and validation subsets. The training engine 220 can evaluate the model designs by performing k-fold cross-validation using the training subset. The training engine 220 can select a model design and evaluate its performance using the validation subset.

[0111] In step 530 of process 500, components of platform 200 (e.g., ETL engine 210, dataset creation engine 215, prediction engine 230, etc.) can acquire baseline data for individual subjects. In some embodiments, components of platform 200 can acquire at least a portion of the individual subject data from a database (e.g., record 201, etc.) or another system. For example, platform 200 can be designed to acquire prediction requests from other systems. In some embodiments, components of platform 200 can create at least a portion of the individual subject data.

[0112] In some embodiments, baseline data for individual subjects (e.g., predictive baseline data) may be the same as baseline data included in training data (e.g., training baseline data). For example, if training baseline data includes a specific combination of demographic data, biomarker data, and cognitive measures, predictive baseline data may include the same demographic data, biomarker data, and cognitive measures. As can be understood, obtaining predictive baseline data may involve rearranging or arranging the predictive baseline data to match the format or arrangement of the training baseline data. Similarly, obtaining predictive baseline data may involve handling missing or error values ​​in the training baseline data. Furthermore, if obtaining training baseline data involves creating specific values ​​(e.g., creating composite values ​​for modules), obtaining predictive baseline data may similarly involve creating these values.

[0113] In step 540 of process 500, a component of platform 200 (e.g., prediction engine 230) can predict the brain amyloid state for a subject in accordance with the disclosed embodiments. In some embodiments, the prediction engine 230 can input prediction baseline data into a trained prediction model. The output of the trained prediction model may be a prediction of the brain amyloid state for the subject. As can be understood, the type of prediction may depend on how the model is trained. If the training baseline data includes brain amyloid data with class values, the output of the prediction model may be a predicted class (or possible class). If the training baseline data includes brain amyloid data with continuously occurring values, the output of the prediction model may be a predicted brain amyloid data value.

[0114] In accordance with the disclosed embodiments, platform 200 can provide predicted cognitive measurements. Platform 200 can provide the predicted cognitive measurements to a user of platform 200 (for example, by providing predicted output classes, class probabilities, or brain amyloid data values ​​to a user display device 299), or it can store the predicted cognitive measurements in a component of platform 200 and provide the predicted cognitive measurements to another system (for example, the system that provided the prediction request). [Examples]

[0115] In accordance with the disclosed embodiments, several studies were conducted on the training and use of the predictive model. These studies concerned both the prediction of cognitive impairment progression and the prediction of cerebral amyloid status.

[0116] Example 1 Figures 6A–6H relate to an investigation into predicting cognitive impairment progression and disease progression in amyloid-positive subjects with mild cognitive impairment (e.g., A+MCI subjects). In particular, the investigation examined whether plasma p-Tau181, a biomarker indicating cerebral amyloid load, can predict disease progression in A+MCI subjects in combination with demographic factors and other inputs. As can be seen, blood-based tests for screening and monitoring subjects in Alzheimer's disease (AD) clinical trials would be faster, easier, and more cost-effective compared to conventional CSF and imaging methods.

[0117] Figure 6A shows a demographic summary of the training and validation cohorts used to create a model of neurological disease progression in patients. The training cohort was used to construct a model incorporating 135 A+MCI placebo subjects from two clinical trials. Two independent validation cohorts to test the performance of these signatures included 115 A+MCI subjects from the placebo arm of another clinical trial (VC-1) and 174 A+MCI subjects from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database (VC-2), respectively. Subjects with slower cognitive impairment progression who dropped out before 18 months were excluded.

[0118] Patient data included biomarker data. Plasma p-Tau181 was measured at three different locations using the Simoa assay. These measurements were normalized to have similar mean and variability, and were made equivalent across the cohort (i.e., by subtracting the mean and dividing by the standard deviation). The training cohort and the first validation cohort included 18 months of clinical follow-up, while the second validation cohort included 3–10 years of follow-up. A threshold for earlier cognitive impairment progression was set by an 18-month change in a Clinical Dementia Rating Scale (CDR-SB) score of 1 or greater.

[0119] Baseline MMSE scores were significantly lower in subjects with earlier cognitive impairment progression across all cohorts (p<0.05). Baseline BMI was significantly lower in subjects experiencing earlier cognitive impairment progression in the training cohort and one of the validation cohorts (p<0.05). Subjects in the training cohort with earlier cognitive impairment progression were significantly older.

[0120] In this study, we constructed several predictive models to forecast the progression of cognitive impairment. These models included a Bayesian elastic network (BEN), a regularized random forest, and a gradient boosting model.

[0121] Additional predictive models were induced to evaluate the added value of ApoE4 status, cognitive function assessments, and brain region measurements from magnetic resonance imaging (MRI). Demographic factors (age, sex, and BMI) were considered in all evaluation models. The performance of these models was first evaluated through 10 iterations of 10-stratified cross-validation within the training cohort, and then tested in the first and second validation cohorts.

[0122] In two independent validation cohorts, BEN was the most frequently used machine learning algorithm among those examined to predict 18-month Alzheimer's disease progression. To predict 18-month cognitive impairment progression, performance similar to baseline cognitive function was achieved in baseline plasma p-Tau181, with receiver operating characteristic area-of-curve (ROC-AUC) values ​​of 64.9% and 70.9% for VC-1 (p=0.199) and 65.3% and 66.7% for VC-2 (p=0.395).

[0123] Overall, predictive models consistent with the disclosed embodiments identified that A+MCI subjects were more likely to experience earlier cognitive impairment progression over 18-month intervals. These predictive models showed improved performance when baseline plasma p-Tau181 levels were used as input and when baseline plasma p-Tau181 levels were combined with baseline cognitive function measures. Furthermore, these predictive models showed improved performance in predicting 36-month progression to AD in MCI subjects when baseline plasma p-Tau181 levels were combined with baseline cognitive function or brain region MRI features.

[0124] Figure 6B shows a comparison of biomarker levels in subjects experiencing slower and faster cognitive impairment progression. Baseline p-Tau181 was significantly increased in subjects with faster cognitive decline (CD) in all three cohorts (training, first validation, and second validation).

[0125] Figure 6C shows a performance summary of predictive models derived from the BEN model for predicting 18-month progression to CD and 36-month progression to AD in VC-1 and VC-2. Demographics (age, sex, BMI) were considered in all predictive models. In addition to demographics, the models used baseline p-Tau181 (column 1), cognitive function (column 2), a combination of baseline p-Tau181 and cognitive function (column 3), a combination of cognitive function and MRI image data (column 4), and a combination of cognitive function, MRI image data, and baseline p-Tau181 (column 5) as inputs. Adding ApoE4 did not improve predictive performance.

[0126] Figure 6D shows the ROC curve predicting 36-month progression from MCI to AD in the second validation cohort. As shown in Figure 6D and Figure 6C, the ROC-AUC was significantly improved by combining baseline plasma p-Tau181 with cognitive function or brain region MRI.

[0127] Figure 6E shows the relative importance (e.g., odds ratio) of specific significant inputs to the BEN model for predicting CD at 18 months, based on biomarkers and cognitive function inputs. Significant inputs in this model included the 13-item ADAS-Cog composite score (ADAS-13), functional assessment questionnaire scores (subparts 2, 5, and 6), a function of patient plasma p-Tau181 levels (e.g., standardized log2-transformed plasma p-Tau181 levels), the ADCNC-digital elimination subscore, and the CDR0106-personal care subscore.

[0128] As the most useful method for predicting the progression of cognitive impairment Figure 6F shows the relative importance (e.g., odds ratio) of specific significant inputs to the BEN model for predicting CD at 18 months, based on biomarkers, cognitive function, and brain imaging data inputs. Significant inputs in this model included a function of patient plasma p-Tau181 levels (e.g., standardized log2-transformed plasma p-Tau181 levels), CDR0106-personal care subscore, and combinations of MRI features, which were automatically selected by the BEN model. These brain regions were associated with the inferior parietal, inferior temporal, middle temporal, and superior temporal sulcus regions. These brain regions are shown in the brain heatmap in Figure 6F. Right lateral (RL) and left lateral (LL) images are shown as right and left lateral views.

[0129] Figures 6G and 6H show heatmaps illustrating the interaction between specific input values ​​and the predicted likelihood of cognitive impairment progression for a gradient boosting machine algorithm. Figure 6G shows the interaction between patient plasma p-Tau181 levels (e.g., standardized log2-transformed plasma p-Tau181 levels) and cognitive levels (e.g., ADAS-13). Figure 6H shows the interaction between patient plasma p-Tau181 levels and combinations of brain region MRI features associated with banks in the inferior parietal, inferior temporal, middle temporal, and superior temporal sulcus regions.

[0130] Example 2 Figures 7A to 7I describe a study on predicting brain amyloid-beta (Aβ) status using blood-based tests. Predictive models were trained using plasma biomarkers (Aβ42, Aβ40, and p-Tau181) for brain Aβ detection. In addition, the accuracy of these models was evaluated based on ApoE4 status, cognitive function measurements, and brain region measurements obtained from imaging data. The study demonstrated that predictive models designed to predict brain Aβ detection probability based on blood-based amyloid and tau signatures can be used as screening tools for detecting brain amyloid load in clinical trials and patient care.

[0131] In this study, Aβ42 and Aβ40 were measured in plasma samples from 513 subjects at screening using immunoprecipitation in conjunction with LC-MS / MS, and p-Tau181 was measured in plasma samples from 398 subjects (n=273 duplicates) using the Simoa Advantage V2 assay kit (immunoassay) by Quanterix. Over 90% of these subjects had mild cognitive impairment. For approximately 80% of the subjects, brain Aβ was assessed based on florbetaben-PET visual readings, while the remainder were assessed using florbetapir or flutemetamol.

[0132] Figure 7A shows cognitive level measurements for subjects with different combinations of demographic and input biomarker data. Some subjects had plasma Aβ measurements, some had plasma p-Tau181 measurements, and some had both types of measurements.

[0133] We trained linear regression-based machine learning models (Monte Carlo Logistic Lasso (MCL) and Bayesian Logistic Elastic Network (BEN)) and tree-based ensemble machine learning models (Regularized Random Forest (RRF) and Stochastic Gradient Boosting Machine (SGB)) to detect brain Aβ. All trained models considered demographic inputs, while some models considered improvements obtained from ApoE4 status, cognitive function levels, and brain region MRI measurements.

[0134] Model performance was evaluated using 10 iterations of stratified cross-validation within a 70% training set. Further evaluation was performed using a 30% holdout set. While plasma marker-based models were tested in the ADNI cohort, cognitive levels could not be examined due to limited data in this cohort.

[0135] Figure 7B shows a summary of model performance for different predictive model inputs. The results for the best-performing model are shown for each input combination. In the table, BA refers to the balanced accuracy (mean of sensitivity and specificity). Aβ refers to the Aβ42 / Aβ40 ratio. Each model includes only the specific optimal features identified by tuning hyperparameters.

[0136] In the best-performing predictive model using demographic data, plasma Aβ42 and Aβ40 levels, cognitive measurements, and ApoE4 status as inputs, an accuracy of 78% was achieved for brain Aβ detection, with a receiver operating characteristic area under the curve of 83.2%. In the best-performing model using demographic data and plasma Aβ42 and Aβ40 levels as inputs, an accuracy of 75.6% was achieved.

[0137] In the best-performing predictive model using demographic data, plasma p-Tau181 levels, cognitive measurements, and ApoE4 status as inputs, an accuracy of 82% was achieved for brain Aβ detection, with a receiver operating characteristic area under the curve of 87.4%. In the best-performing model using demographic data and plasma p-Tau181 levels as inputs, an accuracy of 76% was achieved.

[0138] The best-performing predictive model, using demographic data, plasma Aβ42 and Aβ40 levels, and plasma p-Tau181 levels as inputs, achieved an accuracy of 80.6% for detecting brain Aβ. The best-performing model, using demographic data, plasma Aβ42 and Aβ40 levels, plasma p-Tau181 levels, and cognitive measurements as inputs, achieved an accuracy of 82% for detecting brain Aβ, with a receiver operating characteristic area of ​​87.3%.

[0139] When further tested in an independent ADNI cohort, the best-performing model using demographic and plasma Aβ42 and Aβ40 levels as inputs achieved an accuracy of 75.7%; the best-performing model using demographic and plasma p-Tau181 levels as inputs achieved an accuracy of 75%, which improved to 78% when combined with plasma Aβ42 and Aβ40 levels.

[0140] Figures 7C through 7E show the input features with the greatest effect (in terms of odds ratios) on the output for three different stochastic gradient boosting machine (SGBM) predictive models, respectively. In these examples, the specific cognitive measures identified are CBBPAC (psychomotor attention), CBBLRAP (one-card learning accuracy), CBBID (identification), CBBMC (memory complex), and CBBDETSP (detection speed). Figure 7C shows the input features with the greatest effect (in terms of odds ratios) on the output for a model trained using plasma Aβ42 and Aβ40 levels, cognitive measures, and ApoE4 state. Figure 7D shows the input features with the greatest effect (in terms of odds ratios) on the output for a model trained using plasma p-Tau181 levels, cognitive measures, and ApoE4 state. Figure 7E shows the input features with the greatest effect (in terms of odds ratios) on the output for a model trained using plasma Aβ42 and Aβ40 levels, plasma p-Tau181 levels, and cognitive measurements.

[0141] Figure 7F shows the nonlinear sigmoid relationship between plasma Aβ42 and Aβ40 levels and the predicted probability of brain Aβ detection for the SGB machine learning model. An increase in the Aβ42 / Aβ40 ratio beyond a threshold resulted in a substantial decrease in the predicted probability of brain Aβ detection.

[0142] Figure 7G shows a nonlinear sigmoid relationship between plasma p-Tau181 levels and the predicted probability of brain Aβ detection for the SGB machine learning model. Increases in plasma p-Tau181 levels above a threshold resulted in a substantial increase in the predicted probability of brain Aβ detection.

[0143] Figures 7H and 7I show heatmaps illustrating the interaction between specific input values ​​and the predicted likelihood of brain Aβ detection. Figure 7H shows that the probability of brain Aβ detection increased with decreasing plasma Aβ42 levels and increasing plasma p-Tau181 levels. Figure 7I shows that the probability (probably) of brain Aβ detection increased with decreasing Aβ42 / Aβ40 ratio and increasing p-Tau181 / Aβ42 ratio.

[0144] Example 3 Figures 8A to 8Q describe a study on predicting the progression of cognitive impairment. Prediction used baseline demographic data, ApoE4 allele counts, and cognitive measures for each subject. Using only screening / baseline data, the long-term cognitive trajectory for each subject was predicted using a model trained on control data from previously conducted clinical trials.

[0145] A predictive model was developed (trained) using historical control (placebo) data from subjects with mild cognitive impairment (MCI) and mild Alzheimer's disease (AD) (n=955) from two clinical trials (NCT02956486 and NCT03036280). This model was constructed using a stochastic gradient boosting machine (SGBM) algorithm. The model was built using this training data with the R package "gbm".

[0146] Figure 8A shows the relative importance of key input features to the trained model. In this example, ADCIP (ideomotor execution), ADCRG (word recognition), and ADCDIF (difficulty finding words) are components of the ADAS13 score. The predictive performance of this model was evaluated in a placebo group (n = 231 participants with no missing input data) in another clinical trial (NCT01767311). A 43.2% Spearman correlation was observed between the predicted CDR-SB change from baseline at 18 months and the observed CDR-SB change.

[0147] Figure 8B shows the observed progression of cognitive impairment compared to the predicted progression of cognitive impairment for two models: a Bayesian long-term mixed-effects model and a stochastic gradient boosting model. Progression of cognitive impairment was measured as the change in CDR-SB from baseline. Participants were evaluated at approximately 3, 6, 9, 12, 15, and 18 months.

[0148] Figure 8C shows the correlation between observed and predicted cognitive impairment progression for individual subjects. As is evident from the figure, predicted and observed changes in cognitive performance from baseline correlated.

[0149] Figure 8D shows the correlation between observed and predicted values ​​for 3, 6, 9, 12, 15, and 18-month evaluations for both the Bayesian long-term mixed-effects model and the stochastic gradient boosting model.

[0150] The predictive model was retrained to use ADAS-13 cognitive measures instead of ADAS-14 cognitive measures. The retrained predictive model was used to predict the progression of cognitive impairment in the ADNI target sample. Figure 8E shows the relationship between predicted and observed cognitive impairment progression as measured at 6, 12, and 24-month assessments. Observed results were stratified by MMSE baseline values ​​(17–24, 25–28, and 29–30). As shown in Figure 8F, predicted and observed changes in cognitive measures correlated at 6, 12, and 24-month assessments.

[0151] Figures 8G to 8J show the relationship between four different predictive model inputs and predicted changes in cognitive measures (e.g., predicted change in CDR-SB from baseline) for a stochastic gradient boosting model. As shown in Figure 8G, the predicted change in cognitive measures shows an increase in the sigmoid response to an increase in the ADAS14 score. As shown in Figure 8H, the predicted change in cognitive measures shows an increase in the response to an increase in the FAQ score. As shown in Figure 8I, the predicted change in cognitive measures shows a decrease in the response to an increase in the MMSE score. As shown in Figure 8J, the predicted change in cognitive measures shows a thresholded increase in the response to an increase in the ADCRL subscore.

[0152] Figures 8K through 8N show heatmaps from a stochastic gradient boosting model illustrating the interaction between specific input values ​​and predicted changes in cognitive measures (e.g., a decrease in the CDR-SB change from baseline). As shown in Figure 8K, predicted changes in cognitive measures increased with increasing periods and ADAS14 scores. As shown in Figure 8L, predicted changes in cognitive measures increased with increasing ADAS14 scores and decreasing MMSE scores. As shown in Figure 8M, predicted changes in cognitive measures increased with increasing ADCRL subscores and increasing ADAS14 scores. As shown in Figure 8N, predicted changes in cognitive measures increased with decreasing CDR-SB baseline scores and increasing ADAS14 scores.

[0153] Figure 8O shows the relationship between observed and predicted CDR-SB scores for subjects at 6, 12, and 24-month assessments, stratified by MMSE baseline, in the ADNI trial. The observed and predicted CDR-SB scores show good alignment. Figure 8P shows the observed and predicted changes from baseline in CDR-SB for subjects in the ADNI trial dataset over time. Figure 8Q shows the observed and predicted CDR-SB values ​​for subjects in the ADNI trial dataset over time.

[0154] Example 4 Figures 9A to 9M relate to a study on predicting the progression of cognitive impairment in early Alzheimer's disease. Such predictions can be used to optimize clinical trials and monitor patients. The study examined predictive models trained using 1) baseline patient demographics and clinical cognitive assessments, and 2) baseline patient demographics, clinical cognitive assessments, and magnetic resonance imaging (MRI) scales of brain regions (volume, surface area, and cortical thickness).

[0155] The study used a training cohort of 905 early-stage AD patients from two clinical trials of the same study, and a validation cohort of 230 early-stage AD patients from another clinical trial. Cognitive performance (CDR-SB) was assessed at baseline and at 3, 6, 9, 12, and 18-month assessments.

[0156] In all three cohorts, the Desikan-Killiany atlas was used to generate brain MRI data (volume, surface area, and cortical thickness) for various target brain regions, yielding 207 region scales. Cortical thickness values ​​were expressed in millimeters (mm), and volume (mm) was also calculated. 3 ) and surface area (mm 2 The values ​​were normalized by intracranial volume to reduce inter-subject variability. Hubs and modules were identified from brain regions using MEGENA.

[0157] Predictive models were trained on a training cohort using baseline cognitive measures and demographics. These predictive models included regularized random forests, support vector machines, Bayesian lasso regression models, and stochastic gradient boosting models. Additional predictive models were then trained using baseline cognitive measures, demographics, and identified hubs and modules.

[0158] First, the predictive performance of the prediction model was evaluated using 10x cross-validation with 10 iterations within the training cohort. Next, the performance of the best-performing prediction model was evaluated in the first validation cohort via Spearman correlation between observed and predicted cognitive trajectories. The results for the stochastic gradient boosting model are reported here because it achieved the highest performance among the test models.

[0159] Figure 9A shows a demographic summary of the training cohort and the first validation cohort. Figure 9B shows a summary of the change in CDR-SB scores for the training cohort and the first validation cohort, who were diagnosed with MCI or mild AD at baseline.

[0160] Figure 8A shows the relative impact of the top 10 most important inputs on the best-performing predictive model when inputs are limited to baseline cognitive measures and demographics. Figure 9C shows the relative impact of the top 10 most important inputs on the best-performing predictive model when the input dataset for training the predictive model further includes hubs and modules (e.g., MRI prognostic features 910). Figure 9D shows heatmaps of the major hubs and modules referenced in Figure 9C. Major hubs and regions included VSMTCR (middle temporal area), VVIPCR (inferior parietal cortical volume), VSITR (inferior temporal cortical area), VCSFR (superior frontal cortical thickness), and the SBN module (SBN module including cortical areas around the inferior parietal, inferior temporal, middle temporal, and superior temporal sulci). Figures 9E to 9G show the impact of specific inputs on predicting cognitive impairment progression. These figures show the individual conditional expectation (ICE) profiles for each target and for the mean target (e.g., mean targets 931, 933, and 935). The vertical axis represents the normalized predicted CDR-SB change (normalized for declines at the minimum value of the predictor variable). The heterogeneity between targets in these ICE profiles is due to strong interactions between the predictor variables. These nonlinear relationships and interactions were explained by a stochastic gradient boosting algorithm without prior hypotheses.

[0161] Figures 9H to 9J show predictive interaction profiles with specific inputs in predicting cognitive impairment progression. These interaction profiles demonstrate the dependencies between these inputs. Figures 9H to 9J show that while all subjects exhibited a greater temporal change in CDR-SB scores, subjects with higher baseline ADAS-14 scores tended to exhibit a greater temporal change in CDR-SB scores. Furthermore, among subjects with higher baseline ADAS-14 scores, those with lower scores in specific components of the ADAS-14 assessment (e.g., particularly weak word recognition) experienced faster cognitive impairment progression. Similarly, subjects with lower middle temporal cortex area scores (e.g., VSMTCR scores) experienced faster cognitive impairment progression if they had higher baseline cognitive impairment (i.e., higher baseline ADAS-14 scores).

[0162] Figures 9K to 9L show a comparison between observed and predicted cognitive impairment for two predictive models using the first validation cohort. Figure 9K shows the predicted mean cognitive impairment using the two predictive models: the first predictive model was trained using baseline demographics and cognitive measures (Model 1), and the second predictive model was trained using baseline demographics, cognitive measures, and baseline nodes and hubs (Model 2). As can be observed in Figure 9K, predicted cognitive impairment was well tracked, with cognitive impairment observed on average for both models. Figure 9L shows the relationship between observed and predicted cognitive impairment for individual subjects, measured at 3, 6, 9, 12, 15, and 18-month assessments, for both models. Observed and predicted cognitive impairment values ​​correlated for subjects in the validation cohort.

[0163] Figure 9M shows the Spearman rank correlation between observed and predicted cognitive impairment for subjects in the first validation cohort. Cognitive impairment was predicted using two predictive models, Figures 9K and 9L. Additional columns include p-values ​​for comparing correlations. For predicted 18-month cognitive impairment, the first predictive model achieved a correlation of 0.425 with observed 18-month assessment (revealing 18.1% of the variance). Adding baseline nodes and hubs to the second model significantly improved overall predictive performance, revealing 25% of the variance at 18-month assessment.

[0164] Example 5 Figures 10A to 10W describe a study on predicting the progression of cognitive impairment in early Alzheimer's disease. This study trained a gradient-boosting predictive model to predict cognitive impairment progression in A+ early AD subjects over a duration of 18–24 months. The predictive model was trained using a training cohort (TC, n=934). The training cohort included historical placebo data from two clinical trials (NCT02956486 and NCT03036280) (both parts of the same Phase 3 program). Over 81% of the placebo subjects in the training cohort had mild cognitive impairment (MCI) attributable to AD, and the remainder had mild AD. The trained predictive models were evaluated in two validation cohorts (VC-1, n=235; VC-2, n=421). The first validation cohort (VC-1) included A+ early AD patients from the placebo arm of the 18-month clinical trial (NCT01767311). The second validation cohort (VC-2) included A+ subjects diagnosed with either MCI or AD, with at least one year of clinical follow-up and relevant clinical and MRI assessments from the ADNI database's ADNI-Stage 2 and ADNI-Stage 3. In the best-performing predictive model using cognitive scales and demographics, R values ​​of 0.21 and 0.31 were obtained for predicting 2-year cognitive decline in VC-1 and VC-2, respectively. 2 These were achieved respectively. In the best-performing predictive model using cognitive scales, demographics, and MRI features, VC-1 achieved an R of 0.29 by using the same preprocessing pipeline as TC. 2This was achieved. By using these model-based predictions for enriching clinical trials, the required sample size was reduced by 20%, to 49%.

[0165] Cognitive impairment was defined in terms of change from baseline on the CDR-SB. Cognitive measures included as input to the predictive model included a composite endpoint: Mini-Mental State Examination (MMSE), Alzheimer's Disease Assessment Scale Cognitive Subscale (ADAS-Cog-13), CDR-SB, and all of these subscores. In the training cohort, cognitive measures were assessed at 3, 6, 9, 12, 15, 18, 21, and 24 months. The clinical follow-up periods considered to evaluate the predictive models in the first and second validation cohorts were 3, 6, 9, 12, 15, and 18 months, and 6, 12, and 24 months, respectively.

[0166] All subjects in TC and VC-1 underwent 3.0 Tesla(T) structural MRI at baseline. Approximately 75% of subjects in VC-2 underwent 1.5T MRI, and the remainder underwent 3T MRI. Brain MRI data (volume, area, and cortical thickness) from all three cohorts were compiled for various target brain regions using the Desikan-Killiany atlas, yielding 207 region scales. Cortical thickness values ​​are expressed in millimeters (mm). Volume (mm 3 ) and area (mm 2 The values ​​were normalized (divided) by intracranial volume to reduce inter-subject variability and to clarify the variability attributable to head size within each cohort.

[0167] Figure 10A shows the demographic and clinical data characteristics of the training and validation cohorts. All demographic and clinical characteristics differed significantly between the cohorts (p<0.05). The training cohort had a significantly larger proportion of MCI and ApoE4-positive subjects. The first validation cohort (VC-1) had a larger proportion of males. Subjects in the second validation cohort (VC-2) were older and had higher body mass index (BMI). These differences in early AD subjects across different clinical trials and observed cohorts contributed to providing a more generalizable assessment of the predictive model's performance between the training and validation cohorts.

[0168] Figure 10B shows the long-term changes in CDR-SB at each evaluation point. The mean and standard deviation (SD) of the change in CDR-SB from baseline in the training cohort and validation cohort for each time point are shown, along with the number of subjects available for evaluation.

[0169] Using MEGENA, modules and hubs were obtained using 207 MRI region measures (volume, area, and cortical thickness) from the training cohort. As described herein, this process requires calculating the correlation of MRI measures across all pairs of regions. Regions with significant correlations were embedded on a spherical surface, representative edges (regions correlated with multiple other regions) were extracted, and a network was created by filtering the plane. Finally, a hierarchy of network modules was constructed by recursively clustering region measures with a coherent structure into network modules. This resulted in a total of 18 SBN modules and 45 hub region measures (displayed as SBN.1 to SBN.18). Depending on the nature of the correlation between adjacent regions, some regions were located within two or more network modules. Using MEGENA, the region measures in each SBN module were aggregated into a single composite eigenvalue for each target. Subsequent predictive modeling attempts focused exclusively on these 18 SBN modules and 45 hub region measures.

[0170] Predictive models were trained to predict the long-term cognitive trajectory for each subject, using baseline cognitive function data, demographic data, genomic data, and cognitive function measurement time as predictor variables. The predictive models were stochastic gradient boosting models. To train the models, up to 1000 decision trees were constructed with interactions of up to three predictor variables. The ranking and relative impact of each predictor variable were obtained by evaluating the decrease in mean squared error at each time point. The decision trees were then split using the predictor variables as root nodes in the SGBM algorithm and normalized to a range of 0-100%. Insights into the relationships between predictor variables and outcomes and interactions between predictor variables were obtained through partial dependency plots of individual conditional expectation (ICE) profiles and predictive profiles.

[0171] Model performance was evaluated using 10 iterations of 10x cross-validation within the time chain. Subsequently, the model was evaluated in VC-1 and VC-2. This evaluation was based on the coefficient of determination (R) of observed and predicted cognitive decline (CDR-SB change from baseline) at each time point. 2 This included measuring the mean squared error (MSE) and the mean absolute error.

[0172] Figure 10C shows the input variables ranked by their relative importance to the measured output variables for the first model. Another such predictive model was trained using the hub and modules in addition to the inputs above. Figure 10D shows the input variables ranked by their relative importance to the measured output variables for the second model. In addition to period, the primary baseline clinical predictors in these models were ADAS-13 score, MMSE, word recall and recognition, ideomotor execution, CDR-SB, and word finder difficulty, along with BMI and age. Ideomotor execution refers to the ability to perform a multilevel task, such as the sequence of steps required to brush one's teeth. Some of the primary MRI-based predictors include hub scales of middle temporal cortex area and inferior parietal cortex volume, along with scales for: i) a module including the bank of the inferior parietal gyrus, inferior temporal gyrus, middle temporal gyrus, and superior temporal sulcus (Figure 10E); ii) a module including the entorhinal cortex and temporal pole (Figure 10G); and iii) a module including the superior parietal gyrus, precuneus, cingulate gyrus, lateral occipital gyrus, postcentral gyrus, supramarginal gyrus, superior temporal gyrus, fusiform gyrus, lingual gyrus, and isthmus of the transverse temporal gyrus + the region shown in Figure 10E (Figure 10F).

[0173] Figures 10H to 10O illustrate the relationship between specific baseline inputs and predicted cognitive impairment progression at individual and mean target levels for two predictive models, using individual conditional expectation (ICE) profiles. ICE profiles were created by plotting individual and mean predicted outcomes for different values ​​of baseline inputs (e.g., thus creating mean target profiles 1011, 1013, 1015, 1017, 1021, 1023, 1025, and 1027), while keeping other input values ​​constant. As shown in Figures 10H to 10O, the ICE predictive profiles exhibit a strong sigmoid-like nonlinear relationship between each baseline input and cognitive impairment progression. These relationships represent an intermediate region between floor and ceiling effects and linear effects. Each target's predictive profile was centered by subtracting the predicted CDR-SB change corresponding to the lowest value of the predictor variable. The gradient-boosted predictive model allowed for modeling without pre-identifying or assuming relationships and inflection nodes of the predictor variables.

[0174] Figures 10P to 10S show interaction profiles between specific baseline inputs and predicted cognitive impairment. Figure 10P shows greater cognitive impairment over time in subjects with higher baseline ADAS-13 scores. Figure 10Q shows greater cognitive impairment in subjects with high baseline ADAS-13 scores and worsened ideomotor execution (ADCIP). Figure 10R shows greater cognitive impairment in subjects with high ADAS-13 scores and lower mid-temporal cortical area (VSMTCR). Figure 10S shows greater cognitive impairment in subjects with lower mid-temporal cortical area (VSMTCR) and lower area, volume, or thickness (SBN.15) in the entorhinal cortex and temporal pole.

[0175] Figures 10T and 10U show the mean and 95% confidence intervals for observed and predicted CDR-SB changes from baseline for the two models for both validation datasets. CDR-SB changes from baseline were predicted using the model based solely on baseline clinical features (Model 1), and also with the addition of hubs and modules (Model 2). Figure 10T shows observed and predicted CDR-SB changes from baseline for validation cohort 1. Figure 10U shows observed and predicted CDR-SB changes from baseline for validation cohort 2. The mean predicted cognitive impairment was well-followed and did not differ significantly from the mean observed cognitive impairment across all time points in both validation cohorts.

[0176] Figures 10V and 10W show the observed and predicted CDR-SB changes from baseline at each time point for each individual subject, using both validation datasets for both models. CDR-SB changes from baseline were predicted, along with a 95% prediction interval, using the models based solely on baseline clinical features (Model 1) and with the addition of hubs and modules (Model 2). Figure 10V shows the observed and predicted CDR-SB changes from baseline for validation cohort 1. Figure 10W shows the observed and predicted CDR-SB changes from baseline for validation cohort 2. Predictions of cognitive impairment for individual subjects from the two models were significantly correlated with observed cognitive impairment (p<0.001).

[0177] Figures 10X and 10Y illustrate the effects of using the disclosed predictive model for enriching clinical trials, consistent with the disclosed embodiments. As can be understood, patients may be selected for inclusion in clinical trials based on the predicted degree of cognitive impairment. Such patients may have a greater need for treatment. Furthermore, detectable treatment effects, trial size, and trial power can be improved by screening or selecting patients for inclusion in clinical trials based on the predicted degree of cognitive impairment.

[0178] In this study, 500 clinical trials were simulated using a bootstrap method (sampling by substitution) based on data from placebo arms of clinical trials used for VC-1, where active treatment and placebo were randomly assigned in a 1:1 ratio. The duration of the clinical trials was set to 18 months. The treatment effect, defined as the difference in the change from baseline in CDR-SB between the treatment and placebo groups at 18 months, was set at 30%. Next, for each simulated clinical trial, the impact of selecting only patients with a predicted 18-month CDR-SB change of at least 0.5 and 1 (enrichment scenarios 1 and 2, respectively) was evaluated by comparing the sample size requirements and power between unenriched and enriched clinical trials for these different enrichment scenarios. Sample size evaluation was based on a two-sample t-test.

[0179] Figure 10X and Figure 10Y show the impact on sample size reduction and increased power when enriching the clinical trial for subjects with predicted 18-month CDR-SB that meets two thresholds (ES1 of at least 0.5 and ES2 of at least 1). Approximately 88% and 65% of the clinical trial subjects used in VC-1 met these two criteria. The results for these two enrichment scenarios were determined by two prediction models. In the first prediction model, only baseline cognitive measurements and demographics were used (Model 1), and in the second prediction model, baseline cognitive measurements, demographics, and image data (e.g., hubs and modules) were used (Model 2).

[0180] In clinical trials that do not use a prediction model trained for patient selection or screening, a total sample size of 718 subjects (359 per group) was required to detect a 30% treatment effect from baseline in 18-month CDR-SB with 80% power.

[0181] Figures 10X and 10Y present the breakdown of prediction model performance. Figure 10X shows a summary of the prediction performance for two models in two validation cohorts (VC-1 and VC-2). In the first prediction model, cognitive scales and demographics were used (Model 1). In the second prediction model, cognitive scales, demographics, and image data (e.g., hubs and modules) were used (Model 2). The prediction metrics include the coefficient of determination (R 2 ), mean squared error (MSE), and mean absolute error (MAE) between the predicted clinical decline and the observed clinical decline (change in CDR-SB from baseline). For Model 1, to predict cognitive function decline at 18 months and 24 months in VC-1 and VC-2, respectively, it achieved R 2 values of 0.21 and 0.31, along with MSE values of 2.28 and 3.34, and MAE values of 1.16 and 1.35. The Model 1 predictions mostly used the same image processing pipeline as the training cohort (TC) for R 2Except for the MSE, it was equivalent to Model 2. In VC-1, Model 2 had a R of 0.29. 2 And it achieved MSE 2.08.

[0182] Figure 10Y shows the Pearson correlation coefficients for two validation cohorts (VC-1 and VC-2) for predicted and observed cognitive decline (CDR-SB change from baseline) for two predictive models. The first predictive model used cognitive measures and demographics (Model 1). The second predictive model used cognitive measures, demographics, and image data (e.g., hubs and modules) (Model 2).

[0183] As shown in Figure 10Z, using the first model increased the power to 88.3% and 96.5% for the two enrichment scenarios, respectively. With the power fixed at 80%, using the first model improved the ability to detect treatment effects from 30% to 26.7% and 22.3%, respectively. As shown in Figure 10AA, using the first model reduced the total sample size required to detect a 30% treatment effect for the two enrichment scenarios from 718 to 568 and 398, respectively (a decrease of 20.9% and 44.6%). Using the second model improved these numbers: for the two enrichment scenarios, the power increased to 89.2% and 97.6%, respectively, and the minimum treatment effect detectable with 80% power improved from 30% to 26.3% and 21.3% (Figure 10Z). The total sample size required to detect a 30% treatment effect decreased from 718 to 552 and 364 for the two enrichment scenarios using Model 2 predictions, respectively (a decrease of 23.2% and 49.4%) (Figure 10AA).

[0184] By using predictive models to screen patients, the number of patients requiring screening can be reduced. For the VC-1 clinical trial population and the use of predictive model 2, approximately 89% and 62% met the enrichment criteria for ES1 and ES2, respectively. The total sample size required to detect 30% treatment effect was reduced from 718 to 552 and 364, respectively (a reduction of 23.2% and 49.4%). Therefore, instead of screening 718 subjects, it is necessary to exclusively screen 620 subjects in ES1 (552 divided by 0.89) and 587 subjects in ES2 (364 divided by 0.62).

[0185] Therefore, in addition to a significant reduction in sample size requirements and an increase in statistical power, screening patients using the trained predictive models described herein may enable screening of subjects reduced by 13.6% and 18.2% (compared to, for example, methods without enrichment). More importantly, there may be other practical merits / needs for such enrichment methods in clinical trials, for example, when a candidate treatment is expected to benefit only subjects likely to experience mild to moderate cognitive decline.

[0186] The foregoing description is provided for illustrative purposes only. It is not exhaustive and is not limited to the exact form or embodiment disclosed. Modifications and adaptations of embodiments will be apparent from the consideration of this specification and the practice of the disclosed embodiments. For example, the implementations described include hardware, but systems and methods consistent with this disclosure can be performed by hardware and software. Furthermore, while certain components are described to work in conjunction with each other, such components can be integrated with each other or distributed in any preferred manner.

[0187] Embodiments herein include systems, methods, and tangible computer-readable non-temporary media. A method may be performed at least in part by at least one processor receiving instructions from, for example, a tangible computer-readable non-temporary storage medium. Similarly, a system consistent with this disclosure may include at least one processor and memory, the memory of which may be a tangible computer-readable non-temporary storage medium. As used herein, tangible computer-readable non-temporary storage medium refers to any type of physical storage in which information or data readable by at least one processor may be stored. Examples include random-access memory (RAM), read-only memory (ROM), volatile memory, non-volatile memory, hard drives, CD-ROMs, DVDs, flash drives, disks, registers, caches, and any other known physical storage media. Singular terms such as “memory” and “computer-readable storage medium” may further refer to multiple structures, such as multiple memories or computer-readable storage media. As used herein, “memory” may include any type of computer-readable storage medium unless otherwise specified. A computer-readable storage medium may store instructions for execution by at least one processor, including instructions for causing a processor to perform steps or stages consistent with the embodiments herein. In addition, one or more computer-readable storage mediums may be used in the execution of a computer execution method. The term “computer-readable non-temporary storage medium” should be understood to include tangible items and exclude carrier and transient signals.

[0188] Furthermore, while exemplary embodiments are described herein, the scope includes all embodiments having equivalent elements, modifications, omissions, combinations (e.g., aspects across various embodiments), adaptations, or modifications based on this disclosure. The elements of the claims should be interpreted broadly based on the terminology used in the claims and should not be limited to the examples described herein or in the correspondence thereof, and the examples hereof should be interpreted non-exclusively. Furthermore, the steps of the disclosed methods may be modified in any way, including rearranging steps or insertion or removal steps.

[0189] The features and advantages of this disclosure are evident from the detailed specification, and it is intended that the appended claims encompass all systems and methods included in the true spirit and scope of this disclosure. As used herein, the indefinite articles “a” and “an” mean “one or more.” Similarly, the use of plural terms does not necessarily imply plural unless it is clear in a given context. Furthermore, since numerous modifications and variations readily arise from the testing of this disclosure, it is not desirable to limit this disclosure to the exact configurations and operations illustrated and described, and all suitable modifications and equivalents may be appealed to be included within the scope of this disclosure. Accordingly, the disclosed embodiments and examples are considered merely illustrative, and the true scope of this disclosure is intended to be indicated by the following claims and their equivalents.

[0190] Embodiments can be further described using the following clauses: 1. A system comprising at least one processor; and a non-temporary medium readable by at least one computer, which, when executed by at least one processor, comprises baseline cognitive data and image data for each first subject, including baseline data for the first subject, which includes one or more brain region measurements for one or more brain regions identified as hubs or one or more composite values ​​for one or more clusters of brain regions identified as modules, wherein one or more hubs or one or more modules are identified using network analysis or multilevel clustering; and A system including at least one computer-readable non-temporary medium, which includes instructions to cause the system to perform an operation comprising: acquiring training data which includes cognitive impairment progression data for a first subject, which is acquired over time and includes repeated measures acquired after baseline cognitive data; training a predictive model using the training data to predict cognitive impairment progression data for the first subject using baseline data for the first subject; acquiring baseline data for a second subject that satisfies a cognitive impairment state; and predicting cognitive impairment progression data for the second subject by inputting the baseline data for the second subject into the trained predictive model. 2. The system described in Clause 1, wherein repeated measures include or rely on Clinical Dementia Severity Assessment Scale (CDR-SB) measures; Alzheimer's Disease Composite Score (ADCOMS) measures; or Alzheimer's Disease Assessment Scale (ADAS) measures. 3. A system described in any one of clauses 1 to 2, in which one or more brain region measurements for one or more hubs or one or more composite values ​​for one or more modules are based on MRI, CT, or PET images. 4. A system described in any one of clauses 1 to 3, in which one or more brain region measurements for one or more hubs or one or more composite values ​​for one or more modules are based on MRI images. 5. A system as described in any one of clauses 1 to 4, in which one or more hubs or one or more modules are identified using network analysis or multilevel clustering. 6. One or more hubs or one or more modules, A system according to any one of Clauses 1 to 5, which is identified by creating a planar filtered network graph including nodes corresponding to brain regions and edges corresponding to correlations between brain region measurements for brain regions; creating a hierarchy of network modules including modules by iteratively clustering nodes into network modules using the planar filtered network graph; and identifying nodes as hubs including hubs using connectivity within clusters of nodes. 7. A system according to any one of clauses 1 to 6, wherein one or more hubs or one or more modules are identified using multiscale embedded gene co-expression network analysis (MEGENA). 8. The system according to any one of the clauses 1 to 7, wherein the one or more brain region measurements for one or more brain regions include volume, surface area, or thickness measurements. 9. The system according to Clause 8, wherein one or more hubs include a first hub containing a middle temporal cortex region, a second hub containing an inferior parietal cortex region, a third hub containing an inferior temporal cortex region, or a fourth hub containing a superior frontal cortex region. 10. A system described in any one of clauses 1 to 9, in which one or more composite values ​​are based on volume, surface area, or cortical thickness measurements for brain regions within one or more modules. 11. The system according to Clause 10, wherein one or more modules include a first network module comprising the inferior parietal region, inferior temporal region, middle temporal region, and cortical region around the superior temporal sulcus; or a second network module comprising the entorhinal cortical region and temporal pole region. 12. A system described in any one of Clauses 1 to 11, wherein baseline cognitive data for the first subject includes at least one of the following: Cogstate Brief Battery score, International Shopping List Test score, ADAS score, Mini-Mental State Examination (MMSE) score, CDR-SB score, or FAQ score. 13. A system described in any one of Clauses 1 to 12, wherein baseline cognitive data for the patient includes at least one of the following: ADCRL word recall score, ADCIP ideomotor performance score, ADCRG word recognition score, ADCDIF word difficulty score, CDR0106 personal care score, ADCCP constructive practice score, ADCNC digit elimination score, ADCOR orientation score, CDR0102 orientation score, CDR0103 judgment and problem solving score, or ADCDRL delayed word recall score. 14. A system described in any one of Clauses 1 to 13, wherein the baseline data for the first subject further includes demographic data, including age, sex, or BMI, for the first subject. 15. A system described in any one of clauses 1 to 14, wherein baseline data for the first subject further includes genomic data. 16. The system described in Clause 15, in which the genomic data includes the number of ApoE4 alleles. 17. The system according to any one of Clauses 1 to 16, wherein the baseline data for the first subject further includes plasma, serum, or cerebrospinal fluid biomarker data. 18. Plasma, serum, or cerebrospinal fluid biomarker data for the first subject showing one or more of the following: cerebrospinal fluid Aβ1-42 score, cerebrospinal fluid Aβ1-40 score, combination of cerebrospinal fluid Aβ1-42 and cerebrospinal fluid Aβ1-40 score, cerebrospinal fluid Aβ1-42 score to Aβ1-40 score ratio, cerebrospinal fluid total tau score, cerebrospinal fluid neurolanin score, cerebrospinal fluid neurofilament light (NfL) peptide score, or cerebrospinal fluid microtubule-binding region (MBTR)-tau score, or serum or plasma level A The system according to Clause 17, further comprising one or more of the following: β1-42 score, serum or plasma level Aβ1-40 score, combination of serum or plasma level Aβ1-42 and Aβ1-40 score, serum or plasma level Aβ1-42 score to Aβ1-40 score ratio, serum or plasma level total tau score, serum or plasma level phosphorylated tau score, serum or plasma level glial fibrillary acidic protein (GFAP) score, or serum or plasma level NfL, peptide score. 19. The system according to Clause 18, wherein the serum or plasma-level phosphorylated tau score includes a tau (p-Tau181) score phosphorylated at serum or plasma level 181, a tau (p-Tau217) score phosphorylated at serum or plasma level 217, or a tau (p-Tau231) score phosphorylated at serum or plasma level 231. 20. A system according to any one of clauses 1 to 19, wherein the image data further includes brain region MRI measurements that depend on one or more of the whole brain volume, cortical thickness, or whole hippocampal volume. 21. A system described in any one of Clauses 1 to 20, wherein the image data further includes one or more of the following: taus score, amyloid PET score, or fluorodeoxyglucose (FDG) PET score. 22. A system described in any one of Clauses 1 to 21, in which the fulfillment of a cognitive impairment state depends on a diagnosis of a primary subject of neurological disorder, functional impairment, or injury. 23. A neurological disorder, impairment, or injury of the system as described in Clause 22, including mild cognitive impairment, Alzheimer's disease, or dementia. 24. A system described in any one of the clauses 1 to 23, wherein the primary subject is amyloid-positive. 25. A system described in any one of Clauses 1 to 24, wherein cognitive impairment progression data for the second subject indicates progression from mild cognitive impairment to Alzheimer's disease. 26. A system described in any one of Clauses 1 to 25, in which cognitive impairment progression data for the second subject includes changes in CDR-SB from baseline. 27. The system according to Clause 26, wherein the predictive model shows a reducing relationship between the change in CDR-SB from baseline and brain region measurements for one or more hubs or composite values ​​for one or more modules. 28. The system according to Clause 27, wherein the hub includes a middle temporal cortex region; the hub includes an inferior parietal cortex region; the module includes a bank of the inferior parietal region, inferior temporal region, middle temporal region, or superior temporal sulcus region; or the module includes the entorhinal cortex or temporal pole region. 29. A system according to any one of clauses 27-28, wherein baseline cognitive data includes ADCRL, ADCIP, or ADAS-13 scores; and the predictive model shows an increasing relationship between ADCRL, ADCIP, or ADAS-13 scores and the change in CDR-SB from baseline. 30. A system described in any one of Clauses 1 to 29, wherein the operation further includes: obtaining trial data, including baseline data, for multiple participants, including a second subject, in a clinical trial of the treatment of Alzheimer's disease; predicting cognitive impairment progression data for multiple participants using a trained predictive model and baseline data for multiple participants; and determining the effectiveness of the treatment of Alzheimer's disease, partially using the cognitive impairment progression data for multiple participants. 31. A system described in any one of Clauses 1 to 29, further comprising screening or selecting candidate patients for inclusion in a clinical trial using a trained predictive model. 32. A system described in any one of clauses 1 to 31, in which the predictive model includes a tree-based model. 33. A tree-based model, including a gradient boosting model, as described in Clause 32. 34. A system described in any one of Clauses 1 to 31, wherein the prediction model includes a Bayesian elastic network model; a Bayesian nonlinear regression model; or a neural network model. 35. A system described in any one of clauses 1 to 34, wherein the time elapsed between the acquisition of baseline cognitive data for the first subject and the acquisition of final repeated measure data for the first subject is between 12 and 36 months. 36. The system described in Clause 35, wherein the time elapsed between the acquisition of baseline cognitive data for the first subject and the acquisition of final repeated measure data for the first subject is between 18 and 24 months. 37. A system described in any one of Clauses 1 to 36, in which repeated measurements are taken at time intervals of 3 to 12 months. 38. A system comprising at least one processor; and at least one computer-readable non-temporary medium, which, when executed by the at least one processor, includes instructions causing the system to perform an operation comprising: acquiring training data for a first subject satisfying a cognitive impairment state, wherein for each first subject, the training data includes baseline data for the first subject, comprising baseline data including plasma biomarker data; and cognitive impairment progression data for the first subject; training a predictive model using the training data to predict cognitive impairment progression data for the first subject using the baseline data for the first subject; acquiring baseline data for a second subject satisfying a cognitive impairment state; and predicting cognitive impairment progression data for the second subject by inputting the baseline data for the second subject into the trained predictive model. 39. The system according to Clause 38, wherein the plasma biomarker is p-Tau181, Aβ1-42, or Aβ1-40 biomarker. 40. A system comprising at least one processor; and at least one computer-readable non-temporary medium, which, when executed by the at least one processor, includes instructions causing the system to perform an operation comprising: acquiring training data for a first subject satisfying a cognitive impairment state, wherein for each first subject, the training data includes baseline data for the first subject, comprising plasma biomarker data and brain amyloid data for the first subject; training a predictive model using the training data to predict the brain amyloid state for the first subject using the baseline data for the first subject; acquiring baseline data for a second subject satisfying a cognitive impairment state; and predicting the brain amyloid state for the second subject by inputting the baseline data for the second subject into the trained predictive model. 41. The system according to Clause 40, wherein the plasma biomarker is p-Tau181, Aβ1-42, or Aβ1-40 biomarker.

[0191] As used herein, unless otherwise specifically stated, the term “or” encompasses all possible combinations unless they are impossible to implement. For example, if it is stated that a component may include A or B, then unless otherwise specifically stated or impossible to implement, the component may include A, or B, or A and B. As a second example, if it is stated that a component may include A, B, or C, then unless otherwise specifically stated or impossible to implement, the component may include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.

[0192] Other embodiments will be apparent from the discussion herein and the practice of the embodiments disclosed herein. This specification and the examples are considered to be illustrative only, and the true scope and spirit of the disclosed embodiments are intended to be shown by the following claims.

Claims

1. It is a system, at least one processor; and A non-temporary medium readable by at least one computer, when executed by the at least one processor, Training data for a first subject that meets the criteria for a cognitive impairment state, for each first subject, Baseline data for the first subject, comprising baseline cognitive data and image data including one or more brain region measurements for one or more brain regions identified as hubs or one or more composite values ​​for one or more clusters of brain regions identified as modules, wherein the one or more hubs or one or more modules are identified using network analysis or multilevel clustering; and Cognitive impairment progression data for the first subject, which is acquired over time and includes repeated measures acquired after the baseline cognitive data. Obtaining training data that includes; The predictive model is trained using the aforementioned training data to predict the progression of cognitive impairment in a first subject using baseline data for the first subject; To obtain the baseline data for a second subject that meets the aforementioned cognitive impairment condition; The cognitive impairment progression data for the second subject is predicted by inputting the baseline data for the second subject into the trained predictive model, A non-temporary medium readable by at least one computer, including an instruction causing the system to perform an operation including the above. A system that includes this.

2. The repeated measurement values ​​are Clinical Dementia Severity Assessment Scale (CDR-SB) measurement values; Alzheimer's Disease Composite Score (ADCOMS) measurement; or Alzheimer's Disease Assessment Scale (ADAS) measurement values The system according to claim 1, which includes or depends on the following.

3. The system according to claim 1, wherein the one or more brain region measurements for the one or more hubs or the one or more composite values ​​for the one or more modules are based on MRI images, CT images, or PET images.

4. The system according to claim 1, wherein the one or more brain region measurements for the one or more hubs or the one or more composite values ​​for the one or more modules are based on MRI images.

5. The system according to claim 1, wherein the one or more hubs or the one or more modules are identified using network analysis or multilevel clustering.

6. The one or more hubs or the one or more modules Create a planar filtered network graph that includes nodes corresponding to brain regions and edges corresponding to the correlation between brain region measurements for said brain regions; A hierarchy of network modules including the aforementioned modules is created by iteratively clustering the nodes in the network modules using the planar filtered network graph; Using the connectivity within the cluster between the aforementioned nodes, the node is identified as a hub, including the aforementioned hub. The system according to claim 1, as identified by...

7. The system according to claim 1, wherein one or more hubs or one or more modules are identified using multiscale embedded gene co-expression network analysis (MEGENA).

8. The system according to claim 1, wherein the one or more brain region measurements for the one or more brain regions include volume, surface area, or thickness measurements.

9. The one or more hubs described above The first hub, including the mesotemporal cortex region, The second hub, including the inferior parietal cortex region, A third hub including the inferior temporal cortex region; or The fourth hub, including the superior frontal cortex region The system according to claim 8, including the above.

10. The system according to claim 1, wherein the one or more composite values ​​are based on volume, surface area, or cortical thickness measurements of brain regions within the one or more modules.

11. The one or more modules described above A first network module including the inferior parietal region, inferior temporal region, middle temporal region, and cortical region around the superior temporal sulcus; or A second network module including the entorhinal cortex and temporal pole regions. The system according to claim 10, including the following:

12. The system according to claim 1, wherein the baseline cognitive data for the first subject includes at least one of the following: Cogstate Brief Battery score, International Shopping List Test score, ADAS score, Mini-Mental State Examination (MMSE) score, CDR-SB score, or FAQ score.

13. The system according to claim 1, wherein the baseline cognitive data for the patient includes at least one of the following: ADCRL word recall score, ADCIP ideomotor execution score, ADCRG word recognition score, ADCIF word discovery difficulty score, CDR0106 personal care score, ADCCP constructive practice score, ADCNC digit elimination score, ADCOR orientation score, CDR0102 orientation score, CDR0103 judgment and problem solving score, or ADCDR delayed word recall score.

14. The system according to claim 1, wherein the baseline data for the first subject further comprises demographic data including age, sex, or BMI for the first subject.

15. The system according to claim 1, wherein the baseline data for the first subject further includes genomic data.

16. The system according to claim 15, wherein the genome data includes the number of ApoE4 alleles.

17. The system according to claim 1, wherein the baseline data for the first subject further comprises plasma, serum, or cerebrospinal fluid biomarker data.

18. The plasma, serum, or cerebrospinal fluid biomarker data for the first subject is One or more of the following: cerebrospinal fluid Aβ1-42 score, cerebrospinal fluid Aβ1-40 score, combination of cerebrospinal fluid Aβ1-42 and cerebrospinal fluid Aβ1-40 score, cerebrospinal fluid Aβ1-42 score to Aβ1-40 score ratio, cerebrospinal fluid total tau score, cerebrospinal fluid neurolanin score, cerebrospinal fluid neurofilament light (NfL) peptide score, or cerebrospinal fluid microtubule-binding region (MBTR)-tau score, or One or more of the following: serum or plasma level Aβ1-42 score, serum or plasma level Aβ1-40 score, combination of serum or plasma level Aβ1-42 and Aβ1-40 scores, serum or plasma level Aβ1-42 score to Aβ1-40 score ratio, serum or plasma level total tau score, serum or plasma level phosphorylated tau score, serum or plasma level glial fibrillary acidic protein (GFAP) score, or serum or plasma level NfL, peptide score. The system according to claim 17, further comprising one or more of the above.

19. The system according to claim 18, wherein the serum or plasma-level phosphorylated tau score includes a tau (p-Tau181) score phosphorylated at serum or plasma level 181, a tau (p-Tau217) score phosphorylated at serum or plasma level 217, or a tau (p-Tau231) score phosphorylated at serum or plasma level 231.

20. The system according to claim 1, wherein the image data further includes brain region MRI measurements that depend on one or more of the total brain volume, cortical thickness, or total hippocampal volume.

21. The system according to claim 1, wherein the image data further comprises one or more of the following: a taus score, an amyloid PET score, or a fluorodeoxyglucose (FDG) PET score.

22. The system according to claim 1, wherein the satisfaction of the cognitive impairment state depends on a diagnosis of the first subject of neurological disease, functional impairment, or injury.

23. The system according to claim 22, wherein the neurological disorder, functional impairment, or injury includes mild cognitive impairment, Alzheimer's disease, or dementia.

24. The system according to claim 1, wherein the first subject is amyloid-positive.

25. The system according to claim 1, wherein the cognitive impairment progression data for the second subject shows progression from mild cognitive impairment to Alzheimer's disease.

26. The system according to claim 1, wherein the cognitive impairment progression data for the second subject includes a change in CDR-SB from baseline.

27. The system according to claim 26, wherein the predictive model shows a reduction relationship between the change in CDR-SB from baseline and the brain region measurement values ​​for the hubs of one or more hubs or the composite values ​​for the modules of one or more modules.

28. The hub includes the middle temporal cortex region; The hub includes the inferior parietal cortex region; The module includes a bank of the inferior parietal region, inferior temporal region, middle temporal region, or superior temporal sulcus region; or The module includes the entorhinal cortex or the temporal pole region, The system according to claim 27.

29. The baseline cognitive data includes ADCRL, ADCIP, or ADAS-13 scores; and The prediction model shows an increasing relationship between the ADCRL, ADCIP, or ADAS-13 score and the change in CDR-SB from baseline. The system according to claim 27.

30. The aforementioned operation, To obtain trial data, including baseline data, for multiple participants, including the second subject, in a clinical trial for the treatment of Alzheimer's disease; Using the trained predictive model and baseline data for the aforementioned participants, predict the progression of cognitive impairment for the aforementioned participants; Partially using the cognitive impairment progression data from the aforementioned multiple participants to determine the effectiveness of the Alzheimer's disease treatment, The system according to claim 1, further comprising:

31. The aforementioned operation, Screening or selecting candidate patients for inclusion in clinical trials using the trained predictive model. The system according to claim 1, further comprising:

32. The system according to claim 1, wherein the prediction model includes a tree-based model.

33. The system according to claim 32, wherein the tree-based model includes a gradient boosting model.

34. The aforementioned prediction model, Bayesian elastic net model; Bayesian nonlinear regression model; or Neural network model The system according to claim 1, including the following:

35. The system according to claim 1, wherein the elapsed time between the acquisition of the baseline cognitive data for the first subject and the acquisition of the final repeated measurement data for the first subject is 12 months to 36 months.

36. The system according to claim 35, wherein the elapsed time between the acquisition of the baseline cognitive data for the first subject and the acquisition of the final repeated measurement data for the first subject is 18 to 24 months.

37. The system according to claim 1, wherein the repeated measurements are acquired at time intervals of 3 to 12 months.