Deep Learning to Model Disease Progression
Patent Information
- Application Number
- JP2024528618
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-11-15
- Filing Date
- 2022-11-14
- Publication Date
- 2025-11-07
AI Technical Summary
Current machine learning methods for predicting the progression of neurological diseases like Alzheimer's disease are limited by the lack of large homogeneous datasets and noisy endpoints, leading to inaccurate and non-personalized predictions.
A multimodal, multitask deep learning model that integrates clinical data and neuroimaging data using adversarial losses and sharpness-aware minimization techniques to predict disease progression, mitigating study-specific biases and improving generalization.
The model provides accurate and personalized predictions of neurological disease progression, enhancing the efficiency of evaluating new treatments and increasing the power of treatment effect estimates in clinical trials.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 279,606, entitled “MODELING THE PROGRESSION OF A DISEASE WITH A MULTI-MODAL NEURAL NETWORK,” filed November 15, 2021, which is incorporated by reference in its entirety.
[0002] The present disclosure relates generally to machine learning, and more particularly, to deep learning for modeling disease progression. [Background technology]
[0003] Alzheimer's disease is the most common cause of dementia in people over 65 years of age, affecting 26.6 million people worldwide. Alzheimer's disease, like several other neurological disorders, is a slowly progressive disease caused by degeneration of brain cells, and patients show clinical symptoms several years after the onset of the disease. Therefore, accurate diagnosis and treatment of Alzheimer's disease, among other neurological disorders, at its early stages (e.g., mild cognitive impairment (MCI)) can help prevent irreversible and fatal brain damage. Despite the demand, little progress has been made in predicting the progression of neurological disorders such as Alzheimer's disease, due to the complexity of modeling the progression, the lack of large uniform datasets that include early-stage Alzheimer's disease patients, and noisy endpoints that can be generally difficult to predict. Summary of the Invention
[0004] Methods, systems, and articles of manufacture, including computer program products, are provided for deep learning to model disease progression. In one aspect, a system is provided. The system may include at least one data processor and at least one memory. The at least one memory may store instructions that, when executed by the at least one data processor, result in operations. The operations may include generating, by the machine learning model, a first feature representation based on clinical data associated with a baseline cognitive state of the patient. The operations may also include generating, by the machine learning model, a second feature representation based on an image of the patient's brain. The operations may also include generating, by the machine learning model, a set representation by at least fusing the first feature representation and the second feature representation. The operations may also include predicting, by the machine learning model, a change in the baseline cognitive state over a period of time based at least on the set representation.
[0005] In another aspect, a method for deep learning to model disease progression is provided. The method may include generating, by a machine learning model, a first feature representation based on clinical data associated with a baseline cognitive state of the patient. The method may also include generating, by the machine learning model, a second feature representation based on an image of the patient's brain. The method may also include generating, by the machine learning model, a set representation by at least fusing the first feature representation and the second feature representation. The method may also include predicting, by the machine learning model, a change in the baseline cognitive state over a period of time based at least on the set representation.
[0006] In another aspect, a computer program product is provided that includes a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium may include program code that, when executed by at least one data processor, causes operations. The operations may include generating, by a machine learning model, a first feature representation based on clinical data associated with a baseline cognitive state of the patient. The operations may also include generating, by a machine learning model, a second feature representation based on an image of the patient's brain. The operations may also include generating, by the machine learning model, a set representation by at least fusing the first feature representation and the second feature representation. The operations may also include predicting, by the machine learning model, a change in the baseline cognitive state over a period of time based at least on the set representation.
[0007] In some variations of the methods, systems, and non-transitory computer-readable media, one or more of the following features may be included, optionally in any feasible combination.
[0008] In some variations, the fusing is performed using one or more fusion techniques including at least one of concatenation, addition, simple attention, scaled dot product attention, applying a tensor fusion network, low-rank fusion, and unidirectional context attention.
[0009] In some variations, the first feature representation is an encoded vector that includes a concatenation of at least one of a current cognitive score representing a baseline cognitive state of the patient, demographic information associated with the patient, and genomic information associated with the patient.
[0010] In some variations, the second feature representation includes at least one domain-invariant embedding feature.
[0011] In some variations, the machine learning models include a first machine learning model trained to generate a first feature representation, a second machine learning model trained to generate a second feature representation, a third machine learning model trained to generate a set representation, and a fourth machine learning model trained to predict change in the baseline cognitive state over a period of time.
[0012] In some variations, the machine learning model is trained based at least on multiple modalities including clinical data associated with the patient's baseline cognitive state and images of the patient's brain.
[0013] In some variations, the machine learning model is pre-trained to predict a patient's baseline cognitive state based at least on multiple brain images acquired at multiple time points and across multiple regions.
[0014] In some variations, the machine learning model is trained by at least adversarially training a region detector of the machine learning model to reduce inter-study region shifts associated with images of the patient's brain.
[0015] In some variations, the adversarial training includes adversarially training a feature extraction network of the machine learning model to learn domain-invariant features for generating the second feature representation based at least on the image of the patient's brain.
[0016] In some variations, the adversarial training includes applying an inverse gradient to the second feature representation to generate the region detector input. The adversarial training of the region detector is based at least on the region detector input.
[0017] In some variations, the region detector accounts for drift in inter-study region shifts in guesses.
[0018] In some variations, the change in baseline cognitive status over time indicates the progression of Alzheimer's disease in the patient.
[0019] In some variations, the time period is 12 months.
[0020] In some variations, the image is a three-dimensional magnetic resonance imaging image that includes the inferred mask.
[0021] In some variations, the baseline cognitive status is represented by at least one cognitive score including at least one of a Clinical Dementia Rating Scale (CDRSB) score, an Alzheimer's Disease Assessment Scale-Cognitive Subscale (ADAS-COG12) score, and a Mini-Mental State Examination (MMSE) score.
[0022] In some variations, a system is provided for generating an output indicative of a prediction of disease progression in a subject. The system performs operations including receiving longitudinal physiological data representative of the subject, the longitudinal physiological data including multiple modalities indicative of disease progression in the subject. The system also performs operations including applying a trained multi-modal neural network to the longitudinal physiological data. The trained multi-modal neural network includes multiple feed-forward neural networks configured to learn modality-specific deep features related to the multiple modalities. The system also performs operations including generating an output indicative of a prediction of disease progression in the subject based at least on the modality-specific deep features related to the multiple modalities of the longitudinal physiological data.
[0023] In some variations, the operations further include performing a plurality of domain shifts of the modality-specific deep features, the plurality of domain shifts configured to normalize the modality-specific deep features to a single domain, combining the normalized modality-specific deep features to form a feature space in the single domain, the feature space indicative of disease progression, and generating an output indicative of a prediction of disease progression in the subject based at least on the normalized modality-specific deep features in the feature space. Additionally, the operations further include debiasing each of the normalized modality-specific deep features using an adversarial loss to mitigate study-specific bias corresponding to a modality of the plurality of modalities.
[0024] In some variations, the operations further include applying a sharpness-aware minimization optimization technique to refine the multiple region shifts of the modality-specific deep features into a single region. Additionally, each modality of the multiple modalities is fed into a separate feedforward neural network of the multiple feedforward neural networks. Additionally, the long-term physiological data includes baseline physiological data representing the subject, the baseline physiological data includes a first set of modalities indicative of disease progression, and the physiological data further includes a baseline image related to the subject's brain. Additionally, the long-term physiological data includes delta physiological data representing the subject, the delta physiological data includes a second set of modalities indicative of disease progression, and the delta physiological data further includes a delta image related to the subject's brain. In some variations, the trained multi-modal neural network includes a feature extraction network, an endpoint prediction network, and a region detector. Additionally, the disease is Alzheimer's disease, and wherein the output is a normalized representation.
[0025] In another aspect, a computer-readable storage medium for generating an output indicative of a prediction of disease progression in a subject is provided. The computer-readable storage medium includes instructions including receiving longitudinal physiological data representative of a subject, the longitudinal physiological data including multiple modalities indicative of disease progression in the subject. The computer-readable storage medium also includes instructions including applying a trained multi-modal neural network to the longitudinal physiological data. The trained multi-modal neural network includes multiple feed-forward neural networks configured to learn modality-specific deep features related to the multiple modalities. The computer-readable storage medium also includes instructions including generating an output indicative of a prediction of disease progression in the subject based at least on the modality-specific deep features related to the multiple modalities of the longitudinal physiological data.
[0026] In some variations, the instructions further include performing a plurality of domain shifts of the modality-specific deep features, the plurality of domain shifts configured to normalize the modality-specific deep features to a single domain, combining the normalized modality-specific deep features to form a feature space in the single domain, the feature space indicative of disease progression, and generating an output indicative of a prediction of disease progression in the subject based at least on the normalized modality-specific deep features in the feature space. Additionally, the instructions further include debiasing each of the normalized modality-specific deep features using an adversarial loss to mitigate study-specific bias corresponding to a modality of the plurality of modalities.
[0027] In some variations, the instructions further include applying a sharpness-aware minimization optimization technique to refine the multiple region shifts of the modality-specific deep features into a single region. Additionally, each modality of the multiple modalities is fed into a separate feedforward neural network of the multiple feedforward neural networks. Additionally, the long-term physiological data includes baseline physiological data representing the subject, the baseline physiological data includes a first set of modalities indicative of disease progression, and the physiological data further includes a baseline image related to the subject's brain. Additionally, the long-term physiological data includes delta physiological data representing the subject, the delta physiological data includes a second set of modalities indicative of disease progression, and the delta physiological data further includes a delta image related to the subject's brain. In some variations, the trained multi-modal neural network includes a feature extraction network, an endpoint prediction network, and a region detector. Additionally, the disease is Alzheimer's disease, and wherein the output is a normalized representation.
[0028] In yet another aspect, a method is provided for generating an output indicative of a prediction of disease progression in a subject. The method includes receiving longitudinal physiological data representative of the subject, the longitudinal physiological data including multiple modalities indicative of disease progression in the subject. The method also includes applying a trained multi-modal neural network to the longitudinal physiological data. The trained multi-modal neural network includes multiple feed-forward neural networks configured to learn modality-specific deep features related to the multiple modalities. The method also includes generating an output indicative of a prediction of disease progression in the subject based at least on the modality-specific deep features related to the multiple modalities of the longitudinal physiological data.
[0029] Implementations of the present subject matter can include methods according to the description provided herein, as well as articles comprising a tangibly embodied machine-readable medium operable to cause one or more machines (e.g., computers, etc.) to perform operations that implement one or more of the described features. Similarly, computer systems are described that can include one or more processors and one or more memories coupled to the one or more processors. The memory, which can include a non-transitory computer-readable or machine-readable storage medium, can include, encode, store, etc., one or more programs that cause the one or more processors to perform one or more of the operations described herein. Computer-implemented methods according to one or more implementations of the present subject matter can be implemented by one or more data processors in a single computing system or in multiple computing systems. Such multiple computing systems can be connected, such as via one or more connections, including, for example, connections over a network (e.g., the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, etc.), and can exchange data and / or instructions or other instructions, etc.
[0030] The details of one or more variations of the subject matter described herein are described in the accompanying drawings and the following description.Other features and advantages of the subject matter described herein will be apparent from the description and drawings, and from the claims.It is easily understood that certain features of the subject matter of the present disclosure are described for illustrative purposes with respect to deep learning to model disease progression, but such features are not limiting.The claims following this disclosure are what define the scope of the subject matter that is protected.
[0031] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate certain aspects of the subject matter disclosed herein and, together with the description, help to explain some of the principles associated with the disclosed implementations. [Brief description of the drawings]
[0032] [Figure 1] FIG. 1 depicts an exemplary disease progression prediction system in accordance with an implementation of the present subject matter. [Figure 2A] 1 depicts an exemplary progression of Alzheimer's disease, in accordance with an implementation of the present subject matter. [Figure 2B] FIG. 1 depicts an exemplary progression plot illustrating the progression of Alzheimer's disease in various patients, in accordance with an implementation of the present subject matter. [Diagram 3] FIG. 1 depicts an exemplary graph illustrating various measures of Alzheimer's disease severity over time, in accordance with an implementation of the present subject matter. [Figure 4A] FIG. 1 depicts a network diagram illustrating a disease progression prediction system in accordance with an implementation of the present subject matter. [Figure 4B] FIG. 2 depicts an example dense layer of a network diagram in accordance with an implementation of the present subject matter. [Figure 4C] 2 depicts an example transition block of a network diagram in accordance with an implementation of the present subject matter. [Diagram 5] 1 depicts a diagram of an exemplary fusion technique in accordance with an implementation of the present subject matter. [Figure 6] FIG. 1 depicts a flowchart illustrating an example of a process for deep learning to model disease progression, in accordance with an implementation of the present subject matter. [Figure 7] FIG. 1 depicts a table summarizing an exemplary dataset for a disease progression prediction system, in accordance with an implementation of the present subject matter. [Figure 8] FIG. 1 depicts a table summarizing the performance comparison of machine learning models and techniques in accordance with an implementation of the present subject matter. [Figure 9] FIG. 1 depicts a table summarizing the performance comparison of machine learning models and techniques in accordance with an implementation of the present subject matter. [Figure 10] FIG. 1 depicts a graph illustrating the performance of a disease progression prediction system according to an implementation of the present subject matter. [Figure 11] FIG. 1 depicts a table summarizing the performance comparison of machine learning models and techniques in accordance with an implementation of the present subject matter. [Figure 12] 1 depicts a graph illustrating a performance comparison of optimization techniques according to an implementation of the present subject matter. [Figure 13] FIG. 1 depicts a table summarizing the performance comparison of machine learning models and techniques in accordance with an implementation of the present subject matter. [Figure 14] FIG. 1 depicts a block diagram illustrating an example of a computing system in accordance with an implementation of the present subject matter. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0033] In practice, like labels are used in the drawings to refer to the same or similar items.
[0034] Diagnostic methods for diagnosing neurological disorders, such as Alzheimer's disease, have focused on the task of classifying patients into broad categories, including cognitively normal (CN), mild cognitive impairment (MCI), or Alzheimer's disease (AD), or predicting the conversion from one category to another (e.g., from MCI to AD). In addition, diagnostic methods have been used to classify patients into finer categories, such as MCI, mild Alzheimer's disease, moderate Alzheimer's disease, and severe Alzheimer's disease. As an example, as shown in Figure 2A, MCI due to Alzheimer's disease may affect the hippocampus and entorhinal cortex of the patient's brain, mild Alzheimer's disease in a patient may spread to the temporal and parietal lobes of the patient's brain, moderate Alzheimer's disease may spread to the frontal lobe of the patient's brain, and severe Alzheimer's disease may present with widespread brain atrophy.
[0035] However, applications in clinical trials use more fine-grained measurement scales since clinical trial populations may generally be narrowly defined (e.g., MCI only). One approach is to instead predict the results of cognitive and functional tests, such as the Clinical Dementia Rating Scale (CDRSB), Alzheimer's Disease Assessment Scale-Cognitive subscale (ADAS-COG12), and Mini-Mental State Examination (MMSE), which are measured by continuous numerical values. For example, FIG. 3 shows a graph 300 illustrating various measures of Alzheimer's disease severity over time. While this approach helps provide more granular estimates of disease progression, it can be highly noisy and subjective across studies, making it difficult to establish treatment efficacy in small patient populations. Furthermore, FIG. 2B depicts a graph 200 illustrating that uniform patients with the same diagnosis may have different progression patterns based on predicted results of cognitive and functional testing. This may be beneficial to provide patients with personalized therapies and treatment plans, since patients with Alzheimer's disease may present different phenotypes and progression patterns. Additionally, predicting diagnosis alone generally does not provide a fine-grained scale measurement of disease progression.
[0036] In some cases, machine learning and deep learning based methods have been used to diagnose, monitor, and treat patients with Alzheimer's disease. However, such methods have focused on underdeveloped single-task and / or single-modality models that are not applicable to personalized medicine for Alzheimer's disease. Moreover, single-task and single-modality based models do not utilize complementary information between modalities or correlations between tasks, leading to less accurate predictions, diagnoses, and treatment plans.
[0037] Therefore, accurate prediction of clinical trajectories for patients with Alzheimer's disease (AD) has the potential to increase the efficiency of evaluating novel therapeutics for this complex and heterogeneous disease. Predicted progression rates from prognostic models can be used for covariate adjustment in the primary analysis of clinical trials to increase the accuracy of treatment effect estimates and therefore increase power. However, as mentioned, current machine learning methods incorporate either single-task or single-modality models that cannot take advantage of diverse and abundant data, such as high-dimensional neuroimaging and features of clinical data. Moreover, these methods are generally trained on a single dataset (e.g., a cohort) that cannot be easily generalized to other cohorts.
[0038] Machine learning models according to implementations of the present subject matter provide a multimodal approach to predicting the progression of neurological diseases, such as Alzheimer's disease. For example, the machine learning models described herein consider multiple modalities, such as clinical data including environmental factors, genomics, demographics, and the like, and medical (e.g., brain) imaging, while modeling the complex interactions between each modality. This allows the machine learning models described herein to integrate various modalities to characterize and contextualize a patient's condition as medical datasets grow in size and complexity, and to accurately and efficiently predict the progression of a patient's condition, such as the patient's cognitive status. With an accurate multimodal forward model of disease progression, the machine learning models according to implementations of the present subject matter take into account the variable rate of disease progression when evaluating the effectiveness of new treatments.
[0039] As described herein, deep learning approaches to modeling disease progression include a multimodal multitask deep learning model that predicts disease (e.g., Alzheimer's disease) progression, for example, by analyzing longitudinal clinical and neuroimaging data from multiple cohorts. The described machine learning model integrates high-dimensional magnetic resonance imaging features, for example, generated by a 3D convolutional neural network, with other data modalities, including clinical data, to predict disease progression (e.g., changes) in patients. As described herein, the machine learning model may employ adversarial losses to mitigate study-specific imaging biases, such as inter-study domain shifts, between sources of clinical data and / or images. Additionally and / or alternatively, the machine learning model may employ sharpness-aware minimization (SAM) optimization techniques to further improve generalization and reduce inter-study domain shifts. As described in more detail below, the present machine learning model provides significant improvements over and outperforms other models.
[0040] As an example, a machine learning model according to an implementation of the present subject matter may generate a first feature representation based on a first modality, such as clinical data, related to a patient's baseline cognitive state. The machine learning model may generate a second feature representation based on a second modality, such as an image of the patient's brain. The machine learning model may also generate a set representation by at least fusing the first feature representation and the second feature representation. The machine learning model may predict changes in the baseline cognitive state over a period of time based at least on the set representation. Thus, the machine learning model accurately and efficiently predicts changes in the patient's cognitive state. This allows for early and accurate diagnosis of diseases, such as neurological diseases including Alzheimer's disease, in patients. Additionally and / or alternatively, the machine learning models described herein make accurate predictions that may be used to treat a diagnosed disease and / or provide an appropriate treatment plan based on a predicted progression of the patient's disease.
[0041] FIG. 1 depicts a system diagram illustrating an example of a disease progression prediction system 100 according to an implementation of the present subject matter. With reference to FIG. 1, the disease progression prediction system 100 may include a machine learning controller 110, a machine learning model 120, a client device 130, and a data store 104. The machine learning controller 110, the machine learning model 120, the client device 130, and the data store 104 may be communicatively coupled via a network 140. The network 140 may be a wired network and / or a wireless network, including, for example, a wide area network (WAN), a local area network (LAN), a virtual local area network (VLAN), a public land mobile network (PLMN), the Internet, and the like. In some implementations, the machine learning controller 110, the machine learning model 120, the data store 104, and / or the client device 130 may be included within and / or operate on the same device.
[0042] It should be appreciated that client device 130 may be a processor-based device including, for example, a smartphone, a tablet computer, a wearable device, a virtual assistant, an Internet of Things (IoT) device, etc. Client device 130 may form part of, include, and / or be coupled to a magnetic resonance imaging machine.
[0043] Referring to FIG. 1 , the data store 104 may store clinical data 106 and / or one or more images 108. The clinical data 106 and / or one or more images 108 may define multimodal inputs from multiple modalities for the machine learning model 120. The clinical data 106 may be associated with a baseline cognitive state of the patient. The baseline cognitive state of the patient may be represented by at least one cognitive score. Thus, the clinical data 106 may include at least one cognitive score including at least one of a Clinical Dementia Rating Scale (CDRSB) score, an Alzheimer's Disease Assessment Scale-Cognitive Subscale (ADAS-COG12) score, a Mini-Mental State Examination (MMSE) score, a Functional Activities Questionnaire (FAQ) score, and the like.
[0044] Additionally and / or alternatively, the clinical data 106 includes demographic information associated with the patient and / or genomic information associated with the patient. In some implementations, the demographic information includes at least one of the patient's age, sex, diagnosis, education level, body mass index, etc. The genomic information may include the presence or absence of apolipoprotein E4 (APOE4), a risk factor gene for predicting Alzheimer's disease in the patient. The clinical data 106 may be collected at a baseline visit.
[0045] The clinical data 106 may be specific to a particular patient and / or may include data corresponding to multiple patients. For example, the clinical data 106 corresponding to a particular patient may be collected at a baseline visit. The clinical data 106 collected at the baseline visit may be used as an input modality to the machine learning model 120 to accurately predict the progression of the patient's cognitive status. Additionally and / or alternatively, the clinical data 106 may correspond to one or more patients. Additionally and / or alternatively, the clinical data 106 may be collected at one or more time points. Thus, the clinical data 106 corresponding to one or more patients and / or one or more time points may be used by the machine learning controller 110 to train the machine learning model 120.
[0046] The clinical data 106 may be represented as a D-length vector that includes a concatenation of one or more cognitive scores, genomic information, and / or demographic information, such as at a baseline visit. The vector may be used by the machine learning model 120 to predict a smoothed change in one or more cognitive scores over a time period following the baseline visit. The time period may be one year (in other words, 12 months), or another time period such as 6 months, 18 months, 24 months, etc.
[0047] The images 108 may include one or more images of the patient's brain. The one or more images of the brain may be acquired by a magnetic resonance imaging machine. The one or more images 108 of the patient's brain may include a three-dimensional magnetic resonance image of the patient's brain. For example, the three-dimensional magnetic resonance image may include slices (sagittal and / or coronal slices) having one or more dimensions, such as height (H), width (W), thickness (T), and channel (C). The one or more images 108 may be raw images and / or pre-processed images. For example, the one or more images 108 may include an inferred brain mask (e.g., an inferred mask) for each magnetic resonance image volume. The brain mask may be generated by at least applying one or more image segmentation techniques to the magnetic resonance image volumes. For example, the brain mask may be predicted based at least on an intensity normalization performed based on the distribution of brain intensities in the images. The brain mask may be applied to the raw images of the patient's brain. The brain mask may highlight the brain in one or more images 108 and / or particular portions of the brain in one or more images 108. This allows the machine learning model 120 to ignore portions of the one or more images 108 that do not correspond to the brain.
[0048] According to implementations of the present subject matter, the machine learning controller 110 may include a processor and a memory that stores instructions that, when executed by the processor, perform one or more operations described herein. For example, as described herein, the machine learning controller 110 may train and / or implement the machine learning model 120, such as based on one or more modalities including clinical data 106 and / or images 108.
[0049] 4A illustrates an example architecture 400 of the machine learning model 120. The machine learning controller 110 may train the machine learning model 120 using the architecture 400. As mentioned, the machine learning model 120 may include a multi-modal multi-task deep neural network that combines three-dimensional magnetic resonance images (e.g., images 108) and clinical features (e.g., clinical data 106) including one or more cognitive scores, genomic information, and demographic information into a single end-to-end model for predicting the patient's future status (e.g., disease progression and / or changes in the patient's cognitive status) from the patient's baseline visit. Each modality (e.g., clinical data 106 and / or images 108) is fed to the machine learning model 120 through a separate branch to learn a modality-specific latent space while leveraging complementary knowledge between different modalities.
[0050] With reference to the architecture 400, the machine learning model 120 may include one or more machine learning models, such as one, two, three, four, or more machine learning models. The one, two, three, four, or more machine learning models may include separate networks and / or may form part of a single machine learning model (e.g., the machine learning model 120). As shown in FIG. 4, the machine learning model 120 may include a first machine learning model or network, such as the encoder 402, a second machine learning model or network, such as the feature extraction network 420, a third machine learning model or network, such as the endpoint prediction network 460, and / or a fourth machine learning model or network, such as the region detector 480. In some implementations, the encoder 402, the feature extraction network 420, the endpoint prediction network 460, and the region detector 480 may form a single machine learning model (e.g., the machine learning model 120). In other implementations, the encoder 402, the feature extraction network 420, the endpoint prediction network 460, and the region detector 480 form separate but connected machine learning models, or in other implementations, the encoder 402 defines a first machine learning model, and the feature extraction network 420, the endpoint prediction network 460, and the region detector 480 form a sub-network of a second machine learning model connected to the first machine learning model.
[0051] According to an implementation of the present subject matter, the machine learning model 120 includes a multimodal multitask deep neural network for predicting disease progression and diagnosis. The machine learning model 120 leverages imaging and clinical data to generate patient-specific progression predictions and to generalize predictions across several domains (e.g., different source datasets, such as cross-study datasets, and / or patient disease states). For example, the machine learning model 120 combines various inputs including clinical data 106 (e.g., one or more cognitive scores, demographic information, and / or genomic information) and images 108 (e.g., three-dimensional magnetic resonance brain images), which allows learning of complementary information across modalities for a large patient population. The tabular data and imaging data are provided (e.g., by the machine learning controller 110) as separate inputs to the machine learning model 120 to generate a feature representation that is fused downstream. As described in more detail below, the machine learning controller 110 may train the machine learning model 120 end-to-end to build a feature space for predicting the progression of a cognitive condition, represented by changes in one or more cognitive scores of a patient.
[0052] 4A , the machine learning model 120 includes an encoder 402. The encoder 402 may receive clinical data 106 associated with a patient's baseline cognitive status. As described herein, the clinical data 106 may include one or more cognitive scores representing a current or baseline cognitive status of the patient, demographic information associated with the patient at the baseline visit, and / or genomic information associated with the patient at the baseline visit.
[0053] The patient's baseline cognitive state may be predicted by the machine learning model 120. For example, the machine learning controller 110 may pre-train the machine learning model 120 to predict the patient's baseline cognitive state based at least on multiple brain images (e.g., images 108) acquired at multiple time points and across multiple domains (e.g., clinical studies). This leverages the longitudinal features of the images 108 acquired at all available time points for pre-training. Integrating the longitudinal features of the images 108 allows the machine learning model 120 to learn the underlying temporal characteristics of the target disease, such as Alzheimer's disease. In this manner, the machine learning controller 110 may pre-train the machine learning model 120 to predict the current cognitive score (e.g., baseline cognitive state) for each visit of the patient, such as at various time points. The machine learning controller 110 may use the pre-trained weights as a starting point for training the machine learning model 120 to predict changes in the baseline cognitive state (e.g., baseline cognitive score) for the patient.
[0054] In some instances, multiple regions and / or clinical data (e.g., current cognitive scores, etc.) may introduce region-shift bias into the inputs (e.g., images 108 and / or clinical data 106) of the machine learning model 120. For example, with respect to images 108, an image of a brain slice from one region may look significantly different than an image of the same slice of the brain from another region. As described in more detail below, the machine learning model 120 may be trained, for example, by the machine learning controller 110, to reduce or eliminate such region shifts and improve the accuracy of predictions made by the machine learning model 120.
[0055] 4A, the encoder 402 may generate a first feature representation 410 based on the clinical data 106. The first feature representation 410 may include one or more embedded clinical features extracted from the clinical data 106, such as baseline clinical data. The one or more embedded clinical features may be represented as an encoded vector that includes a concatenation of at least one of a current cognitive score representing a baseline cognitive state of the patient, demographic information associated with the patient, and genomic information associated with the patient.
[0056] 4A, the encoder 402 may include multiple layers. The multiple layers may include a linear layer 404, a leaky ReLU layer 406, and a linear layer 408. The linear layer 404, the leaky ReLU layer 406, and the linear layer 408 may be stacked to improve extraction of embedded clinical features that define a first feature representation 410.
[0057] 4A , the machine learning model 120 includes a feature extraction network 420. The feature extraction network 420 may receive images 108 associated with a baseline cognitive state of the patient. As described herein, the images 108 may include one or more baseline (e.g., current) brain images of the patient. The images 108 may be acquired at a baseline visit. The images 108 may include one or more baseline magnetic resonance images received from one or more regions (e.g., directly from a magnetic resonance imaging machine, scanner, etc.).
[0058] The feature extraction network 420 may generate a second feature representation 421 based on the image 108 of the patient's brain as an output of the feature extraction network 420. The second feature representation 421 may include one or more embedding features, such as one or more region-invariant embedding features (described in more detail below), extracted from the image 108. The one or more embedding features of the second feature representation 421 may include one or more pixel values, spatio-temporal features, etc., extracted from the image 108. The second feature representation 421 may be output by the feature extraction network 420 and used by the machine learning controller 110 as an input to the endpoint prediction network 460 and / or the region detector 480.
[0059] The feature extraction network 420 includes one or more stacked layers, including a convolutional layer 422 (e.g., a 3D convolutional layer), a batch normalization layer 424 (e.g., a 3D batch normalization layer), a hidden layer 426 (e.g., a leaky ReLU layer), and a max pooling layer 428 (e.g., a 3D max pooling layer). The feature extraction network 420 may include one or more dense layers or blocks 440 (see FIG. 4B) and / or one or more transition blocks 450 (see FIG. 4C) following the one or more stacked layers.
[0060] The one or more dense layers 440 may include a first dense layer 430, a second dense layer 434, a third dense layer 462, a fourth dense layer 482, etc. With reference to Figure 4B, each of the one or more dense layers 440 includes six dense layers (e.g., two sets of three stacked dense layer blocks) or twelve dense layers (e.g., four sets of three stacked dense layer blocks). For example, each set of three stacked dense layer blocks includes a batch normalization layer, a hidden layer, and a convolutional layer. As shown in FIG. 4B, the dense layer 440 includes six dense layers, such as a batch normalization layer 441 (e.g., a 3D batch normalization layer), a hidden layer 442 (e.g., a leaky ReLU layer), a convolution layer 443 (e.g., a 3D convolution layer), a batch normalization layer 444 (e.g., a 3D batch normalization layer), a hidden layer 445 (e.g., a leaky ReLU layer), and a convolution layer 447 (e.g., a 3D convolution layer). The first dense layer 430 includes six dense layers, such as two sets of three stacked dense layer blocks. The second dense layer 434 includes twelve dense layers, such as four sets of three stacked dense layer blocks. The third and fourth dense layers 462, 482 include sixteen stacked dense layer blocks. Each of the one or more dense layers generates a feature map by extracting one or more features from the image 108 after the image 108 has passed through the previous block.
[0061] 4A, the one or more transition blocks 450 may include a first transition block 432 and a second transition block 436. Referring to FIG. 4B, each of the first transition blocks 450 includes four dense layers, such as a batch normalization layer 451 (e.g., a 3D batch normalization layer), a hidden layer 452 (e.g., a leaky ReLU layer), a convolution layer 453 (e.g., a 3D convolution layer), and an average pooling layer 454 (e.g., a 3D average pooling layer). Each of the one or more layers of the one or more transition blocks 450 generates a feature map by extracting one or more features from the image 108 after the image 108 passes through the previous block.
[0062] 4A , the feature extraction network 420 includes one or more stacked layers, followed by a first dense layer 430, a first transition block 432, a second dense layer 434, and a second transition block 436. The feature map size varies between the first and second dense layers 430, 434 and through the first and second transition blocks 432, 436. As stated, the feature extraction network 420 generates a second feature representation based on the image 108, the second feature representation including one or more extracted features.
[0063] The machine learning controller 110 provides the second feature representation 421 as input to an endpoint prediction network 460 and a region detector 480. The endpoint prediction network 460 includes a dense layer 462, a hidden layer 463 (e.g., a ReLU layer), an average pooling layer 464 (e.g., a 3D average pooling layer), a dropout layer 466, a feature fusion layer 468, a linear layer 470, a hidden layer 472 (e.g., a ReLU layer), and / or a linear layer 474.
[0064] The endpoint prediction network 460 may generate one or more regression endpoints 476. For example, the endpoint prediction network 460 may predict a change in baseline cognitive status over a period of time (e.g., 12 months from baseline, 24 months from baseline, 48 months from baseline, etc.). For example, the change in baseline cognitive status may be expressed as a change in one or more cognitive scores. The change in baseline cognitive status over time may indicate the progression of a corresponding neurological disorder, such as Alzheimer's disease, in the patient.
[0065] As mentioned, the endpoint prediction network 460 includes a feature fusion layer 468. Although the feature fusion layer 468 is shown as being included in the endpoint prediction network 460, the feature fusion layer 468 may form at least a portion of one or more of the encoder 402, the feature extraction network 420, the region detector 480, the endpoint prediction network 460, or another location as part of the machine learning model 120.
[0066] The feature fusion layer 468 may generate the set representation 467 by fusing at least the first feature representation 410 from the encoder 402 and / or the second feature representation 421 from the feature extraction network 420. The feature fusion layer 468 may fuse the first feature representation 410 and the second feature representation 421 using one or more fusion techniques including concatenation, addition, simple attention, scaled dot product attention, applying a tensor fusion network, low-rank fusion, unidirectional context attention, etc.
[0067] As an example, the machine learning model 120 is trained to extract a first feature representation 410 and a second feature representation 421 from different modalities (e.g., clinical data 106 and images 108, respectively). Multimodal fusion can be achieved by a feature fusion layer 468, for example, by one or more joint representation techniques including concatenating and / or adding the first feature representation 410 and the second feature representation 421 to generate a set representation 467. The set representation 467 is then fed to a fully connected layer of the endpoint prediction network 460 to predict changes in the patient's cognitive status. In this example, gradients are backpropagated upstream through the fully connected layer to each branch of the machine learning model 120 to train the machine learning model 120 end-to-end. This allows the joint representation learning to leverage complementary information across modalities (e.g., clinical data 106 and images 108) and model cross-modality interactions. In some implementations, prior to generating the set representation 467 in the feature fusion layer 468, the first feature representation 410 and / or the second feature representation 421 are normalized so that the first feature representation 410 and / or the second feature representation 421 have the same scale.
[0068] In some implementations, the set representation 467 may be generated using simple attention, in which the input first feature representation 410 and second feature representation 421 are concatenated and fed through a single block including two fully connected layers, i.e., ReLU layers and / or sigmoids, using a nonlinear combination of features within and between modalities (e.g., clinical data 106 and / or images 108) to reweight the concatenated feature vector of the set representation 467.
[0069] In some implementations, the set representation 467 may be generated using scaled dot product attention, in which the input first feature representation 410 and second feature representation 421 are first concatenated into a single vector, which is then used as the query, key, and value in the scaled dot product calculation. In one implementation, the machine learning controller 110 may instantiate one or more weight matrices of the input first feature representation 410 and second feature representation 421 to linearly vary each query, key, and value to generate the set representation 467.
[0070] In some implementations, the set representation 467 may be generated using a tensor fusion network fusion technique. The feature fusion layer 468 may determine a tensor output between the first feature representation 410 and the second feature representation 421.
[0071] In some implementations, the set representation 467 may be generated using low-rank fusion, which decomposes the input tensors including the first feature representation 410 and / or the second feature representation and the learned weight tensors into low-rank factors. The feature fusion layer 468 then reorders the decomposed input tensors to generate the set representation 467.
[0072] In some implementations, the set representation 467 may be generated using unidirectional context attention. FIG. 5 illustrates an example network diagram 500 showing a feature fusion layer 468. In this example, the feature fusion layer 468 assigns a primary modality from multiple modalities (e.g., clinical data 106 and / or images 108) and one or more remaining auxiliary data modalities from multiple modalities (e.g., clinical data 106 and / or images 108) to create a directed attention mechanism. The feature representations from the one or more auxiliary modalities may be input to the feature fusion layer 468 at 512, concatenated, and passed through the feature fusion layer 468, which may include two fully connected layers, such as a ReLU layer 510 and a sigmoid 508. The output from the feature fusion layer 468 (e.g., sigmoid) is used to reweight the feature representations 502 generated based on the primary modality in a channel recalibration 506 to generate a reweighted feature representation 504.
[0073] As an example, the clinical data 106 may be the auxiliary modality and the image 108 may be the primary modality (or vice versa). In this example, the first feature representation 410 (e.g., input at 512) may pass through the fully connected layers 510, 508 of the feature fusion layer 468. The output may be used by the machine learning controller 110 (e.g., machine learning model 120) in channel recalibration 506 to reweight the second feature representation 421 (e.g., feature representation 502 in this example) generated based on the image 108. The reweighted feature representation (e.g., the reweighted first feature representation 410 and / or the reweighted second feature representation 421 shown as feature representation 504 in FIG. 5) may define a set representation 467 that is generated and used as an input to the endpoint prediction network 460. In this manner, the feature representation generated based on the auxiliary modality (e.g., clinical data 106 in this example) provides context to the feature representation generated based on the primary modality (e.g., images 108 in this example) without dominating the optimization. Hence, the first feature representation 410 and the second feature representation 421 may be efficiently fused to generate a set representation 467 for predicting one or more endpoints (e.g., change in cognitive state, regression endpoint 476, etc.) with improved accuracy.
[0074] Referring again to FIG. 4A , the region detector 480 includes a dense layer 482, a batch normalization layer 484 (e.g., a 3D batch normalization layer), a hidden layer 486 (e.g., a leaky ReLU layer), an average pooling layer 488 (e.g., a 3D average pooling layer), and a linear layer 490.
[0075] The region detector 480 enables the machine learning model 120 to learn a region-invariant imaging feature representation, such as a region-invariant embedding feature of the second feature representation 421. In general, convolutional neural networks trained on multi-study images (e.g., magnetic resonance imaging data) often have problems due to region shift and heterogeneity, since scanners and protocols may vary from one study to another, and patient populations and their medical conditions may vary between studies. Notably, there are two common forms of region shift (e.g., bias) in medical image analysis between different studies: population shift and acquisition shift. Population shift occurs when a cohort of subjects exhibits different demographic or clinical characteristics, while acquisition shift is observed due to differences in imaging protocols, modalities, or scanners. This may cause the image 108 to introduce region shifting due to population shift and / or acquisition shift. To mitigate possible inter-study bias, the machine learning controller 110 trains the machine learning model 120 to minimize the mutual information shared between the extracted features and the inter-study distribution shift, and to minimize bias 492, such as inter-study bias in the generated feature representations (e.g., the first feature representation 410 and / or the second feature representation 421) and / or the regression endpoints 476.
[0076] In particular, the machine learning controller 110 adversarially trains the region detector 480 against the feature extraction network 420 to predict the bias distribution (e.g., bias 492). In other words, the machine learning controller 110 may train the machine learning model 120 using an adversarial loss to extract neuroimaging features (e.g., second feature representation 421) based on the images 108 and / or embedded clinical features (e.g., first feature representation 410) while remaining invariant to the lab domain. As a result of the adversarial training, the ability of the region detector 480 to predict regions associated with the images 108 is minimized.
[0077] For example, the feature extraction network 420 and / or the region detector 480 are trained in an adversarial manner such that the feature extraction network 420 learns domain-invariant features for generating the second feature representation 421. The feature extraction network 420 is adversarially trained to increase the ability of the feature extraction network 420 to generalize across multiple domains when generating the second feature representation 421. Hence, the feature extraction network 420 may be trained (e.g., by the machine learning controller 110) until the region detector 480 is unable to correctly identify regions of the second feature representation 421 generated by the feature extraction network 420.
[0078] As mentioned, the machine learning controller 110 adversarially trains the region detector 480 to reduce inter-study and / or inter-modality region shifts associated with the patient's brain images 108. For example, the machine learning controller 110 trains the region detector 480 to distinguish regions associated with the images 108 and / or the clinical data 106 based at least on the first feature representation 410 and / or the second feature representation 421. At the same time, the machine learning controller 110 trains the feature extraction network 420 to generalize regions associated with the images 108 in generating the second feature representation 421 and / or the encoder 402 to generalize regions associated with the clinical data 106 in generating the first feature representation 410. The region detector 480 is trained to minimize a region classification loss and the feature extraction network 420 is trained to maximize a region classification loss.
[0079] The machine learning controller 110 may adversarially train the region detector 480, for example, by applying an inverse gradient 481 to the second feature representation 421 (and / or the first feature representation 410) to generate input to the region detector 480. As stated, adversarially training the region detector 480 results in the second feature representation 421 and / or the first feature representation 410 being region invariant. In other words, the region associated with the second feature representation 421 (or the first feature representation 410) is generalized such that the influence of the region on the extracted features that define the second feature representation 421 (or the first feature representation 410) is minimized or eliminated. In some implementations, the region detector 480 may additionally and / or alternatively be adversarially trained against the encoder 402 to minimize bias in the clinical data 106 and / or the first feature representation 410. Thereby, the region detector 480 may be adversarially trained to produce the region-invariant first feature representation 410 and / or the region-invariant second feature representation 421. In inference, the region detector 480 may additionally and / or alternatively detect drift in the bias 492.
[0080] As mentioned, the machine learning controller 110 may train the machine learning model 120 to predict changes in the baseline cognitive state over time based at least on multiple modalities including clinical data 106 associated with the baseline cognitive state of the patient p and brain images 108 of the patient p. p (e.g., Clinical Data 106) and 3D Magnetic Resonance Imaging (MRI) p An input “patient p” with (e.g., image 108) is provided as input to the machine learning model 120, where each p ) (e.g., the first feature representation 410) and the patient's image embedding f(MRI p ) (e.g., the second feature representation 421), p enters encoder 102, and MRI pgoes into the feature extraction network 420. Then, f(MRI p ) is fed to the area detector 480, which extracts the features, i.e., a(Clin p ) and f(MRI p ) are fed forward through the endpoint prediction network 460. The parameters of each network are a , (e.g., corresponding to feature extraction network 420) θ f , θ (e.g., corresponding to the endpoint prediction network 460) g , (e.g., corresponding to area detector 480) θ h and the subscripts indicate their intrinsic networks as shown in FIG. 4A.
[0081] Although the machine learning model 120 is trained on a mixture of data sources, the machine learning model 120 operates robustly on test data from unseen regions. For this purpose, a mutual information-based loss is added to the objective function for training the machine learning model 120. The training procedure therefore amounts to optimizing the following equation: TIFF2024544147000002.tif14170 where Loss(.,.) and I(.,.) denote the loss function and mutual information (e.g., bias492), respectively, and λ is a hyperparameter for balancing terms. Substituting the mutual information, the above equation becomes TIFF2024544147000003.tif39170
[0082] Here, L エンドポイント (.), L バイアス(.), and H(h○f(.)) respectively represent the end-point prediction loss, the bias prediction loss, and the entropy of the bias acting as a regularizer. A set of networks including the encoder 402, the feature extraction network 420, the end-point prediction network 460, and the region detector 480 are trained end-to-end using adversarial losses and gradient inversion techniques (e.g., based on inverse gradient 481), and the weights θ f , θ h At the beginning of learning, g∧f is trained to rapidly predict the endpoint by using the bias 492. Then, h (e.g., region detector 480) learns to predict the bias 492, and f (e.g., feature extraction network 420) learns to extract feature embeddings (e.g., second feature representation 421) that are invariant to the imaging domain.
[0083] Additionally, during training of the machine learning model 120, the machine learning controller 110 may apply Sharpness Aware Minimization (SAM) to avoid overfitting and improve generalization. SAM performs two forward-backward passes to estimate the smoothness of the loss landscape and improve the final prediction accuracy. The smoothing is performed by at least fitting a linear regression for each patient to each cognitive score for all visits in the first 24 months (or another time period). The slope of the fitted linear regression is then used to determine the smoothed change in cognitive score (e.g., the patient's cognitive status) at the end of the desired time period (e.g., 12 months). Training the machine learning model 120 using the smoothed change in cognitive score helps to mitigate missing values and measurement noise.
[0084] 6 depicts a flowchart illustrating an example of a process 600 for deep learning to model disease progression, such as the progression of one or more neurological disorders, including Alzheimer's disease, according to an implementation of the present subject matter. With reference to FIG. 6, process 600 may be performed by machine learning model 120 and / or machine learning controller 110 to predict changes in a patient's cognitive status over a period of time, such as 12 months. This allows for efficient and accurate prediction of the progression of a neurological disorder, such as Alzheimer's disease, in a patient over time.
[0085] As described herein, the predictions may be domain invariant. In other words, the machine learning model 120 (e.g., the machine learning controller 110) may reduce the effect of domain shift biases derived from inputs (e.g., images 108 and / or clinical data 106) received from one or more modalities (e.g., one or more scanners, machines, studies, clinicians, etc.). Hence, the machine learning controller 110 may train the machine learning model 120 using multi-modal modalities, including the clinical data 106 and / or images 108, with improved accuracy, memory efficiency, and speed. According to an implementation of the present subject matter, the process 600 refers to the exemplary architecture 400 shown in FIG. 4A.
[0086] The machine learning model 120 may receive one or more input modalities, such as clinical data 106 associated with the patient's baseline cognitive state and images 108 of the patient's brain. The images 108 may include three-dimensional magnetic resonance images including an inferred mask. The baseline cognitive state may be represented by at least one cognitive score including at least one of a Clinical Dementia Rating Scale (CDRSB) score, an Alzheimer's Disease Assessment Scale-Cognitive Subscale (ADAS-COG12) score, and a Mini-Mental State Examination (MMSE) score.
[0087] At 602, the machine learning model 120 (e.g., via the machine learning controller 110) may generate a first feature representation, such as a first feature representation 410, based on clinical data, such as clinical data 106 associated with a baseline cognitive state of the patient. The first feature representation 410 may be an encoded vector that includes a concatenation of at least one of a current cognitive score representing the baseline cognitive state of the patient, demographic information associated with the patient, and genomic information associated with the patient.
[0088] In some implementations, the machine learning model 120 may be pre-trained (e.g., by the machine learning controller 110) to predict a patient's baseline cognitive state based at least on multiple brain images acquired at multiple time points and across multiple regions. The machine learning model 120 may be trained to reduce inter-study region shifts associated with multiple regions. For example, the machine learning model 120 may be trained (e.g., by the machine learning controller 110) by at least adversarially training a region detector (e.g., region detector 480) of the machine learning model 120 to reduce inter-study region shifts (e.g., bias 492) associated with the image 108 associated with the patient's brain. The adversarial training includes applying an inverse gradient to the second feature representation 421 to generate a region detector input to the region detector 480. The adversarial training of the region detector 480 is based at least on the region detector input. Additionally, as mentioned, the region detector 480 may exhibit drift in inter-study region shifts in inference. Training adversarially may include training the feature extraction network 420 in an adversarial manner such that the feature extraction network 420 learns domain-invariant features for generating the second feature representation 421. In other words, the feature extraction network 420 may be adversarially trained to increase the ability of the feature extraction network 420 to generalize across multiple domains when generating the second feature representation 421.
[0089] At 604, the machine learning model 120 (e.g., via the machine learning controller 110) may generate a second feature representation, such as a second feature representation 421, based at least on an image, such as image 108, of the patient's brain. The second feature representation 421 may include at least one domain-invariant embedding feature, as described herein, which helps provide predictions with increased accuracy.
[0090] At 606, the machine learning model 120 (e.g., via the machine learning controller 110) may generate a set representation by at least fusing the first feature representation 410 and the second feature representation 421. The machine learning model 120 may perform the fusing using one or more fusion techniques. The one or more fusion techniques may include at least one of concatenation, addition, simple attention, scaled dot product attention, applying a tensor fusion network, low-rank fusion, and unidirectional context attention.
[0091] At 608, the machine learning model 120 (e.g., via the machine learning controller 110) may predict a change in the baseline cognitive state over a period of time based at least on the set representation. The change in the baseline cognitive state over time may indicate the progression of a disease, such as Alzheimer's disease, in the patient. The time period may be 12 months, 24 months, 48 months, etc. Thus, the machine learning model 120 may accurately and efficiently predict the progression of the disease in the patient over time.
[0092] experiment The performance of the machine learning model 120 according to the implementation of the present subject matter was tested based on multiple modalities for multiple patients. FIG. 7 depicts a table 700 summarizing the patient composition for the testing of the machine learning model 120. In the table 700, CN refers to cognitively normal patients, MCI refers to patients with mild cognitive impairment, and AD refers to patients with Alzheimer's disease. As described herein, the input clinical data (e.g., clinical data 106) included patient demographic information (e.g., age, sex, diagnosis, education level, and body mass index), genomic information, one or more cognitive scores, such as CDRSB, MMSE, ADAS-Cog12, and FAQ. The input images (e.g., images 108) included raw three-dimensional magnetic resonance images and volumetric magnetic resonance imaging features. The raw three-dimensional magnetic resonance images were normalized by inferring a brain mask. During training of the machine learning model, the 3D magnetic resonance images and segmentations were isotropically resampled to 1 mm voxel size and standardized to the canonical (RAS+) orientation, with intensities rescaled to 0,1, and with Z-values normalized using only voxels within the brain mask. Finally, the magnetic resonance image volumes were cropped or padded.
[0093] Weighted R 2 is the R of each data set, as shown in Equation 3 below. 2 The machine learning model 120 was defined separately for each endpoint by a weighted average of TIFF2024544147000004.tif35170
[0094] Here, SS res and S.S. tot and denote the residual sum of squares and the total sum of squares, respectively. This weighting is R 2 is proportional to the dataset size, the weighted average R 2 We guarantee that we will contribute to
[0095] Effective sample size increase (ESSI) was calculated for two setups: (1) comparing ESSI in MMMT (multimodality multitask modeling) adjusted analysis with respect to unadjusted analysis; (2) comparing ESSI in MMMT adjusted analysis with baseline linear regression model. FIG. 8 depicts a table 800 summarizing the performance comparison of machine learning models and methods based on the validation set according to an implementation of the present subject matter. For example, as shown in table 800, MMMT (e.g., machine learning model 120) outperformed other models, such as single modality single task (SMST) model, single modality multitask (SMMT) model, and multimodality single task (MMST) model. In table 800, Clin refers to clinical information modality (e.g., clinical data 106), and MRI refers to imaging modality (e.g., image 108). FIG. 9 depicts a table 900 summarizing another performance comparison of machine learning models and methods according to an implementation of the present subject matter. In particular, table 900 shows the ESSI comparison for the machine learning model shown in table 800. As shown in table 800 and table 900, the machine learning model 120 performs best when using all modalities to predict all three endpoints (e.g., three cognitive scores). When used for covariate adjustment in clinical trials, this shows that the machine learning model 120 can achieve up to a 35% increase in sample size compared to unadjusted trials, or a 12% increase in sample size compared to adjusted trials using the regression Clin model shown in table 900.
[0096] 10 depicts a graph 1000 illustrating the performance of a disease progression prediction system according to an implementation of the present subject matter. The graph 1000 shows improved performance of the machine learning model 120 in predicting disease progression based on raw image inputs compared to extracted volumetric features, due at least in part to encoded structural features included in the raw image inputs.
[0097] FIG. 11 depicts a table 1100 summarizing the performance comparison of machine learning models and methods according to the implementation of the present subject matter. As shown in table 1100, the robustness of the machine learning model 120 is assessed by using two different external test sets, including an in-study test set and an out-study test set. The trained multi-task machine learning model 120 was tested by using either clinical information or magnetic resonance images, or both. As shown in table 1100, the machine learning model 120 applied to both input modalities significantly outperformed other models on the in-study test set. On the out-study test set, the machine learning model 120 applied to both modalities significantly outperformed other models for predicting CDRSB, while reaching comparable performance for MMSE and ADAS-COG12 prediction when applied only to clinical information. Moreover, stratifying patients' cognitive trajectories by their predicted CDSB changes reveals significant differences between groups, as the machine learning model 120 can separate stable and declining patients.
[0098] 12 depicts a graph 1200 illustrating a performance comparison of optimization techniques according to an implementation of the present subject matter. As shown in graph 1200, application of the SAM optimization method leads to more stable convergence and better minima compared to other optimization methods.
[0099] 13 depicts a table 1300 summarizing the performance comparison of machine learning models and techniques according to implementations of the present subject matter. Table 1300 illustrates the effectiveness of adversarial training of the region detector 480 of the machine learning model 120. For example, as shown in table 1300, the incorporation of an adversarial loss in training the machine learning model 120 significantly reduces the R 2 This results in improvements of up to 0.22% and 0.09% for
[0100] Computing Systems 14 depicts a block diagram illustrating a computing system 1400 according to an implementation of the present subject matter. Referring to FIGS. 1-13, the computing system 1400 may be used to implement the machine learning controller 110, the machine learning model 120, the disease progression prediction system 100, and / or any components therein.
[0101] As shown in FIG. 14, the computing system 1400 may include a processor 1410, a memory 1420, a storage device 1430, and an input / output device 1440. The processor 1410, the memory 1420, the storage device 1430, and the input / output device 1440 may be interconnected via a system bus 1450. The computing system 1400 may additionally or alternatively include a graphics processing unit (GPU), such as for image processing, and / or associated memory for the GPU. The GPU and / or associated memory for the GPU may be interconnected with the processor 1410, the memory 1420, the storage device 1430, and the input / output device 1440 via the system bus 1450. The memory associated with the GPU may store one or more images described herein, and the GPU may process one or more of the images described herein. The GPU may be coupled to and / or form part of the processor 1410. The processor 1410 can process instructions for execution within the computing system 1400. Such executed instructions can implement one or more components, such as, for example, the machine learning controller 110, the machine learning model 120, the disease progression prediction system 100, etc. In some implementations of the present subject matter, the processor 1410 can be a single-threaded processor. Alternatively, the processor 1410 can be a multi-threaded processor. The processor 1410 can process instructions stored in the memory 1420 and / or in the storage device 1430 to display graphical information for a user interface provided via the input / output device 1440.
[0102] The memory 1420 is a computer-readable medium, such as volatile or nonvolatile, that stores information within the computing system 1400. The memory 1420 can store, for example, data structures representing a configuration object database. The storage device 1430 can provide persistent storage for the computing system 1400. The storage device 1430 can be a floppy disk device, a hard disk device, an optical disk device, or a tape device, or other suitable persistent storage means. The input / output device 1440 provides input / output operations for the computing system 1400. In some implementations of the present subject matter, the input / output device 1440 includes a keyboard and / or a pointing device. In various implementations, the input / output device 1440 includes a display unit for displaying a graphical user interface.
[0103] According to some implementations of the present subject matter, the I / O device(s) 1440 may provide I / O operations for network devices. For example, the I / O device(s) 1440 may include Ethernet ports or other networking ports for communicating with one or more wired and / or wireless networks (e.g., a local area network (LAN), a wide area network (WAN), the Internet).
[0104] In some implementations of the present subject matter, the computing system 1600 may be used to execute various interactive computer software applications (e.g., Microsoft Excel, and / or other types of software) that may be used for organizing, analyzing, and / or storing data in various (e.g., tabular) formats. Alternatively, the computing system 1600 may be used to execute any type of software application. These applications may be used to perform various functionalities, for example, planning functionality (e.g., creating, managing, editing spreadsheet documents, word processing documents, and / or any other objects, etc.), computing functionality, communication functionality, etc. Applications may include various add-in functionalities or may be stand-alone computing products and / or functionality. When active within an application, the functionality may be used to generate a user interface that is provided via the input / output devices 1640. The user interface may be generated by the computing system 1600 and presented to a user (e.g., on a computer screen monitor, etc.).
[0105] One or more aspects or features of the subject matter described herein may be implemented in digital electronic circuitry, integrated circuits, specially designed ASICs, field programmable gate array (FPGA) computer hardware, firmware, software, and / or combinations thereof. These various aspects or features may include implementation in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which may be special purpose or general purpose, coupled to receive data and instructions from a storage system, at least one input device, and at least one output device, and to transmit data and instructions to the storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0106] These computer programs, sometimes referred to as programs, software, software applications, applications, components, or codes, include machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language and / or in assembly / machine language. As used herein, the term "machine-readable medium" refers to any computer program product, apparatus, and / or device used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal, such as, for example, magnetic disks, optical disks, memory, and programmable logic devices (PLDs). The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor. A machine-readable medium may store such machine instructions non-temporarily, such as, for example, in a non-transitory solid-state memory or a magnetic hard drive or any equivalent storage medium. A machine-readable medium may alternatively or additionally store such machine instructions in a transitory manner, such as, for example, in a processor cache or other random access memory associated with one or more physical processor cores.
[0107] To provide for user interaction, one or more aspects or features of the subject matter described herein may be implemented on a computer having a display device, such as, for example, a cathode ray tube (CRT) or liquid crystal display (LCD) or light emitting diode (LED) monitor for displaying information to a user, and a keyboard and a pointing device, such as, for example, a mouse or trackball, by which a user may provide input to the computer. Other types of devices may also be used to provide for user interaction. For example, feedback provided to the user may be any form of sensory feedback, such as, for example, visual feedback, auditory feedback, or haptic feedback, and input from the user may be received in any form, including acoustic, voice, or tactile input. Other possible input devices include touch screens or other touch-sensitive devices, such as single or multi-point resistive or capacitive trackpads, voice recognition hardware and software, optical scanners, optical pointers, digital image capture devices, and associated interpretation software, and the like.
[0108] The subject matter described herein may be embodied in systems, devices, methods, and / or articles, depending on the desired configuration. The implementations described in the above description do not represent all implementations according to the subject matter described herein. Instead, the implementations described in the above description are merely some examples according to aspects related to the described subject matter. Although several variations have been described in detail above, other modifications or additions are possible. In particular, further features and / or variations may be provided in addition to those described herein. For example, the implementations described above may be directed to various combinations and subcombinations of the disclosed features, and / or combinations and subcombinations of several additional features disclosed above. In addition, the logic flows depicted in the accompanying figures and / or described herein do not necessarily require the particular order or sequential order shown to achieve the desired results. For example, the logic flows may include different and / or additional operations than those shown without departing from the scope of the present disclosure. One or more operations of the logic flows may be repeated and / or omitted without departing from the scope of the present disclosure. Other implementations may be within the scope of the following claims.
Claims
1. a processor; a memory storing instructions that, when executed by the processor, generating, by a machine learning model, a first feature representation based on clinical data associated with the patient's baseline cognitive state; generating, by the machine learning model, a second feature representation based on the image of the patient's brain; and generating a set representation by at least fusing the first feature representation and the second feature representation with the machine learning model; predicting, with the machine learning model, a change in the baseline cognitive state over a period of time based at least on the set representation; and a memory for storing instructions that cause operations including A system comprising:
2. 2. The system of claim 1, wherein the fusing is performed using one or more fusion techniques including at least one of concatenation, addition, simple attention, scaled dot product attention, applying a tensor fusion network, low-rank fusion, and unidirectional context attention.
3. 2. The system of claim 1, wherein the first feature representation is an encoded vector comprising a concatenation of at least one of a current cognitive score representing the baseline cognitive state of the patient, demographic information associated with the patient, and genomic information associated with the patient.
4. The system of claim 1 , wherein the second feature representation includes at least one domain-invariant embedding feature.
5. 2. The system of claim 1, wherein the machine learning models include a first machine learning model trained to generate the first feature representation, a second machine learning model trained to generate the second feature representation, a third machine learning model trained to generate the set representation, and a fourth machine learning model trained to predict the change in the baseline cognitive state over the time period.
6. 2. The system of claim 1, wherein the machine learning model is trained based on at least a plurality of modalities including the clinical data associated with the baseline cognitive state of the patient and the images of the brain of the patient.
7. 10. The system of claim 1, wherein the machine learning model is pre-trained to predict the baseline cognitive state of the patient based at least on multiple brain images acquired at multiple time points and across multiple regions.
8. 10. The system of claim 1, wherein the machine learning model is trained by at least adversarially training a region detector of the machine learning model to reduce inter-study region shift associated with the images of the brain of the patient.
9. 10. The system of claim 8, wherein the adversarial training comprises adversarially training a feature extraction network of the machine learning model to learn domain-invariant features for generating the second feature representation based at least on the image of the brain of the patient.
10. 9. The system of claim 8, wherein the adversarial training comprises applying an inverse gradient to the second feature representation to generate a region detector input, and wherein the adversarial training of the region detector is based at least on the region detector input.
11. The system of claim 8 , wherein the region detector indicates a drift in the inter-study region shift in estimation.
12. 10. The system of claim 1, wherein the change in the baseline cognitive status over time indicates progression of Alzheimer's disease in the patient.
13. The system of claim 1 , wherein the time period is 12 months.
14. The system of claim 1 , wherein the image is a three-dimensional magnetic resonance imaging image including an estimated mask.
15. 2. The system of claim 1, wherein the baseline cognitive status is represented by at least one cognitive score comprising at least one of a Clinical Dementia Rating Scale (CDRSB) score, an Alzheimer's Disease Assessment Scale-Cognitive Subscale (ADAS-COG12) score, and a Mini-Mental State Examination (MMSE) score.
16. generating, by a machine learning model, a first feature representation based on clinical data associated with the patient's baseline cognitive state; generating, by the machine learning model, a second feature representation based on the image of the patient's brain; and generating a set representation by at least fusing the first feature representation and the second feature representation with the machine learning model; predicting, with the machine learning model, a change in the baseline cognitive state over a period of time based at least on the set representation; and 11. A computer-implemented method comprising:
17. 17. The method of claim 16, wherein the fusing is performed using one or more fusion techniques including at least one of concatenation, addition, simple attention, scaled dot product attention, applying a tensor fusion network, low-rank fusion, and unidirectional context attention.
18. 17. The method of claim 16, wherein the first feature representation is an encoded vector comprising a concatenation of at least one of a current cognitive score representing the baseline cognitive state of the patient, demographic information associated with the patient, and genomic information associated with the patient.
19. The method of claim 16 , wherein the second feature representation includes at least one domain-invariant embedding feature.
20. 17. The method of claim 16, wherein the machine learning models include a first machine learning model trained to generate the first feature representation, a second machine learning model trained to generate the second feature representation, a third machine learning model trained to generate the set representation, and a fourth machine learning model trained to predict the change in the baseline cognitive state over the time period.
21. 17. The method of claim 16, wherein the machine learning model is trained based on at least a plurality of modalities including the clinical data associated with the baseline cognitive state of the patient and the images of the brain of the patient.
22. 17. The method of claim 16, wherein the machine learning model is pre-trained to predict the baseline cognitive state of the patient based at least on multiple brain images acquired at multiple time points and across multiple regions.
23. 17. The method of claim 16, wherein the machine learning model is trained by at least adversarially training a region detector of the machine learning model to reduce inter-study region shift associated with the images of the brain of the patient.
24. 24. The method of claim 23, wherein the adversarial training comprises adversarially training a feature extraction network of the machine learning model to learn domain-invariant features for generating the second feature representation based at least on the image of the brain of the patient.
25. 24. The method of claim 23, wherein the adversarial training comprises applying an inverse gradient to the second feature representation to generate a region detector input, and wherein the adversarial training of the region detector is based at least on the region detector input.
26. 24. The method of claim 23, wherein the region detector exhibits drift in the inter-study region shift in estimation.
27. 17. The method of claim 16, wherein the change in the baseline cognitive status over time indicates progression of Alzheimer's disease in the patient.
28. 17. The method of claim 16, wherein the time period is 12 months.
29. The method of claim 16 , wherein the image is a three-dimensional magnetic resonance imaging image including an estimated mask.
30. 17. The method of claim 16, wherein the baseline cognitive status is represented by at least one cognitive score comprising at least one of a Clinical Dementia Rating Scale (CDRSB) score, an Alzheimer's Disease Assessment Scale-Cognitive Subscale (ADAS-COG12) score, and a Mini-Mental State Examination (MMSE) score.
31. To one or more processors: generating, by a machine learning model, a first feature representation based on clinical data associated with the patient's baseline cognitive state; generating, by the machine learning model, a second feature representation based on the image of the patient's brain; and generating a set representation by at least fusing the first feature representation and the second feature representation with the machine learning model; predicting, with the machine learning model, a change in the baseline cognitive state over a period of time based at least on the set representation; and A computer program for causing a computer to perform operations including:
32. 32. The computer program product of claim 31 , wherein the fusing is performed using one or more fusion techniques including at least one of concatenation, addition, simple attention, scaled dot product attention, applying a tensor fusion network, low-rank fusion, and unidirectional context attention.
33. 32. The computer program of claim 31 , wherein the first feature representation is an encoded vector comprising a concatenation of at least one of a current cognitive score representing the baseline cognitive state of the patient, demographic information associated with the patient, and genomic information associated with the patient.
34. 32. The computer program product of claim 31 , wherein the second feature representation comprises at least one domain-invariant embedding feature.
35. 32. The computer program of claim 31 , wherein the machine learning models include a first machine learning model trained to generate the first feature representation, a second machine learning model trained to generate the second feature representation, a third machine learning model trained to generate the set representation, and a fourth machine learning model trained to predict the change in the baseline cognitive state over the time period.
36. 32. The computer program of claim 31 , wherein the machine learning model is trained based on at least a plurality of modalities including the clinical data associated with the baseline cognitive state of the patient and the images of the brain of the patient.
37. 32. The computer program of claim 31 , wherein the machine learning model is pre-trained to predict the baseline cognitive state of the patient based at least on multiple brain images acquired at multiple time points and across multiple regions.
38. 32. The computer program of claim 31 , wherein the machine learning model is trained by at least adversarially training a region detector of the machine learning model to reduce inter-study region shift associated with the images of the brain of the patient.
39. 39. The computer program product of claim 38, wherein the adversarial training comprises adversarially training a feature extraction network of the machine learning model to learn domain-invariant features for generating the second feature representation based at least on the image of the brain of the patient.
40. 39. The computer program product of claim 38, wherein the adversarial training comprises applying an inverse gradient to the second feature representation to generate a region detector input, and wherein the adversarial training of the region detector is based at least on the region detector input.
41. 39. The computer program of claim 38, wherein the region detector exhibits a drift in the inter-study region shift in estimation.
42. 32. The computer program of claim 31, wherein the change in the baseline cognitive status over time indicates progression of Alzheimer's disease in the patient.
43. 32. The computer program of claim 31, wherein the time period is 12 months.
44. 32. The computer program of claim 31, wherein the image is a three-dimensional magnetic resonance imaging image including an estimated mask.
45. 32. The computer program of claim 31, wherein the baseline cognitive status is represented by at least one cognitive score comprising at least one of a Clinical Dementia Rating Scale (CDRSB) score, an Alzheimer's Disease Assessment Scale-Cognitive Subscale (ADAS-COG12) score, and a Mini-Mental State Examination (MMSE) score.
46. A non-transitory computer-readable medium storing a computer program according to any one of claims 31 to 45.