Method for processing images of a brain
The method uses a pre-trained brain age model to predict and correct bias in brain age, enabling accurate dementia risk stratification through volumetric features, addressing the challenge of early dementia identification and intervention.
Patent Information
- Application Number
- PCT/GB2025/050518
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-04
- Filing Date
- 2025-03-14
- Publication Date
- 2025-09-18
AI Technical Summary
Existing methods fail to accurately and early identify individuals at risk for dementia by measuring biological age, which is crucial for effective preventative measures before the disease reaches an advanced stage.
A method using a pre-trained brain age model, such as a linear regression model, to predict brain age by extracting volumetric features from brain images, correcting for bias, and stratifying patients into dementia risk groups through a classification model, such as logistic regression, to provide reliable and accurate patient stratification.
The method effectively predicts dementia risk by accurately determining brain age and stratifying patients into risk groups, enabling early identification and targeted interventions, outperforming other machine learning approaches and providing a reliable predictor across all age groups.
Smart Images

Figure GB2025050518_18092025_PF_FP_ABST
Abstract
Description
[0001] Method for processing images of a brain
[0002] FIELD
[0003]
[0001] The present techniques relate to a method for processing images of a brain to extract at least one volumetric feature. The method comprises inputting the extracted information into a pre-trained brain age model to determine a user’s brain age and optionally using the determined brain age to stratify patients into groups according to dementia risk.
[0004] BACKGROUND
[0005]
[0002] Chronological age, i.e. how long a person has existed, can differ significantly from biological age, i.e. how old a person’s cells are. While chronological age is easily measured, the same cannot be said for biological age. A large number of factors may determine a person’s biological age, including lifestyle choices and genetic predispositions. Thus, biological age may be defined in a number of ways, for example, by analysing properties of a user’s brain.
[0006]
[0003] Changes in a user’s brain physiology can be indicators of certain diseases, such as dementia. However, such physiological changes are often only pronounced when a disease has already reached an advanced stage. Because measures such as lifestyle changes are usually most effective when implemented before any disease reaches an advanced stage, it is desirable to identify any changes in brain physiology as early as possible.
[0007]
[0004] The present applicant has identified the need for an improved method for stratifying patients into dementia risk groups.
[0008] SUMMARY
[0009]
[0005] In a first approach to the present techniques, there is provided a computer-implemented method for determining a patient’s brain age, the method comprising: extracting at least one volumetric feature from an image of a brain (i.e. from an image of the patient’s brain) by: obtaining (e.g. receiving or determining) at least one volume value for at least part of the patient’s brain, and normalising the at least one obtained volume value to obtain the at least one volumetric feature; and predicting brain age by: inputting the at least one extracted volumetric feature into a pre-trained brain age model, wherein the brain age model is a linear regression model, and wherein a bias of the linear regression model is corrected when predicting the brain age.
[0010]
[0006] The method may further comprise stratifying patients into dementia risk groups by inputting the predicted brain age into a classification model, e.g. a classification model comprising at least one logistic regression binary classifier. For example, each patient may be assigned to a dementia risk group based on their predicted brain age.
[0011]
[0007] In a second approach to the present techniques, there is provided a computer-implemented method for stratifying patients into dementia risk groups, the method comprising: extracting at least one volumetric feature from an image of a brain by: obtaining (e.g. receiving or determining) at least one volume value for at least part of a patient’s brain, and normalising the at least one obtained volume value to obtain the at least one volumetric feature; predicting brain age by inputting the at least one extracted volumetric feature into a pre-trained brain age model; and stratifying patients into dementia risk groups by inputting the predicted brain age into a classification model, wherein the classification model comprises at least one logistic regression binary classifier.
[0012]
[0008] The pre-trained brain age model may be a linear regression model. A linear regression model is a model that estimates the linear relationship between a scalar response (dependent variable) and one or more explanatory variables (regressor or independent variable). Dementia risk increases with chronological age. A linear regression model can successfully be used to model this relationship and use this to predict dementia risk in patients of all age groups. Predicting brain age may further comprise correcting a bias of the linear regression model. Correcting for age bias ensures that brain age is an effective predictor of dementia risk in all age groups.
[0013]
[0009] In a third approach to the present techniques, there is provided a computer-implemented method for stratifying patients into dementia risk groups, the method comprising: extracting at least one volumetric feature from an image of a brain by: receiving at least one volume value for at least part of a patient’s brain, and normalising the at least one received volume value to obtain the at least one volumetric feature; predicting brain age by inputting the at least one extracted volumetric feature into a pre-trained brain age model; and stratifying patients into dementia risk groups by inputting the predicted brain age into a classification model.
[0014]
[0010] The following features apply to any of the first to third approaches.
[0015]
[0011] Brain age acts as a reliable predictor of dementia status of a patient, and thus provides a fast and accurate method of stratifying patients into dementia risk groups.
[0016]
[0012] The pre-trained brain age model may be a linear regression model. A linear regression model is a model that estimates the linear relationship between a scalar response (dependent variable) and one or more explanatory variables (regressor or independent variable). Dementia risk increases with chronological age. A linear regression model can successfully be used to model this relationship and use this to predict dementia risk in patients of all age groups. Advantageously, linear regression models have been found to outperform other model types for this task. That is, linear regression models have been found to result in better performance than nonlinear machine learning approaches such as XGBoost, LGBM, AdaBoost, CatBoost, KNN, and Random Forests.
[0017]
[0013] Predicting brain age may further comprise correcting a bias of the linear regression model before inputting the predicted brain age into the classification model. Correcting for age bias ensures that brain age is an effective predictor of dementia risk in all age groups.
[0018]
[0014] Alternatively, predicting brain age may further comprise using a linear regression model comprising at least one penalty term. In this case, bias correction is inherent in the linear regression model and no further correction is needed.
[0019]
[0015] Correcting a bias of the linear regression model may comprise: obtaining an initial prediction by inputting the at least one extracted volumetric feature into the linear regression model; and bias correcting the initial prediction by subtracting an intercept of a regression line of the linear regression model and dividing by a slope of the regression line to obtain the predicted brain age which may then be input into the classification model. Dementia risk increases approximately linearly with age. This bias can be corrected by correcting the prediction using this linear relationship.
[0020]
[0016] Alternatively, wherein a bias of the linear regression model may be corrected by adding, to the linear regression model, penalty terms that counteract any bias. In other words, the correction may be done simultaneously when predicting the brain age using the brain age model.
[0021]
[0017] Obtaining at least one value of a volume of at least part of a patient’s brain may comprise obtain ing / receiving a volume value for at least one of hippocampus, ventricles, fusiform, medial temporal lobe, entorhinal cortex, fusiform gyrus and the whole brain, preferably wherein volume values are received for at least brainstem, 3rd-Ventricle, Left-choroid-plexus, Right-Caudate, white matter hypointensities, insular cortex left hemisphere, Right-VentralDC, Left-Thalamus, Left-Cerebellum- White-Matter and / or Left-Accumbens-Area.
[0022]
[0018] Obtaining at least one value of a volume of at least part of a patient’s brain may comprise obtaining sums of left and right brain features. That is, values for both the corresponding left and right brain features may be combined into a single volumetric value which is then input into the model.
[0023]
[0019] Extracting volumetric features from an image of a brain may further comprise: partitioning the image into cortical structures and / or cortical substructures to obtain a partitioned image; labelling each partition in the partitioned image according to the corresponding cortical structure and / or cortical substructure; outputting an image showing the labelled cortical structures and / or cortical substructures, wherein the at least one value of a volume corresponds to at least one of the labelled cortical structures and / or substructures in the output image. In other words, an image of a user’s brain may be processed to generate an edited version of the image which may then be output on a user interface. The edited version may comprise labels as described above and / or may comprise the predicted brain age and / or dementia risk group.
[0024]
[0020] Partitioning the image may comprise partitioning the image by finding a surface enclosing each of the cortical structure and / or cortical substructure.
[0025]
[0021] Alternatively, partitioning the image may comprise partitioning the image by using a Convolutional Neural Network, CNN, to find the cortical structures and / or substructures.
[0026]
[0022] Determining a normalised value of the volume may comprise obtaining a scaling function for scaling the volume value to obtain a normalised volume.
[0027]
[0023] Obtaining a scaling function may further comprise obtaining a scaling function according to statistical properties of a training dataset used to train the pre-trained model. This ensures that any unseen image is comparable to the data the brain age model has been trained on.
[0028]
[0024] Determining a normalised value of the volume may comprise mapping the volume value to a pre-determined statistical distribution. The pre-determined statistical distribution may, for example, be a Gaussian distribution.
[0025] In a fourth approach to the present techniques, there is provided a computer-implemented method for stratifying patients into dementia risk groups using a pre-trained brain age deep learning, DL, model, the method comprising: obtaining an image of a brain; processing the image of a brain using the brain age DL model to predict a brain age; and stratifying patients into dementia risk groups by inputting the predicted brain age into a classification model.
[0029]
[0026] The brain age DL model may be any of a fully convolutional network, a 3D Residual Network and / or a 3D Convolutional Neural Network.
[0030]
[0027] The following features apply to any of the first to fourth approaches.
[0031]
[0028] The brain age model and / or brain age DL model may have been trained using brain images from patients that are between 55 and 90 years old, preferably wherein the patients are between 55 and 85 years old. The brain age model and / or brain age DL model may have been trained using images of brains of individuals that were classified as healthy at the time the image of the brain was taken. Some of the patients included in the training data may have later been diagnosed with some kind of cognitive impairment, such as dementia. That is the training data may include images of brains that were healthy at the time the image of the brain was taken, but later developed into unhealthy brains, for example, by developing dementia.
[0032]
[0029] The classification model may comprise at least one binary classifier. A binary classifier may classify the elements of a set into one of two groups (each called c / ass). Binary classifiers that may be used, can, for example, be decision trees, random forests, Bayesian networks, support vector machines, neural networks, Probit models, genetic programming, multi expression programming and / or linear genetic programming. Any other suitable binary classifier may be used. To stratify patients into more than one dementia risk group (i.e. into low, medium and high risk groups), several binary classifiers may need to be used. Thus, the classification model may comprise a plurality of binary classifiers.
[0033]
[0030] Each binary classifier may be a logistic regression model. Logistic regression is a binary classification method. Several logistic regression models may be used to stratify patients into dementia risk groups. Logistic regression, or a logistic model, is a statistical model that models the log-odds of an event as a linear combination of one or more independent variables.
[0034]
[0031] Advantageously, logistic regression models have high interpretability and a reduced risk of overfitting, especially when there is only a limited number of minority-class samples in a binary classification problem. This may not be the case for some more complex classification models.
[0035]
[0032] The classification model may use the predicted brain age and / or a difference between a chronological age and the predicted brain age. Thus, classification may mean using a binary classification model, such as a logistic regression model and using either of the predicted brain age and / or the difference between chronological and predicted brain age in the binary classification model.
[0036]
[0033] Alternatively, this may mean simply using the predicted brain age and / or a difference between a chronological age and the predicted brain age itself as a classification. That is, the predicted brain age and / or a difference between a chronological age and the predicted brain age may themselves be associated with a dementia risk and may consequently be used as such.
[0037]
[0034] The classification model may comprise a first classification model for male patents and a second or further classification model for female patients. This is because brain age may be related differently to dementia risk for male and female patients. Different classification models may take this bias into account, resulting in more accurate predictions for both groups. When two separate models are used, it will be appreciated that references to the classification model throughout the specification may be interpreted as the first classification model and / or the second classification model appropriately.
[0038]
[0035] Alternatively, the classification model (i.e. a single classification model) may stratify both male and female patients into dementia risk groups. In this case, the model may be trained such that any differences are accounted for, and taking both groups of patients into account may result in a larger and more diverse set of training date which in turn leads to more accurate results.
[0039]
[0036] The or each classification model may have been trained using 10-fold cross validation and / or resampling prior to training. This may lead to the best classification results
[0040]
[0037] Patients may be stratified into dementia risk groups which indicate their current risk of dementia.
[0041]
[0038] Alternatively, or additionally, patients may be stratified into dementia risk groups which indicate their future risk of dementia.
[0042]
[0039] The image may be a magnetic resonance imaging, MRI, image of a brain. MRI images have been shown to be effective and accurate in determining volume values for the whole brain and / or parts of the brain. However, other suitable methods for obtaining an image of a user’s brain may be used. For example, a brain image may be obtained using a computer tomography (CT) or positron emission tomography (PET) scan.
[0043]
[0040] Stratifying patients into dementia risk groups comprises stratifying patients into high, medium and / or low dementia risk groups. High dementia risk may correspond to a “dementia” classification in some datasets used for training and verification of the present techniques, whereas medium risk may correspond to “mild cognitive impairment - MCI”, and low risk may correspond to “cognitively normal - CN”. Equally, patients may be stratified into further categories, such as “healthy”, which may for example, encompass the “cognitively normal” category, or “not healthy” which may encompass several categories, such as “mild cognitive impairment”, “dementia” or other categories such as “subjective memory complaints”, “Parkinson’s disease” or “Multiple Sclerosis”.
[0044]
[0041] The method may further comprise outputting a diagnosis of dementia according to the stratification into dementia risk groups. That is, when a patient is determined to be in a high dementia risk group, the patient may be diagnosed with dementia.
[0045]
[0042] The method may further comprise identifying patients for treatment depending on the stratification into dementia risk groups. That is, patients in the high or medium dementia risk groups may be identified for treatment. This treatment may, for example, include medication or other treatments. Some patients may be identified for treatment targeting certain disease factors, such as genetics and epigenetics, inflammation, misfolded proteins, mitochondria and metabolic function, neuroprotection, vascular and / or other factors.
[0046]
[0043] In a fifth approach to the present techniques, there is provided a method of treatment of a subject with a high risk of dementia comprising the steps of predicting a level of dementia risk using the stratification identified above. For example, the method may comprise: extracting at least one volumetric feature from an image of a brain by: receiving at least one volume value for at least part of a patient’s brain, and normalising the at least one received volume value to obtain the at least one volumetric feature; predicting brain age by inputting the at least one extracted volumetric feature into a pre-trained brain age model; determining a dementia risk of the subject based on the predicted brain age; selecting a treatment according to the determined dementia risk; and administering the treatment. The method may further comprise correcting for bias in the brain age model by: determining a slope and an intercept of a regression line between predicted brain age and age label for the set of training images; and bias correcting the predicted brain age before comparing the predicted brain age with the brain age label. Similarly, the method may comprise obtaining an image of a brain; processing the image of a brain using the brain age DL model to predict a brain age; determining a dementia risk based on the predicted brain age into a classification model, selecting a treatment according to the determined dementia risk; and administering the treatment.
[0047]
[0044] When the dementia risk is high, the treatment may be selected from one or more of: AAV- hTERT, Emtricitabine, LX1001 , Vorinostat, AL003, Bacillus Calmette-Guerin (BCS), JNJ- 40346527, Salsalate, Sirolimus, XPro1595, BEY2153, E2814, Lu AF87908, LY3372993, Trehalose, Grape seed polyphenolic extract, Nicotinamide riboside, Allopregnanolone, Dasatinib + Quercetin, Efavirenz, Mesenchymal Stem Cells, NNI-362, REM0046127, SNK01 , Dabigatran, MK-1942 + Donepezil, MK-4334, 3TC, AL002, ALZT-OP1 (cromolyn and ibuprofen), Bacillus Calmette-Guerin (BCG), Daratumumab, GV1001 , Lenalidomide, Montelukast, Tacrolimus, VX- 745, ABBV-8E12, ABvac40, ACI-35.030 + JACI-35.054, ALZ-801 , APH-1105, BIIB092, IGNIS MAPTRx, JNJ-63733657, LY3303560, Meganatural-Az Grapeseed Extract, NewGam 10% IVIG, Nilotinib, Posiphen, PQ912, RO7126209, Semorinemab, TEP, Benfotiamine, Dapagliflozin, L- Serine Gummy, Liraglutide, Metabolic Cofactor Supplementation, Nicotinamide, Pepinemab, T3D-959, AMX0035, AstroStem and Donepezil, CB-AC-02, CORT108297, HB-adMSCs, Human Mesenchymal Stem Cells, Leuprorelin, Mesenchymal Stem Cells, PTI-125, Rapamycin, Regulatory T cells, S-equol, Sovateltide, T-817MA, Valacyclovir, AD-35, AD-35 60mg, ATH-1017, BPN14770, Bromocriptine, Bryostatin 1 , CT1812, DAOI-A, intermittent Theta Burst Stimulation (iTBS), Levetiracetam, MemorEM, Nicotine Transdermal Patch, SAGE-718, Transcranial Alternating Current Stimulation (tACS), AR1001 , Perindopril|Telmisartan, NE3107, Aducanumab, Donanemab, Gantenerumab, Lecanemab, Trx0237, Extended-release metformin, Ginkgo biloba, GV-971 , Tricaprilin, AGB101 , ANAVEX2-73, Donepezil, Guanfacine, Octohydroaminoacridine, succinate, Troriluzole, Renew NCP-5, BPDO-1603, COR388. Additionally or alternatively, the treatment may be selected from one of more of acetylcholinesterase (AChE) inhibitors, memantine, antipsychotic medicines, antidepressants, cognitive stimulation therapy (CST), cognitive rehabilitation, reminiscence work (talking about things and events from the patient’s past) and / or life story work (compilation of photos, notes and keepsakes from the patient’s childhood to the present day). Other suitable treatment options may be used. Additionally or alternatively, lifestyle changes may be suggested to a patient. These may, for example, include changes in diet and / or exercise pattern, as well as cognitive exercises.
[0048]
[0045] Some treatments are only effective when given early on, while dementia is still developing. In this case, only patients in the low or medium dementia risk group may be identified as suitable forthis treatment. Such treatments may, for example, include Icosapent ethyl, Memantine, Omega 3 treatment, Losartan, an Active NIR-PBM device, Gantenerumab, Crenezumab, Deferiprone, Omega 3 PUFA, Solanezumab, Telmisartan, AGB101 and / or DHA. Additionally, or alternatively to treatment, patients may be identified for other interventions, such as, for example, dietary changes, increasing exercise and / or reducing stress. Lifestyle changes may be particularly helpful if a patient is classified as having a medium dementia risk, as this may then slow down progression to a high dementia risk.
[0049]
[0046] When the dementia risk is medium or low, the patient may be selected for further monitoring. Further monitoring may comprise repeating the brain age prediction and dementia risk classification process as described above. This may be done after a set time interval. For example, further monitoring may take place every 6 months, or at any other suitable time interval, such as every 12 months, every 18 months or every 24 months. If, subsequently, it is determined that a patient has a high risk of dementia, the patient may be selected for treatment as described above.
[0050]
[0047] As will be appreciated by the skilled person, the terms “treating”, “treats” and “treatment” include both preventative and curative treatment of a condition, disease or disorder. These terms also include slowing, interrupting, controlling or stopping the progression of a condition, disease or disorder and preventing, curing, slowing, interrupting, controlling or stopping the symptoms of a condition, disease or disorder. As used herein, "treat", "treating" or "treatment" means inhibiting or relieving a disease or disease. For example, treatment can include a postponement of development of the symptoms associated with a disease or disease, and / or a reduction in the severity of such symptoms that will, or are expected, to develop with said disease. The terms include ameliorating existing symptoms, preventing additional symptoms, and ameliorating or preventing the underlying causes of such symptoms. Thus, the terms denote that a beneficial result is being conferred on at least some of the mammals, e.g., canine patients, being treated. Many medical treatments are effective for some, but not all, patients that undergo the treatment.
[0051]
[0048] The term "subject" or "patient" refers to an animal or human which is the object of treatment, observation, or experiment. By way of example only, a subject includes, but is not limited to, a mammal, including, but not limited to, a human or a non-human mammal, such as a non-human primate, murine, bovine, equine, canine, ovine, or feline.
[0049] In a sixth approach to the present techniques, there may be provided a method of therapy monitoring in a subject receiving treatment for dementia, the method comprising: extracting at least one volumetric feature from an image of a brain by: receiving at least one volume value for at least part of a patient’s brain, and normalising the at least one received volume value to obtain the at least one volumetric feature; predicting brain age by inputting the at least one extracted volumetric feature into a pre-trained brain age model; and determining a dementia risk of the subject based on the predicted brain age; determining, using the dementia risk, whether there is an improvement in the dementia risk compared to a reference dementia risk. It will be appreciated that this approach can also be adapted to the method of predicting brain age using the DL model described above.
[0052]
[0050] That is, a number of brain images may be analysed overtime. These analysed images may be used to determine if a treatment that is being administered to the subject is producing the desired effect. Thus, a reference dementia risk may be a dementia risk for a first brain image of the subject. A first brain image may be a reference image which may be obtained from a healthy ordiseased individual, the healthy ordiseased individual may or may not have received treatment. Dementia risks may be determined for subsequent brain images and these subsequent dementia risks may be compared against the reference dementia risk, i.e. the dementia risk associated with the first brain image. In this way, the efficacy of the treatment can be monitored over time.
[0053]
[0051] The method may further comprise generating training data comprising associated response data (e.g. outcome from a particular treatment) for each of a plurality of subjects. This training data may then be used to train a treatment selection model that recommends a particular treatment or other measure (such as particular diet). Such a method may be used to create a predictive signature of treatment response, to help stratify patients for the most appropriate treatment regime. The method may thus further comprise training a model to recommend a treatment.
[0054]
[0052] In a seventh approach to the present techniques, there is provided a computer- implemented method for training a brain age model to stratify patients into dementia risk groups, the method comprising: obtaining a set of brain images and a set of age labels corresponding to each image in the set of brain images; extracting at least one volumetric feature from each brain image by: receiving at least one volume value for at least part of a patient’s brain, and normalising the at least one received volume value to obtain the at least one volumetric feature; predicting brain age by inputting the at least one extracted volumetric feature into the brain age model; comparing the predicted brain age with the brain age label and using the comparison to train the brain age model; and training a risk prediction model to stratify patients into dementia risk groups based on the brain age prediction by the brain age model. It will be appreciated that this approach can also be adapted to the method of predicting brain age using the DL model described above.
[0055]
[0053] The method may further comprise correcting for bias in the brain age model by: determining a slope and an intercept of a regression line between predicted brain age and age label for the set of training images; and bias correcting the predicted brain age before comparing the predicted brain age with the brain age label.
[0056]
[0054] Obtaining a set of brain images and a set of brain age labels may comprise obtaining a set of cognitively normal brain images and corresponding set of age labels. This ensures that brain age is an accurate predictor of dementia risk. If diseased brain images were used, it would not be possible to determine deviation from the norm in these diseased images. That is, brain age prediction is built on the assumption that degradation due to dementia looks similar to degradation due to ageing.
[0057]
[0055] In an eighth approach to the present techniques, there is provided a computer-readable storage medium comprising instructions which, when executed by a processor, causes the processor to carry out any of the methods described above.
[0058]
[0056] In a ninth approach to the present techniques, there is provided a image processing system comprising: an image capture device which is configured to capture an image of a patient’s brain; an image processor which is configured to receive an image from the image capture device and carry out any of the methods described above; and a user interface which is configured to display an output result generated by the image processor.
[0059]
[0057] As will be appreciated by one skilled in the art, the present techniques may be embodied as a system, method or computer program product. Accordingly, present techniques may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects.
[0060]
[0058] Furthermore, the present techniques may take the form of a computer program product embodied in a computer readable medium having computer readable program code embodied thereon. The computer readable medium may be a computer readable signal medium or a computer readable storage medium. A computer readable medium may be, for example, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing.
[0061]
[0059] Computer program code for carrying out operations of the present techniques may be written in any combination of one or more programming languages, including object oriented programming languages and conventional procedural programming languages. Code components may be embodied as procedures, methods or the like, and may comprise subcomponents which may take the form of instructions or sequences of instructions at any of the levels of abstraction, from the direct machine instructions of a native instruction set to high-level compiled or interpreted language constructs.
[0062]
[0060] Embodiments of the present techniques also provide a non-transitory data carrier carrying code which, when implemented on a processor, causes the processor to carry out any of the methods described herein.
[0063]
[0061] The techniques further provide processor control code to implement the above-described methods, for example on a general purpose computer system or on a digital signal processor (DSP). The techniques also provide a carrier carrying processor control code to, when running, implement any of the above methods, in particular on a non-transitory data carrier. The code may be provided on a carrier such as a disk, a microprocessor, CD- or DVD-ROM, programmed memory such as non-volatile memory (e.g. Flash) or read-only memory (firmware), or on a data carrier such as an optical or electrical signal carrier. Code (and / or data) to implement embodiments of the techniques described herein may comprise source, object or executable code in a conventional programming language (interpreted or compiled) such as C, or assembly code, code for setting up or controlling an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array), or code for a hardware description language such as Verilog (RTM) or VHDL (Very high speed integrated circuit Hardware Description Language). As the skilled person will appreciate, such code and / or data may be distributed between a plurality of coupled components in communication with one another. The techniques may comprise a controllerwhich includes a microprocessor, working memory and program memory coupled to one or more of the components of the system.
[0064]
[0062] It will also be clear to one of skill in the art that all or part of a logical method according to embodiments of the present techniques may suitably be embodied in a logic apparatus comprising logic elements to perform the steps of the above-described methods, and that such logic elements may comprise components such as logic gates in, for example a programmable logic array or application-specific integrated circuit. Such a logic arrangement may further be embodied in enabling elements for temporarily or permanently establishing logic structures in such an array or circuit using, for example, a virtual hardware descriptor language, which may be stored and transmitted using fixed or transmittable carrier media.
[0065]
[0063] In an embodiment, the present techniques may be implemented using multiple processors or control circuits. The present techniques may be adapted to run on, or integrated into, the operating system of an apparatus.
[0066]
[0064] In an embodiment, the present techniques may be realised in the form of a data carrier having functional data thereon, said functional data comprising functional computer data structures to, when loaded into a computer system or network and operated upon thereby, enable said computer system to perform all the steps of the above-described method.
[0067] BRIEF DESCRIPTION OF DRAWINGS
[0068]
[0065] For a better understanding of the invention, and to show how embodiments of the same may be carried into effect, reference will now be made, by way of example only, to the accompanying diagrammatic drawings in which:
[0069]
[0066] Figure 1 is a flowchart showing steps of the method at inference time;
[0070]
[0067] Figure 2 is a flowchart showing steps of the method at training time;
[0071]
[0068] Figures 3a to 3f are charts plotting the volume of different brain regions (whole brain, ventricles, entorhinal cortex, fusiform gyrus, medial temporal lobe and hippocampus, respectively) against age for patients with and without dementia;
[0069] Figures 4a and 4b show predicted brain age plotted against chronological age for test and training data with and without scaling respectively;
[0072]
[0070] Figure 5a is a graph plotting the results by the classification assigned by the classification model (risk prediction model);
[0073]
[0071] Figure 5b shows predicted brain ages plotted against chronological age for participants of the UKB study;
[0074]
[0072] Figure 6a is a bar diagram showing how the probability of having dementia or a cognitively normal (CN) brain are distributed with chronological age;
[0075]
[0073] Figure 6b is a bar diagram showing how the probability of having dementia or a cognitively normal brain are distributed with biological / brain age;
[0076]
[0074] Figure 7 shows a receiver operating characteristic (ROC) curve which plots true positive rate against false positive rate for four risk prediction models;
[0077]
[0075] Figures 8a and 8b plot values for the area under the receiver operating characteristic (AUROC) against months from baseline for different risk prediction models;
[0078]
[0076] Figure 9 is a diagram of a system for implementing the methods described above; and
[0079]
[0077] Figures 10a and 10b show a comparison of feature rank for the 4 variations of Cole model trained with FastSurfer features.
[0080] DETAILED DESCRIPTION OF DRAWINGS
[0081]
[0078] Broadly speaking, embodiments of the present techniques provide a method for determining biological age (also termed brain age) from an MRI image of a brain. In particular, the present techniques may be used to estimate a brain age and a risk of developing a neurodegenerative disease based on the brain age.
[0082]
[0079] Magnetic Resonance Imaging (MRI) of the brain is one of the most widespread tests in neurology and neurosurgery, providing insights into the brain anatomy and various physiological processes occurring within this region. T1 -weighted (T1w) MRI sequences are one of the most common MRI scan types in clinical settings and provide highly detailed anatomical views of brain structures due to their ability to differentiate between T1 relaxation times of different tissues (enhancing the signal of fatty tissue and suppressing the water signal).
[0083]
[0080] In the context of neurodegenerative diseases, T1w MRI is critical for risk prediction and early diagnosis. For instance, in Alzheimer's disease, T1w scans have been used to study the development of atrophy in both healthy and diseased patients (Blanche et al., 2022). Similarly, in Parkinson's disease, T1w MRI can provide insight into the specific regions responsible for levodopa-a dopamine precursor-response of Parkinson’s disease patients which might help in determining treatment strategies (Yan et al., 2024). In the context of multiple sclerosis, the combination of information extracted from T1w scans and another type of structural MRI known as T2w (the T1w / T2w intensity ratio), has been proposed as a clinical biomarker to assess tissue integrity and a decrease in said biomarker has been observed to precede lesion formation (Boaventura et al., 2022). The observed early structural changes can provide a window for early intervention, which is essential for slowing the progression of neurodegenerative disorders.
[0081] T1w MRI has also gained significant attention in the emerging field of brain age prediction. Research in this context seeks to estimate an individual's age-a.k.a. their brain age-based on their neuroimaging data (e.g., from structural characteristics observed in T1w brain scans). While there is not much value to be gained by a model predicting the chronological age of an individual, given that this information can be easily accessed for any patient without the need for an MRI, the real value of brain age models lies in the fact that brain age is trained to be a representation of healthy brain ageing. Thus, an increased brain age compared to chronological age can serve as a biomarker for accelerated brain ageing, which is often associated with cognitive decline and an increased risk of neurodegenerative conditions (Wrigglesworth et al., 2022, Wagen et al., 2022). The ability of T1w MRI to provide precise measurements of brain structures makes it particularly suitable for this application, as small deviations in brain morphology can be quantitatively assessed and correlated with ageing-related pathologies (Cole et al., 2017). Therefore, the application of T1w MRI in brain age prediction underscores the importance of this MRI modality in understanding the complex relationship between brain structure, ageing, and neurodegeneration in addition to its crucial role in the early detection and prediction of neurodegenerative diseases.
[0084]
[0082] The present techniques may be used to determine a brain age in a user and use the determined brain age to distinguish users with dementia, mild cognitive impairment (MCI) and no cognitive impairment (cognitive normal - CN). That is, brain age may be used to stratify patients into dementia risk groups (for example high risk of dementia, medium risk of dementia (MCI) or low risk of dementia (CN)). The biological motivation behind using a brain age is that several changes in the brain associated with neurodegenerative diseases are also observed in normal ageing, but the phenomenon is enhanced or accelerated in those users that have a disease. That is, in users without a disease, the same changes may take place later on or more slowly than in those users who have Alzheimer’s disease or another neurodegenerative disease. Thus, brain age is a suitable predictor for Alzheimer’s disease and other neurodegenerative diseases.
[0085]
[0083] In studies by Cole et al., 2017; Cole & Franke, 2017; Cole et al., 2018; Cole, 2020; Zhang et al., 2023, the primary objective was to develop brain age models (also known as brain clocks), either using the UKBB dataset exclusively to create and validate the brain clock or by combining the UKBB with other well-established MRI-related databases. The utility of brain age as a biomarker of a specific process or indicator of a particular condition is often investigated next. For example, in Cumplido-Mayoral et al., 2023 XGBoost-based brain age models trained exclusively on UKBB data were built and then evaluated on four external datasets (ALFA+, ADNI, EPAD, and OASIS) to investigate associations between brain age deltas and various biomarkers of Alzheimer’s disease pathology, such as CSF A|3, p-tau, APOE-e4, and plasma NfL. Generally, the brain clock is treated as a regression model that, given an MRI scan or a set of features extracted from it, aims to accurately predict the age of a healthy or cognitively normal individual. As previously mentioned, the focus on healthy individuals for model training is paramount, especially for model utility, but the way in which individuals are categorised as healthy is by no means unique. In fact, different approaches can be found in the literature and this task can become particularly challenging when trying to merge databases where the information has been collected from participants recruited with different criteria and for which different clinical information exists. We will consider as reference the seminal work by Cole and colleagues, and employ the corresponding definition of healthy in the analysis. However, due to the fact that subjects enrolled in a specific study may change health status over subsequent follow-ups, a less restrictive definition will be also proposed and evaluated.
[0086]
[0084] As outlined in detail by Zhang et al. (2023), the development of a brain clock typically involves two training steps: first, the regressor is trained, often using various deep learning (DL) architectures, if working directly with images (Peng et al., 2021), or standard machine learning (ML) techniques, where linear models, e.g., LASSO, Ridge models, are frequently employed due to the nature of the features used. Afterthis, an age correction model is learned (Cole et al., 2020; Zhang et al., 2023) and later applied in the attempt to eliminate the systematic bias observed in age prediction problems. A thorough review of these correction methods can be found in (Zhang et al., 2023) but in general terms these corrections fall into two categories: those that do not require knowledge of the individual’s chronological age, such as Cole’s method (Cole et al., 2018; Peng et al., 2021 ; Smith et al., 2019), and those that do require chronological age for application, such as Beheshti's / Lange’s method (de Lange et al., 2020) and Zhang's age-level correction (Zhang et al., 2023). Here we summarise these approaches and further investigate the statistical impact of the corrections.
[0087]
[0085] The present techniques focus on generating various types of brain clocks, utilising a range of DL and traditional ML approaches, as well as developing corresponding age-bias correction terms using the UKBB database. Our objectives for the presented work are threefold. First, we want to demonstrate that models developed using UKBB generalise well to other MRI-related databases, such as the Alzheimer's Disease Neuroimaging Initiative (ADNI) (Jack et al., 2008) and the National Alzheimer's Coordinating Center (NACC) (Beekly et al., 2007). These datasets are critical in neuroimaging research, with ADNI providing extensive data on individuals at various stages of cognitive health, from normal ageing to Alzheimer's disease, and NACC offering a diverse range of clinical and neuroimaging data across different cognitive statuses. Validating our models on these datasets is essential to ensure their robustness and applicability beyond the specific characteristics of the UKBB. Second, we aim to provide a more in-depth comparison of brain clock models and age-bias correction methods, with respect to the standards frequently observed in the literature, by exploring alternatives for developing age-bias corrections through customised preprocessing steps and by extending the set of metrics used to evaluate the performance of our approaches. This includes not only traditional metrics like MAE but also assessments of model performance across specific age brackets and assessment of their utility in distinguishing healthy individuals from different subgroups of interest, an area with limited existing research. Finally, we will provide recommendations on the most promising approaches for both accurate chronological age prediction and condition prediction.
[0086] Changes that may be observed in the brain include changes in volume, texture, shape of brain substructure, as well as change composition, connectivity, metabolism and functionality. Changes in brain substructure may, for example, be quantified using structural magnetic resonance imaging (MRI). In particular, changes in brain substructure may be quantified using T1 and T2-weighted structural MRI images. Change composition may, for example, be determined using MRI images that include diffusion tensor imaging. Metabolism and functionality may, for example, be monitored using functional magnetic resonance imaging (fMRI). Other suitable methods for obtaining an image of a user’s brain may be used. For example, a brain image may be obtained using a computer tomography (CT) or positron emission tomography (PET) scan.
[0088]
[0087] Figure 1 is a flowchart showing steps of the method at inference time. In a first step S100 an image of a brain is received. The image may be an MRI image as described above. Depending on the features of the image that are to be analysed in subsequent steps, the image may be a different type of MRI image or comprise different features, as described above. At step S102, the image is partitioned into different components. Each of these components may represent a separate structure and / or substructure of the user’s brain as shown on the image.
[0089]
[0088] T1w MRI scans of the entire head were downloaded from the databases mentioned below in either DICOM or NlfTI format. DICOM files were converted to NlfTI using dcm2niix (Li et al., 2016), a robust tool for efficient and accurate format conversion. After conversion the images were processed with FastSurfer (Henschel et al., 2020), which provides a fast and accurate means of standardising MRI data, facilitating downstream neuroimaging analyses. By utilising DL algorithms this approach dramatically reduces the time of whole brain segmentation and preprocessing to as little as 3-5 minutes per scan, without compromising accuracy (Henschel et al., 2020). The software has an Apache 2.0 licence, and provides a convenient and cost-effective option where large datasets have to be processed and-more generally-for scaling neuroimaging research. The pipeline performs accurate volumetric segmentation of the brain, but also generates a set of output images in .mgz format with desirable properties for DL training, such as transformation into a “conformed space”, intensities normalisation and skull removal.
[0090]
[0089] The transformation to a conformed space includes several crucial steps: turning intensity values into unsigned 8-bit integers (UCHAR), reslicing the images into a standardised orientation, padding the images to a uniform 256x256x256 matrix size, and enforcing a voxel size of 1 mm isotropic. Conformation of the MRI images is a standardisation step which is critical for constant data input across ML and DL approaches, as it eliminates potential variability due to different scan orientations, resolutions, or fields of view. This step ensures that the images are aligned and normalised, enhancing the reliability and comparability of features extracted from the images.
[0091]
[0090] While our DL-based approaches used the conformed versions of the MRI scans directly, for our ML models we rely on volumetric features (i.e., estimations of the volume of a predefined set of brain regions) derived from the brain segmentation output of the FastSurfer pipeline.
[0092]
[0091] Partitioning may take place using, for example, software such as FreeSurfer (http: / / surfer.nmr.mgh.harvard.edu / ) and / or FastSurfer (Henschel, Leonie, et al. "Fastsurfer-a fast and accurate deep learning based neuroimaging pipeline." NeuroImage 219 (2020): 1 17012.). These are software packages that can be used for the analysis and visualisation of structural and functional neuroimaging data. That is, these programs reconstruct the surface area of different parts of a user’s brain and may use prior knowledge about the topology of a typical human brain to do so. This information may then be used to partition an image of a user’s brain and label cortical and subcortical structures in a user’s brain. Partitioning may mean assigning each pixel in an image of a user’s brain to at least one partition. Labelling may mean comparing each partition to a known cortical structure and / or cortical substructure and consequently assigning a cortical structure / substructure to each partition. Partitioning the cortical structures may be referred to as “parcellation” whereas partitioning subcortical structures may be referred to as “segmentation”. Both parcellation and segmentation may be performed when partitioning the image. The whole image of the brain may be parcellated and segmented, or only certain parts of the image of the brain may be parcellated and segmented. For example, a certain cortical structure and corresponding substructures may be considered more important for determining a user’s brain age than other structures. In that case, the whole brain may be parcellated and only certain cortical structures may be segmented into their cortical substructures.
[0093]
[0092] In other words, an image of a user’s brain is processed to generate an edited version of the image. The edited version of the image comprises labelled portions of the image that each represent a different part of the brain. A different part of the brain may mean a cortical structure and / or substructure. A different part of the brain may also mean a combination of more than one cortical structure and / or substructure. A different part of the brain may also mean only part of a cortical structure or substructure. This part of the cortical structure or substructure may be determined such that a specific anatomic part of the structure or substructure is always used. It will be appreciated that the image of the user’s brain may be also edited to provide additional information, e.g. the predicted brain age and / or risk classification, which are generated as described below.
[0094]
[0093] In a next step, the volume of each partition S104 of the user’s brain may be determined, using any suitable technique for example programs such as FreeSurfer or FastSurfer. In other words, the volume may be considered to be received, obtained and / or calculated using appropriate techniques or programs. When using such programs, this can be done by finding the volume enclosed by the surface that separates each partition from another. FreeSurfer may find the volume by finding a plurality of surfaces that each enclose a different cortical structure or substructure using minimisation techniques. FastSurfer may partition the image of the user’s brain using, for example, Convolutional Neural Networks (CNNs). Other suitable techniques may be used to partition the image of the user’s brain into different anatomical features (i.e. cortical structures or substructures). For example, other deep learning techniques capable of performing image partitioning may be used, such as transformer models.
[0095]
[0094] Next, the volume of each partition is determined S104. This may, for example, be done using the same tools that partition the image, e.g. FreeSurfer and / or FastSurfer tools. In FreeSurfer, volume may be determined using the surfaces enclosing each cortical structure and / or substructure. In particular, volumetric segmentation in FreeSurfer may done using the Desikan-Killiany atlas. Other suitable techniques, such as deep learning / machine learning techniques may be used to find the volume of each of the cortical structures / substructures. While partitioning the image S102 and finding the volume of each partition S104 may be separate steps, they may equally be performed in one step. This may, for example, be the case when machine learning techniques are employed. The output of machine learning and / or deep learning techniques is usually determined by the training data that is being used. In particular, such models are usually trained to output a specific category of result. In other words, a trained model usually outputs data according to how the training data was labelled. In that case, if training data labels contain the volume of specific cortical structures and / or substructures, the trained model may simply output a value for the volume of each of these structures. A step of identifying the cortical structures and / or substructures themselves may then no longer be necessary. Thus, steps S102 and S104 may be performed simultaneously.
[0096]
[0095] In some embodiments, steps S100, S102 and S104 may already have been performed. This is, for example, the case if partitioning of the brain image has already happened remotely. For example, partitioning of the brain image may take place when the brain image is taken using, for example, an MRI scanner. In this case, determining a brain age and image classification may take place in a different location and / or using a different computer or server. This may, for example, be advantageous when brain image files are very large and would need to be transmitted via a network connection for brain age determination and image classification. Then, only values for the volumes of each partition may be transmitted and steps S106 to S116 may be performed in a different location and / or using a different device.
[0097]
[0096] The determined volume of each partition is then normalised S106 to ensure that images of the brains of the different users are comparable. Normalised volumes are referred to as volumetric features. Normalisation is done because some brains may have an overall larger volume than other brains. For example, the average volume of a male brain is around 1260 cm3whereas the average volume of a female brain is around 1130 cm3. Additionally, there can be substantial variation of volume between individuals. Thus, the only way to reliably compare volumetric features in a brain image is to normalise the volume that is calculated for each of these features.
[0098]
[0097] Any suitable normalisation technique may be used. For example, a simple method is to normalise the volumes of different cortical structures and / or substructures by dividing the volume of each partition by a total volume of the user’s brain. Thus, when determining a volume for each partition in step S104, this may include determining a volume of the whole brain. Alternatively, normalisation may be done by scaling each determined volume V to a given range r to obtain a normalised volume Vnormallsed. For example, the normalisation may be determined by the properties of other data, such as training data used to train a trained brain age model. Then, normalisation may, for example, mean transforming the determined volume according to: t—(V ~ mlri) / max ~ ^min) ^normalised ~ ^std ^Jmax ^"min) + Tnin where Vmaxis the largest volume value in the training data and Vminis the smallest volume value in the training data, rmaxis the higher end of the range r and rminis the lower end of the range r. The range r may, for example, take values from 0 to 1 . Then, rmln= 0 and rmax= 1. For example, a volume value may refer to the volume of a particular cortical structure or substructure. This type of normalisation may, for example, be performed using the MinMaxScaler function in the preprocessing module of scikit-learn .documentation).
[0099]
[0098] When normalisation as above is performed, only the first equation may be applied to normalise the data. That is, only Vstdmay be calculated and this may then be used to perform the methods described herein.
[0100]
[0099] Alternatively, normalisation may be percentile based (for example, by using the scikit-learn RobustScaler function) or normalisation may mean mapping the determined volume to a statistical distribution, such as a Gaussian distribution. For example, volumetric features input into the brain age model may be scaled using scikit-learns “StandardScaler”. Thus, scaling the volumetric features may mean standardising features by removing the mean and scaling the features to unit variance. The volumetric features are then centred and scaled. This means that the volumetric features may be scaled in such a way that the data is standard normally distributed (i.e. according to a Gaussian distribution with zero mean and unit variance). This scaling may take place in addition to the scaling described above, or the volumetric features may only be scaled to conform to a Gaussian distribution.
[0101]
[0100] In each of the above described cases, the normalisation / scaling may happen using properties of a training data distribution. That is, the brain age model may be trained using training data and the statistical properties of that training data may be used to normalise any subsequent data which the brain age model classifies or otherwise processes. Only one type of normalisation, i.e. one type of scaler may be used, or several normalisation processes may be applied to the data.
[0102]
[0101] Next, the volumetric features are input into a trained brain age model S108 to find a user’s brain age. The model is trained on data from the Alzheimer’s Disease Neuroimaging Initiative (ADNI) database (htt p ADNI study participants are stratified into CN, MCI and Dementia groups. CN, MCI and Dementia groups may correspond to low, medium and high risk of dementia respectively. This classification is done based on the corresponding ADNI labels. In particular, this classification is done based on the “DX” label (i.e. CN, MCI, dementia). That is, ADNI already provides a classification according to whether a subject suffers from a neurodegenerative disease or is otherwise cognitively impaired. These classifications have been translated into the classifications used in the present application. The model may only be trained on study participants that have been determined to be “cognitively normal”. The brain age model is trained to predict a subject’s chronological age from a corresponding MRI image. For cognitively normal subject, chronological age is taken to correspond to biological / brain age. Thus, for the brain age models, the subject’s chronological age are used to label the training brain images.
[0103]
[0102] All volumetric features may be input into the trained brain age model or only some volumetric features may be input into the trained brain age model. Thus, there may be an optional step of selecting which volumetric features to input into the trained brain age model. Alternatively, this selection may take place prior to partitioning the image of the brain and determining the volume of each partition. Selecting which partitions, i.e cortical structures and / or substructures, are to be input into the trained brain age model before partitioning the whole brain may be advantageous when the partitioning is computationally intensive. For example, partitioning an MRI image of a user’s brain using FreeSurfer may take several hours. Such computational time may be reduced by preselecting which structures and / or substructures are relevant and reducing the number of structures and / or substructures that need to be identified as far as possible.
[0104]
[0103] In some cases, preselection of which cortical structures / substructures should be identified is not possible. This may, for example, be the case for some deep learning methods for partitioning an image of the brain. Such methods may automatically partition the whole brain and little computational resources can be saved by only selecting certain areas of the brain that are to be partitioned. In that case, certain cortical structures / substructures may be selected from the output produced by the partitioning that has taken place using deep learning methods.
[0105]
[0104] When selecting which volumetric features are input into the pre-trained brain age model, some cortical structures and / or substructures may be more relevant than others. Examples of structures and / or substructures that may be input into the pre-trained brain age model are: hippocampus, ventricles, fusiform, medial temporal lobe, entorhinal cortex, fusiform gyrus and / or the whole brain. Other suitable volumetric features may be input into the pre-trained brain age model. Alternatively, or additionally features representing white-matter microstructure, such as diffusion MRI images, may be input into the pre-trained model.
[0106]
[0105] The volumetric features input into the brain age model may be weighted in some manner. Weighting may take place by training the brain age model to prioritise some features over other features. Additionally or alternatively, volumetric features may be multiplied or otherwise modified by a weight. Such a weight may simply be a multiplication factor. A weight may also mean a more complex means of modifying each volumetric feature.
[0107]
[0106] The trained brain age model may be any suitable trained brain age model. For example, the trained brain age model may be a linear regression model. The linear regression model may be implemented using a number of different methods. For example, the linear regression model may be implemented using the scikit-learn library (Scikit-learn: Machine Learning in Python, Pedregosa et al., JMLR 12, pp. 2825-2830, 2011. - scikit-learn.org) for Python. Other suitable models may be used. For example, other type of regression models, such as Bayesian elastic nets, Gradient Boosting models, such as XGBoost, LASSO (least absolute shrinkage and selection operator), support vector machines, Convolutional Neural Networks (CNNs), and / or Artificial Neural Networks (ANNs) may be used.
[0107] Among the established approaches in ML, we implemented the model proposed by Cole et al. (2020), that is based on a LASSO regression, for benchmarking the capability of FastSurfer features in predicting brain age. Feature selection is performed by fixing the penalty parameter gathered with 10-fold cross-validation and then repeating the regression by bootstrapping the training data 1000 times. In this way, a distribution for each feature coefficient is generated, and features that cross zero at 95 confidence level are excluded. The mean values of the population of the coefficients is taken for the selected features in the models labelled as “bootstrap”. As an alternative, a linear model is also re-trained by considering only the selected features for the models we label as “bootstrap+linear”. The features of the model, which are the image derived phenotypes (IDPs), have been standardised with respect to the training set before training the LASSO models. The standardisation model is trained as the first step in the pipeline and is therefore part of the model itself.
[0108]
[0108] Beyond these approaches, we considered a broader range of penalised linear regression models, including LASSO, Ridge, and Elastic Net, to further explore how FastSurfer-extracted IDPs can be exploited for accurate prediction of brain age by increasing the model complexity and optimising the parameters space. These models were trained using a stratified 10-fold cross- validation process (stratified according to the feature used for the train / test split using StratifiedKFold and GridSearchCV of the Scikit-learn package in Python). The hyperparameter search space for LASSO and Ridge included {'alpha': {1.0 x 10A{-n}, 5.0 x 10A{-n}} for n = -4 to 3}, while Elastic Net utilised {'alpha': {1.0 x 10A{-n}, 5.0 x 10A{-n}} for n = -4 to 3, 'I1_ratio': {n x 10A{-1 }} for n = 1 to 9}. Additionally, we included an Ordinary Least Squares (OLS) method, using p-values from t-tests to recursively remove features with p-values greater than 0.05 until all coefficients p-values would fall below such a threshold.
[0109]
[0109] To further broaden our exploration of non-linear models, we employed two AutoML libraries: TPOT (Olson & Moore, 2016) and FLAML (Wang et al., 2021). These tools automate the exploration of diverse models, optimise hyperparameters, and construct ML pipelines, making them invaluable for developing high-performing models efficiently with minimal manual intervention. TPOT, a genetic programming-based AutoML library, evolves ML pipelines by considering a variety of regression models, including decision trees, random forests, gradient boosting, LASSO, Ridge, and support vector regression (SVR). It automates the selection of preprocessing steps, model choice, and hyperparameters, evolving these over multiple generations to maximise performance. On the other hand, FLAML is a lightweight AutoML library optimised for fast and efficient model tuning. In regression tasks, it leverages models such as LightGBM, random forests, extra trees, and linear models. The focus of FLAML on quick convergence and low computational overhead makes it particularly suitable for time-sensitive or resource-constrained projects. For FLAML, we set a time limit of 5 minutes to identify the bestperforming ML pipeline, while for TPOT, we used 5 generations with a population size of 10.
[0110]
[0110] Although AutoML libraries like TPOT and FLAML already include algorithms such as XGBoost and LightGBM, given their strong performance in regression tasks with tabular data (Shwartz-Ziv & Amitai, 2022), we also provided an additional training pathway for these models using Hyperopt (Bergstra et al., 2015) for hyperparameter tuning. Before training, we have also conducted several preprocessing steps, including removing constant features and applying feature selection via pairwise correlation, with a threshold of 0.80, keeping only the features most correlated with chronological age. Additionally, to address potential age bias in the dataset, we applied a resampling step in some models, binning the data into 5-year intervals. This strategy ensured a more uniform chronological age distribution across the age bins, helping to reduce over- or under-predictions in age groups with fewer samples.
[0111]
[0111] Deep Learning Approaches. DL has shown promise in neuroimaging tasks such as brain age prediction and modelling (Cole 2017; Kawahara et al., 2017) but challenges remain, particularly due to the high memory demands of 3D neuroimaging data, which make it difficult to apply successful 2D models like those used in ImageNet classification (Krizhevsky et al., 2017; Simonyan & Zisserman, 2014). Strategies like downsampling, patch-based methods (Kamnitsas 2017; Liu 2018), or the use of 2D slices (Bashyam 2020; Lin 2018) address this at the expense of trade-offs in performance. Furthermore, DL models often require large datasets as small size neuroimaging datasets can lead to overfitting (Raghu 2019; Russakovsky 2015). Brain age prediction shares these challenges and while various methods, including DL approaches (Cole 2017; Feng et al., 2020; Kolbeinsson et al., 2020), have been explored, achieving unbiased, high- performing models remains challenging due to biases toward the group mean (Smith 2019).
[0112]
[0112] With this in mind we employ three different DL models for brain age prediction using 3D T1w MRI images: a Simple Fully Convolutional Network, a 3D ResNet-18, and a custom 3D CNN model.
[0113]
[0113] The Simple Fully Convolutional Network (SFCN) is inspired by the VGGNet architecture (Simonyan and Zisserman, 2014) and employs a fully convolutional design (Long et al., 2015). This model has a relatively low parameter count (~3 million), which reduces computational complexity and memory requirements. The SFCN uses soft labels to estimate brain age, and the final age prediction is obtained by computing a weighted average over predicted age bins. We optimise the model using the Kullback-Leibler Divergence (KLDiv) loss function and the Stochastic Gradient Descent (SGD) optimizer.
[0114]
[0114] Following the SFCN, we utilise a 3D ResNet-18, a residual learning-based architecture originally developed for image classification tasks (He et al., 2016). The model comprises four main residual blocks, each incorporating shortcut connections to facilitate gradient flow and efficient training. The 3D ResNet-18 architecture has a moderate parameter count (~33.2 million) and is also optimised using the KLDiv loss function and SGD optimizer. Like the SFCN, this model predicts brain age by using soft labels.
[0115]
[0115] Lastly, we introduce a custom 3D CNN model (Custom3DCNN), designed specifically for regression-based brain age prediction. This model directly estimates chronological age without relying on classification-based soft labels. To improve accuracy across different age ranges, the Custom3DCNN uses a custom Adaptive Weighted Dynamic (AWD) loss function. The model is optimised using the Adam optimizer and aims to achieve balanced performance across all age groups.
[0116]
[0116] To enhance model generalisation and reduce overfitting, we applied data augmentation techniques such as random shifts and random mirroring during training. All models were trained on approximately 3,100 MRI scans and validated with another 400 MRI scans of different subjects, to avoid data leakage. They were trained for 200 epochs, with a batch size of either 4 or 8, using dynamic learning rate scheduling.
[0117]
[0117] The trained brain age model S110 outputs a brain age according to the volumetric features derived from the brain image. When the brain age model is a deep learning (DL) model, brain age may be output without first analysing volumetric features. When the brain age model is a DL model, the DL model may be trained to directly output a brain age value.
[0118]
[0118] While the brain age model may at least use the volumetric features, other data may also be utilised. For example, the brain age model may utilise other health information, such as a cognitive assessment test like the Montreal Cognitive Assessment (MoCA), blood test results, blood pressure, body mass index (BMI), information on proteins in the blood or brain and / or any other suitable health information. There may be different brain age models for male and female patients. Each of these brain age models may be trained on male or female training data only. Male and female brains may be of different size. Therefore, obtaining brain age using different models may make sense. The models may be trained and used in the same way as described herein for a general brain age model. Only the training data may only comprise male and female patients to train a model for male patients and a model for female patients.
[0119]
[0119] When outputting a brain age according to the volumetric features, a prediction of the brain age model may further be corrected to remove any model bias S112. Such bias correction may take place because there has been shown to be a negative correlation between chronological age and the difference in chronological age and predicted brain age. Thus, there may be systematic bias in the trained brain age model.
[0120]
[0120] To correct such bias, a model bias may be determined during training and a prediction by the brain age model may be modified according to: where ycorris the corrected predicted brain age, ypredis the predicted brain age resulting from the linear regression model, a and b are the coefficient / slope and intercept of the linear regression model and e represents a random fluctuation around the expected value. In other words, correcting a bias may comprise subtracting an intercept of the regression line from the initial predicted brain age and then dividing by a slope of the regression line. Each of a and b may be determined during training as described below.
[0121]
[0121] In step S114, the determined brain age may be used to classify the brain image into a category according to a current dementia status. For example, one category may be “cognitive normal (CN)” or “low risk” for brain images that show no sign of dementia / Alzheimer’s disease and are thus considered cognitively normal. Another category may be “mild cognitive impairment (MCI)” or “moderate risk”. MCI may be considered as a stage between expected decline in memory and thinking due to age and cognitive decline attributable to dementia. Thus, MCI may be considered as a cognitive state in which a person starts to have issues with memory orthinking. Usually, MCI is defined as not causing difficulties that interfere with everyday tasks. Yet another category may be “Dementia” or “high risk”, i.e. a cognitive state / performance associated with a decline of brain functioning. Decline in cognitive performance associated with dementia is generally considered serious enough to interfere with everyday tasks. Dementia may, for example, be due to Alzheimer’s disease and / or other neurodegenerative diseases. “Not dementia” may be a category including both “cognitive normal” and “mild cognitive impairment (MCI)”. “Not normal” may be a category including both “MCI” and “dementia”.
[0122]
[0122] Classifying the image according to a dementia status may be done using a pre-trained risk prediction model which may also be termed a classification model. The risk prediction model may be trained on the same ADNI data as the brain age model. Brain age is input into the risk prediction model and the cognitive status / dementia risk labels from the ADNI database are used as ground truth labels in training. While the brain age model may only be trained on cognitively normal data, the risk prediction model is trained on data falling into all risk groups, i.e. low, medium, high dementia risk I CN, MCI, Dementia. That is, the risk prediction model received a brain age from the brain age model and is then trained using the corresponding dementia risk label taken from the training dataset for that subject.
[0123]
[0123] The risk prediction model may, for example, be a logistic regression model. Other suitable risk prediction models, such as a Cox proportional hazards models or any other suitable regression and / or binary classification model may be used to predict a risk. Thus, the risk prediction model may output a probability that a brain image / user falls into either of “not dementia” or “dementia”. Additionally or alternatively, the risk prediction model may output a probability that the user falls into either of “MCI” or “dementia”, either of “CN” or “MCI”, or either of “CN” or “not normal”. For example, a logistic regression model may model the odds of an event and perform binary classification to determine a risk for each of the above categories.
[0124]
[0124] Different risk prediction models may be trained for male and female patients. This can be done by using different brain age models for male and female patients and then training the risk prediction model using the output from such a gender / sex specific brain age model. Alternatively, the brain age model may be the same model for both male and female patients and only the risk prediction model is trained on male and female patient data. Brain age may correlate with risk of dementia in a different way for male and female patients. Thus, this may be corrected using separate risk prediction models for male and female patients.
[0125]
[0125] In addition to classifying the brain images, the risk prediction model may also be used to determine a probability of a user developing dementia later in time S116. Thus, the risk prediction model may stratify users into risk groups based on an estimated risk of developing dementia or mild cognitive impairment and this may help inform therapeutic decision making. Additionally, stratification into risk groups may help inform recommendations for lifestyle changes. Thus, a user may be told that they have an increased risk of developing dementia or mild cognitive impairment and lifestyle changes to mitigate this risk may be proposed.
[0126]
[0126] Stratification into risk groups, or prediction of a percentage risk of developing dementia, may be given for time points 0 months, 24 months, 60 months and / or 120 months after the brain image is taken. Risk may also be calculated for any other suitable point in time. That is, the risk prediction model may be used to discriminate transition from one disease state into another, e.g. from CN to dementia, from CN to MCI, or from MCI to either CN or dementia.
[0127]
[0127] To learn dementia risk over time, the risk prediction model may also be trained on changes in MRI images and consequently, measured changes in volumetric features over time. That is, risk prediction may take into account how a specific patient’s brain age has changed over time when making a dementia risk prediction.
[0128]
[0128] When the risk prediction model is a logistic regression model, logistic regression may, for example, be implemented using scikit-learn’s Logistic Regression function (https: / / scikil- learn.org / siable / modules / generaied / sklearn.linear model.LogisticRegression.html) . This function may be trained using training data and the trained model may then be applied to an unseen brain image as described above.
[0129]
[0129] The method as described with reference to Figure 1 above may further comprise determining a treatment for a subject based on the brain image classification and / or risk prediction. The treatment may, for example, include medication or other treatments.
[0130]
[0130] When the dementia risk is high, the treatment may be selected from: AAV-hTERT, Emtricitabine, LX1001 , Vorinostat, AL003, Bacillus Calmette-Guerin (BCS), JNJ-40346527, Salsalate, Sirolimus, XPro1595, BEY2153, E2814, Lu AF87908, LY3372993, Trehalose, Grape seed polyphenolic extract, Nicotinamide riboside, Allopregnanolone, Dasatinib + Quercetin, Efavirenz, Mesenchymal Stem Cells, NNI-362, REM0046127, SNK01 , Dabigatran, MK-1942 + Donepezil, MK-4334, 3TC, AL002, ALZT-OP1 (cromolyn and ibuprofen), Bacillus Calmette- Guerin (BCG), Daratumumab, GV1001 , Lenalidomide, Montelukast, Tacrolimus, VX-745, ABBV- 8E12, ABvac40, ACI-35.030 + JACI-35.054, ALZ-801 , APH-1 105, BIIB092, IGNIS MAPTRx, JNJ- 63733657, LY3303560, Meg a natural- Az Grapeseed Extract, NewGam 10% IVIG, Nilotinib, Posiphen, PQ912, RO7126209, Semorinemab, TEP, Benfotiamine, Dapagliflozin, L-Serine Gummy, Liraglutide, Metabolic Cofactor Supplementation, Nicotinamide, Pepinemab, T3D-959, AMX0035, AstroStem and Donepezil, CB-AC-02, CORT108297, HB-adMSCs, Human Mesenchymal Stem Cells, Leuprorelin, Mesenchymal Stem Cells, PTI-125, Rapamycin, Regulatory T cells, S-equol, Sovateltide, T-817MA, Valacyclovir, AD-35, AD-35 60mg, ATH-1017, BPN14770, Bromocriptine, Bryostatin 1 , CT1812, DAOI-A, intermittent Theta Burst Stimulation (iTBS), Levetiracetam, MemorEM, Nicotine Transdermal Patch, SAGE-718, Transcranial Alternating Current Stimulation (tACS), AR1001 , Perindopril|Telmisartan, NE3107, Aducanumab, Donanemab, Gantenerumab, Lecanemab, Trx0237, Extended-release metformin, Ginkgo biloba, GV-971 , Tricaprilin, AGB101 , ANAVEX2-73, Donepezil, Guanfacine, Octohydroaminoacridine, succinate, Troriluzole, Renew NCP-5, BPDO-1603, COR388. Additionally or alternatively, the treatment may be selected from any of acetylcholinesterase (AChE) inhibitors, memantine, antipsychotic medicines, antidepressants, cognitive stimulation therapy (CST), cognitive rehabilitation, reminiscence work (talking about things and events from the patient’s past) and / or life story work (compilation of photos, notes and keepsakes from the patient’s childhood to the present day). Other suitable treatment options may be used. Additionally or alternatively, lifestyle changes may be suggested to a patient. These may, for example, include changes in diet and / or exercise pattern, as well as cognitive exercises.
[0131]
[0131] Some treatments are only effective when given early on, while dementia is still developing. In this case, only patients in the low or medium dementia risk group may be identified as suitable forthis treatment. Such treatments may, for example, include Icosapent ethyl, Memantine, Omega 3 treatment, Losartan, an Active NIR-PBM device, Gantenerumab, Crenezumab, Deferiprone, Omega 3 PUFA, Solanezumab, Telmisartan, AGB101 and / or DHA. Additionally, or alternatively to treatment, patients may be identified for other interventions, such as, for example, dietary changes, increasing exercise and / or reducing stress. Lifestyle changes may be particularly helpful if a patient is classified as having a medium dementia risk, as this may then slow down progression to a high dementia risk. For example, a patient may be asked about their current lifestyle and treatment options may be recommended on the basis of that current lifestyle. For example, when a patient responds that they are not usually physically active, a recommendation may be going for walks or incorporating other forms of physical exercise into their routine.
[0132]
[0132] When a treatment option is applied, the method may further comprise generating training data comprising associated response data (e.g. outcome from a particular treatment) for each of a plurality of subjects. This training data may then be used to train a treatment selection model that recommends a particular treatment or other measure (such as particular diet). Such a method may be used to create a predictive signature of treatment response, to help stratify patients for the most appropriate treatment regime. The method may thus further comprise recommending a treatment and / or treating a patient using the recommended treatment.
[0133]
[0133] Figure 2 is a flowchart showing steps of the method at training time. In a first step S200, a set of training brain images and corresponding labels is obtained. The training images may be MRI images of brains as described above, for example, MRI images and corresponding patient data from the ADNI study. That is, the images may be T1 and T2-weighted structural MRI images, MRI images including diffusion tensor imaging and / or MRI images using functional magnetic resonance imaging. Any other suitable imaging techniques, such as computer tomography (CT) or positron emission tomography (PET) may be used. When the ADNI database is used for training, ADNI study participants are stratified into CN, MCI and Dementia groups. CN, MCI and Dementia groups may correspond to low, medium and high risk of dementia respectively. This classification is done based on the corresponding ADNI labels. In particular, this classification is done based on the “DX” label. That is, ADNI already provides a classification according to whether a subject suffers from a neurodegenerative disease or is otherwise cognitively impaired. These classifications have been translated into the classifications used in the present application. The model may only be trained on study participants that have been determined to be “cognitively normal”. The brain age model is trained to predict a subject’s chronological age from a corresponding MRI image. For cognitively normal subject, chronological age is taken to correspond to biological / brain age. Thus, for the brain age models, the subject’s chronological age are used to label the training brain images.
[0134]
[0134] The brain images may then be partitioned S202, the volume of each partition may be determined next at S204 and the volume of each partition may be normalised to obtain volumetric features S206. These steps may be carried out as described above in relation to Figure 1 . The same techniques should be used at inference time as were used at training time. While partitioning the image S202 and finding the volume of each partition S204 may be separate steps, they may equally be performed in one step, depending on how the image is being partitioned.
[0135]
[0135] The brain age model may, for example, be a linear regression model, such as scikit-learn’s LinearRegression model. As described above with reference to step S106 of Figure 1 , any data input into the model may first be normalised at step S206. This normalisation may be based on the available training data. Thus, not only are the parameters of the brain age linear regression model themselves determined during the training process, but normalisation parameters are also determined while training the brain age model.
[0136]
[0136] The brain age model may then be trained at step S208 using the volumetric features that are input into the model. Each brain image in the training data is labelled with a chronological age of the user whose brain is depicted in the image. This label serves as ground truth that the model uses to learn how volumetric features correspond to brain age. Because this is the case, the model is trained on cognitively normal data only. For cognitively normal data, the chronological age is assumed to correspond to the biological age. Any suitable training method may be used to minimise the difference between the predicted value and the labelled value for a particular image.
[0137]
[0137] All volumetric features may be input into the brain age model or only some volumetric features may be input into the brain age model. Once the brain age model is trained to predict a user’s brain age based on a certain set of volumetric features, the same set of volumetric features is used at inference time. That is, the training and test data should relate to the same information. This may then influence which volumetric features of a brain image are obtained and input into the brain age model at inference time.
[0138]
[0138] When selecting which volumetric features are input into the brain age model, some cortical structures and / or substructures may be more relevant than others. Examples of structures and / or substructures that may be input into the brain age model are: hippocampus, ventricles, fusiform, medial temporal lobe, entorhinal cortex, fusiform gyrus and / or the whole brain. Other suitable volumetric features may be selected for training the brain age model.
[0139]
[0139] In age prediction models there is a systematic bias issue that must be addressed to avoid overestimating the ages of younger individuals and underestimating the ones of older individuals. Assuming a linear relationship between chronological age (CA) and brain age (BA), BA = a * CA + b ,
[0140]
[0140] the slope (a) and intercept (b) can be estimated via linear regression from the training set. Clearly, a perfect prediction would correspond to a = 1 and b = 0, i.e. BA = CA.
[0141]
[0141] Cole’s method for age-bias correction (Cole et al., 2018; Peng et al., 2021 ; Smith et al., 2019) is based on the definition of corrected brain age (BAc) through the inversion of the previous relation:
[0142] BAC= BA - b) / a .
[0143] This approach increases the variance of the brain age prediction by a factor of 1 / a but does not require information on the chronological age (only the slope and intercept derived from the training set). Lange's method (de Lange et al., 2020), on the other hand, includes CA in the correction, by defining the corrected brain age as
[0144] BAC= BA - a * CA + b~) + CA and thus maintaining the same amount of variation for predicted age before and after correction.
[0145]
[0142] If parameterised in terms of PAD (the difference between BA and CA), the corrected PAD (PADc) can be derived for both Cole’s and Lange’s models from the above expressions via
[0146] PADC= BAC- CA.
[0147]
[0143] Finally, Zhang's age-level correction (Zhang et al., 2023) consists in the standardisation of PAD leading to a variance of the corrected brain age of order 1 . Let us consider a cohort with uniform CA = x, and let PADx be the PAD prediction for this cohort with mean px and standard deviation ox. Then, according to Zhang, the age-level corrected PAD for sample i, PADixc, is given by
[0148] PADixc— (PADixix Xwhich leads to
[0149] Ex[PADXC] = 0.
[0150]
[0144] We have incorporated all three types of corrections to ensure comprehensive evaluation and robust results across different scenarios.
[0151]
[0145] Here we provide some further details about the impact of the age-bias correction to the statistics of brain age prediction. Let us define the brain age of a population with fixed chronological age CA as a random variable with expectation <BA> and standard deviation o. We can write it as:
[0152] BA = < BA > +E , where s is a random variable centred in zero with the same standard deviation o. If we assume that o does not depend on CA but is instead a constant, then the relationship between BA and CA becomes:
[0153] BA = a * CA + b + E .
[0154]
[0146] Following the same approach, we may assume that the corrected brain age BAc is:
[0155] BAC= < BAc > TC where ECis analogous to e but with a different standard deviation oc.
[0156]
[0147] The Cole correction can therefore be rewritten as:
[0157] BAC= BA — b) / a = (< BA > —b) / a + s / a = CA + s / a Therefore, if a < 0, we will have that EC= e / a > e, which means that the standard deviation crcincreases by a value 1 / a. On the other hand, the Lange definition reads:
[0158] BAC= BA - (a * CA + b) + CA
[0159] — < BA > + E — < BA > + CA
[0160] — CA + E
[0161]
[0148] Therefore, the standard deviation remains the same before and after the correction, i.e., crc= o. Finally, in the case of Zhang, each of the random variables BA is standardised with respect to its own mean <BA> and standard deviation, and then shifted to be centred into the corresponding CA:
[0162] BAC= (BA — < BA >) / a + CA.
[0163]
[0149] The distribution of brain age representing the variability of the healthy population is then artificially constrained to fit into a unitary standard deviation crc=1 . Since the transformation does not affect the rank of each value with respect to the entire population, this approach can still be employed as a biomarker for predicting a significant deviation from the healthy population, but the associated PAD would be of order 1. Indeed, having PAD of 3 years corresponds to being 3crcfrom the expectation. Although few people would feel worried about appearing 3 years older, in terms of brain this would be a very significant deviation from a biological perspective. Note that parameters a and b are computed in the training set, and when we use them to correct the test set we may observe deviations from the theoretical expectation for both parameters <BAC> and CTC.
[0164]
[0150] Thus, there is a determination of any model bias dy at step S210. Bias correction may be performed according to the equations above or according to the following equations: yPrea = ax + b + e where ytrueis the labelled age; ycorris the corrected brain age; ypredis the initial prediction of the linear regression model, a and b are the coefficient / slope and intercept of the linear regression model and e represents a random fluctuation around the expected value. The magnitude and type of random fluctuation e may be determined while training the brain age model. In particular, the magnitude and type of random fluctuation may be determined based on a difference between the predicted and true brain age values. This ensures that e accurately represents random deviations from the brain age model’s predictions. For example, the e may be drawn from a Gaussian distribution with mean zero and unit variance. However, e may be drawn from any other suitable distribution with any suitable mean and variance, such as a uniform distribution. For example, e may be generated using a pseudo-random number generator as those available in software packages such as Numpy (for Python). Then, values for e may be limited to a particular interval, for example, between -1 and 1 . Any other suitable interval may be used.
[0165]
[0151] Bias correction ensures that there is unitary correlation between the predictions made by the brain age model and the chronological age labels that correspond to the training data. That is, any systematic bias that appears in the brain age model’s predictions is identified and corrected by applying the above equations. Using the parameter e takes into account inevitable random fluctuations around the correct value and ensures that these do not influence the bias correction process.
[0166]
[0152] Next, bias corrected brain age predictions ycorrbased on the initial predictions ypredfor the training data are obtained S212 using for example: where a and b and e are defined above. These bias corrected predictions are then input into a risk prediction model. This ensures that the risk prediction model is trained on previous outputs of the pipeline and can thus work effectively at inference time. The bias corrected age predictions and labels from the training data set are then used to train a risk prediction model S214. As described with respect to Figure 1 above, the risk prediction model may, for example, be a logistic regression model.
[0167]
[0153] Alternatively, age bias may be corrected in the linear model directly by adding penalty terms that counteract any model bias. Advantageously, no bias correction is then necessary, but bias is already corrected in the linear model’s prediction.
[0168]
[0154] Depending on the risk prediction task, the risk prediction model may output different classification results and / or risk predictions. These outputs may, for example, include a brain image classification into different categories, such as: “cognitive normal (CN)” which may also be termed “low or no risk”, “mild cognitive impairment (MCI)” which may also be termed “moderate risk” or “Dementia” which may also be termed “high risk”. The categories may be grouped as defined with respect to Figure 1 above.
[0169]
[0155] Again, as described with respect to Figure 1 above classification into different categories may also include determining a probability of the brain image falling into the various categories. The category with the higher / highest probability is the category which is assigned to the brain image. In addition to classifying the brain images (i.e. the current risk), the risk prediction model may also be trained to determine a probability of a user developing dementia later in time. Thus, the risk prediction model may be trained to stratify users into future risk groups based on estimated risk of developing dementia or mild cognitive impairment. Stratification into risk groups, or prediction of a percentage risk of developing dementia, may be given for time points 0 months, 24 months, 60 months and / or 120 months after the brain image is taken. Risk may also be calculated for any other suitable point in time. That is, the risk prediction model may be trained to discriminate transition from one disease state into another, e.g. from CN to dementia, from CN to MCI, or from MCI to either CN or dementia.
[0170]
[0156] The risk prediction model may be trained on the same ADNI data as the brain age model. Brain age is input into the risk prediction model and the cognitive status / dementia risk labels from the ADNI database are used as ground truth labels in training. While the brain age model may only be trained on cognitively normal data, the risk prediction model is trained on data falling into all risk groups, i.e. low, medium, high dementia risk I CN, MCI, Dementia. That is, the risk prediction model received a brain age from the brain age model and is then trained using the corresponding dementia risk label taken from the training dataset for that subject.
[0171]
[0157] When the risk prediction model is a logistic regression model, logistic regression may, for example, be implemented using scikit-learn’s Logistic Regression function (hltps: / / scikit- learn.org / stabte / modules / generated / skleam.linear model.LogisticRegression.html). This logistic regression function may be trained using suitable training data.
[0172]
[0158] The risk prediction model may comprise several submodels to perform each of the classification and / or stratification tasks described above. For example, one logistic regression model may be trained to determine if a brain image either falls into the CN or the Dementia category, while another logistic regression model may be trained to determine if a brain image falls into the MCI or Dementia category. In this way, the risk prediction model may comprise a submodel for each possible binary classification between all of the categories mentioned above. Thus, a risk prediction submodel may be trained for each possible combination of two categories mentioned above. This is, for example, the case for a logistic regression model which performs binary classification. Other risk prediction models may be used and these may be capable of classifying images into more than two categories. Accordingly, the number of risk prediction submodels may be adapted to the model type used.
[0173]
[0159] Several studies have explored deep learning approaches for brain age prediction using MRI data from the UKBB. Peng et al. (2021) developed a lightweight architecture, the Simple Fully Convolutional Network (SFCN), using 12,949 subjects-not necessarily healthy-for training and achieved a Mean Absolute Error (MAE) of 2.14 years. Dinsdale et al. (2021) applied a 3D CNN using T1w MRI data, trained on 12,802 subjects, to investigate the effect of linear and nonlinear registration on the model’s predictions, and demonstrate correlations between age prediction deltas and clinical measurements. Their best performing model achieved, a MAE of 2.71 years, on the nonlinearly warped data but performance worsened when considering only healthy subjects for model training. Kolbeinsson et al. (2020) trained a CNN to predict brain age using a set of 3067 healthy UKBB participants, achieving an MAE of 3.42 years, and explored the association between brain age differences and over 1 ,400 clinical and lifestyle factors. Ning et al. (2021) trained a CNN on 16,998 subjects-not all healthy-and compared it with a linear regression model, achieving an MAE of around 3.5 years, while identifying genetic factors linked to brain ageing through a genome-wide association study. Dartora et al. (2023), explored brain clocks based on a customised ResNet-architecture with 3D convolutional layers under various data splitting strategies for training, validation, and testing. These models were trained combining UKBB data with various other databases and model evaluation was confined to age prediction accuracy without examining brain-age deltas for differentiating clinical groups. The authors highlight the benefits of cross-validation approaches (e.g., CNN2, CNN4) over hold-out methods (e.g., CNN1) when assessing external datasets such as AddNeuroMed and J-ADNI. Nevertheless, only Lange’s age-bias correction is explored although it is not thoroughly discussed, and in most cases, MAEs on external datasets (e.g., AddNeuroMed, J-ADNI) exceeded those for ADNI, AIBL, GENIC, and UKBB by up to three years (except for CNN3, where AddNeuroMed and J-ADNI were included during cross-validation).
[0174]
[0160] While nearly all previously mentioned studies focused exclusively on T1w MRI scans, some, like (Cole, 2020), also explored brain age prediction from other modalities, such as FLAIR, T2, Diffusion MRI, Task / Resting-state fMRI, and various multimodal combinations. Despite this broader approach, models built solely on T1w MRIs consistently achieved the lowest MAE, slightly above 4.140 during validation, and an R-squared value of 0.468 with respect to any other single modality brain age predictor. When all modalities were used together, these metrics could be improved, reducing the MAE to 3.515 and increasing the R-squared to 0.618.
[0175]
[0161] Additional risk prediction submodels may be trained to stratify users / brain images into risk groups or predict a percentage risk of developing dementia within a given time frame. Again, this risk prediction may be performed by a logistic regression model.
[0176]
[0162] Thus, any reference to a risk prediction model may mean either a single risk prediction model, or a risk prediction model comprising several risk prediction submodels each performing a binary classification and / or stratification task as described above.
[0177]
[0163] T1w data was collected from three widely used MRI databases: UKBB, NACO, and ADNI. One of the main objectives of this work is to demonstrate that models trained solely on UKBB data can generalise well to other datasets. Therefore, the UKBB is the only dataset used to train, validate, and internally test the models, while NACC and ADNI will serve as external test datasets to evaluate the developed brain clocks. The previously described methods and models have been tested using two databases comprising appropriate data for training and testing the models.
[0178]
[0164] The National Alzheimer’s Coordinating Center (NACC) is a central data repository for the National Institute on Aging’s Alzheimer’s Disease Research Centers (ADRC) program. The NACC holds one of the largest and oldest Alzheimer’s disease datasets, built in collaboration with over 42 ADRCs across the United States. The dataset includes comprehensive information on demographics, family history, medications, medical conditions, clinical symptoms, diagnosis, and neuropsychological test results across four cognitive domains.
[0179]
[0165] One of these datasets is the Alzheimer’s Disease Neuroimaging Initiative (ADNI) database The ADNI was launched in 2003 as a public-private partnership, led by Principal Investigator Michael W. Weiner, MD. The primary goal of ADNI has been to test whether serial magnetic resonance imaging (MRI), positron emission tomography (PET), other biological markers, and clinical and neuropsychological assessment can be combined to measure the progression of mild cognitive impairment (MCI) and early Alzheimer’s disease (AD). Data obtained from the ADNI study was obtained, in particular the data was obtained from the table ADNIMERGE, in which a selected group of variables has been already harmonised between different waves of collections of the ADNI study.
[0180]
[0166] Contrary to the UKBB, these databases collect MRI data obtained from a wide range of manufacturers and MRI machines, including Siemens, GE, and Philips models like Skyra, Verio, Biograph mMR, Prisma, amongst others. These external datasets provide diverse data (with different levels of curation), making them ideal for evaluating the robustness of brain clocks developed using UKBB to other clinically relevant scenarios.
[0181]
[0167] ADNI study participants are stratified into CN, MCI and Dementia groups. This classification is done based on the corresponding ADNI labels. In particular, this classification is done based on the “DX” label. That is, ADNI already provides a classification according to whether a subject suffers from a neurodegenerative disease or is otherwise cognitively impaired. These classifications have been translated into the classifications used in the present application.
[0182]
[0168] Features from the ADNI study that have been used to test the methods described in the present application are: ‘AGE’ (i.e. age of study participant), 'Ventricles_bl (volume of ventricles)', 'Hippocampus_br (volume of hippocampus), 'WholeBrain_bl' (volume of the whole brain), 'Entorhinal_bl' (volume of the entorhinal cortex), 'Fusiform_bl' (volume of the fusiform gyrus), 'MidTemp_bl' (volume of the medial temporal lobe). I.e. the age of study participants and volume of different cortical structures and substructures are being used.
[0183]
[0169] The models trained and internally validated with the ADNI dataset are then tested on the independent UK Biobank (UKB) dataset. One of the key datasets driving brain age prediction research is the UK Biobank (UKBB) longitudinal database (Biobank, 2007), which contains an extensive collection of MRI scans and associated health data. Papers leveraging this dataset, such as in the studies of Cole and colleagues (Cole et al., 2017; Cole & Franke, 2017; Cole et al., 2018; Cole, 2020; Zhang et al., 2023), have significantly advanced the development of mathematical models that can predict an individual's brain age based on imaging data.
[0184]
[0170] UK Biobank is a large-scale biomedical database and research resource containing genetic, lifestyle and health information from half a million UK participants. UK Biobank’s database includes blood samples, heart and brain scans and genetic data of the 500,000 volunteer participants. UK Biobank recruited 500,000 people aged between 40-69 years in 2006-2010 from across the UK. With their consent, they provided detailed information about their lifestyle, physical measures and had blood, urine and saliva samples collected and stored for future analysis.
[0185]
[0171] The UKBB is a large prospective cohort study involving 500,000 individuals aged 40 to 69 from across the UK recruited between 2006 and 2010. Participants agreed to have a wide range of data collected for baseline health assessment and to have their health followed up through linkage to their health-related records. Data collected from these individuals includes blood, urine, and saliva samples, physical measurements, and MRI scans, along with extensive health and lifestyle questionnaires. Since its initial assessment the UKBB database has grown significantly and is currently recognised internationally as a powerful resource providing valuable long-term health data to support research into prevention, diagnosis and treatment of a wide range of diseases. All T1w MRI scans in the UKBB dataset are acquired using the same protocol and a Siemens Skyra 3T scanner.
[0186]
[0172] The following MRI fields from UKB are selected and combined to obtain the equivalent of the ADNI MRI features, both obtained by using Freesurfer cortical reconstruction and volumetric segmentation with Desikan-Killiany atlas (except the “whole brain” in UKB):
[0187] 'Ventricles_bl' = '26523-2.0' + '26524-2.0' + '26525-2.0' + '26585-2.0' + '26554-2.0'
[0188] 'Hippocampus_bl' = '26562-2.0' + '26593-2.0'
[0189] 'WholeBrain_bl' = '26514-2.0'
[0190] 'Entorhinal_bl' = '26793-2.0' + '26894-2.0'
[0191] 'Fusiform_bl' = '26794-2.0' + '26895-2.0'
[0192] 'MidTemp_bl' = '26802-2.0' + '26903-2.0'
[0193] The healthy cohort of UKB has been selected as described in Cole 2020 (Cole, James H. "Multimodality neuroimaging brain-age in UK biobank: relationship to biomedical, lifestyle, and cognitive factors." Neurobiology of aging 92 (2020): 34-42.), the MCI cohort is absent, while the demented cohort has been composed by the union of the subjects which registered a first occurrence for Alzheimer’s disease and / or dementia due to Alzheimer’s disease, UKB fields: '131036-0.0' and '130836-0.0'.
[0194]
[0173] We built our UK Biobank healthy cohort by considering the most restrictive criteria outlined in (Cole, 2020) that we will refer to as “healthy_cole”. Specifically, we examined the medical records of individuals for which T1w MRI scans were available and we excluded from the healthy cohort any subject having an ICD-10 diagnosis (category #41270), a self reported long-standing illness disability or infirmity (field #2188), self-reported diabetes (field #2443), stroke history (field #4056) and those not having good or excellent self-reported health (field #2178). Following this definition a total of 4,167 images were available (2,227 from females and 1 ,940 from males), representing 3,714 individuals with an average age of 64.97 years.
[0195]
[0174] As one of the goals of this work is to also assess the utility of brain age models built on the UKBB database for stratifying patients according to different conditions of interest, we also considered for the analysis individuals for which a date of diagnosis was available (either before or after enrolment in the UKBB study) for the following neurodegenerative conditions: Alzheimer Dementia (field #130836), Vascular Dementia (field #130838), other causes or Unspecified Dementia (fields #130840 and #130842), Huntington’s disease (field #131012), Parkinson’s disease (field #131022 and 131024), Alzheimer’s disease (field #131036), Multiple Sclerosis (field #131042), and other spinal cord disorders (field #131118). While all these patients (and the corresponding MRIs) are labelled as non-healthy according to Cole’s criteria and hence not used in model training but only in assessing model utility, we wanted to also investigate the effect of relaxing the health definition criteria and distinguish between individuals that were already diagnosed at the time of MRI scan and the ones that were not. We hence introduced an alternative definition of healthy, namely the “healthy_orx” group, which was defined as the set of individuals belonging to the “healthy_cole” group as well as all images available from individuals diagnosed with one of the above-mentioned neurodegenerative conditions prior to their diagnosis. By doing so, individuals who developed a neurodegenerative condition in the future are not discarded but rather labelled as healthy at the time of scan.
[0196]
[0175] In the case of the NACC and ADNI datasets, the definition of healthy individuals was extracted from their corresponding cognitive status fields. The cognitive status was assigned during the visit (baseline and / or follow-up) through a clinical evaluation of the participants and of their scores in a set of cognitive and functional assessments, e.g. MoCa, MMSE, CDR, and / or others, when available. The cognitive assessment outcome was set to Cognitively Normal (CN), Subjective Memory Concern (SMC), Mild Cognitive Impairment (MCI) or with signs of Dementia. Note that in order for an individual to be clinically assessed as CN, neurodegenerative diseases such as the ones introduced above had to have been excluded (according to exclusion criteria that are specific to each of these studies). Therefore, by definition, these individuals belong to the healthy group according to the healthy_orx criterion. If an individual was assessed as cognitively normal at a certain visit date, the corresponding MRI was labelled as healthy. To mimic Cole’s criteria in the case of ADNI and NACC datasets, any subject that at any follow up visit changed from being CN to any other label, such as MCI or Dementia, was excluded from the healthy_cole category. The fact that different criteria are used to classify individuals from different databases as healthy is an inherent consequence of the fact that each clinical study is set up with a potentially different set of ground rules. While neurological conditions are typically absent in cognitively normal individuals in both ADNI and NACC, some comorbidities unrelated to neurological disorders might still be present. This might ultimately have an impact on the generalisation of the models over these individuals that should be further investigated but is out of the purpose of the present work. Indeed, considering that the training is performed on a stricter definition of healthy (where we all ICD-10 diagnoses are excluded), we expect results on ADNI and NACC individuals without comorbidities to be at least as good as the present techniques and hence view the outcomes as a “worse-case scenario”.
[0197]
[0176] For model training, validation, and testing of age prediction, we used only the MRIs connected to individuals classified as healthy according to one of the two definitions above (the impact of both definitions is considered in when considering results). However, MRIs from individuals with neurodegenerative diagnoses were retained for evaluating the brain clock’s utility in detecting specific conditions or more generally identifying individuals who are not belonging to the healthy class. The final number of MRI images available for training was 3,338 for healthy_cole and 3,497 for healthy_orx. For the age-prediction performance test, an additional 6,665 images from healthy_cole and 6,699 from healthy_orx were used, along with 7,273 images associated with the neurodegenerative conditions of interest for specific condition performance testing.
[0177] Finally, regardless of the considered database, for those labelled as healthy, we implemented a further stratification called “healthy_type”, which takes into account the length of time the individual remained healthy post-MRI extraction (when follow-up information was available). Clinical records were used to verify whether the individual was healthy at baseline (t=0) and if they remained healthy up to 5 or 10 years after MRI extraction.
[0198]
[0178] Figures 3a to 3f shows how the volume of different brain regions varies with age and dementia. Age is plotted against volume of different brain regions for both ADNI and UKB participants. Age in years is plotted on the x-axis, whereas volume in mm3is plotted on the y-axis. Figure 3a shows results for the whole brain, Figure 3b shows results for the ventricles, Figure 3c shows results for the entorhinal cortex, Figure 3d shows results for the fusiform gyrus, Figure 3e shows results for the medial temporal lobe, and Figure 3f shows results for the hippocampus. Round blue labels show the brain age for ADNI participants without dementia (i.e. CN or MCI), round yellow labels show the brain age for ADNI participants with dementia. Square purple labels show the brain age for UKB participants without dementia. Square red labels show the brain age for UKB participants with dementia. It can be seen that participants of the same age have lower volumes of the cortical (sub)structures if they also have a dementia diagnosis. Thus, there is a clear correlation between reduced brain volume and dementia. Additionally, there is a reduction of brain volume with age. It is also clear that ADNI and UKB data differ in the distribution of datapoints that are available. The UKB participants tend to be younger on average, whereas many ADNI participants are 80 and older. This is something that needs to be taken into account when training the model.
[0199]
[0179] The following table shows the statistics of the CN (normal) and Dementia populations in the UKB dataset:
[0200]
[0180] In the example used herein, the cognitively normal (CN) UKB cohort has been resamples to obtain an age distribution that is comparable to that of the ADNI cohort. The table above thus shows the original, unchanged distribution of the UKB cohort.
[0201]
[0181] The following table shows the results of a t-test between dem and CN population in UKB for the variables age and brainage: While CN and Dementia age distribution are comparable, the brain age distributions obtained from the corresponding MRI data show a significant difference.
[0202]
[0182] To train the brain age model and risk prediction model, the training data may be first pre- processed. Preprocessing may, for example, comprise resampling or boot strapping. For the UKB dataset, the healthy cohort has been under-sampled randomly to obtain an age distribution comparable to the cohort with dementia, as is present for the ADNI dataset. This ensures that both datasets have similar statistical properties and an appropriate age distribution.
[0203]
[0183] Normalisation of the data may be performed as described with reference to Figures 1 and 2 above. This means that the scikit-learn Minmax scaler is fit to the training set (the ADNI dataset) and then applied to the test set (UKB). No further processing of the UKB data to make the data comparable with the ADNI dataset has been done.
[0204]
[0184] The normalised data obtained from both datasets was then split into a training and test set. To prevent data leakage, each individual in the UKBB dataset was included exclusively in either the training or the test set. Those individuals with MRIs not labelled as healthy / cognitively normal were placed in the test set. For the remaining healthy individuals, we created a unique label combining the individual’s sex and the latest “healthy_type” recorded. We then applied a stratified split, maintaining an 80:20 ratio for training and testing, based on the previously defined unique labels. Both the brain age model and risk prediction model(s) were trained (for example as described in relation to Figure 2) and validated using the volumetric features of the ADNI dataset. The models were then tested on the independent UKB set (for example as described in relation to Figure 1).
[0205]
[0185] All preprocessing and model training were performed on AWS EC2 instances equipped with NVIDIA T4 Tesla GPUs, with configurations of 16 GB and 32 GB of RAM. This cloud-based infrastructure allowed for scalable and flexible computational power, facilitating the efficient handling of large datasets. The instances operated on Ubuntu and utilised FastSurfer v2.2.0, and additional ML frameworks such as TensorFlow 2, Scikit-learn and SciPy Python libraries. This setup reduced the time required for neuroimaging preprocessing and model training, enabling the analysis of extensive datasets within a reasonable timeframe.
[0206]
[0186] In our modelling framework, we employed a comprehensive array of ML and DL techniques, incorporating established methods from the literature and exploring several promising new approaches. For each of the models tested we considered certain setting variations: i) resampling the age categories in the training set (except when bootstrap is performed); ii) employing healthy_cole or healthy_orx definitions; iii) applying 3 types of age-bias correction (Cole, Lange, Zhang), or none. For each of the ML models we additionally considered the following variations: iv) using the volumetric features as left right component or merged (by their sum); v) considering biological sex as a feature itself, not considering it, and training a separate model for females and males.
[0207]
[0187] Given the broad age ranges present in the UKBB (46 to 82), NACC (18 to 96), and ADNI (50 to 97) datasets, and the fact that the UKBB data was used to train the brain clock models, we chose to focus on the 55 to 85 age bracket. This age range is particularly relevant for studying neurodegenerative conditions, as these diseases typically become more prevalent in mid to late life.
[0208]
[0188] Neurodegenerative conditions like Alzheimer's disease, Parkinson's disease, and other dementias are age-related, with risk increasing significantly after the age of 55. Alzheimer’s disease primarily affects individuals over 65, with prevalence nearly doubling every five years beyond that age (Alzheimer's Association, 2019; Prince et al., 2015). Similarly, the incidence of Parkinson’s disease increases sharply after the age of 60 (de Lau & Breteler, 2006). Limiting the age range to the interval between 55 and 85 allows us to develop a brain clock more accurately tailored to detecting the early stages of these conditions. Including younger individuals could introduce variability related to normal ageing, potentially obscuring the brain clock’s ability to identify abnormal brain ageing or early neurodegeneration. This age restriction ensures the clock’s sensitivity and specificity for neurodegenerative disease detection (Frisoni et al., 2010; Jack et al., 2018).
[0209]
[0189] Considering these selections, the final dataset comprises 2,501 female and 2,205 male MRI scans from the UKBB (mean age ± standard deviation: 64.93 ± 5.76 years for females and 65.81 ± 6.24 years for males), 3,487 female and 2,532 male scans from NACC (70.37 ± 7.38 years for females and 71.43 ± 7.52 years for males), and 3,119 female and 3,626 male scans from ADNI (72.91 ± 6.65 years for females and 74.28 ± 6.09 years for males).
[0210]
[0190] To provide a comprehensive view of performance, in this work models were assessed not only for their ability to predict the age of healthy individuals but also fortheir capacity to distinguish images belonging to healthy individuals from the ones of specific non healthy groups (e.g., individuals affected by dementia). To do so, we considered the list of metrics given in the table below, where metrics are categorised as age prediction, age bias, or disease prediction metrics.
[0191] Note that the wide array of metrics considered was evaluated on different sets depending on the particular scenario to be assessed (male / female, database or databases selected, full or bounded age range, healthy_cole or healthy_orx definition).
[0211]
[0192] In addition to the metrics in the table above, when analysing generalisation of performance we introduced two additional features: 'generalization_median' and 'generalization_max'. These features capture the median and maximum pairwise differences in the absolute values of MAE across all possible case pairs within each database, MRI machine manufacturer, MRI machine, and ethnicity. This approach enables us to monitor variations in model performance and identify potential issues with generalizability across different conditions or subgroups.
[0212]
[0193] A pipeline was developed to conduct appropriate statistical tests for comparing the performance metrics of the different model configurations outlined below and identifying statistically significant differences in average performance. Specifically, the Shapiro-Wilk test was employed to assess the normality of the data groups, and Levene’s test was used to evaluate the assumption of equal variances, both implemented using the SciPy library in Python. For comparisons between two groups, we applied the independent t-test when normality and equal variances were confirmed; Welch’s t-test in cases of unequal variances; and the Mann-Whitney U-test when neither normality nor equal variances were satisfied. A significance level of p < 0.05 was adopted, with Bonferroni correction applied to multiple Shapiro-Wilk tests. For comparisons involving more than two groups, we used ANOVA if normality and homogeneity of variances were upheld, and the Kruskal-Wallis test otherwise. Post-hoc analyses following ANOVA were conducted using Tukey’s HSD test, while Dunn’s test was applied after Kruskal-Wallis. The presented criteria is summarised in Table 2 below. All tests were executed using SciPy, with the exception of ANOVA and pairwise Tukey’s HSD tests, which were performed using the Python statsmodels library.
[0213]
[0194] The table below shows a summary of criteria used to select the appropriate statistical test, depending on the type of comparison considered and the assumptions satisfied.
[0214] 195] Beyond looking for the best performance on average for different model configurations, as part of our analysis we want to identify the most robust individual models developed. To achieve this, we implemented a Pareto-based model selection tool, which allows us to balance the tradeoffs across the following three main objectives: 1) the accuracy of the developed models in predicting chronological age in healthy individuals, 2) the models’ ability to maintain consistent performance across different age groups, and 3) the effectiveness of these models as biomarkers for identifying individuals likely to be affected by neurodegenerative conditions, such as dementia, Parkinson's, or multiple sclerosis, or prone to develop them in the future.
[0215]
[0196] Pareto-based selection is a multi-objective optimisation technique used to identify solutions that provide the best possible trade-offs between multiple performance metrics. This method identifies Pareto-optimal solutions, where improvement in one metric cannot occur without sacrificing another (Deb, 2011). These solutions form the Pareto front, representing the optimal trade-offs among the competing objectives. Given that the number of models maximising multiple metrics may still be large, especially when several hundreds of model options are evaluated, we employed a stepwise Pareto-based selection approach.
[0216]
[0197] The stepwise Pareto-based approach involves successive Pareto steps, where the models selected at each step must belong to the Pareto front derived from the previous iteration. At each step, we select performance metrics of interest to maximise as well as specific test sets for evaluation. More specifically, in this work we propose a two-step approach which we outline below. In the first step, we focus only on the UKBB test set and select three metrics on which to base the Pareto selection in order to address the aforementioned three objectives. In the second iteration, we assess how the performance of models selected in the first step generalises to the ADNI and NACO datasets, not only in age and disease prediction but also in terms of the performance in different subgroups. The selection of certain metrics for the identification of the second Pareto front reflects this by including also a fourth metric aimed at ensuring that the model performs consistently across diverse subpopulations, regardless of demographic factors or the specific MRI equipment used.
[0217]
[0198] Figures 4a and 4b show predicted brain age plotted against chronological age. Figure 4a shows results for data that has been scaled and normalised before being processed using the brain age model. Both the predicted age and actual age ranges from 55 to 90, incrementing in 5 years. These results show predictions for the training (larger yellow dots) and test (smaller blue dots) set. As explained above, the data only comprises healthy patients and no patients labelled as having dementia in the datasets. For these patients, chronological age should correspond to predicted brain age. Figure 4a shows a clear age bias in the predicted brain age results (blue solid line - labelled 40). The black line (labelled 42) shows what perfect correlation between chronological age and predicted brain age would look like. The data has the following statistical properties: PearsonR (statistic = 0.432, pvalue = 1 .43e-12 ), R2=0.151 , MAE= 4.702.
[0218]
[0199] Figure 4b shows results for data that has only been standardised before being processed using the brain age model. In this case the predicted age ranges from 40 to 140, incrementing in 10 years whereas the actual age ranges from 55 to 90, incrementing in 5 years. Yellow dots represent data from the training set, whereas blue dots represent data from the test set. The black line 42 shows what perfect correlation between predicted brain age and chronological age would look like.
[0200] Once we have a predicted brain age, as described above in Figures 1 and 2, the predicted brain age can be used in a risk prediction model (s). The target of the risk prediction model(s) is to determine the risk for an individual (male or female of age 50 to 90) of having a neurodegenerative condition (i.e. MCI or Dementia) at the present time and the risk of developing the condition in different time windows. The following binary risk prediction models (classification models) were defined: “not dementia” or “dementia”, “MCI” or “dementia”, “normal” or “MCI”, “normal” or “not normal”. Then the following scenarios were defined: discriminate the two classes at baseline (i.e. at current time), discriminate transition from cognitive normal to the two target classes and discriminate transition from MCI to the two target classes. The risk prediction submodels used in this case are logistic regression models.
[0219]
[0201] Figure 5a plots the results by the classification assigned by the classification model (risk prediction model). Red squares mean a patient has been classified as having dementia, small purple dots mean classification as MCI, and larger blue dots mean classification as CN. In this case the predicted age ranges from 40 to 140, incrementing in 10 years whereas the actual age ranges from 55 to 90, incrementing in 5 years. The black line shows what perfect correlation with age would look like. The patients classified as having dementia have a significantly increased predicted age relative to their actual age.
[0220]
[0202] Figure 5b shows predicted brain ages plotted against chronological age for participants of the UKB study. The data shown has been resampled so that the age distribution of the UKB participants is comparable to the age distribution of the ADNI participants. In this example, there are two classes “not dementia” and “dementia”. Blue dots represent participants that are cognitively normal (CN). Red dots represent participants that have dementia. The black dotted line shows what unitary correlation with chronological age would look like. It is clear that there is a trend for people with dementia to have a higher predicted brain age than their chronological age. Looking at this graph, it is also apparent that for an ageing related disease such as Alzheimer’s, the healthy cohort must have the same age distribution of the positive (i.e. dementia) cohort to effectively stratify people into CN and dementia groups (or other groups, as described above). As explained above, this may be achieved by pre-processing of the training data.
[0221]
[0203] Figures 6a and 6b are bar diagram showing how the probability of having dementia or a cognitively normal (CN) brain are distributed with chronological age and predicted brain age respectively. Red bars show the probability of having dementia, blue bars show the probability of having a CN brain. In both graphs, the probability of having dementia increases with age but the separation is clearer in Figure 6b.
[0222]
[0204] Figure 7 shows a receiver operating characteristic (ROC) curve which plots true positive rate against false positive rate for four risk prediction models each of which stratify participants of the UKB study into “Dementia” and “not dementia” groups. The different models (and hence curves) correspond to the risk at different target times, i.e. 0 months or baseline (blue line), 24 months (orange line), 60 months (green line), and 120 months (red line). In summary, the brain age and risk prediction model as described in the present application allow for a good prediction of dementia when tested on an independent validation data set made up of the UKB data as described above.
[0223]
[0205] Figures 8a and 8b plot values for the area under the receiver operating characteristic (AUROC) against months from baseline for different risk prediction models that take different factors into account. These factors include one or more of: age, gap between age (also termed chronological age) and predicted brain age (also termed biological age), and MoCa (Montreal Cognitive Assessment) total score. The blue line 52 shows performance for a risk prediction model which only stratifies patients based on their chronological age. The orange line 54 shows performance for a risk prediction model which stratifies patients based on their chronological age and the gap between a patient’s chronological and biological age. The green line 56 shows performance for a risk prediction model which stratifies patients based on their MoCa score (at time = 0). The red line 58 shows performance for a risk prediction model which only stratifies stratifying patients based on their predicted brain age. The purple line 60 shows performance for a risk prediction model which only stratifies stratifying patients based on the gap between their biological and chronological age. For both Figures 8a and 8b, the receiver operating characteristic is a curve such as that shown in Figure 7 that plots true positive rate against false positive rate at a threshold setting. For each model, evaluation of the classification using AUROC is done using the test data.
[0224]
[0206] In Figure 8a, each risk prediction model stratifies patients into “Dementia” and “not dementia” classifications. That is, each risk prediction model classifies each test data sample into one of the two classes. In other words, each model is a binary classifier. Each risk prediction model may also output a probability of a patient falling into one of the classes. Follow-ups shorter than 4 years (48 months) have too small transition groups to produce reliable results. Five stratified train-test splits were generated to overcome fluctuation due to the low number of cases in certain classes. A total of 22 (followup) * 5 (split) * 5 (feature sets) = 550 models were trained.
[0225]
[0207] The results of Figure 8a are summarised in the following table:
[0226]
[0208] The following table shows AUROC values for different age classes of the UKB dataset for the tested model:
[0227]
[0209] The following table shows how the number of participants used for model testing in the UKB dataset is distributed over different age classes:
[0228]
[0210] In Figure 8b, each risk prediction model stratifies patients into “Dementia” and “MCI” classifications.
[0229]
[0211] Figures 8a and 8b show that brain age (red line 58) is overall superior as a criterion for stratifying patients in different groups according to their dementia status. These results also show that while the prevalence of dementia / Alzheimer’s disease may be estimated by considering age (blue line 52) as a feature in the entire cohort of healthy people, as has, for example be done in a study by Wang (Wang, ZT., Fu, Y., Zhang, YR. et al. Modified dementia risk score as a tool for the prediction of dementia: a prospective cohort study of 239745 participants. Transl Psychiatry 12, 509 (2022). https: / / doi.org / 10.1038 / s41398-022-02269-2), age alone does not allow discrimination of healthy and unhealthy participants of the same age.
[0212] Figure 9 is a diagram of a system 1000 which may be used to train the relevant ML models and use the trained models. The system comprises an image capture device 910 and is configured to capture an image. The image capture device 910 may be an MRI device.
[0230]
[0213] The system comprises a user device such as an image processing device 900 which is configured to receive an image from the image capture device 910 and carry out the patient stratification and / or brain age prediction methods described herein (e.g. as described with respect to Figure 1). The image processing device 900 comprises at least one processor 902. The at least one processor 902 may comprise one or more of: a microprocessor, a microcontroller, and an integrated circuit. The image processing device 900 comprises memory 904 coupled to the at least one processor 902. The memory 904 may comprise volatile memory, such as random access memory (RAM), for use as temporary memory, and / or non-volatile memory such as Flash, read only memory (ROM), or electrically erasable programmable ROM (EEPROM), for storing data, programs, or instructions, for example.
[0231]
[0214] In this example, the image processing device 900 stores a brain age model 906 (e.g. a brain age DL model) and at least one risk prediction model 908 (classification model), although it will be appreciated that these may be stored in a database which is accessible by the image processing device 900 when an image is to be processed. These models may be trained on the image processing device 900 or may be trained on a separate server (not shown).
[0232]
[0215] The system comprises a user interface 912 which is configured to display an output result generated by the image processing device 900. The output result may be a modified version of the input image(s), e.g. modified to include labels for different volumes and / or predicted age and / or risk. The user interface 912 may be part of the image processing device 900 or may be separate but connected to the user processing device 900 as illustrated. The user interface 912 may be a display device, for example.
[0233]
[0216] FastSurfer Features and Cole’s Approach Cole’s brain age model. We showed the capability of the IDPs extracted by the UKB pipeline (Alfaro-Almagro et al., 2018) for predicting brain age and the association of PAD with lifestyle and health. Cole’s work highlighted the utility of a multimodal approach as well as the significance of each MRI type separately while addressing the age-bias phenomena and emphasised the small number of significant features selected by the model (32) in comparison with the initial amount of features considered (approximately one thousand). Last but not least, this was based on features available to all researchers that have access to the UKBB database making the model easily reproducible. However, when it comes to generating the same array of features in a new dataset, the UKB pipeline requires a large computation time, up to several hours just for one T1w scan, making it difficult to extend this approach to large databases different from the UKBB.
[0234]
[0217] Figures 10a and 10b show a comparison of feature rank for the 4 variations of Cole model trained with FastSurfer features: with age resampled, 10-fold, bootstrap and bootstrap + linear. In Figure 4c, the coefficients follow the 10-fold ranking; in Figure 4d, the coefficients follow the age resampled ranking.
[0218] Based on these considerations and in order to expand our analysis beyond the UKBB dataset, we resorted to the use of FastSurfer to generate a similar set of MR I -de rived features to build ML-based brain age prediction models (something that to the best of our knowledge has not been done previously on such a large scale). For testing the capability of this approach we replicated the methodology proposed by Cole with FastSurfer IDP inputs and compared the resulting performance with the one reported in the original publication.
[0235]
[0219] In the table below, we show the performance of a 10-fold cross-validation model with and without resampling in comparison with the corresponding model published by Cole. We include for completeness the performance of both T1w-only and multimodal approaches from Cole and the performance of the “bootstrap”, “bootstrap+linear” an “resampled” models derived with FastSurfer features. The performances of FastSurfer based models are comparable or overperforming the published model, validating the capacity of this tool to provide meaningful features for brain age estimation with a very short computation time.
[0236]
[0220] FastSurfer generates 95 volumetric features that only partially overlap with the 165 features derived from T1w images employed in the original model. The overlapping features show, on average, high correlation with the corresponding FastSurfer volume estimations (Pearson R coefficient = 0.67 ± 0.14), however a perfect one-to-one correspondence is not possible for the entire set, since the atlas annotation is different.
[0237]
[0221] In relation to features selection we studied the impact of bootstrapping, subsequent linear fit, and the resampling of age groups. The bootstrap coherently removes features with small coefficients while keeping a similar ranking. The models with resampling of age groups and 10- fold select a certain number of different features but have in common the majority of most significant features. The differences in the selected features and their coefficients for the model with age resampling are responsible for improving the correlation between chronological and predicted age. The four variations of the model have in common 23 selected features with a consistent coefficient sign. The most significant features with a positive coefficient were the brainstem, together with 3rd-Ventricle, Left-choroid-plexus, Right-Caudate, WM-hypointensities (white matter hypointensities) and ctx-lh-insula (insular cortex of left hemisphere), while the most significant features with negative coefficients were Right-VentralDC (right ventral diencephalon), Left-Thalamus, Left-Cerebellum-White-Matter and Left-Accumbens-Area. A few of these were in common with the ones selected in the original model of Cole: Brain stem + 4th ventricle (field 25025), and Right-Thalamus (field 25012).
[0238]
[0222] The table below shows results from a LASSO model trained and tested over UKBB, ADNI, NACC and joint test (before bias correction):
[0239] 223] Given the broad range of modelling approaches explored, the results are presented by distinguishing between three modelling modalities: Cole's LASSO and its variants (CL), additional standard ML, and Deep Learning models (DL).
[0240]
[0224] Age Prediction Performance of uncorrected approaches. Initially, we focus on analysing the models’ ability to predict chronological age in healthy individuals. The results are presented in terms of MAE across the full test set (no distinction between UKBB, NACC and ADNI), as commonly reported in the literature (Cole et al., 2020; Zhang et al., 2023) but the MAE in different age bins is also analysed via the max bin MAE metric. This statistical analysis leverages the variety of modelling configurations explored including factors such as the definition of a healthy individual, whether models are age-specific, the use of individual or merged IDPs (sum of left / right versions of features), and whether oversampling / resampling was applied prior to training the corresponding model. The objective is to identify which configurations, on average, are most likely to produce a high-performing brain age model, without applying age-bias corrections (simply to disregard the effects of corrections in this first part of the analysis). A summary of the key results from this analysis is provided in the table below and summed up in the text below. This table shows a summary of statistical tests outcomes per configuration type against age-prediction performance measured via MAE and max bin MAE. Results are given separately for males and females, and the best option (in average) is reported when a statistically significant difference is observed:
[0241]
[0225] The effect of the model type. To emphasise only strong statistical differences across configurations, we focus on cases where the p-value is below 1 e-2. An initial observation regarding the modelling strategies for ML / CL approaches is that there are no substantial differences in age-prediction performance between linear models (e.g., Ridge, Elastic Nets, LASSO, OLS) and non-linear models (e.g., TPOT, FLAML, XGBoost / LGBM Hyperopt), as indicated by p-values above 0.04 for both sexes. A multigroup Kruskal-Wallis test comparing all models individually yielded p-values exceeding 0.06, suggesting no significant differences in individual ML models either. However, for CL approaches, the standard LASSO trained with 10- fold cross-validation consistently outperformed Cole’s LASSO variant with bootstrap (p-values < 1 e-3). In the case of DL, the SFCN architecture provided the best performance (p-values < 1 e-3) even though, for females, performance is comparable to that of ResNet-18.
[0242]
[0226] Notably, when shifting our focus from overall MAE to examining the maximum MAE within each 5-year age bin (max bin MAE), we identified significant differences across the models (p < 1 e-6 for both genders, Kruskal-Wallis). This discrepancy arises because linear models (specifically Ridge, Elastic Nets, LASSO and OLS) exhibit markedly lower max bin MAE compared to non-linear methods (such as TPOT, FLAML and XGBoost / LGBM Hyperopt), suggesting better worst-case predictions for linear models (p < 1 e-4 for both genders, Mann- Whitney U-test). This phenomenon is attributed to the limited number of MRIs available for individuals over 80 years of age in the part of the test set belonging to UKBB, where the maximum age reached is 82. In contrast, the test set contains a higher volume of MRIs for the oldest age bin derived from the NACC and ADNI datasets. When we constrained the upper age limit of the test set to 80 and repeated our analysis, we found no significant differences in max bin MAE between linear and non-linear models. This observation suggests that, while linear models do not demonstrate a significant advantage over non-linear counterparts in global MAE, they perform better in forecasting unobserved cases, particularly within the over-80 age group. For the CL- based / DL-based approaches, we observe a significant difference in max bin MAE (p-values < 1 e- 2, Kruskal-Wallis / ANOVA), once again favouring the standard LASSO trained with a 10-fold cross-validation over Cole's bootstrapped version for CL and the SFCN architecture for DL.
[0243]
[0227] The effect of the healthy definition and sex specification. Further results indicate no significant difference in MAE between ML / CL / DL models trained on “healthy_orx” versus “healthy_cole” datasets for either females or males (p-values > 0.01 , Mann-Whitney U-test). Additionally, when comparing models trained on each sex separately against those trained on both sexes with an added feature for sex specification, a statistical difference was observed only in the case of ML models for females where training on both sexes led to a better global MAE (p- values < 1 e-2, Mann-Whitney U-test).
[0244]
[0228] The effect of the feature type and oversampling. Lastly, when evaluating the impact of feature types, significant differences were found only in CL models (p-values < 1 e-2, Mann- Whitney U-test), with the best performance observed when merging IDPs obtained from FastSurfer for the same brain region in the left and right hemispheres. Additionally, the application of default oversampling (i.e., resampling within each 5-year age bin to equalise the number of cases) provided a clear advantage for ML / CL approaches. Both MAE and max bin MAE significantly improved following resampling, with p-values < 1 e-6 (Mann-Whitney U-test).
[0245]
[0229] In summary, we find that linear models trained with standard 10-fold cross-validation, using merged IDPs from FastSurfer, considering both sexes in training, and incorporating resampling prior to training, generally achieve the best average performance in terms of age prediction. In the case of DL, the SFCN architecture in general seemed to provide the best age-prediction capabilities for this type of model. More importantly, the combination of linear approaches with resampling enhances the generalizability of the developed brain age clock, allowing for improved forecasting — especially crucial given the differences in age distribution between the UKBB, NACC, and ADNI datasets.
[0246]
[0230] Disease Prediction Performance. Our disease prediction analysis focused on binary classifications between healthy individuals and those diagnosed with various conditions, using data from NACC, ADNI, and UKBB. We concentrated on conditions with at least 50 samples in the minority class. Healthy individuals were defined using the "healthy_orx" criteria, while diagnosed individuals were identified through prior diagnoses and clinical cognitive assessments. The binary classification tasks included comparisons such as Healthy vs. non-Healthy, Healthy vs. Dementia, Healthy vs. Mild Cognitive Impairment, Healthy vs. individuals with Subjective Memory Complaints, Healthy vs. Parkinson's Disease, and Healthy vs. Multiple Sclerosis.
[0247]
[0231] Brain age was used as the baseline predictor, while PAD (the difference between predicted and chronological age) was also evaluated as a potential indicator. The area under the receiver operating characteristic curve (AUROC) was used to assess the performance of these binary classifications. In the table below, we present key findings from our statistical analysis highlighting in the “Above Threshold” column whether the chosen predictor (brain age or PAD) achieved on average an AUROC of 0.70 or higher for the binary classification considered. This value was selected as threshold since the baseline metric, i.e., the AUROC computed when using chronological age as a predictor, does not exceed 0.70 for any of the binary tasks. In fact, it consistently falls below 0.60. Therefore, for the conditions highlighted in this column, there is a significant improvement in using our predicted age over chronological age for distinguishing individuals affected by these diseases. Below, there is provided a detailed breakdown of AUROCs for each disease, comparing chronological age and the selected models.
[0248]
[0232] The table below shows a summary of statistical tests outcomes per configuration type against disease-prediction performance measured via AUROC. Results are given separately for males and females, and the best option is reported when a statistically significant difference is observed. The “Above Threshold” column highlights cases in which the average AUROC computed exceeds 0.70:
[0249]
[0233] The results in the above table demonstrate that the developed brain age models perform competitively as biomarkers, particularly in distinguishing healthy individuals from those with dementia (for ML / CL approaches) and multiple sclerosis (for ML / CL / DL models), showing a massive AUROC improvement of over 0.30 on average when compared to chronological age. This highlights the potential of brain age as a valuable biomarker for both neurological conditions. The performance of the predictor varied across conditions, with brain age outperforming PAD in differentiating between Healthy vs. non-Healthy and Healthy vs. Dementia, while PAD showed stronger results for distinguishing Healthy vs. Multiple Sclerosis. Additionally, with the exception of the Healthy vs. Multiple Sclerosis classification case, linear ML models trained via 10-fold cross-validation and the SFCN architecture for DL consistently outperformed their counterparts.
[0250]
[0234] Corrected Approaches. As discussed above, three prominent age-bias correction methods are incorporated: Cole's, Lange's, and Zhang's approaches. For every model developed, we also trained each of these three correction methods. Here, we present the results of these approaches.
[0251]
[0235] Before examining the results, it is important to address the challenges encountered with the Zhang correction in non-linear ML approaches, specifically TPOT, FLAML, and XGBoost / LGBM Hyperopt. These methods faced difficulties due to the limited number of cases in the training set for individuals over 80 years old, which often led to models exhibiting overfitting. In these instances, the variance parameter associated with Zhang's correction approached zero, resulting in numerical instability. Consequently, this analysis focuses on performance metrics for the population under 80 years of age and excludes the non-linear ML approaches from consideration. In the table below, we present the most relevant findings. This table includes an additional column, “Average Gap”, that computes the difference between the average performance of each alternative and the top-performing approach; i.e., a value of 0 indicates the best alternative. The table below shows a summary of statistical tests outcomes per age-bias methodology considered against age & disease-prediction performance metrics:
[0252]
[0236] Concerning the performance of age prediction. From the table above, it is evident that Zhang-corrected models significantly outperform all other alternatives for age prediction as was expected, regardless of the approach (ML / CL / DL), by reducing the MAEs by more than 1 , 3, and 4 years compared to Lange-corrected, uncorrected, and Cole-corrected models, respectively. A similar trend is observed in the max bin MAE, indicating that Zhang-corrected models demonstrate a high level of age-prediction performance for healthy individuals regardless of the age group to which they belong. Here, we used the "healthy_orx" definition for our average estimates.
[0253]
[0237] Such an exceptional age-prediction capability appears to influence the performance of Zhang-corrected models in disease prediction since, as expected, they exhibit similar performance to chronological age prediction as an indicator. As we will explore, this observation is true for the average performance of these models. However, when comparing all developed approaches, some Zhang-corrected alternatives also reveal competitive capabilities in disease prediction.
[0254]
[0238] Concerning the performance of disease prediction. Before exploring the disease prediction performance metrics, it is important to note that, by design, the Cole-corrected versions yields identical AUROC values when using brain age as an indicator across all binary problems. In this analysis, we specifically focused on the binary settings that appeared most promising based on our results. Notably, in all binary comparisons for ML / CL models (Healthy vs. non-Healthy, Healthy vs. Dementia, Healthy vs. Multiple Sclerosis), both Cole-corrected and uncorrected approaches consistently achieved the highest AUROCs, exceeding those of their counterparts (Lange and Zhang) by at least 0.05 on average. The only exception was observed in the multiple sclerosis comparison, where PAD was used as the indicator, and uncorrected approaches demonstrated the best performance. In the case of DL no significant differences were observed across the various corrections evaluated.
[0255]
[0239] The results presented here indicate that while corrections utilising chronological age (Lange and Zhang in particular) excel in age prediction for healthy individuals, they tend to be outperformed by Cole-corrected and uncorrected approaches in their disease prediction capabilities. This observation underscores the need for a more detailed analysis to understand the trade-offs between these two properties at the model level, allowing for the selection of the most promising methodologies.
[0256]
[0240] Cole Correction vs. Uncorrected Models with Resampling. The resampling step considered in this study aimed to balance the representation of all age bins and subgroups by augmenting the minority bins with existing samples prior to model training. Although the ideal approach would involve incorporating new MRIs for these underrepresented age groups to avoid the risks typically associated with oversampling, this method effectively compensates for the concentration of MRIs in specific age groups, thereby mitigating the age-bias issue. Here, we compare uncorrected models that employed resampling against Cole-corrected models that did not include any resampling, focusing on their overall age-prediction capabilities and their performance by age bin. This comparison is particularly relevant as other standard age-bias correction methods, such as Lange’s and Zhang’s, rely on chronological age fortheir adjustments, while the methods analysed here are independent of that factor. Notably, the resampled approach conducts the entire training in a single phase, unlike Cole's two-step method. The main findings from this analysis are summarised in the table below, where we also included uncorrected models without resampling, labelled here “Uncorrected”, as a baseline for comparison.
[0257]
[0241] The table below shows a summary of statistical tests outcomes of comparison of resample preprocessing as an alternative to Cole’s correction:
[0258] 242] According to posthoc analysis of the multigroup test, even though Cole has the best mean for ML, it is not significantly different from the Resample case. Nevertheless, both are significantly different from the uncorrected case.
[0259]
[0243] We observe that implementing a resampling step priorto training the brain age model yields notable improvements in age-prediction performance as intended. This preprocessing not only decreases the overall MAE compared to models without resampling and Cole-corrected models (with average reductions of up to 2 years), but also significantly reduces the max bin MAE, with average reductions up to 3 years with respect to the baseline uncorrected approach.
[0260]
[0244] However, when considering alternative metrics commonly used to assess age-bias, such as the absolute correlation between PAD and chronological age (Peng et al., 2021), Cole- corrected models demonstrate an advantage by more effectively minimising the correlation between prediction errors and actual age. In this sense, while resampling presents a straightforward and effective approach to enhancing overall MAE performance and ensuring greater consistency in age-prediction across age bins, it may still exhibit residual age-bias, albeit reduced, in both ML and CL methodologies.
[0261]
[0245] Individual Models Comparison. We leveraged various modelling configurations to assess significant differences on average in age- and disease-prediction performance. Here, we focus instead on the selection of the top-performing individual models developed in our work.
[0262]
[0246] We implemented a stepwise Pareto-based model selection framework aimed at identifying a small number of models / corrections that excel in both age-prediction and disease-prediction tasks, while also ensuring strong generalisation to external datasets. This multi-stage approach consists of two main steps:
[0263]
[0247] Step 1 : Initial Pareto-Based Selection
[0264]
[0248] The first step involves selecting models — irrespective of whether they belong to ML / CL / DL categories — that are positioned on the Pareto front when optimising three key performance metrics: MAE_bounded, max_bin_MAE_unbounded, and AUROC_unbounded_HvsNoH. These metrics were chosen for their relevance in capturing different dimensions of model performance, as explained below:
[0265]
[0249] MAE_bounded (bounded MAE): This metric focuses on the standard age-prediction performance. In most studies, model selection is based on minimising overall MAE across the age range considered in the training set. The bounded version of MAE excludes individuals with ages above 80 (a group underrepresented in the UKBB MRI dataset) to provide a more meaningful comparison of model performance in age groups where training data is sufficiently available.
[0266]
[0250] max_bin_MAE_unbounded (max bin MAE, unbounded): This metric assesses the forecasting capacity of the models by tracking their worst-case performance across different age bins, particularly in the over-80 group, where model accuracy often deteriorates. This ensures that models are not only effective in the average case but also robust in age groups where prediction is more challenging.
[0267]
[0251] AUROC_unbounded_HvsNoH (Area Under the Receiver Operating Characteristic Curve for distinguishing Healthy vs. non-Healthy individuals, unbounded): This metric assesses the models' performance in predicting disease by evaluating their ability to differentiate between healthy individuals and those with neurodegen erative conditions, without initially categorising by specific diseases. For each model, we select the highest AUROC value achieved when using either brain age or PAD as the predictive feature. This provides an essential measure of the model's effectiveness as a diagnostic tool, beyond simply predicting age.
[0268]
[0252] Step 2: Generalisation and External Validation
[0269]
[0253] After the initial Pareto-based selection — which focuses on performance on the UKBB test set — we consider a second-stage Pareto-based selection by looking at four metrics (namely, MAE_bounded, max_bin_MAE_unbounded, AUROC_unbounded_HvsNoH, and generalization_median) computed on the external datasets of interest (ADNI and NACC), to test the generalisation capabilities of the selected models. The fourth metric introduced here, generalization_median, is defined by first calculating the median difference in MAE for healthy individuals across key demographic and technical subpopulations — such as ethnicity, MRI manufacturer, and MRI machine type — and then computing the maximum of these median values. The purpose of this metric is to ensure that the model performs consistently across diverse subpopulations, regardless of demographic factors or the specific MRI equipment used. By tracking the median across these subgroups, we assess the robustness of the model in real-world scenarios where such factors vary widely, and by controlling the worst-case performance, we ensure that the identified models maintain strong performance even in more challenging or less familiar data environments.
[0270]
[0254] As a result, the selected models are not only optimised for the UKBB dataset but also demonstrate generalizability and robustness when applied to other databases. This approach aims to find the models that strike the best balance between high accuracy in both age- and disease-prediction tasks, minimal bias across subpopulations, and strong generalisation to new data sources.
[0271]
[0255] The Pareto fronts generated through the model selection process described above, applied to over 800 models for males and an additional 800 models for females, are presented in the provided .csv files, see pareto_final_male / female.csv.
[0272]
[0256] The majority of selected models via the Pareto-based approach described above (i.e., those offering the optimal trade-off between age prediction, disease classification, and generalizability) consist primarily of linear models trained using standard 10-fold cross-validation and were corrected with Zhang’s method. Notably, none of the LASSO models proposed by Cole were included in the final selection. The selected models consistently delivered an outstanding overall MAE below 1.12 years for healthy individuals and a max bin MAE under 1.5 years, demonstrating high-level performance across all age groups. These figures represent the worstcase performance observed across the external datasets, namely NACC and ADNI.
[0273]
[0257] It is also noteworthy that for the male cohort, a small number of linear models corrected via Lange's method were included among the top-performing models, owing to their superior capacity for distinguishing non-healthy individuals — which will be discussed in more detail later. Additionally, in both the male and female cohorts, a limited number of DL models, based on the ResNet-18 architecture and corrected with Lange’s method, were selected. While these models exhibited strong age-prediction performance, they were largely ineffective in differentiating between healthy and non-healthy individuals.
[0274]
[0258] When considering the initial Pareto fronts, based solely on performance within the UKBB dataset (see pareto_initial_male / female.csv), a more diverse set of approaches emerged. This included several non-linear models, predominantly those utilising boosting techniques and corrected via Zhang or Cole, which initially demonstrated strong performance in distinguishing non-healthy individuals. However, these models did not generalise effectively to the external datasets (NACC and ADNI) and were ultimately excluded from the final selection.
[0275]
[0259] Disease Prediction Performance Across Selected Models. The table below outlines the top- performance metrics achieved per sex and for each disease across the entire dataset by the Pareto-selected brain age models. These results are compared against chronological age, which is considered here as the baseline metric. The table below shows Summary of best performance in disease-prediction using selected models and its comparison against chronological age as baseline score:
[0276]
[0260] The results presented in the table above highlight the potential of predicted brain age as a meaningful biomarker for various neurodegenerative conditions. Across all conditions analysed, the AUROC of the selected models consistently surpasses by a substantial margin that of chronological age, which in principle should serve itself as an indicator, given the clear association between increasing age and these conditions. The performance for dementia stratification is particularly notable, with an AUROC close to 0.90 for both males and females.
[0277]
[0261] If we were to extend the analysis on disease prediction performance to include all developed models (regardless of whether they made it into the Pareto selection or not), models with higher AUROCs-as high as 0.942 for Multiple Sclerosis and 0.885 for Parkinson's disease — emerge. This difference arises from the particular set of metrics considered in the Pareto selection, which prioritised distinguishing healthy individuals from non-healthy ones in the NACC and ADNI datasets, neither of which contain multiple sclerosis or Parkinson's disease cases. This suggests that for disease-specific brain clocks, different disease-specific metrics should be given priority and the presence of relevant cases in the dataset over which the selection is applied ought to be ensured.
[0278]
[0262] Age Prediction Performance in Specific Subgroups Across Selected Models. As previously noted during the Pareto selection stage, all models on the Pareto front demonstrated highly competitive performance in age prediction, typically achieving a MAE of less than 1 year, even when considering the worst performance across all external databases.
[0279]
[0263] The analysis is extended to additional subgroups, specifically examining the ethnicity of the healthy individuals analysed, the manufacturer of the MRI machines, and the specific MRI machines utilised. The aim is to determine whether the observed performance is indeed consistent across all relevant scenarios. The table below presents a summary of the key findings from this analysis, highlighting the MAEs for the complete test set associated with the models that exhibited the smallest MAE variations across the corresponding groups. The table below shows a summary of age-prediction performance for the subgroups of interest utilising the models selected above, focusing on those with the smallest gaps for each group:
[0280]
[0264] According to the outcomes presented in the table above, the selected models here, although trained exclusively on the UKBB dataset, exhibit a remarkable ability to predict the age of healthy individuals across all subgroups considered. The MAE across all subgroups, including different ethnicities, MRI machine manufacturers, and specific MRI machines, remained below 1 year. This indicates that the models demonstrate a high level of robustness and generalizability, effectively extending their predictive power beyond the original training set.
[0281]
[0265] The fact that the models maintain such a consistent level of accuracy across diverse subgroups reinforces their versatility and potential utility in real-world applications. Additionally, these models have proven capable in distinguishing non-healthy individuals from healthy ones, particularly excelling in the detection of dementia cases. This is especially noteworthy because the models consistently outperformed chronological age in binary classification tasks, which is often considered a strong baseline due to the natural association between age and many neurodegenerative conditions. The ability of these models to generalise across different subgroups while maintaining high prediction accuracy makes them not only valuable for ageprediction but also effective tools for identifying individuals at risk of various neurodegenerative diseases.
[0282]
[0266] Conclusions. We present a novel and comprehensive analysis focused on developing reliable brain clocks, utilising UKBB data as the training set. Our approach encompasses widely used ML methods (Cole et al., 2020) and DL techniques (Peng et al., 2021), and it extends to additional modelling approaches, including various penalised linear regressions and boosting models. We have also integrated the most commonly used age-bias correction techniques, as illustrated by Zhang et al. (2023).
[0267] Our analysis advances beyond standard practices in several ways:
[0283]
[0268] Predictive Assessment: We do not only evaluate how effectively these approaches predict an individual's age, but we also assess the performance of the resulting brain clock as a biomarker for neurodegenerative diseases.
[0284]
[0269] Robust Statistical Framework: We have established a solid statistical framework to evaluate whether different model configurations significantly affect age and disease prediction performance. This includes examining the impact of a less restrictive definition of healthy, sexspecific model considerations, the treatment of volumetric features (whether to keep them separate or merge them before training), and the effect of resampling underrepresented age bins on MRI data.
[0285]
[0270] Stepwise Pareto-Based Model Selection: We introduce an innovative model selection strategy based on a stepwise Pareto front approach. This strategy allows us to identify effective brain clocks that serve as strong biomarkers for neurodegenerative conditions while ensuring robust age prediction performance that generalises well to external databases.
[0286]
[0271] The key findings include:
[0287]
[0272] Model Performance: Models and associated age-bias correction methods developed exclusively using the UKBB dataset can yield top-performing brain clocks. These are often based on penalised linear models adjusted via Zhang correction, demonstrating strong age prediction capabilities (with mean absolute errors below one year on external datasets) and effective disease prediction (outperforming chronological age across all neurodegenerative conditions in the UKBB, ADNI, and NACC datasets, particularly for dementia and multiple sclerosis, where the AUROC often exceeds 0.80).
[0288]
[0273] Model Configuration Analysis: When assessing the average performance of various model configurations, the best uncorrected age and disease prediction capability appears to be achieved using linear models trained with K-fold cross-validation without bootstrapping, as suggested in (Cole et al., 2020). A less restrictive health definition than Cole's, termed "healthy_orx", did not enhance performance, and sex-specific models performed similarly to those trained with data from both males and females. For feature-based models, merging image-derived phenotypes proved beneficial, especially for penalised approaches.
[0289]
[0274] Deep Learning Approaches: We observed that generalisation to external databases proved more challenging for DL models, highlighting the necessity for fine-tuning to improve their generalizability. Notably, when certain architectures, especially ResNet-18, were corrected using the Lange approach, age prediction performance for external datasets significantly improved, although performance in disease stratification remained limited.
[0290]
[0275] Correction Techniques: It is important to emphasise that Zhang’s correction stands out as the most effective method for age prediction. However, it can lead to numerical issues if applied to models suffering from overfitting, particularly in non-linear contexts. For disease prediction, Cole’s correction and uncorrected approaches generally provide the best alternatives, although individual cases suggest that Zhang- and Lange-corrected models can also perform well. This underscores the need for developing correction methods that minimise age bias while maintaining the disease prediction capabilities of uncorrected approaches.
[0291]
[0276] Resampling: Simple techniques, such as resampling underrepresented age bins to normalise target distributions prior to model training, can lead to reduced overall and age-bin mean absolute errors compared to both Cole-corrected and uncorrected alternatives.
[0292]
[0277] While the results are promising, showing that the UKBB dataset provides a strong foundation for developing brain age models with effective age- and disease-prediction capabilities that typically generalise well to external datasets like ADNI and NACC, we believe that more careful handling of dataset-specific characteristics could significantly affect both image quality and derived phenotypes, ultimately influencing model performance. To further enhance results, especially in DL models, incorporating a larger number of images during training would be beneficial. For instance, using the healthy_orx definition across the full UKBB dataset, rather than relying on Cole's definition to select subjects, could provide valuable insights into the impact of this approach. Importantly, the DL models here faced challenges with generalisation, underscoring their sensitivity to variations in imaging protocols, patient demographics, and other dataset-specific factors. In contrast, the software used to extract anatomical features fortraditional ML methods may have helped reduce these effects, offering more stable, high-level inputs compared to using raw 3D MRI scans alone.
[0293]
[0278] In our future work we plan to investigate not only the use of a broader, more diverse set of images for model training but also the possibility of fine-tuning DL approaches (e.g., by freezing specific layers or adjusting them for tasks like disease stratification). When looking at merging different datasets, the application of harmonisation techniques, such as ComBat (Fortin et al., 2018; Marzi et al., 2024), prior to model training for both corrected and uncorrected FastSurfer feature-based models will also be analysed. Finally, the development of new alternatives to agebias correction that do not require prior knowledge of an individual's age is one of our current areas of focus. Our aim is to retain strong disease prediction performance while ensuring effective age prediction across different age groups. This may involve creating custom loss functions that emphasise underrepresented age bins and / or allow the incorporation of non-healthy MRI images during training to penalise predictions for undesired cases.
[0294]
[0279] Those skilled in the art will appreciate that while the foregoing has described what is considered to be the best mode and where appropriate other modes of performing present techniques, the present techniques should not be limited to the specific configurations and methods disclosed in this description of the preferred embodiment. Those skilled in the art will recognise that present techniques have a broad range of applications, and that the embodiments may take a wide range of modifications without departing from any inventive concept as defined in the appended claims.
[0295] References: • Alfaro-Almagro, Fidel, et al. “Image processing and quality control for the first 10,000 brain imaging datasets from UK Biobank.” NeuroImage 166 (2018): 400-424.
[0296] • Alzheimer's Association. "2019 Alzheimer's disease facts and figures." Alzheimer's & dementia 15.3 (2019): 321-387.
[0297] • Bashyam, Vishnu M., et al. “MRI signatures of brain age and disease over the lifespan based on a deep brain network and 14468 individuals worldwide.” Brain 143.7 (2020): 2312-2324.
[0298] • Beekly, Duane L., et al. "The National Alzheimer's Coordinating Center (NACC) database: the uniform data set." Alzheimer Disease & Associated Disorders 21 .3 (2007): 249-258.
[0299] • Bergstra, James, et al. "Hyperopt: a python library for model selection and hyperparameter optimization." Computational Science & Discovery 8.1 (2015): 014008.
[0300] • Biobank, U. K. "Protocol for a large-scale prospective epidemiological resource." (2007). Retrieved from https: / / www.ukbiobank.ac.uk
[0301] • Boaventura, Mateus, et al. “T1 / T2-weighted ratio in multiple sclerosis: A longitudinal study with clinical associations.” NeuroImage: Clinical (2022) 34:102967.
[0302] • Cole, James H., et al. "Predicting brain age with deep learning from raw imaging data results in a reliable and heritable biomarker." NeuroImage 163 (2017): 115-124.
[0303] • Cole, James H., and Franke, Katja. "Predicting age using neuroimaging: innovative brain ageing biomarkers." Trends in neurosciences 40.12 (2017): 681-690.
[0304] • Cole, James H., et al. "Brain age predicts mortality." Molecular psychiatry 23.5 (2018): 1385-1392.
[0305] • Cole, James H. "Multimodality neuroimaging brain-age in UK biobank: relationship to biomedical, lifestyle, and cognitive factors." Neurobiology of aging 92 (2020): 34-42.
[0306] • Cumplido-Mayoral, Irene, et al. "Biological brain age prediction using machine learning on structural neuroimaging data: Multi-cohort validation against biomarkers of Alzheimer’s Disease and neurodegeneration stratified by sex." Elife 12 (2023): e81067.
[0307] • Dartora, Caroline, et al. "A deep learning model for brain age prediction using minimally preprocessed T1w images as input." Frontiers in Aging Neuroscience 15 (2024): 1303036.
[0308] • de Lange, Ann-Marie G., and Cole, James H. "Commentary: Correction procedures in brain-age prediction." NeuroImage: Clinical 26 (2020).
[0309] • De Lau, Lonneke ML, and Breteler, Monique MB. "Epidemiology of Parkinson's disease." The Lancet Neurology 5.6 (2006): 525-535.
[0310] • Deb, Kalyanmoy. "Multi-objective optimisation using evolutionary algorithms: an introduction." Multi-objective evolutionary optimisation for product design and manufacturing. London: Springer London, 2011. 3-34.
[0311] • Dinsdale, Nikola K., et al. “Learning patterns of the ageing brain in MRI using deep convolutional networks.” NeuroImage 224 (2021): 117401.
[0312] • Feng, Xinyang, et al. “Estimating brain age based on a uniform healthy population with deep learning and structural magnetic resonance imaging.” Neurobiology of Aging 91 (2020): 15-25.
[0313] • Fortin, Jean-Philippe, et al. “Harmonization of cortical thickness measurements across scanners and sites.” NeuroImage 167 (2018): 104-120.
[0314] • Frisoni, Giovanni B., et al. "The clinical use of structural MRI in Alzheimer disease." Nature reviews neurology 6.2 (2010): 67-77.
[0315] • He, Kaiming, et al. “Deep residual learning for image recognition.” Proceedings of the IEEE conference on computer vision and pattern recognition (2016): 770-778.
[0316] • Henschel, Leonie, et al. "FastSurfer - A fast and accurate deep learning based neuroimaging pipeline." NeuroImage 219 (2020): 117012. • Jack Jr, Clifford R., et al. "The Alzheimer's disease neuroimaging initiative (ADNI): MRI methods." Journal of Magnetic Resonance Imaging: An Official Journal of the International Society for Magnetic Resonance in Medicine 27.4 (2008): 685-691 .
[0317] • Jack Jr, Clifford R., et al. "NIA-AA research framework: toward a biological definition of Alzheimer's disease." Alzheimer's & dementia 14.4 (2018): 535-562.
[0318] • Kamnitsas, Konstantinos, et al. “Efficient multi-scale 3D CNN with fully connected CRF for accurate brain lesion segmentation.” Medical Image Analysis 36 (2017): 61-78.
[0319] • Kawahara, Jeremy, et al. “BrainNetCNN: convolutional neural networks for brain networks; towards predicting neurodevelopment.” Neuroimage 146 (2017): 1038-1049.
[0320] • Kolbeinsson, Arinbjbrn, et al. “Accelerated MRI-predicted brain ageing and its associations with cardiometabolic and brain disorders.” Scientific Reports 10.1 (2020): 19940.
[0321] • Krizhevsky, Alex, Sutskever, Ilya, and Hinton, Geoffrey E. “ImageNet classification with deep convolutional neural networks.” Advances in Neural Information Processing Systems (2017): 1097-1105.
[0322] • Li, Xiangrui, et al. “The first step for neuroimaging data analysis: DICOM to NlfTI conversion.” Journal of Neuroscience Methods 1.64 (2016): 47-56.
[0323] • Lin, Weiming, et al. “Convolutional neural networks-based MRI image analysis for the Alzheimer’s disease prediction from mild cognitive impairment.” Frontiers in Neuroscience 12 (2018): 777.
[0324] • Liu, Mingxia, et al. “Landmark-based deep multi-instance learning for brain disease diagnosis.” Medical Image Analysis 43 (2018): 157-168.
[0325] • Long, Jonathan, Shelhamer, Evan, and Darrell, Trevor. “Fully convolutional networks for semantic segmentation.” The IEEE Conference on Computer Vision and Pattern Recognition (2015): 3431-3440.
[0326] • Marzi, Chiara, et al. “Efficacy of MRI data harmonization in the age of machine learning: a multicenter study across 36 datasets.” Scientific Data 11 (2024): 115.
[0327] • Ning, Kaida, et al. “Improving brain age estimates with deep learning leads to identification of novel genetic factors associated with brain aging.” Neurobiology of Aging 105 (2021): 199-204.
[0328] • Olson, Randal S., and Moore, Jason H. "TPOT: A tree-based pipeline optimization tool for automating machine learning." Workshop on automatic machine learning. Proceedings of Machine Learning Research, 2016.
[0329] • Peng, Han, et al. "Accurate brain age prediction with lightweight deep neural networks." Medical image analysis 68 (2021): 101871.
[0330] • Planche, Vincent, et al. “Structural progression of Alzheimer’s disease over decades: the MRI staging scheme.” Bain Communications (2022) 4.3: fcac109.
[0331] • Prince, Martin, et al. “World Alzheimer Report 2015. The Global Impact of Dementia: An analysis of prevalence, incidence, cost and trends.” Diss. Alzheimer's Disease International, 2015.
[0332] • Raghu, Maithra, et al. “Transfusion: understanding transfer learning for medical imaging.” Proceedings of the 33rd International Conference on Neural Information Processing Systems 301 (2019): 3347-3357.
[0333] • Russakovsky, Olga, et al. “ImageNet large scale visual recognition challenge.” International Journal of Computer Vision 115 (2015): 211-252.
[0334] • Simonyan, K, and Zisserman, Andrew. “Very Deep Convolutional Networks for Large- Scale Image Recognition.” International Conference on Learning Representations (2014).
[0335] • Shwartz-Ziv, Ravid, and Armon, Amitai. "Tabular data: Deep learning is not all you need." Information Fusion 81 (2022): 84-90. • Smith, Stephen M., et al. "Estimation of brain age delta from brain imaging." Neuroimage 200 (2019): 528-539.
[0336] • Wang, Chi, et al. "Flaml: A fast and lightweight automl library." Proceedings of Machine Learning and Systems 3 (2021): 434-447.
[0337] • Wagen, Aaron Z, et al. “Life course, genetic, and neuropathological associations with brain age in the 1946 British Birth Cohort: a population-based study.” The Lancet Healthy Longevity (2022) 3.9: E607-E616.
[0338] • Wrigglesworth, Jo, et al. “Brain-predicted age difference is associated with cognitive processing in later life.” Neurobiology of Aging (2022) 109: 195-203.
[0339] • Yan, Juni, et al. “Unlocking the potential: T1 -weighted MRI as a powerful predictor of levodopa response in Parkinson’s disease.” Insights into Imaging (2024) 15: 141.
[0340] • Zhang, Biao, et al. "Age-level bias correction in brain age prediction." NeuroImage: Clinical 37 (2023): 103319.
Claims
CLAIMS1 . A computer-implemented method for processing an image to determine a patient’s brain age, the method comprising: extracting at least one volumetric feature from an image of a brain by: obtaining at least one volume value for at least part of the patient’s brain, and normalising the at least one obtained volume value to obtain the at least one volumetric feature; and predicting the patient’s brain age by: inputting the at least one extracted volumetric feature into a pre-trained brain age model, wherein the brain age model is a linear regression model, and wherein a bias of the linear regression model is corrected when predicting the patient’s brain age.
2. The method as claimed in claim 1 further comprising stratifying the patient into a dementia risk group by inputting the patient’s predicted brain age into a classification model comprising at least one logistic regression binary classifier.
3. A computer-implemented method for stratifying patients into dementia risk groups, the method comprising: extracting at least one volumetric feature from an image of a brain by: obtaining at least one volume value for at least part of a patient’s brain, and normalising the at least one obtained volume value to obtain the at least one volumetric feature; predicting brain age by inputting the at least one extracted volumetric feature into a pretrained brain age model; and stratifying patients into dementia risk groups by inputting the predicted brain age into a classification model wherein the classification model comprises at least one logistic regression binary classifier.
4. The method as claimed in claim 3 wherein predicting brain age further comprises correcting a bias of the brain age model before inputting the predicted brain age into the classification model, wherein the pre-trained brain age model is a linear regression model.
5. The method as claimed in any one of claims 1 , 2 or 4 wherein a bias of the linear regression model is corrected by: obtaining an initial prediction by inputting the at least one extracted volumetric feature into the linear regression model; andbias correcting the initial prediction by subtracting an intercept of a regression line of the linear regression model and dividing by a slope of the regression line to obtain the predicted brain age which is input into the classification model.
6. The method as claimed in any one of claims 1 , 2 or 4 wherein a bias of the linear regression model is corrected by adding, to the linear regression model, penalty terms that counteract any bias.
7. The method as claimed in any preceding claim wherein obtaining at least one volume value for at least part of a patient’s brain comprises obtaining a volume value for at least one of hippocampus, ventricles, fusiform, medial temporal lobe, entorhinal cortex, fusiform gyrus and the whole brain, preferably wherein volume values are obtained for at least brainstem, 3rd-Ventricle, Left-choroid-plexus, Right-Caudate, white matter hypointensities, insular cortex left hemisphere, Right-VentralDC, Left-Thalamus, Left-Cerebellum-White-Matter and / or Left-Accumbens-Area.
8. The method as claimed in any of claims 1 to 6 wherein obtaining at least one volume value for at least part of a patient’s brain comprises obtaining a volume value representing a sum of values for left and right brain features.
9. The method as claimed in any preceding claim wherein extracting volumetric features from an image of a brain further comprises: partitioning the image into cortical structures and / or cortical substructures to obtain a partitioned image; labelling each partition in the partitioned image according to the corresponding cortical structure and / or cortical substructure; outputting an image showing the labelled cortical structures and / or cortical substructures, wherein the at least one value of a volume corresponds to at least one of the labelled cortical structures and / or substructures in the output image.
10. The method as claimed in claim 9 wherein partitioning the image comprises partitioning the image by finding a surface enclosing each of the cortical structure and / or cortical substructure.11 . The method as claimed in claim 9 wherein partitioning the image comprises partitioning the image by using a Convolutional Neural Network, CNN, to find the cortical structures and / or substructures.
12. The method as claimed in any preceding claim wherein determining a normalised value of the volume comprises obtaining a scaling function for scaling the volume value to obtain a normalised volume.
13. The method as claimed in claim 12 wherein obtaining a scaling function further comprises obtaining a scaling function according to statistical properties of a training dataset used to train the pre-trained model.
14. The method as claimed in any preceding claim wherein determining a normalised value of the volume comprises mapping the volume value to a pre-determined statistical distribution.
15. A computer-implemented method for stratifying patients into dementia risk groups using a pre-trained brain age deep learning, DL, model, the method comprising, for each patient: obtaining an image of a brain; processing the image of a brain using the brain age DL model to predict a brain age; and stratifying patients into dementia risk groups by inputting the predicted brain age into a classification model.
16. The method as claimed in claim 15 wherein the brain age DL model is any of a fully convolutional network, a 3D Residual Network and / or a 3D Convolutional Neural Network.
17. The method as claimed in claim 15 or claim 16 wherein the classification model comprises at least one binary classifier.
18. The method as claimed in any one of claims 2 to 14 and 17, wherein each binary classifier is a logistic regression model.
19. The method as claimed in any of claims 2 to 18 further comprising processing, using the classification model, the predicted brain age and / or a difference between a chronological age and the predicted brain age.
20. The method as claimed in any of claims 2 to 19 wherein stratifying patients into dementia risk groups using the classification model comprises using a first classification model for male patients and a second classification model for female patients.21 . The method as claimed in any of claims 2 to 19 wherein stratifying patients into dementia risk groups using the classification model comprises stratifying both male and female patients into dementia risk groups using a single classification model.
22. The method as claimed in any of claims 2 to 21 wherein the classification model has been trained using 10-fold cross validation and / or resampling prior to training.
23. The method as claimed in any of claims 2 to 22 comprising stratifying patients into dementia risk groups which indicate their current risk of dementia.
24. The method as claimed in any of claims 2 to 23 comprising stratifying patients into dementia risk groups which indicate their future risk of dementia.
25. The method as claimed in any preceding claim wherein the image is a magnetic resonance imaging, MRI, image of a brain.
26. The method as claimed in any of claims 2 to 25 wherein stratifying patients into dementia risk groups comprises stratifying patients into high, medium and / or low dementia risk groups.
27. The method as claimed in any preceding claim wherein the brain age model and / or the brain age DL model has been trained using brain images from patients that are between 55 and 90 years old, preferably wherein the patients are between 55 and 85 years old.
28. The method as claimed in any preceding claim wherein the brain age model and / or the brain age DL model has been trained using images of brains of individuals that were classified as healthy at the time the image of the brain was taken.
29. A method of treatment of a subject with a high risk of dementia comprising the steps of predicting a level of dementia risk, the method comprising: extracting at least one volumetric feature from an image of a brain by: obtaining at least one volume value for at least part of a patient’s brain, and normalising the at least one obtained volume value to obtain the at least one volumetric feature; predicting brain age by inputting the at least one extracted volumetric feature into a pretrained brain age model; determining a dementia risk of the subject based on the predicted brain age; selecting a treatment according to the determined dementia risk; and administering the treatment.
30. A method of therapy monitoring in a subject receiving treatment for dementia, the method comprising: extracting at least one volumetric feature from an image of a brain by: obtaining at least one volume value for at least part of a patient’s brain, and normalising the at least one obtained volume value to obtain the at least one volumetric feature;predicting brain age by inputting the at least one extracted volumetric feature into a pretrained brain age model; and determining a dementia risk of the subject based on the predicted brain age; determining, using the dementia risk, whetherthere is an improvement in the dementia risk compared to a reference dementia risk.
31. A computer-implemented method for training a brain age model to stratify patients into dementia risk groups, the method comprising: obtaining a set of brain images and a set of age labels corresponding to each image in the set of brain images; extracting at least one volumetric feature from each brain image by: obtaining at least one volume value for at least part of a patient’s brain, and normalising the at least one obtained volume value to obtain the at least one volumetric feature; predicting brain age by inputting the at least one extracted volumetric feature into the brain age model; comparing the predicted brain age with the brain age label and using the comparison to train the brain age model; and training a risk prediction model to stratify patients into dementia risk groups based on the brain age prediction by the brain age model.
32. The method as claimed in claim 30 further comprising correcting for bias in the brain age model by: determining a slope and an intercept of a regression line between predicted brain age and age label for the set of training images; and bias correcting the predicted brain age before comparing the predicted brain age with the brain age label.
33. The method of claim 31 or 32 wherein obtaining a set of brain images and a set of brain age labels comprises obtaining a set of cognitively normal brain images and corresponding set of age labels.
34. A computer-readable storage medium comprising instructions which, when executed by a processor, causes the processor to carry out the method of any of claims 1 to 28 or any of claims 31 to 33.
35. An image processing system comprising: an image capture device which is configured to capture an image of a patient’s brain;an image processor which is configured to receive an image from the image capture device and carry out the method of any of claims 1 to 28 and of claims 31 to 33; and a user interface which is configured to display an output result generated by the image processor.
Citation Information
Patent Citations
Method of evaluating concomitant clinical dementia rating and its future outcome using predicted age difference and program thereof
US20230086483A1