System and method for determining a dementia prediction indicator based on a candidate image
An AI-based system automates the scoring of clock drawing tests by preprocessing and cropping images, addressing clinician workload and non-standardized data collection issues, enhancing the efficiency of cognitive impairment assessment and regulatory evaluations.
Patent Information
- Application Number
- PCT/IB2025/054962
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-24
- Filing Date
- 2025-05-12
- Publication Date
- 2025-11-27
AI Technical Summary
The existing methods for assessing cognitive impairment through clock drawing tests face challenges such as high clinician workload, non-standardized data collection, and sub-optimal image capture, which hinder efficient processing and scoring, particularly in regulatory contexts like driver licensing.
An AI-based system and method for automating the scoring of drawing tests, including clock drawing tests, using machine learning models to generate a dementia prediction indicator by preprocessing and cropping images, leveraging a dementia prediction model and an image-cropping model to analyze candidate images.
Facilitates efficient and standardized assessment of cognitive impairment, reducing clinician burden and wait times, enabling early access to healthcare services and regulatory evaluations.
Smart Images

Figure IB2025054962_27112025_PF_FP_ABST
Abstract
Description
SYSTEM AND METHOD FOR DETERMINING A DEMENTIA PREDICTION INDICATOR BASED ON A CANDIDATE IMAGECROSS-REFERENCE TO RELATED APPLICATIONS
[0001] The present application claims priority to US provisional application 63 / 651 ,856 filed May 24th, 2024 the entire contents of which are incorporated herein by reference.FIELD
[0002] The present invention relates to computer-implemented systems and methods of drawing test classification and prediction, and specifically, to predicting cognitive disorder indicators through the evaluation of images by machine learning models and training said models.INTRODUCTION
[0003] Dementia is a rapidly growing problem in today’s society. The number of people with dementia worldwide is expected to triple to 150 million by 2050.
[0004] Conventionally, the identification of dementia is based heavily on cognitive testing. Numerous different cognitive test assessments exist, including the Toronto Cognitive Assessment (TorCA), which is a 30-40-minute cognitive assessment tool designed to detect mild cognitive impairment to early / mild stages of dementia The TorCA includes cognitive tests of orientation, immediate verbal recall, delayed verbal and visual recall, delayed verbal and visual recognition, visuospatial function, working memory / attention / executive control, and language.
[0005] Another cognitive assessment is the Behavioral Neurology Assessment - Short Form (BNA-SF) which is a test that overlaps with the TorCA and is a 20-30-minute cognitive assessment tool designed to detect moderate to severe cognitive impairment.
[0006] A major component of the TorCA and other cognitive assessment tools is the clock drawing test. The clock drawing test is a cognitive test that can help with the evaluation of how well the brain is working. The clock drawing test can be performed in under a minute, and since it taps into many cognitive functions (e.g., planning, organization, attention, visuospatial function, memory, language), it is possible to obtain a variety of information from it. An individual may be administered this test as part of a larger cognitive assessment battery, which may be conducted, for example, when theindividual (or someone close to them) notices a change in their cognition (e.g., memory loss, trouble finding words) or behavior (e.g., disinhibition).
[0007] Free drawn clock drawing tests (CDTs) are considered to be more sensitive than other solutions such as digital clock drawing tests. In this test, a person is given a blank piece of paper and asked to 1) draw the face of a clock, 2) put in all the numbers, and 3) set the hands to a certain time. The time setting is extremely important; various times are used, but ten after eleven is most common. In order to draw this time correctly, the ten needs to be recoded to a two for the minute hand.
[0008] Clock drawing tests have been in use since the 1900s and are well known. Several variations of clock drawing tests are known. In one example, free drawn clocks may be used, in which a clock contour, numbers on the clock face, and clock hands are all hand-drawn by a participant. In another example, pre-drawn clocks, may be used, in which the clock contour is provided but numbers and hands are drawn by the individual being tested. In yet another example, examiner clocks may be used, in which the contour and numbers are provided, and only the hands need to be drawn. In other embodiments, parts or all of the tests may be administered via digital means. The variations of the clock drawing test may all generally be used for the assessment of cognitive impairment, albeit with variations to the degrees of freedom available. A variety of scoring systems may be used to assess cognitive function based on the drawn clocks.
[0009] Through examining a person’s clock, a clinician can identify various cognitive impairments, or a lack thereof. Analyzing the clock involves the skill of looking at what the person does to draw their clock (i.e., the process), and then figuring out what cognitive functions go into the process. Different types of errors point to different problems with brain functions. The problems most commonly identified using the clock drawing test are executive dysfunctions (e.g., with planning, attention, perseveration) and visuospatial deficits (e.g., inability to identify visual and spatial relationships among objects).
[0010] The clock drawing test provides neurologists and other clinicians with a lot of information in a very short time. Brevity is hugely beneficial in cognitive testing. A lot goes into drawing a clock, including planning, attention, monitoring, language, memory, and visuospatial functions among others. Individual diagnosis not only for dementia, but for a wide range of brain disorders.
[0011] One application of the clock drawing test is as a standardized test for activities that are particularly dangerous involving individuals with dementia and other brain disorders. For example, the clock drawing test is used in regulatory circumstances related to driver licensing for automobiles. In this way, road safety may be improved. The skilled clinicians required to analyze and score clock drawings are in high demand, as there are generally insufficient numbers of qualified clinicians to perform scoring in a timeefficient manner. Due to these challenges, individuals may face lengthy wait times for an initial assessment.
[0012] Clinical uses of the clock drawing test, as well as other regulatory uses such as for driver licensing, however, suffer from several challenges. These include the large number of clock drawing tests that must be scored by a clinician. In particular, in regulatory applications, the number of clock drawing tests may be extremely high and may burden the specialized clinicians required who are tasked with scoring them.
[0013] The collection of clock drawing tests is conventionally done on paper (as noted above - it is preferable to use free drawn clock drawing tests). The collection of these test results may occur in many different clinical settings. For example, the testing environment may not be standardized in terms of lighting and image collection of the free drawn clocks. An administrator who collects the data may take a picture using a mobile device that may skew, rotate, or stretch the collected image of the free drawn clock. The lighting may also be sub-optimal.
[0014] Other drawing tests exist, including for example the Benson Complex Figure Test. The Benson test is an individually administered that test assesses visuospatial function, visual memory, and executive abilities, allowing the detection of multiple mechanisms of cognitive impairment. Such tests may be used in similar applications as clock drawing tests such as driver licensing.
[0015] The collection of images associated with drawing tests therefore raises technical challenges related to the incoming data generated by the person that is required for assessment or scoring which require a technical solution.
[0016] There is a need therefore for improved systems and methods that can streamline the processing of drawing tests as part of an individual’s care, and as part of other regulatory mechanisms that the score may be used for such as driver licensing.SUMMARY
[0017] Provided are improved systems and methods that automate and improve the scoring of drawing tests, including clock drawing tests. The vision is to create a future in healthcare where artificial intelligence (Al) plays a major role in aiding diagnosis and treatment, similar to the medical tricorder used in Star Trek and provide subspecialty medical expertise to any physician at all corners of this world at any time of day. This vision further extends the analogy to the medical tricorder to allow other, non-skilled clinicians or even users to access early predictions related to cognitive impairment.
[0018] This project represents a step towards the vision and is focused on a rapidly growing problem in today’s society, i.e. , dementia and related disorders. The number of people with dementia worldwide is expected to triple to 150 million by 2050. However, there are insufficient numbers of clinicians to meet this challenge. To partially address this problem, the development of an Al-based tool is proposed to determine potentially required services for individuals who are facing lengthy wait times prior to an initial assessment and well in advance of this initial assessment. For example, the initial assessment may occur at the time of referral. The initial assessment may include as a component the drawing tests, such as clock drawing tests, and as described herein.
[0019] The Toronto Dementia Research Alliance (TDRA) has developed a unique participant database platform that embeds research in clinical care through collection of real-world characterization for each participant and provides a dataset representative of common clinical practice. This platform serves to maximize clinical assessment and management of persons with cognitive impairment and to facilitate collaborative multi- institutional research in dementia and related disorders.
[0020] The TDRA platform provides real time electronic capture of data at point of care and offers up-to-date access of data through an electronic dashboard. This platform captures the following:- A comprehensive standardized intake form for diagnostic assessment of participants.- Cognitive assessments administered on an iPad, i.e., Toronto Cognitive Assessment (TorCA) for mild-to-moderate cognitive complaints and the BehavioralNeurology Assessment - Short Form (BNA-SF) for moderate to severe cognitive impairment, including free drawn clock drawing tests.- Results of medical, psychiatric, and neurological examinations.- Results of ancillary testing such as neuroimaging.- Diagnosis and management plan based on expert opinion and published guidelines.
[0021] The present embodiments provide Al systems and methods to triage individuals and determine if key services are required that can be accessed without clinician intervention. This may include processing portions of the initial assessment such as drawing tests including clock drawing tests. The Al systems and methods may operate under the supervision of a clinician. The triage capability may facilitate and expedite early access to healthcare services while clients are waiting for their physician appointment. By leveraging a research database, machine learning can be used to identify clinical variables in our database to develop an expert system that can triage individuals. For example, this may be done at the time of referral.
[0022] Individuals referred to memory clinics demonstrate the full spectrum of cognitive disorders, from subjective cognitive decline and mild cognitive impairment to dementia. Potential urgent services required - prior to the formal memory clinic assessment - include occupational therapy for home safety assessments, social work for urgent long-term care placement, falls prevention clinics for gait disorders, audiology for hearing loss and driving assessments. A diagnosis is often not required for services like these that can be accessed prior to the assessment. In some cases, the triage system may provide efficient care by eliminating the need for an assessment since only the triage services may be required. The expert Al system may be scalable to other specialty clinics, and to primary care. The Al system may further be applied in regulatory contexts such as driver licensing, or other situations where individuals are assessed for cognitive function.
[0023] In a first aspect there is provided a computer-implemented method for generating a dementia prediction indicator based on a candidate image received, the method comprising: providing, at a memory, a dementia prediction model and an imagecropping model for cropping an input image for the dementia prediction model; receiving, at a processor in communication with the memory, the candidate image; generating, atthe processor, a cropped candidate image based on the candidate image and the imagecropping model; and generating, at the processor, the dementia prediction indicator based on the cropped candidate image and the dementia prediction model.
[0024] In at least one embodiment, the method may further comprise: receiving, at the processor, one or more prediction features, and wherein the generating the dementia prediction indicator may further be based on the one or more prediction features.
[0025] In at least one embodiment, the one or more prediction features may comprise at least one of: a socio-economic indicator, a driving test score, and a prior medical history.
[0026] In at least one embodiment, the candidate image may comprise an image generated based on a standards-based test.
[0027] In at least one embodiment, the candidate image may comprise a plurality of different shapes.
[0028] In at least one embodiment, the candidate image may comprise a clock drawing test.
[0029] In at least one embodiment, the candidate image may be captured using a user device camera in the clinical environment.
[0030] In at least one embodiment, the generating of the cropped candidate image based on the candidate image and the image-cropping model may comprise: processing the candidate image using the image-cropping model to produce a heatmap of a probable location of a clock relative to the candidate image; cropping the candidate image to produce an intermediate candidate image such that a minimum proportion of the heatmap is contained within the intermediate candidate image; and cropping the intermediate candidate image to produce the cropped candidate image such that a minimum proportion of a total pixel value sum of the intermediate candidate image is contained within the cropped candidate image.
[0031] In at least one embodiment, the method may further comprise: preprocessing the candidate image at the processor, the pre-processing comprising: converting, at the processor, the candidate image to a greyscale candidate image; applying, at the processor, one or more line filters to the greyscale candidate image toproduce a plurality of filtered images, the plurality of filtered images comprising at least one selected from the group of: a first filtered image, a second filtered image, and a third filtered image; and generating, at the processor, the pre-processed candidate image by averaging the plurality of filtered images based on a line weight associated with each of the filtered images.
[0032] In at least one embodiment, the method may further comprise padding the cropped candidate image prior to processing the cropped image using the dementia prediction model.
[0033] In at least one embodiment, each of the one or more line filters may be applied to the greyscale candidate image in accordance with one or more pixel values associated with the each of the one or more line filters.
[0034] In at least one embodiment, the image-cropping model may comprise: at least one downsample layer; at least one layer receiving the output of the at least one downsample layer; and at least one upsample layer receiving the output of the at least one layer.
[0035] In at least one embodiment, the at least one downsample layer may comprise at least one convolution downsample layer; the at least one layer comprises at least one convolution layer; and the at least one upsample layer comprises at least one convolution upsample layer.
[0036] In at least one embodiment, the dementia prediction model may comprise: at least one median fully connected layer, the median fully connected layer comprising an equal number of inputs and outputs; and an output layer connected to the output of the at least one median fully connected layer, the output layer comprising two outputs.
[0037] In at least one embodiment, the dementia prediction indicator may comprise a dementia prediction and a clock drawing test measure.
[0038] In at least one embodiment, the image-cropping and the dementia prediction models may comprise machine-learning models.
[0039] In at least one embodiment, the dementia prediction model may further comprise a transformer.
[0040] In a second aspect there is provided a system for generating a dementia prediction indicator based on a candidate image, the system comprising: a memory comprising: a dementia prediction model; and an image-cropping model for cropping an input image for the dementia prediction model; and a processor in communication with the memory, the processor configured to: receive the candidate image; generate a cropped candidate image based on the candidate image and the image-cropping model; and generate the dementia prediction indicator based on the cropped candidate image and the dementia prediction model.
[0041] In at least one embodiment, the processor may be further configured to receive one or more prediction features; and the processor generating the dementia predictor is further based on the one or more prediction features.
[0042] In at least one embodiment, the one or more prediction features may comprise at least one of: a socio-economic indicator, a driving test score, and a prior medical history.
[0043] In at least one embodiment, the candidate image may comprise an image generated based on a standards-based test.
[0044] In at least one embodiment, the candidate image may comprise a plurality of different shapes.
[0045] In at least one embodiment, the candidate image may comprise a clock drawing test.
[0046] In at least one embodiment, the candidate image may be captured using a user device camera in the clinical environment.
[0047] In at least one embodiment, the processor may be further configured to generate the cropped candidate image based on the candidate image and the imagecropping model by: processing the candidate image using the image-cropping model to produce a heatmap of a probable location of a clock relative to the candidate image; cropping the candidate image to produce an intermediate candidate image such that a minimum proportion of the heatmap is contained within the intermediate candidate image; and cropping the intermediate candidate image to produce the cropped candidate image such that a minimum proportion of a total pixel value sum of the intermediate candidate image is contained within the cropped candidate image.
[0048] In at least one embodiment, the processor may be further configured to pre- process the candidate image at the processor, the pre-processing comprising: converting, at the processor, the candidate image to greyscale candidate image; applying, at the processor, one or more line filters to the greyscale candidate image to produce a plurality of filtered images, the plurality of filtered images comprising at least one selected from the group of: a first filtered image, a second filtered image, and a third filtered image; and generating, at the processor, the pre-processed candidate image by averaging the set of filtered images based on a line weight associated with each of the filtered images.
[0049] In at least one embodiment, the processor may further be configured to pad the cropped candidate image prior to processing the cropped image using the dementia prediction model.
[0050] In at least one embodiment, each of the one or more line filters may be applied to the greyscale candidate image in accordance with one or more pixel values associated with the each of the one or more line filters.
[0051] In at least one embodiment, the image-cropping model may comprise: at least one downsample layer; at least one layer receiving the output of the at least one downsample layer; and at least one upsample layer receiving the output of the at least one layer.
[0052] In at least one embodiment, the at least one downsample layer may comprise at least one convolution downsample layer; the at least one layer comprises at least one convolution layer; and the at least one upsample layer comprises at least one convolution upsample layer.
[0053] In at least one embodiment, the dementia prediction model may comprise: at least one median fully connected layer, the median fully connected layer comprising an equal number of inputs and outputs; and an output layer connected to the output of the at least one median fully connected layer, the output layer comprising two outputs.
[0054] In at least one embodiment, the dementia prediction indicator may comprise a dementia prediction and a clock drawing test measure.
[0055] In at least one embodiment, the image-cropping and the dementia prediction models may comprise machine-learning models.
[0056] In at least one embodiment, the dementia prediction model may further comprise a transformer.
[0057] In a third aspect, there is provided a computer-implemented method for generating a dementia prediction model, comprising: receiving, at a processor, a plurality of training images; generating, at the processor, a plurality of cropped training images based on an image-cropping model and the plurality of training images; and fine-tuning, at the processor, a machine learning model using the cropped training images and at least one feature associated with each of the cropped training images, the machine learning model comprising a pre-trained model.
[0058] In at least one embodiment, the at least one feature may comprise at least one of: a manual label, a socio-economic indicator, a driving test score, and a prior medical history.
[0059] In at least one embodiment, the manual label may comprise a clock drawing test measure.
[0060] In at least one embodiment, the method may further comprise: preprocessing, at the processor, each training image in the plurality of training images at the processor prior to generating the plurality of cropped training images, the pre-processing comprising: converting, at the processor, the training image to a greyscale training image; applying, at the processor, one or more line filters to the greyscale training image to produce a plurality of filtered images, the plurality of filtered images comprising at least one selected from the group of: a first filtered image, a second filtered image, and a third filtered image; and generating, at the processor, the pre-processed training image by averaging the set of filtered images based on a line weight associated with each of the filtered images.
[0061] In at least one embodiment, each of the one or more line filters may be applied to the greyscale training image in accordance with one or more pixel values associated with the each of the one or more line filters.
[0062] In at least one embodiment, the method may further comprise: dividing, at the processor, the plurality of training images into a training data set and a validation data set.
[0063] In at least one embodiment, the method may further comprise: replicating, at the processor, one or more images associated with one or more image groups in the validation data set and the training data set, each image group being associated with a value of a manual label, until all of the image groups contain an equal quantity of samples.
[0064] In at least one embodiment, the manual label may comprise a clock drawing test measure.
[0065] In at least one embodiment, the method may further comprise: augmenting, at the processor, each training image in the plurality of training images by performing at least one of the following: scaling the training image; translating the training image; rotating the training image; and shearing the training image.
[0066] In at least one embodiment, the fine-tuning the machine learning model may comprise: modifying, at the processor, only one or more weights associated with a head of the machine learning model until a mean squared error associated with the validation data set has plateaued; and modifying, at the processor, one or more weights associated with a body of the machine learning model and the one or more weights associated with the head of the machine learning model, the body of the machine learning model comprising the pre-trained model.
[0067] In at least one embodiment, the pre-trained model may comprise a transformer.
[0068] In at least one embodiment, the image-cropping model may comprise: at least one downsample layer; at least one layer receiving the output of the at least one downsample layer; and at least one upsample layer receiving the output of the at least one layer.
[0069] In at least one embodiment, the at least one downsample layer may comprise at least one convolution downsample layer; the at least one layer comprises at least one convolution layer; and the at least one upsample layer comprises at least one convolution upsample layer.
[0070] In at least one embodiment, the pre-trained model may further comprise: at least one median fully connected layer, the median fully connected layer comprising an equal number of inputs and outputs; and an output layer connected to the output of the at least one median fully connected layer, the output layer comprising two outputs.
[0071] In at least one embodiment, the plurality of training images may comprise drawings of a clock.
[0072] In at least one embodiment, the method may further comprise generating the image-cropping model by: receiving an image-cropping dataset, the image-cropping dataset comprising images comprising clock drawing tests; generating a manually cropped training dataset based on the image-cropping dataset; generating a pre- processed dataset based on the manually cropped dataset; producing a first target bitmap dataset based on the pre-processed dataset by setting one or more first target pixels of the first target bitmap to a 1 value where the target pixel comprises a part of a clock image and 0 otherwise; producing a second target bitmap dataset based on the pre-processed dataset by setting one or more second target pixels of the second target bitmap to 0.2, 0.4, 0.6, 0.8, or 1 , based on a clock rating of the clock drawing test; and training, at the processor, a cropping machine learning model based on the image-cropping dataset, the first target bitmap, and the second target bitmap.
[0073] In at least one embodiment, the method may further comprise dividing the manually cropped training dataset into a cropping training dataset, a cropping validation dataset, and a cropping testing dataset.
[0074] In at least one embodiment, the method may further comprise augmenting, at the processor, each image in the cropping training dataset, the cropping validation dataset, and the cropping testing dataset by at least one of: adding noise to the image; adding lines to the image; adding writing to the image; and adding rectangles in a random orientation to the image.
[0075] In at least one embodiment, the generating a pre-processed dataset based on the manually cropped training dataset may comprise, for each image in the manually cropped training dataset: rotating the image by a random angle; resizing the image by a random factor; and placing the image onto a black background image.
[0076] In a fourth aspect there is provided a system for generating a dementia prediction model, comprising: a memory comprising a machine learning model; a processor configured to: receive a plurality of training images; generate, at a processor, a plurality of cropped training images based on an image-cropping model and the plurality of training images; and fine-tune the machine learning model using the cropped trainingimages and at least one feature associated with each of the cropped training images, the machine learning model comprising a pre-trained model.
[0077] In at least one embodiment, the at least one feature may comprise at least one of: a manual label, a socio-economic indicator, a driving test score, and a prior medical history.
[0078] In at least one embodiment, the manual label may comprise a clock drawing test measure.
[0079] In at least one embodiment, the processor may be further configured to: pre- process each training image in the plurality of training images at the processor prior to generating the plurality of cropped training images, the pre-processing comprising: converting the training image to a greyscale training image; applying one or more line filters to the greyscale training image to produce a plurality of filtered images, the plurality of filtered images comprising at least one selected from the group of: a first filtered image, a second filtered image, and a third filtered image; and generating the pre-processed training image by averaging the set of filtered images based on a line weight associated with each of the filtered images.
[0080] In at least one embodiment, each of the one or more line filters may be applied to the greyscale training image in accordance with one or more pixel values associated with the each of the one or more line filters.
[0081] In at least one embodiment, the processor may be further configured to divide, at the processor, the plurality of training images into a training data set and a validation data set.
[0082] In at least one embodiment, the processor may be further configured to: replicate one or more images associated with one or more image groups in the validation data set and the training data set, each image group being associated with a value of a manual label, until all of the image groups contain an equal quantity of samples.
[0083] In at least one embodiment, the manual label may comprise a clock drawing test measure.
[0084] In at least one embodiment, the processor may be further configured to: augment each training image in the plurality of training images, prior to fine-tuning thepre-trained model, by performing at least one of the following: scaling the training image; translating the training image; rotating the training image; and shearing the training image.
[0085] In at least one embodiment, the fine-tuning the pre-trained model may comprise: modifying only one or more weights associated with a head of the machine learning model until a mean squared error associated with the validation data set has plateaued; and modifying one or more weights associated with a body of the machine learning model and the one or more weights associated with the head of model, the body of the machine learning model comprising the pre-trained model.
[0086] In at least one embodiment, the pre-trained model may comprise a transformer.
[0087] In at least one embodiment, the image-cropping model may comprise: at least one downsample layer; at least one layer receiving the output of the at least one downsample layer; and at least one upsample layer receiving the output of the at least one layer.
[0088] In at least one embodiment, the at least one downsample layer may comprise at least one convolution downsample layer; the at least one layer may comprise at least one convolution layer; and the at least one upsample layer comprises at least one convolution upsample layer.
[0089] In at least one embodiment, the head of the machine learning model may further comprise: at least one median fully connected layer, the median fully connected layer comprising an equal number of inputs and outputs; and an output layer connected to the output of the at least one median fully connected layer, the output layer comprising two outputs.
[0090] In at least one embodiment, the plurality of training images may comprise drawings of a clock.
[0091] In at least one embodiment, the processor may be further configured to generate the image-cropping model by training a cropping machine learning model based on an image-cropping dataset, a first target bitmap, and a second target bitmap, wherein: the image-cropping dataset comprises images comprising clock drawing tests; the first target bitmap dataset is produced based on a pre-processed dataset by setting one or more first target pixels of the first target bitmap to a 1 value where the target pixelcomprises a part of a clock image and 0 otherwise; and the second target bitmap dataset is produced based on the pre-processed dataset by setting one or more second target pixels of the second target bitmap to 0.2, 0.4, 0.6, 0.8, or 1 , based on a clock rating of the clock drawing test; a manually cropped training dataset is based on the image-cropping dataset; and the pre-processed dataset is based on a manually cropped dataset, the manually cropped dataset being based on the image-cropping dataset.
[0092] In at least one embodiment, the manually cropped training dataset may be divided into a cropping training dataset, a cropping validation dataset, and a cropping testing dataset.
[0093] In at least one embodiment, the processor may be further configured to augment each image in the cropping training dataset, the cropping validation dataset, and the cropping testing dataset by at least one of: adding noise to the image; adding lines to the image; adding writing to the image; and adding rectangles in a random orientation to the image.
[0094] In at least one embodiment, the generating a pre-processed dataset based on the manually cropped training dataset may comprise, for each image in the manually cropped training dataset: rotating the image by a random angle; resizing the image by a random factor; and placing the image onto a black background image.DRAWINGS
[0095] FIG. 1 shows a system diagram of an example environment in which the presently disclosed systems and methods may be used in accordance with one or more embodiments.
[0096] FIG. 2 shows a device diagram of an example processing device in accordance with one or more embodiments.
[0097] FIG. 3 shows a model diagram of an example dementia prediction model architecture in accordance with one or more embodiments.
[0098] FIGs. 4A - 4B show another model diagram of an example image-cropping model architecture in accordance with one or more embodiments.
[0099] FIG. 5 is a method diagram for generating a dementia prediction indicator in accordance with one or more embodiments.
[0100] FIG. 6 is a method diagram for generating a dementia prediction model in accordance with one or more embodiments.
[0101] FIG. 7A shows example augmented clock drawing test images in accordance with one or more embodiments.
[0102] FIG. 7B shows target bitmaps of the example augmented clock drawing test images of FIG. 7A in accordance with one or more embodiments.
[0103] FIG. 7C shows heatmaps of the example augmented clock drawing test images of FIG. 7A in accordance with one or more embodiments.
[0104] FIG. 8 shows example clock drawing tests and corresponding cropped and pre-processed images in accordance with one or more embodiments.
[0105] FIG. 9 shows graphs of mean squared error over time during training of the dementia prediction model in accordance with one or more embodiments.
[0106] FIG. 10 shows example receiver operator characteristics (ROC) of various methods of predicting dementia in accordance with one or more embodiments.DESCRIPTION OF VARIOUS EMBODIMENTS
[0107] Various apparatuses or methods will be described below to provide an example of the claimed subject matter. No example described below limits any claimed subject matter and any claimed subject matter may cover methods or apparatuses that differ from those described below. The claimed subject matter is not limited to apparatuses or methods having all of the features of any one apparatus or methods described below or to features common to multiple or all of the apparatuses or methods described below. It is possible that an apparatus or methods described below is not an example that is recited in any claimed subject matter. Any subject matter disclosed in an apparatus or methods described below that is not claimed in this document may be the subject matter of another protective instrument, for example, a continuing patent application, and the applicants, inventors or owners do not intend to abandon, disclaim or dedicate to the public any such invention by its disclosure in this document.
[0108] Furthermore, it will be appreciated that for simplicity and clarity of illustration, where considered appropriate, reference numerals may be repeated among the figures to indicate corresponding or analogous elements. In addition, numerous specific detailsare set forth in order to provide a thorough understanding of the examples described herein. However, it will be understood by those of ordinary skill in the art that the examples described herein may be practiced without these specific details. In other instances, well- known methods, procedures and components have not been described in detail so as not to obscure the examples described herein. Also, the description is not to be considered as limiting the scope of the examples described herein.
[0109] It should also be noted that the terms “coupled”, or “coupling” as used herein can have several different meanings depending on the context in which these terms are used. For example, the terms “coupled”, or “coupling” can have a mechanical, electrical, or communicative connotation. For example, as used herein, the terms “coupled”, or “coupling” can indicate that two elements or devices can be directly connected to one another or connected to one another through one or more intermediate elements or devices via an electrical element, electrical signal or a mechanical element depending on the particular context. Furthermore, the term “communicative coupling” indicates that an element or device can electrically, optically, or wirelessly send data to another element or device as well as receive data from another element or device.
[0110] It should also be noted that, as used herein, the wording “and / or” is intended to represent an inclusive-or. That is, “X and / or Y” is intended to mean X or Y or both, for example. As a further example, “X, Y, and / or Z” is intended to mean X or Y or Z or any combination thereof.
[0111] It should be noted that terms of degree such as "substantially", "about" and "approximately" as used herein mean a reasonable amount of deviation of the modified term such that the end result is not significantly changed. These terms of degree may also be construed as including a deviation of the modified term if this deviation would not negate the meaning of the term it modifies.
[0112] Furthermore, the recitation of numerical ranges by endpoints herein includes all numbers and fractions subsumed within that range (e.g., 1 to 5 includes 1 , 1.5, 2, 2.75, 3, 3.90, 4, and 5). It is also to be understood that all numbers and fractions thereof are presumed to be modified by the term "about" which means a variation of up to a certain amount of the number to which reference is being made if the end result is not significantly changed.
[0113] Some elements herein may be identified by a part number, which is composed of a base number followed by an alphabetical or subscript-numerical suffix (e.g., 112a, or 112i). Multiple elements herein may be identified by part numbers that share a base number in common and that differ by their suffixes (e.g., 112i, 1122, and 112s). All elements with a common base number may be referred to collectively or generically using the base number without a suffix (e.g., 112).
[0114] The example systems and methods described herein may be implemented in hardware or software, or a combination of both. In some cases, the examples described herein may be implemented, at least in part, by using one or more computer programs, executing on one or more programmable devices comprising at least one processing element, a data storage element (including volatile and non-volatile memory and / or storage elements), and at least one communication interface. These devices may also have at least one input device (e.g., a keyboard, a mouse, a touchscreen, and the like), and at least one output device (e.g., a display screen, a printer, a wireless radio, and the like) depending on the nature of the device. For example, and without limitation, the programmable devices (referred to below as computing devices) may be a server, network appliance, embedded device, computer expansion module, a personal computer, laptop, personal data assistant, cellular telephone, smart-phone device, tablet computer, a wireless device or any other computing device capable of being configured to carry out the methods described herein.
[0115] In some examples, the communication interface may be a network communication interface. In examples in which elements are combined, the communication interface may be a software communication interface, such as those for inter-process communication (IPC). In still other examples, there may be a combination of communication interfaces implemented as hardware, software, and a combination thereof.
[0116] Program code may be applied to input data to perform the functions described herein and to generate output information. The output information is applied to one or more output devices, in known fashion.
[0117] Each program may be implemented in a high-level procedural, declarative, functional or object-oriented programming and / or scripting language, or both, to communicate with a computer system. However, the programs may be implemented inassembly or machine language, if desired. In any case, the language may be a compiled or interpreted language. Each such computer program may be stored on a storage media or a device (e.g., ROM, magnetic disk, optical disc) readable by a general or special purpose programmable computer, for configuring and operating the computer when the storage media or device is read by the computer to perform the procedures described herein. Examples of the system may also be considered to be implemented as a non- transitory computer-readable storage medium, configured with a computer program, where the storage medium so configured causes a computer to operate in a specific and predefined manner to perform the functions described herein.
[0118] Furthermore, the example system, processes and methods are capable of being distributed in a computer program product comprising a computer readable medium that bears computer usable instructions for one or more processors. The medium may be provided in various forms, including one or more diskettes, compact disks, tapes, chips, wireline transmissions, satellite transmissions, internet transmission or downloads, magnetic and electronic storage media, digital and analog signals, and the like. The computer useable instructions may also be in various forms, including compiled and noncompiled code.
[0119] Various examples of systems, methods and computer programs products are described herein. Modifications and variations may be made to these examples without departing from the scope of the invention, which is limited only by the appended claims. Also, in the various user interfaces illustrated in the figures, it will be understood that the illustrated user interface text and controls are provided as examples only and are not meant to be limiting. Other suitable user interface elements may be used with alternative implementations of the systems and methods described herein.
[0120] Herein, the terms subject, patient, participant, and individual may be used interchangeably, but refer to the individual who completed the drawing test. The participant or patient may be a user of a mobile device collecting the image of the drawing test or may be another individual.
[0121] Herein, the term drawing test can include clock drawing tests or other drawn tests that are known in the art. Clock drawing tests (CDTs) have been in use since the 1900s and are well known. Several variations of clock drawing tests are known, for example, free-drawn clocks (FD), which require participants to draw the contour,numbers, hands, and center of the clock; pre-drawn clocks (PD), in which numbers, hands, and center are to be drawn in in a pre-drawn contour; examiner-drawn clock (ED), in which only hands and center are to be inserted in a template including a pre-drawn contour and numbers.
[0122] In other embodiments, parts or all of the tests may be administered via digital means. The variations of the drawing tests may all generally be used for the assessment of cognitive impairment, albeit with variations to the degrees of freedom available.
[0123] As well, different variations of clock drawing test scoring systems are known. These may vary in administration procedures, instructions (e.g., with or without a predrawn circle, different time settings), and scoring methods, for example, the Freedman Clock drawing test.
[0124] Artificial intelligence (Al) may play a major role in aiding diagnosis and treatment of conditions in the future. Al may enable the provision of subspeciality medical expertise to any physician at all corners of the world at any time of day. The systems and methods described herein may address the challenges posed by the continual accumulation of vast amounts of scientific medical information by incorporating the new knowledge into its algorithms, such as biomarkers.
[0125] Patients referred to memory clinics can demonstrate the full spectrum of cognitive disorders, from subjective cognitive decline and mild cognitive impairment to dementia. Potential urgent services required- prior to the formal memory clinic assessment - can include occupational therapy for home safety assessments, social work for urgent long-term care placement, falls prevention clinics for gait disorders, and audiology for hearing loss. A diagnosis is often not required for these services that can be accessed prior to the assessment. In some cases, the triage system may provide efficient care by eliminating the need for an assessment since only the triage services may be required.
[0126] The systems and methods described herein may, in one example usage, form a part of an Al system to automatically triage patients and determine if key services are required that can be accessed without physician or nurse intervention. This automated triage capability may facilitate and expedite early access to healthcare services whileclients await their specialist appointment. Existing subject databases may be leveraged by using machine learning techniques to develop an expert system that can triage patients accurately. Such systems may be scalable to other specialty clinics and to primary care.
[0127] In an alternate example, the systems and methods provided herein may form part of a software system related to regulatory approval for an individual such as a driver licensing system. The systems and methods may operate in this case to provide improved and efficient evaluation of drawing test drawings submitted in association with a regulatory application such as for a driver’s license, for example, clock drawing tests.
[0128] However, the systems and methods described herein are in no way limited only to such usages.
[0129] The systems and methods described herein may be a part of an innovative model to facilitate precision medicine that will be scalable across specialty and primary care clinics.
[0130] Reference is first made to FIG. 1 , which shows a system diagram 100 of an example environment 120 in which systems and methods for generating a dementia prediction indicator based on a candidate image received may be used. The environment 120 includes a user 102, a candidate image 104, and a user device 106. The system includes the user 102, the candidate image 104, the user device 106, a network 108, and a processing device 200 such as a server. The environment 120 may be a clinical environment such as a hospital, a clinic, a retirement, nursing home, or residence in the community. Alternatively, the environment 120 may be a regulatory location such as a drivers testing centre. User 102 may upload candidate image 104 using user device 106 through network 108 to be processed by processing device 200.
[0131] Processing device 200 may receive candidate image 104, process the image, and return a dementia prediction indicator based on the processing thereof. Alternatively, a pre-trained model may be sent from processing device 200 to the user device 106 and the user device 106 may perform the processing of image 104 in the clinical environment 120. The methods associated with prediction of a dementia indicator may be performed either at device 200 or user device 106.
[0132] User 102 may be a clinician, a test participant, a patient, a researcher, or any other party that desires to use system 100 to receive a dementia prediction indicatorbased on an image. For example, user 102 may be a technician or a nurse at a memory clinic who is tasked with triaging a patient prior to formal memory clinic assessment. Alternatively, the user 102 may be an employee of a regulatory organization such as a driver’s licensing centre. In a second example, the user 102 may be an employee at a driver’s test centre that licenses drivers of automobiles. User 102 may be the patient themselves, using the system to perform a preliminary screening for cognitive impairment to determine if there is a need to visit a healthcare provider. User 102 can also be a researcher at a research institute conducting research into persons with cognitive impairment.
[0133] Candidate image 104 may be any image that can form the basis for generating a dementia prediction indicator. For example, candidate image 104 can be a hand drawing on a sheet of paper of an analog clock with hands set at a specific position (i.e. a freely drawn clock drawing test) collected using an image collection device such as a camera or a scanner. Candidate image 104 can also be a particular shape, a line drawing, or a specific figure such as a Rey-Osterreith Complex Figure.
[0134] In some embodiments, candidate image 104 may be digitally created. For example, candidate image 104 may be drawn using a touch sensitive input device from a digital tablet. In another example, candidate image 104 may be drawn using a stylus on a digital input device or be created using other types of input devices such as mouse or keyboard. As another example, candidate image 104 may be assembled from pre-existing shapes on a software application.
[0135] User device 106 can be used by user 102 to receive the candidate image 104 and transmit the image to the processing device 200. Alternatively, a pre-trained model may be transmitted from the device 200 and the processing may occur in the environment 120 at the user device 106. The user device 106 may be any device with the capacity to capture or receive images and to communicate with other devices. A user device 106 may be, for example, a mobile device such as mobile devices running the Google® Android® operating system or Apple® iOS® operating system. A user device 102 may also be, for example, a personal computer operating the Windows® or MacOS® operating system. User device 106 may receive the candidate image 104 through the use of an image capture device as a digital image or through receiving the image as a digital file. For example, user device 106 may be a mobile phone equipped with a camera, andimage 104 may be a hand drawing of a clock. User device 106 may capture a digital image of the image 104 using its embedded camera. User device 106 may then transmit the image 104 to processing device 200. As another example, a user 102 may upload an existing digital file containing image 104 via an external storage device to user device 106. Other image collection devices may be used, such as a scanner or photocopier.
[0136] Network 108 may be any network or network components capable of carrying data including the Internet, Ethernet, fiber optics, satellite, mobile, wireless (e.g. Wi-Fi, WiMAX), SS7 signaling network, fixed line, local area network (LAN), wide area network (WAN), a direct point-to-point connection, mobile data networks (e.g., Universal Mobile Telecommunications System (UMTS), 3GPP Long-Term Evolution Advanced (LTE Advanced), Worldwide Interoperability for Microwave Access (WiMAX), etc.) and others, including any combination of these.
[0137] Processing device 200 can receive candidate image 104 and process the image 104 to produce a dementia prediction indicator. The processing device 200 can be a personal computer, a server, a user device, or any other similar device with processing and communication capabilities. Machine learning models for processing and analyzing the image 104 may be stored in memory on device 200. The processing device 200 may perform the methods of FIG. 5 in order to perform the prediction. The processing device 200 may also train or create one or more models based on the methods for FIG. 6 and may transmit the pre-trained model to the user device 106.
[0138] In some embodiments, user device 106 and processing device 200 may be combined into a single device. For example, processing device 200 may be a mobile phone containing an embedded camera capable of capturing a digital image of image 104 and a processor capable of operating a machine learning model for analyzing the candidate image 104 to determine a dementia prediction indicator. In such embodiments, network 108 may not be needed for the transfer of candidate image 104.
[0139] Reference is next made to FIG. 2 in conjunction with FIG. 1. FIG. 2 shows a schematic diagram of a processing device 200 in accordance with one or more embodiments. Processing device 200 may function as a system for generating a dementia prediction indicator based on a candidate image received in a clinical environment (as described in FIG. 5), or for generating a model for generating a dementia prediction indicator (as described in FIG. 6) which may be transmitted to the user device 106 (seeFIG. 1). Processing device 200 includes a communication unit 204, a display 206, a processor unit 208, a memory unit 210, an I / O unit 212, and a power unit 202. As noted above, the functions of the processing 200 may be combined with user device 106 (see FIG. 1).
[0140] The communication unit 204 can include wired or wireless connection capabilities. The communication unit 204 can include a radio that communicates using standards such as IEEE 802.11a, 802.11 b, 802.11g, or 802.11 n. The communication unit 204 can be used by the device 200 to communicate with other devices such as user device 106 or other computers. Communication unit 204 may communicate with a network, such as network 108 of FIG. 1.
[0141] The display 206 may be an LED or LCD based display and may be a touch sensitive user input device that supports gestures.
[0142] The processor unit 208 controls the operation of the processing device 200. The processor unit 208 can be any suitable processor, controller or digital signal processor that can provide sufficient processing power depending on the configuration, purposes, and requirements of the device 200 as is known by those skilled in the art. For example, the processor unit 208 may be a high-performance general processor. In alternative embodiments, the processor unit 208 can include more than one processor with each processor being configured to perform different dedicated tasks. The processor unit 208 may include a standard processor, such as an Intel® processor or an AMD® processor.
[0143] The memory unit 210 comprises software code for implementing an operating system 220, programs 222, dementia prediction model 224, dementia prediction model training unit 228, image cropping model 226, image cropping model training unit 230. The memory unit 210 can include RAM, ROM, one or more hard drives, one or more flash drives or some other suitable data storage elements such as disk drives, etc. The memory unit 210 is used to store an operating system 220 and programs 222 as is commonly known by those skilled in the art.
[0144] The I / O unit 212 can include at least one of a mouse, a keyboard, a touch screen, a thumbwheel, a trackpad, a trackball, a card-reader, an audio source, a microphone, voice recognition software and the like again depending on the particularimplementation of the server 106. In some cases, some of these components can be integrated with one another.
[0145] The power unit 202202 can be any suitable power source that provides power to the server 106 such as a power adaptor or a rechargeable battery pack depending on the implementation of the server 106 as is known by those skilled in the art.
[0146] The operating system 220 may provide various basic operational processes for the server 200. For example, the operating system 220 may be a server operating system such as Ubuntu® Linux, Microsoft® Windows Server® operating system, or another operating system.
[0147] The programs 222 include various user programs. They may include several hosted applications delivering services to users over the network, for example, medical record management system.
[0148] The dementia prediction model 224 may be a machine learning model that accepts image data as an input and produces an indicator based on the image data. The dementia prediction model 224 can be a classification model that produces one or more classifications as an output. Dementia prediction model 224 may receive candidate image 104 as an input, which may be a hand drawing of a clock. However, the dementia prediction model 224 can be configured to accept a variety of drawing types, such as other kinds of shapes, figures, or line drawings. The dementia prediction model 224 may produce a dementia prediction indicator as an output. The dementia prediction indicator may be an integer score corresponding to the candidate image 104. The score may be indicative of a level of dementia progression or cognitive decline. For example, the dementia prediction indicator could be an integer score from 0-5 indicating a rating of a clock drawing test, the rating corresponding to the degree of cognitive impairment that can be inferred from the clock drawing test.
[0149] Dementia prediction model 224 may be constructed, for example, in accordance with the machine learning model architecture 300 depicted in FIG. 3. Architecture 300 may be comprised of a transformer 304, a median network 330, and an output network 340. The transformer 304 may accept an image 302 as input and may pass its output to median network 330, which further passes its output to network 340. Network 340 may produce an output 322, which may be a dementia prediction indicator.In some embodiments, the image input 302 may be size 3x512x512. However, the image input may be of any size and dimension. The image input 302 may be the candidate image 104 of FIG. 1.
[0150] Transformer 304 may accept an input of a specified size and produce an output to the median network 330. For example, the input may be of size 3x512x512 and the output of size 1024. The transformer 304 may be, for example, the PyT orch torchvision implementation of the vision transformer (ViT) "vit_l_16" architecture with "ViT_L_16_Weights.lMAGENET1 K_SWAG_E2E_V1" pretrained weights. However, the vision transformer can be any suitable pre-trained vision transformer. The head of the model may be removed and replaced with a small fully connected network designed to output a clock score (“cg<round>dclkdraw”; scored from 0-5) and a previous dementia / Alzheimer’s diagnosis as reported by a respondent (“hc<round>disescn9”; with “yes” and “previously reported” set to 1 , 0 otherwise). The input may be modified to conform with input parameters. For example, the single-channel preprocessed image may be converted to 3-channel to conform with the expected input of vit_l_16. The head of vit_l_16 may be replaced with fully-connected linear layers, which may be the median network 330 and output network 340.
[0151] The median network 330 may contain linear layers 306, 310, and 314. The linear layers 306, 310, 314 may accept the output of a previous layer as an input, apply a linear transformation to the input, and pass the output to the next layer. The median network 330 may also contain activation layers 308, 312, 316. The activation layers may accept the output of a previous layer, apply an activation function to each element, and pass the result to the next layer. The activation function may be leaky rectified linear unit (Leaky ReLU), but can be other functions such as ReLLI, sigmoid, tanh, ELU, softmax, or any other suitable activation function. The linear layers and activation layers may be fully connected layers, consisting of the same size of input and output. The layers may, in some embodiments, accept inputs of size 1024 and produce outputs of size 1024. The output network 340 may contain a linear layer 318 and an output layer 320. The linear layer 318 may accept an input of size 1024 and produce an output of size 1 or 2. The output layer 320 may apply a sigmoid function to the output of the previous layer and produce an output 322, which may be a dementia prediction indicator. The output may be 2-dimensional and may comprise a first dimension representing a clock score and asecond dimension representing a dementia diagnosis. In some embodiments, the median network 330 may be comprised of three fully connected layers (1024 input and output features) with the leaky ReLu activation function (0.2 negative slope) followed by one fully connected layer (1024 feature input and 2 feature output) with the sigmoid activation function. The clock score output may range from 0-1 and then may be multiplied by 5 to match the score range, i.e. , 0-5.
[0152] In some embodiments, the dementia prediction model 224 can be a regression model that produces a numerical output within a range of numbers. For example, the dementia prediction indicator could be a score within a continuous range of numbers. In some embodiments, to implement the linear regression, the median network 330 and output network 340 of dementia prediction model 224 may be replaced with a linear regression network. The linear regression network can comprise a single densely connected layer with a 1024 feature input and one output, the one output being a number within a continuous range of numbers.
[0153] The dementia prediction model training unit 228 may perform the method of FIG. 6 to generate and store the dementia prediction model 224 in memory 210.
[0154] Referring back to FIG. 2, the image cropping model 226 may be a machine learning model that is configured to accept an image as input and produce a heatmap as an output. The image may be the candidate image 104. The heatmap may indicate the location in the image where a feature of interest is detected. For example, the heatmap may indicate the precise location the clock of the clock drawing test is in the candidate image 104. The heatmap may be in the form of a bitmap where the pixel value is 1 where a portion of a clock is present at the pixel location relative to the original clock drawing test image, and 0 otherwise.
[0155] The cropping model 206 may be constructed, for example, in accordance with the architecture 400 in shown in FIGs. 4A - 4B. Architecture 400 may contain one or more convolution downsample blocks 410, one or more convolution blocks 420, one or more convolution upsample blocks 430, an output convolution layer 404, and an output activation layer 406. Architecture 400 is designed to receive a preprocessed image 402 as input and produce a heatmap 408 as output. The preprocessed image to be used as input 402 can be a specified size, for example, 1x512x512. The output heatmap may be a different size than the input image. For example, the output heatmap can be size1x64x64. In the embodiment shown in FIG. 4A, convolution downsample blocks 451 , 452, 453, 454, 455, 456, 457, convolution blocks 461 , 462, and convolution upsample blocks 471 , 472, 473, 474 are used in architecture 400.
[0156] Each convolution downsample block 410 may receive an input 412. The input may be the output of a previous block or layer. The input may be processed by a first convolution layer 413, a first activation layer 414, a second convolution layer 415, a second convolution layer 416, and the resulting output 417 may be passed on to the next block or layer. The first convolution layer 413 may apply a convolution to its input with kernel 3x3, stride 1 , and padding 1. The first and second activation layers 414, 416 may be leaky ReLU functions with negative slope 0.2. The second convolution layer 415 may apply a convolution to its input with kernel 4x4, stride 2, padding 1. The input to each convolution downsample block may also include a copy of the preprocessed image 411 , which may be downsampled using max-pool and bilinear downsampling techniques to match the size of the feature maps for each block.
[0157] For each convolution downsample block 410 shown in FIGs. 4A - 4B, the numbers in brackets indicate the number of features with the first number referring to the last layer output and the “+2” referring to the downsampled (using max-pool and bilinear downsampling) image being added as two extra features. The number of features is increased relative to the input at the second convolution layer in each block, except for the first convolution downsample block 451 wherein the features are increased from 1 to 25 in the first convolution layer and 25 to 48 in the second convolution layer.
[0158] Each convolution block 420 may receive an input 421 , which is fed into convolution layer 422 and activation layer 423, to produce an output 424. Convolution layer 422 may apply a convolution with kernel 3x3, stride 1 , and padding 1. Activation layer 423 may be a leaky ReLU with negative slope 0.2.
[0159] Each convolution downsample block 430 may receive an input 431 , which is fed into a convolution transpose layer 432 and an activation layer 433 to produce output 434. Convolution transpose layer may apply transposed convolution to the input with kernel 4x4, stride 2, padding 1. The activation layer 433 may be a leaky ReLU with negative slope 0.2.
[0160] The image cropping model training unit 230 may perform the method of FIG. 6 to generate and store the image cropping model 226 in memory 210.
[0161] Reference is next made to FIG. 5, which shows a method diagram 500 for generating a dementia prediction indicator based on a candidate image received in a clinical environment. The method can be implemented on processing device 200 or user device 106 (see e.g. FIG. 2 and FIG. 1 respectively).
[0162] At step 502, the method begins by providing a dementia prediction model and an image-cropping model for cropping an input image for the dementia prediction model. For example, device 200 may provide a dementia prediction model 224 and an image-cropping model 226 at memory 210.
[0163] The image-cropping and the dementia prediction models may be machinelearning models, such as convolutional neural networks, deep neural networks, variational autoencoders, decision trees, support vector machines, and any other such suitable model. In some embodiments, the dementia prediction model may further comprise a transformer. For example, the dementia prediction model 224 may be a machine learning model with architecture as shown in FIG. 3. The image-cropping model 226 may be a machine learning model with architecture as shown in FIGs. 4A - 4B.
[0164] In some embodiments, the image-cropping model may comprise at least one downsample layer, at least one layer receiving the output of the at least one downsample layer, and at least one upsample layer receiving the output of the at least one layer. In some embodiments, the at least one downsample layer comprises at least one convolution downsample layer, the at least one layer comprises at least one convolution layer, and the at least one upsample layer comprises at least one convolution upsample layer. For example, with reference to FIGs. 4A - 4B, the at least one downsample layer may be the convolution downsample block 410, the at least one layer may be the convolution block 420, and the at least one downsample layer may be the convolution downsample layer 430.
[0165] In some embodiments, the dementia prediction model may comprise at least one median fully connected layer, the median fully connected layer comprising an equal number of inputs and outputs, and an output layer connected to the output of the at least one median fully connected layer, the output layer comprising two outputs.
[0166] In some embodiments, the dementia prediction model may comprise a linear regression layer, the linear regression layer comprising one output, the one output being a number within a continuous range of numbers.
[0167] At step 504, the method proceeds by receiving the candidate image. In one or more embodiments, processing device 200 may receive candidate image 104 from user device 106 through communication unit 204. In some embodiments, the candidate image may comprise a clock drawing test. However, the candidate image could be any drawing, figure, or shape that is suitable for use in dementia prediction. In some embodiments, the candidate image is captured using a user device camera in the clinical environment. For example, the image may be captured using a mobile phone or tablet with camera capabilities in a memory clinic setting.
[0168] In some embodiments, method 500 includes pre-processing the candidate image at the processor. The pre-processing may facilitate improved consistency and maximal detail transfer when passing the images as inputs to a machine learning model. The pre-processing may involve converting the candidate image to a greyscale candidate image, applying one or more line filters to the greyscale candidate image to produce a plurality of filtered images, the plurality of filtered images comprising at least one of a first filtered image, a second filtered image, and a third filtered image, and generating the pre- processed candidate image by averaging the plurality of filtered images based on a line weight associated with each of the filtered images. In some embodiments, each of the one or more line filters may be applied to the greyscale candidate image in accordance with one or more pixel values associated with the each of the one or more line filters. For example, for each image, a first line filter may be applied for line widths of 1 pixel to the image, producing a first filtered image, to which a weight of 2 / 5 is applied. A second line filter may be applied for line widths of 3 pixels to the image, producing a second filtered image, to which a weight of 2 / 5 is applied. A third line filter may be applied for line widths of 5 pixels to the image, producing a third filtered image, to which a weight of 1 / 5 is applied. A weighted average of the pixel values of the filtered images may be produced for each image. A pre-processed image may be produced by setting all pixels with value less than 0.2 to zero. The pre-processed images may be assembled to produce a pre-processed image set.
[0169] At step 506, method 500 proceeds by generating a cropped candidate image based on the candidate image and the image-cropping model. In one or more embodiments, the processor unit 208 may pass the candidate image 104 as an input to image cropping model 226. The processor unit 208 then operates image cropping model 226 to generate a heatmap indicating the location of the clock within the clock drawing test. Various example heatmap outputs produced by the image-cropping model are shown in FIG. 7c. In some embodiments, method 500 further comprises padding the cropped candidate image prior to processing the cropped image using the dementia prediction model.
[0170] In some embodiments, generating the cropped candidate image based on the candidate image and the image-cropping model may comprise processing the candidate image using the image-cropping model to produce a heatmap of a probable location of a clock relative to the candidate image, cropping the candidate image to produce an intermediate candidate image such that a minimum proportion of the heatmap is contained within the intermediate candidate image, and cropping the intermediate candidate image to produce the cropped candidate image such that a minimum proportion of a total pixel value sum of the intermediate candidate image is contained within the cropped candidate image. For example, the generating the cropped candidate image may entail using a multi-step cropping script on the candidate image. The cropping script may follow the following example steps: First, the preprocessed image may be reshaped to be 512x512 and then passed to the image-cropping model, which outputs a heatmap of the probable location of the clock. The 512x512 heatmap was reshaped to the original image size. A simple cropping technique was then applied to the original preprocessed image (not resized to 512x512) such that 84% of the total heatmap value sum was contained within the crop, then expanded by 80 pixels at all edges. The cropped image was then cropped further, such that 96% of the total pixel value sum (of the cropped image) was contained within the crop and then expanded by 30 pixels at all edges. The image was padded with zeros to make the image square and then resized to 512x512. The above process, from heatmap generation to resizing, was repeated two more times, taking the output of the last iteration as the input of the current iteration, effectively narrowing-in on the location of the clock with each iteration.
[0171] At step 508, the method proceeds by generating, at the processor, the dementia prediction indicator based on the cropped candidate image and the dementia prediction model.
[0172] In some embodiments, the dementia prediction indicator can include a clock drawing test measure and a dementia prediction. A clock drawing test measure may be a score related to the clock drawing test that is indicative of an accuracy of the depiction of the clock contained therein.
[0173] In some embodiments, processing device 200 may receive one or more prediction features, and generating the dementia prediction indicator is further based on the one or more prediction features. The prediction features may be one or more of: a socio-economic indicator, a driving test score, and a prior medical history. For example, the one or more prediction features can include the demographic information of patient, and this may be further used to produce the dementia prediction indicator. The one or more prediction features could also, for example, include the clinical history of a patient, which may also be used in generating the dementia prediction indicator.
[0174] Reference is next made to FIG. 6, which shows a block diagram of a method 600 for generating a dementia prediction model. Specifically, the dementia prediction model may be the dementia prediction model 224 and the image cropping model 226 of FIG. 2. Generating model 224 may involve the use of supervised learning training techniques.
[0175] The method begins at step 602, with receiving a plurality of training images.
[0176] The plurality of training images may be collected from research institution databases, from conducting a study, or from any other source of training images containing an adequate number of samples to train a machine learning model. For example, the training images including clock drawing images may be obtained from the Toronto Dementia Research Alliance (TDRA) and the National Health and Aging Trends Study (NHATS).
[0177] The TDRA clinical research database consists of participant’s demographic information, clinical history, cognitive assessment data and diagnoses provided by neurologists, geriatric psychiatrists, or geriatricians from four memory clinics in Toronto located at Baycrest, CAMH, Sunnybrook and UHN. The clinical histories were sourcedfrom either participants, care partners, healthcare records, or combinations of all three. Specifically, data were obtained from August 2017 and August 2021 , totaling 1986 samples with clock drawing tests. Participants were included in the study if written informed consent was obtained and documented. Exclusions were made when it was in the participant’s best interest not to participate, as determined by the healthcare team, physician, or caregiver, or if the participant was unaccompanied and unable to provide consent or if consent was withdrawn at any time.
[0178] The NHATS database was started in 2011. NHATS is a longitudinal study of disability trends and trajectories among Medicare beneficiaries aged 65 years and older in the United States. The anonymized publicly available NHATS database is comprised of annual interviews and cognitive assessments, with samples selected using a stratified random sampling of community-dwelling older adults. Data was obtained from rounds 1 to 11 (collected from 2011 to 2021) from the NHATS website in February 2023, totaling 54027 samples with clock drawing tests.
[0179] In some embodiments, the plurality of training images may comprise drawings of a clock. As an example, the training images may be images taken from the Clock Drawing Test (CDT). The Clock Drawing Test is part of both the TorCA and the BNA-SF and generally proceeds as follows. During the assessment, participants are given an 8 1 x 11-inch paper in portrait orientation and are instructed to "Draw the face of a clock, put in all the numbers and set the hands at 10 after 11." The instructions may be repeated to the participant as needed. As participants draw the clock, the administrator may use an iPad® to record the sequence of elements. In this embodiment, the iPad® provides options for different elements, allowing the administrator to select the element drawn by the participant as they draw the clock. After the participant completes the drawing, the administrator uses the iPad's camera to capture an image of the clock drawing test, creating a digital copy. The assessment may be conducted by trained clinicians, including healthcare professionals, medical students, medical fellows, certified psychological associates, nurses, and psychologists. The clinicians at all sites undergo thorough TorCA and BNA-SF training, with the aim of ensuring correct administration, data collection and scoring consistency. As another example, CDT images may simply be obtained from various participants by having them draw a picture of an analog clock with the hands showing 11 :10 on a blank A4 size paper.
[0180] In some embodiments, the method 600 includes pre-processing each training image in the plurality of training images. The pre-processing may comprise converting the training image to a greyscale training image, applying one or more line filters to the greyscale training image to produce a plurality of filtered images, the plurality of filtered images comprising at least one selected from the group of a first filtered image, a second filtered image, and a third filtered image, and generating the pre-processed training image by averaging the set of filtered images based on a line weight associated with each of the filtered images. Similar to in method 500, the pre-processing may facilitate improved consistency and maximal detail transfer when passing the images as inputs to a machine learning model. For example, image from TDRA and NHATS databases may be used as training images. Both TDRA (varies, approximately 960x1280, RGB) and NHATS (2560x3312, single channel) clock images are either scanned or photographed with cellphones, resulting in a large variance of quality. To address this issue, multiple steps may be taken to preprocess the images so that they are of suitable quality and size to pass to the dementia classification neural network.
[0181] In some embodiments, each of the one or more line filters is applied to the greyscale training image in accordance with one or more pixel values associated with the each of the one or more line filters. For example, for the TDRA images, the first step may be line filtering to remove irrelevant information such as color and shadows while retaining the clock line drawings. After converting the images to black-and-white, the weighted average of line filters with line widths of 1 , 3 and 5 pixels may be applied to the image, with weights of 2 / 5, 2 / 5, and 1 / 5, respectively. The range of the resultant images may then be individually adjusted such that the maximum was 1 and minimum value was 0. Additionally, all values less than 0.2 may be set to zero to reduce noise. For the NHATS images, line filtering may already have been pre-applied, so the images were only inverted. Additionally, since many of the NHATS images used black rectangles to censor the names of the participants, a filter may be applied to remove the rectangles.
[0182] In some embodiments, method 600 may further comprise dividing the plurality of training images into a training data set and a validation data set prior to fine- tuning the pre-trained model. In some embodiments, the method 600 may further include replicating one or more images associated with one or more image groups in the validationdata set and the training data set, each image group being associated with a value of a manual label, until all of the image groups contain an equal quantity of samples.
[0183] In some embodiments, the manual label comprises a clock drawing test measure. The clock drawing test measure may be a numerical value assigned to the image in accordance with a system of scoring. For example, the TDRA 15-point CDT scoring system may be used. The TDRA scoring system was derived from the Freedman scoring system (Freedman, 2018). The Freedman scoring system may demonstrate higher levels of reliability compared to other CDT scoring systems in past studies. The system uses 15 binary (1 if true, 0 otherwise) clock drawing test features scored by trained health practitioners. As another example, a 0-5 scale may be used as the scoring system for the clock drawing test measure. For instance, CDT images may be scored by the interviewers according to pre-defined criteria, using the range of 0-5 as follows: (0) not recognizable as a clock, (1) severely distorted depiction, (2) moderately distorted depiction, (3) mildly distorted depiction, (4) reasonably accurate depiction, and (5) accurate depiction.
[0184] For example, in one embodiment, the plurality of training images may be clock drawing test images from the NHATS dataset comprising 48485 samples. The training data base may be divided into training (47485 samples) and validation (1000 samples) sets. Both sets may be balanced such that all clock scores (0-5) have an equal number of samples by replicating samples of underrepresented scores in the training set. The training images may be preprocessed and cropped as described herein. To improve generalizability, data augmentation may be applied to the preprocessed training images during training and validation, with scaling (e.g., by 0.7 to 1 times original size), translation (e.g., by 0 to 0.5 times the respective dimension size in all directions), rotation (e.g., by - 180 to 180 degrees), and shearing (e.g., by -10 to 10 degrees; all uniformly sampled).
[0185] In some embodiments, the method may further comprise augmenting each training image in the plurality of training images by performing at least one of the following: scaling the training image, translating the training image, rotating the training image, and shearing the training image.
[0186] At step 604, the method 600 proceeds with generating a plurality of cropped training images based on an image-cropping model and the plurality of training images.
[0187] Cropping the training images may improve the consistency and quality of the images for the purposes of passing the images to a machine learning model. For example, many images in the TDRA and NHATS may not be properly cropped with respect to the clock drawing test. Given the limited image size that can be passed to the classification network, "vit_l_16" (512 by 512), it may be beneficial to crop the images because downsampling the original without cropping would result in some clock drawing tests having resolutions that would be too low to make out the details of the drawing. To crop on the clocks, the location of the clocks must first be located. Simply cropping on the lines in the image may not be sufficient due to the presence of non-clock lines in many of the images (e.g. written notes, paper edges, text on page, etc.). To locate the clock drawing test, an image-cropping model may be used. For example, the image-cropping model could be a convolutional neural network that is trained to convert a clock image into a “heatmap” of the probable location of the clock. Pytorch may be used for neural network related code. The image-cropping model may be image-cropping model 226 stored in the memory of processing device 200 as shown in FIG. 2, constructed in accordance with architecture 400 as shown in FIGs. 4A - 4B.
[0188] In some embodiments, the cropping may involve using the image-cropping model in a multi-step cropping script. The cropping script may contain the following example steps: First, the preprocessed image was reshaped to be 512x512 and then passed to Crop-Net which output a heatmap of the probable location of the clock. The 512x512 heatmap was reshaped to the original image size. A simple cropping technique was then applied to the original preprocessed image (not resized to 512x512) such that 84% of the total heatmap value sum was contained within the crop, then expanded by 80 pixels at all edges. The cropped image was then cropped further, such that 96% of the total pixel value sum (of the cropped image) was contained within the crop and then expanded by 30 pixels at all edges. The image was padded with zeros to make the image square and then resized to 512x512. The above process, from heatmap generation to resizing, was repeated two more times, taking the output of the last iteration as the input of the current iteration, effectively narrowing-in on the location of the clock with each iteration. The image-cropping model and the multi-step cropping script described above may be used to crop all the training images. For example, NHATS and TDRA clockimages. FIG. 8 shows examples of cropped and pre-processed images for the TDRA dataset.
[0189] In some embodiments, the image-cropping model may comprise at least one downsample layer, at least one layer receiving the output of the at least one downsample layer, and at least one upsample layer receiving the output of the at least one layer. In some embodiments, the at least one downsample layer comprises at least one convolution downsample layer, the at least one layer comprises at least one convolution layer, and the at least one upsample layer comprises at least one convolution upsample layer. For example, the image-cropping model may follow architecture 400 from FIGs. 4A - 4B.
[0190] At step 606, the method proceeds with fine-tuning a machine learning model using the cropped training images and at least one feature associated with each of the cropped training images, the machine learning model comprising a pre-trained model.
[0191] In some embodiments, the pre-trained model comprises a transformer.
[0192] In some embodiments, the head of the machine learning model may further comprise at least one median fully connected layer, the median fully connected layer comprising an equal number of inputs and outputs, and an output layer connected to the output of the at least one median fully connected layer, the output layer comprising two outputs. For example, the median fully connector layer can be median network 330 of FIG. 3, and the output layer can be network 340 of FIG. 3. In some embodiments, the head can comprise a linear regression layer. The linear regression layer may be implemented, for example, using a densely connected linear layer connected to the output of the body and producing one output, the output being a number within in a continuous range of numbers.
[0193] In some embodiments, the fine-tuning the machine learning model may comprise modifying only one or more weights associated with a head of the machine learning model until a mean squared error associated with the validation data set has plateaued, and modifying one or more weights associated with a body of the machine learning model and the one or more weights associated with the head of the machine learning model, the body of the machine learning model comprising the pre-trained model. For example, the pre-trained model may be model 304 of FIG. 3 comprising a vit_l_16vision transformer. The loss function for training may be mean squared error, and the Adam algorithm may be used. The parameters may be, for example, learning rate = 0.00002, betas = (0.5,0.999), and other parameters set to default value. The learning rate may be divided by 5 when validation MSE plateaus on the validation data. To terminate the learning process, early stopping may be applied when the MSE subjectively remained plateaued after decreasing the learning rate. Where the pre-trained model comprises a head and a body, the head of the model may be trained first, with the weights of the body frozen, until MSE plateaus on the validation data (50 batch size). The body weights may then be unfrozen and training may continue until MSE plateaued on the validation data (4 batch size with gradients accumulated over 10 iterations due to VRAM constraints). FIG. 9 shows example graphs 902 and 904 of the mean squared error over the course of fine- tuning an example pre-trained model. The MSE is shown at each weight update, calculated using 50 random samples drawn from the validation set. The data was smoothed by averaging over a 51-width uniform kernel. After training is complete, the head may be removed, and the 1024-dimensional output of the finetuned vit_l_16 body may be used as the ViT features.
[0194] In some embodiments, the at least one feature may comprise at least one of a manual label, a socio-economic indicator, a driving test score, and a prior medical history. The manual label may comprise a clock drawing test measure.
[0195] In some embodiments, the method may further comprise generating the image-cropping model. The image-cropping model may be a machine-learning model. The model architecture may be constructed in accordance with architecture 400 from FIGs. 4A - 4B. The model may be trained using machine learning techniques such as supervised learning, reinforcement learning, or any other applicable technique.
[0196] The training may involve receiving an image-cropping dataset, the imagecropping dataset comprising images comprising clock drawing tests, generating a manually cropped training dataset based on the image-cropping dataset, generating a pre-processed dataset based on the manually cropped dataset, producing a first target bitmap dataset based on the pre-processed dataset by setting one or more first target pixels of the first target bitmap to a 1 value where the target pixel comprises a part of a clock image and 0 otherwise, producing a second target bitmap dataset based on the pre- processed dataset by setting one or more second target pixels of the second target bitmapto 0.2, 0.4, 0.6, 0.8, or 1 , based on a clock rating of the clock drawing test, and training a cropping machine learning model based on the image-cropping dataset, the first target bitmap, and the second target bitmap. In some embodiments, training the image-cropping model may further comprise augmenting, at the processor, each image in the cropping training dataset, the cropping validation dataset, and the cropping testing dataset by at least one of: adding noise to the image, adding lines to the image, adding writing to the image, and adding rectangles in a random orientation to the image. In some embodiments, generating a pre-processed dataset based on the manually cropped training dataset comprises, for each image in the manually cropped training dataset, rotating the image by a random angle, resizing the image by a random factor, and placing the image onto a black background image. In some embodiments, training the model may further comprise dividing the manually cropped training dataset into a cropping training dataset, a cropping validation dataset, and a cropping testing dataset.
[0197] For example, the training may involve receiving a set of images of clock drawing tests, for example from the NHATS dataset. The dataset may be prepared for use as a training dataset per the following steps. First, the raw images may be manually cropped in such a way that 96% of the total pixel value sum was contained within the crop, producing a first cropped dataset. Then, the images of the first cropped dataset may be expanded by 50 pixels at all edges, producing a second cropped dataset. The second cropped dataset may be divided into three separate datasets, comprising a training dataset (comprising 52527 images), a validation dataset (comprising 500 images), and a testing dataset (comprising 1000 images). For each image in the training and validation datasets, the image may be rotated by a random angle between 0 to 180 degrees, resized by a random factor ranging from 1 to 1 / 8, and placed onto a 512x512 black background image. For each image in all three datasets, noise, lines (comprising single lines and parallel lines), writing (comprising images of handwriting from internet sources), and rectangles may be added to the images, in randomly determined quantities and in varying sizes and orientations. Any portion of added writing and rectangles that overlapped with the main clock image are removed, as noise and lines can overlap with the clock, but writing and rectangles cannot. A first target bitmap with pixels set to 1 where a clock was and 0 for the rest of the image not containing any portions of a clock may be produced. A second target bitmap with pixels set between 0 and 1 , in 0.2 increments may be produced.The cropping model may then be trained on the training images and the first and second target bitmaps. FIG. 7A shows various images of example augmented training images.
[0198] FIG. 7B shows examples of images rendered in accordance with the first and second bitmaps corresponding to the clock drawing test images shown in FIG. 7A. In the embodiment shown in FIG. 7B, a target output was generated by manually setting the location of the clock image to 1 and the rest of the 512x512 augmented image to zero. An additional 512x512 target dimension may be added with the location of the clock set to score / 5 (instead of 1 for all scores) to account for the varying levels of “clockness” of the images, resulting in a 2x512x512 target with values between 0 and 1. In the embodiment shown, the image-cropping model was trained on the NHATS dataset using the following data augmentation technique. A simple cropping technique was initially applied to the images such that 96% of the total pixel value sum was contained within the crop. To avoid removing the edges of the clock, the crop was expanded by 50 pixels at all edges. The cropped images were then placed into training (52527 images), validation (500 images) and test (1000 images) sets. For training and validation datasets, the cropped images were randomly rotated (up to 180 degrees), resized (proportionally from 1 to 1 / 8 times the original size), and placed within a 512x512 black image. Noise, lines (both single and in parallel to mimic paper / desk edges and lined paper, respectively), writing (drawn from publicly available sources on the internet), and rectangles (to mimic any missed censoring) were randomly added in different sizes and orientations to mimic the various non-clock markings in the datasets. Noise and lines could overlap with the clock but writing and rectangles could not (any portion that did overlap was removed).
[0199] In some embodiments, mean squared error (MSE) may be used as the loss function. For minimizing the loss function, the Adam algorithm may be used with learning rate = 0.00002, betas = (0.5,0.999), and all other parameters set to default. The minibatch size was set to 50. The learning rate was divided by 5 when validation MSE plateaued. To terminate the learning process, early stopping may be applied when the MSE subjectively remained plateaued after decreasing the learning rate.
[0200] The cropping script may be used to crop the images to perform fine-tuning on the model. For example, the cropping script may be used on NHATS images. The clock images may then be divided into successful and unsuccessful crops, as manually determined by a human. The network may then be retrained as described above fromrandom weights using the successfully cropped images divided into training and validation sets. The network may be trained as above until MSE subjectively plateaued on the validation data. This two-step training approach may have the advantage of greatly reducing the human-hours required for generating the target outputs (i.e. the optimal crop for each image) required for supervised learning.Example 1
[0201] Reference is next made to FIG. 10, which shows example receiver operating characteristic (ROC) curves 1000 for the performance of various machine learning models in evaluating clock drawing test drawings, as well as the performance of a human. ROC curves 1000 indicate the relative performance of each of the models in scoring a clock drawing test for predicting dementia. ROC 1002 corresponds to the performance of the presently disclosed dementia prediction model, whereas ROC 1004 corresponds to human scoring performance. Other models tested include MiniVGG, MNv2, and RF-VAE.
[0202] The different various methods herein were tested on the same dataset of images in Example 1 , the images being preprocessed and cropped in accordance with the preprocessing and cropping steps in method 500 of FIG. 5. The human-scored clock drawing test features are comprised of 15 binary variables scored by trained health practitioners. Each feature was scored as 1 or 0, depending upon whether the criterion for the feature was satisfied or not.
[0203] The MiniVGG features were based upon the CNN method used in Sato et al. (2022). The network architecture outlined in Sato et al. (2022) was used, with the exception that the softmax in the output layer was replaced with the sigmoid function. The network was trained to predict the clock score and dementia diagnosis on the NHATS data with the same approach used for ViT, except the body and head were trained simultaneously from the start because pretrained weights were not used. After training was complete, the final layer was removed, and the 512-dimension output was used as the MiniVGG features.
[0204] The MNv2 features were based upon the CNN method outlined in Amini et al. (2021), using a PyTorch torchvision implementation of the MobileNetV2 architecture (Sandler et al., 2018) with the "MobileNet_V2_Weights.lMAGENET1 K_V2" pretrained weights(https: / / pytorch.org / vision / main / models / generated / torchvision.models.mobilenet_v2.html ). No additional training was performed. The fully connected layer was removed and replaced with a global average pooling layer, reducing the 7 x 7 spatial dimensions to 1. The resultant 1280-dimension output was used as the MNv2 features.
[0205] The RF-VAE features were based upon the relevance factor variational autoencoder (RF-VAE) method used in Bandyopadhyay et al. (2023). The 10 features output by the pretrained encoder (https: / / github.com / iheallab / Clock-Drawing- Classification-With-RF_VAE) were used as the RF-VAE features.
[0206] For the test samples, dementia diagnosis (N = 522) included those diagnosed by TDRA clinicians with vascular dementia, dementia with Lewy bodies, dementia NYD (cause not yet determined), and Alzheimer’s disease. Those in the normal cognition group (N = 340) performed within the normal range on the TDRA clinical tests. Basic demographic information for both groups is listed in Table 1. Table 1 shows the number of participants, number of female participants, average age, and years of education (YOE) for both groups together and separately.Table 1 : Demographicscognition (coded as 0) from clock drawing test features produced using the present model, ViT, and four comparison methods: Human, MiniVGG, MNv2, and RF-VAE.
[0208] To improve generalizability in the context of photographed clock drawing tests, the TDRA training dataset was augmented with 50 permutations of each drawing by scaling (0.7 to 1 times original size), translating (by 0 to 0.5 times the respective dimension size in all directions), rotating (-45 to 45 degrees), and shearing (-10 to 10 degrees; all uniformly sampled). This augmentation was not performed for the Human or the RF-VAE features, with the former excluded because the individual scores are not affected by such transformations and the later excluded because the model was trainedon digital clock drawing tests that were not subject to such sources of variability (consequently, the RF-VAE linear model performed worse when trained using the 50 permutations: ROC AUC 0.730 vs. 0.741).
[0209] For the Human and RF-VAE features (15 and 10, respectively), linear regression models were trained using leave-one-out cross-validation to estimate out-of- sample performance. For the ViT, MiniVGG and MNv2 features, due to the high dimensionality of the features relative to the dataset sample size (1024, 512, and 1280 vs. 862), the feature dimensionality was reduced to 100 using principal component analysis (PCA) trained on the NHATS dataset. Lasso linear regression models were then trained using nested cross-validation, with leave-one-out cross-validation in the outer loop and 5-fold cross-validation in the inner loop to select the penalty value (a), from 1 / 10A3 to 1 / 10A1 in exponent increments of 0.2.
[0210] To measure the performance of the five different methods, ROC AUC, balanced accuracy (maximized over the decision threshold), F1 scores (maximized over the decision threshold), and true positive rates (TPR) for 5% and 10% false positive rates (FPR) were used. ViT was statistically compared to the other four methods using paired- samples bootstrapping (1000 samples) to generate p(comparison model >= ViT).
[0211] Using linear models, the clock drawing test features produced using the presently disclosed dementia prediction model architecture and four comparison methods (Human, MiniVGG, MNv2, RF-VAE) were trained, validated, and tested in the context of predicting dementia (N = 522) vs. normal cognition (N = 340) in the TDRA clinical dataset (N = 862).
[0212] The results for each method are summarized in Table 2 and the ROC curves are displayed in Figure 10. In Table 2, ROC AUC, balanced accuracy (maximized over the decision threshold), F1 score (maximized over the decision threshold), and true positive rates (TPR) for 5% and 10% false positive rates (FPR) are shown. For each measure, the performance of the presently disclosed model is statistically compared to the four comparison methods, reporting the probability of the null hypothesis, p(comparison method >= ViT), in brackets.
[0213] Overall, the presently disclosed model architecture appeared to be superior at predicting dementia vs. normal cognition compared to human-scored features and allother deep-learning-based methods on all measures. Moreover, the superior performance was statistically significant in all cases except when compared to human-scored features for balanced accuracy (p = 0. 099) and TPR at 5% (p = 0. 284) and 10% FPR (p = 0.082).
[0214] Table 2 shows recorded performance of the different methods tested. Measures shown for the presently disclosed model (ViT), a model based on human- scored features (Human), a model based on Sato et al. (2022; MiniVGG), a model based on Amini et al. (2021 ; MNv2), and a model based on Bandyopadhyay et al. (2023; RF- VAE). The thresholds for balanced accuracy and the F1 scores are those that maximized the respective measure for the given method. Scores with the best performance between models are in bold. Models sorted by balanced accuracy. Numbers in brackets are p(model >= ViT) for the respective measure (*** = p < 0.001 ; ** = p < 0.01 ; * = p < 0.05; ■ = p < 0.1). ROC AUC = receiver operating characteristic area under the curve, TPR = true positive rate, FPR = false positive rate.Table 2: Dementia prediction measures of performance.00215] The present invention has been described here by way of example only. Various modification and variations may be made to these exemplary embodiments without departing from the spirit and scope of the invention, which is limited only by the appended claims.REFERENCES1. Amini, S., Zhang, L., Hao, B., Gupta, A., Song, M., Karjadi, C., Lin, H., Kolachalama, V.B., Au, R. and Paschalidis, I.C. (2021). An artificial intelligence-assisted method for dementia detection using images from the clock drawing test. Journal of Alzheimer's Disease, 83(2), 581-589.2. Bandyopadhyay, S., Wittmayer, J., Libon, D. J., Tighe, P., Price, C., & Rashidi, P. (2023). Explainable semi-supervised deep learning shows that dementia is associated with small, avocado-shaped clocks with irregularly placed hands. Scientific Reports, 13(1), 7384.3. Cao, Q., Tan, C. C., Xu, W, Hu, H., Cao, X. P., Dong, Q., Tan, L. and Yu, J. T. (2020). The prevalence of dementia: a systematic review and meta-analysis. Journal of Alzheimer's Disease, 73(3), 1157-1166.4. Chen, S., Stromer, D., Alabdalrahim, H. A., Schwab, S., Weih, M., & Maier, A. (2020). Automatic dementia screening and scoring by applying deep learning on clock-drawing tests. Scientific Reports, 10(1), 20854.5. Davoudi, A., Dion, C., Amini, S., Tighe, P. J., Price, C. C., Libon, D. J., & Rashidi, P. (2021). Classifying non-dementia and Alzheimer’s disease / vascular dementia patients using kinematic, time-based, and visuospatial parameters: the digital clock drawing test. Journal of Alzheimer's Disease, 82(1), 47-57.6. Deng, J., Dong, W, Socher, R., Li, L.-J. , Li, K., & Fei-Fei, L. (2009). Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition (pp. 248-255).7. Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Geliy, S., & Uszkoreit, J. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929.8. Freedman, M., Leach, L., Kaplan, E., Winocur, G., Shulman, K. I., & Delis, D. C. (1994). Clock drawing: A neuropsychological analysis. Oxford: Oxford University Press.9. Freedman, V. A., & Kasper, J. D. (2019). Cohort profile: the National Health and aging trends study (NHATS). International journal of epidemiology, 48(4), 1044- 1045g.10. Hinton, G. E., Osindero, S., & Teh, Y. W (2006). A fast learning algorithm for deep belief nets. Neural computation, 18(7), 1527-1554.11. Jiang, H., Zhang, Y, Zeng, Z., Ji, J., Wang, Y, Chi, Y, & Miao, C. (2021 , May). Mobile-based Clock Drawing Test for Detecting Early Signs of Dementia. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 35, No. 18, pp. 16048-16050).12. Kingma, D. P., & Ba, J. (2014). Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980.13. Libon, D. J., Malamut, B. L., Swenson, R., Sands, L. P., & Cloud, B. S. (1996). Further analyses of clock drawings among demented and nondemented older subjects. Archives of Clinical Neuropsychology, 11(3), 193-2014. Park, I., & Lee, U. (2021). Automatic, qualitative scoring of the clock drawing test (CDT) based on u-net, CNN and mobile sensor data. Sensors, 21(15), 5239.15. Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L. & Desmaison, A. (2019). Pytorch: An imperativestyle, high-performance deep learning library. Advances in neural information processing systems, 32.16. Raghu, M., Unterthiner, T., Kornblith, S., Zhang, C., & Dosovitskiy, A. (2021). Do vision transformers see like convolutional neural networks?. Advances in Neural Information Processing Systems, 34, 12116-12128.17. Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., & Chen, L. C. (2018). Mobilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 4510-4520).18. Sato, K., Niimi, Y, Mano, T., Iwata, A., & Iwatsubo, T. (2022). Automated evaluation of conventional clock-drawing test using deep neural network: Potential as a mass screening tool to detect individuals with cognitive decline. Frontiers in neurology, 13, 896403.19. Singh, M., Gustafson, L., Adcock, A., de Freitas Reis, V., Gedik, B., Kosaraju, R.P., Mahajan, D., Girshick, R., Dollar, P, & Van Der Maaten, L. (2022). Revisiting weakly supervised pre-training of visual perception models. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (pp. 804-814).20. Souder, E., O'Sullivan, P, & Pechenik, G. (1999). Comparison of scoring criteria for clock drawing test. Journal of Clinical Geropsychology, 5, 139-145.21. Souillard-Mandar, W., Davis, R., Rudin, C., Au, R., Libon, D. J., Swenson, R., Price, C. C., Lamar, M. and Penney, D. L. (2016). Learning classification models of cognitive conditions from subtle behaviors in the digital clock drawing test. Machine learning, 102, 393-441.22. Suhr, J., Grace, J., Allen, J., Nadler, J., & McKenna, M. (1998). Quantitative and qualitative performance of stroke versus normal elderly on six clock drawing systems. Archives of Clinical Neuropsychology, 13(6), 495-502.23. Tang, E. Y, Harrison, S. L., Errington, L., Gordon, M. F, Visser, P. J., Novak, G., Dufouil, C., Brayne, C., Robinson, L., Launer, L. J. and Stephan, B. C. (2015). Current developments in dementia risk prediction modelling: an updated systematic review. PloS one, 10(9), e0136181.24. Wimo, A., Seeher, K., Cataldi, R., Cyhlarova, E., Dielemann, J.L., Frisell, O., Guerchet, M., Jonsson, L., Malaha, A.K., Nichols, E. and Pedroza, P. (2023). The worldwide costs of dementia in 2019. Alzheimer's & Dementia.25. Caffarra, P, Gardini, S., Zonato, F, Concari, L., Dieci, F., Copelli, S., Freedman, M., Stracciari, A., Venneri, A. (2011). Italian norms for the Freedman version of the Clock Drawing Test. Journal of Clinical and Experimental Neuropsychology, 2011 , 33 (9), 982-988.26. Darvesh, S., Leach, L., Black, S. E., Kaplan, E., & Freedman, M. (2005). The Behavioural Neurology Assessment. Canadian Journal of Neurological Sciences, 32(2), 167-177.27. Freedman, M., Leach, L., Carmela Tartaglia, M., Stokes, K. A., Goldberg, Y, Spring, R., et al. (2018a). Correction to: The Toronto cognitive assessment (TorCA): normative data and validation to detect amnestic mild cognitive impairment. Alzheimer's Research & Therapy, 10(1), 120.Freedman, M., Leach, L., Carmela Tartaglia, M., Stokes, K. A., Goldberg, Y, Spring, R., et al. (2018b). The Toronto Cognitive Assessment (TorCA): normative data and validation to detect amnestic mild cognitive impairment. Alzheimer's Research & Therapy, 10(1), 65.
Claims
CLAIMS:1 . A computer-implemented method for generating a dementia prediction indicator based on a candidate image, the method comprising: providing, at a memory, a dementia prediction model and an imagecropping model for cropping an input image for the dementia prediction model; receiving, at a processor in communication with the memory, the candidate image; generating, at the processor, a cropped candidate image based on the candidate image and the image-cropping model; and generating, at the processor, the dementia prediction indicator based on the cropped candidate image and the dementia prediction model.
2. The method of claim 1 , further comprising: receiving, at the processor, one or more prediction features, and wherein the generating the dementia prediction indicator is further based on the one or more prediction features.
3. The method of claim 2, wherein the one or more prediction features comprises at least one of: a socio-economic indicator, a driving test score, and a prior medical history.
4. The method of claim 1 , wherein the candidate image comprises an image generated based on a standards-based test. .
5. The method of claim 4, wherein the candidate image comprises a plurality of different shapes.
6. The method of claim 5, wherein the candidate image comprises a clock drawing test.
7. The method of claim 6, wherein the candidate image is captured using a user device camera in the clinical environment.
8. The method of claim 7, wherein generating the cropped candidate image based on the candidate image and the image-cropping model comprises: processing the candidate image using the image-cropping model to produce a heatmap of a probable location of a clock relative to the candidate image; cropping the candidate image to produce an intermediate candidate image such that a minimum proportion of the heatmap is contained within the intermediate candidate image; and cropping the intermediate candidate image to produce the cropped candidate image such that a minimum proportion of a total pixel value sum of the intermediate candidate image is contained within the cropped candidate image.
9. The method of claim 8, further comprising: pre-processing the candidate image at the processor, the pre-processing comprising: converting, at the processor, the candidate image to a greyscale candidate image; applying, at the processor, one or more line filters to the greyscale candidate image to produce a plurality of filtered images, the plurality of filtered images comprising at least one selected from the group of: a first filtered image, a second filtered image, and a third filtered image; and generating, at the processor, the pre-processed candidate image by averaging the plurality of filtered images based on a line weight associated with each of the filtered images.
10. The method of claim 9, further comprising padding the cropped candidate image prior to processing the cropped image using the dementia prediction model.
11. The method of claim 10, wherein each of the one or more line filters is applied to the greyscale candidate image in accordance with one or more pixel values associated with the each of the one or more line filters.
12. The method of claim 11 , wherein the image-cropping model comprises: at least one downsample layer; at least one layer receiving the output of the at least one downsample layer; and at least one upsample layer receiving the output of the at least one layer.
13. The method of claim 12, wherein: the at least one downsample layer comprises at least one convolution downsample layer; the at least one layer comprises at least one convolution layer; and the at least one upsample layer comprises at least one convolution upsample layer.
14. The method of claim 13, wherein the dementia prediction model comprises: at least one median fully connected layer, the median fully connected layer comprising an equal number of inputs and outputs; and an output layer connected to the output of the at least one median fully connected layer, the output layer comprising two outputs.
15. The method of claim 14, wherein the dementia prediction indicator comprises a dementia prediction and a clock drawing test measure.
16. The method of claim 15, wherein the image-cropping and the dementia prediction models comprise machine-learning models.
17. The method of claim 16, wherein the dementia prediction model further comprises a transformer.
18. A system for generating a dementia prediction indicator based on a candidate image received, the system comprising: a memory comprising: a dementia prediction model; and an image-cropping model for cropping an input image for the dementia prediction model; and a processor in communication with the memory, the processor configured to: receive the candidate image; generate a cropped candidate image based on the candidate image and the image-cropping model; and generate the dementia prediction indicator based on the cropped candidate image and the dementia prediction model.
19. The system of claim 18, wherein: the processor is further configured to receive one or more prediction features; and the processor generating the dementia predictor is further based on the one or more prediction features.
20. The system of claim 19, wherein the one or more prediction features comprises at least one of: a socio-economic indicator, a driving test score, and a prior medical history.
21. The system of claim 18, wherein the candidate image comprises an image generated based on a standards-based test.
22. The system of claim 21, wherein the candidate image comprises a plurality of different shapes.
23. The system of claim 18, wherein the candidate image comprises a clock drawing test.
24. The system of claim 23, wherein the processor is further configured to receive the candidate image from a user device in the clinical environment.
25. The system of claim 24, wherein the processor is further configured to generate the cropped candidate image based on the candidate image and the imagecropping model by: processing the candidate image using the image-cropping model to produce a heatmap of a probable location of a clock relative to the candidate image; cropping the candidate image to produce an intermediate candidate image such that a minimum proportion of the heatmap is contained within the intermediate candidate image; and cropping the intermediate candidate image to produce the cropped candidate image such that a minimum proportion of a total pixel value sum of the intermediate candidate image is contained within the cropped candidate image.
26. The system of claim 25, wherein the processor is further configured to pre- process the candidate image at the processor, the pre-processing comprising: converting, at the processor, the candidate image to greyscale candidate image; applying, at the processor, one or more line filters to the greyscale candidate image to produce a plurality of filtered images, the plurality of filtered images comprising at least one selected from the group of: a first filtered image, a second filtered image, and a third filtered image; and generating, at the processor, the pre-processed candidate image by averaging the set of filtered images based on a line weight associated with each of the filtered images.
27. The system of claim 26, wherein the processor is further configured to pad the cropped candidate image prior to processing the cropped image using the dementia prediction model.
28. The system of claim 27, wherein each of the one or more line filters is applied to the greyscale candidate image in accordance with one or more pixel values associated with the each of the one or more line filters.
29. The system of claim 28, wherein the image-cropping model comprises: at least one downsample layer; at least one layer receiving the output of the at least one downsample layer; and at least one upsample layer receiving the output of the at least one layer.
30. The system of claim 29, wherein: the at least one downsample layer comprises at least one convolution downsample layer; the at least one layer comprises at least one convolution layer; and the at least one upsample layer comprises at least one convolution upsample layer.
31. The system of claim 30, wherein the dementia prediction model comprises: at least one median fully connected layer, the median fully connected layer comprising an equal number of inputs and outputs; and an output layer connected to the output of the at least one median fully connected layer, the output layer comprising two outputs.
32. The system of claim 31, wherein the dementia prediction indicator comprises a dementia prediction and a clock drawing test measure.
33. The system of claim 32, wherein the image-cropping and the dementia prediction models comprise machine-learning models.
34. The system of claim 33, wherein the dementia prediction model further comprises a transformer.
35. A computer-implemented method for generating a dementia prediction model, comprising: receiving, at a processor, a plurality of training images; generating, at the processor, a plurality of cropped training images based on an image-cropping model and the plurality of training images; and fine-tuning, at the processor, a machine learning model using the cropped training images and at least one feature associated with each of the cropped training images, the machine learning model comprising a pretrained model.
36. The method of claim 35, wherein the at least one feature comprises at least one of: a manual label, a socio-economic indicator, a driving test score, and a prior medical history.
37. The method of claim 36, wherein the manual label comprises a clock drawing test measure.
38. The method of claim 35, further comprising: pre-processing, at the processor, each training image in the plurality of training images at the processor prior to generating the plurality of cropped training images, the pre-processing comprising: converting, at the processor, the training image to a greyscale training image; applying, at the processor, one or more line filters to the greyscale training image to produce a plurality of filtered images, the plurality of filtered images comprising at least one selected from the group of: a first filtered image, a second filtered image, and a third filtered image; andgenerating, at the processor, the pre-processed training image by averaging the set of filtered images based on a line weight associated with each of the filtered images.
39. The method of claim 38, wherein each of the one or more line filters is applied to the greyscale training image in accordance with one or more pixel values associated with the each of the one or more line filters.
40. The method of claim 35 further comprising dividing, at the processor, the plurality of training images into a training data set and a validation data set.
41. The method of claim 40, further comprising: replicating, at the processor, one or more images associated with one or more image groups in the validation data set and the training data set, each image group being associated with a value of a manual label, until all of the image groups contain an equal quantity of samples.
42. The method of claim 41 , wherein the manual label comprises a clock drawing test measure.
43. The method of claim 35, further comprising: augmenting, at the processor, each training image in the plurality of training images by performing at least one of the following: scaling the training image; translating the training image; rotating the training image; and shearing the training image.
44. The method of claim 35, wherein the fine-tuning the machine learning model comprises:modifying, at the processor, only one or more weights associated with a head of the machine learning model until a mean squared error associated with the validation data set has plateaued; and modifying, at the processor, one or more weights associated with a body of the machine learning model and the one or more weights associated with the head of the machine learning model, the body of the machine learning model comprising the pre-trained model.
45. The method of claim 44, wherein the pre-trained model comprises a transformer.
46. The method of claim 35, wherein the image-cropping model comprises: at least one downsample layer; at least one layer receiving the output of the at least one downsample layer; and at least one upsample layer receiving the output of the at least one layer.
47. The method of claim 46, wherein: the at least one downsample layer comprises at least one convolution downsample layer; the at least one layer comprises at least one convolution layer; and the at least one upsample layer comprises at least one convolution upsample layer.
48. The method of claim 45, wherein the head of the machine learning model further comprises: at least one median fully connected layer, the median fully connected layer comprising an equal number of inputs and outputs; and an output layer connected to the output of the at least one median fully connected layer, the output layer comprising two outputs.
49. The method of claim 35, wherein the plurality of training images comprises drawings of a clock.
50. The method of claim 35, further comprising generating the image-cropping model by: receiving an image-cropping dataset, the image-cropping dataset comprising images comprising clock drawing tests; generating a manually cropped training dataset based on the imagecropping dataset; generating a pre-processed dataset based on the manually cropped dataset; producing a first target bitmap dataset based on the pre-processed dataset by setting one or more first target pixels of the first target bitmap to a 1 value where the target pixel comprises a part of a clock image and 0 otherwise; producing a second target bitmap dataset based on the pre-processed dataset by setting one or more second target pixels of the second target bitmap to 0.2, 0.4, 0.6, 0.8, or 1 , based on a clock rating of the clock drawing test; and training, at the processor, a cropping machine learning model based on the image-cropping dataset, the first target bitmap, and the second target bitmap.
51. The method of claim 50, further comprising dividing the manually cropped training dataset into a cropping training dataset, a cropping validation dataset, and a cropping testing dataset.
52. The method of claim 51 , further comprising augmenting, at the processor, each image in the cropping training dataset, the cropping validation dataset, and the cropping testing dataset by at least one of: adding noise to the image; adding lines to the image; adding writing to the image; andadding rectangles in a random orientation to the image.
53. The method of claim 50, wherein the generating a pre-processed dataset based on the manually cropped training dataset comprises, for each image in the manually cropped training dataset: rotating the image by a random angle; resizing the image by a random factor; and placing the image onto a black background image.
54. A system for generating a dementia prediction model, comprising: a memory comprising a machine learning model; a processor configured to: receive a plurality of training images; generate, at a processor, a plurality of cropped training images based on an image-cropping model and the plurality of training images; and fine-tune the machine learning model using the cropped training images and at least one feature associated with each of the cropped training images, the machine learning model comprising a pre-trained model.
55. The system of claim 54, where the at least one feature comprises at least one of: a manual label, a socio-economic indicator, a driving test score, and a prior medical history.
56. The system of claim 55, wherein the manual label comprises a clock drawing test measure.
57. The system of claim 54, wherein the processor is further configured to: pre-process each training image in the plurality of training images at the processor prior to generating the plurality of cropped training images, the pre-processing comprising:converting the training image to a greyscale training image; applying one or more line filters to the greyscale training image to produce a plurality of filtered images, the plurality of filtered images comprising at least one selected from the group of: a first filtered image, a second filtered image, and a third filtered image; and generating the pre-processed training image by averaging the set of filtered images based on a line weight associated with each of the filtered images.
58. The system of claim 57, wherein each of the one or more line filters is applied to the greyscale training image in accordance with one or more pixel values associated with the each of the one or more line filters.
59. The system of claim 35, wherein the processor is further configured to divide, at the processor, the plurality of training images into a training data set and a validation data set.
60. The system of claim 59, wherein the processor is further configured to: replicate one or more images associated with one or more image groups in the validation data set and the training data set, each image group being associated with a value of a manual label, until all of the image groups contain an equal quantity of samples.
61. The system of claim 60, wherein the manual label comprises a clock drawing test measure.
62. The system of claim 35, wherein the processor is further configured to: augment each training image in the plurality of training images, prior to fine-tuning the pre-trained model, by performing at least one of the following: scaling the training image; translating the training image;rotating the training image; and shearing the training image.
63. The system of claim 54, wherein the fine-tuning the pre-trained model comprises: modifying only one or more weights associated with a head of the machine learning model until a mean squared error associated with the validation data set has plateaued; and modifying one or more weights associated with a body of the machine learning model and the one or more weights associated with the head of the machine learning model, the body of the machine learning model comprising the pre-trained model.
64. The system of claim 54, wherein the pre-trained model comprises a transformer.
65. The system of claim 54, wherein the image-cropping model comprises: at least one downsample layer; at least one layer receiving the output of the at least one downsample layer; and at least one upsample layer receiving the output of the at least one layer.
66. The system of claim 65, wherein: the at least one downsample layer comprises at least one convolution downsample layer; the at least one layer comprises at least one convolution layer; and the at least one upsample layer comprises at least one convolution upsample layer.
67. The system of claim 64, wherein the head of the machine learning model further comprises: at least one median fully connected layer, the median fully connected layer comprising an equal number of inputs and outputs; andan output layer connected to the output of the at least one median fully connected layer, the output layer comprising two outputs.
68. The system of claim 54, wherein the plurality of training images comprises drawings of a clock.
69. The system of claim 54, wherein the processor is further configured to generate the image-cropping model by training a cropping machine learning model based on an image-cropping dataset, a first target bitmap, and a second target bitmap, wherein: the image-cropping dataset comprises images comprising clock drawing tests; the first target bitmap dataset is produced based on a pre-processed dataset by setting one or more first target pixels of the first target bitmap to a 1 value where the target pixel comprises a part of a clock image and 0 otherwise; and the second target bitmap dataset is produced based on the pre-processed dataset by setting one or more second target pixels of the second target bitmap to 0.2, 0.4, 0.6, 0.8, or 1 , based on a clock rating of the clock drawing test; a manually cropped training dataset is based on the image-cropping dataset; and the pre-processed dataset is based on a manually cropped dataset, the manually cropped dataset being based on the image-cropping dataset.
70. The system of claim 69, wherein the manually cropped training dataset is divided into a cropping training dataset, a cropping validation dataset, and a cropping testing dataset.
71. The system of claim 70, wherein the processor is further configured to augment each image in the cropping training dataset, the cropping validation dataset, and the cropping testing dataset by at least one of: adding noise to the image;adding lines to the image; adding writing to the image; and adding rectangles in a random orientation to the image.
72. The system of claim 69, wherein the generating a pre-processed dataset based on the manually cropped training dataset comprises, for each image in the manually cropped training dataset: rotating the image by a random angle; resizing the image by a random factor; and placing the image onto a black background image.
Citation Information
Patent Citations
AD scale hand-drawn cross pentagon classification method based on convolutional deep neural network
CN111652287A
Self-learning-based pentagonal drawing test intelligent evaluation system and method
CN113989588A