System for clinical assessment of a user
Patent Information
- Application Number
- EP2024804932
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-02
- Filing Date
- 2024-11-01
- Publication Date
- 2026-09-09
AI Technical Summary
Current systems for clinical assessment of patients using human activity recognition (HAR) are limited by inadequate training data and hardware, relying on stationary sensors or smartphone data, which are prone to limitations such as recall bias and restricted mobility tracking.
A system that utilizes one or more processors to obtain and process movement data from wearable devices, employing machine learning models to determine movement features and clinical assessments. This system records movement data during natural behavior, reducing the need for self-reporting and intermittent observations.
The system effectively monitors changes in patient behavior, predicting clinical deterioration and guiding rehabilitation by providing a continuous and accurate assessment of patient movement, thus improving early warning systems and reducing preventable deaths and complications.
Smart Images

Figure IMGF000020_0001 
Figure IMGF000020_0002 
Figure IMGF000023_0001
Abstract
Description
[0001] System for clinical assessment of a user
[0002] Field
[0003] The present invention relates to systems for clinical assessment of patients including prediction of health risks / disease progression, clinical diagnosis and monitoring. In particular, the invention relates to systems that use patient movement as an indicator for clinical assessment.
[0004] Background 30% of preventable deaths in hospitals are caused by staff monitoring failures and 70% of complications in care homes are preventable. Changes in patient behaviour are important early indicators of illnesses, reflect progression of illness / rehabilitation, and act as an important early warning sign for abrupt changes in patient condition. For example, patient deterioration occurs in 30% of acute hospital and 50% of care home patients, resulting in increased death, disability, length of stay, and expense. Deterioration can lead to an increase morbidity, mortality, or length of hospital stay and requires active medical management. Therefore, the monitoring of at-risk patients’ movement is a key component to predicting clinical deterioration and guiding rehabilitation. While current hospital mandated deterioration early warning systems automatically and continuously record vitals, such as heart rate, patient movement is not automatically recorded. Current standards for recording patient movement are limited by reliance on self-reports, visual observations and instructed motor assessments recorded at intermittent time points. These are human-dependent, and suffer from recall bias. Movement monitoring for other conditions or risks suffers similar issues.
[0005] Human activity recognition (HAR) from movement data, determined by machine learning (ML), can be used to monitor changes in patient behaviour. However, current systems are limited both by the training data and the hardware. Current training data is normally obtained from an instructed movement task, performed by either a healthy population or a population of people with chronic illness. Current hardware often relies on stationary sensors such as cameras or on the patient’s smart phone. Camera based systems rely on the patient being in the field of view of the camera, the patient not being wholly or partially obscured, and the patient being correctly identified in the video feed by the ML algorithm. Other stationary systems have similar limitations. Using sensors of the patient’s smart phone relies on the smart phone being on the patient’s body, and is further limited by the potentially changing location of where the smart phone is placed on the patient. Aspects of the present invention aim to address some or all of the issues inherent to the prior art methods for clinical assessment of patients using HAR.
[0006] Summary
[0007] According to a first aspect of the present disclosure a system for clinical assessment of a user is provided. The system comprises one or more processors configured to: obtain movement data of a user, the movement data recorded during the user’s behaviour; determine, based on processing the recorded movement data using a first machine learning model, one or more movement features, wherein the one or more movement features comprise a sequence of activity classes and a duration of time the user spends in each of the activity classes, wherein each activity class represents one or more activities performed by the user during the user’s behaviour; determine, using a second machine learning model, a metric indicative of a clinical assessment of the user, wherein the second machine learning model takes as inputs at least the one or more movement features.
[0008] Movement data can be recorded during any of the user’s behaviour, such as natural behaviour, or during specific assessments, exercises, or routines. Many examples in the following use movement data obtained during the user’s natural behaviour. All disclosed systems and methods may be equally utilized during structured or controlled movement tasks such as specific assessments, exercises or routines.
[0009] Optionally, the movement data is recorded during the user’s natural behaviour, and each activity class represents one or more activities performed by the user during the user’s natural behaviour. Using movement data obtained during the user’s natural behaviour means the user’s movement can be recorded during their daily life. Therefore, recording of movement information is less intermittent than methods which rely on controlled movement tasks. Further, the need for self-reporting is reduced and associated recall-bias can be eliminated. Movement data obtained from natural behaviour is less structured than data from controlled movement tasks and can therefore be more difficult to analyse; however, it is recognised that using machine learning models for determining movement features, as described herein, can overcome this difficulty by allowing for efficient pattern recognition / extraction from the data. Using a sequence of activity classes and the period spent in each activity class as movement features increases expandability of system outputs by linking outputs to changes in user behaviour.
[0010] Optionally, the system further comprises at least one device configured to record the movement data, wherein the movement data comprises position, velocity or acceleration data in the linear or rotational domain. In some implementations, each at least one device comprises any one of: an inertial measurement unit, a radar sensor, a camera, a depth sensor, or an ultrasound sensor. Optionally, the at least one device is placed on a limb of the user and comprises an inertial measurement unit or other movement measurement unit (e.g., as discussed above). Optionally, the at least one device is a smartwatch comprising the inertial measurement unit or other measurement unit. In some specific implementations, the at least one device is two smart watches, one placed on each of the user’s arms, optionally the at least one device is four smart watches, one placed on each of the user’s limbs.
[0011] Optionally, obtaining (from the at least one device) the movement data recorded during the user’s behaviour comprises: obtaining, from an inertial measurement unit, IMU, of each at least one device, triaxial acceleration data and triaxial angular velocity data.
[0012] Optionally, the activity classes comprise: being in a bed, sitting on a chair, and / or mobilizing. In some implementations, the first machine learning model comprises a first neural network, the first neural network comprising a hybrid convolutional neural network and long short-term memory with recurrent attention layers.
[0013] Optionally, the first machine learning model is trained by supervised learning using labelled, raw movement data from a first reference population, wherein the raw movement data from the first reference population is obtained from each particular user in the first reference population by obtaining, for each particular user in the first reference population, movement data recorded during the particular user in the first reference population’s (optionally natural) behaviour, and wherein each label corresponds to one of the activity classes. In some implementations, the second machine learning model comprises a second neural network. Optionally, the second neural network comprises a global local attention model. Optionally, using the global local attention model comprises: applying a local attention mechanism using multi-head attention to each of a plurality of data windows, wherein a null token is embedded in each window, extracting the null token for each window, incorporating the null tokens for each window in a sequence with a global token, and applying a global attention mechanism using multi-head attention to the sequence. In other examples, the second neural network is an artificial neural network with one hidden layer containing eight hidden nodes using a ReLU (rectified linear unit) activation function. The output layer uses a sigmoid activation function.
[0014] Optionally, the second machine learning model is trained by supervised learning on a second reference population, wherein a training data set comprises: for each particular user in the second reference population, one or more movement features of the user in the second reference population, wherein the one or more movement features of the particular user in the second reference population comprise a sequence of activity classes and a duration of time the particular user in the second reference population spends in each of the activity classes, and for each particular user in the second reference population, a metric indicative of the outcome of a clinical assessment of the particular user in the second reference population. Optionally, the training data set further comprises, for each particular user in the second reference population, one or more clinical markers for the particular user in the second reference population, wherein the same clinical markers are used for all users in the second reference population.
[0015] Optionally, the one or more processors are further configured to: obtain one or more clinical markers, wherein the one or more clinical markers comprise at least one feature of a plurality of features describing the user, wherein to determine, using the second machine learning model, the metric indicative of a clinical assessment of the user, the second machine learning model further takes the one or more clinical markers as further inputs.
[0016] Optionally, the one or more movement features further comprise one or more of: statistical movement data obtained from the movement data, movement pattern data extracted from the sequence of activity classes and the duration of time the user spends in each of the activity classes, or data extracted from the movement data and / or from the sequence of activity classes and the duration of time the user spends in each of the activity classes using a further machine learning model. Optionally, wherein the further machine learning model is neural network comprising convolutional, recurrent and attention layers.
[0017] Optionally, the metric indicative of a clinical assessment of the user is a clinical deterioration risk prediction, a prediction of disease progression, or a clinical diagnosis. Optionally, clinical deterioration is defined as an adverse change in a user’s health likely to increase morbidity, mortality, or length of hospital stay and which requires active medical management, optionally, wherein an adverse change in the user’s health comprises any one of: new cases of illness or infection such as sepsis, delirium, renal failure, stroke, urinary tract infection, thrombo-embolism, myocardial infection, or early neurological deterioration. Optionally, the user’s clinical deterioration risk prediction comprises the user’s clinical deterioration risk within 69 hours after the movement data is recorded during the user’s behaviour.
[0018] Optionally, the one or more processors are further configured to: obtain one or more additional metrics, the one or more additional metrics including at least one of heart rate sensor data, step count, distance walked, or fall detection information. The second machine learning model can be configured to take the one or more additional metrics as input.
[0019] Optionally, the system further comprising a receiving unit configured to receive the metric indicative of a clinical assessment of the user from the one or more processors, the receiving unit further configured to at least display information based on the metric indicative of a clinical assessment of the user.
[0020] The processors can be implemented in any suitable way, using any suitable means or modules. For instance, the system for clinical assessment of a user can comprise: a first processing module configured to obtain movement data of a user, the movement data recorded during the user’s (optionally natural) behaviour; a second processing module configured to determine, based on processing the recorded movement data using a first machine learning model, one or more movement features, wherein the one or more movement features comprise a sequence of activity classes and a duration of time the user spends in each of the activity classes, wherein each activity class represents one or more activities performed by the user during the user’s (optionally natural) behaviour; and a third processing module configured to determine, using a second machine learning model, a metric indicative of a clinical assessment of the user, wherein the second machine learning model takes as inputs at least the one or more movement features.
[0021] Also disclosed is a method for clinical assessment of a user, the method comprising: obtaining movement data of a user, the movement data recorded during the user’s (optionally natural) behaviour; determining, based on processing the recorded movement data using a first machine learning model, one or more movement features, wherein the one or more movement features comprise a sequence of activity classes and a duration of time the user spends in each of the activity classes, wherein each activity class represents one or more activities performed by the user during the user’s (optionally natural) behaviour; determining, using a second machine learning model, a metric indicative of a clinical assessment of the user, wherein the second machine learning model takes as inputs at least the one or more movement features.
[0022] Brief Description of the figures
[0023] The detailed description is with reference to the following figures. Like reference numerals refer to like features.
[0024] Figure 1 shows an example system for clinical assessment of a user according to the present disclosure.
[0025] Figure 2 shows (a) an example for the positioning of devices for collection of movement data and (b) corresponding example movement data according to the present disclosure.
[0026] Figure 3 shows an example neural network for classifying movement data, obtained during the user’s behaviour, into activity classes.
[0027] Figure 4 shows the results of an experiment comparing different neural networks for movement data classification. Figure 5 shows the results of an experiment examining the effectiveness of movement data classification, using the neural network of Figure 3.
[0028] Figure 6 shows the attention mechanism of an example neural network for determining a marker or metric indicative of a clinical assessment of the user.
[0029] Figure 7 shows the architecture of the example neural network implementing the attention mechanism of Figure 6. Figure 8 shows the results of an experiment assessing the performance of the neural network of Figure 7.
[0030] Figure 9 shows the results of another experiment assessing the performance of the neural network of Figure 7. Figure 10 shows an example configuration of the system for clinical assessment of a user according to aspects of the present disclosure.
[0031] Detailed Description
[0032] The present disclosure relates to a human activity recognition system configured to classify the movement of a user during their behaviour and then to determine a clinical assessment of the user based on said movement. A clinical assessment as used herein can refer to diagnosis, monitoring, prediction of disease progression or health risk prediction. The user’s behaviour can be any behaviour. Optionally the behaviour is the user’s natural behaviour, and many examples provided herein utilize the user’s natural behaviour in daily life. Further examples of the user’s behaviour are their behaviour during specific assessments, exercises, or routines.
[0033] A processor (optionally a remote server, or any other suitable computing device) receives movement data from one or more devices. At the processor, this movement data is processed by two machine learning models; the processing is configured to first classify the activities the user has engaged in during their behaviour into a sequence of activity classes, and then to determine a metric indicative of the user’s clinical assessment based on the activity class information. The metric indicative of the user’s clinical assessment can be further based on otherwise processed aspects of the movement data, and optionally additional clinical data or otherwise relevant data. The processor then sends / transmits the metric indicative of the user’s clinical assessment (or other data indicative of the user’s clinical assessment), and optionally additional data such as activity class information or other movement features, to a receiving unit and / or a web-interface for display. The processor can additionally or alternatively output the metric for display / rendering on a display device associated with the processor (optionally integral with the processor, e.g. as part of a smart phone or tablet or other personal computer).
[0034] Figure 1 shows an example embodiment of a system too for clinical assessment of a user according to aspects of the present disclosure. The system obtains raw movement data 110 from at least one device 170. In some examples, movement data 110 comprises position, velocity or acceleration data of the user (optionally the user’s limbs) in the linear or rotational domain, or such data is derivable from the movement data. In one example, movement data no is obtained from a device placed on one of the user’s limbs. In another example, movement data is obtained from four devices 170 placed on each of the user’s wrists and ankles. These sensor locations are unobtrusive to the user and capture overall movement of the user during natural behaviour and / or instructed behaviour.
[0035] The device(s) 170 may be any device suitable for collecting movement data. In some examples, device(s) 170 collects movement data using an inertial movement sensor, a radar sensor, a camera, a depth sensor, an ultrasound sensor and / or any other sensor suitable for collecting movement data 110. The sensor(s) for collecting movement data 110 maybe an active or passive sensor and maybe remote to the user and / or placed on the user. In some examples, the device(s) 170 is a wearable device that may be clothes, glasses, watches, or any other suitable wearable. A device 170 may comprise multiple sensors (e.g., a camera and an inertial movement sensor) or only one sensor. If multiple devices 170 are used, the sensors in each device need not be the same. For example, one device 170 may include a depth sensor, and another device 170 may include radar and a camera.
[0036] In further, specific examples, the device(s) 170 is a smart watch with an app installed to extract movement data 110. In other examples, the device(s) 170 is a smart watch that sends or provides the movement data 110 directly from the sensors to a server or other further processing unit 1020 without any local processing. In some cases, a smart watch records movement data 110 using built in inertial movement sensors. Smart watches are highly portable, have high sending range (as long as Wi-Fi is available) and short post-processing latency. In further specific examples, the device(s) 170 include a research-grade inertial sensor such as Xsens or Axivity. In many examples, the movement data 110 is recorded during the user’s activities such as sleeping, reading, watching TV, eating, walking around, doing exercise, etc, and represents movement during the user’s natural behaviour. Collecting data during the user’s natural behaviour allows for continuous monitoring of the user and reduces the need for a supervised movement task to assess the user. Further, the automated data collection avoids issues involved by direct human assessment of movement, such as recall-bias. Any suitable communication method between the device(s) 170 and the processor(s) 1020 maybe used such that processors 1020 obtain the movement data 110. In some examples, short range communication methods (such as Bluetooth or BLE) maybe used to provide the movement data from the device(s) 170. In other examples, long range communication methods (such as Wi-Fi) may be used. As described herein, the communication is wireless to facilitate ease of data collection during the user’s natural behaviour, but wired communication could be used in some implementations.
[0037] The raw movement data 110 is processed by a movement feature determination module 120 to determine movement features 130. The movement features 130 comprise a sequence of activity classes corresponding to each of the user’s daily activities as they progress through their day, as well as the duration of time the user spent in each activity class each time they engaged in an activity belonging to that class and / or the cumulative duration the user has spent in each activity class throughout the recording period. Each activity class represents one or more different activities performed by the user during natural behaviour. In other words, each activity class comprises or contains a set of activities. For example, sleeping may fall under the activity class “being in bed”, walking around may fall under the activity class “mobilizing” and watching TV may fall either under the activity class “being in bed” or “sitting on a chair”. Movement feature determination module 120 includes a first machine learning algorithm, for example a first neural network, for movement feature 130 determination. Activities can fall into different activity classes, depending on the particular classification approach taken and / or the algorithm used. In some examples, movement data of a reference population is used to train the movement feature determination module 120.
[0038] Optionally, movement feature determination module 120 may determine further movement features 130 in addition to the sequence of activity classes and the amount of time the user spends in each activity class. In some examples, the movement feature determination module 120 determines further movement features 130 from raw movement data 110. In an example of this, movement feature determination module 120 determines statistical quantities of the movement data from raw movement data 110 to be used as further movement features 130. In further examples, movement feature determination module 120 analyses the sequence of activity classes and the amount of time the user spends in each activity class to determine further movement features 130. For example, a movement feature 130 of this type maybe a pattern in the sequency of activity classes and the amount of time the user spends in each activity class, such as the frequency and length of wakeful periods the user experiences during the night. Movement feature determination module 120 can comprise a further machine learning model for analysing the sequence of activity classes, the amount of time the user spends in each activity class and / or raw movement data 110 to determine additional movement features 130. In yet further examples, the movement feature determination module 120 comprises a plurality of further machine learning models trained to extract various pieces of data from movement data 110 and previously determined movement features 130. In some examples, movement feature determination module 120 determines further movement features 130 by using a combination of the above methods or other methods. In further examples, other relevant data such as clinical marker data 180 (or other data about the user received by the system) is analysed by one or more of the further machine learning models to determine further markers / features describing the user. In other examples, the first machine learning model may be used to determine additional / further movement features in addition to the sequence of activity classes and duration of time. The movement features 130 including the sequence of activity classes, the duration of time the user spent in each activity class, and any optional movement features 130 are sent to clinical assessment module 150. Clinical assessment module 150 includes a second machine learning algorithm, for example a second neural network. In some examples, movement data previously obtained or determined for each user in a reference population, is used to train clinical assessment module 150.
[0039] Optionally, clinical assessment module 150 may also receive clinical marker data 180 for the user and be trained on clinical marker data for the reference population. Clinical marker data 180 is data obtained by self- reporting, clinicians, and / or care providers about the user of system too. Clinical marker data 180 is discussed in more detail below.
[0040] The clinical assessment module 150 processes the movement features 130 (and optionally the clinical marker data 180 or other relevant markers / features) to determine or predict a metric indicative of a clinical assessment of the user 160. In some examples, a metric indicative of a clinical assessment 160 is a risk prediction such as a user’s deterioration risk, a user’s risk of falling or any other prediction. In further examples, a metric indicative of a clinical assessment 160 is a diagnosis of a condition or disease. In further examples, a metric indicative of a clinical assessment 160 is an evolving status of a user as the system receives new movement data 110 (or other data) from the user. In some examples, the metric may update in set (periodic) time intervals (e.g. every minute, every hour, every day) based on the new movement data 110 received for the user. In such implementations, system too can be used for monitoring a user’s health. In some examples, the metric indicative of a clinical assessment 160 comprises a combination of multiple risk predictions, diagnoses and / or monitored states of the user. For example, the metric indicative of a clinical assessment 160 may be a diagnosis of sepsis and a prediction of a high deterioration risk.
[0041] In some examples, the metric indicative of a clinical assessment of the user 160 comprises the user’s risk of clinical deterioration. In some examples, the metric indicative of a clinical assessment of the user 160 comprises the risk of deterioration within 96 hours of device 170 recording movement data 110. In further examples, metric indicative of a clinical assessment of the user 160 comprises an estimate of the user’s National Early Warning Score 2 (NEWS2), which is a clinical scoring system used to predict clinical deterioration and guide appropriate escalation of acutely ill patients. Clinical deterioration for purposes of this disclosure is defined as any adverse change in a user’s health likely to increase morbidity, mortality, or length of hospital stay and which requires active medical management. Examples of such adverse changes in health include new cases of illness or infection such as sepsis, delirium, renal failure, stroke, urinary tract infection, thrombo-embolism, myocardial infection, or early neurological deterioration. Such a broad definition of clinical deterioration generalizes the applicability of the disclosed system for clinical assessment to a broad range medical conditions and illnesses, as well as a range of risk groups (such as the elderly).
[0042] The clinical assessment module 150 outputs a metric indicative of a clinical assessment 160 and / or data indicative of the metric indicative of a clinical assessment 160. The metric indicative of a clinical assessment 160 and / or indicative data can be made accessible to the user and / or health care professionals monitoring or treating the user (authorized persons or authorized users). In one example, the indicative data and / or metric indicative of a clinical assessment 160 may be made accessible via a web- interface. In another example, the indicative data and / or metric indicative of a clinical assessment 160 is sent to a receiving unit 190. The receiving unit 190 maybe a computer, tablet, mobile phone, or any other device suitable for displaying information. For example, the metric indicative of a clinical assessment 160 maybe displayed on a health dashboard used by care providers to monitor the user’s health. In some implementations, receiving unit 190 accesses a web-interface to display the metric indicative of a clinical assessment 160.
[0043] In some examples, the system too further makes movement features 130 or movement data 110 accessible to the user and / or health care professional monitoring or treating the user. This additional data may help a health care professional to assess rehabilitation progress and decide if there is a further need for treatment. In addition, making the sequence of activity classes and the duration the user spends in each activity class accessible via the receiving unit 190 and / or web-interface (or other display) allows authorized persons to correlate behavioural patterns of a user with a prediction / diagnosis / monitoring notification provided by the system too about that user, thereby increasing the expandability of the systems outputs. For example, after the system too outputs a risk of deterioration as a metric of clinical assessment 160, an authorized person may see from the duration of time the user spent in each activity class that the user spent more time in bed and less time mobilizing than previously. The authorized person may then determine that this is the reason why the system output a risk of deterioration for the user.
[0044] In the following description, these aspects and embodiments of system too will be discussed in greater detail. Device(s) 170 for recording of movement data 110 may be any device implementing a sensor for collection of movement data. Numerous sensors, both active and passive, for determination of a user’s movement data are known in the art, such as inertial movement units, radar, cameras, depth sensors, ultrasound etc. In some implementations, the movement data 110 comprises any of position, velocity, acceleration and / or rotation data representing the user’s movement through space. In one example, device 170 is a stationary camera that records the user’s movement over time in a specific area and the position, velocity, acceleration and / or rotation data is extracted from the sequence of images. In another example, device 170 is a camera secured to the user which records the user’s environment, and the position, velocity, acceleration and / or rotation data is extracted from the relative movement of the environment in the sequence of images. Movement data 110 maybe recorded using any suitable sensor. Methods of extracting raw movement data, and specifically any combination of position, velocity, acceleration and / or rotational data, from such sensors are known to the skilled person. One specific embodiment, in which device(s) 170 are securable to a user’s limb, is shown in Figure 2. Figure 2a shows the example placement of devices 170 on a user 200 for collection of movement data 110, and Figure 2b shows example acceleration data 210 received from said devices 170. A user 200 wears at least one device 170 from which movement data 110 is recorded. In some examples, recordings may between 1 to 5 hours long and obtained during natural behaviour of the user 200. In other examples, the recordings may be longer than 5 hours. Any suitable length of recording can be used. A device 170 may be placed on a user’s (user 200) limbs in any one of the positions shown in Figure 2, i.e. on one or more of the user’s wrists and / or ankles. Any suitable device 170 may be used.
[0045] In one example, wrist sensors are placed near the end of the ulna and radius, and ankle sensors are placed just above the lateral malleolus. These device positions are unobtrusive and do not impact user mobility. User 200 may wear one, two, three or four devices 170. Using four devices 170 allows the recording of asymmetrical patterns of upper and lower limb weakness (right, left, or bilateral). Asymmetrical patterns are indicative of, for example, strokes. However, reducing the number of sensing devices 170 reduces cost of the system, while continuing to enable movement feature determination and clinical assessment of the user 200, for example deterioration risk determination. In one example, using two devices 170 placed on either wrist records left and right limb asymmetry and provides comparable accuracy in deterioration prediction to using four devices.
[0046] Device 170 maybe any device which is suitable for recording movement data. The device 170 can be a wearable device. In some particular examples, wearable device 170 is a consumer smartwatch. Some advantages of using smartwatches include their high portability (no laptops or receiver stations), high range (unlimited in areas with Wi-Fi access), and short post-processing latency (direct Wi-Fi transmission to cloud servers or remote processors 1020). Consumer smartwatches are also low cost, with a discreet design, high user acceptance and adoption by the public, and are easy to program and to provide custom software updates through associated App stores. In some other particular examples, device 170 utilizes an inertial movement unit (IMU) to record movement data 110. The IMU may be in a smartwatch, or in any other suitable sensor or device. Compared to other motion tracking instruments, IMUs have increased computational efficiency, low cost, high accuracy, and are compact. In some other examples, device 170 is a research-grade inertial sensor such as Xsens or Axivity.
[0047] In some examples, research-grade IMU’s maybe secured to the user, for example as shown in Figure 2a. In further examples, device 170 can utilize any other sensor suitable for tracking motion such as radar, cameras, depth sensors, ultrasound, and many others.
[0048] Device 170 maybe secured to user 200 as shown in Fig. 2a, device 170 maybe secured to the user in a different position or device 170 may not be secured to the user at all. Although the following is described with respect to a specific example of movement data from an IMU secured to a user’s limbs, similar principles can be applied to the output from any other sensor for tracking motion. In other words, the movement data described herein can comprise data obtained from sensors (IMUs, radar, cameras, depth sensors, ultrasound, etc.,) of the device 170 and the device 170 need not be placed on the user. In some specific example implementations, movement data 110, specifically acceleration data and rotation data, recorded by the respective IMU of each device 170 comprises at least acceleration data and angular velocity data, in all three dimensions or orientations (e.g. X, Y, Z), over the measurement duration. This triaxial acceleration and triaxial angular velocity data allows three-dimensional determination of the movement of the user. In the embodiment shown in Figure 2a, the triaxial acceleration and triaxial angular velocity data allows three-dimensional determination of the movement of each of the user 200’s limbs which are equipped with a device 170.
[0049] Figure 2b shows an example of triaxial acceleration data 210 obtained from four smartwatches placed on each of user 200’s limbs, one on each wrist and ankle as shown in Figure 2a. Data is recorded for 4 hours. The recorded time-acceleration traces 220 are shown for each limb and for each axis. Blue (or black) indicates the X-axis, orange (or dark grey) the Y-axis and green (or light grey) the Z-axis of the triaxial acceleration data. In this particular example, the smart watch is an Apple Watch. Apple Watches have a triaxial acceleration measurement range of ± 8 g for Series 3, and ± 16 g for
[0050] Series 5 (wherein 1 g corresponds to 9.80665 m / s2) and triaxial angular velocity measurement range of ± 1000 degree / s for Series 3, and ± 2000 degree / s for Series 5. The triaxial acceleration and triaxial angular velocity data is sampled at a rate of too Hz. In some examples, an app extracts the movement data 110 including triaxial acceleration data 210 and, when connected to Wi-Fi, sends the movement data 110 to a server running at least movement feature determination module 120. In other examples, the movement data 110 is sent to processor 1020 (optionally a remote server) without an app as intermediary.
[0051] In another example, device 170 may be a research-grade inertial sensor such as an Xsens. A research grade inertial sensor may record packet-stamped triaxial acceleration (range: ± 16 g), triaxial angular velocity (range: ± 2000 degree / s), and triaxial magnetic fields (range: ± 1.9 x io-4T) data at a sampling rate of too Hz. The recorded movement data 110 is transmitted to a base station and ultimately to movement feature determination module 120. Though specific examples are given, any IMU maybe used whose range and sampling rate enables measurement of typical human movement parameters (typical limb acceleration range is below <6g and the frequency <15HZ).
[0052] Movement feature determination module 120 receives processed / unprocessed movement data 110 (as required, based on the particular algorithm and type of device 170) representing the user’s natural behaviour and determines, using a first machine learning model, the sequence of activity classes the user’s natural behaviour belongs to, as well as the duration the user spent in each activity class. In the following examples, the movement data is unprocessed movement data 110. An advantage in using unprocessed movement data is that all patient movement information is available without bias towards any expected change in movement behaviours. In one example, a small-time segment (e.g. sub-second or a few seconds) of the movement data 110 is analysed by movement feature determination module and the user’s activity class for that duration determined. In some examples, the time segment is 2 seconds long, as this time segment has been found to be long enough for the user to transition from one activity to another, while limiting the instances of two activities in one time segment.
[0053] The thus obtained sequence of activity classes comprises the activity class determined for each time interval. In some examples, there may be time intervals where the user’s activities cannot be classified, such as when the user is transitioning between two activities belonging to different activity classes. From this sequence of activity classes, the duration the user spends in each activity class is determined. The duration the user spends consecutively in each activity class without interruption by spending time in other activity classes is determined, for each instance of the user engaging in any of the activity classes. Further, the duration the user spends in each activity class throughout the whole recording time is determined.
[0054] The user’s natural behaviour includes activities such as sleeping, watching TV, reading, getting dressed, playing a game, going for a walk, working, or any other activity the user might engage in. Activity classes group these activities together, and can be considered as sets of activities. For example, an activity class “in bed” includes any of the different activities the user performs while in bed. Similarly, an activity class “sitting in chair” includes any of the different activities the user performs while sitting.
[0055] Examples of activity classes are “sitting in bed”, “lying in bed”, “sitting in chair”, “standing”, or “walking” (sometimes referred to herein as “Posture” activity classes). Further examples of activity classes are “eating”, “drinking”, “reading”, “dressing”, “writing”, “electronic” (i.e. using an electronic device such as a phone or tablet), or “communication” (sometimes referred to herein as upper-body activity classes). Further activity classes of any type may be defined. Activity classes may also be broader than those described here. For example, a broader activity class “in bed” may contain both the activity class “sitting in bed” and “lying in bed”. Similarly: “meal-time” may contain “eating” and “drinking”; “mobilizing” may contain “standing” and “walking”; and “daily activities” may contain the four narrower activity classes of “reading”, “dressing”, “writing” and “electronic”. In some examples, movement feature determination module 120 classifies the movement data into three activity classes: “being in bed”, “sitting on chair” and “mobilizing”. In other examples, movement feature determination module 120 classifies the movement data 110 into the activity classes: “sitting in bed”, “lying in bed”, “sitting on chair”, “standing” and “walking”. In still further examples, the movement feature determination module 120 classifies the movement data 110 into a different group of activity classes. Any suitable combination of activity classes can be used.
[0056] Movement feature determination module 120 may comprise any appropriate first machine learned model to classify movement data 110 collected during natural behaviour of the user into activity classes. In some examples, the first machine learned model is a first neural network which has been trained to classify movement data 110 collected during natural behaviour of the user into activity classes. Since the first machine learned model is trained on natural behaviour, which is more varied and noisier than instructed behaviour (and therefore can be more difficult to interpret), the model can be used to classify both natural and instructed behaviour, as required by a given use case.
[0057] In some embodiments, to automate the extraction of features from multivariate time series, such as movement data no, deep learning techniques such as Convolutional, Recurrent, and Attention layers are utilized. Convolutional layers, which specialise in extracting repeating features from data with lattice-like structures, are employed to hierarchically extract spatial patterns by analysing the relationships between different measurement axes and identifying localised repeating structures. The Recurrent layers are used to process sequential data through the memory-like structure to model short- to medium-term temporal dynamics from the movement data no through the layer's natural sequence comprehension and long-term temporal dependencies through a hierarchical convolutional condensing.
[0058] Using a multisensor modality for recording movement data introduces additional complexity to the analysis. One example of a movement recording using a multisensor modality is shown in Figure 2a, however any suitable sensor combination may be used.
[0059] Employing a sensor attention mechanism addresses the additional complexity. This mechanism dynamically weighs the features generated by each sensor, considering their relative importance. By doing so, it is ensured that crucial information from each sensor is appropriately considered and propagated to subsequent layers in the deep learning model. This attention mechanism enhances the feature extraction process, providing a more comprehensive representation of the multisensor time series data. By hierarchically combining Convolutional, Recurrent, and Attention layers, a robust deep learning architecture is created that systematically and automatically extracts essential features from the complex multisensor time series data.
[0060] In one embodiment, shown in Figure 3, the first machine learning model of the movement feature determination module 120 is a first neural network comprising a hybrid Convolutional Neural Network (CNN) and Long Short-Term Memory (LSTM) with Recurrent Attention (RA) and Skip connections. The combination model / neural network unites the various advantages of the different component network types. In addition to the advantages laid out in the above embodiment, CNNs are a very effective tool in image recognition and classification tasks. LSTM, as a specialized form of recurrent neural network RNN, are suited to learning any temporal relationship while avoiding the “vanishing gradient” problem that arises when long input sequences are used to train an RNN. An Attention component models the point-to-point relationship between points of data, independent of their temporal relationship, and are often used in conjunction with RNNs. Consequently, a hybrid CNN and LSTM model with Attention is well suited to multimodal time series data, automatically learns features and semantics of activities, improves model generalisability, and captures spatial and temporal relationships in the data. This is just one example of the first neural network, and any other suitable model / network can be used instead, as would be understood by the skilled person.
[0061] An example structure for the specific implementation of a hybrid CNN-LSTM is shown in Figure 3. In the hybrid CNN-LSTM model, first each of the movement data from the four limb IMUs 310 are processed by the neural network individually. Second, 1) the right and left for the upper limb data streams are merged and processed further, and 2) the right and left lower limb are merged and processed further. Third, the upper and lower limb IMU features are merged and processed further. Fourth, the resulting features are fed into a LSTM 340 with a Softmax Dense Layer 350 that outputs a Class Distribution 360. While the hybrid CNN-LSTM takes movement data from up to four
[0062] IMUs, fewer IMUs may be used in other implementations, as shown for example in Figure 5. Alternatively, the movement data can be movement data from any suitable sensor of device 170 (camera, radar, ultrasound, etc.), as discussed above. Alternatively, the device 170 may be placed elsewhere on the user or may be remote from the user.
[0063] The model 300 utilises spatial-temporal convolutional blocks (S-T Conv) 320 as hierarchical feature extractors, which provide base feature maps to the recurrent layers 330. In spatial-temporal convolutional block 320, convolutional layers 322 - 326 are grouped together and residual connections 328 are implemented. Residual or skip connections are unique connections in a neural network model introduced by He et al. in 2015 (He, K., Zhang, X., Ren, S., & Sun, J.. (2015). Deep Residual Learning for Image Recognition, abs / 1512.03385. http: / / arxiv.org / abs / 1512.03385). In residual connections, the input can flow through the block to the output unobstructed, which results in a “gradient highway” in which gradients can backpropagate with minimum loss. Consequently, the neural network can be deeper, allowing for more complex models. Batch normalisation (ReLu BatchNorm) is implemented in convolutional blocks 320, as batch normalisation can help reduce gradient shift during training and speeds up the training process. A 2-Stage convolutional kernel is utilized to extract both temporal and spatial abstract representations efficiently. Lastly, a 1x1 convolution 322 invented by
[0064] Lin et al. is utilized. 1x1 convolution or Network-In-Network layers is a convolution layer with a kernel size of 1. Since the kernel is of that size, Network-In-Network layers essentially learn a point-to-point nonlinear mapping (Lin, M., Chen, Q., & Yan, S.
[0065] (2013). Network in network. arXiv preprint arXiv: 1312.4400.), thus allowing the models with Network-In-Network layers to artificially increase their complexities nearly costlessly. Furthermore, Network-In-Network layers also grant greater information sharing via cross-channel pooling.
[0066] The recurrent attention layer 330 is defined as the following: y(t) = citxt
[0067] Where ytis the output of the layer at time t, cttis the recurrent attention weight computed across all sensors at time t, and xtis the input from all the sensors at time t.
[0068] To compute the recurrent attention weight, the softmax function is utilized as follows: Where ctt,i represent the attention weight at time t for sensor i and et,i represent the energy value at time t for sensor i. The energy function can be defined as follows:
[0069] Where e is an unnormalized energy, Woutis the output weight, Wh is the recurrent weight, and Wxis the input weight. This recurrent attention layer 330 can process cross-sensor temporal attention, which allows the attention layer to not just attend to the sensor data vector internally but also access information from the adjacent sensor vectors and thereby learn the temporal relationship within the data as efficiently.
[0070] After the convolutional stage, the long-term temporal dependencies are modelled using LSTM layers 340. The resulting hidden states are then aggregated before presenting to the classification head 360 with softmax as an output function. In some examples, the hybrid CNN-LSTM model 300 or any other suitable machine learning model / neural network (first machine learned model 300) is trained on data obtained from a suitable reference population. Such a suitable reference population contains users with characteristics relevant to the specific application. For example, if the metric indicative of a clinical assessment of the user is a diagnosis of a medical condition, then the training reference population may contain users with this condition. In some examples, the first machine learned model 300 is trained on a population at risk of deterioration. In one specific example described herein, the first machine learned model 300 was trained on a stroke ward population, as is described below. Often, stroke ward patients are at risk of deterioration. The stroke ward population included patients with both strokes and non-strokes. Non-strokes include conditions such as vestibular neuronitis, migraine, COVID-19 pneumonia, delirium, included cognitive disorder, urinary tract infection, seizure, benign paroxysmal positional vertigo (BPPV), Multiple Sclerosis relapse, and unknown aetiology. The diversity of diagnosis, and diversity of stroke ward patients in terms of medical histoiy, age and other markers, can help to ensure that neural networks trained on this dataset are generalizable to a large range of users, such as users in other acute medical wards, nursing homes, or community care. The training data was collected directly from stroke ward recovery patients using four devices implementing IMU sensors arranged as shown in Figure 2, i.e., left wrist, right wrist, left ankle, and right ankle. To obtain ground truth data, a video showing the patient’s activities was also recorded using a stationary camera. Two human annotators independently watched the videos and classified the movement into activity classes, where only time segments where the annotators agree are labelled. Video ground truth data was aligned with the recorded movement data by performing a visually identifiable movement, that is also visible in the movement data (e.g. shaking the sensors for 3 seconds). Two types of movement data were recorded: a task-controlled dataset and a natural behaviour “in-the-wild” dataset.
[0071] • In a task-controlled dataset, patients perform five everyday activities as instructed, e.g., Lying down, Sitting in Bed, Sitting in Chair, Standing (if possible), and Walking (if possible). Each activity lasts for 1 minute and repeats two times. The annotation is based on the time each activity starts, coupled with video verification. A total of 24 patients participated in this dataset; 18 equipped with research-grade Xsens devices and 6 equipped Apple Watches. • In a natural behaviour “in-the-wild” dataset, instead of dictating patients’ activities, patients move freely with sensors attached. The activities are annotated based on the patients’ video feeds using two annotators. The quality of the annotations is ensured by keeping Cohen’s kappa agreement score above
[0072] 0.7. Additionally, no label will be given if the patients are out of view or engaging in an undeterminable activity. A total of 68 patients participated in this dataset; 34 from Xsens and 34 from Apple Watches cohort. In both datasets, data was collected using research-grade IMU sensors, Xsens MTw Awinda, and an Apple Watch System with an app installed for data collection from the IMU sensors. Furthermore, both “Posture” activity classes (e.g., lying in bed, sitting in bed, sitting in chair, standing, and walking) and upper-body activity classes (e.g., eating, drinking, reading, dressing, writing, electronic device usage, communication) are annotated. In addition, Xsens IMU’s and Apple Watch’s internal sampling rate is set to too Hz to collect the data at the highest resolution possible.
[0073] For training, any combination of the above data sets may be used. For example, training may use: only the natural behaviour data of either the research grade sensor or smart watch; only the controlled task data of either the research grade sensor or smart watch; or any combination of the datasets, including all four datasets. It is understood that the above datasets are exemplary only, and a new dataset having the same characteristics may be generated and used for training. Furthermore, a wide variety of populations maybe used for obtaining a training dataset. Some examples of populations suitable for training data collection may be a healthy population, an intensive care ward population, a stroke ward population, a care home population, a chronic or acute disease population, an population of elderly people, any other population with characteristics putting them at higher risk of experiencing deterioration or any other population having characteristics reflecting the clinical assessment to be made by system too. In other words, training may be based on any suitable population needing monitoring, requiring a clinical diagnosis, or having a risk of unfavourable disease progression.
[0074] In preparation for training, the recording of movement data is down sampled from too Hz to 40 Hz, double the Nyquist frequency of the upper limit of human limb movement. This avoids high frequency artifacts. The continuous movement data stream is segmented using a fixed size sliding window method with a half-size step. A 2 second window is used, as this gives stroke ward patients enough time to finish transitioning from one activity to another while keeping temporal artefacts such as multiple activities transition to a minimum. The labels are assigned to the segment based on a majority vote on the stream of labels.
[0075] For evaluation, leave-one-subject-out cross-validation is used for the task-controlled datasets. For the natural movement task, 6-fold cross-validation is used. Additionally, since the natural tasks dataset can be unbalanced, the training dataset is both subsampled and oversampled to ensure that every label has an equal representation in the training dataset.
[0076] The trained model’s performance is evaluated using weighted F-i measurement that weights each label’s classification results according to their sample proportion. The implementation of a weighted F-i measurement is as follows:
[0077] Figure 4 shows the outcome of benchmarking experiments of the hybrid CNN-LSTM model, in which it is compared to alternative neural networks suitable for activity classification in movement feature determination module 120. The benchmark neural networks are DeepConvLSTM [2] by Ordonez et al. (Ordonez, F. J., & Roggen, D. (2016). Deep convolutional and LSTM recurrent neural networks for multimodal wearable activity recognition. Sensors, 16(1), 115.) and LSTM with Continuous Temporal and Continuous Sensor Attention [3] by Zeng et al. (Zeng, M., Gao, H., Yu, T., Mengshoel, O. J., Langseth, H., Lane, L, & Liu, X. (2018, October). Understanding and improving recurrent networks for human activity recognition by continuous attention. In Proceedings of the 2018 ACM international symposium on wearable computers (pp. 56-63).), both of which are related to the hybrid CNN-LSTM of Figure 3, and two default baseline models: a baseline CNN with four convolutional layers and two dense layers and a baseline LSTM model with two LSTM layers and two dense layers. For benchmarking experiments all neural network models are implemented in PyTorch 1.0. All the models are trained on the above-described dataset using NVIDIA Quadro P6000 with 24GB VRAM and 3840 CUDA cores. All models are trained for at least ten epochs using an ADAM optimiser with a 0.001 learning rate and 0.001 weight decay. The F-i score is shown for all five trained models for determining “Posture” activity classes from the task-controlled dataset (first data column), determining “Posture” activity classes from the natural behaviour dataset (second data column), and determining upper-body activity classes from the natural behaviour dataset (third data column). Task controlled “postures” activity classes (lying in bed, sitting in bed, sitting in chair, standing, and walking) can be classified with relatively high accuracy with all models. Out of all models, the hybrid CNN-LSTM model has the highest F-i score. However, it can be seen that for natural behaviour “postures” and upper-body activities (eating by hand, eating with cutlery, drinking, reading, dressing, writing, electronic device usage, communication), all models’ performance drops considerably compared to the task-controlled dataset, with significant drops in performance in classifying upper body activities. This is attributed to more varied and noisy data when classifying natural movement in general, and the often minimal movements involved in upper body activities such as writing reduce performance further. While the hybrid CNN- LSTM model’s performance for classifying natural “postures” and upper-body activities also decreased, the hybrid CNN-LSTM F-i scores are still ahead of the baseline and benchmark models. In a further experiment, the hybrid CNN-LSTM also performs well on the OPPORTUNITY benchmark dataset, achieving an Fl of 84±2% for four modes of locomotion (stand, sit, lie, and null class).
[0078] Figure 5 shows the weighted Fl scores of further experiments using the hybrid CNN- LSTM. The columns represent, from right to left, 1) Apple Watch smartwatch natural behaviour task model for 4 activity classes (being in bed, sitting in chair, standing, and walking), 2) Apple Watch natural movement task model for 5 activity classes (lying in bed, sitting in bed, sitting in chair, standing, and walking), and 3) for the researchgrade Xsens MTw Awinda natural movement task cohort for 5 activity classes (lying in bed, sitting in bed, sitting in chair, standing, and walking). Each device 170 containing a sensor for recording movement data 110 is labelled according to the position on the user’s limbs (shown to the left of the figure): A = right-hand sensor, B = left-hand sensor, C = right-leg sensor, and D = left-leg sensor. The sensor combination labels on the y axis have not been ordered in any specific way. Merging activity classes from 5 to 4 classes, by merging the activity classes “sitting in bed” and “lying in bed” together into “being in bed”, reduces the difference in model performance between instructed task (Fl 97 ± 2%) and natural movement tasks (Fl 90 ± 1%) (instructed movement task results are not shown). Moreover, model performance for 5 activity classes during natural movement tasks is higher for research-grade sensors (Fl 75 ± 3%) compared to consumer smartwatch sensors (Fl 64 ± 1 %). In reducing the activity classes to 4, all possible sensor combinations achieve moderate performance (Fl 50-70%) and most sensor combinations achieve strong performance (Fl > 70%).
[0079] A particular advantage of the approach described herein is that a single sensor is sufficient for activity class recognition. The 4 and 5 activity class smartwatch models can accurately recognise mobility using only a single sensor on the left leg. For research-grade sensors, four or three sensors (placed on the upper limbs and the right leg) produced the highest performance. Analyses correcting for patient handedness and side of weakness to assess the effect on performance was performed, and found these confounds had an insignificant impact. Subsequent to this experiment, to account for left-handedness, acceleration data between right and left limbs is switched and the inverse of the acceleration y-axis is taken.
[0080] In some examples, movement feature determination module 120 implements a first machine learning model other than a neural network for activity classification. In some examples, support vector machine, random forest, or XGboost models are used. Such models are known to the skilled person. In experiments, XGboost performs with Fl 72% ± 4 for recognizing functional versus non-functional activities based on upper arm movement data. Functional activities include, for example, “Posture” activity classes. Further experiments show that an XGboost model can distinguish 1) communicating and using an electronic device (Fl 87% ± 3) and 2) communicating and eating with cutlery (Fl 81% ± 2) effectively.
[0081] In some examples, movement feature determination module 120 determines additional movement features 130. These features can be any kind of feature obtained by processing of the raw movement data 110 or already determined movement features 130. For example, movement features 130 comprising statistical movement data maybe obtained by calculating statistical quantities of raw movement data 110. In another example, movement features 130 comprise movement patterns extracted from the sequence of activity classes and the duration of time the user spends in each activity class. In other words, the already existing movement features 130 determined by the first machine learning model are processed further to extract movement patterns. In further examples, movement features 130 comprise data extracted from the movement data and / or from the sequence of activity classes and the duration of time the user spends in each of the activity classes using further machine learning models. In some examples, other movement features 130 may also be processed by further machine learning models to determine further movement features. The further machine learning models may be multiple machine learning models determining different movement features or a single machine learning model. In such examples, care should be taken to not over fit the data by repeated processing. Optionally, the first machine learning model is used to determine the additional movement features 130.
[0082] A specific example of how to determine movement features 130 comprising statistical movement data is discussed in the following. It is to be understood that this statistical movement data is exemplary only, and that other statistical movement data may be determined and used as movement features 130.
[0083] In one specific example, in addition to movement data 110 from the user, the system too receives movement data from a reference population. The reference population may be any suitable reference population where movement data has been collected for each user in the reference population according to the present disclosure. In one example, the reference population may be the population used in training the first machine learning model implemented in movement feature determination module 120. In another example, the reference population may be the population used in training the clinical assessment module 150 (which population may be the same as or different than the population used for training the movement feature determination module 120). In further examples, both training populations maybe used as reference populations or new data may be collected from a population at risk of deterioration or other population with the desirable characteristics to use as a reference population.
[0084] Statistical movement data may be calculated or determined for a range of movement data 110 recording durations (e.g., 10 min, 30 min, th, 3I1 and / or 4I1). In some implementations the mean is determined, but other statistical information of the user’s rotation data and acceleration data for each device / sensor (e.g. IMU, radar, ultrasound, camera, etc) maybe determined as well as, or instead of, the mean. For example, the median and the standard deviation may be determined in addition to, or instead of, the mean. Any other suitable statistical movement data may be determined, and a subset of the statistical data maybe selected for use in the second machine learned model.
[0085] The desired set / subset of the determined statistical movement data is sent to clinical assessment module 150 to be used as inputs into the second machine learning model.
[0086] In some implementations, pre-selection of the most relevant of statistical movement data may further improve model performance.
[0087] In some specific examples, as discussed above, the movement data 110 comprises triaxial acceleration data and triaxial angular velocity data for each of the devices 170 placed on the user’s limb. In some of these examples, the first step of determining this statistical movement data is determining the magnitude of the acceleration and angular velocity for each device placed on the user’s limbs, to reduce dimensionality of the data to only magnitude and time. Subsequently, the mean of the user’s acceleration magnitude and acceleration angular velocity for each device / limb is determined.
[0088] Further, the percentage of the user’s acceleration magnitude data and angular velocity magnitude data crossing subject-specific threshold(s) are determined.
[0089] These subject-specific thresholds maybe any percentage or threshold of the subject’s mean acceleration magnitude and mean angular velocity magnitude, such as one or more of >25%, >50%, or >75%. Similarly, the percentage of user’s data samples crossing population specific threshold(s) are determined. The population specific thresholds maybe any percentage, optionally one or more of >25%, >50%, or >75% of the references population’s mean acceleration magnitude and mean angular velocity magnitude, or any other suitable percentage or threshold. Further, a percentage of movement data signal peaks of the reference population that are above a reference threshold are determined. This reference threshold may be one or more of >50% or >75% over the reference population’s mean acceleration magnitude and mean angular velocity magnitude, or any other suitable percentage or threshold. The population’s reference values are obtained analogously to the user’s acceleration magnitude and angular velocity magnitude. These values form part of the statistical movement data in some specific implementations. However, the mean need not be used. Any suitable statistical movement data may be calculated. In some embodiments, movement feature determination module 120 uses the first machine learning model and / or further machine learning models to extract further movement features 130 from the movement data 110 and / or from already determined movement features 130 (such as the sequence of activity classes and the duration of time the user spends in each of the activity classes). Each further machine learning model may be trained to identify different features in the movement data.
[0090] In some examples, further machine learning models to automatically extract features (e.g. movement features 130) from multivariate time series, such as human movement data 110, utilise deep learning techniques such as Convolutional, Recurrent, and Attention layers.
[0091] In one example, a further neural network model for extraction of further movement features 130 employs Convolutional layers, which specialised in extracting repeating features from data with lattice-like structures to hierarchically extract spatial patterns by analysing the relationships between different measurement axes and identifying localised repeating structures. The model further uses Recurrent layers designed to process sequential data through the memory-like structure to model short- to mediumterm temporal dynamics from human movement data 110 through the layer's natural sequence comprehension and long-term temporal dependencies through a hierarchical convolutional condensing.
[0092] When using a multisensor modality for recording movement data 110 (e.g., as shown in Figure 2, though any sensor combination may be used) additional complexity is introduced into the analysis. To address this, in some examples, the neural network model employs a sensor attention mechanism. This mechanism dynamically weighs the features generated by each sensor, considering their relative importance. Thus, it is ensured that crucial information from each sensor is appropriately considered and propagated to subsequent layers in the deep learning model. This attention mechanism enhances the feature extraction process, providing a more comprehensive representation of the multisensor time series data. By hierarchically combining Convolutional, Recurrent, and Attention layers, a robust deep learning neural network architecture is created that systematically and automatically extracts essential features from the complex multisensor time series data. Movement features 130 may be extracted from any suitable layer of the further neural network model. In some examples, the further machine learning models are trained on the movement data of a reference population such as the reference populations described in the context of the first machine learning model or the second machine learning model. The reference population maybe the same reference population as used for training the first or second machine learning model or a different reference population. Each further machine learning model may be trained on the same reference population or on different reference populations. For example, different further machine learning models maybe trained on populations having different chronic diseases or being in different risk groups. Any reference population with characteristics of interest maybe used. The movement data of the reference population may be used as inputs to the further machine learning models either in raw form or processed into movement features 130 (e.g., a sequence of activity classes for the user in the reference population and a duration of time the user spends in each of the activity classes). What features each of the further machine learning model is trained to identify is determined by the ground truth data provided to the respective further machine learning models. In some examples, one or more of the further machine learning models are trained to take multivariate time series other than raw human movement data as inputs. In some examples, one or more of the further machine learning models takes heart rate sensor data, step count, distance walked, or fall detection information received from one or more of devices 170 as inputs. In some examples, step count, distance walked, or fall detection information is comprised in movement features 130 without further processing. In some examples, heart rate sensor data is comprised in clinical marker data 180 (discussed below).
[0093] Once movement feature determination module 120 has determined all desired movement features 130, the movement features are sent to clinical assessment module 150. The movement features 130 include user’s sequence of activity classes and duration spent in each activity class determined by the first machine learning model by classifying the user’s activities during natural behaviour. The number of activity classes used maybe any suitable number of activity classes, such as 4 activity classes (e.g. in bed, sitting on chair, standing, walking) or 3 activity classes (e.g. in bed, sitting on chair, mobilizing). The number of activity classes may be reduced by merging activity classes into a larger activity class, or may be increased by separating out activities into different classes. In some examples, the movement features 130 optionally include statistical movement data obtained from movement data 110, for example mean, median, standard deviation and relevant thresholds. In further examples, the movement features 130 optionally include movement pattern data extracted from the sequence of activity classes and the duration of time the user spends in each of the activity classes. In some examples, movement features 130 optionally include data extracted from the movement data and / or from the sequence of activity classes and the duration of time the user spends in each of the activity classes using a further machine learning model.
[0094] In some embodiments, clinical assessment module 150 additionally receives clinical marker data 180 as an input to the second machine learned model. Clinical marker data 180 can be any suitable information. The clinical marker data 180 is in these examples obtained from health care professionals, patient records, self-reporting, or devices 170, but the data 180 can be obtained or recorded in any suitable manner.
[0095] For example, if the user is a hospital patient, clinical markers may include one or more of patient delay to hospital admission and treatment, demographics, past medical history, comorbidities, severity of condition, lab results, neuroimaging features, vitals, movement change, functional independence, blood markers, mobility aids, mood, handedness and bed-rest. An example set of clinical markers for a user on a stroke ward is provided in Table 1. Non-applicable clinical markers may not be provided to clinical assessment module 150. For example, if the user is a care home resident the clinical markers “patient delay to hospital admission and treatment” or “neuroimaging features” may not be provided. Further, in some examples, the second machine learning model can be trained on alternative clinical markers.
[0096] Features Details
[0097] Demographics Age, sex, weight, and height.
[0098] Past medical Migraine, dyslipidaemia, hypertension, heart failure, history and diabetes, atrial fibrillation, smoker, psychiatric illness, comorbidities COPD, carotid stenosis, vascular disease, previous haemorrhage, obesity, and prior stroke.
[0099] Type and location The presence of hemorrhagic stroke, posterior circulation of the stroke involvement, and side of stroke (right or left) at admission.
[0100] Functional Modified Rankin Score (mRS) at preadmission, admission, independence and sensor recording. Barthel index-activities of daily living
[0101] (BI-ADLS) at the sensor recording. Transfer and mobility evaluations (independent, assistance of 1 person, assistance of 2 persons / hoisting requirement for mobilizing and transfers) at sensor recording and preadmission.
[0102] Stroke severity National institute for Stroke Health Scale (NIHSS) total and subscores at admission and sensor recording
[0103] Vital signs Temperature, diastolic blood pressure, systolic blood pressure, saturated oxygen level, heart rate, respiratory rate, and consciousness level at hospital admission.
[0104] Blood markers Glucose, C-reactive protein, leukocyte count, bilirubin, and creatinine at hospital admission.
[0105] Mobility aids Number and type of aids including a walking stick, Zimmer frame, wheelchair, and no support
[0106] Mood Hospital Anxiety and Depression Scale (HADS) at the sensor recording.
[0107] Handedness Left, right, or ambidextrous using the Edinburgh
[0108] Handedness Scale at sensor recording.
[0109] Bed-rest Whether the patient is on prescribed bed rest at the sensor recording.
[0110] Table 1 Example clinical markers for a user that is a stroke ward patient.
[0111] Once clinical assessment module 150 has received the user’s movement features 130 and optionally clinical marker data 180, the clinical assessment module determines a metric indicative of a clinical assessment 160. This metric may be any suitable indication of a user’s current health status. For example, the metric may be a clinical diagnosis of a condition, a predicted risk of a condition, a prediction of disease progression. The metric maybe updated regularly such that a user’s health status may be monitored. The metric may contain multiple indications of for example, different predicted risks (e.g. deterioration risk, risk of falling) or disease diagnosis.
[0112] In specific example, the metric indicative of a clinical assessment is the user’s clinical deterioration risk. In a further example, the deterioration risk is a user’s risk of deteriorating withing 96 hours of device 170 recording the movement data 110. Additionally or alternatively, the deterioration risk is a predicted National Early Warning Score 2 (NEWS2), which is a clinical scoring system used to predict clinical deterioration and guide appropriate escalation of care for acutely ill patients. In further examples, clinical assessment module 150 may predict additional or alternative measures of deterioration or provide measures related to other conditions.
[0113] Clinical deterioration for purposes of this disclosure is defined as any adverse change in a user’s health likely to increase morbidity, mortality, or length of hospital stay and which requires active medical management. Examples of such adverse changes in health are new cases of illness or infection such as sepsis, delirium, renal failure, stroke, urinary tract infection, thrombo-embolism, myocardial infection, or early neurological deterioration. Such a broad definition of clinical deterioration generalizes the applicability of the disclosed system to a broad range of risk medical conditions and illnesses, including monitoring of the deterioration of the elderly in a domestic setting.
[0114] Clinical assessment module 150 may comprise any appropriate second neural network trained to predict a metric indicative of a clinical assessment of the user based the movement features 130 which include the user’s sequence of activity classes, the duration of time the user spends in each of activity classes (and optionally clinical marker data including features describing the user, such as medical history or mobility aids). Optionally, movement features 130 comprise additional movement features such as statistical movement data, movement patters or other movement information.
[0115] In addition to neural networks, other machine learning methods may be trained to determine a metric indicative of a clinical assessment of the user based on movement features 130 and optionally clinical marker data 180. Suitable machine learning algorithms include alone or in combination: tree-based models, bootstrap decision forest, linear regression, k-nearest neighbour, support vector machine, random forest,
[0116] XGboost and ensemble machine learning. Such algorithms are known to the skilled person.
[0117] In one specific embodiment, clinical assessment module 150 comprises a global local attention model, such as model 700 shown in Figure 7. This model is based on Multi¬
[0118] Head Attention first developed by Vaswani et al., 2017 (Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems, 30.). Multi-Head Attention computes attention vectors across multiple dimensions, allowing the layers to encode multiple relationships in parallel into the final attention vectors. This feature enables Multi-Head Attention layers to efficiently attend to all the data points, thus enabling them to learn the temporal relationships of the sequence. Multi-Head Attention and its underlying Scaled Dot-Product Attention can be implemented as follows: Where
[0119] Where Q, K, and V are the query, key, and value vectors. WiQ , WiK, Wf' are the weight matrices for the query, key, and value vectors, respectively. Concat refers to a tensor concatenation operation. W° is the weight matrix for the concatenated heads, dk is the dimension of the key vector, n is the number of heads.
[0120] Pure Multi-Head Attention-based models, such as Transformer (Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017).
[0121] Attention is all you need. Advances in neural information processing systems, 30.), are computationally and memory intensive due to performing multiple large-matrix multiplications in succession, as seen in the equations above. This limitation hard- limits the input sequence length. The global local attention mechanism 700, unlike pure Multi-Head attention, splits attention into local attention 610 and global attention 620, in which the local attention attends localised data windows while the global attention captures the overall relationships, as depicted in Figure 6. This allows global local attention model to work with much longer data sequences than Multi-Head attention models.
[0122] Figure 7 depicts the specific architecture of example global local attention model 700, which can be used as the second machine learned model. Similarly to a convolution layer, the local attention mechanism attends a locally fixed-sized sliding window of the data using a standard Multi-Head Attention mechanism, resulting in multiple unique local context vectors. While the global attention mechanism is similar to a fully- connected layer, in which a standard Multi-Head Attention mechanism attends multiple local contexts through abstracts of local context vectors. To enable the exchange of information from one localized context to another, a concept known as “global attention” is employed by embedding a specialised null token in each windowed input. A null token is a unique token that only conveys abstract attention information from local to global contexts. Once the local attention mechanism has attended to each input window, the null token that contains contextual information is extracted and incorporated it into a sequence along with a global token. A standard Multi-Head Attention mechanism is then applied to the global token sequence, resulting in a global attention context. Since each global token contains the local context information of each corresponding windowed local input, the global attention mechanism allows the entire input sequence to be virtually attended to with minimal memory footprint due to the reduction of multiplication matrix size. After multiple layers 710 of global local attention the data is processed by a multilayer perceptron (MLP) classifier 720 and the metric of clinical assessment output. In other words, first the global local attention model applies a local attention mechanism using Multi-Head Attention to each of a plurality of data windows, wherein a null token is embedded in each window. Next, the null token for each window is extracted and all null tokens incorporated into a sequence with a global token. Finally, a global attention mechanism using multi-head attention is applied to the new sequence. In one specific example, the global local attention model 700 or any other suitable second neural network, such as a neural network utilizing a different attention mechanism, is trained to predict clinical deterioration of a user. In some examples, such a machine learning model is trained on data obtained from a stroke ward population. The stroke ward population includes patients with both strokes and non- strokes, as discussed above. Non-strokes may include vestibular neuronitis, migraine, COVID-19 pneumonia, delirium, included cognitive disorder, urinary tract infection, seizure, benign paroxysmal positional vertigo (BPPV), Multiple Sclerosis relapse, and unknown aetiology. The diversity of diagnosis and diversity of stroke ward patients in terms of medical history, age and other markers ensures neural networks trained on this dataset are generalizable to a large range of users, such as users in other acute medical wards, nursing homes, or community care. The stroke ward population is here the same type of population as is used in the example for training the first machine learned model, but different training data sets may be used with different reference populations. The training data was collected directly from stroke ward recovery patients using four Apple Watch smartwatches implementing IMU sensors arranged as shown in Figure 2, i.e., left wrist, right wrist, left ankle, and right ankle. Natural behaviour movement data was collected during daytime from 117 patients within 48 hours of admission to the acute stroke ward. To account for left-handedness, Acceleration data was switched between the right and left limbs, and the inverse of the Acceleration y-axis taken. The movement data is sampled at too Hz and approximately 5 Hr long.
[0123] Ground truth data for deterioration prediction was generated by 2 hospital doctors independently, by retrospectively reading the electronic health records for the users included in the training data and assessing whether a deterioration event occurred within 96 hours of sensor recording and the type. A third doctor reviewed and adjudicated discrepancies between assessors. Ground truth data for NEWS2 score prediction was generated by an attending medical professional determining or collecting the NEWS2 scores within 48 hours of admission.
[0124] While a specific training dataset is described, any suitable training dataset collected according to methods of this disclosure may be used to train a second machine learning algorithm suitable for use in clinical assessment module 150, such as the example global local attention model 700. For example, the training dataset described in the context of training movement feature determination module 120, when supplemented with deterioration risk prediction ground truth data may be used. In another example, a training dataset collected from any population at risk of deterioration may be collected. Some examples of populations at risk of deterioration maybe an intensive care ward population, a stroke ward population, a care home population, a chronic or acute disease population, an elderly population or any other population with characteristics putting them at higher risk of experiencing clinical deterioration.
[0125] In further examples, training data may be collected form healthy populations, populations with specific diseases, populations with other risk factors. In yet more examples, training populations may be selected to best enable the system to learn to determine a specific metric indicative of a clinical assessment of the user 160. In such cases, the ground truth is the outcome of a clinical assessment of each user in the reference population corresponding to the metric of interest. Such a metric maybe any suitable metric. For example, in addition to deterioration risk a metric of interest might be a diagnosis, a risk prediction, a disease progression prediction, or other clinical outcomes.
[0126] For training, a training dataset is processed in accordance with the disclosed methods to generate, for each user in the training data, movement features 130 and optionally clinical marker data 180. Subsequently, the second machine learning model is trained using the desired movement features 130 and (optionally) clinical marker data 180 as inputs and the clinical outcome for each user in the training dataset as ground truth. A specific example of training the second machine learning model implemented in clinical assessment module 150 is described in the following. The above-described training dataset is processed in accordance with the disclosed methods to generate, for each user in the training data, movement features 130 (in this specific example, a sequence of activity classes, a duration of time the user spends in each of the activity classes and statistical movement data of the type described in the example above) and optionally clinical marker data 180 (in this specific example, clinical marker data of Table 1). The ground truth for each user in the training data is whether the user experienced deterioration in the 96 hours after recording movement data. In the following experiments are discussed using different subsets of the movement features and clinical marker data as the inputs for clinical assessment module 150. Unless stated otherwise all movement features and clinical markers are included. For processing in the global local attention model, the recorded movement data is down sampled from too Hz to 40 Hz. In addition, a fixed-size sliding window with a half-window size step is used to segment the data into segments, each 4800 samples long (2 minutes at 40 Hz).
[0127] For evaluation, the global local attention model is implemented in PyTorch 1.0 using NVIDIA Quadro P6000 with 24GB VRAM and 3840 CUDA cores. All models are trained for 50 epochs using an ADAM optimiser with a 0.000005 learning rate and 0.001 weight decay. The model is evaluated using a handpicked hold-out dataset. Since the dataset was imbalanced in terms of both deterioration outcome and NEWS2, both subsampling and oversampling techniques are employed to achieve equal representations of each label in the training dataset. The effectiveness of the global local attention model is assessed by utilizing the weighted F-i metric due to the imbalanced distribution of labels in the hold-out dataset. As shown in Figure 8, given two minutes of movement data, the global local attention model (example of the second machine learning model) can predict patient deterioration with a high F-i score of 0.825. However, within the same setting, the model prediction of the patients’ NEWS2 is only moderately good, with an average F-i score of 0.531.
[0128] In another specific implementation that is an alternative to the global local attention mechanism described above, the second machine learning model is an artificial neural network with one hidden layer containing eight hidden nodes using a ReLU (rectified linear unit) activation function. The output layer uses a sigmoid activation function.
[0129] The artificial neural network is trained using the training dataset described in the context of the global local attention model at a learning rate of 0.005 for 500 epochs.
[0130] It is determined, using the second machine learning model, a metric indicative of a clinical assessment of the user, wherein the second machine learning model takes as inputs at least the one or more movement features.
[0131] Table 2 shows an overview of all deteriorated users in the training data (by subject number). For patients indicated in light grey or dark grey, the trained above-described artificial neural network predicted user deterioration in advance of NEWS score escalation or drop in consciousness, demonstrating that in one implementation the trained system too can be used as a deterioration early warning system.
[0132] Subject Details of Sensor Clinically NEWS Drop complication recording determined escalation consciousness event SG9029 Gram negative 03 / 04 / 2019 02 / 04 / 20i9a03 / 04 / 20i9abacteremia, delirium SG9040 Stro e em pares s 11 / 04 / 2019 09 / 04 / 20i9 worsening (lacunar / capsular warning syndrome) 001-052 Hypox a, new 01 / 02 / 2021 31 / 01 / 2021 28 / 01 / 2021 atelectasis
[0133] 001-058 New stroke 05 / 02 / 2021 04 / 02 / 2021a, 03 / 02 / 2021arequiring 09 / 02 / 2021a
[0134] Table 2: Patient deteriorations. * = clinically determined deterioration event recorded in medical notes within 96 hours of the sensor recording determined by 2 doctors and adjudicated by Consultant Neurologist,a= amber or red escalation of National Early
[0135] Warning Score (NEWS) and drops in consciousness ± 4 days from the sensor recording with recorded dates. NIHSS =National Institute for Health Stroke Scale, TIA transient ischemic attack, Na =sodium, CRP =C-reactive protein. Light grey complication detected by the above-described artificial neural network and not by the standard NEWS2 or drop in consciousness, dark grey = complication detected by the above- described artificial neural network first and subsequently by NEWS2 or drop in consciousness.
[0136] Predictor selection models such as univariate filter statistics and recursive feature elimination (RFE) can be used to pre-select markers for inputting into clinical assessment module 150 out of the clinical marker data 180 and movement features 130. Limiting the inputs to the second machine learning model for deterioration prediction to only particularly relevant markers can improve model performance compared to inputting all markers without feature selection. In other words, selecting appropriate clinical markers and / or movement features can improve performance. Different markers and features may be selected for different applications or uses, or for different types of patient condition.
[0137] On the training dataset described above, RFE ranks of the clinical marker data several National Institute for Stroke Health Scale (NIHSS) sub-scores on admission, functional disability measures, previous medical history, type of stroke, NIHSS severity at the sensor recording, and vital signs as important for predicting a higher risk of deterioration. Further, of the movement features including the statistical movement data and activity classes, RFE ranks several subject and population-specific thresholds, standard deviation of movement data, and activity data (e.g., duration in bed and duration mobilising) as important for predicting patients at risk of deterioration.
[0138] For the above described dataset, both RFE and univariate filters select Modified Rankin Score (mRS) at admission, NIHSS visual field and motor left arm at admission, and whether the patient was on bed rest as important clinical markers of deterioration. Both selection methods find duration of the activity classes “in bed” and “mobilising”, left-hand subject threshold >25%, left leg population threshold >25% and 75%, and right leg standard deviation movement features as predictive smartwatch movement features. In some examples, machine learning models for clinical assessment according to this disclosure may be trained on datasets more specific to a single illness or medical condition. In such examples, other clinical markers and movement features may be more relevant than the markers and features found most relevant for the present example. In other words, instead of general deterioration events, models could be trained for a specific subset of the deterioration outcomes e.g., pneumonia or sepsis. This can result in a different set of features which are optimal for predicting these outcomes.
[0139] Figure 9 shows a further experiment analysing the above-described artificial neural network’s performance for stratifying in-patient risk of deterioration. Receiver Operating Curves (ROC) and area under the curve (AUC) for models using different feature subsets are shown: Blue curve 910 uses smartwatch movement features and clinical markers, green curve 930 uses univariate and recursive feature elimination (RFE) filtered clinical markers, orange curve 920 uses standard clinical markers from earlier studies [Fernandez-Lozano, C., et al., Random forest-based prediction of stroke outcome. Sci Rep, 2021. 11(1): p. 10071; Sung, S.M., et al., Prediction of early neurological deterioration in acute minor ischemic stroke by machine learning algorithms. Clin Neurol Neurosurg, 2020. 195: p. 105892; Kwan, J. and P. Hand, Early neurological deterioration in acute stroke: clinical characteristics and impact on outcome. QJM: An International Journal of Medicine, 2006. 99(9): p. 625-633], red curve 940 uses smartwatch movement features, purple curve 950 uses the smartwatch step counter, brown curve 960 represents prediction by chance. The shaded area is the standard error across the 10 runs. The above-described artificial neural network combining smartwatch movement features and clinical markers (Fl 72 ± 2%) improved model performance compared to standard clinical markers used in earlier studies (Fl 45 ± 4%) (the same studies are used as above), optimised ML clinical markers with RFE and filter selection (Fl 62 ± 3%), smartwatch movement features alone (Fl 55 ± 3%), and simple step counters (Fl 35 ± 5%).
[0140] A large range of recording times may be used to achieve accurate deterioration prediction. The above-described artificial neural network achieves near-identical accuracy with smartwatch movement features derived from > 3-hour compared to 1- hour sensor recordings (Fl 72 ± 3%, 86 ± 1% of AUC versus Fl 72 ± 2%, 86 ± 1%). Wilcoxon Sign-Ranked tests showed that most smartwatch movement features were not significantly different between 3-hour compared to 1-hour sensor recordings. In some examples, one or more devices 170 are used to determine a metric indicative of a clinical assessment 160. In experiments where the user’s deterioration risk is predicted, the above-described artificial neural network’s performance when dropping smartwatch movement features from the lower limbs is explored. Keeping smartwatch markers from the upper limbs but removing smartwatch movement features from both lower limbs resulted in similar performance (Fl 72 ± 3%, AUC 86 ± 1%) to that derived from all four limbs. Keeping activity model markers (i.e. activity classes and durations) but removing all smartwatch movement features reduced accuracy (Fl 66 ± 2%, 84 ± 1% of AUC). Keeping smartwatch activity markers but using only the smartwatch movement features from either the left hand or right hand resulted in similar performance (Fl 71 ± 4%, 85 ±1% of AUC and Fl 69 ± 2%, 85 ± 1%, respectively).
[0141] Once the metric indicative of a clinical assessment 160 has been determined, the assessment is made accessible to the user and / or other persons authorized to view the user’s deterioration risk (authorized persons). For example, the metric indicative of a clinical assessment 160 maybe made accessible to a health care professional monitoring or treating the user. In another example, the metric indicative of a clinical assessment 160 may be made accessible to the user’s carer or family. In yet another example, the metric indicative of a clinical assessment 160 may be made accessible to a person designated by the user.
[0142] The metric indicative of a clinical assessment 160 (or data indicative of it) may be made accessible in several ways. In one example, the metric indicative of a clinical assessment 160 is displayed in a web-interface, where only persons authorized to view the user’s information can see the user’s the metric indicative of a clinical assessment 160. In another example, the metric indicative of a clinical assessment 160 is made available via a receiving unit 190.
[0143] Receiving unit 190 maybe a computer, tablet, mobile phone, smart watch or any other device suitable for displaying information visually or audibly. Receiving unit may be remote from the processors, or integrated with the processors in a single device. Receiving unit 190 may be configured to display the clinical assessment as part of a health dashboard used by care providers to monitor a user’s health. In some examples, receiving unit 190 is one of devices 170. For example, information regarding the metric indicative of a clinical assessment 160 can be displayed on the device 170 on a user’s limb. In some other examples, receiving unit 190 displays a web-interface containing the metric indicative of a clinical assessment 160. The means by which the metric indicative of a clinical assessment 160 is made accessible, such as by receiving unit 190 or a web-interface, can provide access to the metric indicative of a clinical assessment for a plurality of different users, each equipped with a system for clinical assessment. For example, only one receiving unit may be required to monitor all users on an intensive care ward. In further examples, multiple receiving units or web-interfaces maybe provided with the user’s metric indicative of a clinical assessment 160. For example, a user may be provided with their metric indicative of a clinical assessment displayed on one of devices 170 or their smart phone, and a health care professional may be provided with the metric in a receiving unit collecting data from multiple patients in a ward.
[0144] In some examples, the receiving unit 190, or any other means for making the metric indicative of a clinical assessment 160 accessible, receives additional information for display to the user or any other authorized person. Different information may be accessible to different persons depending on which information the person is authorized to view. Some examples of additional information that may be made available is movement data 110, movement features 130 (e.g., the sequence of activity classes, duration spent therein, movement patters, statistical movement data, etc.), and / or clinical marker data 180.
[0145] Additionally or alternatively, information from intermediate processing steps may be made available, or partial aspects of the data maybe made available. For example, movement data 110 labelled with activity classes may be made available. In another example, where movement data 110 is recorded from multiple devices 170, movement data from a subset of devices 170 may be made available. In yet another example, statistical movement data may be made available for every recording day, week, month and so on. Such information may aid the user or other authorized persons to assess disease severity, rehabilitation progress, further need for treatment, a need for more exercise or similar considerations. Furthermore, providing access to such data increases the expandability of the clinical assessment module’s 150 outputs. For example, a health care professional may be able to see that a prediction of risk of deterioration coincides with changes in the movement features representing the user’s natural behaviour. And it may be known to the health professional that these changes in behaviour mean an increased risk of deterioration. In some embodiments, receiving unit 190 or a suitable web-interface, may be used to set-up a new user with a system for clinical assessment. For example, each of the one or more devices 170 recording movement data 110 may be allocated to a position in a room or on the user (e.g., a limb of the user), start and end times of recording maybe logged or medical history from which some or all of clinical marker data 180 can be obtained or selected may be recorded. In another example, further information may be provided, such as the user’s name, other identification information, other clinical data, next of kin data etc. Aspects of the present disclosure may be implemented in any suitable computer hardware or software. Figure 10 shows an example system 1000 for clinical assessment of a user according to aspects of the present disclosure. Example system 1000 comprises of device(s) 1010 for recording movement data (examples of device(s) 170), one or more processors 1020 or processing movement data and determining a metric indicative of a clinical assessment, and one or more receiving unit(s) 1030 for receiving the deterioration risk and displaying the deterioration risk to authorized persons.
[0146] While the system 1000 is drawn as including both device(s) 1010 and receiving unit(s) 1030, these components maybe external to the system. In other words, system 1000 may comprise only the processor(s) 1020, which are configured to receive movement data from the device(s) 1010 and output the user’s metric indicative of a clinical assessment (optionally to the receiving unit 1030). In such implementations, the processor(s) 1020 can be remote from the device(s) 1010 and the receiving unit 1030. The processor(s) may be implemented as a server (or servers 1020). Any suitable remote processing means can be used, such as cloud processing. Device(s) 1010 is any device suitable for detecting motion of the user. Such a device 1010 maybe a smart phone, smartwatch or any other device implementing a sensor for determining user motion such as camera, inertial movement unit, depth sensors, ultrasound, radar or any other suitable sensor.
[0147] Processor(s) 1020 are configured to implement aspects of the present disclosure. Any suitable data processing architecture maybe used to implement processors 1010. In one example in which the processors are remote from the device(s) 1010, the processors are implemented at a server (understood to be just one, non-limiting, example of remote processing). Server 1020 comprises one or more processors configured implement aspects of the present disclosure. In some examples, server 1020 implements movement features determination module 120, clinical assessment module 150 and any other processing required for implementation of aspects of the disclosure. In other examples, some or all of the processing may be executed by device(s) 1010 or receiving unit 1030. Sever 1020 maybe implemented as a single server in one location or as multiple servers in communication with each other either at a single location or at multiple locations. Server 1020 may store data, such as reference population or user data, locally or obtain data from external locations. Any suitable server architecture maybe used to implement server 1020. Receiving unit 1030 may be any suitable computing device for receiving or obtaining the user’s metric indicative of a clinical assessment and / or data indicative of the user’s clinical assessment (and optionally any other data about the user), and conveying such information to one or more healthcare workers and / or any other authorized persons. Such conveying may include displaying information on any kind of visual display, audibly rendering the information to the authorized persons (e.g., a beeping sound or spoken phrase), a flashing light etc. In other words, receiving unit may output information and / or an alert based on the user’s clinical assessment. Receiving unit 1030 may be a smart phone, tablet, desktop computer, wearable device, or any other device suitable for conveying information to authorized persons. Receiving unit 1030 may receive the information to be displayed directly from server 1020 or may access a web-interface which is providing the information in order to obtain and display the information.
Claims
Claims1. A system for clinical assessment of a user, the system comprising one or more processors configured to: obtain movement data of a user, the movement data recorded during the user’s behaviour; determine, based on processing the recorded movement data using a first machine learning model, one or more movement features, wherein the one or more movement features comprise a sequence of activity classes and a duration of time the user spends in each of the activity classes, wherein each activity class represents one or more activities performed by the user during the user’s behaviour; determine, using a second machine learning model, a metric indicative of a clinical assessment of the user, wherein the second machine learning model takes as inputs at least the one or more movement features.
2. The system according to claim 1, the system further comprising at least one device configured to record the movement data, wherein the movement data comprises one or more of: position, velocity, or acceleration data in the linear and / or rotational domain, optionally wherein each at least one device comprises any one of: an inertial measurement unit, a radar sensor, a camera, a depth sensor, or an ultrasound sensor.
3. The system according to claim 2, wherein the at least one device is placed on a limb of the user and comprises an inertial measurement unit or other movement measurement unit.
4. The system according to claim 3, wherein the at least one device is a smartwatch comprising the inertial measurement unit or other movement measurement unit, optionally wherein the at least one device is two smart watches, one placed on each of the user’s arms, optionally, wherein the at least one device is four smart watches, one placed on each of the user’s limbs.
5. The system according to any preceding claim 2-4, wherein obtaining, from the at least one device, the movement data recorded during the user’s behaviour comprises:obtaining, from an inertial measurement unit, IMU, of each at least one device, triaxial acceleration data and triaxial angular velocity data.
6. The system according to any preceding claim, wherein the activity classes comprise one or more of: being in a bed, sitting on a chair, and mobilizing.
7. The system according to any preceding claim, wherein the first machine learning model comprises a first neural network, the first neural network comprising a hybrid convolutional neural network and long short-term memory with recurrent attention layers.
8. The system according to any preceding claim, wherein the first machine learning model is trained by supervised learning using labelled, raw movement data from a first reference population, wherein the raw movement data from the first reference population is obtained from each particular user in the first reference population by obtaining, for each particular user in the first reference population, movement data recorded during the particular user in the first reference population’s behaviour, and wherein each label corresponds to one of the activity classes.
9. The system according to any preceding claim, wherein the second machine learning model comprises a second neural network, the second neural network optionally comprising a global local attention model.
10. The system according to any preceding claim, wherein the second machine learning model is trained by supervised learning on a second reference population, wherein a training data set comprises: for each particular user in the second reference population, one or more movement features of the user in the second reference population, wherein the one or more movement features of the particular user in the second reference population comprise a sequence of activity classes and a duration of time the particular user in the second reference population spends in each of the activity classes, and for each particular user in the second reference population, a metric indicative of the outcome of a clinical assessment of the particular user in the second reference population.
11. The system of claim 10, wherein the training data set further comprises: for each particular user in the second reference population, one or more clinical markers for the particular user in the second reference population, wherein the same clinical markers are used for all users in the second reference population.
12. The system of any one of claims 1 to 11, wherein the one or more processors are further configured to: obtain one or more clinical markers, wherein the one or more clinical markers comprise at least one feature of a plurality of features describing the user, wherein to determine, using the second machine learning model, the metric indicative of a clinical assessment of the user, the second machine learning model further takes the one or more clinical markers as further inputs.
13. The system according to any preceding claim, wherein the one or more movement features further comprise one or more of: statistical movement data obtained from the movement data, movement pattern data extracted from the sequence of activity classes and the duration of time the user spends in each of the activity classes, or data extracted from the movement data and / or from the sequence of activity classes and the duration of time the user spends in each of the activity classes using a further machine learning model.
14. The system according to claim 13, wherein the further machine learning model is neural network comprising convolutional, recurrent and attention layers.
15. The system of any preceding claim, wherein the metric indicative of a clinical assessment of the user is a clinical deterioration risk prediction, a prediction of disease progression, or a clinical diagnosis.
16. The system of claim 15, wherein clinical deterioration is defined as an adverse change in a user’s health likely to increase morbidity, mortality, or length of hospital stay and which requires active medical management, optionally, wherein an adverse change in the user’s health comprises any one of: new cases of illness or infection such as sepsis, delirium, renal failure, stroke, urinarytract infection, thrombo-embolism, myocardial infection, or early neurological deterioration.
17. The system of any one of claims 15 and 16, wherein the user’s clinical deterioration risk prediction comprises the user’s clinical deterioration risk within 69 hours after the movement data is recorded during the user’s behaviour.
18. The system of any preceding claim, wherein the one or more processors are further configured to: obtain one or more additional metrics, the one or more additional metrics including at least one of heart rate sensor data, step count, distance walked, or fall detection information, and wherein the second machine learning model is configured to take the one or more additional metrics as input.
19. The system of any preceding claim, further comprising a receiving unit configured to receive the metric indicative of a clinical assessment of the user from the one or more processors, the receiving unit further configured to at least display information based on the metric indicative of a clinical assessment of the user.
20. A method for clinical assessment of a user, the method comprising: obtaining movement data of a user, the movement data recorded during the user’s behaviour; determining, based on processing the recorded movement data using a first machine learning model, one or more movement features, wherein the one or more movement features comprise a sequence of activity classes and a duration of time the user spends in each of the activity classes, wherein each activity class represents one or more activities performed by the user during the user’s behaviour; determining, using a second machine learning model, a metric indicative of a clinical assessment of the user, wherein the second machine learning model takes as inputs at least the one or more movement features.