Method and system for analysing body movement

CA3324030A1Pending Publication Date: 2025-09-18KINEPHONICS IP PTY LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CA3324030
Authority / Receiving Office
CA · CA
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-15
Filing Date
2025-03-07
Publication Date
2025-09-18

AI Technical Summary

Technical Problem

Existing human motion recognition technologies face challenges in accurately tracking and classifying human movements due to issues such as partial or full obstruction of body parts, changes in lighting, movement within the frame, differences in perspective, and variations in human physique, which affect the analysis of movements for applications like sports, medicine, and fitness.

Method used

A computer-implemented method that utilizes landmark detection and tracking to analyze movements of body parts, classifying and generating a data summary of these movements, including determining angles, ratios, and cumulative sums to track movements over time, using machine learning models for facial and hand landmark detection.

Benefits of technology

Enables precise tracking and classification of human movements, generating actionable data summaries that can be used for behavioral analysis, diagnosis, and fitness assessment by accurately identifying and quantifying movements of eyes, hands, and body parts.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

A computer implemented method is described that includes receiving a video recording of a person, analysing movements of body parts of the person based on identified landmarks, classifying and tracking movements over time, generating a data summary of the classified movements. An action may be performed based on the data summary.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND SYSTEM FOR ANALYSING BODY MOVEMENTTECHNICAL FIELD

[0001] Aspects of the present disclosure are directed to human movement recognition. Particular embodiments relate to the utilisation of human movement recognition to detect behavioural characteristics of a person during an activity and to utilisation of the detected behavioural characteristics.BACKGROUND

[0002] Human motion recognition is a branch of computer vision that is directed to understanding and analysing the movement of people in video. In general human motion recognition may involve extracting and describing characteristics of human motion by a computer and realising the recognition of an action type.

[0003] There are many applications of human motion recognition in a wide variety of contexts. For example, human motion recognition has applications in sports and exercise science, in medicine, speech pathology, occupational therapy, neurology, behavioural biometrics, robotics, smart surveillance, art and entertainment. In particular, in the field of medicine human motion recognition may be used for the study and analysis of orthopaedics, neurology, musculoskeletal disorders, body posture and perhaps even fitness. In general, the application of human movement recognition helps determine the relevant body parts involved in a process and possibly the duration of the movement, number of repetitions of the movement or the frequency of movement, or a combination thereof. For example, humancomputer interaction may only involve hand gestures while sports or dance may involve the whole body.SUMMARY

[0004] A computer implemented method is described that includes receiving a video recording of a person, analysing movements of body parts of the person based on identified landmarks, classifying and tracking movements over time, generating a data summary of the classified movements. An action may be performed based on the data summary. Computer systems and computer storage configured to perform the method are also described.

[0005] A computer implemented method comprises: receiving a video recording of a person; analysing movements of body parts of the person based on identified landmarks;classifying and tracking movements over time; generating a data summary of the classified movements; and performing an action based on the data summary.

[0006] In some embodiments, the movements analysed are one or more of: eye movement, hand movement, mouth movement, head movement or body movement.

[0007] In some embodiments, the data summary generated includes a percent motor change over time.

[0008] In some embodiments, prior to generating the data summary, the method further includes the steps of: receiving at least one additional video recording of the user; analysing movements of the user in the at least one additional video recording using machine learning methods to detect and identify landmarks; and determining and tracking patterns in movement over time. The data summary generated may include a comparison of the patterns of movements of the user between video recordings. The data summary may include a percent motor change for each of the video recordings analysed.

[0009] In some embodiments, the identified landmarks comprise landmarks of a face, the landmarks of the face comprising a combination of two or more of right inner eyebrow, right outer eyebrow, top right eye, right pupil, outer right eye, bottom right eye, inner right eye, top upper-lip, bottom upper-lip, right mouth, mouth left, top lower-lip, bottom lower-lip, nose tip, right nose bridge, left nose bridge, left inner eyebrow, left outer eyebrow, top left eye, left pupil, outer left eye, bottom left eye and inner left eye. The analysing movements of body parts of the person based on identified landmarks may comprise determining head pose and wherein determining head pose is based on angles of lines between the identified landmarks. The identified landmarks may comprise the nose tip and the angles of lines between identified landmarks comprise an angle of a line between the nose tip and another identified landmark. The method may further comprise determining an angle of rotation of a head based on the identified landmarks.

[0010] In some embodiments, the identified landmarks comprise landmarks of a hand, the landmarks of the hand comprising a combination of two or more of wrist, lower thumb, first thumb knuckle, second thumb knuckle, tip of thumb, first index knuckle, second index knuckle, third index knuckle, tip of index finger, first knuckle middle finger, second knuckle middle finger, third knuckle middle finger, tip of middle finger, first knuckle of the third finger, second knuckle of the third finger, third knuckle of the third finger, tip of third finger,first knuckle of the fourth finger, second knuckle of the fourth finger, third knuckle of the fourth finger and tip of the fourth finger.[Oil] In some embodiments, the identified landmarks comprise landmarks of an upper body, the landmarks of the upper body comprising a combination of two or more of right ear, left ear, right shoulder, left shoulder, right elbow, left elbow, right wrist, left wrist, right hip and left hip.

[0012] In some embodiments, said classifying and tracking movements over time comprises determining a frequency of blinking of an eye. Determining the frequency of blinking may comprise determining a ratio of width and height of the eye and determining whether the eye is an open or closed state based on the determined ratio. Determining the frequency of blinking may comprise determining a number of times the ratio of width and height traverses a threshold value over time.

[0013] In some embodiments, said classifying and tracking movements over time comprises determining when a rate of change of body movement of a said body part exceeds a threshold value.

[0014] In some embodiments, said classifying and tracking movements over time comprises determining when a change in position of a said body part exceeds a threshold value.

[0015] In some embodiments, the identified landmarks comprise landmarks of a mouth and said classifying the tracking movements over time comprises determining opening, closing or opening and closing of the mouth. Determining opening, closing or opening and closing of the mouth may comprise determining when a ratio of mouth height and mouth width traverses a threshold value.

[0016] In some embodiments, the identified landmarks comprise landmarks of a mouth, the landmarks of the mouth comprising a combination of two or more of the top upper-lip, the bottom lower-lip, right mouth edge and left mouth edge. Determining opening, closing or opening and closing of the mouth may comprise determining when a ratio of mouth height and mouth width traverses a threshold value.

[0017] In some embodiments, the identified landmarks comprise at least one landmark of a wrist and wherein said classifying and tracking movements over time comprises determining wrist movement. Said determining wrist movement may comprise determining a cumulative sum of wrist movement over a period of time. Said determining wrist movementmay comprise determining a difference function between the cumulative sum of wrist movement and a polynomial best fit function of the cumulative sum and determining local maxima or peaks in the difference function. Determining local maxima or peaks in the difference function may comprise determining local maxima in a smoothed form of the difference function.

[0018] In some embodiments, said classifying and tracking movements over time comprises determining movement of a midpoint between the left should and the right shoulder. Said classifying and tracking movements over time may comprise determining a cumulative movement of the midpoint between the left shoulder and the right shoulder. Said determining movement of the midpoint may comprise determining a difference function between the cumulative movement of the midpoint and a polynomial best fit function of the cumulative sum and determining local maxima or peaks in the difference function. Determining local maxima or peaks in the difference function may comprise determining local maxima in a smoothed form of the difference function.

[0019] In some embodiments, said classifying and tracking movements over time comprises determining a cumulative sum of movement based on the identified landmarks, determining a best fit to the cumulative sum and performing local maxima or peak analysis of a difference between the cumulative sum and the best fit.

[0020] In some embodiments, said classifying and tracking movements over time comprises determining a cumulative sum of movement based on the identified landmarks, determining a best fit to the cumulative sum and performing local maxima or peak analysis of a difference between a smoothed form of the cumulative sum and the best fit.

[0021] In some embodiments, said classifying and tracking movements over time comprises determining a speed of movement. Said classifying and tracking movements over time may comprise determining a count of peaks or local maxima in the determined speed of movement. The speed of movement may be a speed of movement of at least one hand, a speed of change in a distance between the left and right shoulders, or a speed of movement of a head.

[0022] In some embodiments, generating a data summary of the classified movements comprises determining a measure of volatility of the classified movements.

[0023] In some embodiments, generating a data summary of the classified movements comprises determining a measure of variation of the classified movements.

[0024] Further aspects of the present disclosure and further details of the aspects described above will be apparent from the following description, including with reference to the accompanying drawings.BRIEF DESCRIPTION OF DRAWINGS

[0025] Fig. 1 is a flowchart showing the machine learning method according to aspects of the present disclosure.

[0026] Fig. 2A shows an example face landmark map.

[0027] Fig. 2B shows an example head with the roll, yaw and pitch.

[0028] Fig. 3 shows an example hand landmark map.

[0029] Fig. 4 shows an example body landmark map for a portion of the upper body.

[0030] Fig. 5A is plot of the left eye ratio over time.

[0031] Fig. 5B is a plot of the right eye ratio over time.

[0032] Fig. 6 is a plot of the mouth ratio over time.

[0033] Fig. 7 A is a plot of the movement of the left hand of the patient over a time period.

[0034] Fig. 7B is a plot of the movement of the right hand of the patient over a time period.

[0035] Fig. 8 A is a plot of the cumulative sum and polynomial fit for the left hand movement data from Fig. 7 A

[0036] Fig. 8B is a plot of the smoother difference function.

[0037] Fig. 9 A is a plot of the calculated change in distance of the shoulder midpoint between successive frames in the video recording over a period of time.

[0038] Fig. 9B is a plot of the cumulative sum and polynomial fit for the left hand movement data from Fig. 9 A.

[0039] Fig. 9C is a plot of the smoothed difference function with the detected peaks shown as circles.

[0040] Fig. 10A is a plot of the calculated change in angle between unit vectors of successive frames in the video recording over a period of time.

[0041] Fig. 10B is a plot of the cumulative sum and polynomial fit for the head movement data from Fig. 10A.

[0042] Fig. 10C is a plot of the smoothed difference function with the detected peaks shown as circles.

[0043] Fig. 11A is a plot of the movement of the patient’s left hand over a time period

[0044] Fig. 1 IB is a plot of the movement of the patient’s right hand over a time period.

[0045] Fig. 12A is a plot of the movement of the patient’s body over a time period.

[0046] Fig. 12B is a plot of the movement of the patient’s head over a time period.

[0047] Fig. 13 A is an example of a data summary showing the total number of counts of eye movement, mouth movement; body movement; head movement and hand movement is shown.

[0048] Fig. 13B shows another example data summary showing the movement count as a function of the duration of the time period.

[0049] Fig. 14A is an example of the data summary of the movement count per minute of a user’s eyes, mouth, head, hands and body over multiple recorded videos.

[0050] Fig. 14B shows another example data summary of the movement count as an average per session.

[0051] Fig. 14C shows a focused version of the plot of Fig. 14B showing only the count per minute for the movement of the eyes and hands of the user plotted over a series of sessions.

[0052] Fig. 15A shows yet another example of data showing the movement counts of the eyes, mouth, hands, head and body as a percentage of the total counts for the time duration.

[0053] Fig. 15B shows yet another example of data showing the movement counts of the eyes, mouth, hands, head and body as a percentage of the total counts for the session.

[0054] Fig. 15C shows further example data.

[0055] Fig. 15D shows the percent motor change over different sessions.

[0056] Fig. 15E shows a table with an example calculation for the percent motor changes for four sessions.

[0057] Fig. 16A is a pie chart showing the movement counts for a first session as a percentage of the total movement counts for the first session.

[0058] Fig. 16B is another pie chart showing the movement counts for a second session as a percentage of the total movement counts for the second session.

[0059] Fig. 17A shows a percent motor change between two sessions shown in Figs. 16A and 14B.

[0060] Fig. 17B is a table showing an example calculation of the percent motor change between a first session and a second session

[0061] Fig. 18 is a plot of the left and right eye ratio over time.

[0062] Fig. 19 provides a block diagram that illustrates one example of a computing device upon which embodiments of the invention may be implemented.DETAILED DESCRIPTION

[0063] The general process of human motion recognition and analysis requires a scene to be captured (a video) with an appropriate system, for example an electronic recording device. Further, movements of a human may be tracked for at least a portion of the video and that tracking can be used for further analysis, for example by processing data generated by the tracking process.

[0064] Human tracking is the identification and estimation of movement of a person or part of a person in a video. Identification and location or orientation estimation of body parts in successive frames of a video for the purposes of movement or motion tracking has many challenges. Example challenges include full or partial obstruction of one or more body parts, changes in lighting, movement of the person within the frame, differences in perspective and / or variation in human physique (e.g., a child vs and adult). Further, classifying body movements of the person may present additional challenges.

[0065] Fig. 1 illustrates an exemplary method 100 for analysing a person’s movement in a video. The analysis is conducted in relation to the persons’ movement while performing an activity, which includes one or more actions. A specific application of the method 100 is analysing a student while performing an activity. The method 100 has other applications, including for example and without limitation analysing a patient for creating data that may be useful for diagnosis or useful for the analysis of movement for diagnosis purposes.

[0066] At step 102 a video recording of a person performing an activity is received. The actions of the activity may include one or more of reading, listening, speaking, typing. If the activity is a student participating in a task, the recording will show the student actively engaged in the task. In some embodiments all the body is in the frame of the video recording. In some embodiments at least the student’s upper body, head, face and hands are in the frame of the video recording. In other embodiments only the user’s head or face is in the frame of the video recording.

[0067] At step 104 the movements of body parts of the student in the video recording are identified. For example, the types of body part movements that are identified may be any one of, or a combination of: eye movement, mouth movement, head movement, hand or hands movement and body position movement. In order to analyse the body part movements of the student, a landmark map of the body part is made. The landmark map is then tracked over time, for example frame-by-frame or with reference to another interval that may correlate to two or more frames, to analyse the movement of that body part.

[0068] Referring specifically to face and head movement, facial landmark detection may be used. Facial landmark detection is a general class of computer vision learning that involves detecting and localising specific landmarks or points on a face. In one simple example, the landmarks may include the left and right eyes, the nose, mouth and chin. In some examples, there are a plurality of landmarks on the face that may be considered together for facial landmark detection. The overall goal of facial landmark detection is to accurately identify the chosen face landmarks in images or in videos. Once these facial landmarks are identified in successive images of a video, they may be analysed to determine movement, for example to determine a change in facial expression.

[0069] Fig. 2 shows an example of a landmark map 200 for a face. Each dot indicates a landmark in the landmark map. In this example, face landmark map 200 includes the following landmarks: right inner eyebrow 202, right outer eyebrow 204, top right eye 206, right pupil 208, outer right eye 210, bottom right eye 212, inner right eye 214, top upper-lip 216, bottom upper-lip 218, right mouth 220, mouth left 222, top lower-lip 224, bottom lower- lip 226, nose tip 228, right nose bridge 230, left nose bridge 232, left inner eyebrow 230, left outer eyebrow 234, top left eye 236, left pupil 238, outer left eye 240, bottom left eye 242 and inner left eye 244. Where the descriptor ‘left’ and ‘right’ here have been used with respect to the user’s perspective.

[0070] Example landmark map 200 as shown is an example of a static face landmark map. Each of the landmarks in this map is defined as the location when the face is static or at rest.

[0071] In some embodiments landmark detection is by a trained machine learning model. An example machine learning model for facial landmark detection currently available include MediaPipe Face Landmarker, available via Google® for Developers. Another example is Microsoft® Face API through Azure®. Both of these examples can detect face landmarks ina video and generate data for analysis of the landmarks over successive frames of a video recording.

[0072] In some examples, head pose estimation of the user may be determined in addition to or as a part of the facial landmark identification and analysis. In general, head pose estimation finds the direction of the user’s head - i.e., which direction they are facing. The head pose of the user in a frame or image can be determined, for example, from the angles of lines between facial landmarks in three-dimensional space - e.g., the Euler angles about an axis through an origin will specify orientation of the head, e.g., roll, pitch, yaw.

[0073] It will be appreciated that for images with depth data, three angles may be determined to estimate the orientation of the head. However, in some examples video image data in two dimensions only may be available. In such examples, two angles may be determined to estimate the orientation of the head of the user.

[0074] In one example, three angles may be determined from a local three dimensional Cartesian axis positioned at a landmark of the user. For example, the local axes may be positioned on the nose tip 228 with the one axis e.g., the X-axis, substantially orthogonal to the tip of the nose such that when the user is substantially face-on in the frame the X-axis is too. The remaining two axis, e.g., the Y-axis and the Z-axis, form a plane orthogonal to the direction of the X-axis. Fig. 2B shows an example of such a local coordinate system. Rotations about the X, Y and Z axis may be called the roll, pitch and yaw respectively.

[0075] The facial landmark detection model may additionally identify and detect the change in position of the Cartesian coordinates of the landmarks by analysing successive frames in a video recording. By analysing the change in position of the Cartesian coordinates and applying geometric calculations, the angles of rotation in the head’s reference frame (i.e., Euler angles) can be determined or the angles of rotation in the space reference frame (i.e., roll pitch and yaw) may be determined.

[0076] Body parts may also be analysed using landmark maps. For example, Fig. 3 shows an example of a landmark map 300 for a hand. Each dot indicates a landmark in the hand landmark map. In this example, hand landmark 300 includes the following landmarks: wrist 302, lower thumb 304, first thumb knuckle 306, second thumb knuckle 308, tip of thumb 310, first index knuckle 312, second index knuckle 314, third index knuckle 316, tip of index finger 318, first knuckle middle finger 320, second knuckle middle finger 322, third knuckle middle finger 324, tip of middle finger 326, first knuckle of the third finger 328,second knuckle of the third finger 330, third knuckle of the third finger 332, tip of third finger 334, first knuckle of the fourth finger 336, second knuckle of the fourth finger 338, third knuckle of the fourth finger 340 and tip of the fourth finger 342. It will be appreciated that in order to track two hands of the user two hand landmark maps may be used - one for each hand of the user. A hand pose estimation model may be used to track the position of the hands in the video recording. In one example, MediaPipe Hand Landmarker model available via Google® for Developers may be used.

[0077] Fig. 4 shows an example of a body landmark map 400 for a portion of the upper body. Each dot indicates a landmark in the body landmark map. In this example, body landmark 400 includes the following landmarks: right ear 402, left ear 402, right shoulder 406, left shoulder 408, right elbow 410, left elbow 412, right wrist 414, left wrist 416, right hip 418 and left hip 420. It will be appreciated that a landmark map may include more or less landmarks than that shown in landmark map 400. For example, body landmark map may also include landmarks corresponding to the hands, facial features or lower body parts. In some example, the body landmark map may include more landmarks and include substantially the whole body also.

[0078] It will be appreciated that some landmark detection models will have more or less landmarks to identify and detect. These models may be considered to be a more dense landmark map or a less dense landmark map, respectively.

[0079] A combination of landmark models may be used together to estimate the position of the landmarks and analyse the person’s movement. For example any two or more of landmarks for the face, hands and body can be used together.

[0080] Returning to Figure 1, at step 106, based on the analyses and detection of the movement of the landmarks identified from the video recording, the method includes classifying detected movements and tracking when the classes of movement occur over time. The classes into which the detected movements are classified are utilised to define behavioural characteristics of the person.

[0081] Examples of movements in the patient that may be analysed are: movement of the landmarks associated with one or more of the eyes, mouth, hands, body and head. In one example, the movement may be the movement of the eyes - the relatively rapid and simultaneous closing and opening of both eyes is a classification - i.e. a class that is commonly referred to as blinking. In a further example an analysis may include the order inwhich two or more detected movements occur. By classifying the detected movement of identified landmarks as blinking and tracking the movements over time, a behavioural characteristic of frequency of blinking may be determined.

[0082] The eye movement may be determined using the eye landmarks. For example, for the left eye the inner left eye 244 and the outer left eye 240 may be used to calculate the width of the left eye. Also the top left eye 226 and bottom left eye 242 may be used to calculate the height of the left eye. Similarly for the right eye, the outer right eye 210 and inner right eye 214 can be used to calculate the width of the right eye. And the top right eye 206 and bottom right eye 212 can be used to calculate the height of the right eye.

[0083] The distances for width and height of the left and right eyes are describable as two-dimensional Euclidean distances. Therefore, the eye ratio (width / height of the eye) can be calculated for each eye. This ratio can then be used to determine whether the left and right eye is in the open or closed state. In one example, the threshold eye ratio is 5.5 and if the left and / or right ratio exceeds this threshold then the eye is determined to be in the open state. Other values of threshold may be used and an effective threshold for a population may be determined experimentally. A ratio measure is useful as it is not affected by a person being closer or further away from the camera.

[0084] Fig. 5A shows a plot 500 of the left eye ratio over time. On the x-axis is time in seconds and on the y-axis is the ratio of the left eye. As shown by the data 502 the ratio of the left eye changes over time. Also shown on plot 500 is the eye ratio threshold 504 of 5.5. The eye state data 505 is also shown in plot 500 and represents when the left eye is in the open state (below threshold 504) or closed state (above the threshold 504). As can be seen the left eye ratio increases over the threshold value three times - see 506, 508 and 510. Thus, it may be determined that the person closes their left eye three times in the approximately 8 second period tracked.

[0085] Fig. 5B shows a plot 530 of the right eye ratio over time. On the x-axis is time in seconds and on the y-axis is the ratio of the right eye. As shown by the data 502 the ratio of the left eye changes over time. Also shown on plot 530 is the eye ratio threshold 534 of 5.5. The eye state data 535 is also shown in plot 530 and represents when the right eye is in the open state (below threshold 534) or closed state (above the threshold 534). As can be seen the right eye ratio increases over the threshold value three times - see 536, 538 and 510. Thus, itmay be determined that the person closes their right eye three times in the approximately 8 second period tracked.

[0086] Returning again to Figure 1, at step 108 a data summary is generated. The data summary may represent the determined movement classes.

[0087] In some embodiments the data summary includes a count of one or more classified movements. For each classified movement a count is determined. For some classified movements the count may be based on a substantially singular event. For example, a count may be determined for an opening of a person’s eye(s) or a rapid closing and opening of a person’s eyes - i.e., one count corresponds to the detection from the video data of a threshold condition associated with a blink. For other classified movements the counts may be related to continuous data. For example, a count may be determined when a rate of change of body movement of a body part exceeds a threshold or when a relatively large change in position of a body part occurs.

[0088] In some examples, the data summary includes time information, for example including the count per time interval, e.g., over a one minute time interval or including data to enable a behavioural characteristic defined by or related to the count per time interval to be determined. Additionally or alternatively, the data summary may include counts per video and / or per session. Additionally or alternatively, the data summary may include an average of counts over a plurality of time intervals. The data summary may include a total count of each movement analysed for a time period. The time period may be a portion of the time of the video recording or may be substantially the whole of the time period of the video recording. Additionally or alternatively, the data summary includes volatility information, for example volatility in one or both of counts of classified movements and measures of extent of body movement. Additionally or alternatively, the data summary includes statistical measures related to the classified and tracked movements over time, such as a standard deviation or variation.

[0089] In other examples, the data summary includes the counts of one or more classified movements per a first time period, e.g., per minute, over a plurality of sessions. Additionally or alternatively, the data summary may include secondary data, such as data defining a trend line corresponding to the counts, which might be within a session or across a plurality of sessions. Alternatively, or additionally a behavioural characteristic based on the counts (e.g. frequency of blinking or a total count) is included in the data summary.

[0090] Continuing with the example of eye movement the data summary may include a count of the number of times the left and / or right eye opened. As discussed already the state of the eye, that is whether it is open or closed, is determined relative to a threshold. For example, the eye closed state can be assigned a ‘0’ and the eye open state can be assigned a ‘ 1’ . Using this binary the eye state over time can be calculated by comparing the eye ratio to the threshold and assigning a 0 if the ratio exceeds the threshold and 1 if the eye ratio does not exceed the threshold. In this way, an eye state for a frame in the video recording can be determined.

[0091] Next, the number of times the patient closes their eye (or eyes) can be determined by counting the number of times the eye state changes from 1 (open eye state) to 0 (closed eye state). Similarly, the number of times a patient opens their eye (eyes) can be determined by counting the number of times the eye state changes from 0 to 1. In some embodiments, it is chosen to count the number of times the eye closes rather than opens.

[0092] Returning to Fig. 5A, it can be seen that the eye close count is three (indicated by 506, 508 and 510). Similarly, for Fig. 5B the eye close count is three.

[0093] In another example, the movement analysed may be the movement of the mouth. In this example the closing and opening of the mouth is the movement pattern. The mouth movement may be determined using the landmarks corresponding to the mouth. For example, using the top upper- lip 216 and the bottom lower-lip 226 landmarks the height of the mouth may be calculated. Also using the right mouth edge 220 and left mouth edge 222 landmarks the width of the mouth may be calculated.

[0094] Again the distances of the mouth height and mouth width are describable as two- dimensional Euclidean distances. From the height and width the mouth ratio may be calculated as width divided by height of the mouth. The mouth ratio is then compared to a threshold to determine if the mouth is in an open state or a closed state. In one example, the threshold is 5.5 and when the mouth ratio is below 5.5 it is determined that the mouth is in the open state. And if the ratio is above 5.5 the mouth is determined to be in the closed state.

[0095] Fig. 6 shows a plot 600 of the mouth ratio over time. On the x-axis is time in seconds and on the y-axis is the ratio of the mouth. As shown by the data 602 the ratio of the mouth changes over time. Also shown on plot 600 is the mouth ratio threshold 604 of 5.5. The mouth state data 605 is also shown in plot 600 and represents when the mouth is in the open state (below threshold 604) or closed state (above the threshold 604). As can be seen themouth ratio decreases over the threshold value three times - see 606, 608 and 610. This example shows a patient with their mouth initially open (at t=0) then closing it three times in the approximately 7 second period tracked. The patient first opens their mouth at approximately 1.5 seconds, then closes it - which is the flat plateau line. Then they open their mouth again at 3.5 seconds, close it at 4.5 seconds, open it again at 5.5 seconds and close it again at 6.5 seconds.

[0096] Similarly to the eye movement, the movement of the mouth may be determined to be in the open or closed state. The count of this change in state may be determined similarly as for the eye movement by assigning the mouth closed state 0 and the mouth open state 1. Then the number of times the mouth closes can be counted by counting the number of times the mouth state changes from 1 to 0. Returning to Fig. 6, this change in state occurs twice - at 606 and 608.

[0097] Other types of movements that are analysed for classification may be more complex than occurring in two-states - i.e., open or closed. In this case, the data summary other than a threshold count is used.

[0098] One example of a class of movement that is not usefully understood as a binary of states is hand movement, for example a pattern of hand movement (or a pattern of movement of another body part). As described herein a hand state may be determined using one or more of the landmarks from the hand landmark map 300. In one example, the hand movement may be assumed to correspond to the movement of the user’s wrists from one position (one hand state) to another position (another hand state). As such, tracking the movement of the user’s wrist using the left and right wrist 302 is sufficient. The movement of each hand is tracked by analysing the position of the corresponding wrist marker in subsequent frames of the video recording. Tracking the wrist(s) of the patient may be useful, since the wrist point is a robust representation of the position of the hand(s).

[0099] In one example a measure of hand movement is the two-dimensional Euclidean distance between wrist coordinates of the left and / or right hand between successive frames in the video recording. Changes in this distance can be tracked over time and movements that satisfy a threshold detected as a count.

[0100] In other examples, the hand movement may be tracked by combining data from two or more landmark models. For example, the hands may be tracked by combining the hand landmark map 300 and the body landmark map 400.

[0101] Fig. 7A shows a plot 700 of the movement of the left hand of a person (e.g. a patient) over a time period. On the x-axis is time in seconds and on the y-axis is distance in pixels. As shown by the data the position or coordinates of the person’s left wrist changes over time - that is, it changes from frame to frame of the video recording.

[0102] Fig. 7B shows a plot 710 of the movement of the right hand of the person over a time period. On the x-axis is time in seconds and on the y-axis is distance in pixels. As shown by the data the position or coordinates of the person’s right wrist changes over time - that is, it changes from frame to frame of the video recording.

[0103] In some embodiments a cumulative sum of the movement of a person’s wrists is determined. The cumulative sum of the movement is useful as it takes into consideration that movement patterns occur over time and not at discrete points in time. For example, a hand may move 10 cm in 1 second.

[0104] The cumulative sum of the movement can be used to count when a peak in movement occurs. For example, when the hand position changes dramatically over time, this represents when a significant hand movement has occurred and may be counted.

[0105] The following provides an example process for implementing a peak-count method according to a first embodiment. The method is described in combination with example Python code for implementing a peak-count method.

[0106] First, the cumulative sum of the data is calculated. For example the cumulative sum of the left hand movement data in Fig. 7A is calculated.# Cumulative Sum cumsum = np.cumsum(np.array(original_list))

[0107] Next, a polynomial function is fitted to the cumulative sum. cumsum_polyfit_coefficients = np.polyfit(index_list, cumsum, 3) cumsum_polyfit_line = np.polyval(cumsum_polyfit_coefficients, index_list)

[0108] Fig. 8 A is a plot of the cumulative sum 802 and polynomial fit 804 for the left hand movement data from Fig. 7A. The x-axis is time in seconds and the y-axis is distance in pixels.

[0109] Next, the polynomial function is subtracted from the cumulative sum to yield the difference function. This has the advantage of enhancing the peaks in the modified data.# Modified Cumulative Sum to exaggerate the peaks mod_cumsum = cumsum - cumsum_polyfit_line

[0110] The difference function output is then smoothed to remove ‘jittery’ outputs. The output is the smoother difference function. mod_cumsum = signal. savgol_filter(mod_cumsum, window_length=10, polyorder=3, mode=“nearest”)

[0111] Fig. 8B is a plot of the smoother difference function 810. The x-axis is time in seconds and the y-axis is distance in pixels.

[0112] Next, a peak detection algorithm is used to detect local maxima in the difference function 810. peaks, > = find_peaks(mod_cumsum)

[0113] The peaks of the smoothed difference function are also shown in Fig. 8B as circles.

[0114] Lastly, the number of peaks are counted. In this example, 25 peaks were counted, being 12 counts for the left hand and 13 counts for the right hand. Either or both the individual hand counts and the total count may be recorded and utilised.

[0115] The data summary for hand movement may include a count of the number of peaks. In some examples the data summary may record the number of peaks for the left and right hand separately. In other examples, the data summary includes the total number of peaks counted for the left and right hands combined. In yet another example, the movement to be analysed is the movement of the patient’s body. In order to analyse the movement of the patient’s body multiple body landmarks can be identified and tracked - see body landmark map 400.

[0116] In one example the body movement may be analysed using the movement of a midpoint between the left shoulder 408 and right shoulder 406. The movement of this midpoint will describe the general movement of the person’s torso in the video recording. In one example, the shoulder midpoint is calculated using the left shoulder landmark 408 and right shoulder 406. The use of the shoulder midpoint for determining body movement may be preferable in video recordings where the person is shown from the shoulder’ upward.

[0117] Alternatively, the body movement may be tracked by a new landmark corresponding to the shoulder midpoint 420. To track the body movement over time, thecoordinate of the midpoint between the shoulders is tracked. This may be done by calculating the Euclidean distance between the shoulder midpoint landmark in successive frames in the video recording. The change in position (i.e., the distance calculated) can then be tracked by plotting this variable over a period of time.

[0118] Fig. 9 A is a plot 900 of the calculated change in distance of the shoulder midpoint between successive frames in the video recording over a period of time. On the x-axis is time in seconds and on the y-axis is distance in pixels.

[0119] The data shown in plot 900 appears to be rather jittery in part because it can be difficult to approximate the shoulder midpoint when a patient is covered by clothing. In some embodiments the cumulative sum of the movement is determined. In general, the cumulative sum is a useful parameter when a body movement occurs over time. For example, the body (e.g., shoulder mid-point) may move over a time period of one second or more.

[0120] In order to track the body movement over time a peak-count method is used. Firstly, a cumulative sum of the body movement data is calculated. Fig. 9B is a plot of the cumulative sum 902 of the body movement corresponding to the movement of the shoulder mid-point data from Fig. 9A. A polynomial function is fitted to the cumulative sum data. Data 904 of Fig. 9B shows the fitted polynomial function.

[0121] Next, a difference function is calculated by subtracting the cumulative sum 902 from the polynomial function 904. In this example, the difference function is smoothed to remove the ‘jittery’ data. Fig. 9C is a plot of the smoothed difference function 910.

[0122] A peak detection algorithm may be used to determine the number of peaks of the (smoothed) difference function. The detected peaks are shown as circles along the smoothed difference function 910. In this example, 28 peaks are detected and counted.

[0123] The data summary for body movement may include a count of the number of peaks detected. The count may be used as a behavioural characteristic, especially where the videos analysed are of fixed duration. A count per time interval of the number of peaks detected may be used as the behavioural characteristic, which may accommodate for different durations of video.

[0124] In yet another example, body movement may be analysed using a plurality of the landmarks from the body landmark map 400. In one example, the body movement may correspond to a collective movement of the user’s body. For example, the body movementmay be a measure of the average movement of a combination of landmarks. A peak in a cumulative measure of the average movement may then indicate a count.

[0125] In other examples, the body movement of the patient may be measured and tracked using an accelerometer worn by the patient during the video recording.

[0126] In yet another example, the movement to be analysed may be the position of the head. As discussed above, the head positon can be estimated based on the calculated Euler angles. To track the head movement, a unit vector is used to represent the angles of rotation about the X, Y and Z axis of the head pose. Where a rotation about the X-axis is the roll angle, a rotation about the Y-axis is the pitch angle and a rotation about the Z-axis is the yaw angle.

[0127] The angle change between successive unit vectors over time, represents the patient’s head moving.

[0128] The roll, pitch and yaw angles is converted to a unit vector. This may be done using the following formula: x = cos(yaw) * cos(pitch) y = sin(yaw) * cos(pitch) z = sin (pitch)

[0129] This formula converts the roll, pitch and yaw angles to a unit vector (a vector of magnitude 1) which starts at the origin (0,0,0) and terminates at the point (x, y, z).

[0130] Thus for each frame of the video recording the head position can be summarised by a single unit vector. The angle between the unit vector of successive frames in the video recording can now be calculated using trigonometry. For example, the angle (0) between vector A and vector B can be determine using the following formula:6 = cos-1(A • B)

[0131] Fig. 10A is a plot of the calculated change in angle between unit vectors of successive frames in the video recording over a period of time. On the x-axis is time in seconds and on the y-axis is angle in degrees.

[0132] In some embodiments a cumulative sum is used to see when a large peak of head movement occurs. The peak then represents when a large head movement occurs. In order to track the head movement over time a peak-count method is used.

[0133] Fig. 1 OB is a plot of the cumulative sum 1002 and polynomial fit 1004 for the head movement data from Fig. 10A.

[0134] Fig. 10C is a plot of the smoothed difference function 1010. The detected peaks are shown as circles along the smoothed difference function 1010.

[0135] The peaks can then be counted. In this example, the number of peaks 25.

[0136] The following provides an example process for implementing a peak-count method according to a second embodiment. The method uses a speed of movement.

[0137] In one example, a measure of hand movement is the Euclidean distance between the left / right wrist (e.g., landmark 302, 414 or 416) and a nose landmark (e.g., landmark 228) between successive frames in the video recording. Changes in this distance can be tracked over time and movements that satisfy a threshold (e.g. a relatively large movement) detected as a count. For example, when a significant distance is observed over multiple successive frames. Tracking the wrist(s) of the patient using a nose landmark as the relative point may be useful because it is robust in detection and it allows the hand movement to be relative to the position of the body.

[0138] The following provides an example process for identifying the patterns of movement for a hand movement.

[0139] First, the distance between the wrist(s) landmark and a nose landmark is calculated for every frame in the tracked period of time.

[0140] Next, the change in distance between successive frames is calculated. This change in distance over time may be understood to be the speed of the hand movement or hand movement speed. The hand movement speed may be calculated using the Python NumPy package diff() function which calculates the change in distance between the wrist(s) landmark and the nose landmark between successive frames.

[0141] Next, the speed data are converted to positive values. This transformation ensures that the positively valued speed data has positive curvature and therefore the peaks can be easily counted. For example, this may be performed using the Python NumPy package abs(s) function.

[0142] Next, the positive speed data are smoothed over time. This data may be called the smoothed data. The positive speed data can be smoothed using mathematical signal processing techniques. For example using a moving average filter or a Savitzky-Golay filter.

[0143] Lastly, the peaks of the smoothed data are counted. In an example, the peaks are counted using the Python SciPy package find_peaks function to find local maxima (i.e., peaks) in the positive speed data. The number of peaks counted represents a summary of the hand movement(s).

[0144] Figure 11 A shows a plot 1100 of the movement of the left hand of a patient over a time period. On the x-axis is the time in seconds and on the y-axis is the speed in units of change in pixels per second. The left hand movement speed data 1102 are shown in the plot 1100. As seen in the plot 1100 data 1102 is rather jittery and may be difficult to determine a count of peaks from this data alone. The vertical lines 1104 represent video frames where no hand movement was detected. The smoothed data 1106 is also shown in plot 1100. The local maxima (peaks) are shown as peaks 1108, 1110 and 1112. In this example, the left hand movement count is three.

[0145] Figure 1 IB shows a plot 1120 of the movement of the right hand of a patient over a time period. On the x-axis is the time in seconds and on the y-axis is the speed in units of change in pixels per second. The right hand movement speed data 1132 are shown in the plot 1130. As seen in the plot 1130 data 1132 is rather jittery and may be difficult to determine a count of peaks from this data alone. The vertical lines 1134 represent video frames where no hand movement was detected. The smoothed data 1136 is also shown in plot 1130. The local maxima (peaks) are shown as peaks 1138, 1140 and 1142. In this example, the right hand movement count is three.

[0146] Again, the data summary for hand movement using this method may include a count of the number of peaks. In some examples the data summary may record the number of peaks for the left and right hand separately. In other examples, the data summary includes the total number of peaks counted for the left and right hands combined.

[0147] Using the speed of movement to trigger a count may be applied to other body movements, in addition to or instead of hand movement.

[0148] In one example, the body movement may be analysed using the two-dimensional Euclidean distance between the left shoulder (e.g., using landmark 408) and right shoulder (e.g., using landmark 406). The distance between the left and right shoulder may be called the shoulder width. The change in the shoulder width will describe the general movement of the person’s torso in the video recording including torso rotation and torso leaning. The shoulder width changes when the body rotates, respective to the position of the camera. When thepatient leans forward or backwards, the shoulder width may also change, altering the depth perception relative to the camera and increasing or decreasing the shoulder width.

[0149] The body movement peaks are counted using the body movement speed, using a similar method as described above with regard to hand moment count.

[0150] To track the body movement over time, the distance between the left and right shoulders is tracked between successive frames. This may be done by calculating the two- dimensional Euclidean distance between the left and right shoulder landmark in successive frames in the video recording. The change in position (i.e., the speed calculated) can then be tracked by plotting this variable over a period of time.

[0151] Next, the speed data values are converted to positive values. This transformation ensures that the positively valued speed data has positive curvature and therefore the peaks can be easily counted. For example, this may be performed using the Python NumPy package abs() function.

[0152] Next, the positive speed data are smoothed over time. This data may be called the smoothed data. The positive speed data can be smoothed using mathematical signal processing techniques. For example using a moving average filter or a Savitzky-Golay filter.

[0153] Eastly, the peaks of the smoothed data are counted. In an example, the peaks are counted using the Python SciPy package find_peaks() function to find local maxima (i.e., peaks) in the smoothed positive speed data. The number of peaks counted represents a summary of the hand movement(s).

[0154] Figure 12A shows a plot 1200 of the movement of the body movement of a patient over a time period. On the x-axis is the time in seconds and on the y-axis is the speed in units of pixels per second. The body movement speed data 1202 are shown in the plot 1200. As seen in the plot 1200 data 1202 is rather jittery and may be difficult to determine a count of peaks from this data alone. The smoothed data 1204 is also shown in plot 1200. The local maxima (peaks) are shown as peaks 1206, 1208, 1210, 1212, 1214, 1216 and 1218. In this example, the body movement count is seven. The data summary for body movement may include a count of the seven peaks in this example.

[0155] Next, the speed data are converted to positive values. This transformation ensures that the positively valued speed data has positive curvature and therefore the peaks can be easily counted. For example, this may be performed using the Python NumPy package abs() function.

[0156] Next, the positive speed data are smoothed over time. This data may be called the smoothed data. The positive speed data can be smoothed using mathematical signal processing techniques. For example using a moving average filter or a Savitzky-Golay filter.

[0157] Lastly, the peaks of the smoothed data are counted. In an example, the peaks are counted using the Python SciPy package find_peaks() function to find local maxima (i.e., peaks) in the smoothed positive speed data. The number of peaks counted represents a summary of the hand movement(s).

[0158] The movement of the patient’s head may also be determined using the speed of movement method.

[0159] Figure 12B shows a plot 1230 of the movement of the patient’s head over a time period. On the x-axis is the time in seconds and the y-axis is the speed of the head angle in degrees. The head movement speed data 1232 are shown in the plot 1230. As seen in the plot 1300 data 1232 is rather jittery and may be difficult to determine a count of peaks from this data alone. The smoothed data 1234 is also shown in plot 1230. The local maxima (peaks) are shown as peaks 1236, 1238, 1240, 1242 and 1244. In this example, the head movement count is five. The data summary for head movement may include a count of the five peaks in this example.

[0160] Next, the positive speed data are smoothed over time. This data may be called the smoothed data. The positive speed data can be smoothed using mathematical signal processing techniques. For example using a moving average filter or a Savitzky-Golay filter.

[0161] Lastly, the peaks of the smoothed data are counted. In an example, the peaks are counted using the Python SciPy package find_peaks() function to find local maxima (i.e., peaks) in the smoothed positive speed data. The number of peaks counted represents a summary of the hand movement(s).

[0162] In another embodiment, the non-binary movement class is determined using a different method.

[0163] Again returning to Fig. 1, at step 108, a data summary of the classified movements is generated. Figs. 13 to 18 show specific example data summaries that may be generated.

[0164] Fig. 13A is an example data summary generated using method 100, presented as a table for ease of illustration. In this example, the data summary of multiple time periods orvideo recordings is shown. The total number of counts of eye movement, mouth movement; body movement; head movement and hand movement is shown. The data summary may also include the duration of the time period, which in some examples may be the length of the video recording, or may be the length of a portion of the video recording that was analysed.

[0165] Fig. 13B shows another example data summary generated using method 100. In this example, the count of the movement is shown as a function of the duration of the time period. The data shown in Fig. 13B may be utilised in addition to or instead of the data shown in Fig. 13A.

[0166] Fig. 14A is an example of a data summary of a movement count per minute of a person’s eyes, mouth, head, hands and body over multiple recorded videos or video portions. In addition to the count data for each video recording there is also a trend line corresponding to each of the eyes, mouth, head, hands and body movements of the user.

[0167] Fig. 14B shows another example data summary of the movement count per minute as an average per group of videos (e.g. a group taken over a single session) a video or video portion that has been analysed. The average count per minute relating to the movement of the eyes, mouth, head, hands and body of the person is plotted as a function of the session.

[0168] By comparing movement between sessions of the same person, a change in motor system across sessions is highlighted. For example, when the videos are of a person completing an activity, comparing the change across sessions can represent changes in the movement of the person.

[0169] Fig. 14C shows a focused version of the plot of Fig. 14B. In particular, Fig. 14C shows only the count per minute for the movement of the eyes and hands of the person plotted over a series of sessions. In this example, it can be seen that the person moves their hands more over time and their eyes less over time.

[0170] Fig. 15A shows yet another example of a data summary generated by method 100. Here, the plot shows the movement counts of the eyes, mouth, hands, head and body as a percentage of the total counts for the time duration, for example the duration of the video recording. This data summary allows for insight into the interrelationship between the different movements. That is to better understand the motor systems at work during activities and how they all work together.

[0171] Fig. 15B shows yet another example of a data summary generated by method 100. Here, the plot shows the movement counts of the eyes, mouth, hands, head and body as apercentage of the total counts for the session. This data summary allows for understanding of the changing motor systems between sessions.

[0172] Fig. 15C is an example of movement data. In table 1510 data for four sessions are shown. For each session the movement of the eyes, mouth, body, head and hands is shown as a percentage of the total detected movements. The movement count percentage is calculated as the total individual movement count (eye, mouth, body, head or hands) divided by the total movement count (eye, mouth, body, head and hands) for a given session. In Fig. 15C the percentages are shown in decimal form - e.g. 0.50 means 50%.

[0173] Fig. 15D shows the percent motor change (y-axis) over different sessions (x- axis). The percent motor change represents the net difference between each session for the combination of movements.

[0174] Fig. 13E shows a table with an example calculation for the percent motor change between four sessions. The percent motor change is calculated by first subtracting the movement count percentage of a first video from the movement count percentage of a second video. In this example, the second video corresponds to a time after the first video. This subtraction is performed for each movement - eyes, mouth, body, head and hands. The resulting movement count difference between the first and second video are shown in row 1512 of the table. The movement count difference between the second and third video are shown in row 1514 of the table. The movement count difference between the third and fourth video are shown in row 1516 of the table.

[0175] Next, the positive movement count differences are summed to yield the percent motor change. The calculated percent motor change is shown in the last column of the table. For example, in row 1312 the positive values are 0.07 (eyes) and 0.06 (head) the remaining values for mouth, body and hands are negative. Summing the positive values together yields the percent motor change of 0.23 %.

[0176] The percent motor change is also shown in Figs. 15A and 15B as trend lines 1502 and 1504.

[0177] Fig. 16A is a pie chart showing the movement counts for a first session as a percentage of the total movement counts for the first session. The movement counts are shown for the eyes, mouth, hands, body and head.

[0178] Fig. 16B is another pie chart showing the movement counts for a second session as a percentage of the total movement counts for the second session. The movement countsare shown for the eyes, mouth, hands, body and head. Here, the second session occurred after the first session. As such, by comparison of the two pie charts it can be seen how the user’s movements have changed between the two sessions. In some embodiments it may be preferable to compare the first session with the latest or more recent session.

[0179] Fig. 17A shows a percent motor change between two sessions. In this example, the percent motor change is calculated based on the data of the first session shown in Fig. 16A and the data from a second session shown in Fig. 16B. The percent motor change is the net positive difference between two sessions.

[0180] Fig. 17B is a table showing an example calculation of the percent motor change between a first session and a second session. The data for Fig. 17B was different to the data for Fig. 17A.

[0181] Fig. 18 is a plot of the eye movement over time. On the x-axis is time in seconds and on the y-axis is eye ratio. The ratio of one eye 1802 and the other eye 1804 are shown separately. As can be seen from this plot the left and right eyes are not synchronised and therefore are not working together and blinking at the same time. This may be indicative of a the eye movement of a stroke patient where different parts of the brain are not working properly.

[0182] As such, based on the generated information summary of the method an action may be taken. The action may be to assess a capability of the person. Alternatively or in addition the action may be to diagnose the person, for example to diagnose the person has had a stroke.

[0183] Returning again to Figure 1, at step 110 an action is performed based on the data summary. In other embodiments step 110 is omitted.

[0184] In one example, the action performed is an assessment of the user’s capability. For example, by comparing the percent motor change over a plurality of sessions, it may be determined if the user is in an upward curve, a downward curve or remaining static. Based on this, it may be determined that the session activity is becoming too easy (or hard) or is too easy (or hard) and therefore the activity can be changed or adjusted.

[0185] If the body movement is detected over sessions, an example of the action performed is an adaption of an activity plan based on the data summary generated. For example, the method may be performed on a number of videos recording a user developing speech capability. Based on the data summary generated at step 108 where the classifiedeye(s), mouth, head, body, hand(s) counts are shown for the videos recorded, it may be determined that the user is progressing well and as such the activity plan should be adapted to extend the user.

[0186] In yet another example, the action performed is a diagnostic action relating to the user. For example, the present techniques enable a group of people with a diagnosis to be analysed and compared with a control group without diagnosis and movement indicators that may be used for diagnosis identified.

[0187] In another example, the movement detection is utilised to indicate proprioception or likely proprioception. The movement detection may be tracked over time and a period of in which is relatively void of movement may indicate proprioception. For example a slow down or absence of blinking may indicate proprioception, or a slow down or absence of mouth movement, or a combination of a slow down or absence of blinking or mouth movement may indicate proprioception. Other body parts, may be utilised to indicate proprioception, either alone or in combination with one or both of the eyes and mouth.

[0188] Aspects of the present disclosure may be implemented on a video analysis tool. The video analysis tool may be a computing device including one or more hardware processors programmed to perform the relevant operations pursuant to program instructions stored in firmware, memory (e.g. a non-transient or non-transitory memory device), other storage, or a combination. In certain embodiments, the analysis tool may be a desktop computer system, a portable computer system such as a laptop, tablet computer or smartphone, a computer server or any other computing device configured to perform the operations described herein.

[0189] In some embodiments the video analysis tool also has capability to record a video for analysis. For example, a camera may be included in or connected to the video analysis tool to record a video.

[0190] In some embodiments, the video analysis tool has the ability to present material to a user to invoke activities in the form of a content or skill outcome tool. For example, in the context of content memorisation, the video analysis tool may include an electronic display to present material to be memorised. The video analysis tool may include an audio output, for example one or more speakers or connections for headphones or earphones. Further, the video analysis tool can record videos of the patient utilising the learning tool. Alternatively or additionally, the video analysis tool may receive a video recording of a person engaged in a task such as a language task or other activity.

[0191] It will be appreciated that if multiple users use the same device, it may retrieve and store content for each of the users. Then depending on the user identified by the video analysis tool (either through user input or recognition from a captured image), the video analysis tool may utilise content suitable for the identified user during an interaction session.

[0192] By way of example, Fig. 19 provides a block diagram that illustrates one example of a computing device upon which embodiments of the invention may be implemented. As described previously, the computing device may be a standalone device or it may be a multipurpose computing device such as a PC, a tablet, or a mobile phone onto which a software program for activities and video analysis can be installed and executed in addition to other software programs.

[0193] Video analysis tool 1900 includes a bus 1902 or other communication mechanism for communicating information, and a hardware processor 1904 coupled with bus 1902 for processing information. Hardware processor 1904 may be, for example, a microprocessor, a graphical processing unit, or other processing unit.

[0194] In terms of storage, the learning tool 1900 includes a main memory 1906, such as a random access memory (RAM); a volatile memory 1908 (e.g., a ROM) for storing static information and instructions, and / or a non-transient or non-transitory memory 1910 (e.g., a storage device). One or more of these memory modules store data such as content 1920 and user profiles 1922, and program modules 1924 for execution by the processor 1904. The memory modules are coupled to bus 1902.

[0195] The program modules 1924 may be stored as independent instruction modules such as a speech recognition module, a facial recognition module, an image capturing and / or processing module, an input module etc. These modules may be invoked in any given order or concurrently depending on the required outcomes. Moreover, one program module may invoke one or more other program modules during operation without departing from the scope of the present disclosure.

[0196] Analytics module 1926 is configured to analyse the movements of the user in a video recording using machine learning models to identify and detect landmarks. The analytics module 1926 is also configured to determine and track patterns in the movement over time. As well as to generate a data summary of the movement patterns over time.

[0197] Further, data may be stored in one or more databases maintained by the video analysis tool 1900. For example, one database may store program modules, another databasemay store content 1920, another database may maintain a user’s required outcomes and a further database may store user profiles 1922, for example. The analysed data and data summaries may also be stored in the analytics modules 1926. These databases may be implemented as relational databases in one example.

[0198] The main memory 1906 may also be used for storing temporary variables or other intermediate information during execution of instructions to be executed by processor 1904. Such instructions, when stored in non-transitory storage media accessible to processor 1904, render video analysis tool 1900 into a special-purpose machine that is configured to perform the operations specified in the instructions.

[0199] The video analysis tool 1900 may also include output devices 1912 to provide an output. The output devices include a display (such as an LCD, LED, touch screen display or other display), for displaying a user interface and an audio device (e.g. a speaker, headphones, or other audio device) configured to output sound in accordance with instructions received from the processor 1904.

[0200] The video analysis tool also includes one or more user input devices 1914 to receive user input. These input devices may be in the form of a touch sensitive panel (not shown) physically associated with the display to collectively form a touch-screen, a keypad (not shown) coupled to the bus 1902 for communicating information and command selections to processor 1904, a cursor control (not shown) for communicating direction information and command selections to processor 1904 and for controlling cursor movement on the display, a microphone for detecting audio inputs, and / or an image capturing device 1928 such as a camera or video recorder (for detecting visual inputs, capturing footage of the user, etc.). The video recording captured by device 1928 is the video recording that is analysed in the analytics module 1926.

[0201] According to one embodiment, the techniques herein are performed by the video analysis tool 1900 in response to processor 1904 executing sequences of one or more instructions (e.g. the method described previously) contained in main memory 1906. Such instructions may be read into main memory 1906 from another storage medium, such as a remote database (not shown) or the non-transient or non-transitory memory 1910. Execution of the sequences of instructions contained in the main memory 1906 causes the processor 1904 to perform the process steps described herein. In alternative embodiments, hard-wired circuitry may be used in place of or in combination with software instructions.

[0202] Video analysis tool 1900 also includes a communication interface 1916 coupled to bus 1902. Communication interface 1916 provides two-way data communication coupling to a network link that is connected to a communication network 1918. For example, communication interface 1916 may be an integrated services digital network (ISDN) card, cable modem, satellite modem, etc. As another example, communication interface 1916 may be a local area network (LAN) card to provide a data communication connection to a compatible LAN. Wireless links may also be implemented. In any such implementation, communication interface 1916 sends and receives electrical, electromagnetic or optical signals that carry digital data streams representing various types of information.

[0203] Video analysis tool 1900 can send messages and receive data, including content and program code, through the network(s) 1918, network link and communication interface 1916. For example, the video analysis tool 1900 may periodically or in response to a trigger condition download / upload content 1920 and / or user profiles 1922 from / to an external server (not shown). The received content 1920 may be executed by processor 1904 as it is received, and / or stored in non-transient or non-transitory memory 1910, or other non-volatile storage for later execution.

[0204] In Fig. 19 the video analysis tool 1900 is shown as a standalone computing device onto which a learning software program (in the form of instructions) can be installed and executed to perform the functions described herein. In other embodiments the learning tool may be part of a server-client architecture, in which the learning tool includes a client in communication with a server that hosts the learning software program and content.

[0205] The term "comprising" (and its grammatical variations) as used herein are used in the inclusive sense of "having" or "including" and not in the sense of "consisting only of".

[0206] It will be appreciated by persons skilled in the art that numerous variations and / or modifications may be made to the invention as shown in the specific embodiments without departing from the spirit or scope of the invention as broadly described. The present embodiments are, therefore, to be considered in all respects as illustrative and not restrictive.

Claims

CLAIMS1. A computer implemented method comprising: receiving a video recording of a person; analysing movements of body parts of the person based on identified landmarks; classifying and tracking movements over time; generating a data summary of the classified movements; and performing an action based on the data summary.

2. The method of claim 1, wherein the movements analysed are one or more of: eye movement, hand movement, mouth movement, head movement or body movement.

3. The method of any one of claims 1 or 2, wherein the data summary generated includes a percent motor change over time.

4. The method of claim 1, wherein prior to generating the data summary, the method further includes the steps of: receiving at least one additional video recording of the user; analysing movements of the user in the at least one additional video recording using machine learning methods to detect and identify landmarks; and determining and tracking patterns in movement over time.

5. The method of claim 4, wherein the data summary generated includes a comparison of the patterns of movements of the user between video recordings.

6. The method of claim 4 or 5, wherein the data summary includes a percent motor change for each of the video recordings analysed.

7. The method of any one of claims 1 to 6, wherein the identified landmarks comprise landmarks of a face, the landmarks of the face comprising a combination of two or more of right inner eyebrow, right outer eyebrow, top right eye, right pupil, outer right eye, bottom right eye, inner right eye, top upper-lip, bottom upper-lip, right mouth, mouth left, top lower-lip, bottom lower-lip, nose tip, right nose bridge, left nose bridge, left inner eyebrow, left outer eyebrow, top left eye, left pupil, outer left eye, bottom left eye and inner left eye.

8. The method of claim 7, wherein said analysing movements of body parts of the person based on identified landmarks comprises determining head pose and wherein determining head pose is based on angles of lines between the identified landmarks.

9. The method of claim 8, wherein the identified landmarks comprise the nose tip and the angles of lines between identified landmarks comprise an angle of a line between the nose tip and another identified landmark.

10. The method of any one of claims 7 to 9, further comprising determining an angle of rotation of a head based on the identified landmarks.

11. The method of any one of claims 1 to 10, wherein the identified landmarks comprise landmarks of a hand, the landmarks of the hand comprising a combination of two or more of wrist, lower thumb, first thumb knuckle, second thumb knuckle, tip of thumb, first index knuckle, second index knuckle, third index knuckle, tip of index finger, first knuckle middle finger, second knuckle middle finger, third knuckle middle finger, tip of middle finger, first knuckle of the third finger, second knuckle of the third finger, third knuckle of the third finger, tip of third finger, first knuckle of the fourth finger, second knuckle of the fourth finger, third knuckle of the fourth finger and tip of the fourth finger.

12. The method of any one of claims 1 to 10, wherein the identified landmarks comprise landmarks of an upper body, the landmarks of the upper body comprising a combination of two or more of right ear, left ear, right shoulder, left shoulder, right elbow, left elbow, right wrist, left wrist, right hip and left hip.

13. The method of any one of claims 1 to 12, wherein said classifying and tracking movements over time comprises determining a frequency of blinking of an eye.

14. The method of claim 13, wherein determining the frequency of blinking comprises determining a ratio of width and height of the eye and determining whether the eye is an open or closed state based on the determined ratio.

15. The method of claim 14, wherein determining the frequency of blinking comprises determining a number of times the ratio of width and height traverses a threshold value over time.

16. The method of any one of claims 1 to 15, wherein said classifying and tracking movements over time comprises determining when a rate of change of body movement of a said body part exceeds a threshold value.

17. The method of any one of claims 1 to 15, wherein said classifying and tracking movements over time comprises determining when a change in position of a said body part exceeds a threshold value.

18. The method of any one of claims 1 to 17, wherein the identified landmarks comprise landmarks of a mouth and said classifying the tracking movements over time comprises determining opening, closing or opening and closing of the mouth.

19. The method of any one of claims 1 to 18, wherein the identified landmarks comprise landmarks of a mouth, the landmarks of the mouth comprising a combination of two or more of the top upper-lip, the bottom lower-lip, right mouth edge and left mouth edge.

20. The method of claim 18 or claim 19, wherein determining opening, closing or opening and closing of the mouth comprises determining when a ratio of mouth height and mouth width traverses a threshold value.

21. The method of any one of claims 1 to 20, wherein the identified landmarks comprise at least one landmark of a wrist and wherein said classifying and tracking movements over time comprises determining wrist movement.

22. The method of claim 21, wherein said determining wrist movement comprises determining a cumulative sum of wrist movement over a period of time.

23. The method of claim 22, wherein said determining wrist movement comprises determining a difference function between the cumulative sum of wrist movement and a polynomial best fit function of the cumulative sum and determining local maxima or peaks in the difference function.

24. The method of claim 23, wherein determining local maxima or peaks in the difference function comprises determining local maxima in a smoothed form of the difference function.

25. The method of any one of claims 1 to 24, wherein said classifying and tracking movements over time comprises determining movement of a midpoint between the left should and the right shoulder.

26. The method of claim 25, wherein said classifying and tracking movements over time comprises determining a cumulative movement of the midpoint between the left shoulder and the right shoulder.

27. The method of claim 26, wherein said determining movement of the midpoint comprises determining a difference function between the cumulative movement of the midpoint and a polynomial best fit function of the cumulative sum and determining local maxima or peaks in the difference function.

28. The method of claim 23, wherein determining local maxima or peaks in the difference function comprises determining local maxima in a smoothed form of the difference function.

29. The method of any one claims 1 to 28, wherein said classifying and tracking movements over time comprises determining a cumulative sum of movement based on the identified landmarks, determining a best fit to the cumulative sum and performing local maxima or peak analysis of a difference between the cumulative sum and the best fit.

30. The method of any one claims 1 to 28, wherein said classifying and tracking movements over time comprises determining a cumulative sum of movement based on the identified landmarks, determining a best fit to the cumulative sum and performing local maxima or peak analysis of a difference between a smoothed form of the cumulative sum and the best fit.

31. The method of any one of claims 1 to 30 wherein said classifying and tracking movements over time comprises determining a speed of movement.

32. The method of claim 31 wherein said classifying and tracking movements over time comprises determining a count of peaks or local maxima in the determined speed of movement.

33. The method of claim 31 or claim 32, wherein the speed of movement is a speed of movement of at least one hand.

34. The method of claim 31 or claim 32, wherein the speed of movement is a speed of change in a distance between the left and right shoulders.

35. The method of claim 31 or claim 32, wherein the speed of movement is a speed of movement of a head.

36. The method of any one of claims 1 to 35, wherein generating a data summary of the classified movements comprises determining a measure of volatility of the classified movements.

37. The method of any one of claims 1 to 35, wherein generating a data summary of the classified movements comprises determining a measure of variation of the classified movements.

38. A computer system comprising a processor and computer storage including instructions that when executed by the processor cause the computer system to perform the method of any one of claims 1 to 37.

39. Non-transitory computer readable storage storing instructions to cause a processor of a computer system to perform the method of any one of claims 1 to 37.