Movement disorder detection and assessment

The method and system for TD assessment using video analysis and machine learning models address the underdiagnosis issue by providing a standardized, remote, and objective TD diagnosis, improving patient access and treatment adherence.

WO2026033526A1PCT designated stage Publication Date: 2026-02-12TEVA PHARMA IND LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/IL2025/050676
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-09
Filing Date
2025-08-07
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Tardive dyskinesia (TD) is often underdiagnosed due to the time-consuming and subjective nature of conventional Abnormal Involuntary Movement Scale (AIMS) tests, which require in-person visits to trained clinicians and can lead to inconsistent results, making it difficult for patients to receive timely and accurate diagnoses.

Method used

A method and system for assessing TD risk using video data analysis, involving facial landmark detection, machine learning models, and automated scoring to determine a total severity score indicative of TD diagnosis, enabling remote and standardized assessment.

Benefits of technology

Facilitates timely and standardized TD diagnosis through remote video analysis, reducing underdiagnosis and improving patient access to appropriate treatment by providing objective and consistent severity scores.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IL2025050676_12022026_PF_FP_ABST
    Figure IL2025050676_12022026_PF_FP_ABST
Patent Text Reader

Abstract

A method for assessing severity of involuntary movement associated with tardive dyskinesia (TD) may include receiving video data that corresponds to facial movements of the patient, and processing the video data to identify facial landmark data of the patient. The method includes generating a plurality of long segments from the video data, such that the plurality of long segments are defined by durations of the video data between the excluded portions of the video data. The method includes generating a plurality of short segments based on the plurality of long segments, and applying, for each short segment, a first trained machine learning model to the facial landmark data of the short segment to generate a movement indicator for each of the plurality of short segments, and applying a second trained machine learning model to the movement indicators of the short segments to determine a total severity score associated with TD.
Need to check novelty before this filing date? Find Prior Art

Description

MOVEMENT DISORDER DETECTION AND ASSESSMENTCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of Provisional U.S. Patent Application No. 63 / 681,255, filed August 9, 2024, the disclosure of which is incorporated herein by reference in its entirety.BACKGROUND

[0002] Tardive dyskinesia (TD) is a movement disorder that affects the nervous system. For example, TD may cause a range of involuntary repetitive muscle movements in the face, arms, and legs. Typically, a person develops TD by using certain psychiatric drugs for an extended period of time. TD is generally underdiagnosed as patients are typically prescribed psychiatric drugs from a psychiatrist who monitors that patient’s response to the drug but is not trained in diagnosing movements disorders caused by the prescribed drug. Further, TD may also be underdiagnosed as trained neurologists are not attentive to the possibility of TD causing the patient’s symptoms.

[0003] For those patients seeking a diagnosis for or being treated for TD, the patient is usually advised to undergo a time consuming Abnormal Involuntary Movement Scale (AIMS) test. A conventional AIMS test may require the patient to visit a trained clinician in person, making the AIMS test even more time consuming. As such, patients may quit seeking a diagnosis or treatment for TD after not receiving a timely TD diagnosis or after not seeing appropriate clinical improvement due to infrequent AIMS testing. Further, the results of an AIMS test may be inconsistent between clinicians, as a portion of the test relies upon at least partially subjective judgments made by the clinician.SUMMARY

[0004] The present disclosure relates generally to detecting and assessing a movement disorder, and more particularly, to assessing a risk of a patient for tardive dyskinesia (TD). Thedisclosed technology relates to a method and systems (e.g., devices) for assessing risk of a patient for TD.

[0005] Methods and systems for assessing severity of involuntary movement associated with TD may be described herein. Methods for assessing severity of involuntary movement associated with tardive dyskinesia (TD) are described herein. In some examples, the method may include receiving video data of the patient. The video data may correspond to facial movements of the patient. For instance, the video data may include video frames of the patient’s head and / or face, and the video frames may include facial movements of the patient. The video data may correspond to facial movements of the patient during an assessment session (e.g., the video data may be captured during an assessment session, such as a TD assessment session, of the patient that is performed by a physician). In some examples, the method may include providing a prompt to the patient to perform an action (e.g., where the video data is captured while the patient is performing the action). Further, in some examples, the method may include providing the prompt to the patient by displaying an instruction to perform the action on a user interface of an electronic device (e.g., via the patient’s smartphone or other electronic device). As such, in some examples, the method includes capturing one or more videos of the patient, where the video data of the patient is derived from (e.g., includes) the one or more videos of the patient.

[0006] The method may include processing the video data to identify facial landmark data of the patient. In some examples, processing the video data may include applying an image processing technique to the video data to identify facial landmarks in one or more frames of the video data to produce labeled frames. The image processing technique may include any combination of a dlib facial recognition method, a Haar Cascade method, a Fisherface method, or an elastic graph matching method. In some instances, the processing of the video data may include identifying the facial landmark data that corresponds to each of the facial landmarks from the labeled frames. The facial landmark data may include, for example, one or more of a landmark identifier, landmark coordinate position, frame source identifier, video source identifier, and / or patient identifier.

[0007] The method may include segmenting the video data into a plurality of long segments. Portions of the video data may be excluded from the plurality of long segments. For example, the method may include filtering the video data to remove portions of the video data that are notincluded within the long segments e.g., manually or using an image processing technique). The plurality of long segments may be defined by the durations of the video data between the excluded portions of the video data. For instance, the long segments may be the video data that was not excluded. The excluded portions of the video data may be portions where the patient is not looking at the camera, where facial landmark data of the patient is not identified within the video data, and / or where the patient is not performing a predefined task.

[0008] The method may include segmenting each of the plurality of long segments into a plurality of short segments. The long segments may be defined by multiple different durations. For example, the long segments be aperiodic and / or the long segments may be defined by varying or non-consistent durations (e.g., the long segments may have durations that are based on the time of the excluded portions that are removed when the video data is filtered). The short segments may be defined by a predefined duration, such as about four seconds. In some examples, the plurality of short segments include video segments that overlap with at least one other short segment. For instance, the short segments (e.g., the short segments associated with a long segment (e.g., a single long segment)) may overlap with one another.

[0009] The method may include applying multiple trained machine learning models. For instance, the method may include applying, for each short segment of the plurality of short segments, a first trained machine learning model to the facial landmark data of the short segment to generate a movement indicator for each of the plurality of short segments. The first trained machine learning model may comprise a plurality of first trained machine learning models, wherein each of the first trained machine learning models generates a movement indicator for each of the plurality of short segments corresponding to a different category of facial movement. Each movement indicator may indicate a probability that movement of a category of facial movement occurred within the short segment. The categories of facial movement may include one or more of eye movement, mouth movement, and / or head movement (e.g., which may include any head movement, including movements based on torso tremor and / or head tilting).

[0010] The method may include determining a total severity score based on the movement indicators associated with each of the categories of facial movement. In some examples, the total severity score is determined by applying the movement indicators associated with each category of facial movement. The total severity score may correspond to the risk of the patient receiving a positive TD diagnosis by a trained physician. The total severity score maycorrespond to an Abnormal Involuntary Movement Scale (AIMS) score (e.g., however, in some examples, the total severity score is not the same as the AIMS score, although it is calculated using the same scale and indicative of the risk of the patient receiving a positive TD diagnosis by a trained physician).

[0011] In some examples, the method may include determining a movement severity score for a long segment (e.g., each of the plurality of long segments) based on the movement indicators of the short segments associated with the long segment. In some examples, the method may include applying, for each long segment of the plurality of long segments, a trained machine learning model (e.g., a second machine learning model) to the movement indicators of the short segments associated with the long segment to generate a movement severity score for each of the plurality of long segments. In such examples, the third machine learning model may generate the total severity score based on the movements severity indicators associated with each long segment of the plurality of long segments. However, it should be appreciated that the inclusion of this second machine learning model is not necessary. Nonetheless, in examples where it is used, each movement severity score may indicate a severity of movement (e.g., involuntary movement) of the category of facial movement within the long segment.

[0012] The scale associated with the plurality of movement indicators may be different from the scale associated with the total severity score. For instance, the scale associated with the plurality of movement indicators may be a range from 0.0- 1.0 or a binary scale. The scale associated with the total severity score may be a range of 0-28. Further, in relevant examples, the scale associated with the plurality of movement indicators may be different from the scale associated with the plurality of movement severity scores. Further, and for example, the scale associated with the plurality of movement severity scores may be in a range of 0.0-10.0. As such, in some instances, the scale associated with the total severity score, the scale associated with the plurality of movement indicators, and / or the scale associated with the plurality of movement severity scores may be different from one another.

[0013] The method may include generating a notification indicating the total severity score. For example, the method may include displaying the notification on a display device of the patient and / or a physician. In some examples, the method may include generating a notification that includes a recommendation to refer the patient to a trained clinician when the total severity score is above a predetermined threshold.

[0014] The method may include providing any combination of the movement indicators, the movement severity scores, and / or the total severity score to a trained physician via a dashboard of a display device. For examples, the method may include generating, via a graphical user interface (GUI) displayed on a display device, an indication of the movement severity scores for each category of facial movement associated with each of the plurality of long segments. In some examples, the indication of the movement severity scores for each category of facial movement may be color coded based on a value of each respective movement severity score (e.g., red to indicate scores in a highest range, yellow to indicates scores in an intermediate range, and green to indicates scores in a lowest range). In some instances, the method may include displaying, in response to a user selection of a long segment of the plurality of long segments, the video data associated with the selected long segment via the display device. For instance, the method may display, with the video data associated with the selected long segment, a point cloud associated with the facial landmark data associated with the selected long segment (e.g., the point cloud throughout the long segment that was used to determine the movement severity score for the long segment). For instance, in some examples, the method may include displaying, with the video data associated with the selected long segment, a graph that indicates the respective movement severity score as well as the progression of respective movement indicators for the short segments that comprise the selected long segment. Finally, in some examples, the method may include displaying a graph that illustrates changes in each of the movement severity scores associated with each category of facial movement of the selected long segment throughout a duration of the selected long segment (e.g., via a graph that indicates the severity score throughout the duration of the selected long segment).

[0015] In some examples, the plurality of first trained machine learning models is trained using training data comprising facial landmark data extracted from videos from a plurality of patients with a positive TD diagnosis, AIMS scores associated with each patient of the plurality of patients, and / or labels corresponding to each of the videos indicating a presence of a plurality of distinct categories of facial movement. The videos may be divided into a plurality of short segments having a predefined duration, wherein the labels are applied to each of the plurality of short segments. The categories of facial movement comprise one or more of mouth and lower face movement, upper face movement, head tremor, other tremor, head movement, tongue movement, other movement, contractures and tilts, and number of blinks.

[0016] In some examples, the method may include treating the patient based on the total severity score and / or the movement severity scores. For instance, the method may include administering to the subject a first daily amount of a medication associated with TD subsequent to receiving the video data. The method may include increasing the daily amount of the medication to a first subsequent daily amount based on the total severity score of the patient being above a threshold. The medication may include a vesicular monoamine transporter 2 (VMAT2) inhibitor. For examples, the VMAT2 inhibitor may include any combination of deutetrabenazine, tetrabenazine, and / or valbenazine. In some examples, the first daily amount is pursuant to a titration schedule associated with the VMAT2 inhibitor. For instance, the first daily amount may be at least about 6 mg / day (e.g., and the medication may be deutetrabenazine), and the first subsequent daily amount may be at least about 6 mg / day more than the first daily amount.BRIEF DESCRIPTION OF THE DRAWINGS

[0017] FIG. 1A illustrates a schematic diagram of an example system environment.

[0018] FIG. IB illustrates a block diagram showing an example of one or more portions of the example system environment of FIG. 1 A.

[0019] FIG. 2 is a flowchart that illustrates an example performance of assessing a risk of a patient for tardive dyskinesia.

[0020] FIG. 3 illustrates an example interface displaying an example prompt to a patient.

[0021] FIG. 4 is a block diagram of an example computing device.

[0022] FIG. 5 is a diagram of an example facial mask that includes one or more landmarks associated with a patient.

[0023] FIG. 6 is a diagram of an example graphical user interface (GUI) that indicates a plurality of patients, a total severity score for each patient, and an indication of a trend associated with the total severity score for each patient.

[0024] FIG. 7 is a diagram of an example GUI that indicates a trend of total severity scores for a patient, an indication of the movement severity score for different categories of facial movement across a plurality of long segments associated with the patient, and a video and point cloud associated with the video data of the patient.DETAILED DESCRIPTION

[0025] The following discussion omits or only briefly describes conventional features of treatment or diagnostic systems, which are apparent to those skilled in the art. It is noted that various embodiments are described in detail with reference to the drawings, in which like reference numerals represent like parts and assemblies throughout the several views. Reference to various embodiments does not limit the scope of the claims attached hereto. Additionally, any examples set forth in this specification are intended to be non-limiting and merely set forth some of the many possible embodiments for the appended claims. Further, particular features described herein can be used in combination with other described features in each of the various possible combinations and permutations.

[0026] Unless otherwise specifically defined herein, all terms are to be given their broadest possible interpretation including meanings implied from the specification as well as meanings understood by those skilled in the art and / or as defined in dictionaries, treatises, etc. It must also be noted that, as used in the specification and the appended claims, the singular forms “a,” “an” and “the” include plural referents unless otherwise specified, and that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof.

[0027] For the diagnosis and assessment of Tardive Dyskinesia (TD), the Abnormal Involuntary Movement Scale (AIMS) is the primary tool utilized by clinicians. The AIMS is a test comprising a combination of items rated on a 5-point anchored scale, including orofacial movements, extremity movements, and truncal movements, as well as global severity of movements as judged by the examiner, in addition to yes-no questions (e.g., graded on a binary scale) concerning dental health. To assess the severity of each of the items assessed on the anchored scale, the clinician instructs the patient to undergo a series of specific activities, such as instructing the patient to protrude their tongue, flex and extend the patient’s arms, etc. The 5- point anchored scale is utilized to indicate a severity of involuntary movements of the patient’s body. For instance, a movement severity score of zero indicates no involuntary movement, a movement severity score of one indicates minimal or normal involuntary movement, a movement severity score of two indicates mild involuntary movement, a movement severityscore of three indicates moderate involuntary movement, and a movement severity score of four indicates severe involuntary movement.

[0028] The AIMS can only be performed by appropriately trained clinicians. TD may generally be underdiagnosed or improperly diagnosed, given that TD develops in patients prescribed certain psychiatric drugs for an extended period from clinicians such as psychiatrists or other medical professionals who often are not certified to administer the AIMS. As such, access to an accurate TD diagnosis can be difficult. Even when TD has been diagnosed by a certified clinician, the complexity of the AIMS test can prevent TD patients from obtaining periodic AIMS testing to track and document treatment progress, which can lead to frustration in treatment and subsequent treatment abandonment. Also, as the AIMS involves implicitly subjective determinations, AIMS scoring can differ even between trained clinicians.

[0029] As a result, there is a need for a diagnostic tool for easily and uniformly assessing TD risks and symptoms in patients.

[0030] FIG. 1A illustrates a schematic diagram of a system environment (or “environment”) 100, for implementing a TD diagnostic system as described herein. As illustrated, the environment 100 includes one or more server device(s) 102 connected to one or more of a device 106 (e.g., a clinician device), a device 108 (e.g., a patient device), and a database 110 via a network 111. FIG. IB illustrates a block diagram showing an example of one or more portions of the example system environment 100 of FIG. 1 A.

[0031] As shown in FIG. 1 A, the server device(s) 102, the device 106, the device 108, and the database 110 may communicate with each other via the network 111. The network 111 may comprise any suitable network over which computing devices can communicate. The network 111 may include a wired and / or wireless communication network. Example wireless communication networks may be comprised of one or more types of radio frequency (RF) communication signals using one or more wireless communication protocols, such as a cellular communication protocol, a wireless local area network (WLAN) or WIFI communication protocol, and / or another wireless communication protocol. Though FIG. 1 A illustrates the components of environment 100 communicating via the network 111, it will be appreciated that the components of environment 100 may communicate directly with each other, for example, bypassing the network 111. For example, the device 108 may communicate directly with the server device(s) 102.

[0032] The server device(s) 102 may generate, receive, analyze, store, and / or transmit digital data, such as video data of a patient. In one or more cases, the server device(s) 102 may communicate with the device 106 and the device 108. For example, the server device(s) 102 may send data to the device 106, including video data or other information, such as movement indicators, movement severity scores and / or a total severity score. In another example, the server device(s) 102 may receive input from the user (e.g., the patient) via device 108 or the clinician via device 106. In some cases, the server device(s) 102 may include a distributed collection of servers, in which the server device(s) 102 include a number of server devices distributed across the network 111. In some examples, the distributed server device(s) 102 may be located in the same location or at different physical locations. In other cases, the server device(s) 102 may comprise a content server, an application server, a communication server, a web-hosting server, or another type of server.

[0033] In one or more cases, the server device(s) 102, device 106, and / or device 108 may include an assessment system 104 or portions thereof. The assessment system 104 may analyze video data of a patient and identify landmarks (e.g., facial landmarks) of the patient. Further, the assessment system 104 may apply one or more trained machine learning models, such as machine learning models 113a-l 13c, 114a-l 14c, to the video data and identified landmarks to determine a spatial arrangement and temporal shift of the identified landmarks. In some cases, the assessment system 104 may identify the landmarks prior to applying the trained machine learning model to the identified landmarks. For instance, the assessment system 104 may utilize one or more image processing techniques to identify one or more landmarks in the video data (e.g., Haar Cascade technique, integral projection technique, Fisher face technique, elastic graph matching technique, radial basis function network recognition technique, hidden Markov model technique, and other like techniques for facial recognition).

[0034] The assessment system 104 may apply the trained machine learning model(s) to video data (e.g., corresponding to one video of the patient or a series of videos of the patient) to determine movement of landmark positions over a period of time. Having determined the spatial arrangement and temporal shift of the identified landmarks, the trained machine learning model(s), such as machine learning models 113a-l 13c, 114a- 114c, may determine one or more movement indicators and / or one or more movement severity scores (e.g., which may indicate a severity of involuntary movements of a patient’s body). Each movement indicator may beassociated with a distinct (e.g., unique) category of movement of the patient, such as eye movement, mouth movement, and / or head movement (e.g., which may include any head movement). In examples where movement severity scores are determined and used, each movement severity score may be associated with a distinct (e.g., unique) movement of the patient, such as eye movement, mouth movement, and / or head movement (e.g., which may include any head movement), and may be associated with and based on a plurality of movement indicators. In some cases, based on at least the movement indicators associated with the categories of facial movement (e.g., and / or in some examples, based on movement severity scores), the assessment system 104 generates a total severity score utilizing a machine learning model 123. The total severity score may indicate a risk or severity of the analyzed movement disorder. In one embodiment, the total severity score may correspond to an AIMS score.

[0035] The assessment system 104, or one or more portions thereof, residing on the device 106 or device 108 may allow for on-device analysis of assessing a risk of the patient for a movement disorder, such as, but not limited to, TD. The assessment system 104, or one or more portions thereof, residing on either the device 106 or device 108 may allow the respective device, such as device 106 or device 108, to assess the risk of a patient for TD or severity of TD for a patient. In one or more cases, the assessment system 104, or one or more portions thereof, may reside on the server device(s) 102, such that the device 106 and / or device 108 may offload certain risk analysis tasks for being performed at the server device(s) 102 and / or request information from the server device(s) 102 to enable the device 106 and / or device 108 to perform as described herein.

[0036] In one or more cases, the assessment system 104 operates on a central server, such as the server device(s) 102, and may be utilized by one or more computing electronic devices, such as devices 106 and 108, via an application downloaded from the central server or a third-party application store and executed on the one or more computing electronic devices. In one or more cases, the assessment system 104 may be a software- based program, downloaded from a central server, such as the server device(s) 102, and installed on one or more computing electronic devices, such as devices 106 and 108. In one or more cases, the assessment system 104 may be utilized as a software service provided by a third-party cloud service provider (not shown). In one or more cases, the assessment system 104 may be preinstalled, as software and / or firmware, on the one or more computing electronic devices. In one or more cases, the assessment system104 may be installed onto the one or more computing electronic devices via an external storage device, such as a universal serial bus (USB) flash drive

[0037] In one or more cases, the device 108 is an electronic computing device, such as a desktop computer, a laptop computer, a tablet computer, a personal digital assistant (PDA), a smart phone, a thin client, or any other electronic device or computing system capable of communicating with the server device(s) 102 through the network 111. The device 108 may be a client to the server device(s) 102. The device 108 can be configured with a camera (e.g., camera 304 illustrated in FIG. 3) or another type of imaging device suitable for acquiring images and video, such as, for example, video of a patient performing an action in response to a prompt from the assessment system 104. In other cases, the device 108 can be any suitable type of mobile device capable of running mobile applications, including smart phones, tablets, slate, or any type of device that runs a mobile operating system. For example, the device 108 may be a mobile device operated by a user, such as a patient, and capable of connecting to a network, such as the network 111, to transmit captured video of the patient. In yet other cases, the device 108 can be any wearable electronic device, such as a head-mounted display, a smartwatch, or the like that is capable of sending, receiving, and processing data.

[0038] In one or more cases, the device 108 can include a user interface for providing an end user with the capability to interact with the assessment system 104. A user interface refers to the information (such as graphics, text, and sound) the assessment system 104 presents to a user and the control sequences the user employs to control the assessment system 104 and respond to prompts generated by the assessment system 104. A user interface can be, for example, a keyboard that allows a user to input text, a camera that can recognize user gestures and / or objects, a touchscreen that accepts input from a user via touch of a body part and / or a stylus, or the like. A user may access the assessment system 104 through the user interface to enable the assessment system 104 to operate on the user's device, such as device 108.

[0039] In one or more cases, the device 106 includes one or more of the same or similar features as discussed with respect to the device 108. Accordingly, a description of such features is not repeated. In other embodiments, the device 106 can comprise a system including a digital camera or other motion recording device in communication with a separate electronic computing device.

[0040] FIG. IB illustrates a block diagram showing an example of one or more portions of the example system environment 100 of FIG. 1 A. The database 110 is configured to store information such as captured video data 112, facial landmark data 118, movement indicator data 121 (e.g., movement indicators), movement severity data 122 (e.g., movement severity scores), total severity data 124 (e.g, total severity scores), acoustic severity data 127 (e.g, acousticspecific total severity score), patient information, and / or the like. Further, the database 110 may be configured to store trained machine learning models, such as, but not limited to machine learning models 113a-l 13c, 114a- 114c, 123, 125. It is noted that the database 110 may store information for more than one patient. It is also noted that the information stored within the database 110 may be tagged with an identifier, that correlates each piece of data associated with the patient. Further, information stored within the database 110 may be stored as anonymized data. As described herein, the information stored within the database 110 may be localized on the database 110, on the server device(s) 102, or distributed on the database 110, server device(s) 102, and other data storage repositories and servers. Additionally, although FIG. IB illustrates machine learning models 113a-113c, 114a-114c, 123, 125 as being localized on database 110, it should be understood that machine learning models 113a-l 13c, 114a- 114c, 123, 125 may also be localized on assessment system 104 or distributed across database 110 and assessment system 104. Further, in some examples, the machine learning models 114a-l 14c and / or 125 may be omitted.

[0041] In one or more cases, the assessment system 104 includes a data preprocessor 116, the features of which may be implemented in hardware and / or software. The data preprocessor 116 may be configured to receive raw signals, such as video data 112, and to process the raw signals to remove invalid data, anomalies, and the like. For example, the data preprocessor 116 may be configured to receive raw signals representing video data 112 captured by camera, such as a camera of the device 106 or 108. The data preprocessor 116 may be configured to process the raw video data 112 to determine facial landmark data 118, which, in some examples, may be used to identify landmarks of a patient. In one or more cases, the data preprocessor 116 may process the raw signals in an online mode, i.e., processing the raw signals as the signals are received, for example, from the camera of the device 108, using one or more data buffers. In one or more other cases, the data preprocessor 116 may process the raw signals in an offline mode, i.e., processing raw signals retrieved from a data storage repository, such as the database 110.For the cases in which the server device(s) 102 utilizes a distributed computing system, the data preprocessor 116 utilizes the distributed computing components of the server device(s) 102 to process the raw signals in parallel.

[0042] In one or more cases, the assessment system 104 includes an assessment engine 120, which may be implemented in hardware. In one or more other cases, the assessment engine 120 may be implemented as an executable program maintained in a tangible, non-transitory memory, such as memory (e.g., memory 404 of FIG. 4), which may be executed by one or processors (e.g., such as processor 402 of FIG. 4). As further described herein, the assessment engine 120 may be configured to implement the machine learning models 113a-l 13c, 114a-l 14c, 123, 125 to generate movement indicator data 121, movement severity data 122, total severity data 124, and / or acoustic-specific total severity data 127 and to determine a risk of the patient for TD or severity of TD symptoms.

[0043] The first trained machine learning models 113 a- 113c may each be configured to determine whether a particular category of facial movement associated with TD (e.g., eye movement, mouth movement, and / or head movement (e.g., which may include any head movement, including movements based on torso tremor and / or head tilting)) is present within the video data. For example, the first trained machine learning models 113a-l 13c may be configured to determine one or more movement indicators. As described herein, movement indicators may indicate a probability that movement of a category of facial movement occurred within a video segment (e.g., such as a short segment of video data). In some examples, the movement indicator may be in the scale of 0.0- 1.0, however other scales may be used. For example, the movement indicator may be binary (e.g., either one or zero).

[0044] Each of the first trained machine learning models 113a-l 13c may determine whether a corresponding category of facial movement associated with TD is present within the video data based on facial landmark data 118. As such, the first trained machine learning models 113a-l 13c can produce a movement indicator (e.g., such as a numerical indication) that indicates whether a particular category of facial movement associated with TD is present within the video segment (e.g., the short segment). In some examples, the assessment system 104 may accordingly apply the trained machine learning models 113a-113c to a reduced amount of data, for example, the facial landmark data 118 that is relevant to the category of facial movement corresponding to a particular algorithm 113a-113c.

[0045] In some examples, the assessment engine 120 may apply each of the second machine learning models 114a-l 14c to the movement indicators produced by a corresponding one of the first machine learning models 113a-l 13c to generate movement severity data for each of the long segments of video data. The movement severity data (e.g., a movement severity score) may indicate a severity of movement for a particular category of facial movement associated with TD for each long segment. The movement severity score may be scored on a scale that indicates a severity of involuntary movements of the patient’s body. For instance, a movement severity score of a lower number may indicate less involuntary movement than a movement severity score of a higher number. The movement severity score may be in the scale of 0.0 to 10.0, but any scales may be utilized (0-1, 0-10, 0-100, etc.). However, it should be appreciated that in some examples, the assessment engine 120 may not use the second machine learning models 114a-l 14b and / or the assessment engine 120 may use the second machine learning models 114a- 114b but the movement severity scores generated by these models may not be used when determining a total severity score for the patient. In other embodiments, movement severity scores may comprise an average (e.g., a scaled-up average) or other numerical generalization of the corresponding movement indicators for a particular long segment of video data.

[0046] The assessment engine 120 may determine a total severity score for the patient. The total severity score may indicate a risk or severity of involuntary movements associated with TD. In one embodiment, the total severity score may correspond to an AIMS score. The assessment engine 120 may determine the total severity score using the third machine learning model based on the movement indicators generated by the plurality of first machine learning models 11 Sal l 3c. However, in some examples, the assessment engine 120 may determine the total severity score based on the movement severity scores generated by the plurality of second machine learning models 114a- 114c.

[0047] The video data may comprise segments that are not useful for assessing involuntary movement according to the methods described in the current disclosure (e.g., portions where the patient is not looking at the camera, where facial landmark data of the patient is not identified within the video data, and / or where the patient is not performing a predefined task). As described in more detail below, the assessment system 104 may segment the video data into a plurality of long segments in order to isolate the useful segments for involuntary movement assessment. For example, the method may include filtering the video data to remove portions ofthe video data that are not needed to generate the long segments. The plurality of long segments may be defined by the durations of the video data between the excluded portions of the video data.

[0048] In one embodiment, the video data can be filtered manually by a trained professional to generate the long segments. Alternatively, the assessment system 104 may filter (e.g., automatically filter) the video data based on an analysis of facial landmarks 117, for example, using one or more facial recognition models. For instance, as described in more detail below, the assessment system 104 may include a detector to identify a face within each frame of the video data, a shape predictor to identify facial landmarks 117 within the video data (e.g., to precisely localize the face), and / or a face recognition model (e.g., such as dlib facial recognition). The assessment system 104 may generate the long segments based on the facial landmarks 117 of the video data. The assessment system 104 may identify portions of the video data 112 where the patient has not performed the prompt based on the facial landmarks 117 and filter the video data 112 based on these portions where the patient has not performed the prompt. That is, the assessment system 104 may remove, or filter out, the portion(s) of video data 112 where the patient is not performing the prompt(s), such that the remaining segments of the video data 112 are the long segments of the video data 112.

[0049] For example, the assessment system 104 may be preprogrammed to associate one or more the identified facial landmarks 117 with each prompt. By associating a particular subset of the identified facial landmarks 117 with a prompt, the assessment system 104 may analyze those facial landmarks 117 to determine whether the performed action corresponds to the prompt. The assessment system 104 may determine that the performed action does not correspond to the prompt based on the movement in the spatial arrangement of the corresponding facial landmarks 117. For example, the assessment system 104 may determine that none of the facial landmarks 117 associated with the patient’s cheeks, jaw, lips, and / or tongue move (i.e., an indication that the patient did not open the patient’s mouth) in the video data 112 (e.g., a particular duration of the video data 112). In other cases, the assessment system 104 may determine that the performed action does not correspond to the prompt if a positional displacement of spatial arrangement of the corresponding facial landmarks 117 does not exceed a threshold value (e.g., an indication that the patient did not open the patient’s mouth enough to analyze the landmarks associated withthe patient’s tongue). As such, the assessment system 104 may determine that the performed action does not correspond to the provided prompt based on the facial landmarks 117.

[0050] Additionally, the assessment system 104 may remove, or filter out, the portion(s) of video data 112 where each, or alternatively a threshold amount, of the facial landmarks 117 are not present within the video data. Such portions of the video data 112 may represent portions of the assessment where the patient is not adequately facing the camera, and thus the facial landmarks 117 cannot be properly assessed.

[0051] The assessment system 104 may segment each of the plurality of long segments into a plurality of short segments. The long segments may be defined by multiple different durations. For example, the long segments may be aperiodic and / or the long segments may be defined by varying or non-consistent durations (e.g., the long segments may have durations that are based on the time of the excluded portions that are removed when the video data is filtered). The short segments may be defined by a predefined duration, such as about four seconds. In some examples, the plurality of short segments includes video segments that overlap with at least one other short segment. For instance, the short segments (e.g., the short segments associated with a long segment (e.g., a single long segment)) may overlap with one another. The overlap of the short segments may help to ensure that a movement (e.g., an involuntary movement indicative of TD) is substantially captured by at least one of the short segments (e.g., even if it occurs at the very beginning or end of a short segment). If the short segments did not overlap (i.e., were simply sequential), an instance of involuntary movement may be split between multiple short segments. Additionally, a short segment duration of about four segments may provide the benefit of capturing an entire instance of involuntary movement (i.e., whereas shorter segments may only capture portions of involuntary movement instances), while being short enough to provide sufficient resolution within the overall length of video data.

[0052] The assessment system 104, such as the assessment engine 120, may apply a first trained machine learning model 113 a- 113c to facial landmark data within each the short segments to generate movement indicators for one or more categories of facial movement. In some examples, a single first machine learning model is used, while in other examples, a plurality of first machine learning models 113a-l 13c are used. When multiple first trained machine learning models 113a-l 13c are used, each of the first machine learning models 113a- 113c may be unique to a specific category of facial movement (e.g., model 113a for analysis ofeye movement, model 113b for analysis of mouth movement, model 113c for analysis of head movement, etc.).

[0053] In some examples, the assessment system 104, such as the assessment engine 120, may apply a second trained machine learning model 114a-l 14c to the movement indicators associated with each of the short segments associated with a long segment to determine a movement severity score for the long segment. For example, a long segment may be associated with (i.e., comprised of) multiple short segments, and the assessment system 104 may apply the second trained machine learning model 114a-l 14c to the movement indicators associated with the short segments of the long segment to determine a movement severity score for each of the long segments. In some examples, a single second machine learning model is used, while in other examples, a plurality of second machine learning models 114a-l 14c are used. When multiple second trained machine learning models 114a-l 14c are used, each of the second machine learning models 114a-l 14c may be unique to a specific category of facial movement (e.g., model 114a for analysis of eye movement, model 114b for analysis of mouth movement, etc.), and each of the second trained machine learning models 114a-l 14c may be applied to the movement indicators produced by the corresponding one of the first trained machine learning models 11 Sal l 3c. Each movement severity score may indicate a severity of movement (e.g., involuntary movement) of a particular category of facial movement within the long segment. However, as noted herein, in some examples, the assessment system 104 may not use the second machine learning models 114a-l 14b and / or the assessment system 104 may use the second machine learning models 114a-l 14b but the movement severity scores generated by these models may not be used when determining a total severity score for the patient. In other embodiments, movement severity scores may comprise an average (e.g., a scaled-up average) or other numerical generalization of the corresponding movement indicators for a particular long segment of video data.

[0054] Finally, the assessment system 104 may determine a total severity score based on the movement indicators associated with each of the categories of facial movement. The total severity score may correspond to the risk of the patient receiving a positive TD diagnosis by a trained physician. Alternatively, the total severity score may correspond to the severity of the patient’s involuntary movement associated with TD. For example, the total severity score may correspond to an Abnormal Involuntary Movement Scale (AIMS) score (e.g., however, in someexamples, the total severity score is not the same as the AIMS score, although it is calculated using the same scale and indicative of the risk of the patient receiving a positive TD diagnosis by a trained physician). In one embodiment, the assessment system 104 may apply the third machine learning model 123 to each of the movement indicators determined by each of the first trained machine learning models 113a-l 13c to determine the total severity score.

[0055] However, in other examples, the assessment system 104 may determine a total severity score based on the movement severity scores associated with the plurality of long segments that are associated with the video data of a patient. In such situations, and in some examples, the assessment system 104 may apply the third machine learning model 123 to each of the movement severity scores determined by each of the second trained machine learning models 114a-l 14c to determine the total severity score.

[0056] There may be benefits to using multiple layers of machine learning models to determine the total severity score in the assessment system 104. For example, the use of multiple layers of machine learning models may simplify the training procedure and / or increase the accuracy of the model to determine a total severity score. That is, it may be more difficult or complex and / or less accurate to train a single machine learning model that can accurately determine a total severity score directly from landmark data than it is to use multiple layers of machine learning models, as described herein. For instance, it may be beneficial to train the first machine learning model(s) to determine movement indicators based on landmark data (e.g., where the landmark data is determined using one or more facial recognition techniques) and train the third machine learning model to determine a total severity score (e.g., based on the movement indicators determined by the first machine learning model). Further, in some examples, it might provide further benefit to train the second machine learning model(s) to determine movement severity scores (e.g., based on the movement indicators determined from the first machine learning model), and in some examples, then train the third machine learning model to determine the total severity score (e.g., based on the movement severity scores determined by the second machine learning model).

[0057] For similar reasons, the use of multiple machine learning models for each layer of analysis (i.e., multiple first machine learning models and / or multiple second machine learning models), where each model within each layer corresponds to a particular category of facial movement may be beneficial. For example, it might be beneficial to train a plurality of firstmachine learning models to determine the movement indicators based on the landmark data where, for example, each of the first machine learning models are trained to generate a movement indicator for a specific category of facial movement. Further, it might be beneficial to train a plurality of second machine learning models to determine the movement severity scores based on the movement indicators where, for example, each of the second machine learning models are trained to generate a movement severity score for a specific category of facial movement based on the movement indicator(s) associated with that same category of facial movement.

[0058] In some examples, the assessment system 104 may receive one or more voice recordings of the patient, generate acoustic data 109 based on the voice recordings, and generate an acoustic-specific total severity score based on the acoustic data. The acoustic data 109 may be characterized by voice-based digital biomarkers of the patient. The assessment system 104 may be configured to analyze the acoustic data 109 (e.g., voice-based digital biomarkers) of the patient. The assessment system 104 may determine an acoustic-specific total severity score based on the acoustic data 109 of the patient, refine the total severity score based on the acoustic data 109 of the patient, and / or determine the progression of TD based on the acoustic data 109 of the patient.

[0059] The assessment system 104 may receive the acoustic data 109 (e.g., a voice recording) of the patient. For example, the patient may be instructed to produce acoustic data 109 in response to a prompt to speak freely and / or in response to a prompt that requests that the patient follow a specific script. The acoustic data 109 may have a predetermined duration (e.g., 1 to 5 minutes). The prompt may include one or more of text and / or an image (e.g., similar to the prompts presented at 202). Similar to the video data above, the acoustic data 109 may be filtered (e.g., segmented) into discrete segments, for example discrete segments of a predetermined duration. The predetermined duration may be different from the durations of the long segments or the short segments. For example, the segments of the acoustic data 109 may be less than the short segments of the video data. In some examples, the segments of the acoustic data 109 may be between about 50-100 ms.

[0060] The acoustic data 109 may be processed (e.g., by data preprocessor 116) to extract acoustic features from the acoustic data 109, such as an amplitude envelope, root mean square deviation, zero crossing count, fundamental frequency, mel-frequency cepstral coefficients(MFCCs), delta, delta-delta , and / or chroma features from the acoustic data 109. Such acoustic features may thus comprise at least a part of the acoustic data 109.

[0061] Amplitude envelope may include an indication of how sound volume (e.g., measured in local, dB) changes over time. For example, the amplitude envelope may include an indication of a maximum amplitude and a minimum amplitude. A root-mean square (RMS) may include a squaring of an amplitude (e.g., instantaneous amplitude) for a (e.g., each) point in time. The points in time may be equally spaced.

[0062] The RMS may be calculated by determining the average of each amplitude value (e.g., for each point in time) and taking the square root of the average. RMS deviation may be the difference between the RMS (e.g., total RMS) and the RMS for any one amplitude value (e.g., for one point in time).

[0063] Zero crossing may include a point where a value crosses a zero level axis. For example, the amplitude (e.g., measured in local, dB) value may cross the zero level axis at the zero crossing.

[0064] A fundamental frequency may include the lowest frequency (e.g., measured in Hz) of a periodic waveform.

[0065] The mel-frequency cepstral coefficients (MFCCs) may include an indication of a shortterm power spectrum of a sound, for example a coefficient of the short-term power spectrum of a sound.

[0066] A chroma feature may include an indication of tonal content of a sound. For example, the chroma feature may include a profile of one or more pitch classes (e.g., of twelve pitch classes). The chroma feature may include harmonic and / or melodic characteristics of a sound.

[0067] The assessment system 104 may use a machine learning model to determine an acoustic-specific total severity score based on the acoustic data 109 of a patient. The acousticspecific total severity score determined by the assessment system 104 using the acoustic data 109 may be different from the total severity score determined using the video data. As such, the assessment system 104 may determine multiple total severity scores, such as a first total severity score based on the video data 112 and a second total severity score based on the acoustic data 109. Alternatively or additionally, the assessment system 104 may refine, based on the acoustic data 109, the total severity score that was determined based on the video data 112. Alternatively or additionally, the assessment system 104 may determine the progression of TD based on theacoustic data 109 of the patient. The machine learning model may determine a digital biomarker of the acoustic data 109. The digital biomarker may correspond to TD severity of the patient using the acoustic data 109. The digital biomarker may be for diagnostic purposes. Further, in some examples, the assessment system 104 may include an assessment model 125 that is specifically trained to analyze the acoustic data 109 of the patient and generate a total severity score, refine a total severity score, and / or determine the progression of TD based on the acoustic data 109 of the patient.

[0068] FIG. 2 is a flowchart that illustrates an example procedure 200 of assessing a risk of a patient for TD and / or a severity of TD in a patient. An assessment system, such as the assessment system 104, may perform one or more steps of the procedure 200. Although described primarily in the context of the procedure being performed by the assessment system 104, in some examples, one or more of the device 106 and / or the device 108 may perform any combination of the steps of the procedure 200 (e.g., in combination with the assessment system 104). The risk of TD may be assessed in a variety of locations. For instance, the patient may perform the assessment at home (e.g., via device 108). In another instance, the patient may perform the assessment while at a clinician’s office (e.g., via device 106). In one or more cases, to access the assessment system 104, the patient may login to an assessment portal, via, for example, but not limited to a telehealth application. It is noted that in the example provided herein, the patient may perform the assessment via device 108. However, it should be understood that one, more, or all of the processes discussed herein for performing the assessment via device 108 may also be implemented via device 106.

[0069] The assessment system 104 may generate one or more preliminary prompts to the respective device, for example, device 108, to prepare the patient to begin the assessment (201). The device 108 may display the preliminary prompts on a user interface of the device 108 (e.g., via the user interface 302). For example, one prompt may instruct the patient to remove any objects (e.g., gum or candy) from the patient’s mouth. In another example, another prompt may instruct the patient to begin the assessment while sitting in a chair that is hard, firm, and one without arms. In another example, another prompt may instruct the patient to remove their shoes and socks. In yet another example, another prompt may instruct the patient to position the device 108 such that the device 108 is in a particular position and a camera 304 of the device 108 can capture video data of the patient. For example, the prompt may specify that the camera 304 ofthe device 108 should be placed at a specific distance from the patient and / or a specific height relative to the patient. Additionally, the prompt may specify that the camera 304 of the device 108 should be in placed in a stable, stationary position. One example of a prompt may be a notification that is generated via graphical user interface (GUI) on a display device of one or more devices described herein, such as a notification provided via a display of the device 106 and / or the device 108.

[0070] A prompt to perform an action is generated (202), preferably by the assessment system 104. For example, the assessment system 104 may provide a prompt to the device 108, which displays the prompt on the user interface. For example, the displayed prompt 306 may instruct the patient to “Please open your mouth for 30 seconds”, for example, as illustrated in FIG. 3. In one or more cases, the user interface 302 may display a timer (e.g., the timer 307) that displays a time corresponding to the displayed prompt (e.g., the displayed prompt 306). When capturing video data of the patient, the timer may provide an indication of the length of time remaining to complete the prompted instruction. For example, when the camera begins capturing video of the patient with an open mouth, the timer may countdown from a predefined time period, such as 30 seconds.

[0071] The assessment system 104 may generate a prompt to perform a scripted action. A prompt for a scripted action may include, for example, but not limited to, an instruction for the patient to perform a series of actions. For example, a prompt may instruct the patient to open the patient’s mouth for a first period of time, briefly close the patient’s mouth at the expiration of the first period of time, and subsequently open the patient’s mouth for a second period of time. In another example, a prompt may instruct the patient to open the patient’s mouth and protrude the patient’s tongue twice. In other examples, a prompt for a scripted action may include, for example, but not limited to, an instruction for the patient to perform a singular activity.

[0072] In addition to or in the alternative to providing a prompt to perform a scripted action, the assessment system 104 may generate a prompt to perform an unscripted action. A prompt for an unscripted action may include, for example, but not limited to, an instruction for the patient to stare at the display screen of the device 108 for a certain period of time.

[0073] Further, in other examples, the prompt may be for the patient to stay still (e.g., not move) for a duration of time, for instance, while the patient maintains their head, face, and / or mouth in a particular orientation.

[0074] The assessment system 104 may determine whether the prompts for the assessment session are complete. For example, an assessment session for certain movement disorders may include a patient responding to a series of prompts to perform respective actions. In response to the patient performing the action corresponding to the provided prompts, the assessment system 104 may determine whether the prompts for the assessment session are complete. For the cases in which it is determined that the prompts for the assessment session are not complete, the assessment system 104 provides another prompt to perform an action, as described with respect to 202.

[0075] It should be appreciated that, in some examples, any combination of 201 and / or 202 may be omitted, and the procedure 200 may begin at 204.

[0076] Video data captured during an assessment session is received (204), such as by the assessment system 104. The assessment system 104 may receive the video data from the camera that captures the video data (e.g., directly or indirectly, such as by retrieving the video data from memory, such as memory that resides on one or more servers 102). In some examples, the assessment system 104 may capture video data during the assessment session via a camera of the device 108 (e.g., camera 304). For instance, during the assessment session, the camera may record video of the patient at rest or performing a bodily movement in response to the prompt displayed on the user interface (e.g., the user interface 302). The device 108 may provide the captured video as video data to the assessment system 104. In one or more cases, the device 108 may begin recording video when the assessment session commences. In one or more other cases, the device 108 may begin recording video in response to receiving a prompt from the assessment system 104. In one or more cases, the assessment system 104 may store the video data (e.g., video data 112) in the database 110. In some examples, the assessment system 104 may receive (e.g., retrieve) the video data from memory (e.g., memory of the database 110, the device 106, the device 108, and / or the server device 102).

[0077] The video data can be captured by a camera, such as a camera of device 108, that is focused on a particular region of the patient’s body. For example, the camera by be configured to capture video data of the patient’s facial region. As such, the video data may correspond to (e.g., include video segments that include) facial movements of the patient. For instance, the video data may include video frames of the patient’s head and / or face, and the video frames may include facial movements of the patient. The video data may correspond to facial movements ofthe patient during an assessment session (e.g., the video data may be captured during an assessment session, such as a TD assessment session, of the patient that is performed by a physician).

[0078] The video data may be processed to identify landmarks of the patient (206), for example, by the assessment system 104. For example, the assessment system 104, and for example the data preprocessor 116, may extract and analyze the video data, such as video data 112, to identify landmarks (e.g., facial landmark data and / or other bodily landmarks, such as landmarks of the hands, feet, or trunk) of the patient by utilizing one or more image processing techniques (e.g., dlib facial recognition, Haar Cascade technique, Fisherface method, elastic graph matching technique, and other like techniques for facial recognition). In some examples, one or more image processing techniques may be applied to the video data to identify facial landmarks in one or more frames of the video data to produce labeled frames. The image processing technique may include any combination of a dlib facial recognition method, a Haar Cascade method, a Fisherface method, and / or an elastic graph matching method. In some instances, the processing of the video data may include identifying the facial landmark data that corresponds to each of the facial landmarks from the labeled frames. The facial landmark data may include, for example, one or more of a landmark identifier, landmark coordinate position, frame source identifier, video source identifier, and / or patient identifier.

[0079] In the example of dlib facial recognition, the assessment system 104 may receive video data of the patient, and the video data may include a plurality of frames of the patient’s face. The assessment system 104 may include a detector to identify the faces within each frame of the video data, a shape predictor to identify facial landmarks 117 within the video data (e.g., to precisely localize the face), and / or a face recognition model (e.g., such as dlib facial recognition). The assessment system 104 may map frames of the video data that include an image of the patient’s face to a multi-dimensional vector space (e.g., a 128-dimensional vector space). In some examples, the assessment system 104 may be configured to identify a bounding box for the patient’s face (e.g., identify a set of facial landmarks 117 associated with the patient’s face within each frame of the video data). The assessment system 104 can perform face recognition by mapping landmarks associated with the patient’s face to the multi-dimensional vector space and then comparing the Euclidean distance of the identified landmarks to a distance threshold (e.g., a distance threshold of 0.6) to ensure that the patient’s face is being accuratelyand consistently identified across frames of the video data. In some examples, the assessment system 104 may determine an accuracy of close to 100% (e.g., 99.38%) on the standard Labeled Faces in the Wild (LFW) face recognition benchmark.

[0080] Referring to FIG. 5, in some embodiments, the data preprocessor 116 may extract and analyze the video data 112 and output a facial mask 115 comprising a plurality of unique, uniquely labeled facial landmarks 117 (e.g., numbered 1-60 in FIG. 5, as an example). In such cases, the data preprocessor 116 may extract and analyze the video data 112 per each frame of video to determine a facial mask 115, and likewise constituent facial landmarks 117 of the patient. Such facial landmarks 117 may correspond to physical landmarks on the patient’s face, such as the facial outline, eyes, nose, lips, tongue, a nasion, an inion, a lateral canthus, an external auditory meatus (e.g., ear attachment point), one or more preauricular points of the human subject, and the like. For example, the points representing the facial landmarks 117 may be indicated as coordinates (e.g., X and Y coordinates for standard facial datasets, or X, Y, and Z coordinates for three-dimensional facial datasets) represented in units of pixels and / or a gradient of pixels corresponding to the predicted regions or landmarks. Using the aforementioned image processing techniques, the assessment system 104 (e.g., the data preprocessor 116) can extract and store facial landmark data 118 corresponding to each of the facial landmarks 117 identified in one or more frames of the video data 112 as a data array and / or data table. For example, facial landmark data 118 corresponding to a particular facial landmark 117 may be in the form of data array 119. Facial landmark data 118 can include landmark identifier, landmark coordinate position, frame source identifier, video source identifier, patient (e.g., anonymized) identifier, etc. The assessment system 104 may store the identified facial landmarks 117 of the patient as facial landmark data 118 in the database 110. Finally, it should be appreciated that the facial landmarks 117 illustrated in FIG. 5 are a subset of the actual number of landmarks that can be identified by the data preprocessor 116 of the assessment system 104.

[0081] In one embodiment, the assessment system 104 may process the captured video data 112 during the assessment session. In other examples, the assessment system 104 may obtain the captured video data 112 and store the video data 112 in the database 110 during the assessment session. Alternatively, the assessment system 104 may receive (e.g., retrieve) the stored video data 112 and process the video data 112 after the patient completed the examination procedure.

[0082] The assessment system 104 may segment the video data into a plurality of long segments (208). In some examples, the assessment system 104 may exclude portions of the video data from the plurality of long segments. For example, the assessment system 104 may filter the video data to identify and remove portions of the video data that will not be used later during the procedure 200, such that the long segments do not include these excluded portions. As such, the plurality of long segments may be defined by the durations of the video data between the excluded portions of the video data. For instance, the long segments may be the video data that was not excluded. The excluded portions of the video data may be portions where the patient is not looking at the camera, where facial landmark data of the patient is not identified within the video data, and / or where the patient is not performing a predefined prompt. Thus, the excluded portions of the video data may be portions that are unusable for assessment of involuntary movement according to the methods described herein.

[0083] The assessment system 104 may determine how to generate the long segments based on the facial landmarks 117 of the video data. The assessment system 104 may identify portions of the video data 112 where the patient has not performed the prompt and filter the video data 112 based on these portions. That is, the assessment system 104 may remove, or filter out, the portion(s) of video data 112 where the patient is not performing the prompt(s), such that the remaining segments of the video data 112 are the long segments of the video data 112. Accordingly, the long segments may be those portions of the video data 112 where the patient is performing the prompt.

[0084] For example, the assessment system 104 may analyze the facial landmarks 117 to determine whether the performed action of the patient corresponds to the prompt (e.g., the prompt generated at 202). The assessment system 104 may be preprogrammed to associate one or more the identified facial landmarks 117 with each prompt. For instance, the assessment system 104 may be preprogrammed to associate the prompt “Please open your mouth for 30 seconds” with a first subset of the facial landmarks (e.g., landmarks labeled 4-14), and the prompt “Please open your mouth and protrude your tongue” with a second subset of the landmarks 118 (e.g., landmarks labeled 49-60). In some instances, for example, by associating a particular subset of the identified facial landmarks 117 with a prompt, the assessment system 104 may analyze those facial landmarks 117 to determine whether the performed action corresponds to the prompt. For example, in response to providing the patient with a prompt “Please open yourmouth for 30 seconds,” the assessment system 104 may analyze the first subset of facial landmarks 117 to determine the spatial arrangement and temporal shifts in those particular facial landmarks 117.

[0085] The assessment system 104 may determine that the performed action does not correspond to the prompt based on the movement in the spatial arrangement of the corresponding facial landmarks 117. For example, the assessment system 104 may determine that none of the facial landmarks 117 associated with the patient’s cheeks, jaw, lips, and / or tongue move (i.e., an indication that the patient did not open the patient’s mouth) in the video data 112 (e.g., a particular duration of the video data 112). In other cases, the assessment system 104 may determine that the performed action does not correspond to the prompt if a positional displacement of spatial arrangement of the corresponding facial landmarks 117 does not exceed a threshold value. For example, the assessment system 104 may detect a slight movement in the patient’s lips and jaw based on the spatial arrangement and shifting of the associated facial landmarks 117 in a series of frames from the captured video data 112. However, the assessment system 104 may determine that the spatial arrangement and shift of the associated facial landmarks 117 did not shift beyond a predetermined threshold value (i.e., an indication that the patient did not open the patient’s mouth enough to analyze the landmarks associated with the patient’s tongue). As such, the assessment system 104 may determine that the performed action does not correspond to the provided prompt.

[0086] In some examples, for the cases in which it is determined that the performed action does not correspond to the prompt, the assessment system 104 may provide the same prompt to the device 108 (202), which displays the prompt on the user interface 302. It is noted that in some cases in which the assessment system 104 detects that the patient is not performing the action corresponding to the prompt, the assessment system 104 may issue another prompt or indication to display on the user interface 302 that further directs the patient to perform the action corresponding to the prompt.

[0087] The assessment system 104 may determine that the performed action corresponds to the provided prompt by applying one or more classifiers associated with the action of the prompt to at least a portion of the temporal shifts of the identified facial landmarks 117. In some cases, each classifier may correspond to one or more orofacial movements of the patient. The assessment system 104 may select one or more orofacial movements of the respective classifiersand compare the selected orofacial movements to the temporal shifts of the identified facial landmarks 117. In one or more cases, the assessment system 104 may determine that the patient performed the correct action based on the temporal shifts of the identified facial landmarks 117 being correctly associated with the selected orofacial movements.

[0088] The assessment system 104 may determine that the performed action corresponds to the prompt if a positional displacement of the facial landmarks 117 exceeds a threshold value. For example, the assessment system 104 may detect a movement in the patient’s lips and jaw based on the spatial arrangement and shifting of the associated facial landmarks 117 in the captured video data 112. Further, the assessment system 104 may determine that the spatial arrangement and shift of the associated landmarks shifted beyond a predetermined threshold value (e.g., an indication that the patient opened the patient’s mouth enough to analyze the landmarks associated with the patient’s tongue). As such, the assessment system 104 may determine that the performed action corresponds to the provided prompt.

[0089] Additionally, the assessment system 104 may remove, or filter out, the portion(s) of video data 112 where each, or alternatively a threshold amount, of the facial landmarks 117 are not present within the video data. Such portions of the video data 112 may represent portions of the assessment where the patient is not adequately facing the camera, and thus the facial landmarks 117 cannot be properly assessed.

[0090] Further, in some examples, the plurality of long segments may be generated based on the video data by a trained professional (e.g., the video data can be filtered manually by a trained professional to generate the long segments). In such examples, 208 may be performed manually and not by the assessment system 104.

[0091] The assessment system 104 may generate one or more short segments based on each of the plurality of long segments of the video data (210). The long segments may be defined by multiple different durations. For example, the long segments may be aperiodic and / or the long segments may be defined by varying or non-consistent durations (e.g., the long segments may have durations that are based on the time of the excluded portions that are removed when the video data is filtered). The short segments may be defined by a predefined duration, such as about four seconds. In some examples, the plurality of short segments include video segments that overlap with at least one other short segment. For instance, the short segments may overlap with one another (e.g., the short segments associated with a long segment (e.g., a single longsegment) may overlap with one another). The overlap of the short segments may help to ensure that a movement (e.g., an involuntary movement indicative of TD) is captured by at least one of the short segments (e.g., even if it occurs at the very beginning or end of a short segment).

[0092] The assessment system 104 may apply a plurality of machine learning models to the obtained video data to generate a total severity score. The assessment system 104 may apply a first trained machine learning model (e.g., a plurality of first trained machine learning models) to determine a movement indicator for each short segment (212). For example, the assessment system 104 may apply, for each short segment of the plurality of short segments, a plurality of first trained machine learning models to the facial landmark data of the short segment, where each of the plurality of first trained machine learning models corresponds to a specific category of facial movement, to generate a plurality of movement indicators for each of the plurality of short segments, where each of the movement indicators corresponds to a respective, specific category of facial movement. Each movement indicator may indicate a probability that movement of a category of facial movement occurred within the short segment. As noted below, the movement indicator (e.g., the probability that movement occurred) may be on a range from 0.0- 1.0 or a binary scale. The categories of facial movement may include one or more of eye movement, mouth movement, and / or head movement (e.g., which may include any head movement, including movements based on torso tremor and / or head tilting). Further, in other examples, the one or more categories of facial movement may include any combination of the following: mouth and lower face movement, upper face movement, head tremor, other tremor, head movement, tongue movement, other movement, contractures and tilts, and / or number of blinks. These additional categories of facial movement, as noted below, may be associated with or defined within one of eye movement, mouth movement, and / or head movement.

[0093] In some examples, the assessment system 104 may apply a second trained machine learning model (e.g., a plurality of second machine learning models) to determine a movement severity score for each long segment (214). For example, the assessment system 104 may apply, for each long segment of the plurality of long segments, each of a plurality of second trained machine learning models to the movement indicators of the short segments associated with the long segment to generate a plurality of movement severity scores for each of the plurality of long segments. As with the first trained machine learning models that are utilized to determine the movement indicators for the plurality of short segments, each of the plurality of second trainedmachine learning models corresponds to a specific category of facial movement. In other words, for each first trained machine learning model utilized to determine movement indicators, there may be a corresponding second trained machine learning model utilized to determine movement severity scores based on the movement indicators determined by the corresponding first trained machine learning model. Each movement severity score may indicate a severity of movement (e.g., involuntary movement) of the category of facial movement within the corresponding long segment. However, it should be appreciated that in some examples, the assessment system 104 may not use the second machine learning models 114a-l 14b at all, and in such examples, 214 may be omitted. Further, in some examples, the assessment system 104 may use the second machine learning models 114a-l 14b to determine the movement severity scores but the movement severity scores generated by these models may not be used when determining a total severity score for the patient. In other embodiments, the movement severity scores may comprise an average (e.g., a scaled-up average) or other numerical generalization of the corresponding movement indicators for a particular long segment of video data.

[0094] The assessment system 104 may determine a total severity score (216). For example, the assessment system 104 may determine the total severity score based on the movement indicators generated by the first machine learning model(s). The assessment system 104 may apply a third machine learning model to the movement indicators to determine the total severity score. The total severity score may correspond to the risk of the patient receiving a positive TD diagnosis by a trained physician and / or a severity of TD in a patient. The total severity score may correspond to an Abnormal Involuntary Movement Scale (AIMS) score (e.g., however, in some examples, the total severity score is not the same as the AIMS score, although it is calculated using the same scale and indicative of the risk of the patient receiving a positive TD diagnosis by a trained physician). Further, as noted herein, in some examples, the assessment system 104 may determine the total severity score based on the movement severity scores associated with the plurality of long segments that were generated using the second machine learning model(s).

[0095] In some examples, the scale associated with the plurality of movement indicators may be different from the scale associated with the plurality of movement severity scores. For instance, the scale associated with the plurality of movement indicators may be a range from 0.0- 1.0 or a binary scale. Further, and for example, the scale associated with the plurality ofmovement severity scores may be in a range of 0.0-10.0. Further, in some instances, the scale associated with the total severity score may be different from the scale associated with the plurality of movement indicators and / or the scale associated with the plurality of movement severity scores. For example, the scale associated with the total severity score may be a range of 0-28.

[0096] In one or more cases, a notification indicating one or more of the movement severity scores and / or the total severity score is generated (218), for example, by the assessment system 104. For example, the assessment system 104 may display the notification on a display device of the patient and / or a physician. In some examples, the assessment system 104 may generate a notification that includes a recommendation to refer the patient to a trained clinician when the total severity score is above a predetermined threshold. Specific examples of notifications and graphical user interfaces that can be generated by the assessment system 104 are shown and described with reference to FIG. 6, 7, and 8.

[0097] FIG. 6 is a diagram of an example graphical user interface (GUI) 600 that can be displayed on a display device that indicates a plurality of patients, a total severity score for each patient, and an indication of a trend associated with the total severity score for each patient. The GUI 600 may include one or more notifications (e.g., the GUI 600 may be a notification). The assessment system 104 may display the GUI 600 on a display device of the patient and / or a physician.

[0098] The GUI 600 may include an indication of a plurality of patients 610, an indication of the last visit date 620 for each of the patients, and indication of a last total severity score 630 for each patient corresponding to the last visit date 620, and an indication of a trend of the patient’s total severity scores 640. The indication of a total severity score 630 may be based on a procedure (e.g., the procedure 200) that the patient performed to assess the risk for and / or severity of TD. The trend of the patient’s total severity scores 640 may be a line graph, for example as shown, that indicates a plurality of the patient’s total severity scores over a period of time (e.g., the previous 3 to 5 visits).

[0099] The assessment system 104 may provide a notification that includes any combination of the movement indicators, the movement severity scores, and / or the total severity score to the patient and / or a trained physician via a dashboard of a display device. For example, the assessment system 104 may generate (e.g., via a GUI displayed on a display device) anindication of the movement severity scores for each category of facial movement associated with each of the plurality of long segments. In some examples, the indication of the movement severity scores for each category of facial movement may be color coded based on a value of each respective movement severity score (e.g., red to indicate scores in a highest range, yellow to indicates scores in an intermediate range, and green to indicates scores in a lowest range). In some instances, the assessment system 104 may display, in response to a user selection of a particular long segment of the plurality of long segments, the video data associated with the selected long segment via the display device. For instance, the assessment system 104 may display, with the video data associated with the selected long segment, a point cloud associated with the facial landmark data associated with the selected long segment (e.g., the point cloud throughout the long segment that was used to determine the movement severity score for the long segment). For instance, in some examples, the assessment system 104 may display, with the video data associated with the selected long segment, changes in each of the movement severity scores associated with each category of facial movement of the selected long segment throughout a duration of the selected long segment (e.g., via a graph that indicates the severity score throughout the duration of the selected long segment). The facial landmarks 117 of FIG. 5 are an example of a point cloud associated with facial landmark data.

[0100] FIG. 7 is a diagram of an example GUI 700 that can be displayed on a display device that indicates a trend of total severity scores for a patient, an indication of the movement severity score for different categories of facial movement across a plurality of long segments associated with the patient, and a video and point cloud associated with the video data of the patient. The GUI 700 may include one or more notifications (e.g, the GUI 700 may be a notification). The assessment system 104 may display the GUI 700 on a display device of the patient and / or a physician.

[0101] The GUI 700 may include a trend of the patient’s total severity scores 710. The trend of the patient’s total severity scores 710 may be a line graph, for example as shown, that indicates a plurality of the patient’s total severity scores over a period of time (e.g., the previous seven visits). The GUI 700 may include a table 720 that provides an indication of the movement severity scores for each category of facial movement (e.g., mouth score, head score, and eye score) associated with each of the plurality of long segments (e.g., in this example, segments one through five) filtered from the complete video data that are part of a total severity score (e.g.,seven) determined using a procedure (e.g., the procedure 200) on a particular date (e.g., October 13, 2014) for a patient. In some examples, the indication of the movement severity scores for each category of facial movement may be color coded based on a value of each respective movement severity score (e.g., red to indicate scores in a highest range, such as from 8 to 10, yellow to indicates scores in an intermediate range, such as from 3 to 7, and green to indicates scores in a lowest range, such as from 0 to 2). In the GUI 700, the color codes are illustrated by different shades of greyscale.

[0102] The GUI 700 may include video data 730 associated with a long segment, a point cloud 740 that is based on the video data 730, and indications 750a, 750b, and 750c of the movement severity scores associated with each category of facial movement (e.g., mouth movement score, as indicated by 750a, head movement score, as indicated by 750b, and eye movement score, as indicated by 750c) for the long segment. The facial landmarks 117 of FIG. 5 are an example of a point cloud associated with facial landmark data. The indications 750a, 750b, and 750c may include a graph that indicates the respective movement severity score as well as the progression of respective movement indicators for the short segments that comprise the selected long segment. The assessment system 104 may, for example, display the video data 730, the point cloud 740, and / or the indications 750a, 750b, and 750c in response to a user selection of one of the long segments of the video data (e.g., a selection of a long segment from the table 720). The assessment system 104 may play the video data 730 and associated point cloud 740 in response to a user selection, for example, so that a user (e.g., a physician or the patient) can review the video data 730 and / or the point cloud 740 for the long segment. As such, the assessment system 104 may display, in response to a user selection of a long segment of the plurality of long segments, the video data 730 associated with the selected long segment via the display device.

[0103] In one or more other cases, the assessment system 104 may provide the clinician or prescribing practitioner of the psychiatric drug with the notification indicating the total severity score. In one or more cases, the assessment system 104 may flag the movement severity scores and / or the total severity score as indicating evidence of TD. For instance, if the assessment system 104 determines that the total severity score exceeds a threshold value, the assessment system 104 determines that there is evidence of TD. In another instance, if the assessment system 104 determines that the total severity score exceeds a threshold value, the assessment system 104 provides a recommendation to refer the patient to a trained clinician.

[0104] The assessment system 104 may store the total severity score in the database 110. Further, the assessment system 104 may store the total severity score as anonymized data in the database 110. One or more subsequent total severity scores for the patient as provided by a clinician can also be provided to and stored within the database 110, such that subsequently the movement severity score and video data for the patient and the clinician-provided AIMS score can be included within the training data for further training the machine learning models, wherein an exemplary process for training the machine learning model is described below.

[0105] In some examples, the assessment system 104 may analyze acoustic data (e.g., voicebased digital biomarkers) of the patient, and determine the progression of TD based on the acoustic data of the patient. Alternatively or additionally, the assessment system 104 may be configured to analyze acoustic data (e.g., voice-based digital biomarkers) of the patient, and determine the total severity score further based on the acoustic data of the patient. The assessment system 104 may receive the acoustic data (e.g., a voice recording) of the patient. For example, the patient may be instructed to produce audio data in response to a prompt to speak freely and / or in response to a prompt that requests that the patient follow a specific script. The acoustic data may have a predetermined duration (e.g., 1 to 5 minutes). The prompt may include one or more of text and / or an image (e.g., similar to the prompts presented at 202). Similar to the video data above, the acoustic data may be filtered (e.g., segmented) into discrete segments, for example discrete segments of a predetermined duration. The predetermined duration may be different from the durations of the long segments or the short segments. For example, the segments of the acoustic data may be less than the short segments of the video data. In some examples, the segments of the acoustic data may be between about 50-100 ms.

[0106] The acoustic data may include one or more acoustic features, such as amplitude envelope, root mean square deviation, zero crossing count, fundamental frequency, mel- frequency cepstral coefficients (MFCCs), delta, delta-delta, a chroma feature. The assessment system 104 may be configured to detect one or more biomarkers, such as acoustic features, in the audio data. Examples of such acoustic features include block, jitter, a shimmer, a tremor, Harmonics-to-noise ratio (HNR), FDR, ADR, quasi-open quotient (QOQ), normalized amplitude quotient (NAQ), a peak slope, Fl mean, F2 mean, Fl variability, F2 variability, Fl range, vowel space, LPC, mel-frequency cepstral coefficients (MFCCs), fO mean, fO variability, fO range, intensity mean, intensity variability, energy variance, energy velocity, maximum phonation time,speech rate, articulation rate, time talking, utterance duration mean, pause duration mean, pause variability, pause rate, and / or pauses total.

[0107] The assessment system 104 may use a machine learning model 125 to determine an acoustic-specific total severity score based on the acoustic severity data 127 of a patient. The total severity score determined by the assessment system 104 using the acoustic data may be different from the total severity score determined using the video data. As such, the assessment system 104 may determine multiple total severity scores, such as a first total severity score based on the video data and a second total severity score based on the acoustic data. Alternatively or additionally, the assessment system 104 may determine the progression of TD based on the acoustic data of the patient.

[0108] The machine learning model 125 may have been trained using training data to assess the progression of TD and / or refinement of the total severity score based on biomarkers identified within acoustic data of a patient. The machine learning model may be trained using a plurality of processed audio recordings. The plurality of audio recordings may be obtained from a plurality of patients, such as the patient data that was captured during one or more studies. The patients may have confirmed TD diagnoses and / or AIMS scores, and an indication of the TD diagnoses and / or AIMS scores may be provided to the model with the acoustic data. The model may be trained using AIMS scores, for example corresponding to the audio recording of each of the plurality of patients.

[0109] The assessment system 104 may be configured to determine a progression of the movement disorder over a period of time. Upon a TD diagnosis, a clinician can prescribe a chemical or biologic medication for the treatment of TD. For example, the clinician can prescribe a treatment comprising deutetrabenazine, valbenazine, tetrabenazine, clonazepam, and / or botulinum toxin.

[0110] Further provided herein, are methods useful in treating TD. In some examples, the assessment system 104 may be configured to determine a TD total severity score as described above to inform the creation, maintenance, and / or revisions of a medical treatment plan for the patient, such as administering to a patient in need of a pharmaceutical composition comprising a pharmacologically active agent described herein, in an effective amount to treat TD. As used herein, the terms “method of treatment” or “therapy” (e.g., as well as different forms thereof) in relation to tardive dyskinesia, include preventative (e.g., prophylactic), curative, or palliativetreatment. As used herein, the term “treating” includes alleviating or reducing at least one adverse or negative effect or symptom of tardive dyskinesia. The term “administering” means providing to a patient the pharmaceutical composition or dosage form (used interchangeably herein). As used herein, the terms “compound”, “drug”, “pharmacologically active agent”, “active agent”, or “medicament” are used interchangeably herein to refer to a compound or compounds or composition of matter which, when administered to a subject (human or animal) induces a desired pharmacological and / or physiologic effect by local and / or systemic action. Possible pharmacologically active agent can be selected from tetrabenazine, deutetrabenazine, valbenazine, deuvalbenazine, clonazepam and / or botulinum toxin. One preferred active agent disclosed herein is deutetrabenazine. “Deutetrabenazine” or “deu-TBZ” is a selectively deuterium-substituted, stable, non-radioactive isotopic form of tetrabenazine in which the six hydrogen atoms on the two O-linked methyl groups have been replaced with deuterium atoms (i.e. -0CD3 rather than -0CH3 moieties).

[0111] A vesicular monoamine transporter 2 (VMAT2) inhibitor may be prescribed for treating uncontrolled, involuntary movements associated with movement disorder, such as TD. The VMAT2 inhibitor may include, for example, deutetrabenazine (e.g., Austedo® or Austedo® XR), tetrabenazine, and / or valbenazine. In some examples, the total daily dose of the VMAT2 inhibitor can be administered on a once daily basis (qd) or on a twice daily basis (bid). In some examples, the total daily dose of the VMAT2 inhibitor is 6 mg, or 12 mg, or 18 mg, or 24 mg, or 30 mg, or 36 mg, or 42 mg or 48 mg of the VMAT2 inhibitor. In other examples, the total daily dose of the VMAT 2 inhibitor is 40 mg or 80 mg. The daily amount of drugs may require periodic reevaluation from clinicians, for example, to update the patient’s titration schedule. For instance, a patient may receive a package of VMAT2 inhibitor pills having different daily amounts, also referred to as a titration kit. For example, the VMAT2 inhibitor pills may be available in pills with different strengths, such as a 6 mg pill, a 12 pill, and / or a 24 mg pill. A clinician may prescribe the patient with an initial daily amount of a VMAT2 inhibitor. The daily amount may be based on the results of an AIMS test, for example, in combination with other factors associated with the patient. In an embodiment, the daily amount may be based, at least in part, on the total severity score determined by the assessment system 104 as described above. The titration kit may comprise a supply of VMAT2 inhibitor pills for a predetermined period of time (e.g., four weeks). The daily amount of VMAT2 inhibitor may progressively increasethroughout the titration kit. For example, the daily amount of VMAT2 inhibitor may increase in a systematic, stepwise manner (e.g., 6 mg / day for week 1, 18 mg / day for week 2, 24 mg / day for week 3, 30 mg / day for week 4, etc.).

[0112] The accurate diagnosis and assessment of TD for the patient may require periodic review and assessment of the patient’s conditions (e.g., such as through the use of AIMS tests and / or the procedure 200). In an embodiment, the assessment system 104 (e.g., the movement severity scores and / or total severity score generated by the assessment system 104) may be used to flag patients for further assessment for potential TD diagnosis by a trained condition. The patient’s clinician may determine an initial daily amount and / or increase the daily amount for the patient after receiving and analyzing a latest periodic evaluation and / or based on a predetermined schedule. The assessment system 104 (e.g., the movement severity scores and / or total severity score generated by the assessment system 104) may be used to determine the titration process for the patient and / or determine the daily amount for the patient to take during any particular interval of treatment (e.g., the initial daily amount). As such, the assessment system 104 may be configured to adjust the patient’s dosing schedule and / or dosing amount based on the movement severity scores generated by the assessment system 104. In using the assessment system 104 as an aid in dosing selection and titration, both the clinician and patient are provided with an objective tool that standardizes and expedites the TD assessment process. Examples of titration regimens for the treatment of TD are described in PCT Publication No. WO 2016 / 0144901, and U.S. Patent Publication No. US 2016 / 0287574, the entireties of which are incorporated by reference herein.

[0113] In one or more cases, in conjunction with a regimen for the treatment of TD, the patient may utilize the assessment system 104 to perform multiple assessments (e.g., a second assessment) during a subsequent time period to determine a progression of the movement disorder over a period of time. The second assessment may be performed in a same or similar manner as described with the example procedure 200 of FIG. 2. As such, a description of such features is not repeated. As a result of the second assessment, the assessment system 104 may generate a second total severity score. In one or more cases, additional movement severity scores may be obtained as described herein. In one or more cases, the assessment system 104 may compare the second total severity score to the previous total severity score to determine a state of the movement disorder of the patient. For instance, for the cases in which the assessmentsystem 104 determines that the second total severity score is less than the previous total severity score, the assessment system 104 may determine that the severity of the movement disorder is regressing. In another instance, for the cases in which the assessment system 104 determines that the second total severity score is greater than the previous total severity score, the assessment system 104 may determine that the severity of the movement disorder is progressing. In one or more cases, the assessment system 104 may provide a notification to the clinician and / or patient via the device 106 and / or device 108 recommending a reevaluation or adjustment of the currently prescribed treatment regimen.

[0114] The assessment system 104 may be configured to determine a progression of a severity of the movement disorder and may generate and provide a notification corresponding to the progression of the severity of the movement disorder, for example, in conjunction with the regimen for the treatment of TD and subsequent assessments that generate corresponding total severity scores. For example, to begin treatment of TD, a patient be administered a total daily dose of 6mg of pharmacologically active agent (e.g., deutetrabenazine) based on an initial total severity score. At a subsequent time (e.g., a week from when the patient performed the first assessment and the assessment system 104 generated the initial total severity score), the patient may perform the second assessment in a same or similar manner as described with the example procedure 200 of FIG. 2 and generate the second total severity score. In one or more cases, additional movement severity scores may be obtained as described herein. In one or more cases, the assessment system 104 or a clinician may compare the second total severity score to the previous total severity score to determine a state of the movement disorder of the patient (i.e., a progression of a severity of the movement disorder), whether a current treatment is effective or non-effective based on the state of the movement disorder, and / or whether the patient is failing to follow the treatment regimen. For example, the assessment system 104 and / or a clinician may determine that the second total severity score is greater than the previous total severity score (e.g., the initial total severity score), and thus, the severity of the movement disorder is progressing. Further, the assessment system 104 may provide a notification (e.g., via displaying the notification on an interface of a device, such as device 108) based on the determination that the severity of the movement disorder is progressing indicating a likelihood that the current treatment is no longer effective. The assessment system 104 may provide the state of the movement disorder such that a clinician may determine as to whether the regimen for thetreatment of TD should be altered. The patient may continue performing assessments at intervals, such as, but not limited to, weekly intervals, such that the assessment system 104 may generate additional total severity scores and determine a state of the movement disorder. By providing an automated and objective method for repeatedly assessing TD severity and progression, patient adherence can be improved through increased efficiency of TD severity assessment, as well as discrete metrics for observing TD improvement during a treatment regimen.

[0115] The machine learning model, such as the machine learning models 113a- 113c, 114a- 114c, 123, may be trained using training data that includes video data from a plurality of patients and total severity scores associated with each patient of the plurality of patients. In some examples, the plurality of trained machine learning models are trained using training data that includes facial landmark data extracted from videos from a plurality of patients with a positive TD diagnosis, clinician-determined Abnormal Involuntary Movement Scale (AIMS) scores associated with each patient of the plurality of patients, and / or clinician-assigned labels corresponding to each of the videos and / or portions thereof indicating a presence of a plurality of distinct category of facial movement. For instance, in some examples, the categories of facial movement comprise eye movement, mouth movement, and / or head movement. In one embodiment, the labels may comprise binary indicators of whether a particular category of facial movement occurred. In another embodiment, the labels may comprise scores rating a severity of facial movement (e.g., 0.0-1.0, etc.). In a further embodiment, the labels may comprise a count of instances of a particular category of facial movement within a video (e.g., number of blinks).

[0116] In some examples, the training data used to train any combination of the machine learning models 113a-l 13c, 114a- 114c, 123 may include any of the following categories of facial movement: mouth and lower face movement, upper face movement, head tremor, other tremor, head movement, tongue movement, other movement, contractures and tilts, and / or number of blinks. In such examples, the models may map the clinician-assigned label categories to the categories associated with the movement indicators and / or movement severity scores for training the corresponding models. For instance, the mouth and lower face movement and the tongue movement may be associated with mouth movement (e.g., may be used to train the first and / or second machine learning models associated with mouth movement). The head tremor, other tremor, head movement, and contractures and tilts may be associated with head movement (e.g., may be used to train the first and / or second machine learning models associated with headmovement). The upper face movement and / or number of blinks may be associated with eye movement (e.g., may be used to train the first and / or second machine learning models associated with eye movement). Finally, in some examples, the other movement category may be associated with a single category (e.g., head movement) or multiple categories (e.g., may be used to train any combination of the first and / or second machine learning models associated with each of head movement, eye movement, and / or mouth movement).

[0117] As described herein, the video data (e.g., the video data for each patient) may be manually or automatically segmented (e.g., filtered or divided) into long segments, and each long segment may be segmented into short segments. The long segments may be defined by multiple different durations. For example, the long segments be aperiodic and / or the long segments may be defined by varying or non-consistent durations (e.g., the long segments may have durations that are based on the time of the excluded portions that are removed when the video data is filtered). The short segments may be defined by a predefined duration, such as about four seconds. In some examples, the plurality of short segments include video segments that overlap with at least one other short segment. For instance, the short segments (e.g., the short segments associated with a long segment (e.g., a single long segment)) may overlap with one another. The long segments and / or the short segments may include facial landmark data associated with facial landmarks of the patient.

[0118] The first machine learning models 113 a- 113c may be trained to determine a movement indicator of a short segment using training data, where the training data includes short segments of video data from a plurality of patients, facial landmarks 117 for the short segments, and labels corresponding to the categories of facial movement described above for each short segment of the training data, and / or clinician-determined AIMS scores. In some examples, the labels for each short segment of the training data may be assigned, for example manually, by a trained clinician.

[0119] When used, the second machine learning models 114a-l 14c may be trained to determine a movement severity score of a long segment using training data, where the training data includes long segments of video data from a plurality of patients, associated movement indicators for each short segments that make up the long segments of the training data, a movement severity score for each long segment, and / or clinician-determined AIMS scores. Insome examples, the movement severity score for each long segment of the training data may be assigned, for example manually, by a trained clinician.

[0120] Finally, the third machine learning model 123 may be trained to determine a total severity score of a plurality of long segments using training data, wherein the training data includes sets of a plurality of long segments of video data from a plurality of different users, labels corresponding to the categories of facial movement described above for the short segments comprising each long segments of each set, and a total severity score for each set of long segments. In some examples, the total severity score for the set of long segments of each patient in the training data may be assigned, for example manually, by a trained clinician. For example, the total severity score for the training data can comprise a clinician-assessed AIMS score.

[0121] In some examples, the training data may include only short segments where a movement of relevant facial landmark data was detected (e.g., by a physician, by a technician, and / or using a facial recognition tool). For example, the training data may include the segments (e.g., only the video segments) where a particular category of facial movement associated with TD (e.g., upper face movement, tongue movement, mouth and lower face movement, head shaking, and head tilting) is present within the video segment. Further, in some examples the training data may include (e.g., only include) the facial landmark data associated with the corresponding category of facial movement from each video segment. Such selective feeding of data into the training data may be done for statistical reasons, such as to avoid overfitting. The training data may include correlations between each landmark and the particular category (e.g., or categories) of facial movement they are relevant to. Further, in some examples, the training data may include a reduced amount of data, i.e., the facial landmark data that is relevant to the category of facial movement assessed by a particular machine learning model. For example, the training data used to analyze mouth movement may not include (e.g., have filtered out) the facial landmark data corresponding to landmarks not relevant to mouth movement.

[0122] The machine learning model(s) described herein (e.g., the first plurality of machine learning models 113a-l 13c, the second plurality of machine learning models 114a-l 14c, and / or the third machine learning model 123) may include any combinations of algorithms. For example, the machine learning model may comprise gradient boosted decision trees, a random forest algorithm, a logarithmic regression model, a RandOm Convolutional KErnel Transform (ROCKET) model, and / or an XGBoost algorithm. In an example, gradient boosted decisiontrees may combine weak learners to minimize the loss function. For example, regression trees may be used to produce real values for splits and / or to be added together. The weak learners may be constrained, for example, to a maximum number of layers, a maximum number of nodes, a maximum number of splits, etc. Trees may be added one at a time to the machine learning algorithm and / or existing trees may remain unchanged. A gradient descent procedure may be used to minimize loss when adding trees. For example, additional trees may be added to reduce the loss (e.g., follow the gradient). In an example, the additional tree(s) may be given parameters and those parameters may be modified to reduce the loss. In an example, the XGBoost algorithm may comprise an implementation of gradient boosted decision trees that may be designed for speed and / or performance.

[0123] The XGBoost algorithm may (e.g., automatically) handle missing data values, support parallelization of tree construction, and / or continued training. The Gradient Boosting Trees technique may produce a prediction model in the form of an ensemble (e.g., multiple learning algorithms) of base prediction models, which are decision trees (e.g., a tree-like model of decisions and their possible consequences). The XGBoost algorithm may build a single strong learner model in an iterative fashion by using an optimization algorithm to minimize some suitable loss function (e.g., a function of the difference between estimated and true values for an instance of data). The optimization algorithm may use a training set of known values of the response variable (e.g., AIMS score) and their corresponding values of predictors (e.g., training data patient landmarks) to minimize the expected value of the loss function. The learning procedure may consecutively fit new models to provide a more accurate estimate of the response variable.

[0124] Although described primarily in the context of an XGBoost algorithm, other machine learning models may be used, such as an unsupervised learning method (e.g., a clustering method, such as a k-means or c- means clustering method) and / or a supervised learning method (e.g., gradient boosted decision trees). As an example, the assessment system 104 may use a gradient descent or a stochastic gradient descent learning method.

[0125] A supervised learning method may use labeled training data to train the machine learning algorithm. As training data is received, the supervised learning method may adjust weights until the machine learning algorithm is appropriately weighted. The supervised learning method may measure the accuracy of the machine learning algorithm using a loss function. Thesupervised learning method may continue adjusting the weights until the error is reduced below a predetermined threshold.

[0126] The model used to train the first machine learning models 113a-l 13c, the second machine learning models 114a-l 14c, and / or the third machine learning model 123 may be the same or different. Further, in some examples, the first machine learning models 113a- 113c may use different models, for example, based on the category of facial feature. For instance, in one example, the first machine learning models that generate movement indicators for mouth movement and eye movement may use an XGBoost model, while the first machine learning model that generates movement indicators for head movement may use a ROCKET model. Finally, in some examples, the third machine learning model 123 may include an XGBoost model.

[0127] FIG. 4 is a block diagram illustrating an example computing device 400. One or more computing devices, such as the computing device 400, may implement one or more features for assessing a risk of a patient for TD, as described herein. For example, the computing device 400 may be an example of one or more of the device 106, the device 108, and / or the server device(s) 102 shown in FIG. 1A. The computing device 400 may comprise a processor 402, a memory 404, a storage device 406, an VO interface 408, and / or a communication interface 410, which may be communicatively coupled by way of a communication infrastructure 412. It should be appreciated that the computing device 400 may include fewer or more components than those shown in FIG. 4. For instance, the computing device may include a camera (e.g., when implemented as the device 108).

[0128] The processor 402 may include hardware for executing instructions, such as those making up a computer program. In examples, to execute instructions for dynamically modifying workflows, the processor 402 may retrieve (or fetch) the instructions from an internal register, an internal cache, the memory 404, or the storage device 406 and decode and execute the instructions.

[0129] The memory 404 may be a volatile or non-volatile memory used for storing data, metadata, computer-readable or machine-readable instructions, and / or programs for execution by the processor(s) for operating as described herein. The storage device 406 may include storage, such as a hard disk, flash disk drive, or other digital storage device, for storing data or instructions for performing the methods described herein.

[0130] The memory 404 may comprise a computer-readable storage media or machine- readable storage media that maintains computer-executable instructions for performing one or more as described herein. For example, the memory 404 may comprise computer-executable instructions or machine-readable instructions that include one or more portions of the procedures described herein. The processor 402 of the device 400 may access the instructions from memory for being executed to cause the processor 402 of the device 400 to operate as described herein. The memory 404 may comprise computer-executable instructions for executing configuration software. For example, the computer-executable instructions may be executed to perform, in part and / or in their entirety, one or more procedures as described herein. Further, the memory 404 may have stored thereon one or more settings and / or control parameters associated with the device 400.

[0131] The I / O interface 408 may allow a user to provide input to, receive output from, and / or otherwise transfer data to and receive data from the computing device 400. The VO interface 408 may include a mouse, a keypad or a keyboard, a touch screen, a camera, an optical scanner, network interface, modem, other known I / O devices or a combination of such I / O interfaces. The I / O interface 408 may include one or more devices for presenting output to a user, including, but not limited to, a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., display drivers), one or more audio speakers, and one or more audio drivers. The I / O interface 408 may be configured to provide graphical data to a display for presentation to a user. The graphical data may be representative of one or more graphical user interfaces and / or any other graphical content (e.g., any combination of the prompts described herein).

[0132] The communication interface 410 may include hardware, software, or both. In any event, the communication interface 410 may provide one or more interfaces for communication (such as, for example, packet-based communication) between the computing device 400 and one or more other computing devices or networks. The communication may be a wired or wireless communication. As an example, and not by way of limitation, the communication interface 410 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as a WI-FI.

[0133] Additionally, the communication interface 410 may facilitate communications with various types of wired or wireless networks. The communication interface 410 may alsofacilitate communications using various communication protocols. The communication infrastructure 412 may also include hardware, software, or both that couples components of the computing device 400 to each other. For example, the communication interface 410 may use one or more networks and / or protocols to enable a plurality of computing devices connected by a particular infrastructure to communicate with each other to perform one or more aspects of the processes described herein.

[0134] In addition to what has been described herein, the methods and systems may also be implemented in a computer program(s), software, or firmware incorporated in one or more computer-readable media for execution by a computer(s) or processor(s), for example. Examples of computer-readable media include electronic signals (transmitted over wired or wireless connections) and tangible / non-transitory computer-readable storage media. Examples of tangible / non-transitory computer-readable storage media include, but are not limited to, a read only memory (ROM), a random-access memory (RAM), removable disks, and optical media such as CD-ROM disks, and digital versatile disks (DVDs).

[0135] While this disclosure has been described in terms of certain embodiments and generally associated methods, alterations and permutations of the embodiments and methods will be apparent to those skilled in the art. Accordingly, the above description of example embodiments does not constrain this disclosure. Other changes, substitutions, and alterations are also possible without departing from the spirit and scope of this disclosure.

Claims

CLAIMSWhat is claimed is:

1. A method for assessing severity of involuntary movement associated with tardive dyskinesia (TD), the method comprising: receiving video data of the patient, wherein the video data corresponds to facial movements of the patient; processing the video data to identify facial landmark data of the patient; generating a plurality of long segments from the video data, wherein portions of the video data are excluded from the plurality of long segments, such that the plurality of long segments are defined by durations of the video data between the excluded portions of the video data; generating a plurality of short segments based on the plurality of long segments; applying, for each short segment of the plurality of short segments, a first trained machine learning model to the facial landmark data of the short segment to generate a movement indicator for each of the plurality of short segments, wherein each movement indicator is indicative of movement of a category of facial movement within the short segment; and applying a second trained machine learning model to the movement indicators of the short segments to determine a total severity score for involuntary movement associated with TD.

2. The method of claim 1, further comprising: filtering the video data using the facial landmark data to generate the plurality of long segments.

3. The method of claim 1, wherein the excluded portions of the video data comprise portions where the patient is not looking at the camera, where facial landmark data of the patient is not identified within the video data, or where the patient is not performing a predefined task.

4. The method of claim 1, wherein the plurality of long segments are defined by multiple different durations, and wherein each short segment is defined by a predefined duration.

5. The method of claim 4, wherein the predefined duration is about four seconds.

6. The method of claim 4, wherein the plurality of short segments comprise short segments that overlap with at least one other short segment.

7. The method of claim 1 , wherein the first trained machine learning model comprises a plurality of first trained machine learning models, wherein each of the first trained machine learning models generates a movement indicator for each of the plurality of short segments corresponding to a different category of facial movement.

8. The method of claim 7, wherein the plurality of trained machine learning models are trained using training data comprising facial landmark data extracted from videos from a plurality of patients with a positive TD diagnosis, Abnormal Involuntary Movement Scale (AIMS) scores associated with each patient of the plurality of patients, and labels corresponding to each of the videos indicating a presence of a plurality of distinct category of facial movement.

9. The method of claim 8, wherein the categories of facial movement comprise mouth and lower face movement, upper face movement, head tremor, other tremor, head movement, tongue movement, other movement, contractures and tilts, and number of blinks.

10. The method of claim 8, wherein the videos are divided into a plurality of short segments having a predefined duration, wherein the labels are applied to each of the plurality of short segments.

11. The method of claim 10, wherein the predefined duration is about four seconds.

12. The method of claim 1, wherein the categories of facial movement comprise one or more of eye movement, mouth movement, or head movement.

13. The method of claim 1, wherein a scale associated with the plurality of movement indicators is different from a scale associated with the total severity score.

14. The method of claim 13, wherein the scale associated with the plurality of movement indicators is a range from 0.0- 1.0 or a binary scale.

15. The method of claim 13, wherein the scale associated with the total severity score is a range of 0-28.

16. The method of claim 1, further comprising: generating a movement severity score for each of the plurality of long segments based on the movement indicators of the short segments associated with the long segment, wherein each movement severity score indicates a severity of movement of the category of facial movement within the long segment.

17. The method of claim 16, wherein the movement severity scores for each of the plurality of long segments is determined using a third trained machine model.

18. The method of claim 16, wherein the second trained machine learning model determines the total severity score for involuntary movement associated with TD based on the movement severity scores associated with the plurality of long segments.

19. The method of claim 16, further comprising: generating, via a graphical user interface (GUI) displayed on a display device, an indication of the movement severity scores for each category of facial movement associated with each of the plurality of long segments.

20. The method of claim 19, wherein the indication of the movement severity scores for each category of facial movement is color coded based on a value of each respective movement severity score.

21. The method of claim 19, further comprising: displaying, in response to a user selection of a long segment of the plurality of long segments, the video data associated with the selected long segment via the display device.

22. The method of claim 21, further comprising: displaying, with the video data associated with the selected long segment, a point cloud associated with the facial landmark data associated with the selected long segment.

23. The method of claim 16, wherein a scale associated with the total severity score, a scale associated with the plurality of movement indicators, and a scale associated with the plurality of movement severity scores are all different from one another.

24. The method of claim 1, wherein processing the video data comprises applying an image processing technique to the video data to identify facial landmarks in one or more frames of the video data to produce labeled frames.

25. The method of claim 24, wherein the image processing technique comprises a dlib facial recognition method, a Haar Cascade method, a Fisherface method, or an elastic graph matching method.

26. The method of claim 24, wherein processing the video data further comprises identifying the facial landmark data that corresponds to each of the facial landmarks from the labeled frames.

27. The method of claim 24, wherein the facial landmark data comprises one or more of a landmark identifier, landmark coordinate position, frame source identifier, video source identifier, or patient identifier.

28. The method of claim 1, wherein the video data corresponds to facial movements of the patient during an assessment session.

29. The method of claim 1, wherein the total severity score corresponds to an Abnormal Involuntary Movement Scale (AIMS) score.

30. The method of claim 1, further comprising: providing a prompt to the patient to perform an action.

31. The method of claim 30, further comprising: providing the prompt to the patient by displaying an instruction to perform the action on a user interface of an electronic device.

32. The method of claim 1, further comprising: capturing one or more videos of the patient, wherein the video data of the patient comprises the one or more videos.

33. The method of claim 1, wherein the total severity score corresponds to a risk of receiving a positive TD diagnosis by a trained physician.

34. The method of claim 33, further comprising: generating a notification, wherein the notification comprises a recommendation to refer the patient to a trained clinician when the total severity score is above a predetermined threshold.

35. The method of claim 1, further comprising generating a notification indicating the total severity score.

36. The method of claim 35, further comprising displaying the notification on the display device.

37. The method of claim 1, wherein each movement indicator indicates a probability that movement of a category of facial movement occurred within the short segment.

38. The method of claim 1, further comprising: administering to the patient a daily amount of a medication associated with TD subsequent to receiving the video data.

39. The method of claim 38, further comprising: determining that the total severity score is above a threshold; and generating a recommendation that a dosing regimen of a medication of the patient be adjusted based on the total severity score of the patient being above the threshold.

40. The method of claim 38, wherein the medication comprises a vesicular monoamine transporter 2 (VMAT2) inhibitor.

41. The method of claim 40, wherein the VMAT2 inhibitor comprises deutetrabenazine, tetrabenazine, or valbenazine.

42. The method of claim 38, further comprising: receiving second video data of the patient, wherein the second video data corresponds to facial movements of the patient that occur at a time subsequent to administering the first subsequent daily amount of the medication; processing the second video data to identify second facial landmark data of the patient; generating a plurality of second long segments from the second video data, wherein portions of the second video data are excluded from the plurality of second long segments, such that the plurality of second long segments are defined by durations of the second video data between the excluded portions of the second video data; generating a plurality of second short segments based on the plurality of second long segments; applying, for each second short segment of the plurality of second short segments, the first trained machine learning model to the second facial landmark data of the second short segment to generate a second movement indicator for each of the plurality of second short segments; applying the second trained machine learning model to the second movement indicators of the second short segments to determine a second total severity score for involuntary movement associated with TD; and increasing the daily amount of the medication to a second subsequent daily amount based on the second total severity score of the patient being above a threshold.

43. The method of claim 42, wherein the daily amount is a first daily amount administered pursuant to a titration schedule associated with the VMAT2 inhibitor, the method further comprising: increasing the daily amount of the medication to a first subsequent daily amount when the second total severity score of the patient is above a threshold.

44. The method of claim 43, wherein the first daily amount is at least about 6 mg / day, and the medication is deutetrabenazine, and the first subsequent daily amount is at least about 6 mg / day more than the first daily amount.

45. The method of claim 1, further comprising: receiving acoustic data of the patient; generating short segments of acoustic data based on the acoustic data; applying, to each short segment of acoustic data, a fourth trained machine learning model to the acoustic data to generate an acoustic-specific total severity score for involuntary movement associated with TD or an indication of a progression of TD.46 The method of claim 45, wherein a length of each short segment of acoustic data is different from a length of each short segment of video data.

47. The method of claim 45, wherein the length of each short segment of acoustic data is in the range of about 50-100 milliseconds, and wherein the length of each short segment of video data is about four seconds.

48. The method of claim 45, wherein the acoustic data is associated with one or more voice biomarkers of the patient, and wherein the fourth trained machine learning model is configured to generate the acoustic-specific total severity score for involuntary movement associated with TD or the indication of a progression of TD based on the one or more voice biomarkers of the patient.

49. A method for titrating a medication of a patient associated with tardive dyskinesia (TD), the method comprising: administering to the patient a daily amount of a vesicular monoamine transporter 2 (VMAT2) inhibitor; receiving video data of the patient, wherein the video data corresponds to facial movements of the patient; processing the video data to identify facial landmark data of the patient; generating a plurality of long segments from the video data, wherein portions of the video data are excluded from the plurality of long segments, such that the plurality of long segments are defined by durations of the video data between the excluded portions of the video data; generating a plurality of short segments based on the plurality of long segments; applying, for each short segment of the plurality of short segments, a first trained machine learning model to the facial landmark data of the short segment to generate a movement indicator for each of the plurality of short segments, wherein each movement indicator indicates a probability that movement of a category of facial movement occurred within the short segment; and applying a second trained machine learning model to the movement indicators of the short segments to determine a total severity score for involuntary movement associated with TD.

50. The method of claim 49, further comprising: increasing the daily amount of the VMAT2 inhibitor based at least in part on the total severity score of the patient being above a threshold.

51. The method of claim 49, wherein the VMAT2 inhibitor comprises deutetrabenazine, tetrabenazine, or valbenazine.

52. The method of claim 49, further comprising: receiving second video data of the patient, wherein the second video data corresponds to facial movements of the patient that occur at a time subsequent to administering the first subsequent daily amount of the medication; processing the second video data to identify second facial landmark data of the patient;generating a plurality of long segments from the video data, wherein portions of the video data are excluded from the plurality of long segments, such that the plurality of long segments are defined by durations of the video data between the excluded portions of the video data; generating a plurality of short segments based on the plurality of long segments; applying, for each second short segment of the plurality of second short segments, the first trained machine learning model to the second facial landmark data of the second short segment to generate a second movement indicator for each of the plurality of second short segments; applying the second trained machine learning model to the second movement indicators of the second short segments to determine a second total severity score for involuntary movement associated with TD; and increasing the daily amount of the medication to a second subsequent daily amount based on the second total severity score of the patient being above a threshold.

53. The method of claim 52, wherein the daily amount is a first daily amount administered pursuant to a titration schedule associated with the VMAT2 inhibitor, the method further comprising: increasing the daily amount of the medication to a first subsequent daily amount when the second total severity score of the patient is above a threshold.

54. The method of claim 53, wherein the first daily amount is at least about 6 mg / day, and the medication is deutetrabenazine, and the first subsequent daily amount is at least about 6 mg / day more than the first daily amount.

55. A system for assessing severity of involuntary movement associated with tardive dyskinesia (TD), the system comprising: one or more processors configured to: receive video data of the patient, wherein the video data corresponds to facial movements of the patient; process the video data to identify facial landmark data of the patient;generate a plurality of long segments from the video data, wherein portions of the video data are excluded from the plurality of long segments, such that the plurality of long segments are defined by durations of the video data between the excluded portions of the video data; generate a plurality of short segments based on the plurality of long segments; apply, for each short segment of the plurality of short segments, a first trained machine learning model to the facial landmark data of the short segment to generate a movement indicator for each of the plurality of short segments, wherein each movement indicator is indicative of movement of a category of facial movement within the short segment; and apply a second trained machine learning model to the movement indicators of the short segments to determine a total severity score for involuntary movement associated with TD.

56. The system of claim 55, wherein the one or more processors are configured to: filter the video data using the facial landmark data to generate the plurality of long segments.

57. The system of claim 55, wherein the excluded portions of the video data comprise portions where the patient is not looking at the camera, where facial landmark data of the patient is not identified within the video data, or where the patient is not performing a predefined task.

58. The system of claim 55, wherein the plurality of long segments are defined by multiple different durations, and wherein each short segment is defined by a predefined duration.

59. The system of claim 58, wherein the predefined duration is about four seconds.

60. The system of claim 55, wherein the plurality of short segments comprise short segments that overlap with at least one other short segment.

61. The system of claim 55, wherein the first trained machine learning model comprises a plurality of first trained machine learning models, wherein each of the first trained machine learning models generates a movement indicator for each of the plurality of short segments corresponding to a different category of facial movement.

62. The system of claim 61, wherein the plurality of trained machine learning models are trained using training data comprising facial landmark data extracted from videos from a plurality of patients with a positive TD diagnosis, Abnormal Involuntary Movement Scale (AIMS) scores associated with each patient of the plurality of patients, and labels corresponding to each of the videos indicating a presence of a plurality of distinct category of facial movement.

63. The system of claim 62, wherein the categories of facial movement comprise mouth and lower face movement, upper face movement, head tremor, other tremor, head movement, tongue movement, other movement, contractures and tilts, and number of blinks.

64. The system of claim 62, wherein the videos are divided into a plurality of short segments having a predefined duration, wherein the labels are applied to each of the plurality of short segments.

65. The system of claim 64, wherein the predefined duration is about four seconds.

66. The system of claim 55, wherein the categories of facial movement comprise one or more of eye movement, mouth movement, or head movement.

67. The system of claim 55, wherein a scale associated with the plurality of movement indicators is different from a scale associated with the total severity score.

68. The system of claim 67, wherein the scale associated with the plurality of movement indicators is a range from 0.0-1.0 or a binary scale.

69. The system of claim 67, wherein the scale associated with the total severity score is a range of 0-28.

70. The system of claim 55, wherein the one or more processors are configured to:generate a movement severity score for each of the plurality of long segments based on the movement indicators of the short segments associated with the long segment, wherein each movement severity score indicates a severity of movement of the category of facial movement within the long segment.

71. The system of claim 70, wherein the movement severity scores for each of the plurality of long segments is determined using a third trained machine model.

72. The system of claim 70, wherein the second trained machine learning model determines the total severity score for involuntary movement associated with TD based on the movement severity scores associated with the plurality of long segments.

73. The system of claim 70, wherein the one or more processors are configured to: generate, via a graphical user interface (GUI) displayed on a display device, an indication of the movement severity scores for each category of facial movement associated with each of the plurality of long segments.

74. The system of claim 73, wherein the indication of the movement severity scores for each category of facial movement is color coded based on a value of each respective movement severity score.

75. The system of claim 73, wherein the one or more processors are configured to: display, in response to a user selection of a long segment of the plurality of long segments, the video data associated with the selected long segment via the display device.

76. The system of claim 75, wherein the one or more processors are configured to: display, with the video data associated with the selected long segment, a point cloud associated with the facial landmark data associated with the selected long segment.

77. The system of claim 70, wherein a scale associated with the total severity score, a scale associated with the plurality of movement indicators, and a scale associated with the plurality of movement severity scores are all different from one another.

78. The system of claim 55, wherein the one or more processors are configured to apply an image processing technique to the video data to identify facial landmarks in one or more frames of the video data to produce labeled frames to process the video data.

79. The system of claim 78, wherein the image processing technique comprises a dlib facial recognition method, a Haar Cascade method, a Fisherface method, or an elastic graph matching method.

80. The system of claim 78, wherein the one or more processors are configured to identify the facial landmark data that corresponds to each of the facial landmarks from the labeled frames to process the video data.

81. The system of claim 78, wherein the facial landmark data comprises one or more of a landmark identifier, landmark coordinate position, frame source identifier, video source identifier, or patient identifier.

82. The system of claim 55, wherein the video data corresponds to facial movements of the patient during an assessment session.

83. The system of claim 55, wherein the total severity score corresponds to an Abnormal Involuntary Movement Scale (AIMS) score.

84. The system of claim 55, wherein the one or more processors are configured to: provide a prompt to the patient to perform an action.

85. The system of claim 84, wherein the one or more processors are configured to:provide the prompt to the patient by displaying an instruction to perform the action on a user interface of an electronic device.

86. The system of claim 55, wherein the one or more processors are configured to: capture one or more videos of the patient, wherein the video data of the patient comprises the one or more videos.

87. The system of claim 55, wherein the total severity score corresponds to a risk of receiving a positive TD diagnosis by a trained physician.

88. The system of claim 87, wherein the one or more processors are configured to: generate a notification, wherein the notification comprises a recommendation to refer the patient to a trained clinician when the total severity score is above a predetermined threshold.

89. The system of claim 55, wherein the one or more processors are configured to generate a notification indicating the total severity score.

90. The system of claim 89, wherein the one or more processors are configured to display the notification on the display device.

91. The system of claim 55, wherein each movement indicator indicates a probability that movement of a category of facial movement occurred within the short segment.

92. The system of claim 55, wherein the one or more processors are configured to: administer to the patient a daily amount of a medication associated with TD subsequent to receiving the video data.

93. The system of claim 92, wherein the one or more processors are configured to: determine that the total severity score is above a threshold; and generate a recommendation that a dosing regimen of a medication of the patient be adjusted based on the total severity score of the patient being above the threshold.

94. The system of claim 92, wherein the medication comprises a vesicular monoamine transporter 2 (VMAT2) inhibitor.

95. The system of claim 94, wherein the VMAT2 inhibitor comprises deutetrabenazine, tetrabenazine, or valbenazine.

96. The system of claim 92, wherein the one or more processors are configured to: receive second video data of the patient, wherein the second video data corresponds to facial movements of the patient that occur at a time subsequent to administering the first subsequent daily amount of the medication; process the second video data to identify second facial landmark data of the patient; generate a plurality of second long segments from the second video data, wherein portions of the second video data are excluded from the plurality of second long segments, such that the plurality of second long segments are defined by durations of the second video data between the excluded portions of the second video data; generate a plurality of second short segments based on the plurality of second long segments; apply, for each second short segment of the plurality of second short segments, the first trained machine learning model to the second facial landmark data of the second short segment to generate a second movement indicator for each of the plurality of second short segments; apply the second trained machine learning model to the second movement indicators of the second short segments to determine a second total severity score for involuntary movement associated with TD; and generate a notification that recommends an increase of the daily amount of the medication to a second subsequent daily amount based on the second total severity score of the patient being above a threshold.

97. The system of claim 96, wherein the daily amount is a first daily amount administered pursuant to a titration schedule associated with the VMAT2 inhibitor, and wherein the one or more processors are configured to:generate a notification that recommends an increase of the daily amount of the medication to a first subsequent daily amount when the second total severity score of the patient is above a threshold.

98. The system of claim 97, wherein the first daily amount is at least about 6 mg / day, and the medication is deutetrabenazine, and the first subsequent daily amount is at least about 6 mg / day more than the first daily amount.

99. The system of claim 55, wherein the one or more processors are configured to: receive acoustic data of the patient; generate short segments of acoustic data based on the acoustic data; apply, to each short segment of acoustic data, a fourth trained machine learning model to the acoustic data to generate an acoustic-specific total severity score for involuntary movement associated with TD or an indication of a progression of TD.100 The system of claim 99, wherein a length of each short segment of acoustic data is different from a length of each short segment of video data.

101. The system of claim 99, wherein the length of each short segment of acoustic data is in the range of about 50-100 milliseconds, and wherein the length of each short segment of video data is about four seconds.

102. The system of claim 99, wherein the acoustic data is associated with one or more voice biomarkers of the patient, and wherein the fourth trained machine learning model is configured to generate the acoustic-specific total severity score for involuntary movement associated with TD or the indication of a progression of TD based on the one or more voice biomarkers of the patient.

Citation Information

Patent Citations

  • Methods for the treatment of abnormal involuntary movement disorders

    US20160287574A1

  • Methods for the treatment of abnormal involuntary movement disorders

    WO2016144901A1

  • Measuring body movement in movement disorder disease

    US20190110736A1

  • Movement Disorder Diagnostics from Video Data Using Body Landmark Tracking

    US20230027320A1

  • Neural network architecture for movement analysis

    US20230410290A1