Machine learning classification of videos to determine movement disorder symptoms

Machine learning models analyze video and audio data to efficiently and accurately identify movement disorders, addressing the inefficiencies of traditional manual screening methods.

JP2025531293APending Publication Date: 2025-09-19VIDERA HEALTH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025516203
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-09-12
Filing Date
2023-09-13
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Traditional manual screening and assessment of movement disorders, such as Huntington's disease, Parkinson's disease, and tardive dyskinesia, are time-consuming, require specialized training, and are prone to test bias, making them inefficient and unreliable.

Method used

Utilizing machine learning models, particularly neural networks, to analyze video and audio data of patients to detect movement disorder symptoms, enabling rapid and accurate identification and scoring of these disorders.

Benefits of technology

The machine learning models provide efficient, unbiased, and timely detection of movement disorders, reducing the need for manual screening and enabling early diagnosis and treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025531293000001_ABST
    Figure 2025531293000001_ABST
Patent Text Reader

Abstract

The method includes acquiring, by a processing device, video data of a patient, the video data including image data and audio data. The method further includes providing, by the processing device, the video data to a first trained machine learning model. The method further includes acquiring output from the first trained machine learning model based on the video data, the output including a first indication that the patient exhibits symptoms of one or more target movement disorders in association with the video data. The method further includes providing a warning to a user indicative of the one or more target movement disorders.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to methods related to machine learning models used for the assessment of movement disorders. In particular, the present disclosure relates to machine learning models used for classification based on recorded data to determine movement disorders. [Background technology]

[0002] Patients may experience a variety of movement disorders that may cause unintended, unwanted, and / or involuntary movement or an inability to move one or more body parts as the patient intends. Identifying such movement disorders can be costly in terms of time, expertise, monetary expense, etc. Summary of the Invention [Means for solving the problem]

[0003] The following is a simplified summary of the disclosure to provide a basic understanding of some aspects of the disclosure. This summary is not an extensive overview of the disclosure. It is not intended to identify key or critical elements of the disclosure or to delineate the scope of any particular embodiments of the disclosure or the scope of the claims. Its sole purpose is to present some concepts of the disclosure in a simplified form as a prelude to the more detailed description that is presented later.

[0004] In one aspect of the present disclosure, a method includes acquiring, by a processing device, video data of a patient, the video data including image data and audio data. The method further includes providing, by the processing device, the video data to a first trained machine learning model. The method further includes acquiring output from the first trained machine learning model based on the video data, the output including a first indication that the patient exhibits symptoms of one or more target movement disorders in association with the video data. The method further includes providing a warning to a user indicative of the one or more target movement disorders.

[0005] In another aspect of the present disclosure, a method includes acquiring a first plurality of video data of a first plurality of patients. The method further includes performing a cropping operation on each of the first plurality of video data to generate a first plurality of video data crops. The method further includes receiving a first plurality of labels associated with each of the first plurality of video data crops. The first plurality of labels include a first indication of the presence or absence of evidence of a movement disorder within the first plurality of video data crops. The method further includes training a first machine learning model by providing the first plurality of video data crops as training inputs and the first plurality of labels as target outputs. The first machine learning model is configured to generate an output indicating whether the input video data crops include an indication of a movement disorder.

[0006] In another aspect of the present disclosure, a non-transitory machine-readable storage medium is disclosed. The storage medium stores instructions that, when executed, cause a processing device to perform operations. The operations include acquiring video data of a patient, the video data including image data and audio data. The operations further include providing the video data to a first trained machine learning model. The operations further include obtaining output from the first trained machine learning model based on the video data. The output includes a first indication that the patient exhibits symptoms of one or more target movement disorders in association with the video data. The operations further include providing a warning to a user indicating the one or more target movement disorders.

[0007] The present disclosure is illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a block diagram illustrating an exemplary system architecture, according to some embodiments. [Figure 2]FIG. 1 illustrates a model training workflow and a model application workflow, according to some embodiments. [Figure 3] 1A-1C illustrate frames of a video recording as portions of the video for use in determining the presence of a movement disorder, according to some embodiments. [Figure 4A] 1 is a flow diagram of a method for generating a dataset for a machine learning model, according to some embodiments. [Figure 4B] 1 is a flow diagram of a method for utilizing machine learning models in determining whether evidence of a movement disorder is captured in video data, according to some embodiments. [Figure 4C] 1 is a flow diagram of a method for training a machine learning model to make decisions related to movement disorders, according to some embodiments. [Figure 5] FIG. 1 illustrates an operational flow for making decisions related to movement disorders, according to some embodiments. [Figure 6] FIG. 1 is a block diagram illustrating a computer system according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0009] Aspects of the present disclosure relate to utilizing a machine learning model to determine whether a patient's video data contains evidence of one or more movement disorder symptoms. Exemplary movement disorders include Huntington's disease, Parkinson's disease, and tardive dyskinesia (TD). The output of the machine learning model may be utilized in screening a patient for further evaluation for a movement disorder. The output of the machine learning model may be utilized in diagnosing a movement disorder. The output of the machine learning model may be utilized in treating a movement disorder. The output of the machine learning model may be utilized in determining the effectiveness of a movement disorder treatment. The output of the machine learning model may be utilized in improving a movement disorder treatment.

[0010] In traditional diagnosis and treatment of movement disorders, manual screening / assessments may be performed. For example, a physician may administer a test to generate a score that indicates the likelihood that a patient is experiencing symptoms of a movement disorder. One such test is the abnormal involuntary movement scale (AIMS) test. Scoring a manual test for movement disorders can be a time-consuming process (e.g., it may take 30 minutes to an hour), may require specialized training of the physician, may be subject to test bias, etc.

[0011]

[0003] Aspects of the present disclosure allow for training and / or utilizing one or more machine learning models to measure and / or determine the presence of movement disorder symptoms from a patient's video. Aspects of the present disclosure may be applicable to disorders other than TD, whose symptoms are detectable in a patient's video. In some cases, movement disorders may be intermittent, e.g., involuntary movements or loss of muscle control may occur infrequently, sporadically, irregularly, etc.

[0012] In some embodiments, a machine learning model is trained for movement disorder symptom detection. The machine learning model may be trained, validated, selected, tested, etc. The machine learning model may be trained using a training dataset, validated using a validation dataset, tested using a test dataset, etc. The machine learning model may be a neural network, such as a convolutional neural network. The machine learning model may be a trained machine learning model; for example, the training dataset may be a labeled dataset (e.g., labeled with a target output).

[0013] The training dataset (as well as the validation and test sets) may be manually labeled. The dataset may include as output a video of the patient, including classifications such as whether the patient is experiencing a dyskinetic symptom, whether the dyskinetic symptom is visible in the video, whether the dyskinetic symptom is detectable by audio, etc. The dataset may be labeled via a scoring system, e.g., with respect to how many instances of a dyskinetic symptom are shown, how severe the patient's dyskinetic symptom is, etc. The dataset may be labeled via a clinical scoring system, e.g., the dataset may be labeled with the AIMS score of the patient or video. Clinical scoring may be performed by an expert, such as a physician.

[0014] The training dataset may include videos with portions manually flagged to demonstrate potential movement disorder symptoms. For example, a video may have portions labeled as including involuntary muscle movements. A video may have portions labeled as including sounds indicative of movement disorder symptoms. A video may have portions labeled as including speech / language indicative of movement disorder symptoms. The labeled portions may be short, e.g., approximately one second in duration. The labeled portions may be of any length. Sections of the video that are not labeled as including symptomatic behavior may be separated into portions. For example, a training dataset may include some portions labeled as including symptomatic behavior and some portions not labeled as including symptomatic behavior. The portions that do not include symptomatic behavior may be separated from the rest of the video at irregular intervals. The portions that do not include symptomatic behavior may be approximately the same length as the portions that include symptomatic behavior (e.g., may match similar length statistics, such as mean deviation, standard deviation, etc., as the labeled portions). The portions that do not contain symptomatic behavior may be generally similar to the portions labeled as containing symptomatic behavior, e.g., in that they are spread across multiple locations throughout the video from which they originate. The portions that do not contain symptomatic behavior may be provided for training, testing, and / or validation to determine whether they meet a target or threshold percentage of the data provided to the model. For example, an equal number of portions that contain symptomatic behavior and portions that do not contain symptomatic behavior may be provided, twice as many portions as portions that do not contain symptomatic behavior may be provided, etc.

[0015] In some embodiments, the training dataset (and / or the validation and / or test dataset) may include videos targeted at a certain symptomatic area. The training videos may include a more limited or focused perspective, e.g., videos belonging to the dataset may be cropped to a particular symptomatic area. The videos belonging to the dataset may be cropped to a landmark area of ​​the body. For example, the training videos may include the patient's eyes, eyebrows, face (e.g., facial muscles), head, chin, lips, hands, upper limbs, lower limbs, and / or torso. The training videos may be cropped to include the target area. For example, the training dataset may include videos that have been cropped to include a single symptomatic area. In some embodiments, multiple machine learning models may be trained in relation to different symptomatic areas. By providing training data including videos of symptomatic areas, machine learning models related to symptomatic regions / areas may be generated.

[0016] In some embodiments, the training dataset (and / or validation and / or test dataset) may include additional information, such as contextual information. The additional data may provide information that influences the onset of a patient's movement disorder symptoms. The additional data may include any information that influences or is suspected to influence the detectable presence of a movement disorder symptom from the video. The additional data may include medical history, medication history, time of video, etc. For example, the additional data may include a history of medications taken by the patient, including medications that may cause movement disorders, medications to treat movement disorders, etc. The additional data may include a historical record of movement disorder symptoms.

[0017] In some embodiments, the training dataset (and / or validation and / or test dataset) may include noise, e.g., to generalize results, generalize use of the model, etc. For example, one or more videos in the dataset may be altered, noise may be introduced, etc. The videos may be altered by cropping, rotating, distorting, recoloring, decolorizing, blurring, filtering, posterizing, smoothing, or otherwise altering the patient video and / or audio associated with the patient video. Introducing noise into the training dataset may reduce the reliability of the trained machine learning model to the quality of the video and / or audio recordings.

[0018] In some embodiments, the one or more trained machine learning models may be configured to encourage further screening of the patient for a movement disorder. The one or more trained machine learning models may be configured to generate a likelihood that the patient is exhibiting symptoms of a movement disorder. The one or more trained machine learning models may be configured to generate a score indicating the likelihood or severity of the movement disorder symptoms. The one or more trained machine learning models may be configured to generate an AIMS score based on one or more videos of the patient.

[0019] Results, recommendations, and / or predictions can be generated by providing one or more videos of a patient to one or more machine learning models. The videos may include prompting the patient to respond to one or more questions or requests. The questions may include diagnostic questions, such as those asked during an AIMS test. The questions may include open-ended questions to elicit a spoken response. The videos may include recordings of the patient's responses to various requests, such as having the patient sit still, stand, walk, position their face, lips, chin, or tongue, clap their hands, etc. The videos may include one or more parts of the participant's body, potentially including hands, torso, face, eyes, chin, tongue, lips, etc. The videos may include one or more target parts of the patient's body.

[0020] One or more videos may be provided to one or more trained machine learning models. The one or more trained machine learning models may perform one or more of several operations. The operations performed by the trained machine learning models may include segmenting the video. The segmentation may include spatial segmentation. For example, a portion of the video may be cropped to highlight or accentuate a particular body part, a symptomatic area, etc. The video recording data may be cropped to include one or more target parts of the patient's body. The segmentation may include temporal segmentation. For example, a portion of the video including the target body part may be separated from a portion that does not include the body part. The portions of the video may be provided to corresponding machine learning models. For example, a portion of the video that has been temporally and / or spatially cropped to the target body part may be provided to a machine learning model associated with the target body part. Multiple temporal and / or spatial portions of the video may be provided to different trained machine learning models associated with different body parts, landmark areas of the body, etc.

[0021] The operations performed by the trained machine learning model may include detecting portions of the video containing evidence of a movement disorder symptom. The machine learning model may detect involuntary movements. The machine learning model may detect patient movements indicative of one or more movement disorders. The machine learning model may isolate, flag, etc., portions of the video containing movements indicative of a movement disorder. The operations performed by the trained machine learning model may further include isolating portions of the video that do not contain evidence of a movement disorder symptom. The machine learning model may utilize both video data that exhibits predicted movement disorder symptoms and video data that does not have predicted movement disorder symptoms, for example, to provide a baseline or comparison.

[0022] The operations performed by the trained machine learning model may include determining whether the video indicates that the patient is experiencing symptoms of a movement disorder. Video data may be provided to the machine learning model to determine whether the video indicates a movement disorder. The machine learning model may provide portions of the video cropped to include, for example, one or more target body parts or target portions of the patient's body, portions of the video flagged as including evidence of a movement disorder, portions of the video flagged as not including evidence of a movement disorder, etc. Additional information that may contribute to the prediction / recommendation may be provided to the machine learning model. For example, additional information that may affect the presence of movement disorder symptoms, such as medical history, medication history, etc., may be provided to the model. The model may use the additional information to improve the accuracy of the prediction, the accuracy of the recommendation, etc.

[0023] Operations performed by the trained machine learning model may include determining a risk, classification, and / or score of a patient experiencing a movement disorder symptom. Predicting risk, classifying symptoms, or determining a score related to a movement disorder may include receiving output from one or more models. For example, a machine learning model (e.g., a fusion model for analyzing the output of several other models) may collect data from different body part-related, voice-related, and / or speech-related models and utilize the machine learning output itself to make a prediction regarding a patient's risk of experiencing a movement disorder symptom. Operations performed by the trained machine learning model may include providing a classification, such as whether to recommend further screening. Execution performed by the trained machine learning model may include providing a risk factor, such as the likelihood that a patient is experiencing a movement disorder symptom. Operations performed by the trained machine learning model may include providing a score, such as a predicted AIMS score.

[0024] Operations for screening, diagnosing, and / or treating patients for movement disorders may include storing historical data, accessing historical data, processing historical data, etc. For example, a patient's AIMS score as determined by the methods and systems of the present disclosure may be tracked over time to determine the severity of a progressive movement disorder, the effectiveness of a treatment plan, etc.

[0025] The operations of the present disclosure may be performed by a model, such as a trained machine learning model. The operations of the present disclosure may be performed by multiple models. The operations of the present disclosure may be performed by multiple models working together, e.g., an ensemble model.

[0026] The disclosed operations may be used as a screen for movement disorder symptoms, e.g., the operations may be used to recommend further investigation for movement disorder symptoms exhibited by a patient. The disclosed operations may be used as a diagnostic tool for movement disorders. The disclosed operations may be used as a therapeutic tool, e.g., in assessing the therapeutic effectiveness of a movement disorder treatment plan.

[0027] In one aspect of the present disclosure, a method includes acquiring, by a processing device, video data of a patient, the video data including image data and audio data. The method further includes providing, by the processing device, the video data to a first trained machine learning model. The method includes acquiring output from the first trained machine learning model based on the video data, the output including a first indication that the patient exhibits symptoms of one or more target movement disorders in association with the video data. The method further includes providing a warning to a user indicative of the one or more target movement disorders.

[0028] In another aspect of the present disclosure, a method includes acquiring a first plurality of video data of a first plurality of patients. The method further includes cropping each of the first plurality of video data to generate a first plurality of video data crops. The method further includes receiving a first plurality of labels associated with each of the first plurality of video data crops. The first plurality of labels include a first indication of the presence or absence of evidence of a movement disorder within the first plurality of video data crops. The method further includes training a machine learning model by providing the first plurality of video data crops as training inputs and the first plurality of labels as target outputs. The first machine learning model is configured to generate an output indicating whether the input video data crops include an indication of a movement disorder.

[0029] In another aspect of the present disclosure, a non-transitory machine-readable storage medium is disclosed. The storage medium stores instructions that, when executed, cause a processing device to perform operations. The operations include acquiring video data of a patient, the video data including image data and audio data. The operations further include providing the video data to a first trained machine learning model. The operations further include obtaining output from the first trained machine learning model based on the video data. The output includes a first indication that the patient exhibits symptoms of one or more target movement disorders in association with the video data. The operations further include providing a warning to a user indicating the one or more target movement disorders.

[0030] 1 is a block diagram illustrating an example system 100 (an example system architecture) according to some embodiments. System 100 includes a client device 120, a prediction server 112, and a data store 140. Prediction server 112 may be part of a prediction system 110. Prediction system 110 may further include server machines 170 and 180.

[0031] Client device 120 includes equipment for image data collection 124 and audio data collection 126. For example, client device 120 may include a camera and a microphone. Client device 120 may be configured to generate video data based on image data collection 124 and audio data collection 126. Client device 120 may be utilized to collect video data of a patient, for example, a patient who may be exhibiting symptoms of a movement disorder.

[0032] Image data collection 124 and audio data collection 126 may be utilized in collecting image data 142 and audio data 160. Both image data 142 and audio data 160 may include video recording data of a patient, for example, a patient who may be exhibiting movement disorder symptoms. The patient's video recording data may include historical recordings and current recordings. The patient's video recording data may include recordings generated to provide for a method performed by system 100 or recordings provided to system 100 for other purposes. For example, a virtual appointment with a healthcare professional may be recorded for further action to be performed by one or more components of system 100. In a further example, a movement disorder assessment may be stored, such as an assessment including prompting speech and movement patterns to determine whether a patient is experiencing a movement disorder.

[0033] The historical image data 144 may be associated with a historical data collection, such as a historical video recording. The historical image data 144 may include a series of image frames of a patient. The historical image data 144 may be utilized in configuring a machine learning model to perform one or more tasks. For example, the historical image data 144 may be used to train, validate, test, etc. the machine learning model. The historical image data 144 may be provided to a machine learning model to adjust one or more model parameters that enable the machine learning model to learn. The historical image data 144 may be labeled, for example, classified according to whether evidence of a movement disorder can be observed in a set of data, e.g., a video clip. The historical audio data 164 may share one or more characteristics with the historical image data 144. A portion of the historical audio data 164 and the historical image data 144 together may form a historical video clip.

[0034] The current image data 146 may be image data associated with the records of a patient to be screened for movement disorders. The current image data 146 may be utilized as input to a trained machine learning model. The current audio data 166 may share one or more features with the current image data 146. The segmented image data 148 may include a spatially and / or temporally cropped image, a series of images, a video clip, or the like. For example, the output of the spatial cropping 118 or the temporal cropping 116 may be or may include the segmented image data 148. The segmented image data 148 may include historical and / or current image data. The segmented image data 148 may be utilized in constructing a machine learning model, such as being provided as input to the trained machine learning model. The segmented audio data 168 may share one or more features with the segmented image data 148. The segmented audio data 168 may be provided as input to the trained machine learning model. In some embodiments, the audio data and corresponding image data (e.g., video data) may be provided as input to one or more trained machine learning models.

[0035] In some embodiments, the image data 142 and / or audio data 160 may be processed (e.g., by the client device 120 and / or by the prediction server 112). The processing of the sensor data 142 may include generating features. In some embodiments, the features are patterns (e.g., slope, width, height, peaks, etc.) within the image data 142 and / or audio data 160 or combinations of values ​​from the image data 142 and audio data 160 (e.g., composite video data, etc.). In some embodiments, the processed data may be spatially and / or temporally cropped data. The processed data may be used to perform signal processing and / or to obtain prediction data 168 for taking corrective actions.

[0036] Each instance (e.g., set) of image data 142 and audio data 160 may correspond to a crop or segment of a video recording, a video recording, a patient (e.g., including multiple recordings), etc. A set of image data 142 and audio data 160 may include video recording data. The data store may further store information associating sets of different data types, such as information indicating that a set of image data 142 and a set of audio data 160 are associated with the same patient.

[0037] In some embodiments, prediction system 110 may generate predicted data 168 using supervised machine learning. For example, predicted data 168 may include output from a machine learning model trained using labeled data, such as video data labeled with a diagnosis of a movement disorder of the subject of the video data. In some embodiments, prediction system 110 may generate predicted data 168 using unsupervised machine learning (e.g., predicted data 168 includes output from a machine learning model trained using unlabeled data, where the output may include clustering results, principal component analysis, anomaly detection, etc.). In some embodiments, prediction system 110 may generate predicted data 168 using semi-supervised learning (e.g., training data may include a mixture of labeled and unlabeled data, etc.).

[0038] Client device 120, prediction server 112, data store 140, server machine 170, and server machine 180 may be coupled to one another via network 130 to generate prediction data 168 for performing corrective actions. In some embodiments, network 130 may provide access to cloud-based services. Operations performed by client device 120, prediction system 110, data store 140, etc. may be performed by a virtual cloud-based device.

[0039] In some embodiments, network 130 is a public network that provides client device 120 with access to prediction server 112, data store 140, and other publicly available computing devices. In some embodiments, network 130 is a private network that provides client device 120 with access to publicly available computing devices, such as data store 140, additional client devices (not shown). Network 130 may include one or more wide area networks (WANs), local area networks (LANs), wired networks (e.g., Ethernet networks), wireless networks (e.g., 802.11 networks or Wi-Fi networks), cellular networks (e.g., Long Term Evolution (LTE) networks), routers, hubs, switches, server computers, cloud computing networks, and / or combinations thereof.

[0040] The client device 120 may include a computing device such as a personal computer (PC), laptop, mobile phone, smartphone, tablet computer, netbook computer, network-connected television ("smart TV"), network-connected media player (e.g., Blu-ray player), set-top box, over-the-top (OTT) streaming device, operator box, etc. The client device 120 may include a corrective action component 122. The corrective action component 122 may receive user input (e.g., via a graphical user interface (GUI) displayed via the client device 120) of instructions related to determining whether a movement disorder is present in the video data. In some embodiments, the corrective action component 122 sends instructions to the prediction system 110, receives output (e.g., predicted data 188) from the prediction system 110, determines corrective actions based on the output, and causes the corrective actions to be implemented. In some embodiments, the corrective action component 122 retrieves data (e.g., current image data 146 from the data store 140, etc.) and provides data related to one or more movement disorder patients to the prediction system 110.

[0041] In some embodiments, the corrective action component 122 receives instructions for corrective actions from the predictive system 110 and causes the corrective actions to be implemented. Each client device 120 may include an operating system that enables a user to one or more of generate, view, or edit data (e.g., instructions related to one or more patients, corrective actions related to one or more patients, etc.).

[0042] In some embodiments, the corrective action includes providing a warning to the user. For example, client device 120 may cause a warning to be displayed indicating to a medical professional that the patient has exhibited symptoms of one or more movement disorders. Client device 120 may indicate that the subject of one or more video data has a greater than threshold likelihood of experiencing tardive dyskinesia, Parkinson's disease, Huntington's chorea, or another movement disorder. In some embodiments, a machine learning model is trained to monitor patient recordings, including, for example, image data and audio data. In some embodiments, the machine learning model may generate a clip, recording, or output indicating a likelihood that the patient is exhibiting symptoms of one or more target movement disorders.

[0043] Prediction server 112, server machine 170, and server machine 180 may each include one or more computing devices such as a rack-mounted server, a router computer, a server computer, a personal computer, a mainframe computer, a laptop computer, a tablet computer, a desktop computer, a graphics processing unit (GPU), an accelerator application-specific integrated circuit (ASIC) (e.g., a tensor processing unit (TPU)), etc. Operations of prediction server 112, server machine 170, server machine 180, data store 140, etc. may be performed by a cloud computing server, a cloud data storage device, etc.

[0044] The prediction server 112 may include a prediction component 114. In some embodiments, the prediction component 114 may receive current image data 146 and / or current audio data 166 (e.g., received from the client device 120 and retrieved from the data store 140) and generate output (e.g., prediction data 168) for performing corrective actions related to one or more patients based on the current data. In some embodiments, the prediction data 168 may include one or more predicted likelihoods that a clip, video recording, or patient is exhibiting symptoms of one or more target movement disorders. In some embodiments, the prediction data 168 may indicate the severity of one or more target movement disorders exhibited in the clip, video recording, or by the patient. The prediction component 114 may utilize one or more trained machine learning models 190 to determine output for performing corrective actions based on the current data.

[0045] One type of machine learning model that can be used to perform some or all of the above tasks is an artificial neural network, such as a deep neural network. An artificial neural network generally includes a feature representation component with a classifier or recurrent layer that maps features to a desired output space. A convolutional neural network (CNN), for example, hosts multiple layers of convolutional filters. The top-layer features extracted by the convolutional layer are mapped to a decision (e.g., a classification output), and pooling can be performed in lower layers, where multilayer perceptrons are typically added, to address nonlinearities.

[0046] A recurrent neural network (RNN) is another type of machine learning model. Recurrent neural network models are designed to interpret a series of inputs, such as time trace data, continuous data, etc., where the inputs are inherently interrelated. The output of a perceptron in an RNN is fed back as an input to that perceptron to generate the next output.

[0047] The Transformer architecture is another type of machine learning model that can be utilized in connection with the present disclosure. The Transformer included an attention mechanism that allows the Transformer to capture relationships between various portions of the input without relying on recurrent layers. The Transformer applies positional coding to various portions of the input data to enable the attention mechanism to determine correlations between various portions of the input data and the importance of the correlations. The attention mechanism increases the importance of some portions of the input data while suppressing signals from unimportant input data.

[0048] Deep learning is a class of machine learning algorithms that uses a cascade of multiple layers of nonlinear processing units for feature extraction and transformation. Each successive layer uses the output from the previous layer as input. Deep neural networks can learn in a supervised (e.g., classification) and / or unsupervised (e.g., pattern analysis) manner. Deep neural networks include a hierarchy of layers, where different layers learn different levels of representation corresponding to different levels of abstraction. In deep learning, each level learns to transform its input data into a more abstract and complex representation. In an image recognition application, for example, the raw input may be a matrix of pixels; the first representation layer may abstract the pixels and encode edges; the second layer may construct and encode the edge configuration; the third layer may encode higher-level shapes (e.g., teeth, lips, gums, etc.); and the fourth layer may recognize the role of the scan. In particular, the deep learning process can independently learn which features are optimally placed within which levels. The "deep" in "deep learning" refers to the number of layers through which data is transformed. More precisely, deep learning systems have significant credit assignment path (CAP) depth. A CAP is a chain of transformations from input to output. A CAP describes the potential causal relationship between input and output. For feedforward neural networks, the CAP depth can be the depth of the network, which can be the number of hidden layers plus one. For recurrent neural networks, where a signal can propagate through layers more than once, the CAP depth is potentially infinite.

[0049] In some embodiments, prediction component 114 receives current image data 146 and / or current audio data 166, performs signal processing to decompose the current data into a set of current data, provides the set of current data as input to trained model 190, and obtains output from trained model 190 indicative of predicted data 168. In some embodiments, prediction server 112 may receive current data (e.g., video data including image data and audio data), provide the current data to temporal cropping 116 to generate a clip and / or provide the current data to spatial cropping 118 to crop an image to a target body part, and provide segmented image data 148 and / or segmented audio data 168 (e.g., segmented by spatial cropping 118 and / or temporal cropping 116) to model 190 to generate predicted data.

[0050] In some embodiments, the various models discussed with respect to model 190 (e.g., supervised machine learning models, unsupervised machine learning models, etc.) may be combined into one model (e.g., an ensemble model) or may be separate models.

[0051] Data may be passed between several separate models included within models 190 and prediction component 114. In some embodiments, some or all of these operations may instead be performed by different devices, e.g., client device 120, server machine 170, server machine 180, etc. Those skilled in the art will understand that variations in data flow, which components perform which processes, which data is provided to which models, etc., are within the scope of this disclosure.

[0052] Data store 140 may be memory (e.g., random access memory), a drive (e.g., a hard drive, a flash drive), a database system, a cloud-accessible memory system, or another type of component or device capable of storing data. Data store 140 may include multiple storage components (e.g., multiple drives or multiple databases) that may span multiple computing devices (e.g., multiple server computers). Data store 140 may store image data 142, audio data 160 (e.g., video data), and prediction data 168.

[0053] In some embodiments, prediction system 110 further includes server machine 170 and server machine 180. Server machine 170 includes dataset generator 172 that is capable of generating datasets (e.g., a set of data inputs and a set of target outputs) for training, validating, and / or testing model 190, including one or more machine learning models. Some operations of dataset generator 172 are described in more detail below with respect to FIGS. 2 and 4A . In some embodiments, dataset generator 172 may partition historical data (e.g., historical image data 144, historical audio data 164) into a training set (e.g., 60% of the historical data), a validation set (e.g., 20% of the historical data), and a test set (e.g., 20% of the historical data).

[0054] In some embodiments, the prediction system 110 generates (e.g., via the prediction component 114) multiple sets of features. For example, a first set of features may correspond to a first set of crops of video data (e.g., a first sequence of included video frames stored in the image data, a first spatial crop targeting a particular part of the patient's body, etc.) corresponding to each of the datasets (e.g., a training set, a validation set, and a test set), and a second set of features may correspond to a second set of crops of video data (e.g., from temporal and / or spatial cropping) corresponding to each of the datasets. The video data may include audio data, image data (e.g., a sequence of images), etc. In some embodiments, the machine learning model may receive as input output from one or more other machine learning models. For example, one or more models may provide information indicative of evidence of a movement disorder in one or more clips of video, a second model may receive the output of the first one or more models and provide an output indicative of the likelihood that the video as a whole will present evidence of a movement disorder, and a third model may receive the output of the second model related to several recorded videos and determine, based on the videos, the likelihood that the patient is experiencing one or more movement disorder symptoms.

[0055] Server machine 180 includes a training engine 182, a validation engine 184, a selection engine 185, and / or a test engine 186. The engines (e.g., training engine 182, validation engine 184, selection engine 185, and test engine 186) may refer to hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, processing device, etc.), software (e.g., instructions executing on a processing device, general-purpose computer system, or dedicated machine), firmware, microcode, or a combination thereof. Training engine 182 may be capable of training a model 190 using one or more sets of features associated with a training set from dataset generator 172. Training engine 182 may generate multiple trained models 190, where each trained model 190 corresponds to a distinct set of features of the training set (e.g., a different collection of video crops, the output of a classification model, etc.). For example, a first trained model may be trained using all features (e.g., X1-X5), a second trained model may be trained using a first subset of features (e.g., X1, X2, X4), and a third trained model may be trained using a second subset of features (e.g., X1, X3, X4, and X5) that may partially overlap with the first subset of features. The dataset generator 172 may receive the output of the trained models, compile the data into training, validation, and test datasets, and use the datasets to train a second model (e.g., a machine learning model configured to output predictions, corrective actions, etc.).

[0056] The validation engine 184 may be capable of validating the trained models 190 using a corresponding set of features of the validation set from the dataset generator 172. For example, a first trained machine learning model 190 trained using a first set of features of the training set may be validated using a first set of features of the validation set. The validation engine 184 may determine the accuracy of each of the trained models 190 based on the corresponding set of features of the validation set. The validation engine 184 may discard trained models 190 having an accuracy that does not meet a threshold accuracy. In some embodiments, the selection engine 185 may be capable of selecting one or more trained models 190 having an accuracy that meets the threshold accuracy. In some embodiments, the selection engine 185 may be capable of selecting the trained model 190 with the highest accuracy of the trained models 190.

[0057] Test engine 186 may be capable of testing trained model 190 using a corresponding set of test set features from dataset generator 172. For example, a first trained machine learning model 190 trained using a first set of training set features may be tested using a first set of test set features. Test engine 186 may determine a trained model 190 that has the highest accuracy of all of the trained models based on the test set.

[0058] In the case of a machine learning model, model 190 may refer to a model artifact created by training engine 182 using a training set that includes data inputs and corresponding target outputs (correct answers for each training input). Patterns in a dataset that map data inputs to target outputs (correct answers) may be found, and machine learning model 190 is provided with a mapping that captures these patterns. Machine learning model 190 may use one or more of support vector machines (SVMs), radial basis functions (RBFs), clustering, supervised machine learning, semi-supervised machine learning, unsupervised machine learning, k-Nearest Neighbor algorithms (k-NNs), linear regression, random forests, neural networks (e.g., artificial neural networks, recurrent neural networks), etc.

[0059] In some embodiments, one or more machine learning models 190 may be trained using historical data (e.g., historical image data 144, historical audio data 164, potentially segmented either spatially and / or temporally).

[0060] The prediction component 114 can provide current data to the model 190 and can run the model 190 on the input to obtain one or more outputs. For example, the prediction component 114 can provide current image data 146 to the model 190 and can run the model 190 on the input to obtain one or more outputs. The prediction component 114 can be capable of determining (e.g., extracting) predicted data 168 from the output of the model 190. The prediction component 114 can determine (e.g., extract) confidence data from the output that indicates a confidence level that the predicted data 168 is an accurate predictor of a process related to the input data for a patient who may be experiencing a movement disorder. The prediction component 114 or the corrective action component 122 can use the confidence data to determine whether to trigger a corrective action based on the predicted data 168.

[0061] The confidence data may include or indicate a confidence level that the prediction data 168 is an accurate prediction of a movement disorder associated with at least a portion of the input data. In one example, the confidence level is a real number between 0 and 1, inclusive, where 0 indicates no confidence that the prediction data 168 is an accurate prediction of one or more movement disorders according to the input data, and 1 indicates absolute confidence that the prediction data 168 will accurately predict one or more movement disorders according to the input data. In response to confidence data indicating a confidence level below a threshold level over a predetermined number of instances (e.g., percentage of instances, frequency of instances, total number of instances, etc.), the prediction component 114 may cause the trained model 190 to be retrained (e.g., based on the current image data 146, the current audio data 166, etc.). In some embodiments, the retraining may include utilizing historical data to generate one or more additional datasets (e.g., via the dataset generator 172).

[0062] Performing comprehensive screening for movement disorders can be costly, inconvenient, training-dependent, and unreliable. For example, screening for tardive dyskinesia can require a trained specialist to schedule an interview with a potential patient and for the specialist to closely observe the patient during the scheduled interview. In addition, an appointment with a healthcare professional who is not specialized in one or more movement disorders may not be sufficient to screen for or diagnose a movement disorder. By providing video data (e.g., from a scheduled or scheduled virtual appointment with a healthcare professional or a video provided by the patient for screening) to a trained machine learning model and receiving an output indicating that the patient may be exhibiting symptoms of a movement disorder, a diagnosis can be made more quickly, a referral to a specialist can be made, and treatment can begin earlier, which can improve the patient's quality of life without the patient needing to regularly meet with a trained specialist in recognizing one or more movement disorders.

[0063] For purposes of illustration, and not limitation, aspects of the present disclosure describe training one or more machine learning models 190 using historical data to determine predictive data 168. In other embodiments, heuristic, physics-based, or rule-based models are used to determine predictive data 168 (e.g., without using a trained machine learning model). In some embodiments, such models may be trained using historical data. In some embodiments, these models may be retrained utilizing a combination of historical and current data. Any of the information described with respect to data input 262 in FIG. 2 may be monitored or otherwise used in a heuristic, physics-based, or rule-based model.

[0064] In some embodiments, the functionality of client device 120, prediction server 112, server machine 170, and server machine 180 may be provided by fewer machines. For example, in some embodiments, server machines 170 and 180 may be combined into a single machine, and in some other embodiments, server machine 170, server machine 180, and prediction server 112 may be combined into a single machine. In some embodiments, client device 120 and prediction server 112 may be combined into a single machine. In some embodiments, the functionality of client device 120, prediction server 112, server machine 170, server machine 180, and data store 140 may be performed by a cloud-based service.

[0065] In general, functionality described in one embodiment as being performed by client device 120, prediction server 112, server machine 170, and server machine 180 may, in other embodiments, be performed on prediction server 112, where appropriate. Additionally, functionality attributed to a particular component may be performed by different or multiple components operating together. For example, in some embodiments, prediction server 112 may determine corrective actions based on prediction data 168. In another example, client device 120 may determine prediction data 168 based on output from a trained machine learning model.

[0066] Additionally, the functionality of a particular component may be performed by different or multiple components working together. One or more of prediction server 112, server machine 170, or server machine 180 may be accessed as a service offered to other systems or devices through an appropriate application programming interface (API).

[0067] Additionally, in some embodiments, multiple devices may perform actions that are attributed to a single component. For example, a patient may record video (e.g., including image data collection 124 and audio data collection 126) on their own device, and a medical professional's device may generate one or more alerts to the user (e.g., via corrective action component 122).

[0068] In embodiments, a "user" may be represented as a single individual. However, other embodiments of the present disclosure include "users" that are entities and / or automated sources controlled by multiple users. For example, a set of individual users organized as a group of administrators may be considered a "user."

[0069] 2 illustrates a model training workflow 205 and a model application workflow 217 for movement disorder determination according to some embodiments of the present disclosure. In embodiments, the model training workflow 205 may be executed on a server that may or may not include a movement disorder prediction data generation application, and the trained model is provided to a prediction component (e.g., on the prediction server 112 of FIG. 1 ), which may execute the model application workflow 217. The model training workflow 205 and the model application workflow 217 may be performed by processing logic executed by a processor of a computing device. One or more of these workflows 205, 217 may be implemented, for example, by one or more machine learning modules implemented by the server 112 of FIG. 1 .

[0070] In some embodiments, the trained machine learning model is a neural network, a decision tree, a random forest model, a support vector machine, or other type of machine learning model.

[0071] In some embodiments, the trained machine learning model may be an artificial neural network (also referred to simply as a neural network). The artificial neural network may be, for example, a convolutional neural network (CNN) or a deep neural network. In some embodiments, the processing logic performs supervised machine learning to train the neural network.

[0072] Artificial neural networks generally include a feature representation component with a classifier or recurrent layer that maps features to a target output space. Convolutional neural networks (CNNs), for example, host multiple layers of convolutional filters. Features extracted by the convolutional layers in the top layer are mapped to decisions (e.g., classification outputs), and pooling is performed in lower layers, where multilayer perceptrons are typically added, to address nonlinearities. Neural networks can be deep networks with multiple hidden layers or shallow networks with zero or a few (e.g., one or two) hidden layers. Deep learning is a class of machine learning algorithms that use a cascade of multiple layers of nonlinear processing units for feature extraction and transformation. Each successive layer uses the output from the previous layer as input. Neural networks can be trained in a supervised (e.g., classification) and / or unsupervised (e.g., pattern analysis) manner. Some neural networks (e.g., deep neural networks) include a hierarchy of layers, where different layers learn different levels of representation corresponding to different levels of abstraction. In deep learning, each level learns to transform its input data into a more abstract and complex representation.

[0073] Training a neural network can be accomplished in a supervised learning fashion, which involves feeding a training dataset consisting of labeled inputs through the network, observing its outputs, defining an error (by measuring the difference between the output and the label values), and using techniques such as deep gradient descent and backpropagation to adjust the network's weights across all its layers and nodes so that the error is minimized. In many applications, repeating this process across many labeled inputs in the training dataset results in a network that can produce correct outputs when presented with inputs that differ from those present in the training dataset. In high-dimensional settings, such as large images, this generalization is achieved when a sufficiently large and diverse training dataset is made available.

[0074] The trained machine learning model may be periodically or continuously retrained to achieve continuous learning and improvement of the trained machine learning model. The model may generate an output based on the input, an action may be performed based on the output, and the result of the action may be measured. In some cases, the result of the action is measured within seconds or minutes, and in some cases, measuring the result of the action takes longer. For example, one or more additional processes may be performed before the result of the action can be measured. The action and the result of the action may indicate whether the output was the correct output and / or the difference between what the output should have been and what the output was. Thus, the action and the result of the action may be used to determine a target output that can be used as a label for the sensor measurement. Once the result of the action is determined, the input to the model (e.g., one or more video clips of the patient), the output of the trained machine learning model (e.g., a prediction of a movement disorder), and the target result (e.g., a recommended corrective action) and the actual measurement result (e.g., a further determination based on a medical professional's assessment of the patient's presence of a movement disorder) may be used to generate new training data items. The new training data items may then be used to further train the trained machine learning model. This retraining process may be performed by the computing device that performed the initial training and / or configuration of the model, or by one or more different computing devices.

[0075] The model training workflow 205 trains one or more machine learning models (e.g., deep learning models) to perform one or more classification, segmentation, detection, recognition, decision, etc. tasks related to predicting one or more movement disorders. The model application workflow 217 applies the one or more trained machine learning models to perform classification, segmentation, detection, recognition, decision, etc. tasks to identify movement disorder symptoms from input data. One or more of the machine learning models can receive and process outcome data (e.g., patient assessments by healthcare professionals) and input data (e.g., recorded video data).

[0076] Various machine learning outputs are described herein. Specific numbers and arrangements of machine learning models are described and shown. However, it should be understood that the number and types of machine learning models used, as well as the arrangements of such machine learning models, may be modified to achieve the same or similar end results. Thus, the arrangements of machine learning models described and shown are merely examples and should not be construed as limiting.

[0077] In an embodiment, one or more machine learning models are trained to perform one or more of the following tasks: Each task may be performed by a separate machine learning model. Alternatively, a single machine learning model (e.g., an ensemble model) may perform all of the tasks or a subset of the tasks. Additionally or alternatively, different machine learning models may be trained to perform different combinations of tasks. In one example, one or a few machine learning models may be trained, where the trained machine learning model is a single shared neural network with multiple shared layers and multiple higher-level separate output layers, where each of the output layers outputs a different prediction, classification, identification, etc. Tasks that the one or more trained machine learning models may be trained to perform are as follows: 1. Cropping Video Data. A patient's recorded or live video may be cropped (e.g., temporally and / or spatially) for further processing. Temporal cropping may be performed to generate clips shorter in duration than the patient's entire video recording. Clips may be generated that approximately correspond to the expected length of a display exhibiting movement disorder symptoms. For example, image frames of a patient being screened for tardive dyskinesia may be separated into periods of approximately 1 second, which may approximately correspond to the involuntary movement patterns caused by tardive dyskinesia. Clips of any temporal length may be generated based on the intended application; for example, approximately 1 second, 0.5 to 5 seconds, 0.1 to 10 seconds, or other lengths may be used. Various temporal crops (potentially overlapping in time), which may be of uniform or variable length, may be taken from a single video recording. In another example, a patient's audio may be temporally cropped into segments that approximately correspond to the expected length of a movement disorder symptom that is detectable in the patient's speech. Spatial cropping may be performed on image data (e.g., video frames) to emphasize, highlight, or zoom in on body parts expected to demonstrate one or more movement disorders. For example, tardive dyskinesia is particularly likely to be demonstrated in movements of the jaw, lips, tongue, eyes, facial muscles, and hands. One or more machine learning models may be utilized to generate cropped videos focused on target body positions. In some embodiments, multiple crops of each patient's video recording may be provided to various machine learning models. Clips may be cropped both temporally and spatially. For example, a video clip may be generated for further processing that is approximately one second long and focused on the patient's jaw. 2. Determining movement disorder symptoms based on video data. One or more machine learning models may be trained to determine whether a movement disorder was indicated within a series of image frames and / or audio of recorded video data. Cropped (spatially and / or temporally) video data may be provided to one or more machine learning models trained to determine the presence of movement disorder symptoms. Crops of the video data, for example, many approximately one-second clips, many spatial crops focused on specific body parts, etc., may be provided to one or more machine learning models trained to determine the presence of movement disorder symptoms. The one or more machine learning models trained to determine the presence of movement disorder symptoms may output a likelihood that the clip contains evidence of a movement disorder. The one or more machine learning models trained to determine the presence of movement disorder symptoms may output a projection into a neural space (e.g., the penultimate layer of a conventional classification network), which may include additional information represented in an N-dimensional vector space. 3. Determining a movement disorder exhibited in a video recording. One or more machine learning models may be trained to determine the likelihood that the entire video exhibits one or more target movement disorders based on the output of one or more machine learning models operating on various data crops. The machine learning model may receive as input the output from one or more other machine learning models, such as the machine learning model trained to determine movement disorder symptoms based on video data described in Task 2 above. The machine learning model may receive an N-dimensional embedding. For example, the final layer of the machine learning model may be a classification layer, with the number of output nodes depending on the number of possible classifications. A layer of nodes before the final classification layer may include more nodes than the classification layer and may carry some additional information encoded as vector values ​​dimensionally equal to the number of nodes in the layer. The N-dimensional embedding provided to the machine learning model may be the values ​​of the nodes of a separate machine learning classifier layer before the classification (e.g., conventional output) layer. The machine learning model may determine whether the video as a whole exhibits one or more target movement disorders based on the data provided based on various temporal and / or spatial crops of the video. The machine learning model may determine a composite indication for the video based on determining the number of crops of the video. The likelihood of a disorder, severity of the disorder, etc. may be further predicted. The model may be configured to predict disorder severity by providing associated patient video data along with a label indicating the severity of the disorder within the patient. 4. Determining a movement disorder present in a patient. One or more machine learning models may be configured to determine whether a patient is experiencing a movement disorder based on multiple video samples. The machine learning model may receive as input outputs from one or more other machine learning models, such as the machine learning model trained to determine a movement disorder shown in a video recording described in task 3 above. The machine learning model may generate composite instructions for the patient based on the machine learning outputs determined from the multiple videos. The machine learning model may be provided with dated information associated with the one or more video data to track the progression of symptoms, disorders, and / or treatments over time. The machine learning model may predict the severity of movement disorder symptoms. The machine learning model may be trained to predict symptom severity by providing patient training data along with associated patient symptom severity labels. The machine learning model may receive as input an N-dimensional embedding, e.g., a layer of machine learning classifiers prior to the classification layer.

[0078] One type of machine learning model that can be used to perform some or all of the above tasks is an artificial neural network, such as a deep neural network. An artificial neural network may generally include a feature representation component with a classifier or recurrent layer that maps features to a desired output space. A convolutional neural network (CNN), for example, hosts multiple layers of convolutional filters. The top-layer features extracted by the convolutional layers are mapped to decisions (e.g., classification outputs), and pooling is performed in lower layers, where multilayer perceptrons are typically added, to address nonlinearities. Deep learning is a class of machine learning algorithms that uses a cascade of multiple layers of nonlinear processing units for feature extraction and transformation. Each successive layer uses the output from the previous layer as input. Deep neural networks can learn in a supervised (e.g., classification) and / or unsupervised (e.g., pattern analysis) manner. Deep neural networks include a hierarchy of layers, where different layers learn different levels of representation corresponding to different levels of abstraction. In deep learning, each level learns to transform its input data into a somewhat abstract and complex representation. In particular, the deep learning process can independently learn which features are best placed within which levels. The "deep" in "deep learning" refers to the number of layers through which data is transformed. More precisely, deep learning systems have significant credit assignment path (CAP) depth. A CAP is a chain of transformations from input to output. A CAP describes the potential causal relationship between input and output. For feedforward neural networks, the CAP depth can be the depth of the network, which can be the number of hidden layers plus one. For recurrent neural networks, where a signal can propagate through layers more than once, the CAP depth is potentially infinite.

[0079] Training a neural network can be accomplished in a supervised learning fashion, which involves feeding a training dataset consisting of labeled inputs through the network, observing its outputs, defining an error (by measuring the difference between the output and the label values), and using techniques such as deep gradient descent and backpropagation to adjust the network's weights across all its layers and nodes so that the error is minimized. In many applications, repeating this process across many labeled inputs in the training dataset results in a network that can produce correct outputs when presented with inputs that differ from those present in the training dataset.

[0080] For the model training workflow 205, a training dataset including hundreds, thousands, tens of thousands, hundreds of thousands, or more of movement disorder label data 210 (e.g., labels for historical video data provided by trained experts) may be used to form the training dataset. In embodiments, the training dataset may include associated video data 212 to form the training dataset, where each data point and / or associated recombination configuration may include various labels or classifications of one or more types of useful information. This data may be processed to generate 236 one or more training datasets for training one or more machine learning models. In some embodiments, additional relevant information may be provided, such as metadata (e.g., date and time of recording), patient data (e.g., prescription medications, medical history, etc.), survey data (e.g., answers to questions posed to the patient, such as the time the patient last took one or more medications, whether the patient is experiencing any unusual conditions that may be altering their physical behavior, etc.).

[0081] In one embodiment, generating one or more training data sets (236) includes generating various crops of the video data. Temporal crops, spatial crops, temporal and spatial crops, etc. may be generated. In some embodiments, a trained labeler may implement or assist in generating the crops. For example, the trained labeler may flag short periods of time during which movements indicative of a movement disorder were made, and the training data set may include the short recorded period as labeled movements in the training data set. As a further example, the trained labeler may input that movements indicative of a movement disorder are present in a particular body part (e.g., the jaw), and the model may spatially crop at least a portion of the video to emphasize the jaw when generating the training data set.

[0082] To accomplish the training, processing logic inputs the training dataset 236 into one or more untrained machine learning models. Prior to inputting the first input into the machine learning models, the machine learning models may be initialized. Processing logic trains the untrained machine learning models based on the training dataset to generate one or more trained machine learning models that perform the various operations described above.

[0083] Training may be performed by inputting one or more of the movement disorder label data 210 and the video data 212 into the machine learning model one at a time. In some embodiments, training the machine learning model includes receiving the video data 212 (e.g., potentially temporally and / or spatially cropped) and adjusting the model to provide the movement disorder label data 210 as output. The machine learning model processes the input and generates an output. The artificial neural network includes an input layer consisting of values ​​in the data points. The next layer is called the hidden layer, and the nodes in the hidden layer each receive one or more of the input values. Each node includes parameters (e.g., weights) to apply to the input values. Each node therefore essentially inputs the input values ​​into a multivariate function (e.g., a nonlinear mathematical transformation) to produce an output value. The next layer may be another hidden layer or an output layer. In either case, the nodes in the next layer receive the output values ​​from the nodes in the previous layer, and each node applies a weight to those values ​​and then generates its own output value. This may be performed at each layer. The final layer is the output layer, where there is one node for each class, prediction, and / or output that the machine learning model can produce.

[0084] Thus, the output may include one or more predictions or inferences. For example, the output prediction or inference may include a clip, video, or prediction of the likelihood that the patient will exhibit symptoms of the target movement disorder. The output may include a classification of the severity of the movement disorder. The output may include a description of the evolution of the movement disorder over time. The output may include a prediction of the future evolution of the movement disorder. The output may include the effectiveness of treatment for the movement disorder. The output may include a recommendation, such as recommending additional screening for the patient.

[0085] The processing logic may compare the classification output of the model with the provided label and determine whether a threshold criterion is met (e.g., the accuracy of the model meets a threshold accuracy). The processing logic determines an error (i.e., a classification error) based on the difference between the processed output and the training label. The processing logic adjusts the weights of one or more nodes in the machine learning model based on the error. An error term or error delta may be determined for each node in the artificial neural network. Based on this error, the artificial neural network adjusts one or more of the parameters (weights for one or more inputs of the node) for one or more of its nodes. The parameters may be updated in a back-propagation manner, with nodes in higher layers being updated first, followed by nodes in the next layer, and so on. The artificial neural network includes multiple layers of "neurons," each layer receiving values ​​as inputs from neurons in the previous layer. The parameters for each neuron include weights associated with the values ​​received from each of the neurons in the previous layer. Adjusting the parameters may therefore include adjusting weights assigned to each of the inputs to one or more neurons in one or more layers within the artificial neural network.

[0086] Once the model parameters are optimized, model validation may be performed to determine whether the model has improved and to determine the current accuracy of the deep learning model. After one or more training rounds, the processing logic may determine whether a stopping criterion has been met. The stopping criterion may be a target level of accuracy, a target number of processed images from the training dataset, a target amount of change to the parameters over one or more previous data points, a combination thereof, and / or other criteria. In one embodiment, the stopping criterion is met when at least a minimum number of data points have been processed and at least a threshold accuracy has been achieved. The threshold accuracy may be, for example, 70%, 80%, or 90% accuracy. In one embodiment, the stopping criterion is met when the accuracy of the machine learning model stops improving. If the stopping criterion is not met, further training is performed. If the stopping criterion is met, training may be complete. Once the machine learning model is trained, a reserved portion of the training dataset may be used to test the model.

[0087] As an example, in one embodiment, a machine learning model (e.g., movement disorder predictor 267) is trained to determine a movement disorder (e.g., classification of one or more movement disorders). Similar processes may be performed to train a machine learning model to perform other tasks, such as those described above. Many (e.g., thousands) sets of combinations of labeled data may be collected, and a movement disorder classification may be determined.

[0088] Once the one or more trained machine learning models 238 are generated, they may be stored in model storage 245 and added to the movement disorder determination application. The movement disorder determination application can then use the one or more trained ML models 238, as well as additional processing logic, to implement an automated mode in which user manual input of information is minimized or, in some cases, even eliminated. In some embodiments, a user (such as a medical professional) can alternatively or additionally select the video data to provide for analysis.

[0089] For the apply model workflow 217, according to one embodiment, input data 262 may be input into a movement disorder predictor 267, which may include a trained neural network. Based on the input data 262, the movement disorder predictor 267 outputs information indicative of a predicted movement disorder decision as predicted movement disorder data 269. The movement disorder decision may include a movement disorder classification, a likelihood of the movement disorder, a movement disorder severity, etc.

[0090] In some embodiments, movement disorder predictor 267 may be a trained machine learning model including multiple trained machine learning models. For example, movement disorder predictor 267 may include various models for generating predictions based on data segmentation. Movement disorder predictor 267 may include models trained to make predictions based on various spatial segmentations (e.g., various body parts). Movement disorder predictor 267 may include models trained to make predictions based on various data types. Movement disorder predictor 267 may include one or more models trained to make predictions based on video data (e.g., including sequences of audio and image data). Movement disorder predictor 267 may include one or more models trained to make predictions based on image data (e.g., sequences of images). Movement disorder predictor 267 may include one or more models trained to make predictions based on audio data (e.g., audio from one or more video recordings). Movement disorder predictor 267 may include one or more models for generating a composite prediction based on the output of several other models, e.g., movement disorder predictor 267 may be an ensemble model, a combination of models, etc. In some embodiments, movement disorder predictor 267 may not be a combination of models. In some embodiments, movement disorder predictor 267 may make a prediction based on video data input.

[0091] 3 illustrates a frame of a video recording 300 as a portion of a video for use in determining the presence of a movement disorder, according to some embodiments. The video (e.g., video recording data including audio data and a series of image frames) may be provided to one or more machine learning models for analysis. The video may be recorded. The video may include the patient's face and may further include upper limbs, lower limbs, torso, etc.

[0092] In some embodiments, the video may be segmented, for example, temporally and / or spatially. Temporal segmentation may include generating one or more video clips. For example, to generate a labeled training dataset, a user may flag short portions of the recorded video containing movement disorder symptoms. In some embodiments, a user may label portions of the video based on observing evidence of a movement disorder. The user may provide timestamps (e.g., via a graphical user interface on a client device) corresponding to the movement disorder evidence. In some embodiments, the dataset used for training (including validation and test operations) may include both temporal crops based on user-provided labels (e.g., clips containing movement disorder evidence) and temporal crops that avoid user-labeled portions (e.g., arbitrarily selected clips that avoid movement disorder evidence). Providing both clips containing and not containing movement disorder evidence from the same patient or recorded video may enable the machine learning model to distinguish between various interfering factors that may be correlated to a patient experiencing a movement disorder and narrow the range of actual movements, speech patterns, or other evidence of a movement disorder.

[0093] Spatial segmentation may include cropping the image to highlight target body parts. Spatial segmentation may include cropping the image to include target parts of the patient's body. The video recording 300 includes a series of exemplary spatial segmentations. An eye crop 302 may be utilized to highlight involuntary movements of the eyes, eyelids, muscles near the eyes, eyebrows, etc. A lip crop 304 may be utilized to highlight involuntary movements of the lips, tongue, jaw, etc. Various other spatial segmentations of the video data may be utilized. For example, one or more of jaw, lip, mouth, tongue, eyes, head, facial muscles, hands, upper and / or lower limbs, and torso crops may be utilized. A spatial crop corresponding to a trained machine learning model, for example, a machine learning model configured to identify evidence of one or more movement disorders based on the target body parts of the spatial crops, may be provided.

[0094] 4A-4C are flow diagrams of methods 400A-C related to training and utilizing a machine learning model, according to some embodiments. Methods 400A-C may be implemented by processing logic, which may include hardware (e.g., circuitry, specialized logic, programmable logic, microcode, a processing device, etc.), software (e.g., instructions executing on a processing device, a general-purpose computer system, or a dedicated machine), firmware, microcode, or a combination thereof. In some embodiments, methods 400A-C may be implemented in part by prediction system 110. Method 400A may be implemented in part by prediction system 110 (e.g., server machine 170 and dataset generator 172 of FIG. 1 ). Prediction system 110 may use method 400A to generate a dataset for at least one of training, validating, or testing a machine learning model, according to embodiments of the present disclosure. Methods 400B-C may be performed by prediction server 112 (e.g., prediction component 114), client device 120, and / or server machine 180 (e.g., training operations, validation operations, and testing operations may be performed by server machine 180). In some embodiments, a non-transitory machine-readable storage medium stores instructions that, when executed by a processing device (e.g., of prediction system 110, of server machine 180, of prediction server 112, etc.), cause the processing device to perform one or more of methods 400A-C.

[0095] For ease of explanation, methods 400A-C are shown and described as a series of operations. However, operations according to the present disclosure may occur in various orders and / or concurrently with other operations not shown and described herein. Moreover, not all shown operations need be performed to implement methods 400A-C in accordance with the disclosed subject matter. In addition, those skilled in the art will understand and appreciate that methods 400A-C may alternatively be represented as a series of interrelated states via a state diagram or events.

[0096] 4A is a flow diagram of a method 400A for generating a dataset for a machine learning model, according to some embodiments. Referring to FIG. 4A, in some embodiments, at block 401, processing logic implementing method 400A initializes a training set T to an empty set.

[0097] At block 402, processing logic generates a first data input (e.g., a first training input, a first validation input, a first test input) that may include one or more of video data, video crop data, N-dimensional embedded data related to the video crop, or other data inputs described herein. In some embodiments, the first data input may include a first set of features related to a type of data, and the second data input may include a second set of features related to a type of data. The input data may include historical data in some embodiments.

[0098] In some embodiments, at block 403, the processing logic optionally generates a first target output for one or more of the data inputs (e.g., the first data input). In some embodiments, the input includes one or more video clips, and the output includes a determination of the presence of evidence of a movement disorder symptom. In some embodiments, the input includes an N-dimensional embedding of a series of video recordings of a patient, and the output includes a determination of the severity of the movement disorder symptom experienced by the patient. In some embodiments, the first target output is predictive data. Any of the functions of the machine learning models described with respect to FIG. 2 (or combinations of functions, e.g., in the case of ensemble machine learning models) may have an associated dataset including corresponding inputs and target outputs. In some embodiments, no target output is generated (e.g., an unsupervised machine learning model that is capable of grouping or finding correlations in input data rather than requiring a target output to be provided).

[0099] At block 404, processing logic optionally generates mapping data indicating an input-output mapping. The input-output mapping (or mapping data) may refer to data inputs (e.g., one or more of the data inputs described herein), target outputs for the data inputs, and associations between the data inputs and the target outputs. In some embodiments, such as in connection with a machine learning model that does not provide a target output, block 404 may not be performed.

[0100] At block 405, processing logic, in some embodiments, adds the mapping data generated at block 404 to dataset T.

[0101] At block 406, processing logic branches based on whether dataset T is sufficient for at least one of training, validating, and / or testing a machine learning model, such as model 190 of FIG. 1. If so, execution proceeds to block 407; if not, execution continues back to block 402. Note that in some embodiments, the sufficiency of dataset T may be determined solely based on the number of inputs mapped to outputs in the dataset, while in some other embodiments, the sufficiency of dataset T may be determined based on one or more other criteria (e.g., a measure of diversity of the data examples, accuracy, etc.) in addition to or instead of the number of inputs.

[0102] At block 407, processing logic provides (e.g., to server machine 180) dataset T for training, validating, and / or testing machine learning model 190. In some embodiments, dataset T is a training set and is provided to training engine 182 of server machine 180 to perform training. In some embodiments, dataset T is a validation set and is provided to validation engine 184 of server machine 180 to perform validation. In some embodiments, dataset T is a test set and is provided to test engine 186 of server machine 180 to perform testing. In the case of a neural network, for example, input values ​​(e.g., numerical values ​​associated with data input 210A) of a given input-output mapping are input to the neural network, and output values ​​(e.g., numerical values ​​associated with target outputs) of the input-output mapping are stored in output nodes of the neural network. Connection weights in the neural network are then adjusted according to a learning algorithm (e.g., backpropagation, etc.), and the procedure is repeated for other input-output mappings in dataset T. After block 407, the model (e.g., model 190) may be at least one of trained using the training engine 182 of the server machine 180, validated using the validation engine 184 of the server machine 180, or tested using the testing engine 186 of the server machine 180. The trained model may be implemented by the prediction component 114 (of the prediction server 112) to generate predictive data and / or to generate predictive data 168 for implementing corrective actions related to movement disorders.

[0103] 4B is a flow diagram of a method 400B for utilizing machine learning models in determining whether evidence of a movement disorder has been captured, according to some embodiments. At block 410, processing logic acquires video data of the patient, including image data and audio data. In some embodiments, only image data may be utilized in determining a movement disorder. In some embodiments, only audio data may be utilized in determining a movement disorder. In some embodiments, the image data (e.g., a series of image frames) and audio data may be provided to the same set of machine learning models, separate sets of machine learning models, or overlapping sets of machine learning models.

[0104] In some embodiments, the video data may be or may include a patient recording. The video data may include video recording data of the patient. In some embodiments, the video data may be a recording of an interview between the patient and a healthcare professional. In some embodiments, the video data may be or may include a recording of the patient responding to movement disorder screening prompts, potentially including prompts, e.g., questions to answer, exercises or positions to perform, etc. The instructions may be provided by a healthcare professional, a screener, or may be provided automatically by the screening system, such as by on-screen text instructions, recorded and / or generated voice instructions (e.g., text-to-speech generation). In some embodiments, the video data may be a recording of an interview between the patient and a healthcare professional, during which the patient is prompted to perform actions that may be used to make decisions related to a movement disorder. For example, the patient may be asked questions that may reveal evidence of a movement disorder, may be prompted to perform certain actions or activate certain muscles, etc.

[0105] In some embodiments, the video data may be pre-processed. For example, the video data may be segmented. The video data may be segmented spatially and / or temporally. In some embodiments, the video data may be temporally cropped (e.g., split into one or more clips of shorter duration) and / or spatially cropped (e.g., divided into components that highlight target body parts or portions of the patient's body). In some embodiments, various segments may be provided to a machine learning model for analysis. For example, several temporal crops, several spatial crops, and the entire recording may all be provided to one or more machine learning models for analysis. In some embodiments, spatial crops may include crops that highlight the patient's chin, lips, mouth, tongue, eyes, head, facial muscles, hands, upper and / or lower limbs, torso, etc. In some embodiments, temporal crops may include clips of approximately 1 second in duration, crops of 0.5 to 5 seconds, clips of 0.1 to 10 seconds, or clips of other lengths. In some embodiments, the length of the clip may correspond approximately to the expected duration of the movement disorder indication (e.g., the duration of involuntary movements typical for one or more target movement disorders).

[0106] At block 412, processing logic provides the video data to a first trained machine learning model. The first trained machine learning model may be an ensemble model. The first trained machine learning model may be configured to receive the video data and provide a decision related to one or more movement disorders. The first trained machine learning model may include several other models (e.g., an ensemble machine learning model). The first trained machine learning model may include one or more models for image data, one or more models for audio data, one or more models for different spatial crops (e.g., to emphasize different body parts), etc.

[0107] In some embodiments, the first trained machine learning model may include several models operating in series on one or more sets of video data (e.g., it may be an ensemble model). For example, the first model may receive a crop of video data and generate an indication of the likelihood that a target movement disorder was shown in the crop. The second model may receive an indication of the likelihood that a target movement disorder was shown in one or more crops from the first model and generate as output a composite indication of the likelihood that the target movement disorder was shown in a video including several analyzed clips. The third model may receive data regarding several video recordings and generate as output a composite indication of a decision related to the patient based on several video recordings of the patient. For example, the third model may generate a composite indication of movement disorder progression, movement disorder severity, etc.

[0108] At block 414, processing logic obtains output from the first trained machine learning model based on the video data. The output includes an indication that the patient exhibits symptoms of one or more target movement disorders in association with the video data. The target movement disorder may be tardive dyskinesia. The target movement disorder may be Parkinson's disease. The target movement disorder may be Huntington's chorea. The target movement disorder may be essential tremor, dystonia, Tourette's syndrome, restless legs syndrome, or another movement disorder that may be indicated via image and / or audio data.

[0109] At block 416, the processing logic provides a warning to the user indicating one or more target movement disorders. In some embodiments, the warning may be provided to a medical professional. In some embodiments, the warning may indicate whether evidence of a movement disorder is detected. In some embodiments, the warning may indicate the likelihood that evidence of a movement disorder is contained within the video data. In some embodiments, the warning may indicate the severity or progression of the movement disorder. In some embodiments, the warning may be provided via a graphical user interface of the client device. In some embodiments, the warning may prompt the user to take additional measures, such as performing additional screening or diagnostic techniques.

[0110] 4C is a flow diagram of a method 400C for training a machine learning model to make decisions related to movement disorders, according to some embodiments. At block 420, processing logic acquires a first plurality of video data of a first plurality of patients. The video data may include recorded image data (e.g., a series of image frames) and audio data. The patients may include both patients experiencing movement disorder symptoms and patients not experiencing movement disorder symptoms.

[0111] At block 422, processing logic performs cropping of each of the first plurality of video data to generate a first plurality of video data crops. In some embodiments, the cropping may be temporal, e.g., separating the video data into clips of a target duration or duration range. In some embodiments, the cropping may be spatial, e.g., emphasizing one or more target body regions and / or parts of the patient's body. In some embodiments, the cropping may be both temporal and spatial. In some embodiments, the cropping may be performed by one or more users. For example, a trained user may review a patient's recording. The user may flag portions of the video that include evidence of a movement disorder. The processing logic may generate clips based on the flagged portions of the recording. The processing logic may generate both clips that include involuntary movement and clips that do not include involuntary movement from the same recording of the same patient. Providing both clips that include and do not include evidence of a movement disorder from the same patient may enable the machine learning model to distinguish between natural movement and movement associated with a movement disorder.

[0112] At block 424, processing logic receives a first plurality of labels associated with each of the first plurality of video data crops. The first plurality of labels include an indication of the presence or absence of evidence of a movement disorder within the first plurality of video data crops. These labels may be generated by one or more users, for example, users trained to recognize movement disorders.

[0113] At block 426, processing logic trains a first machine learning model by providing the first plurality of video data crops as training inputs and the first plurality of labels as target outputs, wherein the first machine learning model is configured to generate an output indicating whether the input video data crops include an indication of a movement disorder.

[0114] Machine learning models configured to perform different tasks may be trained in a similar manner, utilizing training data targeted to the intended use of the machine learning model. For example, a machine learning model configured to receive indications of a movement disorder from one or more clips of video recordings and provide a composite indication regarding a determined likelihood of evidence of the movement disorder within the entire clip may be provided as training input indications of the movement disorder from the training clips and labels associated with the video recordings of the training clips as a target output. As a further example, the model may be configured to provide a composite indication of a determination of whether a patient is experiencing symptoms based on data from a series of video recordings. The model may be trained based on labels indicating the severity of the patient's movement disorder as a target output and instructions based on the series of recordings as a training input.

[0115] FIG. 5 illustrates an operational flow 500 for making a determination related to a movement disorder, according to some embodiments. Patient video data is generated to generate data for both training and inference. In some embodiments, the video data may be a recording of a virtual or in-person interview with a healthcare professional. In some embodiments, the patient may be provided with a series of target questions and / or instructions used to determine whether a movement disorder is present. In one example, questions targeting facial responses are asked of the patient and / or instructions targeting facial responses are given to the patient, with responses recorded for facial analysis 502, for example. Further instructions may be provided to the patient targeting areas that indicate other symptoms related to the movement disorder, such as measuring verbal responses 504 and physical analysis 506. The questions may be similar to those asked during traditional movement disorder screening, such as the AIMS test, which is administered to determine whether a patient is experiencing symptoms of tardive dyskinesia. The patient may be asked questions and may also be asked to perform actions such as opening their mouth, snapping their fingers, or staring into a camera.

[0116] A video timeline 508 is generated based on the patient recordings. In practice, many patients may be recorded, and multiple recordings may be made for each patient. Portions of the video timeline 508 may be labeled. For example, a user trained to recognize movement disorders may highlight portions of the video (e.g., portions 510 and 512) where evidence of the target movement disorder is present. The user may further enter a description of the body part where the movement disorder is demonstrated.

[0117] The labeled video timeline 508 is then provided to a temporal segmentation tool 514 and a spatial segmentation tool 516. The segmentation tools utilize user-provided labels to segment the video into segments. Segmentation may be spatial, for example, based on user-provided body part labels. Segmentation may be temporal, for example, based on user-provided timestamp flagging for movement disorders. In some embodiments, segments are generated that are segmented both spatially and temporally, for example, one-second clips cropped to focus on target body parts. Segmentation may be guided by user labels. Segmentation may include separating (temporally and / or spatially) both portions that provide evidence of a movement disorder and portions that do not. Providing both segments that indicate and do not indicate a movement disorder may aid in training a reliable model for distinguishing between movement disorders.

[0118] The segments are provided to a data augmentation tool 518. The data augmentation tool may provide some variability to the data clips. The variability may be used to protect the model from overfitting based on spurious data correlations. For example, the video segments may be adjusted in terms of color, brightness, contrast, sharpness, size, frame rate, aspect ratio, rotation, etc. Such variability can make the trained model more robust to variability in the input data, as well as effectively increasing the available size of the training data.

[0119] The augmented data is used to generate a training data set 520 (e.g., as described with respect to FIGS. 2 and 4A). The data flow proceeds further to model training 522 (e.g., as described with respect to FIGS. 2 and 4C) and model inference 524 (e.g., as described with respect to FIGS. 2 and 4B).

[0120] 6 is a block diagram illustrating a computer system 600, according to some embodiments. In some embodiments, computer system 600 may be connected to other computer systems (e.g., via a network, such as a local area network (LAN), an intranet, an extranet, or the Internet). Computer system 600 may operate as a server or a client in a client-server environment, or as a peer computer in a peer-to-peer or distributed network environment. Computer system 600 may be provided by a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a cellular phone, a web appliance, a server, a network router, a switch, or a bridge, or any device capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by that device. Furthermore, the term “computer” is intended to include any collection of computers that individually or together execute a set (or sets) of instructions to perform any one or more of the methodologies described herein.

[0121] In a further aspect, computer system 600 may include a processing device 602, a volatile memory 604 (e.g., random access memory (RAM)), a non-volatile memory 606 (e.g., read-only memory (ROM) or electrically erasable programmable ROM (EEPROM)), and a data storage device 618, which may communicate with each other via a bus 608.

[0122] The processing device 602 may be provided by one or more processors, such as a general-purpose processor (e.g., a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a microprocessor implementing other types of instruction sets, or a microprocessor implementing a combination of instruction set types) or a special-purpose processor (e.g., an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), or a network processor).

[0123] Computer system 600 may further include a network interface device 622 (e.g., coupled to a network 674). Computer system 600 may include a video display unit 610 (e.g., an LCD), an alphanumeric input device 612 (e.g., a keyboard), a cursor control device 614 (e.g., a mouse), and a signal generating device 620.

[0124] In some embodiments, the data storage device 618 may include a non-transitory computer-readable storage medium 624 (e.g., a non-transitory machine-readable medium) capable of storing instructions 626 for encoding any one or more of the methods or functions described herein, including the instruction-encoding components of FIG. 1 (e.g., the prediction component 114, the corrective action component 122, the model 190, etc.), and for implementing the methods described herein.

[0125] The instructions 626 may reside, completely or partially, within the volatile memory 604 and / or within the processing device 602 during execution thereof by the computer system 600; thus, the volatile memory 604 and the processing device 602 may also constitute machine-readable storage media.

[0126] Although computer-readable storage medium 624 is shown as a single medium in the illustrative example, the term "computer-readable storage medium" is intended to include a single medium or multiple media (e.g., centralized or distributed databases, and / or associated caches and servers) that store one or more sets of executable instructions. The term "computer-readable storage medium" is also intended to include any tangible medium that is capable of storing or encoding a set of instructions for execution by a computer that cause the computer to perform any one or more of the methodologies described herein. The term "computer-readable storage medium" is intended to include, but is not limited to, solid-state memory, optical media, and magnetic media.

[0127] The methods, components, and features described herein may be implemented by discrete hardware components or integrated into the functionality of other hardware components, such as an ASIC, FPGA, DSP, or similar device. In addition, the methods, components, and features may be implemented by firmware modules or functional circuits within a hardware device. Furthermore, the methods, components, and features may be implemented within any combination of hardware devices and computer program components, or within a computer program.

[0128] Unless otherwise specified, terms such as "receive," "execute," "provide," "obtain," "cause," "access," "determine," "add," "use," "train," "reduce," "generate," "correct," and the like refer to actions and processes performed or implemented by a computer system that manipulate and convert data represented as physical (electronic) quantities in computer system registers and memory into other data similarly represented as physical quantities in computer system memory or registers, or other such information storage, transmission, or display devices. Also, as used herein, terms such as "first," "second," "third," "fourth," and the like are intended to be labels for distinguishing different elements and may not have a sequential meaning according to their numerical designation.

[0129] The examples described herein also relate to apparatus for performing the methods described herein. The apparatus may be specifically constructed to perform the methods described herein, or the apparatus may include a general-purpose computing system selectively programmed by a computer program stored within the computer system. Such a computer program may be stored in a computer-readable tangible storage medium.

[0130] The methods and illustrative examples described herein are not inherently related to any particular computer or other apparatus. Various general-purpose systems may be used in accordance with the teachings described herein, or it may prove convenient to construct more specialized apparatus to perform the methods described herein and / or each of their individual functions, routines, subroutines, or operations. Examples of constructions for a variety of these systems are set forth in the description above.

[0131] The above description is intended to be illustrative, not limiting. While the present disclosure has been described with reference to particular illustrative examples and embodiments, it will be recognized that the disclosure is not limited to the described examples and embodiments. The scope of the present disclosure should be determined with reference to the scope of the following claims, along with the full scope of equivalents to which such claims are entitled. [Explanation of symbols]

[0132] 100 systems 110 Prediction System 112 Prediction Server 114 Prediction Components 116 Temporal Cropping 118 Spatial Cropping 120 client devices 122 Corrective Action Component 124 Image Data Collection 126 Audio Data Collection 130 Network 140 data stores 142 Image data, sensor data 144 Historical Image Data 146 Current Image Data 148 Segmented Image Data 160 audio data 164 historical voice data 166 Current Audio Data 168 Segmented voice data, prediction data 170 Server Machine 172 Dataset Generator 180 Server Machine 182 Training Engine 184 Validation Engine 185 Selection Engine 186 Test Engine 190 First Trained Machine Learning Models 205 Model Training Workflow 210 Movement Disorders Label Data 212 Video Data 217 Model Application Workflow 236 training datasets 238 trained machine learning models 245 Model Storage 262 input data 267 Movement Disorder Predictor 269 ​​Predicted Movement Disorders Data 300 video recordings 302 Eye Crop 304 Lip Crop 500 Operational Flow 502 Facial Analysis 504 Verbal Response 506 Body Analysis 508 Video Timeline 510 parts 512 parts 514 Temporal Segmentation Tools 516 Spatial Segmentation Tools 518 Data Augmentation Tools 520 training datasets 522 Model Training 524 Model Inference 600 Computer Systems 602 Processing Device 604 Volatile Memory 606 Non-volatile memory 608 Bus 610 Video Display Unit 612 Alphanumeric Input Device 614 Cursor Control Device 618 Data Storage Devices 620 Signal Generating Device 622 Network Interface Device 624 Non-transitory computer-readable storage medium 626 command 674 Network

Claims

1. acquiring, with a processing device, video data of the patient, including image data and audio data; providing, by the processing device, the video data to a first trained machine learning model; obtaining an output from the first trained machine learning model based on the video data, the output including a first indication that the patient exhibits symptoms of one or more target movement disorders in association with the video data; providing a warning to a user indicative of the one or more target movement disorders; A method comprising:

2. obtaining record data of the patient; generating a first video crop of the recorded data by performing temporal cropping of the recorded data, wherein the video data of the patient includes the recorded data and the first video crop of the recorded data; The method of claim 1 further comprising:

3. 3. The method of claim 2, further comprising generating a plurality of video crops of the recorded data, each of the plurality of video crops having a length of 0.1 seconds to 10 seconds, and the video data of the patient further comprising the plurality of video crops.

4. obtaining record data of the patient; generating a first video crop of the recorded data by performing spatial cropping of the recorded data, wherein the video data of the patient includes the recorded data and the first video crop of the recorded data; The method of claim 1 further comprising:

5. 5. The method of claim 4, further comprising generating a plurality of video crops of the recorded data, each of the plurality of video crops including a spatial portion of the recorded data that includes a target portion of the patient's body, and the video data of the patient further comprising the plurality of video crops.

6. 10. The method of claim 1, wherein the first trained machine learning model comprises a first model configured to receive the image data as an input and a second model configured to receive the audio data as an input.

7. the first trained machine learning model: a second trained machine learning model configured to receive as input the crop of the video data of the patient and to generate as output a second indication of the likelihood that a first target movement disorder of the one or more target movement disorders was exhibited in the crop of the video data of the patient; and a third trained machine learning model configured to receive as input one or more indications of the likelihood that the first target movement disorder was exhibited in one or more crops of the video data of the patient, and to generate as output a composite indication of the likelihood that the first target movement disorder was exhibited in the video data of the patient; and 2. The method of claim 1, comprising:

8. the first trained machine learning model: a second trained machine learning model configured to generate as an output a second indication of the likelihood that a first target movement disorder of the one or more target movement disorders was exhibited in the video data of the patient; and a third trained machine learning model configured to receive as input the second indication of the likelihood that the first target movement disorder was exhibited in the video data of the patient and multiple indications of the likelihood that the first target movement disorder was exhibited in multiple video data of the patient, and to generate as output a third indication of a severity of the patient's symptoms relative to the first target movement disorder; and 2. The method of claim 1, comprising:

9. a first target movement disorder of the one or more target movement disorders; tardive dyskinesia, Huntington's chorea, or Parkinson's disease The method of claim 1 , comprising one of:

10. acquiring, with a processing device, a first plurality of video data of a first plurality of patients; performing cropping of each of the first plurality of video data to generate a first plurality of video data crops; receiving a first plurality of labels associated with each of the first plurality of video data crops, the first plurality of labels including a first indication of the presence or absence of evidence of a movement disorder within the first plurality of video data crops; training a first machine learning model by providing the first plurality of video data crops as training inputs and providing the first plurality of labels as target outputs, the first machine learning model being configured to generate an output indicative of whether the input video data crops include an indication of the movement disorder; A method comprising:

11. receiving a second plurality of labels, each label of the second plurality of labels including a second indication of the presence or absence of evidence of the movement disorder within the first plurality of video data; training a second machine learning model by providing an output of the first machine learning model as a training input and providing the second plurality of labels as a target output, the second machine learning model being configured to generate an output indicating whether a video including the input video data crop contains a third indication of the movement disorder; 11. The method of claim 10, further comprising:

12. receiving a third plurality of labels, each label of the third plurality of labels including a fourth indication of the severity of the movement disorder in an associated patient; training a third machine learning model by providing the output from the second machine learning model as a training input and providing the third plurality of labels as a target output, the third machine learning model being configured to generate an output indicative of a prediction of a severity of a movement disorder symptom of the patient based on one or more videos of the patient; 12. The method of claim 11, further comprising:

13. The first plurality of video data crops comprises: Temporal crop, or spatial crops, each spatial crop including a target portion of the patient's body; 11. The method of claim 10, comprising one or more of:

14. A non-transitory machine-readable storage medium having instructions stored thereon, The instructions, when executed, cause a processing device to: acquiring video data of the patient, including image data and audio data; providing the video data to a first trained machine learning model; obtaining an output from the first trained machine learning model based on the video data, the output including a first indication that the patient exhibits symptoms of one or more target movement disorders in association with the video data; and providing a warning to a user indicative of the one or more target movement disorders; causing an action including Non-transitory machine-readable storage medium.

15. The operation is obtaining record data of the patient; generating a first video crop of the recorded data by performing temporal cropping of the recorded data, wherein the video data of the patient includes the recorded data and the first video crop of the recorded data; 15. The non-transitory machine-readable storage medium of claim 14, further comprising:

16. The operation is obtaining record data of the patient; generating a first video crop of the recorded data by performing spatial cropping of the recorded data, wherein the video data of the patient includes the recorded data and the first video crop of the recorded data; 15. The non-transitory machine-readable storage medium of claim 14, further comprising:

17. The operation is 17. The non-transitory machine-readable storage medium of claim 16, further comprising generating a plurality of video crops of the recorded data, each of the plurality of video crops comprising a spatial portion of the recorded data that includes a target portion of the patient's body, and wherein the video data of the patient further comprises the plurality of video crops.

18. the first trained machine learning model: a second trained machine learning model configured to receive as input the crop of the video data of the patient and to generate as output a second indication of the likelihood that a first target movement disorder of the one or more target movement disorders was exhibited in the crop of the video data of the patient; and a third trained machine learning model configured to receive as input one or more indications of the likelihood that the first target movement disorder was exhibited in one or more crops of the video data of the patient, and to generate as output a composite indication of the likelihood that the first target movement disorder was exhibited in the video data of the patient; and 15. The non-transitory machine-readable storage medium of claim 14, comprising:

19. the first trained machine learning model: a second trained machine learning model configured to generate as an output a second indication of the likelihood that a first target movement disorder of the one or more movement disorders was exhibited in the video data of the patient; and a third trained machine learning model configured to receive as input the second indication of the likelihood that the first target movement disorder was exhibited in the video data of the patient and multiple indications of the likelihood that the first target movement disorder was exhibited in multiple video data of the patient, and to generate as output a third indication of a severity of the patient's symptoms relative to the first target movement disorder; and 15. The non-transitory machine-readable storage medium of claim 14, comprising:

20. a target movement disorder of the one or more target movement disorders: tardive dyskinesia, Huntington's chorea, or Parkinson's disease 15. The non-transitory machine-readable storage medium of claim 14, comprising one or more of: