Systems and methods for predicting and assessing motion of humans and / or objects in an environment
Patent Information
- Application Number
- PCT/US2025/018785
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-06
- Filing Date
- 2025-03-06
- Publication Date
- 2025-10-02
AI Technical Summary
Conventional neural network-based pose estimation techniques for human motion analysis suffer from low accuracy due to limited 'in-the-wild' datasets, particularly when occlusions occur, and lack real-time data processing capabilities, making them unsuitable for industrial environments. There is a need for improved methods to track multiple people and provide real-time safety alerts and ergonomic assessments.
A method utilizing a device that captures image data, determines 3-D coordinates and kinematic parameters using an extended Kalman filter, classifies user activities, predicts future intent and motion, and generates ergonomic risk assessments, even in the presence of occlusions, by integrating multiple sensors and machine learning algorithms.
Enables accurate, real-time prediction and assessment of human motion, providing ergonomic feedback and safety alerts, enhancing human-machine collaboration and worker safety in industrial settings.
Smart Images

Figure US2025018785_02102025_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS FOR PREDICTING AND ASSESSING MOTION OF HUMANS AND / OR OBJECTS IN AN ENVIRONMENT TECHNICAL FIELD
[0001] The present invention relates to systems and methods for predicting and assessing motion of humans and / or objects in an environment. In particular, the present invention relates to systems and methods that utilize three-dimensional (3D) pose estimation, kinematic modeling, activity classification and user intent recognition, future motion prediction, and motion assessment including an ergonomic risk assessment. RELATED APPLICATION
[0002] This application is claiming priority to U.S. Provisional Application Serial No. 63 / 562,165 filed March 6, 2024, which is incorporated herein in its entirety. BACKGROUND OF THE INVENTION
[0003] Many industrial processes involve collaboration between humans and machines. While a lot of work has been carried out in developing digital twins of machines, development of human digital twins is still in its infancy. Developing a high-fidelity human digital twin can assist with designing effective human-machine collaboration mechanisms, improve human worker safety and performance, and assist with training the workforce. As such, the human digital twin is an important part of the futuristic industrial metaverse. Human digital twin development involves considerations such as user personality, behavioral characteristics, skills, and decision-making.
[0004] Human motion behavior in industrial applications has been studied in ergonomics. That study focused on human safety and comfort, but does not address human-machine interaction research. Typical technology used to study human motion includes motion trackers and wearables, where the intended outcome has been to obtain safe postures for carrying out a variety of work. The recent advances in computer vision, camera technology, artificial intelligence (AI) / machine learning (ML) techniques, and computing technology have made it possible to carry out human pose estimation and motion tracking effectively and efficiently.
[0005] Human pose estimation often involves using computer vision to estimate joint locations of the human body. In recent years, thanks to the enormous advancement in machine learning techniques, human pose estimation has made significant progress, especially in two- dimensional (2-D) pose estimation. 3-D pose estimation has been attempted using differentneural network architectures with diverse types of input data. Examples of these networks include OpenPose and HRNet. In one approach, 3-D poses are predicted by using monocular images with heatmap data to estimate the 3-D joints of a target person directly. See, e.g., “Human pose regression by combining indirect part detection and contextual information”, by D. Luvizon, H. Tabia, and D. Picard, Computers & Graphics, 85, 15–2 (2019). Another approach is to reconstruct 3-D joint locations from 2-D joint predictions through a convolutional neural network (CNN), as is discussed in Tome et al. See, e.g., “Lifting from the Deep: Convolutional 3D Pose Estimation from a Single Image”, by D. Tome, C. Russell, and L. Agapito, from Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2500–2509 (2017).
[0006] An “in-the-wild” 3-D pose dataset refers to data that is collected in uncontrolled, real- world environments, rather than data collected in a structured, laboratory setting. This means the dataset may include natural variations such as different lighting conditions, diverse backgrounds, varying camera angles, occlusions, and / or the like. One issue with 3-D pose estimation techniques is that due to a limited “in-the-wild” 3-D pose dataset, many conventional neural network-based pose estimation techniques are subjected to bias and cannot be generalized for multiple purposes. One of the largest limitations is low accuracy, particularly when an occlusion occurs. An occlusion refers to a situation where a portion of a user is blocked from view of a camera. For example, if a user in a warehouse is walking around holding a box, portions of the user may not be visible to the camera, making it difficult to estimate joint positions for joints found in the occluded region. For this reason, deep neural network human pose estimation through a single camera is not suitable for several working environments in industries such as manufacturing workspace, which includes partial, full, and / or occasional occlusions of the body parts of workers. Hence, there is a need for another source of information to augment the single camera pose estimation.
[0007] Further, many conventional techniques have been applied for the post-processing of data to obtain information retroactively, and thus there is a need for a solution capable of efficiently and effectively processing data in real-time. Furthermore, there is a need to provide safety alerts and feedback to users (e.g., workers, athletes, etc.) in real-time based on current and predicted motions to minimize worker injury in the workplace. Still further, there is a need to develop an algorithmic means to combine information from multiple cameras to track multiple people.
[0008] In addition to these, there is also a need to develop a custom dataset using motion capture along with AI-driven text descriptions for a diverse range of manufacturing tasks.Further, there is a need for fine-tuning AI model hyperparameters and loss functions to consider ergonomics. Furthermore, there is a need to develop a framework for comparative analysis of real-time motions using AI-driven results. Still further, there is a need to carry out recognition of human intent and predict human motion using deep learning networks and generative AI. Additionally, there is a need to utilize a human digital twin model to predict human physiological information such as energy consumption, heart rate, and temperature. Moreover, there is a need to integrate component recognition to validate task completion and to provide optical feedback using motion to text (including ergonomics) and to generate next task procedures.
[0009] It is an object of the present invention to overcome one or more of the problems and / or needs described above. SUMMARY
[0010] In an aspect of the invention, a method for predicting and assessing motion of a user in an environment is provided. The method includes receiving, by a device, image data depicting the user in the environment. The image data is captured by one or more sensor devices. The method further includes determining, by the device, a first set of three-dimensional (3-D) coordinates representing estimated positions of key points of the user at a time ^^. The method further includes determining, by the device, a set of link lengths for linkages connecting the key points. The method further includes determining, by the device and for a time ^^^^, a second set of 3-D coordinates representing predicted positions of the key points and a set of kinematic parameters representing predicted kinematics for the key points. The second set of 3-D coordinates and the set of kinematic parameters are part of a state vector determined by modeling kinematics of the key points using the first set of 3-D coordinates for the time ^^and the link lengths.
[0011] The method further includes updating, by the device, the state vector by using an extended Kalman filter (EKF) to compare the second set of 3-D coordinates representing the predicted positions of the key points for the time ^^^^with a third set of 3-D coordinates representing observed positions of the key points at the time ^^^^. The updated state vector is provided back to the 3-D kinematic model for iterative learning. The 3-D kinematic model and the extended Kalman filter are used to determine and update state vectors periodically over a time period ^^^.
[0012] The method further includes classifying, by the device, an activity being performed bythe user by using machine learning to process the state vectors determined over the time period^^^. The method further includes predicting, by the device, a future intent of the user for atime period ^^^. The future intent is predicted by using machine learning to process the state vectors and the classified activity being performed by the user. The method further includes predicting, by the device, a future motion path of the user for the time period ^^^by using machine learning to process the user activity and the future intent of the user. The method further includes generating, by the device, a motion assessment for the user by processing the predicted future motion path of the user using machine learning. The method further includes making, by the device, the motion assessment available to a device or account associated with the user.
[0013] In an embodiment of the invention, the one or more sensors that capture the image data are a plurality of sensors and the image data received depicts the user and a second user. In this embodiment, the method further comprises determining a unified 3-D coordinate plane by using a calibration and orientation module to process the image data. The unified 3-D coordinate plane is used to determine the first set of 3-D coordinates representing the estimated positions of the key points of the user and another set of 3-D coordinates representing estimated positions of key points of the second user.
[0014] In another embodiment of the invention, the first set of 3-D coordinates at the time ^^include one or more null or empty set values based on a portion of the user being occluded from view of the one or more sensors. In this embodiment, when determining the second set of 3-D coordinates for the time ^^^^, the method includes determining one or more 3-D coordinates representing predicted positions of one or more key points occluded from view of the one or more sensors.
[0015] In another embodiment of the invention, the set of kinematic parameters of the state vector include velocity parameters and acceleration parameters. In this embodiment, the state vector further includes angular data identifying angular positions of the key points. In this embodiment, the state vector is used to classify the activity being performed by the user, to predict the future intent of the user, and to predict the future motion of the user.
[0016] In another embodiment of the invention, the motion assessment is an ergonomic risk assessment. In this embodiment, generating the ergonomic risk assessment includes generating the ergonomic risk assessment by using an ergonomic assessment model to process the predicted future motion path of the user. The ergonomic risk assessment includes one or moreof: a safety assessment of past motion of the user, and a recommendation providing one or more corrective motions for the user.
[0017] In another embodiment of the invention, generating the ergonomic risk assessment includes generating a video including a 3-D model of the user performing the one or more corrective motions, and including the video in the recommendation.
[0018] In another embodiment of the invention, the method further includes determining a third set of 3-D coordinates representing estimated positions of one or more objects in the environment of the user. In this embodiment, the method further includes determining a second set of kinematic parameters representing predicted kinematics for the one or more objects. The third set of 3-D coordinates and second set of kinematic parameters are part of a new state vector. In this embodiment, the method further includes predicting a future motion path of the one or more objects by processing the third set of 3-D coordinates and the second set of kinematic parameters using machine learning. In this embodiment, the predicted future motion path is incorporated into the motion assessment.
[0019] In another embodiment of the invention, the method further includes detecting an anomaly relating to motion of the user by processing the updated state vector. In this embodiment, generating the motion assessment includes generating an alert indicating that the anomaly has been detected.
[0020] In another aspect of the invention, a device is provided. The device includes one or more memories and one or more processors, communicatively coupled to the one or more memories. The one or more processors are to receive image data captured by one or more sensor devices. The image data depicting a user in an environment. The one or more processors are further to determine a first set of three-dimensional (3-D) coordinates representing estimated positions of key points of a user at a time ^^. The one or more processors are further to determine a set of link length for linkages connecting the key points.
[0021] The one or more processors are further to determine, for a time ^^^^, a second set of 3- D coordinates representing predicted positions of the key points and a set of kinematic parameters representing predicted kinematics for the key points. The second set of 3-D coordinates and the set of kinematic parameters are part of a state vector determined by modeling kinematics of the key points using the first set of 3-D coordinates for the time ^^and the link lengths. State vectors are then determined periodically over a time period ^^^.
[0022] The one or more processors are further to classify an activity being performed by the user by using machine learning to process the state vectors determined over the time period^^^. The one or more processors are further to predict a future intent of the user for a final time period defined as ^^^(^^^), ^^^(^^^), …, ^^^(^^^). The future intent is predicted by using machine learning to process the state vectors and the classified activity being performed by the user. The one or more processors are further to predict a future intent of the user for a time period ^^^. The future intent is predicted by using machine learning to process the state vectors and the classified activity being performed by the user. The one or more processors are further to predict a future motion path of the user for the time period ^^^by using machine learning to process the user activity and the future intent of the user. The one or more processors are further to generate a motion assessment for the user by processing the predicted future motion path of the user using machine learning. The motion assessment includes at least one of: an ergonomic risk assessment relating to past motions of the user, a recommendation providing one or more corrective motions, and an alert indicating that an anomaly or unsafe motion has been detected. The one or more processors are further to make the motion assessment available to a device or account associated with the user.
[0023] In an embodiment of the invention, the one or more processors, when updating the state vector, are to update the state vector by using an extended Kalman filter (EKF) to compare thesecond set of 3-D coordinates representing the predicted positions of the key points for the time^^^^ with a third set of 3-D coordinates representing observed positions of the key points atthe time ^^^^. The updated state vector is provided back to the 3-D kinematic model foriterative learning. The EKF is used to update the state vectors periodically over the time period^^^.
[0024] In another embodiment of the invention, the first set of 3-D coordinates representing the estimated positions of the key points of the user at the time ^^include one or more null or empty set values based on a portion of the user being occluded from view of the one or more sensors. In this embodiment, the one or more processors, when determining the second set of 3-D coordinates for the time ^^^^, are to determine one or more 3-D coordinates representing predicted positions of one or more key points occluded from view of the one or more sensors.
[0025] In another embodiment of the invention, the one or more processors are further to: detect an anomaly relating to motion of the user by processing the updated state vector. In this embodiment, when generating the motion assessment, the one or more processors are to generate the alert based on the detected anomaly.
[0026] In another embodiment of the invention, the one or more sensors that capture the image data are a plurality of sensors and the image data received depicts the user and a second user.In this embodiment, the one or more processors are further to: determine a unified 3-D coordinate plane that is static relative to a center of the user and relative to a center of the second user, and wherein the unified 3-D coordinate plane is used to determine the first set of 3-D coordinates representing the estimated positions of the key points of the user and another set of 3-D coordinates representing estimated positions of key points of the second user.
[0027] In another embodiment of the invention, the set of kinematic parameters of the state vector include velocity parameters and acceleration parameters. In this embodiment, the state vector further includes angular data identifying angular positions of the key points. In this embodiment, the state vector is used to classify the activity being performed by the user, to predict the future intent of the user, and to predict the future motion of the user.
[0028] In another aspect of the invention, a non-transitory computer-readable medium storing instructions is provided. The instructions include one or more instructions that, when executed by one or more processors, cause the one or more processors to receive image data captured by one or more sensor devices, the image data depicting a user in an environment. The one or more instructions, when executed by one or more processors, further cause the one or more processors to determine a first set of three-dimensional (3-D) coordinates representing estimated positions of key points of a user at a time ^^. The first set of 3-D coordinates include one or more null or empty set values based on a portion of the user being occluded from view of the one or more sensor devices. The one or more instructions, when executed by one or more processors, further cause the one or more processors to determine a set of link lengths for linkages connecting the key points.
[0029] The one or more instructions, when executed by one or more processors, further cause the one or more processors to determine, for a time ^^^^, a second set of 3-D coordinates representing predicted positions of the key points and a set of kinematic parameters representing predicted kinematics for the key points. The second set of 3-D coordinates include one or more 3-D coordinates representing predicted positions of one or more key points occluded from the view of the one or more sensors. The second set of 3-D coordinates and the set of kinematic parameters are part of a state vector determined by modeling kinematics of the key points using the first set of 3-D coordinates for the time ^^and the link lengths. The state vectors are determined periodically over a time period ^^^.
[0030] The one or more instructions, when executed by one or more processors, further cause the one or more processors to classify an activity being performed by the user by using machine learning to process the state vectors determined over the time period ^^^. The one or moreinstructions, when executed by one or more processors, further cause the one or more processors to predict a future intent of the user for a time period ^^^. The future intent is predicted by using machine learning to process the state vectors and the classified activity being performed by the user. The one or more instructions, when executed by one or more processors, further cause the one or more processors to predict a future motion path of the user for the time period ^^^by using machine learning to process the user activity and the future intent of the user. The one or more instructions, when executed by one or more processors, further cause the one or more processors to generate a motion assessment for the user by processing the predicted future motion path of the user using machine learning. The one or more instructions, when executed by one or more processors, further cause the one or more processors to make the motion assessment available to a device or account associated with the user.
[0031] In an embodiment of the invention, the one or more instructions, when executed by the one or more processors, further cause the one or more processors to update the state vector by using an extended Kalman filter (EKF) to compare the second set of 3-D coordinates representing the predicted positions of the key points for the time ^^^^with a third set of 3-D coordinates representing observed positions of the key points at the time ^^^^. The updated state vector is provided back to the 3-D kinematic model for iterative learning. The EKF is used to update the state vectors periodically over the time period ^^^.
[0032] In another embodiment of the invention, the one or more instructions, that cause the one or more processors to generate the motion assessment, cause the one or more processors to generate, as the motion assessment, an ergonomic risk assessment by using an ergonomic assessment model to process the predicted future motion path of the user. The ergonomic risk assessment includes one or more of: a safety assessment of past motion of the user, and a recommendation providing one or more corrective motions to improve user safety.
[0033] In another embodiment of the invention, the one or more instructions, when executed by the one or more processors, further cause the one or more processors to determine a third set of 3-D coordinates representing estimated positions of one or more objects in the environment of the user. In this embodiment, the one or more instructions, when executed by the one or more processors, further cause the one or more processors to determine a third set of kinematic parameters representing predicted kinematics for the one or more objects. The third set of 3-D coordinates and third set of kinematic parameters are part of a second state vector. In this embodiment, the one or more instructions, when executed by the one or moreprocessors, further cause the one or more processors to predict a future motion path of the one or more objects using the second state vector. In this embodiment, the one or more instructions, that cause the one or more processors to generate the motion assessment, cause the one or more processors to use the predicted future motion path of the one or more objects when generating the motion assessment.
[0034] In another embodiment of the invention, the one or more instructions, when executed by the one or more processors, further cause the one or more processors to detect an anomaly relating to motion of the user by processing the state vector. In this embodiment, the one or more processors, when generating the motion assessment, are to generate, as part of the motion assessment, a real-time alert based on the detected anomaly.
[0035] In another embodiment of the invention, the set of kinematic parameters of the state vector include velocity parameters and acceleration parameters. In this embodiment, the state vector further includes angular data identifying angular positions of the key points. In this embodiment, the state vector is used to classify the activity being performed by the user, to predict the future intent of the user, and to predict the future motion of the user.
[0036] The above summary presents a simplified overview of some embodiments of the invention to provide a basic understanding of certain aspects of the invention discussed herein. The summary is not intended to provide an extensive overview of the invention, nor is it intended to identify any key or critical elements or delineate the scope of the invention. The sole purpose of the summary is merely to present some concepts in a simplified form as an introduction to the detailed description presented below. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The objects and advantages of the invention disclosed will be further appreciated in light of the following detailed descriptions and drawings in which:
[0038] Fig. 1 is a diagram showing an example environment that includes sensor devices and components of a motion prediction platform according to the principles of the present disclosure.
[0039] Fig. 2A is a diagram showing the sensor devices capturing image data of users and providing the image data to the motion prediction platform.
[0040] Fig.2B is a diagram showing processing performed by a pose estimation module of the motion prediction platform.
[0041] Fig. 2C is a diagram showing processing performed by a point mass filtering model, a 3-D kinematic model, and an extended Kalman filter of the motion prediction platform.
[0042] Fig. 2D is a diagram showing processing performed by an activity recognition model of the motion prediction platform.
[0043] Fig.2E is a diagram showing processing performed by an object recognition model of the motion prediction platform.
[0044] Fig. 2F is a diagram showing processing performed by a motion prediction model of the motion prediction platform.
[0045] Fig. 2G is a diagram is a diagram showing processing performed by an anomaly detection model of the motion prediction platform.
[0046] Fig.2H is a diagram is a diagram showing processing performed by an ergonomic risk assessment model of the motion prediction platform.
[0047] Fig. 2I is a diagram showing real-time outputs generated by an assessment-alert- recommendation (AAR) engine of the motion prediction platform.
[0048] Fig.2J is a diagram showing non real-time outputs generated by the AAR engine of the motion prediction platform.
[0049] Fig.3 is a diagram showing a 3-D skeletal model of a user with labeled key points based on the processing performed by the pose estimation module.
[0050] Fig.4A is a diagram showing the 3-D skeletal model of the user with labeled kinematics based on the processing performed by the 3-D kinematic model.
[0051] Fig. 4B is a diagram showing the 3-D kinematic model predicting positions and kinematics of one or more key points corresponding to a portion of the user that occluded from view of the sensor devices.
[0052] Fig.5 is a diagram showing an architecture of the activity recognition model when the activity recognition model is a hierarchical dynamic graph convolutional network (HD-GCN).
[0053] Fig. 6 is a diagram showing an architecture of the object recognition model when the object recognition model is a bi-directional feature pyramid network (BiFPN).
[0054] Fig. 7 is a diagram showing an architecture of the motion prediction model when the motion prediction model is a graph-conditional generative adversarial network (GCN-GAN).
[0055] Fig. 8A is a diagram showing a first half of the process for training a custom motion model of the motion prediction platform.
[0056] Fig. 8B is a diagram showing the second half of the process for training the custom motion model of the motion prediction platform.
[0057] Fig. 9A is a diagram showing the processing steps performed by a conversion module of the motion prediction platform.
[0058] Fig. 9B is a diagram showing the processing steps performed by a biomechanical analysis model of the motion prediction platform.
[0059] Fig.9C is a diagram showing the processing steps performed by the motion assessment model and the custom motion model of the motion prediction platform.
[0060] Fig. 9D is a diagram showing diagram showing a text-based ergonomic recommendation being provided to a user device.
[0061] Fig. 10 is a diagram of an example environment in which systems and / or methods described herein may be implemented.
[0062] Fig.11 is a diagram of example components of one or more devices of Fig.10.
[0063] Fig.12 is a flowchart of an example process for predicting and assessing motion of one or more users in an environment.
[0064] Fig. 13 is a flowchart of another example process for predicting and assessing motion of one or more users in an environment. DETAILED DESCRIPTION OF THE INVENTION
[0065] The following detailed description of example embodiments refers to the accompanying drawings. The same reference numbers in different drawings may identify the same or similar elements.
[0066] Fig. 1 is a diagram showing an example system 10 that includes one or more sensor devices 12 and components of a motion prediction platform 14 according to the principles of the present disclosure. The motion prediction platform 14 may include a pose estimation module 16, a point mass filter 17, a three-dimensional (3-D) kinematic model 18, an extended Kalman filter (EKF) 20, an activity recognition model 22, a motion prediction model 24, an assessment-alert-recommendation (AAR) engine 26, a motion assessment model 28, a custom motion model 30, an object recognition model 32, a conversion module 34, a biomechanical analysis model 36, and an anomaly detection model 38.
[0067] Each sensor device 12 may include a camera adapted to capture video data, image data, and / or audio data. For example, a camera may be adapted to record an environment, such as a workplace (e.g., a factory, a facility, an office, etc.), an environment in which a user performs an exercise or sport, and / or another type of environment where a motion assessment may be applicable.
[0068] In some embodiments, a sensor device 12 may include an inertia measurement unit (IMU) adapted to measure one or more kinematic parameters, including acceleration data,velocity data (e.g., angular velocity via a gyroscope), and / or the like. In some embodiments, an environment may be configured with multiple sensor devices 12. In some embodiments, the environment may be configured using only non-invasive cameras, such that the users are not required to have wearable devices. In some embodiments, the environment may be configured with a combination of cameras and wearable devices such as IMUs. An example of an environment utilizing multiple cameras is shown and described in connection with Fig. 2A.
[0069] The motion prediction platform 14 may include one or more server devices adapted to analyze past motion of one or more users, to predict future motion of the one or more users, and to generate an assessment based on past and / or predicted future motion of the one or more users. While each of the modules and / or models described below are described as being part of the motion prediction platform 14, it is to be understood that this is provided by way of example, and that in practice, any number of modules and / or models may be hosted by a device other than the motion prediction platform 14, such as a different server device, a user device (e.g., a desktop computer, a mobile device, etc.), etc.
[0070] The pose estimation module 16 may utilize one or more data models adapted to determine a set of three-dimensional (3-D) coordinates representing estimated positions of key points of users in an environment. As used herein, the term “key point” refers to any joint, body part, or area of the user, such as a user’s pelvis, right hip, right knee, right ankle, left hip, left knee, left ankle, spine, thorax, neck, head, left shoulder, left elbow, left wrist, right shoulder, right elbow, right wrist, and / or the like. Example data models of the pose estimation module 16 are shown and described in connection with Fig.2B.
[0071] The point mass filter 17 may be adapted to determine a state vector for an upcoming time ^^^^using physics-based state transition equations. The 3-D kinematic model 18 may be adapted to first determine a set of link length values for linkages of the user, where a linkage represents a part of the user that connects two adjoining key points. Next, the 3-D kinematic model 18 may be adapted to determine a state vector for the upcoming time ^^^^, using kinematic relationships and / or biomechanical constraints. In some embodiments, the system and methods described herein may be implemented without the point mass filter 17. In some embodiments, one or more features described as being performed by the point mass filter 17 may be performed by the 3-D kinematic model 18.
[0072] The EKF 20 may be adapted to update the state vector for the current time ^^. For example, the EKF 20 may compare the 3-D coordinates representing predicted positions of keypoints of the user for the upcoming time ^^^^(received from 3-D kinematic model 18) with a set of 3-D coordinates representing observed positions of the key points for the time ^^^^(received from the pose estimation module 16). That is to say, while the 3-D kinematic model 18 predicts key point positions and kinematics for an upcoming time ^^^^that has yet to occur, the EKF 20 eventually also receives observed position estimates when the current time matches the time ^^^^that the 3-D kinematic model 18 used for making its future predictions. The EKF 20 may update the state vector to account for missing data, noisy data, different values for the same key point, and so forth. The point mass filter 17, 3-D kinematic model 18, and EKF 20 are further shown and described in connection with Fig.2C and Figs.4A and 4B.
[0073] The activity recognition model 22 may be adapted to classify user activity and / or to predict a future intent of the user. The activity recognition model 22 is further shown and described in connection with Fig.2D and Fig.5.
[0074] The motion prediction model 24 may be adapted to predict a future motion path of one or more users. The motion prediction model 24 is further shown and described in connection with Fig.2F and Fig.7.
[0075] The AAR engine 26 may be adapted to generate a motion assessment, such as an ergonomic risk assessment, a performance assessment, and / or another type of assessment. For example, the AAR engine 26 may generate an ergonomic risk assessment that assesses the motion of the user over a time period (e.g., such as the duration of a work shift). Additionally, or alternatively, the AAR engine 26 may generate a motion assessment in the form of an alert. Motion assessments can be generated in real-time and / or in a non-real-time manner. The AAR engine 26 is further shown and described in connection with Figs.2I and 2J.
[0076] The motion assessment model 28 may be adapted to perform a motion assessment based on past and / or predicted future movement of users. The motion assessment model 28 is further shown and described in connection with Fig.2H and Fig.8B.
[0077] The custom motion model 30 may be adapted to generate text based on a labeled 3-D skeletal model of the user. For example, custom motion model 30 may generate a textual description of a motion of one or more users and / or may generate text providing ergonomic feedback on the motion of the one or more users. In addition to having video (motion)-text as the input-output, the custom motion model 30 can also be adapted to support text-video (motion) and / or motion (video) - (motion) video as the input-output. In some embodiments, the custom motion model 30 may be a custom motionGPT model. The custom motion model 30 is further shown and described in connection with Figs.8A and 8B and Fig.9C.
[0078] The object recognition model 32 may be adapted to determine 3-D coordinates representing estimated positions of key points of one or more objects in the environment with the one or more users at a time ^^. The object recognition model 32 may be further adapted to communicate with the motion prediction model 24 such that the motion prediction model 24 can be used to determine 3-D representing predicted positions of the key points of the one or more objects for an upcoming time ^^^^. The object recognition model 32 is further shown and described in connection with Fig.2E and Fig.6.
[0079] The conversion module 34 may be adapted to convert updated estimated positions of key points of users to a format suitable for the custom motion model 30. For example, the EKF 20 may provide updated estimated positions of key points of users for x number of key points. However, the custom motion model 30 may utilize y number of key points, where x is different from y. As such, the conversion module 34 may convert the received key points to a number of key points that is suitable for the custom motion model 30. The conversion module 34 is further shown and described in connection with Fig.9A.
[0080] The biomechanical analysis model 36 may be adapted to determine state vectors and link length orientations. The biomechanical analysis model 36 may be further adapted to determine time series joint angle data. The biomechanical analysis model 36 is further shown and described in connection with Fig.9B.
[0081] The anomaly detection model 38 may be adapted to detect an anomaly or unsafe motion of a user. This can be used to provide real-time feedback to the user. The anomaly detection model 38 is further shown and described in connection with Fig.2G.
[0082] Figs. 2A-2J are diagrams showing example embodiments of the motion prediction platform 14 being used to assess past motion of users, to predict future motion of users, and to generate a motion assessment based on the past and / or predicted future motion of the users. Fig. 2A is a diagram showing sensor devices (e.g., cameras) 12-1, 12-2, and 12-3 capturing image data of users in an environment. While shown as image data, it is noted that cameras 12-1, 12-2, and 12-3 are generating video data by recording the users. However, the term “image data” is used herein because a video comprises a collection of frames (i.e., images) and the motion prediction platform 14 analyzes those frames. As shown by reference number 40, the cameras 12-1, 12-2, and 12-3 may provide the image data to the motion prediction platform 14.
[0083] In some embodiments, the environment may be equipped with cameras and the cameras may capture image data, video data, and / or audio data of the users. For example, camera 12- 1, camera 12-2, and camera-12-3 may be positioned in different locations in a workenvironment (e.g., a warehouse, a facility, an office space, etc.) and may capture image data of the users over a time period such as a work shift. Notably, in this scenario, users are not required to put on wearable devices because all image capturing is done using non-invasive cameras 12-1, 12-2, 12-3.
[0084] In some embodiments, the sensor devices 12 may further include wearable devices, such as inertia measurement units (IMUs). In this case, the IMUs may generate IMU data, including kinematic parameters such as acceleration data, velocity data (e.g., indicating an angular velocity), and so forth. In some cases, a user may have multiple IMUs, where each IMU is dedicated to a particular key point of the user. In this case, the kinematic parameters determined by the IMUs may provide measurements used to iteratively train the 3-D kinematic model 18.
[0085] To the extent that one or more embodiments described herein involve recording a user, it is to be understood that the user has provided consent to such recording. Moreover, any image data collected of a user may be stored in a manner compliant with all applicable privacy laws in the jurisdiction in which the data is generated.
[0086] In some embodiments, the sensors 12 may be configured to provide image data to the motion prediction platform 14 based on a trigger condition being satisfied. For example, the sensors 12 may be configured to provide image data to the motion prediction platform 14 periodically over a configured time period and may transmit the image data to the motion prediction platform 14 after the configured time period.
[0087] Fig.2B is a diagram showing processing performed by the pose estimation module 16 of the motion prediction platform 14. Notably, processing occurs in real-time while the users are moving around in the environment. As shown by reference number 42, the motion prediction platform 14 may determine a set of 3-D coordinates representing estimated positions of key points of users. To make this determination, a series of processing operations are performed using components of the pose estimation module 16, such as a convolutional neural network (CNN) 44, a calibration and orientation module 46, a cuboid proposal network (CPN) 48, and a pose regression network (PRN) 50.
[0088] As shown by reference number 52, the motion prediction platform 14 (e.g., using CNN 44) may process the sensor data to generate two-dimensional (2-D) heatmaps. For example, each image may be passed through a convolutional neural network (CNN) backbone (e.g., HR- Net) to determine feature maps and to produce heatmaps used for key point identification. To provide a specific example, a heatmap may be generated that corresponds to an image capturedby camera 12-1. The heatmap may include bright spots, where each bright spot indicates a high confidence level that a key point is detected.
[0089] In some embodiments, a 2-D heatmap may be generated for each frame received from each respective camera. However, in other embodiments, one or more of the motion analysis steps described herein (e.g., including 2-D heatmap generation) may involve processing only a subset of the total frames received by the respective cameras 12-1, 12-2, 12-3. For example, the motion prediction platform 14 may perform a frame skipping prediction, a frame reduction technique, and / or a sliding window prediction such that only a portion of the frames are processed. By only performing a motion analysis on a subset of the frames, the motion prediction platform 14 is able to provide dynamic, real-time outputs.
[0090] As shown by reference number 54, the motion prediction platform 14 (e.g., using calibration and orientation module 46) may process the sensor data (e.g., image or video data) to determine a unified 3-D coordinate plane. First, the motion prediction platform 14 may use one or more computer vision techniques to calibrate the cameras 12-1, 12-2, 12-3, such as a structure-from-motion (SfM) technique, a technique identifying calibration patterns (e.g., a checkerboard or structured light project), and / or using a similar technique. By using computer vision to process the image data, the motion prediction platform 14 can determine intrinsic calibration parameters (e.g., a camera focal length, an optical center, etc.) and extrinsic calibration parameters (e.g., position, rotation, etc.) of each camera 12.
[0091] Next, the motion prediction platform 14 (e.g., using calibration and orientation module 44) may map a local coordinate system of each camera 12 to a common global 3-D space. For example, the motion prediction platform 14 may use the image data and the calibration parameters to determine a relative position and orientation of each camera in the environment and may define a global reference point (e.g., a center of the observed workspace). This ensures that a key point observed from each respective camera refers to the same fixed 3-D location, regardless of the viewing angle.
[0092] As shown by reference number 56, the motion prediction platform 14 (e.g., using CPN 48) may aggregate the 2-D heatmaps, map the 2-D heatmaps to a 3-D voxel grid, and provide the 3-D voxel grid to PRN 50. For example, the CPN 48 may aggregate 2-D heatmaps such that feature embeddings (e.g., features indicative of key points) are aggregated in a 3-D voxel grid. Feature embeddings may include spatial and / or contextual information about the detected key points. In some embodiments, CPN 48 may define cuboid search regions around detected key point positions and may refine and / or localize the pose estimates.
[0093] Next, the motion prediction platform 14 (e.g., using PRN 50) may determine the 3-D set of coordinates representing the estimated positions of key points of the users. For example, PRN 50 may receive the 2-D heatmaps and the aggregated feature embeddings in the 3-D voxel grid from the CPN 48. To determine the 3-D set of coordinates representing the estimated positions of the key points of the users, the PRN 50 may use a direct regression technique (e.g., the network directly outputs 3-D coordinates for key points) or a voxel-based method (e.g., the system predicts 3-D heatmaps or volumetric grid to refine 3-D localization).
[0094] To provide an example, Fig.3 shows an example diagram 100 of a 3-D skeletal model of the user with labeled key points 102 based on the processing performed by the pose estimation module 16. The 3-D skeletal model represents the detected pose of the user in a 3- D coordinate space. The axes (x, y, z) are labeled, showing the user’s posture relative to a coordinate system.
[0095] The bottom half of Fig. 3 shows a top-down view of the detected pose, a front-facing projection of the 3-D pose, and a side view of the pose estimation. In the top down view, the XY plane represents how the skeletal model is positioned when viewed from above. This projection helps analyze horizontal positioning, such as foot placement and body alignment from an overhead perspective. The XZ project provides a view of the human figure from the front, showing leg and arm alignment. The YZ projection provides insights into forward / backward bending, leaning, and joint angles along the depth axis (z-axis).
[0096] Fig.2C is a diagram showing processing performed by the point mass filter model 17, the 3-D kinematic model 18, and the extended Kalman filter (EKF) 20 of the motion prediction platform 14. As shown by reference number 58, the point mass filter model 17 may determine a state vector for an upcoming time ^^^^using physics-based state transition equations. The state vector for the upcoming time ^^^^may include 3-D coordinates representing predicted positions of key points of users and predicted kinematic parameters for respective key points (e.g., velocities, accelerations, etc.). An illustrative example is provided below. For example, the point mass filtering model 17 may be configured with the following equations: x(t+1) = x(t) + ∆t * v(t) +^ ^^ a^ (1)
[0097] Equations (1), (2), and (3) are used to compute the next time state. In Equation (1), x (t+1) represents a new key point position for the next state. The expression ∆t * v(t) is used to compute displacement due to the current velocity over the time interval ∆t. It represents thedistance traveled if the velocity v(t) remains constant during this interval. The expression^^ a^^accounts for the additional displacement resulting from acceleration (a) over the time interval ∆t. This reflects how acceleration influences the change in position, considering that acceleration affects velocity over time, leading to a change in displacement.
[0098] Equation (2) can be used to predict the velocity at the next time step v(t+1), considering acceleration and elapsed time. The term v(t) refers to the current velocity at time t. This represents how fast a key point is moving at the present moment. Regarding the expression ∆t * a, the change in velocity is due to constant acceleration a over the time interval ∆t. The expression^^ a^^is an additional correction term that accounts for the fact that acceleration may not be perfectly constant over the time step. Equation (2) updates the velocity at each time step t+1 based on the prior velocity, acceleration, and elapsed time. This ensures smooth transitions in velocity estimation which can be used for motion prediction.
[0099] The state vector determined at time t+1 can be represented using a matrix equation. Equation (4) is shown below as an example:
[0100] Equation (4) above represents an example of the state transition model which predicts the next state. In Equation (4), the term ^^is the current state vector at time n. The term ^^^^is the predicted state vector at the next time step n+1. The first matrix represents the current state vector at time n and includes predicted position (x̂), velocity (ẋ), and acceleration (ẍ) values in the x and y directions at the next time step n+1. The second matrix represents a state transition matrix F which governs how the system evolves from the current state ^^to the predicted next state ^^^^. The first row shows that a predicted position of a key point in the x direction is determined based on the current position, plus a term for velocity (∆t * v(t)) and a term for acceleration (^^ a^^). The second row shows that a predicted velocity in the x directionis determined based on v(t) and ∆t * a. The third row shows that acceleration remains constant. The fourth through sixth rows provide the same kinematic relationships applied in the y direction. The third matrix represents the current state vector at time n and includes position, velocity, and acceleration values in the x and y directions.
[0101] A larger state transition model (e.g., matrix) could be implemented to account for the predicted position, velocity, and acceleration in the z direction. In some embodiments, one or more features described as being performed by the point mass filter 17 may be performed by the 3-D kinematic model 18.
[0102] As shown by reference number 60, the motion prediction platform 14 (e.g., using the 3-D kinematic model 18 may determine a state vector for the upcoming time ^^^^using kinematic relationships and biomechanical constraints. Notably, while the point mass filter 17 determines a state vector that incorporates equations indicative of known laws of physics, the 3-D kinematic model 18 determines a state vector that accounts for kinematic relationships and biomechanical constraints relating to joint angles, link lengths, and / or segment orientations.
[0103] First, the 3-D kinematic model 18 may determine a set of link length values for linkages of each respective user. A linkage, or link, refers to a part of the user that connects two adjoining key points. For example, an ankle and a knee may be key points and a lower leg of the user may represent the link between the ankle and knee. In this case, the 3-D kinematic model 18 may determine link lengths by computing a distance between adjoining key points. To provide a specific example, the 3-D kinematic model 18 may determine link lengths using the following equation: L= sqrt [(^ − ^ )^ + ( ^ ^^ ^ ^^ − ^^) + (^^ − ^^) ] (5)
[0104] In Equation (5),by computing the distance between adjacent key points. The variables ^^, ^^, and ^^represent coordinates for a first key point. The variables ^^, ^^, and ^^represent coordinates for a second key point that is adjacent to the first key point. In some embodiments, the set of link lengths may have been determined by the pose estimation module 16 or another component described herein.
[0105] In some embodiments, the 3-D kinematic model 18 may determine an orientation of each respective link. The 3-D kinematic model 18 may determine the orientation of each link by using inverse kinematic equations to process the 3-D coordinates representing the estimated positions of key points of the users and the determined link lengths. For example, the 3-D kinematic model 18 may determine the orientation of each link using the following equations: θ^^= ^^^^^ (^^^ ^^)(6)θ^^= ^^^^^ (^^^ ^^)(^^^ ^^)(7)
[0106] Equations (6), (7), compute the orientation of links based on relative 3-This is a form of inverse kinematics, where key point angles are inferred from position data. The terms θ^^, θ^^, and θ^^represented the computed angles along the x, y, and z axes respectively. The superscript s indicates that these angles belong to a specific key point, such as a shoulder, elbow, or other joint. The angles are derived using the arctangent function which helps determine the orientation of the link in 3-D space. The terms ^^, ^^, and ^^represent the 3-D coordinates of the starting point of a link (e.g., a shoulder joint in an arm segment / link or a hip joint in a leg segment / link). The terms ^^, ^^, and ^^represent 3-D coordinates of the ending point of thelink (e.g., elbow in an arm segment / link or knee in a leg segment / link). The expression ^^ −^^ represents the vertical displacement between adjoining key points. The expression ^^ − ^^represents the lateral displacement between adjoining key points. The expression ^^ − ^^represents the forward-backward displacement between two adjoining key points. The inverse tangent function computes the angle of rotation by measuring the slope of the vector connecting the two points.
[0107] Referring back to Fig. 4A, the middle portion of Fig. 4A illustrates key point (limb) connections and coordinate frames used to compute joint angles for motion analysis. Thediagram depicts a lower limb (^^) and an upper limb (^^), connected at key points (^^^^ !"^#,^^!$^%, ^%#&^'). These segments form a kinematic chain, meaning movement of one jointaffects the position and orientation of the others. The elbow connects the upper and lower limb, forming an articulation point where angles can be calculated.
[0108] Still referring to Fig. 4A, the left side of Fig. 4A shows a 3-D skeletal model of the human body with key point angles labeled as α1 to α12. For example, α1 is a shoulder angle, α2 is a hip joint angle, α3 is a knee joint angle, α4 is an ankle joint angle, α5 is another knee joint angle (opposite leg), α6 is an ankle joint (opposite leg), α7 is a toe joint, α8 is a neck movement or upper spine rotation, α9, α10, α11 are elbow, wrist, and upper body angles, and α12 is head rotation or tilt.
[0109] Referring back to Fig. 2C, the state vector determined by the 3-D kinematic model 18 can be represented as a matrix equation, such as example matrix Equation (9) shown below:
[0110] Equatio d by the 3-D kinematic model 18. For example, Equation (9) can be used to predict key point (joint) positions, angles, and link (limb) lengths by applying a transformation matrix H to the state vector containing kinematic parameters.
[0111] The first matrix, which is shown on the left hand side in Equation (9), representsobserved measurements of the pose estimated by the pose estimation module 16. The variables^(, ^^, and ^( refer to a root position of a key point. This matrix further includes squareddistance variables which represent differences between key point positions in x, y, and zcoordinates. For example, (^( − ^^)^, (^( − ^^)^, and (^( − ^^)^ represent squareddistances between adjacent joints which can be used to compute link lengths. These distances are invariant and help ensure that link length constraints are satisfied throughout motion tracking.
[0112] The second matrix in Equation (9) is an observation matrix H which maps the observed joint positions, distances, and angles (the first matrix), to the internal kinematic state vector (the third matrix). Diagonal elements )(, )^, …, )^*represent weighting factors applied to corresponding state variables. Each&applies transformations to account for human- imposed limitations to key point movement and / or rotational motion. This filters out unrealistic motions, ensuring predicted poses adhere to human biomechanics.
[0113] The third matrix in Equation (9) represents the estimated internal kinematic state of a user. The variables ^(, ^^, and ^(refer to a root position of a key point. The variables ẋ(, ẏ^, and ż(refer to velocity components. The variables ẍ(, ÿ^, and z̈(refer to acceleration components. The variable 2 represents key point orientation in a 3-D space. The variables2^^3 , 2^^3 , 2^^3 refer to joint angles in x-, y-, and z- axes for a first key point. The variables 2^^34 ,2^ , 2^^34 refer to joint angles in x-, y-, and z- axes for a sixteenth key point.In some embodiments, as is shown in Fig.4B, the motion prediction platform 14 (e.g., using the 3-D kinematic model 18 may determine 3-D coordinates representing predicted positions of key points, and predicted kinematic parameters, of users that are partially or fully occluded from view of the cameras 12-1, 12-2, 12-3. For example, and as is shown in Fig.4B, a user may be carrying a box which results in occluded regions 110 which are not visible to the cameras 12-1, 12-2, 12-3. In this case, as is shown by reference number 112, the 3-D kinematic model 18 may receive 3-D coordinates representing the estimated positions of key points of the users for a time t^. The 3-D coordinates may include one or more null or empty set values based on a portion of the user being occluded from view of the cameras 12-1, 12-2, 12-3. As shown by reference number 114, the 3-D kinematic model 18 may determine a state vector for an upcoming time t^^^, where the state vector includes 3-D coordinates representing predicted positions of key points and predicted kinematic parameters for respective key points. In this case, the state vector includes 3-D coordinates representing predicted positions of key points corresponding to occluded region 110 and predicted kinematic parameters for those key points.
[0115] As shown by reference number 62, the motion prediction platform 14 (e.g., using the EKF 20) may update the state vector to include refined predicted positions of key points and / or refined predicted kinematic parameters. For example, the EKF 20 may receive an observed set of 3-D coordinates for key points of the user at time t^^^(i.e., where the current time is now the time which was originally predicted by the 3-D kinematic model 18). At this point, the EKF 20 now has (1) 3-D coordinates representing predicted positions of key points at time t^^^(i.e., at time t^, the 3-D kinematic model 18 provided the predicted positions for time t^^^to the EKF 20)) and (2) 3-D coordinates representing estimated positions of key points for the observed time of t^^^(i.e., where the current time is now the time which was originally predicted by the 3-D kinematic model 18). The EKF 20 may then compare these values to compute an error between predicted and actual values and may apply a correction to improve the accuracy of the data.
[0116] In some embodiments, the EKF 20 processes the above-identified inputs using a diagonal covariance matrix A and a process noise covariance matrix Q. Example matrices are provided below:∆78 ∆7: 9^ ∆tA =∆7:
[0117] In the diagonal covariance A represents the state transition model which predicts how theover time. That is to say, it captures how the position, velocity, and acceleration change from one time step to the next. Each term in matrix A is derived from kinematic motion equations such as Newtonian mechanics. In theexample shown, higher-order terms (∆t9, ∆t^, ∆t^) account for acceleration effects over time.Matrix A models how position, velocity, and acceleration interact at each time step. For example, the first row of matrix A corresponds to a position prediction, the second row of matrix A corresponds to a velocity prediction, and the third row of matrix A corresponds to 8 acceleration tracking. For example, in the first row, the matrix value∆79 derives from kinematic equations of motion that predict positional uncertainty based on acceleration. The fourth power (∆t9) arises because acceleration influences position over time squared and the motion prediction platform 14 computes the effect over multiple time steps. That is to say, this term accounts for the effect of acceleration compounded over time when predicting position. ∆7:Continuing with the first row, the matrix value^represents the contribution of velocityuncertainty to position estimation. It arises acceleration effects into velocity over time (∆t^factor). Continuing with the first row, the matrix value ∆t represents direct influence of velocity on position. In the second row, the matrix value∆ :
[0118] 7^ tracks the effect of acceleration but now in relation to a velocity prediction. Continuing with the second row, the matrix value ∆t^represents that the velocity at the next time step depends on the acceleration acting over time. Continuing with the second row, the matrix value ∆t ensures that velocity accumulates correctly over time.
[0119] In the third row, the matrix value a is the acceleration component acting on the system. This defines how acceleration contributes to velocity updates. Continuing with the third row, the matrix value ∆t accounts for the fact that acceleration influences velocity over time.Continuing with the third row, the matrix value 1 ensures that acceleration itself remains constant over small time steps.
[0120] The process noise covariance matrix Q represents uncertainty in the state transition model due to unmodeled dynamics, sensor noise, and / or external disturbances. This accounts for real-world variations that are not captured by the motion equations. In the process noise covariance matrix Q, each matrix value captures how random noise affects various parts of the state vector. That is to say, each matrix value is associated with different sources of noise affecting position, velocity, and acceleration. For example, each value <&represents process noise in a specific state variable. This allows the motion prediction to remain flexible while reducing the likelihood of inaccurate predicted values.
[0121] In the example shown in matrix Q, the matrix value <(models uncertainty in position measurements. For example, if a camera misreads a position due to environmental noise, this value captures the expected variance in that measurement. The matrix value <^captures uncertainty in velocity prediction. For example, if the velocity of a key point is predicted but has measurement noise, <^accounts for this. The matrix value <^*captures noise causing fluctuations in acceleration. For example, external forces such as unexpected user movements, impacts, or vibrations can cause variations in acceleration. Off-diagonal matrix values are typically zero because noise in one state component (e.g., position) does not directly affect another state component (e.g., acceleration). If correlated noise occurs, such as a shaky camera causing both velocity and acceleration errors, then off-diagonal matrix values may be non-zero.
[0122] In some embodiments, the EKF 20 may use matrix A and Q to update the state vector to include refined 3-D coordinates representing predicted positions of key points and refined kinematic parameters representing predicted kinematics of respective key points. For example, the EKF 20 may predict the next state, may predict covariance, and may update the predicted state using a Kalman gain.
[0123] In some embodiments, as is shown in Fig. 2C, the updated state vector determined by EKF 20 may be provided back to the 3-D kinematic model 18 as part of a feedback loop. This allows the 3-D kinematic model 18 to be further trained in real time by being provided with data that can be compared against older predictions to determine if any weights of the 3-D kinematic model 18 need to be adjusted.
[0124] To provide an example, the 3-D kinematic model 18 may receive pose estimation data from pose estimation module 16 for a current state ^^at time ^^. The 3-D kinematic model 18 may predict the next state ^^^^ | ^at upcoming time ^^^^using the following equation:x^^^^ | 'A– A*x^^(10)
[0125] In Equation (10), the variable A is the state transition matrix which models motionusing kinematic equations. The variables x^^^^ | 'A represent a predicted state at upcoming time^^^^ based on the state at time ^^. The “| ^^” portion signifies that this value is a predictionmade using only the information available at time ^^, before incorporating any new measurements at time ^^^^. Next, the EKF 20 may update the state vector using the following equation: x^^^^= x^^^^ | 'A+ K (^^^^^– H * x^^^^ | 'A) (11)
[0126] In Equation (11), x^^^^is the updated state at time ^^^^. The variable K is the Kalmangain which determines how much the prediction is adjusted based on new data. The variables^^^^^ refer to the actual camera measurement at time ^^^^. H is an observation matrix thatmaps the state into a measurement space. The expression ^^^^^– H * x^^^^ | 'Arepresents the difference between predicted and observed values.
[0127] In some embodiments, the motion prediction platform 14 may generate a 3-D skeletal model of each respective user. For example, the motion prediction platform 14 may generate a 3-D skeletal model which labels key points and linkages for each user. The 3-D skeletal model may be displayed on a user interface and / or used in one or more of the additional processing steps described further herein. An example of a 3-D skeletal model is shown in connection with Fig.4B.
[0128] By utilizing one or more of the features described above, the motion prediction platform 14 ensures biomechanical realism by enforcing joint angle limits and body segment constraints. Further, the 3 motion prediction platform 14 uses inverse kinematics to refine key point orientations. Furthermore, using one or more features described above allows the motion prediction platform 14 to reduce errors from noisy detections. Still further, by using the 3-D kinematic model 18 to make key point and kinematic predictions for occluded regions, the motion prediction platform 14 provides a real-time motion prediction even when portions of a user are not visible to the cameras (e.g., cameras 12-1, 12-2, 12-3, etc.). Moreover, by using EKF 20 as part of a feedback loop, the 3-D kinematic model 18 is able to iteratively learn in real-time, further improving the accuracy of the predictions.
[0129] Fig. 2D is a diagram showing processing performed by the activity recognition model 22 of the motion prediction platform 14. As shown by reference number 62, the activity recognition model 22 may receive updated state vectors from the EKF 20. For example, updated state vectors determined by EKF 20 may be stored in a cache and may be periodicallyprovided to the activity recognition model 22. To provide a specific example, the cache may store updated state vectors for over a time period TP defined as ^^^B(, ^^^9C, …, ^^^^, ^^.Notably, ^^ in time period TP refers to a current time and thus is a different time than the time^^ discussed in connection with the pose estimation module 16 and point mass filter 17.
[0130] As shown by reference number 64, the motion prediction platform 14 (e.g., using the activity recognition model 22) may classify an activity being performed by each respective user.
[0131] In some embodiments, the activity recognition model 22 may be a high-dimensional graph convolutional neural network (HD-GCN). In this case, the activity recognition model 22 may classify an activity of the user by using one or more classification layers to process the updated state vectors collected over the time period TP. The activity recognition model 22 may include a set of GCN blocks where each respective GCN block is adapted to process the updated state vectors at increasing levels of abstraction. This allows the model to learn meaningful spatial and temporal patterns over time. For example, a first subset of GCN blocks may identify low-level features, such as local key point movements, while a second subset of GCN blocks focus on higher-level semantics, such as full-body motion patterns.
[0132] Each GCN block may include an HD-graph convolution feature, an attention-guided hierarchy aggregation (A-HA) feature, and a temporal convolution feature. The HD-graph convolution feature processes spatial information from the updated 3-D coordinates representing refined estimated positions of the key points of users. The HD-graph convolution feature treats the human body as a graph, where key points are nodes and linkages are edges. The HD-graph convolution feature performs graph convolutions to identify relationships between key points. For example, assume the following matrix represents the refined estimated positions of key points of a user: ^é ^G^HI ^^G^HI ^^G^HIê^^^GJKLI^M ^^^GJKLI^M ^^^GJKLI^Mùúê^^^ ^^ ^^ úêLNJO ^LNJO ^LNJO úë ^^OMP^Q ^^OMP^Q ^^OMP^Q û
[0133] In the matrix above, each row represents a joint position in 3-D space. A graph adjacency matrix may be constructed based on skeletal connectivity (e.g., elbow is connected to the wrist and shoulder). A convolution features from neighboring key (e.g., the wrist’s motion is influenced by the set of key point featureembeddings, where each row contains learned spatial features rather than raw 3-D coordinates. A matrices representation of this output is provided below. Ué ^^V"ù êU^^^ !"^#ê U^!$^%úú êU%#&^' úë … û
[0134] The A-HA feature refines spatial relationships by learning which key points are more important for a given activity. The A-HA feature uses attention mechanisms to dynamically assign weights to different key points based on motion context. A-HA may also aggregate key point information, focusing on key output from the HD-graph convolution feature may be provided as input to The A-HA feature uses a learnable attention matrix to assign weights Key points contributing more to activity recognition (e.g., wrists for walking, etc.) receive higher attention scores. Features are aggregated individual key point features into limb- level and body-level features.outputs a refined hierarchical motion representation that emphasize the most important key points. An example matrix representation of the output is provided below. XYY^# $^"^W X!^%^# $^"^ [XZ^#^ ^'V$&!&'^
[0135] The temporal convolution feature captures temporal dependencies in motion, analyzes how key point positions change over time, and applies one-dimensional (1-D) convolution operations over the temporal dimension. A hierarchical motion representation across multiple time steps may be provided as input to the temporal convolution feature. An example matrix representation of the input is provided below. éX'YA ' 'Y\3^# $^"^X AYY^# $^"^ X A]3YY^# $^"^' ' ' ùêX A\3 A A]3! X Xúê^%^# $^"^ !^%^# $^"^ !^%^# $^"^''A úëX A\3 ]3Z^#^ ^'V ^'V$&!&'^ û
[0136] The temporal convolution temporal convolutional kernel slides over the time dimension. The outputs a compressed time-series feature that captures motionin walking versus a sudden stop). This enables the activity recognition model 22 to predict ongoing activity and future movements.
[0137] Lastly, a fully connected (FC) layer may be used to classify an activity of the user. The classified activity may be running, sitting, lifting a box, and / or any other activity that may be performed by a user in the environment. In the FC layer, every input neuron may be connected to every output neuron. The FC layer may receive the output of the temporal convolution as input (e.g., which encodes spatiotemporal motion patterns). The FC layer may process this data using linear transformation (e.g., matrix multiplication and bias addition) and / or may apply an activation function such as Softmax. The activation function converts raw scores into probabilities. The FC layer may output a probability distribution over different activity classes and the highest probability may be selected as the classified activity of the user.
[0138] To provide another example, assume there are nine GCN blocks, such as is shown in the example architecture 116 of an HD-GCN in Fig. 5. In this case, a first subset of GCN blocks (e.g., blocks 1-3) may be dedicated to low-level processing. For example, blocks 1-3 may identify local spatial relationships between key points, identify small movements of key points, and may focus on short-term dependencies (e.g., frame-by-frame motion). A second subset of GCN blocks (e.g., blocks 4-6) may be dedicated to intermediate processing. For example, blocks 4-6 may capture motion sequences rather than static frames. Blocks 4-6 may aggregate information across time using A-HA and may identify representative movement patterns (e.g., walking cycles). A third subset of GCN blocks may be dedicated to high-level processing. For example, blocks 7-9 may recognize complete activity patterns (e.g., running, lifting a box, etc.), may perform motion reasoning by leveraging long-term dependencies, and may determine semantic meaning of movement (e.g., distinguishing between “walking with a box” and “placing a box down”).
[0139] As shown by reference number 66, the motion prediction platform 14 (e.g., using the activity recognition model 22) may predict a future intent of the users. For example, the activity recognition model 22 may include one or more intent prediction layers that may be used to predict a future intent of the users. The classified activity of the user may be provided as an input to the one or more intent prediction layers. Motion features may be aggregated to help determine the future intent of the user. In some embodiments, the same GCN blocks may be used to predict future intent, but these GCN blocks may be optimized for future motion modeling. In other embodiments, another machine learning technique or model may be implemented.
[0140] To provide an example output, if the classified activity is “lifting a box”, the predicted intent of the user may be “placing the box on a shelf.” If the classified activity is “walking toward a box”, the predicted intent of the user may be “picking up the box”. Outputsdetermined by the activity recognition model 22 may be provided to the motion prediction model 24, as is shown and described in connection with Fig.2F.
[0141] In this way, the motion prediction platform 14 classifies activities and predicts future intent of the users. As will be described further herein, this may be used to predict a future motion path of each user and / or may be used for performing a motion assessment of each user.
[0142] Fig.2E is a diagram showing processing performed by the object recognition model 32 of the motion prediction platform 14. As shown by reference number 68, the motion prediction platform 14 may determine a set of 3-D coordinates representing estimated positions of one or more objects in the environment of the users.
[0143] The object recognition model 32 may receive, as input, the image data captured by the cameras 12-1, 12-2, and 12-3. The object recognition model 32 may also receive, as input, the 3-D coordinates representing estimated positions of key points of users (e.g., the output of the pose estimation module 16. In some embodiments, rather than receiving the output of the pose estimation module 16 as an input, the object recognition model 32 may receive the updated state vector determined by EKF 20.
[0144] Next, the motion prediction platform 14 (e.g., using the object recognition model 32) may determine the set of 3-D coordinates representing the estimated positions of one or more objects (or one or more key points of objects) in the environment of the users. For example, the object recognition model 32 may determine 3-D coordinates representing estimated positions of one or more key points of the objects or may determine 3-D coordinates representing boundaries of the objects.
[0145] In some embodiments, the object recognition model 32 may be a feature pyramid network (FPN), such as a bidirectional feature pyramid network (BiFPN). The BiFPN includes multi-scale feature identification, bi-directional fusion, weighted feature fusion, and efficient scaling for real-time use. Multi-scale feature identification involves identifying feature maps at multiple resolutions. Bi-directional fusion allows for both top-down and bottom-up feature aggregation, reinforcing feature consistency across scales. The weighted feature fusion assigns different weights to each input feature map based on its importance for the final estimation. Unlike traditional FPNs, BiFPN removes unnecessary connections and balances computational cost with accuracy. The BiFPN outputs the set of 3-D coordinates representing the estimated positions of one or more objects (or one or more key points of objects) of the users.
[0146] An example architecture of a BiFPN is shown in Fig.6 using reference number 118. In this example, the image data may be processed by a backbone network (e.g., ResNet, EfficientNet, etc.) to identify feature maps. As can be seen in Fig.6, the BiFPN includes featurepyramid levels3̂^ ,_̂9 , …,^̀^^a.3̂^ is the highest resolution, resulting in the largest feature map including low-level features.^̀^^ais the lowest resolution, resulting in the smallest feature map that includes high-level semantic features. Each feature map is processed to identify relevant information at multiple scales.
[0147] The middle portion of the BiFPN provides bi-directional feature fusion. For example, the multi-scale feature maps undergo a bi-directional information exchange, where the top- down pathway (solid black arrows) uses high-resolution details refine low-resolution feature maps, and where bottom-up pathway uses lower-resolution, high-level semantic features to reinforce the high-resolution feature maps. The network dynamically re-weights feature contributions, improving efficiency and accuracy.
[0148] Colored (or differently shaded) feature maps represent different levels of abstraction.For example, red features ( :̂ ) represent high detail, low-level spatial f 8̂a eatures. Purple (^*)represents intermediate maps capturing structural detailb̂s. Green features (^^) representsemantic-rich feature maps. Blue features ( 4̂*9) represent more abstract representations ofobjects. Yellow features (^̀^^a) represent the highest-level features, capturing global object details.
[0149] The subnets can be seen on the right side of the BiFPN. Each subnet processes BiFPN outputs for different tasks. Class Net predicts the object class (e.g., chair, box, forklift, etc.) and is used for object recognition. Box Net determines a bounding box of the detected object and is useful for object localization in a 3-D space. Rotation Net estimates a rotation angle of an object. Translation Net computes the spatial movement of an object (i.e., a change in position over time). These values can be used for motion tracking and / or future motion prediction. In some embodiments, the BiFPN (or a related subnet) may classify the type of object and / or may predict an intent of the object for an upcoming time ^^^^. In some embodiments, outputs from the BiFPN may be provided to one or more kinematic models described herein and / or to activity recognition model 22 and these models may be trained to determine the kinematic parameters associated with the object and / or may be trained to classify the object and to predict the future use of the object.
[0150] In some embodiments, the object recognition model 32 may be based on the component recognition described in Deep Learning-Based Recognition of Manufacturing Components Using Augmented Reality For Worker Training of Assembly Tasks, the contents of which is incorporated herein by reference. See, e.g., Deep Learning-Based Recognition ofManufacturing Components Using Augmented Reality For Worker Training of Assembly Tasks, by S. Deshpanda, M. Raj Aryal, Sam Anand, Proceedings of the ASME 2024 19thInternational Manufacturing Science and Engineering Conference (June 71-21, 2024). In this embodiment, the object recognition model 32 is used for identifying any tools that a worker is interacting with assembly tasks and human-machine interaction. The model 32 leverages physics-based rendering methods to generate image segmentation masks and labels for training a deep model to automatically recognize the components. This allows the model 32 to identify objects in real-time.
[0151] In some embodiments, the object recognition model 32 may perform pixel-by-pixel analysis. This allows the model 32 to capture miniscule details such as sub-components of a tool, wires, and so forth, thereby enabling ergonomic motion predictions and recommendations that are extremely detailed and helpful for a user who may need to correct their form while using a particular tool
[0152] Additionally, or alternatively, the object recognition model 32 may be trained to estimate positions of other objects in the environment of the users. For example, model 32 may be trained to estimate positions of 3-D objects in motion. This enables motion predictions and recommendations such as warning a user of a potential collision with a moving object in the user’s environment.
[0153] In this way, the motion prediction platform 14 (e.g., using the object recognition model 32) determines 3-D coordinates representing estimated positions of objects in the environment with the users and may determine the relative positions of these objects to the users in the environment. Further, the motion prediction platform 14 (e.g., using the object recognition model 32 and / or one or more of the other models described herein) may determine kinematic parameters for the objects, may classify the activity in which the object is being used for, and / or may predict a future use and / or position of the object.
[0154] Fig.2F is a diagram showing processing performed by the motion prediction model 24 of the motion prediction platform 14. As shown by reference number 70, the motion prediction platform 14 (e.g., using the motion prediction model 24) may predict a future motion path of the users and / or objects in the environment. The future motion path of the users and / or objects can be predicted in real-time so that potentially risky contact with any objects, machines, or vehicles in the environment, can be avoided. The future motion path may be predicted for an upcoming time period. The key difference between the future intention prediction of the activity recognition model 22 and the motion prediction of the motion prediction model 22 isthat the activity recognition model 22 predicts ‘what’ a user will do in the future, while the motion prediction model 24 predicts ‘how’ that user will move in the future.
[0155] The motion prediction model 24 may receive, as input, the updated state vectors determined by the EKF 20 over the time period TP. Additionally, or alternatively, the motion prediction model 24 may receive, as input, an activity classification for the users and / or the predicted future intent of the users as determined by the activity recognition model 22. Additionally, or alternatively, the motion prediction model 24 may receive, as input, a state vector representing estimated positions and / or kinematics of one or more objects in the environment as determined by the component recognition model 32.
[0156] In some embodiments, the motion prediction model 24 may be trained to predict a future motion path based on the activity classification of a user and the updated state vectors which include at least the 3-D coordinates representing the estimated positions of key points of the user. For example, assume the motion prediction model 24 is a graph-conditional generative adversarial network (GCN-GAN). In this case, the GCN-GAN may use classification labels to generate motion sequences for key points using a GAN. Graph-based data generation may occur using a graph-conditional GAN (graph-cGAN) to ensure the preservation of node features, edge indices, and Cartesian positions. This allows the GCN- GAN to output the predicted future motion path of the user for the future time period.
[0157] Reference number 120 in Fig.7 provides an example architecture of a GCN-GAN. As shown in Fig.7, the GCN-GAN includes a conditional graph generator. Initial key point data (e.g., determined by pose estimation module 16), along with labeled data, such as the classified activity being performed by the users (determined by activity recognition model 22), the predicted future intent of the user (also determined by activity recognition model 22), and / or the updated state vectors (determined by the EKF 20) are provided as input to the conditional graph generator. The conditional graph generator may also receive random noise as input. This introduces a stochastic input vector to encourage diverse motion outputs. Noise may be sampled from a Gaussian distribution or using another sampling technique known in the art. The random noise is provided because generative adversarial networks (GANs) need variability in their outputs. Without noise, the GAN might produce only a single type of motion per activity class. Noise helps the GAN produce a variety of plausible motion paths, such as different styles of walking, lifting, or turning.
[0158] To process the input data, the conditional graph generator may generate a graphical representation of human motion. That is to say, the input data, such as key points, velocities, accelerations, classifications labels, intent labels, etc., are structured as a graph, where nodesrepresent key points and edges define relationships between adjoining key points. Next, the conditional graph generator may use one or more GCN layers which learn spatial dependencies between key points, such as how the knee moves relative to the hip. Information is propagated between nodes to model biomechanical constraints. Since motion is a sequential process, the conditional graph generator performs temporal encoding by processing frame-by-frame motion evolution. For example, the conditional graph generator uses one or more recurrent graph neural networks (RGNNs) or one or more transformers to learn how key points evolve over time. The output is a predicted motion graph for upcoming time steps ^^^^, ^^^^, … ^^^^. Each node (key point) has an updated position, velocity, and acceleration at each upcoming time step.
[0159] Furthermore, training is enhanced using true graph data and a GCN discriminator. The true graph data (e.g., a Nanyang Technological University (NTU) red green blue (RGB) 120 dataset) may be provided to the GCN discriminator as an input. The NTU RGB 120 dataset includes RGB video, depth maps, 3-D skeletal model data, and infrared sequences for improving action recognition models. Users shown in the videos of the NTU dataset are treated as graph structured key points (using pose estimation module 16). This allows the GAN to be trained using the NTU RGB 120 dataset to predict a spatial-temporal evolution of key points when their initial positions are provided.
[0160] As can be seen in Fig. 7, bidirectional communication occurs between the GCN discriminator and conditional graph generator for training purposes. For example, the GCN discriminator may be part of an adversarial training setup and may evaluate the quality of generated motion sequences and may provide feedback to refine predictions made by the conditional graph generator.
[0161] To provide a specific example, the conditional graph generator may take labeled data, initial key points, and random noise to generate a predicted motion sequence for future time steps. The motion sequence is passed to the GCN discriminator. The GCN discriminator evaluates the predicted motion by comparing it to real motion data (e.g., from the NTU RGB 120 dataset). For example, the GCN discriminator may analyze motion dynamics, smoothness, biomechanical feasibility, and / or the like. The GCN discriminator may provide loss gradients to the conditional graph generator based on how much the generated motion deviates from real- world motion patterns. The conditional graph generator adjusts its parameters to reduce this loss, thereby improving the ability to generate realistic motion sequences over time. This allows the conditional graph generator to be trained to predict future motion of a user.
[0162] Referring back to Fig.2F, in some embodiments, the first motion model 24 may receive, as input, 3-D coordinates representing estimated positions and / or kinematic parameters of one or more objects in the environment of the users. The first motion model 24 may determine a predicted future motion path for the one or more objects in the same or a similar manner as that described in connection with the predicted future motion path of the users. Furthermore, the first motion model 24 may determine the predicted future motion path of both the users and the one or more objects so as to know the relative positioning of each user and object in the environment.
[0163] Fig. 2G is a diagram showing processing performed by the anomaly detection model 38 of the motion prediction platform 14. As shown by reference number 72, the motion prediction platform 14 (e.g., using the anomaly detection model 38) may detect an anomaly relating to motion of a user. An anomaly, as used herein, may refer to an unusual motion by a user, a dangerous motion by a user, an atypical motion by a user, and / or any other motion not typically performed by a user in the context of the user’s environment and / or in the context of a task or action being performed by the user.
[0164] The anomaly detection model 38 may receive, as input, the 3-D coordinates of the estimated positions of key points of the users (e.g., from pose estimation module 16) and / or may receive the predicted future motion paths for the users (e.g., from the motion prediction model 24).
[0165] In some embodiments, to detect an anomaly, the anomaly detection model 38 may identify unexpected or dangerous behavior by comparing detected motion patterns to normal motion. For example, anomaly detection may be used to detect a fall when the user’s joint trajectory suddenly changes by a threshold amount. Additionally, or alternatively, the anomaly detection model 38 may detect unusual or high-risk postures, such as excessive bending, unsafe lifting, and / or the like. Additionally, or alternatively, the anomaly detection model 38 may detect a dangerous collision, such as a collision between two users or a collision between a user and an object. Additionally, or alternatively, the anomaly detection model 38 may detect that movement exceeds a recommended ergonomic limit. In some embodiments, the anomaly detection model 38 may be part of the motion assessment model 28.
[0166] Notably, the anomaly detection model 38 is reactionary in the sense that it reacts to completed actions. However, the motion prediction model 24 can predict that a worker will step into a hazardous zone, overbend, or reach beyond a certain range, resulting in an early warning issued to notify the worker that they have exceeded a safe bending angle.
[0167] Fig. 2H is a diagram showing processing performed by the motion assessment model 28 of the motion prediction platform 14. As shown by reference number 74, the motion prediction platform 14 (e.g., using the motion assessment model 28) may perform a motion assessment of the users. One or more embodiments herein refer specify that the motion assessment is an ergonomic risk assessment. The ergonomic risk assessment may include a risk or safety assessment of a user and / or a recommendation providing one or more corrective (e.g., safer) motions that can be performed by the user. In other embodiments, the motion assessment may include a performance assessment. The performance assessment may include a detailed assessment of performance metrics relating to the task being performed and / or a recommendation for one or more actions capable of improving performance.
[0168] In some embodiments, the motion assessment model 28 may receive, as input, 3-D coordinates of estimated positions of key points of the users. The motion assessment model 28 may also receive, as input, the future motion path prediction for the users.
[0169] In some embodiments, such as when the motion assessment model 28 is an ergonomic risk assessment model, equations may be implemented to model muscle and external forces affecting key points. For example, the motion assessment model 28 may be configured withthe following equation:cdB^e^ = ∑ (X^^' * f^^') + ∑ (X^ ^Z!^ * f^ ^Z!^) – inertial force (12)
[0170] In Equation (12), the variable cdB^e^represents the sum of all forces (e.g., allcontributing muscle moments and external moments) acting on the L5-S1 joint. The variableX^^' represents external forces (e.g., ground reaction forces, weight of an object being lifted,etc.). The variable f^^'is the perpendicular distance from the point of force application to the L5-S1 joint. The greater the distance, the larger the external moment exerted on the lower back. The variable X^ ^Z!^represents the force produced by muscles crossing the L5-S1 joint to support the back. The variable f^ ^Z!^is the moment arm of the muscle forces relative to the L5-S1 joint. A larger moment arm means the muscles need to exert more force to stabilize the body. Inertial forces represent forces resulting from the acceleration of linkages while in motion. When lifting or moving quickly, these forces add to the total moment on the lower back.
[0171] To summarize, Equation (12) calculates the lower back moment (torque) at the L5-S1 joint. This can be used as part of an ergonomic risk assessment to evaluate the potential strain on a user’s lower back when performing physical tasks such as lifting. In particular, Equation (12) may be used to determine whether the forces exerted on the lower back exceed safeergonomic thresholds. For example, if the moment exceeds 200 Newton meters (Nm) during heavy lifting, the risk of lower back injury increases significantly. The ergonomic risk assessment model may use such calculations to analyze safe lifting techniques, posture adjustments, and workload recommendations. Equation (12) is provided by way of example. In practice, the risk assessment model may be configured with any number of similar equations each corresponding to one or more key points.
[0172] In some embodiments, the motion assessment model 28, such as the ergonomic risk assessment model, may output values from a biomechanical and postural analysis. For example, the ergonomic risk assessment model may output values such as those determined using Equation (12). To provide a specific example, the biomechanical and postural analysis may calculate the stress on a key point in motion, may monitor key point movements and angles to ensure the movements and angles are safe (as determined by threshold values), may determine muscle force estimates which estimate the force exerted by muscles while a user is completing a task, may determine inertial forces which capture the effects of acceleration on key point loads, and / or the like. In some embodiments, the biomechanical and postural analysis performed by the ergonomic risk assessment model corresponds to the analysis performed by the biomechanical analysis model 36, which is shown and described in connection with Fig. 9B.
[0173] Additionally, or alternatively, the motion assessment model 28, such as the ergonomic risk assessment model, may output a set of ergonomic risk scores. For example, the ergonomic risk assessment model may generate ergonomic risk scores for key points and may determine a risk level based on the risk scores. For example, key point movements may be determined to be safe, moderate risk, or high risk, depending on the ergonomic risk scores assigned to key points or key point movements. To provide a specific example, a wrist flexion of 0-5 degrees may be marked as being safe or having a low risk level, a wrist flexion of 5-30 degrees may be marked as having a moderate risk level, and wrist flexion with more than 30 degrees may be marked as a high risk level.
[0174] Additionally, or alternatively, the motion assessment model 28, such as the ergonomic risk assessment model, may output posture and / or motion assessments. For example, the ergonomic risk assessment model may generate a posture assessment that indicates that unsafe postures are detected. As another example, the ergonomic risk assessment model may generate a motion assessment indicating that high repetition may cause repetitive strain injury (RSI) risks. As another example, the ergonomic risk assessment model may generate a task safetyassessment indicating proper techniques or form for completing a task and / or that indicates a recommended change to the technique or form for completing the task.
[0175] Additionally, or alternatively, the motion prediction platform 14 may perform a Rapid Entire Body Assessment (REBA). For example, the motion prediction platform 14 (e.g., using the motion assessment model 28) may perform a REBA assessment that considers real-time key point angle computations. In this case, the motion prediction platform 14 determines key point angles while the user is moving around in the environment. For example, the motion prediction platform 14 compares key point angles with accepted threshold values for each key point angle. If the key point angle is within a safe threshold, no action is taken and the motion prediction platform 14 continues determining key point angles. If a key point angle is not within a safe threshold, the motion prediction platform 14 may trigger a real-time.
[0176] Continuing with the example, the motion prediction platform 14 may store all key point angle data time data until the end of a work shift for further analysis. The motion prediction platform 14 may then separately analyze each linkage (segment) of the user. The motion prediction platform 14 may then determine an initial risk score for each linkage based on a deviation of the motion from a configured (e.g., neutral) posture. In some cases, the motion prediction platform 14 may also consider additional factors such as load handling and static postures.
[0177] Continuing with the example, the motion prediction platform 14 may then combine each linkage score into a final risk score, i.e., a REBA score. Next, the motion prediction platform 14 may compare the final risk score to configured risk levels. For example, the motion prediction platform 14 may be configured with risk levels relating to musculoskeletal disorders (MSDs). To provide a specific example, final risk score / score ranges may be 1, 2-3, 4-7, 5-10, and 11+. A score of 1 may correspond to a negligible risk level, where no action is required. A score of 2-3 may correspond to a low risk level, where a change may be needed. A score of 4-7 may correspond to a medium risk level, requiring further investigation and relatively fast posture correction. A score of 8-10 may correspond to a high risk level, requiring immediate investigation and fast posture correction. A score of 11+ may corresponds to a very high risk level, where immediate posture correction is required. This is provided by way of example, and in practice, the motion prediction platform 14 may be configured with any number of different risk level ranges.
[0178] Continuing with the example, the motion prediction platform 14 may then analyze a highest contributing key point angle and may generate custom improvement suggestions. For example, the motion prediction platform 14 may determine the most critical key pointcontributing to the final risk score is the left wrist with an angle of 167.32 degrees. This indicates high wrist strain according to the guideline stating that wrist angles beyond 15 degrees suggest high wrist strain. Example improvements may include adjusting the height of the desk or chair to ensure the wrists are in a neutral position and not excessively flexed, using a wrist rest or ergonomic keyboard / mouse to maintain a more natural wrist alignment during work, and taking regular breaks to stretch the wrists and perform wrist exercises to prevent strain.
[0179] Continuing with the example, the motion prediction platform 14 may integrate the final risk score (e.g., from the REBA analysis) with the personalized ergonomic guidance (e.g., the custom improvement suggestions). While outputs are discussed in connection with Figs. 2I and 2J, it is noted that the results for this embodiment may be provided to a user (e.g., for display on a user interface) and may include the final risk (e.g., REBA) score, risk level(s), and the AI-generated ergonomic feedback.
[0180] In this way, the motion prediction platform 14 (e.g., using motion assessment model 28) performs a motion assessment, such as an ergonomic risk assessment, which can be used to assess the safety and / or performance of the users. As will be explained below, the ergonomic risk assessment also allows real-time alerts to be generated, and allows recommendations and reports to be generated for offline consumption, thereby improving user safety and / or user performance.
[0181] Fig.2I is a diagram showing real-time outputs generated by the AAR engine 26 of the motion prediction platform 14. As shown by reference number 76, the motion prediction platform 14 (e.g., using AAR engine 26) may generate an alert and / or a recommendation in real-time.
[0182] In some embodiments, an alert and / or recommendation may be generated based on the detected anomaly determined by the anomaly detection model 38. For example, assume a user picks up a heavy box and, rather than bending their knees, bends straight over to pick the box up. This may cause the anomaly detection model 38 to detect an unsafe motion and to provide anomaly data indicative of the unsafe motion to the AAR engine 26. The AAR engine 26 may process the anomaly data and may generate an alert and / or recommendation that can be provided to the user in real-time. For example, the AAR engine 26 may generate an alert indicating that it is unsafe to pick up heavy materials using improper form. The AAR engine 26 may also generate a recommendation that includes instructions for the user on how to complete the task or motion using proper form.
[0183] In some embodiments, an alert and / or recommendation may be generated based on the motion assessment determined by the motion assessment model 28. For example, theergonomic risk assessment model may output a set of ergonomic risk scores and one of the risk scores may correspond to a wrist flexion of more than degrees. This may cause the ergonomic risk assessment model to determine that the wrist flexion is a high risk motion, which may caus the ergonomic risk assessment model to provide the risk assessment to the AAR engine 26. The AAR engine 26 may process the risk assessment and may generate an alert and / or recommendation that can be provided to the user in real-time. For example, the AAR engine 26 may generate an alert indicating that it is unsafe to have a wrist flexion of more than 30 degrees due to injury risk. The AAR engine 26 may also generate a recommendation that includes a recommended wrist flexion angle for the given task, or that includes a different recommended way of completing the task that does not require a wrist flexion angle of more than 30 degrees.
[0184] In some embodiments, the AAR engine 26 may generate one or more alerts and / or recommendations based in part on the 3-D coordinates representing the estimated positions of one or more objects in the environment. For example, the AAR engine 26 may generate an alert if an unsafe movement is detected that involves an object in the environment, may generate a recommendation indicating a corrected action to take involving how the user should handle the object, and / or any other alert or recommendation involving an object in the environment.
[0185] In some embodiments, because the first motion prediction platform 24 predicts the future motion path of the users and / or objects for an upcoming time period FTP, the AAR 26 may generate an alert and / or recommendation based on a predicted action that has yet to occur. For example, assume the first motion prediction platform 24 predicts a future motion path of a user and that the motion assessment model 28 (or the anomaly detection model 38) generates a motion assessment (or detects an anomaly) that includes a risky or dangerous motion and / or an unsafe predicted action involving the user and / or an object in the environment. In this case, the AAR engine 26 may generate an alert and / or a recommendation that can be provided to the user before the risky or dangerous motion or unsafe predicted action occurs.
[0186] The AAR engine 26 may generate any number of different types of alerts and / or recommendations depending on the motion, the environment, the user, and / or the task being completed. For example, the AAR engine 26 may generate an alert and / or a recommendation for a repetitive user motion to warn the user about repetitive strain injury (RSI) risks, for incorrect posture, for lifting form adjustments, for optimal load positioning, for task redesign suggestions, for recommended rest periods, for form suggestions during task completion, for a risky or unsafe movement, and / or the like. Further, one or more of these alerts and / orrecommendations may be generated and included in one or more of the reports described in connection with Fig.2J.
[0187] As shown by reference number 78, the alert and / or recommendation may be provided to another device, such as a site manager device 80, a user device 82, and / or another device. For example, a site manager may be managing users A and B, and the alert and / or recommendation may be provided to the site manager device 80. The alert and / or recommendation may be provided for display on a user interface of the site manager device 80. By providing real-time, the site manager is notified of the potentially dangerous or high injury risk situation and has an opportunity to engage with the users to discuss the situation.
[0188] As another example, the alert and / or recommendation may be provided to the user device 82, which may be a mobile device of user A or user B. The alert and / or recommendation may be provided for display on a user interface of the user device 82. By providing the user with real-time alerts and / or recommendations, the user is quickly made aware of the dangerous or high injury risk situation and can take corrective actions before injury occurs.
[0189] In some embodiments, the user device 82 may be a wearable device. For example, assume the wearable device is a smart belt. In this case, the AAR engine 26 may generate an alert which may cause the user’s smart belt to vibrate (e.g., if poor posture is detected). As another example, the wearable device may have a user interface and the alert and / or recommendation may be displayed on the user interface. The wearable device may also include data logging for identifying long term trends and for tracking motion patterns over time for risk analysis. To provide a specific example, the data logging may indicate that posture deviation increased by 15% this week and that the user should monitor this closely.
[0190] In this way, the motion prediction platform 14 (e.g., using AAR 26) provides real-time alerts and / or recommendations to user based on the motion estimations and motion predictions. This allows the user to benefit from real time ergonomic feedback, improving user safety and / or efficiency while completing one or more tasks.
[0191] Fig.2J is a diagram showing non real-time outputs generated by the AAR engine 26 of the motion prediction platform 14. As shown by reference number 84, the motion prediction platform 14 (e.g., using AAR engine 26) may generate a report assessing the motion of the users. The report may include assessment information and / or one or more recommendations.
[0192] In some embodiments, the AAR engine 26 may generate a workplace risk assessment report that summarizes high-risk activities for each worker. For example, the workplace risk assessment report may indicate that 30% of lifting tasks in a warehouse exceed safe back stress limits.
[0193] In some embodiments, the AAR engine 26 may generate a personalized ergonomic profile that provides tailored insights to the user or worker. For example, the personalized ergonomic profile may indicate that the user exceeded safe lifting limits five times today an that additional training is recommended
[0194] In some embodiments, the AAR engine 26 may generate a report that includes task modification recommendations. For example, the report may recommend lowering a height of the conveyor belt to align with ergonomic standards.
[0195] In some embodiments, the AAR engine 26 may generate a report that includes a video of a 3-D model of the user performing a recommended motion or task. For example, the motion prediction platform 14 may generate a video with a 3-D model of the user performing a corrective motion. This allows the user to view the corrective motion in a format that is easy to understand and that is easier for the user to duplicate.
[0196] In some embodiments, the AAR engine 26 may generate a report that includes one or more alerts and / or recommendations described in connection with Fig. 2I. That is to say, one or more example alerts and / or recommendations that are described as being performed in real- time may also be performed in a non-real-time manner. Similarly, one or more example reports described as being performed in a non-real-time manner may be performed in real-time.
[0197] As shown by reference number 86, the motion prediction platform 14 may provide the report to the site manager device 80, to the user device 82, and / or to another device. This allows a site manager and / or a user to view the report via a user interface display.
[0198] In this way, the motion prediction platform 14 provides reports detailing the ergonomic risks and / or ergonomic performance of the users in the environment.
[0199] Figs.8A and 8B show an example process for training the custom motion model 30 and Figs.9A-9D show an example process for using the trained custom motion model 30 as part of a process for performing an ergonomic risk assessment and for using the trained custom motion model 30 to determine a text-based ergonomic recommendation.
[0200] Overall, the embodiments described herein build on an existing natural language-based motion language model (MLM) that can act as a virtual assistant to guide workers, provide advisory feedback, and reduce injuries in factory floor activities. As would be understood by one skilled in the art, the same functionality applied to in the manufacturing space can be used to provide similar services to users in a different environment, different workplace, different career field, etc.
[0201] One or more embodiments described below describe features carried out by the motion prediction platform 14. In practice, any feature carried out by the motion prediction platform14 can also be carried out by a different device, such as developer device 124 and / or user device 158. Moreover, any feature carried out by developer device 124 or the user device 158 can be carried out by another device, such as the motion prediction platform 14.
[0202] Fig.8A is a diagram showing a first half of the process for training the custom motion model 30. As shown by reference number 122, historical motion capture (MoCap) data may be made accessible to a developer who is using developer device 124.
[0203] The historical MoCap data may include motion capture data collected from a diverse range of manufacturing tasks, including material handling, assembly tasks, maintenance tasks, environment interaction tasks, and / or the like. The materials handling tasks may involve lifting, carrying, pushing, loading, and / or unloading, and may include lifting tasks, carrying tasks, pushing / pulling tasks, loading / unloading tasks, and / or the like. The assembly tasks may involve precise hand movements, repetitive motions, tool use, and coordination, and may include manual assembly tasks, tool-based assembly tasks, precision assembly tasks, robot- assisted assembly tasks, and / or the like. The maintenance tasks may involve repairing, adjusting, inspecting, and / or servicing equipment, and may include mechanical maintenance tasks, electrical maintenance tasks, facility maintenance tasks, inspection and safety checks tasks, and / or the like. The environment interaction tasks may involve navigating, adjusting, and interacting with the workplace environment, and may include navigating workspaces, adjusting workspaces, using workstation tools and fixtures, and / or the like. The historical MoCap data may include hundreds, thousands, tens of thousands, or more, motion sequences, including long motion sequences and short motion sequences. In a preferred embodiment, the MoCap data may include high-fidelity manufacturing centric motion capture data.
[0204] As shown by reference number 126, the developer device 124 may generate a parametric 3-D model. For example, the developer device 124 may convert the raw MoCap data to a skinned multi-person linear model (SMPL) format. This mesh-based 3-D model allows for standardized human representation and consistent motion representation across different datasets. At this point, the parametric 3-D model includes positions of key points but does not include kinematic and / or ergonomic parameters. Also, using the SMPL format requires key point conversion from a current number of key points (e.g., 15 key points) to a different number of key points required by custom motion model 30 (e.g., 24 key points). A description of key point conversion is provided in connection with Fig.9A.
[0205] As shown by reference number 128, the developer device 124 may determine kinematic parameters for key points of the parametric 3-D model. For example, the developer device 124 may use a simulation application (e.g., OpenSim, etc.) to determine flexion or extension anglesof key points and moments or forces at the key points. The simulation application may determine these values using inverse kinematics and inverse dynamics. For example, inverse kinematics may be used to estimate joint angles for key points. The simulation application (e.g., OpenSim, etc.) adjusts the kinematic skeleton to match MoCap positions while enforcing biomechanical constraints.
[0206] In some embodiments, prior to determining the kinematic parameters, the developer device 124 may perform one or more preprocessing operations to normalize noisy or missing MoCap data using one or more filtering techniques, such as EKF 20, a median filtering technique, and / or the like. Linkages such as bone segment lengths may be standardized based on anthropometric scaling or via another standardization technique.
[0207] In some embodiments, the developer device 124 may estimate joint angles using pose constraints. For example, the developer device 124 may use linkage (segment) vectors between key points to define limb orientations. To provide a specific example, the developer device 124 may determine joint angles using XYZ Euler decomposition, rotation matrices, and / or another technique known in the art.
[0208] In some embodiments, the developer device 124 may determine the flexion / extension angles for key points. For example, the developer device 124 may determine key point angles using arctan-based inverse kinematics equations. Some key points (e.g., shoulder, hip, etc.) have multiple DoF (e.g., abduction, rotation, etc.). Optimization techniques such as least squares minimization may be implemented to ensure the best-fit angles.
[0209] In some embodiments, the developer device 124 may determine joint angle trajectories over time while defining flexion / extension angles for each key point. For example, knee flexion over time may be represented as 2^^^^(t) = [15∘, 30∘, 45∘, 60∘]. As another example, shoulder extension over time may be represented as 2^^^ !"^#(t) = [20∘, 10∘, −5∘].
[0210] In some embodiments, the developer device 124 may estimate moments and key point forces using inverse dynamics. Inverse dynamics calculates forces and moments required to produce the observed key point motions. The key point angles determined using inverse kinematics may be provided as an input to an inverse dynamics’ equation or function. Next, the inverse dynamics equation or function may determine linkage mass and inertia properties. For example, an anthropometric model may estimate mass and center of mass for each linkage. To provide a specific example, a thigh mass may be defined as being 10% of a body weight of the user. Next, the developer device 124 may apply Newton-Euler equations and may account for one or more external forces. If external forces (e.g., an object lifting) are applied, these are factored into the dynamics. The output from inverse dynamics may be key point (e.g., joint)reaction forces (e.g., knee force = 500 N), key point (e.g., joint) moments (Nm) (e.g., wrist moment = 20 Nm), key point (e.g., muscle) force estimation (e.g., quadriceps force = 800 N, and / or the like.
[0211] In conclusion, the developer device 124 may determine kinematic parameters of key points such as joint angle trajectories (e.g., flexion / extension angle data) and joint moments and forces (e.g., represented using torque).
[0212] Fig. 8B is a diagram showing the second half of the process for training the custom motion model 30. As shown by reference number 130, the developer device 124 may use an annotation tool box to annotate motions. For example, the user may interact with a user interface of an application with a custom motion annotation toolbox to annotate motions with textual descriptions. The interface may have a drop-down menu and the user may select and annotate motions using pre-configured ergonomic data, task descriptions data (e.g., material handling, assembly tasks, machine operation, etc.), direction information (e.g., movement type, key point placement, interaction with equipment, etc.), environment details (e.g., workspace type, surface conditions, temperature / humidity, etc.), productivity metrics (e.g., time to complete motion, efficiency scores, error rate, fatigue index, etc.), and / or the like. The annotated data is pre-configured. For example, human experts have created annotated motion data and, as will be explained below, these annotations serve as the “ground truth” textualized description of the motion.
[0213] As shown by reference number 132, the developer device 124 may provide, to a third party generative AI server 134, the annotated motion selections and kinematic parameters. For example, the generative AI server 134 may support an AI chatbot that uses generative AI to answer questions. In some embodiments, the user may provide (e.g., upload) the annotated motion selections and kinematic parameters and may input text into a questions box to request that the AI chatbot provide a textual description of motions of respective key points.
[0214] As shown by reference number 136, the generative AI server 134 may generate the text description of motions of a user. As shown by reference number 138, the generative AI server 134 may provide, to the developer device 124, the text descriptions of the motions. In some embodiments, the text description generated by the generative AI server 134 may be compared against the ground truth (expert-annotated motion descriptions).
[0215] If the result is not accurate (e.g., fails to satisfy a configured threshold accuracy level), then iterative training may take place to improve the overall accuracy of the textual descriptions of the motions. If the result is accurate (e.g., satisfies the configured threshold accuracy level), the, as is shown by reference number 140, the developer device 124 may train the custommotion model 30 using image-text pairs. An image-text pair may include a structured motion representation and a textualized description of the motion. The structured motion representation may be a parametric 3-D model representation, a motion capture frame or skeleton sequence, kinematic data including 3-D key point angles, and / or the like.
[0216] To provide a specific example, the generative AI server 134 may generate the following: “The task involves backward movement, bending, reaching, and lifting tools with both one-handed and two-handed lifting techniques, supported by appropriate workstation design and carrying aids. there is no repetition, and proper containers and work techniques are utilized to pick items up from the ground efficiently. the flexion / extension plot shows values ranging outside the safe threshold with both positive and negative readings. in lower back moment, high peaks indicate significant bending forces exceeding the recommended 200 nm value for heavy lifting. the left wrist flexion exceeds 30 degrees, putting it in the risk zone. the right wrist flexion also exceeds 40 degrees multiple times, clearly indicating a risk for injury”.
[0217] Comparatively, the ground truth description may state the following: “The task involves back-to-front movements such as walking, lifting, and bending with medium to heavy boxes, ensuring there is no repetition and proper workstation design and techniques are used. The person maintains a normal productivity level while moving from the left to right and vice versa. The flexion / extension degree plot shows significant fluctuations with values exceeding the safe threshold, indicating a risk. The L4_L5 lower back moment plot displays extreme peaks well beyond the 200 Nm limit, highlighting a potential for harm. Both left and right wrist flexion plots show a trend of oscillating movement, with the left wrist staying within the warning threshold while the right wrist flexion plot stays mostly within the warning threshold”. In this example, the model was fairly precise in accurately describing the motion, meaning the description may be used as part of an image-text pair to train the custom motion model 30.
[0218] As can be seen from the example above, the text labels above contain long descriptions (e.g., more than one sentence). Accordingly, in some embodiments, these motions may be relabeled to contain shorter, user-friendly labels created using ergonomic rules. For example, an ergonomic description of “lower back flexion or extension is within the safe range (-5 to 60 degrees) may be assigned a technical description of “normal posture” and a human-friendly response of “normal posture.” To provide another example, an ergonomic description may state that “excessive flexion of the lower back beyond 60 degreesbut up to 100 degrees, requiring caution”. This may be assigned a technical description of “overbending” and a human-friendly response of “bending too much.”
[0219] In some embodiments, the custom motion model 30 may be trained in three stages. The first stage is motion encoding and discretization. The motion and text pairs are provided as input. A motion encoder may convert raw motion data to a compact representation. This can be done using vector quantized variational autoencoder (VQ-VAE). A transformer attention module may map the encoded motion to a discrete space using cosine similarity. For example, the motion of lifting a box may be assigned a motion token index that represents the movement. A motion decoder may reconstruct sequences from the motion tokens. The output is a sequence of motion tokens representing the input motion. This ensures that motion can be discretized efficiently for downstream processing.
[0220] The second stage is motion-language pre-training. Discrete motion tokens from the first stage and text tokens (from the annotated descriptions) are provided as input. Motion tokens are mixed with words as discrete indices to create a shared embedding space between motion and text. A transformer based language model may be used. An encoder maps mixed motion-text inputs into a common representation. A decoder generates contextual outputs (e.g., text or motion predictions). The output of the second stage is a pre-trained model that understands the relationship between motion sequences and textual descriptions.
[0221] The third stage involves instruction tuning for motion tasks. The pre-trained model from the second stage is provided as an input. The model is tuned for three primary tasks: (1) motion generation to generate 3-D motion sequences from textual input, (2) motion captioning describing a given motion sequence in natural language, and (3) motion prediction to predict future motion based on past movement patterns. The output from the third stage is a tuned model that can generate, describe, and predict human motions in a given environment.
[0222] While one or more embodiments described herein refer to the custom motion model 30 as being trained using motion (video)-text pairs, it is to be understood that this is provided by way of example. In practice, the custom motion model 30 may be trained using a different pair, such as a text-motion pair or a motion-motion pair. For example, the custom motion model 30 may be trained in a “text-to-motion” manner to generate 3-D motions based on a textual description provided by a user. This can be useful for generating a 3-D skeletal model of a user and for providing the user with video data illustrating how a recommended motion is performed.
[0223] By training the custom motion model 30 using the above-identified pairs, multi-modal learning occurs by connecting numerical motion data with human language. This reduces the amount of generalization that occurs and allows the model to predict unseen movements. Further, traditional motion models output numerical values which are not intuitive for non- expert users. By using the above-identified pairs, the model learns how to describe motion in natural language, making it easier for users to interpret the results.
[0224] Fig.9A is a diagram showing the processing steps performed by the conversion module 34 of the motion prediction platform 14. As shown by reference number 142, the motion prediction platform 14 may convert the 3-D coordinates to a format suitable for motion model 30.
[0225] As a preliminary matter, the motion prediction platform 14 is first provided with the 3- D coordinates representing the estimated positions of the key points of the users. In some embodiments, this can be the output provided by the pose estimation module 16 or the output provided by the EKF 20. Alternatively, this data may be MoCap data captured based on one or more users moving around in an environment.
[0226] Next, the motion prediction platform 14 may perform one or more normalization and / or pre-processing operations. For example, a filtering technique (e.g., filtering from EKF 20, median filtering, etc.) may reduce noise and / or outliers in the data. As another example, a normalization technique may transform data to a uniform scale. As another example, key point alignment may be performed to fit the dataset to the format suitable for the custom motion model 30.
[0227] To convert the 3-D coordinates representing estimated positions of key points to the format compatible with the custom motion model 30, the motion prediction platform 14 may map an n-key point human dataset (where n is between 15 and 21) to a 22 key point SMPL model. This conversion uses a feed-forward neural network to transform the n-key point dataset to a 22 key point dataset. In some embodiments, the feed-forward neural network may include one or more loss functions for improving the accuracy of the network. For example, the feed-forward neural network may be configured with a positional loss function, a distance consistency loss function, an angle consistency loss function, and / or a combined loss function. The positional loss function ensures predicted key points are close to a ground truth. The distance consistency loss function preserves pairwise distances between specified key points (e.g., key point pair 8, 11, key point pair 9,14, etc.). The angle consistency loss function maintains angular relationships between vectors formed by key points and axes (e.g., key pointpair 9, 14, key point pair 6, 3, etc.). The combined loss function balances all three losses with weights. Example loss function equations are provided below. ^!^^^ = ^m ∑m&o^ || ^'# ^,& - ^Y#^",& ||² (13)^
[0228] , Euclideandistance between the predicted and true key points. N is number of key points. ^'# ^,&refers to the ground truth 3-D coordinate of the i-th key point. ^Y#^",& refers to the predicted3-D coordinate of the i-th key point. || ^'# ^,& - ^Y#^",& ||² refers to squaring the Euclideandistance between the ground truth and predicted key point positions. This ensures that the predicted key points are as close as possible to the ground truth key points.
[0229] In Equation (14), ^!^^^refers to consistency loss which ensure that the relative distances between specified key points are preserved. M refers to the number of selected key point pairs. The variables i, j are indices of key points in the dataset. The term “pair” denotes the set ofkey point pairs used for distance consistency (e.g., wrist-elbow, shoulder-hip, etc.). ^'# ^,& -^Y#^",q refers to the distance between two true key points. ^Y#^",& - ^Y#^",q refers to the distancebetween two predicted key points. The subtraction ensures that the relative spatial structure of the body is maintained during the key point transformation.
[0230] In Equation (15), ^V^r!^ refers to angular loss which ensures that the anglesbetween certain key points remain consistent. K refers to the number of key point pairsused for angle consistency. The variables i, j represent indices of key points used for anglemeasurement. The term 2'# ^, represents the ground truth angle formed between keypoints. The term 2Y#^", represents the predicted angle formed between key points. Theterm cos (2) represents the cosine of the angle formed between key points, ensuring rotational consistency. The absolute difference between the cosine values ensures that predicted angles remain close to the true angles.
[0231] In Equation (16), the total loss is the final loss function used to train the feed-forward neural network. The positional loss ensures that individual key points are close to the ground truth locations The distance loss ensures that relative distances between key points remain consistent. The angle loss ensures that key point angles are preserved. The term u"&^'is a weight factor for the distance consistency loss. The term uV^r!^is a weight factor for the angleconsistency loss. The weighted combination of all three losses ensures balance between position, distance consistency, and angle consistency. The loss function framework helps to convert an n-key point dataset, where n is between 15 and 21, into a 22-key point SMPL model while ensuring spatial and angular consistency. It is noted that the one or more loss functions may be used in connection with any of the embodiments described herein, including those described in connection with Figs.2A-2J.
[0232] Reference number 144 shows a parameterized 3-D skeletal model with 15 key points. Reference number 146 shows a parameterized 3-D skeletal model with 24 key points, which is the output of the key point conversion performed by the motion prediction platform 14.
[0233] Fig.9B is a diagram showing the processing steps performed by biomechanical analysis model 36 of the motion prediction platform 14. As shown by reference number 148, the motion prediction platform 14 (e.g., using the biomechanical analysis model 36) may determine linkage vectors and link length orientations.
[0234] First, and as shown in Fig. 9B, the biomechanical analysis model 36 may be provided with the converted key point data. Next, the biomechanical analysis model 36 may determine linkage vectors based on the vector difference between two connected key points. For example,a linkage vector for a bone may be determined using the following equation:^$^^^ = (^q^&^'3 , ^q^&^'3, ^q^&^'3) – c (17)
[0235] In Equation (17), the expression (^q^&^'3, ^q^&^'3, ^q^&^'3) represents 3-D coordinates fora first key point (e.g., a first joint) and the expression (^q^&^'3 , ^q^&^'3, ^q^&^'3) represents 3-Dcoordinates for a second adjoining key point (e.g., a second joint). Subtracting these from each other provides the linkage vector for the bone.
[0236] In some embodiments, the biomechanical analysis model 36 may determine link length orientations for the links. For example, the biomechanical analysis model 36 may determine link lengths using Equation (5). Next, the biomechanical analysis model 36 may determine link orientation using one or more rotation matrices. The orientation of each linkage vector is determined with respect to a fixed coordinate frame.
[0237] As shown by reference number 150, the motion prediction platform 14 (e.g., using the biomechanical analysis model 36) may determine key point angular data. For example, the biomechanical analysis model 36 may determine key point angular data using inverse kinematics. Key point angular data may include data indicative of flexion, extension, rotation, and / or the like. The key point angular data, as well as the linkage vectors and link lengthorientation, may be referred to collectively as biomechanical parameters. To relate this to the description of Figs.2A-2J, these are also types of kinematic parameters.
[0238] In some embodiments, angles describing rotational movement in each axis (x, y, and z) may be determined using ZXY Euler decomposition. The key position angular data may be determined over multiple frames to identify trends in movement. This creates a time series dataset of key point motion, which is later used as part of an ergonomic risk assessment.
[0239] Fig.9C is a diagram showing the processing steps performed by the motion assessment model 28 and the custom motion model 30 of the motion prediction platform 14. As shown by reference number 152, the motion prediction platform 14 may perform an ergonomic risk assessment. In some embodiments, the motion assessment model may be an ergonomic risk assessment model which evaluates the physical strain of a user’s movements by analyzing biomechanical parameters (e.g., such as those determined in Fig. 9B) and / or one or more external force inputs (e.g., ground reaction forces and load).
[0240] To perform the ergonomic risk assessment, the ergonomic risk assessment model may first compare key point angles to corresponding threshold key point angle values. In some embodiments, each key point may have a safe range, a warning range, and a stop range. For example, for wrist flexion risk levels, the safe zone may be between 0 and 30 degrees, the warning zone may be between 30 and 45 degrees, and the stop zone may be more than 45 degrees. If a key point angle exceeds a corresponding threshold angle value, the ergonomic risk assessment data may be provided to AAR engine 26 so that the AAR engine 26 can generate and provide an alert to the user.
[0241] Next, the ergonomic risk assessment model may determine key point moments. The moment (torque) on each key point may be estimated using inverse dynamics. For example, the moment for a key point may be estimated using Equation (12) as discussed in Fig.2H.
[0242] Next, the ergonomic risk assessment model may determine a risk score based on a number of times a key point angle exceeded a corresponding threshold value, based on a duration of time spent in high-risk postures, based on a frequency of a particular repetitive motion, and / or based on any other considerations such as those discussed in connection with Fig. 2H. Notably, while the remaining steps relate to the use of the custom motion model 30 (which generates a textual assessment), it is to be understood that the ergonomic risk assessment that has been determined may be used to provide any number of different outputs, such as the outputs shown and described in connection with Figs.2I-2J.
[0243] As shown by reference number 154, the motion prediction platform 14 may determine text-based ergonomic recommendations. For example, the motion assessment, which may bean ergonomic risk assessment, may be provided to the custom motion model 30. This may cause the custom motion model 30 to determine text-based ergonomic recommendations. Examples are provided in the description of Fig.9D.
[0244] Fig. 9D is a diagram showing the text-based ergonomic recommendation being provided to a user device. As shown by reference number 156, the motion prediction platform 14 may provide the text-based ergonomic recommendation to the user device 158. As shown by reference number 160, the user device 158 may display the text-based ergonomic recommendation via a user interface. Reference number 162 provides an example of a text- based ergonomic recommendation.
[0245] Furthermore, one or more of the outputs described in connection with Figs.2H-2J may be implemented in the process shown in Figs. 9A-9D. For example, there may be both real- time outputs and non-real-time outputs. One example of this is real-time outputs while a user is working and performs a dangerous movement. Real-time alerts are available, while the motion prediction platform 14 may also provide non-real-time outputs such as by analyzing motion for the user’s entire work shift.
[0246] While one or more embodiments described herein refer to the custom motion model 30 as outputting text using video (motion)-based inputs, it is to be understood that this is provided by way of example. In practice, the custom motion model 30 may output a video (motion) using text-based inputs or may output a video (motion) using video-based inputs.
[0247] Fig.10 is a diagram of an example environment 164 in which systems and / or methods described herein may be implemented. As shown in Fig. 10, environment 164 may include a sensor device 12, motion prediction platform 14 supported within a cloud computing environment 166, a user device 168, and / or a network 170. Devices of environment 164 may interconnect via wired connections, wireless connections, or a combination of wired and wireless connections.
[0248] Sensor device 12 includes one or more devices capable of receiving, capturing, sensing, processing, and / or providing sensor data. For example, a sensor device 12 may include a camera (e.g., a video camera, etc.), a wearable device, an internet of things (IoT) device, an inertia measurement unit (IMU) sensor, and / or another type of sensor known in the art. Sensor data may include image data, video data, audio data, measurement data, IoT data, and / or the like.
[0249] In some embodiments, multiple sensor devices 12 may be configured in the same environment to capture sensor data from different positions. For example, multiple different sensor devices 12 may capture image / video data and may transmit the image / video data to themotion prediction platform 14. Data transmissions may occur at predetermined intervals, based on another configurable trigger condition being satisfied, and / or based on a request provided by a user.
[0250] In some embodiments, an environment may be configured using only image / video capturing sensor devices 12. In this embodiment, one or more components of the motion prediction platform 14 (e.g., 3-D kinematic model 18) may be used to determine kinematic parameters. As such, the user does not need to equip a wearable device such as an IMU.
[0251] In other embodiments, the user may be equipped with a wearable device such as an IMU. In this embodiment, the IMU may be adapted to capture acceleration data and / or velocity data (e.g., angular velocity via a gyroscope).
[0252] Motion prediction platform 14 includes one or more devices capable of receiving, storing, processing, and / or providing information associated with motion of one or more users. For example, motion prediction platform 14 may include a server device (e.g., a host server, a web server, an application server, etc.), a data center device, or a similar device. In some embodiments, the motion prediction platform 14 (e.g., using the pose estimation module 16, the 3-D kinematic model 18, the motion prediction model 24, etc.) may generate a 3-D skeletal model of each respective user based on the 3-D set of coordinates representing the estimated positions of key points the user.
[0253] In some embodiments, the motion prediction platform 14 may have access to a database and / or data structure used to sort, organize, and filter one or more types of data described herein. In some embodiments, the database and / or data structure may be local to the motion prediction platform 14. In some embodiments, the database and / or data structure may be a third-party storage provider.
[0254] In some embodiments, the motion prediction platform 14 may train a data model, such as one of the data models described herein, using machine learning. The data model may be used to make predictions, classifications, and / or recommendations in accordance with the principles of the present disclosure. In some embodiments, the data model may be trained by an external device or server and the trained data model may be provided to or made accessible to the motion prediction platform 14.
[0255] In some embodiments, as shown, the motion prediction platform 14 may be hosted in the cloud computing environment 166. Notably, while embodiments described herein describe the motion prediction platform 14 as being hosted in the cloud computing environment 166, in some embodiments, the motion prediction platform 14 may not be cloud-based (i.e., may be implemented outside of a cloud computing environment) or may be partially cloud-based.
[0256] Cloud computing environment 166 includes an environment that hosts motion prediction platform 14. Cloud computing environment 166 may provide computation, software, data access, storage, etc. services that do not require end-user knowledge of a physical location and configuration of system(s) and / or device(s) that hosts the motion prediction platform 14. As shown, the cloud computing environment 166 may include a group of computing resources 172 (referred to collectively as “computing resources 172” and individually as “computing resource 172”).
[0257] Computing resource 172 includes one or more personal computers, workstation computers, server devices, or another type of computation and / or communication device. In some embodiments, the computing resource 172 may host the motion prediction platform 14. The cloud resources may include compute instances executing in the computing resource 172, storage devices provided in the computing resource 172, data transfer devices provided by the computing resource 172, and / or the like. In some embodiments, the computing resource 172 may communicate with other computing resources 172 via wired connections, wireless connections, or a combination of wired and wireless connections.
[0258] As further shown in Fig. 10, computing resource 172 may include a group of cloud resources, such as one or more applications (“APPs”) 172-1, one or more virtual machines (“VMs”) 172-2, virtualized storage (“VSs”) 172-3, one or more hypervisors (“HYPs”) 172-4, and / or the like.
[0259] Application 172-1 may include one or more software applications that may be provided to or accessed by user device 168. Application 172-1 may eliminate a need to install and execute the software applications on these devices. In some embodiments, one application 172-1 may send / receive information to / from one or more other applications 172-1, via virtual machine 172-2. In some embodiments, application 172-1 may be a motion prediction and analysis application. In some embodiments, the motion prediction and analysis application may include one or more user interfaces that are accessible by users.
[0260] Virtual machine 172-2 may include a software implementation of a machine (e.g., a computer) that executes programs like a physical machine. Virtual machine 172-2 may be either a system virtual machine or a process virtual machine, depending upon use and degree of correspondence to any real machine by virtual machine 172-2. A system virtual machine may provide a complete system platform that supports execution of a complete operating system (“OS”). A process virtual machine may execute a single program and may support a single process. In some embodiments, virtual machine 172-2 may execute on behalf of anotherdevice, and may manage infrastructure of the cloud computing environment 166, such as data management, synchronization, or long-duration data transfers.
[0261] Virtualized storage 172-3 may include one or more storage systems and / or one or more devices that use virtualization techniques within the storage systems or devices of the computing resource 172. In some embodiments, within the context of a storage system, types of virtualizations may include block virtualization and file virtualization. Block virtualization may refer to abstraction (or separation) of logical storage from physical storage so that the storage system may be accessed without regard to physical storage or heterogeneous structure. The separation may permit administrators of the storage system flexibility in how the administrators manage storage for end users. File virtualization may eliminate dependencies between data accessed at a file level and a location where files are physically stored. This may enable optimization of storage use, server consolidation, and / or performance of non-disruptive file migrations.
[0262] Hypervisor 172-4 may provide hardware virtualization techniques that allow multiple operating systems (e.g., “guest operating systems”) to execute concurrently on a host computer, such as computing resource 172. Hypervisor 172-4 may present a virtual operating platform to the guest operating systems and may manage the execution of the guest operating systems.
[0263] User device 168 includes one or more devices capable of receiving, generating, storing, processing, and / or providing information associated with motion of one or more users. For example, user device 168 may include a device, such as a tablet computer (e.g., an iPad, etc.), a mobile phone (e.g., a smart phone, a radiotelephone, etc.), a laptop computer, a handheld computer, a server computer, a gaming device, a wearable communication device (e.g., a smart wristwatch, a pair of smart eyeglasses, etc.), or a similar type of device. User device 168 may be utilized by a user to view a motion assessment generated by the motion prediction platform 14. User device 168 corresponds to site manager device 80, user device 82, developer device 124, and / or user device 158 as described in connection with other figures.
[0264] Network 170 includes one or more wired and / or wireless networks. For example, network 104 may include a cellular network (e.g., a fifth generation (5G) network, a fourth generation (4G) network, such as a long-term evolution (LTE) network, a third generation (3G) network, a code division multiple access (CDMA) network, a public land mobile network (PLMN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a telephone network (e.g., the Public Switched Telephone Network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber optic-based network, a cloud computing network, or the like, and / or a combination of these or other types of networks.
[0265] The number and arrangement of devices and networks shown in Fig. 10 are provided as an example. In practice, there may be additional devices and / or networks, fewer devices and / or networks, different devices and / or networks, or differently arranged devices and / or networks than those shown in Fig. 10. Furthermore, two or more devices shown in Fig. 10 may be implemented within a single device, or a single device shown in Fig. 10 may be implemented as multiple, distributed devices. Additionally, or alternatively, a set of devices (e.g., one or more devices) of environment 140 may perform one or more functions described as being performed by another set of devices of environment 140.
[0266] Fig. 11 is a diagram of example components of a device 174. Device 174 may correspond to the sensor device 12, the user device 168, and / or the motion prediction platform 14. In some embodiments, the sensor device 12, the user device 168, and / or the motion prediction platform 14 may include one or more devices 156 and / or one or more components of device 174. As shown in Fig. 11, device 174 may include a bus 176, a processor 178, a memory 180, a storage component 182, an input component 184, an output component 186, and / or a communication interface 188.
[0267] Bus 176 includes a component that permits communication among multiple components of device 174. Processor 178 is implemented in hardware, firmware, and / or a combination of hardware and software. Processor 178 includes a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a microprocessor, a microcontroller, a digital signal processor (DSP), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), and / or another type of processing component. In some embodiments, processor 178 includes one or more processors capable of being programmed to perform a function. Memory 180 includes a random-access memory (RAM), a read only memory (ROM), and / or another type of dynamic or static storage device (e.g., a flash memory, a magnetic memory, and / or an optical memory) that stores information and / or instructions for use by processor 178.
[0268] Storage component 182 stores information and / or software related to the operation and use of device 174. For example, storage component 182 may include a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optic disk, and / or a solid-state disk), a compact disc (CD), a digital versatile disc (DVD), a floppy disk, a cartridge, a magnetic tape, and / or another type of non-transitory computer-readable medium, along with a corresponding drive.
[0269] Input component 184 includes a component that permits device 174 to receive information, such as via user input (e.g., a touch screen display, a keyboard, a keypad, a mouse, a button, a switch, and / or a microphone). Additionally, or alternatively, input component 184may include a sensor for sensing information (e.g., a global positioning system (GPS) component, an accelerometer, a gyroscope, and / or an actuator). Output component 186 includes a component that provides output information from device 174 (e.g., a display, a speaker, and / or one or more light-emitting diodes (LEDs)).
[0270] Communication interface 188 includes a transceiver-like component (e.g., a transceiver and / or a separate receiver and transmitter) that enables device 174 to communicate with other devices, such as via a wired connection, a wireless connection, or a combination of wired and wireless connections. Communication interface 188 may permit device 174 to receive information from another device and / or provide information to another device. For example, communication interface 188 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi interface, a cellular network interface, an application programming interface (API), and / or the like.
[0271] Device 174 may perform one or more processes described herein. Device 174 may perform these processes based on processor 178 executing software instructions stored by a non-transitory computer-readable medium, such as memory 180 and / or storage component 182. A computer-readable medium is defined herein as a non-transitory memory device. A memory device includes memory space within a single physical storage device or memory space spread across multiple physical storage devices.
[0272] Software instructions may be read into memory 180 and / or storage component 182 from another computer-readable medium or from another device via communication interface 188. When executed, software instructions stored in memory 180 and / or storage component 182 may cause processor 178 to perform one or more processes described herein. Additionally, or alternatively, hardwired circuitry may be used in place of or in combination with software instructions to perform one or more processes described herein. Thus, embodiments described herein are not limited to any specific combination of hardware circuitry and software.
[0273] The number and arrangement of components shown in Fig. 11 are provided as an example. In practice, device 174 may include additional components, fewer components, different components, or differently arranged components than those shown in Fig. 11. Additionally, or alternatively, a set of components (e.g., one or more components) of device 174 may perform one or more functions described as being performed by another set of components of device 174.
[0274] Fig.12 is a diagram of an example process 190 for predicting and assessing motion of one or more users in an environment. Example process 190 includes receiving image datadepicting the one or more users (block 192). For example, the motion prediction platform 14 may receive image data depicting one or more users the an environment. The image data may be captured by one or more sensors 12.
[0275] Example process 190 further includes determining estimated positions of key points of each respective user for a time ^^(block 194). For example, the motion prediction platform (e.g., using pose estimation module 16) 14 may determine 3-D coordinates representing estimated positions of key points of each user for the time ^^. Example process 190 further includes determining link lengths for linkages connecting the key points (block 196). For example, the motion prediction platform 14 may determine link lengths for linkages connecting the key points.
[0276] Example process 190 further includes determining predicted kinematic parameters for the key points for an upcoming time ^^^^(block 198). For example, the motion prediction platform 14 may determine estimated kinematic parameters for the key points for the upcoming time ^^^^. Example process 190 further includes refining the estimated positions of key points and estimated kinematic parameters at a current time ^^^^(block 200). For example, the motion prediction platform 14 may refine the estimated positions of key points and the estimated kinematic parameters at the current time ^^^^.
[0277] Example process 190 further includes classifying an activity being performed by the users (block 202). For example, the motion prediction platform 14 (e.g., using the activity recognition model 22) may classify an activity being performed by the users.
[0278] Example process 190 further includes predicting a future intent of the users for an upcoming time period (block 204). For example, the motion prediction platform 14 (e.g., using the activity recognition model 22) may predict a future intent of the users for an upcoming time period. The future intent may be predicted by using machine learning to process the estimated positions of key points, the estimated kinematic parameters, and the classified activity being performed by the user.
[0279] Example process 190 further includes predicting a future motion path of the users for the upcoming time period (block 206). by using machine learning to process the user activity and the future intent of the user. For example, the motion prediction platform 14 (e.g., using motion prediction model 24) may predict a future motion path of the users for the upcoming time period.
[0280] Example process 190 further includes generating a motion assessment for the users (block 208). For example, the motion prediction platform 14 (e.g., using motion assessmentmodel 28) may generate a motion assessment for the users that is based on past motions and future motion predictions.
[0281] Example process 190 further includes making the motion assessment available to a device or account associated with the respective users (block 210). For example, the motion prediction platform 14 (e.g., using AAR engine 26) may make the motion assessment available to a device or account associated with the respective users.
[0282] Fig.13 is a diagram of an example process 212 for predicting and assessing motion of one or more users in an environment. Example process 212 includes generating a 3-D skeletal model that includes estimated positions of key points and linkages connecting the key points (block 214). For example, the motion prediction platform 14 may generate a 3-D skeletal model that includes estimated positions of key points and linkages connecting the key points.
[0283] Example process 212 further includes determining biomechanical parameters using machine learning (block 216). For example, the motion prediction platform 14 may determine biomechanical parameters using machine learning.
[0284] Example process 212 further includes performing an ergonomic risk assessment using the 3-D skeletal model and the biomechanical parameters (block 218). For example, the motion prediction platform 14 may perform an ergonomic risk assessment using the 3-D skeletal model and the biomechanical parameters.
[0285] Example process 212 further includes providing the ergonomic risk assessment for display on a user interface (block 220). For example, the motion prediction platform 14 may provide the ergonomic risk assessment for display on a user interface.
[0286] The foregoing disclosure provides illustration and description but is not intended to be exhaustive or to limit the embodiments to the precise form disclosed. Modifications and variations may be made in light of the above disclosure or may be acquired from practice of the embodiments.
[0287] Some embodiments are described herein in connection with thresholds. As used herein, satisfying a threshold may refer to a value being greater than the threshold, more than the threshold, higher than the threshold, greater than or equal to the threshold, less than the threshold, fewer than the threshold, lower than the threshold, less than or equal to the threshold, equal to the threshold, etc., depending on the context.
[0288] Certain user interfaces have been described herein and / or shown in the figures. A user interface may include a graphical user interface, a non-graphical user interface, a text-based user interface, etc. A user interface may provide information for display. In some embodiments, a user may interact with the information, such as by providing input via an inputcomponent of a device that provides the user interface for display. In some embodiments, a user interface may be configurable by a device and / or a user (e.g., a user may change the size of the user interface, information provided via the user interface, a position of information provided via the user interface, etc.). Additionally, or alternatively, a user interface may be pre-configured to a standard configuration, a specific configuration based on a type of device on which the user interface is displayed, and / or a set of configurations based on capabilities and / or specifications associated with a device on which the user interface is displayed.
[0289] It will be apparent that systems and / or methods, described herein, may be implemented in different forms of hardware, firmware, and / or a combination of hardware and software. The actual specialized control hardware or software code used to implement these systems and / or methods is not limiting of the embodiments. Thus, the operation and behavior of the systems and / or methods were described herein without reference to specific software code - it being understood that software and hardware can be used to implement the systems and / or methods based on the description herein.
[0290] No element, act, or instruction used herein should be construed as critical or essential unless explicitly described as such. Also, as used herein, the articles “a” and “an” are intended to include one or more items and may be used interchangeably with “one or more.” Furthermore, as used herein, the term “set” is intended to include one or more items (e.g., related items, unrelated items, a combination of related and unrelated items, etc.), and may be used interchangeably with “one or more.” Where only one item is intended, the phrase “only one” or similar language is used. Also, as used herein, the terms “has,” “have,” “having,” or the like are intended to be open-ended terms. Further, the phrase “based on” is intended to mean “based, at least in part, on” unless explicitly stated otherwise.
[0291] While all the invention has been illustrated by a description of various embodiments, and while these embodiments have been described in considerable detail, it is not the intention of the Applicant to restrict or in any way limit the scope of the appended claims to such detail. Additional advantages and modifications will readily appear to those skilled in the art. The invention in its broader aspects is therefore not limited to the specific details, representative apparatus and method, and illustrative examples shown and described. Accordingly, departures may be made from such details without departing from the spirit or scope of the Applicant’s general inventive concept.
Claims
CLAIMS What is claimed is:
1. A method for predicting and assessing motion of a user in an environment, the method comprising: receiving, by a device, image data depicting the user in the environment, the image data being captured by one or more sensor devices; determining, by the device, a first set of three-dimensional (3-D) coordinates representing estimated positions of key points of the user at a time ^^; determining, by the device, a set of link lengths for linkages connecting the key points; determining, by the device and for a time ^^^^, a second set of 3-D coordinates representing predicted positions of the key points and a set of kinematic parameters representing predicted kinematics for the key points, wherein the second set of 3-D coordinates and the set of kinematic parameters are part of a state vector determined by modeling kinematics of the key points using the first set of 3-D coordinates for the time ^^and the link lengths; updating, by the device, the state vector by using an extended Kalman filter (EKF) to compare the second set of 3-D coordinates representing the predicted positions of the key points for the time ^^^^with a third set of 3-D coordinates representing observed positions of the key points at the time ^^^^, wherein the updated state vector is provided back to the 3-D kinematic model for iterative learning, wherein the 3-D kinematic model and the extendedKalman filter are used to determine and update state vectors periodically over a time period^^^;classifying, by the device, an activity being performed by the user by using machine learning to process the state vectors determined over the time period ^^^; predicting, by the device, a future intent of the user for a time period ^^^, wherein the future intent is predicted by using machine learning to process the state vectors and the classified activity being performed by the user; predicting, by the device, a future motion path of the user for the time period ^^^by using machine learning to process the user activity and the future intent of the user; generating, by the device, a motion assessment for the user by processing the predicted future motion path of the user using machine learning; andmaking, by the device, the motion assessment available to a device or account associated with the user.
2. The method of claim 1, wherein the one or more sensors that capture the image data are a plurality of sensors, wherein the image data received depicts the user and a second user, and wherein the method further comprises: determining a unified 3-D coordinate plane by using a calibration and orientation module to process the image data, wherein the unified 3-D coordinate plane is used to determine the first set of 3-D coordinates representing the estimated positions of the key points of the user and another set of 3-D coordinates representing estimated positions of key points of the second user.
3. The method of claim 1, wherein the first set of 3-D coordinates at the time ^^include one or more null or empty set values based on a portion of the user being occluded from view of the one or more sensors, and wherein determining the second set of 3-D coordinates for the time ^^^^comprises: determining one or more 3-D coordinates representing predicted positions of one or more key points occluded from view of the one or more sensors.
4. The method of claim 1, wherein the set of kinematic parameters of the state vector include velocity parameters and acceleration parameters, wherein the state vector further includes angular data identifying angular positions of the key points, and wherein the state vector is used to classify the activity being performed by the user, to predict the future intent of the user, and to predict the future motion of the user.
5. The method of claim 1, wherein the motion assessment is an ergonomic risk assessment, and wherein generating the ergonomic risk assessment comprises: generating the ergonomic risk assessment by using an ergonomic assessment model to process the predicted future motion path of the user, wherein the ergonomic risk assessment includes one or more of: a safety assessment of past motion of the user, and a recommendation providing one or more corrective motions for the user.
6. The method of claim 5, wherein generating the ergonomic risk assessment comprises:generating a video including a 3-D model of the user performing the one or more corrective motions, and including the video in the recommendation.
7. The method of claim 1, further comprising: determining a third set of 3-D coordinates representing estimated positions of one or more objects in the environment of the user; determining a second set of kinematic parameters representing predicted kinematics for the one or more objects, wherein the third set of 3-D coordinates and second set of kinematic parameters are part of a new state vector; predicting a future motion path of the one or more objects by processing the third set of 3-D coordinates and the second set of kinematic parameters using machine learning; and wherein generating the motion assessment comprises: incorporating the predicted future moth path into the motion assessment.
8. The method of claim 1, further comprising: detecting an anomaly relating to motion of the user by processing the updated state vector; and wherein generating the motion assessment comprises: generating, as the motion assessment, an alert indicating that the anomaly has been detected.
9. A device, comprising: one or more memories; and one or more processors, communicatively coupled to the one or more memories, to: receive image data captured by one or more sensor devices, the image data depicting a user in an environment; determine a first set of three-dimensional (3-D) coordinates representing estimated positions of key points of a user at a time ^^; determine a set of link length for linkages connecting the key points; determine, for a time ^^^^, a second set of 3-D coordinates representing predicted positions of the key points and a set of kinematic parameters representingpredicted kinematics for the key points, wherein the second set of 3-D coordinates and the set of kinematic parameters are part of a state vector determined by modeling kinematics of the key points using the first set of 3-D coordinates for the time ^^and the link lengths, and wherein state vectors are determined periodically over a time period ^^^; classify an activity being performed by the user by using machine learning to process the state vectors determined over the time period ^^^; predict a future intent of the user for a final time period defined as ^^^(^^^),^^^(^^^), …, ^^^(^^^), wherein the future intent is predicted by using machinelearning to process the state vectors and the classified activity being performed by the user; predict a future intent of the user for a time period ^^^, wherein the future intent is predicted by using machine learning to process the state vectors and the classified activity being performed by the user; predict a future motion path of the user for the time period ^^^by using machine learning to process the user activity and the future intent of the user; generate a motion assessment for the user by processing the predicted future motion path of the user using machine learning, wherein the motion assessment includes at least one of: an ergonomic risk assessment relating to past motions of the user, a recommendation providing one or more corrective motions, and an alert indicating that an anomaly or unsafe motion has been detected; and make the motion assessment available to a device or account associated with the user.
10. The device of claim 9, wherein the one or more processors are further to: update the state vector by using an extended Kalman filter (EKF) to compare the second set of 3-D coordinates representing the predicted positions of the key points for the time ^^^^with a third set of 3-D coordinates representing observed positions of the key points at the time ^^^^, wherein the updated state vector is provided back to the 3-D kinematic model for iterative learning, wherein the EKF is used to update the state vectors periodically over the time period ^^^.
11. The device of claim 9, wherein the first set of 3-D coordinates representing the estimated positions of the key points of the user at the time ^^include one or more null or empty set values based on a portion of the user being occluded from view of the one or more sensors, and wherein the one or more processors, when determining the second set of 3-D coordinates for the time ^^^^, are to: determine one or more 3-D coordinates representing predicted positions of one or more key points occluded from view of the one or more sensors.
12. The device of claim 9, wherein the one or more processors are further to: detect an anomaly relating to motion of the user by processing the updated state vector; and wherein the one or more processors, when generating the motion assessment, are to: generate the alert based on the detected anomaly.
13. The device of claim 9, wherein the one or more sensors that capture the image data are a plurality of sensors, wherein the image data received depicts the user and a second user, and wherein the one or more processors are further to: determine a unified 3-D coordinate plane that is static relative to a center of the user and relative to a center of the second user, and wherein the unified 3-D coordinate plane is used to determine the first set of 3-D coordinates representing the estimated positions of the key points of the user and another set of 3-D coordinates representing estimated positions of key points of the second user.
14. The device of claim 9, wherein the set of kinematic parameters of the state vector include velocity parameters and acceleration parameters, wherein the state vector further includes angular data identifying angular positions of the key points, and wherein the state vector is used to classify the activity being performed by the user, to predict the future intent of the user, and to predict the future motion of the user.
15. A non-transitory computer-readable medium storing instructions, the instructions comprising:one or more instructions that, when executed by one or more processors, cause the one or more processors to: receive image data captured by one or more sensor devices, the image data depicting a user in an environment; determine a first set of three-dimensional (3-D) coordinates representing estimated positions of key points of a user at a time ^^, wherein the first set of 3-D coordinates include one or more null or empty set values based on a portion of the user being occluded from view of the one or more sensor devices; determine a set of link lengths for linkages connecting the key points; determine, for a time ^^^^, a second set of 3-D coordinates representing predicted positions of the key points and a set of kinematic parameters representing predicted kinematics for the key points, wherein the second set of 3-D coordinates include one or more 3-D coordinates representing predicted positions of one or more key points occluded from the view of the one or more sensors, wherein the second set of 3-D coordinates and the set of kinematic parameters are part of a state vector determined by modeling kinematics of the key points using the first set of 3-D coordinates for the time ^^and the link lengths, and wherein state vectors are determined periodically over a time period ^^^; classify an activity being performed by the user by using machine learning to process the state vectors determined over the time period ^^^; predict a future intent of the user for a time period ^^^, wherein the future intent is predicted by using machine learning to process the state vectors and the classified activity being performed by the user; predict a future motion path of the user for the time period ^^^by using machine learning to process the user activity and the future intent of the user; generate a motion assessment for the user by processing the predicted future motion path of the user using machine learning; and make the motion assessment available to a device or account associated with the user.
16. The non-transitory computer-readable medium of claim 15, wherein the one or more instructions, when executed by the one or more processors, further cause the one or more processors to:update the state vector by using an extended Kalman filter (EKF) to compare the second set of 3-D coordinates representing the predicted positions of the key points for the time ^^^^with a third set of 3-D coordinates representing observed positions of the key points at the time ^^^^, wherein the updated state vector is provided back to the 3-D kinematic model for iterative learning, wherein the EKF is used to update the state vectors periodically over the time period ^^^.
17. The non-transitory computer-readable medium of claim 15, wherein the one or more instructions, that cause the one or more processors to generate the motion assessment, cause the one or more processors to: generate, as the motion assessment, an ergonomic risk assessment by using an ergonomic assessment model to process the predicted future motion path of the user, wherein the ergonomic risk assessment includes one or more of: a safety assessment of past motion of the user, and a recommendation providing one or more corrective motions to improve user safety.
18. The non-transitory computer-readable medium of claim 15, wherein the one or more instructions, when executed by the one or more processors, further cause the one or more processors to: determine a third set of 3-D coordinates representing estimated positions of one or more objects in the environment of the user; determine a third set of kinematic parameters representing predicted kinematics for the one or more objects, wherein the third set of 3-D coordinates and third set of kinematic parameters are part of a second state vector; predict a future motion path of the one or more objects using the second state vector; and wherein the one or more instructions, that cause the one or more processors to generate the motion assessment, cause the one or more processors to: use the predicted future motion path of the one or more objects when generating the motion assessment.
19. The non-transitory computer-readable medium of claim 15, wherein the one or more instructions, when executed by the one or more processors, further cause the one or more processors to: detect an anomaly relating to motion of the user by processing the state vector; and wherein the one or more processors, when generating the motion assessment, are to: generate, as part of the motion assessment, a real-time alert based on the detected anomaly.
20. The non-transitory computer-readable medium of claim 15, wherein the set of kinematic parameters of the state vector include velocity parameters and acceleration parameters, wherein the state vector further includes angular data identifying angular positions of the key points, and wherein the state vector is used to classify the activity being performed by the user, to predict the future intent of the user, and to predict the future motion of the user.