Method for training machine learning model to determine body posture and position
Through the graph attention converter model, the machine learning model is trained under the global coordinate system, and combined with space-time encoding and attention mechanism, the efficiency and accuracy of motion trajectory and posture prediction in the existing technology are solved, and efficient and real-time prediction effects are achieved, suitable for safe navigation and collaborative manufacturing of autonomous vehicles and mobile robots.
Patent Information
- Application Number
- CN202510163701.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-16
- Filing Date
- 2025-02-14
- Publication Date
- 2025-08-19
AI Technical Summary
The prior art is difficult to efficiently and accurately predict human motion trajectories and postures in social robots and autonomous systems, especially in terms of computing resources and data efficiency, which affects the safety and efficiency of robots' interactions with humans.
Using the Graph-Attention converter model, the machine learning model is trained to predict motion trajectories and poses simultaneously by converting the input data into a unified global coordinate system, combining space-time encoding and attention mechanisms, and optimize model performance using a combined loss function.
It realizes efficient and real-time motion trajectory and posture prediction in mobile robots and autonomous vehicles, improves the safety and efficiency of robots' interactions with humans, and is suitable for autonomous navigation and collaborative manufacturing environments.
Smart Images

Figure CN120509446A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to methods for training a machine learning model to determine body pose and position of a body having multiple body parts. Background Art
[0002] The ability to predict human motion and recognize human activities is an essential component of social robotics and autonomous systems, which must operate in human spaces and interact with humans in domestic and industrial environments. For example, in collaborative manufacturing environments, robots accurately track and predict the future postures and activities of human workers to provide safe and efficient support in collaborative assembly tasks. Autonomous vehicles benefit from tracking the presence and attention of pedestrians to safely navigate urban roads. As a result, automated systems, including mobile robots, manipulators, and vehicles, increasingly rely on predicting motion processes for safe navigation, improved human-robot collaboration, personnel tracking, event observation, and simulation purposes.
[0003] Predicting human (or animal) motion involves both trajectory and pose prediction. Trajectory prediction focuses on the coarse prediction of the motion paths of dynamic units. On the other hand, pose prediction (or whole-body pose prediction) involves the fine-grained prediction of the human 3D skeleton joints relative to a fixed reference point on the body (typically the pelvis or torso).
[0004] For both, accurate and efficient methods (in terms of memory requirements, (training) data efficiency, and computational effort) are needed. Summary of the Invention
[0005] According to various embodiments, a method for training a machine learning model to determine body posture and position of a body having multiple body parts is provided, the method having:
[0006] for each training trajectory of a plurality of training trajectories, wherein each training trajectory indicates, for each time point of a corresponding sequence of time points, a position of each body part of a plurality of body parts in a global coordinate system,
[0007] o converting the training trajectory into a transformed training trajectory such that, for a time point of the sequence of time points defined as a prediction start time point, the position of the specified reference body part coincides with the specified point of the global coordinate system, and the direction of movement of the reference body part corresponds to the specified direction in the global coordinate system,
[0008] o predicting, with the aid of a machine learning model, the positions of the body parts at one or more time points after the prediction start time point by feeding the positions of the body parts indicated by the converted training trajectories up to the prediction start time point to the machine learning model, and
[0009] o determining a loss by comparing the predicted position with the positions of the body parts indicated by the training trajectory for one or more time points after the prediction start time point; and
[0010] Adjusting the machine learning model to reduce the total loss including the identified loss.
[0011] The method described above enables joint prediction of trajectory and pose in scenarios that are most important for mobile robots, while balancing the accuracy of trajectory and pose prediction. Furthermore, the method is computationally efficient, which makes it feasible for real-time use, particularly for human-centric motion planning of mobile robots, such as autonomous vehicles and intralogistics robots.
[0012] This prediction can be used as part of a pipeline for (particularly mobile) autonomous systems, but also for stationary manipulators and co-production robots: detection, tracking, prediction, planning, and control. This allows an autonomous system navigating a shared environment to identify and track other dynamic agents, plan its own navigation path, and execute it through a series of control actions. For example, the prediction module provides information for tracking humans and, therefore, for path planning and control actions.
[0013] According to various embodiments, a graph-attention-based transformer model (i.e., a neural network with a transformer architecture) with input data transformation is provided, which can predict motion trajectories and poses in the form of 3D positions of a collection of body parts in a global coordinate system. This enables real-time predictions for robotic applications that require immediate response.
[0014] The method is able to predict the dynamics of body posture and motion trajectories during different walking styles, including running, acceleration, deceleration, and turning – types of motion that are important for socially acceptable robot navigation.
[0015] Various embodiments are described below.
[0016] Example 1 is the method as described above.
[0017] Embodiment 2 is a method according to embodiment 1, wherein the machine learning model has a graph attention network, with the help of which the positions of these body parts indicated by the converted training trajectory until the prediction start time point are processed in the following manner: for each time point until the prediction start time point, a node with the corresponding indicated position as a node feature is assigned to each of these body parts and the nodes assigned to the connected body parts (for example, according to a skeletal model) are connected with the help of edges, the corresponding postures are represented as graphs, and these graphs are processed with the help of the graph attention network.
[0018] In this way, the relationships between the body parts (ie, connections via bones, for example) are effectively taken into account.
[0019] Embodiment 3 is a method according to embodiment 1 or 2, wherein the machine learning model has a converter architecture.
[0020] The transformer is able to effectively consider the temporal sequence required to predict a pose from a sequence of previous poses. Using an attention mechanism, it is also possible to consider the semantic relationships between different poses (such as looking left and right before crossing the street: this look may indicate that the corresponding person then moves forward).
[0021] Embodiment 4 is a method according to any one of embodiments 1 to 3, which comprises: determining the space-time codes of the positions indicated by the converted training trajectory until the prediction start time point, and delivering these space-time codes together with the positions of the body parts indicated by the converted training trajectory until the prediction start time point to the machine learning model.
[0022] The combination of spatial position encoding (which is typically provided by transformers used for processing text, for example) and temporal encoding enables the machine learning model to account for positional variations in the input sequence, which improves predictions.
[0023] Embodiment 5 is a method according to any one of embodiments 1 to 4, wherein the loss includes a first loss component, which includes, for each of one or more time points after the prediction start time point and each of these body parts, the difference between the predicted position of the body part and the position of the body part indicated by the training trajectory or the converted training trajectory as a loss contribution; and / or wherein the one or more time points after the prediction start time point have one or more pairs of consecutive time points, and the loss includes a second loss component, which includes, for each of the one or more pairs and each of these body parts, the difference between the predicted position for the pair of time points and the difference between the position of the body part indicated by the training trajectory or the converted training trajectory for the pair of time points as a loss contribution.
[0024] In particular, a combined loss can be used that includes both a pose loss (first loss component) and a trajectory loss (second loss component), so that the machine learning model predicts good absolute and relative trajectories (ie consistent trajectories).
[0025] Embodiment 6 is a method for predicting one or more body poses and one or more positions of a body having multiple body parts, the method having:
[0026] Training a machine learning model according to any one of embodiments 1 to 5;
[0027] a trajectory of the detected body, the trajectory indicating, for each detection time point of the sequence of detection time points, a position of each body part of the plurality of body parts in the global coordinate system;
[0028] transforming the detected trajectory into a transformed detected trajectory such that, for a last of the detection time points, the position of the specified reference body part coincides with the specified point of the global coordinate system, and the direction of movement of the reference body part corresponds to the specified direction in the global coordinate system; and
[0029] Predicting the positions of the body parts by means of a machine learning model, by feeding the positions of the body parts indicated by the converted detected trajectories to the machine learning model.
[0030] Embodiment 7 is a method according to embodiment 6, further comprising: controlling a robotic device according to the predicted positions of the body parts (ie, the method in this case is a method for controlling a robotic device).
[0031] Embodiment 8 is a data processing system (in particular a control device), which is configured to execute the method according to any one of embodiments 1 to 7.
[0032] Embodiment 9 is a computer program having instructions, which, when executed by a processor, cause the processor to perform the method according to any one of embodiments 1 to 7.
[0033] Embodiment 10 is a computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform the method according to any one of embodiments 1 to 7. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In the accompanying drawings, like reference numerals generally refer to the same parts throughout the different views. These drawings are not necessarily to scale, with emphasis instead generally being placed on presenting the principles of the invention. In the description below, various aspects are described with reference to the following drawings.
[0035] Figure 1 A vehicle is shown.
[0036] Figure 2 The pose prediction and motion trajectory prediction are clarified.
[0037] Figure 3 Input data conversion according to an embodiment is explained.
[0038] Figure 4 A machine learning model for gesture trajectory prediction is shown in accordance with an embodiment.
[0039] Figure 5 A flowchart is shown presenting a method for training a machine learning model to determine body posture and position of a body having multiple body parts in accordance with an embodiment. DETAILED DESCRIPTION
[0040] The detailed description below refers to the accompanying drawings, which, for purposes of illustration, illustrate specific details and aspects of the present disclosure in which the present invention may be implemented. Other aspects may be used and structural, logical, and electrical modifications may be performed without departing from the scope of the present invention. The different aspects of the present disclosure are not necessarily mutually exclusive, as some aspects of the present disclosure may be combined with one or more other aspects of the present disclosure to form new aspects.
[0041] Various examples are described in more detail below.
[0042] Figure 1 A vehicle 101 is shown.
[0043] exist Figure 1 In the example of FIG. 1 , a vehicle 101 , such as a passenger vehicle (PKW) or a truck (LKW), is equipped with a vehicle control device (also referred to as an electronic control unit, such as a control device, such as an Electronic Control Unit (ECU)) 102 .
[0044] Vehicle control device 102 has data processing components, such as a processor (eg, CPU (Central Processing Unit)) 103 and a memory 104 for storing control software 107 according to which vehicle control device 102 operates and data processed by processor 103. Processor 103 executes control software 107.
[0045] For example, the stored control software (computer program) has commands which, when executed by the processor, cause the processor 103 to perform driver assistance functions, i.e. functions of an ADAS (Advanced Driver Assistance System) or even autonomously control the vehicle (AD (Autonomous Driving)).
[0046] The control software 107 is transferred, for example, from the computer system 105 via the communication network 106 (or also by means of a storage medium such as a memory card) to the vehicle 101. This can also take place during operation (or at least when the vehicle 101 is with the user), because the control software 107 is updated to a new version over time, for example.
[0047] The control software 107 determines control actions for the vehicle (such as steering actions, braking actions, etc.) based on input data that are available to the control software and contain information about the environment or the control software derives information about the environment from this input data (such as through the detection of other traffic participants, such as other vehicles, pedestrians, cyclists, etc.). These input data are, for example, sensor data from one or more sensor devices 109, such as from a camera of the vehicle 101, which are connected to the vehicle control device 102 via a communication system 110 (for example, a vehicle bus system such as a CAN (Controller Area Network)).
[0048] The control software 107 can be trained at least partially, for example, by means of machine learning (ML), i.e., the control software 107 implements, for example, a machine learning model 108 (e.g., a neural network (NN)), which is trained in this example based on training data by the computer system 105. Thus, the computer system 105 implements an ML training algorithm for training one (or more) ML models 108.
[0049] For example, ML model 108 (e.g., a neural network) is an ML model for predicting the behavior of other traffic participants, such as pedestrians. This includes, for example, predictions of motion trajectories (e.g., a sequence of 2D positions of a walking person) and predictions of pose (e.g., predictions of whole-body joint positions relative to a designated "central" body part or joint (e.g., pelvis or hips). The latter is of interest because, for example, the direction in which a pedestrian's head rotates provides information about the direction in which the pedestrian intends to walk. For example, a typical left-right look may indicate that the pedestrian intends to cross the road.
[0050] Motion trajectory prediction and pose prediction are also of interest for other applications, such as other mobile robots. In collaborative manufacturing environments, for example, it is desirable for robots to accurately predict the future positions and movements (e.g., arm movements) of human workers in order to provide safe and efficient support in collaborative assembly tasks.
[0051] Figure 2 The prediction of the pose 201 is illustrated, and the motion trajectory prediction is illustrated in diagram 202 (ground truth (GroundTruth) and prediction, respectively, starting from the prediction start time point 0s).
[0052] Motion trajectory prediction and posture prediction can be treated as separate problems, but then two machine learning models are required accordingly. Here, for example, the motion trajectory is represented as the displacement [dx, dy] (or velocity) of the 2D position above the ground plane. This encoding of the 2D motion trajectory provides a similar representation of various motion trajectories, regardless of their absolute (X, Y) coordinates, which enables generalization for motion trajectory prediction by the machine learning model. For posture prediction, in addition to the 2D position coordinates, the 3D joint coordinates (or body part coordinates) are also encoded relative to the 2D position.
[0053] It is also possible to merge the decoupled (separate) pose and motion trajectory prediction modules at the end of the prediction pipeline to achieve simultaneous prediction of pose and motion trajectory. However, this also leads to a large (and therefore memory-intensive) and computationally intensive machine learning model (e.g., a neural network), which performs poorly due to the separate processing of walking dynamics and body movement dynamics, which are actually tightly coupled.
[0054] Therefore, according to various embodiments, a coupled method (for joint prediction of motion trajectory and posture) is provided, wherein training is performed with the aid of a transformation (of the input data) in global coordinates, i.e., the entire sequence of joint coordinates of the 3D skeleton is represented relative to a unique reference point (for example, based on this, the position of the pelvic joint is predicted at the "prediction time point", i.e., the time point of the last state of the input state sequence).
[0055] Through this joint prediction, the number of parameters of the machine learning model can be approximately halved relative to the method in which there are two prediction modules and the results of the two prediction modules are combined as described above, and faster inference time and more accurate motion trajectory prediction can be achieved (this is a key part of the problem when it comes to mobile robots or autonomous vehicles).
[0056] According to various embodiments, coupled prediction of motion trajectory and posture is performed with the aid of a machine learning model with a converter architecture (which corresponds, for example, to the machine learning model 108). Here, the 3D body part positions (in the global coordinate system) are used directly, which are obtained, for example, from a 3D position estimation pipeline for humans. That is, instead of processing the position of a reference body part (e.g., the hip) separately from the positions of other body parts, all body part positions are used together as input. Therefore, the input of the machine learning model is a (posture) trajectory in the form of an (input) sequence of states (or "frames") of a sequence of time points up to the prediction start time point, wherein each state contains the 3D positions of a specified set of body parts (e.g., joint positions or limb center positions, etc.). Here, the term "posture trajectory" is used for the posture sequence (series) (but hereinafter also referred to as "trajectory" for short). In contrast, a "motion trajectory" refers only to the position sequence of the corresponding body (i.e., the position of a central body part such as the hip, for example), but does not refer to the positions of multiple body parts in the global coordinate system (i.e., the reference frame).
[0057] According to various embodiments, the input (i.e., the pose trajectory) is transformed by means of a (input data) transformation in order to couple the pose and trajectory prediction flows in a unified architecture and thus simplify the machine learning model for simultaneously predicting poses and trajectories as a coupled task. For training trajectories that, in addition to the input sequence of states (i.e., past trajectories), also contain states of one or more other time points (after the prediction start time point, i.e., future trajectories) as ground truth, these states are also transformed for training or converted back to predictions of the machine learning model, which are then compared with the ground truth for loss calculation.
[0058] Figure 3The input data conversion for four training tracks 301 , 302 , 303 , 304 according to one embodiment is explained.
[0059] The four training trajectories 301-304 are shown in global coordinates in a first diagram 305 and are shown as 2D position sequences for simplicity. The dashed lines are past trajectories, and the solid lines are future trajectories (the future trajectories are shown here for illustration purposes but omitted for inference).
[0060] The conversion comprises: the training trajectories 301-304 are oriented at the origin (X=0, Y=0) and in the positive X-axis direction, which is shown in more detail in the second diagram 306 for the trajectories: the trajectory 307 is converted into a converted trajectory 308, so that the prediction start time point is set to the origin and the direction of movement at the prediction start time point corresponds to the X-axis direction.
[0061] That is, the input data transformation is used to transform the input trajectory (especially the training trajectory) into a common reference frame (especially for learning) by means of rotation and translation. The result is the transformed input trajectory 309. This enables generalization beyond various training trajectories in global coordinates and enables training using absolute 3D coordinates without the need for decomposition into separate pose and trajectory prediction data streams.
[0062] In addition to input data transformation, various embodiments employ a Graph Attention Network (GAT) to generate graph embeddings that capture the spatial skeleton architecture and inform the machine learning model about the skeleton hierarchy and the relative dependencies between individual joints. To this end, the GAT encoder is applied to the states of the input trajectory using the skeleton's adjacency matrix, resulting in more accurate and realistic predictions of body dynamics.
[0063] Figure 4 A machine learning model 400 for gesture trajectory prediction (ie, joint prediction of motion trajectory and gesture) is shown, in accordance with an embodiment.
[0064] In the following, represents a body pose (e.g., a human pose) at time point t, which includes N 3D body part positions (e.g., joint positions): Among them, each The three spatial coordinates (x, y, z) representing the position of the i-th body part (e.g., the i-th joint) in the robot's coordinate system at time t. The input trajectory (input sequence of states) is the sequence of postures from time 0 until the prediction start time T1: They all refer to the same person:
[0065]
[0066] The goal of the machine learning model 400 is to use global translation to predict the sequence of postures from time point T1+1 to time point T1+T2
[0067]
[0068] As mentioned above, for the input trajectory First, input data conversion 401 is performed.
[0069] This input data transformation is used to normalize the input trajectory to a common reference frame in order to generalize and predict human motion in global coordinates beyond various motion directions. Figure 3 As described, global invariance is established by ensuring that all predictions start from a consistent origin by performing a translation using a vector v that is the negative counterpart of the position of the reference body part (with index r (for the “root”, e.g., hip) in the last state of the input sequence, i.e., v = -j r (T1). Applies a translation using vector v to the input sequence For each state of , this produces the sequence for Each state (frame) of The relative displacement pose in is given by the following formula:
[0070]
[0071] By this displacement, the human pose of the last state of the input sequence is anchored at the origin, which represents a unified starting point for subsequent motion prediction.
[0072] Furthermore, directional invariance is established by transforming the input data 401 in such a way that the direction of motion is canonically oriented on the positive x-axis. To this end, a rotation angle θ is calculated based on the direction of motion of the input trajectory (at the prediction start time) and the positive x-axis. This angle is calculated as the inverse tangent of the ratio of the difference between the y and x coordinates of the position of the reference body part at the last state of the input sequence T1 and the position of the reference body part at the previous time point (T1-w), where w (standing for "window") is a specified time interval:
[0073]
[0074] Finally, the corresponding rotation matrix for rotation about the z-axis is applied to the shifted sequence S', which produces the rotated sequence S":
[0075]
[0076] The posture sequence S″ is the result of the input data transformation 401.
[0077] Now, as already briefly described above, GAT (Graph Attention Network) 402 generates a spatial graph embedding from a pose sequence S ″. Any human pose can be represented as a graph 403, where each joint in the pose corresponds to a node and the connections between them (bones or body parts) are edges. The input of GAT 402 is a redesigned version of the pose sequence S ″: each joint (or in general each body part) is a node in the graph 403, which has three spatial features (given by the position of the body parts). The edges E are determined based on the kinematic chain of the used body skeleton, which may vary from one dataset to another. From this kinematic chain, the adjacency matrix A is derived such that when joint j i (t) and j j (t) When there is a connection between ij =1, otherwise 0. GAT402 calculates attention scores e between joint pairs, which record the importance of one joint compared to another:
[0078] e ij =LeakyReLU(a T [Wj i (t)||Wj j (t)]) (6)
[0079] Where || represents a link, and a and W are learnable parameters (vectors or matrices). When using normalized attention coefficients, the common features are updated by aggregating information from neighboring joints, which produces a joint embedding.
[0080] Therefore, GAT 402 generates a new set of node features as output, and thereby generates a joint embedding for all poses of the input sequence.
[0081] Then, the joint embedding of the input sequences is flattened so that Creates a holistic pose embedding for the input sequence. GAT 402 is used to facilitate attention between body parts within a frame and effectively detect spatial relationships. The output of GAT 402 is fed into transformer (network) 404. The transformer is designed to identify and learn temporal relationships between poses, ensuring a comprehensive understanding of the spatial and temporal dynamics in the data.
[0082] In addition to the embedding generated by the GAT 402, an additional (spatio-temporal) position encoding 405 is used in order to record human dynamics in detail.
[0083] To this end, sinusoidal spatial position codes are generated that are specifically designed to distinguish different joints in each image. For each joint, the method generates a dimension J dim The position encoding of the position encoding takes into account all N joints. In addition, a time encoding is generated to record the time-varying process of the sequence. These time encodings have dimension J dim ×N, taking into account all frames and accounting for the continuous dynamics from one frame to the next. The spatial position encoding is flattened and then merged with the temporal encoding in order to obtain a unified 2D position encoding 405.
[0084] The transformer 404 has a classic transformer architecture with an encoder 406 and a decoder 407. The encoder 406 takes the output of the GAT 205 and the (spatial-temporal) position encoding 405 and processes them through L layers 408. Each layer contains a self-attention block (masked single-headed self-attention with RPP (Relative Position Representations)) 409 and a feedforward network 411 (followed by corresponding addition and normalization operations 410, 412, respectively). Layer 408 processes the input into a sequence of elements Z = [z1, z2, ..., z in the latent space. T ].
[0085] The decoder 407 then uses this sequence of elements of the latent space to generate a sequence of output positions 413. The decoder comprises a plurality of layers 414, each having a self-attention block (masked single-headed self-attention with RPP) 415, a multi-headed cross-attention block 417 and a feedforward network 419 (also followed by corresponding addition and normalization operations 416, 418, 420, respectively), and after these layers is a final multi-headed attention block 421, which is followed by a final feedforward network 422.
[0086] In order to ensure that the predicted pose depends only on the previous pose (i.e. until T1) and not on future poses, “Casual Masked Self-Attention” is used not only for the encoder 406 but also for the decoder 407. The core principle of masked self-attention is that the weights determined during the calculation of the attention weights are set to extremely negative values, such as 10, by using a mask. -9 The mask is designed such that for a particular sequence position, all future positions in the sequence are marked as irrelevant. If a Softmax function is subsequently applied to these weights during the attention mechanism, the values corresponding to the masked positions are ignored.
[0087] In self-attention with relative position representation, the attention value of a particular sequence position is given greater weight towards closely adjacent poses. This is beneficial for sequences with human poses, where not only the order of positions is important, but also the relative transitions between images. This focus on relative distances can lead to a better understanding of human movement processes.
[0088] The decoder's self-attention query is initialized with x(T1) (as Query 423), i.e., with the final input state, which is repeated multiple times depending on the number of desired output poses.
[0089] After the decoding phase of layers 414, the final multi-head attention block 421 causes a multi-head attention mechanism, also called end-attention, where the outputs of these layers 414 are used as queries and the outputs of the GAT 402 are considered not only as keys but also as values.
[0090] The embedding output by the multi-head attention block 421 is then propagated through the linear layer of the final feed-forward network 422 and returned to the original motion direction and global coordinate space via an inverse transformation 424 with the aid of v and θ, which produces the predicted 3D pose and associated trajectory, i.e., a sequence of output positions:
[0091]
[0092] The machine learning model 400 can be trained using a combined pose and trajectory loss. Since each pose has a dimension of 3N (where N represents the number of body parts), the predicted pose sequence is a vector sequence And the ground truth position sequence is the vector sequence yT1+1, yT1+2, ..., yT1+T2. If L1 denotes the pose loss and L2 denotes the trajectory loss, then the combined loss (of the training trajectory) is L=L1+L2, where
[0093]
[0094] and
[0095]
[0096] Over a batch of training trajectories, this combined loss can be aggregated (or averaged) into a total (batch) loss, and the parameters (weights, etc.) of the machine learning model can be adjusted in a direction that reduces the total loss.
[0097] In summary, according to various embodiments, there is provided a Figure 5 The method shown in .
[0098] Figure 5 A flowchart 500 is shown presenting a method for training a machine learning model to determine body posture and position of a body having multiple body parts (e.g., a person or an animal, but also, for example, a mechanical system having one or more degrees of freedom with respect to the relative positions of components) in accordance with an embodiment.
[0099] At 501, for each training trajectory of a plurality of training trajectories, wherein each training trajectory indicates a position of each body part (e.g., a joint or a connection between two joints) of a plurality of body parts in a global coordinate system for each time point of a corresponding sequence of time points,
[0100] At 502, converting the training trajectory into a converted training trajectory such that, for a time point of the sequence of time points defined as a prediction start time point, a position of a specified reference body part coincides with a specified point (e.g., an origin) of a global coordinate system, and a movement direction of the reference body part corresponds to a specified direction in the global coordinate system;
[0101] At 503, predicting (i.e., determining) the positions of the body parts at one or more time points after the prediction start time point by means of a machine learning model, by feeding the positions of the body parts indicated by the converted training trajectories up to the prediction start time point to the machine learning model;
[0102] At 504 , a (single) loss is determined by comparing the predicted position with the positions of the body parts indicated by the training trajectory for one or more time points after the predicted start time point in the training trajectory.
[0103] At 505 , the machine learning model is adjusted to reduce the total loss including the determined losses (i.e., (trainable) parameters (e.g., weights) are adjusted in a direction such that the total loss including the individual losses decreases).
[0104] Figure 5 The method can be performed by one or more computers having one or more data processing units. The term "data processing unit" can be understood as any type of entity capable of processing data or signals. For example, these data or signals can be processed according to at least one (that is, one or more than one) specific function, and the function is performed by the data processing unit. The data processing unit may include an analog circuit, a digital circuit, a logic circuit, a microprocessor, a microcontroller, an integrated circuit of a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a field programmable gate array (FPGA), or any combination thereof or be constructed by these. Any other means for implementing the corresponding functions described in more detail herein may also be understood as a data processing unit or a logic circuit device. One or more of the method steps described in detail here can be implemented (e.g., realized) by a data processing unit through one or more specific functions, and these functions are performed by the data processing unit.
[0105] Thus, according to various embodiments, the method is particularly computer-implemented.
[0106] After training, the machine learning model can be used to generate control signals for a robotic device by feeding the machine learning model with sensor data about its environment or (past) gesture trajectories (e.g., of a person) derived therefrom, and thereby generating predictions for one or more gestures (in a global coordinate system). The term "robotic device" is to be understood as referring to any technical system (having mechanical parts whose movement is controlled), such as a computer-controlled machine, a vehicle, a household appliance, a power tool, a manufacturing machine, a personal assistant, or an access control system.
[0107] Various embodiments may receive time series of sensor data from various sensors such as video, radar, LiDAR, ultrasound, motion, thermal imaging, etc., and use these time series to determine past gestures.
Claims
1. A method for training a machine learning model (108, 400) to determine body posture and position of a body having a plurality of body parts, the method comprising: For each training trajectory (307) in a plurality of training trajectories, where Each training trajectory indicates, for each time point of a corresponding sequence of time points, a position of each body part of the plurality of body parts in a global coordinate system, converting (502) the training trajectory (307) into a converted training trajectory (308) such that for a time point of the sequence of time points defined as a prediction start time point, the position of the specified reference body part coincides with the specified point of the global coordinate system, and the direction of movement of the reference body part corresponds to the specified direction in the global coordinate system, predicting (503) the position of the body part at one or more time points after the prediction start time point by feeding the position of the body part indicated by the converted training trajectory (308) up to the prediction start time point to the machine learning model (108, 400), and determining (504) a loss by comparing the predicted position with a position of the body part indicated by the training trajectory (307) for one or more time points after the predicted start time point in the training trajectory (307); and The machine learning model (108, 400) is adjusted (505) to reduce the total loss including the determined loss.
2. The method according to claim 1, wherein The machine learning model (108, 400) has a graph attention network (402), with the help of which the position of the body part indicated by the converted training trajectory (308) until the prediction start time point is processed in the following manner: for each time point until the prediction start time point, a node with the corresponding indicated position as a node feature is assigned to each body part and the nodes assigned to the connected body parts are connected by edges, the corresponding posture is represented as a graph (403), and the graph (403) is processed with the help of the graph attention network (402).
3. The method according to claim 1 or 2, wherein The machine learning model (108, 400) has a transformer architecture.
4. The method according to any one of claims 1 to 3, comprising: determining a space-time encoding (405) of the position indicated by the converted training trajectory (308) until the prediction start time point, and delivering the space-time encoding (405) together with the position of the body part indicated by the converted training trajectory (308) until the prediction start time point to the machine learning model (108, 400).
5. The method according to any one of claims 1 to 4, wherein The loss includes a first loss component, which includes, for each of the one or more time points after the prediction start time point and each of the body parts, a difference between the predicted position of the body part and the position of the body part indicated by the training trajectory or the converted training trajectory as a loss contribution; and / or wherein the one or more time points after the prediction start time point have one or more pairs of consecutive time points, and the loss includes a second loss component, which includes, for each of the one or more pairs and each of the body parts, a difference between the predicted position for the pair of time points and the position of the body part indicated by the training trajectory or the converted training trajectory for the pair of time points as a loss contribution.
6. A method for predicting one or more body poses and one or more positions of a body having a plurality of body parts, the method comprising: Training a machine learning model (108, 400) according to any one of claims 1 to 5; detecting a trajectory of the body, the trajectory indicating, for each detected time point of a sequence of detected time points, a position of each body part of the plurality of body parts in the global coordinate system; converting the detected trajectory into a converted detected trajectory such that, for a last one of the detection time points, the position of the specified reference body part coincides with the specified point of the global coordinate system, and the direction of movement of the reference body part corresponds to the specified direction in the global coordinate system; and With the aid of the machine learning model (108, 400), the position of the body part is predicted by feeding the position of the body part indicated by the converted detected trajectory to the machine learning model (108, 400).
7. The method according to claim 6, further comprising: controlling a robotic device (101) as a function of the predicted position of the body part.
8. A data processing system (102, 105) configured to carry out the method according to any one of claims 1 to 7. 9 . A computer program comprising instructions which, when executed by a processor, cause the processor to perform the method according to claim 1 .
10. A computer-readable medium storing instructions which, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 7.