A Human Posture Prediction Method, Device and Medium

By constructing the undirected space-time graph structure of the human body and designing corresponding prediction algorithm models, the problem that the human posture prediction method in the prior art cannot effectively combine the characteristics of the human body skeleton and inertial motion device is solved, and efficient human posture prediction is achieved.

CN119670813BActive Publication Date: 2025-05-27东方电气长三角(杭州)创新研究院有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510188382.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-27
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

The existing human posture prediction method based on inertial motion data cannot effectively combine the characteristics of the human skeleton and inertial motion device, and its design is not sufficient to meet the timing prediction needs.

Method used

By constructing an undirected space-time graph structure of the human body, the motion acceleration, angular velocity and undirected space-time graph adjacency matrix of various parts of the human body are input into the human body posture prediction algorithm model, and the analysis module, interaction prediction module and optimization module are used to perform feature analysis, space-time interaction and data optimization to predict the joint angles of various parts of the human body.

Benefits of technology

It effectively solves the problem of complex and irrelevant data interference, realizes the aggregation of multi-scale spatial information and the flow of complex spatial and temporal information, and improves the accuracy and efficiency of human posture prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119670813B_ABST
    Figure CN119670813B_ABST
Patent Text Reader

Abstract

The present invention discloses a human body pose prediction method, device and medium, including: obtaining the motion acceleration and angular velocity of each part of the human body; constructing an undirected spatio-temporal graph structure of the human body: obtaining the adjacency matrix of the undirected spatio-temporal graph of the human body based on the undirected spatio-temporal graph structure of the human body; inputting the motion acceleration, angular velocity of each part of the human body and the adjacency matrix of the undirected spatio-temporal graph of the human body into a human body pose prediction algorithm model to predict the joint angles of each part of the human body; obtaining the human body pose based on the predicted joint angles of each part of the human body; by introducing the undirected spatio-temporal graph structure of the human body, establishing an association between the motion data of the inertial sensors worn by the human body and the human body limb segments, and jointly using them as the input of the human body pose prediction algorithm model, providing key spatial feature information for the human body pose prediction algorithm model, and effectively solving the problem of interference from complex irrelevant data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and particularly to a human body posture prediction method, device and medium. Background Art

[0002] On the one hand, traditional human body posture recognition based on inertial motion data uses kinematic theory of the human body to recognize real-time human body postures, and cannot predict and estimate human body postures. On the other hand, the methods for predicting human body postures based on inertial motion data mainly use convolutional neural networks, but the deficiencies of such network models are mainly manifested as follows: 1. The design of the network model does not combine the characteristics of the human skeleton and inertial motion devices; 2. The network design does not carry out design for time series prediction requirements. Summary of the Invention

[0003] The purpose of the present invention is to provide a human body posture prediction method, device and medium for the deficiencies of the prior art.

[0004] The purpose of the present invention is achieved by the following technical solutions: A human body posture prediction method includes:

[0005] Obtaining the motion acceleration and angular velocity of each part of the human body; the human body parts include the head and neck, spine, pelvis, left and right upper arms, left and right forearms, left and right thighs, left and right calves, left and right feet;

[0006] Constructing an undirected spatio-temporal graph structure of the human body, regarding human body nodes as undirected spatio-temporal graph nodes, and the edges of the undirected spatio-temporal graph include two types: (1) the connections between graph nodes in each time frame; (2) the connections of each graph node among multiple time frames; the human body nodes include the cervical vertebra, lumbar vertebra, sacrum, left and right shoulder joints, left and right elbow joints, left and right hip joints, left and right knee joints, left and right ankle joints; obtaining the undirected spatio-temporal graph adjacency matrix of the human body based on the undirected spatio-temporal graph structure of the human body;

[0007] Inputting the motion acceleration, angular velocity of each part of the human body and the undirected spatio-temporal graph adjacency matrix of the human body into a human body posture prediction algorithm model to predict the joint angles of each part of the human body;

[0008] Obtaining the human body posture based on the predicted joint angles of each part of the human body;

[0009] As Figure 1 shown, the human body posture prediction algorithm model includes an analysis module, an interaction prediction module, and an optimization module;

[0010] The analysis module is used to analyze key feature information from the motion acceleration and angular velocity of each part of the human body;

[0011] The interaction prediction module predicts the motion state through feature interaction in time and space;

[0012] The optimization module is used to optimize the predicted motion state data sequence.

[0013] Further, the obtaining of the motion acceleration and angular velocity of each part of the human body includes:

[0014] The positions where the human body wears inertial motion sensors are the head and neck, spine, pelvis, left and right upper arms, left and right forearms, left and right thighs, left and right calves, and left and right feet; inertial motion data of each part of the human body in the sensor coordinate system is collected through the inertial motion sensors, and the inertial motion data includes inertial motion acceleration, angular velocity, and magnetic field strength;

[0015] Using the human body pose calibration method, the inertial motion data of each part of the human body in the sensor coordinate system is transformed into the human body coordinate system to obtain the motion acceleration and angular velocity of each part of the human body.

[0016] Further, the parsing module includes two comprehensive encoder modules; the input of the first comprehensive encoder module is the original input data , that is, the motion acceleration and angular velocity of each part of the human body; the input of the second comprehensive encoder module is the original input data combined with the residual of the output of the first comprehensive encoder module ; the output of the parsing module is the original input data , the original input data combined with the residual of the output of the first comprehensive encoder module , and the output of the second comprehensive encoder module, and the residual combination of the three;

[0017] The comprehensive encoder module includes a spatial position encoding module, a temporal convolutional network module, and an attention mechanism encoder module; the input of the spatial position encoding module is the original input data , the input of the temporal convolutional network module is the combination of the output of the spatial position encoding module and the input of the comprehensive encoder module, and the input of the attention mechanism encoder module is the output of the temporal convolutional network module;

[0018] Among them, the temporal convolutional network module uses causal convolution and dilated convolution to process the time series task of the original input data ;

[0019] Among them, the spatial position encoding module uses absolute position encoding to generate an absolute position encoding matrix for the input data of the comprehensive encoder and assigns a unique encoding to each position data in the input sequence;

[0020] Among them, the attention mechanism encoder module includes multiple layers of encoders. Each layer of encoder includes multi-head self-attention, a feed-forward neural network, and a normalization layer, which are connected in sequence. Among them, multi-head self-attention is composed of multiple self-attention modules. That is, the input vector is respectively passed into n self-attention modules to obtain n outputs, and then the outputs of multiple self-attention modules are combined and passed into the feed-forward neural network and the normalization layer to obtain the final output Z.

[0021] Furthermore, the input of the interaction prediction module includes the original input data , the output of the parsing module and the undirected spatio-temporal graph adjacency matrix . After being processed by the interaction prediction module, the output is obtained;

[0022] The interaction prediction module includes a spatial interaction module and a temporal interaction module;

[0023] Among them, based on the undirected spatio-temporal graph, the spatial interaction module uses a spatial segmentation strategy to spatially divide the limb segments of the human body graph structure. For various spatial segmentation strategies, using the original input data as the input, using the self-attention module, first solve the attention weights between joints within each spatial domain respectively, and then perform summation normalization to obtain the inter-domain attention weights of the entire human body spatial domain ;

[0024] At the same time, the spatial interaction module uses the original input data as the input, introduces a time window mechanism, slides in the time dimension to obtain the joint information of multiple time windows, uses the original input data as the query volume, uses the joint data of each time window as the comparison volume, uses the self-attention module to solve the attention weights between the original input data and the joint data of each time window, and then performs summation normalization to obtain the attention weights of the entire time domain ;

[0025] Aggregate and the two attention weights to extract the weight information of spatio-temporal information. Using this weight information, update the adjacency matrix of the undirected graph to obtain the adjacency matrix that fuses spatio-temporal information;

[0026] Finally, use the updated adjacency matrix and the output of the parsing module to perform dimension fusion using Einstein summation convention to obtain the output ;

[0027] The temporal interaction module takes Take the input, perform interactions in the time dimension using dilated causal convolutions and residual operations, and obtain the output .

[0028] Furthermore, the spatial segmentation strategy includes the following four types:

[0029] Spatial segmentation strategy one: Consider all limb segments in the human body diagram structure as a whole without spatial segmentation;

[0030] Spatial segmentation strategy two: Divide the human body diagram structure into two regions. The upper limb region includes the pelvis, L5 segment, L3 segment, T12 segment, T8 segment, head and neck, left and right hands, left and right forearms, left and right upper arms, and left and right shoulders; the lower limb region includes the pelvis, left and right thighs, left and right calves, and left and right feet;

[0031] Spatial segmentation strategy three: Divide the human body diagram structure into three regions: upper limb trunk, upper limb arms, and lower limbs. The upper limb trunk includes the pelvis, L5 segment, L3 segment, T12 segment, T8 segment, and head and neck; the upper limb arms include the left and right hands, left and right forearms, left and right upper arms, left and right shoulders, and T8 segment; the lower limbs include the pelvis, left and right thighs, left and right calves, and left and right feet;

[0032] Spatial segmentation strategy four: Divide all limb segments into five regions: upper limb trunk, left upper limb, right upper limb, left lower limb, and right lower limb; the upper limb trunk includes the pelvis, L5 segment, L3 segment, T12 segment, T8 segment, and head and neck, the left upper limb includes the left hand, left forearm, left upper arm, left shoulder, and T8 segment; the right upper limb includes the right hand, right forearm, right upper arm, right shoulder, and T8 segment, the left lower limb includes the pelvis, left thigh, left calf, and left foot, and the right lower limb includes the pelvis, right thigh, right calf, and right foot.

[0033] Furthermore, the input of the optimization module is the output of the interaction prediction module ; the optimization module is composed of several single-stage optimization modules and a time series model based on the attention mechanism connected together;

[0034] The single-stage optimization module includes a normal convolutional layer, a dilated convolutional layer, a batch normalization layer, and an activation function layer connected in sequence; the convolutional layer is used to adjust the dimension of the input features and match the number of feature maps in the network, use dilated convolutions, and use a doubling dilation rate in different single-stage optimization modules, and then use batch normalization and activation functions to obtain , finally, use residual connections to combine the input and to obtain the output .

[0035] Furthermore, the time series model based on the attention mechanism includes an encoder module and a decoder module;

[0036] Use a recurrent neural network as the encoder, combine the input vector of the current time step and the hidden state of the previous time step to transform and output the hidden state of the current time step Perform a weight distribution process on the hidden layer state quantity of the encoder and use the weights of the hidden state quantity to obtain the context vector ; The decoder uses a recurrent neural network and uses the output of the previous time step , the encoder context vector and the hidden state of the previous time step to determine the hidden state of the current time step ;

[0037] Finally, based on the output of the previous time step , the decoder hidden state vector , and the context vector , use a fully connected feedforward neural network to output the final predicted value to predict the joint angles of each part of the human body.

[0038] The present invention also provides a human body posture prediction device, including:

[0039] A motion data acquisition module for acquiring the motion acceleration and angular velocity of each part of the human body; the human body parts include the head and neck, spine, pelvis, left and right upper arms, left and right forearms, left and right thighs, left and right calves, and left and right feet;

[0040] A joint angle prediction module that constructs an undirected spatio-temporal graph structure of the human body, regards the human body nodes as undirected spatio-temporal graph nodes, and the edges of the undirected spatio-temporal graph include two types: (1) the connections between graph nodes in each time frame; (2) the connections between each graph node among multiple time frames; the human body nodes include the cervical vertebra, lumbar vertebra, sacrum, left and right shoulder joints, left and right elbow joints, left and right hip joints, left and right knee joints, and left and right ankle joints; obtain the adjacency matrix of the undirected spatio-temporal graph of the human body based on the undirected spatio-temporal graph structure of the human body; input the motion acceleration, angular velocity, and adjacency matrix of the undirected spatio-temporal graph of each part of the human body into the human body posture prediction algorithm model to obtain the joint angles of each part of the human body;

[0041] The human body posture prediction algorithm model includes an analysis module, an interactive prediction module, and an optimization module;

[0042] The analysis module is used to analyze key feature information from the motion acceleration and angular velocity of each part of the human body;

[0043] The interaction prediction module predicts the motion state through feature interaction in time and space;

[0044] The optimization module is used to optimize the predicted motion state data sequence;

[0045] The posture transformation module is used to obtain the human body posture based on the joint angles of each part of the human body; it also includes obtaining the ergonomic characteristic parameters based on the joint angles of each part of the human body to evaluate and determine the potential risk of musculoskeletal injury.

[0046] The present invention also provides a human body posture prediction device, including one or more processors for implementing the above-mentioned human body posture prediction method.

[0047] The present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it is used to implement the above-mentioned human body posture prediction method.

[0048] Compared with the prior art, the beneficial effects of the embodiments of the present invention are:

[0049] 1. By introducing the human body undirected spatio-temporal graph structure, the present invention establishes an association between the motion data of the inertial sensors worn by the human body and the human body limb segments, and jointly uses them as the input of the human body posture prediction algorithm model. Compared with traditional methods, this method provides key spatial feature information for the human body posture prediction algorithm model and effectively solves the problem of interference from complex irrelevant data;

[0050] 2. In the present invention, the joint angle prediction module is based on the human body undirected spatio-temporal graph structure. Through the spatial interaction strategy and the time interaction strategy, it synchronously extracts and focuses on information at different spatial scales and different time spans, realizes the aggregation of multi-scale spatial information, and promotes the flow of complex spatio-temporal information;

[0051] 3. The posture transformation module in the present invention uses the predicted human joint angle data to predict the ergonomic characteristic parameters, realizes the end-to-end prediction of ergonomic evaluation parameters, and provides data support for evaluating and determining the potential risk of musculoskeletal injury. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained without creative efforts based on these drawings.

[0053] Figure 1 It is a schematic diagram of a human body posture prediction algorithm model provided by an embodiment of the present invention;

[0054] Figure 2 Schematic diagram of an analysis module structure provided by an embodiment of the present invention;

[0055] Figure 3 Schematic diagram of a cross-interference measurement module structure provided by an embodiment of the present invention;

[0056] Figure 4 Schematic diagram of an optimization module structure provided by an embodiment of the present invention;

[0057] Figure 5 Schematic diagram of a hardware structure provided by an embodiment of the present invention. Detailed implementation manners

[0058] The present invention will be described in detail below with reference to the accompanying drawings. Without conflict, the features in the following embodiments and implementation manners can be combined with each other.

[0059] A human body posture prediction method of the present invention includes the following steps:

[0060] (1) Obtain the motion acceleration and angular velocity of each part of the human body: The positions where the human body wears the inertial motion sensors are the head and neck, spine, pelvis, left and right upper arms, left and right forearms, left and right thighs, left and right calves, and left and right feet. Collect the inertial motion data of each part of the human body in the sensor coordinate system and convert it into the motion data of each part of the human body in the human body coordinate system;

[0061] Specifically, the obtaining of the motion acceleration and angular velocity of each part of the human body includes:

[0062] Collect the inertial motion data of each part of the human body in the sensor coordinate system through the inertial motion sensor, and the inertial motion data includes inertial motion acceleration, angular velocity, and magnetic field strength;

[0063] Convert the inertial motion data of each part of the human body in the sensor coordinate system to the human body coordinate system to obtain the motion acceleration and angular velocity of each part of the human body.

[0064] Specifically, use the calibration posture method to convert the inertial motion data collected in real time for each part of the human body to the human body coordinate system to obtain the motion acceleration and angular velocity of each part of the human body;

[0065] In one embodiment, after a human body wears a multi-node inertial motion sensor, an attitude calibration action needs to be completed. Among them, the attitude calibration action is that the wearer's arms hang naturally on both sides of the body, the legs are together and straight, and stand upright facing the geomagnetic north direction (determined by the magnetic field intensity in the inertial motion data). Obtain the motion acceleration and angular velocity data of the inertial motion sensors placed at the head and neck, spine, pelvis, left and right upper arms, left and right forearms, left and right thighs, left and right calves, and left and right feet in the sensor coordinate system under this calibrated attitude; then, in combination with the acceleration and angular velocity of each part of the human body in the known human body coordinate system under the calibration attitude, obtain the attitude transformation matrix of each part of the human body in the real-time state; use this attitude transformation matrix to convert the inertial motion data of each part of the human body in the sensor coordinate system to obtain the motion acceleration and angular velocity of each part of the human body in the human body coordinate system.

[0066] (2) Construct an undirected spatio-temporal graph structure of the human body. Regard the human body nodes (including the cervical vertebra, lumbar vertebra, sacrum, left and right shoulder joints, left and right elbow joints, left and right hip joints, left and right knee joints, left and right ankle joints) as undirected spatio-temporal graph nodes. The edges of the undirected spatio-temporal graph include two types: (1) the connections between graph nodes in each time frame; (2) the connections of each graph node between multiple time frames. Construct an undirected spatio-temporal graph structure of the human body to obtain the initial undirected spatio-temporal graph adjacency matrix; input the motion acceleration, angular velocity, and initial undirected spatio-temporal graph adjacency matrix of each part of the human body into the human body pose prediction algorithm model to obtain the joint angles of each part of the human body; the human body pose prediction algorithm model includes an analysis module, an interaction prediction module, and an optimization module;

[0067] The analysis module is used to analyze key feature information from the input inertial motion data;

[0068] The interaction prediction module predicts the motion state through feature interaction in time and space;

[0069] The optimization module is used to optimize the predicted motion state data sequence.

[0070] In one embodiment, as Figure 2 shown, the analysis module includes two comprehensive encoder modules; the input of the first comprehensive encoder module is the original input data , that is, the motion acceleration and angular velocity of each part of the human body; the input of the second comprehensive encoder module is the original input data and the residual combination of the output of the first comprehensive encoder module ; the output of the analysis module is the residual combination of the original input data , the original input data , the output of the first comprehensive encoder module, and the output of the second comprehensive encoder module Residual connection;

[0071] The comprehensive encoder module includes a spatial position encoding module, a temporal convolutional network module, and an attention mechanism encoder module; the input of the spatial position encoding module is the original input data , the input of the temporal convolutional network module is the combination of the output of the spatial position encoding module and the input of the comprehensive encoder module, and the input of the attention mechanism encoder module is the output of the temporal convolutional network module;

[0072] Among them, the temporal convolutional network module uses causal convolution and dilated convolution to process the time series task of the original input data .

[0073] Among them, the spatial position encoding module uses absolute position encoding to generate an absolute position encoding matrix for the input data of the comprehensive encoder, and assigns a unique encoding to each position data in the input sequence.

[0074] Among them, the attention mechanism encoder module includes multiple layers of encoders connected in sequence. Each layer of encoder includes multi-head self-attention, a feed-forward neural network, and a normalization layer. The multi-head self-attention is composed of multiple self-attention modules. That is, the input vector is respectively passed into n self-attention modules to obtain n outputs, namely Z1, Z2, Z3,..., Zn; then the outputs of multiple self-attention modules are combined and passed into the feed-forward neural network and the normalization layer to obtain the final output Z of the multi-head self-attention.

[0075] Two comprehensive encoder modules are used in the parsing module. The comprehensive encoder module uses residual connection as the output. The residual connection in the first comprehensive encoder module uses the original input data and the data processed by the spatial position encoding module, the temporal convolutional network module, and the attention mechanism encoder module , the residual connection of the second comprehensive encoder module uses the original input data , the output data of the first comprehensive encoder module , the output data of the second comprehensive encoder module , to obtain the output .

[0076] In one embodiment, according to the positions where the multi-node inertial motion sensors are arranged on the human body, it is mapped into an undirected spatio-temporal graph structure. The cervical vertebra, lumbar vertebra, sacrum, left and right shoulder joints, left and right elbow joints, left and right hip joints, left and right knee joints, and left and right ankle joints of the human body are regarded as graph nodes. The edges of the graph include two types: (1) the connections between graph nodes in each frame; (2) the connections of each graph node between multiple frames. It is the adjacency matrix of the graph structure. Each weight coefficient represents the spatial relationship between two nodes, and the diagonal node degree matrix D is used to normalize it.

[0077]

[0078]

[0079] The input of the interaction prediction module includes the original input data , the output data of the parsing module and the initial undirected spatio-temporal graph adjacency matrix . After being processed by the interaction prediction module, the output is obtained.

[0080] As Figure 3 shown, the interaction prediction module includes a spatial interaction module and a temporal interaction module.

[0081] Among them, based on the undirected spatio-temporal graph, the spatial interaction module uses the spatial segmentation strategy to spatially divide the limb segments of the human body graph structure. For various spatial segmentation strategies, with the original input data as the input, using the self-attention module, first solve the attention weights between joints within each spatial domain respectively, and then perform summation normalization to obtain the inter-domain attention weights of the entire human body spatial domain .

[0082]

[0083] At the same time, the spatial interaction module takes the original input data as the input, introduces the time window mechanism, slides in the time dimension to obtain the joint information of multiple time windows. With the original input data as the query volume and the joint data of each time window as the comparison volume, using the self-attention module, solve the attention weights between the original input data and the joint data of each time window, and then perform summation normalization to obtain the attention weights of the entire time domain .

[0084] Aggregate the above two attention weights to extract the weight information of spatio-temporal information. Using this weight information, update the adjacency matrix of the undirected graph to obtain the adjacency matrix that fuses spatio-temporal information.

[0085]

[0086] Finally, use the updated adjacency matrix and the output of the parsing module to perform dimension fusion using the Einstein summation convention to obtain the output 。

[0087] The time interaction module takes as input, uses dilated causal convolution and residual operations to perform interactions in the time dimension, and obtains the output 。

[0088] Among them, the spatial segmentation strategy includes the following four types:

[0089] Spatial segmentation strategy one: All limb segments in the human body diagram structure are regarded as a whole without spatial segmentation.

[0090] Spatial segmentation strategy two: The human body diagram structure is divided into two regions. The upper limb region includes the pelvis, L5 segment, L3 segment, T12 segment, T8 segment, head and neck, left and right hands, left and right forearms, left and right upper arms, left and right shoulders; the lower limb region includes the pelvis, left and right thighs, left and right calves, left and right feet.

[0091] Spatial segmentation strategy three: The human body diagram structure is divided into three regions: upper limb trunk, upper limb arms, and lower limbs. The upper limb trunk includes the pelvis, L5 segment, L3 segment, T12 segment, T8 segment, head and neck; the upper limb arms include the left and right hands, left and right forearms, left and right upper arms, left and right shoulders, T8 segment; the lower limbs include the pelvis, left and right thighs, left and right calves, left and right feet.

[0092] Spatial segmentation strategy four: All limb segments are divided into five regions: upper limb trunk, left upper limb, right upper limb, left lower limb, and right lower limb. The upper limb trunk includes the pelvis, L5 segment, L3 segment, T12 segment, T8 segment, head and neck, the left upper limb includes the left hand, left forearm, left upper arm, left shoulder, T8 segment; the right upper limb includes the right hand, right forearm, right upper arm, right shoulder, T8 segment, the left lower limb includes the pelvis, left thigh, left calf, left foot, and the right lower limb includes the pelvis, right thigh, right calf, right foot.

[0093] In one embodiment, the input of the optimization module is the output of the interaction prediction module 。As Figure 4 shown, the optimization module includes a number of single-stage optimization modules (in this embodiment, 3 single-stage optimization modules are used for illustration, but it is not limited), and 1 time series model based on the attention mechanism.

[0094]

[0095]

[0096] Each single-stage optimization module includes a convolutional layer, a dilated convolutional layer, a batch normalization layer, and an activation function layer and is connected in sequence; the convolutional layer is used to adjust the input The dimension of the feature is matched with the number of feature maps in the network. Dilated convolutions are used, and doubling dilation rates are used in different single-stage optimization modules. Then, batch normalization and activation functions are used to obtain , and finally, residual connections are used to combine the input and to obtain the output .

[0097] The input of the previous stage of the first single-stage optimization module is , and the output is obtained;

[0098] The input of the previous stage of the second single-stage optimization module is , and the output is obtained;

[0099] The input of the previous stage of the third single-stage optimization module is , and the output is obtained.

[0100] Three single-stage optimization modules are continuously stacked and used, and a time series model based on the attention mechanism is used at the end to complete sequence-to-sequence learning. In the time series model based on the attention mechanism, a recurrent neural network is used as the encoder, which combines the input vector at the current time step and the hidden state at the previous time step, and transforms and outputs the hidden state at the current time step. Weight distribution processing is performed on the hidden layer state quantity of the encoder, and the weight of the hidden state quantity is used to obtain the context vector . The decoder uses a recurrent neural network, and uses the output at the previous time step, the encoder context vector and the hidden state at the previous time step to determine the hidden state at the current time step;

[0101] Finally, based on the output at the previous time step, the decoder hidden state vector , and the context vector , a fully connected feedforward neural network is used to output the final predicted value , and the joint angles of the main limb segments of the human body are obtained.

[0102] (3) Obtain human ergonomic characteristic parameters based on the joint angles of each part of the human body;

[0103] Specifically, by using the joint angles of various parts of the human body and based on ergonomic analysis, the human body posture is predicted to evaluate and determine the potential ergonomic state of the human body.

[0104] The present invention also provides a human body posture prediction device, including:

[0105] A motion data acquisition module, configured to acquire the motion acceleration and angular velocity of various parts of the human body; the human body parts include the head and neck, spine, pelvis, left and right upper arms, left and right forearms, left and right thighs, left and right calves, and left and right feet;

[0106] A joint angle prediction module, configured to construct an undirected spatio-temporal graph structure of the human body to obtain an undirected spatio-temporal graph adjacency matrix; input the motion acceleration, angular velocity of various parts of the human body and the undirected spatio-temporal graph adjacency matrix into a human body posture prediction algorithm model to predict the joint angles of various parts of the human body; construct the undirected spatio-temporal graph structure: regard the human body nodes as undirected spatio-temporal graph nodes, and the edges of the undirected spatio-temporal graph include two types: (1) the connections between graph nodes in each time frame; (2) the connections of each graph node between multiple time frames; the human body nodes include the cervical vertebra, lumbar vertebra, sacrum, left and right shoulder joints, left and right elbow joints, left and right hip joints, left and right knee joints, and left and right ankle joints;

[0107] A posture conversion module, configured to obtain human ergonomic characteristic parameters based on the joint angles of various parts of the human body. An embodiment provided by the present invention is to convert the predicted joint angle data into musculoskeletal injury characteristic parameters of the wrist, elbow, shoulder, neck, back and legs, such as posture, duration and action frequency, and evaluate and determine the potential musculoskeletal injury risk based on ergonomic methods.

[0108] It should be noted that the device embodiment shown in this embodiment matches the content of the above method embodiment, and the content of the above method embodiment can be referred to and will not be elaborated here.

[0109] Corresponding to the embodiment of the above-mentioned human body posture prediction method, the present invention also provides an embodiment of a human body posture prediction device.

[0110] See Figure 5 , a human body posture prediction device provided by an embodiment of the present invention includes one or more processors, configured to implement a human body posture prediction method in the above embodiment.

[0111] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0112] An embodiment of a human body posture prediction device of the present invention can be applied to any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented by software, or can be implemented by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities where it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. From a hardware level, as Figure 5 shown, it is a hardware structure diagram of any device with data processing capabilities where a human body posture prediction device of the present invention is located. Except for Figure 5 the shown processor, memory, network interface, and non-volatile memory, usually according to the actual functions of the any device with data processing capabilities where the device in the embodiment is located, other hardware may also be included, which will not be elaborated here.

[0113] For the implementation processes of the functions and roles of each unit in the above device, please refer to the implementation processes of the corresponding steps in the above method for details, which will not be elaborated here.

[0114] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0115] An embodiment of the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, a human body posture prediction method in the above embodiment is implemented.

[0116] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit of any device with data processing capabilities and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or will be output.

[0117] The above embodiments are only used to illustrate the design concept and features of the present invention, and their purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made according to the principles and design concepts disclosed by the present invention are within the protection scope of the present invention.

Claims

1. A human body posture prediction method, characterized in that: include: Obtain the motion acceleration and angular velocity of each part of the human body; The human body parts include the head and neck, spine, pelvis, left and right upper arms, left and right lower arms, left and right thighs, left and right calves, and left and right feet; Constructing the structure of the human body undirected spatiotemporal graph: The human body nodes are regarded as undirected spatiotemporal graph nodes. The edges of the undirected spatiotemporal graph include two types: (1) the connection between graph nodes in each time frame; (2) the connection between each graph node in multiple time frames; the human body nodes include cervical vertebrae, lumbar vertebrae, sacrum, left and right shoulder joints, left and right elbow joints, left and right hip joints, left and right knee joints, and left and right ankle joints; Based on the structure of the human body undirected spatiotemporal graph, an adjacency matrix of the human body undirected spatiotemporal graph is obtained; The motion acceleration, angular velocity and adjacency matrix of the human body's undirected space-time graph are input into the human body posture prediction algorithm model to predict the joint angles of the human body. Obtaining human body posture based on the predicted joint angles of each part of the human body; The human posture prediction algorithm model includes an analysis module, an interactive prediction module, and an optimization module; The analysis module is used to analyze key feature information from the motion acceleration and angular velocity of various parts of the human body; The interactive prediction module predicts the motion state through feature interaction in time and space; the input of the interactive prediction module includes the original input data X0, the output X3 of the parsing module and the undirected spatiotemporal graph adjacency matrix W, and the output X4 is obtained through the interactive prediction module processing; The interaction prediction module includes a spatial interaction module and a temporal interaction module; Among them, based on the undirected space-time graph, the spatial interaction module uses the spatial segmentation strategy to spatially divide the limb segments of the human body graph structure. For various spatial segmentation strategies, the original input data X0 is used as input, and the self-attention module is used to first solve the attention weights between the joints in each spatial domain, and then add and normalize them to obtain the inter-domain attention weights attn of the entire human body spatial domain. s ; At the same time, the spatial interaction module takes the original input data X0 as input, introduces the time window mechanism, slides on the time dimension, obtains the joint information of multiple time windows, takes the original input data X0 as the query amount, and uses the joint data of each time window as the comparison amount. The self-attention module is used to solve the attention weight between the original input data and the joint data of each time window, and then adds and normalizes it to obtain the attention weight attn of the entire time domain. t ; Attn s and attn t The two attention weights are aggregated to extract the weight information of spatiotemporal information. The adjacency matrix W of the undirected graph is updated using the weight information to obtain the adjacency matrix W_st that integrates spatiotemporal information. Finally, the updated adjacency matrix W_st is used to perform dimension fusion with the output X3 of the parsing module using the Einstein summation convention to obtain the output X′3; The temporal interaction module takes X′3 as input, uses dilated causal convolution and residual operations to interact in the temporal dimension, and obtains the output X4; The optimization module is used to optimize the predicted motion state data sequence.

2. A human body posture prediction method according to claim 1, characterized in that: The obtaining of the motion acceleration and angular velocity of each part of the human body includes: The positions where the human body wears inertial motion sensors include the head and neck, spine, pelvis, left and right upper arms, left and right lower arms, left and right thighs, left and right calves, and left and right feet; the inertial motion sensors are used to collect inertial motion data of various parts of the human body in the sensor coordinate system, and the inertial motion data includes inertial motion acceleration, angular velocity, and magnetic field strength; By using the human body posture calibration method, the inertial motion data of each part of the human body in the sensor coordinate system is converted to the human body coordinate system to obtain the motion acceleration and angular velocity of each part of the human body.

3. A human body posture prediction method according to claim 1, characterized in that: The analysis module includes two integrated encoder modules; the input of the first integrated encoder module is the original input data X0, i.e., the motion acceleration and angular velocity of various parts of the human body; the input of the second integrated encoder module is the residual combination of the original input data X0 and the output X1 of the first integrated encoder module; the output X3 of the analysis module is the residual combination of the original input data X0, the original input data X0 and the output X1 of the first integrated encoder module, and the output X2 of the second integrated encoder module; The comprehensive encoder module includes a spatial position encoding module, a temporal convolutional network module and an attention mechanism encoder module; the input of the spatial position encoding module is the original input data X0, the input of the temporal convolutional network module is the combination of the output of the spatial position encoding module and the input of the comprehensive encoder module, and the input of the attention mechanism encoder module is the output of the temporal convolutional network module; Among them, the temporal convolutional network module uses causal convolution and dilated convolution to process the time series task of the original input data X0; Among them, the spatial position encoding module uses absolute position encoding to generate an absolute position encoding matrix for the input data of the integrated encoder, and assigns a unique code to each position data in the input sequence; Among them, the attention mechanism encoder module includes multiple layers of encoders, each layer of encoders includes multi-head self-attention, feedforward neural network, normalization layer and connected in sequence; among them, multi-head self-attention is composed of multiple self-attention modules, that is, the input vector is passed into n self-attention modules respectively, n outputs are obtained, and then the outputs of multiple self-attention modules are combined, passed into the feedforward neural network and normalization layer, and the final output Z is obtained.

4. A human body posture prediction method according to claim 1, characterized in that: There are four spatial partitioning strategies: Spatial segmentation strategy 1: All limb segments in the human body graph structure are considered as one, without spatial segmentation; Spatial segmentation strategy 2: The human body structure is divided into two regions. The upper limb region includes the pelvis, L5 segment, L3 segment, T12 segment, T8 segment, head and neck, left and right hands, left and right forearms, left and right upper arms, and left and right shoulders; the lower limb region includes the pelvis, left and right thighs, left and right calves, and left and right feet. Spatial segmentation strategy three: Divide the human body structure into three areas: upper limbs and trunk, upper limbs and arms, and lower limbs; the upper limbs and trunk include the pelvis, L5 segment, L3 segment, T12 segment, T8 segment, and head and neck; the upper limbs and arms include left and right hands, left and right forearms, left and right upper arms, left and right shoulders, and T8 segment; the lower limbs include the pelvis, left and right thighs, left and right calves, and left and right feet; Spatial segmentation strategy 4: Divide all limb segments into five areas: upper limb trunk, upper limb left arm, upper limb right arm, lower limb left leg, lower limb right leg; the upper limb trunk includes pelvis, L5 segment, L3 segment, T12 segment, T8 segment, head and neck; the upper limb left arm includes left hand, left forearm, left upper arm, left shoulder, and T8 segment; the upper limb right arm includes right hand, right forearm, right upper arm, right shoulder, and T8 segment; the lower limb left leg includes pelvis, left thigh, left calf, and left foot; the lower limb right leg includes pelvis, right thigh, right calf, and right foot.

5. A human body posture prediction method according to claim 1, characterized in that: The input Y of the optimization module 0 It is the output X4 of the interactive prediction module; the optimization module is composed of several single-stage optimization modules and a time series model based on the attention mechanism; The single-stage optimization module includes a normal convolution layer, a dilated convolution layer, a batch normalization layer, and an activation function layer, which are connected in sequence; the convolution layer is used to adjust the input Y t-1 The feature dimension matches the number of feature maps in the network, using dilated convolutions and doubling the dilation rate in different single-stage optimization modules, and then using batch normalization and activation functions to obtain Finally, a residual connection is used to connect the input Y t-1 and Combine and get the output Y t .

6. A human body posture prediction method according to claim 5, characterized in that: The time series model based on the attention mechanism includes an encoder module and a decoder module; Use a recurrent neural network as an encoder, combined with the current time step input vector X i and the hidden state h at the previous time step i-1 , transform outputs the hidden state h of the current time step i , for the hidden layer state h of the encoder i Do weight distribution processing and use the hidden state h i The weight α ji , get the context vector c i ; The decoder uses a recurrent neural network and uses the output Y of the previous time step j-1 , encoder context vector c i and the hidden state s at the previous time step j-1 , determine the hidden state s of the current time step j ; Finally, based on the output Y of the previous time step j-1 , decoder hidden state vector s j , context vector c i , using a fully connected feedforward neural network to output the final predicted value Y j , predict the joint angles of various parts of the human body.

7. A human body posture prediction device, characterized in that: include: The motion data acquisition module is used to obtain the motion acceleration and angular velocity of various parts of the human body; The human body parts include the head and neck, spine, pelvis, left and right upper arms, left and right lower arms, left and right thighs, left and right calves, and left and right feet; The joint angle prediction module constructs an undirected spatiotemporal graph structure of the human body, and regards the human body nodes as undirected spatiotemporal graph nodes. The edges of the undirected spatiotemporal graph include two types: (1) connections between graph nodes in each time frame; (2) connections between each graph node in multiple time frames; the human body nodes include cervical vertebrae, lumbar vertebrae, sacrum, left and right shoulder joints, left and right elbow joints, left and right hip joints, left and right knee joints, and left and right ankle joints; based on the undirected spatiotemporal graph structure of the human body, the undirected spatiotemporal graph adjacency matrix of the human body is obtained; the motion acceleration, angular velocity, and undirected spatiotemporal graph adjacency matrix of each part of the human body are input into the human posture prediction algorithm model to obtain the joint angles of each part of the human body; The human posture prediction algorithm model includes an analysis module, an interactive prediction module, and an optimization module; The analysis module is used to analyze key feature information from the motion acceleration and angular velocity of various parts of the human body; The interactive prediction module predicts the motion state through feature interaction in time and space; the input of the interactive prediction module includes the original input data X0, the output X3 of the parsing module and the undirected spatiotemporal graph adjacency matrix W, and the output X4 is obtained through the interactive prediction module processing; The interaction prediction module includes a spatial interaction module and a temporal interaction module; Among them, based on the undirected space-time graph, the spatial interaction module uses the spatial segmentation strategy to spatially divide the limb segments of the human body graph structure. For various spatial segmentation strategies, the original input data X0 is used as input, and the self-attention module is used to first solve the attention weights between the joints in each spatial domain, and then add and normalize them to obtain the inter-domain attention weights attn of the entire human body spatial domain. s ; At the same time, the spatial interaction module takes the original input data X0 as input, introduces the time window mechanism, slides on the time dimension, obtains the joint information of multiple time windows, takes the original input data X0 as the query amount, and uses the joint data of each time window as the comparison amount. The self-attention module is used to solve the attention weight between the original input data and the joint data of each time window, and then adds and normalizes it to obtain the attention weight attn of the entire time domain. t ; Attn s and attn t The two attention weights are aggregated to extract the weight information of spatiotemporal information. The adjacency matrix W of the undirected graph is updated using the weight information to obtain the adjacency matrix W_st that integrates spatiotemporal information. Finally, the updated adjacency matrix W_st is used to perform dimension fusion with the output X3 of the parsing module using the Einstein summation convention to obtain the output X′3; The temporal interaction module takes X′3 as input, uses dilated causal convolution and residual operations to interact in the temporal dimension, and obtains the output X4; The optimization module is used to optimize the predicted motion state data sequence; The posture conversion module is used to obtain the human posture based on the joint angles of various parts of the human body; it also includes obtaining ergonomic characteristic parameters based on the joint angles of various parts of the human body to evaluate and determine the potential risk of musculoskeletal injuries.

8. A human body posture prediction device, characterized in that: It comprises one or more processors, and is used to implement a human body posture prediction method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, it is used to implement a human body posture prediction method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Three-dimensional human body posture estimation method fusing multi-scale spatial-temporal characteristics

    CN116229304A

  • Human motion posture prediction method based on adaptive space-time diagram convolution hybrid network

    CN117373126A