A human motion posture tracking method and device
By integrating thermal imaging infrared sensors and inertial measurement units on short-sleeved clothing and combining them with neural network models for multimodal information fusion, the problems of insufficient data accuracy and user comfort in smart sensing clothing are solved, and high-precision, low-cost, and low-privacy-threat human motion posture tracking is achieved.
Patent Information
- Application Number
- CN202411923572.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-12-25
AI Technical Summary
Existing smart sensing clothing has shortcomings in ensuring data accuracy, improving robustness and user comfort, and optical camera equipment is expensive, has high environmental requirements and poses a great threat to privacy.
By integrating thermal imaging infrared sensor devices and inertial measurement units on short-sleeved clothing, combined with residual network models, long short-term memory network models and multi-layer perceptron models, the movement posture of the human upper body is predicted through multimodal information fusion, reducing environmental dependence and privacy threats.
It achieves high-precision, low-interference, and low-cost human motion posture tracking, improves the stability and accuracy of posture determination, enhances user comfort, and reduces privacy threats and environmental requirements.
Smart Images

Figure CN119818058B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of human motion capture and perception technology, and particularly relates to a human motion posture tracking method and device. BACKGROUND
[0002] With the rapid development of technology, human motion capture and perception technology has been widely used in many fields, including virtual reality, augmented reality, motion analysis and rehabilitation training, intelligent monitoring and human-computer interaction, etc. Human motion capture and perception technology needs to be applied to motion capture systems, which usually rely on expensive optical camera equipment and complex site arrangement, and have many limitations, such as high cost, high requirements for the environment, and potential threats to user privacy.
[0003] In recent years, the application technology of smart sensing clothing integrates micro sensors into daily wear clothes, which can capture human motion information in real time. However, most of the application technology of smart sensing clothing on the current market still has a lot of room for improvement in ensuring data accuracy, improving the robustness of smart sensing clothing, and optimizing user comfort.
[0004] Therefore, the present application needs to find a human motion posture tracking method with low cost, low environmental requirements, low privacy threats, high posture determination accuracy, high device robustness, and high user comfort, and a human motion posture tracking device applying the method. SUMMARY
[0005] The present application provides a control method and device of a phase-shifted full-bridge converter to improve the accuracy of human motion posture determination, the robustness of the device, and the comfort of the user, while reducing the cost, the environmental requirements, and the privacy threats.
[0006] A first aspect of the present application provides a method for tracking human motion posture, which is applied to a human motion posture tracking device, the human motion posture tracking device comprising a short-sleeved garment and a control device; a plurality of thermal imaging infrared sensor devices and a plurality of inertial measurement units are provided on the short-sleeved garment; the plurality of thermal imaging infrared sensor devices sense a thermal image sequence of an elbow joint, and the plurality of inertial measurement units sense an initial motion information vector sequence of a non-elbow joint; the control device converts the thermal image sequence into an infrared optical flow vector sequence of the elbow joint, and converts the initial motion information vector sequence into a motion information vector sequence of the non-elbow joint; based on a first neural network model, the rotation angle vector of the elbow joint corresponding to the infrared optical flow vector sequence of the elbow joint is predicted; the rotation angle vector of the elbow joint and the motion information vector sequence of the non-elbow joint are spliced to obtain a multimodal information vector sequence; based on a second neural network model combination, the rotation angle vector of the non-elbow joint corresponding to the multimodal information vector sequence is predicted; and the motion posture of the upper body of the human body is determined according to the rotation angle vector of the elbow joint and the rotation angle vector of the non-elbow joint.
[0007] In some embodiments of the first aspect, the short-sleeved garment includes two sleeves, and the plurality of thermal imaging infrared sensor devices are respectively disposed on the two sleeves.
[0008] In some embodiments of the first aspect, the short-sleeved garment further includes two short-sleeved shoulder lines and two armpit covering points, a first connecting line is formed between the two short-sleeved shoulder lines, a second connecting line is formed between the two armpit covering points, and multiple inertial measurement units are respectively arranged at the midpoint between the first connecting line and the second connecting line, and at the hem of the short-sleeved garment near the pelvis.
[0009] In some embodiments of the first aspect, the initial motion information vector sequence includes an angular velocity vector sequence and an acceleration vector sequence, the angular velocity vector sequence includes angular velocity vectors sorted by time, and the acceleration vector sequence includes acceleration vectors sorted by time. The control device also converts the angular velocity vector into a rotational acceleration vector based on the interval time between adjacent angular velocity vectors; adjusts the acceleration vector based on the proportion of the acceleration vector in the rotational acceleration vector to obtain a target acceleration vector; and obtains a motion information vector sequence of the non-elbow joint based on the spliced target acceleration vector and angular velocity vector.
[0010] In some embodiments of the first aspect, the control device also obtains the weight of the acceleration vector based on the proportion of the acceleration vector in the rotational acceleration vector, the preset linear layer weights and parameters, wherein the proportion of the acceleration vector in the rotational acceleration vector is the ratio of the quadratic norm of the vector of the acceleration vector to the norm of the vector of the rotational acceleration vector; and adjusts the acceleration vector according to the weight of the acceleration vector.
[0011] In some embodiments of the first aspect, the control device further converts the thermal image sequence into an infrared optical flow vector sequence according to an optical flow algorithm; and aligns the time stamps of the infrared optical flow vector sequence of the elbow joint and the motion information vector sequence of the non-elbow joint.
[0012] In some embodiments of the first aspect, the first neural network model is trained by taking the infrared optical flow vector sequence of the elbow joint as input data and taking the optical image sequence of the elbow joint as label data, the optical image sequence of the elbow joint being used to indicate the true value of the rotation angle vector of the elbow joint; and the second neural network model is trained by taking the multi-modal information vector sequence as input data and taking the optical image sequence of the non-elbow joint as label data, the optical image sequence of the non-elbow joint being used to indicate the true value of the rotation angle vector of the non-elbow joint.
[0013] When the sum of the first loss value of the residual network model and the second loss value of the long short-term memory network model and the multi-layer perception model is within a preset range, the training of the residual network model, the long short-term memory network model and the multi-layer perception model is completed, wherein the first loss value is calculated according to the rotation angle vector of the elbow joint and the true value of the rotation angle vector of the elbow joint; and the second loss value is calculated according to the rotation angle vector of the non-elbow joint and the true value of the rotation angle vector of the non-elbow joint.
[0014] The second aspect of the present application provides a human motion posture tracking device, which comprises a short-sleeved garment and a control device; the short-sleeved garment is provided with a plurality of thermal imaging infrared sensing devices and a plurality of inertial measurement units; the plurality of thermal imaging infrared sensing devices are used to sense a thermal image sequence of an elbow joint, and the plurality of inertial measurement units are used to sense an initial motion information vector sequence of a non-elbow joint; the control device is used to convert the thermal image sequence into an infrared optical flow vector sequence of the elbow joint and convert the initial motion information vector sequence into a motion information vector sequence of the non-elbow joint; based on a first neural network model, a rotation angle vector of the elbow joint corresponding to the infrared optical flow vector sequence of the elbow joint is predicted; the rotation angle vector of the elbow joint and the motion information vector sequence of the non-elbow joint are spliced to obtain a multi-modal information vector sequence; based on a second neural network model combination, a rotation angle vector of the non-elbow joint corresponding to the multi-modal information vector sequence is predicted; and the motion posture of the upper body of the human body is determined according to the rotation angle vector of the elbow joint and the rotation angle vector of the non-elbow joint.
[0015] In some embodiments of the second aspect, the short-sleeved garment comprises two sleeves, and the plurality of thermal imaging infrared sensing devices are arranged on the two sleeves respectively.
[0016] In some embodiments of the second aspect, the short-sleeved garment further includes two short-sleeved shoulder lines and two armpit covering points, a first connecting line is formed between the two short-sleeved shoulder lines, a second connecting line is formed between the two armpit covering points, and multiple inertial measurement units are respectively arranged at the midpoint between the first connecting line and the second connecting line, as well as at the hem of the short-sleeved garment near the pelvis.
[0017] The human motion posture tracking method and device disclosed in this application utilize multiple thermal imaging infrared sensors and multiple inertial measurement units (IMUs) mounted on a short-sleeved garment to respectively sense a thermal image sequence of the elbow joint and an initial motion information vector sequence of non-elbow joints. A control device then integrates the two modal information—the infrared light flow sequence of the elbow joint derived from the thermal image sequence and the motion information sequence of non-elbow joints derived from the initial motion information vector sequence—to determine human motion posture based on a first neural network model and a second neural network model. This overcomes limitations such as ambient light interference caused by using a single modal information source and data drift caused by long-term motion posture determination, thereby improving the stability and accuracy of posture determination and the robustness of the device. Furthermore, posture determination can be performed while the user is wearing short-sleeved clothing, enhancing comfort. This method achieves high-precision, low-interference, low-cost, and low-privacy-threat (no optical camera equipment required) capture of upper body joint motion, while also requiring minimal environmental requirements (no complex site setup required). BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0019] Figure 1 A schematic diagram of the structure of a human body motion posture tracking device provided in an embodiment of the present application;
[0020] Figure 2 A schematic diagram of a flow chart of a method for tracking human motion posture provided in an embodiment of the present application;
[0021] Figure 3 A diagram illustrating an application scenario of the method for tracking human motion posture provided in an embodiment of the present application;
[0022] Figure 4 Another application scenario diagram of the human body motion posture tracking method provided in an embodiment of the present application;
[0023] Figure 5 Another schematic diagram of the process of tracking human body motion posture provided by an embodiment of the present application;
[0024] Figure 6 A diagram illustrating another application scenario of the human motion posture tracking method provided in an embodiment of the present application;
[0025] Figure 7 and Figure 8 This is another application scenario diagram of the human motion posture tracking method provided in an embodiment of the present application.
[0026] Figure 9 A schematic structural diagram of a control device provided in an embodiment of the present application.
[0027] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION
[0028] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0029] The terms "first", "second", etc. involved in this application are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features.
[0030] The following describes in detail the technical solution of the present application and how the technical solution of the present application solves the technical problem using specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0031] The data vector involved in this application refers to data in the form of a vector. For example, the infrared light flow vector refers to an infrared light flow in the form of a vector.
[0032] Please combine Figures 1 to 3 , Figure 1 Schematic diagram of a human body motion posture tracking device according to an embodiment of the present application. Figure 2 This is a flow chart of a method for tracking human motion posture according to an embodiment of the present application. Figure 3 Schematic diagram of a scenario of a human motion posture tracking method. The human motion posture tracking method of an embodiment of the present application is applied to a human motion posture tracking device.
[0033] like Figure 1As shown, the human motion posture tracking device 100 includes a control device 10 and a short-sleeved garment 20, on which a plurality of thermal imaging infrared sensor devices 30 and a plurality of inertial measurement units 40 are provided. The thermal imaging infrared sensor device 30 is used to sense the thermal image of the elbow joint and generate a thermal image sequence. The inertial measurement unit 40 is used to sense the initial motion information of the non-elbow joint and generate an initial motion information sequence. The two modal information are provided to the control device 10 to determine the motion posture of the upper body of the human body. In some embodiments, the control device 10 can be a laptop computer, a mobile phone, a server, etc.
[0034] In some embodiments, a short-sleeved garment 20 includes two sleeves 21, each of which is provided with two thermal imaging infrared sensor devices 30. The camera of the thermal imaging infrared sensor device 30 on one sleeve 21 is oriented toward the left elbow joint to sense a thermal image of the left elbow joint. The camera of the thermal imaging infrared sensor device 30 on the other sleeve 21 is oriented toward the right elbow joint to sense a thermal image of the right elbow joint.
[0035] It can be understood that the embodiment of the present application provides two thermal imaging infrared sensor devices 30 on the two sleeves 21 respectively, so that when the human arm performs various movements, the thermal imaging infrared sensor device 30 can capture the thermal image of the elbow joint, and the thermal image indicates the movement status of the elbow joint.
[0036] In some embodiments, the short-sleeved garment 20 further includes two short-sleeved shoulder lines 22 and two armpit coverage points 23. A first connecting line 24 is formed between the two short-sleeved shoulder lines 22, and a second connecting line 25 is formed between the two armpit coverage points 23. An inertial measurement unit 40 is disposed at the midpoint 26 between the first connecting line 24 and the second connecting line 25. Another inertial measurement unit 40 is disposed at the hem of the short-sleeved garment 20 near the pelvis.
[0037] It can be understood that the midpoint 26 between the first connecting line 24 between the two short-sleeved shoulder lines 22, the second connecting line 25 between the two armpit-covering points 23, and the hem of the short-sleeved garment 20 has a stronger load-bearing capacity, intersecting at the sleeve 21, and does not conflict with the placement of the thermal imaging infrared sensor device 30. Therefore, in this embodiment of the present application, by disposing an inertial measurement unit 40 at the midpoint 26 and the hem of the short-sleeved garment 20, the inertial measurement unit 40 can stably sense the motion information of non-elbow joints (i.e., other upper body joints other than the elbow joints) when the human body performs various movements, thereby reflecting the motion status of the non-elbow joints of the human body.
[0038] As can be understood, compared to related technologies, which require users to wear tight-fitting clothing (with integrated sensors) or tight-fitting sensing equipment for posture tracking, this application installs a thermal imaging infrared sensor device 30 and an inertial measurement unit 40 on a short-sleeved garment 20, achieving for the first time the deployment of multimodal sensors on short-sleeved clothing and tracking of upper-body motion. This allows users to exercise freely while wearing short-sleeved clothing, improving user comfort. The absence of optical camera equipment reduces the threat of privacy leakage.
[0039] It should be noted that Figure 1 This is only a schematic diagram of a structure provided in the embodiment of the present application. Figure 1 The actual form of the various devices included in the Figure 1 The interaction mode between devices is limited, and in the specific application of the solution, it can be set according to actual needs.
[0040] like Figure 2 As shown, the human body motion posture tracking method may include the following steps:
[0041] Step S110: a plurality of thermal imaging infrared sensor devices sense a thermal image sequence of the elbow joint.
[0042] The thermal image sequence of the elbow joints includes a thermal image sequence of the left elbow joint and a thermal image sequence of the right elbow joint. The thermal image sequence refers to thermal images sorted by time.
[0043] Step S120: Multiple inertial measurement units sense initial motion information vector sequences of non-elbow joints.
[0044] The expression of the initial motion information vector sequence can be The initial motion information vector sequence refers to the motion information vectors sorted by time, through which the real-time motion status of the non-elbow joint can be understood.
[0045] Step S130: The control device converts the thermal image sequence into an infrared light flow vector sequence of the elbow joint, and converts the initial motion information vector sequence into a motion information vector sequence of the non-elbow joint.
[0046] Specifically, the control device 10 obtains the thermal image sequence of the elbow joint sent by the thermal imaging infrared sensor device 30, and converts the thermal image sequence into an infrared optical flow vector sequence according to the optical flow algorithm. The expression of the infrared optical flow vector sequence can be , represented by the thermal image sequence of the left elbow joint , and thermal image sequences of the right elbow joint Splicing to get infrared optical flow vector sequence .
[0047] The optical flow algorithm can be a Horn-Schunck algorithm. The sequence of infrared optical flow vectors refers to infrared optical flow vectors sorted by time, each infrared optical flow vector including the movement acceleration, movement angle and movement direction of each pixel point of the elbow joint in the thermographic image, so that through the sequence of infrared optical flow vectors, the real-time movement state of the elbow joint can be understood.
[0048] Then, the sequence of initial movement information vectors of non-elbow joints sent by the inertial measurement unit 40 is acquired, and the sequence of initial movement information vectors of non-elbow joints is optimized to obtain the sequence of movement information vectors of non-elbow joints.
[0049] Finally, after aligning the time stamps of the sequence of infrared optical flow vectors of the elbow joint and the sequence of movement information vectors of non-elbow joints, step S140 is performed.
[0050] Step S140: The control device predicts, based on a first neural network model, a rotation angle vector of the elbow joint corresponding to the sequence of infrared optical flow vectors of the elbow joint.
[0051] The expression of the rotation angle vector of the elbow joint can be .
[0052] Step S150: The control device splices the rotation angle vector of the elbow joint and the sequence of movement information vectors of non-elbow joints to obtain a sequence of multi-modal information vectors.
[0053] After obtaining the rotation angle vector of the elbow joint, the rotation angle vector of the elbow joint and the sequence of movement information vectors of non-elbow joints are spliced to obtain a sequence of multi-modal information vectors. That is, the sequence of multi-modal information vectors includes two types of information, i.e., the rotation angle vector of the infrared elbow joint obtained based on the thermographic infrared sensing device 30 and the sequence of movement information vectors of non-elbow joints obtained based on the inertial measurement unit 40.
[0054] Step S160: The control device predicts, based on a second neural network model combination, a rotation angle vector of non-elbow joints corresponding to the sequence of multi-modal information vectors.
[0055] The expression of the rotation angle vector of non-elbow joints can be .
[0056] Step S170: The control device determines the movement posture of the upper body of the human body according to the rotation angle vector of the elbow joint and the rotation angle vector of non-elbow joints.
[0057] Specifically, the rotation angle vector of the elbow joint and the rotation angle vector of the non-elbow joint are spliced to obtain the movement posture vector of the upper body of the human body, thereby predicting the movement posture of the upper body of the human body in real time.
[0058] The expression of the upper body motion posture vector can be ,in Represents the motion posture vector of the upper body.
[0059] It is understandable that related technologies include motion capture systems based on single sensor applications, such as inertial measurement units (IMUs) or visual sensors. Due to the single type of sensor and the single type of sensing data obtained, various limitations arise, resulting in low accuracy in determining motion states and insufficient robustness of the motion capture system. Limitations include: motion capture systems based on IMUs are susceptible to noise interference and data drift. Motion capture systems based on visual sensors are susceptible to environmental and obstructive factors when sensing data, resulting in inaccurate data acquisition. Furthermore, when visual sensors are placed on clothing, there is also the issue of privacy leakage.
[0060] Compared to the use of a single sensor in related art, the above-described technical solution utilizes multiple thermal imaging infrared sensing devices 30 and multiple inertial measurement units 40 installed on a short-sleeved garment 20 to respectively sense a thermal image sequence of the elbow joint and an initial motion information vector sequence of non-elbow joints. Subsequently, a control device 10 integrates the two modal information components: the infrared light flow sequence of the elbow joint obtained from the thermal image sequence and the motion information sequence of the non-elbow joint obtained from the initial motion information vector sequence. Based on the first and second neural network models, human motion posture is determined. This overcomes limitations such as ambient light interference caused by using a single modal information source and data drift caused by long-term motion posture determination, thereby improving the stability and accuracy of posture determination and enhancing the robustness of the device. Furthermore, posture determination can be performed while the user is wearing the short-sleeved garment 20, enhancing comfort. This system achieves high-precision, low-interference, low-cost, and low-privacy-threat (no optical camera equipment required) capture of upper body joint motion, while also requiring minimal environmental requirements (no complex site setup required).
[0061] In some embodiments, the human body motion posture tracking method further includes controlling the device to construct a combination of the first neural network model and the second neural network model. Figure 4 First, the control device obtains training data as a training set, and the training data includes input data and label data. That is, the human body motion posture tracking 100 also includes an optical device 50, and the optical device 50 includes a video camera or a digital camera.
[0062] The control device 10 obtains an optical image sequence of the elbow joint and an optical image sequence of the non-elbow joint through the optical device 50. An optical image sequence refers to an optical image sorted by time. After processing, the optical image sequence of the elbow joint can indicate the true value of the rotation angle vector of the elbow joint, and after processing, the optical image sequence of the non-elbow joint can indicate the true value of the rotation angle vector of the non-elbow joint. The optical image sequence of the elbow joint and the optical image sequence of the non-elbow joint can be used as label data. At the same time, the control device 10 retrieves the infrared light flow vector sequence of the elbow joint and the motion information vector sequence of the non-elbow joint.
[0063] Secondly, the control device 10 uses the infrared light flow vector sequence of the elbow joint as input data and the optical image sequence of the elbow joint as label data to train the residual network model (Residual Network, ResNet). The residual network model is a deep convolutional neural network model. By introducing jump connections (residual blocks), it solves the problem of gradient disappearance in deep networks, thereby enabling more efficient training of neural networks with many layers and that need to learn complex features.
[0064] During training, the residual network model outputs the elbow joint's rotation angle vector. The control device 10 then concatenates the elbow joint's rotation angle vector with the non-elbow joint's motion information vector sequence to generate a multimodal information vector sequence. The control device 10 uses the multimodal information vector sequence as input data and the non-elbow joint's optical image sequence as label data to train a Long Short-Term Memory (LSTM) network model and a Multilayer Perceptron (MLP) model. The LSTM network model is a special recursive neural network model that utilizes a gating mechanism to effectively capture long-term dependencies in sequence data, solving the problems of vanishing and exploding gradients in traditional recurrent neural networks (RNNs).
[0065] It is understood that the infrared optical flow vector sequence and multimodal information vector sequence of the elbow joint during the training phase are used as training data. The infrared optical flow vector sequence of the elbow joint generated in step S130 and the multimodal information vector sequence generated in step S150 are used as prediction objects and can be generated in the training phase or newly generated.
[0066] Until the total loss value of the total loss function of the residual network model, the long short-term memory network model, and the multi-layer perceptron model falls within a preset range, a first neural network model is constructed from the residual network model, and a second neural network model combination is constructed from the long short-term memory network model and the multi-layer perceptron model. That is, when the sum of the first loss value of the residual network model and the second loss values of the long short-term memory network model and the multi-layer perceptron model falls within a preset range, the training of the first neural network model and the second neural network model combination is completed.
[0067] Exemplarily, the formula for determining the first loss value of the residual network model is as follows:
[0068]
[0069] in, represents the first loss value, represents the rotation angle vector of the elbow joint, The expression can be , represents the rotation angle vector of the left elbow joint, represents the rotation angle vector of the right elbow joint, The rotation angle vector has six dimensions. That is, the rotation angle vector of the elbow joint is , the rotation angle vector of the left elbow joint and the rotation angle vector of the right elbow joint Data splicing is performed. Represents the true value of the rotation angle vector of the elbow joint. For specific expressions, please refer to .
[0070] It can be understood that the rotation angle vector of the elbow joint during the training phase is The true value of the rotation angle vector of the elbow joint The average of the squares of the differences is the first loss value .
[0071] Exemplarily, the formula for determining the second loss value of the long short-term memory network model and the multi-layer perceptron model is as follows:
[0072]
[0073] in, represents the second loss value, Represents the rotation angle vector of the non-elbow joints during the training phase.
[0074] The expression for can be, , represents the rotation angle vector of the hip joint, a rotation angle vector of the cervical vertebrae, a rotation angle vector of the cervical vertebrae, a rotation angle vector of the elbow joint. The rotation angle vector has fifty-four dimensions. It can be understood that the rotation angle vector of the non-elbow joint is obtained by data splicing from the other human upper body joint vector except the elbow joint. The true value of the rotation angle vector of the non-elbow joint.
[0075] It can be understood that the rotation angle vector of the non-elbow joint in the training stage The average of the square of the difference between the true value of the rotation angle vector of the non-elbow joint is the second loss value .
[0076] The determination formula of the first loss value of the residual network model and the sum of the second loss values of the long short-term memory network model and the multi-layer perception model is as follows:
[0077]
[0078] wherein, the total loss value, the total loss value is the sum of the first loss value of the residual network model and the second loss value of the long short-term memory network model and the multi-layer perception model .
[0079] After completing the construction of the second neural network model combination and the first neural network model, the infrared optical flow vector sequence of the elbow joint obtained in step S130 is input into the first neural network model, and the rotation angle vector of the elbow joint is output, so as to predict the real-time motion state of the elbow joint. At the same time, the multi-modal information vector sequence obtained in step S150 is input into the second neural network model combination, and the rotation angle vector of the non-elbow joint is output, so as to predict the real-time motion state of the non-elbow joint.
[0080] It can be understood that, compared with other traditional machine learning algorithms, the residual network model as a convolutional neural network model has good performance in the task of processing the infrared optical flow vector sequence of the elbow joint.
[0081] Understandably, since in real-world scenarios, clothing can wrinkle and deform as the human body moves, mapping partial body images to human motion postures is extremely challenging. Therefore, the present application utilizes deep learning technologies such as residual network models, long short-term memory network models, and multi-layer perceptron models to achieve a mapping of the thermal image sequence of the elbow joint sensed by the thermal imaging infrared sensor device and the initial motion information vector sequence of the non-elbow joint sensed by the inertial measurement unit to human motion postures, thereby solving the problem of mapping body images to human motion postures.
[0082] Since the data structures and feature dimensions sensed by the thermal imaging infrared sensor device and the inertial measurement unit are different, it is difficult to directly splice the two types of data and input them into the neural network for prediction. Therefore, in an embodiment of the present application, the residual network model is first used to predict the rotation angle vector of the elbow joint corresponding to the infrared optical flow vector sequence of the elbow joint. Then, through a combination of a long short-term memory network model and a multi-layer perceptron model, the multimodal information vector sequence obtained by splicing the rotation angle vector of the elbow joint and the motion information vector sequence of the non-elbow joint is predicted, and the corresponding rotation angle vector of the non-elbow joint is predicted. This solves the problem that the two types of data cannot be directly spliced and simultaneously input into the neural network for prediction.
[0083] Furthermore, if the thermal imaging infrared sensor and inertial measurement unit were to be directly used to predict their respective joint positions, they would be unable to fully utilize the correlation between the user's different joint movements to further improve prediction results. Therefore, the present invention has designed a multimodal feature fusion framework based on a combination of a residual network model, a long short-term memory network model, and a multilayer perceptron model to fully utilize the correlation between the user's different joint movements and further improve prediction results.
[0084] In some embodiments, see Figure 5 and 6 The human body motion posture tracking method also includes the process of the control device using an attention mechanism model to optimize the initial motion information vector sequence of the non-elbow joint to obtain the motion information vector sequence of the non-elbow joint. Figure 5 As shown, the method further includes the following steps:
[0085] Step S210: The control device converts the angular velocity vector into a rotational acceleration vector according to the interval between adjacent angular velocity vectors.
[0086] The initial motion information vector sequence of the non-elbow joints obtained includes an angular velocity vector sequence and an acceleration vector sequence. Figure 6 As shown, the initial motion information vector sequence includes initial motion information vectors sorted in time , the angular velocity vector sequence includes the angular velocity vectors sorted by time , the acceleration vector sequence includes the acceleration vectors sorted by time That is, the inertial measurement unit 40 includes an accelerometer and a gyroscope. The accelerometer is used to sense an acceleration vector sequence, and the gyroscope is used to sense an angular velocity vector sequence.
[0087] like Figure 6 As shown, according to the angular velocity vector Get the rotation acceleration vector , specifically, the formula for determining the rotational acceleration vector is as follows:
[0088]
[0089] in, represents the rotation acceleration vector, represents the angular velocity vector, represents another adjacent angular velocity vector. That is, the rotational acceleration vector can be determined based on the interval time between adjacent angular velocity vectors (the interval time can be a unit time) and the adjacent angular velocity vectors. In one embodiment, t+1 represents the current time, and t represents the previous time.
[0090] Step S220: The control device adjusts the acceleration vector based on the proportion of the acceleration vector in the rotation acceleration vector to obtain a target acceleration vector.
[0091] like Figure 6 As shown, the rotation acceleration vector and the acceleration vector The input acceleration vector ratio determines the model and the output acceleration vector The rotational acceleration vector The proportion of .
[0092] The acceleration vector ratio model is as follows:
[0093]
[0094] It can be understood that the quadratic norm of the acceleration vector is Norm of the vector with the rotational acceleration vector The ratio is the proportion of the acceleration vector to the rotation acceleration vector .
[0095] Next, the acceleration vector In the acceleration vector The proportion of , input acceleration vector weight to determine the model, output acceleration vector weight , the weight of the acceleration vector is the attention matrix of the attention mechanism model. The acceleration vector weight determination model is as follows:
[0096]
[0097] in, represents the linear layer weights, , Indicates the parameters, which are preset. That is, according to the proportion of the acceleration vector in the rotation acceleration vector , the preset linear layer weights and parameters Get the weight of the acceleration vector It can be understood that when the human body is assumed to be a rigid body, by increasing the attention matrix This linear layer (weight of the acceleration vector) allows the residual network model, long short-term memory network model, and multi-layer perceptron model to have a certain degree of freedom to fine-tune the weights during training, making the constructed combination of the first neural network model and the second neural network model more adaptable to the real world.
[0098] Next, according to the weight of the acceleration vector Adjust the acceleration vector , get the target acceleration vector .
[0099] It can be understood that the embodiment of the present application is based on the attention mechanism model, according to the acceleration vector The rotational acceleration vector The proportion of , determine whether the current motion is dominated by the rotational acceleration caused by rotation (that is, the larger the proportion, the less it is dominated by rotational acceleration). If not, give the acceleration vector of the accelerometer of the current frame (that is, the current moment) a larger weight for regression. If so, give it a smaller weight.
[0100] Step S230: The control device obtains a motion information vector sequence of the non-elbow joint based on the spliced target acceleration vector and angular velocity vector.
[0101] Specifically, if Figure 6 As shown, the target acceleration vector and the angular velocity vector Data splicing , get the motion information vector of the non-elbow joint In this way, based on the motion information vectors of multiple non-elbow joints, a motion information vector sequence of the non-elbow joints is obtained.
[0102] It can be understood that the process of using the attention mechanism model to optimize the initial motion information vector sequence of the non-elbow joint to obtain the motion information vector sequence of the non-elbow joint is suitable for both the training stage and the prediction stage.
[0103] It's understandable that due to differences in user size and user movement during data collection, the IMU produces varying displacements and records noise, resulting in reduced accuracy in tracking human motion. Assuming the human body is rigid, the IMU records translational acceleration, acceleration relative to gravity, and rotational acceleration. Of these three components, the first two are unaffected by wearer position; only the rotational acceleration caused by rotation is sensitive to the IMU's displacement within the user's body. Therefore, the greater the proportion of the accelerometer's acceleration vector to the gyroscope's rotational acceleration vector (i.e., the greater the proportion of the accelerometer's acceleration vector), the greater the weight assigned to the acceleration vector, making the target acceleration vector larger. This weakens the gyroscope's rotational acceleration vector and reduces the impact of the IMU's displacement within the user's body. This allows the combined first and second neural network models to be more robust to IMU offsets while retaining most relevant information, thereby reducing noise from IMUs worn by different users.
[0104] Understandably, multimodal motion monitoring solutions exist in related technologies that fuse data from multiple sensors. These typically utilize multiple sensing devices, such as accelerometers, gyroscopes, and magnetometers, combined with data fusion algorithms (such as Kalman filters) to estimate human posture and motion trajectory. These solutions further integrate data from visual sensors, pressure sensors, or depth cameras, using complex algorithms to achieve more comprehensive motion capture. Deep learning techniques are also used to optimize model performance to accommodate different user body types and motion styles.
[0105] However, the data fusion algorithms used in multimodal motion monitoring solutions still face challenges such as long-term monitoring data, which leads to reduced tracking accuracy and strong environmental dependence. In the embodiments of this application, an attention mechanism model is used to calibrate the noise of the inertial measurement unit data in real time, and a streamer algorithm is used to calibrate the data of the thermal imaging infrared sensor device in real time, thereby enhancing the robustness of the combination of the first neural network model and the second neural network model for different user wearing scenarios.
[0106] See also Figure 7 and Figure 8 The above embodiments are further described through the following application scenario. The human body motion posture tracking device 100 performs the following steps:
[0107] a1: a sequence of thermal images of the left and right elbow joints sensed by the thermal imaging infrared sensor device 30, a sequence of initial motion information vectors of the non-elbow joints sensed by the inertial measurement unit 40, and a sequence of optical images of the elbow joints and non-elbow joints sensed by the optical device 50;
[0108] a2: The control device 10 converts the thermal image sequence into an infrared optical flow vector sequence according to the optical flow algorithm, and converts the initial motion information vector sequence of the non-elbow joint into a motion information vector sequence of the non-elbow joint according to the attention mechanism model;
[0109] a3: The control device 10 uses the infrared light flow vector sequence of the elbow joint as input data and the optical image sequence of the elbow joint as label data to train the residual network model, and outputs the rotation angle vector of the elbow joint and calculates the first loss value;
[0110] a4: The control device 10 combines the rotation angle vector of the elbow joint and the motion information vector sequence of the non-elbow joint to obtain a multimodal information vector sequence;
[0111] a5: The control device 10 uses the multimodal information vector sequence as input data and the optical image sequence of the non-elbow joint as label data to train the long short-term memory network model and the multilayer perceptron model, and outputs the rotation angle vector of the elbow joint and calculates the second loss value;
[0112] a6: When the sum of the first loss value and the second loss value is within a preset range, the control device 10 constructs a combination of the first neural network model and the second neural network model;
[0113] a7: The control device 10 obtains a real-time infrared optical flow vector sequence and an initial motion information sequence, inputs the real-time infrared optical flow vector sequence and the initial motion information sequence into a combination of a first neural network model and a second neural network model, and outputs the motion posture of the upper body of the human body.
[0114] The human motion posture tracking device 100 further comprises an ESP32 circuit board 60 and a battery (not shown in the figure). The infrared thermal imaging sensor device 30 is produced by Heimann Company, and the model is HTPA80x864dR2L3.9 / 0.8. The height of each infrared thermal imaging sensor device 30 is 12.2mm, the diameter is 20mm, the field of view angle is 120x90 degrees, the resolution is 64x80, and the frame rate is 20FPS. Each infrared thermal imaging sensor device 30 is connected to the respective ESP32 circuit board 60 (specification model is ESP-WROOM-32) through a wire, and is powered by a 5V, 900mAh battery (which can last for 2.45 hours). The infrared thermal imaging sensor device 30 sends and saves the measurement data to the control device 10 in the data acquisition field through WiFi. The inertial measurement unit 40 for obtaining the initial motion information vector sequence of the non-elbow joint is produced by Movella Company, and the model is Xsens DOT. The length and width of each inertial measurement unit 40 is 32mmx20mm, and the thickness is 10mm. The data acquisition frame rate is 60FPS, and is equipped with a 45mAh LIR2032 rechargeable button cell battery (which can last for 6 hours) and a Bluetooth module (not shown in the figure) that can send measurement data to the control device 10.
[0115] Figure 9 The control device provided in the present application is shown in the structural schematic diagram. As shown in the figure, Figure 9 the control device 10 comprises:
[0116] a processor 11, a memory 12 and a bus 13;
[0117] The memory 12 is used for storing the computer program code of the processor 11.
[0118] The processor 11 is configured to execute the technical solutions of the human motion posture tracking method in any one of the preceding method embodiments by executing the computer program code.
[0119] Optionally, the memory 12 can be independent or integrated with the processor 11.
[0120] The memory 12 is connected with the processor 11 through the bus 13 and completes the communication between each other.
[0121] Optionally, the memory 12 can contain a random access memory (RAM), and can also include a non-volatile memory, such as at least one disk memory.
[0122] Bus 13 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, and control buses. For ease of illustration, the figure uses only one thick line, but this does not imply that there is only one bus or only one type of bus.
[0123] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components.
[0124] The control device 10 is used to execute the technical solution provided in any of the aforementioned method embodiments. Its implementation principles and technical effects are similar and will not be repeated here.
[0125] Those skilled in the art will appreciate that all or part of the steps in the above-described method embodiments can be implemented using hardware associated with program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0126] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for tracking human body motion posture, applied to a human body motion posture tracking device, characterized in that: The human motion posture tracking device includes a short-sleeved garment and a control device; the short-sleeved garment is provided with a plurality of thermal imaging infrared sensor devices and a plurality of inertial measurement units; the plurality of thermal imaging infrared sensor devices sense a thermal image sequence of an elbow joint, and the plurality of inertial measurement units sense an initial motion information vector sequence of a non-elbow joint; the short-sleeved garment further includes two short-sleeved shoulder lines and two armpit covering points, a first connecting line is formed between the two short-sleeved shoulder lines, and a second connecting line is formed between the two armpit covering points, and the plurality of inertial measurement units are respectively provided at the midpoint between the first connecting line and the second connecting line, and at the hem of the short-sleeved garment near the pelvis; The control device converts the thermal image sequence into an infrared light flow vector sequence of the elbow joint, and converts the initial motion information vector sequence into a motion information vector sequence of the non-elbow joint; Predicting, based on the first neural network model, a rotation angle vector of the elbow joint corresponding to the infrared optical flow vector sequence of the elbow joint; splicing the rotation angle vector of the elbow joint and the motion information vector sequence of the non-elbow joint to obtain a multimodal information vector sequence; Predicting, based on the second neural network model combination, a rotation angle vector of a non-elbow joint corresponding to the multimodal information vector sequence; Determining a movement posture of the upper body of the human body according to the rotation angle vector of the elbow joint and the rotation angle vector of the non-elbow joint; The initial motion information vector sequence includes an angular velocity vector sequence and an acceleration vector sequence, wherein the angular velocity vector sequence includes angular velocity vectors sorted by time, and the acceleration vector sequence includes acceleration vectors sorted by time. The control device further: converting the angular velocity vector into a rotational acceleration vector according to the interval time between adjacent angular velocity vectors; Adjusting the acceleration vector based on the proportion of the acceleration vector in the rotation acceleration vector to obtain a target acceleration vector; Based on the spliced target acceleration vector and the angular velocity vector, a motion information vector sequence of the non-elbow joint is obtained.
2. The method according to claim 1, characterized in that The short-sleeved clothing includes two sleeves, and the multiple thermal imaging infrared sensor devices are respectively arranged on the two sleeves.
3. The method according to claim 1, characterized in that The control device also The weight of the acceleration vector is obtained according to the proportion of the acceleration vector in the rotation acceleration vector, the preset linear layer weight and the parameter, wherein the proportion of the acceleration vector in the rotation acceleration vector is the ratio of the quadratic norm of the vector of the acceleration vector to the norm of the vector of the rotation acceleration vector; The acceleration vector is adjusted according to a weight of the acceleration vector.
4. The method according to claim 1, wherein The control device also Converting the thermal image sequence into the infrared optical flow vector sequence according to an optical flow algorithm; Align the timestamps of the infrared optical flow vector sequence of the elbow joint and the motion information vector sequence of the non-elbow joint.
5. The method according to claim 1, wherein The first neural network model is obtained by training a residual network model using the infrared optical flow vector sequence of the elbow joint as input data and the optical image sequence of the elbow joint as label data, wherein the optical image sequence of the elbow joint is used to indicate the true value of the rotation angle vector of the elbow joint; The second neural network model combination is obtained by training a long short-term memory network model and a multi-layer perceptron model using the multimodal information vector sequence as input data and an optical image sequence of the non-elbow joint as label data, wherein the optical image sequence of the non-elbow joint indicates a true value of the rotation angle vector of the non-elbow joint; When the sum of the first loss value of the residual network model and the second loss values of the long short-term memory network model and the multi-layer perceptron model is within a preset range, the training of the residual network model and the long short-term memory network model and the multi-layer perceptron model is completed, wherein the first loss value is calculated based on the rotation angle vector of the elbow joint and the true value of the rotation angle vector of the elbow joint; the second loss value is calculated based on the rotation angle vector of the non-elbow joint and the true value of the rotation angle vector of the non-elbow joint.
6. A human body motion posture tracking device, characterized in that: The human motion posture tracking device includes a short-sleeved garment and a control device; the short-sleeved garment is provided with a plurality of thermal imaging infrared sensor devices and a plurality of inertial measurement units; the plurality of thermal imaging infrared sensor devices are used to sense a thermal image sequence of an elbow joint, and the plurality of inertial measurement units are used to sense an initial motion information vector sequence of a non-elbow joint; the short-sleeved garment further includes two short-sleeved shoulder lines and two armpit covering points, a first connecting line is formed between the two short-sleeved shoulder lines, and a second connecting line is formed between the two armpit covering points, and the plurality of inertial measurement units are respectively provided at the midpoint between the first connecting line and the second connecting line, and at the hem of the short-sleeved garment near the pelvis; The control device is used to convert the thermal image sequence into an infrared light flow vector sequence of the elbow joint, and convert the initial motion information vector sequence into a motion information vector sequence of the non-elbow joint; Predicting, based on the first neural network model, a rotation angle vector of the elbow joint corresponding to the infrared optical flow vector sequence of the elbow joint; splicing the rotation angle vector of the elbow joint and the motion information vector sequence of the non-elbow joint to obtain a multimodal information vector sequence; Predicting, based on the second neural network model combination, a rotation angle vector of a non-elbow joint corresponding to the multimodal information vector sequence; Determining a movement posture of the upper body of the human body according to the rotation angle vector of the elbow joint and the rotation angle vector of the non-elbow joint; The initial motion information vector sequence includes an angular velocity vector sequence and an acceleration vector sequence, wherein the angular velocity vector sequence includes angular velocity vectors sorted by time, and the acceleration vector sequence includes acceleration vectors sorted by time. The control device further: converting the angular velocity vector into a rotational acceleration vector according to the interval time between adjacent angular velocity vectors; Adjusting the acceleration vector based on the proportion of the acceleration vector in the rotation acceleration vector to obtain a target acceleration vector; Based on the spliced target acceleration vector and the angular velocity vector, a motion information vector sequence of the non-elbow joint is obtained.
7. The human body motion posture tracking device according to claim 6, characterized in that: The short-sleeved clothing includes two sleeves, and the multiple thermal imaging infrared sensor devices are respectively arranged on the two sleeves.
Citation Information
Patent Citations
Upper body posture reconstruction method and device, electronic equipment and storage medium
CN114092639A
Hand action detection method and device, storage medium and terminal
CN118781662A