Human Motion Capture Method, System and Storage Medium
By using UWB technology to communicate between the base station and the converged sensor, and combining with the IMU module to acquire motion attitude data, the problems of poor anti-interference ability and low positioning accuracy of data transmission in inertial motion capture technology are solved, and high-precision real-time position and motion attitude calculation are achieved, and attitude drift is suppressed.
Patent Information
- Application Number
- CN202410914722.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-09
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-07-09
AI Technical Summary
The existing inertial motion capture technology has problems such as poor anti-interference ability of data transmission, close transmission distance, low bandwidth, large latency and inability to accurately determine the absolute position of the IMU, which has affected the accuracy of positioning and motion estimation.
The integrated sensor is adopted, including the UWB module and the IMU module, and the data packet is transmitted through UWB technology. The base station receives and adds a timestamp and forwards it to the computer device. The computer device uses channel impulse response and position information to determine the target data packet, thereby determining the real-time position and motion posture of the target user.
It enhances the system's signal anti-interference, reduces sensor data acquisition delay, reduces system power consumption, improves system bandwidth, significantly improves the calculation accuracy of real-time position and motion posture, and suppresses the attitude drift of the IMU module under long-term use.
Smart Images

Figure CN118870291B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human motion capture, and in particular, to a human motion capture method, system, and storage medium. Background Art
[0002] Human motion capture technology uses different sensor devices to measure, track, and record the motion information (including speed, position, posture, etc.) of the human body limbs in three-dimensional space, and then uses algorithms such as inverse kinematics analysis to reconstruct the human motion process. This technology has been widely applied in interdisciplinary fields such as film and animation production, human-computer interaction, and sports training, and has broad market prospects and application space.
[0003] Currently, the mainstream high-precision motion capture technologies include optical motion capture and inertial motion capture. Inertial motion capture has the advantages of low cost, convenient installation, and easy use compared with optical motion capture, and can be used in indoor or outdoor scenes with complex optical environments. The principle of inertial motion capture technology is mainly: by fixing multiple wearable inertial sensors (i.e., IMUs) on each human limb to collect the acceleration and angular velocity information of each limb, and then separately calculating the posture of each limb, and then constructing a complete human motion through a set of constraint algorithms. Wearable inertial motion capture devices have no occlusion and specific space limitations, and can monitor the activities of users in any weather and venue, so they have a wide range of application scenarios in the fields of sports health and rehabilitation medicine.
[0004] However, the existing wireless IMU data transmission is generally through WIFI, Bluetooth, or Zigbee, etc., which has problems such as poor anti-interference ability, short transmission distance, low bandwidth, and large latency, and cannot accurately determine the absolute position of the IMU. For inertial motion capture, three-dimensional space positioning is an important function second only to human posture solution. Inertial motion capture relies on the double integration of accelerometer data to calculate displacement, which will cause drift over time. Especially in the case of long-term use, this will affect the accuracy of positioning and motion estimation. In addition, the existing IMU-based motion capture technology often needs to wear more than ten IMUs on the human body to achieve full-body motion capture. At the same time, the wearing error of the IMU will also cause errors in human motion restoration. The motion capture is difficult, and the configuration and wearing time of the IMU are relatively cumbersome. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a human motion capture method, system, and storage medium to alleviate the above problems existing in the related technologies.
[0006] In a first aspect, an embodiment of the present invention provides a human motion capture method, which is applied to a human motion capture system. The human motion capture system includes a base station, a fusion sensor, and a computer device. The human motion capture method includes: the base station receives a data packet from the fusion sensor, adds a timestamp to the data packet, and then forwards the timestamped data packet to the computer device; wherein, the fusion sensor includes a UWB module and an IMU module. The IMU module is used to collect the motion posture data of the target user, and the UWB module is used to transmit the data packet by using UWB technology. The data packet contains the motion posture data, and the timestamp is the reception moment corresponding to when the base station receives the data packet; the computer device obtains the channel impulse response corresponding to when the base station receives the data packet and the location information of the location where the base station is located, and determines a target data packet from the received data packets based on the channel impulse response; the computer device determines the first positioning information of the target user based on the location information and the timestamp corresponding to the target data packet, and determines the real-time position and real-time motion posture of the target user based on the first positioning information and the motion posture data included in the target data packet.
[0007] In a second aspect, an embodiment of the present invention further provides a human motion capture system. The human motion capture system includes a base station, a fusion sensor, and a computer device. The fusion sensor includes a UWB module and an IMU module. The IMU module is used to collect the motion posture data of the target user to form a data packet, and the UWB module is used to transmit the data packet by using UWB technology; the base station is used to receive the data packet from the fusion sensor, add a timestamp to the data packet, and then forward the timestamped data packet to the computer device; wherein, the data packet contains the motion posture data, and the timestamp is the reception moment corresponding to when the base station receives the data packet; the computer device is used to: obtain the channel impulse response corresponding to when the base station receives the data packet and the location information of the location where the base station is located, and determine a target data packet from the received data packets based on the channel impulse response; determine the first positioning information of the target user based on the location information and the timestamp corresponding to the target data packet, and determine the real-time position and real-time motion posture of the target user based on the first positioning information and the motion posture data included in the target data packet.
[0008] In a third aspect, an embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the human motion capture method described in the first aspect above.
[0009] A human motion capture method, system and storage medium provided by an embodiment of the present invention. A base station receives data packets from a fusion sensor, adds a timestamp to the data packets, and then forwards the data packets with timestamps to a computer device; the fusion sensor includes a UWB module and an IMU module. The IMU module is used to collect motion attitude data of a target user, and the UWB module is used to transmit data packets by using UWB technology. The data packets contain motion attitude data, and the timestamp is the reception moment corresponding to when the base station receives the data packets; the computer device obtains the channel impulse response corresponding to when the base station receives the data packets and the location information of the location where the base station is located, and determines a target data packet from the received data packets based on the channel impulse response; the computer device determines the first positioning information of the target user based on the location information and the timestamp corresponding to the target data packet, and determines the real-time position and real-time motion attitude of the target user based on the first positioning information and the motion attitude data included in the target data packet. By adopting the above technology, through communicating between the base station and the fusion sensor by using UWB technology and combining the IMU module to collect motion attitude data, the signal anti-interference ability of the system can be enhanced, the sensor data acquisition delay can be reduced, the system power consumption can be reduced, and the system bandwidth can be improved, so as to ensure the calculation accuracy of subsequent real-time position and real-time motion attitude; through the fusion of UWB and IMU, and determining the target data packet according to the channel impulse response data, and then using the location information of the location where the base station is located, the timestamp corresponding to the target data, and the motion attitude data included in the target data to determine the real-time position and real-time motion attitude of the target user, the calculation accuracy of the real-time position and real-time motion attitude can be significantly improved, and the attitude drift of the IMU module during long-term use can be suppressed, and the stability of long-term motion capture can be improved; in addition, positioning can be achieved by fusing UWB communication data with IMU data, or IMU data can be used for positioning when the UWB signal is blocked for a short time, so as to improve the stability of positioning.
[0010] Other features and advantages of the present invention will be described in the following description, and, in part, will be obvious from the description, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the description, the claims and the drawings.
[0011] To make the above objectives, features and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given, and detailed descriptions are made in conjunction with the accompanying drawings as follows. Description of the Drawings
[0012] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0013] Figure 1 Structural schematic diagram of a human motion capture system in an embodiment of the present invention;
[0014] Figure 2 Flow schematic diagram of a human motion capture method in an embodiment of the present invention;
[0015] Figure 3 Flow example diagram for determining a target data packet in an embodiment of the present invention;
[0016] Figure 4 Hardware architecture example diagram of a fusion sensor in an embodiment of the present invention;
[0017] Figure 5 Flow example diagram of a human motion capture method in an embodiment of the present invention;
[0018] Figure 6 Example diagram of using an RNN neural network to solve human postures in an embodiment of the present invention. Specific embodiments
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0020] The embodiments of the present invention provide a human motion capture method, which can be applied to a human motion capture system. Refer to Figure 1 As shown, the human motion capture system may include a base station 120, a fusion sensor 110, and a computer device 130.
[0021] Refer to Figure 2 As shown, the human motion capture method may include the following steps:
[0022] Step S202, the base station 120 receives a data packet from the fusion sensor 110, adds a timestamp to the data packet, and then forwards the data packet with the timestamp to the computer device 130.
[0023] Among them, the fusion sensor 110 includes a UWB module and an IMU module. The IMU module is used to collect the motion attitude data of the target user, and the UWB module is used to transmit data packets by UWB technology. The data packets contain motion attitude data, and the second timestamp is the receiving moment corresponding to when the base station 120 receives the data packet.
[0024] Step S204, the computer device 130 obtains the Channel Impulse Response (CIR) corresponding to when the base station 120 receives the data packet and the location information of the location where the base station 120 is located, and determines the target data packet from the received data packets based on the channel impulse response.
[0025] Step S206, the computer device 130 determines the first positioning information of the target user based on the location information and the timestamp corresponding to the target data packet, and determines the real-time position and real-time motion attitude of the target user based on the first positioning information and the motion attitude data included in the target data packet.
[0026] In the actual application process, the number of base stations 120 set is usually not less than three. The base stations 120 can communicate with each other and with the computer either through wired connections or through wireless connections (such as Wi-Fi, Bluetooth, etc.), and this is not limited.
[0027] A human motion capture method provided by an embodiment of the present invention can enhance the signal anti-interference of the system, reduce the sensor data acquisition delay, reduce the system power consumption, and increase the system bandwidth by communicating between the base station 120 and the fusion sensor 110 using UWB technology and combining the IMU module to collect motion attitude data, so as to ensure the calculation accuracy of the subsequent real-time position and real-time motion attitude; through the fusion of UWB and IMU, and determining the target data packet according to the channel impulse response data, and then using the location information of the location where the base station is located, the timestamp corresponding to the target data, and the motion attitude data included in the target data to determine the real-time position and real-time motion attitude of the target user, which can significantly improve the calculation accuracy of the real-time position and real-time motion attitude, and can suppress the attitude drift of the IMU module during long-term use, improving the stability of long-term motion capture; in addition, positioning can be achieved either by fusing IMU data with UWB communication data or by using IMU data when the UWB signal is blocked for a short time, thereby improving the stability of positioning.
[0028] In the actual application process, the above-mentioned channel impulse response may include the total received (RX) power, the first path (FP) power, and the time delay between the first path and the signal peak.
[0029] The accuracy of UWB positioning is affected by LOS (Line Of Sight) and NLOS (Non-Line Of Sight). LOS refers to the propagation mode where the UWB signal has no obstacle between the transmitter and the receiver; NLOS refers to the propagation mode where the UWB signal encounters obstacles during propagation. Under LOS conditions, the propagation path of the UWB signal is relatively simple, so the positioning accuracy is high. Under NLOS conditions, the UWB signal is affected by reflection and scattering of obstacles, resulting in a complex propagation path and thus reducing the positioning accuracy.
[0030] The above UWB module can use a UWB chip. Based on the hardware characteristics of the UWB chip, the channel conditions of LOS and NLOS can be classified to improve the accuracy of the three-dimensional positioning process. The process can be as follows: Use the RX power, FP power, and the time delay between the first path and the signal peak obtained from channel estimation to calculate the probability that the corresponding channel is a reliable signal under the LOS path and / or an unreliable signal under the NLOS path.
[0031] The calculation process of RX power is:
[0032]
[0033] In this formula, C refers to the overall power value of the channel impulse response estimated from the UWB chip, N refers to the cumulative count value of the preamble, and A is a predefined constant value.
[0034] The calculation process of FP power is:
[0035]
[0036] In this formula, F 1 、F 2 、F 3 refer to the first, second, and third harmonics of the signal amplitude of the first path estimated from the UWB chip.
[0037] The calculation process of the time delay between the first path and the signal peak is:
[0038] TimeDelay=(PeakPathIndex-FirstPathIndex)*1
[0039] In this formula, the difference between the index of the first path (FirstPathIndex) and the peak index (PeakPathIndex) is calculated and unit conversion is performed. 1 refers to the conversion from 1 index unit to 1 nanosecond.
[0040] Based on this, in the above step S204, the computer device 130 determining the target data packet from the received data packets based on the channel impulse response may include: for each received data packet, the computer device 130 determines the power difference between the total received power included in the data packet and the first path power included in the data packet, and determines the probability that the data packet belongs to non-line-of-sight based on the power difference corresponding to the data packet, the delay included in the data packet, and preset power threshold, scaling factor, delay threshold, and delay constant. If the probability corresponding to the data packet is less than the preset probability threshold, the data packet is determined as the target data packet.
[0041] Exemplarily, referring to Figures 1 to 2 as shown, after the computer device 130 determines the power difference between the total received power included in a certain received data packet and the first path power included in the data packet, the following operations can be performed:
[0042] Step 1, if the power difference corresponding to the data packet is greater than the power threshold, the probability that the data packet belongs to non-line-of-sight is determined to be 0.
[0043] Step 2, if the power difference corresponding to the data packet is not greater than the power threshold and greater than the product of the power threshold and the scaling factor, determine the probability that the data packet belongs to non-line-of-sight based on the power difference corresponding to the data packet, the power threshold, and the scaling factor.
[0044] Step 3, if the power difference corresponding to the data packet is not greater than the product of the power threshold and the scaling factor, determine the probability that the data packet belongs to non-line-of-sight based on the delay corresponding to the data packet, the delay threshold, and the delay constant.
[0045] Exemplarily, the above delay threshold may include a first delay threshold and a second delay threshold greater than the first delay threshold. For a certain data packet, the operation mode of the above step 3 can be: if the delay corresponding to the data packet is not greater than the first delay threshold, the probability that the data packet belongs to non-line-of-sight is determined to be 0; if the delay corresponding to the data packet is greater than the first delay threshold and less than the second delay threshold, determine the probability that the data packet belongs to non-line-of-sight based on the delay corresponding to the data packet and the delay constant; if the delay corresponding to the data packet is not less than the second delay threshold, the probability that the data packet belongs to non-line-of-sight is determined to be 1.
[0046] Among them, the above first delay threshold and second delay threshold can be defined according to the actual situation and are not limited thereto.
[0047] In an experimental scenario using at least four base stations 120, data collection can be performed on both NLOS signals and LOS signals between the base stations 120 and the computer device 130 for three parameters (i.e., the RX power, FP power of the CIR included in the data packet, and the time delay between the first path and the signal peak), and binary classification based on a threshold can be carried out, which is divided into two categories: reliable signals under the LOS path and unreliable signals under the NLOS path. See Figure 3 As shown, the binary classification process based on the threshold is as follows:
[0048] For a certain UWB signal, calculate the power difference between the RX power and the FP power; determine whether the power difference is greater than the power threshold; if the power difference is greater than the power threshold, then return that the probability value of the signal being NLOS is 0, that is, the signal is definitely not NLOS (definitely LOS); if the power difference is not greater than the power threshold, then determine whether the power difference is greater than the power threshold * scaling factor; if the power difference is greater than the power threshold * scaling factor, then return that the probability value of the signal being NLOS = (power difference / power threshold * scaling factor) / (1 - scaling factor); if the power difference is not greater than the power threshold * scaling factor, then determine whether the time delay is greater than the minimum time delay threshold (i.e., the first time delay threshold); if the time delay is not greater than the minimum time delay threshold, then return that the probability value of the signal being NLOS is 0, that is, the signal is definitely not NLOS (definitely LOS); if the time delay is greater than the minimum time delay threshold, then determine whether the time delay is less than the maximum time delay threshold (i.e., the second time delay threshold); if the time delay is less than the maximum time delay threshold, then return that the probability value of the signal being NLOS = first time delay constant * time delay - second time delay constant; if the time delay is not less than the maximum time delay threshold, then return that the probability value of the signal being NLOS is 1, that is, the signal is definitely NLOS (definitely not LOS).
[0049] Among them, the power threshold, minimum time delay threshold, maximum time delay threshold, scaling factor, time delay constant A, and time delay constant B can be determined and adjusted according to the actual scenario, and no limitations are imposed on this. The return value is the probability of the signal being NLOS.
[0050] After calculating the probability values of the UWB signals of all fusion sensors being NLOS, the UWB signals for subsequent positioning are those with probability values less than the preset probability threshold. The probability threshold can be defined according to the actual scenario, and generally, a relatively small probability value (such as 0.3, 0.2, 0.1, etc.) can be taken, and no limitations are imposed on this.
[0051] As a possible implementation, there are at least four such base stations; based on this, in step S206 above, the computer device 130 determining the first positioning information of the target user based on the location information and the timestamp corresponding to the target data packet may include: the computer device 130 determines one of the base stations 120 that receives the target data packet as the first base station, and determines each of the base stations other than the first base station as the second base station. Then, it determines the timestamp difference between the timestamp corresponding to the target data packet received by each second base station and the timestamp corresponding to the target data packet received by the first base station, and uses a preset TDOA (Time Difference of Arrival) positioning algorithm to solve the location information and the determined timestamp difference to determine the first positioning information.
[0052] Among them, the basic principle of the TDOA positioning algorithm is: to determine the position of the tag by measuring the time difference between multiple base stations 120 and the tag (i.e., the fusion sensor 110 in this article). Each base station 120 has an accurate clock. When the tag emits a UWB pulse carrying the pulse emission moment, the base station 120 records the pulse arrival moment. By comparing the time differences corresponding to different base stations 120 (i.e., the duration between the pulse emission moment and the pulse arrival moment), and combining the speed of light, the distance difference between the tag and each base station 120 can be calculated. Based on the distance difference and the location information of the position where the base station 120 is located, the position of the tag can be determined.
[0053] In the actual application process, the above preset TDOA positioning algorithm can adopt algorithms such as the Chan algorithm, the Fang algorithm, the Foy algorithm (i.e., the Taylor series algorithm), etc., and there is no limitation to this.
[0054] As a possible implementation, the above motion attitude data may include angular velocity and acceleration; based on this, in step S204 above, the computer device 130 determining the real-time position and real-time motion attitude of the target user based on the first positioning information and the motion attitude data included in the target data packet may include: the computer device 130 performs first filtering on the angular velocity and acceleration to obtain the angle information of the target user; the computer device 130 performs second filtering on the first positioning information and the acceleration to obtain the second positioning information; the computer device 130 uses a preset neural network algorithm to predict the second positioning information, the acceleration, and the angle information to obtain the real-time motion attitude.
[0055] Among them, the above-mentioned first filtering, the above-mentioned second filtering, and the above-mentioned neural network algorithm can be specifically selected according to the actual scenario. For example, the above-mentioned first filtering can adopt complementary filtering, Machony filtering, Madgwick filtering, Kalman Filtering (KF), etc., the above-mentioned second filtering can adopt Extended Kalman Filtering (EKF), unscented Kalman filtering, particle filtering, etc., and the above-mentioned neural network algorithm can adopt standard RNN (Recurrent Neural Network), LSTM, recursive RNN, etc., and no limitation is imposed on this.
[0056] Exemplarily, the above-mentioned IMU module includes a gyroscope and an accelerometer. A plurality of fusion sensors 110 can be worn on different parts of the user's body, so as to collect the angular velocity and acceleration corresponding to the respective fusion sensors 110 of the user during movement through the gyroscope and accelerometer of the same fusion sensor 110 as the user's motion attitude data; after the computer device 130 calculates the first positioning information by using the preset TDOA positioning algorithm, it can calculate the angle information of the user corresponding to each fusion sensor 110 by performing first filtering on the angular velocity and acceleration, and calculate the second positioning information by performing second filtering on the obtained first positioning information and acceleration. Then, the obtained second positioning information, acceleration, and angle information are input into a pre-trained neural network model to calculate and output the real-time motion attitude of the user through the neural network model.
[0057] Exemplarily, there can be six above-mentioned fusion sensors 110, which are respectively fixed on the first target part and the second target part of the target user. The first target part is the pelvis, and the second target part includes the left forearm, the right forearm, the left calf, the right calf, the head, and the pelvis; based on this, the steps for the above-mentioned computer device 130 to predict the second positioning information, acceleration, and angle information by using the preset neural network algorithm to obtain the real-time motion attitude can include: the computer device 130 generates the first input information based on the acceleration and angle information corresponding to the six fusion sensors, and generates the second input information based on the second positioning information corresponding to the five fusion sensors fixed on the second target part; the computer device 130 uses the above-mentioned neural network algorithm to predict the first input information and the second input information to obtain the above-mentioned real-time position and the above-mentioned real-time motion attitude.
[0058] In the actual application process, for the convenience of calculating the above-mentioned real-time position and the above-mentioned real-time motion posture, the above neural network algorithm may include a first neural network algorithm and a second neural network algorithm; based on this, when the computer device 130 uses the above neural network algorithm to predict the first input information and the second input information, it may use the above first neural network algorithm to perform a first prediction on the first input information and the second input information to obtain the above real-time position, and then use the above second neural network algorithm to perform a second prediction on the first input information and the above real-time position to obtain the above real-time motion posture.
[0059] For the convenience of understanding, the operation mode of the above human motion capture method is exemplarily introduced as follows by taking a specific application as an example.
[0060] See Figure 4 As shown, the hardware architecture of the fusion sensor 110 mainly includes a power management module, a UWB module, a main control module, and an IMU module.
[0061] Among them, the power management module includes a PMIC power management chip. The built-in charge and discharge management module controls the charging current of the battery through DCDC3. The power supply selection module can intelligently switch between power supply and battery supply. The power supply part is divided into three paths: LDO1, DCDC1, and DCDC2. LDO1 is responsible for supplying power to the RGB indicator light, DCDC2 is responsible for supplying power to the UWB chip radio frequency module, DCDC1 is the main power supply, and DCDC1 is responsible for supplying power to other remaining modules such as the main control module, the UWB chip (i.e., the UWB module), and the IMU chip (i.e., the IMU module). At the same time, the power management module will be connected to the main control chip (i.e., the main control MCU or the main control module), the UWB chip, and the control button through the control bus, and control the voltage regulation and on / off control of each power supply through the signals of the control bus. The main control module is connected to the IMU module and the UWB module through SPI. One high-precision crystal oscillator provides a clock signal for the main control MCU, and another high-precision crystal oscillator provides a clock signal for the UWB chip. The RTC crystal oscillator ensures stable timing when the main circuit of the main control module is in sleep. Two power filter modules are responsible for power noise reduction and decoupling. One power filter module is responsible for power filtering of the main control MCU, and the other power filter module is responsible for power filtering of the main power supply of the UWB chip). The main control chip obtains the acceleration, angular velocity, and time information of the IMU through the IMU module and sends it to the UWB chip. The UWB chip encapsulates the received information into a data packet and sends it to the external base station through the radio frequency part (responsible for radio frequency power filtering of the UWB chip) and the antenna part (connected to the UWB chip through the antenna interface). The UWB chip can also receive corresponding information from the external base station 120.
[0062] In the actual application process, both the fusion sensor 110 and the base station 120 can be powered by POE (Power Over Ethernet) or batteries, and the power supply methods of the fusion sensor 110 and the base station 120 are not limited herein.
[0063] The main process of the above human motion capture method is as follows: The UWB positioning position is obtained through TDOA positioning calculation, and the Kalman filter (KF) is used to fuse the calculated UWB positioning position with the IMU data to obtain accurate positioning information. Finally, the accurate positioning information and the IMU data are input into the RNN neural network to calculate the human body posture. See Figure 5 As shown, the above human motion capture method can be divided into five parts: LOS / NLOS classification judgment, Machony filtering, TDOA positioning calculation (CHAN algorithm), UWB-IMU fusion KF, and RNN neural network calculation. The specific process is as follows:
[0064] The UWB data is the data timestamp (TS) and the channel impulse response (CIR). The CIR is judged by the LOS / NLOS judgment module for the stability of the TS of different base stations, and the stable TS is used to calculate the UWB positioning position (i.e., the first positioning information) through the CHAN algorithm of TODA positioning. The IMU data is the acceleration and angular velocity. The acceleration and angular velocity are filtered by Machony to obtain the angle information, and the first positioning information and the acceleration are filtered by the Kalman filter (KF) to obtain more accurate positioning information (i.e., the second positioning information). Finally, the second positioning information, the acceleration, and the angle information are input into the RNN neural network to calculate and output the human body posture (i.e., the above real-time motion posture) through the RNN neural network. The RNN neural network will also output the human body position (i.e., the above real-time position) during the calculation process.
[0065] The implementation processes of LOS / NLOS classification judgment, Machony filtering, TDOA positioning calculation, UWB-IMU fusion KF, and RNN neural network calculation will be introduced in detail below respectively.
[0066] (1) LOS / NLOS classification judgment
[0067] After each time the fusion sensor 110 sends a UWB pulse wave to the base station 120, when receiving at the antenna end of the base station 120, the RX power, FP power obtained from channel estimation, and the time delay between the first path and the signal peak can be read out from the UWB chip; based on this, the process of LOS / NLOS classification and determination can be as follows: Calculate the probability value that the corresponding UWB signal is NLOS based on the RX power, FP power, and the time delay between the first path and the signal peak obtained from channel estimation, and use the TS corresponding to the UWB signal with a probability value less than 0.1 as the basic data for subsequent TDOA positioning calculation.
[0068] The probability value that the corresponding UWB signal is NLOS can be specifically calculated according to the threshold-based binary classification method shown in the previous text Figure 3 and will not be elaborated here.
[0069] (2) Mahony Filtering
[0070] The basic principle of Mahony filtering includes two steps: prediction and correction, as follows:
[0071] Prediction: According to the angular velocity collected by the gyroscope of the IMU module, use the differential equation of quaternion to predict the attitude change of the rigid body (i.e., the fusion sensor 110), that is:
[0072]
[0073] Among them, is the attitude quaternion of the current rigid body, q is the attitude quaternion of the previous moment, ω is the angular velocity quaternion of the gyroscope, is the multiplication of quaternions.
[0074] Correction: According to the gravity direction collected by the accelerometer of the IMU module, use the rotation equation of quaternion to calculate the attitude error of the rigid body, and then use the proportional-integral controller to correct the attitude estimation of the rigid body, that is:
[0075]
[0076] Among them, e is the attitude error quaternion of the rigid body, K p and K i are the proportional and integral gain parameters respectively, ∫edt is the integral term of the error, and the attitude error quaternion e can be calculated according to the formula e = sine = a × g, is the gravity acceleration vector of the accelerometer, is the gravity acceleration vector integrated by the gyroscope.
[0077] The proportional term K p is used to control the "confidence" from different sensor sources, and the integral term Ki used to eliminate static errors. K p The larger the value of K, the more significant the error compensation from the accelerometer, that is, the more the accelerometer data is trusted. Conversely, when K p is smaller, the data from the gyroscope is more trusted, and the error compensation from the accelerometer is less. And the integral term K i is used to eliminate the bias noise in the angular velocity measurement value. However, for the gyroscope data that has been zero-bias corrected, a very small value of K i can meet the usage requirements. The selection of K p and K i should match the movement of the carrier (human limb), that is, the values of K p and K i need to be determined according to the actual movement of the human limb.
[0078] The last step is to apply the compensation values (i.e., K p e and K i ∫edt) to the angular velocity measurement value (i.e., ω), and substitute it into the quaternion difference equation (i.e., ) to update the current quaternion (i.e., q) to obtain the angle information (i.e., ).
[0079] (III) TDOA Positioning Solution
[0080] The TDOA positioning solution can be implemented through the CHAN algorithm. The specific solution principle is as follows:
[0081] TDOA calculates the position of the tag (i.e., the fusion sensor 110) by measuring the distance differences from the tag to all base stations 120. The tag only broadcasts one message to each base station 120 for position measurement. For example, for two base stations (i.e., Base Station 1 and Base Station 2) and one tag, the measured distance difference can be expressed by the following formula (which is the difference between the distances from Base Station 2 and Base Station 1 to the tag respectively):
[0082]
[0083] where, b 1 represents Base Station 1, b 2 represents Base Station 2, t represents the tag,, is the distance from Base Station 2 to the tag, is the distance from Base Station 1 to the tag, T represents the time difference, is the time difference between the reception time (i.e., the timestamp) corresponding to the data received by Base Station 2 from the tag and the transmission time (i.e., the timestamp) corresponding to the data sent by the tag to Base Station 2, It is the time difference between the reception time (i.e., timestamp) corresponding to the data sent by the tag received by Base Station 1 and the transmission time (i.e., timestamp) corresponding to the tag sending the corresponding data to Base Station 1.
[0084] For a scenario with N base stations, the position of the tag can be solved using the following CHAN algorithm:
[0085]
[0086] where:
[0087] K S =||X S || 2
[0088] Among them, represents the difference in the distances from Base Station N and Base Station 1 to the tag respectively. S ∈ {t, b 1 , b 2 ,..., b N}, X S =(x S , y S , z S ), t represents the tag, b 1 , b 2 ,…, b N represent Base Stations 1 to N, X S represents the position of S. The position of the tag can be solved by first assuming that there is no relationship between x S , y S , z S and using the least squares method or the maximum likelihood method.
[0089] (IV) UWB-IMU Fusion KF
[0090] The algorithm for this part is mainly divided into five steps: state prediction, state covariance prediction, Kalman gain calculation, state update, and state covariance matrix update. Among them, the IMU acceleration data is used for state prediction, and the UWB ranging data is used for state update. The specific algorithm process can be expressed as:
[0091] Init (Initialization stage): Initial values of the state equation X(0), state covariance matrix P(0), accelerometer control input matrix B, covariance matrix Q, UWB sensor observation noise covariance matrix R, n×n identity matrix I n
[0092] Input (Input stage): Accelerometer data u(k)
[0093] Begin (Start stage):
[0094] 1. State prediction
[0095]
[0096] 2. State covariance matrix prediction
[0097] P(k|k - 1) = FP(k - 1|k - 1)F T + Q
[0098] 3. Kalman gain calculation
[0099] K = P(k|k - 1)H T (k)[H(k)P(k|k - 1)H T (k)+R] -1
[0100] 4. State update
[0101]
[0102] 5. State covariance matrix update
[0103] P(k|k) = [I n - KH(k)]P(k|k - 1)
[0104] Where F is the state transition matrix, H is the measurement matrix, and F, B, and H can be set in the following forms respectively:
[0105]
[0106] Where k is the time, T is the acquisition duration, X(k) is the system state vector, and X(k) includes the position x(k) and velocity v(K), then X(k) can be expressed as:
[0107] X(k) = [x x (k), x y (k), x z (k), v x (k), v y (k), v z (k)] T
[0108] Z is the observation matrix of UWB data, representing the three-dimensional position of the UWB tag.
[0109] (5) RNN neural network calculation
[0110] The RNN neural network has three inputs, namely the acceleration contained in the IMU data, the angle information obtained through the above-mentioned Machony filtering, and the UWB positioning position obtained after the above-mentioned UWB-IMU fusion KF; the output of the RNN neural network is the human body posture.
[0111] See Figure 6 As shown, the human body posture can be restored through six fusion sensors 110 (respectively installed on the user's left forearm, right forearm, left calf, right calf, head, and pelvis). Among them, the fusion sensor 110 installed on the pelvis can be regarded as the root node where the human body remains stationary, and the remaining five fusion sensors 110 can be regarded as leaf nodes that move relative to the root node. Since the IMU module can obtain the acceleration of 3 axes (R 3 ) and can obtain the rotation of 3 axes represented in the form of a rotation matrix through the above-mentioned Machony filtering (R 3*3 ), the data obtained by the six IMU modules can be normalized and connected together to form a 72-dimensional input variable x (R 6*(3+3*3) ). At the same time, the leaf node fusion positions (i.e., the above-mentioned precise positioning information or the above-mentioned second positioning information) obtained by fusing all the leaf nodes through the above-mentioned UWB-IMU fusion positioning (including LOS / NLOS classification judgment, TDOA positioning calculation, and UWB-IMU fusion KF) are normalized and connected together to form a 15-dimensional input variable pL (R 5*3 ).
[0112] See Figure 6 As shown, the RNN neural network can adopt a two-level RNN. The first-level RNN is denoted as RNN1, and the second-level RNN is denoted as RNN2. The matrix formed by the two variables x and pL is the input matrix of RNN1. The output of RNN1 is the full node position (including the position of each node relative to the root node, where the output value of the root node corresponding to RNN1 is 0). The obtained full node position and the input variable x are input to RNN2 together to calculate and output the human body posture through RNN2.
[0113] We use the DIP-IMU dataset, which is a dataset based on the SMPL model. The DIP-IMU dataset contains IMU data collected by multiple participants wearing multiple (usually more than a dozen) IMU sensors. The DIP-IMU dataset aims to overcome the differences in sensor data characteristics between synthetic data and real data and supplement the activity types in existing motion capture (Mocap) datasets. Participants are required to repeat five different categories of actions, including controlled movements of limbs (arms, legs), sports, and more natural whole-body activities (such as jumping, boxing) and interaction tasks with daily objects. The DIP-IMU dataset can be used to train and evaluate models for tasks such as action recognition, pose estimation, and motion generation. Since the leaf node positions can be obtained using the DIP-IMU dataset and the errors of the above-mentioned human motion capture system positioning can be obtained, UWB data (used to simulate precise positioning information) can be generated using the known leaf node positions and errors, and this part of the UWB data and the IMU information (i.e., acceleration and rotation) are used as the original input of the RNN neural network. The human poses corresponding to the original input in the DIP-IMU dataset are used as labels. Then, the training set and test set of the RNN neural network are constructed using the original input and its corresponding labels to train and test the RNN neural network. After the training and testing are completed, the trained RNN neural network can be used to perform the above RNN neural network calculations to solve the human pose.
[0114] In summary, with the above human motion capture method, since only the fusion sensor 110, the base station 120, and the computer device are required for the hardware part to achieve human pose calculation, the hardware cost is lower than that of optical motion capture. Compared with the sensor system used in inertial motion capture, in addition to having the basic functions of the IMU, the UWB communication method has stronger anti-interference ability, and positioning can also be achieved using the UWB communication method. Since the human pose can be calculated directly using the IMU data or by fusing the IMU data with the UWB data, dual-modal calculation of the human pose can be achieved, thereby improving the calculation accuracy of the motion pose and avoiding the pose drift problem of the IMU module during long-term use. By deploying multiple fusion sensors 110 on the human body simultaneously, the distance information of multiple fusion sensors 110 can be combined to reduce the influence of abnormal distance information caused by partial occlusion of some fusion sensors 110. Combining UWB positioning and IMU integration, temporary positioning information can be obtained using IMU integration when the UWB signal is occluded for a short time, thereby improving the stability of positioning.
[0115] Based on the above human motion capture method, an embodiment of the present invention further provides a human motion capture system. See Figure 1As shown in the figure, the human body motion capture system may include a base station 120, a fusion sensor 110, and a computer device 130. The fusion sensor 110 includes a UWB module and an IMU module. The IMU module is used to collect the motion posture data of the target user to form a data packet, and the UWB module is used to transmit the data packet by using UWB technology. The base station 120 is used to receive the data packet from the fusion sensor 110, add a timestamp to the data packet, and then forward the data packet with the timestamp to the computer device 130. Wherein, the data packet contains motion posture data, and the timestamp is the reception moment corresponding to when the base station 120 receives the data packet. The computer device 130 is used to: obtain the channel impulse response corresponding to when the base station 120 receives the data packet and the location information of the location where the base station 120 is located, and determine the target data packet from the received data packets based on the channel impulse response; determine the first positioning information of the target user based on the location information and the timestamp corresponding to the target data packet, and determine the real-time location and real-time motion posture of the target user based on the first positioning information and the motion posture data included in the target data packet.
[0116] For the human body motion capture system provided by the embodiments of the present invention, its implementation principle and the technical effects produced are the same as those of the foregoing embodiments of the human body motion capture method. For a brief description, for the parts not mentioned in the system embodiments, reference may be made to the corresponding content in the foregoing method embodiments.
[0117] The embodiments of the present invention further provide a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the foregoing human body motion capture method. For the specific implementation, reference may be made to the foregoing method embodiments, and details are not described herein again.
[0118] Unless otherwise specifically stated, the relative steps, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the present invention.
[0119] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0120] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0121] Finally, it should be noted that the above-mentioned embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting them. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions described in the foregoing embodiments, or can easily think of changes, or make equivalent replacements for some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be determined by the protection scope of the claims.
Claims
1. A human motion capture method, characterized in that: The invention is applied to a human motion capture system, the human motion capture system comprises a base station, a fusion sensor and a computer device, and the human motion capture method comprises: The base station receives a data packet from the fusion sensor, adds a timestamp to the data packet, and then forwards the data packet with the timestamp to the computer device; wherein the fusion sensor includes a UWB module and an IMU module, the IMU module is used to collect motion posture data of the target user, the UWB module is used to transmit the data packet using UWB technology, the data packet includes the motion posture data, and the timestamp is the receiving time corresponding to when the base station receives the data packet; The computer device obtains the channel impulse response corresponding to when the base station receives the data packet and the location information of the base station, and determines the target data packet from the received data packets based on the channel impulse response; The computer device determines the first positioning information of the target user based on the position information and the timestamp corresponding to the target data packet, and determines the real-time position and real-time motion posture of the target user based on the first positioning information and the motion posture data included in the target data packet; The channel impulse response includes a total received power, a first path power, and a time delay between the first path and a signal peak; determining a target data packet from received data packets based on the channel impulse response, including: for each received data packet, determining a power difference between a total received power contained in the data packet and a first path power contained in the data packet, and determining a probability that the data packet belongs to non-line-of-sight based on the power difference corresponding to the data packet and the time delay contained in the data packet as well as a preset power threshold, a scaling factor, a time delay threshold, and a time delay constant; if the probability corresponding to the data packet is less than a preset probability threshold, determining the data packet as a target data packet.
2. The human motion capture method according to claim 1, characterized in that: There are at least four base stations; Determining first positioning information of the target user based on the location information and a timestamp corresponding to the target data packet includes: Determine one of the base stations that receives the target data packet as the first base station, and determine each base station other than the first base station as a second base station, and then determine the timestamp difference between the timestamp corresponding to the target data packet received by each second base station and the timestamp corresponding to the target data packet received by the first base station; The preset TDOA positioning algorithm is used to solve the position information and the determined timestamp difference to determine the first positioning information.
3. The human motion capture method according to claim 1, characterized in that: The motion posture data includes angular velocity and acceleration; determining the real-time position and real-time motion posture of the target user based on the first positioning information and the motion posture data included in the target data packet, including: Performing a first filtering on the angular velocity and the acceleration to obtain angle information of the target user; Performing a second filtering on the first positioning information and the acceleration to obtain second positioning information; The second positioning information, the acceleration and the angle information are predicted using a preset neural network algorithm to obtain the real-time motion posture.
4. The human motion capture method according to claim 1, characterized in that: Determining the probability that the data packet belongs to non-line-of-sight based on the power difference corresponding to the data packet, the delay included in the data packet, and a preset power threshold, a scaling factor, a delay threshold, and a delay constant, including: If the power difference corresponding to the data packet is greater than the power threshold, the probability that the data packet belongs to non-line-of-sight is determined to be 0; If the power difference corresponding to the data packet is not greater than the power threshold and greater than the product of the power threshold and the scaling factor, determining the probability that the data packet belongs to non-line-of-sight based on the power difference corresponding to the data packet, the power threshold and the scaling factor; If the power difference corresponding to the data packet is not greater than the product of the power threshold and the scaling factor, the probability that the data packet belongs to non-line-of-sight is determined based on the delay corresponding to the data packet, the delay threshold and the delay constant.
5. The human motion capture method according to claim 4, characterized in that: The delay threshold includes a first delay threshold and a second delay threshold greater than the first delay threshold; Determining the probability that the data packet belongs to non-line-of-sight based on the delay corresponding to the data packet, the delay threshold, and the delay constant includes: If the delay corresponding to the data packet is not greater than the first delay threshold, the probability that the data packet belongs to non-line-of-sight is determined to be 0; If the delay corresponding to the data packet is greater than the first delay threshold and less than the second delay threshold, determining the probability that the data packet belongs to non-line-of-sight based on the delay corresponding to the data packet and the delay constant; If the delay corresponding to the data packet is not less than the second delay threshold, the probability that the data packet belongs to non-line-of-sight is determined to be 1.
6. The human motion capture method according to claim 3, characterized in that: There are six fusion sensors, which are respectively fixed on the first target part and the second target part of the target user, the first target part is the pelvis, and the second target part includes the left forearm, the right forearm, the left calf, the right calf and the head; The method uses a preset neural network algorithm to predict the second positioning information, the acceleration and the angle information to obtain the real-time motion posture, including: Generate first input information based on acceleration and angle information corresponding to the six fusion sensors, and generate second input information based on second positioning information corresponding to the five fusion sensors fixed on the second target part; The neural network algorithm is used to predict the first input information and the second input information to obtain the real-time position and the real-time motion posture.
7. The human motion capture method according to claim 6, characterized in that: The neural network algorithm includes a first neural network algorithm and a second neural network algorithm; using the neural network algorithm to predict the first input information and the second input information to obtain the real-time position and the real-time motion posture includes: Using the first neural network algorithm to perform a first prediction on the first input information and the second input information to obtain the real-time position; The second neural network algorithm is used to perform a second prediction on the first input information and the real-time position to obtain the real-time motion posture.
8. A human motion capture system, characterized in that: The human motion capture system includes a base station, a fusion sensor and a computer device, the fusion sensor includes a UWB module and an IMU module, the IMU module is used to collect the motion posture data of the target user to form a data packet, and the UWB module is used to transmit the data packet using the UWB technology; The base station is used to receive a data packet from the fusion sensor, add a timestamp to the data packet, and then forward the data packet with the timestamp to the computer device; wherein the data packet includes the motion posture data, and the timestamp is the receiving time corresponding to when the base station receives the data packet; The computer device is used to: obtain the channel impulse response corresponding to the base station when receiving the data packet and the location information of the base station, and determine the target data packet from the received data packet based on the channel impulse response; determine the first positioning information of the target user based on the location information and the timestamp corresponding to the target data packet, and determine the real-time position and real-time motion posture of the target user based on the first positioning information and the motion posture data contained in the target data packet; The channel impulse response includes a total received power, a first path power, and a time delay between the first path and a signal peak; determining a target data packet from received data packets based on the channel impulse response, including: for each received data packet, determining a power difference between a total received power contained in the data packet and a first path power contained in the data packet, and determining a probability that the data packet belongs to non-line-of-sight based on the power difference corresponding to the data packet and the time delay contained in the data packet as well as a preset power threshold, a scaling factor, a time delay threshold, and a time delay constant; if the probability corresponding to the data packet is less than a preset probability threshold, determining the data packet as a target data packet.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by the processor, the computer-executable instructions prompt the processor to implement the human motion capture method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Motion posture evaluation system based on UWB / IMU fusion
CN116019442A
Human body motion posture capturing method and system and storage medium
CN117213495A
Device determining method, electronic device,and computer-readable storage medium
US20230341543A1