A method for estimating spatial position of human body in RGB video and its step speed derivation application
Patent Information
- Application Number
- CN202410422380.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-09
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-04-09
AI Technical Summary
这会导致测量的步速受到相机参数、图像质量和视角等因素的影响,进而产生较大误差
[0034]1、本发明提出了一种针对RGB视频中人体空间位置估计方法,利用2D人体姿态估计与3D人体姿态估计技术从RGB视频中提取2D人体关节点位置与无深度信息的相对3D人体姿态,并且采集摄像机的相机内参,利用相机内参对提取出的2D人体关节点位置进行投影编码操作,将其投影到归一化投影平面上,以消除相机成像所带来的误差;之后将归一化投影后的2D人体关节点位置与相对3D人体姿态合并作为输入,经由空间位置估计网络,输出视频每一帧中人体骨盆关节点在相机坐标系下的3D位置,具有较高估计准确性;
Smart Images

Figure CN118314207B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of computer vision and speed measurement, specifically relating to a method for estimating the spatial position of the human body in RGB video and its application for deriving walking speed. Background Technology
[0002] Human spatial position estimation estimates the 3D position of a fixed joint point (e.g., the hip joint) in each frame of a video, and is an extension of human pose estimation. Human pose estimation is mainly divided into 2D human pose estimation and 3D human pose estimation, which respectively estimate the 2D pixel coordinates and 3D spatial positions of human joint points in an image or video in the camera coordinate system. 3D human pose estimation is further divided into relative 3D human pose estimation and absolute 3D human pose estimation; relative 3D human pose estimation estimates the 3D human pose in the camera coordinate system without depth information, while absolute 3D human pose estimation additionally estimates the specific position of the pelvic joint points in the camera coordinate system.
[0003] The main methods for absolute 3D human pose estimation include direct estimation and minimizing reprojection error. For direct estimation, Pavllo et al. proposed a method in "3D Human Pose Estimation in Video With Temporal Convolutions and Semi-Supervised Training" that directly regresses 3D human pose from 2D human pose. However, due to the lack of consideration for factors such as camera intrinsic parameter variations, its estimation accuracy is poor. Moon et al., in "Camera distance-aware top-down approach for 3d multi-person pose estimation from a single RGB image," designed a network to estimate the root joint depth from a cropped single-person image, which inevitably loses contextual information about the subject. Minimizing reprojection error involves reprojecting the relative 3D human pose onto the pixel plane and minimizing the error between the estimated depth value and the 2D human pose by adjusting the depth value to be estimated, ultimately obtaining the 3D position of the human pelvic joints in the camera coordinate system within a single frame. However, this method is constrained by the instability of the minimization error algorithm, and the estimation results often vary greatly between different frames, leading to discontinuities and jitter in the position estimation results.
[0004] Pedestrian speed measurement is of crucial value in practical applications, especially in sports and athletic events such as track and field. There are two main methods for measuring gait speed: sensor-based and video-based.
[0005] Sensor-based methods often employ handheld or wearable devices to acquire data. For example, Bishop used two accelerometers mounted on the lower legs to calculate gait, treating the inverted pendulum posture as an independent stride cycle; Park used a handheld device to estimate gait, requiring participants to hold a triaxial accelerometer for speed measurement; Gomez used a wearable, visualized sensor device with test points mounted on the user's glasses to calculate the user's gait; Lu Yongle et al. proposed a multi-motion pattern recognition algorithm for humans based on MEMS (Micro-Electro-Mechanical Systems) inertial sensors, selecting the temporal features of the MEMS accelerometer as pattern recognition features and extracting the temporal features of the MEMS angular velocity sensor as secondary recognition features to achieve gait detection. However, these methods have limitations, such as the possibility that the carried devices may affect the natural movement of pedestrians, and data transmission delays.
[0006] In recent years, advancements in sensor technology, particularly the widespread adoption of MEMS sensors and IoT (Internet of Things) technologies, have begun to overcome these challenges. Highly integrated and low-power sensors can measure stride speed more accurately and in real-time, while cloud computing and edge computing technologies allow for faster data processing and real-time feedback. However, for some smaller sports teams, the high cost of sensors and their accompanying software remains prohibitive.
[0007] Video-based methods require only a single camera to collect data, offering advantages such as low cost and ease of deployment. Numerous innovations have also emerged in video-based gait measurement. Beyond traditional monocular RGB cameras and multi-camera systems, modern solutions are incorporating deep learning and computer vision algorithms to improve the accuracy of 2D and 3D coordinate tracking. For example, convolutional neural networks (CNNs) are used for object detection and tracking, followed by time-series analysis using recurrent neural networks (RNNs) or long short-term memory networks (LSTMs) to more accurately estimate gait.
[0008] Nevertheless, a major problem remains in video-based gait measurement: the lack of distance references in the images. This causes the measured gait speed to be affected by factors such as camera parameters, image quality, and viewpoint, resulting in significant errors. Therefore, there is an urgent need for a high-precision and reliable image-based method for human spatial position estimation and gait measurement to meet the needs of various practical application scenarios. Summary of the Invention
[0009] To address the shortcomings of the existing technology, this invention provides a method for estimating human spatial position in RGB video and its application for deriving walking speed, which can eliminate errors caused by camera imaging and improve the accuracy of human spatial position estimation.
[0010] A method for estimating the spatial position of the human body in RGB video includes the following steps:
[0011] Step 1: Construct a training set consisting of multiple RGB videos and obtain the camera intrinsics of the camera that captured the RGB videos, including a 3×3 projection matrix and a distortion matrix consisting of multiple meridional distortion parameters and tangential distortion parameters.
[0012] Step 2: Calculate the human spatial position matrix for each RGB video, specifically:
[0013] Step 2.1: Based on the 2D human pose estimation network, estimate the pixel coordinates of human joints in each frame of the RGB video in the pixel coordinate system to obtain an N×2 dimensional 2D human joint matrix before correction; where N is the total number of human joints.
[0014] Step 2.2: Based on the distortion matrix, perform distortion correction on the 2D human joint matrix before correction to obtain an N×2 dimensional 2D human joint matrix after correction.
[0015] Step 2.3: Based on the inverse of the projection matrix, perform intrinsic parameter normalization on the corrected 2D human joint matrix to obtain an N×2 dimensional intrinsic parameter normalized 2D human joint matrix.
[0016] Step 2.4: Based on the relative 3D human pose estimation network, perform relative 3D human pose estimation on the input intrinsic parameter normalized 2D human joint matrix to obtain an N×3 human 3D pose matrix in the camera coordinate system with the human hip joint as the origin; concatenate the N×2 dimensional intrinsic parameter normalized 2D human joint matrix with the N×3 dimensional human 3D pose matrix to obtain an N×5 human spatial position matrix.
[0017] Step 3: Construct a spatial location estimation network. Take the human spatial location matrix corresponding to each RGB video in the training set as input, the 3D position of the human hip joint in each frame of each RGB video in the camera coordinate system as the estimation target, and minimize the mean squared error (MSE) loss error between the network output and the true value of the estimation target as the training objective. Train the spatial location estimation network to obtain the trained spatial location estimation network.
[0018] Step 4: After the RGB video to be estimated is processed sequentially in Step 2, the N×5 human spatial position matrix corresponding to the RGB video to be estimated is obtained. This matrix is then input into the trained spatial position estimation network, which outputs the human spatial position in the RGB video to be estimated, i.e., the 3D position of the human hip joint in each frame in the camera coordinate system.
[0019] Furthermore, N=17, and the joints of each human body are the hip joint, right hip joint, right knee joint, right ankle joint, left hip joint, left knee joint, left ankle joint, thoracic spine joint, neck joint, nose joint, head joint, left shoulder joint, left elbow joint, left hand joint, right shoulder joint, right elbow joint, and right hand joint.
[0020] Furthermore, the 2D human pose estimation network in step 2.1 is a publicly available network for 2D human pose estimation, such as the ViT (Vision Transformer) network, the stacked hourglass network, the Openpose network, etc.
[0021] Furthermore, step 2.2 uses distortion correction functions from open-source toolkits such as OpenCV to perform distortion correction operations.
[0022] Furthermore, the formula for the intrinsic parameter normalization operation in step 2.3 is:
[0023]
[0024] In the formula, f is the inverse of the projection matrix. x f represents the actual physical length of each pixel in the horizontal direction of the camera's imaging in the normalized coordinate system, in meters. y c represents the actual physical length of each pixel along the vertical axis in the normalized coordinate system of the camera image, in meters. x c is the x-coordinate of the pixel at the center of the image. y y is the pixel ordinate of the imaging center; u and v are the length and height coordinates of each human joint in the corrected 2D human joint matrix, respectively; X / Z and Y / Z are the length and height coordinates of each human joint in the 2D human joint matrix after intrinsic parameter normalization, respectively.
[0025] Furthermore, the 3D human pose estimation network in step 2.4 is specifically a publicly available network for 3D human pose estimation, such as the MixSTE network, MotionBERT network, PoseFormer network, etc.
[0026] Furthermore, in step 3, the spatial location estimation network adopts the Transformer Encoder architecture, which includes a location encoding layer, a multi-head self-attention (MSA) layer, and a multilayer perceptron (MLP).
[0027] Furthermore, in step 3, the input during training is a T×5N input sequence formed by concatenating the human spatial position matrices corresponding to each RGB video. The specific concatenation method is as follows: the N×5 human spatial position matrix is expanded into a 1×5N unit sequence, and then the 1×5N unit sequences of adjacent frames are concatenated to form a multi-frame input sequence, namely the T×5N input sequence; where T is the preset number of sequences.
[0028] Afterwards, the RGB video to be estimated in step 4 is processed sequentially in step 2 to obtain the N×5 human spatial position matrix corresponding to the RGB video to be estimated. The resulting T×5N input sequence is then fed into the trained spatial position estimation network to output the human spatial position in the RGB video to be estimated, that is, the 3D position of the human hip joint in each frame in the camera coordinate system.
[0029] This invention also proposes the application of the aforementioned method for estimating human spatial position in RGB video in gait rate extraction.
[0030] A method for deriving step speed, the specific process is as follows:
[0031] Based on the N×3 human 3D pose matrix corresponding to the RGB video to be estimated, the distance between the two ankles of the human body in each frame is calculated, resulting in a curve showing the distance changing with the frame number. All maxima of the curve are solved, and the specific frame and frame number corresponding to each maxima are obtained. Based on the 3D position of the human hip joint in the camera coordinate system output by the trained spatial position estimation network for each frame, the 3D position of the human hip joint in the camera coordinate system in a specific frame is obtained according to the frame number corresponding to each maxima. The distance traveled between adjacent specific frames is calculated, and then the step speed for each step is calculated.
[0032]
[0033] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0034] 1. This invention proposes a method for estimating the spatial position of the human body in RGB video. It utilizes 2D and 3D human pose estimation techniques to extract 2D human joint positions and relative 3D human pose (without depth information) from the RGB video. Furthermore, it collects camera intrinsic parameters and uses these parameters to project the extracted 2D human joint positions onto a normalized projection plane to eliminate errors caused by camera imaging. Then, the normalized projected 2D human joint positions and relative 3D human pose are combined as input and passed through a spatial position estimation network to output the 3D position of the human pelvic joints in the camera coordinate system for each frame of the video, achieving high estimation accuracy.
[0035] 2. Based on human spatial position estimation, this invention also proposes a method for deriving walking speed, with a walking speed error of about 1%, which is extremely low and has stable performance;
[0036] 3. Preferably, in practical applications, the input sequence obtained by input splicing can ensure the continuity of the estimation results and greatly reduce the instability of the output position caused by the position estimation method that minimizes the reprojection error. Attached Figure Description
[0037] Figure 1 This is a diagram of the human posture structure used in Embodiment 1 of the present invention;
[0038] Figure 2 This is a schematic diagram of the network architecture of the spatial location estimation network in Embodiment 1 of the present invention;
[0039] Figure 3 This is a schematic diagram of the multi-head self-attention layer in Embodiment 1 of the present invention;
[0040] Figure 4 This is a schematic diagram of the structure of the multilayer sensor layer in Embodiment 1 of the present invention. Detailed Implementation
[0041] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0042] Example 1
[0043] This embodiment proposes a method for estimating the spatial position of the human body in RGB video, including the following steps:
[0044] Step 1: Construct a training set consisting of multiple RGB videos and obtain the camera intrinsic parameters of the camera that captured the RGB videos (at arbitrary resolution), including a 3×3 projection matrix and a distortion matrix consisting of multiple meridional distortion parameters and tangential distortion parameters.
[0045] The camera's intrinsic parameters are obtained using methods such as Zhang's calibration method, or by consulting the camera manufacturer.
[0046] The Zhang calibration method is a method that calculates the optimal camera intrinsic parameters by taking multiple shots of the chessboard from different angles using a camera. The chessboard is a calibration board composed of alternating black and white squares. The side lengths of the black and white squares are uniform and known, and it is used as the calibration object for camera calibration (an object mapped from the real world into the digital image).
[0047] Step 2: Calculate the human spatial position matrix for each RGB video, specifically:
[0048] Step 2.1: Based on a 2D human pose estimation network (specifically a ViT network), estimate the pixel coordinates of human joints in each frame of the RGB video in the pixel coordinate system, and obtain the N×2 dimensional 2D human joint matrix X before correction. 2D Where N = 17 is the total number of joints in the human body;
[0049] Figure 1 This is a diagram of human posture and structure. The joints of the human body include: 0 hip joint, 1 right hip joint, 2 right knee joint, 3 right ankle joint, 4 left hip joint, 5 left knee joint, 6 left ankle joint, 7 thoracic spine joint, 8 neck joint, 9 nose joint, 10 head joint, 11 left shoulder joint, 12 left elbow joint, 13 left hand joint, 14 right shoulder joint, 15 right elbow joint, and 16 right hand joint.
[0050] 2D human pose estimation is a common computer vision task. Its purpose is to locate and identify the pixel coordinates of human joints (such as the head, left hand, right foot, etc.) in an image or video to be estimated. This coordinate system is a pixel coordinate system, with the origin (0,0) usually at the top left corner of the image, and the bottom right corner representing the pixel's length and height (w,h). These joints are connected in sequence to form the human torso, such as... Figure 1 As shown;
[0051] The 2D human pose estimation network can be implemented using any network or method. The estimation result of 2D pose estimation will affect the final estimation result to a certain extent. In this embodiment, it is recommended to use the ViT network.
[0052] Step 2.2: Based on the distortion matrix, adjust the pre-correction 2D human joint matrix X. 2D Distortion correction can be performed using distortion correction functions in open-source toolkits such as OpenCV. This corrects the pixel coordinates of distorted human joints to their correct primitive coordinates, resulting in an N×2 dimensional corrected 2D human joint matrix X. 2D-DC ;
[0053] Step 2.3: Based on the inverse of the projection matrix, process the corrected 2D human joint matrix X. 2D-DC Perform intrinsic parameter normalization to obtain an N×2 dimensional 2D human joint matrix X after intrinsic parameter normalization. 2D-Norm The operating formula is:
[0054]
[0055] In the formula, f is the inverse of the projection matrix. xf represents the actual physical length of each pixel in the horizontal direction of the camera's imaging in the normalized coordinate system, in meters. y c represents the actual physical length of each pixel along the vertical axis in the normalized coordinate system of the camera image, in meters. x c is the x-coordinate of the pixel at the center of the image. y y is the pixel ordinate of the imaging center; u and v are the length and height coordinates of each human joint in the corrected 2D human joint matrix, respectively, with u ranging from [0, w] and v ranging from [0, h]; X / Z and Y / Z are the length and height coordinates of each human joint in the 2D human joint matrix after intrinsic parameter normalization, respectively.
[0056] Step 2.4: Based on the relative 3D human pose estimation network (specifically using the MixSTE network as the network architecture for the 2D to 3D pose enhancement network), normalize the input intrinsic parameters to obtain the 2D human joint matrix X. 2D-Norm Relative 3D human pose estimation is performed to obtain an N×3 human 3D pose matrix X in the camera coordinate system with the human hip joint as the origin. 3D ;
[0057] The 2D human joint matrix X is normalized with N×2 dimensional intrinsic parameters. 2D-Norm On the left, the N×3 dimensional human 3D pose matrix X 3D On the right, the 2D human joint matrix X after internal parameter normalization. 2D-Norm and human 3D pose matrix X 3D By concatenating the left and right sides, we obtain an N×5 human spatial position matrix X. input ;
[0058] 3D human pose estimation is a common computer vision task. This invention specifically refers to relative human pose estimation, which aims to locate and identify the three-dimensional coordinates (x, y, z) of human joints (such as head, left hand, right foot, etc.) in an estimated image or video. The coordinate system is the camera coordinate system, and the origin of the coordinate system is the position of the human hip joint (0, 0, 0).
[0059] The relative 3D human pose estimation network can be implemented using any network or method. The estimation result of 3D pose estimation will affect the final estimation result to a certain extent. This invention recommends using the MixSTE network for implementation.
[0060] Step 3: Calculate the human spatial position matrix X corresponding to each RGB video. inputThe T×5N input sequence is constructed by concatenating the following steps: The N×5 human spatial position matrix is expanded into a 1×5N unit sequence; then, the 1×5N unit sequences of adjacent frames are concatenated to form a multi-frame input sequence, i.e., the T×5N input sequence X. T-input Using an input sequence allows the network to make full use of temporal location information; where T is the preset number of sequences, and in this embodiment, the value is 91.
[0061] Step 4: Construct a spatial location estimation network (specifically using a Transform Encoder architecture network, such as...) Figure 2 As shown), the input sequence X is T×5N corresponding to each RGB video in the training set. T-input The input is the 3D position of the human hip joint in the camera coordinate system of each frame in each RGB video as the estimation target. The training objective is to minimize the mean square error loss between the network output and the true value of the estimation target. The spatial position estimation network is trained to obtain the trained spatial position estimation network.
[0062] like Figure 2 As shown, the spatial location estimation network includes a linear projection layer, a temporal location code, multiple identical Transformer Encoder blocks, and a linear regression layer in sequence.
[0063] The input sequence X of T×5N T-input The input is fed into a linear projection layer, mapping the 5N dimensions to a higher dimension. Temporal information is then provided through temporal location encoding. The resulting feature vector is processed by stacked Transformer Encoder blocks (where the attention operation has a T dimension) before being fed into a linear regression layer, regressing the high-dimensional features to a T×3 network output. The Transformer Encoder block includes a first normalization layer, a multi-head self-attention (MSA) layer, a second normalization layer, and a multilayer perceptron (MLP) layer. The structure of the multi-head self-attention layer is as follows: Figure 3 As shown, the structure of the multilayer perceptron is as follows: Figure 4 As shown;
[0064] The formula for calculating the mean squared error loss (L2 error) used during training is as follows:
[0065]
[0066] In the formula, P and P * These represent the true location and the network output estimated location of the human hip joint in the camera coordinate system, respectively.
[0067] Step 5: After processing the RGB video to be estimated in Step 2, the N×5 human spatial position matrix X corresponding to the RGB video to be estimated is obtained. input The T×5N input sequence obtained by splicing is input into the trained spatial location estimation network, which outputs the T×3 human spatial location in the RGB video to be estimated, that is, the 3D position of the human hip joint in the camera coordinate system in each frame.
[0068] Step 6: Calculate the step speed, specifically:
[0069] Based on the N×3 human 3D pose matrix corresponding to the RGB video to be estimated, the distance between the two ankles of the human body in each frame is calculated, resulting in a curve showing the distance changing with the frame number. All maxima of the curve are solved, and the specific frame and frame number corresponding to each maxima are obtained. Based on the 3D position of the human hip joint in the camera coordinate system output by the trained spatial position estimation network for each frame, the 3D position of the human hip joint in the camera coordinate system in a specific frame is obtained according to the frame number corresponding to each maxima. The distance traveled between adjacent specific frames is calculated, and then the step speed for each step is calculated.
[0070]
[0071] This embodiment uses the Human3.6M dataset as an example for verification and illustration, following the commonly used Protocol 1 rule, i.e., training is performed on a 3D human pose dataset filmed by five actors {S1, S5, S6, S7, S8}, and testing and error evaluation are performed on a dataset filmed by two actors {S9, S11}. The average position error of the human spatial position estimation method proposed in this embodiment is 23.2 cm. The spatial distance of human positions in the dataset is about 4-5 m, resulting in a relative position error of about 5%. The gait speed error derived from the position estimation is about 1%, indicating that the method in this embodiment has high estimation accuracy.
[0072] Although the illustrative specific embodiments of the present invention have been described above to enable those skilled in the art to understand the invention, it should be understood that the invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the invention as defined and determined by the appended claims, and all inventions utilizing the concept of the present invention are protected.
Claims
1. A method for estimating the spatial position of the human body in RGB video, characterized in that, Includes the following steps: Step 1: Construct a training set consisting of multiple RGB videos and obtain the camera intrinsics of the camera that captured the RGB videos, including a 3×3 projection matrix and a distortion matrix consisting of multiple meridional distortion parameters and tangential distortion parameters. Step 2: Calculate the human spatial position matrix for each RGB video, specifically: Step 2.1: Based on the 2D human pose estimation network, estimate the pixel coordinates of human joints in each frame of the RGB video in the pixel coordinate system to obtain an N×2 dimensional 2D human joint matrix before correction; where N is the total number of human joints. Step 2.2: Based on the distortion matrix, perform distortion correction on the 2D human joint matrix before correction to obtain an N×2 dimensional 2D human joint matrix after correction. Step 2.3: Based on the inverse of the projection matrix, perform intrinsic parameter normalization on the corrected 2D human joint matrix to obtain an N×2 dimensional intrinsic parameter normalized 2D human joint matrix. The formula for the intrinsic parameter normalization operation is as follows: In the formula, It is the inverse of the projection matrix. This represents the actual physical length of each pixel along the horizontal axis in the normalized coordinate system used for camera imaging. This represents the actual physical length of each pixel along the vertical axis in the normalized coordinate system used for camera imaging. The x-coordinate of the pixel at the center of the image. The ordinate of the pixel at the center of the image; and These are the length and height coordinates of each human joint in the corrected 2D human joint matrix, respectively. and These are the length and height coordinates of each human joint in the 2D human joint matrix after internal parameter normalization; Step 2.4: Based on the relative 3D human pose estimation network, perform relative 3D human pose estimation on the input intrinsic parameter normalized 2D human joint matrix to obtain an N×3 human 3D pose matrix in the camera coordinate system with the human hip joint as the origin; concatenate the N×2 dimensional intrinsic parameter normalized 2D human joint matrix with the N×3 dimensional human 3D pose matrix to obtain an N×5 human spatial position matrix. Step 3: Construct a spatial location estimation network, specifically using a Transformer Encoder architecture network, including a location encoding layer, a multi-head self-attention layer, and a multilayer perceptron. The human spatial location matrix corresponding to each RGB video in the training set is used as input, and the 3D position of the human hip joint in each frame of each RGB video in the camera coordinate system is used as the estimation target. The training objective is to minimize the mean squared error loss between the network output and the true value of the estimation target. The spatial location estimation network is trained to obtain the trained spatial location estimation network. Specifically, during training, the input is a T×5N input sequence composed of the human spatial position matrix corresponding to each RGB video. The specific splicing method is as follows: the N×5 human spatial position matrix is expanded into a 1×5N unit sequence, and then the 1×5N unit sequences of adjacent frames are spliced to form a multi-frame input sequence, namely the T×5N input sequence; where T is the preset number of sequences. Step 4: After the RGB video to be estimated is processed sequentially in Step 2, the N×5 human spatial position matrix corresponding to the RGB video to be estimated is obtained. The resulting T×5N input sequence is then fed into the trained spatial position estimation network, which outputs the human spatial position in the RGB video to be estimated, i.e., the 3D position of the human hip joint in each frame in the camera coordinate system.
2. The method for estimating the spatial position of a human body in RGB video according to claim 1, characterized in that, N=17, and the joints of each human body are the hip joint, right hip joint, right knee joint, right ankle joint, left hip joint, left knee joint, left ankle joint, thoracic spine joint, neck joint, nose joint, head joint, left shoulder joint, left elbow joint, left hand joint, right shoulder joint, right elbow joint, and right hand joint.
3. The method for estimating human spatial position in RGB video according to claim 1, characterized in that, In step 2.1, the 2D human pose estimation network is specifically a ViT network, a stacked hourglass network, or an Openpose network.
4. The method for estimating the spatial position of a human body in RGB video according to claim 1, characterized in that, In step 2.4, the relative 3D human pose estimation network is specifically the MixSTE network, the MotionBERT network, or the PoseFormer network.
5. The application of the human spatial position estimation method in RGB video as described in any one of claims 1 to 4 in gait rate derivation.
6. A method for deriving step speed, characterized in that, The specific process is as follows: According to any one of claims 1 to 4, the method for estimating the spatial position of the human body in RGB video obtains an N×3 human body 3D pose matrix corresponding to the RGB video to be estimated, and the 3D position of the human hip joint point in the camera coordinate system in each frame. Based on the N×3 human 3D pose matrix corresponding to the RGB video to be estimated, the distance between the two ankles of the human body in each frame is calculated, resulting in a curve showing the distance changing with the frame number. All maxima of the curve are solved, and the specific frame and frame number corresponding to each maxima are obtained. Based on the 3D position of the human hip joint in the camera coordinate system for each frame, and according to the frame number corresponding to each maxima, the 3D position of the human hip joint in the camera coordinate system for a specific frame is obtained. The distance traveled between adjacent specific frames is calculated, and then the step speed for each step is calculated. 。
Citation Information
Patent Citations
3D human body posture estimation model building method based on single-frame image and application of 3D human body posture estimation model building method
CN113192186A
Human body three-dimensional posture estimation method based on multi-view fusion
CN114529605A