A 3D Human Pose Detection Method Based on Monocular Vision

By using the YOLOX model and regression mapping neural network in 3D human posture detection, the 2D human joint node coordinates are converted into 3D joint node coordinates, which solves the problem of missed detection and small object detection in the existing methods, and achieves higher detection accuracy.

CN115497162BActive Publication Date: 2025-06-13NANTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211149451.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-21
Publication Date
2025-06-13
Estimated Expiration
2042-09-21

AI Technical Summary

Technical Problem

The existing 3D human posture detection methods have problems such as missed detection and inaccurate detection of small objects, and deep learning-based methods have poor detection of small objects and have failed detection.

Method used

The 3D human posture detection method based on monocular vision is adopted, and the YOLOX model is used for human detection, the 2D human joint node coordinate sequence is extracted, and the regression mapping neural network is constructed to convert it into 3D human joint node coordinates. The method includes multiple convolutional layers, attention mechanisms, and jump connections to extract and convert joint node coordinates.

Benefits of technology

It effectively solves the problems of missed detection and inaccurate detection of small targets, and improves the accuracy and reliability of 3D human posture detection, especially when the character target is small.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115497162B_ABST
    Figure CN115497162B_ABST
Patent Text Reader

Abstract

The present invention discloses a 3D human pose detection method based on monocular vision, including: 1. Using the YOLOX model to detect the people existing in the input image and obtaining the corresponding human detection frames; 2. Calculating the center point coordinates of the detection frames based on the human detection frames detected by YOLOX, and then magnifying the detected people to obtain a new human image; 3. Extracting the 2D human joint point coordinate sequence in the obtained new human image; 4. Constructing a regression mapping neural network to convert the 2D human joint point coordinate sequence into a 3D human joint point coordinate sequence; 5. Training the above regression mapping neural network on the publicly available Human3.6M dataset; 6. Testing the regression mapping neural network to detect the 3D human pose. This method uses the regression method to detect the human pose, solving the missed detection situation in the original method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a 3D human pose detection method based on monocular vision. Background Art

[0002] The development of human pose detection technology has become closer to reality, and related technologies are widely used in gait analysis, human-computer interaction, video surveillance and other fields, but there are still some challenges: First, the method of obtaining human pose data using wearable devices has strict requirements for the scene where the detected person is located, and the related wearable devices are expensive, so it is difficult to promote; Next, although the 2D human pose detection method based on deep learning technology can obtain relatively accurate human pose data, due to the lack of depth information, it is impossible to judge the distance between the detected person and the camera, and there is also a problem that the predicted 2D human pose joint points cannot accurately correspond to the actual 3D human pose joint points.

[0003] Currently, the methods for 3D human pose detection are mainly divided into two categories: one is to use wearable devices to obtain motion data by equipping relevant sensors on the detected person, but this method has extremely high requirements for the detection environment and equipment, and the cost is high and the process is complex; the other is the detection method based on deep learning, which can make good use of the deep neural network, learn data features and continuously optimize the network model to perform human pose detection, but currently this method has poor detection effect on small target poses and there are cases of detection failure. Summary of the Invention

[0004] The purpose of the present invention is to provide a 3D human pose detection method based on monocular vision to solve the problems of missed detection and inaccurate small target detection in current 3D human pose detection.

[0005] To solve the above problems, the specific technical solution of the present invention is as follows:

[0006] A 3D human pose detection method based on monocular vision, comprising the following steps:

[0007] Step 1, use the YOLOX model to detect the people existing in the input image and obtain the corresponding human detection boxes;

[0008] Step 2, calculate the center point coordinates of the human detection box, and then enlarge the detected person to obtain a new human image;

[0009] Step 3, extract the 2D human joint point coordinate sequence in the obtained new human image;

[0010] Step 4, construct a regression mapping neural network to convert the 2D human joint point coordinate sequence into a 3D human joint point coordinate sequence;

[0011] Step 4.1: Use a conventional one-dimensional convolution with a kernel size of 3 to extract the features of the input 2D human joint coordinate sequence;

[0012] Step 4.2: Input the features into the scaling module containing the CoordAttention attention mechanism that is sensitive to both position and direction to further extract features; this module contains two depthwise convolutions with a kernel size of 3, where an improved one-way CoordAttention mechanism module is added after the first convolution to obtain features that strengthen cross-channel information, direction information, and position information. The dilation coefficient of the second convolution is 3, and two one-dimensional convolutions with a kernel size of 1 are added between the two depthwise convolutions to adjust the number of channels of the features;

[0013] Step 4.3: Input the features obtained in Step 4.1 into the temporal convolution module containing the spatial attention mechanism to further extract features; this module contains a one-dimensional convolution with a kernel size of 3 and a dilation coefficient of 3, a one-dimensional convolution with a kernel size of 1, and a spatial attention mechanism module, which are connected in sequence;

[0014] Step 4.4: Concatenate the features obtained in Step 4.1 with the features obtained in Steps 4.2 and 4.3 through skip connection, and then use a convolution with a kernel size of 1 to adjust the number of channels of the concatenated features;

[0015] Step 4.5: Stack several block structures (each block structure consists of the above Steps 4.2, 4.3, and 4.4) according to different receptive fields to form the backbone of the regression mapping neural network;

[0016] Step 4.6: Use a one-dimensional convolution with a kernel size of 1 after the last block to achieve the prediction of 3D human joint coordinates;

[0017] Step 5: Train the above regression mapping neural network on the Human3.6M public dataset. The regression mapping neural network uses the mean per-joint position error (MPJPE) of each joint, which is commonly used in the field of 3D human pose detection, as the loss function. The formula is as follows:

[0018]

[0019] where N represents the number of skeleton joint points in the detected video frame; represents the predicted value of the i-th joint point of the skeleton in the detected video frame, J i represents the true value of the corresponding joint point;

[0020] Step 6: Test the regression mapping neural network to obtain the 3D human pose.

[0021] Furthermore, step 1 specifically includes the following steps:

[0022] Step 1.1: Use the weights trained on the ImageNet dataset as pre-trained weights, and then train the YOLOX model on the self-owned dataset for 100 epochs with a learning rate of 0.001, and select Adam as the optimizer;

[0023] Step 1.2: Use the trained YOLOX model to detect the people existing in the input image and obtain the corresponding person detection boxes.

[0024] Furthermore, step 2 specifically includes the following steps:

[0025] Calculate the center point according to the person detection box, and then calculate a new person detection box according to the aspect ratio of 16:9 based on the height, and enlarge it to 1920*1080 to obtain a new person image; the height and width of the new person detection box are respectively:

[0026] h = h, w = (16h) / 9.

[0027] Use the classic stacked hourglass model to extract the 2D human joint point coordinate sequence from the obtained new person image.

[0028] Furthermore, in step 3, use the classic stacked hourglass model to extract the 2D human joint point coordinate sequence from the obtained new person image.

[0029] Furthermore, step 5 specifically includes the following steps: The regression mapping neural network is trained on the Human3.6M dataset, which contains a total of seven groups of data, namely S1, S5, S6, S7, S8, S9, and S11. Among them, S1, S5, S6, S7, and S8 are used as the training set, and S9 and S11 are used as the test set. The network is trained for 80 epochs in total, Adam is selected as the optimizer, the initial learning rate is 0.001, and an exponentially decaying learning rate strategy is adopted. A shrinkage factor of 0.95 is applied for each epoch; after passing through the depth convolution in the scaling module, two convolutions with a kernel size of 1, the dilated convolution in the temporal convolution module, and the convolution with a kernel size of 1, batch normalization, ReLU, and Dropout are used for processing in turn. The momentum of batch normalization is 0.1 and decays exponentially, reaching 0.001 in the last epoch, and the Dropout value is 0.25.

[0030] A 3D human pose detection method based on monocular vision of the present invention has the following advantages:

[0031] The present invention proposes a 3D human pose detection method based on monocular vision, which uses an object detection algorithm for human detection to solve the problem of missed detection in the original method. At the same time, different from the method of directly obtaining the 3D human pose from the input image, this method adopts a two-stage approach. First, it uses an excellent 2D human pose detection network to estimate the 2D human joint point coordinate sequence, and then uses a regression mapping neural network to convert it into the 3D human joint point coordinates. Description of the Drawings

[0032] Figure 1 It is the processing flow chart of the 3D human pose detection method based on monocular vision of the present invention;

[0033] Figure 2 It is the structure diagram of the regression mapping neural network when the receptive field of the present invention is 27;

[0034] Figure 3 It is the schematic diagram of the unidirectional CoordAttention module of the present invention;

[0035] Figure 4 It is the schematic diagram of the spatial attention module of the present invention;

[0036] Figure 5 It is the comparison diagram of the test results of the present invention;

[0037] Figure 6 It is the comparison result of the present invention with the classical Videopose algorithm when the receptive field is 27. Detailed Embodiment

[0038] In order to better understand the purpose, structure and function of the present invention, the following further describes in detail a 3D human pose detection method based on monocular vision of the present invention with reference to the drawings.

[0039] Step 1: Use the weights trained on the ImageNet dataset as the pre-trained weights, and then train the YOLOX model on the self-owned dataset for 100 epochs with a learning rate of 0.001, and select Adam as the optimizer;

[0040] Use the trained YOLOX model to detect the people existing in the input image and obtain the corresponding person detection boxes;

[0041] Step 2: Calculate the center point of the person detection box according to the person detection box, and then calculate a new person detection box according to the aspect ratio of 16:9 based on the height, and enlarge it to 1920*1080 to obtain a new person image; the height and width of the new person detection box are respectively:

[0042] h = h, w = (16h) / 9;

[0043] Step 3: Use the classic stacked hourglass model to extract the 2D human joint point coordinate sequence from the obtained new human image;

[0044] Step 4: Construct a regression mapping neural network to convert the 2D human joint point coordinate sequence into a 3D human joint point coordinate sequence;

[0045] Step 4.1: Use a conventional one-dimensional convolution with a kernel size of 3 to extract the feature x of the obtained 2D human joint point coordinate sequence, and then perform operations such as batch normalization, ReLU, and Dropout;

[0046] Step 4.2: Use a scaling module containing the improved one-way CoordAttention mechanism to extract deep information from the input feature x. First, use a depth convolution with a kernel size of 3 to further extract deep features from the input feature; then use the improved one-way CoordAttention mechanism to obtain direction and position information while capturing cross-channel information, that is, perform a global average pooling operation, and then use two one-dimensional convolutions with a kernel size of 1. After the first convolution, perform batch normalization and ReLU operations, and then use the Sigmoid function to convert it into a weight; then multiply the weight by the input feature and input it into the scaling structure to adjust the channel number of the input feature. This structure consists of two one-dimensional convolutions with a kernel size of 1. The first convolution reduces the channel number of the input feature to 1 / 64 of the original, and the second convolution expands the channel number back to the original channel number; finally, use a depth convolution with a kernel size of 3 and a dilation coefficient of 3 to extract features to obtain the feature x';

[0047] Step 4.3: Use a time-domain convolution module containing a spatial attention mechanism to further extract features from the input feature x. First, use a one-dimensional convolution with a kernel size of 3 and a dilation coefficient of 3 to extract features from the input feature x; then perform a one-dimensional convolution with a kernel size of 1; then use the spatial attention module to strengthen the spatial distribution information of the 3D human joint coordinates, that is, first predict a 3D human joint coordinate from the input feature, then use a two-dimensional convolution with a kernel size of 1 to extract features, and then obtain two feature maps through global average pooling and global max pooling. Then, splice the two feature maps and input them into a two-dimensional convolution with a kernel size of 3, and then use the Sigmoid function to convert it into a weight; then multiply the weight by the predicted 3D human joint point coordinate sequence and restore the result to temporal information; finally, adjust the channel number through a one-dimensional convolution with a kernel size of 1 to obtain the feature x″;

[0048] Step 4.4: Concatenate the feature x obtained in the first step with the features x' and x″ obtained in the second and third steps in the channel dimension through a skip connection, and then use a one-dimensional convolution with a kernel size of 1 to reduce the channel number of the concatenated feature to 1 / 3;

[0049] Step 4.5: Different numbers of block structures composed of the above second, third, and fourth steps can be stacked according to different sizes of receptive fields, so as to obtain the backbone of the regression mapping neural network. For example, when the receptive field is 27, the backbone of the network contains two block structures, where the dilation coefficient in the first block is 3 and the dilation coefficient in the second block is 9;

[0050] Step 4.6: After the last block, a one-dimensional convolution with a kernel size of 1 is used to predict the 3D human joint coordinates.

[0051] Step 5: Train the above regression mapping neural network on the Human3.6M public dataset, and use the 2D human joint coordinate sequence extracted from the new human images obtained by the classic stacked hourglass model as the input. The regression mapping neural network uses the mean per-joint position error (MPJPE) commonly used in the field of 3D human pose detection as the loss function, and the formula is as follows:

[0052]

[0053] where N represents the number of skeleton joint points in the detected video frame; represents the predicted value of the i-th joint point of the skeleton in the detected video frame, and J i represents the true value of the corresponding joint point;

[0054] Step 6: Test the regression mapping neural network to obtain the 3D human pose. The test results are as Figure 5 shown. Compared with other detection algorithms, in the case of a smaller person target, the algorithm proposed by the present invention has higher accuracy and more accurate recognition. The comparison results of the present invention with the classic Videopose algorithm when the receptive field is 27 are as Figure 6 shown.

[0055] It can be understood that the present invention is described through some embodiments. Those skilled in the art know that without departing from the spirit and scope of the present invention, various changes or equivalent replacements can be made to these features and embodiments. In addition, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.

Claims

1. A 3D human pose detection method based on monocular vision, characterized in that, it includes the following steps: Step 1: Use the YOLOX model to detect the people existing in the input image and obtain the corresponding person detection boxes; Step 2: Calculate the center point coordinates of the person detection boxes, and then enlarge the detected people to obtain new person images; Step 3: Extract the 2D human joint point coordinate sequence in the obtained new person images; Step 4: Construct a regression mapping neural network to convert the 2D human joint point coordinate sequence into a 3D human joint point coordinate sequence; Step 4.1: Use a conventional one-dimensional convolution with a kernel size of 3 to extract the features of the input 2D human joint point coordinate sequence; Step 4.2: Input the features into a scaling module containing the CoordAttention attention mechanism that is sensitive to both position and direction to further extract features; This module contains two depthwise convolutions with a kernel size of 3. After the first convolution, an improved one-way CoordAttention mechanism module is added to obtain features that strengthen cross-channel information, direction information, and position information. The dilation coefficient of the second convolution is 3, and two one-dimensional convolutions with a kernel size of 1 are added between the two depthwise convolutions to adjust the number of channels of the features; Step 4.3: Input the features obtained in Step 4.1 into a dilated convolution module containing a spatial attention mechanism to further extract features; This module contains a one-dimensional convolution with a kernel size of 3 and a dilation rate of 3, a one-dimensional convolution with a kernel size of 1, and a spatial attention mechanism module, which are connected in sequence; Step 4.4: Concatenate the features obtained in Step 4.1 with the features obtained in Steps 4.2 and 4.3 through skip connections, and then use a convolution with a kernel size of 1 to adjust the number of channels of the concatenated features; Step 4.5: Stack the block structures composed of the above Steps 4.2, 4.3, and 4.4 according to different receptive fields to obtain the backbone of the regression mapping neural network; Step 4.6: Use a one-dimensional convolution with a kernel size of 1 after the last block to achieve the prediction of the 3D human joint point coordinates; Step 5: Train the above regression mapping neural network on the Human3.6M public dataset. The regression mapping neural network uses the mean per-joint position error (MPJPE) of each joint commonly used in the field of 3D human pose detection as the loss function, and the formula is as follows: Where N represents the number of skeleton joint points in the detected video frame; represents the predicted value of the i-th joint point of the skeleton in the detected video frame, J i represents the true value of the corresponding joint point; Step 6: Test the regression mapping neural network to obtain the 3D human pose.

2. The 3D human pose detection method based on monocular vision according to claim 1, characterized in that, Step 1 specifically includes the following steps: Step 1.1: Use the weights trained on the ImageNet dataset as the pre-trained weights, and then train the YOLOX model on the self-owned dataset for 100 epochs with a learning rate of 0.001, and select Adam as the optimizer; Step 1.2: Use the trained YOLOX model to detect the people existing in the input image and obtain the corresponding person detection boxes.

3. The 3D human pose detection method based on monocular vision according to claim 1, It is characterized in that Step 2 specifically includes the following steps: Calculate the center point according to the person detection box, then calculate a new person detection box according to the aspect ratio of 16:9 based on the height, and enlarge it to 1920*1080 to obtain a new person image; the height and width of the new person detection box are respectively: h = h, w = (16h) / 9; Use the classic stacked hourglass model to extract the 2D human joint point coordinate sequence from the obtained new person image.

4. The monocular vision-based 3D human pose detection method according to claim 1, It is characterized in that In step 3, the classic stacked hourglass model is used to extract the 2D human joint point coordinate sequence from the obtained new person image.

5. The monocular vision-based 3D human pose detection method according to claim 1, It is characterized in that Step 5 specifically includes the following steps: The regression mapping neural network is trained on the Human3.6M dataset, which contains a total of seven groups of data, namely S1, S5, S6, S7, S8, S9, and S11. Among them, S1, S5, S6, S7, and S8 are used as the training set, and S9 and S11 are used as the test set. The network is trained for 80 epochs in total. The optimizer is selected as Adam, the initial learning rate is 0.001, and an exponentially decaying learning rate strategy is adopted, with a shrinkage factor of 0.95 applied to each epoch; after passing through the depth convolution in the scaling module, two convolutions with a kernel size of 1, the dilated convolution in the temporal convolution module, and the convolution with a kernel size of 1, batch normalization, ReLU, and Dropout are used for processing in sequence. The momentum of batch normalization is 0.1 and decays exponentially, reaching 0.001 in the last epoch, and the Dropout value is 0.25.