A gesture estimation method based on dynamic and static anchor points
By setting static anchor points on depth map samples and building gesture estimation networks for dynamic static anchor points, the existing 3D gesture estimation methods are solved, and efficient gesture estimation and spatial relationship constraints are achieved.
Patent Information
- Application Number
- CN202211462570.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-22
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-11-22
AI Technical Summary
The existing 3D gesture estimation method takes a long time to process point clouds and requires high performance requirements for computer equipment, and ignores the positional relationship between hand joint nodes.
A gesture estimation method based on dynamic and static anchor points is proposed. By setting static anchor points on the depth map sample and building a gesture estimation network of dynamic and static anchor points, including feature extraction module, static anchor point estimation module, static anchor point weight estimation module, dynamic anchor point weight estimation module and dynamic anchor point correction module, we learn the mutual position relationship between hand joint nodes.
This method effectively constrains the gesture estimation process spatial relationship, improves estimation efficiency, and reduces the requirements for computer equipment performance.
Smart Images

Figure CN115862058B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of 3D gesture estimation, and specifically to a gesture estimation method based on dynamic and static anchor points. Background Art
[0002] 3D gesture estimation is a hot topic in the field of computer vision. It has a wide range of applications in human-computer interaction, behavior analysis and other fields. With the development of depth cameras, 3D gesture estimation based on depth maps has attracted more and more attention from scholars. Compared with 3D gesture estimation methods based on RGB images, 3D gesture estimation methods based on depth maps do not obtain information about a person's appearance and are less likely to cause privacy leakage. At the same time, compared with RGB images, depth maps are less susceptible to interference from environmental factors and can provide richer 3D information.
[0003] Most existing methods design deep convolutional neural networks with various structures to extract features contained in depth maps and perform 3D hand gesture estimation. Most of these methods focus on converting depth maps into point clouds and extracting hand features from multiple perspectives, while ignoring the positional relationship between hand joints. Moreover, the process of processing point clouds takes a lot of time and has high requirements on the performance of computer equipment.
[0004] Therefore, to address the problems of the above 3D gesture estimation algorithm, a gesture estimation method based on dynamic and static anchor points is proposed. Summary of the invention
[0005] In order to solve the problems existing in the background technology, the present invention aims to provide a gesture estimation method based on dynamic and static anchor points to solve the above-mentioned situation.
[0006] Solutions used to solve the problem:
[0007] A gesture estimation method based on dynamic and static anchor points, the method comprising:
[0008] S1: Setting a static anchor point on the depth map sample to obtain the depth map sample after the static anchor point is set;
[0009] S2: Based on the depth map samples after setting the static anchor points, a gesture estimation network of dynamic and static anchor points is constructed; the gesture estimation network of dynamic and static anchor points includes: a feature extraction module, a static anchor point estimation module, a static anchor point weight estimation module, a dynamic anchor point weight estimation module and a dynamic anchor point correction module;
[0010] S3: Train the gesture estimation network based on dynamic and static anchor points until convergence, and obtain a trained dynamic and static anchor point gesture estimation network;
[0011] S4: Input the tested depth map samples into the trained dynamic and static anchor gesture estimation network to achieve gesture estimation.
[0012] Furthermore, the gesture estimation network of constructing dynamic and static anchor points specifically includes:
[0013] Input the depth map sample into the feature extraction module to obtain the depth map features;
[0014] The depth map features are input into the static anchor point estimation module to obtain the estimation results of the static anchor point for the 2D plane coordinates of the joint point and the estimation results of the static anchor point for the depth coordinates of the joint point;
[0015] Input the depth map features into a preset static anchor point weight estimation module to obtain the normalized weight of the static anchor point to the joint point;
[0016] Based on the normalized weights of the static anchor points to the joint points, weighted summation is performed on the estimation results of the static anchor points to the 2D plane coordinates of the joint points to obtain a preliminary estimation result of the 2D plane coordinates of the joint points;
[0017] Based on the normalized weights of the static anchor points to the joint points, the estimated results of the depth coordinates of the static anchor points to the joint points are weighted summed to obtain the final estimated results of the depth coordinates of the joint points;
[0018] The estimation result of the 2D plane coordinates of the joint points by the static anchor points is input into the preset dynamic anchor point correction model to obtain the correction result of the 2D plane coordinates of the joint points by the dynamic anchor points;
[0019] Input the preliminary estimation result of the 2D plane coordinates of the joint point into the preset dynamic anchor point weight estimation model to obtain the normalized weight of the dynamic anchor point to the joint point;
[0020] Based on the normalized weights of the dynamic anchor points to the joint points, weighted summation is performed on the correction results of the dynamic anchor points to the 2D plane coordinates of the joint points to obtain the correction results of the 2D plane coordinates of the joint points;
[0021] Based on the normalized weights of the static anchor points on the joint points, the correction results of the 2D plane coordinates of the joint points are weighted summed to obtain the final estimation results of the 2D plane coordinates of the joint points.
[0022] Further, based on any input depth map sample I, is the matrix representation of the depth map sample I, W and H correspond to the number of rows and columns of the depth map sample matrix representation, respectively. The representation matrix is a real number matrix; the depth map sample I is marked by the 2D plane coordinates and depth coordinates of N joint points:
[0023]
[0024]
[0025] in, Represents the true value of the 2D plane coordinates of the nth joint point, Represents the true value of the depth coordinate of the nth joint point;
[0026] On the depth map sample I, a static anchor point is set every h pixels, and a total of M static anchor points are set; the depth map sample I is input into the feature extraction module to obtain the depth map features; the structure of the feature extraction module is consistent with that of ResNet-50, which consists of a series of convolutional layers and pooling layers; the output of the feature extraction module is the depth map features Where W 1 , H 1 , D represent the height, width and number of channels of the output feature map respectively; the depth map feature F has a total of W 1 ×H 1 pixels, the feature vector f of each pixel i The dimension is D, i = 1, 2, ..., W 1 ×H 1 ;
[0027] The input of the static anchor estimation module is the depth map feature F, and the output is the estimation result S of the 2D plane coordinates of the static anchor point to the joint point xy And the estimation result S of the depth coordinate of the joint point by the static anchor point z ; is the matrix representation of the estimation result of the static anchor point on the 2D plane coordinates of the joint point, M, N and 2 correspond to the number of rows, columns and channels of the matrix representation of the estimation result of the static anchor point on the 2D plane coordinates of the joint point respectively; (S xy ) a,b For S xy The value of row a and column b in (S xy ) a,b It represents the estimated result of the 2D plane coordinates of the bth joint point by the ath static anchor point. It is the matrix representation of the estimated result of the 2D plane coordinates of the bth joint point by the ath static anchor point; is the matrix representation of the estimation result of the depth coordinates of the joint points by the static anchor points, and M, N and 1 correspond to the number of rows, columns and channels of the matrix representation of the estimation result of the depth coordinates of the joint points by the static anchor points, respectively; (S z ) a,b For S z The value of row a and column b in (S z ) a,b Represents the estimated result of the depth coordinate of the bth joint point by the ath static anchor point. It is the matrix representation of the estimated result of the depth coordinate of the bth joint point by the ath static anchor point;
[0028] The structure of the static anchor point estimation module: includes a 2D plane coordinate estimation module and a depth estimation module; the 2D plane coordinate estimation module contains convolution layer 1, convolution layer 2, convolution layer 3, convolution layer 4, convolution layer 5 and transformation layer 1; convolution layer 1 has 256 convolution kernels, and the size of the convolution kernel is 3×3; the input of convolution layer 1 is the depth map feature F, and the output is P 1 ; Convolutional layer 2 has 256 convolution kernels, and the size of the convolution kernel is 3×3; The input of convolutional layer 2 is P 1 , the output is P 2 ; Convolution layer 3 has 256 convolution kernels, and the size of the convolution kernel is 3×3; the input of convolution layer 3 is P 2 , the output is P 3 ; Convolution layer 4 has 256 convolution kernels, and the size of the convolution kernel is 3×3; the input of convolution layer 4 is P 3 , the output is P 4 ; Convolutional layer 5 has H 1 convolution kernels, the size of the convolution kernel is 3×3; the input of convolution layer 5 is P 4 , the output is P 5 ; Conversion layer 1 is used to change P 5 The shape of the transformation layer 1 is P 5 , the output is S xy ; The depth estimation module includes convolution layer 6, convolution layer 7, convolution layer 8, convolution layer 9, convolution layer 10 and transformation layer 2; convolution layer 6 has 256 convolution kernels, and the size of the convolution kernel is 3×3; the input of convolution layer 6 is the depth map feature F, and the output is P 6 ; Convolution layer 7 has 256 convolution kernels, and the size of the convolution kernel is 3×3; the input of convolution layer 7 is P 6 , the output is P 7 ; Convolution layer 8 has 256 convolution kernels, and the size of the convolution kernel is 3×3; the input of convolution layer 8 is P 7 , the output is P 8 ; Convolution layer 9 has 256 convolution kernels, and the size of the convolution kernel is 3×3; the input of convolution layer 9 is P 8 , the output is P 9 ; Convolutional layer 10 has H 2 convolution kernels, the size of the convolution kernel is 3×3; the input of convolution layer 10 is P 9 , the output is P 10 ; Conversion layer 2 is used to change P 10 The shape of the transformation layer 2 is P 10 , the output is S z .
[0029] Furthermore, the input of the static anchor weight estimation module is the depth map feature F, and the output is the normalized weight W of the static anchor to the joint point, where: It is the matrix representation of the normalized weights of the static anchor points to the joint points;
[0030] The structure of the static anchor weight estimation module includes convolution layer 11, convolution layer 12, convolution layer 13, convolution layer 14, convolution layer 15, transformation layer 3 and normalization layer; convolution layer 11 has 256 convolution kernels, the size of the convolution kernel is 3×3, the input of convolution layer 11 is the depth map feature F, and the output is P 11 ; Convolution layer 12 has 256 convolution kernels, the size of the convolution kernel is 3×3, and the input of convolution layer 12 is P 11 , the output is P 12 ; Convolution layer 13 has 256 convolution kernels, the size of the convolution kernel is 3×3, and the input of convolution layer 13 is P 12 , the output is P 13 ; Convolution layer 14 has 256 convolution kernels, the size of the convolution kernel is 3×3, and the input of convolution layer 14 is P 13 , the output is P 14 ; Convolutional layer 15 has H 3 convolution kernels, the size of the convolution kernel is 3×3, and the input of convolution layer 15 is P 14 , the output is P 15 ; Conversion layer 3 is used to change P 15 The shape of the transformation layer 3 is P 15 , the output is the weight of the static anchor point to the joint point in, is the matrix representation of the weight of the static anchor point to the joint point, for The value of column b in represents the weight of M static anchor points to the bth joint point, is the matrix representation of the weights of M static anchor points to the bth joint point, for The value of row a and column b in represents the weight of the a-th static anchor point to the b-th joint point, is the matrix representation of the weight of the a-th static anchor point to the b-th joint point; the activation function of the normalization layer is the softmax function, and the input of the normalization layer is The output is the normalized weight W of the static anchor point to the joint point. is the normalized weight matrix representation of the static anchor point to the joint point, W a,b is the value of row a and column b in W, W a,b represents the normalized weight of the a-th static anchor point to the b-th joint point, is the matrix representation of the normalized weight of the a-th static anchor point to the b-th joint point, W :,b is the value of column b in W, W :,bis the normalized weight of M static anchor points to the bth joint point, is the matrix representation of the normalized weights of M static anchor points on the bth joint point, W :,b The calculation method is as follows:
[0031]
[0032] in, Represents the softmax activation function.
[0033] Furthermore, the estimation results of the 2D plane coordinates of the joint points by the static anchor points are weighted and summed to obtain the preliminary estimation result G of the 2D plane coordinates of the joint points. is the matrix representation of the preliminary estimation result of the 2D plane coordinates of the joint point, (G) b,: is the value of the b-th row in G, which represents the preliminary estimate of the 2D plane coordinates of the b-th joint point, and is calculated as follows:
[0034]
[0035] Where M represents the number of static anchor points, W a,b *(S xy ) a,b Represents the weighted result of the static anchor point a on the 2D plane coordinates of the joint point b;
[0036] The estimated results of the depth coordinates of the joint points by the static anchor points are weighted summed to obtain the final estimated result Z of the depth coordinates of the joint points; is the matrix representation of the final estimated result of the joint point depth coordinate, (Z) b,: is the value of the b-th row in Z, which represents the final estimated result of the depth coordinate of the b-th joint point, and is calculated as follows:
[0037]
[0038] Where M represents the number of static anchor points, W a,b *(S z ) a,b Represents the weighted result of the depth coordinate of the static anchor point a to the joint point b.
[0039] Furthermore, the estimation result of the static anchor point on the 2D plane coordinates of the joint point is input into the dynamic anchor point correction module to obtain the correction result of the dynamic anchor point on the 2D plane coordinates of the joint point. In a depth map sample, each joint point is defined as a dynamic anchor point. The input of the dynamic anchor point correction module is the estimation result S of the static anchor point on the 2D plane coordinates of the joint point xy , the output is the correction result Q of the dynamic anchor point to the 2D plane coordinates of the joint point;
[0040] The structure of the dynamic anchor correction module includes a conversion layer 4, a copy layer, a cascade layer, a convolution layer 16, and a merging layer. The conversion layer 4 is used to change S xy The shape of the transformation layer 4 is S xy , the output is R, where is the matrix representation of the output of transformation layer 4, W 2 , N and D 2 The number of rows, columns, and channels of the matrix representation corresponding to the output of transformation layer 4, R :,:,d represents the value of the dth channel in R, is the matrix representation of the dth channel in R. Convolutional layer 16 has H 4 convolution kernels, the size of the convolution kernel is 3×3, the input of convolution layer 16 is R, and the output is P 16 , is the matrix representation of the output of convolutional layer 16, W 2 , N and D 2 ×N represent the number of rows, columns and channels of the matrix representation of the output of convolution layer 16 respectively; the input of the replication layer is R, and the replication layer is used to replicate each channel of R N times. The replication layer has D 2 The outputs are Among them, R d Indicates that the value R of the dth channel in R :,:,d Copy the result N times, is the matrix representation of the result of replicating the value of the dth channel in R by N times; the input of the cascade layer is The output is O, and the expression of O is as follows:
[0041]
[0042] Among them, [·,·] means connection by channel, is the matrix representation of the output of the cascade layer; the input of the merging layer is O and P 16 , the output is the correction result Q of the dynamic anchor point to the 2D plane coordinates of the joint point. The calculation process of Q is as follows:
[0043] Q=O+P 16 ;
[0044] Among them, + means adding elements at corresponding positions. is the matrix representation of the correction result of the dynamic anchor point on the 2D plane coordinates of the joint point, W 2 , N and (D 2 ×N) is the number of rows, columns and channels of the matrix representation of the correction result of the dynamic anchor point on the 2D plane coordinates of the joint point.
[0045] Furthermore, the preliminary estimation results of the 2D plane coordinates of the joint points are input into the dynamic anchor point weight estimation module to obtain the normalized weights of the dynamic anchor points to the joint points;
[0046] The structure of the dynamic anchor weight estimation module includes: linear layer, 1D convolution layer 1, 1D convolution layer 2, 1D convolution layer 3, 1D convolution layer 4, 1D convolution layer 5 and normalization layer 1; the number of neurons in the linear layer is 256, the input of the linear layer is G, and the output is Q 1 , Q 1 The calculation steps are as follows:
[0047]
[0048] Among them, W 1 are the weights of the linear layer, is the bias vector of the linear layer; the input of 1D convolutional layer 1 is Q 1 , the output is Q 2 , 1D convolution layer 1 has 256 convolution kernels, and the size of the convolution kernel is 3×1; the input of 1D convolution layer 2 is Q 2 , the output is Q 3 , 1D convolution layer 2 has 128 convolution kernels, and the size of the convolution kernel is 3×1; the input of 1D convolution layer 3 is Q 3 , the output is Q 4 , 1D convolution layer 3 has 64 convolution kernels, and the size of the convolution kernel is 3×1; the input of 1D convolution layer 4 is Q 4 , the output is Q 5 , 1D convolution layer 4 has 32 convolution kernels, the size of the convolution kernel is 3×1; the input of 1D convolution layer 5 is Q 5 , the output is the weight Q of the dynamic anchor point to the joint point 6 , 1D convolution layer 5 has 16 convolution kernels, and the size of the convolution kernel is 3×1; is the matrix representation of the weight of the dynamic anchor point to the joint point (Q 6 ) :,b Indicates Q 6 The value of column b of is the matrix representation of the weights of 16 dynamic anchor points on the bth joint point (Q 6 ) f,b Indicates Q 6 The value of the fth row and bth column of The activation function of the normalization layer 1 is the softmax function, and the input of the normalization layer 1 is Q 6 , the output is the normalized weight Q of the dynamic anchor point to the joint point 7 , is the matrix representation of the normalized weights of the dynamic anchor points to the joint points, (Q 7) f,b Indicates Q 7 The value of the fth row and bth column of represents the normalized weight of the fth dynamic anchor point to the bth joint point, (Q 7 ) :,b Q 7 The value of column b of is the matrix representation of the normalized weights of N dynamic anchor points to the bth joint point, and its calculation process is as follows:
[0049]
[0050] in, Represents the softmax activation function.
[0051] Further, weighted summation is performed on the correction results of the dynamic anchor points on the 2D plane coordinates of the joint points to obtain the correction result H of the 2D plane coordinates of the joint points; is the matrix representation of the correction result of the 2D plane coordinates of the joint point. The calculation steps of H are as follows:
[0052] 1) The matrix shape of the correction result Q of the dynamic anchor point to the 2D plane coordinates of the joint point is transformed into N×N×(M×2). The calculation steps are as follows:
[0053] Q 1 =Re(Q),
[0054] in, is the matrix representation of the correction result of the dynamic anchor point on the 2D plane coordinates of the joint point after the shape transformation; Re(.) is the shape transformation function; (Q 1 ) f,b Q 1 The value of the fth row and the bth column, (Q 1 ) f,b It represents the correction result of the 2D plane coordinates of the bth joint point by the fth dynamic anchor point. is the matrix representation of the correction result of the 2D plane coordinates of the bth joint point by the fth dynamic anchor point;
[0055] 2) Correction result Q of the 2D plane coordinates of the joint point for the dynamic anchor point after the shape transformation 1 The weighted summation is performed to obtain the correction result H of the 2D plane coordinates of the joint point. The calculation steps are as follows:
[0056]
[0057] in, It is the matrix representation of the correction result of the 2D plane of the b-th joint point.
[0058] Further, the correction results of the 2D plane coordinates of the joint points are weighted and summed to obtain the final estimation result J of the 2D plane coordinates of the joint points; is the matrix representation of the final estimated result of the 2D plane coordinates of the joint point. The calculation process of J is as follows:
[0059]
[0060] Among them, (H) a,b is the value in row a and column b of H, (J) b is the value of the b-th row in J, and is also the final estimated result of the 2D plane coordinates of the b-th joint point.
[0061] Furthermore, the gesture estimation network based on dynamic and static anchors is trained until convergence. The input of the network is the depth map sample I, and the output is the final estimation result J of the 2D plane coordinates of the joint point and the final estimation result Z of the depth coordinate. For V training samples, the loss function L used by the training network is:
[0062]
[0063] Among them, α represents the balance coefficient, [L xy ] v is the 2D plane loss function of the vth sample, [L z ] v is the depth loss function, [L xy ] v and [L z ] v The calculation method is as follows:
[0064]
[0065]
[0066] Where N represents the number of joint points in each training sample, [(J) n ] v represents the final estimated result of the 2D plane coordinates of the nth joint point in the vth sample, Represents the true value of the 2D plane coordinates of the nth joint point in the vth sample, [(Z) n,: ] v represents the final estimated result of the depth coordinate of the nth joint point in the vth sample, represents the true value of the depth coordinate of the nth joint point in the vth sample; L τ1 (.) represents the L1-smooth loss function with parameter τ1, which is defined as follows:
[0067]
[0068] L τ2 (.) represents the L1-smooth loss function with parameter τ2, which is defined as follows:
[0069]
[0070] Beneficial effects:
[0071] The beneficial effects of the above technical solution are: the proposed gesture estimation network based on dynamic and static anchor points can learn the relative position relationship between the hand joints, and effectively constrain the gesture estimation process with spatial relationships; during the network training process, only one depth map is used as the input of the network, and the performance requirements of the computer equipment are not high. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 This is the static anchor point estimation module diagram;
[0073] Figure 2 It is the static anchor weight estimation module;
[0074] Figure 3 It is a dynamic anchor point correction module;
[0075] Figure 4 It is a dynamic anchor weight estimation module;
[0076] Figure 5 A gesture estimation network based on dynamic and static anchors;
[0077] Figure 6 Build a flow chart for the model of the present invention;
[0078] Figure 7 is the input depth map sample 1;
[0079] Figure 8 is a schematic diagram of the final estimation result of the 2D plane coordinates and the final estimation result of the depth coordinates of the joint point in the input depth map sample 1;
[0080] Fig. 9 is the input depth map sample 2;
[0081] Fig.10 Schematic diagram of the final estimation result of the 2D plane coordinates and the final estimation result of the depth coordinates of the joint point in the input depth map sample 2. DETAILED DESCRIPTION
[0082] The specific implementation mode of the present invention is described below in conjunction with embodiments:
[0083] It should be noted that the structures, proportions, sizes, etc. shown in this specification are only used to match the contents disclosed in the specification so that people familiar with this technology can understand and read them, and are not used to limit the conditions under which the present invention can be implemented. Any structural modification, change in proportional relationship or adjustment of size should still fall within the scope of the technical content disclosed in the present invention without affecting the effects and purposes that can be achieved by the present invention.
[0084] At the same time, the terms such as "upper", "lower", "left", "right", "middle" and "one" cited in this specification are only for the convenience of description and are not used to limit the scope of implementation of the present invention. Changes or adjustments to their relative relationships should be regarded as the scope of implementation of the present invention without substantially changing the technical content.
[0085] Embodiment 1:
[0086] 1. For any input depth map sample I, is the matrix representation of the depth map sample I, W and H correspond to the number of rows and columns of the depth map sample matrix representation, respectively. The representation matrix is a real number matrix. The depth map sample I is marked by the 2D plane coordinates and depth coordinates of N joint points:
[0087]
[0088]
[0089] in, Represents the true value of the 2D plane coordinates of the nth joint point, Indicates the true value of the depth coordinate of the nth joint point.
[0090] On the depth map sample I, a static anchor point is set every h pixels, and a total of M static anchor points are set. The depth map sample I is input into the feature extraction module to obtain the depth map features. The structure of the feature extraction module is consistent with that of ResNet-50, which consists of a series of convolutional layers and pooling layers. The output of the feature extraction module is the depth map feature Where W 1 , H 1 , D represent the height, width and number of channels of the output feature map respectively. The depth map feature F has a total of W 1 ×H 1 pixels, the feature vector f of each pixel i The dimension is D, i = 1, 2, ..., W 1 ×H 1 .
[0091] 2. Input the depth map features into the static anchor point estimation module to obtain the estimation results of the static anchor point for the 2D plane coordinates of the joint point and the estimation results of the static anchor point for the depth coordinates of the joint point.
[0092] The input of the static anchor estimation module is the depth map feature F, and the output is the estimation result S of the 2D plane coordinates of the static anchor point to the joint point xy And the estimation result S of the depth coordinate of the joint point by the static anchor point z . is the matrix representation of the estimated result of the static anchor point on the 2D plane coordinates of the joint point, and M, N and 2 correspond to the number of rows, columns and channels of the matrix representation of the estimated result of the static anchor point on the 2D plane coordinates of the joint point, respectively. (S xy ) a,b For S xy The value of row a and column b in (S xy ) a,b It represents the estimated result of the 2D plane coordinates of the bth joint point by the ath static anchor point. It is the matrix representation of the estimated result of the 2D plane coordinates of the bth joint point by the ath static anchor point. is the matrix representation of the estimation result of the depth coordinates of the joint points by the static anchor points, and M, N and 1 correspond to the number of rows, columns and channels of the matrix representation of the estimation result of the depth coordinates of the joint points by the static anchor points, respectively. (S z ) a,b For S z The value of row a and column b in (S z ) a,b Represents the estimated result of the depth coordinate of the bth joint point by the ath static anchor point. It is the matrix representation of the estimated result of the depth coordinate of the bth joint point by the ath static anchor point.
[0093] The structure of the static anchor point estimation module is as follows Figure 1 As shown in the figure, it includes a 2D plane coordinate estimation module and a depth estimation module. The 2D plane coordinate estimation module includes convolution layer 1, convolution layer 2, convolution layer 3, convolution layer 4, convolution layer 5 and transformation layer 1. Convolution layer 1 has 256 convolution kernels, and the size of the convolution kernel is 3×3. The input of convolution layer 1 is the depth map feature F, and the output is P 1 Convolutional layer 2 has 256 convolution kernels, and the size of the convolution kernel is 3×3. The input of convolutional layer 2 is P 1 , the output is P 2 Convolution layer 3 has 256 convolution kernels, and the size of the convolution kernel is 3×3. The input of convolution layer 3 is P 2 , the output is P 3 Convolutional layer 4 has 256 convolution kernels, and the size of the convolution kernel is 3×3. The input of convolutional layer 4 is P 3, the output is P 4 . Convolutional layer 5 has H 1 The convolution kernel size is 3×3. The input of convolution layer 5 is P 4 , the output is P 5 . Conversion layer 1 is used to change P 5 The shape of the transformation layer 1 is P 5 , the output is S xy The depth estimation module includes convolutional layer 6, convolutional layer 7, convolutional layer 8, convolutional layer 9, convolutional layer 10 and transformation layer 2. Convolutional layer 6 has 256 convolution kernels, and the size of the convolution kernel is 3×3. The input of convolutional layer 6 is the depth map feature F, and the output is P 6 Convolution layer 7 has 256 convolution kernels, and the size of the convolution kernel is 3×3. The input of convolution layer 7 is P 6 , the output is P 7 Convolution layer 8 has 256 convolution kernels, and the size of the convolution kernel is 3×3. The input of convolution layer 8 is P 7 , the output is P 8 Convolution layer 9 has 256 convolution kernels, and the size of the convolution kernel is 3×3. The input of convolution layer 9 is P 8 , the output is P 9 . Convolutional layer 10 has H 2 The convolution kernel size is 3×3. The input of convolution layer 10 is P 9 , the output is P 10 . Conversion layer 2 is used to change P 10 The shape of the transformation layer 2 is P 10 , the output is S z .
[0094] 3. Input the depth map features into the static anchor weight estimation module to obtain the normalized weight of the static anchor to the joint point. The input of the static anchor weight estimation module is the depth map feature F, and the output is the normalized weight W of the static anchor to the joint point. It is the matrix representation of the normalized weights of the static anchor points to the joint points.
[0095] The structure of the static anchor weight estimation module is as follows Figure 2 As shown, it includes convolution layer 11, convolution layer 12, convolution layer 13, convolution layer 14, convolution layer 15, transformation layer 3 and normalization layer. Convolution layer 11 has 256 convolution kernels, and the size of the convolution kernel is 3×3. The input of convolution layer 11 is the depth map feature F, and the output is P 11 Convolution layer 12 has 256 convolution kernels, and the size of the convolution kernel is 3×3. The input of convolution layer 12 is P 11 ,lose
[0096] Output is P 12Convolution layer 13 has 256 convolution kernels, and the size of the convolution kernel is 3×3. The input of convolution layer 13 is P 12 , the output is P 13 Convolution layer 14 has 256 convolution kernels, and the size of the convolution kernel is 3×3. The input of convolution layer 14 is P 13 , the output is P 14 . Convolutional layer 15 has H 3 convolution kernels, the size of the convolution kernel is 3×3. The input of convolution layer 15 is P 14 , the output is P 15 . Conversion layer 3 is used to change P 15 The shape of the transformation layer 3 is P 15 , the output is the weight of the static anchor point to the joint point in, It is the matrix representation of the weight of the static anchor point to the joint point. for The value of column b in represents the weight of M static anchor points to the bth joint point, It is the matrix representation of the weights of M static anchor points to the bth joint point. for The value of the ath row and bth column in represents the weight of the a-th static anchor point to the b-th joint point, is the matrix representation of the weight of the a-th static anchor point to the b-th joint point. The activation function of the normalization layer is the softmax function, and the input of the normalization layer is The output is the normalized weight W of the static anchor point to the joint point. W is the normalized weight matrix representation of the static anchor point to the joint point. a,b is the value of row a and column b in W, W a,b represents the normalized weight of the a-th static anchor point to the b-th joint point, W is the matrix representation of the normalized weight of the a-th static anchor point to the b-th joint point. :,b is the value of column b in W, W :,b is the normalized weight of M static anchor points to the bth joint point, W is the matrix representation of the normalized weights of M static anchor points on the bth joint point. :,b The calculation method is as follows:
[0097]
[0098] in, Represents the softmax activation function.
[0099] 4. Take the weighted sum of the estimation results of the static anchor points on the 2D plane coordinates of the joint points to obtain the preliminary estimation result G of the 2D plane coordinates of the joint points. It is the matrix representation of the preliminary estimation result of the 2D plane coordinates of the joint points. (G) b,: is the value of the b-th row in G, which represents the preliminary estimate of the 2D plane coordinates of the b-th joint point, and is calculated as follows:
[0100]
[0101] Where M represents the number of static anchor points, W a,b *(S xy ) a,b Represents the weighted result of the static anchor point a on the 2D plane coordinates of the joint point b.
[0102] 5. Take the weighted sum of the estimation results of the depth coordinates of the joint points by the static anchor points to obtain the final estimation result Z of the depth coordinates of the joint points. It is the matrix representation of the final estimated result of the joint point depth coordinate. (Z) b,: is the value of the b-th row in Z, which represents the final estimated result of the depth coordinate of the b-th joint point, and is calculated as follows:
[0103]
[0104] Where M represents the number of static anchor points, W a,b *(S z ) a,b Represents the weighted result of the depth coordinate of the static anchor point a to the joint point b.
[0105] 6. Input the estimation result of the static anchor point to the joint point 2D plane coordinates into the dynamic anchor point correction module to obtain the correction result of the dynamic anchor point to the joint point 2D plane coordinates. In a depth map sample, each joint point is defined as a dynamic anchor point. The input of the dynamic anchor point correction module is the estimation result S of the static anchor point to the joint point 2D plane coordinates. xy , the output is the correction result Q of the dynamic anchor point to the 2D plane coordinates of the joint point.
[0106] The structure of the dynamic anchor correction module is as follows: Figure 3 As shown, it includes a conversion layer 4, a copy layer, a cascade layer, a convolution layer 16 and a merging layer. The conversion layer 4 is used to change S xy The shape of the transformation layer 4 is S xy , the output is R. Among them, is the matrix representation of the output of transformation layer 4, W 2 , N and D 2The number of rows, columns and channels of the matrix representation corresponding to the output of the conversion layer 4. :,:,d represents the value of the dth channel in R, is the matrix representation of the dth channel in R. Convolutional layer 16 has H 4 The convolution kernel size is 3×3. The input of convolution layer 16 is R and the output is P 16 . is the matrix representation of the output of convolutional layer 16, W 2 , N and D 2 ×N represent the number of rows, columns, and channels of the matrix representation of the output of convolutional layer 16. The input of the replication layer is R, and the replication layer is used to replicate each channel of R N times. The replication layer has D 2 outputs, respectively R 1 ,…,R d ,…,R D2 , d∈[1,D 2 ]. Among them, R d Indicates that the value R of the dth channel in R :,:,d Copy the result N times, is the matrix representation of the result of replicating the value of the dth channel in R by N times. The input of the cascade layer is The output is O, and the expression of O is as follows:
[0107]
[0108] Among them, [·,·] means connection by channel, is the matrix representation of the output of the cascade layer. The input of the merge layer is O and P 16 , the output is the correction result Q of the dynamic anchor point to the 2D plane coordinates of the joint point. The calculation process of Q is as follows:
[0109] Q=O+P 16 ,
[0110] Among them, + means adding elements at corresponding positions. is the matrix representation of the correction result of the dynamic anchor point on the 2D plane coordinates of the joint point, W 2 , N and (D 2 ×N) is the number of rows, columns and channels of the matrix representation of the correction result of the dynamic anchor point on the 2D plane coordinates of the joint point.
[0111] 7. Input the preliminary estimation results of the 2D plane coordinates of the joint points into the dynamic anchor weight estimation module to obtain the normalized weights of the dynamic anchor points to the joint points. The structure of the dynamic anchor weight estimation module is as follows: Figure 4As shown, it includes a linear layer, a 1D convolution layer 1, a 1D convolution layer 2, a 1D convolution layer 3, a 1D convolution layer 4, a 1D convolution layer 5, and a normalization layer 1. The number of neurons in the linear layer is D 3 , the input of the linear layer is G and the output is Q 1 , Q 1 The calculation steps are as follows:
[0112]
[0113] Among them, W 1 are the weights of the linear layer, is the bias vector of the linear layer. The input of 1D convolutional layer 1 is Q 1 , the output is Q 2 1D convolution layer 1 has 256 convolution kernels, and the size of the convolution kernel is 3×1. The input of 1D convolution layer 2 is Q 2 , the output is Q 3 1D convolution layer 2 has 128 convolution kernels, and the size of the convolution kernel is 3×1. The input of 1D convolution layer 3 is Q 3 , the output is Q 4 . 1D convolution layer 3 has 64 convolution kernels, and the size of the convolution kernel is 3×1. The input of 1D convolution layer 4 is Q 4 , the output is Q 5 . 1D convolution layer 4 has 32 convolution kernels, and the size of the convolution kernel is 3×1. The input of 1D convolution layer 5 is Q 5 , the output is the weight Q of the dynamic anchor point to the joint point 6 The 1D convolution layer 5 has 16 convolution kernels, and the size of the convolution kernel is 3×1. It is the matrix representation of the weight of the dynamic anchor point to the joint point. (Q 6 ) :,b Indicates Q 6 The value of column b of It is the matrix representation of the weights of N dynamic anchor points on the bth joint point. (Q 6 ) f,b Indicates Q 6 The value of the fth row and bth column of represents the weight of the fth dynamic anchor point to the bth joint point. The activation function of the normalization layer 1 is the softmax function, and the input of the normalization layer 1 is Q 6 , the output is the normalized weight Q of the dynamic anchor point to the joint point 7 . It is the matrix representation of the normalized weight of the dynamic anchor point to the joint point. (Q 7 ) f,b Indicates Q 7 The value of the fth row and bth column of represents the normalized weight of the fth dynamic anchor point to the bth joint point. (Q 7 ) :,b Q 7 The value of column b of is the matrix representation of the normalized weights of N dynamic anchor points to the bth joint point, and its calculation process is as follows:
[0114]
[0115] in, Represents the softmax activation function.
[0116] 8. Take the weighted sum of the correction results of the dynamic anchor points on the 2D plane coordinates of the joint points to obtain the correction result H of the 2D plane coordinates of the joint points. is the matrix representation of the correction result of the 2D plane coordinates of the joint point. The calculation steps of H are as follows:
[0117] 1) The matrix shape of the correction result Q of the dynamic anchor point to the 2D plane coordinates of the joint point is transformed into N×N×(M×2). The calculation steps are as follows:
[0118] Q 1 =Re(Q),
[0119] in, is the matrix representation of the correction result of the dynamic anchor point to the 2D plane coordinates of the joint point after the shape transformation. Re(.) is the shape transformation function. (Q 1 ) f,b Q 1 The value of the fth row and the bth column, (Q 1 ) f,b It represents the correction result of the 2D plane coordinates of the bth joint point by the fth dynamic anchor point. It is the matrix representation of the correction result of the 2D plane coordinates of the bth joint point by the fth dynamic anchor point.
[0120] 2) Correction result Q of the 2D plane coordinates of the joint point for the dynamic anchor point after the shape transformation 1 The weighted summation is performed to obtain the correction result H of the 2D plane coordinates of the joint point. The calculation steps are as follows:
[0121]
[0122] in, It is the matrix representation of the correction result of the 2D plane of the b-th joint point.
[0123] 9. Take the weighted sum of the correction results of the 2D plane coordinates of the joint points to obtain the final estimated result J of the 2D plane coordinates of the joint points. is the matrix representation of the final estimated result of the 2D plane coordinates of the joint point. The calculation process of J is as follows:
[0124]
[0125] Among them, (H) a,b is the value in row a and column b of H, (J) b is the value of the b-th row in J, and is also the final estimated result of the 2D plane coordinates of the b-th joint point.
[0126] 10. Construct a gesture estimation network based on dynamic and static anchor points, such as Figure 5 As shown. The input of the network is the depth map sample I, and the output is the final estimated result J of the 2D plane coordinates of the joint point and the final estimated result Z of the depth coordinate. For V training samples, the loss function L used by the training network is:
[0127]
[0128] Among them, α represents the balance coefficient, [L xy ] v is the 2D plane loss function of the vth sample, [L z ] v is the depth loss function, [L xy ] v and [L z ] v The calculation method is as follows:
[0129]
[0130]
[0131] Where N represents the number of joint points in each training sample, [(J) n ] v represents the final estimated result of the 2D plane coordinates of the nth joint point in the vth sample, Represents the true value of the 2D plane coordinates of the nth joint point in the vth sample, [(Z) n,: ] v represents the final estimated result of the depth coordinate of the nth joint point in the vth sample, Represents the true value of the depth coordinate of the nth joint point in the vth sample. τ1 (.) represents the L1-smooth loss function with parameter τ1, which is defined as follows:
[0132]
[0133] L τ2(.) represents the L1-smooth loss function with parameter τ2, which is defined as follows:
[0134]
[0135] 11. Input each trained depth map sample into the gesture estimation network based on dynamic and static anchor points, and train the network until convergence.
[0136] 12. Input the tested depth map sample into the trained dynamic and static anchor gesture estimation network to obtain the 3D coordinates of the N joint points in the current tested depth map sample to achieve gesture estimation.
[0137] Embodiment 2:
[0138] As the algorithm flow Figure 6 As shown, the present invention provides a gesture estimation method based on dynamic and static anchor points, which specifically includes:
[0139] 1. The total number of depth map sample sets is 240. Two-thirds of the samples are randomly selected as the training set, and the remaining one-third are selected as the test set, resulting in a total of 160 training samples and 80 test samples. For any depth map sample I, is the matrix representation of the depth map sample I, the number of rows and columns of the depth map sample are both 176, The representation matrix is a real number matrix. The depth map sample I is marked by the 2D plane coordinates and depth coordinates of the 16 joint points:
[0140]
[0141]
[0142] in, Represents the true value of the 2D plane coordinates of the nth joint point, Indicates the true value of the depth coordinate of the nth joint point.
[0143] On the depth map sample I, a static anchor point is set every 4 pixels, and a total of 1936 static anchor points are set. The depth map sample I is input into the feature extraction module to obtain the depth map features. The structure of the feature extraction module is consistent with that of ResNet-50, which consists of a series of convolutional layers and pooling layers. The output of the feature extraction module is the depth map feature. The height and width of the depth map feature F are both 11, and the number of channels is 1024. The depth map feature F has a total of 121 pixels, and the feature vector f of each pixel is i The dimension is 1024, i=1,2,…,1024.
[0144] 2. Input the depth map features into the static anchor estimation module to obtain the estimation results of the static anchor points for the 2D plane coordinates of the joint points and the estimation results of the static anchor points for the depth coordinates of the joint points. The input of the static anchor estimation module is the depth map feature F, and the output is the estimation result S of the static anchor points for the 2D plane coordinates of the joint points. xy And the estimation result S of the depth coordinate of the joint point by the static anchor point z . is the matrix representation of the estimation result of the static anchor point on the 2D plane coordinates of the joint point, and 1936, 16 and 2 correspond to the number of rows, columns and channels of the matrix representation of the estimation result of the static anchor point on the 2D plane coordinates of the joint point, respectively. (S xy ) a,b For S xy The value of row a and column b in (S xy ) a,b It represents the estimated result of the 2D plane coordinates of the bth joint point by the ath static anchor point. It is the matrix representation of the estimated result of the 2D plane coordinates of the bth joint point by the ath static anchor point. is the matrix representation of the estimation result of the depth coordinates of the joint points by the static anchor points, and 1936, 16 and 1 correspond to the number of rows, columns and channels of the matrix representation of the estimation result of the depth coordinates of the joint points by the static anchor points, respectively. (S z ) a,b For S z The value of row a and column b in (S z ) a,b Represents the estimated result of the depth coordinate of the bth joint point by the ath static anchor point. It is the matrix representation of the estimated result of the depth coordinate of the bth joint point by the ath static anchor point.
[0145] The structure of the static anchor point estimation module is as follows Figure 1 As shown in the figure, it includes a 2D plane coordinate estimation module and a depth estimation module. The 2D plane coordinate estimation module includes convolution layer 1, convolution layer 2, convolution layer 3, convolution layer 4, convolution layer 5 and transformation layer 1. Convolution layer 1 has 256 convolution kernels, and the size of the convolution kernel is 3×3. The input of convolution layer 1 is the depth map feature F, and the output is P 1 Convolutional layer 2 has 256 convolution kernels, and the size of the convolution kernel is 3×3. The input of convolutional layer 2 is P 1 , the output is P 2 Convolution layer 3 has 256 convolution kernels, and the size of the convolution kernel is 3×3. The input of convolution layer 3 is P 2 , the output is P 3 Convolutional layer 4 has 256 convolution kernels, and the size of the convolution kernel is 3×3. The input of convolutional layer 4 is P 3 , the output is P 4Convolution layer 5 has 512 convolution kernels, and the size of the convolution kernel is 3×3. The input of convolution layer 5 is P 4 , the output is P 5 . Conversion layer 1 is used to change P 5 The shape of the transformation layer 1 is P 5 , the output is S xy The depth estimation module includes convolutional layer 6, convolutional layer 7, convolutional layer 8, convolutional layer 9, convolutional layer 10 and transformation layer 2. Convolutional layer 6 has 256 convolution kernels, and the size of the convolution kernel is 3×3. The input of convolutional layer 6 is the depth map feature F, and the output is P 6 Convolution layer 7 has 256 convolution kernels, and the size of the convolution kernel is 3×3. The input of convolution layer 7 is P 6 , the output is P 7 Convolution layer 8 has 256 convolution kernels, and the size of the convolution kernel is 3×3. The input of convolution layer 8 is P 7 , the output is P 8 Convolution layer 9 has 256 convolution kernels, and the size of the convolution kernel is 3×3. The input of convolution layer 9 is P 8 , the output is P 9 Convolution layer 10 has 256 convolution kernels, and the size of the convolution kernel is 3×3. The input of convolution layer 10 is P 9 , the output is P 10 . Conversion layer 2 is used to change P 10 The shape of the transformation layer 2 is P 10 , the output is S z .
[0146] 3. Input the depth map features into the static anchor weight estimation module to obtain the normalized weight of the static anchor to the joint point. The input of the static anchor weight estimation module is the depth map feature F, and the output is the normalized weight W of the static anchor to the joint point. It is the matrix representation of the normalized weights of the static anchor points to the joint points.
[0147] The structure of the static anchor weight estimation module is as follows Figure 2 As shown, it includes convolution layer 11, convolution layer 12, convolution layer 13, convolution layer 14, convolution layer 15, transformation layer 3 and normalization layer. Convolution layer 11 has 256 convolution kernels, and the size of the convolution kernel is 3×3. The input of convolution layer 11 is the depth map feature F, and the output is P 11 Convolution layer 12 has 256 convolution kernels, and the size of the convolution kernel is 3×3. The input of convolution layer 12 is P 11 , the output is P 12 Convolution layer 13 has 256 convolution kernels, and the size of the convolution kernel is 3×3. The input of convolution layer 13 is P 12 , the output is P 13Convolution layer 14 has 256 convolution kernels, and the size of the convolution kernel is 3×3. The input of convolution layer 14 is P 13 , the output is P 14 Convolution layer 15 has 256 convolution kernels, and the size of the convolution kernel is 3×3. The input of convolution layer 15 is P 14 , the output is P 15 . Conversion layer 3 is used to change P 15 The shape of the transformation layer 3 is P 15 , the output is the weight of the static anchor point to the joint point in, It is the matrix representation of the weight of the static anchor point to the joint point. for The value of column b in represents the weight of 1936 static anchor points to the bth joint point, It is the matrix representation of the weights of 1936 static anchor points on the bth joint point. for The value of row a and column b in represents the weight of the a-th static anchor point to the b-th joint point, is the matrix representation of the weight of the a-th static anchor point to the b-th joint point. The activation function of the normalization layer is the softmax function, and the input of the normalization layer is The output is the normalized weight W of the static anchor point to the joint point. W is the normalized weight matrix representation of the static anchor point to the joint point. a,b is the value of row a and column b in W, W a,b represents the normalized weight of the a-th static anchor point to the b-th joint point, W is the matrix representation of the normalized weight of the a-th static anchor point to the b-th joint point. :,b is the value of column b in W, W :,b is the normalized weight of 1936 static anchor points to the bth joint point, W is the matrix representation of the normalized weights of 1936 static anchor points on the bth joint point. :,b The calculation method is as follows:
[0148]
[0149] in, Represents the softmax activation function.
[0150] 4. Take the weighted sum of the estimation results of the static anchor points on the 2D plane coordinates of the joint points to obtain the preliminary estimation result G of the 2D plane coordinates of the joint points. It is the matrix representation of the preliminary estimation result of the 2D plane coordinates of the joint points. (G) b,:is the value of the b-th row in G, which represents the preliminary estimate of the 2D plane coordinates of the b-th joint point, and is calculated as follows:
[0151]
[0152] Among them, the number of static anchor points is 1936, W a,b *(S xy ) a,b Represents the weighted result of the static anchor point a on the 2D plane coordinates of the joint point b.
[0153] 5. Take the weighted sum of the estimation results of the depth coordinates of the joint points by the static anchor points to obtain the final estimation result Z of the depth coordinates of the joint points. It is the matrix representation of the final estimated result of the joint point depth coordinate. (Z) b,: is the value of the b-th row in Z, which represents the final estimated result of the depth coordinate of the b-th joint point, and is calculated as follows:
[0154]
[0155] Among them, the number of static anchor points is 1936, W a,b *(S z ) a,b Represents the weighted result of the depth coordinate of the static anchor point a to the joint point b.
[0156] 6. Input the estimation result of the static anchor point to the joint point 2D plane coordinates into the dynamic anchor point correction module to obtain the correction result of the dynamic anchor point to the joint point 2D plane coordinates. In a depth map sample, each joint point is defined as a dynamic anchor point. The input of the dynamic anchor point correction module is the estimation result S of the static anchor point to the joint point 2D plane coordinates. xy , the output is the correction result Q of the dynamic anchor point to the 2D plane coordinates of the joint point.
[0157] The structure of the dynamic anchor correction module is as follows: Figure 3 As shown, it includes a conversion layer 4, a copy layer, a cascade layer, a convolution layer 16 and a merging layer. The conversion layer 4 is used to change S xy The shape of the transformation layer 4 is S xy , the output is R. Among them, is the matrix representation of the output of the conversion layer 4. The number of rows and columns of the matrix representation of the output of the conversion layer 4 are both 16, and the number of channels is 242. :,:,d represents the value of the dth channel in R, is the matrix representation of the dth channel in R. Convolution layer 16 has 3872 convolution kernels, and the size of the convolution kernel is 3×3. The input of convolution layer 16 is R, and the output is P 16 . is the matrix representation of the output of convolution layer 16. The number of rows and columns of the matrix representation of the output of convolution layer 16 is 16, and the number of channels is 242×16. The input of the replication layer is R, and the replication layer is used to replicate each channel of R 16 times. The replication layer has 242 outputs, namely R 1 ,…,R d ,…,R 242 , d∈[1,242]. Among them, R d Indicates that the value R of the dth channel in R :,:,d Copy the result 16 times, is the matrix representation of the result of replicating the value of the dth channel in R by 16 times. The input of the cascade layer is R 1 ,…,R d ,…,R 242 , d∈[1,242], the output is O, and the expression of O is as follows:
[0158] O=[R 1 ,…,R d ,…,R 242 ],d∈[1,242],
[0159] Among them, [·,·] means connection by channel, is the matrix representation of the output of the cascade layer. The input of the merge layer is O and P 16 , the output is the correction result Q of the dynamic anchor point to the 2D plane coordinates of the joint point. The calculation process of Q is as follows:
[0160] Q=O+P 16 ,
[0161] Among them, + means adding elements at corresponding positions. It is the matrix representation of the correction result of the dynamic anchor point on the 2D plane coordinates of the joint point. The number of rows and columns of the matrix representation of the correction result of the dynamic anchor point on the 2D plane coordinates of the joint point are both 16, and the number of channels is 242×16.
[0162] 7. Input the preliminary estimation results of the 2D plane coordinates of the joint points into the dynamic anchor weight estimation module to obtain the normalized weights of the dynamic anchor points to the joint points. The structure of the dynamic anchor weight estimation module is as follows: Figure 4 As shown, it contains a linear layer, a 1D convolution layer 1, a 1D convolution layer 2, a 1D convolution layer 3, a 1D convolution layer 4, a 1D convolution layer 5, and a normalization layer 1. The number of neurons in the linear layer is 256, the input of the linear layer is G, and the output is Q 1 , Q 1 The calculation steps are as follows:
[0163]
[0164] Among them, W1 are the weights of the linear layer, is the bias vector of the linear layer. The input of 1D convolutional layer 1 is Q 1 , the output is Q 2 1D convolution layer 1 has 256 convolution kernels, and the size of the convolution kernel is 3×1. The input of 1D convolution layer 2 is Q 2 , the output is Q 3 1D convolution layer 2 has 128 convolution kernels, and the size of the convolution kernel is 3×1. The input of 1D convolution layer 3 is Q 3 , the output is Q 4 . 1D convolution layer 3 has 64 convolution kernels, and the size of the convolution kernel is 3×1. The input of 1D convolution layer 4 is Q 4 , the output is Q 5 . 1D convolution layer 4 has 32 convolution kernels, and the size of the convolution kernel is 3×1. The input of 1D convolution layer 5 is Q 5 , the output is the weight Q of the dynamic anchor point to the joint point 6 The 1D convolution layer 5 has 16 convolution kernels, and the size of the convolution kernel is 3×1. It is the matrix representation of the weight of the dynamic anchor point to the joint point. (Q 6 ) :,b Indicates Q 6 The value of column b of It is the matrix representation of the weights of 16 dynamic anchor points on the bth joint point. (Q 6 ) f,b Indicates Q 6 The value of the fth row and bth column of represents the weight of the fth dynamic anchor point to the bth joint point. The activation function of the normalization layer 1 is the softmax function, and the input of the normalization layer 1 is Q 6 , the output is the normalized weight Q of the dynamic anchor point to the joint point 7 . It is the matrix representation of the normalized weight of the dynamic anchor point to the joint point. (Q 7 ) f,b Indicates Q 7 The value of the fth row and bth column of represents the normalized weight of the fth dynamic anchor point to the bth joint point. (Q 7 ) :,b Q 7 The value of column b of is the matrix representation of the normalized weights of the 16 dynamic anchor points to the bth joint point, and its calculation process is as follows:
[0165]
[0166] in, Represents the softmax activation function.
[0167] 8. Take the weighted sum of the correction results of the dynamic anchor points on the 2D plane coordinates of the joint points to obtain the correction result H of the 2D plane coordinates of the joint points. is the matrix representation of the correction result of the 2D plane coordinates of the joint point. The calculation steps of H are as follows:
[0168] 1) The matrix shape of the correction result Q of the dynamic anchor point to the 2D plane coordinates of the joint point is transformed into 16×16×1936×2). The calculation steps are as follows:
[0169] Q 1 =Re(Q),
[0170] in, is the matrix representation of the correction result of the dynamic anchor point to the 2D plane coordinates of the joint point after the shape transformation. Re(.) is the shape transformation function. (Q 1 ) f,b Q 1 The value of the fth row and the bth column, (Q 1 ) f,b It represents the correction result of the 2D plane coordinates of the bth joint point by the fth dynamic anchor point. It is the matrix representation of the correction result of the 2D plane coordinates of the bth joint point by the fth dynamic anchor point.
[0171] 2) Correction result Q of the 2D plane coordinates of the joint point for the dynamic anchor point after the shape transformation 1 The weighted summation is performed to obtain the correction result H of the 2D plane coordinates of the joint point. The calculation steps are as follows:
[0172]
[0173] Among them, H b Represents the value of the b-th column in H, which is also the corrected result of the 2D plane coordinates of the b-th joint point. It is the matrix representation of the correction result of the 2D plane of the b-th joint point.
[0174] 9. Take the weighted sum of the correction results of the 2D plane coordinates of the joint points to obtain the final estimated result J of the 2D plane coordinates of the joint points. is the matrix representation of the final estimated result of the 2D plane coordinates of the joint point. The calculation process of J is as follows:
[0175]
[0176] Among them, (H) a,b is the value in row a and column b of H, (J) bis the value of the b-th row in J, and is also the final estimated result of the 2D plane coordinates of the b-th joint point.
[0177] 10. Construct a gesture estimation network based on dynamic and static anchor points, such as Figure 5 As shown. The input of the network is the depth map sample I, and the output is the final estimated result J of the 2D plane coordinates of the joint point and the final estimated result Z of the depth coordinate. For 160 training samples, the loss function L used by the training network is:
[0178]
[0179] Among them, 0.5 represents the balance coefficient, [L xy ] v is the 2D plane loss function of the vth sample, [L z ] v is the depth loss function, [L xy ] v and [L z ] v The calculation method is as follows:
[0180]
[0181]
[0182] The number of joint points in each training sample is 16, [(J) n ] v represents the final estimated result of the 2D plane coordinates of the nth joint point in the vth sample, Represents the true value of the 2D plane coordinates of the nth joint point in the vth sample, [(Z) n,: ] v represents the final estimated result of the depth coordinate of the nth joint point in the vth sample, Represents the true value of the depth coordinate of the nth joint point in the vth sample. 1 (.) represents the L1-smooth loss function with parameter 1, which is defined as follows:
[0183]
[0184] L 3 (.) represents the L1-smooth loss function with a parameter of 3, which is defined as follows:
[0185]
[0186] 11. Input each trained depth map sample into the gesture estimation network based on dynamic and static anchor points, and train the network until convergence.
[0187] 12. Input the tested depth map sample into the trained dynamic and static anchor gesture estimation network to obtain the 3D coordinates of the 16 joint points in the current tested depth map sample to achieve gesture estimation.
[0188] Figure 7 is the input depth map sample 1, Figure 8 is a schematic diagram of the final estimation result of the 2D plane coordinates and the final estimation result of the depth coordinates of the joint point in the input depth map sample 1; Fig. 9 is the input depth map sample 2, Fig.10 Schematic diagram of the final estimation result of the 2D plane coordinates and the final estimation result of the depth coordinates of the joint point in the input depth map sample 2.
[0189] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0190] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0191] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
Claims
1. A gesture estimation method based on dynamic and static anchor points, It is characterized in that The method comprises: S1: Setting a static anchor point on the depth map sample to obtain the depth map sample after the static anchor point is set; S2: Based on the depth map samples after setting the static anchor points, a gesture estimation network of dynamic and static anchor points is constructed; the gesture estimation network of dynamic and static anchor points includes: a feature extraction module, a static anchor point estimation module, a static anchor point weight estimation module, a dynamic anchor point weight estimation module and a dynamic anchor point correction module; S3: Train the gesture estimation network based on dynamic and static anchor points until convergence, and obtain a trained dynamic and static anchor point gesture estimation network; S4: Input the tested depth map sample into the trained dynamic and static anchor gesture estimation network to achieve gesture estimation; The gesture estimation network for constructing dynamic and static anchor points specifically includes: Input the depth map sample into the feature extraction module to obtain the depth map features; The depth map features are input into the static anchor point estimation module to obtain the estimation results of the static anchor point for the 2D plane coordinates of the joint point and the estimation results of the static anchor point for the depth coordinates of the joint point; Input the depth map features into a preset static anchor point weight estimation module to obtain the normalized weight of the static anchor point to the joint point; Based on the normalized weights of the static anchor points to the joint points, weighted summation is performed on the estimation results of the static anchor points to the 2D plane coordinates of the joint points to obtain a preliminary estimation result of the 2D plane coordinates of the joint points; Based on the normalized weights of the static anchor points to the joint points, the estimated results of the depth coordinates of the static anchor points to the joint points are weighted summed to obtain the final estimated results of the depth coordinates of the joint points; The estimation result of the 2D plane coordinates of the joint points by the static anchor points is input into the preset dynamic anchor point correction model to obtain the correction result of the 2D plane coordinates of the joint points by the dynamic anchor points; Input the preliminary estimation result of the 2D plane coordinates of the joint point into the preset dynamic anchor point weight estimation model to obtain the normalized weight of the dynamic anchor point to the joint point; Based on the normalized weights of the dynamic anchor points to the joint points, weighted summation is performed on the correction results of the dynamic anchor points to the 2D plane coordinates of the joint points to obtain the correction results of the 2D plane coordinates of the joint points; Based on the normalized weights of the static anchor points on the joint points, the correction results of the 2D plane coordinates of the joint points are weighted summed to obtain the final estimation results of the 2D plane coordinates of the joint points.
2. A gesture estimation method based on dynamic and static anchor points according to claim 1, It is characterized in that Based on any input depth map sample I, is the matrix representation of the depth map sample I, W and H correspond to the number of rows and columns of the depth map sample matrix representation, respectively. The representation matrix is a real number matrix; the depth map sample I is marked by the 2D plane coordinates and depth coordinates of N joint points: in, Represents the true value of the 2D plane coordinates of the nth joint point, Represents the true value of the depth coordinate of the nth joint point; On the depth map sample I, a static anchor point is set every h pixels, and a total of M static anchor points are set; the depth map sample I is input into the feature extraction module to obtain the depth map features; the structure of the feature extraction module is consistent with that of ResNet-50, which consists of a series of convolutional layers and pooling layers; the output of the feature extraction module is the depth map features Where W 1 , H 1 , D represent the height, width and number of channels of the output feature map respectively; the depth map feature F has a total of W 1 ×H 1 pixels, the feature vector f of each pixel i The dimension is D, i = 1, 2, ..., W 1 ×H 1 ; The input of the static anchor estimation module is the depth map feature F, and the output is the estimation result S of the 2D plane coordinates of the static anchor point to the joint point xy And the estimation result S of the depth coordinate of the joint point by the static anchor point z ; is the matrix representation of the estimation result of the static anchor point on the 2D plane coordinates of the joint point, M, N and 2 correspond to the number of rows, columns and channels of the matrix representation of the estimation result of the static anchor point on the 2D plane coordinates of the joint point respectively; (S xy ) a,b For S xy The value of row a and column b in (S xy ) a,b It represents the estimated result of the 2D plane coordinates of the bth joint point by the ath static anchor point. It is the matrix representation of the estimated result of the 2D plane coordinates of the bth joint point by the ath static anchor point; is the matrix representation of the estimation result of the depth coordinates of the joint points by the static anchor points, and M, N and 1 correspond to the number of rows, columns and channels of the matrix representation of the estimation result of the depth coordinates of the joint points by the static anchor points, respectively; (S z ) a,b For S z The value of row a and column b in (S z ) a,b Represents the estimated result of the depth coordinate of the bth joint point by the ath static anchor point. It is the matrix representation of the estimated result of the depth coordinate of the bth joint point by the ath static anchor point; The structure of the static anchor point estimation module: includes a 2D plane coordinate estimation module and a depth estimation module; the 2D plane coordinate estimation module contains convolution layer 1, convolution layer 2, convolution layer 3, convolution layer 4, convolution layer 5 and transformation layer 1; convolution layer 1 has 256 convolution kernels, and the size of the convolution kernel is 3×3; the input of convolution layer 1 is the depth map feature F, and the output is P 1 ; Convolutional layer 2 has 256 convolution kernels, and the size of the convolution kernel is 3×3; The input of convolutional layer 2 is P 1 , the output is P 2 ; Convolution layer 3 has 256 convolution kernels, and the size of the convolution kernel is 3×3; the input of convolution layer 3 is P 2 , the output is P 3 ; Convolution layer 4 has 256 convolution kernels, and the size of the convolution kernel is 3×3; the input of convolution layer 4 is P 3 , the output is P 4 ; Convolutional layer 5 has H 1 convolution kernels, the size of the convolution kernel is 3×3; the input of convolution layer 5 is P 4 , the output is P 5 ; Conversion layer 1 is used to change P 5 The shape of the transformation layer 1 is P 5 , the output is S xy ; The depth estimation module includes convolution layer 6, convolution layer 7, convolution layer 8, convolution layer 9, convolution layer 10 and transformation layer 2; convolution layer 6 has 256 convolution kernels, and the size of the convolution kernel is 3×3; the input of convolution layer 6 is the depth map feature F, and the output is P 6 ; Convolution layer 7 has 256 convolution kernels, and the size of the convolution kernel is 3×3; the input of convolution layer 7 is P 6 , the output is P 7 ; Convolution layer 8 has 256 convolution kernels, and the size of the convolution kernel is 3×3; the input of convolution layer 8 is P 7 , the output is P 8 ; Convolution layer 9 has 256 convolution kernels, and the size of the convolution kernel is 3×3; the input of convolution layer 9 is P 8 , the output is P 9 ; Convolutional layer 10 has H 2 convolution kernels, the size of the convolution kernel is 3×3; the input of convolution layer 10 is P 9 , the output is P 10 ; Conversion layer 2 is used to change P 10 The shape of the transformation layer 2 is P 10 , the output is S z .
3. The method for estimating gestures based on dynamic and static anchor points according to claim 1, It is characterized in that The input of the static anchor weight estimation module is the depth map feature F, and the output is the normalized weight W of the static anchor to the joint point, where It is the matrix representation of the normalized weights of the static anchor points to the joint points; The structure of the static anchor weight estimation module includes convolution layer 11, convolution layer 12, convolution layer 13, convolution layer 14, convolution layer 15, transformation layer 3 and normalization layer; convolution layer 11 has 256 convolution kernels, and the size of the convolution kernel is 3×3; the input of convolution layer 11 is the depth map feature F, and the output is P 11 ; Convolution layer 12 has 256 convolution kernels, and the size of the convolution kernel is 3×3; the input of convolution layer 12 is P 11 , the output is P 12 ; Convolution layer 13 has 256 convolution kernels, and the size of the convolution kernel is 3×3; the input of convolution layer 13 is P 12 , the output is P 13 ; Convolution layer 14 has 256 convolution kernels, and the size of the convolution kernel is 3×3; the input of convolution layer 14 is P 13 , the output is P 14 ; Convolutional layer 15 has H 3 convolution kernels, the size of the convolution kernel is 3×3; the input of convolution layer 15 is P 14 , the output is P 15 ; Conversion layer 3 is used to change P 15 The shape of the transformation layer 3 is P 15 , the output is the weight of the static anchor point to the joint point in, is the matrix representation of the weight of the static anchor point to the joint point, for The value of column b in represents the weight of M static anchor points to the bth joint point, is the matrix representation of the weights of M static anchor points on the bth joint point for The value of row a and column b in represents the weight of the a-th static anchor point to the b-th joint point, is the matrix representation of the weight of the a-th static anchor point to the b-th joint point; The activation function of the normalization layer is the softmax function, and the input of the normalization layer is The output is the normalized weight W of the static anchor point to the joint point. is the normalized weight matrix representation of the static anchor point to the joint point, W a,b is the value of row a and column b in W, W a,b represents the normalized weight of the a-th static anchor point to the b-th joint point, is the matrix representation of the normalized weight of the a-th static anchor point to the b-th joint point, W :,b is the value of column b in W, W :,b is the normalized weight of M static anchor points to the bth joint point, is the matrix representation of the normalized weights of M static anchor points on the bth joint point, W :,b The calculation method is as follows: in, Represents the softmax activation function.
4. The method for estimating gestures based on dynamic and static anchor points according to claim 1, It is characterized in that The weighted summation of the estimated results of the static anchor points on the 2D plane coordinates of the joint points is performed to obtain the preliminary estimated result G of the 2D plane coordinates of the joint points. is the matrix representation of the preliminary estimation result of the 2D plane coordinates of the joint point, (G) b,: is the value of the b-th row in G, which represents the preliminary estimate of the 2D plane coordinates of the b-th joint point, and is calculated as follows: Where M represents the number of static anchor points, W a,b *(S xy ) a,b Represents the weighted result of the static anchor point a on the 2D plane coordinates of the joint point b; The estimated results of the depth coordinates of the joint points by the static anchor points are weighted summed to obtain the final estimated result Z of the depth coordinates of the joint points. It is the matrix representation of the final estimation result of the depth coordinates of the joint points; (Z) b,: is the value of the b-th row in Z, which represents the final estimated result of the depth coordinate of the b-th joint point, and is calculated as follows: Where M represents the number of static anchor points, W a,b *(S z ) a,b Represents the weighted result of the depth coordinate of the static anchor point a to the joint point b.
5. The method for estimating gestures based on dynamic and static anchor points according to claim 1, It is characterized in that The estimation result of the 2D plane coordinates of the joint point by the static anchor point is input into the dynamic anchor point correction module to obtain the correction result of the 2D plane coordinates of the joint point by the dynamic anchor point; In a depth map sample, each joint point is defined as a dynamic anchor point; The input of the dynamic anchor correction module is the estimated result S of the static anchor point on the 2D plane coordinates of the joint point xy , the output is the correction result Q of the dynamic anchor point to the 2D plane coordinates of the joint point; The structure of the dynamic anchor correction module: including conversion layer 4, copy layer, cascade layer, convolution layer 16 and merging layer; Conversion layer 4 is used to change S xy The shape of the transformation layer 4 is S xy , the output is R, where is the matrix representation of the output of transformation layer 4, W 2 , N and D 2 The number of rows, columns, and channels of the matrix representation corresponding to the output of transformation layer 4, R :,:,d represents the value of the dth channel in R, is the matrix representation of the dth channel in R; convolutional layer 16 has H 4 convolution kernels, the size of the convolution kernel is 3×3, the input of convolution layer 16 is R, and the output is P 16 , is the matrix representation of the output of convolutional layer 16, W 2 , N and D 2 ×N represent the number of rows, columns and channels of the matrix representation of the output of convolution layer 16 respectively; the input of the replication layer is R, and the replication layer is used to replicate each channel of R N times. The replication layer has D 2 The outputs are Among them, R d Indicates that the value R of the dth channel in R :,:,d Copy the result N times, is the matrix representation of the result of replicating the value of the dth channel in R by N times; the input of the cascade layer is The output is O, and the expression of O is as follows: Among them, [·,·] means connection by channel, The matrix output by the cascade layer represents the input of the merge layer as O and P 16 , the output is the correction result Q of the dynamic anchor point to the 2D plane coordinates of the joint point. The calculation process of Q is as follows: Q=O+P 16 ; Among them, + means adding elements at corresponding positions. is the matrix representation of the correction result of the dynamic anchor point on the 2D plane coordinates of the joint point, W 2 , N and D 2 ×N) is the number of rows, columns and channels of the matrix representation of the correction result of the dynamic anchor point on the 2D plane coordinates of the joint point.
6. The method for estimating gestures based on dynamic and static anchor points according to claim 1, It is characterized in that The preliminary estimation result of the 2D plane coordinates of the joint point is input into the dynamic anchor point weight estimation module to obtain the normalized weight of the dynamic anchor point to the joint point; The structure of the dynamic anchor weight estimation module includes: linear layer, 1D convolution layer 1, 1D convolution layer 2, 1D convolution layer 3, 1D convolution layer 4, 1D convolution layer 5 and normalization layer 1; the number of neurons in the linear layer is 256, the input of the linear layer is G, and the output is Q 1 , Q 1 The calculation steps are as follows: Among them, W 1 is the weight of the linear layer, θ 1 is the bias vector of the linear layer; the input of 1D convolutional layer 1 is Q 1 , the output is Q 2 , 1D convolution layer 1 has 256 convolution kernels, and the size of the convolution kernel is 3×1; the input of 1D convolution layer 2 is Q 2 , the output is Q 3 , 1D convolution layer 2 has 128 convolution kernels, and the size of the convolution kernel is 3×1; the input of 1D convolution layer 3 is Q 3 , the output is Q 4 , 1D convolution layer 3 has 64 convolution kernels, and the size of the convolution kernel is 3×1; the input of 1D convolution layer 4 is Q 4 , the output is Q 5 , 1D convolution layer 4 has 32 convolution kernels, the size of the convolution kernel is 3×1; the input of 1D convolution layer 5 is Q 5 , the output is the weight Q of the dynamic anchor point to the joint point 6 , 1D convolution layer 5 has 16 convolution kernels, and the size of the convolution kernel is 3×1; is the matrix representation of the weight of the dynamic anchor point to the joint point (Q 6 ) :,b Indicates Q 6 The value of column b of is the matrix representation of the weights of 16 dynamic anchor points on the bth joint point (Q 6 ) f,b Indicates Q 6 The value of the fth row and bth column of The activation function of the normalization layer 1 is the softmax function, and the input of the normalization layer 1 is Q 6 , the output is the normalized weight Q of the dynamic anchor point to the joint point 7 , is the matrix representation of the normalized weights of the dynamic anchor points to the joint points, (Q 7 ) f,b Indicates Q 7 The value of the fth row and bth column of represents the normalized weight of the fth dynamic anchor point to the bth joint point, (Q 7 ) :,b Q 7 The value of column b of is the matrix representation of the normalized weights of N dynamic anchor points to the bth joint point, and its calculation process is as follows: in, Represents the softmax activation function.
7. The method for estimating gestures based on dynamic and static anchor points according to claim 1, It is characterized in that The correction result H of the 2D plane coordinates of the joint point is obtained by weighted summing the correction results of the dynamic anchor point on the 2D plane coordinates of the joint point; is the matrix representation of the correction result of the 2D plane coordinates of the joint point. The calculation steps of H are as follows: 1) The matrix shape of the correction result Q of the dynamic anchor point to the 2D plane coordinates of the joint point is transformed into N×N×(M×2). The calculation steps are as follows: Q 1 =Re(Q), in, is the matrix representation of the correction result of the dynamic anchor point on the 2D plane coordinates of the joint point after the shape transformation; Re(.) is the shape transformation function; (Q 1 ) f,b Q 1 The value of the fth row and the bth column, (Q 1 ) f,b It represents the correction result of the 2D plane coordinates of the bth joint point by the fth dynamic anchor point. is the matrix representation of the correction result of the 2D plane coordinates of the bth joint point by the fth dynamic anchor point; 2) Correction result Q of the 2D plane coordinates of the joint point for the dynamic anchor point after the shape transformation 1 The weighted summation is performed to obtain the correction result H of the 2D plane coordinates of the joint point. The calculation steps are as follows: in, It is the matrix representation of the correction result of the 2D plane of the b-th joint point.
8. The method for estimating gestures based on dynamic and static anchor points according to claim 1, It is characterized in that The weighted sum of the correction results of the 2D plane coordinates of the joint points is performed to obtain the final estimation result J of the 2D plane coordinates of the joint points; is the matrix representation of the final estimated result of the 2D plane coordinates of the joint point. The calculation process of J is as follows: Among them, (H) a,b is the value in row a and column b of H, (J) b is the value of the b-th row in J, and is also the final estimated result of the 2D plane coordinates of the b-th joint point.
9. The method for estimating gestures based on dynamic and static anchor points according to claim 1, It is characterized in that The gesture estimation network based on dynamic and static anchors is trained until convergence. The input of the network is the depth map sample I, and the output is the final estimation result J of the 2D plane coordinates of the joint point and the final estimation result Z of the depth coordinate. For V training samples, the loss function L used by the training network is: Among them, α represents the balance coefficient, [L xy ] v is the 2D plane loss function of the vth sample, [L z ] v is the depth loss function, [L xy ] v and [L z ] v The calculation method is as follows: Where N represents the number of joint points in each training sample, [(J) n ] v represents the final estimated result of the 2D plane coordinates of the nth joint point in the vth sample, Represents the true value of the 2D plane coordinates of the nth joint point in the vth sample, [(Z) n,: ] v represents the final estimated result of the depth coordinate of the nth joint point in the vth sample, represents the true value of the depth coordinate of the nth joint point in the vth sample; L τ1 (.) represents the L1-smooth loss function with parameter τ1, which is defined as follows: L τ2 (.) represents the L1-smooth loss function with parameter τ2, which is defined as follows:
Citation Information
Patent Citations
Gesture identification method based on depth information
CN102789568A
Three-dimensional gesture estimation method and three-dimensional gesture estimation system based on depth data
CN105389539A