A method for skeletal key point detection using a hollow transposed convolutional hourglass structure

Through the hollow transposed convolutional hourglass structure, 3D bone key point detection is solved on the RGB video stream, and the problem of insufficient spatial depth information modeling in the existing technology is achieved, real-time and efficient 3D bone key point detection on the mobile terminal.

CN114821776BActive Publication Date: 2025-07-08HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210414609.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-20
Publication Date
2025-07-08
Estimated Expiration
2042-04-20

AI Technical Summary

Technical Problem

In the prior art, in human posture estimation, 2D bone key point detection based on a single RGB image cannot effectively model spatial depth information, and RGBD images are difficult to obtain and have large calculations, making it difficult to meet real-time and deployment requirements.

Method used

The hollow transposed convolution hourglass structure is used to expand the timing receptive field through hollow convolution, and the timing and implicit modeling capabilities between the space are increased by transposed convolution, reducing the amount of model training parameters, and using GPU parallel acceleration calculations to generate accurate 3D skeletal key points.

Benefits of technology

It realizes implicitly modeling timing and depth information on RGB video streams, improves spatial perception capabilities, reduces the computational volume and training data requirements, and is suitable for mobile deployment and real-time prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114821776B_ABST
    Figure CN114821776B_ABST
Patent Text Reader

Abstract

The object of the present invention is to provide a method for detecting skeletal key points using a hollow transposed convolutional hourglass structure in view of the deficiencies of the prior art. The method expands the temporal receptive field by utilizing the characteristics of dilated convolution, and increases the implicit modeling ability between time series and spatial domain by using transposed convolution, reduces the requirements for the number of model training parameters, the amount of training data and the input data format, and makes full use of the characteristics of parallel acceleration calculation of general-purpose platforms such as GPUs to generate accurate and smooth 3D skeletal key points. After obtaining the set of RGB videos to be measured, the steps 1-6 of the present invention are sequentially performed to obtain the final 3D skeletal key point results. The present invention has four advantages: 1) the input does not need to contain depth information, 2) the use of convolutional calculation improves the computational parallelism and has a fast running speed, 3) the requirement for the quality of training data is low, and 4) it is possible to replace dilated conventional convolution for mobile deployment and real-time prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, especially the field of human pose estimation in video understanding, and relates to a method for detecting skeletal key points based on a dilated transposed convolutional hourglass structure. Background Art

[0002] Human pose estimation is a technology that uses data captured by sensors (usually cameras) for human pose analysis. In recent years, with the breakthrough progress of deep learning in computer vision fields such as image classification, object detection, and semantic segmentation, human pose estimation has also achieved good development, mainly reflected in aspects such as datasets and network structures.

[0003] As one of the basic tasks in computer vision, human pose estimation has a wide range of applications and can be used in broad fields such as action recognition, detection, pedestrian tracking, film and animation production, virtual reality, medical assistance, autonomous driving, and sports analysis. In the field of film and animation production, low-cost, stable, and accurate human pose estimation can assist designers in quickly constructing character models and endowing them with more vivid actions. In the field of monitoring and security, human pose estimation can track and analyze human actions and perform pedestrian re-identification to assist security personnel in identifying potential dangers. In the field of autonomous driving, human pose estimation technology can assist autonomous vehicles in making more intelligent responses to pedestrians and providing more effective information to traffic management personnel. In the field of sports training, human pose estimation can accurately analyze athletes' postures, such as information on body postures, running distances, jumping times, and heights, and provide quantitative sports statistics to assist sports training.

[0004] According to the classification of input information, the inputs of human pose estimation are mainly divided into RGB (Red Green Blue) images, RGBD (Red Green Blue Depth) images, and RGB video streams. The acquisition method of RGB images is relatively simple, but the information contained is less than that of RGBD images. RGBD images require special depth cameras to capture due to their inclusion of depth information. In contrast, RGB video streams can provide temporal information while containing RGB image information and are relatively simple to obtain.

[0005] With the development of deep learning technology, the 2D skeleton keypoint detection technology based on image input and Convolutional Neural Network (CNN) has developed rapidly. It represents features in an end-to-end manner and implicitly models the spatial position relationships of keypoints, outputting a tensor with spatial information. The channels of this tensor correspond to human keypoints. However, due to the lack of input information dimension, its application scenarios are still limited. The 3D skeleton keypoint detection technology based on RGBD images and convolutional neural networks can accurately obtain human pose information, but due to the difficulty of obtaining RGBD images, this method has great limitations. With the in-depth research of application scenarios and pose estimation, 3D skeleton keypoint detection methods based on video input have been continuously proposed. In order to make up for the lack of depth space information in a single 2D image, related work uses methods such as Attention, Long-Short Time Memory (LSTM), and Gated Recurrent Units (GRU) to process continuous video frames to achieve the modeling of depth space information; in addition, relevant researchers have proposed a Sequence-To-Sequence model, which encodes the 2D pose sequence in the video into a fixed-size vector and then decodes it into a 3D pose sequence; there are also researchers who have proposed the Histograms Of Oriented Gradients method to reduce the sensitivity of the algorithm to noise.

[0006] The deficiencies of the above methods are mainly reflected in the following aspects: (1) Since the directly input single RGB image lacks depth space information and cannot effectively model spatial information, it is only applicable to 2D skeleton keypoint detection tasks and performs poorly in 3D skeleton keypoint detection tasks. And the method using RGBD as input is often limited by the dataset and is difficult to be applied to actual scenarios; (2) The methods based on RNN and LSTM still have significant problems of gradient disappearance and gradient explosion, making it difficult to train deep networks. Moreover, due to their computational characteristics, it is difficult to accelerate in parallel in time series, and they have low affinity for graphics processors; (3) The existing methods have large amounts of computation and parameters, require high training data and deployment platforms, and are difficult to meet the requirements of training and deployment in actual application scenarios. Based on the above considerations, for application scenarios with limited data volume and high requirements for real-time performance, such as action recognition, medical assistance, and motion analysis, there is an urgent need to design a skeleton keypoint detection method that can use common RGB video streams to implicitly model time series and depth information to achieve parallel processing of information on mobile computing platforms. Summary of the Invention

[0007] The object of the present invention is to provide a 3D human pose estimation method based on a dilated transposed convolutional hourglass structure in view of the deficiencies of the above-mentioned prior art. The method expands the temporal receptive field by utilizing the characteristics of dilated convolution, increases the implicit modeling ability between time series and spatial domain by using transposed convolution, reduces the requirements for the number of model training parameters, the amount of training data, and the input data format, and makes full use of the parallel acceleration calculation characteristics of general-purpose platforms such as GPUs to generate accurate and smooth 3D bone key points.

[0008] Before using the method of the present invention, it is necessary to first obtain a set of RGB videos to be measured, and then use the method of the present invention.

[0009] A method for detecting bone key points using a dilated transposed convolutional hourglass structure includes:

[0010] Step 1. Construct a 2D bone key point detection model, input an RGB video stream, and obtain a 2D key point sequence;

[0011] Step 2. According to the set dilation factor, convolution kernel size, padding, downsampling layer number, etc., construct a dilated regular convolutional layer for downsampling, a dilated transposed convolutional layer for upsampling, and an activation operator;

[0012] Step 3. Perform a temporal downsampling feature extraction operation, implicitly model the temporal and spatial depth relationship, and obtain noisy 3D bone key points and an intermediate layer feature set;

[0013] Step 4. Repeat Step (3) to implement skip connections and upsampling dilated transposed convolutions, and extract the intermediate feature sets of spatial and temporal information during the reverse implicit modeling process;

[0014] Step 5. According to the set number of hourglasses, repeat Steps 2 to 4 to obtain the output intermediate features of each hourglass, construct residual skip connections within and between hourglasses, and calculate the mean square error of the hourglass output;

[0015] Step 6. Repeat Steps 2 to 3 to obtain the final 3D bone key point result, calculate the mean square error, and optimize the detection model using the stochastic gradient descent method.

[0016] Further, Step 1 is specifically:

[0017] Sub-step 1-1. Input a set of RGB videos containing the human motion to be detected where N represents the total number of video frames included in I, H is the height of the input image, W is the width of the input image, and 3 represents an RGB three-channel image;

[0018] Sub-step 1-2. Input the number of bone key points K ∈ N * , generally, K = 17. Construct a 2D bone key point detector D 2d, where D 2d can be any 2D skeletal key point detector that satisfies the input of I and the output of ;

[0019] Sub-step 1-3. Input I frame by frame into D 2d , for D 2d All the temporal outputs are 2D key point sequences where N is the total number of video frames of I, and 2 represents the X and Y dimensions of the 2D skeletal key points.

[0020] Furthermore, step 2 is specifically:

[0021] Sub-step 2-1. Input the set of dilation factors d all ={d i |d i ∈[3,9], and i∈[1,n]}, convolution kernel ks, padding p, number of downsampling layers n. Generally, ks is a 2D matrix and p = 1, n = 6. For the input of ks, d all , p and n, output the set of dilated convolutional layers where indicates that this is the i-th dilated transposed convolution in the first hourglass structure; Sub-step 2-2. Input the set of dilation factors d all ={d i |d i ∈[3,9], and i∈[1,n]}, convolution kernel ks, padding p, number of upsampling layers n, and output the dilated transposed convolution where indicates that this is the i-th dilated transposed convolution in the first hourglass structure, corresponds one-to-one with . The same operation as (2-1), each layer shares the activation operator AL;

[0022] Sub-step 2-3. Instantiate the activation operator AL through the custom Mish(x) activation layer and output it. Since AL does not need to be trained, all activation operations share this activation operator;

[0023] Sub-step 2-4. Further, for any input element the Mish(x) activation output is Mish(x) = x·tang(ln(1 + e x ));

[0024] Sub-step 2-5. Even further, for any input element the tanh(x) output is

[0025] Further, step 3 is specifically as follows:

[0026] Sub-step 3-1. Input DCL and LA, pair the two, that is and there is l ∈ [1, n]. Multiple Stack, that is (Downsample module, Down sample), output DCA 1 ;

[0027] Sub-step 3-2. Input and successively pass through where l ∈ [1, n], then the output is the calculation information of each intermediate layer and the 3D skeleton key point coordinate information with noise At this time The superscript 0 above it represents the 3D key point information obtained by the first hourglass structure;

[0028] (3-3). Input is the number of feature map channels of the l-th layer of dilated regular convolution and l ∈ [0, n]. If the number of input channels is The number of output channels is Then according to the equation

[0029] Sub-step 3-4. Input DTCL and AL, pair DTCL and the activation operator AL as DTCA, that is l ∈ [1, n]. Stack to form (Downsample module, Downsample), output DTCA 1 ;

[0030] The number of input channels is Then the number of output channels is And

[0031] Furthermore, step 4 is specifically as follows:

[0032] Sub-step 4-1. Input where That is the output of. Does not participate in the residual skip connection within the hourglass structure where it is located, but is directly input to to obtain the output And let

[0033] Sub-step 4-2. Input and According to the residual skip connection formula Perform the residual skip connection;

[0034] Sub-step 4-3. Input to Perform the upsampling operation and calculate, output the upsampled intermediate layer feature FU n-2 , according to the residual skip connection formula and input it to And so on, the intermediate feature set during the upsampling process can be obtained

[0035] Sub-step 4-5. For those participating in the residual skip connection and l ∈ [1, n]. Furthermore, if a complete hourglass structure contains n downsampling layers, then there are n corresponding upsampling layers. Therefore, in the complete hourglass structure and Are paired one by one.

[0036] Furthermore, step 5 is specifically:

[0037] Sub-step 5-1. Input The set of dilation factors d set for each layer all ={d i |d i ∈ [3, 9], and i ∈ [1, n]}, convolution kernel ks, padding p, number of downsampling layers n (the same as the number of upsampling layers), and repeat step (2). Output the set of dilated convolutional layers where DCL 2 The superscript 2 indicates that this is the second hourglass structure, the set of dilated transposed convolutions Similarly, there is DCL all ={DCL 1 , DCL 2 , …, DCL h}, DTCL all ={DTCL 1 , DTCL 2 , …, DTCLh};

[0038] Sub-step 5-2. Input DCL all , DTCL all and AL, repeat step (3) to output Similarly, repeating the step gives DCA all ={DCA 1 , DCA 2 , …, DCA h}, DTCA all ={DTCA1 , DTCA 2 , …, DTCA h};

[0039] Sub-step 5-3. Input That is the output of is used to perform a residual skip connection with and is expressed by the formula as

[0040] Sub-step 5-4. Input Similarly to steps 2, 3, 4, sub-step 5-1, sub-step 5-2, and sub-step 5-3, we can obtain and where

[0041] Sub-step 5-5. Input and Using the input 2D key points as the training target for inverse implicit modeling, that is then there is Using the mean squared error calculation as the 2D key point position loss Using the training label as the training target for implicit modeling of time domain and spatial domain information, calculate the mean squared error of the 3D key point position loss

[0042] Furthermore, step 6 is specifically:

[0043] Sub-step 6-1. Input d all ={d i |d i ∈[3, 9], and i ∈ [1, n]}, convolution kernel ks, padding p, downsampling layer number n (same as the upsampling layer number), and repeat step (2). Output the dilated convolution layer set where DCL h+1 The superscript h + 1 indicates that this is the (h + 1)-th hourglass structure;

[0044] Sub-step 6-2. Input and DCL h+1 , and similarly to step (3), output the prediction result and calculate the corresponding 3D key point position loss

[0045] Sub-step 6-3. Overall loss of the key point detection model where λ > 0 is The trade-off coefficient between the output results with intermediate noise, and the key point model composed of multiple stacked dilated regular convolutional layers and dilated transposed convolutional layers is optimized using the stochastic gradient descent method.

[0046] The method of the present invention addresses the problem of implicit modeling of temporal and spatial information of skeletal key points, and has the following advantages: 1) The input does not need to contain depth information. Only the skeletal key point information between the 2D and 3D dimensions extracted by the dilated regular convolution is used to implicitly model the temporal information and depth information, improving the ability of depth space perception; 2) Using convolutional calculation to improve computational parallelism and facilitate acceleration; 3) Increasing the position loss of the intermediate layer of the calculation to facilitate model training optimization and reduce the requirements for the quality of training data; 4) The dilated regular convolution can be replaced with causal convolution operations, which is convenient for mobile deployment and real-time prediction.

[0047] The present invention has the ability to capture temporal information and refine skeletal key points, which can be used for any existing downstream tasks of 2D key point detectors to achieve real-time 3D skeletal key point detection operations, and can be widely applied to fields such as action recognition detection, pedestrian tracking, movie and animation production, virtual reality, medical assistance, autonomous driving, and motion analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is a flowchart of the complete model of the dilated transposed hourglass convolution of the method of the present invention;

[0049] Figure 2 is a schematic diagram of the information flow of the dilated transposed convolution hourglass structure of the method of the present invention;

[0050] Figure 3 is the internal detail of the dilated transposed convolution hourglass structure of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] The technical solutions of the present invention will be further specifically described below through specific embodiments and in conjunction with the accompanying drawings.

[0052] Embodiment 1

[0053] As Figure 1 and Figure 2 shown, a method for skeletal key point detection using a dilated transposed convolution hourglass structure includes:

[0054] Step (1). Construct a 2D skeletal key point detection model, input an RGB video stream, and obtain a 2D key point sequence;

[0055] Step (2). According to the set dilation factor, convolution kernel size, padding, and number of downsampling layers, etc., construct a dilated regular convolutional layer for downsampling, a dilated transposed convolutional layer for upsampling, and an activation operator;

[0056] Step (3). Perform temporal downsampling feature extraction operation, implicitly model the temporal and spatial depth relationship, and obtain noisy 3D bone key points and intermediate layer feature sets;

[0057] Step (4). Repeat Step (3) to implement skip connection and upsampling dilated transposed convolution, and extract intermediate feature sets of spatial and temporal information in the reverse implicit modeling process;

[0058] Step (5). According to the set number of hourglasses, repeat Steps (2) to (4) to obtain the output intermediate features of each hourglass, construct residual skip connections within and between hourglasses, and calculate the mean square error of the hourglass output;

[0059] Step (6). Repeat Steps (2) to (3) to obtain the final 3D bone key point results, calculate the mean square error, and optimize the detection model using the stochastic gradient descent method.

[0060] Furthermore, Step (1) is specifically:

[0061] (1-1). As Figure 3 , input the RGB video set containing the human motion to be detected where N represents the total number of video frames in I, H is the height of the input image, W is the width of the input image, and 3 represents the RGB three-channel image;

[0062] (1-2). Input the number of bone key points K ∈ N * , generally, K = 17. Construct a 2D bone key point detector D 2d , where D 2d can be any 2D bone key point detector that satisfies the input of I and the output of ;

[0063] (1-3). Input I into D frame by frame 2d , and for all temporal outputs of D 2d to obtain a 2D key point sequence where N is the total number of video frames in I, and 2 represents the X and Y dimensions of the 2D bone key points;

[0064] Furthermore, Step (2) is specifically:

[0065] (2-1). The input is the set of dilation factors d for each layer all ={d i |d i ∈[3,9], and i ∈ [1, n]}, the convolution kernel ks, the padding p, and the number of downsampling layers n. Here, the set d all will effectively expand the receptive field of the network model in the temporal dimension. Generally, ks is a 2D matrix and p = 1, n = 6. Corresponding to ks, dall , p and n inputs, output set of dilated convolutional layers where refers to the i-th dilated transposed convolution in the first hourglass structure;

[0066] (2 - 2). The input is the set of dilation factors d all = {d i | d i ∈ [3, 9], and i ∈ [1, n]}, convolution kernel ks, padding p, number of upsampling layers n, where the set d all will effectively expand the receptive field of the network model in the temporal dimension, outputting dilated transposed convolution where refers to the i-th dilated transposed convolution in the first hourglass structure, corresponds one-to-one with . Similar to

[0067] (2 - 1) operation, each layer shares the activation operator AL. The calculation processes of DCL and DTCL are opposite. Transposed convolution is used to implement the upsampling operation, reverse model the 3D information, that is, implicitly model the spatial depth information as temporal information to achieve the effects of refining information and self-supervision;

[0068] (2 - 3). Obtain the activation operator AL by instantiating the custom Mish(x) activation layer and output it. Since AL does not need to be trained, all activation operations share this activation operator, and the model will obtain non-linear expression ability accordingly;

[0069] (2 - 4). Further, the activation operator will act on any input tensor. For any input then the Mish(x) activation output is Mish(x) = x · tanh(ln(1 + e x )), and Mish(x) is a composite function of tanh((x);

[0070] (2 - 5). Still further, for the tanh(x) activation operator applied in Mish(x). For any input the output of tanh(x) is

[0071] Even further, step (3) is specifically:

[0072] (3 - 1). Input DCL and LA, pair the two, that is and l ∈ [1, n]. Multiple are stacked, that is (downsampling module, Down sample), output DCA 1 ;

[0073] (3 - 2). Input And successively pass through where \(l\in[1,n]\), then the output is the calculation information of each intermediate layer and the noisy 3D skeleton key point coordinate information At this time The superscript 0 indicates the 3D key point information obtained from the first hourglass structure;

[0074] (3 - 3). Input is the number of feature map channels of the dilated regular convolution in the \(l\)-th layer and \(l\in[0,n]\). If the number of input channels is The number of output channels is Then according to the equation Reduce the redundancy of channel information during this process and realize the implicit modeling of temporal information to spatial information;

[0075] (3 - 4). Input DTCL and AL, pair DTCL and the activation operator AL as DTCA, that is \(l\in[1,n]\). Stack to form (Downsampling module, Down sample), output DTCA 1 ;

[0076] (3 - 5). The number of input channels is Then the number of output channels is And Increase the redundancy of channel information during this process and realize the implicit modeling of spatial information to temporal information;

[0077] Furthermore, step (4) is specifically:

[0078] (4 - 1). Input where That is the output of Does not participate in the residual skip connection within the hourglass structure where it is located, but is directly input to to obtain the output And let

[0079] (4 - 2). Input and According to the residual skip connection formula Execute the residual skip connection. Such as Figure 3, this operation can effectively increase information fusion and alleviate problems such as vanishing gradients and exploding gradients caused by the excessive depth of the network model. During the training process, the network will autonomously weight and learn multi-layer information;

[0080] (4-3). Input to Perform upsampling operation and calculate, output the intermediate feature FU of upsampling n-2 , according to the residual skip connection formula and input it to By analogy, the intermediate feature set during the upsampling process can be obtained

[0081] (4-5). For the residual skip connection and l∈[1,n]. Further, if a complete hourglass structure contains n downsampling layers, then there are n corresponding upsampling layers. Therefore, in the complete hourglass and are paired one by one.

[0082] Furthermore, step (5) is specifically:

[0083] (5-1). Input The set of dilation factors d set for each layer all ={d i |d i ∈[3,9], and i∈[1,n]}, convolution kernel ks, padding p, number of downsampling layers n (the same as the number of upsampling layers), and repeat step (2).

[0084] Output the set of dilated convolutional layers where DCL 2 The superscript 2 indicates that this is the second hourglass structure, the set of dilated transposed convolutions Similarly, there is DCL all ={DCL 1 , DCL 2 ,…, DCL h}, DTCL all ={DTCL 1 , DTCL 2 ,…, DTCL h};

[0085] (5-2). Input DCL all , DTCL all and AL, repeat step (3) to output Similarly, there is DCA all ={DCA 1 , DCA 2,…,DCA h},DTCA all ={DTCA 1 ,DTCA 2 ,…,DTCA h};

[0086] (5 - 3) Input That is the output of is used to perform residual skip connection with , which is expressed by the formula

[0087] (5 - 4) Input Similarly, for steps (2), (3), (4), (5 - 1), (5 - 2) and (5 - 3), we can get and where

[0088] (5 - 5). Input and Using the input 2D key points as the training target for inverse implicit modeling, that is then we have Using the mean squared error calculation as the 2D key point position loss Using the training label as the training target for implicit modeling of time domain and spatial domain information, and calculating the mean squared error of the 3D key point position loss

[0089] Finally, step (6) is specifically:

[0090] (6 - 1). Input d all ={d i |d i ∈[3, 9], and i ∈ [1, n]}, convolution kernel ks, padding p, downsampling layer number n (same as the upsampling layer number), and repeat step (2). Output the dilated convolution layer set where DCL h+1 The superscript h + 1 indicates that this is the (h + 1)-th hourglass structure;

[0091] (6 - 2). Input and DCL h+1 , and similarly to step (3), output the final result and calculate the 3D key point position loss of the final result

[0092] (6 - 3). Overall loss of the key point detection model where λ > 0 is a trade-off coefficient between the and the output result with intermediate noise, and the key point model composed of multiple stacked dilated regular convolutional layers and dilated transposed convolutional layers is optimized using the stochastic gradient descent method.

[0093] The content described in this embodiment is only an enumeration of the implementation forms of the inventive concept. The protection scope of the present invention should not be regarded as limited to the specific forms stated in the embodiments. The protection scope of the present invention also covers equivalent technical means that can be conceived by those skilled in the art based on the inventive concept of the present invention.

Claims

1. A method for detecting skeletal key points using a hollow transposed convolution hourglass structure, characterized in that It includes the following steps: Step 1: Construct a 2D skeleton keypoint detection model, input an RGB video stream, and obtain a 2D keypoint sequence; Step 2: According to the set dilation factor, convolutional kernel size, padding, and number of downsampling layers, construct a dilated regular convolutional layer for downsampling, a dilated transposed convolutional layer for upsampling, and an activation operator; Sub-step 2-1, the input is the set of dilation factors d set for each layer all ={d i |d i ∈[3, 9], and i ∈ [1, n]}, convolutional kernel ks, padding p, number of downsampling layers n; p = 1, n = 6; corresponding to ks, d all , p and n are input, and the output is the set of dilated convolutional layers where indicates that this is the i-th dilated transposed convolution in the first hourglass structure; Sub-step 2-2, the input is the set of dilation factors d set for each layer all ={d i |d i ∈[3, 9], and i ∈ [1, n]}, convolution kernel ks, padding p, number of upsampling layers n, and the output is dilated transposed convolution where indicates that this is the i-th dilated transposed convolution in the first hourglass structure, corresponding one-to-one with ; each layer shares the activation operator AL; Sub-step 2-3: Instantiate the activation operator AL through a custom Mish(x) activation layer and output it. Since AL does not require training, all activation operations share this activation operator; Sub-step 2-4, for any input element the activation output of Mish(x) is Mish(x) = x · tanh(ln(1 + e x )); Sub-step 2-5, for any input element the output of tanh(x) is Step 3: Perform a temporal downsampling feature extraction operation, implicitly model the temporal and spatial depth relationship, and obtain noisy 3D skeleton keypoints and an intermediate layer feature set; Step 4: Repeat Step 3 to implement skip connections and upsampling dilated transposed convolutions, and extract the intermediate feature sets of spatial and temporal information during the reverse implicit modeling process; Step 5: According to the set number of hourglasses, repeat Steps 2 to 4 to obtain the output intermediate features of each hourglass, construct residual skip connections within and between hourglasses, and calculate the mean square error of the hourglass output; Step 6: Repeat Steps 2 and 3 to obtain the final 3D skeleton keypoint result; Calculate the mean square error and optimize the detection model using the stochastic gradient descent method.

2. The bone key point detection method using the hollow transposed convolutional hourglass structure according to claim 1, wherein The said Step 1 includes the following sub-steps: Sub-step 1-1: Input an RGB video set containing the human motion to be detected where N represents the total number of video frames contained in I, H is the height of the input image, W is the width of the input image, and 3 represents an RGB three-channel image; Sub-step 1-2, input the number of key points of the skeleton \(K\in N\). * , construct a 2D skeleton key point detector \(D\). 2d , where \(D\). 2d can be any 2D skeleton key point detector that satisfies the input of \(I\) and the output of ; Sub-step 1-3, input \(I\) frame by frame into \(D\). 2d , for \(D\). 2d all the temporal outputs are 2D key point sequences where \(N\) is the total number of video frames of \(I\), and 2 represents the \(X\) and \(Y\) dimensions of the 2D skeleton key points.

3. The bone key point detection method using the hollow transposed convolutional hourglass structure according to claim 2, characterized in that, The said Step 3 includes the following sub-steps: Sub-step 3-1: Input DCL and LA, pair them up, i.e., and there is l ∈ [1, n]; multiple stack them up, i.e., Output DCA 1 ; Sub-step 3-2, input and successively pass through where l ∈ [1, n], then the output is the calculation information of each intermediate layer and the 3D skeleton key point coordinate information with noise At this time The superscript 0 on it represents the 3D key point information obtained by the first hourglass structure; Sub-step 3-3, input is the number of feature map channels of the l-th layer of conventional hollow convolution, and l ∈ [0, n]. If the number of input channels is the number of output channels is then according to the equation Sub-step 3-4, input DTCL and AL, pair DTCL with the activation operator AL as DTCA, that is l ∈ [1, n]; Stacked structure Output DTCA 1 ; Sub-step 3-5, the number of input channels is then the number of output channels is and 4. A method for detecting skeletal key points using a hollow transposed convolutional hourglass structure according to claim 1, characterized in that The said Step 4 includes the following sub-steps: Sub-step 4-1, input wherein that is the output of does not participate in the residual skip connection within the hourglass structure where it is located, but is directly input into to obtain the output and let Sub-step 4-2, input and According to the residual skip connection formula Perform the residual skip connection; Sub-step 4-3, input to Perform upsampling operation and calculate, output the intermediate feature FU of upsampling n-2 , according to the residual skip connection formula and input to By analogy, the intermediate feature set in the upsampling process can be obtained Sub-step 4-5, for the participation in the residual skip connection and where \(l\in[1,n]\), in the complete hourglass structure and are paired one by one.

5. A method for detecting skeletal key points using a hollow transposed convolutional hourglass structure according to claim 1, characterized in that, The said Step 5 includes the following sub-steps: Sub-step 5-1, input The set d of dilation factors set for each layer all ={d i |d i ∈[3, 9], and i ∈ [1, n]}, convolutional kernel ks, padding p, number of downsampling layers n, and repeat step 2; output the set of dilated convolutional layers where DCL 2 The superscript 2 indicates that this is the second hourglass structure, the set of dilated transposed convolutional layers Similarly, there is DCL all ={DCL 1 , DCL 2 , …, DCL h}, DTCL all ={DTCL 1 , DTCL 2 , …, DTCL h}; Sub-step 5-2, input DCL all , DTCL all and AL, repeat step 3 to output Similarly, repeat step 4 to obtain DCA all ={DCA 1 , DCA 2 ,…, DCA h}, DTCA all ={DTCA 1 , DTCA 2 ,…, DTCA h}; Sub-step 5-3, input That is Output, let And Perform residual skip connection, which is expressed by the formula Sub-step 5-4, input According to Step 2, Step 3, Step 4, Step 5-1, Step 5-2 and Step 5-3, it can be obtained that and wherein Sub-step 5-5, input and Use the input 2D key points as the training target for inverse implicit modeling, that is Then there is Use the mean squared error calculation as the 2D key point position loss Use the training label As the training target for implicit modeling of time domain and spatial domain information, calculate the mean squared error of the 3D key point position loss 6. The bone key point detection method using a hollow transposed convolution hourglass structure according to claim 1, characterized in that The said Step 6 includes the following sub-steps: Sub-step 6-1, input d all = {d i | d i ∈ [3, 9], and i ∈ [1, n]}, convolutional kernel ks, padding p, number of downsampling layers n, and repeat step 2; output the set of dilated convolutional layers where DCL h+1 The superscript h + 1 indicates that this is the (h + 1)-th hourglass structure; Sub-step 6-2, input and DCL h+1 , and output the prediction result according to Step 3 and calculate the corresponding 3D key point position loss Sub-step 6-3, the overall loss of the key point detection model where λ > 0 is the trade-off coefficient between and the output result with intermediate noise, and optimize the key point model composed of multiple stacked dilated regular convolutional layers and dilated transposed convolutional layers using the stochastic gradient descent method.