A Fast Human Pose Estimation Method Based on Thumbnail Representation
Through the scale shrinker and dual-branch automatic coding training strategy in the ThumbPose framework, the accuracy reduction caused by small-sized images is solved, and it is suitable for various network models, improving the accuracy and computing efficiency of human pose estimation.
Patent Information
- Application Number
- CN202211184075.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-27
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-09-27
AI Technical Summary
The prior art leads to a problem of decreasing the accuracy of human pose estimation when reducing the amount of model calculations through small-sized images, and existing lightweight models are difficult to apply to advanced networks with non-stacked structures.
Using a fast human pose estimation method based on abbreviation representation, a training strategy of scale-reducing and dual-branch automatic encoding is introduced, and a framework called ThumbPose is designed, including scale-reducing, backbone network, upsampling module and network head, using style supervision and knowledge distillation technology to ensure that key information is retained in small-sized images.
It realizes the accuracy of human posture estimation while reducing the calculation amount. It is suitable for various network models, simplifies the implementation process and improves prediction accuracy.
Smart Images

Figure CN115601787B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision, and particularly to a fast human pose estimation method based on thumbnail representation. Background Art
[0002] The goal of human pose estimation is to determine the positions or spatial positions of the body key points (parts / joints) of a person from a given image or video. This technology uses image-based observations to obtain the poses of jointed humans, which are composed of joints and rigid parts. As an important basic research in the field of computer vision, human pose estimation is the technical basis for tasks such as behavior recognition, human intention prediction, and intelligent monitoring. Current human pose estimation research is mostly based on deep learning, and higher accuracy is achieved by using a large-scale fully convolutional network (FCN) to make inferences on larger-sized images. Although good performance has been achieved, the huge computational amount makes it difficult to be applied to real-time systems.
[0003] To accelerate the human pose estimation network, existing work compresses deep or wide network models into narrow (fewer channels) or shallow (fewer layers) networks and uses knowledge distillation technology to maintain the estimation accuracy. Although this method has achieved certain results, such a method is limited in the following aspects: (1) Simplicity. The powerful fitting ability of the FCN is highly coupled with its delicate structure. Therefore, it is necessary to carefully design the network contraction strategy to maintain its structural characteristics. (2) Universality. Existing lightweight models are obtained by directly reducing the number of Hourglass stacks. For example, an Hourglass stacked 8 levels repeatedly is reduced to be stacked 4 levels repeatedly. This method can be effectively applied to specific stacked structure networks, but cannot be generalized to other advanced non-stacked structure networks, such as SimpleBaseline and HRNet, etc.
[0004] Another effective method to reduce the computational complexity of deep models is to make inferences on small-sized images. As the spatial size of the feature map decreases, the required computational amount will decrease accordingly. This method can solve the limitations of the network compression method. First, this method only needs to simply reduce the size of the input image, avoiding weakening the representation ability of the network. In addition, this method can be generalized to any advanced model because the FCN allows inputs of any size. However, directly reducing the image size will inevitably lose some key information, resulting in a sharp decline in the final estimation performance. Summary of the Invention
[0005] The objective of the present invention is: in view of the problem that when reducing the computational amount of a model by using small-size images in the prior art, the accuracy of the model will decrease, to propose a fast human pose estimation method based on thumbnail representation.
[0006] The technical solution adopted by the present invention to solve the above technical problem is:
[0007] A fast human pose estimation method based on thumbnail representation, comprising:
[0008] The step of obtaining the image to be detected and getting the thumbnail of the image to be detected, and
[0009] The step of inputting the thumbnail of the image to be detected into the trained human pose estimation network for human pose estimation;
[0010] The human pose estimation network sequentially includes a downscaler, a backbone network, an upsampling module, and a network head;
[0011] Select an intermediate feature map in the backbone network as the monitoring point, the network layers before this monitoring point are the first half of the backbone network, and the network layers after this monitoring point are the second half of the backbone network;
[0012] The downscaler and the first half of the backbone network constitute an encoder;
[0013] The second half of the backbone network, the upsampling module, and the network head constitute a heatmap decoder;
[0014] The training process of the human pose estimation network is assisted by introducing a teacher network. The teacher network sequentially includes a backbone network and a network head, and the structures of the backbone network and the network head included in the teacher network are the same as those of the backbone network and the network head in the human pose estimation network;
[0015] The training process of the human pose estimation network includes two training stages:
[0016] Training stage 1:
[0017] The downscaler is supervised and trained using the style of the original-size image, and the encoder is supervised and trained using the intermediate feature map from the teacher network;
[0018] Training stage 2:
[0019] The encoder and the heatmap decoder are end-to-end supervised and trained using the joint heatmap from the teacher network and the real joint heatmap.
[0020] Further, the specific content of training stage 1 is:
[0021] First, by supervising the gap between the distribution of pixel gray values of the thumbnail image and the distribution of pixel gray values of the original-size image, a thumbnail with a distribution of pixel gray values consistent with that of the original-size image is obtained, where the distribution of pixel gray values includes the mean and variance of the pixel gray values;
[0022] A feature decoder is introduced, which is used to map the intermediate feature map in the encoder to a space with the same dimension as the intermediate feature map in the teacher network;
[0023] The network level of the intermediate feature map in the teacher network in the corresponding backbone network is the same as the network level of the intermediate feature map in the decoder in the corresponding backbone network;
[0024] By supervising the mean square error of the pixel gray values between the intermediate feature map in the teacher network and the intermediate feature map in the decoder, an intermediate feature map of the human pose estimation network consistent with the intermediate feature in the teacher network is obtained, so that the information extracted by the encoder is consistent with the information extracted by the first half of the backbone network in the teacher network. Figure 1 Specifically, the training phase 2 is as follows:
[0025] Specifically, the training phase 2 is as follows:
[0026] By supervising the mean square error of the pixel gray values and the KL divergence of the pixel gray value distribution between the joint heatmap from the teacher network and the joint heatmap from the human pose estimation network, a joint heatmap of the human pose estimation network consistent with the joint heatmap of the teacher network is obtained; Figure 1 By supervising the mean square error of the pixel gray values between the real joint heatmap and the joint heatmap from the human pose estimation network, an accurate joint heatmap of the human pose estimation network is obtained.
[0027] By supervising the mean square error of the pixel gray values between the real joint heatmap and the joint heatmap from the human pose estimation network, an accurate joint heatmap of the human pose estimation network is obtained.
[0028] Furthermore, the loss function for supervised training using the style of the original-size image is:
[0029]
[0030] where i represents the channel number, φ s (·) represents the operation process of the downscaler, μ(z i ) and σ 2 (z i ) represent the mean and variance operations respectively, λ represents a hyperparameter used to balance the mean and variance, and z represents the input image of the original size.
[0031] Furthermore, the loss function for supervised training using the intermediate feature map from the teacher network is:
[0032]
[0033] Among them, N represents the number of channels of the intermediate layer feature map, represents the mapping function from z to F of the teacher network, F t represents the intermediate feature map in the encoder, t represents the front half of the backbone network with the network layer before the monitoring point as the main part, and φ (·) represents the feature decoder, and the structure of the feature decoder is 1 deconvolution layer. d Furthermore, the loss function of the mean square error of the pixel gray values between the joint heatmaps from the teacher network and the joint heatmaps from the human pose estimation network is:
[0034]
[0035]
[0036] Among them, represents the process in which the teacher network maps the input image z to H t , H t represents the joint heatmap output by the teacher network, and respectively represent the embedding functions of the backbone network and the network head in the human pose estimation network, and φ u (·) represents the upsampling module, K represents the number of joint points on a person, and k represents the serial number of the joint point.
[0037] Furthermore, the loss function of the KL divergence is:
[0038]
[0039]
[0040]
[0041] Among them, W and L respectively represent the width and height of the joint heatmap, and respectively represent the probability distribution maps obtained after smoothing and normalizing H t and H s using the Softmax function containing the temperature coefficient, H s represents the joint heatmap output by the human pose estimation network, and respectively represent P t and P sThe grayscale value of the pixel at the spatial position (k, i, j) in the middle, τ represents a parameter for adjusting the softening degree of the distribution, and the subscripts k, i, j and k, o, v respectively represent the values of the elements at the spatial positions (k, i, j) and (k, o, v) in the three-dimensional tensor.
[0042] Further, the loss function of the mean square error of the pixel grayscale values between the true joint heatmap and the joint heatmap from the human pose estimation network is:
[0043]
[0044] Among them, represents the training label of the heatmap of the k-th joint point.
[0045] Further, the thumbnail of the image to be detected is obtained through a scaler;
[0046] The scaler includes: 1 interpolation operation and 2 convolutional layers, where each convolutional layer contains 1 convolutional operation with a kernel size of 5×5 and 1 ReLU activation function;
[0047] The first convolutional layer uses multiple convolutional kernels to expand the 3-channel RGB image into a 12-channel feature map;
[0048] The second convolutional layer restores the 12-channel feature map to a 3-channel RGB image;
[0049] The scaler uses a style-based loss function to supervise the pixel values of the output channels.
[0050] Further, the upsampling module includes 3 network layers. Among them, the first layer includes a bilinear interpolation layer, the second layer includes 1 convolutional operation and batch normalization, and the third layer sequentially includes 1 convolutional operation, BN, and ReLU activation function.
[0051] The beneficial effects of the present invention are:
[0052] This application adaptively filters out redundant information in the input image using an online image scaler based on style supervision, and compresses the remaining relevant information into a small-sized thumbnail, greatly reducing the network computation amount and ensuring the prediction accuracy; using a dual-branch autoencoder training strategy to introduce the knowledge distillation technique into the training of the lightweight network enables the thumbnail feature map to extract more key information and further improves the prediction accuracy. The trained scaler can directly replace the traditional image scaling technology. This application solves the problems of complex implementation, difficult generalization, and low accuracy of existing human pose estimation acceleration methods, promotes the development of lightweight human pose estimation technology, and provides certain technical support for the intelligent development of systems such as security monitoring and human-computer interaction. Description of the Drawings
[0053] Figure 1 Overall structure diagram of the ThumbPose, a fast human pose estimation framework based on thumbnail representation;
[0054] Figure 2 Structure diagram of the style-supervised online downscaler;
[0055] Figure 3 Flowchart of the training strategy for the dual-branch autoencoder. Detailed Implementation Modes
[0056] It should be specifically noted that, without conflict, the various implementation modes disclosed in this application can be combined with each other.
[0057] Detailed Implementation Mode 1: Refer to Figure 1 This detailed implementation mode specifically describes a fast human pose estimation method based on thumbnail representation, which includes:
[0058] Steps of obtaining the image to be detected and obtaining the thumbnail of the image to be detected, and
[0059] Steps of inputting the thumbnail of the image to be detected into the trained human pose estimation network for human pose estimation;
[0060] The human pose estimation network sequentially includes a downscaler, a backbone network, an upsampling module, and a network head;
[0061] Select an intermediate feature map in the backbone network as the monitoring point. The network layers before this monitoring point are the first half of the backbone network, and the network layers after this monitoring point are the second half of the backbone network;
[0062] The downscaler and the first half of the backbone network constitute the encoder;
[0063] The second half of the backbone network, the upsampling module, and the network head constitute the heatmap decoder;
[0064] The training process of the human pose estimation network is assisted by introducing a teacher network. The teacher network sequentially includes a backbone network and a network head. The backbone network and the network head included in the teacher network have the same structure as the backbone network and the network head in the human pose estimation network;
[0065] The training process of the human pose estimation network includes two training stages:
[0066] Training Stage 1:
[0067] The downscaler is supervised and trained using the style of the original-sized image, and the encoder is supervised and trained using the intermediate feature maps from the teacher network;
[0068] Training stage 2:
[0069] The encoder and the heatmap decoder are end-to-end supervised and trained using the joint heatmaps from the teacher network and the real joint heatmaps.
[0070] Aiming at the problems existing in the prior art, this application proposes a fast human pose estimation technology based on thumbnail representation, which is essentially a fast human pose estimation framework named ThumbPose based on thumbnail representation. Among them, aiming at the problems of complex implementation route and inability to be applied to various network models in the existing network compression-based methods, this application designs a simple and highly generalizable human pose estimation acceleration method based on the idea of small-size input. Aiming at the problem of low prediction accuracy of the small-size input model, this application first designs an online image downscaler based on style supervision (hereinafter referred to as the downscaler) to shrink the input image into a thumbnail as losslessly as possible. The downscaler removes the redundant information in the original-sized image, so that most of the key information helpful for accurate prediction is included in the small-size thumbnail. The input of the thumbnail will reduce the feature maps in the network to thumbnail feature maps. At the same time, in order to obtain more informative thumbnail feature maps, this application also designs a training strategy of dual-branch autoencoding, which aligns the feature maps to make the thumbnail feature maps contain the same information as their corresponding items in the original input network, thereby explicitly integrating knowledge distillation into ThumbPose to achieve further improvement in prediction accuracy. In addition, a well-trained downscaler can directly replace the traditional image downscaling technology, improving the model accuracy of small-size input.
[0071] To overcome the deficiencies in the current research on human pose estimation acceleration methods, that is, the existing methods have complex implementation routes; they cannot be applied to various network models; the prediction accuracy of the small-size input model is low. In view of this, this application starts from two key links of human pose estimation based on small-size input, namely image downscaling and training process. Based on the current advanced human pose estimation model as the basic architecture, it introduces an online image downscaler based on style supervision and a training strategy of dual-branch autoencoding, and finally proposes a fast human pose estimation framework ThumbPose based on thumbnail representation, as Figure 1 shown. This application divides the key technologies of the fast human pose estimation framework ThumbPose based on thumbnail representation into two modules: an online image downscaler based on style supervision and a training strategy of dual-branch autoencoding. The brief descriptions of each module are as follows.
[0072] In the online image downscaler based on style supervision, to adaptively filter redundant information in images and achieve image downscaling with any rational ratio, the basic structure of the downscaler is designed based on CNN. The downscaler consists of one interpolation operation and two convolutional layers. Each convolutional layer contains a convolutional operation with a kernel size of 5×5 and a ReLU activation function. To enhance the filtering ability, the first convolutional layer uses multiple convolutional kernels to expand the 3-channel RGB image into a 12-channel feature map, while the second convolutional layer restores the 12-channel feature map to 3 channels to match the RGB characteristics. By setting the stride factor, the second convolutional layer can achieve the image downsampling function. Through equivalent transformation, the downscaler can be regarded as composed of three modules: interpolation, 12-channel composite low-pass filter (Composite LPF), and downsampling. To ensure that the output thumbnail is in the same RGB format as the input, a style-based loss function is designed to supervise the pixel values of the output channels.
[0073] In the training strategy of the dual-branch autoencoder, to obtain more informative thumbnail feature maps, based on the knowledge distillation technique, the teacher network with the original-size input is used to guide the training of the inference network (i.e., the human pose estimation network) with the thumbnail input. The supervision of the teacher network guiding the inference network includes feature supervision based on the L2 distance, joint heatmap supervision based on the L2 distance, and joint heatmap supervision based on distribution perception. During the training process, the inference network is divided into three modules: Encoder, Feature Decoder (FDecoder), and Heatmap Decoder (HDecoder). The training for these three modules is divided into two stages: Encoder-FDecoder network training and Encoder-HDecoder network training.
[0074] This application aims to better evaluate human pose estimation technology. Taking the MPII and MS COCO 2017 datasets widely used in the field of human pose estimation as the experimental benchmarks, users can construct corresponding databases according to actual application requirements. For MPII, this dataset contains 40,000 human samples, and each sample is labeled with 16 joint point categories and their coordinates; the train set is used for training, while the val set is used for testing; the evaluation metrics use the mAP and mAR scores based on PCKh. For MS COCO 2017, this dataset contains 250,000 human samples, and each sample is labeled with 17 joint point categories and their coordinates; the train set is used for training, while the val set is used for testing; the evaluation metrics use the mAP and mAR scores based on Object Keypoint Similarity (OKS). The operation speed of the human pose estimation algorithm is measured by the number of floating-point operations GFLOPs. After establishing the training database, select some existing network models as examples, and a lightweight human pose estimation model can be trained based on the method proposed in this application.
[0075] As Figure 1 shown, the HRNet network is selected as an example for elaboration. This application first uses to define the joint heatmap of the training label, where K represents the number of joint points on a person, which represents the training label of the heatmap of the k-th joint point. ThumbPose includes a teacher network, an inference network, a supervision from the original size image style a supervision from the intermediate feature map of the teacher network a supervision from the joint heatmap of the teacher network and a supervision from the softened distribution of this heatmap a supervision from the ground truth of the joint heatmap The backbones of both the teacher network and the inference network are selected as HRNet, where the teacher network takes the original size image as input (taking 256×192 pixels as an example) and has been trained. For the teacher network, the original size image (z) is input into the network, and during its operation, an intermediate feature map (F t ) is generated, and finally, a joint heatmap (H t ) is output. These two types of feature maps are both used to online guide the training of the inference network. For the inference network, the Downscaler downsizes the input image to a thumbnail image (taking 128×96 pixels as an example), and then this thumbnail image is input into the backbone of the inference network and an Upsampling module, and finally, a joint heatmap (H s)。In addition, the intermediate feature map (F s ) of the inference network is extracted and calculated Therefore, a transposed convolutional layer is used to increase the size of F s to be consistent with F t .
[0076] Design an online downscaler based on style supervision. The current mainstream image downscaling method is affine transformation, which can be decoupled into an interpolation and a downsampling operation. The factorization of an affine downscaling operation (T s ) is as follows:
[0077]
[0078] where h and w represent the length and width of the original image, S* represents the downscaling coefficient, Ti and I* represent the interpolation operation and coefficient respectively; Td and D* represent the downsampling process and coefficient respectively. For an S*-fold downscaling operation, the input image is first interpolated by I* times and then downsampled by D* times. The Shannon sampling theorem has proven that spectral aliasing will cause distortion of the downsampled image. Therefore, image downscaling based on affine transformation will lose key information, resulting in a decrease in the accuracy of human pose estimation. To solve the aliasing distortion problem, the classical method is to use a low-pass filter before downsampling, but this requires designing the filter and adjusting the parameters, making the method relatively complex. To incorporate an advanced low-pass filter into the downscaling operation, this application uses CNN to achieve adaptive filtering of the aliasing components and image downsampling, and proposes an online downscaler based on style supervision, as shown in Figure 2As shown. The online downscaler based on style supervision consists of one interpolation operation and two convolutional layers. Each convolutional layer contains one convolutional operation with a kernel size of 5×5 and one ReLU activation function. To enhance the filtering ability, the first convolutional layer uses multiple convolutional kernels to expand the 3-channel RGB image into a 12-channel feature map, while the second convolutional layer restores the 12-channel feature map to 3 channels to match the RGB characteristics. By setting the stride factor, the second convolutional layer can achieve the image downsampling function. Through equivalent transformation, the downscaler can be regarded as consisting of three modules: interpolation, 12-channel Composite Low-Pass Filter (CompositeLPF), and downsampling. The trainable feature of CNN enables the 12-channel composite filter to adaptively learn advanced low-pass characteristics, which helps to reasonably filter out the aliasing components and preserve the key information of the original-size image with as low distortion as possible. Although the downscaler can acquire advanced filtering ability, this module only contains a few network layers and is at the last position of backpropagation, which makes it vulnerable to interference during training. Moreover, due to the lack of effective supervision, the output small-size image is not in the RGB form of the original input. Generating a realistic and low-distortion small-size image can limit the subsequent network training to optimize in the direction of more fully understanding the thumbnail information. To achieve this goal, the similarity between the thumbnail image and the original-size image is supervised to ensure the realism of the thumbnail. Inspired by the successful application of style transfer technology, the distribution (mean and variance) of pixel gray values is used to describe the style of the image. Therefore, the style-supervised loss function is expressed as:
[0079]
[0080] In the above formula, i represents the channel number; φ s (·) represents the operation process of the downscaler; μ(z i ) and σ 2 (z i ) represent the mean and variance operations respectively. These two statistics are independently calculated for each channel of each image; λ is a hyperparameter used to balance the two statistics. In addition to being supervised, the downscaler is also added to the entire network of human pose estimation and trained end-to-end with other components simultaneously. This online training method provides a result-oriented supervision, enabling the downscaler to filter out more information irrelevant to accurate detection, and thus preserving more key information in the thumbnail intact.
[0081] Design the training strategy of dual-branch auto-encoding. Although using an online downscaler based on style supervision can obtain low-distortion thumbnails, it cannot guarantee that all the key information relevant to accurate prediction can be included in such small-sized images, and the network cannot smoothly learn the ability to process feature maps with such high information density. To maintain the performance of the model, a teacher network (using the original-sized image input) is used to guide the training of the inference network in the way of knowledge distillation, aiming to make the thumbnail feature map contain the same information as the teacher network and let the inference network learn a more appropriate method for processing thumbnail information. Specifically, the intermediate layer feature map F of the inference network s is first selected as a monitoring point, and then the training strategy of dual-branch auto-encoding is designed accordingly to introduce all 5 supervisions and into the network training, as shown in Figure 3 . Define the encoder (Encoder) to be composed of the downscaler φ s (·) and the first half of the inference network , and the encoder maps the original-sized input image z to the latent code F s ; define the feature decoder (FDecoder) as a transposed convolutional layer φ d (·) to map F s to the space with the same dimension as F t ; define the heatmap decoder (HDecoder) to be composed of the upsampling module φ of the second half of the backbone network u and the network head to map F s to the joint heatmap H s of the predicted output. We divide the training of the inference network into two stages, where stage 1 trains the Encoder-FDecoder network and stage 2 trains the Encoder-HDecoder network.
[0082] Design training stage 1 to regularize the Encoder. In the Encoder-FDecoder network, the transposed convolutional layer φ d (·) consists of a transposed convolutional operation and a ReLU activation function, and is used to make the size of F s the same as that of F t , as shown in Figure 3 (a). By aligning the enlarged F s with F t , the information difference between the feature maps of the inference network and the teacher network can be reduced, and the corresponding loss function is calculated as follows:
[0083]
[0084] In the above formula, N represents the number of channels of the intermediate feature map; denotes the mapping function from z to F by the teacher network. This stage is essentially the regularization of the Encoder, enabling the Encoder to extract as much information as possible that is the same as the content contained in F t . To assist in information extraction, a style-based loss function is simultaneously used in this stage t to supervise the training of the downscaler.
[0085] Design training stage 2 to reconstruct the heatmap through the Encoder-FDecoder network in the way of knowledge distillation. Although the regularization in stage 1 makes the thumbnail feature map in the inference network contain almost the same information as the teacher network, some key information may still be lost. This is because the capacity of F s is much smaller than that of F t (such as 1 / 4 in the example). To make the thumbnail representation contain more valuable information, we further refine the information by enabling the HDecoder to learn to reconstruct H s from F t . As shown in (a) of Figure 3 , we insert an upsampling module between the backbone network and the network head to increase the size of H s to be the same as that of H t . The upsampling module consists of 3 network layers. The first layer is a bilinear interpolation layer, the second layer consists of 1 convolutional operation and batch normalization (BatchNormalization, BN), and the third layer consists of 1 convolutional operation, BN, and ReLU activation function in sequence. The loss function measures the difference between H s and H t using the mean square error and is used to supervise the reconstruction training. Its calculation is as follows:
[0086]
[0087] In the above formula denotes the process of the teacher network mapping the input image z to H t ; and represent the embedding functions of the backbone network and the network head in the inference network respectively; φ u (·) represents the upsampling module. To more finely distill the heatmap knowledge, the loss function is designed to supervise the alignment of the heatmap after distribution softening. For the heatmap corresponding to each joint point, it is defined as Where W and L represent width and height respectively. For reasonable knowledge transfer, the heatmaps of the inference network need to be subjected to the same softening process. For H t and H s The softening process is as follows:
[0088]
[0089]
[0090] In the above formula, and represent the probability distributions after softening and normalization of H t and H s respectively. It should be noted that these response regions belonging to background knowledge do not interfere with the joint point prediction accuracy, because the final joint point coordinates are calculated through the argmax function. KL divergence is used to measure the gap between two probability distributions, and its corresponding loss function is as follows:
[0091]
[0092] The training strategy of the double-branch autoencoder includes all 5 supervisions in the training of the inference network. Except for the and mentioned above, the Encoder-HDecoder network is also supervised by the alignment error between the Encoder-HDecoder network and the label joint point heatmap Hl, and its corresponding loss function is as follows:
[0093]
[0094] The overall training process of the inference network is shown in Algorithm 1. In the first stage of training, we train the Encoder-FDecoder network to regularize the encoder, as shown in Algorithm 1 (lines 2 to 4). In this stage, the hyperparameter α is used to balance and In the second stage of training, we train the Encoder-HDecoder network to generate the final joint point heatmap H s , as shown in Algorithm 1 (lines 5 to 7). In this stage, the network is fine-tuned by reducing the encoder learning rate by 10 times, and the hyperparameters β and γ are used to balance
[0095] Train a network model based on the fast human pose estimation technology using thumbnail representation proposed in this application. Adam is used as the optimizer during training. The standard deviation of the Gaussian distribution of the heatmap is 2, λ is set to 0.1, and τ is set to 4. α, β, and γ are set to 1, 0.5, and 0.1 respectively. In Figure 3 In the training strategy shown, the total number of loops is 236, where the first stage is the first 26 loops and the second stage is the remaining loops. The basic learning rate Lr is set to 1e-3 and is reduced to 1e-4 and 1e-5 at the 196th and 226th loops respectively. The reduction factor S* is set to 0.5. For experiments on MPII, the size of the original input is set to 256×256 pixels, and the size of the thumbnail is 128×128 pixels. For experiments on the MS COCO2017 dataset, the size of the original input is set to 256×192 pixels, and the size of the thumbnail is 128×96 pixels.
[0096] Through the above training, a fast human pose estimation network based on thumbnail representation can be obtained. Table 1 proves that this application can be applied to various network models and achieve an effective improvement in prediction accuracy. Table 2 compares this application with existing human pose estimation acceleration methods, proving that accelerating the model from the perspective of reducing the size of the feature map is an effective and competitive method compared to the method of reducing the network. Table 3 compares this application with existing human pose estimation models based on small-size inputs, proving that higher prediction performance can be obtained using the technology proposed in this application. As an accessory of this application, Table 4 uses a trained scaler to directly replace the bilinear interpolation scaling technology in other models, proving that the scaler can effectively replace traditional technologies and obtain better performance. In summary, this application breaks through the complexity and non-universality limitations of existing human pose estimation model acceleration methods, solves the problem of a significant reduction in human pose estimation prediction accuracy when using small-size inputs, and obtains a scaler that can directly replace traditional image scaling technologies.
[0097] Table 1 Universality Test (Training on the train set of MS COCO2017 and testing on the val set, input size 128×96 pixels)
[0098]
[0099] Table 2 Comparison with Existing Human Pose Estimation Model Acceleration Methods (Training on the train set of MPII and testing on the val set)
[0100]
[0101]
[0102] Table 3 Comparison with existing small-size input methods for human pose estimation (trained on the train set of MS COCO2017 and tested on the val set, input size 128×96 pixels)
[0103]
[0104] Table 4 Replacing the traditional image scaling technology in other models with a trained scaler (trained on the train set of MSCOCO2017 and tested on the val set, input size 128×96 pixels)
[0105]
[0106]
[0107]
[0108] It should be noted that the specific implementation manners are only explanations and illustrations of the technical solutions of the present invention, and the scope of the claimed rights cannot be limited thereby. Those that are only partial changes made according to the claims and the description of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A fast human pose estimation method based on thumbnail representation, characterized in that Including: The steps of obtaining the image to be detected and getting the thumbnail of the image to be detected, and The step of inputting the thumbnail of the image to be detected into the trained human pose estimation network for human pose estimation; The human pose estimation network sequentially includes a downscaler, a backbone network, an upsampling module, and a network head; Select an intermediate feature map in the backbone network as the monitoring point. The network layers before this monitoring point are the first half of the backbone network, and the network layers after this monitoring point are the second half of the backbone network; The downscaler and the first half of the backbone network constitute an encoder; The second half of the backbone network, the upsampling module, and the network head constitute a heatmap decoder; The training process of the human pose estimation network is assisted by introducing a teacher network. The teacher network sequentially includes a backbone network and a network head. The backbone network and the network head included in the teacher network have the same structure as the backbone network and the network head in the human pose estimation network; The training process of the human pose estimation network includes two training stages: Training stage 1: The downscaler is supervised and trained using the style of the original-size image, and the encoder is supervised and trained using the intermediate feature map from the teacher network; Training stage 2: The encoder and the heatmap decoder are supervised and trained end-to-end using the joint heatmap from the teacher network and the real joint heatmap.
2. The fast human pose estimation method based on thumbnail representation according to claim 1, wherein Specifically, the training stage 1 is as follows: First, by supervising the gap between the distribution of pixel gray values of the thumbnail image and the distribution of pixel gray values of the original-size image, a thumbnail with a distribution of pixel gray values consistent with that of the original-size image is obtained. The distribution of pixel gray values includes the mean and variance of pixel gray values; Introduce a feature decoder, which is used to map the intermediate feature map in the encoder to a space with the same dimension as the intermediate feature map in the teacher network; The network level of the intermediate feature map in the teacher network in the corresponding backbone network is the same as the network level of the intermediate feature map in the decoder in the corresponding backbone network; By supervising the mean square error of pixel gray values between the intermediate feature map in the teacher network and the intermediate feature map in the decoder, an intermediate feature map of the human pose estimation network consistent with the intermediate feature map in the teacher network is obtained, so that the information extracted by the encoder is consistent with the information extracted by the first half of the backbone network in the teacher network.
3. A fast human pose estimation method based on thumbnail representation according to claim 2, characterized in that Specifically, the training stage 2 is as follows: By supervising the mean square error of pixel gray values and the KL divergence of the pixel gray value distribution between the joint heatmap from the teacher network and the joint heatmap from the human pose estimation network, a joint heatmap of the human pose estimation network consistent with the joint heatmap of the teacher network is obtained; By supervising the mean square error of pixel gray values between the real joint heatmap and the joint heatmap from the human pose estimation network, an accurate joint heatmap of the human pose estimation network is obtained.
4. A fast human pose estimation method based on thumbnail representation according to claim 3, characterized in that The loss function for the supervised training using the style of the original-size image is: Among them, i represents the channel number, φ s (·) represents the operation process of the downscaler, μ(z i ) and σ 2 (z i ) represent the mean and variance operations respectively, λ represents the hyperparameter used to balance the mean and variance, and z represents the input image of the original size.
5. A fast human pose estimation method based on thumbnail representation according to claim 4, characterized in that The loss function for the supervised training using the intermediate feature map from the teacher network is: Among them, N represents the number of channels of the intermediate layer feature map, represents the mapping function from z to F of the teacher network, F t represents the intermediate feature map in the encoder, t represents that the network layer before the monitoring point is the first half of the backbone network, φ d (·) represents the feature decoder, and the structure of the feature decoder is 1 transposed convolution layer. 6. A fast human pose estimation method based on thumbnail representation according to claim 5, characterized in that The loss function of the mean squared error of the pixel gray values between the joint heatmap from the teacher network and the joint heatmap from the human pose estimation network is as follows: Among them, denotes the process in which the teacher network maps the input image z to H t where H t denotes the joint heatmap output by the teacher network, and denote the embedding functions of the backbone network and the network head in the human pose estimation network respectively, and φ u (·) denotes the upsampling module, K denotes the number of joint points on a person, and k denotes the serial number of the joint point.
7. A fast human pose estimation method based on thumbnail representation according to claim 6, characterized in that The loss function of the KL divergence is as follows: Where W and L represent the width and height of the joint heatmap respectively, and respectively represent the probability distribution maps obtained by smoothing and normalizing H using the Softmax function including the temperature coefficient, t and H s After that, the obtained probability distribution map, H s represents the joint heatmap output by the human pose estimation network, and respectively represent P t and P s The grayscale values of the pixel points at the spatial position (k, i, j) in, τ represents a parameter used to adjust the softening degree of the distribution, and the subscripts k, i, j and k, o, v respectively represent the values of the elements at the spatial positions (k, i, j) and (k, o, v) in the three-dimensional tensor.
8. A fast human pose estimation method based on thumbnail representation according to claim 7, characterized in that The loss function of the mean squared error of the pixel gray values between the true joint heatmap and the joint heatmap from the human pose estimation network is as follows: Among them, represents the heatmap training label of the k-th joint point.
9. A fast human pose estimation method based on thumbnail representation according to claim 1, characterized in that The thumbnail of the image to be detected is obtained through a downscaler; The downscaler consists of 1 interpolation operation and 2 convolutional layers, where each convolutional layer contains 1 convolutional operation with a kernel size of 5×5 and 1 ReLU activation function; The first convolutional layer uses multiple convolutional kernels to expand the 3-channel RGB image into a 12-channel feature map; The second convolutional layer restores the 12-channel feature map to a 3-channel RGB image; The downscaler uses a style-based loss function to supervise the pixel values of the output channels.
10. A fast human pose estimation method based on thumbnail representation according to claim 1, characterized in that The upsampling module includes 3 network layers. Among them, the first layer includes a bilinear interpolation layer, the second layer includes 1 convolutional operation and batch normalization, and the third layer sequentially includes 1 convolutional operation, BN, and ReLU activation function.
Citation Information
Patent Citations
Generative adversarial network training method and device and image style migration method and device
CN111862274A
Underwater binocular depth estimation method based on unsupervised adaptive network
CN114299130A