A lightweight real-time video pose estimation method based on Anchor-free
Through the lightweight real-time video pose estimation method based on Anchor-free, the NAFNet framework and the denoising-recovery-global characteristic algorithm are used to solve the problem of difficult to take into account real-time and accuracy in the prior art, and efficient and accurate video human pose estimation is achieved.
Patent Information
- Application Number
- CN202211343627.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-29
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-10-29
AI Technical Summary
The existing video pose estimation technology is difficult to balance between real-time and accuracy, the model is large and time-consuming, and cannot meet real-time needs.
A lightweight real-time video pose estimation method based on Anchor-free is used to perform pose estimation of keyframes through the NAFNet framework, and the remaining frames are completed using the denoising-recovery-global characteristic algorithm to complete the pose estimation of the entire video.
It realizes efficient identification of human postures in video, improves the accuracy and efficiency of estimation of human postures in video, and can increase the number of parameters in a small amount, while greatly improving the far and near targets.
Smart Images

Figure CN115661933B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pattern recognition and computer vision, and in particular to a lightweight real-time video pose estimation method based on Anchor-free. Background Art
[0002] With the development of artificial intelligence technology, computer vision has more and wider applications in real life, such as face recognition, autonomous driving, etc. The mature development of neural networks has further expanded people's imagination, making computer vision-related applications more practical and convenient in various application scenarios. Video pose estimation technology has great applications in scenarios such as sports competitions, fitness exercises, and learning concentration. For the application to be implemented, both real-time performance and accuracy need to be considered. Currently, many video pose estimation technologies have high accuracy, but the models are large and time-consuming, unable to meet the real-time requirements. In order to improve the comfort of more users using AI applications, we need a lightweight video pose estimation technology with stronger real-time performance. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide a lightweight real-time video pose estimation method based on Anchor-free, which can effectively recognize human poses in videos.
[0004] To achieve the above purpose, the present invention adopts the following technical solutions: A lightweight real-time video pose estimation method based on Anchor-free, comprising the following steps:
[0005] Step S1: Obtain a video dataset, and select key frames through a key frame extraction algorithm;
[0006] Step S2: Use the lightweight method NAFNet based on the Anchor-free framework to perform pose estimation on the key frames;
[0007] Step S3: Obtain a human pose estimation dataset and perform lightweight model training;
[0008] Step S4: Use the denoising-restoration-global feature algorithm to complete the pose estimation of the entire video by complementing the remaining frames.
[0009] In a preferred embodiment, the specific steps of step S1 are as follows:
[0010] Step S11: Obtain a video human pose estimation dataset from the Internet and obtain its annotations;
[0011] Step S12: For a video in the dataset, given It is represented that the video V contains a total of T frames, I tRepresents the video image at time t; approximately one fifth of the n key frames are selected from the video sequence at equal intervals, and the key frame sequence is F = {f 1 , f 2 , …, f n};
[0012] Step S13: Use the two-frame difference method to fine-tune the extracted key frame, record now as the current time, and the current frame f now (x, y) and the next frame f in the key frame sequence now+1 (x, y), where x and y represent points on the horizontal and vertical axes respectively, and the pixel difference between the two frames is D now (x, y); the formula is:
[0013] D now (x, y) = |f now (x, y)-f now+1 (x, y)|
[0014] Let the threshold be T, if D now (x, y) is less than the threshold, then f n+1 Frame is replaced by f n+1 Frame corresponding to I t The next frame I t+1 , and finally the final key frame sequence S is obtained.
[0015] In a preferred embodiment, the specific method of step S2 is:
[0016] Step S21: Based on the Anchor-free human posture estimation framework NAFNet, a lightweight Backbone is used for feature extraction. The backbone network is based on the MicroNet network, and the undecomposed convolution part of the network is replaced with a lightweight Ghost convolution block. In the lightweight Ghost module, the number of channels of all convolution layers starting from the second layer is 1 / 2 of the number of output channels. The output features of the remaining 1 / 2 channels are generated by cheap operations of all previous convolution layers. At the same time, a lightweight global attention module is integrated into the network; the attention module uses a 1×1 depth-separable convolution DW 1×1 For each set of feature FMap k The channel is extracted as a new feature, k represents the group index value, and after the maximum pooling Maxpool, the attention feature map with the number of channels as 1 is formed by point convolution PointConv, and the attention feature map is scaled using the Softmax activation function; the formula is:
[0017] AMap k =Softmax(PointConv(Maxpool(DW 1×1 (FMap k))))
[0018] Obtain the attention matrix AMap k After that, for each group of features FMap k Multiply each element of the attention matrix corresponding to the features, and then add it to the original features of each group to obtain the processed features F of this group k , the formula is:
[0019] F k =(AMap k ·FMap k )+FMap k
[0020] After concatenating k groups of features through the Concat operation, extract the final feature map;
[0021] Step S22: The Neck part adopts a simplified version of the PAN structure, and replaces the convolution in the original PAN network structure with the same lightweight Ghost module as in Step S21;
[0022] Step S23: The detection part uses an Anchor point in the middle and a vectorized prediction method to predict human key points. The loss part uses a more efficient version of Focalloss that can balance positive and negative samples. The formula is:
[0023] EFL(σ)=-|μ - σ| β ((μ - 1)log(1 - μ)(1 - σ)+μ(1 - σ)logσ)
[0024] Where μ is the quality label from 0 to 1, σ is the predicted value, and β is the adjustment factor; first train a group of initial feature points to obtain a group of initial vectors V first , and then use the dynamic deformable convolution DyCconv to obtain another group of vectors V at the key points DyCconv , add the two obtained features to form the final key point feature V last , and use these predicted vectors to obtain the final pose estimation result.
[0025] In a preferred embodiment, the step S3 specifically includes the following steps:
[0026] Step S31: Obtain the pose estimation dataset COCO from the Internet and obtain its human key point annotations;
[0027] Step S32: Use the novel Anchor-free human pose estimation framework NAFNet proposed in step S2 for training to obtain a human pose training model for quickly detecting single-frame images;
[0028] Step S33: After obtaining the training model, detect the extracted key frame sequence S to obtain the detected result P pre , P pre = NAFNet(S);
[0029] In a preferred embodiment, in step S4, the following steps are specifically included:
[0030] Step S41: Input the detected result P pre into the denoising network IndenoiseNet to obtain the posture P after denoising processing clean ; P Clean = IndenoiseNet(P pre ); The denoising network IndenoiseNet first converts the noisy image into a low-resolution image and high-frequency encoding, and the training formula is:
[0031]
[0032] where γ represents the noisy image, low represents the low frequency, and g(γ) low refers to the low-frequency part learned by the network, x low is the true low-frequency part of the image, w rorw is the number of pixel points of the entire image, index represents the index value of the pixel point, and this operation g is reversible;
[0033] Step S42: After obtaining the denoised image, it is necessary to restore the entire image sequence according to the sparse frames and use RecoverNet for restoration; the training formula is:
[0034]
[0035] where w back represents the number of pixel points of the entire image after processing, P Clean represents the denoised posture result, G~N(0, 1) is a variable randomly sampled from the normal distribution to supplement the detailed information of the high-frequency part; g -1 represents the reversible operation in step S41, and g and g -1 should be trained simultaneously; at the same time, 1D time information also exists in RecoverNet to restore the postures of neighboring frames, and the formula is:
[0036] P = RecoverNet(Conv1d(P clean ), L Recover (g(γ), x low ))
[0037] Use this operation to restore the remaining frame sequences, and finally output the video human pose estimation result P.
[0038] Compared with the prior art, the present invention has the following beneficial effects:
[0039] 1. It can efficiently recognize human poses in videos and improve the accuracy of video human pose estimation.
[0040] 2. It can utilize multi-scale information to process features, and while slightly increasing the number of parameters, it can greatly improve both near and far targets.
[0041] 3. Compared with the traditional Top-down pose estimation idea, the present invention is based on the Anchor-Free idea and proposes a lightweight framework NAFNet to more efficiently complete the pose estimation task.
[0042] 4. Aiming at the problem of slow frame-by-frame detection and high time consumption in videos, a denoising-restoration-global feature algorithm is proposed to improve the efficiency and accuracy of video human pose estimation. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is a flowchart of the method of the preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0044] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0045] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0046] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0047] As Figure 1 shown, the present invention provides a lightweight real-time video pose estimation and recognition method based on Anchor-free, including the following steps:
[0048] Step S1: Obtain a video data set and select key frames through a key frame extraction algorithm;
[0049] Step S2: Use the lightweight method NAFNet based on the Anchor-free framework to perform pose estimation of the key frame;
[0050] Step S3: Obtain a human posture estimation dataset and perform lightweight model training;
[0051] Step S4: Use the denoising-restoration-global characteristic algorithm to complete the remaining frames and complete the posture estimation of the entire video.
[0052] Furthermore, the step S1 specifically includes the following steps:
[0053] Step S11: Obtain a video human posture estimation dataset from the Internet and obtain its annotations;
[0054] Step S12: For a video in the dataset, given Indicates that the video V contains T frames, I t Represents the video image at time t. From the video sequence, approximately one fifth of the n key frames are selected at equal intervals. The key frame sequence is F = {f 1 , f 2 , …, f n};
[0055] Step S13: Use the two-frame difference method to fine-tune the extracted key frame, record now as the current time, and the current frame f now (x, y) and the next frame f in the key frame sequence now+1 (x, y), where x and y represent points on the horizontal and vertical axes respectively, and the pixel difference between the two frames is D now (x, y). The formula is:
[0056] D now (x, y) = |f now (x, y)-f now+1 (x, y)|
[0057] Let the threshold be T, if D now (x, y) is less than the threshold, then f n+1 Frame is replaced by f n+1 Frame corresponding to I t The next frame I t+1 , and finally obtain the final key frame sequence S;
[0058] Furthermore, the specific method of step S2 is:
[0059] Step S21: The proposed Anchor-free human pose estimation framework NAFNet uses a lightweight Backbone for feature extraction. The backbone network is based on the MicroNet network, and the non-decomposed convolutional part in the network is replaced with lightweight Ghost convolution blocks. In the lightweight Ghost module, starting from the second layer, the number of channels of all convolutional layers is 1 / 2 of the output channels, and the output features of the remaining 1 / 2 channels are generated by cheap operations performed by all previous convolutional layers respectively. At the same time, a lightweight global attention module is fused in the network. This attention module uses a 1×1 depthwise separable convolution DW 1×1 For each group of features FMap k Extract the channels as new features. k represents the group index value. After performing max pooling Maxpool, then form an attention feature map with 1 channel through point convolution PointConv, and use the Softmax activation function to scale the attention feature map. The formula is:
[0060] AMap k = Softmax(PointConv(Maxpool(DW 1×1 (FMap k ))))
[0061] Obtain the attention matrix AMap k After that, multiply each group of features FMap k element-wise with the attention matrix, and then add it to the original each group of features to obtain the processed group of features F k , and the formula is:
[0062] F k =(AMap k ·FMap k )+FMap k
[0063] After that, through the Concat operation, connect the k groups of features, and then extract to obtain the final feature map.
[0064] Step S22: The Neck part adopts a simplified version of the PAN structure, replacing the convolutions in the original PAN network structure with the same lightweight Ghost modules as in Step S21. Compared with traditional methods, it reduces the computational amount while hardly reducing the fusion efficiency;
[0065] Step S23: In the detection part, we use an Anchor point in the middle and a vectorized prediction method to predict human key points. The loss part uses a more efficient version of Focalloss that can balance positive and negative samples better. The formula is:
[0066] EFL(σ)= -|μ - σ|β ((μ - 1)log(1 - μ)(1 - σ) + μ(1 - σ)logσ)
[0067] Where μ is a quality label from 0 to 1, σ is the predicted value, and β is a regulation factor, generally taken as 2. We first train a set of initial feature points to obtain a set of initial vectors V first , and then use the dynamic deformable convolution DyCconv to obtain another set of vectors V at the key points DyCconv . Add the two obtained features to form the final key point feature V last . Use these predicted vectors to obtain the final pose estimation result
[0068] Furthermore, the step S3 includes the following steps:
[0069] Step S31: Obtain the pose estimation dataset COCO from the Internet and obtain its human key point annotations
[0070] Step S32: Use the novel Anchor-free human pose estimation framework NAFNet proposed by us in step S2 for training to obtain a human pose training model for quickly detecting single-frame images
[0071] Step S33: After obtaining the training model, detect the extracted key frame sequence S to obtain the detected result P pre , P pre = NAFNet(S)
[0072] Furthermore, in the step S4, it specifically includes the following steps:
[0073] Step S41: To reduce the noise of single-frame detection, input the detected result P pre into the denoising network IndenoiseNet to obtain the pose P after denoising processing Clean . P Clean = IndenoiseNet(P pre ). This network is different from traditional neural networks. The IndenoiseNet designed by us first converts the noisy image into a low-resolution image and high-frequency encoding. The training formula is:
[0074]
[0075] Where γ represents the noisy image, low represents the low frequency, g(γ) low refers to the low-frequency part learned by the network, x low is the true low-frequency part of the image, w forwis the number of pixel points in the entire image, index represents the index value of the pixel point, and this operation g is reversible;
[0076] Step S42: After obtaining the denoised image, the entire image sequence needs to be restored according to the sparse frames. We designed a RecoverNet for restoration. The training formula is:
[0077]
[0078] where w back represents the number of pixel points in the entire image after processing, P Clean represents the pose result after denoising, G~N(0, 1) is a variable randomly sampled from the normal distribution to supplement the detailed information of the high-frequency part. g -1 represents the reversible operation in step S41, and g and g -1 should be trained simultaneously. At the same time, 1D time information also exists in RecoverNet to restore the poses of neighboring frames. The formula is:
[0079] P = RecoverNet(Conv1d(P clean ), L Recover (g(γ), x low ))
[0080] Use this operation to restore the remaining frame sequences, and finally output the video human pose estimation result P.
[0081] The above are the preferred embodiments of the present invention. All changes made according to the technical solution of the present invention, when the functions and effects produced do not exceed the scope of the technical solution of the present invention, shall fall within the protection scope of the present invention.
Claims
1. A lightweight real-time video pose estimation method based on Anchor-free, characterized in that, it includes the following steps: Step S1: Obtain a video dataset, and select key frames through a key frame extraction algorithm; Step S2: Use the lightweight method NAFNet based on the Anchor-free framework to perform pose estimation on the key frames; Step S3: Obtain a human pose estimation dataset and perform lightweight model training; Step S4: Use the denoising-restoration-global feature algorithm to complete the pose estimation of the entire video by completing the remaining frames; The specific method of step S2 is: Step S21: Based on the Anchor-free human pose estimation framework NAFNet, use a lightweight Backbone for feature extraction. The backbone network is based on the MicroNet network, and the non-decomposed convolutional part in the network is replaced with lightweight Ghost convolutional blocks. In the lightweight Ghost module, the number of channels of all convolutional layers starting from the second layer is 1 / 2 of the output channels, and the output features of the remaining 1 / 2 channels are generated by cheap operations performed by all previous convolutional layers respectively. At the same time, a lightweight global attention module is fused in the network; The attention module uses a 1×1 depthwise separable convolution DW 1×1 for each group of features FMap k extracts channels as new features. k represents the group index value. After performing max pooling Maxpool, it then forms an attention feature map with 1 channel through point convolution PointConv, and scales the attention feature map using the Softmax activation function; the formula is: AMap k = Softmax(PointConv(Maxpool(DW 1×1 (FMap k )))) Obtain the attention matrix AMap k After that, for each group of features FMap k Multiply it element-wise with the attention matrix, and then add the result to the original features of each group to obtain the processed features F of this group k , and the formula is: F k = (AMap k · FMap k ) + FMap k After concatenating k groups of features through the Concat operation and extracting, the final feature map is obtained; Step S22: The Neck part uses a simplified version of the PAN structure, and replaces the convolutions in the original PAN network structure with the same lightweight Ghost modules as in step S21; Step S23: In the detection part, one Anchor point in the middle and a vectorized prediction method are used to predict human key points. The loss part uses a more efficient version of Focalloss that can balance positive and negative samples. The formula is: EFL(σ) = -|μ - σ| β ((μ - 1)log(1 - μ)(1 - σ) + μ(1 - σ)logσ) where μ is a mass label between 0 and 1, σ is the predicted value, and β is a tuning factor; first, a set of initial feature points are trained to obtain a set of initial vectors V first , and then another set of vectors V at the key points are obtained using the dynamic deformable convolution DyCconv DyCconv . The two features obtained are added together to form the final key point feature V last . Using these predicted vectors, the final pose estimation result is obtained.
2. A lightweight real-time video pose estimation method based on Anchor-free according to claim 1, characterized in that, the specific steps of step S1 include the following steps: Step S11: Obtain a video human pose estimation dataset from the Internet and obtain its annotations; Step S12: For a video in the dataset, given represents that video V contains a total of T frames, and I t represents the video image at time t; Approximately one-fifth of n key frames are initially selected equidistantly from the video sequence, and the key frame sequence is F = {f 1 , f 2 , …, f n}; Step S13: Use the two-frame difference method to fine-tune the extracted key frames. Denote now as the current moment and the current frame as f now (x, y), and the next frame f now+1 (x, y) in the key frame sequence, where x and y respectively represent points on the horizontal and vertical axes, and the pixel difference between the two frames is D now (x, y); The formula is: D now (x,y) = |f now (x,y) - f now+1 (x,y)| Let the threshold be T. If D now (x, y) is less than the threshold, then replace the f n+1 frame with the f n+1 frame corresponding to the I t next frame I t+1 , and finally obtain the final key frame sequence S.
3. A lightweight real-time video pose estimation method based on Anchor-free according to claim 1, characterized in that, the specific steps of step S3 include the following steps: Step S31: Obtain the pose estimation dataset COCO from the Internet and obtain its human key point annotations; Step S32: Use the Anchor-free human pose estimation framework NAFNet in step S2 for training to obtain a human pose training model for quickly detecting single-frame images; Step S33: After obtaining the trained model, detect the extracted key-frame sequence S to obtain the detected result P pre , P pre = NAFNet(S).
4. A lightweight real-time video pose estimation method based on Anchor-free according to claim 1, characterized in that, in step S4, it specifically includes the following steps: Step S41: The detected result P pre is input into the denoising network IndenoiseNet to obtain the pose P after denoising processing clean ; P clean = IndenoiseNet(P pre ); The denoising network IndenoiseNet first converts the noisy image into a low-resolution image and high-frequency coding, and the training formula is: Among them, γ represents the noisy image, low represents the low frequency, and g(γ) low refers to the low-frequency part learned by the network, x low is the true low-frequency part of the image, w forw is the number of pixels in the entire image, index represents the index value of the pixel, and the operation g is reversible; Step S42: After obtaining the denoised image, it is necessary to restore the entire image sequence according to the sparse frames and use RecoverNet for restoration. The training formula is: where w back represents the number of pixels in the entire image after processing, P clean represents the pose result after denoising, G ∼ N(0, 1) is a variable randomly sampled from a normal distribution, and details of the high-frequency part are supplemented; g -1 represents the reversible operation in step S41, and g and g -1 should be trained simultaneously; at the same time, 1D time information also exists in RecoverNet to recover the poses of neighboring frames. The formula is: P = RecoverNet(Conv1d(P clean ), L Recover (g(γ), x low )) Use this operation to restore the remaining frame sequences, and finally output the video human pose estimation result P.
Citation Information
Patent Citations
Model constraint-based on-orbit 3D space target attitude estimation method and system
CN104748750A