Method, apparatus, device and storage medium for estimating human body posture
Through the combination of adaptive attitude pooling module, multi-layer perceptron and information fusion module, the problem of insufficient boundary information in traditional three-dimensional human posture estimation is solved, the continuity and accuracy of three-dimensional human posture estimation is improved, and complex action scenes are adapted to, and a natural and coherent attitude sequence is generated.
Patent Information
- Application Number
- CN202510601980.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-12
AI Technical Summary
The traditional three-dimensional human posture estimation method ignores boundary information when processing video clips, resulting in discontinuous time and incoherent movements, especially when complex movements or rapid changes, the action details are lost or the transition is unnatural.
Through the combination of adaptive attitude pooling module, multi-layer perceptron and information fusion module, feature extraction and edge information compensation are performed to improve the continuity and accuracy of three-dimensional human posture estimation.
The action consistency and accuracy of three-dimensional human posture estimation is realized, adapting to complex action scenes, and generating a naturally coherent posture sequence.
Smart Images

Figure CN120108005B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image recognition technology. More specifically, this application relates to a method, apparatus, device, and storage medium for estimating human body postures. Background Art
[0002] Three-dimensional human body pose estimation technology has been widely used in fields such as motion analysis, virtual reality, and motion capture. Traditional methods for estimating three-dimensional human body postures usually divide a video into video segments of a fixed length. Then, all video segments are processed in two stages. Specifically, in the first stage, a two-dimensional human body pose estimator is used to extract two-dimensional human body postures, and in the second stage, the two-dimensional human body postures are mapped into three-dimensional space to form three-dimensional human body postures. Although this method can capture local temporal dynamics within a single segment, it ignores the problem of insufficient boundary information of video segments, resulting in the estimation results of three-dimensional human body postures being prone to temporal discontinuity and unsmooth actions. Summary of the Invention
[0003] The objective of the embodiments of this application is to provide a method, apparatus, device, and storage medium for estimating human body postures, which can improve the continuity and accuracy of three-dimensional human body pose estimation. The embodiments of this application are mainly implemented through the following technical solutions:
[0004] In the first aspect of the embodiments of this application, a method for estimating human body postures is provided, including:
[0005] Obtain a video sequence;
[0006] Input the video sequence into a video segmentation module for segmentation processing to obtain multiple video segments;
[0007] Input a target video segment into a two-dimensional human body pose estimator for feature extraction processing to obtain two-dimensional human body pose information corresponding to the target video segment and image features of multiple resolutions, where the target video segment is any one of the multiple video segments;
[0008] Input the two-dimensional human body pose information and the image features of multiple resolutions into a human body pose estimation model for feature extraction and edge information compensation processing to obtain a target fusion feature corresponding to the target video segment;
[0009] Input the target fusion feature into a first multi-layer perceptron for pose estimation processing to obtain a three-dimensional human body pose estimation result corresponding to the target video segment;
[0010] Merge all the three-dimensional human body pose estimation results to form a three-dimensional human body pose estimation sequence.
[0011] According to an embodiment of the present application, the step of inputting the two-dimensional human pose information and the image features of multiple resolutions into a human pose estimation model for feature extraction and edge information compensation processing to obtain target fusion features corresponding to the target video segment includes:
[0012] Input the two-dimensional human pose information and the image features of multiple resolutions into the adaptive pose pooling module of the human pose estimation model for feature extraction processing to obtain first compensation features;
[0013] Input the two-dimensional human pose information into the second multi-layer perceptron of the human pose estimation model, and the second multi-layer perceptron projects the two-dimensional human pose information into a high-dimensional hidden layer feature space to obtain high-dimensional hidden layer features;
[0014] Input the high-dimensional hidden layer features into the boosting model of the human pose estimation model for edge information compensation processing to obtain second compensation features;
[0015] Input the first compensation features and the second compensation features into the information fusion module of the human pose estimation model for fusion processing to obtain the target fusion features.
[0016] According to an embodiment of the present application, the step of inputting the two-dimensional human pose information and the image features of multiple resolutions into the adaptive pose pooling module of the human pose estimation model for feature extraction processing to obtain first compensation features includes:
[0017] Input the two-dimensional human pose information and the image features of multiple resolutions into the pose-aware offset generation module of the adaptive pose pooling module for sampling processing to generate sampling points;
[0018] The pose-aware sampling module of the adaptive pose pooling module samples the image features of multiple resolutions based on the sampling points to obtain third compensation features;
[0019] The feature update module of the adaptive pose pooling module updates the third compensation features using a preset algorithm to obtain the first compensation features.
[0020] According to an embodiment of the present application, the calculation formula for inputting the two-dimensional human pose information and the image features of multiple resolutions into the pose-aware offset generation module of the adaptive pose pooling module for sampling processing to generate sampling points is:
[0021] ;
[0022] Wherein, is the sampling point; A function for implementing the function of generating pose perception offset in the pose perception offset generation module; For the image features of the resolution; Is the two-dimensional human pose information.
[0023] According to an embodiment of the present application, the formula for calculating the step of obtaining the third compensation feature by sampling the image features of multiple resolutions by the pose perception sampling module of the adaptive pose pooling module based on the sampling points is:
[0024] ;
[0025] Wherein, Is the third compensation feature; Is a function for implementing the function of pose perception sampling in the pose perception sampling module; For the image features of the resolution; Is the sampling point.
[0026] According to an embodiment of the present application, the formula for calculating the step of obtaining the first compensation feature by updating the third compensation feature by the feature update module of the adaptive pose pooling module using a preset algorithm is:
[0027] ;
[0028] Wherein, Is the first compensation feature corresponding to the ; Is a positive real number in [0, 1]; For the third compensation feature corresponding to the personal body pose estimation model; third compensation feature corresponding to the
[0029] According to an embodiment of the present application, the steps of inputting the first compensation feature and the second compensation feature into the information fusion module of the human body pose estimation model for fusion processing to obtain the target fusion feature include:
[0030] The information fusion module calculates the correlation between the first compensation feature and the second compensation feature using the multi-head cross-attention mechanism to obtain the attention weight;
[0031] Fuse the second compensation feature and the attention weight to obtain the target fusion feature.
[0032] In a second aspect of the embodiments of the present application, there is provided an apparatus for estimating human body postures, including:
[0033] An acquisition module, configured to acquire a video sequence;
[0034] A segmentation module, configured to input the video sequence into a video segmentation module for segmentation processing to obtain a plurality of video segments;
[0035] A feature extraction module, configured to input a target video segment into a two-dimensional human body posture estimator for feature extraction processing to obtain two-dimensional human body posture information corresponding to the target video segment and image features of multiple resolutions, where the target video segment is any one of the plurality of video segments;
[0036] A human body posture estimation model processing module, configured to input the two-dimensional human body posture information and the image features of multiple resolutions into a human body posture estimation model for feature extraction and edge information compensation processing to obtain a target fusion feature corresponding to the target video segment;
[0037] A posture estimation module, configured to input the target fusion feature into a first multi-layer perceptron for posture estimation processing to obtain a three-dimensional human body posture estimation result corresponding to the target video segment;
[0038] A merging module, configured to merge all the three-dimensional human body posture estimation results to form a three-dimensional human body posture estimation sequence.
[0039] In a third aspect of the embodiments of the present application, there is provided a terminal device, including: a processor and a memory. The memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the steps of the method for estimating human body postures provided in the first aspect of the embodiments of the present application.
[0040] In a fourth aspect of the embodiments of the present application, there is provided a computer-readable storage medium, and the computer-readable storage medium is used to store a computer program, and the computer program enables a computer to execute the steps of the method for estimating human body postures provided in the first aspect of the embodiments of the present application.
[0041] The beneficial effects of the embodiments of the present application include:
[0042] The human pose estimation model designed in the embodiments of the present application has an edge information compensation function, and can perform feature extraction and edge information compensation processing on a video sequence, so that the three-dimensional human pose estimation sequence corresponding to the video sequence has the coherence (i.e., continuity) of actions. Specifically, in the embodiments of the present application, the video sequence is input into a video segmentation module for segmentation processing to obtain a plurality of video segments; the target video segment is input into a two-dimensional human pose estimator for feature extraction processing to obtain two-dimensional human pose information corresponding to the target video segment and image features of multiple resolutions, where the target video segment is any one of the plurality of video segments; the two-dimensional human pose information and the image features of multiple resolutions are input into the human pose estimation model for feature extraction and edge information compensation processing to obtain a target fusion feature corresponding to the target video segment; the target fusion feature is input into a first multi-layer perceptron for pose estimation processing to obtain a three-dimensional human pose estimation result corresponding to the target video segment; all the three-dimensional human pose estimation results are combined to form a three-dimensional human pose estimation sequence. Thus, compared with the prior art, the embodiments of the present application can solve the problem of insufficient boundary information of video segments and improve the continuity and accuracy of three-dimensional human pose estimation. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application or in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0044] Figure 1 It is a flowchart of the method for estimating the human pose of the present application in some embodiments;
[0045] Figure 2 It is a flowchart of the method for estimating the human pose of the present application in other embodiments;
[0046] Figure 3 It is a schematic flowchart of the feature extraction process by the two-dimensional human pose estimator in the present application;
[0047] Figure 4 It is a schematic block diagram of the principle of the human pose estimation device of the present application in some embodiments;
[0048] Figure 5 It is a schematic block diagram of the principle of the terminal device of the present application in some embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] To make the above objects, features, and advantages of the present application more apparent and understandable, the following provides a detailed description of the specific implementation manners of the present application with reference to the accompanying drawings. Many specific details are set forth in the following description to facilitate a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present application. Therefore, the present application is not limited by the specific embodiments disclosed below.
[0050] It should be noted that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0051] The term "exemplary" or "for example" is used to indicate an example, illustration, or explanation. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of the term "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0052] The term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these process, method, product, or device.
[0053] Unless otherwise defined, all technical and scientific terms used in the specification of the present application have the same meaning as commonly understood by those skilled in the art belonging to the technical field of the present application. The terms used in the specification of the present application are only for the purpose of describing specific implementation manners and are not intended to limit the present application. The term "and / or" used in the specification of the present application includes any and all combinations of one or more of the related listed items.
[0054] Three-dimensional human pose estimation technology has been widely used in fields such as action analysis, virtual reality, and motion capture. Traditional methods for estimating three-dimensional human poses usually divide a video into video segments of a fixed length. Then, all video segments are processed in two stages. Specifically, in the first stage, a two-dimensional human pose estimator is used to extract two-dimensional human poses, and in the second stage, the two-dimensional human poses are mapped into three-dimensional space to form three-dimensional human poses. Although this method can capture local temporal dynamics within a single segment, it ignores the problem of insufficient boundary information of video segments, resulting in the estimation results of three-dimensional human poses being prone to temporal discontinuity and unsmooth actions. For example, when a human action contains complex continuous rotational actions or rapidly changing actions, the individually processed video segments may not be able to correctly predict the pose of the next frame, resulting in the loss of action details or unnatural transition phenomena.
[0055] In addition, the independent processing of video segments increases the dependence on the length and boundaries of video segments. When the video segments are divided unreasonably, the estimation accuracy may be reduced.
[0056] To solve the above problems, this application proposes a method for estimating human poses. The following further describes the specific implementation manners of this application with reference to the accompanying drawings.
[0057] Refer to Figure 1 As shown, it is a flowchart of a method for estimating human poses provided in the first aspect of an embodiment of this application. In Figure 1 it, the method for estimating human poses includes:
[0058] S1. Obtain a video sequence.
[0059] S2. Input the video sequence into a video segmentation module for segmentation processing to obtain multiple video segments.
[0060] The video segmentation module is implemented by a time-based segmentation algorithm or a key-frame-based segmentation algorithm. Specifically, the time-based segmentation algorithm can use Python (Python is a computer programming language) combined with FFmpeg (FFmpeg is an open-source computer program that can be used to record, convert digital audio and video, and convert them into streams) to set the duration of each video segment. Then, the video sequence is intercepted according to the duration in chronological order to obtain the multiple video segments. The key-frame-based segmentation algorithm calculates the optical flow field between adjacent frames in the video sequence using the optical flow method, then detects key frames by analyzing the changes in the optical flow field, and finally segments the video sequence according to the key frames to obtain the multiple video segments.
[0061] In other embodiments, the video segmentation module may be implemented by other algorithms, which can be specifically set by those skilled in the art according to actual needs.
[0062] In the embodiments of the present application, in order to be consistent with MotionAGFormer (Motion Attention-GCNFormer, Action Attention Graph Neural Transformation Network), the video segmentation module divides the video sequence into video segments of 243 frames.
[0063] S3. Input the target video segment into a two-dimensional human pose estimator for feature extraction processing to obtain two-dimensional human pose information corresponding to the target video segment and image features of multiple resolutions, where the target video segment is any one of the multiple video segments.
[0064] The step S3 can refer to Figure 2 the steps of "Image I", "two-dimensional human pose estimator", and "image features of multiple resolutions" in
[0065] The two-dimensional human pose estimator is HRNet (High-Resolution Network). HRNet is a deep neural network architecture for visual tasks such as pose estimation, semantic segmentation, or object detection. In other embodiments, the two-dimensional human pose estimator may also be SHN (Stacked Hourglass Networks), CPN (Cascaded Pyramid Network), or ViTPose (ViTPose is a human pose estimation model based on Vision Transformer), which can be specifically set by those skilled in the art according to actual needs.
[0066] The target video segment can be expressed as , where is the target video segment; is a real-valued tensor of shape (that is, a set of real numbers of dimensions); is the number of frames of the target video segment; is the height of each frame image in the target video segment; is the width of each frame image in the target video segment.
[0067] Further, as shown in Figure 3 , the two-dimensional human pose estimator processes the target video segment (that is, Figure 3The process of performing feature extraction processing on "Image I" in it can be generally divided into four stages (that is, Figure 3 "Stage 1", "Stage 2", "Stage 3" and "Stage 4" in it). After four stages, the two-dimensional human pose estimator will output image features of four resolutions, and the image features of the four resolutions can be referred to Figure 3 in , , , .
[0068] Furthermore, the calculation formula of step S3 is: ; where is the image feature of the th resolution; is the total number of resolutions of the image features; is the two-dimensional human pose information, , is a real-valued tensor with a shape of (that is, a set of real numbers in dimensions), is the number of frames of the target video segment, is the number of human key points; is the two-dimensional human pose estimator; is the target video segment. In , is a real-valued tensor with a shape of (that is, a set of real numbers in dimensions), is the height of each frame image in the target video segment; is the width of each frame image in the target video segment, takes a value of 4, representing 4 different resolution image features.
[0069] S4. Input the two-dimensional human pose information and the image features of multiple resolutions into a human pose estimation model for feature extraction and edge information compensation processing to obtain target fusion features corresponding to the target video segment.
[0070] In the embodiments of the present application, the human pose estimation model includes an adaptive pose pooling module, a second multi-layer perceptron, an enhancement model, and an information fusion module. Each module in the human pose estimation model complements each other to jointly improve the accuracy and robustness of pose estimation.
[0071] Furthermore, step S4 includes:
[0072] S41. Input the two-dimensional human pose information and the image features of multiple resolutions into the adaptive pose pooling module of the human pose estimation model for feature extraction processing to obtain the first compensation feature. The step S41 can refer to Figure 2 the "Adaptive Pose Pooling" step in
[0073] The adaptive pose pooling module includes a pose-aware offset generation module, a pose-aware sampling module, and a feature update module. The adaptive pose pooling module can adaptively optimize feature extraction according to the changes in human poses, enabling the model to exhibit higher detail capture capabilities in complex action scenarios.
[0074] Further, the step S41 includes:
[0075] S411. Input the two-dimensional human pose information and the image features of multiple resolutions into the pose-aware offset generation module of the adaptive pose pooling module for sampling processing to generate sampling points. The step S411 can refer to Figure 2 the "Pose-Aware Offset Generation" and "Sampling Points" steps in
[0076] Further, the calculation formula of the step S411 is:
[0077] ;
[0078] where is the sampling point, , is a real-valued tensor of shape (i.e., a set of real numbers in dimensions), is the number of frames compensating for edge information, , is the much less than symbol in mathematics, is the number of frames of the target video segment, is the number of sampling points, must be divisible by , is the number of human key points; is the function that implements the pose-aware offset generation function in the pose-aware offset generation module; is the th resolution of the image feature; is the total number of resolutions of the image feature; is the two-dimensional human pose information. In the embodiments of the present application, .
[0079] S412. The pose-aware sampling module of the adaptive pose pooling module samples the image features of the multiple resolutions based on the sampling points to obtain a third compensation feature. The step S412 can refer to Figure 2 the "pose-aware sampling" step and the "third compensation feature" step in
[0080] The third compensation feature is the compensated image feature. And, ; is the third compensation feature, is a real-valued tensor of shape (i.e., a set of real numbers of dimensions), is the number of frames for compensating edge information, , is the much less than symbol in mathematics, is the number of frames of the target video segment, is the number of human key points, is the feature dimension.
[0081] Further, the calculation formula for the step S412 is:
[0082] ;
[0083] where, is the third compensation feature; is the function that implements the pose-aware sampling function in the pose-aware sampling module; is the th image feature of the resolution; is the total number of resolutions of the image feature; is the sampling point.
[0084] S413. The feature update module of the adaptive pose pooling module updates the third compensation feature using a preset algorithm to obtain the first compensation feature. The step S413 can refer to Figure 2 the "feature update" step in
[0085] The first compensation feature is the image feature obtained after updating the compensated image feature (i.e., the third compensation feature).
[0086] Further, the calculation formula for the step S413 is:
[0087] ;
[0088] where, is the th first compensation feature corresponding to the human pose estimation model, ; is a positive real number on [0, 1] and is used to control the amplitude of the update of the third compensation feature ; is the third compensation feature corresponding to the th human body pose estimation model;
[0089] S42. Input the two-dimensional human body pose information into the second multi-layer perceptron of the human body pose estimation model. The second multi-layer perceptron projects the two-dimensional human body pose information into a high-dimensional hidden layer feature space to obtain high-dimensional hidden layer features. The high-dimensional hidden layer features can be expressed as .
[0090] The step S42 maps the two-dimensional human body pose information to a three-dimensional space.
[0091] The second multi-layer perceptron (MLP, Multilayer Perceptron) is a feedforward artificial neural network model.
[0092] S43. Input the high-dimensional hidden layer features into the boosting model of the human body pose estimation model for edge information compensation processing to obtain second compensation features. The step S43 can refer to Figure 2 the steps of "boosting model", "edge information compensation" and "second compensation features" in
[0093] The boosting model is MotionAGFormer. The boosting model is a variant of a Transformer model stacked with several layers, and the structure of each layer in the boosting model is the same.
[0094] The MotionAGFormer extracts the spatio-temporal information of the human body pose sequence through the self-attention mechanism, models the dependence relationship of human body key points in the time and space dimensions. Thus, the problem of ignoring the global time dependence in traditional methods can be solved, and the coherence and naturalness of actions can be effectively improved. The second compensation feature is the compensated high-dimensional hidden layer feature. The second compensation feature can be expressed as .
[0095] S44. Input the first compensation feature and the second compensation feature into the information fusion module of the human body pose estimation model for fusion processing to obtain the target fusion feature. The step S44 can refer to Figure 2 the steps of "information fusion" and "target fusion feature" in
[0096] The information fusion module can enhance the ability to understand complex action semantics and dynamic changes in the scene, making the estimation results of 3D poses more coherent and accurate. Moreover, the information fusion module can also improve the ability to capture complex actions and rapidly changing scenes.
[0097] Furthermore, the step S44 includes:
[0098] S441. The information fusion module uses a multi-head cross-attention mechanism to calculate the correlation between the first compensation feature and the second compensation feature, obtaining attention weights.
[0099] The multi-head cross-attention mechanism can deeply interact the spatio-temporal information extracted by MotionAGFormer with image features. The multi-head cross-attention mechanism can dynamically focus on the correlation between image features and pose features, thereby achieving efficient feature fusion.
[0100] Furthermore, the calculation formula for the step S441 is: ; where is the attention weight; is the multi-head cross-attention mechanism; is the second compensation feature; is the first compensation feature.
[0101] S442. Fuse the second compensation feature and the attention weights to obtain the target fusion feature.
[0102] Furthermore, the calculation formula for the step S442 is: ; where is the target fusion feature; is the second compensation feature; is the attention weight.
[0103] S5. Input the target fusion feature into the first multi-layer perceptron for pose estimation processing to obtain the 3D human pose estimation result corresponding to the target video segment. The step S5 can be understood as the "regression head" and "estimation result" steps in the Figure 2 .
[0104] The first multi-layer perceptron (MLP, Multilayer Perceptron) is also a feedforward artificial neural network model.
[0105] S6. Combine all the 3D human pose estimation results to form a 3D human pose estimation sequence.
[0106] The human pose estimation model designed in the embodiments of the present application has an edge information compensation function, and can perform feature extraction and edge information compensation processing on a video sequence, so that the three-dimensional human pose estimation sequence corresponding to the video sequence has the coherence (i.e., continuity) of actions. Thus, compared with the prior art, the embodiments of the present application can solve the problem of insufficient boundary information of video clips, improve the continuity and accuracy of three-dimensional human pose estimation, and enable the embodiments of the present application to generate more natural and coherent three-dimensional pose sequences in complex human action scenarios. The continuity described herein can be understood as the continuity of actions.
[0107] In the embodiments of the present application, the high-precision feature extraction of the two-dimensional human pose estimator ensures the quality of the input data, the adaptive pose pooling module enhances the attention to key parts, the enhancement model solves the problem of discontinuous actions through global spatio-temporal modeling, and the information fusion module further improves the adaptability to dynamic actions. The embodiments of the present application can adapt to different scenario requirements and can be applied to fields such as motion capture, virtual reality, and action analysis.
[0108] In some embodiments, multiple human pose estimation models can be set to obtain the final hidden layer features, and the final hidden layer features are used as the target fusion features. Then, the target fusion features are output through a first multi-layer perceptron to obtain the three-dimensional human pose estimation result predicted by the model.
[0109] Exemplarily, the number of the multiple human pose estimation models is N. The output of the first human pose estimation model is used as the input of the enhancement model of the second human pose estimation model, the output of the second human pose estimation model is used as the input of the enhancement model of the third human pose estimation model, and so on. The output of the (N - 1)th human pose estimation model is used as the input of the enhancement model of the Nth human pose estimation model. Finally, the output of the Nth human pose estimation model is used as the final hidden layer feature.
[0110] When the number of human pose estimation models is N, the calculation formula of step S413 ; in The value of is .
[0111] In some embodiments, the method for estimating the human pose further includes: the adaptive pose pooling module dynamically adjusts the pooling range and pooling method based on the two-dimensional human pose information. Thus, detailed information related to the key parts of the human body can be extracted from the image features.
[0112] Further, the steps of the adaptive pose pooling module dynamically adjusting the pooling range and pooling method based on the two-dimensional human pose information include:
[0113] S7. The adaptive pose pooling module analyzes the human pose features of the two-dimensional human pose information.
[0114] The human pose features include the spatial distribution of human key points in the two-dimensional human pose information, the constraint relationship between human key points, the sensitivity of human key points to resolution, and the sensitivity of human key points to the overall two-dimensional human pose information. In other embodiments, the human pose features may also be other information, which can be specifically set by those skilled in the art according to actual needs.
[0115] S8. Adjust the pooling range and the pooling method based on the human pose features.
[0116] In some embodiments, the method for estimating the human pose further includes the training steps of the human pose estimation model. Specifically, the training steps of the human pose estimation model include:
[0117] S91. Obtain a training data set and a true label set, where each training data in the training data set has a one-to-one correspondence with one of the true labels in the true label set.
[0118] Each training data is a video sequence for training. Each true label in the true label set is a true three-dimensional human pose estimation sequence for training.
[0119] S92. Input the target training data into the original video segmentation module for segmentation processing to obtain a plurality of training segments, where the target training data is any training data in the training data set.
[0120] The difference between the original video segmentation module and the video segmentation module lies only in the different model parameters.
[0121] S93. Input the target training segment into the original two-dimensional human pose estimator for feature extraction processing to obtain two-dimensional human pose training information corresponding to the target training segment and image training features of multiple resolutions, where the target training segment is any training segment in the plurality of training segments. The difference between the original two-dimensional human pose estimator and the two-dimensional human pose estimator lies only in the different model parameters.
[0122] S94. Input the two-dimensional human pose training information and the image training features of multiple resolutions into the original human pose estimation model for feature extraction and edge information compensation processing to obtain target fusion training features, where the target training data is any training data in the training dataset.
[0123] The original human pose estimation model includes an original adaptive pose pooling module, an original second multi-layer perceptron, an original boosting model, and an original information fusion module. Among them, the difference between the original adaptive pose pooling module and the adaptive pose pooling module lies only in the model parameters, the difference between the original second multi-layer perceptron and the second multi-layer perceptron lies only in the model parameters, the difference between the original boosting model and the boosting model lies only in the model parameters, and the difference between the original information fusion module and the information fusion module lies only in the model parameters.
[0124] Further, the S94 step includes:
[0125] S941. Input the two-dimensional human pose training information and the image training features of multiple resolutions into the original adaptive pose pooling module for feature extraction processing to obtain first compensated training features.
[0126] S942. Input the two-dimensional human pose training information into the original second multi-layer perceptron, and the original second multi-layer perceptron projects the two-dimensional human pose training information into a high-dimensional hidden layer feature space to obtain high-dimensional training hidden layer features.
[0127] S943. Input the high-dimensional training hidden layer features into the original boosting model for edge information compensation processing to obtain second compensated training features.
[0128] S944. Input the first compensated training features and the second compensated training features into the original information fusion module for fusion processing to obtain the target fusion training features.
[0129] S95. Input the target fusion training features into the original first multi-layer perceptron for pose estimation processing to obtain a training result corresponding to the target training segment.
[0130] The difference between the original first multi-layer perceptron and the first multi-layer perceptron lies only in the model parameters.
[0131] S96. Combine all the training results to form a prediction result. The prediction result is a three-dimensional human pose estimation sequence.
[0132] S97. Calculate a loss function based on the prediction result and the true label corresponding to the target training data.
[0133] Further, the calculation formula of the loss function is as follows:
[0134] ;
[0135] wherein, is the loss function to calculate the norm between the prediction result and the true label; is the length of the training data set, that is, the total number of all training data; is the true label corresponding to the target training data, that is, the true label corresponding to the th training data in the training data set; is the prediction result, that is, the prediction result corresponding to the target training data.
[0136] S98. Adjust the model parameters of the original human body pose estimation model based on the loss function to obtain the human body pose estimation model.
[0137] S99. Adjust the parameters of the original video segmentation module based on the loss function to obtain the video segmentation module; adjust the parameters of the original two-dimensional human body pose estimator based on the loss function to obtain the two-dimensional human body pose estimator; adjust the parameters of the original first multi-layer perceptron based on the loss function to obtain the first multi-layer perceptron. Refer to Figure 4 shown, which is a schematic block diagram of an apparatus for estimating human body pose provided in the second aspect of the embodiments of the present application. In Figure 4 , the apparatus 100 for estimating human body pose includes:
[0138] An acquisition module 101, configured to acquire a video sequence;
[0139] A segmentation module 102, configured to input the video sequence into the video segmentation module for segmentation processing to obtain a plurality of video segments;
[0140] A feature extraction module 103, configured to input the target video segment into the two-dimensional human body pose estimator for feature extraction processing to obtain two-dimensional human body pose information corresponding to the target video segment and image features of multiple resolutions, wherein the target video segment is any one of the plurality of video segments;
[0141] A human body pose estimation model processing module 104, configured to input the two-dimensional human body pose information and the image features of multiple resolutions into the human body pose estimation model for feature extraction and edge information compensation processing to obtain target fusion features corresponding to the target video segment;
[0142] The pose estimation module 105 is configured to input the target fusion feature into a first multi-layer perceptron for pose estimation processing, so as to obtain a three-dimensional human pose estimation result corresponding to the target video segment.
[0143] The merging module 106 is configured to merge all the three-dimensional human pose estimation results to form a three-dimensional human pose estimation sequence.
[0144] In a third aspect of the embodiments of the present application, a terminal device is provided. The principle block diagram of the terminal device can be as Figure 5 shown. The terminal device includes a processor, a memory, a network interface, a display screen, and a temperature sensor connected through a system bus. Among them, the processor is used to provide computing and control capabilities. The memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the terminal device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for estimating human poses is implemented. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the temperature sensor is pre-set inside the terminal device to detect the operating temperature of the internal device.
[0145] Those skilled in the art can understand that Figure 5 the principle block diagram shown in
[0146] is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the terminal device to which the solution of the present invention is applied. The specific terminal device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0147] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0148] Without changing the basic principles of the present application, the technical features of the above embodiments can be combined. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0149] The above embodiments only represent several implementation manners of the present application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the patent protection scope of the present application should be subject to the appended claims.
Claims
1. A method for estimating human body postures, characterized in that, Including: Obtain a video sequence; Input the video sequence into a video segmentation module for segmentation processing to obtain multiple video clips; Input a target video clip into a two-dimensional human pose estimator for feature extraction processing to obtain two-dimensional human pose information corresponding to the target video clip and image features of multiple resolutions, where the target video clip is any one of the multiple video clips; Input the two-dimensional human pose information and the image features of multiple resolutions into a human pose estimation model for feature extraction and edge information compensation processing to obtain a target fusion feature corresponding to the target video clip; Input the target fusion feature into a first multi-layer perceptron for pose estimation processing to obtain a three-dimensional human pose estimation result corresponding to the target video clip; Merge all the three-dimensional human pose estimation results to form a three-dimensional human pose estimation sequence; The step of inputting the two-dimensional human pose information and the image features of multiple resolutions into a human pose estimation model for feature extraction and edge information compensation processing to obtain a target fusion feature corresponding to the target video clip includes: inputting the two-dimensional human pose information and the image features of multiple resolutions into an adaptive pose pooling module of the human pose estimation model for feature extraction processing to obtain a first compensation feature; inputting the two-dimensional human pose information into a second multi-layer perceptron of the human pose estimation model, and the second multi-layer perceptron projects the two-dimensional human pose information into a high-dimensional hidden layer feature space to obtain high-dimensional hidden layer features; inputting the high-dimensional hidden layer features into an enhancement model of the human pose estimation model for edge information compensation processing to obtain a second compensation feature; inputting the first compensation feature and the second compensation feature into an information fusion module of the human pose estimation model for fusion processing to obtain the target fusion feature.
2. The method for estimating human body postures according to claim 1, wherein, The step of inputting the two-dimensional human pose information and the image features of multiple resolutions into an adaptive pose pooling module of the human pose estimation model for feature extraction processing to obtain a first compensation feature includes: Input the two-dimensional human pose information and the image features of multiple resolutions into a pose-aware offset generation module of the adaptive pose pooling module for sampling processing to generate sampling points; A pose-aware sampling module of the adaptive pose pooling module samples the image features of multiple resolutions based on the sampling points to obtain a third compensation feature; A feature update module of the adaptive pose pooling module updates the third compensation feature using a preset algorithm to obtain the first compensation feature.
3. The method for estimating a human body posture according to claim 2, wherein The calculation formula for inputting the two-dimensional human pose information and the image features of multiple resolutions into a pose-aware offset generation module of the adaptive pose pooling module for sampling processing to generate sampling points is: ; Among them, is the sampling point; is a function that realizes the function of generating pose perception offset in the pose perception offset generation module; is the image feature of the total number of resolutions of the image feature; is the two-dimensional human body pose information.
4. The method for estimating a human body posture according to claim 2, wherein The calculation formula for the step of a pose-aware sampling module of the adaptive pose pooling module sampling the image features of multiple resolutions based on the sampling points to obtain a third compensation feature is: ; Among them, is the third compensation feature; is a function that implements the attitude perception sampling function in the attitude perception sampling module; is the image feature at the total number of resolutions of the image feature; is the sampling point.
5. The method for estimating a human body posture according to claim 2, wherein The calculation formula for the step of the feature update module of the adaptive pose pooling module to update the third compensation feature using a preset algorithm to obtain the first compensation feature is as follows: ; Among them, is the first compensation feature corresponding to the ; is a positive real number in [0, 1]; is the third compensation feature corresponding to the personal human pose estimation model; is the third compensation feature corresponding to the personal human pose estimation model.
6. The method for estimating a human body posture according to claim 1, wherein The steps of inputting the first compensation feature and the second compensation feature into the information fusion module of the human pose estimation model for fusion processing to obtain the target fusion feature include: The information fusion module uses a multi-head cross-attention mechanism to calculate the correlation between the first compensation feature and the second compensation feature to obtain attention weights; The second compensation feature and the attention weights are fused to obtain the target fusion feature.
7. An apparatus for estimating a human body posture, characterized in that, It includes: An acquisition module for acquiring a video sequence; A segmentation module for inputting the video sequence into a video segmentation module for segmentation processing to obtain a plurality of video segments; A feature extraction module for inputting a target video segment into a two-dimensional human pose estimator for feature extraction processing to obtain two-dimensional human pose information corresponding to the target video segment and image features of multiple resolutions, where the target video segment is any one of the plurality of video segments; A human pose estimation model processing module for inputting the two-dimensional human pose information and the image features of multiple resolutions into a human pose estimation model for feature extraction and edge information compensation processing to obtain a target fusion feature corresponding to the target video segment; A pose estimation module for inputting the target fusion feature into a first multi-layer perceptron for pose estimation processing to obtain a three-dimensional human pose estimation result corresponding to the target video segment; A merging module for merging all the three-dimensional human pose estimation results to form a three-dimensional human pose estimation sequence; The human pose estimation model processing module is further configured to input the two-dimensional human pose information and the image features of multiple resolutions into the adaptive pose pooling module of the human pose estimation model for feature extraction processing to obtain a first compensation feature; input the two-dimensional human pose information into the second multi-layer perceptron of the human pose estimation model, and the second multi-layer perceptron projects the two-dimensional human pose information into a high-dimensional hidden layer feature space to obtain high-dimensional hidden layer features; input the high-dimensional hidden layer features into the enhancement model of the human pose estimation model for edge information compensation processing to obtain a second compensation feature; input the first compensation feature and the second compensation feature into the information fusion module of the human pose estimation model for fusion processing to obtain the target fusion feature.
8. A terminal device, characterized in that, It includes: A processor and a memory, where the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the steps of the method for estimating human pose according to any one of claims 1 to 6 above.
9. A computer-readable storage medium, characterized in that, For storing a computer program, the computer program causes a computer to execute the steps of the method for estimating human pose according to any one of claims 1 to 6 above.
Citation Information
Patent Citations
Road construction cone barrel falling detection method based on attitude estimation
CN119723518A
Method, System and Device for Direct Prediction of 3D Body Poses from Motion Compensated Sequence
US20170316578A1