Human body posture estimation method and device, equipment and storage medium

By segmenting and feature extraction of video sequences, combining adaptive pose pooling and information fusion modules, the time discontinuity problem of three-dimensional human pose estimation in traditional methods is solved, and a more natural and coherent pose estimation is achieved.

CN120108005AActive Publication Date: 2025-06-06PEKING UNIV SHENZHEN GRADUATE SCHOOL

Patent Information

Application Number
CN202510601980.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-06-06
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

The traditional three-dimensional human posture estimation method is insufficient in the processing of video clip boundaries, resulting in discontinuous time and incoherent movements, especially when complex movements or rapid changes, the action details are lost or the transition is unnatural.

Method used

The video segmentation module is used to split the video sequence into multiple segments, combining a two-dimensional human pose estimator and human pose estimation model for feature extraction and edge information compensation, and using an adaptive pose pooling module, multi-layer perceptron and information fusion module to improve the continuity and accuracy of feature extraction and pose estimation.

Benefits of technology

It improves the continuity and accuracy of three-dimensional human posture estimation, can generate naturally coherent posture sequences in complex action scenes, and enhances the coherence and detail capture ability of movement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108005A_ABST
    Figure CN120108005A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image recognition. The invention discloses a human body posture estimation method and device, equipment and a storage medium. The continuity and accuracy of three-dimensional human body posture estimation can be improved. The method comprises the following steps: acquiring a video sequence; inputting the video sequence into a video segmentation module for segmentation processing to obtain a plurality of video clips; the target video clip is input into a two-dimensional human body posture estimator for feature extraction processing, two-dimensional human body posture information and image features of multiple resolutions are obtained, and the target video clip is any one of the multiple video clips; the two-dimensional human body posture information and the image features are input into a human body posture estimation model for compensation processing, and target fusion features are obtained; inputting the target fusion feature into a first multi-layer perceptron for attitude estimation processing to obtain a three-dimensional human body attitude estimation result corresponding to the target video clip; and combining all three-dimensional human body posture estimation results to form a three-dimensional human body posture estimation sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image recognition technology, and more specifically, to a method, device, equipment and storage medium for estimating human posture. Background Art

[0002] 3D human pose estimation technology has been widely used in motion analysis, virtual reality, motion capture and other fields. Traditional 3D human pose estimation methods usually divide the video into video clips of fixed length, and then process all video clips in two stages. Specifically, the first stage is to use a 2D human pose estimator to extract the 2D human pose, and the second stage is to map the 2D human pose to the 3D space to form a 3D human pose. Although this method can capture local temporal dynamics in a single clip, it ignores the problem of insufficient boundary information of the video clip, resulting in the 3D human pose estimation results being prone to time discontinuity and motion incoherence. Summary of the invention

[0003] The purpose of the embodiments of the present application is to provide a method, device, equipment and storage medium for estimating human body posture, which can improve the continuity and accuracy of three-dimensional human body posture estimation. The embodiments of the present application are mainly implemented through the following technical solutions: According to a first aspect of an embodiment of the present application, a method for estimating a human body posture is provided, comprising: Get video sequence; Inputting the video sequence into a video segmentation module for segmentation processing to obtain multiple video clips; Inputting a target video segment into a two-dimensional human posture estimator for feature extraction processing to obtain two-dimensional human posture information and image features of multiple resolutions corresponding to the target video segment, wherein the target video segment is any one of the multiple video segments; Inputting the two-dimensional human body posture information and the image features of the multiple resolutions into a human body posture estimation model for feature extraction and edge information compensation processing to obtain a target fusion feature corresponding to the target video clip; Inputting the target fusion feature into a first multi-layer perceptron for posture estimation processing to obtain a three-dimensional human posture estimation result corresponding to the target video clip; All 3D human pose estimation results are combined to form a 3D human pose estimation sequence.

[0004] According to one embodiment of the present application, the step of inputting the two-dimensional human body posture information and the image features of the multiple resolutions into a human body posture estimation model for feature extraction and edge information compensation processing to obtain a target fusion feature corresponding to the target video clip includes: Inputting the two-dimensional human body posture information and the image features of the multiple resolutions into the adaptive posture pooling module of the human body posture estimation model for feature extraction processing to obtain a first compensation feature; Inputting the two-dimensional human body posture information into a second multi-layer perceptron of the human body posture estimation model, wherein the second multi-layer perceptron projects the two-dimensional human body posture information into a high-dimensional hidden layer feature space to obtain high-dimensional hidden layer features; Inputting the high-dimensional hidden layer features into the lifting model of the human posture estimation model to perform edge information compensation processing to obtain a second compensation feature; The first compensation feature and the second compensation feature are input into the information fusion module of the human posture estimation model for fusion processing to obtain the target fusion feature.

[0005] According to one embodiment of the present application, the two-dimensional human body posture information and the image features of the multiple resolutions are input into the adaptive posture pooling module of the human body posture estimation model for feature extraction processing, and the step of obtaining the first compensation feature includes: Inputting the two-dimensional human body posture information and the image features of the multiple resolutions into the posture perception offset generation module of the adaptive posture pooling module for sampling processing to generate sampling points; The gesture perception sampling module of the adaptive gesture pooling module performs sampling processing on the image features of the multiple resolutions based on the sampling points to obtain a third compensation feature; The feature updating module of the adaptive posture pooling module uses a preset algorithm to update the third compensation feature to obtain the first compensation feature.

[0006] According to one embodiment of the present application, the two-dimensional human body posture information and the image features of the multiple resolutions are input into the posture perception offset generation module of the adaptive posture pooling module for sampling processing, and the calculation formula for generating the sampling points is: ; in, is the sampling point; A function for realizing the gesture perception offset generation function in the gesture perception offset generation module; For the Image features with resolutions of 100; is the total number of resolutions of the image features; is the two-dimensional human body posture information.

[0007] According to one embodiment of the present application, the gesture perception sampling module of the adaptive gesture pooling module performs sampling processing on the image features of the multiple resolutions based on the sampling points, and the calculation formula of the step of obtaining the third compensation feature is: ; in, is the third compensation feature; A function that implements the gesture sensing sampling function in the gesture sensing sampling module; For the Image features with resolutions of 100; is the total number of resolutions of the image features; is the sampling point.

[0008] According to one embodiment of the present application, the feature updating module of the adaptive posture pooling module uses a preset algorithm to update the third compensation feature, and the calculation formula for the step of obtaining the first compensation feature is: ; in, For the the first compensation feature corresponding to the human body posture estimation model, ; is a positive real number in [0, 1]; For the The third compensation feature corresponding to the personal human posture estimation model; For the The third compensation feature corresponding to the personal body posture estimation model.

[0009] According to an embodiment of the present application, the first compensation feature and the second compensation feature are input into the information fusion module of the human posture estimation model for fusion processing, and the step of obtaining the target fusion feature includes: The information fusion module uses a multi-head cross attention mechanism to calculate the correlation between the first compensation feature and the second compensation feature to obtain an attention weight; The second compensation feature and the attention weight are fused to obtain the target fusion feature.

[0010] According to a second aspect of an embodiment of the present application, a human body posture estimation device is provided, comprising: An acquisition module, used for acquiring a video sequence; A segmentation module, used for inputting the video sequence into a video segmentation module for segmentation processing to obtain multiple video segments; a feature extraction module, configured to input a target video segment into a two-dimensional human posture estimator for feature extraction processing, and obtain two-dimensional human posture information and image features of multiple resolutions corresponding to the target video segment, wherein the target video segment is any one of the multiple video segments; A human posture estimation model processing module, used for inputting the two-dimensional human posture information and the image features of the multiple resolutions into a human posture estimation model for feature extraction and edge information compensation processing, so as to obtain a target fusion feature corresponding to the target video segment; A posture estimation module, used for inputting the target fusion feature into a first multi-layer perceptron for posture estimation processing to obtain a three-dimensional human posture estimation result corresponding to the target video clip; The merging module is used to merge all the 3D human posture estimation results to form a 3D human posture estimation sequence.

[0011] The third aspect of the embodiments of the present application provides a terminal device, including: a processor and a memory, the memory is used to store a computer program, the processor is used to call and run the computer program stored in the memory, and execute the steps of the human body posture estimation method provided in the first aspect of the embodiments of the present application.

[0012] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium is used to store a computer program, wherein the computer program enables a computer to execute the steps of the human body posture estimation method provided in the first aspect of the embodiment of the present application.

[0013] The beneficial effects of the embodiments of the present application include: The human posture estimation model designed in the embodiment of the present application has the function of edge information compensation, and can perform feature extraction and edge information compensation processing on the video sequence, so that the three-dimensional human posture estimation sequence corresponding to the video sequence has the continuity of action (that is, continuity). Specifically, the embodiment of the present application obtains multiple video clips by segmenting the video sequence input into the video segmentation module; inputs the target video clip into the two-dimensional human posture estimator for feature extraction processing, and obtains the two-dimensional human posture information and image features of multiple resolutions corresponding to the target video clip, wherein the target video clip is any one of the multiple video clips; inputs the two-dimensional human posture information and the image features of multiple resolutions into the human posture estimation model for feature extraction and edge information compensation processing, and obtains the target fusion feature corresponding to the target video clip; inputs the target fusion feature into the first multi-layer perceptron for posture estimation processing, and obtains the three-dimensional human posture estimation result corresponding to the target video clip; and merges all the three-dimensional human posture estimation results to form a three-dimensional human posture estimation sequence. Therefore, compared with the prior art, the embodiment of the present application can solve the problem of insufficient boundary information of video clips and improve the continuity and accuracy of three-dimensional human posture estimation. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the conventional technology, the drawings required for use in the embodiments or the conventional technology descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0015] Figure 1 A flowchart of a method for estimating a human body posture of the present application in some embodiments; Figure 2 A flowchart of a method for estimating a human body posture in some other embodiments of the present application; Figure 3 A schematic diagram of the process of feature extraction processing for a two-dimensional human posture estimator in this application; Figure 4 A principle block diagram of a human body posture estimation device of the present application in some embodiments; Figure 5 This is a functional block diagram of the terminal device of the present application in some embodiments. DETAILED DESCRIPTION

[0016] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are described in detail below in conjunction with the accompanying drawings. In the following description, many specific details are set forth to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without violating the connotation of the present application, so the present application is not limited by the specific embodiments disclosed below.

[0017] It should be noted that the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0018] The terms "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.

[0019] The terms "comprises," "comprising," or any other variations thereof are intended to cover a non-exclusive inclusion, for example, a process, method, system, product, or apparatus that comprises a series of steps or elements is not necessarily limited to those steps or elements explicitly listed, but may include other steps or elements not explicitly listed or inherent to such process, method, product, or apparatus.

[0020] Unless otherwise defined, all technical and scientific terms used in the specification of this application have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The term "and / or" used in the specification of this application includes any and all combinations of one or more related listed items.

[0021] 3D human pose estimation technology has been widely used in motion analysis, virtual reality, motion capture and other fields. Traditional 3D human pose estimation methods usually divide videos into video clips of fixed length, and then process all video clips in two stages. Specifically, the first stage is to extract 2D human pose using a 2D human pose estimator, and the second stage is to map the 2D human pose into 3D space to form 3D human pose. Although this method can capture local temporal dynamics in a single clip, it ignores the problem of insufficient boundary information of the video clip, resulting in the 3D human pose estimation results being prone to time discontinuity and motion incoherence. For example, when human motion contains complex continuous rotation or rapidly changing motion, the video clip processed separately may not be able to correctly predict the pose of the next frame, resulting in loss of motion details or unnatural transition.

[0022] Moreover, the independent processing of video segments increases the dependence on the length and boundaries of the video segments, which may reduce the estimation accuracy when the video segments are not divided reasonably.

[0023] In order to solve the above problems, the present application proposes a method for estimating human body posture. The specific implementation of the present application is further described below in conjunction with the accompanying drawings.

[0024] refer to Figure 1 FIG. 1 is a flowchart of a method for estimating a human posture provided by the first aspect of the embodiment of the present application. Figure 1 In the method, the human body posture estimation method includes: S1. Obtain a video sequence.

[0025] S2. Input the video sequence into a video segmentation module for segmentation processing to obtain multiple video clips.

[0026] The video segmentation module is implemented by a time-based segmentation algorithm or a key-frame-based segmentation algorithm. Specifically, the time-based segmentation algorithm can be implemented by Python (Python is a computer programming language) combined with FFmpeg (FFmpeg is an open source computer program that can be used to record, convert digital audio and video, and convert them into streams) to set the duration of each video segment, and then intercept and process the video sequence in chronological order according to the duration, so as to obtain the multiple video segments; the key-frame-based segmentation algorithm uses the optical flow method to calculate the optical flow field between adjacent frames in the video sequence, and then detects the key frames by analyzing the changes in the optical flow field. Finally, the video sequence is segmented according to the key frames to obtain the multiple video segments.

[0027] In other implementations, the video segmentation module may be implemented by other algorithms, which may be specifically configured by those skilled in the art according to actual requirements.

[0028] In the embodiment of the present application, in order to be consistent with MotionAGFormer (Motion Attention-GCNFormer, action attention graph neural transformation network), the video segmentation module divides the video sequence into video segments of 243 frames.

[0029] S3. Input the target video segment into a two-dimensional human posture estimator for feature extraction processing to obtain two-dimensional human posture information corresponding to the target video segment and image features of multiple resolutions, wherein the target video segment is any one of the multiple video segments.

[0030] The S3 step can refer to Figure 2 The "Image I", "2D Human Pose Estimator", and "Image Features at Multiple Resolutions" steps in .

[0031] The two-dimensional human posture estimator is HRNet (High-Resolution Network). HRNet is a deep neural network architecture for visual tasks (such as posture estimation, semantic segmentation or target detection, etc.). In other embodiments, the two-dimensional human posture estimator can also be SHN (Stacked Hourglass Networks), CPN (Cascaded Pyramid Network) or ViTPose (ViTPose is a human posture estimation model based on visual Transformer), which can be specifically set by technicians in this field according to actual needs.

[0032] The target video segment can be expressed as ,in, is the target video clip; The shape is A real-valued tensor of -dimensional real number set); is the number of frames of the target video segment; is the height of each frame of the target video clip; is the width of each frame image in the target video segment.

[0033] Further, refer to Figure 3 As shown, the two-dimensional human posture estimator estimates the target video segment (i.e. Figure 3 The process of feature extraction can be divided into four stages (i.e. Figure 3 After four stages, the two-dimensional human posture estimator will output image features of four resolutions. The image features of the four resolutions can refer to Figure 3 In , , , .

[0034] Furthermore, the calculation formula of step S3 is: ;in, It is Image features with resolutions of 100; is the total number of resolutions of the image features; is the two-dimensional human body posture information, , The shape is A real-valued tensor of dimensional set of real numbers), is the number of frames of the target video segment, is the number of key points of the human body; is the two-dimensional human posture estimator; is the target video segment. middle, The shape is A real-valued tensor of dimensional set of real numbers), is the height of each frame of the target video clip; is the width of each frame of the target video segment, The value is 4, which means there are 4 image features with different resolutions.

[0035] S4. Input the two-dimensional human body posture information and the image features of the multiple resolutions into a human body posture estimation model to perform feature extraction and edge information compensation processing to obtain a target fusion feature corresponding to the target video clip.

[0036] In an embodiment of the present application, the human body posture estimation model includes an adaptive posture pooling module, a second multi-layer perceptron, a lifting model and an information fusion module. The modules in the human body posture estimation model complement each other and jointly improve the accuracy and robustness of posture estimation.

[0037] Furthermore, the S4 step includes: S41, inputting the two-dimensional human posture information and the image features of the multiple resolutions into the adaptive posture pooling module of the human posture estimation model for feature extraction processing to obtain a first compensation feature. The step S41 can be referred to Figure 2 The "Adaptive Pose Pooling" step in .

[0038] The adaptive posture pooling module includes a posture perception offset generation module, a posture perception sampling module and a feature update module. The adaptive posture pooling module can adaptively optimize feature extraction according to changes in human posture, so that the model can show higher detail capture ability in complex action scenes.

[0039] Furthermore, the step S41 includes: S411, input the two-dimensional human body posture information and the image features of multiple resolutions into the posture perception offset generation module of the adaptive posture pooling module for sampling processing to generate sampling points. The step S411 can be referred to Figure 2 The "Posture-aware offset generation" and "Sampling points" steps in the .

[0040] Furthermore, the calculation formula of step S411 is: ; in, is the sampling point, , The shape is A real-valued tensor of dimensional set of real numbers), is the number of frames to compensate for edge information, , It is the much less than symbol in mathematics. is the number of frames of the target video segment, is the number of sampling points, Must be divisible, is the number of key points of the human body; is a function in the posture perception offset generation module that implements the posture perception offset generation function; It is Image features with resolutions of 100; is the total number of resolutions of the image features; is the two-dimensional human body posture information. In the embodiment of the present application, .

[0041] S412: The gesture perception sampling module of the adaptive gesture pooling module performs sampling processing on the image features of the multiple resolutions based on the sampling points to obtain a third compensation feature. The step S412 can be referred to as Figure 2 The "posture perception sampling" step and the "third compensation feature" step in .

[0042] The third compensation feature is a compensated image feature. And, ; is the third compensating characteristic, The shape is A real-valued tensor of dimensional set of real numbers), To compensate for the number of frames of edge information, , is the much less than symbol in mathematics, is the number of frames of the target video segment, is the number of key points of the human body, is the feature dimension.

[0043] Furthermore, the calculation formula of step S412 is: ; in, is the third compensation feature; A function that implements the gesture sensing sampling function in the gesture sensing sampling module; For the Image features with resolutions of 100; is the total number of resolutions of the image features; is the sampling point.

[0044] S413: The feature updating module of the adaptive posture pooling module updates the third compensation feature using a preset algorithm to obtain the first compensation feature. Figure 2 The Feature Update step in .

[0045] The first compensation feature is an image feature obtained after the compensated image feature (ie, the third compensation feature) is updated.

[0046] Furthermore, the calculation formula of step S413 is: ; in, For the the first compensation feature corresponding to the human body posture estimation model, ; is a positive real number in [0, 1], used to control the third compensation characteristic the magnitude of the update; For the The third compensation feature corresponding to the personal human posture estimation model; For the The third compensation feature corresponding to the personal body posture estimation model.

[0047] S42, inputting the two-dimensional human posture information into the second multi-layer perceptron of the human posture estimation model, and the second multi-layer perceptron projects the two-dimensional human posture information into a high-dimensional hidden layer feature space to obtain high-dimensional hidden layer features. The high-dimensional hidden layer features can be expressed as .

[0048] The step S42 is to map the two-dimensional human body posture information to a three-dimensional space.

[0049] The second multilayer perceptron (MLP) is a feed-forward artificial neural network model.

[0050] S43, inputting the high-dimensional hidden layer features into the lifting model of the human posture estimation model to perform edge information compensation processing to obtain a second compensation feature. The step S43 can be referred to as Figure 2 The "enhanced model", "edge information compensation" and "second compensation feature" steps in .

[0051] The lifting model is MotionAGFormer. The lifting model is a deformation of the Transformer model with several layers stacked, and the structure of each layer in the lifting model is the same.

[0052] The MotionAGFormer extracts the spatiotemporal information of the human posture sequence through the self-attention mechanism and models the dependencies between the key points of the human body in the temporal and spatial dimensions. This can solve the problem of ignoring the global temporal dependency in traditional methods, thereby effectively improving the coherence and naturalness of the action. The second compensation feature is the compensated high-dimensional hidden layer feature. The second compensation feature can be expressed as .

[0053] S44, inputting the first compensation feature and the second compensation feature into the information fusion module of the human posture estimation model for fusion processing to obtain the target fusion feature. Figure 2 The “information fusion” and “target fusion feature” steps in .

[0054] The information fusion module can enhance the ability to understand the semantics of complex actions and dynamic changes in scenes, making the estimation results of three-dimensional posture more coherent and accurate. In addition, the information fusion module can also improve the ability to capture complex actions and rapidly changing scenes.

[0055] Furthermore, the step S44 includes: S441. The information fusion module uses a multi-head cross-attention mechanism to calculate the correlation between the first compensation feature and the second compensation feature to obtain an attention weight.

[0056] The multi-head cross attention mechanism can deeply interact the spatiotemporal information extracted by MotionAGFormer with image features. The multi-head cross attention mechanism can dynamically focus on the correlation between image features and posture features, thereby achieving efficient feature fusion.

[0057] Furthermore, the calculation formula of step S441 is: ;in, is the attention weight; The multi-head cross attention mechanism; is the second compensation feature; is the first compensation feature.

[0058] S442: Fusing the second compensation feature and the attention weight to obtain the target fusion feature.

[0059] Furthermore, the calculation formula of step S442 is: ;in, fusing features for the target; is the second compensation feature; is the attention weight.

[0060] S5, inputting the target fusion feature into the first multi-layer perceptron for posture estimation processing, and obtaining a three-dimensional human posture estimation result corresponding to the target video clip. The step S5 can be understood as Figure 2 The "Regression Head" and "Estimate Results" steps in .

[0061] The first multilayer perceptron (MLP) is also a feedforward artificial neural network model.

[0062] S6. Combining all the 3D human posture estimation results to form a 3D human posture estimation sequence.

[0063] The human posture estimation model designed in the embodiment of the present application has an edge information compensation function, and can perform feature extraction and edge information compensation processing on the video sequence, so that the three-dimensional human posture estimation sequence corresponding to the video sequence has the coherence of the action (that is, continuity). Therefore, compared with the prior art, the embodiment of the present application can solve the problem of insufficient boundary information of the video clip, improve the continuity and accuracy of the three-dimensional human posture estimation, and enable the embodiment of the present application to generate a more natural and coherent three-dimensional posture sequence in complex human action scenes. The continuity described in this article can be understood as the continuity of the action.

[0064] In the embodiment of the present application, the high-precision feature extraction of the two-dimensional human posture estimator ensures the quality of the input data, the adaptive posture pooling module enhances the focus on key parts, the lifting model solves the problem of incoherent movements through global spatiotemporal modeling, and the information fusion module further improves the adaptability to dynamic movements. The embodiment of the present application can adapt to different scene requirements and can be applied to fields such as motion capture, virtual reality, and motion analysis.

[0065] In some embodiments, multiple human body posture estimation models can be set to obtain the final hidden layer features, and the final hidden layer features are used as the target fusion features. Then, the target fusion features are passed through a first multi-layer perceptron output model to predict the three-dimensional human body posture estimation results.

[0066] Exemplarily, the number of the multiple human posture estimation models is N. The output of the first human posture estimation model is used as the input of the lifting model of the second human posture estimation model, the output of the second human posture estimation model is used as the input of the lifting model of the third human posture estimation model, and so on, the output of the N-1th human posture estimation model is used as the input of the lifting model of the Nth human posture estimation model, and finally, the output of the Nth human posture estimation model is used as the final hidden layer feature.

[0067] When the number of the human body posture estimation models is N, the calculation formula of step S413 is ;middle The value of is .

[0068] In some embodiments, the human body posture estimation method also includes: the adaptive posture pooling module dynamically adjusts the pooling range and pooling method based on the two-dimensional human body posture information, thereby extracting detailed information related to key parts of the human body from the image features.

[0069] Furthermore, the step of the adaptive posture pooling module dynamically adjusting the pooling range and the pooling mode based on the two-dimensional human posture information includes: S7. The adaptive posture pooling module analyzes the human posture features of the two-dimensional human posture information.

[0070] The human body posture features include the spatial distribution of human body key points in the two-dimensional human body posture information, the constraint relationship between human body key points, the sensitivity of human body key points to resolution, and the sensitivity of human body key points to the two-dimensional human body posture information as a whole. In other embodiments, the human body posture features can also be other information, which can be specifically set by those skilled in the art according to actual needs.

[0071] S8. Adjust the pooling range and the pooling mode based on the human body posture feature.

[0072] In some implementations, the method for estimating human body posture further includes a step of training the human body posture estimation model. Specifically, the step of training the human body posture estimation model includes: S91. Obtain a training data set and a real label set, wherein each training data in the training data set has a one-to-one correspondence with one of the real labels in the real label set.

[0073] Each training data is a video sequence used for training. Each real label in the real label set is a real 3D human pose estimation sequence used for training.

[0074] S92: Input the target training data into an original video segmentation module for segmentation processing to obtain a plurality of training segments, wherein the target training data is any one of the training data sets.

[0075] The difference between the original video segmentation module and the video segmentation module is only the difference in model parameters.

[0076] S93, inputting the target training segment into the original two-dimensional human pose estimator for feature extraction processing, obtaining two-dimensional human pose training information corresponding to the target training segment and image training features of multiple resolutions, wherein the target training segment is any one of the multiple training segments. The difference between the original two-dimensional human pose estimator and the two-dimensional human pose estimator is only the difference in model parameters.

[0077] S94. Input the two-dimensional human posture training information and the image training features of multiple resolutions into the original human posture estimation model to perform feature extraction and edge information compensation processing to obtain target fusion training features, wherein the target training data is any one of the training data sets.

[0078] The original human posture estimation model includes an original adaptive posture pooling module, an original second multi-layer perceptron, an original lifting model and an original information fusion module. The difference between the original adaptive posture pooling module and the adaptive posture pooling module is only the difference in model parameters, the difference between the original second multi-layer perceptron and the second multi-layer perceptron is only the difference in model parameters, the difference between the original lifting model and the lifting model is only the difference in model parameters, and the difference between the original information fusion module and the information fusion module is only the difference in model parameters.

[0079] Furthermore, the step S94 includes: S941: Input the two-dimensional human body posture training information and the image training features of multiple resolutions into the original adaptive posture pooling module for feature extraction processing to obtain a first compensation training feature.

[0080] S942: Input the two-dimensional human posture training information into the original second multi-layer perceptron, and the original second multi-layer perceptron projects the two-dimensional human posture training information into a high-dimensional hidden layer feature space to obtain high-dimensional training hidden layer features.

[0081] S943: Input the high-dimensional training hidden layer features into the original lifting model to perform edge information compensation processing to obtain second compensated training features.

[0082] S944: Input the first compensation training feature and the second compensation training feature into the original information fusion module for fusion processing to obtain the target fusion training feature.

[0083] S95: Input the target fusion training feature into the original first multi-layer perceptron for posture estimation processing to obtain a training result corresponding to the target training segment.

[0084] The difference between the original first multilayer perceptron and the first multilayer perceptron is only the difference in model parameters.

[0085] S96: All training results are combined to form a prediction result. The prediction result is a three-dimensional human posture estimation sequence.

[0086] S97. Calculate a loss function based on the prediction result and the true label corresponding to the target training data.

[0087] Furthermore, the calculation formula of the loss function is: ; in, is the loss function to calculate the difference between the predicted result and the true label norm; is the length of the training data set, that is, the total number of all training data; is the true label corresponding to the target training data, that is, The true labels corresponding to the training data; is the prediction result, that is, the prediction result corresponding to the target training data.

[0088] S98. Adjust the model parameters of the original human body posture estimation model based on the loss function to obtain the human body posture estimation model.

[0089] S99, adjusting the parameters of the original video segmentation module based on the loss function to obtain the video segmentation module; adjusting the parameters of the original two-dimensional human posture estimator based on the loss function to obtain the two-dimensional human posture estimator; adjusting the parameters of the original first multilayer perceptron based on the loss function to obtain the first multilayer perceptron. Figure 4 , which is a principle block diagram of a human body posture estimation device provided in the second aspect of the embodiment of the present application. Figure 4 In the embodiment, the human body posture estimation device 100 comprises: An acquisition module 101 is used to acquire a video sequence; A segmentation module 102, configured to input the video sequence into a video segmentation module for segmentation processing to obtain multiple video segments; The feature extraction module 103 is used to input the target video segment into the two-dimensional human posture estimator for feature extraction processing to obtain two-dimensional human posture information and image features of multiple resolutions corresponding to the target video segment, wherein the target video segment is any one of the multiple video segments; A human posture estimation model processing module 104 is used to input the two-dimensional human posture information and the image features of multiple resolutions into a human posture estimation model to perform feature extraction and edge information compensation processing to obtain a target fusion feature corresponding to the target video segment; The posture estimation module 105 is used to input the target fusion feature into the first multi-layer perceptron for posture estimation processing to obtain a three-dimensional human posture estimation result corresponding to the target video segment.

[0090] The merging module 106 is used to merge all the 3D human body posture estimation results to form a 3D human body posture estimation sequence.

[0091] A third aspect of the embodiments of the present application provides a terminal device. The principle block diagram of the terminal device can be as follows: Figure 5 As shown. The terminal device includes a processor, a memory, a network interface, a display screen and a temperature sensor connected via a system bus. The processor is used to provide computing and control capabilities. The memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the terminal device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a method for estimating a human body posture is implemented. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the temperature sensor is pre-set inside the terminal device to detect the operating temperature of the internal device.

[0092] Those skilled in the art will understand that Figure 5 The principle block diagram shown in the figure is only a block diagram of a partial structure related to the scheme of the present invention, and does not constitute a limitation on the terminal device to which the scheme of the present invention is applied. The specific terminal device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0093] In some embodiments, an embodiment of the present application provides a terminal device, the terminal device includes a processor and a memory, the memory is used to store a computer program, the processor is used to call and run the computer program stored in the memory, and execute the steps of the human body posture estimation method provided in the first aspect of the embodiment of the present application. A fourth aspect of the embodiment of the present application provides a computer-readable storage medium, the computer-readable storage medium is used to store a computer program, and the computer program enables a computer to execute the steps of the human body posture estimation method provided in the first aspect of the embodiment of the present application.

[0094] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0095] The technical features of the above embodiments can be combined without changing the basic principles of the present application. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0096] The above embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the patent application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the scope of patent protection of the present application shall be subject to the attached claims.

Claims

1. A method for estimating a human body posture, characterized in that: include: Get video sequence; Inputting the video sequence into a video segmentation module for segmentation processing to obtain multiple video clips; Inputting a target video segment into a two-dimensional human posture estimator for feature extraction processing to obtain two-dimensional human posture information and image features of multiple resolutions corresponding to the target video segment, wherein the target video segment is any one of the multiple video segments; Inputting the two-dimensional human body posture information and the image features of the multiple resolutions into a human body posture estimation model for feature extraction and edge information compensation processing to obtain a target fusion feature corresponding to the target video clip; Inputting the target fusion feature into a first multi-layer perceptron for posture estimation processing to obtain a three-dimensional human posture estimation result corresponding to the target video clip; All 3D human pose estimation results are combined to form a 3D human pose estimation sequence.

2. The method for estimating human body posture according to claim 1, characterized in that: The step of inputting the two-dimensional human body posture information and the image features of the multiple resolutions into a human body posture estimation model for feature extraction and edge information compensation processing to obtain a target fusion feature corresponding to the target video segment comprises: Inputting the two-dimensional human body posture information and the image features of the multiple resolutions into the adaptive posture pooling module of the human body posture estimation model for feature extraction processing to obtain a first compensation feature; Inputting the two-dimensional human body posture information into a second multi-layer perceptron of the human body posture estimation model, wherein the second multi-layer perceptron projects the two-dimensional human body posture information into a high-dimensional hidden layer feature space to obtain high-dimensional hidden layer features; Inputting the high-dimensional hidden layer features into the lifting model of the human posture estimation model to perform edge information compensation processing to obtain a second compensation feature; The first compensation feature and the second compensation feature are input into the information fusion module of the human posture estimation model for fusion processing to obtain the target fusion feature.

3. The method for estimating human body posture according to claim 2, characterized in that: The step of inputting the two-dimensional human body posture information and the image features of the multiple resolutions into the adaptive posture pooling module of the human body posture estimation model for feature extraction processing to obtain the first compensation feature includes: Inputting the two-dimensional human body posture information and the image features of the multiple resolutions into the posture perception offset generation module of the adaptive posture pooling module for sampling processing to generate sampling points; The gesture perception sampling module of the adaptive gesture pooling module performs sampling processing on the image features of the multiple resolutions based on the sampling points to obtain a third compensation feature; The feature updating module of the adaptive posture pooling module uses a preset algorithm to update the third compensation feature to obtain the first compensation feature.

4. The method for estimating human body posture according to claim 3, characterized in that: The two-dimensional human body posture information and the image features of the multiple resolutions are input into the posture perception offset generation module of the adaptive posture pooling module for sampling processing, and the calculation formula for generating sampling points is: ; in, is the sampling point; A function for realizing the gesture perception offset generation function in the gesture perception offset generation module; For the Image features with resolutions of 100; is the total number of resolutions of the image features; is the two-dimensional human body posture information.

5. The method for estimating human body posture according to claim 3, characterized in that: The gesture perception sampling module of the adaptive gesture pooling module performs sampling processing on the image features of the multiple resolutions based on the sampling points, and the calculation formula of the step of obtaining the third compensation feature is: ; in, is the third compensation feature; A function that implements the gesture sensing sampling function in the gesture sensing sampling module; For the Image features with resolutions of 100; is the total number of resolutions of the image features; is the sampling point.

6. The method for estimating human body posture according to claim 3, characterized in that: The feature updating module of the adaptive posture pooling module uses a preset algorithm to update the third compensation feature, and the calculation formula of the step of obtaining the first compensation feature is: ; in, For the the first compensation feature corresponding to the human body posture estimation model, ; is a positive real number in [0, 1]; For the The third compensation feature corresponding to the personal human posture estimation model; For the The third compensation feature corresponding to the personal body posture estimation model.

7. The method for estimating human body posture according to claim 2, characterized in that: The step of inputting the first compensation feature and the second compensation feature into the information fusion module of the human posture estimation model for fusion processing to obtain the target fusion feature includes: The information fusion module uses a multi-head cross attention mechanism to calculate the correlation between the first compensation feature and the second compensation feature to obtain an attention weight; The second compensation feature and the attention weight are fused to obtain the target fusion feature.

8. A human body posture estimation device, characterized in that: include: An acquisition module, used for acquiring a video sequence; A segmentation module, used for inputting the video sequence into a video segmentation module for segmentation processing to obtain multiple video segments; a feature extraction module, configured to input a target video segment into a two-dimensional human posture estimator for feature extraction processing, and obtain two-dimensional human posture information and image features of multiple resolutions corresponding to the target video segment, wherein the target video segment is any one of the multiple video segments; A human posture estimation model processing module, used for inputting the two-dimensional human posture information and the image features of the multiple resolutions into a human posture estimation model for feature extraction and edge information compensation processing, so as to obtain a target fusion feature corresponding to the target video segment; A posture estimation module, used for inputting the target fusion feature into a first multi-layer perceptron for posture estimation processing to obtain a three-dimensional human posture estimation result corresponding to the target video clip; The merging module is used to merge all the 3D human posture estimation results to form a 3D human posture estimation sequence.

9. A terminal device, characterized in that: include: A processor and a memory, the memory being used to store a computer program, the processor being used to call and run the computer program stored in the memory to execute the steps of the method for estimating a human body posture as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Used to store a computer program, wherein the computer program enables a computer to execute the steps of the method for estimating a human body posture as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Road construction cone barrel falling detection method based on attitude estimation

    CN119723518A

  • Method, System and Device for Direct Prediction of 3D Body Poses from Motion Compensated Sequence

    US20170316578A1

Cited By

  • Motion feature extraction method, human body posture estimation method, human body network reconstruction method, equipment and medium

    CN122244472A

  • A method for motion feature extraction, a method for human pose estimation, a method for human network reconstruction, equipment, and media.

    CN122244472B