This invention discloses a 3D human
pose estimation method, device, and storage medium, belonging to the field of
computer vision technology. To address the problems of existing technologies ignoring inconsistencies in
human body part motion, susceptibility to
noise interference from auxiliary views, and disruption of spatiotemporal entanglement structures, this invention proposes a cross-view feature adaptive enhancement and spatiotemporal collaborative modeling network. First, multi-view 2D keypoints are acquired; then, enhanced
pose features are generated by jointly modeling the overall
human body and part-specific features through a hierarchical
knowledge extraction module; auxiliary view features are extracted and filtered using the
spatial correlation of the current view as weights, achieving cross-view adaptive enhancement and fusion; subsequently, dual normalization along the channel and time dimensions is performed, and global
spatiotemporal correlation is calculated by combining spatiotemporal identifier encoding to extract spatiotemporal collaborative features that retain the original motion entanglement structure; finally, the 3D coordinates are regressed and output. This invention effectively improves the
estimation accuracy and robustness of 3D human
pose.