Human pose estimation method and device based on multi-modal perception, terminal and medium

By clustering standard datasets and constructing contextual cue pools, the problem of insufficient generalization performance of 3D human pose estimation models in cross-domain data and complex motion tasks is solved, achieving higher robustness and generalization ability.

CN122290218BActive Publication Date: 2026-07-24PEKING UNIV SHENZHEN GRADUATE SCHOOL
2 Cites 0 Cited by

Patent Information

Application Number
CN202610740605.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-27
Publication Date
2026-07-24
Estimated Expiration
2046-05-27

AI Technical Summary

Technical Problem

Existing 3D human pose estimation models lack generalization performance and robustness when faced with cross-domain data or complex motion tasks, mainly due to the use of random selection in the prompt retrieval process, which leads to mismatch in context logic and imbalance in motion patterns.

Method used

By clustering standard datasets, we determine each standard data class and cluster center, construct a context cue pool, and generate context cue logically related to the query sample by using the embedding matching of human motion sequences and cluster centers. We then perform context learning to determine the 3D human pose.

Benefits of technology

It improves the generalization performance and robustness of the 3D human pose estimation model in cross-domain data and complex motion tasks, ensures the relevance and balance of contextual cues, and enhances the model's generalization ability.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The application discloses a human posture estimation method and device based on multi-modal perception, a terminal and a medium, and relates to the field of computer vision.The method extracts a two-dimensional skeleton sequence based on a human motion sequence corresponding to a target object; clusters standard data sets to determine each standard data class and a corresponding cluster center; determines a context prompt according to the human motion sequence, each standard data class and each cluster center; uses the context prompt as prior information, and uses a human posture estimation model to perform context learning according to the two-dimensional skeleton sequence to determine a corresponding three-dimensional human posture.Because the application determines the context prompt according to each cluster center, each standard data class and the human motion sequence after clustering the standard data sets, the problem that the three-dimensional human posture estimation model has insufficient generalization performance and robustness when facing cross-domain data or complex motion tasks due to the random selection mode of the prompt retrieval link in the prior art can be effectively solved.
Need to check novelty before this filing date? Find Prior Art