Human pose estimation method and device based on multi-modal perception, terminal and medium
By clustering standard datasets and constructing contextual cue pools, the problem of insufficient generalization performance of 3D human pose estimation models in cross-domain data and complex motion tasks is solved, achieving higher robustness and generalization ability.
Patent Information
- Application Number
- CN202610740605.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-27
- Publication Date
- 2026-07-24
- Estimated Expiration
- 2046-05-27
AI Technical Summary
Existing 3D human pose estimation models lack generalization performance and robustness when faced with cross-domain data or complex motion tasks, mainly due to the use of random selection in the prompt retrieval process, which leads to mismatch in context logic and imbalance in motion patterns.
By clustering standard datasets, we determine each standard data class and cluster center, construct a context cue pool, and generate context cue logically related to the query sample by using the embedding matching of human motion sequences and cluster centers. We then perform context learning to determine the 3D human pose.
It improves the generalization performance and robustness of the 3D human pose estimation model in cross-domain data and complex motion tasks, ensures the relevance and balance of contextual cues, and enhances the model's generalization ability.