Monocular 3D Pose Tokenization for Biomechanical Joint Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for 3D pose estimation from 2D images face challenges with occlusion, depth ambiguity, and domain shift, particularly in capturing biomechanically accurate joint locations and movements, and lack the ability to integrate text and music inputs for diverse and precise motion generation.
Innovation Solution
A computer-implemented method using a pose tokenizer, image conditioned masked transformer, and multi-scale features to generate 3D mesh reconstruction, incorporating text and music inputs for biomechanically accurate pose estimation and motion generation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If marker-based motion capture systems are used for biomechanically accurate pose estimation, then measurement precision and manufacturing precision are improved, but device complexity and cost increase significantly
Solution Approach 1:
The patent replaces the mechanical marker-based motion capture system with a computational vision system using transformers and neural networks to estimate 3D pose from 2D images, eliminating the need for physical markers and specialized hardware while achieving comparable or superior biomechanical accuracy
Solution Approach 2:
The patent creates a virtual copy of the physical motion capture system by training transformer models on marker-based data, allowing the learned representations to reproduce biomechanically accurate pose estimates without requiring the actual marker infrastructure
2Ease of operation
If existing 3D pose estimation models are used, then ease of operation is improved, but measurement precision deteriorates due to oversimplification of anatomical structures
Solution Approach 1:
The patent segments the pose estimation task into multiple components: transformer-based 3D pose estimation, anatomical structure modeling, and biomechanical constraint enforcement, allowing each component to be optimized independently while maintaining overall accuracy
Solution Approach 2:
The patent changes the representation parameters from simple 2D keypoints to 3D pose parameters with anatomical constraints, enabling the model to capture complex joint locations and movements while maintaining ease of operation through automated processing
3Adaptability or versatility
If text-driven human motion generation is used, then adaptability is improved, but precision control over specific human joints deteriorates
Solution Approach 1:
The patent adds a spatial control dimension to text-driven motion generation by enabling precise 3D joint position control while maintaining text-based semantic guidance, allowing simultaneous control of both overall motion semantics and specific joint locations
Solution Approach 2:
The patent introduces 3D pose representations as an intermediary between text descriptions and final motion output, allowing the model to translate semantic guidance into precise spatial control through the intermediate pose estimation stage
4Productivity
If current dance generation models are used, then productivity is improved, but adaptability deteriorates due to lack of text input options
Solution Approach 1:
The patent makes the dance generation model universal by accepting multiple input modalities including text descriptions, music audio signals, and pose control signals, allowing the same model to handle diverse choreography creation tasks without requiring separate specialized models
5Productivity
If existing dance generation models are used, then productivity is improved, but adaptability deteriorates due to inability to allow user edits
Solution Approach 1:
The patent introduces dynamic editability to the dance generation process by allowing users to modify generated dance sequences and have the model iteratively refine the output based on user feedback, transforming a static generation process into a dynamic iterative collaboration
Data Source
AI summary
A computer-implemented method includes converting by a pose tokenizer, based on a learned codebook, pose parameters of a body into a sequence of discrete pose tokens; randomly masking a portion of the sequence of discrete pose tokens; predicting the randomly masked sequence of discrete pose tokens based on multi-scale features extracted from a monocular image by an image conditioned masked transformer; optimizing the sequence of discrete pose tokens by aligning a re-projected three-dimensional (3D) pose with an estimated two-dimensional (2D) pose; directly regressing, from the multi-scale features, a shape parameter of the body and a weak perspective camera parameter; and generating a 3D mesh reconstruction of the body based on the shape parameter and the weak perspective camera parameter.


