Monocular 3D Pose Tokenization for Biomechanical Joint Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for 3D pose estimation from 2D images face challenges with occlusion, depth ambiguity, and domain shift, particularly in capturing biomechanically accurate joint locations and movements, and lack the ability to integrate text and music inputs for diverse and precise motion generation.

Innovation Solution

A computer-implemented method using a pose tokenizer, image conditioned masked transformer, and multi-scale features to generate 3D mesh reconstruction, incorporating text and music inputs for biomechanically accurate pose estimation and motion generation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If marker-based motion capture systems are used for biomechanically accurate pose estimation, then measurement precision and manufacturing precision are improved, but device complexity and cost increase significantly

Engineering Contradiction:
Improvebiomechanical accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces the mechanical marker-based motion capture system with a computational vision system using transformers and neural networks to estimate 3D pose from 2D images, eliminating the need for physical markers and specialized hardware while achieving comparable or superior biomechanical accuracy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent creates a virtual copy of the physical motion capture system by training transformer models on marker-based data, allowing the learned representations to reproduce biomechanically accurate pose estimates without requiring the actual marker infrastructure

Inventive Principle:
Principle #26Copying

2Ease of operation

If existing 3D pose estimation models are used, then ease of operation is improved, but measurement precision deteriorates due to oversimplification of anatomical structures

Engineering Contradiction:
Improveease of useVSAvoidjoint location accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent segments the pose estimation task into multiple components: transformer-based 3D pose estimation, anatomical structure modeling, and biomechanical constraint enforcement, allowing each component to be optimized independently while maintaining overall accuracy

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the representation parameters from simple 2D keypoints to 3D pose parameters with anatomical constraints, enabling the model to capture complex joint locations and movements while maintaining ease of operation through automated processing

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If text-driven human motion generation is used, then adaptability is improved, but precision control over specific human joints deteriorates

Engineering Contradiction:
Improvesemantic guidanceVSAvoidspatial control precision
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent adds a spatial control dimension to text-driven motion generation by enabling precise 3D joint position control while maintaining text-based semantic guidance, allowing simultaneous control of both overall motion semantics and specific joint locations

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent introduces 3D pose representations as an intermediary between text descriptions and final motion output, allowing the model to translate semantic guidance into precise spatial control through the intermediate pose estimation stage

Inventive Principle:
Principle #24Intermediary (Mediator)

4Productivity

If current dance generation models are used, then productivity is improved, but adaptability deteriorates due to lack of text input options

Engineering Contradiction:
Improvedance generation efficiencyVSAvoidinput modality diversity
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent makes the dance generation model universal by accepting multiple input modalities including text descriptions, music audio signals, and pose control signals, allowing the same model to handle diverse choreography creation tasks without requiring separate specialized models

Inventive Principle:
Principle #6Universality (Multi-functionality)

5Productivity

If existing dance generation models are used, then productivity is improved, but adaptability deteriorates due to inability to allow user edits

Engineering Contradiction:
Improveautomatic dance creation speedVSAvoiditerative refinement capability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent introduces dynamic editability to the dance generation process by allowing users to modify generated dance sequences and have the model iteratively refine the output based on user feedback, transforming a static generation process into a dynamic iterative collaboration

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20260065565A1Methods, systems, and computer program products for generating 3D human pose and movement estimation from monocular image information
Publication Date: 2026.03.05 THE UNIV OF NORTH CAROLINA AT CHAPEL HILL
  • US20260065565A1 patent drawing
  • US20260065565A1 patent drawing
  • US20260065565A1 patent drawing

AI summary

A computer-implemented method includes converting by a pose tokenizer, based on a learned codebook, pose parameters of a body into a sequence of discrete pose tokens; randomly masking a portion of the sequence of discrete pose tokens; predicting the randomly masked sequence of discrete pose tokens based on multi-scale features extracted from a monocular image by an image conditioned masked transformer; optimizing the sequence of discrete pose tokens by aligning a re-projected three-dimensional (3D) pose with an estimated two-dimensional (2D) pose; directly regressing, from the multi-scale features, a shape parameter of the body and a weak perspective camera parameter; and generating a 3D mesh reconstruction of the body based on the shape parameter and the weak perspective camera parameter.