3D Skeleton Generation via Joint-Based Calibration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for extracting a precise 3D skeleton for virtual reality and augmented reality applications are costly and limited to laboratory settings, with low precision when estimating human posture from photographs without attached devices.
Innovation Solution
A three-dimensional skeleton generation method using calibration based on a joint acquired from a multiview camera, which involves acquiring a multiview color-depth video, generating a three-dimensional skeleton from each viewpoint, optimizing extrinsic parameters, aligning and integrating partial skeletons into a single 3D skeleton, and refining the skeleton using a three-dimensional mesh model.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If various devices such as sensors are attached to a person to extract skeleton, then movement recognition precision is improved, but device cost and complexity increase
Solution Approach 1:
The patent replaces mechanical sensor systems attached to the body with an optical-based deep learning system that processes images from standard cameras. The skeleton extraction is achieved through neural network processing of visual data rather than mechanical sensing, thereby reducing device complexity while maintaining measurement precision.
Solution Approach 2:
The patent creates a virtual 3D skeleton model as a copy of the actual human body structure by processing 2D image data through deep learning. This digital replica allows for precise movement recognition without requiring physical sensors on the body, resolving the contradiction between precision and device complexity.
2Measurement precision
If various devices such as sensors are attached to a person to extract skeleton, then movement recognition precision is improved, but implementation cost increases
Solution Approach 1:
The patent uses standard, inexpensive cameras instead of expensive specialized sensor systems. The deep learning model processes images from these affordable devices to achieve high-precision skeleton extraction, significantly reducing implementation cost while maintaining measurement precision.
Solution Approach 2:
By replacing expensive mechanical sensor systems with software-based deep learning processing on standard cameras, the patent eliminates the need for costly hardware while achieving comparable or superior precision through algorithmic processing.
3Device complexity
If feature extraction from photographs is used to estimate human posture, then device requirements are reduced, but measurement precision deteriorates
Solution Approach 1:
The patent transforms the processing approach by changing from traditional feature extraction methods to deep learning-based parameter estimation. The neural network learns optimal parameters directly from image data, achieving high precision posture estimation while maintaining simple device requirements.
Solution Approach 2:
The patent creates accurate 3D skeleton copies from 2D photographs through deep learning, enabling precise posture estimation without requiring complex imaging devices. The learned model reconstructs three-dimensional body structure from two-dimensional images, resolving the precision-device complexity trade-off.
4Ease of manufacture
If conventional skeleton extraction methods are used, then implementation simplicity is improved, but skeleton precision for 3D content generation deteriorates
Solution Approach 1:
The patent changes the extraction methodology from conventional 2D feature-based approaches to deep learning-based 3D parameter estimation. This allows the system to directly output precise 3D skeleton data suitable for volumetric content generation while keeping the implementation relatively simple through automated neural network processing.
Solution Approach 2:
The patent transitions from 2D image processing to 3D skeleton generation by using deep learning models that infer three-dimensional body structure from two-dimensional photographs. This dimensional transformation enables high-precision 3D content generation while maintaining implementation simplicity through end-to-end learning.
Data Source
AI summary
Proposed is a three-dimensional skeleton generation method using calibration based on a joint acquired from a multiview camera, capable of extracting a partial skeleton of each viewpoint from a distributed RGB-D camera, calculating a camera parameter by using a joint of each partial skeleton as a feature point, and integrating each partial skeleton into a three-dimensional skeleton based on the parameter. The three-dimensional skeleton generation method includes: (a) acquiring a multiview color-depth video; (b) generating a three-dimensional skeleton of each viewpoint from a color-depth video of each viewpoint, and generating a joint of the skeleton of each viewpoint as a feature point; (c) performing extrinsic calibration for optimizing an extrinsic parameter by using the joint of the skeleton of each viewpoint; and (d) aligning and integrating the three-dimensional skeleton of each viewpoint by using the extrinsic parameter.


