3D Pose-Shape-Clothing Recognition via ML Feature Descriptors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image processing technologies fail to accurately distinguish body shape, pose, and clothing in 3D input images, especially with complex clothing, leading to sub-optimal results in generating and animating digital avatars.
Innovation Solution
A machine learning algorithm is trained using synthetic training images to recognize combinations of body shape, pose, and clothing, generating feature descriptors that allow for accurate matching of input images to example images, even with varying viewpoints.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If existing scanning methods treat the scanned human being as a single static object, then the scanning process is simple and fast, but the system fails to distinguish between clothing and body features, resulting in loss of information about body shape and pose
Solution Approach 1:
The patent segments the scanned figure into distinct components: body shape, body pose, and clothing. This is achieved through a machine learning algorithm that processes the depth map to separately identify and estimate these three elements, allowing the system to maintain scanning efficiency while recovering detailed body information that would otherwise be lost in a single static object representation
Solution Approach 2:
The patent introduces an intermediary machine learning algorithm that acts as a mediator between the raw depth map and the final avatar generation. This algorithm processes the depth map to extract and separate body shape, pose, and clothing information, enabling accurate body feature recognition without requiring manual intervention or sacrificing scanning speed
2Measurement precision
If a virtual environment platform needs to recognize certain portions of the human body, then accurate body recognition is achieved, but the user must manually identify key portions of the body, increasing operation time
Solution Approach 1:
The system implements self-service by enabling the machine learning algorithm to automatically identify and recognize key body portions without requiring manual user input. The algorithm processes the depth map and autonomously estimates body shape and pose, eliminating the need for users to manually identify knees, elbows, and other key anatomical features while maintaining high recognition accuracy
3Device complexity
If existing solutions provide estimates of only body shape or only pose, then the processing complexity is reduced, but the system fails to provide comprehensive information needed for accurate avatar animation and redressing
Solution Approach 1:
The patent merges the estimation of body shape, body pose, and clothing into a unified machine learning processing framework. The algorithm simultaneously processes all three elements from the depth map and generates comprehensive information that enables accurate avatar generation, animation, and redressing, avoiding the need for separate processing steps while reducing overall system complexity
Data Source
AI summary
Certain embodiments involve recognizing combinations of body shape, pose, and clothing in three-dimensional input images. For example, synthetic training images are generated based on user inputs. These synthetic training images depict different training figures with respective combinations of a body pose, a body shape, and a clothing item. A machine learning algorithm is trained to recognize the pose-shape-clothing combinations in the synthetic training images and to generate feature descriptors describing the pose-shape-clothing combinations. The trained machine learning algorithm is outputted for use by an image manipulation application. In one example, an image manipulation application uses a feature descriptor, which is generated by the machine learning algorithm, to match an input figure in an input image to an example image based on a correspondence between a pose-shape-clothing combination of the input figure and a pose-shape-clothing combination of an example figure in the example image.


