Hand Pose Estimation Using Local Area Feature Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for estimating poses and models of human hands primarily rely on global features that lack location information, leading to inadequate discrimination and accuracy in describing joint points and model vertices.
Innovation Solution
A method that acquires global features and location codes from input images, divides them into local area features, and uses transformer networks to determine precise location information for joint points and model vertices, enhancing the accuracy of pose and model estimation by incorporating location information and geometric structure.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If global features are used to describe all regions of the hand part, then the method is simple and computationally efficient, but the accuracy and location discrimination power for joint points and model vertices deteriorates
Solution Approach 1:
The patent divides the global feature into multiple local area features corresponding to different regions of the hand (palm, fingers, thumb). Each local area feature is processed separately to extract location information for joint points and model vertices, thereby improving location discrimination power while maintaining computational efficiency through localized processing.
Solution Approach 2:
The patent applies different feature processing strategies to different local areas of the hand. By creating location-specific feature representations for each hand region, the system achieves high accuracy in locating joint points and vertices without requiring complex global processing, thus resolving the contradiction between precision and complexity.
2Reliability
If global features without location information are used, then the computational cost is low, but the ability to discriminate specific positions such as joint points and model vertices deteriorates
Solution Approach 1:
The patent segments the hand feature space into multiple local areas, each processed independently to extract pose information. This segmentation allows the system to achieve reliable pose estimation by focusing computational resources on specific regions, reducing the total number of parameters needed compared to processing the entire hand globally.
Solution Approach 2:
The patent processes only the necessary local areas required for accurate pose estimation rather than analyzing the entire hand globally. By applying feature processing selectively to relevant local regions, the system achieves high reliability in pose estimation while minimizing the quantity of calculation parameters required.
3Productivity
If shared global features are used for all hand regions, then the method is computationally efficient, but the partial discrimination power for feature information deteriorates
Solution Approach 1:
The patent divides the global feature representation into multiple local area features, each preserving location-specific information. This segmentation prevents information loss by ensuring that location details are maintained in dedicated local feature vectors, while still achieving processing efficiency through localized operations on smaller feature subsets.
Solution Approach 2:
The patent enhances local quality by creating specialized feature representations for different hand regions. Each local area feature contains optimized information for its specific region, preventing the loss of location discrimination power while maintaining overall processing efficiency through region-specific feature extraction.
Data Source
AI summary
An object pose and model estimation method includes acquiring a global feature of an input image, and a location code of an object including location information for a joint point of the object and location information for a model vertex in a template model; determining a local area feature of the object based on the global feature of the input image and based on the location code of the object in the template model; and acquiring location information for the joint point of the object in the input image and location information for the model vertex in the input image based on the local area feature of the object.


