Object Positioning Using Depth Images and Edge Distance Weights
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing object positioning technologies using depth images face challenges such as insufficient image recognition due to varying distances between the user and the camera, complex backgrounds, and ambient light effects, which can result in incomplete skeleton information and lower recognition rates.
Innovation Solution
A method and apparatus that convert depth image information into real-world coordinates, compute distances of pixels to edges in multiple directions, assign weights based on these distances, and select extremity positions using a weight limit, allowing for accurate object positioning without establishing a user skeleton and unaffected by ambient light or shelter.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple pre-defined contour shapes are used to track user local area, then the tracking accuracy for different body parts is improved, but the device complexity and processing time increase
Solution Approach 1:
The patent extracts and uses only the depth information component to represent the user's local area, eliminating the need for multiple pre-defined contour shapes. By taking out the essential depth data and processing it through distance calculation to edges, the system achieves tracking accuracy without the complexity of managing multiple contour shape definitions and matches.
Solution Approach 2:
The patent creates a universal approach using a single depth image processing method that can identify different body parts (hands, feet, head, etc.) without requiring separate pre-defined contour shapes for each part. The distance-to-edge calculation serves multiple functions: it detects extremities, determines body part locations, and provides spatial relationships all through one unified algorithm.
2Reliability
If depth images are used to track user local area, then the recognition rate is improved under varying lighting conditions, but the processing time and computational complexity increase
Solution Approach 1:
The patent segments the depth image processing into distinct functional steps: converting depth values to real-world coordinates, calculating distances from each pixel to edges in multiple directions, assigning weights based on distance, and selecting extremity positions. This segmentation allows for optimized processing at each stage and enables parallel computation where applicable, reducing overall processing time while maintaining high recognition accuracy.
Solution Approach 2:
The patent changes the parameter representation from raw depth image values to real-world coordinates and then to weighted distance metrics. By transforming the data through these parameter changes, the system achieves more efficient comparison and selection operations, reducing computational complexity while improving recognition reliability under varying lighting conditions.
3Measurement precision
If color and depth information are combined to locate hand and face areas, then the positioning accuracy is improved, but the device complexity and processing requirements increase
Solution Approach 1:
The patent takes out and uses only the depth information component, extracting the essential spatial relationships needed for positioning. By eliminating the need to process and integrate color information, the system achieves positioning accuracy through depth alone, reducing the complexity of multi-information processing while maintaining the ability to accurately locate hands, faces, and other body parts.
Data Source
AI summary
According to an exemplary embodiment, a method for object positioning by using depth images is executed by a hardware processor as following: converting depth information of each of a plurality of pixels in each of one or more depth images into a real world coordinate; based on the real world coordinate, computing a distance of each pixel to an edge in each of a plurality of directions; assigning a weight to the distance of each pixel to each edge; and based on the weight of the distance of each pixel to each edge and a weight limit, selecting one or more extremity positions of an object.


