Gesture-Based Object Labeling for Sign Language Datasets
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems for associating signs from signing languages with objects are not interactive, require manual adjustment of bounding boxes and polygons, are not agile, and lack automated and dynamic labeling capabilities, especially in environments like those used by the deaf community.
Innovation Solution
A system that automatically and dynamically labels objects using machine learning classification techniques, allowing users to interact with gestures, eliminating the need for manual adjustments and mouse/keyboard controls, and enabling the creation of datasets for sign language classifiers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual adjustment of bounding boxes and polygons is used to label objects, then labeling accuracy can be improved, but time consumption and operational complexity increase significantly
Solution Approach 1:
The system performs preliminary automatic labeling using machine learning models to generate initial bounding boxes and polygons before manual adjustment. This preliminary action provides a close approximation that requires minimal manual refinement, thereby maintaining high accuracy while significantly reducing the time and effort needed for manual adjustment.
Solution Approach 2:
The patent replaces the mechanical manual adjustment process with an automated machine learning-based labeling system. The ML models automatically generate and adjust bounding boxes and polygons, substituting human manual operations with intelligent algorithms that can process data much faster while maintaining or improving accuracy.
2Stability of the object's composition
If traditional labeling systems are used, then system stability is maintained, but user interaction capability and agility are reduced
Solution Approach 1:
The system introduces a gesture recognition intermediary layer between the user and the labeling process. Users interact through natural hand gestures that are captured by computer vision systems, which then translate these gestures into labeling commands. This intermediary enables intuitive interaction while maintaining system stability through controlled processing of gesture inputs.
Solution Approach 2:
The system enables users to interact and label objects through their own natural gestures without requiring learning of complex controls or keyboard shortcuts. The gesture recognition system adapts to user movements and provides immediate feedback, making the system highly agile and easy to operate while maintaining stability through consistent recognition algorithms.
3Productivity
If automated machine learning classification is implemented, then productivity and speed are improved, but system complexity increases
Solution Approach 1:
The system segments the labeling process into distinct functional modules: gesture capture, gesture recognition, object detection, bounding box generation, and labeling output. Each module handles a specific task independently, which manages system complexity by breaking down the automated ML classification process into manageable, well-defined components that can be developed and maintained separately.
Data Source
AI summary
Image data are received into a system. The image data include an object. A first gesture of a human hand is detected in the image data. The first gesture identifies the object for processing. A second gesture of the human hand is detected. The second gesture indicates a readiness to accept a third gesture, and then the third gesture is detected. The object is labeled with the third gesture, and the labeled object is stored in a database. The third gesture can be a sign from a sign language.


