Gesture-Based Object Labeling for Sign Language Datasets

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems for associating signs from signing languages with objects are not interactive, require manual adjustment of bounding boxes and polygons, are not agile, and lack automated and dynamic labeling capabilities, especially in environments like those used by the deaf community.

Innovation Solution

A system that automatically and dynamically labels objects using machine learning classification techniques, allowing users to interact with gestures, eliminating the need for manual adjustments and mouse/keyboard controls, and enabling the creation of datasets for sign language classifiers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual adjustment of bounding boxes and polygons is used to label objects, then labeling accuracy can be improved, but time consumption and operational complexity increase significantly

Engineering Contradiction:
Improvelabeling accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary automatic labeling using machine learning models to generate initial bounding boxes and polygons before manual adjustment. This preliminary action provides a close approximation that requires minimal manual refinement, thereby maintaining high accuracy while significantly reducing the time and effort needed for manual adjustment.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces the mechanical manual adjustment process with an automated machine learning-based labeling system. The ML models automatically generate and adjust bounding boxes and polygons, substituting human manual operations with intelligent algorithms that can process data much faster while maintaining or improving accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Stability of the object's composition

If traditional labeling systems are used, then system stability is maintained, but user interaction capability and agility are reduced

Engineering Contradiction:
Improvesystem stabilityVSAvoiduser interaction capability
Core Design Contradiction:
Stability of the object's compositionVSEase of operation

Solution Approach 1:

The system introduces a gesture recognition intermediary layer between the user and the labeling process. Users interact through natural hand gestures that are captured by computer vision systems, which then translate these gestures into labeling commands. This intermediary enables intuitive interaction while maintaining system stability through controlled processing of gesture inputs.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system enables users to interact and label objects through their own natural gestures without requiring learning of complex controls or keyboard shortcuts. The gesture recognition system adapts to user movements and provides immediate feedback, making the system highly agile and easy to operate while maintaining stability through consistent recognition algorithms.

Inventive Principle:
Principle #25Self-service

3Productivity

If automated machine learning classification is implemented, then productivity and speed are improved, but system complexity increases

Engineering Contradiction:
Improvelabeling speedVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system segments the labeling process into distinct functional modules: gesture capture, gesture recognition, object detection, bounding box generation, and labeling output. Each module handles a specific task independently, which manages system complexity by breaking down the automated ML classification process into manageable, well-defined components that can be developed and maintained separately.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12536837B2Object identification and labeling with sign language
Publication Date: 2026.01.27 LENOVO (SINGAPORE) PTE LTD
  • US12536837B2 patent drawing
  • US12536837B2 patent drawing
  • US12536837B2 patent drawing

AI summary

Image data are received into a system. The image data include an object. A first gesture of a human hand is detected in the image data. The first gesture identifies the object for processing. A second gesture of the human hand is detected. The second gesture indicates a readiness to accept a third gesture, and then the third gesture is detected. The object is labeled with the third gesture, and the labeled object is stored in a database. The third gesture can be a sign from a sign language.