Deep Learning Gesture Recognition Using Contour Normalization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing gesture recognition technologies, particularly those using deep learning for sign language recognition, require a large number of training images and are sensitive to environmental conditions, making them inefficient and unreliable with less training data.
Innovation Solution
A method that extracts contours from input images, normalizes contour information, generates training data by adding reliability information, and uses a deep learning model for gesture recognition, which includes feature data extraction and data augmentation to enhance recognition performance without relying on depth information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If deep learning-based sign language recognition is implemented using End-to-End training method, then automatic gesture recognition capability is achieved, but a large number of training images (more than one million) are required
Solution Approach 1:
The patent extracts and utilizes depth information from RGB-D images as a separate feature source, rather than relying solely on RGB image data. This extraction of depth information allows the system to achieve better recognition performance with fewer training images by adding a complementary data dimension that provides structural and spatial context.
Solution Approach 2:
The patent combines multiple data types (RGB image data and depth information) to form a composite training dataset. This composite approach leverages the strengths of both data sources - the color and texture information from RGB images and the spatial and structural information from depth data - thereby improving recognition accuracy while reducing the overall quantity of training images needed.
2Measurement precision
If traditional gesture recognition technology using depth information is used, then recognition accuracy is improved, but environmental limitations and dependency on depth information increase
Solution Approach 1:
The patent merges RGB image data and depth information in a fused training approach, combining the environmental robustness of RGB data with the precision of depth data. This merging allows the system to maintain high recognition accuracy while reducing dependency on any single data source, thereby improving environmental adaptability.
Solution Approach 2:
The patent employs data augmentation techniques that transform and vary training data parameters (such as rotating, scaling, and adjusting RGB and depth image pairs). This parameter transformation creates diverse training samples from limited data, improving the model's ability to generalize across different environmental conditions while maintaining recognition accuracy.
Data Source
AI summary
Deep learning-based automatic gesture recognition method and system are provided. The training method according to an embodiment includes: extracting a plurality of contours from an input image; generating training data by normalizing pieces of contour information forming each of the contours; and training an AI model for gesture recognition by using the generated training data. Accordingly, robust and high-performance automatic gesture recognition can be performed, without being influenced by an environment and a condition even while using less training data.


