Deep Learning Gesture Recognition Using Contour Normalization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing gesture recognition technologies, particularly those using deep learning for sign language recognition, require a large number of training images and are sensitive to environmental conditions, making them inefficient and unreliable with less training data.

Innovation Solution

A method that extracts contours from input images, normalizes contour information, generates training data by adding reliability information, and uses a deep learning model for gesture recognition, which includes feature data extraction and data augmentation to enhance recognition performance without relying on depth information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If deep learning-based sign language recognition is implemented using End-to-End training method, then automatic gesture recognition capability is achieved, but a large number of training images (more than one million) are required

Engineering Contradiction:
Improveautomatic gesture recognition capabilityVSAvoidnumber of training images
Core Design Contradiction:
Extent of automationVSQuantity of substance

Solution Approach 1:

The patent extracts and utilizes depth information from RGB-D images as a separate feature source, rather than relying solely on RGB image data. This extraction of depth information allows the system to achieve better recognition performance with fewer training images by adding a complementary data dimension that provides structural and spatial context.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent combines multiple data types (RGB image data and depth information) to form a composite training dataset. This composite approach leverages the strengths of both data sources - the color and texture information from RGB images and the spatial and structural information from depth data - thereby improving recognition accuracy while reducing the overall quantity of training images needed.

Inventive Principle:
Principle #40Composite materials

2Measurement precision

If traditional gesture recognition technology using depth information is used, then recognition accuracy is improved, but environmental limitations and dependency on depth information increase

Engineering Contradiction:
Improvegesture recognition accuracyVSAvoidenvironmental adaptability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent merges RGB image data and depth information in a fused training approach, combining the environmental robustness of RGB data with the precision of depth data. This merging allows the system to maintain high recognition accuracy while reducing dependency on any single data source, thereby improving environmental adaptability.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent employs data augmentation techniques that transform and vary training data parameters (such as rotating, scaling, and adjusting RGB and depth image pairs). This parameter transformation creates diverse training samples from limited data, improving the model's ability to generalize across different environmental conditions while maintaining recognition accuracy.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10846568B2Deep learning-based automatic gesture recognition method and system
Publication Date: 2020.11.24 KOREA ELECTRONICS TECH INST
  • US10846568B2 patent drawing
  • US10846568B2 patent drawing
  • US10846568B2 patent drawing

AI summary

Deep learning-based automatic gesture recognition method and system are provided. The training method according to an embodiment includes: extracting a plurality of contours from an input image; generating training data by normalizing pieces of contour information forming each of the contours; and training an AI model for gesture recognition by using the generated training data. Accordingly, robust and high-performance automatic gesture recognition can be performed, without being influenced by an environment and a condition even while using less training data.