Eyewear Sign Language Translation via CNN Gesture Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current portable eyewear devices, such as smart glasses, lack effective solutions for real-time object recognition, translation of sign language, and assistance for visually impaired users, as they do not efficiently convert visual information into audible feedback or text.

Innovation Solution

The eyewear device employs a convolutional neural network (CNN) to identify hand gestures and objects within the user's field of view, converting this information into speech through a camera-based system, enabling real-time translation of sign language and reading of visual content for users with disabilities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If portable eyewear devices integrate cameras and see-through displays, then the device functionality is enhanced, but the ability to provide real-time sign language translation and object recognition is insufficient

Engineering Contradiction:
Improvedevice functionalityVSAvoidreal-time translation capability
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent introduces an intermediary processing system that captures video input, processes it through neural networks for sign language recognition and object detection, and outputs translated text or speech. This intermediary layer bridges the gap between basic camera/display integration and sophisticated real-time translation capabilities, resolving the contradiction by adding specialized processing functionality without requiring complete system redesign.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The eyewear device is designed to perform multiple functions including sign language translation, object recognition, and visual information-to-audio conversion within a single platform. By making the device universal and multi-functional, the patent enhances adaptability while maintaining reliable core translation capabilities through integrated processing.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If the eyewear device uses CNN to identify hand gestures and objects in real-time, then translation accuracy is improved, but processing speed and computational load increase

Engineering Contradiction:
Improvegesture recognition accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-processing video frames, pre-training neural network models for specific gesture recognition tasks, and pre-establishing translation dictionaries. This preliminary preparation enables faster real-time processing while maintaining high accuracy, as the computationally intensive work is done in advance rather than during real-time interaction.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the complex gesture recognition task into distinct components: hand detection, gesture classification, and translation generation. By dividing the processing into separate stages with specialized algorithms for each, the system achieves both high accuracy through focused processing and improved speed by parallelizing independent tasks.

Inventive Principle:
Principle #1Segmentation

3Adaptability or versatility

If the device converts visual information to audible feedback, then accessibility for visually impaired users is improved, but device complexity increases

Engineering Contradiction:
ImproveaccessibilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent replaces complex mechanical or hardware-based visual-to-audio conversion systems with software-based neural network processing and digital signal processing. This substitution maintains the accessibility function while reducing physical complexity, as the conversion is achieved through algorithmic processing rather than mechanical components.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

An intermediary software layer is introduced that handles the conversion from visual input to audible output through text-to-speech synthesis. This mediator simplifies the overall system architecture by creating a clear processing pipeline: camera captures visual information, neural network processes it, and speech synthesis converts it to audio, making the system more manageable despite its multifunctionality.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11900729B2Eyewear including sign language to speech translation
Publication Date: 2024.02.13 SNAP INC
  • US11900729B2 patent drawing
  • US11900729B2 patent drawing
  • US11900729B2 patent drawing

AI summary

An eyewear having an electronic processor configured to identify a hand gesture including a sign language, and to generate speech that is indicative of the identified hand gesture. The electronic processor uses a convolutional neural network (CNN) to identify the hand gesture by matching the hand gesture in the image to a set of hand gestures, wherein the set of hand gestures is a library of hand gestures stored in a memory. The hand gesture can include a static hand gesture, and a moving hand gesture. The electronic processor is configured to identify a word from a series of hand gestures.