Automatic Sign Language Framing via Motion Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional communication systems for hard-of-hearing individuals often require manual adjustment of camera views to capture sign language, leading to inefficiencies and missed gestures, as they focus primarily on framing the user's face rather than the signing area.

Innovation Solution

A video endpoint with a processor-controlled camera that automatically frames the field of view to include a determined signing area of the user, using real-time image processing and motion tracking to ensure that sign language gestures are consistently captured, regardless of the user's movement or position.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If manual adjustment of camera view is used to capture sign language, then the user can control the framing, but it leads to inefficiencies and missed gestures due to focusing on face rather than signing area

Engineering Contradiction:
Improvemanual camera adjustmentVSAvoidcommunication efficiency
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The system performs automatic framing without requiring manual user intervention. The processor analyzes video data to identify the signing area and automatically adjusts camera parameters (zoom, pan, tilt) to frame the signing area, enabling the system to serve itself rather than requiring continuous manual adjustment by the user

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical camera adjustment with an automated image processing system. The processor uses algorithms to detect the signing area in video frames and computationally determines the appropriate camera framing, substituting mechanical user manipulation with electronic image analysis and automated control

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If the camera focuses on the user's face, then facial features are captured clearly, but sign language gestures in the signing area are missed or poorly captured

Engineering Contradiction:
Improvefacial feature captureVSAvoidsign language gestures
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The system extracts the signing area as a distinct region of interest from the overall video frame. By identifying and isolating the signing area through image processing, the system can frame and capture this specific region separately from the face, ensuring that gesture information is not lost while maintaining the ability to capture facial features when needed

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent shifts the framing focus from the traditional face-centered view to a signing-area-centered view. This dimensional shift in the camera's field of view allows the system to capture gestures in the signing area while the face remains visible in peripheral regions, effectively adding spatial dimensionality to the capture strategy

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Productivity

If automatic framing of signing area is implemented, then gesture capture is improved, but device complexity increases due to image processing requirements

Engineering Contradiction:
Improvegesture capture accuracyVSAvoidimage processing system
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The processor in the video endpoint performs multiple functions: it processes audio data, manages video encoding, and now also performs signing area detection and automatic framing control. By making the processor multi-functional, the system avoids adding separate dedicated hardware for each function, thereby managing complexity while achieving improved gesture capture

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10460159B1Video endpoint, communication system, and related method for automatic framing a signing area of a sign language user
Publication Date: 2019.10.29 SORENSON IP HOLDINGS LLC
  • US10460159B1 patent drawing
  • US10460159B1 patent drawing
  • US10460159B1 patent drawing

AI summary

Communication systems and methods are disclosed for enabling a first user at a video endpoint to communicate with a far-end user at a communication device via a relay service providing translation services for the first user. The video endpoint may include a camera and may be configured to frame a view of the camera to include a signing area of a user. The video endpoint may be configured to determine the signing area of the user by taking measurements of the user's body and framing a region around the user to include the signing area based on the measurements, by monitoring a range of motion for the signing area of the user, and other methods.