Automated Sign Language Video Synthesis for Caption Translation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for sign language translation in video media are inaccurate and limited, particularly for hearing-impaired individuals, due to differences in grammar and expression systems between languages, leading to information distortion and the high cost of using sign language interpreters.

Innovation Solution

A system and method for sign language translation that extracts character strings from video captions, translates them into machine language, matches them with sign language videos, and synchronizes the original video with the sign language video using a neural network algorithm, ensuring high accuracy and user-editable output.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a sign language interpreter is used for communication, then communication accuracy is improved, but cost increases significantly

Engineering Contradiction:
Improvecommunication accuracyVSAvoidcost
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent uses automated caption generation to create text copies of spoken content, which are then translated into sign language videos. This copying approach replaces expensive human interpreters with automated systems that generate accurate captions and corresponding sign language representations, significantly reducing cost while maintaining communication accuracy.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical system of human sign language interpreters with an automated computational system consisting of audio processing, caption generation, and sign language video synthesis. This substitution eliminates the need for human interpreters while providing consistent, accurate translation services at lower cost.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Quantity of substance

If automated audio recognition and translation is used, then cost is reduced, but translation accuracy decreases

Engineering Contradiction:
ImprovecostVSAvoidtranslation accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent segments the translation process into distinct stages: audio processing, caption generation, and sign language video creation. By dividing the complex translation task into manageable segments, each stage can be optimized independently, improving overall accuracy while maintaining cost-effectiveness compared to using human interpreters for the entire process.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces captions as an intermediary between audio and sign language translation. The audio is first converted to text captions, which then serve as the basis for generating sign language videos. This intermediary step improves accuracy by providing a structured text representation that captures the meaning of the audio before translating it into sign language, reducing errors that would occur in direct audio-to-sign-language translation.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Quantity of substance

If text is used for communication with hearing-impaired persons, then cost is reduced, but information distortion and loss occur

Engineering Contradiction:
ImprovecostVSAvoidinformation distortion
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent creates video copies of sign language translations based on the original audio content. These sign language videos preserve the visual and gestural information that text cannot convey, eliminating information distortion and loss while remaining cost-effective compared to human interpreters. The sign language videos serve as accurate visual copies of the communicated information.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS9800955B2Method and system for sign language translation and descriptive video service
Publication Date: 2017.10.24 SAMSUNG ELECTRONICS CO LTD
  • US9800955B2 patent drawing
  • US9800955B2 patent drawing
  • US9800955B2 patent drawing

AI summary

A method and a system for a sign language translation and descriptive video service are disclosed. The method and system enables an easy preparation of video including a descriptive screen and a sign language so that a hearing-impaired person and a visually impaired person can receive a help for using a video media. The method includes extracting a character string in a text form from a caption of an original video; translating the character string in the text form extracted from the caption of the original video to a machine language; matching the character string translated to the machine language with a sign language video in a database; synchronizing the original video with the sign language video, and mixing the original video and the synchronized sign language video; and editing the sign language video with a sign language video editing tool.