System and method for automatically generating a sign language video with an input speech using a machine learning model

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing solutions for generating sign language videos from text inputs fail to capture the speaker's emotion and other attributes, limiting effective communication for deaf and hard-of-hearing individuals.

Innovation Solution

A system and method using a machine learning model that extracts spectrograms from input speech, generates pose sequences, and automatically creates sign language videos by correlating historical pose sequences with spectrograms, employing a generative adversarial network (GAN) to enhance the accuracy and inclusion of speaker's emotion.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If text is used as input modality for sign language generation, then the system can process speech content, but the speaker's emotion and other attributes are lost

Engineering Contradiction:
Improvespeaker's emotionVSAvoidinput processing complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent merges multiple input modalities (speech audio signals and text transcripts) into a unified processing framework. The speech-to-text conversion module processes audio signals to generate transcripts, while the emotion detection module analyzes acoustic features of the speech to extract emotional attributes. These combined inputs are then fed to the sign language generation model, preserving both content and emotional information that would be lost with text-only input.

Inventive Principle:
Principle #5Merging (Combining)

2Quantity of substance

If gloss level annotation is used for sign language datasets, then the dataset can be created, but the annotation process becomes tedious and the dataset size is limited

Engineering Contradiction:
Improvedataset sizeVSAvoidannotation ease
Core Design Contradiction:
Quantity of substanceVSEase of manufacture

Solution Approach 1:

The system employs automated machine learning models to perform tasks that would otherwise require manual human annotation. The speech-to-text conversion module automatically transcribes speech to text, the emotion detection module automatically extracts emotional attributes from speech signals, and the sign language generation model automatically creates corresponding sign language sequences. This self-service automation eliminates the need for tedious manual gloss-level annotation, enabling the creation of large-scale datasets without proportional increases in annotation effort.

Inventive Principle:
Principle #25Self-service

3Loss of information

If existing sign language platforms use text as input, then they can convert text to sign language, but they miss the speaker's emotion and other attributes

Engineering Contradiction:
Improvespeaker's emotionVSAvoidcommunication effectiveness
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The system incorporates feedback loops where the generated sign language output is evaluated against the original speech input's emotional attributes. The emotion detection module continuously monitors speech signals for emotional cues, and this emotional information is fed back to the sign language generation model to adjust the generated gestures and expressions. This feedback mechanism ensures that the emotional content of the original speech is preserved and reflected in the generated sign language, enhancing communication effectiveness.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12475910B2System and method for automatically generating a sign language video with an input speech using a machine learning model
Publication Date: 2025.11.18 INT INST OF INFORMATION THCHNOLOGY HYDERABAD
  • US12475910B2 patent drawing
  • US12475910B2 patent drawing
  • US12475910B2 patent drawing

AI summary

Embodiments herein provide a system and method for automatically generating a sign language video from an input speech using the machine learning model. The method includes (i) extracting a plurality of spectrograms of an input speech by (a) encoding, using an encoder, a time domain series of the input speech to a frequency domain series, and (b) decoding, using a decoder, a plurality of tokens for time steps of the frequency domain series, (ii) generating a plurality of pose sequences for a current time step of the plurality of spectrograms using a first machine learning model, and (iii) automatically generating, using a discriminator of a second machine learning model, a sign language video for the input speech using the plurality of pose sequences and the plurality of spectrograms when the plurality of pose sequences are matched with corresponding the plurality of spectrograms that are extracted.