System and method for automatically generating a sign language video with an input speech using a machine learning model
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing solutions for generating sign language videos from text inputs fail to capture the speaker's emotion and other attributes, limiting effective communication for deaf and hard-of-hearing individuals.
Innovation Solution
A system and method using a machine learning model that extracts spectrograms from input speech, generates pose sequences, and automatically creates sign language videos by correlating historical pose sequences with spectrograms, employing a generative adversarial network (GAN) to enhance the accuracy and inclusion of speaker's emotion.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If text is used as input modality for sign language generation, then the system can process speech content, but the speaker's emotion and other attributes are lost
Solution Approach 1:
The patent merges multiple input modalities (speech audio signals and text transcripts) into a unified processing framework. The speech-to-text conversion module processes audio signals to generate transcripts, while the emotion detection module analyzes acoustic features of the speech to extract emotional attributes. These combined inputs are then fed to the sign language generation model, preserving both content and emotional information that would be lost with text-only input.
2Quantity of substance
If gloss level annotation is used for sign language datasets, then the dataset can be created, but the annotation process becomes tedious and the dataset size is limited
Solution Approach 1:
The system employs automated machine learning models to perform tasks that would otherwise require manual human annotation. The speech-to-text conversion module automatically transcribes speech to text, the emotion detection module automatically extracts emotional attributes from speech signals, and the sign language generation model automatically creates corresponding sign language sequences. This self-service automation eliminates the need for tedious manual gloss-level annotation, enabling the creation of large-scale datasets without proportional increases in annotation effort.
3Loss of information
If existing sign language platforms use text as input, then they can convert text to sign language, but they miss the speaker's emotion and other attributes
Solution Approach 1:
The system incorporates feedback loops where the generated sign language output is evaluated against the original speech input's emotional attributes. The emotion detection module continuously monitors speech signals for emotional cues, and this emotional information is fed back to the sign language generation model to adjust the generated gestures and expressions. This feedback mechanism ensures that the emotional content of the original speech is preserved and reflected in the generated sign language, enhancing communication effectiveness.
Data Source
AI summary
Embodiments herein provide a system and method for automatically generating a sign language video from an input speech using the machine learning model. The method includes (i) extracting a plurality of spectrograms of an input speech by (a) encoding, using an encoder, a time domain series of the input speech to a frequency domain series, and (b) decoding, using a decoder, a plurality of tokens for time steps of the frequency domain series, (ii) generating a plurality of pose sequences for a current time step of the plurality of spectrograms using a first machine learning model, and (iii) automatically generating, using a discriminator of a second machine learning model, a sign language video for the input speech using the plurality of pose sequences and the plurality of spectrograms when the plurality of pose sequences are matched with corresponding the plurality of spectrograms that are extracted.


