Synthetic Video Model for Hearing-Impaired Speech Correction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Hearing-impaired individuals face difficulties in improving their speech as they cannot hear themselves talk, lacking the feedback loop that hearing-abled people use to correct pronunciation, which is essential for motor memory development.

Innovation Solution

A synthetic video model that analyzes a user's speech and generates a realistic visualization of correct articulations, allowing hearing-impaired individuals to compare and learn proper speech through visual guidance, combined with real-time feedback on tone, volume, and cadence.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If hearing-impaired individuals use conventional speech training methods, then they can access speech improvement resources, but they cannot receive auditory feedback to correct pronunciation

Engineering Contradiction:
Improvefeedback informationVSAvoidhearing impairment
Core Design Contradiction:
Loss of informationVSObject-affected harmful factors

Solution Approach 1:

The patent introduces visual feedback as an intermediary to bridge the gap caused by hearing impairment. A camera captures the user's mouth movements and compares them to correct articulation models, providing visual feedback that compensates for the lack of auditory feedback. This intermediary mechanism allows hearing-impaired individuals to learn speech correction without relying on hearing.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the auditory feedback mechanism with a visual feedback system. Instead of using sound waves to convey pronunciation information, the system uses visual images of mouth articulations captured by a camera and processed by image recognition algorithms. This substitution transforms the feedback modality from acoustic to optical, making it accessible to hearing-impaired users.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If speech feedback systems are implemented, then pronunciation correction is possible, but the system complexity increases

Engineering Contradiction:
Improvearticulation accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates visual copies of correct mouth articulations and compares them with the user's actual articulations. The system captures images of correct pronunciations, processes them through image recognition to extract feature points, and compares these features with the user's speech images. This copying and comparison approach enables precise measurement of articulation accuracy without requiring overly complex analysis systems.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent divides the speech analysis task into discrete segments - capturing individual phonemes or syllables rather than analyzing continuous speech. The image recognition system processes speech images by identifying key feature points (lips, cheeks, jaw) in segmented frames, making the complexity manageable by breaking down the continuous speech stream into analyzable discrete units.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20240144918A1Synthetic video model for instructing a user to improve speech
Publication Date: 2024.05.02 CISCO TECHNOLOGY INC
  • US20240144918A1 patent drawing
  • US20240144918A1 patent drawing
  • US20240144918A1 patent drawing

AI summary

A method, computer system, and computer program product are provided for improving user speech. A data sample of a user speaking one or more words is received, wherein the data sample includes video data and audio data of the user speaking. The data sample is analyzed to determine a correct articulation of a mouth when speaking the one or more words. A synthetic video of the user performing the correct articulation is generated. The synthetic video of the user is presented to the user. A live video of the user is presented to the user while the synthetic video is presented.