Articulation Training System Using Mel-Frequency Cepstral Representations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Hearing impaired individuals face challenges in communicating verbally due to their inability to speak clearly, as speech articulation skills are often linked to hearing ability, and existing devices primarily focus on sign language interpretation rather than speech training.

Innovation Solution

A system and method for articulation training that uses a database of mel-frequency cepstral representations to convert audible inputs into text or images, allowing hearing impaired individuals to practice and improve their speech articulation skills without relying on sign language, through a mobile web application with training, testing, and free talk modes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If sign language learning applications are used, then communication ability is improved, but speech articulation skills are not developed

Engineering Contradiction:
Improvecommunication abilityVSAvoidspeech articulation skills
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The system segments speech training into individual phoneme-level exercises, allowing hearing impaired users to practice and master each sound separately before combining them into words and sentences, thereby developing speech articulation skills while maintaining communication ability

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces visual feedback and haptic guidance as intermediary tools between the user's speech attempts and the target pronunciation, enabling users to see and feel their articulation in real-time without requiring auditory feedback

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If speech training is implemented, then speech articulation is improved, but reliance on sign language decreases

Engineering Contradiction:
Improvespeech articulationVSAvoiddependence on sign language
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The system dynamically adapts the training program based on user progress, gradually increasing difficulty and reducing visual/haptic support as speech articulation improves, enabling users to become more independent in their speech communication over time

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If mel-frequency cepstral representation is used for speech recognition, then speech accuracy is improved, but system complexity increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system uses mel-frequency cepstral coefficients to create simplified spectral representations of speech sounds that capture essential articulation features while reducing computational complexity, enabling accurate speech recognition with efficient processing

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11361677B1System for articulation training for hearing impaired persons
Publication Date: 2022.06.14 KING ABDULAZIZ UNIV
  • US11361677B1 patent drawing
  • US11361677B1 patent drawing
  • US11361677B1 patent drawing

AI summary

A computing device, method, and a non-transitory computer readable medium for articulation training for hearing impaired persons is disclosed. The computing device comprises a database including stored mel-frequency cepstral representations of audio recordings associated with text and/or images related to the audio recordings, a microphone configured to receive audible inputs and a display. The computing device is operatively connected to the database, the microphone and the display. The computing device includes circuitry and program instructions stored therein which when executed by one or more processors, cause the system to receive an audible input from the microphone, convert the audible input to a mel-frequency cepstral representation, search the database for a match of the mel-frequency cepstral representation to a stored mel-frequency cepstral representation and display the text and/or images related to the stored mel-frequency cepstral representation when the match is found.