Multimodal Speech Recognition for Accents and Speech Impediments
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Speech recognition systems in vehicles face challenges in accurately recognizing speech from users with speech impediments or accents, leading to suboptimal performance.
Innovation Solution
A system and method that integrates auditory and non-auditory inputs, using a combination of audible data and user data, such as facial cues, to improve speech recognition through data fusion techniques like Bayesian networks or neural networks, and fine-tunes a personalized speech recognition model for each user.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional speech recognition systems are used, then the system is simple to operate, but speech recognition accuracy deteriorates for users with speech impediments or accents
Solution Approach 1:
The patent combines auditory input (microphone data) with non-auditory input (camera data capturing facial cues, lip movements, and head gestures) into a unified speech recognition system. This merging of multiple data sources enables accurate recognition of users with speech impediments or accents by compensating for degraded auditory signals with visual information, thereby improving speech recognition accuracy without requiring completely separate systems
Solution Approach 2:
The patent adds a visual dimension to traditional auditory-only speech recognition by incorporating camera captures of facial cues, lip movements, and head gestures. This dimensional expansion from single-sense (auditory) to multi-sense (auditory-visual) input allows the system to interpret speech intent through multiple channels, improving accuracy for users whose speech patterns deviate from norms while maintaining system usability
2Measurement precision
If auditory input alone is used, then the system has low complexity, but speech recognition quality deteriorates for specific users with speech-related special needs
Solution Approach 1:
The patent introduces a data fusion module as an intermediary that processes and integrates auditory data from microphones with non-auditory data from cameras. This intermediary component combines multiple input streams using techniques such as Bayesian networks or neural networks to produce a unified interpretation of user intent, improving speech recognition quality for users with speech impediments while managing processing complexity through structured integration rather than separate parallel systems
Solution Approach 2:
The patent segments the speech recognition input into distinct auditory and non-auditory components, processing each through specialized modules before integration. Auditory data is processed through traditional speech recognition pathways while non-auditory visual data is processed through separate computer vision pathways that identify facial cues, lip movements, and gestures. This segmentation allows optimized processing of each data type while maintaining overall system manageability
3Measurement precision
If a personalized speech recognition system is implemented, then speech recognition accuracy improves for individual users, but system adaptability complexity increases
Solution Approach 1:
The patent implements feedback mechanisms where the system continuously monitors user interactions with speech commands and uses this information to adapt and refine personalized recognition models. The system learns from successful and unsuccessful recognition attempts, adjusting its interpretation of individual user patterns over time. This feedback-driven adaptation improves personalized accuracy while managing complexity through iterative learning rather than requiring complete system reconfiguration for each user
Solution Approach 2:
The patent incorporates preliminary training phases where the system collects and analyzes user-specific speech and visual data patterns before full deployment of personalized recognition. During this preliminary action phase, the system builds user profiles capturing individual speech characteristics, facial expression patterns, and gesture conventions. This advance preparation enables accurate personalized recognition from the outset while managing adaptability complexity through structured onboarding procedures
Data Source
AI summary
A speech recognition method includes receiving audible data and user data. The audible data includes information about an utterance by the user. The user data includes information about movements by the user. The method further includes fusing the audible data and the user data to obtain fused data and determining at least one spoken word of the utterance based on the fused data.


