Keyword Pronunciation Correction in Real-Time Videoconferences
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The mispronunciation of words during videoconferencing can be distracting and irritating, especially when correcting someone's name or specific keywords, due to technical challenges in real-time communication systems.
Innovation Solution
A method and system for automatically correcting mispronounced keywords by identifying them at a server using a database of correct pronunciations, generating corrected audio portions, and transmitting these corrections to participant devices, either through re-encoding at the server or updating at the receiver end.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If real-time audio processing is implemented to correct mispronunciations, then pronunciation accuracy is improved, but system complexity and processing time increase
Solution Approach 1:
The patent introduces an intermediary audio processing system that sits between the audio input and output, automatically detecting mispronunciations and generating corrected audio portions. This intermediary layer handles the complex processing tasks of comparing audio against a database, identifying errors, and synthesizing corrections, thereby improving pronunciation accuracy without requiring direct modification of the core conferencing system architecture.
Solution Approach 2:
The system performs preliminary actions by pre-processing audio data to identify mispronunciations before the audio is fully transmitted to all participants. The server detects erroneous pronunciations and prepares corrected portions in advance, allowing for timely intervention and correction during the conference call without causing noticeable delays.
2Measurement precision
If audio data is processed and corrected at the server, then pronunciation accuracy is improved, but transmission time and network bandwidth increase
Solution Approach 1:
The patent extracts only the portions of audio data that contain mispronunciations for separate processing and correction. Instead of re-transmitting the entire audio stream, the system identifies specific erroneous segments, generates corrected versions of only those segments, and transmits them separately to participants. This extraction approach minimizes the amount of data that needs to be processed and transmitted, reducing the impact on transmission time and network bandwidth.
3Ease of operation
If mispronunciations are corrected in real-time, then user experience is improved, but processing load and computational resources increase
Solution Approach 1:
The system applies partial action by focusing computational resources only on detecting and correcting specific mispronunciations rather than analyzing and processing every audio segment. The audio processing system uses targeted detection methods that identify only erroneous pronunciations against a database of correct pronunciations, generating corrected portions only when needed. This selective approach reduces the overall processing load while still providing the user experience benefit of corrected mispronunciations.
Data Source
AI summary
The present disclosure relates to automatically correcting mispronounced keywords during a conference session. More particularly, the present invention provides methods and systems for automatically correcting audio data generated from audio input having indications of mispronounced keywords during an audio/videoconferencing system. In some embodiments, the process of automatically correcting the audio data may require a re-encoding process of the audio data at the conference server. In alternative embodiments, the process may require updating the audio data at the receiver end of the conferencing system.


