Keyword Pronunciation Correction in Real-Time Videoconferences

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The mispronunciation of words during videoconferencing can be distracting and irritating, especially when correcting someone's name or specific keywords, due to technical challenges in real-time communication systems.

Innovation Solution

A method and system for automatically correcting mispronounced keywords by identifying them at a server using a database of correct pronunciations, generating corrected audio portions, and transmitting these corrections to participant devices, either through re-encoding at the server or updating at the receiver end.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If real-time audio processing is implemented to correct mispronunciations, then pronunciation accuracy is improved, but system complexity and processing time increase

Engineering Contradiction:
Improvepronunciation accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary audio processing system that sits between the audio input and output, automatically detecting mispronunciations and generating corrected audio portions. This intermediary layer handles the complex processing tasks of comparing audio against a database, identifying errors, and synthesizing corrections, thereby improving pronunciation accuracy without requiring direct modification of the core conferencing system architecture.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary actions by pre-processing audio data to identify mispronunciations before the audio is fully transmitted to all participants. The server detects erroneous pronunciations and prepares corrected portions in advance, allowing for timely intervention and correction during the conference call without causing noticeable delays.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If audio data is processed and corrected at the server, then pronunciation accuracy is improved, but transmission time and network bandwidth increase

Engineering Contradiction:
Improvepronunciation accuracyVSAvoidtransmission time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent extracts only the portions of audio data that contain mispronunciations for separate processing and correction. Instead of re-transmitting the entire audio stream, the system identifies specific erroneous segments, generates corrected versions of only those segments, and transmits them separately to participants. This extraction approach minimizes the amount of data that needs to be processed and transmitted, reducing the impact on transmission time and network bandwidth.

Inventive Principle:
Principle #2Taking out (Extraction)

3Ease of operation

If mispronunciations are corrected in real-time, then user experience is improved, but processing load and computational resources increase

Engineering Contradiction:
Improveuser experienceVSAvoidprocessing load
Core Design Contradiction:
Ease of operationVSUse of energy by moving object

Solution Approach 1:

The system applies partial action by focusing computational resources only on detecting and correcting specific mispronunciations rather than analyzing and processing every audio segment. The audio processing system uses targeted detection methods that identify only erroneous pronunciations against a database of correct pronunciations, generating corrected portions only when needed. This selective approach reduces the overall processing load while still providing the user experience benefit of corrected mispronunciations.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12437766B2Autocorrection of pronunciations of keywords in audio/videoconferences
Publication Date: 2025.10.07 ADEIA GUIDES INC
  • US12437766B2 patent drawing
  • US12437766B2 patent drawing
  • US12437766B2 patent drawing

AI summary

The present disclosure relates to automatically correcting mispronounced keywords during a conference session. More particularly, the present invention provides methods and systems for automatically correcting audio data generated from audio input having indications of mispronounced keywords during an audio/videoconferencing system. In some embodiments, the process of automatically correcting the audio data may require a re-encoding process of the audio data at the conference server. In alternative embodiments, the process may require updating the audio data at the receiver end of the conferencing system.