Verbal Conference Control via Speech Recognition and Speaker ID
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Contemporary conferencing platforms rely on cumbersome manual user input mechanisms, such as dual-tone multi-frequency (DTMF) or keyed input, to activate conference call features, which are not accessible to users with rotary phones and are inconvenient for participants.
Innovation Solution
Deployment of automatic speech recognition functionality in conferencing platforms, where 'hot' words are configured to invoke conference features upon recognition, allowing participants to control features verbally by analyzing a mixed audio stream for hot words and identifying speakers to authorize feature invocations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If manual user input mechanisms (DTMF or keyed input) are used to activate conference call features, then feature control is achieved, but user convenience deteriorates and accessibility is limited
Solution Approach 1:
The patent replaces mechanical input mechanisms (DTMF key presses, rotary phone dialing) with an acoustic field-based system (speech recognition). The conference bridge analyzes spoken words directly from the audio stream, eliminating the need for physical key presses or specialized input devices. This substitution dramatically improves ease of operation while reducing the complexity of required user actions.
Solution Approach 2:
The system enables participants to control conference features through their own natural speech without requiring external control mechanisms. The speech recognition system processes the participant's own voice commands directly, allowing self-service operation of conference features like muting, locking, and roll call activation.
2Ease of operation
If speech recognition is deployed to enable verbal control, then user convenience is improved, but system complexity increases
Solution Approach 1:
The conference bridge is enhanced to perform multiple functions: traditional audio mixing plus speech recognition plus speaker identification. By making the conference bridge a multi-functional platform that handles both audio processing and speech analysis, the patent avoids adding separate dedicated systems, thereby managing overall system complexity while enabling verbal control.
Solution Approach 2:
The patent combines speech recognition functionality and speaker identification capabilities within the existing conference bridge infrastructure. Rather than adding separate independent systems, the conference bridge is upgraded to integrate multiple functions (audio mixing, hot word detection, speaker verification) into a single unified platform, reducing overall system complexity.
3Productivity
If hot words are configured to invoke features, then feature activation becomes simpler, but system reliability may be affected by misrecognition
Solution Approach 1:
The system implements a two-stage verification process: first detecting hot words in the audio stream, then verifying speaker identity through speaker identification technology. This feedback mechanism ensures that only authorized speakers can invoke features, significantly improving reliability while maintaining fast activation through the hot word detection stage.
Solution Approach 2:
The system performs preliminary hot word detection on the audio stream before final feature invocation. By pre-identifying potential feature activation commands and then verifying speaker authority, the system ensures accurate and reliable feature invocation while maintaining rapid response times.
4Reliability
If speaker identification is implemented to authorize features, then access control is improved, but processing time increases
Solution Approach 1:
Speaker identification is performed as a preliminary verification step immediately following hot word detection. By quickly analyzing speaker characteristics and comparing against stored profiles, the system provides rapid authorization decisions that minimize processing time while ensuring reliable access control.
Solution Approach 2:
The system replaces complex manual authorization verification with automated speaker identification technology that analyzes acoustic characteristics. This substitution enables rapid, automated verification of speaker identity and authority, reducing processing time while improving the reliability of access control compared to manual methods.
Data Source
AI summary
A system, method, and computer readable medium that facilitate verbal control of conference call features are provided. Automatic speech recognition functionality is deployed in a conferencing platform. Hot words are configured in the conference platform that may be identified in speech supplied to a conference call. Upon recognition of a hot word, a corresponding feature may be invoked. A speaker may be identified using speaker identification technologies. Identification of the speaker may be utilized to fulfill the speaker's request in response to recognition of a hot word and the speaker. Particular participants may be provided with conference control privileges that are not provided to other participants. Upon recognition of a hot word, the speaker may be identified to determine if the speaker is authorized to invoke the conference feature associated with the hot word.


