Verbal Conference Control via Speech Recognition and Speaker ID

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Contemporary conferencing platforms rely on cumbersome manual user input mechanisms, such as dual-tone multi-frequency (DTMF) or keyed input, to activate conference call features, which are not accessible to users with rotary phones and are inconvenient for participants.

Innovation Solution

Deployment of automatic speech recognition functionality in conferencing platforms, where 'hot' words are configured to invoke conference features upon recognition, allowing participants to control features verbally by analyzing a mixed audio stream for hot words and identifying speakers to authorize feature invocations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If manual user input mechanisms (DTMF or keyed input) are used to activate conference call features, then feature control is achieved, but user convenience deteriorates and accessibility is limited

Engineering Contradiction:
Improveuser convenienceVSAvoidinput mechanism complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent replaces mechanical input mechanisms (DTMF key presses, rotary phone dialing) with an acoustic field-based system (speech recognition). The conference bridge analyzes spoken words directly from the audio stream, eliminating the need for physical key presses or specialized input devices. This substitution dramatically improves ease of operation while reducing the complexity of required user actions.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables participants to control conference features through their own natural speech without requiring external control mechanisms. The speech recognition system processes the participant's own voice commands directly, allowing self-service operation of conference features like muting, locking, and roll call activation.

Inventive Principle:
Principle #25Self-service

2Ease of operation

If speech recognition is deployed to enable verbal control, then user convenience is improved, but system complexity increases

Engineering Contradiction:
Improveverbal control capabilityVSAvoidsystem complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The conference bridge is enhanced to perform multiple functions: traditional audio mixing plus speech recognition plus speaker identification. By making the conference bridge a multi-functional platform that handles both audio processing and speech analysis, the patent avoids adding separate dedicated systems, thereby managing overall system complexity while enabling verbal control.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent combines speech recognition functionality and speaker identification capabilities within the existing conference bridge infrastructure. Rather than adding separate independent systems, the conference bridge is upgraded to integrate multiple functions (audio mixing, hot word detection, speaker verification) into a single unified platform, reducing overall system complexity.

Inventive Principle:
Principle #5Merging (Combining)

3Productivity

If hot words are configured to invoke features, then feature activation becomes simpler, but system reliability may be affected by misrecognition

Engineering Contradiction:
Improvefeature activation speedVSAvoidfeature invocation accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system implements a two-stage verification process: first detecting hot words in the audio stream, then verifying speaker identity through speaker identification technology. This feedback mechanism ensures that only authorized speakers can invoke features, significantly improving reliability while maintaining fast activation through the hot word detection stage.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary hot word detection on the audio stream before final feature invocation. By pre-identifying potential feature activation commands and then verifying speaker authority, the system ensures accurate and reliable feature invocation while maintaining rapid response times.

Inventive Principle:
Principle #10Preliminary action

4Reliability

If speaker identification is implemented to authorize features, then access control is improved, but processing time increases

Engineering Contradiction:
Improveauthorized access controlVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

Speaker identification is performed as a preliminary verification step immediately following hot word detection. By quickly analyzing speaker characteristics and comparing against stored profiles, the system provides rapid authorization decisions that minimize processing time while ensuring reliable access control.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system replaces complex manual authorization verification with automated speaker identification technology that analyzes acoustic characteristics. This substitution enables rapid, automated verification of speaker identity and authority, reducing processing time while improving the reliability of access control compared to manual methods.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS8380521B1System, method and computer-readable medium for verbal control of a conference call
Publication Date: 2013.02.19 WEST TECH GRP LLC
  • US8380521B1 patent drawing
  • US8380521B1 patent drawing
  • US8380521B1 patent drawing

AI summary

A system, method, and computer readable medium that facilitate verbal control of conference call features are provided. Automatic speech recognition functionality is deployed in a conferencing platform. Hot words are configured in the conference platform that may be identified in speech supplied to a conference call. Upon recognition of a hot word, a corresponding feature may be invoked. A speaker may be identified using speaker identification technologies. Identification of the speaker may be utilized to fulfill the speaker's request in response to recognition of a hot word and the speaker. Particular participants may be provided with conference control privileges that are not provided to other participants. Upon recognition of a hot word, the speaker may be identified to determine if the speaker is authorized to invoke the conference feature associated with the hot word.