Voice Command Processing for Conference Calls
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current conferencing systems require manual input to execute functions during conference calls, leading to inefficiencies and pauses, as they lack the ability to predictively apply conferencing functions based on voice commands.
Innovation Solution
Implementing a conference system that uses natural language processing (NLP) to parse voice commands and invoke corresponding conferencing functions, allowing for automatic or semi-automatic execution of functions such as muting participants, sharing screens, and managing conference sessions without the need for explicit user input.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual input is required to execute conferencing functions, then system complexity is reduced, but productivity and ease of operation deteriorate due to inefficiencies and pauses
Solution Approach 1:
The conferencing system automatically executes functions by monitoring and analyzing participant voice communications. The system self-services by detecting keywords and contexts in spoken language, determining appropriate conferencing actions, and executing them without requiring manual user input through interfaces or controls.
Solution Approach 2:
The patent replaces manual mechanical interaction (buttons, switches, menu navigation) with voice-based natural language processing. The system uses speech recognition and natural language understanding to substitute the mechanical control system with an acoustic and linguistic processing system.
2Ease of operation
If manual input is required for conferencing functions, then ease of operation worsens due to pauses and inefficiencies, but measurement precision is maintained
Solution Approach 1:
The system performs preliminary analysis of voice communications continuously during the conference, preparing to execute functions as soon as relevant keywords or contexts are detected. This preliminary monitoring and preparation eliminates delays by having the system ready to act immediately when execution conditions are met.
Solution Approach 2:
The voice monitoring and function execution operates continuously throughout the conference call without interruption. The system maintains constant surveillance of audio inputs and can execute functions at any moment based on real-time analysis, ensuring continuous useful action without pauses or breaks in the conferencing flow.
3Productivity
If predictive application of functions based on voice commands is implemented, then productivity improves, but device complexity increases due to NLP requirements
Solution Approach 1:
The voice processing system is segmented into distinct functional modules: audio capture, speech recognition, natural language processing, keyword detection, context analysis, and function execution. This segmentation allows each component to be optimized independently and simplifies the overall system architecture by dividing complex tasks into manageable segments.
Solution Approach 2:
The natural language processing system serves multiple functions simultaneously: it transcribes speech, detects keywords, determines speaker intent, identifies relevant conferencing contexts, and triggers appropriate functions. This multi-functionality reduces the need for separate specialized systems and decreases overall complexity despite the advanced capabilities required.
Data Source
AI summary
An example method includes receiving at a conference bridge media from a plurality of participants during a conference session and mixing the media received from the plurality of participants to provide mixed media. At least one utterance of the mixed media is parsed using natural language processing to determine a command and at least one subject or object associated with the command. The method also includes invoking a selected conference function during the conference session based on the determined command and each identified subject or object.


