Conference System Voice Duplication Prevention via Conversation State Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In conference systems where multiple users participate from different locations, users often experience the issue of hearing the same voice twice, once directly and once through their terminal, leading to audio duplication.
Innovation Solution
A conference system that allocates microphones and speakers to users, where a first microphone acquires and outputs voice to a second speaker, and a second microphone does the same, with a conversation state determiner controlling whether to output voices based on direct conversation state determination to prevent duplication.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the speech system outputs the first acquired voice from the second speaker, then the first user's voice can be heard by the second user, but voice duplication occurs when users are in direct conversation
Solution Approach 1:
The system dynamically adjusts the speaker output based on the detected conversation state. When users are in direct conversation, the speech system suppresses or stops outputting voices through speakers. When users are not in direct conversation, the system enables normal voice output. This dynamic adaptation resolves the contradiction by making the system behavior flexible rather than fixed.
Solution Approach 2:
The system uses a conversation state determiner that continuously monitors and detects whether users are engaged in direct conversation. This feedback mechanism provides real-time information about the conversation state, which is then used by the speech system to control speaker output. The feedback loop enables the system to automatically adjust its behavior to prevent voice duplication while maintaining reliable voice transmission when needed.
2Adaptability or versatility
If the speech system continuously outputs voices, then all users can hear each other, but direct conversation becomes disrupted due to hearing voices twice
Solution Approach 1:
The system transitions between different operational modes based on the conversation state. In direct conversation mode, the system suppresses speaker output to maintain natural audio quality. In non-direct conversation mode, the system enables full speech system functionality. This dynamic mode switching allows the system to adapt to different usage scenarios while maintaining ease of operation in each mode.
Solution Approach 2:
The conversation state determiner automatically detects the conversation state without requiring user input or manual configuration. The system self-adjusts its behavior based on the detected state, eliminating the need for users to manually control the speech system. This self-service approach maintains ease of operation while providing adaptive functionality.
3Quantity of substance
If speakers are allocated to all users, then voice distribution is comprehensive, but audio duplication occurs in same-location conversations
Solution Approach 1:
The system applies different quality characteristics to different spatial contexts. For users at the same location engaged in direct conversation, the system suppresses speaker output to prevent duplication. For users at different locations or not engaged in direct conversation, the system enables normal speaker output. This local quality differentiation allows comprehensive speaker allocation while preventing harmful duplication in specific contexts.
Solution Approach 2:
The speech system dynamically controls speaker output based on real-time detection of direct conversation states. When a direct conversation is detected between users at the same location, the system selectively suppresses output from the involved speakers. When no direct conversation is detected, the system enables full speaker output. This dynamic control maintains comprehensive voice distribution while preventing audio duplication.
Data Source
AI summary
A conference system includes a conversation state determiner that determines whether or not the state of first and second users is a direct conversation state in which direct conversation is possible without using a speech system, and an output controller that controls whether or not to cause the speech system to output a first acquired voice from a second speaker, based on the determination result of the conversation state determiner.


