Group Call Voice Synthesis for Overlap Resolution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional group call services face issues with overlapping voices during simultaneous utterances, leading to voice loss and poor communication quality.
Innovation Solution
An electronic device with a communication module and processor that senses simultaneous utterances and generates a synthesized voice by connecting overlapping utterances, adjusting reproduction speeds to prevent overlap and ensure clear transmission.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If uttered voices of multiple speakers are transmitted simultaneously during group call, then communication efficiency is improved, but voice overlap and loss occur
Solution Approach 1:
The server acts as an intermediary that receives uttered voices from multiple speakers, processes them through voice activity detection and synthesis, and transmits the synthesized voice to participants. This mediator resolves the conflict by preventing direct simultaneous transmission that causes overlap while maintaining efficient group communication through centralized processing.
Solution Approach 2:
The system changes the parameter of voice transmission from direct simultaneous transmission to synthesized sequential transmission. By detecting voice activity periods and synthesizing voices based on these parameters, the system eliminates overlap while preserving the efficiency of group call functionality.
2Reliability
If synthesized voice is generated by connecting overlapping utterances, then voice overlap is prevented, but transmission time increases
Solution Approach 1:
The server performs preliminary voice activity detection and synthesis processing before transmission. By detecting the voice activity period in advance and pre-synthesizing the voice content, the system prepares the processed voice data ready for immediate transmission, reducing actual transmission time while maintaining high voice quality.
3Measurement precision
If voice activity detection is performed to identify speaking periods, then accurate voice segmentation is achieved, but processing complexity increases
Solution Approach 1:
The server serves as an intermediary that centralizes the complex voice activity detection and processing tasks. Instead of each participant's device performing complex detection, the server handles the sophisticated analysis while participants' devices only need to transmit and receive audio data, reducing individual device complexity while maintaining high detection accuracy.
Data Source
AI summary
An electronic device includes a communication module and a processor operatively connected to the communication module. The processor is configured to: receive and store a first speech voice related to at least a first external device, and a second speech voice related to a second external device; if individual speech is detected, transmit the first speech voice or the second speech voice having a first playback speed to at least a first external device and a second external device; and, if simultaneous speech is detected, convert, into a second playback speed different from the first playback speed, at least a part of a synthesized voice in which at least first overlap speech of the first speech voice and at least second overlap speech of the second speech voice are successively connected, and transmit the synthesized voice to the at least first external device and the second external device.


