Group Call Voice Synthesis for Overlap Resolution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional group call services face issues with overlapping voices during simultaneous utterances, leading to voice loss and poor communication quality.

Innovation Solution

An electronic device with a communication module and processor that senses simultaneous utterances and generates a synthesized voice by connecting overlapping utterances, adjusting reproduction speeds to prevent overlap and ensure clear transmission.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If uttered voices of multiple speakers are transmitted simultaneously during group call, then communication efficiency is improved, but voice overlap and loss occur

Engineering Contradiction:
Improvecommunication efficiencyVSAvoidvoice transmission quality
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The server acts as an intermediary that receives uttered voices from multiple speakers, processes them through voice activity detection and synthesis, and transmits the synthesized voice to participants. This mediator resolves the conflict by preventing direct simultaneous transmission that causes overlap while maintaining efficient group communication through centralized processing.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system changes the parameter of voice transmission from direct simultaneous transmission to synthesized sequential transmission. By detecting voice activity periods and synthesizing voices based on these parameters, the system eliminates overlap while preserving the efficiency of group call functionality.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If synthesized voice is generated by connecting overlapping utterances, then voice overlap is prevented, but transmission time increases

Engineering Contradiction:
Improvevoice transmission qualityVSAvoidtransmission time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The server performs preliminary voice activity detection and synthesis processing before transmission. By detecting the voice activity period in advance and pre-synthesizing the voice content, the system prepares the processed voice data ready for immediate transmission, reducing actual transmission time while maintaining high voice quality.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If voice activity detection is performed to identify speaking periods, then accurate voice segmentation is achieved, but processing complexity increases

Engineering Contradiction:
Improvevoice activity detection accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The server serves as an intermediary that centralizes the complex voice activity detection and processing tasks. Instead of each participant's device performing complex detection, the server handles the sophisticated analysis while participants' devices only need to transmit and receive audio data, reducing individual device complexity while maintaining high detection accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20230410788A1Method for providing group call service, and electronic device supporting same
Publication Date: 2023.12.21 SAMSUNG ELECTRONICS CO LTD
  • US20230410788A1 patent drawing
  • US20230410788A1 patent drawing
  • US20230410788A1 patent drawing

AI summary

An electronic device includes a communication module and a processor operatively connected to the communication module. The processor is configured to: receive and store a first speech voice related to at least a first external device, and a second speech voice related to a second external device; if individual speech is detected, transmit the first speech voice or the second speech voice having a first playback speed to at least a first external device and a second external device; and, if simultaneous speech is detected, convert, into a second playback speed different from the first playback speed, at least a part of a synthesized voice in which at least first overlap speech of the first speech voice and at least second overlap speech of the second speech voice are successively connected, and transmit the synthesized voice to the at least first external device and the second external device.