Video Conference Speech Bubbles for Overlapping Conversation Flow

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Video conferences often experience overlapping speech and sound cutoffs, leading to communication inefficiencies due to delayed two-way communication and difficulty in understanding the conversation flow.

Innovation Solution

A method and system that displays voice conversations as speech bubbles on a screen, automatically activating the cartoon mode based on similarity, network conditions, and participant interaction, allowing for clearer communication by organizing speech texts and providing a whisper function.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If multiple participants speak simultaneously during a video conference, then the conversation flow becomes chaotic with overlapping speech, but the system cannot determine the order of speech and communication efficiency decreases

Engineering Contradiction:
Improveconversation flow informationVSAvoidcommunication efficiency
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The patent segments the audio stream into individual speech segments by detecting speech boundaries and separating overlapping speech signals. Each participant's speech is isolated and processed independently, allowing the system to reconstruct the conversation flow accurately even when multiple participants speak simultaneously.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds a visual dimension to the audio communication by displaying speech bubbles with participant avatars and speech text on the screen. This visual representation provides an additional channel for conveying conversation flow information, allowing participants to understand who is speaking and in what order, complementing the audio experience.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If voice signals are transmitted simultaneously from multiple participants, then network bandwidth is consumed and transmission reliability decreases, but two-way communication becomes difficult due to overlapping

Engineering Contradiction:
Improvetransmission reliabilityVSAvoidtwo-way communication
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent performs preliminary speech processing and analysis before transmission, identifying which participants are speaking and separating their speech signals in advance. This preliminary segmentation allows the system to manage multiple voice signals more effectively, reducing network congestion and improving transmission reliability by prioritizing and organizing speech data before it enters the network.

Inventive Principle:
Principle #10Preliminary action

3Loss of information

If speech bubbles are displayed for all participants, then visual information is provided to clarify conversation flow, but screen space is consumed and interface complexity increases

Engineering Contradiction:
Improvespeech order informationVSAvoidinterface complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent applies local quality by displaying speech bubbles with varying levels of detail based on the participant's relevance and speech characteristics. Active speakers receive prominent speech bubbles with full text and avatar display, while less relevant participants may have simplified or condensed representations. This selective detail approach provides necessary information without uniformly cluttering the entire interface.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12563159B2Method for providing speech bubble in video conference and system thereof
Publication Date: 2026.02.24 SAMSUNG SDS CO LTD
  • US12563159B2 patent drawing
  • US12563159B2 patent drawing
  • US12563159B2 patent drawing

AI summary

Provided is a method for providing a speech bubble in a video conference. The method is performed by a user terminal and includes: receiving a first speech text converted from a voice signal of a first conference participant participating in a video conference into text; determining whether to activate a cartoon mode; displaying, based on determining to activate the cartoon mode, a conference screen including a first participant object and a first speech bubble, wherein the first participant object indicates the first conference participant and the first speech bubble is generated using the first speech text; and displaying, in response to a user input to select the first speech bubble, a sequence of speech texts of the video conference, the sequence including a speech text corresponding to the first speech bubble.