Distributed Telepresence Main Speaker Determination

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Distributed telepresence systems face limitations in implementing high-definition video and audio due to the absence of a central server, leading to increased media traffic congestion and reduced immersion for participants.

Innovation Solution

A method for determining a main speaker in a distributed telepresence service by analyzing audio input signals for feature information such as likelihood ratios, pitch, and energy changes, allowing each terminal to request and transmit high-definition video from the main speaker, while displaying it prominently and reducing bandwidth usage by downgrading other participants' videos.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If a centralized telepresence system is used, then it is easy to implement various functions, but media traffic congestion increases and service capacity is limited

Engineering Contradiction:
Improveease of implementationVSAvoidmedia traffic load
Core Design Contradiction:
Ease of manufactureVSQuantity of substance

Solution Approach 1:

The patent segments the telepresence system into multiple distributed terminals, each capable of independently determining main speakers and processing media packets. This distributes the previously centralized media processing functions across multiple nodes, reducing traffic concentration on a single server while maintaining functional capabilities through peer-to-peer collaboration among terminals.

Inventive Principle:
Principle #1Segmentation

2Quantity of substance

If a distributed telepresence system is used, then media traffic congestion is reduced, but implementation of various functions becomes limited

Engineering Contradiction:
Improvemedia traffic loadVSAvoidfunction implementation capability
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The patent implements multi-functionality by enabling each distributed terminal to perform multiple roles: audio analysis for main speaker determination, video processing at high definition, and media packet handling. Each terminal becomes a multi-functional node that can both consume and provide services, allowing the distributed system to maintain functional versatility without requiring a centralized server.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If high-definition video is provided for all participants in a distributed system, then immersion is enhanced, but bandwidth requirements increase significantly

Engineering Contradiction:
Improveimmersion qualityVSAvoidbandwidth consumption
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent applies local quality by providing high-definition video selectively to the terminal displaying the main speaker, while other terminals receive standard definition video. This differentiated quality approach ensures that HD video bandwidth consumption is localized to where it provides the most value (main speaker display), reducing overall system bandwidth requirements while maintaining immersion quality where needed.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS9368127B2Method and device for providing distributed telepresence service
Publication Date: 2016.06.14 TSTONE IP LLC
  • US9368127B2 patent drawing
  • US9368127B2 patent drawing
  • US9368127B2 patent drawing

AI summary

There is provided a method of determining a main speaker that is performed by a first terminal participating in a distributed telepresence service. The method of determining a main speaker according to an embodiment of the invention includes obtaining first feature information for determining a main speaker from an audio input signal, obtaining second feature information for determining a main speaker of a second terminal from the second terminal participating in the distributed telepresence service, and determining a main speaker terminal for providing a video and an audio of a main speaker who is participating in a telepresence and is speaking based on the first feature information for determining a main speaker and the second feature information for determining a main speaker.