Distributed Telepresence Main Speaker Determination
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Distributed telepresence systems face limitations in implementing high-definition video and audio due to the absence of a central server, leading to increased media traffic congestion and reduced immersion for participants.
Innovation Solution
A method for determining a main speaker in a distributed telepresence service by analyzing audio input signals for feature information such as likelihood ratios, pitch, and energy changes, allowing each terminal to request and transmit high-definition video from the main speaker, while displaying it prominently and reducing bandwidth usage by downgrading other participants' videos.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If a centralized telepresence system is used, then it is easy to implement various functions, but media traffic congestion increases and service capacity is limited
Solution Approach 1:
The patent segments the telepresence system into multiple distributed terminals, each capable of independently determining main speakers and processing media packets. This distributes the previously centralized media processing functions across multiple nodes, reducing traffic concentration on a single server while maintaining functional capabilities through peer-to-peer collaboration among terminals.
2Quantity of substance
If a distributed telepresence system is used, then media traffic congestion is reduced, but implementation of various functions becomes limited
Solution Approach 1:
The patent implements multi-functionality by enabling each distributed terminal to perform multiple roles: audio analysis for main speaker determination, video processing at high definition, and media packet handling. Each terminal becomes a multi-functional node that can both consume and provide services, allowing the distributed system to maintain functional versatility without requiring a centralized server.
3Reliability
If high-definition video is provided for all participants in a distributed system, then immersion is enhanced, but bandwidth requirements increase significantly
Solution Approach 1:
The patent applies local quality by providing high-definition video selectively to the terminal displaying the main speaker, while other terminals receive standard definition video. This differentiated quality approach ensures that HD video bandwidth consumption is localized to where it provides the most value (main speaker display), reducing overall system bandwidth requirements while maintaining immersion quality where needed.
Data Source
AI summary
There is provided a method of determining a main speaker that is performed by a first terminal participating in a distributed telepresence service. The method of determining a main speaker according to an embodiment of the invention includes obtaining first feature information for determining a main speaker from an audio input signal, obtaining second feature information for determining a main speaker of a second terminal from the second terminal participating in the distributed telepresence service, and determining a main speaker terminal for providing a video and an audio of a main speaker who is participating in a telepresence and is speaking based on the first feature information for determining a main speaker and the second feature information for determining a main speaker.


