Hub Apparatus Speaker Identification via Audio Timbre Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video communication systems struggle to effectively and reliably identify speakers in multi-user conferences, especially when using mobile terminals with limited camera angles and complex video processing requirements, leading to poor participant visibility and voice recognition issues.
Innovation Solution
A mobile terminal and hub apparatus system that uses voice timbre pattern recognition to generate and transmit a current speaker indicator, allowing the hub to create a single output video stream that highlights the active speaker, with wireless connectivity and efficient processing to manage speaker identification and video stream composition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If video analysis techniques are used to identify the speaker by analyzing lip movement, then speaker identification can be achieved, but the processing resources required become excessively high during the entire conference duration
Solution Approach 1:
The patent extracts the speaker identification function from video analysis and implements it through audio analysis instead. The mobile terminal analyzes audio signals to generate a current speaker indicator, which is then transmitted to the hub apparatus. This extraction of the identification function from video to audio domain significantly reduces processing resource requirements while maintaining identification capability.
Solution Approach 2:
The patent replaces the mechanical/video-based lip movement analysis system with an audio-based signal processing system. By substituting the visual analysis mechanism with audio timbre analysis, the system achieves speaker identification with much lower computational complexity and energy consumption.
2Adaptability or versatility
If a PC is connected to a television with a webcam to enable video conferencing, then video conference functionality is achieved, but the limited shooting angle of the webcam (maximum 90 degrees) causes most participants to either be too far to see clearly or not captured at all
Solution Approach 1:
The patent makes each mobile terminal serve multiple functions: it acts as both a video camera and an audio analysis device for speaker identification. By utilizing the audio capabilities already present in mobile terminals, the system achieves speaker identification without requiring additional specialized equipment, thereby improving versatility while maintaining participation quality.
Solution Approach 2:
The mobile terminal uses its own built-in audio acquisition unit and processing unit to generate the speaker indicator. Each terminal independently performs audio analysis on its own microphone input, eliminating the need for external audio equipment or complex centralized audio processing, thus achieving self-service functionality.
3Reliability
If known endpoint apparatus with video cameras, screens, microphones, and loudspeakers are mounted in meeting rooms, then video conference functionality is achieved, but the apparatus are very expensive and poorly flexible
Solution Approach 1:
The patent merges the speaker identification functionality into the existing mobile terminal hardware that users already possess. By combining audio acquisition, audio processing, and speaker indicator generation within the mobile terminal itself, the system eliminates the need for separate expensive endpoint apparatus while maintaining conference effectiveness.
Solution Approach 2:
The patent utilizes inexpensive mobile terminals that users already own instead of requiring investment in expensive dedicated conference equipment. The system leverages the existing audio and processing capabilities of consumer mobile devices, dramatically reducing system cost while maintaining functional reliability.
Data Source
AI summary
A hub apparatus (20) is designated to be used in a video communication system comprising the hub apparatus (20) and a plurality of mobile terminals (10a-10d) configured to be wirelessly connectable to the hub apparatus (20). The hub apparatus (20) comprises: a receiving unit (24) configured to receive from each mobile terminal (10) of the plurality of mobile terminals (10a-10d) a video stream, a current speaker indicator to indicate whether the user of the mobile terminal is speaking and an association information which associates the current speaker indicator transmitted by the mobile terminal with the video stream transmitted from such mobile terminal (10), and a generation unit (40) operatively connected to said receiving unit (24) and configured to generate an output video communication stream (6) based on the plurality of video streams received from each mobile terminal (10) of the plurality of mobile terminals (10a-10d), on the plurality of current speaker indicators received from each mobile terminal (10) of the plurality of mobile terminals (10a-10d) and on the plurality of association information received from each mobile terminal (10) of the plurality of mobile terminals (10a-10d).


