Hub Apparatus Speaker Identification via Audio Timbre Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video communication systems struggle to effectively and reliably identify speakers in multi-user conferences, especially when using mobile terminals with limited camera angles and complex video processing requirements, leading to poor participant visibility and voice recognition issues.

Innovation Solution

A mobile terminal and hub apparatus system that uses voice timbre pattern recognition to generate and transmit a current speaker indicator, allowing the hub to create a single output video stream that highlights the active speaker, with wireless connectivity and efficient processing to manage speaker identification and video stream composition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If video analysis techniques are used to identify the speaker by analyzing lip movement, then speaker identification can be achieved, but the processing resources required become excessively high during the entire conference duration

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidprocessing resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent extracts the speaker identification function from video analysis and implements it through audio analysis instead. The mobile terminal analyzes audio signals to generate a current speaker indicator, which is then transmitted to the hub apparatus. This extraction of the identification function from video to audio domain significantly reduces processing resource requirements while maintaining identification capability.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces the mechanical/video-based lip movement analysis system with an audio-based signal processing system. By substituting the visual analysis mechanism with audio timbre analysis, the system achieves speaker identification with much lower computational complexity and energy consumption.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If a PC is connected to a television with a webcam to enable video conferencing, then video conference functionality is achieved, but the limited shooting angle of the webcam (maximum 90 degrees) causes most participants to either be too far to see clearly or not captured at all

Engineering Contradiction:
Improvevideo conference setup flexibilityVSAvoidparticipant visibility quality
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent makes each mobile terminal serve multiple functions: it acts as both a video camera and an audio analysis device for speaker identification. By utilizing the audio capabilities already present in mobile terminals, the system achieves speaker identification without requiring additional specialized equipment, thereby improving versatility while maintaining participation quality.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The mobile terminal uses its own built-in audio acquisition unit and processing unit to generate the speaker indicator. Each terminal independently performs audio analysis on its own microphone input, eliminating the need for external audio equipment or complex centralized audio processing, thus achieving self-service functionality.

Inventive Principle:
Principle #25Self-service

3Reliability

If known endpoint apparatus with video cameras, screens, microphones, and loudspeakers are mounted in meeting rooms, then video conference functionality is achieved, but the apparatus are very expensive and poorly flexible

Engineering Contradiction:
Improvevideo conference system effectivenessVSAvoidsystem cost and adaptability
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges the speaker identification functionality into the existing mobile terminal hardware that users already possess. By combining audio acquisition, audio processing, and speaker indicator generation within the mobile terminal itself, the system eliminates the need for separate expensive endpoint apparatus while maintaining conference effectiveness.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent utilizes inexpensive mobile terminals that users already own instead of requiring investment in expensive dedicated conference equipment. The system leverages the existing audio and processing capabilities of consumer mobile devices, dramatically reducing system cost while maintaining functional reliability.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

Data Source

PatentUS11950019B2Mobile terminal and hub apparatus for use in a video communication system
Publication Date: 2024.04.02 HUDDLE ROOM TECH SRL
  • US11950019B2 patent drawing
  • US11950019B2 patent drawing
  • US11950019B2 patent drawing

AI summary

A hub apparatus (20) is designated to be used in a video communication system comprising the hub apparatus (20) and a plurality of mobile terminals (10a-10d) configured to be wirelessly connectable to the hub apparatus (20). The hub apparatus (20) comprises: a receiving unit (24) configured to receive from each mobile terminal (10) of the plurality of mobile terminals (10a-10d) a video stream, a current speaker indicator to indicate whether the user of the mobile terminal is speaking and an association information which associates the current speaker indicator transmitted by the mobile terminal with the video stream transmitted from such mobile terminal (10), and a generation unit (40) operatively connected to said receiving unit (24) and configured to generate an output video communication stream (6) based on the plurality of video streams received from each mobile terminal (10) of the plurality of mobile terminals (10a-10d), on the plurality of current speaker indicators received from each mobile terminal (10) of the plurality of mobile terminals (10a-10d) and on the plurality of association information received from each mobile terminal (10) of the plurality of mobile terminals (10a-10d).