Robot Video Call Speaker Detection and Noise Filtering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Video calls with multiple users suffer from decreased quality due to voice signals from users not appearing on the screen being transmitted, making it difficult to determine the main speaker and resulting in poor video call quality.
Innovation Solution
A robot equipped with a camera, multi-channel microphone, and processor that identifies the main user by analyzing voice signal positions and intensities over time, prioritizing the user who spoke most frequently or is closest to the robot, and centers them in the video frame, while filtering out voice signals from users not in the camera's view.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If all voice signals from multiple users are transmitted during a video call, then the video call can accommodate multiple users, but the video call quality decreases due to voice signals from users not appearing on screen
Solution Approach 1:
The patent extracts and removes voice signals from users who are not visible in the camera view. The processor identifies which users are within the camera's field of view and selectively transmits only their voice signals, while filtering out voice signals from users outside the view. This resolves the contradiction by maintaining multi-user capability while ensuring only relevant voice signals are transmitted, thus preserving video call quality.
2Adaptability or versatility
If multiple users are present on the screen, then the video call supports group communication, but it becomes difficult to determine which user is the main speaker
Solution Approach 1:
The patent applies local quality by differentiating the treatment of different users based on their position and visibility. The processor determines which user is the main speaker by analyzing factors such as proximity to the robot and visibility in the camera view, then prioritizes transmitting that user's voice signal. This allows the system to maintain group video call support while clearly identifying and prioritizing the main speaker, preventing loss of speaker identification information.
3Ease of operation
If voice signals from all users are transmitted regardless of camera view, then the system简单易operate, but the video call quality deteriorates due to noise from off-screen users
Solution Approach 1:
The patent implements feedback by continuously monitoring the camera view and dynamically adjusting which voice signals are transmitted. The processor receives video input, determines which users are visible, and uses this information to control the transmission of voice signals in real-time. This feedback mechanism automatically filters out noise from off-screen users while maintaining simple operation, as the system autonomously makes the filtering decisions based on camera input.
Data Source
AI summary
Provided are a video communication method and a robot implementing the same. The robot includes a camera configured to acquire a first video of a space for a video call, a multi-channel microphone configured to receive a sound signal output to the space, a memory storing one or more instructions, and a processor configured to execute the one or more instructions. The processor determines a first user among N users in the first video based on the sound signal received in a previous time period prior to a first time point, wherein the first user is a main user of the video call at the first time point and N is an integer greater than or equal to 2.


