Robot Video Call Speaker Detection and Noise Filtering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Video calls with multiple users suffer from decreased quality due to voice signals from users not appearing on the screen being transmitted, making it difficult to determine the main speaker and resulting in poor video call quality.

Innovation Solution

A robot equipped with a camera, multi-channel microphone, and processor that identifies the main user by analyzing voice signal positions and intensities over time, prioritizing the user who spoke most frequently or is closest to the robot, and centers them in the video frame, while filtering out voice signals from users not in the camera's view.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If all voice signals from multiple users are transmitted during a video call, then the video call can accommodate multiple users, but the video call quality decreases due to voice signals from users not appearing on screen

Engineering Contradiction:
Improvemulti-user video call capabilityVSAvoidvideo call quality
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent extracts and removes voice signals from users who are not visible in the camera view. The processor identifies which users are within the camera's field of view and selectively transmits only their voice signals, while filtering out voice signals from users outside the view. This resolves the contradiction by maintaining multi-user capability while ensuring only relevant voice signals are transmitted, thus preserving video call quality.

Inventive Principle:
Principle #2Taking out (Extraction)

2Adaptability or versatility

If multiple users are present on the screen, then the video call supports group communication, but it becomes difficult to determine which user is the main speaker

Engineering Contradiction:
Improvegroup video call supportVSAvoidmain speaker identification
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent applies local quality by differentiating the treatment of different users based on their position and visibility. The processor determines which user is the main speaker by analyzing factors such as proximity to the robot and visibility in the camera view, then prioritizes transmitting that user's voice signal. This allows the system to maintain group video call support while clearly identifying and prioritizing the main speaker, preventing loss of speaker identification information.

Inventive Principle:
Principle #3Local quality

3Ease of operation

If voice signals from all users are transmitted regardless of camera view, then the system简单易operate, but the video call quality deteriorates due to noise from off-screen users

Engineering Contradiction:
Improvesystem operation simplicityVSAvoidnoise from off-screen users
Core Design Contradiction:
Ease of operationVSObject-affected harmful factors

Solution Approach 1:

The patent implements feedback by continuously monitoring the camera view and dynamically adjusting which voice signals are transmitted. The processor receives video input, determines which users are visible, and uses this information to control the transmission of voice signals in real-time. This feedback mechanism automatically filters out noise from off-screen users while maintaining simple operation, as the system autonomously makes the filtering decisions based on camera input.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10878822B2Video communication method and robot for implementing the method
Publication Date: 2020.12.29 LG ELECTRONICS INC
  • US10878822B2 patent drawing
  • US10878822B2 patent drawing
  • US10878822B2 patent drawing

AI summary

Provided are a video communication method and a robot implementing the same. The robot includes a camera configured to acquire a first video of a space for a video call, a multi-channel microphone configured to receive a sound signal output to the space, a memory storing one or more instructions, and a processor configured to execute the one or more instructions. The processor determines a first user among N users in the first video based on the sound signal received in a previous time period prior to a first time point, wherein the first user is a main user of the video call at the first time point and N is an integer greater than or equal to 2.