SoC Speaker Highlighting via Gesture and Audio Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In remote video conferences, participants often struggle to identify the current speaker due to multiple participants in the image, leading to communication inefficiencies.

Innovation Solution

A System on Chip (SoC) with person recognition, hand gesture detection, and sound detection circuits processes image and audio data to highlight the current speaker, using deep learning methods to determine speaker identity and region, and applying a gesture lock for accurate tracking.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If multiple participants are displayed in the video conference image, then the conference can include more participants, but it becomes difficult to identify the current speaker

Engineering Contradiction:
Improvenumber of participantsVSAvoidspeaker identification accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent applies visual highlighting (color change) to the current speaker by adding a semi-transparent background or border around the speaker's region in the video image. This allows multiple participants to remain visible while clearly distinguishing the current speaker through color differentiation, resolving the contradiction between showing multiple participants and enabling speaker identification.

Inventive Principle:
Principle #32Color changes

2Area of stationary object

If the image size is reduced to fit the display, then more participants can be seen, but the ability to correctly identify the speaker decreases

Engineering Contradiction:
Improveimage display areaVSAvoidspeaker identification accuracy
Core Design Contradiction:
Area of stationary objectVSMeasurement precision

Solution Approach 1:

By applying color-based highlighting (such as a colored border or semi-transparent background) to the speaker's region, the patent enables accurate speaker identification even when the overall image is compressed to fit smaller display areas. The visual marker compensates for the reduced image quality and size.

Inventive Principle:
Principle #32Color changes

Solution Approach 2:

The patent adds a new visual dimension (color/ transparency overlay) to the existing video image to encode speaker identity information. This additional dimension allows speaker identification without requiring changes to the spatial dimensions of the image, thus maintaining compatibility with reduced image sizes.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If hand gesture detection is added to ensure accurate speaker identification, then speaker tracking precision improves, but system complexity increases

Engineering Contradiction:
Improvespeaker tracking accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The video processing system performs multiple functions using the same image data: it simultaneously displays the video feed, detects hand gestures, identifies speakers, and applies highlighting. By making the system multi-functional and reusing the same processing pipeline for multiple purposes, the patent reduces the need for separate dedicated systems, thereby mitigating the increase in complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240037993A1Video processing method arranged to perform partial highlighting with aid of hand gesture detection and associated system on chip
Publication Date: 2024.02.01 REALTEK SEMICON CORP
  • US20240037993A1 patent drawing
  • US20240037993A1 patent drawing
  • US20240037993A1 patent drawing

AI summary

A video processing method for performing partial highlighting with the aid of hand gesture detection and an associated SoC are provided. The SoC includes a person recognition circuit, a hand gesture detection circuit, a sound detection circuit and a processing circuit. The person recognition circuit obtains image data from an image capturing device, and performs person recognition on the image data to generate a recognition result. The hand gesture detection circuit performs hand gesture detection on hand gesture image data to generate a hand gesture detection result. The sound detection circuit receives multiple sound signals from multiple microphones, and determines a voice characteristic value of a main sound. The processing circuit determines a specific region in the image data according to the recognition result, the hand gesture detection result, and the voice characteristic value, and processes the image data to highlight the specific region.