Image Sound Pickup Device Audio Video Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing sound pickup systems require high directionality to identify sound source locations, which is challenging in large rooms with many participants, and they rely solely on audio data, lacking effective methods to enhance specific sound sources.

Innovation Solution

An image and sound pickup device that combines audio signal phase differences with face recognition technology to generate location information and display estimated utterer images, allowing for the enhancement of specific sound sources by controlling audio output based on user selection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If sound source location is calculated based only on audio data, then directionality requirement becomes very high, but this makes it difficult to identify sound sources in large rooms with many participants

Engineering Contradiction:
Improvesound source location identification accuracyVSAvoiddirectionality requirement
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines audio data analysis with video image analysis to identify sound source locations. The sound source location calculation unit integrates phase difference information from multiple microphones with face recognition results from video images, allowing accurate sound source identification without requiring high directionality from the microphone array alone.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces video image data as an intermediary to bridge the gap between audio signals and sound source location identification. By using face recognition on video frames, the system can visually identify participants and correlate their positions with audio phase differences, enabling accurate localization without high directionality requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If high directionality is required for sound pickup, then sound source location can be identified accurately, but this becomes challenging in large rooms with many participants

Engineering Contradiction:
Improvesound source location accuracyVSAvoidapplicability in large rooms with many participants
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system merges audio phase difference calculation with video face recognition to identify sound sources. This combination allows the system to handle large rooms with many participants effectively, as the video component provides spatial context that complements the audio data, removing the need for high directionality that would be required in purely audio-based systems.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent adds the visual dimension by incorporating video image analysis into the sound source identification process. By using face recognition on video frames, the system creates a two-dimensional visual reference that complements the audio phase difference information, enabling effective sound source localization in complex environments without requiring high directional sensitivity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Device complexity

If only audio data is used for sound pickup, then system complexity is reduced, but the ability to enhance specific sound sources is limited

Engineering Contradiction:
Improvesystem simplicityVSAvoidsound source enhancement capability
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The patent combines audio data processing with video image processing to enable sound source enhancement. By integrating face recognition results with audio phase difference information, the system can identify and enhance specific participants' voices in the output, providing selective sound enhancement capability while maintaining reasonable system complexity through unified processing architecture.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11227423B2Image and sound pickup device, sound pickup control system, method of controlling image and sound pickup device, and method of controlling sound pickup control system
Publication Date: 2022.01.18 YAMAHA CORP
  • US11227423B2 patent drawing
  • US11227423B2 patent drawing
  • US11227423B2 patent drawing

AI summary

Provided is a method of controlling an image and sound pickup device, which is includes obtaining a plurality of audio signals and a participant image, which shows a plurality of participants, and generating location information about a sound source location by using comparison information about a comparison among the plurality of audio signals and face recognition that is performed on the participant image; and generating an estimated utterer image, which displays an estimated utterer, by using the location information.