Subject-Aware Audio Processing in Imaging for Clear Voice Capture

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing imaging apparatuses face difficulties in accurately following the directivity of a microphone to clearly capture the sound of a specific subject, especially when the subject is moving, making it challenging to obtain clear voice during shooting.

Innovation Solution

An imaging apparatus that integrates image recognition techniques with sound extraction, using a trained neural network to detect specific subjects like persons or animals, and processes audio data to enhance or suppress sounds based on the detected type, ensuring clear sound capture.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If voice detection-based directivity adjustment is used, then sound capture accuracy improves, but ease of operation deteriorates because users must manually follow moving subjects

Engineering Contradiction:
Improvesound capture accuracyVSAvoidease of operation
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The imaging apparatus automatically performs subject tracking and microphone directivity adjustment without requiring user intervention. The system uses the imager to detect subject position and movement, then autonomously adjusts the microphone array's directivity to follow the subject, making the device self-serve the sound capture function.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system continuously monitors subject position through image recognition and uses this feedback to dynamically adjust microphone directivity. The detected subject position information feeds back to the signal processing unit, which real-time adjusts the microphone array's sensitivity pattern to maintain optimal sound capture of the moving subject.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If manual subject following is required, then sound capture precision improves, but device complexity increases due to coordination requirements

Engineering Contradiction:
Improvesound capture precisionVSAvoiddevice complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges the subject tracking function (originally for imaging) with the sound capture control function. The same subject detection and tracking results from the imager are utilized to control microphone directivity, combining multiple functions into a unified system that reduces overall complexity while improving sound capture precision.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The subject detection mechanism serves multiple purposes: it provides imaging information for the display and simultaneously provides control information for microphone directivity adjustment. This multi-functional use of the image recognition system eliminates the need for separate sensing mechanisms, reducing device complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP3709215B1Imaging apparatus
Publication Date: 2026.03.11 PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
  • EP3709215B1 patent drawingFigure 1
  • EP3709215B1 patent drawingFigure 2
  • EP3709215B1 patent drawingFigure 3A~3B

AI summary

An imaging apparatus (100) includes: an imager (115) configured to capture a subject image to generate image data; an audio input device (165) configured to receive audio data indicating sound during; a detector (122) configured to detect a subject and a its type based on the image data generated by the imager (115); an audio processor (170) configured to process the audio data received by the audio input device (165) based on the type of subject detected by the detector (122); and an operation member (150) configured to set a target type to be processed by the audio processor (170) among a plurality of types including first and second types different from each other, based on a user operation, wherein the audio processor (170) is configured to process the audio data to emphasize or suppress specific sound corresponding to the target type in audio data received when a subject of the target type is detected in the image data.