Imaging Apparatus Sound Extraction for Ambisonic Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current imaging technologies that record three-dimensional sound fields, such as ambisonics, often result in narration or intentional sounds having directivity, making content uncomfortable for viewers, and there is a need to eliminate unintentional sounds during image capturing.
Innovation Solution
An imaging apparatus is configured to record stereophonic sound and convert specific sound data, like narration or unintended noise, into omnidirectional sound or remove it from the recorded data, using a combination of microphone elements and sound processing units.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If ambisonic microphone records sounds in all directions, then three-dimensional sound field reproduction is achieved, but narration and unintended sounds become directional and uncomfortable for viewers
Solution Approach 1:
The extraction unit separates and extracts specific sound sources (narration, unintended sounds) from the ambient sound field recorded by the ambisonic microphone. By isolating these problematic directional sounds, the system can process them differently from the rest of the three-dimensional sound field, thereby eliminating viewer discomfort while preserving spatial audio accuracy.
Solution Approach 2:
The system applies different processing qualities to different sound components: the extraction unit identifies and extracts specific sound sources with directional characteristics, while the combination processing unit integrates them back with omnidirectional properties. This local differentiation allows narration and unintended sounds to be treated specially without affecting the overall three-dimensional sound field reproduction quality.
2Measurement precision
If all sounds are recorded with directional information, then spatial accuracy is improved, but viewer control over specific sounds (like narration) is lost
Solution Approach 1:
The extraction unit separates specific sound sources from the mixed ambient recording, enabling independent control of extracted sounds (like narration) from the three-dimensional ambient sound field. This extraction capability provides viewers with versatility to control specific sounds while preserving the spatial accuracy of the remaining ambient audio.
3Reliability
If ambisonic technique is used for realistic remote location content, then immersive experience is enhanced, but unintended noise sounds are also captured and become directional
Solution Approach 1:
The extraction unit identifies and extracts unintended noise sounds from the ambient recording captured by the ambisonic microphone. By separating these harmful sounds from the immersive three-dimensional sound field, the system can eliminate or process them independently, preserving the realistic immersive experience while removing unwanted directional noise.
Solution Approach 2:
The system converts the harmful effect of captured unintended noises into a benefit by using the extraction unit to identify and separate these sounds. This allows the narration and unintended sounds to be extracted and processed to have omnidirectional characteristics, transforming the problem of captured noise into an opportunity to enhance overall sound quality and viewer comfort.
Data Source
AI summary
An imaging apparatus includes an imaging circuit and at least one CPU or at least one circuit configured to realize the functions of the following units: a conversion unit configured to convert a sound signal acquired by a microphone including a plurality of microphone elements disposed at respective predetermined angles into stereophonic sound data; an extraction unit configured to extract a sound signal of a specific sound source from the sound signal acquired by the microphone and convert the sound signal of the specific sound source into omnidirectional data of the stereophonic sound data; a combination processing unit configured to combine the stereophonic sound data and the omnidirectional data; and a recording unit configured to record sound data output from the combination processing unit and image data acquired by the imaging circuit.


