Stereo Audio Compensation Using Autofocus Position Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Most digital cameras have a single microphone, resulting in monaural audio during video playback, which fails to provide a spatial audio experience due to the lack of multiple sound capture points.
Innovation Solution
A recorder system that includes an optical assembly, audio system, and a compensation system to determine the position of a subject relative to the camera, allowing for the creation of a stereophonic sound track from a monaural audio source by adjusting sound based on the subject's position along one, two, or three axes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single microphone is used to capture sound, then the device complexity and cost are reduced, but the audio quality becomes monaural and lacks spatial representation
Solution Approach 1:
The patent creates a virtual copy of the audio signal by generating a second audio signal that is a copy of the first audio signal captured by the single microphone. This virtual copy is then processed to simulate the effect of multiple microphones, allowing the system to reconstruct spatial audio information without physically adding multiple microphones to the device.
Solution Approach 2:
The patent transitions from one-dimensional monaural audio to two-dimensional or three-dimensional stereophonic audio by processing the single audio signal through spatial algorithms. The system adds spatial dimensionality to the audio representation by calculating position information and adjusting the audio signals accordingly, effectively creating a multi-dimensional audio experience from a single microphone input.
2Loss of information
If multiple microphones are added to capture spatial audio, then the stereophonic sound quality is improved, but the device complexity and cost increase
Solution Approach 1:
Instead of physically adding multiple microphones, the patent creates virtual copies of the audio signal through processing. The system generates a second audio signal that replicates the characteristics of a multi-microphone setup by algorithmically processing the single microphone input, thereby achieving the same spatial audio效果 without the physical complexity of multiple microphones.
Solution Approach 2:
The patent replaces the mechanical system of multiple physical microphones with a digital signal processing system. Instead of using additional hardware components to capture spatial audio, the system uses computational algorithms to process the single microphone signal and synthesize the spatial audio experience, substituting mechanical complexity with digital processing.
3Loss of information
If position information is used to adjust the sound track, then the spatial audio representation is improved, but the processing time and complexity increase
Solution Approach 1:
The patent performs preliminary processing by capturing position information of the sound source relative to the microphone at the time of audio recording. This pre-acquired position data is then used during the audio processing stage to adjust the sound track, allowing the system to leverage already-available spatial information rather than requiring real-time complex calculations during audio processing.
Data Source
AI summary
A recorder (10) for recording a scene (12) includes an apparatus frame (218), an optical assembly (220), an image system (222), a position assembly (243), an audio system (224), and a compensation system (248). The image system (222) captures an image (252) of the scene (12). The position assembly (243) can be an autofocusing assembly (244) that focuses the optical assembly (220) on a subject (16) of the scene (12). The position assembly (243) generates position information relating to the position of the subject (16) relative to the recorder (10). The audio system (224) captures a captured sound from the scene (12). The compensation system (248) evaluates the position information and the captured sound from the scene (12) and provides an adjusted sound track in view of the position information.


