Virtual Stereo Audio Generation Using Camera-Object Localization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Portable electronic devices face challenges in generating high-performance stereophonic sounds due to space constraints and the difficulty in mounting high-performance microphones.
Innovation Solution
An electronic apparatus equipped with a camera, a general microphone, and a processor that identifies objects and allocates sounds to corresponding objects, generates sounds of multiple channels by adjusting sound characteristics based on the audio source and object location, and mixes these sounds to produce stereo or surround sounds.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If a general microphone is used in portable electronic devices, then the device can be made compact and easier to manufacture, but the sound quality and stereophonic performance deteriorate
Solution Approach 1:
The system changes parameters of the audio signal by adjusting sound characteristics (volume, pitch, timbre) based on object identity and location. This allows a general microphone to produce high-quality stereophonic sound through software-based parameter modification rather than relying on expensive hardware
Solution Approach 2:
The system creates virtual sound copies by generating stereophonic sound fields from a single microphone input. It copies and distributes audio signals to multiple output channels with different spatial characteristics, simulating the effect of multiple microphones without physically installing them
2Manufacturing precision
If multiple microphones are mounted to achieve high-performance sound output, then the sound quality improves, but the device size increases and space constraints are violated
Solution Approach 1:
The system segments the audio processing function from the physical microphone array. Instead of using multiple physical microphones, it processes the single microphone input into multiple virtual sound channels, dividing the audio signal processing into separate spatial components that simulate multiple source positions
Solution Approach 2:
The system adds a spatial dimension to the audio output by generating stereophonic sound fields from a single-point microphone input. It transforms one-dimensional audio data into three-dimensional spatial audio experience through virtual positioning and characteristic adjustment based on object location
3Manufacturing precision
If a surround microphone is used to receive various sound inputs, then the stereophonic sound quality improves, but the device complexity and manufacturing difficulty increase
Solution Approach 1:
The system makes a general-purpose microphone perform the function of a specialized surround microphone through software processing. The single microphone is universally applied to capture all sound sources, and the processor universally handles different sound types (voice, music, noise) by adjusting characteristics based on object identification and location
Solution Approach 2:
The system replaces the mechanical complexity of multiple physical microphones with software-based audio processing. Instead of mechanically installing multiple sensors, it uses digital signal processing to simulate the acoustic field capture that would require complex hardware arrangements
Data Source
AI summary
An electronic apparatus and/or a controlling method are provided. The electronic apparatus may include a camera for photographing an image, a microphone for receiving an input of a sound of a first channel, and a processor for generating sounds of a plurality of channels based on the input sound, wherein the processor is configured to identify an object and the location of the object from the photographed image, classify the input sound based on an audio source, and allot the sound to the corresponding identified object, copy the classified sound and generate sounds of two channels, adjust characteristics of the generated sounds of two channels based on the audio source allotted to the identified object and the location of the identified object, and mix the sounds of two channels wherein the characteristics were adjusted according to the audio source and generate a stereo sound of two channels.


