Display Audio Spatial Positioning via Virtual Speaker Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current stereophonic sound services in display apparatuses fail to simulate the effect of audio being output from specific image locations, providing only stereo sound through multiple channels without accurately correlating audio with video images.
Innovation Solution
A display apparatus equipped with a controller and audio processor that detects vocalized positions within video frames and adjusts audio signals accordingly, dividing and recombining audio signals based on distance and audio characteristics to create output signals for multiple speakers, simulating audio output from image locations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Illumination intensity
If audio signals are output through multiple speakers using multi-channel technology, then audio quality and stereo effect are improved, but the ability to simulate audio output from specific image locations is lost
Solution Approach 1:
The patent segments the audio signal processing by dividing it into multiple virtual speakers corresponding to different regions of the display screen. Each virtual speaker handles audio from specific image locations, allowing the system to maintain spatial positioning information while using multi-speaker output. The audio signal is segmented and routed to appropriate virtual speakers based on the vocalized position in the video frame.
Solution Approach 2:
The patent introduces virtual speakers as intermediary elements between the physical speakers and the image content. These virtual speakers act as mediators that associate audio signals with specific spatial positions on the screen, enabling the system to simulate audio output from image locations while utilizing multiple physical speakers for high-quality audio reproduction.
2Loss of information
If audio signals are processed to simulate output from image locations, then spatial audio effect is improved, but system complexity increases
Solution Approach 1:
The controller in the patent performs multiple functions: it detects vocalized positions in video frames, determines distances between vocalized positions and display regions, processes audio signals accordingly, and controls multiple speakers. By making the controller universal and multi-functional, the patent avoids adding separate dedicated components for each processing step, thereby managing system complexity while achieving spatial audio effects.
Solution Approach 2:
The system uses the existing display apparatus components (controller, speakers, display screen) for audio processing rather than requiring entirely separate dedicated audio equipment. The controller that already manages video display also handles audio processing and spatial positioning, making the system self-sufficient and reducing overall complexity.
3Loss of information
If the system adjusts audio signals based on vocalized position detection, then audio-visual correlation is improved, but processing time and computational load increase
Solution Approach 1:
The system detects vocalized positions and calculates distances in advance before audio signal processing. By performing position detection and distance calculation preliminarily, the patent prepares the spatial information needed for audio processing, enabling more efficient real-time audio signal adjustment and reducing overall processing time.
Data Source
Figure 1
Figure 2
Figure 3(a)~3(c)
AI summary
Embodiments disclose a display apparatus including a controller configured to detect a vocalized position in the video frame; and an audio processor configured to process an audio signal corresponding to the video frame differently according to a distance between the vocalized position and each of the plurality of speakers, create a plurality of audio output signals, and provide each created audio output signal to each of the plurality of speakers, and the controller controls the audio processor to change the each created audio output signal provided to the each of the plurality of speakers according to the moved vocalized position in response to the vocalized position being moved within the video frame.