Video-Audio Processing Apparatus for Object Sound Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sound source separation techniques fail to accurately and simply separate the desired object sound from background noise, especially when using mobile devices like camcorders or smartphones, due to challenges in determining the object sound and background sound.
Innovation Solution
A video-audio processing apparatus that includes a display control portion, an object selecting portion, and an extraction portion to display and select video objects, extract audio signals based on object position information, and perform sound source separation using fixed beam forming, enabling the separation of object and background sounds.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If sound source separation techniques are used to separate object sound from background sound, then sound separation capability is improved, but accuracy and simplicity are worsened
Solution Approach 1:
The patent segments the audio signal processing by separating object sound extraction from general sound source separation. It uses video object recognition results to identify target regions, then extracts audio signals specifically from those regions using spatial information, rather than attempting to separate all sound sources simultaneously. This segmentation approach improves both accuracy and simplicity.
Solution Approach 2:
The patent introduces video object recognition results and spatial position information as intermediary elements between the audio signal and the separation process. These intermediaries provide accurate spatial cues that guide the audio extraction, enabling precise object sound separation without relying on complex sound source separation algorithms.
2Extent of automation
If automatic determination of object sound and background sound is implemented, then processing automation is improved, but calculation resource consumption is worsened
Solution Approach 1:
The patent performs video object recognition and spatial position determination as preliminary actions before audio signal extraction. By pre-identifying target objects and their locations in the video frame, the system prepares the necessary spatial information in advance, enabling efficient and automated audio extraction without requiring intensive real-time calculation resources during the audio processing stage.
3Adaptability or versatility
If existing sound source separation techniques are used, then sound separation function is improved, but simplicity of operation is worsened
Solution Approach 1:
The patent merges video object recognition technology with audio signal extraction, combining visual and auditory processing into a unified system. By integrating these functions, the system achieves accurate object sound separation through a single, simple operation rather than requiring separate complex sound source separation processes, thereby improving ease of operation.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The apparatus allows for the simple and accurate separation of desired object sounds by displaying video objects, extracting audio signals based on user selection, and encoding metadata for improved sound source separation.
Implementation Method 1
The extraction portion extracts an audio signal of a selected video object as an audio object signal. The extraction portion extracts a signal other than an audio object signal of the selected video object as a background sound signal from the audio signal.
Implementation Method 2
The object selecting portion produces object position information exhibiting a position of the selected video object on a space, and causes the extraction portion to extract the audio object signal based on the object position information.
Data Source
AI summary
The present technique relates to an apparatus and a method for video-audio processing, and a program each of which enables a desired object sound to be more simply and accurately separated.A video-audio processing apparatus includes a display control portion configured to cause a video object based on a video signal to be displayed; an object selecting portion configured to select the predetermined video object from the one video object or among a plurality of the video objects; and an extraction portion configured to extract an audio signal of the video object selected by the object selecting portion as an audio object signal. The present technique can be applied to a video-audio processing apparatus.


