Signal Processing Apparatus for Virtual Viewpoint Video
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In virtual viewpoint video generation systems, it is challenging to avoid the reflection of sound acquisition operators and shotgun microphones on cameras, and existing sound acquisition techniques struggle to control directivity based on the three-dimensional position of targets, including depth and height.
Innovation Solution
A signal processing apparatus that selects and combines delayed acoustic signals from multiple sound acquisition units based on the estimated position of a target, using a combination of image and sound wave reception units to generate a high-quality acoustic signal while avoiding unnecessary foreground elements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Area of stationary object
If cameras are arranged to surround a target for virtual viewpoint video generation, then 360-degree coverage is achieved, but reflection of sound acquisition operators and microphones on cameras becomes difficult to avoid
Solution Approach 1:
The patent extracts and removes sound acquisition operators and microphones from the virtual viewpoint video generation process by using automated sound acquisition units positioned around the target, eliminating the need for human operators who cause reflections
Solution Approach 2:
The patent uses multiple sound acquisition units positioned around the target to capture sound from different directions, creating a comprehensive sound field that compensates for the inability to avoid reflections from any single position
2Ease of operation
If only azimuth angle estimation is used for directivity control, then simple control is achieved, but three-dimensional position control including depth and height cannot be performed
Solution Approach 1:
The patent transitions from two-dimensional azimuth angle control to three-dimensional position control by incorporating depth and height information from multiple captured images, enabling full spatial directivity control
Solution Approach 2:
The patent creates a universal position estimation system that processes multiple captured images to extract three-dimensional position information, which can be applied to various sound acquisition scenarios requiring spatial control
3Measurement precision
If sound acquisition units are positioned close to the target for high-quality sound capture, then sound quality improves, but the risk of reflection and interference increases
Solution Approach 1:
The patent divides the sound acquisition function into multiple distributed sound acquisition units positioned around the target, each capturing sound from its own position without requiring close proximity that would cause reflections
Solution Approach 2:
The patent merges the acoustic signals from multiple sound acquisition units through signal processing to create a unified high-quality sound output that combines the advantages of multiple positions while avoiding the disadvantages of individual positions
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enables the acquisition of high-quality acoustic signals by selecting and combining delayed sound signals from strategically positioned sound acquisition units, effectively addressing the challenges of reflection and directivity control in virtual viewpoint video generation.
Implementation Method 1
combining delayed acoustic signals obtained by delaying acoustic signals from each of the selected sound acquisition units, based upon a delay amount based upon a distance between the selected sound acquisition unit and the target
Data Source
AI summary
A signal processing apparatus comprises one or more processors, and a memory storing executable instructions which, when executed by the one or more processors, cause the image processing apparatus to function as a selection unit configured to select, as selected sound acquisition units, two or more sound acquisition units from a plurality of sound acquisition units, based upon a position of a target estimated based upon a plurality of captured images including the target, a combining unit configured to combine delayed acoustic signals obtained by delaying acoustic signals from each of the selected sound acquisition units, based upon a delay amount based upon a distance between the selected sound acquisition unit and the target, and an output unit configured to output, as an acoustic signal of the target, a combination result combined by the combination unit.


