Sound Emphasis Processing for Virtual Viewpoint Video
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current information processing systems fail to effectively generate and emphasize sounds from specific regions corresponding to a target subject's position in a virtual viewpoint video, leading to inefficient sound processing and potential viewer discomfort due to frequent changes in observation direction.
Innovation Solution
An information processing apparatus that acquires sound information from multiple devices, specifies target sounds based on device and subject positions, and generates emphasis sound information using viewpoint, visual line, and angle-of-view data to focus sound on the target subject, while adjusting output based on observation direction changes and angle-of-view thresholds.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If sound information from multiple sound collection devices is processed to generate virtual viewpoint video, then the completeness of sound information is improved, but the complexity of sound processing increases
Solution Approach 1:
The sound processing is segmented by spatial region. The imaging region is divided into multiple regions based on the positions of sound collection devices, and sound information is processed separately for each region. This allows the system to handle multiple sound sources without processing all sound information uniformly, thereby reducing overall processing complexity while maintaining completeness.
Solution Approach 2:
Different sound processing is applied to different regions. The system determines the region corresponding to the target subject based on viewpoint position and visual line direction, then applies emphasis processing specifically to sound information from that region. This local quality approach allows selective processing of sound information, improving efficiency while maintaining regional accuracy.
2Measurement precision
If sound emphasis is generated based on target subject position, then the accuracy of sound localization is improved, but the computational load increases
Solution Approach 1:
The system preliminarily determines the region corresponding to the target subject based on viewpoint position and visual line direction before generating sound emphasis. This preliminary spatial mapping allows the system to pre-identify which sound information requires emphasis, reducing the computational load during the actual sound emphasis generation process while maintaining accurate localization.
Solution Approach 2:
The system changes processing parameters based on spatial parameters. When the angle of view is within a predetermined range, the system applies different processing parameters (emphasis generation) compared to when the angle of view exceeds the range (integration sound generation). This parameter adaptation reduces computational load by applying intensive processing only when necessary while maintaining localization accuracy.
3Object-affected harmful factors
If sound processing adapts to observation direction changes, then the viewer comfort is improved, but the responsiveness requirement increases
Solution Approach 1:
The sound processing system is made dynamic by adjusting the generation process based on real-time observation direction changes. When the angle of view changes within a predetermined range, the system dynamically switches between emphasis generation and integration sound generation. This dynamic adaptation improves viewer comfort by synchronizing sound with visual attention while maintaining responsive processing through predefined response thresholds.
Data Source
AI summary
An information processing apparatus acquires a plurality of pieces of sound information, sound collection device position information, and target subject position information. In addition, the information processing apparatus specifies a target sound of a region corresponding to a position of a target subject from the plurality of pieces of sound information based on the acquired sound collection device position information and the acquired target subject position information. Further, the information processing apparatus generates target subject emphasis sound information indicating a sound including a target subject emphasis sound in which the specified target sound is emphasized more than a sound emitted from a region different from the region corresponding to the position of the target subject indicated by the acquired target subject position information in a case in which a virtual viewpoint video is generated.


