Sound Field Alignment in 360-Degree Video Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

When collecting sound signals with multiple microphones, the direct mixing results in a divergence between the visual video image range and the auditory sound field range, causing a mismatch in user experience during interactive viewing, especially in VR and AR applications.

Innovation Solution

An apparatus that processes collected sound signals using processors and memory devices to perform frequency analysis, beamforming, and synthesized acoustic signal generation, allowing users to set an angle section and adjust sound field range to match the visual image range, using a spherical microphone array and omnidirectional camera to capture 360-degree video and sound, and outputting stereo acoustic signals that align with the visual-auditory range.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If collected sound signals from multiple microphones are mixed directly, then the sound field range covers all directions, but the user experiences a divergence between the visual video image range and the auditory sound field range

Engineering Contradiction:
Improveomnidirectional sound collectionVSAvoidalignment between visual and auditory ranges
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent applies local quality by making the sound field characteristics direction-dependent. Instead of uniform omnidirectional mixing, the system adjusts the sound field range and sound image positions based on the specific angle section the user is viewing. This allows the auditory output to locally adapt to the visual content being displayed, resolving the divergence between visual and auditory ranges while maintaining omnidirectional collection capability.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent implements dynamics by making the sound field range adjustable based on user-selected angle sections. The system dynamically reconfigures the beamforming parameters and sound image positions according to the current viewing direction, rather than maintaining a fixed omnidirectional sound field. This dynamic adaptation ensures alignment between visual and auditory ranges as the user interacts with the content.

Inventive Principle:
Principle #15Dynamics

2Manufacturing precision

If beamforming is applied to adjust sound field range, then alignment with visual image range is improved, but the complexity of signal processing increases

Engineering Contradiction:
Improvealignment between visual and auditory rangesVSAvoidsignal processing complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the 360-degree sound field into discrete angle sections, each associated with specific visual content. Instead of continuously adjusting beamforming parameters across all directions, the system segments the sound field into manageable sections and applies appropriate beamforming configurations to each. This reduces processing complexity while maintaining precise alignment between visual and auditory ranges.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements preliminary action by pre-calculating and storing beamforming parameters for different angle sections. Rather than performing complex real-time optimization during playback, the system prepares the beamforming configurations in advance based on the video content structure. This preliminary processing reduces the computational burden during actual sound field adjustment while maintaining alignment precision.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12022267B2Apparatus, method and computer-readable storage medium for mixing collected sound signals of microphones
Publication Date: 2024.06.25 KDDI CORP
  • US12022267B2 patent drawing
  • US12022267B2 patent drawing
  • US12022267B2 patent drawing

AI summary

An apparatus comprising: one or more processors; and one or more memory devices configured to store one or more computer programs executable by the one or more processors. The one or more programs, when executed by the one or more processors, cause the apparatus to function as: a setting unit configured to set an angle section at a single sound collection position, selected by a user; a analysis unit configured to convert each of M collected sound signals into a frequency component; a beamforming unit configured to multiply M frequency components obtained through conversion by the analysis unit by respective beamforming matrices to generate a plurality of acoustic signals of two channels; and a signal generation unit configured to synthesize the acoustic signals per channel and outputting an acoustic signal for every channel.