Audio-Visual Modulation for Spatial Sound Source Positioning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio processing technologies fail to effectively create a realistic and immersive listening experience by simply mixing sound effect elements into music, lacking the ability to simulate a specific sound source position, which affects user immersion.

Innovation Solution

A method and apparatus that perform audio-visual modulation on sound effect elements to determine their position, converting them into dual-channel audio that simulates the sound source's location, enhancing the sense of presence and immersion by rendering the audio as if it originates from a specific scene.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If sound effect elements are simply mixed into original music, then the playing device can play music with added sound effects, but the user cannot feel the artistic conception constructed by the sound effect element, affecting the sense of realism and immersion

Engineering Contradiction:
Improverealism and immersion of listening experienceVSAvoidaudio processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The audio signal is divided into multiple channels (dual-channel audio) to represent different spatial positions. The sound effect element is segmented and assigned to specific channels based on its intended position in the virtual scene, allowing independent control of each channel's audio characteristics to create realistic spatial perception.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a spatial dimension to audio playback by mapping sound effect elements to specific positions in a virtual scene. Dual-channel audio is used to represent different spatial locations, transforming the traditional single-dimension audio playback into a multi-dimensional spatial audio experience that enhances realism and immersion.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If fixed sound effect elements are added to music, then various sound effect elements can be artificially added to achieve special playing effects, but it is difficult for the user to feel the artistic conception constructed by the sound effect element

Engineering Contradiction:
Improvesound effect element selectionVSAvoidartistic conception and immersion
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent implements dynamic sound effect processing where the sound effect element's audio characteristics are adjusted in real-time based on its position in the virtual scene. The dual-channel audio parameters are dynamically modified to reflect the sound source's location, movement, and environmental interactions, creating an adaptive and immersive artistic conception.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes audio parameters (such as pan position, volume, and spatial characteristics) of the sound effect element based on its position in the virtual scene. By dynamically adjusting these parameters according to the sound source location and scene context, the system creates a more realistic and immersive artistic conception that adapts to different listening scenarios.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12089021B2Method and apparatus for listening scene construction and storage medium
Publication Date: 2024.09.10 TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
  • US12089021B2 patent drawing
  • US12089021B2 patent drawing
  • US12089021B2 patent drawing

AI summary

A method and an apparatus for virtual listening scene construction and a storage medium are provided. The method includes the following. Target audio is determined, where the target audio is used to characterize a sound feature in a target scene. A position of a sound source of the target audio is determined. Dual-channel audio of the target audio is obtained by performing audio-visual modulation on the target audio according to the position of the sound source, where the dual-channel audio of the target audio during simultaneous output is able to produce an effect that the target audio is from the position of the sound source. The dual-channel audio of the target audio is rendered into target music to produce an effect that the target music is played in the target scene.