Audio Agent Sound Field Control via Wavefront Composition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio-based agent systems face challenges in effectively processing and controlling audio outputs to prevent mixing with other audio content, leading to difficulties in localizing sound images and maintaining user interaction without interfering with other audio sources.

Innovation Solution

An information processing device and method that acquires audio information from an agent device and other content sources, using sound field control processing to isolate and localize audio outputs through wavefront composition, ensuring distinct audio experiences for users.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If audio output from multiple sources is played simultaneously, then the system provides comprehensive information to the user, but the audio outputs mix together causing interference and making it difficult to distinguish agent audio from other content

Engineering Contradiction:
Improvemulti-source audio playback capabilityVSAvoidaudio clarity and distinguishability
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent segments the audio output space by creating distinct sound fields for different audio sources. The sound field control unit divides the physical space into multiple regions, assigning each audio source (agent device and other content) to a specific spatial zone. This spatial segmentation allows simultaneous playback without mixing, as each source occupies its own acoustic region.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from two-dimensional audio mixing (stereo left-right channels) to three-dimensional sound field control by utilizing vertical and depth dimensions. Through wavefront composition and phase control, the system creates spatially separated audio zones in 3D space, allowing multiple audio sources to coexist without interference by positioning them at different spatial coordinates.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Loss of information

If sound field control processing is applied to separate audio sources, then audio clarity and localization are improved, but the device complexity and processing requirements increase

Engineering Contradiction:
Improveaudio clarityVSAvoidsound field control system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent introduces a sound field control unit as an intermediary component between the audio sources and the speakers. This intermediary processes the audio signals and coordinates the output across multiple speakers to create the desired sound field. By centralizing the control function, the system manages complexity through a dedicated control module rather than requiring complex modifications to each audio source.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent manipulates acoustic parameters such as phase, amplitude, and timing of audio signals to achieve sound field separation. By adjusting these parameters dynamically, the system creates distinct acoustic zones without requiring physical separation of speakers or audio sources. This parameter-based control achieves the desired effect while maintaining a relatively simple hardware configuration.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If wavefront composition is used to localize sound images, then audio positioning accuracy is improved, but the computational load and processing time increase

Engineering Contradiction:
Improvesound image localization accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary calculations of speaker timing and phase relationships before audio playback. The sound field control unit pre-computes the required signal adjustments for each speaker based on the desired sound field configuration. This preliminary action allows the system to quickly switch between different audio sources and positions without performing complex real-time calculations during playback, thus reducing processing time.

Inventive Principle:
Principle #10Preliminary action

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This solution enables clear and localized audio experiences for users, preventing interference between agent audio and other content, thereby enhancing user interaction and audio clarity in multi-source environments.

Implementation Method 1

the controller performs wavefront composition of pieces of audio output from a plurality of speakers, controls a sound field of the agent device

Methodology Applied
Scientific EffectWavefront composition:

Data Source

PatentUS11234094B2Information processing device, information processing method, and information processing system
Publication Date: 2022.01.25 SONY GROUP CORP
  • US11234094B2 patent drawing
  • US11234094B2 patent drawing
  • US11234094B2 patent drawing

AI summary

Provided is an information processing device that processes a dialogue of an audio agent.The information processing device includes an acquisition unit that acquires audio information of an agent device that is played back through interaction with a user and audio information of other contents different from the audio information of the agent device, and a controller that performs sound field control processing on an audio output signal based on the audio information of the agent device acquired by the acquisition unit. The controller performs wavefront composition of pieces of audio output from a plurality of speakers, controls a sound field of the agent device, and avoids mixing with the other contents different from the audio information of the agent device.