Dynamic Sound Localization for Mixed Audio Streams

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies do not adequately enhance the user experience during live or video distributions that include sound, particularly in scenarios where multiple users are simultaneously listening to content and conversing, as they struggle to clearly distinguish between sound sources.

Innovation Solution

An information processing apparatus and method that analyzes both content data and user situation data to output sound control information, allowing for dynamic control of sound image localization of voices and content sounds on user terminals, improving the overall viewing experience by optimizing sound output based on the content and user interactions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If users simultaneously listen to content sound and talk voice, then the user can enjoy content while communicating with other users, but the user cannot clearly distinguish between different sound sources

Engineering Contradiction:
Improvecapability to enjoy content and communicate simultaneouslyVSAvoidsound source discrimination accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent segments the mixed audio signal into separate content sound and talk voice components using spatial separation processing. By dividing the audio signal processing into distinct channels with different localization characteristics, the system enables users to clearly distinguish between content sound and talk voice while simultaneously enjoying both, thus resolving the contradiction between simultaneous enjoyment capability and sound source discrimination accuracy.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If sound control information is output based on analysis of content data and user situation data, then the sound image localization can be dynamically optimized, but the processing complexity increases

Engineering Contradiction:
Improvesound image localization accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system employs automatic analysis of content data and user situation data to generate sound control information without requiring manual intervention. The processing apparatus autonomously performs spatial separation processing and adjusts sound image localization based on the analyzed data, reducing the need for complex manual configuration while achieving optimized sound localization accuracy.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If spatial separation processing is performed on audio content and talk voice, then the call sound can be heard clearly, but the overall viewing experience may be compromised

Engineering Contradiction:
Improvetalk voice clarityVSAvoidviewing experience quality
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies different sound control characteristics to different sound sources based on their spatial properties. Talk voice is localized with higher clarity for communication purposes, while content sound maintains its original audio quality for viewing experience. This local differentiation of sound quality allows clear talk voice perception without compromising the overall viewing experience.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20250008195A1Information processing apparatus, information processing method, and program
Publication Date: 2025.01.02 SONY GROUP CORP
  • US20250008195A1 patent drawing
  • US20250008195A1 patent drawing
  • US20250008195A1 patent drawing

AI summary

[Problem] Provided is a new and improved information processing apparatus capable of further improving a viewing experience of a user in a content including a sound. [Solution] An information processing apparatus includes an information output unit configured to output sound control information on a basis of an analysis result of first time-series data included in content data and an analysis result of second time-series data indicating a situation of a user, in which the sound control information includes information for controlling sound image localization of a voice of another user output to a user terminal used by the user or a sound included in the content data.