3D Spatial Audio Reproduction for Video Conferencing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video conference systems fail to consider the distance between communication partners, leading to an unrealistic sense of presence and inconvenient communication experiences.

Innovation Solution

An information processing apparatus and method that recreates a virtual three-dimensional space to control sound output based on the separation distance between communication partners, allowing for a more comfortable and realistic communication experience by adjusting sound levels and visuals accordingly.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If video conference systems allow real-time communication between distant locations, then communication connectivity is improved, but the sense of spatial distance and presence is lost

Engineering Contradiction:
Improvecommunication speedVSAvoidspatial distance information
Core Design Contradiction:
SpeedVSLoss of information

Solution Approach 1:

The patent applies 3D spatial positioning to audio output, transforming 2D flat audio reproduction into 3D spatial audio. By calculating separation distances between communication partners in a virtual three-dimensional space and adjusting audio output accordingly, the system restores the perception of spatial distance that was lost in traditional video conferencing.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Loss of information

If content sharing is enabled during communication, then information exchange is improved, but privacy protection becomes more difficult

Engineering Contradiction:
Improveinformation exchange efficiencyVSAvoidprivacy invasion
Core Design Contradiction:
Loss of informationVSObject-affected harmful factors

Solution Approach 1:

The patent applies different output values to different sound source types based on their spatial characteristics. By selectively adjusting the audio output for different communication partners according to their separation distances and spatial positions, the system enables differentiated information exchange while maintaining privacy protection for appropriate content.

Inventive Principle:
Principle #3Local quality

3Productivity

If automatic call connection is implemented, then communication efficiency is improved, but user convenience deteriorates due to inconvenient timing

Engineering Contradiction:
Improvecommunication efficiencyVSAvoidcommunication comfort
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent uses state information of users as feedback to dynamically adjust communication connection timing. By monitoring user states and using this information to determine appropriate call timing, the system avoids inconvenient automatic calls while maintaining high communication efficiency through context-aware connection management.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10834359B2Information processing apparatus, information processing method, and program
Publication Date: 2020.11.10 SONY GROUP CORP
  • US10834359B2 patent drawing
  • US10834359B2 patent drawing
  • US10834359B2 patent drawing

AI summary

To provide an information processing apparatus, an information processing method, and a program capable of aurally producing distance in a virtual three-dimensional space by using the space for a connection to a communication partner, and realizing more comfortable communication. An information processing apparatus including: a reception unit configured to receive data from a communication destination; and a reproduction control unit configured to perform control such that sound data of a space of the communication destination is reproduced from a sound output unit in a space of a communication source with an output value in accordance with separation distance between the communication destination and the communication source disposed in a virtual three-dimensional space, the output value being different for each sound source type.