6DoF Media Guide Voice Generation via Viewpoint Adaptation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In 6DoF media distribution, generating an appropriate guide voice that explains the content of a scene viewed by the user is challenging due to the user's ability to freely move and change their viewpoint, making it difficult to prepare voice commentary in advance.

Innovation Solution

A media distribution device and method that include a guide voice generation unit, which uses scene descriptions and user viewpoint information to generate a guide voice describing the rendered image viewed from a specific viewpoint in a virtual space, and an audio encoding unit that mixes and encodes the guide voice with original audio.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If a guide voice is prepared in advance for 6DoF media distribution, then the cost is reduced and preparation is simpler, but the guide voice cannot be appropriately matched to the user's current viewpoint and scene content

Engineering Contradiction:
Improveguide voice preparationVSAvoidguide voice adaptability to user viewpoint
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The system performs preliminary actions by pre-generating guide voices for multiple possible viewpoints and scene configurations before the user actually views them. These pre-generated guide voices are stored and then selectively retrieved based on the user's actual viewpoint, combining the benefits of advance preparation with viewpoint-specific customization.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system changes parameters by generating different guide voices based on varying viewpoint parameters (position, direction, distance) and scene parameters (visible objects, lighting conditions). By dynamically adjusting the guide voice content according to these parameter changes, the system achieves adaptability while maintaining efficient pre-processing.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If dedicated guide voices are created for each of a large number of positions and directions, then the guide voice appropriateness is improved, but the cost becomes enormous and the process is not realistic

Engineering Contradiction:
Improveguide voice viewpoint matching accuracyVSAvoidguide voice production cost
Core Design Contradiction:
Manufacturing precisionVSEase of manufacture

Solution Approach 1:

The system segments the continuous 6DoF viewpoint space into a finite number of discrete viewpoint categories (e.g., front, back, left, right, and intermediate positions). Guide voices are created for each segment rather than for every possible position, reducing the total number of guide voices needed while maintaining sufficient viewpoint matching accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system creates guide voices that serve multiple viewpoints and scene configurations simultaneously. A single guide voice can be applicable to multiple similar viewpoints or scene types, reducing the total number of unique guide voices required while maintaining appropriateness across diverse viewing conditions.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If guide voice generation is performed in real-time based on user viewpoint, then the guide voice relevance is improved, but the system complexity and processing requirements increase

Engineering Contradiction:
Improveguide voice relevance to current sceneVSAvoidguide voice generation system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system performs the complex computation of guide voice generation in advance for multiple possible viewpoints and scene configurations, storing the results for rapid retrieval during actual use. This shifts the computational complexity from real-time operation to offline pre-processing, reducing real-time system complexity while maintaining high adaptability.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12217366B2Media distribution device and media distribution method
Publication Date: 2025.02.04 SONY GROUP CORP
  • US12217366B2 patent drawing
  • US12217366B2 patent drawing
  • US12217366B2 patent drawing

AI summary

The present disclosure relates to a media distribution device, a media distribution method, and a program enabling to generate a guide voice more appropriately. A media distribution device includes a guide voice generation unit that generates a guide voice describing a rendered image viewed from a viewpoint in a virtual space by using a scene description as information describing a scene in the virtual space and a user viewpoint information indicating a position and a direction of the viewpoint of a user; and an audio encoding unit that mixes the guide voice with original audio, and encodes the guide voice. The present technology can be applied to, for example, a media distribution system that distributes 6DoF media.