Video Generation Using Mesh and Texture Data for Viewpoint Interaction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Pre-generated video content limits user interaction and immersion, as viewers cannot easily change viewpoints or explore virtual environments, unlike interactive experiences that respond to their location and orientation.

Innovation Solution

Encoding video frames with mesh and texture data instead of fully-rendered images, allowing users to change viewpoints by transmitting data that describes the scene from different angles, enabling exploration and interaction within pre-generated content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If pre-generated video content is used, then production efficiency is improved, but user interaction capability deteriorates

Engineering Contradiction:
Improvecontent production efficiencyVSAvoidviewpoint change capability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The video content is segmented into two distinct data types: mesh data representing the three-dimensional scene structure and texture data representing visual appearance. This segmentation allows the scene to be stored in a structured format that enables later viewpoint changes without regenerating the entire video, thus maintaining production efficiency while enabling interaction.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from storing only two-dimensional rendered video frames to storing three-dimensional mesh data with associated texture data. This dimensional upgrade allows users to navigate and view the scene from multiple angles and viewpoints after production, solving the contradiction between efficient pre-generation and post-production interaction flexibility.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Manufacturing precision

If fully-rendered video frames are transmitted, then rendering quality is improved, but data transmission efficiency deteriorates

Engineering Contradiction:
Improverendering qualityVSAvoiddata transmission efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent extracts only the essential geometric and visual information needed to reconstruct scenes from different viewpoints, rather than transmitting complete rendered video frames. By separating mesh data (geometric structure) from texture data (visual appearance), the system transmits compact representations that can be rendered on-demand for various viewpoints, achieving both quality and efficiency.

Inventive Principle:
Principle #2Taking out (Extraction)

3Adaptability or versatility

If interactive viewpoint changes are enabled, then user immersion is improved, but computational complexity deteriorates

Engineering Contradiction:
Improveviewpoint interaction capabilityVSAvoidcomputational complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent performs the computationally intensive scene construction work in advance during content production, creating mesh and texture data representations. This preliminary action stores the scene in a format that requires minimal computation during playback, allowing interactive viewpoint changes to be achieved by simply re-rendering from stored data rather than generating new content in real-time.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11115644B2Video generation method and apparatus using mesh and texture data
Publication Date: 2021.09.07 SONY INTERACTIVE ENTERTAINMENT LLC
  • US11115644B2 patent drawing
  • US11115644B2 patent drawing
  • US11115644B2 patent drawing

AI summary

A display system for displaying video content, the display unit including a display controller operable to control the display of at least a region of the video content and a visual acuity identification unit operable to identify at least regions of high visual acuity and low visual acuity of a viewer at display positions corresponding to particular regions of the video content, where the display controller is operable to control the display of video content such that regions of high visual acuity are displayed differently to regions of low visual acuity.