Virtual Viewpoint Generation Using Scene Type and Time Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques for generating virtual viewpoint content from multiple camera viewpoints face challenges in determining the correct time and position, leading to difficulties in identifying scenes and enhancing user convenience, particularly in applications like soccer games where quick scene recognition is necessary.

Innovation Solution

An information processing apparatus that generates scene information including type and time data, and provides virtual viewpoint auxiliary information such as position, orientation, and moving speed settings, allowing for efficient generation and operation of virtual viewpoint video images without the need for continuous user adjustment.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If virtual viewpoint content is generated from multiple camera viewpoints, then the ability to view content from any virtual viewpoint is improved, but the difficulty in determining the correct time and position increases

Engineering Contradiction:
Improvevirtual viewpoint flexibilityVSAvoidscene identification complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system performs preliminary analysis of captured images from multiple cameras to generate scene information including type information (identifying key events like goals, fouls, or interesting moments) and time information (temporal location of events) before virtual viewpoint rendering. This pre-processing creates a structured database that guides subsequent viewpoint selection, eliminating the need for complex real-time analysis during virtual viewpoint generation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces scene information (comprising type information and time information) as an intermediary layer between the raw multi-camera images and the virtual viewpoint content. This intermediary structure acts as a mediator that translates complex multi-viewpoint data into simplified guidance for viewpoint determination, making the system more manageable and less complex.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If virtual viewpoint content is generated without pre-defined settings, then the customization option is improved, but the time and effort required for generation increases

Engineering Contradiction:
Improveviewpoint customizationVSAvoidviewpoint setup time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs preliminary analysis of captured images from multiple cameras to generate scene information including type information (identifying key events like goals, fouls, or interesting moments) and time information (temporal location of events) before virtual viewpoint rendering. This pre-processing creates a structured database that guides subsequent viewpoint selection, eliminating the need for complex real-time analysis during virtual viewpoint generation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses generated scene information as feedback to automatically determine appropriate virtual viewpoint positions and orientations. The type information and time information feed back into the viewpoint selection process, enabling the system to automatically adjust viewpoints based on identified scenes without requiring manual user configuration for each new content generation task.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11734931B2Information processing apparatus, information processing method, and storage medium
Publication Date: 2023.08.22 CANON KK
  • US11734931B2 patent drawing
  • US11734931B2 patent drawing
  • US11734931B2 patent drawing

AI summary

An information processing apparatus that provides information about a virtual viewpoint image includes: a generation unit configured to generate scene information including type information and time information, the type information indicating a type of an event occurring in an image-capturing region in which an image is captured by a plurality of cameras, the time information indicating a time when the event has occurred; and a provision unit configured to provide an output destination of material data with the scene information generated by the generation unit, the material data being generated from a plurality of captured images obtained by the plurality of cameras capturing images of the image-capturing region from different directions, the material data being used to generate the virtual viewpoint image depending on a position and an orientation of a virtual viewpoint.