Virtual Viewpoint Image Generation Using Scene Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users face difficulties in setting appropriate viewpoints for virtual viewpoint images, especially in events like sports and concerts, as they require time and expertise to select optimal viewing positions, which can be a barrier for those not familiar with the process.
Innovation Solution
An image processing system that includes a storage device, an image processing device, and a user terminal, which generates and adjusts virtual viewpoint images by analyzing multiple camera shots, allowing users to select scenes and composition scenarios to automatically set virtual camera paths, enabling users to easily choose desired viewpoints without extensive configuration.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If users manually set viewpoints for virtual viewpoint images, then the viewing experience can be optimized for specific scenes, but the operation becomes complex and time-consuming for users not familiar with the process
Solution Approach 1:
The system automatically analyzes the captured image to identify the main subject and determine an appropriate viewpoint without requiring user configuration. The viewpoint setting unit autonomously performs subject detection, position determination, and camera parameter calculation based on the identified subject, enabling the system to serve itself rather than requiring user intervention for viewpoint selection
Solution Approach 2:
The system pre-calculates and stores multiple candidate viewpoints around the captured scene before user viewing. By preparing various viewpoint options in advance based on the main subject position and scene characteristics, the system eliminates the need for users to perform complex real-time viewpoint configuration, presenting them with pre-processed viewpoint options instead
2Adaptability or versatility
If multiple cameras are used to capture events from different positions, then comprehensive coverage of the event is achieved, but the complexity of processing and generating virtual viewpoint images increases
Solution Approach 1:
The system divides the complex task of virtual viewpoint generation into distinct functional modules: image acquisition from multiple cameras, main subject detection, position information extraction, viewpoint determination, and image synthesis. This segmentation allows each module to handle a specific aspect of the process independently, reducing overall system complexity while maintaining comprehensive event coverage
Solution Approach 2:
The system introduces an intermediary processing layer that automatically extracts position information of the main subject from multiple camera feeds and uses this information to determine optimal viewpoint settings. This intermediary layer acts as a mediator between the raw multi-camera input and the final virtual viewpoint output, simplifying the processing pipeline by centralizing the decision-making logic for viewpoint selection
Data Source
Figure 1
Figure 2A~2C
Figure 3
AI summary
An information processing device (300) decides a viewpoint position and generates a virtual viewpoint image based on the decided viewpoint position by using a plurality of images shot by a plurality of imaging apparatuses. The information processing device (300) includes determining means arranged to determine a scene related to the virtual viewpoint image to be generated, and deciding means arranged to decide the viewpoint position related to the virtual viewpoint image in the scene determined by the determining means, based on the scene determined by the determining means.