Virtual Viewpoint Image Synthesis for Natural Line-of-Sight Alignment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Bidirectional communication systems, such as video conferencing, face issues where the line-of-sight direction of users displayed on a screen does not match their actual gaze direction, leading to a strange and unnatural viewing experience, especially when multiple users are present.
Innovation Solution
An information processing apparatus that generates virtual viewpoint images from multiple camera angles and combines them to create a single image displayed on a screen, ensuring that each user's line-of-sight direction matches their actual viewpoint, using a virtual viewpoint image generation unit and an image combining unit to adjust the image based on the relative position of viewers to the screen.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If a single camera viewpoint is used to photograph users, then the system is simple and cost-effective, but the line-of-sight direction of users displayed on the screen does not match their actual gaze direction, creating an unnatural viewing experience
Solution Approach 1:
The patent divides the single viewpoint into multiple virtual viewpoints corresponding to different viewer positions. By segmenting the viewing experience into multiple perspective images generated from one camera, the system maintains simplicity while providing natural viewing for multiple users at different positions.
Solution Approach 2:
The patent adds a spatial dimension by generating multiple virtual viewpoint images from a single camera viewpoint. This dimensional transformation allows different viewers to see images appropriate to their positions, resolving the contradiction between system simplicity and viewing naturalness.
2Ease of operation
If multiple cameras are used to capture different viewpoints, then the line-of-sight direction matches actual gaze direction for multiple viewers, but the system complexity and cost increase significantly
Solution Approach 1:
The patent creates virtual copies of camera viewpoints through image processing. Instead of using multiple physical cameras, the system generates multiple virtual viewpoint images from a single camera, achieving the effect of multiple perspectives without the complexity of multiple devices.
Solution Approach 2:
The patent replaces the mechanical system of multiple physical cameras with an optical/image processing system. By using image synthesis and viewpoint transformation algorithms, the system achieves multi-perspective functionality without the mechanical complexity of multiple camera installations.
3Ease of manufacture
If conventional display methods are used, then the implementation is simple and cost-effective, but users at different positions experience unnatural line-of-sight directions, reducing communication effectiveness
Solution Approach 1:
The patent makes the display dynamic by adapting the viewpoint image to each viewer's position. Instead of a static single-perspective display, the system dynamically selects and presents appropriate virtual viewpoint images based on viewer location, maintaining implementation simplicity while improving communication effectiveness.
4Ease of operation
If special displays are used to correct viewpoint directions, then the line-of-sight direction matches actual gaze direction, but the cost increases and conventional displays cannot be used
Solution Approach 1:
The patent uses virtual copying of viewpoint information through image processing rather than special display hardware. By synthesizing multiple viewpoint images computationally, the system achieves natural viewing experiences using conventional displays, avoiding the high cost of specialized display technology.
Data Source
AI summary
Achieving a configuration that reduces the artificiality to give a strange feeling about the viewpoint of the user displayed on the display unit appearing different from the actual viewpoint. Photographed images from a plurality of different viewpoints are input to generate a plurality of virtual viewpoint images, and then, the plurality of virtual viewpoint images is combined to generate a combined image to be output on a display unit. The virtual viewpoint image generation unit generates a plurality of user viewpoint-corresponding virtual viewpoint images each corresponding to each of viewpoints of each of a plurality of viewing users viewing the display unit, while the image combining unit extracts a portion from each of the plurality of user viewpoint-corresponding virtual viewpoint images in accordance with a relative position between the viewing user and the display unit, and combines the extracted image to generate a combined image. The combined image is generated by extracting a display region image located at a front position of the viewing user at the viewpoint corresponding to the virtual viewpoint image from among the user viewpoint-corresponding virtual viewpoint images corresponding to individual viewing users.


