Mirror-Assisted Depth Detection for Occluded Scene Capture

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies for providing immersive experiences in captured video content, such as movies or sports events, face challenges in allowing users to freely navigate viewpoints due to limitations in content capturing systems, including occlusions, image resolution, and camera calibration, leading to discomfort and reduced immersion in VR contexts.

Innovation Solution

The implementation of a volumetric approach in free viewpoint content generation, which involves capturing and processing three-dimensional data to enable users to explore environments freely, using multiple cameras and advanced processing techniques like image fusion and depth estimation to overcome occlusions and provide complete environmental information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple cameras are used to capture content from different viewpoints, then the ability to provide free viewpoint navigation is improved, but the complexity of the capturing system and difficulty of camera calibration increase

Engineering Contradiction:
Improveviewpoint navigation freedomVSAvoidcapturing system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the capturing system into two distinct parts: a simplified camera array for capturing content from multiple viewpoints, and a separate mirror arrangement for capturing occluded regions. This segmentation allows each subsystem to be optimized independently, reducing overall system complexity while maintaining the ability to provide free viewpoint navigation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces mirrors as intermediary devices that redirect light paths to capture occluded regions without requiring additional complex camera positioning or movement. The mirrors act as mediators between the limited camera viewpoints and the full 360-degree environment, enabling comprehensive capture with a simpler camera arrangement.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If mirrors are added to capture occluded regions, then the completeness of captured content is improved, but the complexity of the capturing system increases

Engineering Contradiction:
Improveoccluded region coverageVSAvoidcapturing system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent introduces mirrors as intermediary devices that redirect light paths to capture occluded regions without requiring additional complex camera positioning or movement. The mirrors act as mediators between the limited camera viewpoints and the full 360-degree environment, enabling comprehensive capture with a simpler camera arrangement.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent uses mirrors to create optical copies of occluded regions, allowing the camera to capture images of areas that would otherwise be invisible directly. This copying approach through reflection enables complete environmental capture without physically moving cameras into occluded spaces or adding complex mechanical systems.

Inventive Principle:
Principle #26Copying

3Device complexity

If pre-defined viewpoints are used in 3D video, then the system complexity is reduced, but user immersion and comfort decrease

Engineering Contradiction:
Improveviewpoint system complexityVSAvoiduser immersion quality
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent transitions from static pre-defined viewpoints to dynamic free viewpoint navigation, where the viewpoint can continuously change based on user movement and orientation. This dynamic approach maintains system simplicity while dramatically improving user immersion by allowing natural head movements and exploration of the virtual environment.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent creates a universal viewpoint system that works for all users and all viewing scenarios within the captured environment. Instead of providing separate pre-defined viewpoints for different positions, the system enables any user to navigate to any viewpoint freely, making the system adaptable to diverse user preferences and movement patterns while maintaining immersion.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enhances user immersion by allowing seamless viewpoint changes and exploration within virtual environments, improving the quality and realism of the experience by generating a comprehensive three-dimensional representation of the scene, even in the presence of occlusions.

Implementation Method 1

a depth detector configured to capture depth representations of the scene

Methodology Applied
Scientific EffectTime of Flight: Time of Flight

Implementation Method 2

providing a mirror arranged to reflect at least some of the non-visible signal emitted by the emitter to one or more features within the scene that would otherwise be occluded by the user and to reflect light from the one or more features

Methodology Applied
Scientific EffectReflection: Reflection

Data Source

PatentEP3742396B1Image processing
Publication Date: 2024.01.31 SONY INTERACTIVE ENTERTAINMENT LLC
  • EP3742396B1 patent drawingFigure 1~2
  • EP3742396B1 patent drawingFigure 3~4a
  • EP3742396B1 patent drawingFigure 4b~5

AI summary

Apparatus comprises a camera configured to capture images of a user in a scene; a depth detector configured to capture depth representations of the scene, the depth detector comprising an emitter configured to emit a non-visible signal; a mirror arranged to reflect at least some of the non-visible signal emitted by the emitter to one or more features within the scene that would otherwise be occluded by the user and to reflect light from the one or more features to the camera; a pose detector configured to detect a position and orientation of the mirror relative to at least one of the camera and depth detector; and a scene generator configured to generate a three-dimensional representation of the scene in dependence on the images captured by the camera and the depth representations captured by the depth detector and the pose of the mirror detected by the pose detector.