360-Degree Immersive Video Overlays from Single Fisheye Images

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for overlaying content in 360-degree immersive videos lack the ability to accurately determine depth and orientation, leading to suboptimal user experiences, especially when objects in the scene are moving or when cameras are not constantly moving, and require complex setups.

Innovation Solution

An object-based 3D aware overlay method using ellipse-ellipsoid constraints for depth estimation and embedding loss functions, combined with convolutional neural networks for feature extraction and computer vision techniques, to determine accurate overlay positions from a single fisheye image.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If complex camera motion tracking setups are used to determine overlay positions, then overlay accuracy is improved, but device complexity and setup requirements worsen

Engineering Contradiction:
Improveoverlay position accuracyVSAvoidcamera motion tracking setup
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and removes the complex camera motion tracking requirement from the system. Instead of using exhaustive camera motion tracking setups, the invention uses a single fisheye image with ellipse-ellipsoid constraints to directly estimate depth and orientation, thereby achieving accurate overlay positioning without the complex tracking infrastructure.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces ellipse-ellipsoid constraints as an intermediary mathematical model between the single fisheye image and the overlay position determination. This intermediary model enables depth and orientation estimation from static images, bridging the gap between simple image capture and accurate 3D overlay positioning without requiring complex motion tracking.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If exhaustive camera motion tracking is used, then depth estimation accuracy is improved, but loss of time and processing complexity worsen

Engineering Contradiction:
Improvedepth estimation accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by establishing ellipse-ellipsoid constraints from a single static fisheye image before any motion tracking would occur. This preliminary depth and orientation estimation from the static image eliminates the need for time-consuming exhaustive camera motion tracking, achieving accurate depth estimation instantly without prolonged processing.

Inventive Principle:
Principle #10Preliminary action

3Ease of operation

If traditional 2D overlay methods are used, then ease of operation is improved, but immersion quality and realism worsen

Engineering Contradiction:
Improveoverlay implementation simplicityVSAvoidimmersion quality
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent transitions from traditional 2D overlay methods to 3D aware overlay by introducing depth estimation through ellipse-ellipsoid constraints. This dimensionality change from 2D to 3D enables realistic and immersive experiences while maintaining operational simplicity, as the system still processes single images but now generates accurate 3D overlay positioning.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentEP4303817B1A method and an apparatus for 360-degree immersive video
Publication Date: 2025.10.15 NOKIA TECHNOLOGIES OY
  • EP4303817B1 patent drawingFigure 1a~1b
  • EP4303817B1 patent drawingFigure 2a~2b
  • EP4303817B1 patent drawingFigure 3

AI summary

The embodiments relate to a method comprising receiving a corrected fisheye image; detecting one or more objects from the corrected fisheye image and indicating the one or more objects with a corresponding bounding volume; predicting an ellipse within each bounding volume; estimating a relative pose for detected objects based on corresponding ellipses; estimating an inverse depth for objects based on the relative poses; estimating object-based normal; and generating a placement for a three-dimensional aware overlay. The embodiments also relate to a technical equipment for implementing the method.