Peripheral Video Composition With 3D Occlusion Inference

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing peripheral video generation devices distort areas shielded by objects around a vehicle, making it difficult to intuitively grasp conditions as the proximity to the screen edges increases, leading to unnatural video appearances.

Innovation Solution

A peripheral video generation device that includes a video input unit, a video composition unit, a three-dimensional shape estimation unit, a shielded area estimation unit, an inference unit using deep learning, and a video superimposition unit to composite and superimpose inferred videos of shielded areas, utilizing techniques like Structure from Motion (SfM) and generative adversarial networks (GAN) to generate a natural-looking composite video.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If peripheral video data is composited to generate a composite video as viewed from a predetermined viewpoint, then the driver can recognize conditions around the vehicle, but areas shielded by objects become distorted and unnatural, especially near screen edges

Engineering Contradiction:
Improverecognition accuracyVSAvoidvideo distortion
Core Design Contradiction:
ReliabilityVSShape

Solution Approach 1:

The patent transitions from 2D video composite to 3D spatial understanding by estimating three-dimensional shapes of peripheral objects. This dimensional enhancement allows the system to calculate accurate shielded areas by considering the spatial relationships and volumes of objects, thereby reducing distortion in the composite video while maintaining recognition accuracy.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent introduces an inference unit that uses deep learning to generate inferred video of shielded areas. This intermediary component fills in the distorted regions by predicting what the shielded areas should look like based on learned patterns, thereby eliminating visual distortion while preserving the reliability of obstacle detection.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Shape

If deep learning is used to infer video of shielded areas, then the video appearance becomes more natural, but the device complexity increases

Engineering Contradiction:
Improvevideo naturalnessVSAvoidsystem complexity
Core Design Contradiction:
ShapeVSDevice complexity

Solution Approach 1:

The patent divides the video processing task into distinct segments: video composition, three-dimensional shape estimation, shielded area estimation, and inferred video generation. Each module handles a specific aspect of the problem, allowing the complex deep learning process to be applied only where necessary (in shielded areas) rather than to the entire video, thereby reducing overall system complexity while maintaining naturalness.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies deep learning-based video inference only to specific shielded areas rather than the entire composite video. By localizing the complex processing to only where needed (areas occluded by objects), the system achieves natural video appearance in critical regions while keeping the overall device complexity manageable through selective application of advanced algorithms.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12450784B2Peripheral video generation device, peripheral video generation method, and storage medium storing program
Publication Date: 2025.10.21 DENSO CORP
  • US12450784B2 patent drawing
  • US12450784B2 patent drawing
  • US12450784B2 patent drawing

AI summary

A peripheral video generation device includes: a video input unit that inputs peripheral video data captured by a plurality of cameras; a video composition unit that composites the peripheral video data to generate a composite video as viewed from a predetermined viewpoint; a three-dimensional shape estimation unit that estimates a three-dimensional shape of a peripheral object based on the peripheral video data; a shielded area estimation unit that uses an estimation result of the three-dimensional shape to estimate a shielded area not visible from the predetermined viewpoint in the composite video; an inference unit that infers a video of the shielded area using deep learning; and a video superimposition unit that superimposes the video inferred by the inference unit on the shielded area in the composite video.