Adaptable Stitching for Electronic Mirror Video Frames
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional electronic mirror systems for vehicles face challenges in stitching video feeds from multiple cameras, leading to blending lines that can cause objects to be missing in the output frame, particularly due to the high computational resource requirements and time sensitivity of stereo image processing for depth calculation.
Innovation Solution
An apparatus that uses a processor to detect objects, determine depth information from a single camera, adjust blending lines to prevent object omission, and generate panoramic video frames for a 3-in-1 electronic mirror display, employing a convolutional neural network for monocular depth estimation and adaptive stitching.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If stereo images are used to calculate depth information and stitch images, then object visibility is improved, but computational resource requirements increase and processing time increases
Solution Approach 1:
The patent replaces expensive stereo camera systems with a single camera equipped with a neural network model for depth estimation. This substitutes costly hardware (stereo cameras) with a more economical solution (single camera + software processing), achieving comparable depth information extraction without the high computational burden of processing two synchronized camera feeds
Solution Approach 2:
The patent replaces the mechanical/optical stereo vision system with a computational approach using neural networks. Instead of relying on physical baseline separation between two cameras, the system uses a trained neural network model to estimate depth from monocular images, substituting mechanical depth sensing with AI-based computational inference
2Reliability
If stereo images are used to calculate depth information and stitch images, then object visibility is improved, but processing time increases
Solution Approach 1:
The patent performs depth estimation using a pre-trained neural network model that has already learned depth relationships during training. This preliminary training phase allows the system to quickly infer depth from single images during runtime without requiring the time-consuming stereo matching process, achieving real-time performance suitable for e-mirror applications
Solution Approach 2:
The patent uses a single camera instead of two, reducing the amount of data that needs to be processed simultaneously. This single-image processing approach with neural network inference is computationally faster than synchronized stereo image processing, decreasing latency while maintaining depth accuracy
3Object-affected harmful factors
If blending lines are used to combine video feeds, then visual artifacts are reduced, but objects may be cropped out and become missing
Solution Approach 1:
The patent dynamically adjusts the blending line position based on detected object locations and depth information. Instead of using fixed blending lines, the system moves and adapts the blending boundaries to accommodate objects at different depths, ensuring objects remain visible while still achieving smooth transitions between camera views
Solution Approach 2:
The patent uses object detection and depth estimation results to provide feedback for adjusting blending line positions. The system continuously monitors object locations and modifies the blending strategy accordingly, creating a closed-loop system that prevents object cropping while maintaining visual artifact reduction
Data Source
AI summary
An apparatus including an interface and a processor. The interface may be configured to receive video frames generated by a plurality of capture devices. The processor may be configured to perform operations to detect objects in the video frames received from a first of the capture devices, determine depth information corresponding to the objects detected, determine blending lines in response to the depth information, perform video stitching operations on the video frames from the capture devices based on the blending lines and generate panoramic video frames in response to the video stitching operations. The blending lines may correspond to gaps in a field of view of the panoramic video frames. The blending lines may be determined to prevent the objects from being in the gaps in the field of view. The panoramic video frames may be generated to fit a size of a display.


