Multi-Frame Depth Cost Volume for Real-Time Moving Object Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing moving object detection techniques, such as Siam, are not suitable for real-time applications like automated driving systems due to their high computational requirements, necessitating a lightweight model for high-speed operation while maintaining accuracy.
Innovation Solution
A detection apparatus using a lightweight model that generates features from multiple frames, calculates a cost volume for depth likelihood, and detects moving objects based on depth direction value distributions through three-dimensional convolutional networks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing moving object detection techniques (e.g., Siam) are used, then detection accuracy is improved, but computational load increases making real-time operation impossible
Solution Approach 1:
The patent extracts and utilizes only the essential depth information from multiple frames through cost volume calculation, rather than processing complete high-dimensional feature maps. This selective extraction of critical depth likelihoods reduces computational load while preserving detection accuracy.
Solution Approach 2:
The patent transforms the detection problem by adding a depth dimension through cost volume computation, organizing matching likelihoods across multiple frames in a three-dimensional space (width, height, depth). This dimensional transformation enables efficient processing by structuring data to exploit spatial and temporal correlations.
2Measurement precision
If existing moving object detection techniques (e.g., Siam) are used, then detection accuracy is improved, but device complexity increases
Solution Approach 1:
The patent segments the detection process into distinct functional modules: feature extraction from individual frames, cost volume computation for depth likelihood, and final detection based on depth distribution analysis. This segmentation allows each module to be optimized independently, reducing overall system complexity.
Solution Approach 2:
By organizing depth information in a structured three-dimensional cost volume, the patent creates a more manageable data representation that simplifies subsequent processing. The explicit depth dimension allows for efficient aggregation and analysis operations compared to handling unstructured multi-frame feature data.
Data Source
AI summary
An apparatus for detecting a moving object is provided. The apparatus generates a feature of each image in a plurality of frames. The apparatus generates a cost volume indicating likelihood of matching between the plurality of frames for each depth based on the feature of the image. The apparatus generates a feature indicating a distribution of values in a depth direction in the cost volume. The apparatus detects a moving object in the image based on the feature indicating the distribution of values in the depth direction.


