Camera-Based Target Motion Detection With Depth Coordinate Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for estimating the motion velocity and direction of objects in unmanned driving, security, and scene understanding scenarios, such as LiDAR, require significant computational processing and fail to meet high real-time requirements due to the need for point cloud data acquisition and target tracking.
Innovation Solution
A method and apparatus using computer vision technology to capture and analyze image sequences, perform target detection, determine depth information, and transform coordinates to calculate motion information without emitting high-frequency laser beams, thereby reducing computational load.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If LiDAR is used to collect point cloud data for target detection and tracking, then measurement precision of target position is improved, but processing time increases and productivity decreases
Solution Approach 1:
The patent replaces the LiDAR mechanical scanning system with a camera-based optical system. Instead of using laser beams to scan and collect point cloud data, the system uses image capturing devices to acquire image data, which is then processed to extract target position and motion information. This substitution eliminates the need for mechanical laser scanning while maintaining measurement capabilities through computational methods.
Solution Approach 2:
The patent creates a virtual copy of the physical scene through image capture and processing. Instead of directly measuring target position with LiDAR, the system captures images of the scene, processes these images to generate virtual representations (including depth map generation through neural networks), and extracts motion information from these virtual models. This copying approach enables efficient processing while maintaining measurement precision.
2Measurement precision
If LiDAR emits high-frequency laser beams to obtain point cloud data, then depth information accuracy is improved, but energy consumption increases
Solution Approach 1:
The patent replaces the high-energy LiDAR laser emission system with a low-energy camera-based optical capture system. The image capturing device uses ambient light rather than emitting high-frequency laser beams, dramatically reducing energy consumption while obtaining sufficient depth and position information through image processing and neural network-based depth estimation.
Solution Approach 2:
The patent enables the system to derive depth information autonomously from captured images without requiring active illumination. The neural network processes the captured image data to generate depth maps, allowing the system to self-service the depth estimation function using passive optical information rather than active laser illumination, thereby reducing energy consumption.
3Measurement precision
If target detection and tracking are performed on point cloud data from multiple time points, then motion information accuracy is improved, but device complexity increases
Solution Approach 1:
The patent replaces the complex point cloud data processing pipeline with a simplified image-based processing system. Instead of performing target detection and tracking on three-dimensional point cloud data from multiple time points, the system processes two-dimensional image sequences, which are computationally less intensive and require simpler processing algorithms while still enabling accurate motion information extraction.
Solution Approach 2:
The patent uses image-based virtual representations to simplify the complexity of physical point cloud processing. By creating and processing virtual image copies of the scene rather than manipulating complex three-dimensional point cloud structures, the system reduces computational complexity and device requirements while maintaining the ability to extract accurate motion information through frame-to-frame comparison.
Data Source
AI summary
Disclosed are a method for detecting motion information of a target, a device and memory. The method includes: performing target detection on a first image to obtain a detection box of a first target; acquiring depth information of first image in a corresponding first camera coordinate system and determining depth information of the detection box therefrom, and determining first coordinates of first target in first camera coordinate system based on a location of the detection box in an image coordinate system and the depth information thereof; transforming second coordinates of a second target in a second camera coordinate system corresponding to the second image into third coordinates in the first camera coordinate system based on pose change information of an image capturing device; and determining motion information of the first target based on the first and third coordinates. The disclosure avoids abundant computational processing and improves processing efficiency.


