Event-Based Pose Estimation Using Flicker Frequency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current 3D object tracking and six-degree of freedom (6-DoF) pose estimation methods face inaccuracies due to noise and insufficient image information, making it challenging to accurately estimate the pose of objects in 3D space using visual signals.
Innovation Solution
A pose estimation method utilizing an event-based vision sensor that captures light-emitting devices flickering at a predetermined frequency, generating an image frame sequence, extracting feature vectors through singular value decomposition (SVD), and applying these features to a deep neural network (DNN) model to estimate the object's pose.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional 3D object tracking and pose estimation methods are used, then the process can be completed with standard imaging, but measurement precision deteriorates due to noise and insufficient image information
Solution Approach 1:
The patent applies periodic action by using light-emitting devices that flicker at a predetermined frequency and detecting polarity changes in event streams at these periodic intervals. This periodic flickering creates distinct temporal signatures that enable accurate identification of target pixels corresponding to light-emitting devices, thereby improving pose estimation accuracy even with sparse image information from event-based sensors.
2Loss of information
If all pixels in the event stream are processed, then complete image information is obtained, but device complexity increases due to processing unnecessary background pixels
Solution Approach 1:
The patent extracts only the necessary subset of pixels from the event stream by identifying target pixels whose polarity change periods correspond to the predetermined flicker frequency of light-emitting devices. This extraction process filters out background pixels that do not exhibit the characteristic periodic polarity changes, thereby reducing processing complexity while maintaining the essential information needed for accurate pose estimation.
3Measurement precision
If high-resolution image frames are generated from event streams, then image quality improves, but processing time increases due to the complex reconstruction process
Solution Approach 1:
The patent takes out only the essential polarized event data corresponding to target pixels with polarity changes matching the light-emitting device flicker frequency. By extracting and processing only this relevant subset of event data rather than reconstructing complete high-resolution image frames from all events, the system achieves sufficient image quality for pose estimation while significantly reducing processing time and computational burden.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enables accurate 6-DoF pose estimation even with sparse image information, improving tracking precision by distinguishing flicker events from noise and reducing computational complexity through feature compression.
Implementation Method 1
obtaining an event stream from an event-based vision sensor configured to capture a target object to which light-emitting devices flickering at a predetermined first frequency are attached
Data Source
AI summary
A pose estimation method includes obtains an event stream from an event-based vision sensor configured to capture a target object to which light-emitting devices flickering at a predetermined first frequency are attached, obtains a polarity change period of at least one pixel based on the event stream, generates an image frame sequence using at least one target pixel having a polarity change period corresponding to the first frequency, among the at least one pixel, extracts a feature sequence including feature vectors corresponding to the at least one target pixel, from the image frame sequence, and estimates a pose sequence of the target object by applying the feature sequence to a deep neural network (DNN) model.


