Sparse Sensor Capture for Object Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional object tracking methods require capturing full images, which are computationally expensive and power-intensive, limiting their use in power-constrained devices like mobile and AR/VR devices due to high pixel data transfer and processing latency.
Innovation Solution
A computing system predicts the object's pose and camera pose, generating pixel-activation instructions to capture a subset of pixels based on a 3D model projection, reducing the number of pixels needed for image capture and processing, thereby minimizing power consumption and latency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If full images are captured for object tracking, then tracking accuracy is maintained, but power consumption increases and latency increases
Solution Approach 1:
The patent segments the image capture process by activating only specific regions of interest (ROIs) within the image sensor instead of capturing the entire image. This is achieved by dividing the sensor into multiple regions and selectively activating only those regions containing tracked objects, thereby reducing the number of pixels that need to be read out and processed, which directly lowers power consumption while maintaining tracking accuracy.
Solution Approach 2:
The patent applies local quality by assigning different capture qualities to different regions of the image sensor. Regions containing tracked objects are captured at full resolution and quality, while regions without objects are either captured at lower resolution or not captured at all. This selective quality approach ensures tracking accuracy is maintained for objects while reducing overall power consumption.
2Loss of information
If full images are captured for object tracking, then complete object information is obtained, but processing time increases
Solution Approach 1:
The patent segments the image data processing by reading out only the pixel data from activated regions of interest instead of the entire image sensor. This segmentation of data transfer and processing significantly reduces the amount of data that needs to be handled, thereby reducing processing latency while ensuring that all necessary object information is captured within the ROIs.
3Loss of information
If full images are captured for object tracking, then all pixel data is available for analysis, but data transfer volume increases
Solution Approach 1:
The patent segments the pixel data output by configuring the image sensor to output data only from selectively activated regions of interest. This is achieved by controlling the sensor's readout circuitry to transfer only the pixel values from ROIs to the processing system, dramatically reducing the volume of pixel data that needs to be transferred and processed while maintaining availability of all relevant object information.
Data Source
AI summary
In one embodiment, a computing system instructs, at a first time, a camera having a plurality of pixel sensors to use the plurality of pixel sensors to capture a first image of an environment comprising an object. The computing system predicts, using at least the first image, a projection of the object appearing in a virtual image plane associated with a predicted camera pose at a second time. The computing system determines, based on the predicted projection of the object, a first region of pixels and a second region of pixels. The computing system generates pixel-activation instructions for the first region of pixels and the second region of pixels. The computing system instructs the camera to capture a second image of the environment at the second time according to the pixel-activation instructions. The pixel-activation instructions are configured to cause a first subset of the plurality of pixel sensors to sample the first region of pixels and a second subset of the plurality of pixel sensors to sample the second region of pixels. The first subset of the plurality of pixel sensors used for sampling the first region of pixels is more dense than the second subset of the plurality of pixel sensors used for sampling the second region of pixels. The computing system tracks, based on the second image, the object at the second time.


