Video-Based Customer Interaction Detection Using Pose Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods struggle to accurately predict customer demand for products when customer interactions do not result in purchases, limiting the ability to assess interaction with products in stores.
Innovation Solution
A method and apparatus that utilize frame images from videos to detect interactions by acquiring pose data from feature points, determining interaction occurrence through neural networks, estimating regions of interest, and extracting product information, with weighted feature points emphasizing arm and hand movements for accurate interaction detection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If purchase record data is used to predict customer demand, then the prediction is based on actual transaction data, but it cannot capture interactions that do not result in purchases
Solution Approach 1:
The patent replaces the traditional mechanical approach of using purchase records (physical transaction data) with an optical system using video cameras and image processing to detect customer interactions. This substitution enables capture of interaction behaviors without requiring actual purchases to occur.
Solution Approach 2:
The patent introduces video footage and image processing algorithms as intermediaries between the customer and the product interaction data. Instead of directly relying on purchase records, the system uses video frames and pose estimation technology to detect and analyze customer-product interactions, serving as a mediator that captures subtle interaction cues.
2Loss of information
If video analysis is used to detect interactions, then non-purchase interactions can be captured, but the system complexity increases significantly
Solution Approach 1:
The patent extracts only the essential features needed for interaction detection from video data, such as pose information of body parts and spatial relationships between customer and product. By extracting only relevant features rather than processing entire video frames in detail, the system reduces computational complexity while maintaining interaction detection capability.
Solution Approach 2:
The patent segments the interaction detection process into distinct stages: video frame acquisition, pose estimation, interaction determination, and product identification. This segmentation allows each component to be optimized independently and processed efficiently, reducing overall system complexity.
3Measurement precision
If detailed pose data from multiple feature points is collected, then interaction detection accuracy improves, but the data processing load increases
Solution Approach 1:
The patent applies different levels of analysis to different parts of the pose data. Rather than uniformly processing all feature points with equal detail, the system focuses computational resources on critical regions such as hands and upper body movements when detecting product interactions, while using coarser analysis for less relevant areas.
Data Source
AI summary
An interaction detection method is provided. The method includes the steps of: acquiring one or more frame images; acquiring pose data of a first object using information on a plurality of feature points detected for the first object from a first frame image; determining occurrence of an interaction of the first object using the pose data of the first object; estimating a region of interest (ROI) of the first object using the information on the plurality of feature points; and acquiring information on a product corresponding to the region of interest of the first object.


