Calibrated Mobile Vision Task Generation for Product Arrangement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Retailers face challenges in arranging products in stores to optimize sales, enhance shopping experiences, and manage inventory efficiently, which existing technologies like RFID tags and smart shelving cannot fully address, and these solutions are costly due to hardware investments.
Innovation Solution
A computer-implemented method using a mobile electronic device calibrated in a virtual representation of a store, employing machine learning to identify actions and locations within captured images, and generating actionable tasks with AR guidance for product arrangement without costly hardware installations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If RFID tags and smart shelving are implemented to monitor product levels and automate reordering, then inventory management efficiency is improved, but hardware cost and system complexity increase significantly
Solution Approach 1:
The patent uses image capture devices to create visual copies of physical inventory instead of using RFID tags on each product. The system captures images of shelves and products, then uses ML models to analyze these images and extract inventory information, replacing the need for physical RFID tags and smart shelving hardware with software-based image analysis
Solution Approach 2:
The patent replaces the mechanical/electronic RFID reading system with an optical system using standard camera devices. Instead of electromagnetic field-based RFID communication, the system uses image capture and ML-based visual analysis to detect product levels and generate reordering tasks
2Reliability
If RFID tags and smart shelving are deployed throughout the store, then product monitoring capability is improved, but implementation cost increases due to hardware investment
Solution Approach 1:
The patent makes the image capture device serve multiple functions: capturing images for inventory monitoring, providing spatial context through pose estimation, and enabling task generation all in one device. This eliminates the need for separate RFID readers, tags, and smart shelving components, reducing implementation cost while maintaining monitoring capability
Solution Approach 2:
The system uses the mobile device's own camera and sensors to perform monitoring tasks without requiring external specialized hardware. The device captures images, determines its pose, and the server processes this information to generate tasks, making the monitoring system self-sufficient and cost-effective
3Measurement precision
If a virtual representation with calibrated mobile device pose is used to generate spatially accurate tasks, then task positioning precision is improved, but system complexity and processing requirements increase
Solution Approach 1:
The patent introduces a virtual representation of the store space as an intermediary between the physical mobile device and the task generation process. The calibrated pose in virtual space serves as a mediator that translates real-world device position into accurate spatial coordinates for task placement, simplifying the mapping process while maintaining precision
Solution Approach 2:
The system performs preliminary calibration of the mobile device pose in the virtual representation before generating tasks. By pre-establishing the device's position and orientation in virtual space, the system prepares accurate spatial context in advance, enabling precise task positioning without complex real-time calculations during task generation
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The present invention relates to a computer implemented method for generating a task to be performed in a real space. The method comprising: obtaining i) a first image captured by a camera of a first mobile electronic device being calibrated in a virtual representation of the real space such that the first mobile electronic device has a known pose in the virtual representation of the real space and ii) an associated first pose of the first mobile electronic device at a moment of capturing the first image; inputting the first image to a machine learning model trained to output an action based on context in an image and a location in the image associated with the action, thereby obtaining i) an action to be performed and ii) a location within the first image associated with the action; for the action to be performed, determining a position of the action to be performed within the virtual representation of the real space as an intersection between a known structure of the virtual representation of the real space and a raycast from the first pose against a screen space coordinate of the location associated with the action to be performed in the first image; and generating a task to be performed, the task comprising the action to be performed and its position within the virtual representation of the real space.