AR Navigation Using Vision-Language Models for Dynamic Retail

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems struggle to provide real-time, accurate navigation and inventory management within dynamic retail environments, such as busy stores where shelves and aisles frequently change.

Innovation Solution

The system utilizes Augmented Reality (AR) and a Vision-and-Language Model (VLM) with multi-modal Artificial Intelligence (AI) to automatically generate in-store product information and navigation guidance. It leverages crowd-sourced images from end-user devices to create up-to-date planograms and inventory maps, enabling real-time localization and navigation within the store.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional inventory systems are used in dynamic retail environments, then system simplicity is maintained, but real-time accuracy of product information and inventory data becomes insufficient due to frequent changes in shelves and aisles

Engineering Contradiction:
Improvereal-time accuracy of product informationVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical inventory tracking systems with a vision-based system using smartphones and vision-and-language models. The system captures images of shelves and products, automatically identifies products and their locations through AI processing, and maintains real-time inventory data without requiring complex mechanical sensors or manual updates.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service inventory management by allowing customers to use their own smartphones to capture images of products and shelves. The vision-and-language model automatically processes these images to update inventory records, eliminating the need for dedicated inventory staff or complex automated scanning systems.

Inventive Principle:
Principle #25Self-service

2Ease of operation

If real-time navigation guidance is provided in dynamic retail environments, then customer experience is improved, but system complexity increases due to need for continuous updates in changing environments

Engineering Contradiction:
Improvecustomer navigation experienceVSAvoidnavigation system complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system continuously captures images of the retail environment using customer smartphones and feeds this visual feedback into vision-and-language models to update the spatial map and product location database in real-time. This feedback loop enables the navigation system to adapt to changes in shelf arrangements and product placements automatically.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent makes the customer's smartphone a multi-functional device that serves both as a navigation tool and as an inventory tracking system. The same phone used for navigation also captures images for product identification and location mapping, eliminating the need for separate specialized devices.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If crowd-sourced image collection is used to maintain up-to-date inventory maps, then real-time accuracy is improved, but data processing volume and system complexity increase

Engineering Contradiction:
Improveaccuracy of inventory mappingVSAvoidvolume of image data
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system extracts only the essential information needed for inventory mapping from the captured images using vision-and-language models. Instead of processing complete image datasets, the AI models identify and extract product names, locations, and spatial relationships, reducing the effective data volume to manageable levels while maintaining high accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The vision-and-language model acts as an intermediary between the raw image data from smartphones and the inventory database. This intermediary layer processes and interprets the visual information, transforming unstructured images into structured inventory records, thereby managing data complexity effectively.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250029170A1Automatic Generation of In-Store Product Information and Navigation Guidance, Using Augmented Reality (AR) and a Vision-and-Language Model (VLM) and Multi-Modal Artificial Intelligence (AI)
Publication Date: 2025.01.23 WE R AUGMENTED REALITY CLOUD LTD
  • US20250029170A1 patent drawing
  • US20250029170A1 patent drawing
  • US20250029170A1 patent drawing

AI summary

Automatic generation of in-store product information and navigation guidance, using Augmented Reality (AR) and a Vision-and-Language Model (VLM) and multi-modal Artificial Intelligence (AI). An automated method includes: providing to the VLM images that are captured within a retailer venue by an electronic device that is a smartphone or an AR device or smart glasses; (b) automatically feeding into the VLM those images, or pre-sliced or pre-cropped image-portions of those images that were sliced or cropped using Machine Learning that performs object boundaries detection and not product recognition; (c) invoking the VLM to generate outputs of VLM analysis of content of those fed images or image-portions. The VLM outputs can be: VLM-based product recognition, VLM-based product-related information, VLM-generated virtual shopping assistance, VLM-generated in-store navigation guidance, or other VLM-generated outputs. Based on the VLM-generated outputs, the electronic device provides real-time information about products depicted in the images, VLM-generated shopping assistance, and VLM-generated in-store navigation guidance.