Voice-Controlled AR Product Presentation for Faster Object Placement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current augmented reality (AR) and virtual reality (VR) systems for home decor preview are cumbersome and time-consuming, requiring users to navigate complex virtual environments to place objects, which is inefficient.

Innovation Solution

A system that uses natural-language processing and computer vision to allow users to navigate and place virtual objects in AR/VR environments through voice commands and image recognition, generating and situating objects in real-world spaces based on spatial coordinates and object recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If users navigate complex virtual environments to place objects using traditional AR/VR systems, then object placement functionality is achieved, but user effort and time consumption increase significantly

Engineering Contradiction:
Improveease of object placementVSAvoidtime for navigation and placement
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent replaces manual navigation and dragging operations with voice command control. Users speak natural language instructions like 'place the sofa in the living room' instead of manually navigating virtual environments and positioning objects, substituting mechanical interaction with speech-based control.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system introduces an intermediate processing layer that translates voice commands into object placement actions. This intermediary layer interprets natural language input, determines spatial relationships, and executes placement without requiring users to directly manipulate virtual objects, thereby simplifying the interaction process.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If traditional AR/VR interfaces are used for product preview, then virtual object placement is enabled, but system complexity and user learning curve increase

Engineering Contradiction:
Improveproduct preview capabilityVSAvoidinterface complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system automatically processes voice commands and performs object placement without requiring users to understand complex interface operations. The AI interpreter handles spatial reasoning and placement decisions autonomously, making the system self-sufficient and reducing the cognitive load on users.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The voice-based interface serves multiple functions: navigating virtual environments, selecting objects, determining placement locations, and adjusting object orientations—all through a single unified command structure. This multi-functional approach eliminates the need for separate controls for each operation.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250316034A1Augmented reality enabled dynamic product presentation
Publication Date: 2025.10.09 SHOPIFY INC
  • US20250316034A1 patent drawing
  • US20250316034A1 patent drawing
  • US20250316034A1 patent drawing

AI summary

Systems and methods described herein allow a customer to employ AR/VR software to generate virtual representations of physical spaces (e.g., house) and sub-spaces (e.g., living room) to preview virtual objects situated in AR/VR virtual environments. A commerce system (or mobile app associated with the commerce system) may generate virtualized environments representing a physical space (e.g., house, apartment) and regions (e.g., living room, kitchen) based on source images uploaded to or otherwise captured by the commerce system. The end-user may operate the software on a client device and interacts with VR or AR presentations of the virtual environment using a voice-based interface recognized by the software. For example, the end-user may say the name of room (region) or an object and the system retrieves data of the identified room or an appropriate room, such as virtual representations of furniture or objects situated in the room.