3D Text Detection Interface for Accessible XR Interaction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods and interfaces for interacting with augmented and virtual reality environments are cumbersome, inefficient, and limited for users with visual, motor, hearing, and cognitive impairments, leading to increased cognitive burden and energy waste, particularly in battery-operated devices.

Innovation Solution

Implementing improved user interfaces that include magnifying regions, ray-based selection, bimanual navigation, guided access modes, audio-visual feedback for sound localization, automatic detection of textual content, and non-visual information provision to enhance interaction efficiency and accessibility for users with impairments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If conventional user interfaces are used in XR environments, then basic interaction functionality is provided, but users with visual, motor, hearing, and cognitive impairments experience increased cognitive burden and interaction difficulty

Engineering Contradiction:
ImproveAccessibility for users with impairmentsVSAvoidInterface complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The interface is segmented into multiple specialized modes (guided access mode, magnifying region, audio-visual feedback mode) that can be activated based on user needs. Each mode provides simplified interactions tailored to specific impairment types, breaking down the complex XR interface into manageable, accessible components without increasing overall system complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The XR system incorporates universal accessibility features that serve multiple functions: the magnifying region assists users with visual impairments while also providing focus enhancement for all users; audio-visual feedback serves both hearing-impaired users (visual) and provides confirmation for all users; guided access mode simplifies navigation for users with cognitive impairments while maintaining standard functionality for others.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If conventional interaction methods are used in XR environments, then standard functionality is maintained, but interaction efficiency is reduced and energy is wasted

Engineering Contradiction:
ImproveInteraction efficiencyVSAvoidEnergy consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary actions by automatically detecting user intent through gaze tracking and body pose analysis before explicit interaction commands are given. The magnifying region proactively appears based on gaze direction, and guided access modes are automatically activated based on detected interaction patterns, eliminating the need for users to manually navigate complex menus or perform redundant actions, thus improving efficiency and reducing energy consumption.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The XR system provides self-service through automatic detection and adaptation: the system automatically detects textual content and provides audio descriptions, automatically adjusts the magnifying region based on gaze tracking, and automatically switches between interaction modes based on detected user needs, reducing the cognitive and physical effort required from users and minimizing energy-wasting manual operations.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250341941A1Devices, Methods, and Graphical User Interfaces for Improving Accessibility of Interactions with Three-Dimensional Environments
Publication Date: 2025.11.06 APPLE INC
  • US20250341941A1 patent drawing
  • US20250341941A1 patent drawing
  • US20250341941A1 patent drawing

AI summary

While a view of a three-dimensional environment is visible via a display generation component, a computer system automatically detects an object in the three-dimensional environment. In response to detecting the object and in accordance with a determination that the object includes textual content, the computer system automatically displays, via the display generation component, a user interface element for generating an audio representation of textual content. Further, an input selecting the user interface element is detected. In response to detecting the input selecting the user interface element, an audio representation of at least a portion of the textual content of the object is generated.