Scene-to-text conversion for visually impaired users

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Visually impaired individuals face challenges in accessing and interacting with graphical environments, such as virtual and augmented reality, due to the lack of accessible descriptions of their surroundings.

Innovation Solution

A device and method for performing scene-to-text conversion, which identifies objects in an environment using environmental data and generates an audio output describing these objects based on user-specific characteristics, such as location, education level, and vision capability, to provide a narrative description.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If graphical environments are presented to users, then visual information is provided, but visually impaired users cannot access or understand the environment

Engineering Contradiction:
Improveaccessibility of environmental informationVSAvoidvisual impairment limitation
Core Design Contradiction:
Loss of informationVSObject-affected harmful factors

Solution Approach 1:

The patent introduces an audio output as an intermediary medium to convey environmental information. The system captures visual environmental data through sensors, processes it to identify objects and characteristics, and then presents this information through audio output that describes the environment. This intermediary audio layer bridges the gap between visual graphical environments and visually impaired users, allowing them to access and understand their surroundings without directly viewing the graphical elements.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If audio output is generated to describe environment, then accessibility is improved, but device complexity increases

Engineering Contradiction:
Improveaccessibility for visually impaired usersVSAvoidprocessing and output system complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent leverages existing multi-functional components of mobile devices to reduce overall system complexity. The device uses its existing camera or environmental sensors to capture visual data, the same processor that handles graphical rendering to analyze and identify objects, and the existing audio output hardware (speakers or headphones) to present descriptive information. By making these existing components serve the additional function of accessibility support, the patent avoids adding dedicated separate hardware systems while still providing comprehensive environmental description capabilities.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12033381B2Scene-to-text conversion
Publication Date: 2024.07.09 APPLE INC
  • US12033381B2 patent drawing
  • US12033381B2 patent drawing
  • US12033381B2 patent drawing

AI summary

Various implementations disclosed herein include devices, systems, and methods for performing scene-to-text conversion. In various implementations, a device includes a non-transitory memory and one or more processors coupled with the non-transitory memory. In some implementations, a method includes obtaining environmental data corresponding to an environment. Based on the environmental data, a plurality of objects that are in the environment are identified. An audio output describing at least a first object of the plurality of objects in the environment is generated based on a characteristic value associated with a user of the device. The audio output is outputted.