Scene-to-text conversion for visually impaired users
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Visually impaired individuals face challenges in accessing and interacting with graphical environments, such as virtual and augmented reality, due to the lack of accessible descriptions of their surroundings.
Innovation Solution
A device and method for performing scene-to-text conversion, which identifies objects in an environment using environmental data and generates an audio output describing these objects based on user-specific characteristics, such as location, education level, and vision capability, to provide a narrative description.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If graphical environments are presented to users, then visual information is provided, but visually impaired users cannot access or understand the environment
Solution Approach 1:
The patent introduces an audio output as an intermediary medium to convey environmental information. The system captures visual environmental data through sensors, processes it to identify objects and characteristics, and then presents this information through audio output that describes the environment. This intermediary audio layer bridges the gap between visual graphical environments and visually impaired users, allowing them to access and understand their surroundings without directly viewing the graphical elements.
2Ease of operation
If audio output is generated to describe environment, then accessibility is improved, but device complexity increases
Solution Approach 1:
The patent leverages existing multi-functional components of mobile devices to reduce overall system complexity. The device uses its existing camera or environmental sensors to capture visual data, the same processor that handles graphical rendering to analyze and identify objects, and the existing audio output hardware (speakers or headphones) to present descriptive information. By making these existing components serve the additional function of accessibility support, the patent avoids adding dedicated separate hardware systems while still providing comprehensive environmental description capabilities.
Data Source
AI summary
Various implementations disclosed herein include devices, systems, and methods for performing scene-to-text conversion. In various implementations, a device includes a non-transitory memory and one or more processors coupled with the non-transitory memory. In some implementations, a method includes obtaining environmental data corresponding to an environment. Based on the environmental data, a plurality of objects that are in the environment are identified. An audio output describing at least a first object of the plurality of objects in the environment is generated based on a characteristic value associated with a user of the device. The audio output is outputted.


