Touch Screen Haptic Feedback for Visually Impaired Image Interaction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques fail to effectively communicate the features and details of digital images to visually impaired individuals, limiting their ability to perceive and interact with images on smart devices, as they rely on vague descriptions and lack tactile feedback.
Innovation Solution
A system using a touch-sensitive screen that employs machine learning models to generate audible captions and unique vibration patterns for objects within digital images, allowing visually impaired users to identify and explore images through tactile feedback, with bounding boxes and object tags enhancing the description and interaction experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If traditional text-based screen readers are used to describe images, then visually impaired users can access image information, but the descriptions are vague and lack detailed spatial relationships and object features
Solution Approach 1:
The patent segments the image into multiple objects with bounding boxes and generates separate captions for each object. This segmentation allows detailed description of individual objects and their spatial relationships without overwhelming the user with a single vague description, thereby reducing information loss while maintaining manageable system complexity through modular processing
Solution Approach 2:
The patent introduces machine learning models as intermediaries between the image and the user. These models automatically generate object captions and spatial relationship descriptions, acting as a mediator that transforms visual information into accessible audio descriptions without requiring manual annotation, thus reducing information loss while avoiding the complexity of manual processing systems
2Loss of information
If detailed object captions are generated for all objects in an image, then visually impaired users can perceive image details, but the processing time and computational resources increase
Solution Approach 1:
The patent generates captions for objects based on user interaction - when a user touches or hovers over a specific object, the system generates a caption for that object rather than pre-generating captions for all objects. This partial action approach provides accurate object descriptions when needed while avoiding the time consumption of generating all possible captions in advance
Solution Approach 2:
The patent performs preliminary processing by detecting objects and generating bounding boxes before user interaction. This preliminary action prepares the system for rapid caption generation upon user request, reducing the perceived processing time while maintaining description accuracy when users need information about specific objects
3Ease of operation
If only auditory feedback is provided for image description, then the system remains simple, but visually impaired users lack tactile feedback to explore and identify objects in the image
Solution Approach 1:
The patent merges auditory feedback with tactile feedback by integrating haptic vibration patterns into the image description system. When users touch objects on the screen, they receive both audio captions describing the object and distinctive vibration patterns, creating a combined sensory experience that enhances interaction capability without requiring entirely separate systems
Solution Approach 2:
The patent applies local quality by providing differentiated tactile feedback - each object in the image can have a unique vibration pattern associated with it. This allows users to distinguish between different objects through tactile means while maintaining system simplicity through the use of standard haptic motor technology already present in mobile devices
Data Source
AI summary
Some implementations include methods for communicating features of images to visually impaired users. An image to be displayed on a touch sensitive screen of a computing device may include one or more objects. Each of the one or more objects may be associated with a bounding box. A contact with the image may be detected via the touch sensitive screen. The contact may be determined to be within a bounding box associated with a first object of the one or more objects. Responsive to detecting the contact to be within the bounding box associated with the first object, a caption of the first object may be caused to become audible and the touch sensitive screen may be caused to vibrate based on a vibration pattern unique to the first object.


