Image Description Generation for Visually Impaired Users
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Social networking systems face challenges in providing visually impaired users with effective access to visual content, such as images and videos, as traditional interfaces are not optimized for users with physical impairments, limiting their engagement and interaction with the platform.
Innovation Solution
The system employs machine learning techniques for automatic image description generation, using object recognition, facial recognition, and optical character recognition to identify concepts in images, assign confidence scores, and filter them based on thresholds, generating descriptions that can be embedded in images for screen readers and allowing users to request additional information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If traditional interfaces are used for visually impaired users, then the system maintains simplicity, but user engagement and interaction are limited
Solution Approach 1:
The patent introduces an intermediary system that includes image analysis software and screen reader integration. This intermediary translates visual content into audible descriptions, enabling visually impaired users to access and engage with image content that would otherwise be inaccessible through traditional visual interfaces alone
Solution Approach 2:
The patent replaces the mechanical/visual interface system with an auditory information delivery system. By substituting visual display mechanisms with text-to-speech and audio description technologies, the system enables visually impaired users to interact with content through their hearing capability rather than sight
2Measurement precision
If machine learning techniques are used to identify concepts in images, then image description accuracy is improved, but system complexity increases
Solution Approach 1:
The patent segments the image analysis task into multiple specialized machine learning components: object recognition models, facial recognition models, and text recognition models. Each component focuses on a specific aspect of image content, improving overall accuracy while allowing independent optimization and management of each recognition subsystem
Solution Approach 2:
The patent implements a universal image analysis system that uses multi-functional machine learning models capable of performing multiple recognition tasks (objects, faces, text) within a single integrated framework, reducing the need for separate specialized systems for each recognition type
3Loss of information
If all identified concepts are included in image descriptions, then information completeness is improved, but description length and processing time increase
Solution Approach 1:
The patent changes the parameter of concept selection by introducing confidence score thresholds. Concepts are filtered based on their confidence scores, allowing the system to adjust the balance between information completeness and processing efficiency by setting appropriate threshold levels dynamically
Solution Approach 2:
The patent applies local quality by differentiating the inclusion criteria for different types of concepts. High-confidence concepts are included in the primary description, while lower-confidence concepts may be included in expanded descriptions or excluded entirely, allowing optimized processing for each concept type based on its importance and reliability
Data Source
AI summary
Systems, methods, and non-transitory computer-readable media can receive an image. One or more concepts depicted in the image are identified based on machine learning techniques. The one or more concepts are filtered based on filtering criteria to identify one or more selected concepts. An image description is generated comprising the one or more selected concepts.


