Head-Mounted Device Text Recognition via Contextual OCR
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Visually impaired users face challenges in navigating real-world environments due to inability to recognize text, with existing solutions either altering the environment, requiring additional devices, or overloading users with information from multiple text sources.
Innovation Solution
A head-mounted device coupled with an assistance engine that performs optical character recognition on video streams from the user's environment, providing auditory feedback based on context and dictionary-specific text recognition, allowing users to direct the OCR process and receive personalized assistance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing text recognition solutions are used, then text can be identified, but users experience cognitive overload from multiple text sources and environmental alterations are required
Solution Approach 1:
The patent combines the camera, microphone, speaker, and OCR processing into a single head-mounted device system. The camera captures video streams, the microphone receives verbal context, the OCR engine processes text recognition, and the speaker provides auditory feedback - all integrated into one wearable system that works seamlessly together to eliminate the need for multiple separate devices or environmental modifications.
Solution Approach 2:
The system uses the user's own verbal context and video streams to automatically perform text recognition and provide relevant information. The OCR engine processes only the text that is relevant to the user's spoken queries, eliminating the need for environmental alterations or manual text source selection, and providing self-directed service without cognitive overload.
2Ease of operation
If environmental alterations are made to assist text recognition, then text can be recognized, but the environment must be modified and additional devices may be required
Solution Approach 1:
The patent replaces mechanical or environmental modifications with an electronic/optical system. Instead of altering the physical environment to make text visible or audible, the system uses a camera to capture video streams, processes them through OCR, and delivers results through auditory feedback - substituting physical environmental changes with electronic processing and digital output.
Solution Approach 2:
The head-mounted device performs multiple functions: capturing video streams, receiving verbal context, performing OCR processing, and providing auditory feedback. This multi-functional system eliminates the need for separate devices or environmental modifications, as one device handles text recognition accessibility in any environment without requiring implementation of environmental alterations.
3Loss of information
If all text sources are processed, then comprehensive information is provided, but users experience cognitive overload
Solution Approach 1:
The system applies different processing quality to different text sources based on relevance. Instead of uniformly processing all text sources with the same detail, the OCR engine prioritizes and processes only the text that is locally relevant to the user's specific verbal context and current needs, providing high-quality information for relevant text while filtering out irrelevant content to prevent cognitive overload.
Solution Approach 2:
The system uses feedback from the user's verbal context to dynamically adjust text processing. The microphone captures what the user is saying or asking about, the OCR engine uses this feedback to identify and process only the relevant text sources, and the speaker provides targeted auditory feedback - creating a closed-loop system that prevents cognitive overload by processing only what the user needs to know.
Data Source
AI summary
A user is assisted in a real-world environment. An assistance engine receives at least one context from the user. The assistance engine also receives a video stream of the real-world environment. The assistance engine performs an optical character recognition process on the video stream based upon the at least one context. The assistance engine generates a response for the user. A microphone on a head-mounted device receives the context from the user. A camera on the head-mounted device captures the video stream of the real-world environment. A speaker on the head-mounted device communicates the response to the user. The user may move in the real-world environment based upon the response to improve the optical character recognition process.


