Gaze-Based Text Recognition for Real-Time Audio Translation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing visual devices used for interpretation or translation systems often block the user's field of view, limiting their usability, especially in outdoor environments. Additionally, gaze recognition technology may incur limitations in specific user environments, and non-visual user feedback methods can result in partial errors and inappropriate feedback.
Innovation Solution
A method, apparatus, and system that utilize visual information to recognize text at a user's gaze position, convert it into a target language, and provide the result in an auditory form, thereby overcoming the limitations of existing visual devices and feedback methods.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If visual devices are used for interpretation or translation, then language conversion capability is improved, but field of view is blocked and usability is limited
Solution Approach 1:
The patent uses the user's gaze position as an intermediary to select which text to translate. Instead of blocking the view with a display, the system captures the user's natural line of sight, identifies text in that region, and provides translation through audio feedback. This mediator approach allows language conversion without visual obstruction.
Solution Approach 2:
The patent replaces the mechanical visual display system with an auditory feedback system. Instead of presenting translated text visually (which blocks field of view), the system converts translation results into audio signals that provide language conversion capability without obstructing the user's natural vision.
2Ease of operation
If auditory or tactile feedback is used to compensate for visual device limitations, then field of view is preserved, but partial errors and inappropriate feedback occur
Solution Approach 1:
The patent implements a feedback loop where the system continuously monitors the user's gaze position, identifies the current text region of interest, and provides audio translation feedback corresponding to that specific text. This targeted feedback mechanism ensures that the auditory information accurately reflects what the user is currently viewing, reducing inappropriate feedback.
Solution Approach 2:
The patent applies local quality by providing translation feedback specifically for the text region where the user's gaze is directed, rather than providing general or global translation. This localized approach ensures that the audio feedback corresponds precisely to the relevant text portion, improving feedback accuracy and reducing errors.
3Loss of information
If text recognition is performed on the entire image, then all text is detected, but processing time increases and real-time interpretation is difficult
Solution Approach 1:
The patent segments the image processing task by dividing it into regions based on the user's gaze position. Instead of processing the entire image, the system identifies the gaze region and performs text recognition only on that specific area. This segmentation dramatically reduces processing time while maintaining detection of relevant text information.
Solution Approach 2:
The patent applies partial action by performing text recognition only on the portion of the image that corresponds to the user's gaze position, rather than processing the entire image. This partial processing approach reduces computational load and processing time while still capturing all text that the user is currently interested in.
Data Source
AI summary
Provided is a method of providing an interpretation result using visual information, and the method includes: acquiring a spatial domain image including line-of-sight information of a user and gaze position information in the spatial domain image; segmenting the acquired spatial domain image into a plurality of images; detecting text areas including text for each of the segmented images; generating text blocks, each of which is a text recognition result for each of the detected text areas, and determining the text block corresponding to the gaze position information; converting a first language included in the determined text block into a second language that is a target language; and providing the user with a conversion result of the second language.


