Intelligent Screen Reading for Context- and Emotion-Aware Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing electronic devices lack the ability to intelligently read displayed content, failing to associate intent, context, and emotion, leading to confusion for visually impaired and general users.
Innovation Solution
An electronic device equipped with an intelligent screen reading engine that analyzes displayed content to extract insights on intent, importance, emotion, and sound representation, using deep neural networks and generative text reading to provide meaningful audio output.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If existing text-to-speech method is used to read displayed content, then the device can read aloud text and emoji definitions, but the reading lacks emotional meaning and context understanding
Solution Approach 1:
The patent introduces an intelligent screen reading engine as an intermediary between the displayed content and the user. This engine analyzes the visual content, extracts emotional meaning and context, and generates enhanced audio output with appropriate emotional tone, thereby recovering the lost emotional information without requiring complete system redesign
Solution Approach 2:
The patent replaces the mechanical text-to-speech reading system with an intelligent system that uses deep learning models (BERT, GPT) to understand and generate emotionally nuanced audio. This substitution transforms the reading process from simple text conversion to intelligent content comprehension and emotional expression
2Loss of information
If the device reads every displayed content element in detail, then complete information is provided, but the user gets confused and loses the actual intent
Solution Approach 1:
The intelligent screen reading engine extracts only the most relevant and meaningful elements from the displayed content, separating essential information from redundant details. It identifies key semantic units, emotional indicators, and contextual importance, presenting only what matters to the user while maintaining complete understanding of the original content
Solution Approach 2:
The patent applies different processing qualities to different parts of the displayed content based on their importance. Critical information receives detailed analysis and prominent audio presentation, while less important elements receive simplified handling. This local differentiation optimizes both information completeness and user comprehension
3Productivity
If the device reads content without understanding meaning and intent, then the reading process is simple and fast, but the output appears mechanical and lacks human-like emotion
Solution Approach 1:
The intelligent screen reading engine performs preliminary analysis of the displayed content before generating audio output. It pre-processes the visual information to extract semantic meaning, emotional tone, and contextual relationships, preparing the groundwork for emotionally accurate reading while maintaining efficient processing throughput
Data Source
AI summary
A method for intelligently reading displayed contents by an electronic device is provided. The method includes obtaining a screen representation based on a plurality of contents displayed on a screen of the electronic device. The method includes extracting a plurality of insights comprising at least one of intent, importance, emotion, sound representation and information sequence of the plurality of contents from the plurality of contents based on the screen representation. The method includes generating audio emulating the extracted plurality of insights.


