Relative Narration for Cognitive Audio Output

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing narration-based applications provide a binary experience, reading textual information from graphical user interfaces verbatim, which can be difficult to comprehend audially due to lack of consideration for human cognitive processes, leading to increased task completion times as they overload listeners with details.

Innovation Solution

A computing device with a mode for relative narration that extracts entities from the user interface and compares them to known information, such as contacts, time, locations, and browsing history, to generate a more understandable audio output, allowing for the transformation of textual information into a narrated string that is easier to comprehend.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing narration-based applications read textual information verbatim from the graphical user interface, then fidelity in the transformation from visual to audible experience is ensured, but the audio output becomes difficult to comprehend and task completion times increase significantly

Engineering Contradiction:
Improvefidelity of information transformationVSAvoidcomprehensibility of audio output
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent introduces an intermediary processing layer between the graphical user interface and the audio output. This layer analyzes the visual context, identifies the user's current task, and generates narrations that are relevant to the task at hand rather than reading all screen elements verbatim. The intermediary filters and transforms information based on task relevance, making the audio output more comprehensible while maintaining fidelity to the user's information needs.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system dynamically changes the parameters of narration delivery based on the user's current task and context. Instead of a fixed verbatim reading approach, the narration content, timing, and detail level are adjusted as parameters to match the user's cognitive processing needs. This allows the system to maintain fidelity by adapting the narration to what the user actually needs to know for their current task, rather than providing all possible information.

Inventive Principle:
Principle #35Parameter changes

2Loss of information

If narration applications provide detailed verbatim reading of all screen elements, then complete information is delivered to the user, but the listener is overloaded with details and task completion times increase by a factor of three to ten times

Engineering Contradiction:
Improvecompleteness of information deliveryVSAvoidtask completion time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent extracts only the most relevant information from the graphical user interface based on the user's current task. Rather than reading all screen elements, the system identifies and extracts the specific information elements that are pertinent to the user's goals. This extraction process filters out redundant or less important information, delivering complete task-relevant information while avoiding overload and reducing task completion time.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system applies partial action by providing narration only for the portions of the interface that are relevant to the current task, rather than reading everything. This selective approach prevents information overload while ensuring that all necessary information for task completion is delivered. The narration is calibrated to provide exactly the right amount of information needed, neither too much nor too little.

Inventive Principle:
Principle #16Partial or excessive action

3Shape

If graphical user interfaces are designed for visual framework, then they present information in a structured visual form, but they do not translate well to an audible experience for users with impaired vision or when hands-free operation is needed

Engineering Contradiction:
Improvevisual structure of information presentationVSAvoidadaptability to audible experience
Core Design Contradiction:
ShapeVSAdaptability or versatility

Solution Approach 1:

The patent introduces dynamic adaptability that allows the interface to switch between visual and audible modes based on user needs. The system dynamically adjusts the narration delivery to match the visual structure of the interface, providing audio descriptions that preserve the hierarchical and contextual relationships present in the visual layout. This dynamic adaptation enables the same information structure to serve both visual and audible users effectively.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system implements a universal interface design that functions effectively for both visual and audible users. By integrating intelligent narration capabilities that understand the visual structure, the interface becomes multi-functional, serving users with different sensory preferences and accessibility needs. The narration system translates visual structural relationships into meaningful audio descriptions, making the same interface universally accessible.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11650791B2Relative narration
Publication Date: 2023.05.16 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11650791B2 patent drawing
  • US11650791B2 patent drawing
  • US11650791B2 patent drawing

AI summary

A computing device and a method for generating relative narration. In one instance, the computing device include a display device displaying a graphical user interface including textual information received from a first application. An electronic processor of the computing device receives a user interface element associated with the textual information scheduled for relative narration. The electronic processor extracts a plurality of entities from the user interface element, converts the plurality of entities into a narrated string using a second application, generates the relative narration of the textual information using the narrated string, and output the relative narration.