Multimodal Browser Document Session Replay

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Multimodal applications on small devices face challenges in replaying document sessions due to limited user interaction capabilities, as existing technologies do not effectively support the replay of varied user responses and visual elements across different interaction modes.

Innovation Solution

A multimodal browser identifies speech prompts and responses from a log produced by a Form Interpretation Algorithm, retrieves and renders associated X+V pages, and replays both speech and visual elements to facilitate document session replay, enabling review of specific interactions in multimodal applications.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If multimodal applications use multiple input modes (speech, keyboard, touch screen) to improve user interaction on small devices, then ease of operation is improved, but device complexity increases

Engineering Contradiction:
Improveuser interaction capabilityVSAvoidinteraction modes
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent implements a universal replay mechanism that handles multiple interaction modes (speech, keyboard, touch screen, stylus) through a single integrated system. The browser's replay functionality universally processes all these different input types using the same log structure and replay architecture, allowing one system to serve multiple interaction functions without requiring separate replay mechanisms for each mode.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If document sessions are replayed to improve review capability, then measurement precision of user interactions is improved, but device complexity increases

Engineering Contradiction:
Improveinteraction review accuracyVSAvoidreplay system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary action by logging all interaction details (speech prompts, responses, visual elements, timing) during the original document session. This pre-capture of complete interaction data enables accurate replay without requiring complex real-time processing during the replay phase, as all necessary information has already been recorded in a structured format.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a detailed copy of the original document session by recording all interactions, speech exchanges, and visual states in a log. This copy contains sufficient information to reconstruct and replay the entire session accurately, allowing review of user interactions without needing to maintain the original session state or implement complex reconstruction algorithms.

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If speech recognition and synthesis are integrated to improve multimodal capability, then adaptability is improved, but device complexity increases

Engineering Contradiction:
Improvemultimodal accessVSAvoidspeech processing
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent merges speech recognition and speech synthesis functionalities into an integrated multimodal system. The speech engine combines both recognition (converting speech to text) and synthesis (converting text to speech) capabilities, allowing the application to handle spoken input and generate spoken output through a unified interface. This merging reduces the need for separate independent systems while maintaining full dual-directional speech capability.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS7801728B2Document session replay for multimodal applications
Publication Date: 2010.09.21 MICROSOFT TECHNOLOGY LICENSING LLC
  • US7801728B2 patent drawing
  • US7801728B2 patent drawing
  • US7801728B2 patent drawing

AI summary

Methods, apparatus, and computer program products are described for document session replay for multimodal applications. including identifying, by a multimodal browser in dependence upon a log produced by a Form Interpretation Algorithm (‘FIA’) during a previous document session with a user, a speech prompt provided by a multimodal application in the previous document session; identifying, by a multimodal browser in replay mode in dependence upon the log, a response to the prompt provided by a user of the multimodal application in the previous document session; retrieving, by the multimodal browser in dependence upon the log, an X+V page of the multimodal application associated with the speech prompt and the response; rendering, by the multimodal browser, the visual elements of the retrieved X+V page; replaying, by the multimodal browser, the speech prompt; and replaying, by a multimodal browser, the response.