Multimodal Browser Document Session Replay
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Multimodal applications on small devices face challenges in replaying document sessions due to limited user interaction capabilities, as existing technologies do not effectively support the replay of varied user responses and visual elements across different interaction modes.
Innovation Solution
A multimodal browser identifies speech prompts and responses from a log produced by a Form Interpretation Algorithm, retrieves and renders associated X+V pages, and replays both speech and visual elements to facilitate document session replay, enabling review of specific interactions in multimodal applications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If multimodal applications use multiple input modes (speech, keyboard, touch screen) to improve user interaction on small devices, then ease of operation is improved, but device complexity increases
Solution Approach 1:
The patent implements a universal replay mechanism that handles multiple interaction modes (speech, keyboard, touch screen, stylus) through a single integrated system. The browser's replay functionality universally processes all these different input types using the same log structure and replay architecture, allowing one system to serve multiple interaction functions without requiring separate replay mechanisms for each mode.
2Measurement precision
If document sessions are replayed to improve review capability, then measurement precision of user interactions is improved, but device complexity increases
Solution Approach 1:
The system performs preliminary action by logging all interaction details (speech prompts, responses, visual elements, timing) during the original document session. This pre-capture of complete interaction data enables accurate replay without requiring complex real-time processing during the replay phase, as all necessary information has already been recorded in a structured format.
Solution Approach 2:
The patent creates a detailed copy of the original document session by recording all interactions, speech exchanges, and visual states in a log. This copy contains sufficient information to reconstruct and replay the entire session accurately, allowing review of user interactions without needing to maintain the original session state or implement complex reconstruction algorithms.
3Adaptability or versatility
If speech recognition and synthesis are integrated to improve multimodal capability, then adaptability is improved, but device complexity increases
Solution Approach 1:
The patent merges speech recognition and speech synthesis functionalities into an integrated multimodal system. The speech engine combines both recognition (converting speech to text) and synthesis (converting text to speech) capabilities, allowing the application to handle spoken input and generate spoken output through a unified interface. This merging reduces the need for separate independent systems while maintaining full dual-directional speech capability.
Data Source
AI summary
Methods, apparatus, and computer program products are described for document session replay for multimodal applications. including identifying, by a multimodal browser in dependence upon a log produced by a Form Interpretation Algorithm (‘FIA’) during a previous document session with a user, a speech prompt provided by a multimodal application in the previous document session; identifying, by a multimodal browser in replay mode in dependence upon the log, a response to the prompt provided by a user of the multimodal application in the previous document session; retrieving, by the multimodal browser in dependence upon the log, an X+V page of the multimodal application associated with the speech prompt and the response; rendering, by the multimodal browser, the visual elements of the retrieved X+V page; replaying, by the multimodal browser, the speech prompt; and replaying, by a multimodal browser, the response.


