Multimodal Task Assistant Context Maintenance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current personal assistants for mobile devices that accept voice input are limited in their ability to provide efficient and seamless multimodal interactions, lacking integration of voice, text, and touch controls, which restricts user convenience and effectiveness in accessing and interacting with applications.

Innovation Solution

A multimodal task assistant system that enables input through voice, typed text, and touch or gross movement controls, and provides output via speech, text, visual displays, and haptic feedback, allowing users to interact with applications in a more intuitive and efficient manner by maintaining context and providing contextual suggestions and navigation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If voice input is used for personal assistance, then user convenience is improved, but interaction effectiveness is limited due to lack of multimodal integration

Engineering Contradiction:
Improveuser convenienceVSAvoidinteraction effectiveness
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The patent combines multiple input modalities (voice, text, touch, gross movement) and output modalities (speech, text, visual display, haptic feedback) into a unified personal assistant system. This merging allows the system to process and respond through multiple channels simultaneously, resolving the contradiction by maintaining ease of voice-based operation while significantly improving interaction effectiveness through complementary modalities.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The personal assistant is designed to perform multiple functions across different modalities - it can receive input through voice, text, touch, or movement sensors, and provide output through speech synthesis, text display, visual interfaces, or haptic feedback. This multi-functionality enables the system to adapt to various user needs and contexts, improving interaction effectiveness without sacrificing the convenience of voice-based operation.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Device complexity

If single-modal input is used, then system simplicity is maintained, but user experience efficiency is restricted

Engineering Contradiction:
Improvesystem simplicityVSAvoidtask execution efficiency
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The patent segments the input and output functions into distinct modalities (voice input, text input, touch input, movement input; speech output, text output, visual output, haptic output). Each modality is processed independently but integrated through a common context-maintaining framework. This segmentation allows the system to manage complexity through modular design while achieving high task execution efficiency through coordinated multimodal operation.

Inventive Principle:
Principle #1Segmentation

3Ease of operation

If voice-only interaction is used, then ease of use is improved, but context maintenance capability is limited

Engineering Contradiction:
Improveease of useVSAvoidcontext maintenance capability
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent introduces a context maintenance mechanism that acts as an intermediary between different input modalities and the core processing system. This intermediary continuously tracks and updates the interaction context across voice, text, touch, and movement inputs, preventing information loss by maintaining a comprehensive state representation that integrates all modalities while preserving the ease of voice-only interaction when sufficient.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9111546B2Speech recognition and interpretation system
Publication Date: 2015.08.18 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9111546B2 patent drawing
  • US9111546B2 patent drawing
  • US9111546B2 patent drawing

AI summary

A method of providing a task assistant comprising starting to receive speech input from a user, and identifying a format associated with a destination for speech input based on a flag associated with the destination field. When the format comprises dictation, converting the speech to text, and inserting it into the destination location, and when the format comprises an intent, determining a meaning of the input, and sending a formatted query to an application. The method further comprising receiving data from the application in response to the intent and providing a response to the user through multimodal output.