Multi-Modal Input Framework for Content Capture

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computing devices and software applications are limited by a single-modality framework for user input, which restricts the creative potential for enhancing content capture experiences.

Innovation Solution

A new multi-mode user input framework that combines voice signals and touch gestures sustained at least partially coincident with each other, allowing for the identification and association of content objects with captured voice signals.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a single-modality framework is used for user input, then the device complexity is reduced and ease of operation is maintained, but the adaptability and versatility of content capture capabilities are limited

Engineering Contradiction:
Improvecontent capture capabilitiesVSAvoidinput framework
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent combines multiple input modalities (voice commands, touch gestures, spatial gestures) into a unified input framework. The system integrates these different input types and processes them together to control content capture events, allowing users to switch between or combine modalities seamlessly. This merging approach enhances adaptability while managing complexity through a unified processing architecture.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The input framework is designed to support multiple functions across different modalities. A single framework handles voice-based content capture, touch-based capture, spatial gesture-based capture, and hybrid combinations thereof. This multi-functional design allows the same underlying system to serve multiple purposes without requiring separate specialized systems for each input type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If multiple input modalities are supported simultaneously, then the adaptability and versatility of content capture are enhanced, but the device complexity and difficulty of operation increase

Engineering Contradiction:
Improvecontent capture capabilitiesVSAvoiduser input
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The system dynamically adapts to user input preferences and context. It can switch between different modalities based on situational requirements, user behavior patterns, and environmental factors. This dynamic behavior allows the system to present the most appropriate input mode at any given moment, maintaining ease of operation while supporting multiple modalities.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent introduces a unified input framework that acts as an intermediary layer between various input modalities and the content capture system. This mediator translates and harmonizes inputs from different sources (voice, touch, spatial gestures) into a common format, simplifying the user experience while enabling multi-modality support.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If traditional single-modality input is used, then the ease of operation is maintained, but the productivity and efficiency of content capture are limited

Engineering Contradiction:
Improvecontent capture efficiencyVSAvoidinput framework
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system performs preliminary processing and preparation of input data from multiple modalities before content capture occurs. It pre-processes voice commands, touch gestures, and spatial gestures to identify user intent and prepare the content capture event. This preliminary action streamlines the capture process and improves efficiency while managing complexity through structured preprocessing.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The unified input framework enables continuous processing of multiple input streams simultaneously. Rather than handling modalities sequentially, the system processes voice, touch, and spatial gesture inputs in parallel, maintaining continuous useful action across all input channels. This continuity improves productivity by eliminating idle time between input modalities.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentEP4058882B1Content capture experiences driven by multi-modal user inputs
Publication Date: 2025.04.23 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP4058882B1 patent drawingFigure 1
  • EP4058882B1 patent drawingFigure 2A
  • EP4058882B1 patent drawingFigure 2B

AI summary

Systems, methods, and software are disclosed herein for enhancing the content capture experience on computing devices. In an implementation, a combined user input comprises a voice signal and a touch gesture sustained at least partially coincident with the voice signal. An occurrence of the combined user input triggers the identification of an associated content object which may then be associated with a captured version of the voice signal. Such an advance provides users with a new framework for interacting with their devices, applications, and surroundings.