Voice Document Interaction via Grammar-Based Command Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current digital document interaction systems require users to manually input data or navigate screens, which can be cumbersome and inaccessible for individuals with disabilities or those engaged in other activities, lacking efficient voice-based interaction capabilities.

Innovation Solution

A computer-implemented method and system that utilizes natural language voice input to perform actions on digital documents by mapping voice-entered terms to computer-executable commands, allowing users to interact with documents using voice commands to enter data or retrieve information, with audible feedback, leveraging document-specific voice grammars and translation data structures.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual input methods (touch gestures, mouse clicks) are used to interact with digital documents, then input precision and control are improved, but ease of operation deteriorates for users with disabilities or those engaged in other activities

Engineering Contradiction:
Improveinput precisionVSAvoidease of operation
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent replaces manual mechanical input methods (touch gestures, mouse clicks) with voice-based acoustic input. The system captures spoken commands through a microphone, processes them through speech recognition, and executes corresponding actions on document elements, eliminating the need for manual dexterity while maintaining input precision through structured grammar-based command interpretation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of operation

If voice input is used to interact with digital documents, then ease of operation is improved for hands-free operation, but device complexity increases due to speech recognition and grammar processing requirements

Engineering Contradiction:
Improveease of operationVSAvoiddevice complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent implements a universal voice command system that can perform multiple functions (navigating, selecting, entering data, reading information) through a single voice interface. The document-specific grammar structure allows the same speech recognition engine to handle diverse document interactions, reducing the need for separate control mechanisms for each function while maintaining ease of operation.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces a document-specific grammar as an intermediary layer between the voice input and the document manipulation system. This grammar structure acts as a translation medium that converts natural language speech into structured commands, simplifying the processing complexity by providing a standardized intermediate representation rather than directly interpreting raw speech for each possible action.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If traditional navigation methods are used to find document elements, then measurement precision of element location is improved, but loss of time increases due to manual searching and scrolling

Engineering Contradiction:
Improveelement location precisionVSAvoidtime to locate element
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements preliminary action by using voice commands to directly navigate to and identify document elements without requiring manual scrolling or searching. The speech recognition system can interpret navigation commands (e.g., 'go to section', 'find element') and automatically position the document view and select the target element, saving time while maintaining precise element location through the structured grammar that specifies element identifiers and positions.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP3195308B1Actions on digital document elements from voice
Publication Date: 2019.07.03 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP3195308B1 patent drawingFigure 1
  • EP3195308B1 patent drawingFigure 2
  • EP3195308B1 patent drawingFigure 3

AI summary

A set of one or more terms can be derived from the voiced user input. It can be determined that the set of one or more terms corresponds to a specific digital document in a computer system, and that the set of one or more terms corresponds to at least one computer-executable command. The determination that the set of one or more terms corresponds to the at least one command can include analyzing the set of one or more terms using a document-specific natural language translation data structure. The translation data structure can be a computer-readable data structure that is identified in the computer system with the specific digital document in the computer system. It can be determined that the set of one or more terms corresponds to an element in the document, and the at least one command can be executed on the element in the document.