Voice Command Accuracy via Physical Interaction Correlation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing natural language processing systems, such as IBM Watson, face inaccuracies in processing voice-activated commands due to language peculiarities and human reasoning, leading to incorrect outcomes in voice-activated artificial intelligence platforms.

Innovation Solution

A system that captures verbal content, converts it into text, and analyzes physical interactions to identify correlations, using a knowledge engine to select appropriate user interface activities and execute corresponding user interface traces within identified applications.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If natural language processing systems process voice-activated commands based on acquired knowledge, then the system can understand and respond to user input, but the outcome can be incorrect or inaccurate due to language peculiarities and human reasoning

Engineering Contradiction:
Improvevoice-activated command processingVSAvoidaccuracy of command execution
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent introduces an intermediary system that captures both verbal content and corresponding physical interactions with the user interface. This intermediary layer correlates speech with actual user actions, creating a bridge between voice commands and intended operations, thereby improving accuracy without compromising ease of operation

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements feedback by capturing physical interactions that occur in response to voice commands and using this information to verify and correct the interpretation of verbal content. The feedback loop allows the system to learn from the discrepancy between spoken commands and actual user actions, improving reliability over time

Inventive Principle:
Principle #23Feedback

2Measurement precision

If the system captures and analyzes both verbal content and physical interactions to identify correlations, then the accuracy of command interpretation improves, but the system complexity increases

Engineering Contradiction:
Improveaccuracy of verbal content interpretationVSAvoidsystem architecture
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies multi-functionality by using a single system component that performs multiple tasks: capturing verbal content, capturing physical interactions, correlating the two data streams, and executing commands. This universal approach improves measurement precision without proportionally increasing device complexity

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system merges the capture and analysis of verbal content with the capture and analysis of physical interactions into a unified processing framework. By combining these previously separate functions into one integrated system, the patent achieves higher interpretation accuracy while minimizing the increase in overall system complexity

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS10620912B2Machine learning to determine and execute a user interface trace
Publication Date: 2020.04.14 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10620912B2 patent drawing
  • US10620912B2 patent drawing
  • US10620912B2 patent drawing

AI summary

A system, computer program product, and method are provided for use with an intelligent computer platform to convert speech intents to one or more physical actions. The aspect of converting speech intent includes receiving audio, converting the audio to text, parsing the text into segments, identifying a physical action and associated application that are proximal in time to the received audio. A corpus is searched for evidence of a pattern associated with the received audio and corresponding physical action(s) and associated application. An outcome is generated from the evidence. The outcome includes identification of an application and production of an affirmative instruction. The instruction is converted to a user interface trace that is executed within the identified application.