Speech Disambiguation Using Depth Camera Visual Cues

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems face difficulties in disambiguating ambiguous speech inputs, often relying on user clarification that can disrupt natural interaction by requiring visual cues, such as gestures or identity verification, to resolve ambiguities.

Innovation Solution

The method involves using depth cameras and microphones to receive depth information and audio inputs, analyzing visual cues like user identity and digital content consumption information to identify unambiguous terms in speech inputs, thereby disambiguating them and taking appropriate actions on a computing device.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech recognition systems ask users to repeat or correct ambiguous inputs, then accuracy of speech recognition is improved, but interaction efficiency deteriorates and natural interaction is disrupted

Engineering Contradiction:
Improveaccuracy of speech recognitionVSAvoidinteraction efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent introduces visual cues (depth information, user identity, gestures) as an intermediary to bridge the gap between speech input and intended meaning. Instead of directly asking users to repeat, the system uses visual context as a mediator to automatically disambiguate ambiguous terms, thus maintaining both accuracy and interaction efficiency

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs self-service by automatically resolving ambiguities using integrated sensors and user context data without requiring user intervention. The computing device uses its own resources (depth camera, microphones, user profiles) to independently disambiguate speech inputs, eliminating the need for users to repeat or correct themselves

Inventive Principle:
Principle #25Self-service

2Measurement precision

If speech recognition systems use multiple sensors and contextual analysis, then accuracy of speech recognition is improved, but device complexity increases

Engineering Contradiction:
Improveaccuracy of speech recognitionVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges multiple sensor types (depth camera, microphones, image sensors) and data sources (user profiles, digital content consumption information) into a unified speech recognition system. By combining these components, the system achieves higher accuracy through multi-modal data fusion while managing complexity through integrated processing

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS9190058B2Using visual cues to disambiguate speech inputs
Publication Date: 2015.11.17 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9190058B2 patent drawing
  • US9190058B2 patent drawing
  • US9190058B2 patent drawing

AI summary

Embodiments related to recognizing speech inputs are disclosed. One disclosed embodiment provides a method for recognizing a speech input including receiving depth information of a physical space from a depth camera, determining an identity of a user in the physical space based on the depth information, receiving audio information from one or more microphones, and determining a speech input from the audio input. If the speech input comprises an ambiguous term, the ambiguous term in the speech input is compared to one or more of depth image data received from the depth image sensor and digital content consumption information for the user to identify an unambiguous term corresponding to the ambiguous term. After identifying the unambiguous term, an action is taken on the computing device based on the speech input and the unambiguous term.