Multimodal Sensor Fusion for Natural Language Understanding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audible input-enabled devices often misinterpret voice commands due to insufficient context, leading to incorrect execution of functions, as they rely solely on audible input without considering additional modalities like gestures or environmental factors.

Innovation Solution

A device with a processor, microphone, and sensors that receives both audible and sensor inputs to perform natural language understanding, augmenting the interpretation with data from sensors such as cameras, GPS, and motion sensors to enhance context understanding and resolve ambiguities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If devices rely solely on audible input for natural language understanding, then device complexity is reduced, but language understanding precision deteriorates due to insufficient context

Engineering Contradiction:
Improvelanguage understanding precisionVSAvoiddevice complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines multiple sensing modalities (audible input from microphone, visual input from camera, location data from GPS, motion data from accelerometers) into a unified natural language understanding system. The processor integrates data from all these sensors to augment context and resolve ambiguities, thereby improving language understanding precision while managing device complexity through systematic data fusion.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The device employs a multi-functional approach where a single processor handles multiple sensing modalities and performs various functions including audio processing, image analysis, location tracking, and motion detection. This universal processing architecture allows the device to improve language understanding precision without proportionally increasing overall device complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If devices use only audible input processing, then processing speed is maintained, but language understanding reliability deteriorates due to lack of contextual information

Engineering Contradiction:
Improvelanguage understanding reliabilityVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system performs preliminary data collection by continuously monitoring multiple sensors (microphone, camera, GPS, motion sensors) before natural language processing is needed. This preliminary action gathers contextual information in advance, improving language understanding reliability when the actual processing occurs, while distributing the processing complexity over time rather than concentrating it all at once.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary processing layer that receives data from multiple sensors and performs data fusion before passing information to the natural language understanding module. This intermediary layer integrates contextual information from various sources, improving reliability by providing more complete context, while managing processing complexity through modular architecture.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10741175B2Systems and methods for natural language understanding using sensor input
Publication Date: 2020.08.11 LENOVO SWITZERLAND INTERNATIONAL GMBH
  • US10741175B2 patent drawing
  • US10741175B2 patent drawing
  • US10741175B2 patent drawing

AI summary

In one aspect, a device includes a processor, a microphone accessible to the processor, at least a first sensor that is accessible to the processor, and storage accessible to the processor. The storage bears instructions executable by the processor to receive first input from the microphone that is generated based on audible input from a user. The instructions are also executable by the processor to receive second input from the first sensor, perform natural language understanding based on the first input, augment the natural language understanding based on the second input, and provide an output based on the augmentation.