Location-Based Voice Recognition for Multi-Device IoT Networks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice command systems face challenges in accurately determining the intended device for voice commands in multi-device setups, leading to undesired operations due to the use of additional sensors like cameras or multiple microphones, which increase costs and complexity.

Innovation Solution

A location-based voice recognition system that uses multiple microphones to determine the utterance direction of a user by calculating relative locations and applying sound attenuation models, allowing for accurate interpretation of voice commands without additional sensors.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If additional sensors like cameras or multiple microphones are installed to determine user gaze or sound source direction, then the accuracy of device selection for voice commands is improved, but the manufacturing cost and device complexity increase

Engineering Contradiction:
Improveaccuracy of device selectionVSAvoiddevice complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces a server as an intermediary that performs the complex calculations for determining user gaze direction and sound source direction. The sensor network devices (cameras and microphones) simply collect data and transmit it to the server, which then processes the information and returns the identified target device. This mediator approach allows high measurement precision without requiring each device to have complex processing capabilities.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system uses existing sensors (cameras and microphones) that are already part of the smart home infrastructure to perform the gaze and sound source detection functions. Rather than requiring specialized dedicated sensors, the system makes the existing sensors serve multiple purposes including voice command recognition, user identification, and device selection, thereby reducing overall system complexity.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If additional sensors like cameras or multiple microphones are installed to determine user gaze or sound source direction, then the accuracy of device selection for voice commands is improved, but the manufacturing cost increases

Engineering Contradiction:
Improveaccuracy of device selectionVSAvoidmanufacturing cost
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent makes existing sensors serve multiple functions. Cameras used for security or monitoring also perform gaze detection for device selection. Microphones used for general audio input also perform sound source direction determination. This multi-functionality approach allows high measurement precision without requiring specialized dedicated sensors, thereby reducing manufacturing costs.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The server acts as a centralized processing unit that receives data from distributed sensors and performs the computationally intensive tasks of gaze and sound source analysis. This allows the sensor network devices to remain simple and inexpensive while achieving high measurement precision through sophisticated server-side processing.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If multiple devices use the same voice command for simple operations like turning on switches, then the ease of operation is improved, but the reliability of device operation decreases due to undesired operations

Engineering Contradiction:
Improveease of operationVSAvoidreliability of device operation
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system provides feedback by determining the user's gaze direction and sound source direction to identify which device the user intends to control. Before executing the voice command, the system uses this spatial information to confirm the target device, providing a form of verification feedback that prevents undesired operations while maintaining simple voice-based operation.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent applies different weighting or attention to different spatial locations relative to the user. The device in the direction of the user's gaze or sound source is identified as the intended target, while other devices are excluded. This local quality approach ensures that the same voice command operates different devices depending on the user's spatial context, preventing undesired operations.

Inventive Principle:
Principle #3Local quality

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enhances user intention recognition and prevents undesired device operations by using location and direction information, providing a cost-effective and efficient solution for smart home and IoT applications.

Implementation Method 1

a microphone receiving a voice command from the user

Methodology Applied
Scientific EffectSound: Sound

Implementation Method 2

calculating relative locations and applying sound attenuation models

Methodology Applied
Scientific EffectSound attenuation: Acoustic Absorption

Data Source

PatentEP3754650B1Location-based voice recognition system through voice command
Publication Date: 2023.08.16 LUXROBO CORP
  • EP3754650B1 patent drawingFigure 1
  • EP3754650B1 patent drawingFigure 2
  • EP3754650B1 patent drawingFigure 3

AI summary

An object of the present invention is to facilitate recognition of a voice command of a user in a situation where multiple devices including microphones are connected through a sensor network. A relative location of each device is determined and a location and a direction of the user are tracked through a time difference in which the voice command is applied. The command is interpreted based on the location and the direction of the user. Such a method as a method for a sensor network, Machine to Machine (M2M), Machine Type Communication (MTC), and Internet of Things (IoT) may be used for an intelligent service (smart home, smart building, etc.), digital education, security and safety related services, and the like.