Location-Based Voice Recognition for Multi-Device IoT Networks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice command systems face challenges in accurately determining the intended device for voice commands in multi-device setups, leading to undesired operations due to the use of additional sensors like cameras or multiple microphones, which increase costs and complexity.
Innovation Solution
A location-based voice recognition system that uses multiple microphones to determine the utterance direction of a user by calculating relative locations and applying sound attenuation models, allowing for accurate interpretation of voice commands without additional sensors.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If additional sensors like cameras or multiple microphones are installed to determine user gaze or sound source direction, then the accuracy of device selection for voice commands is improved, but the manufacturing cost and device complexity increase
Solution Approach 1:
The patent introduces a server as an intermediary that performs the complex calculations for determining user gaze direction and sound source direction. The sensor network devices (cameras and microphones) simply collect data and transmit it to the server, which then processes the information and returns the identified target device. This mediator approach allows high measurement precision without requiring each device to have complex processing capabilities.
Solution Approach 2:
The system uses existing sensors (cameras and microphones) that are already part of the smart home infrastructure to perform the gaze and sound source detection functions. Rather than requiring specialized dedicated sensors, the system makes the existing sensors serve multiple purposes including voice command recognition, user identification, and device selection, thereby reducing overall system complexity.
2Measurement precision
If additional sensors like cameras or multiple microphones are installed to determine user gaze or sound source direction, then the accuracy of device selection for voice commands is improved, but the manufacturing cost increases
Solution Approach 1:
The patent makes existing sensors serve multiple functions. Cameras used for security or monitoring also perform gaze detection for device selection. Microphones used for general audio input also perform sound source direction determination. This multi-functionality approach allows high measurement precision without requiring specialized dedicated sensors, thereby reducing manufacturing costs.
Solution Approach 2:
The server acts as a centralized processing unit that receives data from distributed sensors and performs the computationally intensive tasks of gaze and sound source analysis. This allows the sensor network devices to remain simple and inexpensive while achieving high measurement precision through sophisticated server-side processing.
3Ease of operation
If multiple devices use the same voice command for simple operations like turning on switches, then the ease of operation is improved, but the reliability of device operation decreases due to undesired operations
Solution Approach 1:
The system provides feedback by determining the user's gaze direction and sound source direction to identify which device the user intends to control. Before executing the voice command, the system uses this spatial information to confirm the target device, providing a form of verification feedback that prevents undesired operations while maintaining simple voice-based operation.
Solution Approach 2:
The patent applies different weighting or attention to different spatial locations relative to the user. The device in the direction of the user's gaze or sound source is identified as the intended target, while other devices are excluded. This local quality approach ensures that the same voice command operates different devices depending on the user's spatial context, preventing undesired operations.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enhances user intention recognition and prevents undesired device operations by using location and direction information, providing a cost-effective and efficient solution for smart home and IoT applications.
Implementation Method 1
a microphone receiving a voice command from the user
Implementation Method 2
calculating relative locations and applying sound attenuation models
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An object of the present invention is to facilitate recognition of a voice command of a user in a situation where multiple devices including microphones are connected through a sensor network. A relative location of each device is determined and a location and a direction of the user are tracked through a time difference in which the voice command is applied. The command is interpreted based on the location and the direction of the user. Such a method as a method for a sensor network, Machine to Machine (M2M), Machine Type Communication (MTC), and Internet of Things (IoT) may be used for an intelligent service (smart home, smart building, etc.), digital education, security and safety related services, and the like.