Adaptive Speech Keyword Detection Using Activity Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech keyword detection systems in electronic devices face challenges in accurately recognizing voice keywords due to high false alarm and miss error rates, particularly in noisy environments, and lack adaptive sensitivity adjustment based on user activity.

Innovation Solution

A system that integrates a speech keyword detector, activity predictor, decision maker, and voice detector, utilizing sensor data from accelerometers, gyroscopes, and other sensors to predict user activities and adjust sensitivity dynamically, enabling more accurate voice keyword recognition by leveraging Dempster-Shafer theory or Gaussian mixture models for improved keyword detection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech keyword detection is performed continuously with high sensitivity, then keyword detection accuracy is improved, but false alarm rate increases

Engineering Contradiction:
Improvekeyword detection accuracyVSAvoidfalse alarm rate
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent dynamically adjusts the sensitivity threshold of speech keyword detection based on detected user activities. When activities such as phone lifting, shaking, or specific gestures are detected, the system temporarily increases sensitivity to capture keywords that might otherwise be missed. This dynamic adjustment resolves the contradiction by making the detection threshold adaptive rather than fixed, allowing high sensitivity only when contextually appropriate.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system uses sensor data (accelerometer, gyroscope, proximity sensor) to provide feedback about user behavior context. This feedback loop allows the system to understand when the user is likely to speak a keyword (e.g., when lifting the phone to ear) and adjust detection parameters accordingly. The feedback mechanism enables the system to distinguish between intentional keywords and random speech, reducing false alarms while maintaining high detection accuracy.

Inventive Principle:
Principle #23Feedback

2Reliability

If speech keyword detection is performed continuously, then keyword detection coverage is improved, but energy consumption increases

Engineering Contradiction:
Improvedetection coverageVSAvoidenergy consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

Instead of continuous speech keyword detection, the system employs periodic detection triggered by specific user activities detected through sensors. The detection is activated periodically when activities such as phone lifting, shaking, or specific gesture patterns are detected. This periodic action approach maintains detection coverage for important moments while dramatically reducing energy consumption during idle periods when keywords are unlikely to be spoken.

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The system performs preliminary activity recognition using low-power sensors (accelerometer, gyroscope, proximity sensor) before activating full speech keyword detection. This preliminary action allows the system to predict when a keyword might be spoken based on user behavior patterns, enabling energy-efficient selective activation of the speech detection module only when contextually appropriate.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If speech keyword detection sensitivity is increased to reduce miss errors, then detection accuracy is improved, but false alarm rate increases

Engineering Contradiction:
Improvedetection accuracyVSAvoidfalse alarm rate
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent applies different detection sensitivity levels to different contextual situations based on local user activity characteristics. When activities such as phone lifting to ear, shaking, or specific gestures are detected, high sensitivity is applied locally to those moments. For other situations, lower sensitivity is used. This local quality approach allows the system to optimize detection accuracy for keyword-prone situations without incurring false alarms in all situations.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically changes detection parameters (sensitivity threshold, detection window duration) based on detected user activities. When activities indicating likely keyword speech are detected, parameters are adjusted to increase sensitivity. When such activities are absent, parameters return to default lower sensitivity settings. This parameter changes mechanism allows the system to maintain high detection accuracy when needed while avoiding false alarms during normal operation.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9747894B2System and associated method for speech keyword detection enhanced by detecting user activity
Publication Date: 2017.08.29 MEDIATEK INC
  • US9747894B2 patent drawing
  • US9747894B2 patent drawing
  • US9747894B2 patent drawing

AI summary

The invention provides a system for speech keyword detection and associated method. The system includes a speech keyword detector, an activity predictor and a decision maker. The activity predictor obtains sensor data provided by a plurality of sensors, and processes the sensor data to provide an activity prediction result indicating a probability for whether a user is about to give voice keyword. The decision maker processes the activity prediction result and a preliminary keyword detection result of the speech keyword detection to provide a keyword detection result.