On-Device Speech Recognition Selective Activation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing automated assistants require explicit user invocation, such as speaking hot-words or performing specific user inputs, before processing spoken commands or queries, which increases response time and resource consumption.
Innovation Solution
Implementing client devices with on-device speech recognition, natural language understanding, and fulfillment capabilities, allowing for selective activation based on explicit or implicit cues, such as spoken hot-words, user presence, or directed speech, to reduce the need for explicit invocation and conserve resources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If explicit invocation (hot-words or specific user inputs) is required to activate automated assistant, then user privacy and resource consumption are controlled, but response time increases and user interaction becomes less efficient
Solution Approach 1:
The system performs preliminary actions by continuously monitoring audio inputs and maintaining readiness state without requiring explicit user invocation. The automated assistant pre-processes audio streams and prepares to execute commands, eliminating the need for hot-words or specific activation gestures, thereby reducing response time while maintaining privacy controls through selective processing.
2Loss of time
If on-device speech recognition and processing are continuously active, then response time is reduced and user interaction is improved, but device resources and energy consumption increase
Solution Approach 1:
The system dynamically adjusts its operational state based on contextual conditions. Speech recognition and processing components are activated only when specific conditions are met (such as detecting potential commands or user presence), rather than running continuously. This dynamic activation reduces energy consumption while maintaining fast response times for actual user commands.
Solution Approach 2:
The system changes operational parameters by adjusting the sensitivity and activation thresholds of speech recognition based on contextual factors. Processing intensity, audio monitoring sensitivity, and component activation levels are dynamically modified according to device state, user context, and environmental conditions, optimizing the balance between response time and energy consumption.
3Reliability
If all audio inputs are processed by automated assistant, then no commands are missed, but false activations increase and resource waste occurs
Solution Approach 1:
The system employs feedback mechanisms where speech recognition results are validated against contextual information, user profiles, and command patterns before triggering full automated assistant execution. Audio inputs are pre-filtered and classified, with only those meeting specific criteria advancing to full processing. This feedback loop reduces false activations while ensuring legitimate commands are reliably detected and executed.
Data Source
AI summary
Implementations can reduce the time required to obtain responses from an automated assistant by, for example, obviating the need to provide an explicit invocation to the automated assistant, such as by saying a hot-word/phrase or performing a specific user input, prior to speaking a command or query. In addition, the automated assistant can optionally receive, understand, and/or respond to the command or query without communicating with a server, thereby further reducing the time in which a response can be provided. Implementations only selectively initiate on-device speech recognition responsive to determining one or more condition(s) are satisfied. Further, in some implementations, on-device NLU, on-device fulfillment, and/or resulting execution occur only responsive to determining, based on recognized text form the on-device speech recognition, that such further processing should occur. Thus, through selective activation of on-device speech processing, and/or selective activation of on-device NLU and/or on-device fulfillment, various client device resources are conserved.


