On-Device Speech Recognition Selective Activation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing automated assistants require explicit user invocation, such as speaking hot-words or performing specific user inputs, before processing spoken commands or queries, which increases response time and resource consumption.

Innovation Solution

Implementing client devices with on-device speech recognition, natural language understanding, and fulfillment capabilities, allowing for selective activation based on explicit or implicit cues, such as spoken hot-words, user presence, or directed speech, to reduce the need for explicit invocation and conserve resources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If explicit invocation (hot-words or specific user inputs) is required to activate automated assistant, then user privacy and resource consumption are controlled, but response time increases and user interaction becomes less efficient

Engineering Contradiction:
Improveresponse timeVSAvoidinvocation requirement
Core Design Contradiction:
Loss of timeVSEase of operation

Solution Approach 1:

The system performs preliminary actions by continuously monitoring audio inputs and maintaining readiness state without requiring explicit user invocation. The automated assistant pre-processes audio streams and prepares to execute commands, eliminating the need for hot-words or specific activation gestures, thereby reducing response time while maintaining privacy controls through selective processing.

Inventive Principle:
Principle #10Preliminary action

2Loss of time

If on-device speech recognition and processing are continuously active, then response time is reduced and user interaction is improved, but device resources and energy consumption increase

Engineering Contradiction:
Improveresponse timeVSAvoiddevice energy consumption
Core Design Contradiction:
Loss of timeVSUse of energy by moving object

Solution Approach 1:

The system dynamically adjusts its operational state based on contextual conditions. Speech recognition and processing components are activated only when specific conditions are met (such as detecting potential commands or user presence), rather than running continuously. This dynamic activation reduces energy consumption while maintaining fast response times for actual user commands.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes operational parameters by adjusting the sensitivity and activation thresholds of speech recognition based on contextual factors. Processing intensity, audio monitoring sensitivity, and component activation levels are dynamically modified according to device state, user context, and environmental conditions, optimizing the balance between response time and energy consumption.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If all audio inputs are processed by automated assistant, then no commands are missed, but false activations increase and resource waste occurs

Engineering Contradiction:
Improvecommand detection accuracyVSAvoidenergy wasted on false activations
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The system employs feedback mechanisms where speech recognition results are validated against contextual information, user profiles, and command patterns before triggering full automated assistant execution. Audio inputs are pre-filtered and classified, with only those meeting specific criteria advancing to full processing. This feedback loop reduces false activations while ensuring legitimate commands are reliably detected and executed.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12315508B2Selectively activating on-device speech recognition, and using recognized text in selectively activating on-device NLU and/or on-device fulfillment
Publication Date: 2025.05.27 GOOGLE LLC
  • US12315508B2 patent drawing
  • US12315508B2 patent drawing
  • US12315508B2 patent drawing

AI summary

Implementations can reduce the time required to obtain responses from an automated assistant by, for example, obviating the need to provide an explicit invocation to the automated assistant, such as by saying a hot-word/phrase or performing a specific user input, prior to speaking a command or query. In addition, the automated assistant can optionally receive, understand, and/or respond to the command or query without communicating with a server, thereby further reducing the time in which a response can be provided. Implementations only selectively initiate on-device speech recognition responsive to determining one or more condition(s) are satisfied. Further, in some implementations, on-device NLU, on-device fulfillment, and/or resulting execution occur only responsive to determining, based on recognized text form the on-device speech recognition, that such further processing should occur. Thus, through selective activation of on-device speech processing, and/or selective activation of on-device NLU and/or on-device fulfillment, various client device resources are conserved.