Automated Assistant Speaker Diarization for Concurrent Utterances

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current automated assistants struggle to differentiate between multiple simultaneous spoken utterances from different users, leading to incorrect actions, resource waste, and prolonged interactions due to unprocessed utterances.

Innovation Solution

An automated assistant that employs speaker diarization processes to distinguish between users' utterances, allowing selective response to intended requests while considering user verification and context, and rendering selectable elements for user selection to confirm actions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If the automated assistant processes multiple concurrent spoken utterances without differentiation, then the processing speed is improved, but the accuracy of request identification deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidrequest identification accuracy
Core Design Contradiction:
SpeedVSMeasurement precision

Solution Approach 1:

The patent segments the audio stream into distinct utterance segments associated with different speakers. The speaker diarization module divides the concurrent speech into separate segments, each tagged with speaker identity, allowing the system to process multiple utterances simultaneously while maintaining accurate identification of which utterance belongs to which speaker and which should be executed.

Inventive Principle:
Principle #1Segmentation

2Reliability

If the automated assistant uses fixed criteria to differentiate between utterances, then the reliability of processing is improved, but the adaptability to various speaking scenarios deteriorates

Engineering Contradiction:
Improveprocessing reliabilityVSAvoidadaptability to speaking scenarios
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic differentiation criteria that adapt to the specific context of each utterance. Rather than using fixed rules, the system dynamically determines which utterances to process based on speaker verification status, user profiles, and contextual analysis. This allows the system to reliably process the intended utterances while adapting to various speaking scenarios such as multiple authorized users, partial authorization, or unauthorized speakers.

Inventive Principle:
Principle #15Dynamics

3Loss of energy

If the automated assistant filters out unprocessed utterances, then the computational resources are saved, but the completeness of request processing deteriorates

Engineering Contradiction:
Improvecomputational resource consumptionVSAvoidrequest completeness
Core Design Contradiction:
Loss of energyVSLoss of information

Solution Approach 1:

The patent implements a feedback mechanism where the system provides verbal acknowledgment for processed utterances and can request clarification for unprocessed ones. When an utterance is not executed due to speaker verification failure or contextual mismatch, the system feeds back to the user to confirm whether the utterance was intended for the assistant, ensuring request completeness while still filtering out clearly unauthorized utterances to save computational resources.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250279094A1Processing concurrently received utterances from multiple users
Publication Date: 2025.09.04 GOOGLE LLC
  • US20250279094A1 patent drawing
  • US20250279094A1 patent drawing
  • US20250279094A1 patent drawing

AI summary

Implementations set forth herein relate to an automated assistant that is responsive to spoken utterances spoken by multiple different users simultaneously, or otherwise within a narrow window of time, in furtherance of initializing one or more actions. For instance, when a primary user is interacting with an automated assistant, a secondary user may also provide a spoken utterance. In response, the automated assistant can selectively determine whether any input from the secondary user should affect a request from the primary user. In some instances, the automated assistant can elect to either disregard the secondary input, separately respond to the secondary input, or use some amount of content of the secondary input to further the request from the primary user. Incorporation of secondary input when fulfilling a primary request can be fluidly performed to resemble human conversation and eliminate any unnecessary engagement between the automated assistant and the users.