Dynamic Speaker Profile Adaptation for Virtual Assistant Voice Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speaker identification systems for virtual assistants often produce unreliable results when users speak unnaturally during enrollment or in different acoustic environments, leading to inaccurate voice modeling.

Innovation Solution

The system generates a speaker profile by continuously monitoring audio inputs and using speech-to-text conversion to identify the speaker, with modes for building, modifying, and using static profiles, incorporating contextual information to enhance identification accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a predetermined enrollment process is used for speaker identification, then the user's voice can be effectively modeled, but the system produces unreliable results when the user speaks in an unnatural manner or in a different acoustic environment

Engineering Contradiction:
Improvespeaker identification reliabilityVSAvoidadaptability to different acoustic environments and speaking styles
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The speaker profile is made dynamic and continuously adaptable. Instead of a static profile created during enrollment, the system continuously monitors audio inputs and updates the speaker profile in real-time, allowing it to adapt to different acoustic environments and speaking styles while maintaining reliable speaker identification

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the parameters of the speaker profile over time by incorporating new audio data. The speaker profile evolves by integrating new acoustic characteristics and speaking patterns, enabling the system to maintain reliability across varying conditions through continuous parameter adjustment

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If the speaker profile is updated frequently to adapt to voice changes, then the system becomes more adaptable, but the complexity of the system increases

Engineering Contradiction:
Improveadaptability to voice changes over timeVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system performs self-updating of the speaker profile without requiring manual intervention. The continuous monitoring and automatic integration of new audio data allows the system to adapt to voice changes autonomously, reducing the complexity burden on users while maintaining high adaptability

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The speaker profile updates occur continuously in the background during normal operation rather than requiring discrete update operations. This continuous adaptation process integrates smoothly into system operation, maintaining adaptability while minimizing the perceived complexity for users

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentEP3201914B1Speaker identification and unsupervised speaker adaptation techniques
Publication Date: 2020.09.23 APPLE INC
  • EP3201914B1 patent drawingFigure 1
  • EP3201914B1 patent drawingFigure 2
  • EP3201914B1 patent drawingFigure 3

AI summary

Systems and processes for generating a speaker profile for use in performing speaker identification for a virtual assistant are provided. One example process can include receiving an audio input including user speech and determining whether a speaker of the user speech is a predetermined user based on a speaker profile for the predetermined user. In response to determining that the speaker of the user speech is the predetermined user, the user speech can be added to the speaker profile and operation of the virtual assistant can be triggered. In response to determining that the speaker of the user speech is not the predetermined user, the user speech can be added to an alternate speaker profile and operation of the virtual assistant may not be triggered. In some examples, contextual information can be used to verify results produced by the speaker identification process.