Digital Assistant Speaker Profile Matching for User Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In multi-user environments, digital assistants face challenges in accurately identifying the current user of a shared electronic device, leading to inefficiencies and potential security breaches in providing personalized responses and managing media content.

Innovation Solution

The method involves receiving speaker profiles from external devices, comparing them to natural language speech inputs to determine the likelihood of matching with a specific user, and providing personalized responses only when the likelihood exceeds a certain threshold, ensuring accurate user identification and secure content management.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speaker profiles are compared to speech inputs to identify users in multi-user environments, then user identification accuracy is improved, but device complexity increases

Engineering Contradiction:
Improveuser identification accuracyVSAvoiddevice complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by collecting and storing speaker profiles in advance before actual user identification is needed. These profiles include various acoustic characteristics and are pre-processed and stored for quick comparison during runtime, eliminating the need for complex real-time analysis

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary component - the speaker profile database - that mediates between the speech input and user identification. This intermediary stores pre-analyzed acoustic characteristics and serves as a reference library, simplifying the identification process by comparing incoming speech against stored profiles rather than performing complex analysis from scratch

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If likelihood threshold comparison is performed to determine user matching, then security of personal information is improved, but processing time increases

Engineering Contradiction:
Improvesecurity of personal informationVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system changes parameters by establishing a likelihood threshold parameter that quantifies the confidence level required for user identification. This threshold parameter allows the system to balance security and processing time by adjusting the confidence level required, enabling quick decisions without complex real-time analysis

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If personalized responses are provided based on user identification, then user experience is improved, but energy consumption increases

Engineering Contradiction:
Improveuser experienceVSAvoidenergy consumption
Core Design Contradiction:
Ease of operationVSUse of energy by moving object

Solution Approach 1:

The system applies partial action by providing personalized responses only when user identification confidence exceeds the likelihood threshold. For low-confidence cases, the system uses generic responses, avoiding the energy cost of full personalization processing while maintaining good user experience for clear identification cases

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11657813B2Voice identification in digital assistant systems
Publication Date: 2023.05.23 APPLE INC
  • US11657813B2 patent drawing
  • US11657813B2 patent drawing
  • US11657813B2 patent drawing

AI summary

Systems and processes for operating an intelligent automated assistant are provided. An example method includes receiving, from one or more external electronic devices, a plurality of speaker profiles for a plurality of users; receiving a natural language speech input; determining, based on comparing the natural language speech input to the plurality of speaker profiles: a first likelihood that the natural language speech input corresponds to a first user of the plurality of users; and a second likelihood that the natural language speech input corresponds to a second user of the plurality of users; determining whether the first likelihood and the second likelihood are within a first threshold; and in accordance with determining that the first likelihood and the second likelihood are not within the first threshold: providing a response to the natural language speech input, the response being personalized for the first user.