Virtual Assistant User Identity Identification via Facial and Speech Embeddings
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional virtual assistants lack the ability to identify users and their relationships, limiting their capacity to provide context-aware and personalized services to multiple users.
Innovation Solution
A virtual assistant system that uses facial and speech embeddings, generated by respective models, to determine user identities and relationships, enabling personalized actions based on input commands.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional virtual assistants are used, then the system is simple and easy to operate, but the assistant cannot identify user identity and relationships
Solution Approach 1:
The system segments user identification into multiple independent components: facial embedding model for visual recognition, speech embedding model for auditory recognition, and sound source localization model for spatial identification. Each component processes specific modalities separately before integrating results, which maintains manageability while achieving comprehensive user identification.
Solution Approach 2:
The virtual assistant system is designed with multi-functionality to handle various identification scenarios through a single unified system. It can identify users based on facial recognition, speech patterns, or sound source location, and can determine relationships between users, making the system universally applicable to different user identification needs without requiring separate specialized systems.
2Adaptability or versatility
If multiple sensors and models are integrated for user identification, then user personalization capability is improved, but the processing time and computational resources increase
Solution Approach 1:
The system performs preliminary actions by pre-extracting and storing facial embeddings, speech embeddings, and sound source location data during normal operation. When user identification is needed, the system queries these pre-processed data structures rather than performing complete analysis from scratch, significantly reducing processing time while maintaining comprehensive identification capability.
Solution Approach 2:
The system applies partial action by selectively activating only the necessary identification models based on the current situation. For example, if a user is already known from previous interactions, the system may only need to verify current presence rather than perform full facial and speech analysis, reducing processing time while maintaining sufficient identification accuracy.
Data Source
AI summary
A method may include receiving, by a virtual assistant of a user device, an input from a user, the virtual assistant being based on software. The method may include obtaining, by the virtual assistant of the user device and via a sensor of the user device, audio information or video information of the user. The method may include determining, by the virtual assistant of the user device, an identity of the user based on the audio information or the video information of the user and a set of facial embeddings and speech embeddings that is correlated with the user, the set of facial embeddings and speech embeddings being generated using a facial embedding model, a speech embedding model, and a sound source localization model. The method may include performing, by the virtual assistant of the user device, an action based on the input and the identity of the user.


