Speech Profile Merging and Splitting for Accurate Talker Enrollment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Active user enrollment for speech profile training is time-consuming and inconvenient, while automatic enrollment can lead to incorrect associations of speech with multiple or single talkers.
Innovation Solution
Systems and methods for automatically merging or splitting speech profiles based on similarity and difference metrics, using audio embeddings to improve accuracy and resource efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If active user enrollment is used for speech profile training, then speech profile accuracy is improved, but user convenience deteriorates due to time-consuming enrollment process
Solution Approach 1:
The system performs preliminary speech profile generation automatically without requiring user action. Audio embeddings are extracted from recorded conversations and speech profiles are generated in advance, so when transcription services are needed, the profiles are already ready to use, eliminating the need for time-consuming active enrollment while maintaining accuracy through pre-processed speech data
Solution Approach 2:
The system enables self-service by automatically generating speech profiles from recorded audio data without requiring user intervention. The profile manager autonomously extracts audio embeddings, compares them with existing profiles, and merges or creates new profiles based on similarity metrics, allowing the system to serve itself rather than requiring active user enrollment
2Ease of operation
If automatic user enrollment is used for speech profile generation, then user convenience is improved, but speech profile accuracy deteriorates due to incorrect associations of speech with talkers
Solution Approach 1:
The system implements feedback through similarity metric comparison. When generating speech profiles automatically, the profile manager compares audio embeddings against existing speech profiles using similarity metrics. This feedback mechanism ensures that speech is correctly associated with the appropriate talker by verifying similarity thresholds, preventing incorrect associations while maintaining automatic enrollment convenience
Solution Approach 2:
The system replaces manual verification mechanisms with automated acoustic analysis. Instead of requiring users to manually verify or correct speech profile associations, the system uses audio embedding extraction and similarity metric computation to automatically and accurately associate speech with the correct talker, maintaining both convenience and accuracy
3Measurement precision
If multiple speech profiles are maintained for different talkers, then speech profile accuracy is improved, but system complexity increases due to profile management overhead
Solution Approach 1:
The system merges speech profiles automatically when similarity metrics indicate they belong to the same talker. The profile manager consolidates audio embeddings from multiple sources into unified speech profiles, reducing the total number of profiles that need to be managed while maintaining accuracy by combining related speech data under a single profile identifier
Solution Approach 2:
The speech profile structure is designed to be universal and multi-functional. Each speech profile can contain audio embeddings from multiple sources and time periods, serving various transcription needs. This universal structure simplifies management by using a single profile type for all talkers rather than requiring different profile structures for different scenarios
Data Source
AI summary
A device includes a memory configured to store enrolled speech profiles. The device also includes one or more processors configured to obtain multiple audio embeddings representing speech that is identified as associated with a single talker in an audio stream. The one or more processors are also configured to determine a first speech profile based on the multiple audio embeddings. The one or more processors are further configured to determine a similarity metric based on a comparison of the first speech profile to a second speech profile of the enrolled speech profiles. The one or more processors are also configured to, based on the similarity metric, determine whether to combine the first speech profile and the second speech profile.


