Speech Synthesis Model Evaluation Using Voiceprint Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for evaluating personalized speech synthesis models are inefficient due to the need for individual verification of synthesized audio signals using voiceprint verification models, leading to low evaluation efficiency.
Innovation Solution
The method involves clustering voiceprint features to obtain central features and calculating cosine distances between them to evaluate the overall reproduction degree of audio signals, thereby increasing evaluation efficiency without relying on voiceprint verification models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If voiceprint verification models are used to evaluate synthesized audio signals one by one, then the evaluation can be performed individually, but the evaluation efficiency becomes low
Solution Approach 1:
The patent merges multiple individual voiceprint verification operations into a single batch processing operation. By clustering voiceprint features and calculating cosine distances between clustered representations, the system evaluates multiple audio signals simultaneously rather than one by one, thereby improving evaluation efficiency while maintaining accuracy
Solution Approach 2:
The patent changes the evaluation parameter from individual voiceprint verification to cosine distance calculation between clustered features. This parameter transformation enables batch processing of multiple audio signals, converting a sequential evaluation process into a parallelizable computation that significantly improves throughput
Data Source
Figure 1
Figure 2~3
Figure 4~5
AI summary
The present disclosure discloses a model evaluation method, a model evaluation device and an electronic device, and relates to the technical field of audio data processing. The model evaluation method includes obtaining M first audio signals synthesized by using a first to-be-evaluated speech synthesis model, and obtaining N second audio signals generated through recording; performing voiceprint extraction on each of the M first audio signals to obtain M first voiceprint features; performing voiceprint extraction on each of the N second audio signals to obtain N second voiceprint features; clustering the M first voiceprint features to obtain K first central features; clustering the N second voiceprint features to obtain J second central features; counting the cosine distances between the K first central features and the J second central features to obtain a first distance; and evaluating the first to-be-evaluated speech synthesis model based on the first distance. With the technical means of the present disclosure, the evaluation efficiency of the first to-be-evaluated speech synthesis model can be improved.