Cross-Correlation Alignment for Parallel Speech Quality Assessment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice conversion systems require manual and subjective quality assurance of parallel speech utterances, which is time-consuming and prone to errors due to differing accented pronunciations and speaker uniqueness, affecting the accuracy and reliability of the collected data and system output.
Innovation Solution
An automated system using cross-correlation and alignment techniques to assess parallel utterances objectively, incorporating acoustic and linguistic features, and employing neural networks to calculate a quality metric based on alignment and deviation from an ideal path, reducing manual inspection and improving accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual inspection is used to assess parallel utterances, then subjective quality evaluation can be performed, but the process becomes time-consuming and prone to errors
Solution Approach 1:
The patent replaces the manual mechanical inspection process with an automated computational system that uses cross-correlation algorithms and neural networks to objectively assess parallel utterances, eliminating human subjectivity and time consumption while maintaining or improving assessment accuracy
Solution Approach 2:
The patent introduces an intermediary automated quality assurance system that acts as a mediator between the parallel utterances and the final quality decision, using acoustic feature extraction, cross-correlation analysis, and alignment algorithms to objectively evaluate utterance quality without direct human intervention
2Reliability
If diverse candidate speakers are used to improve voice conversion quality, then output quality enhances, but the quality assurance process becomes more complex and expensive
Solution Approach 1:
The patent changes the parameters of quality assessment by using automated acoustic feature extraction and cross-correlation metrics instead of manual evaluation, enabling efficient handling of diverse candidate speakers through standardized, objective measurements that reduce complexity while maintaining quality standards
Solution Approach 2:
The patent applies local quality assessment by evaluating specific acoustic features and alignment characteristics of individual utterance segments rather than requiring comprehensive manual review of entire utterances, enabling efficient quality control across multiple candidate speakers
3Adaptability or versatility
If accented pronunciations are accommodated to improve speaker adaptability, then system versatility improves, but determining acceptance criteria becomes more difficult
Solution Approach 1:
The patent inverts the traditional approach by not requiring accented pronunciations to match reference utterances exactly, but rather measuring the deviation from alignment and using cross-correlation to objectively determine acceptable variations, thereby accommodating diversity while maintaining measurable acceptance criteria
Data Source
AI summary
The disclosed technology relates to methods, voice conversion systems, and non-transitory computer readable media for determining quality assurance of parallel speech utterances. In some examples, a candidate utterance and a reference utterance in obtained audio data are converted into first and second time series sequence representations, respectively, using acoustic features and linguistic features. A cross-correlation of the first and second time series sequence representations is performed to generate a result representing a first degree of similarity between the first and second time series sequence representations. An alignment difference of path-based distances between the reference and candidate speech utterances is generated. A quality metric is then output, which is generated based on the result of the cross-correlation and the alignment difference. The quality metric is indicative of a second degree of similarity between the candidate and reference utterances.


