Pseudo-Label Training for Robust Multitask Audio Quality Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing perceptual quality models lack robustness in predicting audio quality levels for diverse types of audio content and artifacts due to reliance on limited labeled training data, and speech-based models are inadequate for non-verbal sounds.
Innovation Solution
A multitask learning model is trained using unlabeled audio clips to jointly predict quality scores from multiple perceptual quality models, generating a trained model that accurately estimates perceived audio quality levels for diverse audio content and distortions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If perceptual quality models are trained using labeled audio clips from subjective listening tests, then the models can provide computationally efficient quality assessment, but the models lack robustness when encountering audio content or artifacts not represented in the training data
Solution Approach 1:
The system performs preliminary action by training multiple perceptual quality models on diverse audio content and artifact types before deployment. This pre-training with varied data ensures the models have already learned robust patterns for handling different audio scenarios, improving reliability when encountering unseen content during actual quality assessment operations
Solution Approach 2:
The system uses composite materials by combining multiple perceptual quality models with different training specializations into an ensemble system. Each model contributes its strengths from training on specific audio types, and their combined output provides more robust and reliable quality assessment across diverse audio content than any single model could achieve alone
2Measurement precision
If speech-based machine learning models are used for perceptual audio quality assessment, then the models can leverage specialized speech processing capabilities, but the models are unable to accurately predict quality levels for non-verbal sounds such as music and sound effects
Solution Approach 1:
The system applies universality by designing perceptual quality models that can handle multiple audio types including speech, music, and sound effects. The models are trained on diverse audio content rather than speech-only data, enabling them to universally assess quality across different audio domains while maintaining specialized capabilities for each type
Solution Approach 2:
The system uses segmentation by dividing the audio assessment task into specialized sub-tasks handled by different models or model components. Each model can be optimized for specific audio types (speech, music, effects) while the ensemble combines their results, allowing precise handling of speech content and accurate assessment of non-verbal sounds separately
Data Source
AI summary
In various embodiments, a training application trains a multitask learning model to assess perceived audio quality. The training application computes a set of pseudo labels based on a first audio clip and multiple models. The set of pseudo labels specifies metric values for a set of metrics that are relevant to audio quality. The training application also computes a set of feature values for a set of audio features based on the first audio clip. The training application trains a multitask learning model based on the set of feature values and the set of pseudo labels to generate a trained multitask learning model. In operation, the trained multitask learning model maps different sets of feature values for the set of audio features to different sets of predicted labels. Each set of predicted labels specifies estimated metric values for the set of metrics.


