Anti-Teacher Learning for Paralinguistic Speech Under Conflicting Labels
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional paralinguistic information estimation models struggle with low accuracy due to the difficulty in identifying correct emotions in speech, especially when multiple listeners provide conflicting judgments, making it challenging to learn the characteristics inherent in those emotions.
Innovation Solution
The proposed solution involves using an anti-teacher decision unit to determine incorrect paralinguistic information based on listener judgments and an anti-teacher estimation model to learn from these incorrect labels, alongside a traditional teacher model to improve accuracy, and optionally incorporating multi-task learning to enhance the estimation process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional techniques use majority listener judgment to determine correct paralinguistic information, then a single correct label can be identified, but estimation accuracy decreases when listener judgments conflict
Solution Approach 1:
Instead of directly learning to identify correct paralinguistic information labels, the patent inverts the approach by learning to identify incorrect labels (anti-teachers). The anti-teacher estimation model is trained to recognize what is wrong with utterances, and the correct label is then determined by excluding the identified incorrect options. This inversion transforms a difficult direct identification problem into a more manageable exclusion problem, improving accuracy when listener judgments conflict.
2Quantity of substance
If the model learns from utterances with conflicting listener judgments, then more training data is available, but the characteristics of correct emotions become difficult to learn
Solution Approach 1:
The patent converts the harmful effect of conflicting listener judgments into a beneficial training resource. By treating utterances with conflicting judgments not as noise to be discarded but as valuable training data for identifying incorrect labels, the system transforms a problem into an opportunity. The anti-teacher model learns from these conflicts by identifying what makes certain labels incorrect, thereby improving overall estimation accuracy while utilizing the full training dataset.
3Measurement precision
If only correct paralinguistic information is used for learning, then the model focuses on accurate labels, but insufficient training data results in poor model learning
Solution Approach 1:
The patent makes the training system multi-functional by enabling the model to learn from both correct and incorrect labels simultaneously. The anti-teacher estimation model serves dual purposes: it identifies incorrect labels while implicitly learning the characteristics of correct labels through exclusion. This multi-functionality allows the system to utilize all available training data effectively, improving model learning efficiency without sacrificing label accuracy.
Data Source
AI summary
Paralinguistic information is estimated with high accuracy even when an utterance for which it is difficult to identify paralinguistic information is used for model learning. An acoustic feature extraction unit 11 extracts an acoustic feature from an utterance. An anti-teacher decision unit 12 decides, based on a paralinguistic information label indicating a determination result of paralinguistic information given by a plurality of listeners for each utterance, an anti-teacher label indicating an anti-teacher serving as incorrect paralinguistic information for the utterance. An anti-teacher estimation model learning unit 13 learns, based on an acoustic feature extracted from the utterance and the anti-teacher label, an anti-teacher estimation model for outputting a posterior probability of anti-teacher for an input acoustic feature.


