Explainable Emotion Prediction Using Counterfactual Vocal Samples
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current machine learning models, particularly in audio prediction, lack relatable explanations that are difficult for users to interpret, as they often rely on technical visualizations like spectrograms and saliency maps that are not semantically meaningful, making it hard for lay users or non-experts to understand why specific predictions are made.
Innovation Solution
A method and system that generate explainable predictions of emotions associated with vocal samples by using a Relatable Explanation Network (RexNet) which includes Contrastive Saliency, Counterfactual Synthetic, and Contrastive Cues explanations, leveraging neural networks and generative adversarial networks to provide relatable and human-interpretable explanations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If current explanation techniques present saliency maps on audiograms or spectrograms, then technical accuracy is maintained, but interpretability for lay users deteriorates
Solution Approach 1:
The patent introduces counterfactual synthetic vocal samples as an intermediary between the technical model and the user. These synthesized examples act as a bridge, translating complex model decisions into relatable audio comparisons that users can intuitively understand while preserving the technical accuracy of the original model predictions.
Solution Approach 2:
The system creates copies of the original vocal sample through generative adversarial networks to produce counterfactual examples. These synthetic copies maintain the structural and acoustic properties of the original while modifying specific features to illustrate why certain predictions were made, making technical explanations accessible without sacrificing precision.
2Ease of operation
If example-based explanations extract or produce examples for users to compare, then relatability improves, but semantic meaning deteriorates as humans must speculate why examples are similar or different
Solution Approach 1:
The system provides feedback by automatically generating and presenting counterfactual examples that directly illustrate the reasoning behind predictions. Instead of requiring users to speculate, the model actively generates explanatory examples with controlled modifications, providing semantic meaning through systematic feature manipulation while maintaining relatability through audio format.
3Loss of information
If spectrograms are used for visual explanation of audio, then technical detail is preserved, but user comprehension deteriorates since sound is not visual and spectrograms are ill-suited for lay users
Solution Approach 1:
The patent replaces the visual-mechanical representation (spectrograms) with an acoustic representation (synthetic vocal samples). This substitution allows users to comprehend explanations through their natural sense of hearing and intuition about voice, while the underlying technical detail is preserved through controlled manipulation of acoustic features in the counterfactual examples.
Data Source
AI summary
A method and a system for generating an explainable prediction of an emotion associated with a vocal sample are disclosed. The method includes receiving, by a processing device, a vector representation (z,) of an initial prediction (y0) of the emotion associated with the vocal sample (x), a counterfactual synthetic vocal sample (xY) associated with the vocal sample (x) and an alternate emotion (y) different from the initial prediction (y0) of the emotion, a vector representation (z,) of an emotion prediction (y0) associated with the counterfactual synthetic vocal sample (xy), vocal cue information (cy, cy) associated with the vocal sample (x) and the counterfactual synthetic vocal sample (xY) and attribution explanation information (iVc7) associated with relative importance of the vocal cue information (cy, cy) in prediction of the emotion. The method also includes determining, using the processing device, numeric cue differences (cyy) between the vocal cue information (cy) associated with the vocal sample (x) and the vocal cue information (cy) associated with the counterfactual synthetic vocal sample (xy), generating, using the processing device, cue difference relations information (r{circumflex over ( )}) based on the attribution explanation information (iv{circumflex over ( )}7), the numeric cue differences (cyy), the vector representation (z,) of the initial prediction (y0) and the vector representation (z,) of the emotion prediction (y0) associated with the counterfactual synthetic vocal sample (xY) using a first neural network (Mr), generating, using the processing device, a final prediction (y) of the emotion based on the numeric cue differences (cyy), the vector representation (z,) of the initial prediction (y0) and the vector representation (z,) of the emotion prediction (y0) associated with the counterfactual synthetic vocal sample (xY) using a second neural network (My), and generating, using the processing device, the explainable prediction of the emotion associated with the vocal sample (x) based on at least the counterfactual synthetic vocal sample (xY), the final prediction (y) of the emotion and the cue difference relations information (r{circumflex over ( )}).


