Sentiment Training Data Augmentation for Negation and Fairness
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing artificial intelligence-based chatbots face challenges in accurately handling negation and fairness in sentiment analysis due to inadequate training data, leading to incorrect sentiment predictions and perpetuation of biases.
Innovation Solution
The method involves data augmentation by creating negation pairs and fairness invariance data sets through rewriting examples with negation cues, sentiment prefixes/suffixes, and demographic word substitutions, followed by batch balancing to enhance model training.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If standard training data is used for machine learning models, then the model can be trained efficiently, but the model incorrectly predicts sentiment class frequently when training data is inadequate
Solution Approach 1:
The patent applies preliminary action by proactively generating synthetic training data with negation pairs and fairness invariance examples before model training. This involves creating augmented datasets with contradictory sentiment labels and demographic word substitutions in advance, so the model is pre-exposed to edge cases and bias scenarios that would otherwise require extensive manual data collection and annotation.
Solution Approach 2:
The patent implements parameter changes by systematically modifying training data parameters including sentiment polarity (through negation cues), demographic attributes (through word substitution), and label contradictions. These parameter transformations create diverse training scenarios that improve model robustness without requiring additional real-world data collection.
2Reliability
If the machine learning model is trained with more comprehensive training data to improve accuracy, then the model can handle negation and fairness better, but the training process becomes more complex
Solution Approach 1:
The patent applies self-service by implementing automated data augmentation systems that generate synthetic training examples without manual intervention. The system automatically creates negation pairs, applies demographic word substitutions, and generates fairness invariance datasets through algorithmic processes, eliminating the need for manual data annotation and reducing training complexity.
Solution Approach 2:
The patent uses parameter changes to systematically transform training data characteristics through automated processes. By programmatically applying negation cues, sentiment prefixes/suffixes, and demographic word substitutions, the system creates complex training scenarios through simple parameter transformations rather than complex manual data curation.
3Measurement precision
If the model is trained to be more accurate in sentiment classification, then the model can distinguish sentiment classes better, but the model may perpetuate biases and become less fair
Solution Approach 1:
The patent applies preliminary anti-action by proactively counteracting bias through fairness invariance training. The system generates training examples with demographic word substitutions that preserve sentiment labels, teaching the model to ignore demographic attributes. This preliminary anti-bias training prevents the model from learning biased patterns during standard sentiment classification training.
Solution Approach 2:
The patent implements parameter changes by systematically substituting demographic words while maintaining sentiment labels constant. This parameter transformation approach creates fairness-constrained training scenarios that teach the model to achieve accurate sentiment classification without relying on demographic information, thereby reducing bias perpetuation.
Data Source
AI summary
Techniques for augmentation and batch balancing of training data to enhance negation and fairness of a machine learning model. In one particular aspect, a method is provided that includes obtaining a training set of labeled examples for training a machine learning model to classify sentiment, searching the training set of labeled examples or an unlabeled corpus of text on target domains for sentiment examples having negation cues, sentiment laden words, words with sentiment prefixes or suffixes, or a combination thereof, rewriting the sentiment examples to create negated versions thereof and generate a labeled negation pair data set, and training the machine learning model using labeled examples from the labeled negation pair data set.


