Automated Sentiment Training Data Construction via Dictionary Pre-labeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current electronic sentiment analysis technologies face challenges in automatically constructing accurate training sets from vast amounts of electronic communications, which are essential for training neural networks to analyze sentiments effectively.
Innovation Solution
A system and method that involve receiving electronic communications, segmenting them into smaller blocks, determining sentiment scores using a sentiment dictionary, aggregating scores to determine overall sentiments, and constructing training data for neural networks, allowing for the training and retraining of these networks based on user feedback.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual construction of training sets is used, then accuracy of sentiment analysis can be maintained, but productivity and efficiency deteriorate due to time-consuming manual labeling
Solution Approach 1:
The system performs preliminary sentiment analysis using a sentiment dictionary to pre-label electronic communications before they are used for training the neural network. This preliminary action creates an initial training set automatically, which is then refined through user feedback, thereby improving productivity while maintaining accuracy through iterative refinement
Solution Approach 2:
The system incorporates user feedback on automatically labeled communications to continuously improve the training set quality. Users can correct mislabeled sentiments, and these corrections are fed back into the system to retrain and refine the neural network, ensuring accuracy improves over time while maintaining high productivity
2Productivity
If automated training set construction is used, then productivity improves, but measurement precision deteriorates due to potential inaccuracies in automatic sentiment determination
Solution Approach 1:
The system dynamically adapts the training set construction process based on user feedback. Initially, it operates in automated mode for high productivity, then incorporates user corrections to dynamically refine the sentiment labeling accuracy, creating a evolving system that balances both productivity and precision
Solution Approach 2:
The system performs preliminary automated labeling to quickly generate training data, then uses user feedback as a refinement step. This two-stage approach maintains high productivity through automation while improving accuracy through selective human review and correction
3Quantity of substance
If large volumes of electronic communications are processed, then quantity of training data increases, but device complexity increases due to computational requirements
Solution Approach 1:
The system segments electronic communications into smaller units (individual communications, batches, or categories) and processes them in manageable portions. This segmentation allows the system to handle large volumes of data by dividing the computational task into smaller, more manageable steps, thereby reducing the complexity burden on the processing device
Solution Approach 2:
The system performs preliminary filtering and preprocessing of electronic communications before full neural network training. It uses sentiment dictionaries and basic analysis to pre-process data, eliminating obviously irrelevant communications and preparing data in advance, which reduces the computational complexity of the main processing stage
Data Source
AI summary
Training data for training a neural network usable for electronic sentiment analysis can be automatically constructed. For example, an electronic communication usable for training the neural network and including multiple characters can be received. A sentiment dictionary including multiple expressions mapped to multiple sentiment values representing different sentiments can be received. Each expression in the sentiment dictionary can be mapped to a corresponding sentiment value. An overall sentiment for the electronic communication can be determined using the sentiment dictionary. Training data usable for training the neural network can be automatically constructed based on the overall sentiment of the electronic communication. The neural network can be trained using the training data. A second electronic communication including an unknown sentiment can be received. At least one sentiment associated with the second electronic communication can be determined using the neural network.


