Autonomous Sentiment Analysis Training Corpus Builder
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Manual selection and labeling of training samples for sentiment analysis computer models is time-consuming and labor-intensive, leading to lower confidence in sentiment determination, especially when dealing with large volumes of data.
Innovation Solution
An autonomous system that extracts semantic and sentiment elements from textual data, generalizes these elements, generates informative ranking scores, selects informative samples, and constructs a training corpus without human intervention, using components like named entity recognizers, feature term detectors, opinion word detectors, and semantic generalizers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual selection and labeling of training samples is used, then high quality training data can be obtained, but the process becomes time-consuming and labor-intensive
Solution Approach 1:
The system performs self-service by automatically selecting and labeling training samples without human intervention. The sentiment analysis model autonomously processes textual data, extracts sentiment elements, generates labels, and ranks samples for inclusion in the training corpus, eliminating the need for manual labor while maintaining data quality
Solution Approach 2:
The manual mechanical process of sample selection and labeling is replaced with an automated computational system. The patent substitutes human operators with a computer-based sentiment analysis model that uses natural language processing, semantic analysis, and machine learning algorithms to perform tasks previously requiring human judgment and effort
2Reliability
If manual selection of training samples is used, then expert support can ensure domain-specific accuracy, but the process requires significant human resources
Solution Approach 1:
The system achieves domain-specific reliability through self-service mechanisms where the automated model continuously learns from and adapts to domain-specific patterns in the data. The sentiment analysis system autonomously identifies domain-relevant features and adjusts its labeling criteria without requiring external expert intervention for each sampling decision
Solution Approach 2:
The patent creates a universal system that can handle multiple domains and types of textual data. The sentiment analysis model is designed to be domain-agnostic initially but can adapt to specific domains through automated learning, making it versatile for product reviews, social media, customer service, and other text-based sentiment analysis tasks without requiring domain-specific manual configuration
3Productivity
If random sampling of training data is used, then processing speed increases, but confidence in sentiment determination decreases
Solution Approach 1:
The system changes the parameter of sample selection from random to informed/ranked selection based on sentiment analysis metrics. By calculating sentiment scores, confidence levels, and diversity metrics for each sample, the system transforms the selection process into a parameter-driven optimization problem where samples are chosen based on their informational value rather than random chance
Solution Approach 2:
The patent introduces an intermediary ranking mechanism that mediates between the large volume of available data and the limited training corpus size. The sentiment analysis model acts as an intermediary layer that evaluates all candidate samples, assigns rankings based on multiple criteria (sentiment clarity, diversity, representativeness), and selects the top-ranked samples, thereby maintaining both speed and reliability
Data Source
AI summary
The methods, systems, and computer program products described herein provide ways to generate an informative training corpus of samples for use in machine training a high-quality sentiment analysis computer model. In some aspects, a method is disclosed including receiving a plurality of training samples, extracting semantic and sentiment elements of one or more of the training samples, generalizing the semantic and sentiment elements of the one or more of the training samples, generating an informative ranking score for one or more of the training samples based on the generalized semantic and sentiment elements, selecting informative training samples from the plurality of training samples based at least in part on the generated informative ranking scores, and adding the selected informative training samples to an informative training corpus.


