Neural Network Opinion Mining with Small-Data POS Vectors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional systems for opinion mining from free-form text responses require large datasets, leading to increased processing resources, memory needs, and complexity, while also struggling with domain specificity and out-of-vocabulary issues, resulting in inaccurate results and high costs.
Innovation Solution
The use of small-data training datasets that employ parts of speech and learned text response patterns to train a neural network, generating word-agnostic vectors to determine opinion, target, or irrelevant words, allowing for efficient and domain-agnostic opinion mining.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If large training datasets are used to train machine-learning algorithms for opinion mining, then the accuracy of extracting opinions from text responses is improved, but the processing resources, memory needs, and system complexity increase significantly
Solution Approach 1:
The patent changes the fundamental parameters of the training approach by using small datasets with carefully curated features (parts of speech, sentiment labels, target objects) instead of large datasets with raw text. This parameter transformation maintains extraction accuracy while dramatically reducing system complexity and resource requirements.
Solution Approach 2:
The patent extracts only the essential linguistic features (parts of speech, sentiment, target objects) from text responses, discarding unnecessary raw text data. This extraction approach achieves accurate opinion mining using minimal training data, resolving the contradiction between accuracy and complexity.
2Measurement precision
If large training datasets are used to train machine-learning algorithms, then the accuracy of opinion extraction is improved, but the memory and storage requirements increase
Solution Approach 1:
The patent transforms the training data from large volumes of raw text to small datasets of structured linguistic features. This parameter change reduces memory and storage requirements from tens of thousands of sentences to merely hundreds or thousands of annotated examples while preserving extraction accuracy.
Solution Approach 2:
The patent extracts only the critical linguistic components (parts of speech, sentiment, targets) needed for opinion extraction, eliminating the need to store and process large volumes of raw text data. This extraction methodology achieves high accuracy with minimal storage requirements.
3Adaptability or versatility
If large training datasets are used, then the coverage of vocabulary is improved, but the processing time and computational resources increase
Solution Approach 1:
The patent changes the approach from learning vocabulary semantics from large datasets to using pre-defined parts of speech categories and sentiment labels. This parameter transformation provides broad vocabulary coverage through linguistic generalization while maintaining fast processing speeds.
Solution Approach 2:
The patent performs preliminary classification of words into parts of speech and sentiment categories before opinion extraction. This preliminary action enables the system to handle diverse vocabulary efficiently without requiring extensive training on each word, reducing processing time while maintaining versatility.
4Measurement precision
If domain-specific training datasets are used, then the accuracy for that specific domain is improved, but the system cannot handle other domains effectively
Solution Approach 1:
The patent creates a universal opinion mining system that functions across multiple domains by using domain-agnostic linguistic features (parts of speech, sentiment, target objects). This universal approach maintains accuracy across different domains without requiring domain-specific training datasets, achieving both precision and versatility.
Solution Approach 2:
The patent inverts the conventional approach by not specializing the system for specific domains through domain-specific training data. Instead, it uses general linguistic features that apply universally across domains, achieving domain-agnostic accuracy through inversion of the specialization strategy.
5Measurement precision
If manual tagging and labeling of training data is performed, then the quality of training data is improved, but the cost and time required increases
Solution Approach 1:
The patent applies partial labeling by annotating only the essential linguistic features (parts of speech, sentiment, target objects) rather than comprehensively labeling all aspects of the text. This partial action achieves sufficient data quality for accurate opinion mining while dramatically reducing annotation time and cost.
Data Source
AI summary
The present disclosure relates to a response analysis system that employs a small-data training dataset to train a neural network that accurately performs domain-agnostic opinion mining. For example, in one or more embodiments, the response analysis system trains a response classification neural network using part of speech information (e.g., syntactic information) to learn and apply response classification labels for opinion text responses. In particular, the response analysis system employs part of speech information patterns without regard to word patterns to determine whether words in a text response correspond to an opinion, the target of the opinion, or neither. In addition, the trained response classification neural network has a significantly reduced learned parameter space, which decreases processing, memory requirements, and overall complexity.


