Unsupervised Sentiment Lexicon Adaptation via Domain Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current sentiment lexicons fail to capture domain and context-dependent sentiment expressions, leading to poor coverage or precision in sentiment analysis due to their inability to adapt to specific contexts and domains.
Innovation Solution
An unsupervised method that updates a source lexicon by selecting a seed set of tokens, generating a candidate set from a target domain corpus using similarity parameters calculated through machine learning algorithms, and interpolating sentiment scores to adjust or add tokens, thereby creating a domain-specific sentiment lexicon.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If sentiment lexicons tag domain and context dependent expressions with overall polarity tendency based on statistics, then coverage is improved, but precision deteriorates
Solution Approach 1:
The patent applies local quality by creating domain-specific sentiment lexicons tailored to different contexts (e.g., technology, healthcare, finance) rather than using a single universal lexicon. Each domain lexicon captures the specific sentiment nuances of that domain, allowing the system to select the appropriate lexicon based on the input text's domain, thereby maintaining both broad coverage across domains and high precision within each domain.
Solution Approach 2:
The patent implements dynamics by making the sentiment lexicon adaptable and updatable. The system continuously learns from new data sources and updates domain-specific lexicons to reflect evolving sentiment expressions and domain-specific nuances over time, allowing the lexicon to dynamically adjust to new contexts and maintain precision as language and domains evolve.
2Measurement precision
If sentiment lexicons exclude domain and context dependent sentiment expressions, then precision is improved, but coverage deteriorates
Solution Approach 1:
The patent applies segmentation by dividing the sentiment lexicon into multiple domain-specific subsets (e.g., technology domain, healthcare domain, finance domain) rather than maintaining a single monolithic lexicon. This segmentation allows the system to exclude ambiguous expressions from general lexicons while capturing them in appropriate domain-specific lexicons, thereby maintaining precision in each domain while achieving comprehensive coverage across all domains through the collection of specialized lexicons.
3Ease of operation
If a comprehensive sentiment lexicon is constructed without domain adaptation, then ease of operation is improved, but adaptability deteriorates
Solution Approach 1:
The patent applies universality by creating a multi-functional sentiment analysis system that can operate across multiple domains using a standardized framework. The system maintains a unified architecture that automatically selects and applies the appropriate domain-specific lexicon based on the input text, providing ease of operation through consistent interface and workflow while achieving adaptability through the collection of domain-specific lexicons that can be applied to various domains as needed.
Data Source
AI summary
A method, system, and computer program product for unsupervised automated generation of lexicons in a specified target domain, comprising tokens having domain-specific sentiment orientation, by selecting a seed set of tokens from a source lexicon; generating a candidate set of tokens from a text corpus in the target domain based on a similarity parameter with the seed set; calculating a sentiment score for each of the tokens in the candidate set; and automatically updating the source lexicon based on the candidate list.

