Entity Sentiment Summarization via N-gram Web Model
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for finding and summarizing user sentiment towards entities, such as brands or products, across the vast and noisy web data are inefficient and lack automation, as they rely on keyword searches and limited linguistic analysis, failing to provide reliable and compact summaries.
Innovation Solution
An entity summarization system that uses a controlled vocabulary list to scan web content, leveraging an N-gram web model for efficient data processing, to produce weighted lists of sentiment terms associated with entities, allowing for automated comparison and summarization of user sentiment across the internet.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If keyword-based search queries are used to find information about entities, then users can access web pages containing the search term, but the system cannot automatically summarize or analyze user sentiment towards entities
Solution Approach 1:
The patent replaces manual sentiment analysis with an automated NLP-based opinion mining system. The system uses computational algorithms to automatically extract, analyze, and summarize sentiment from web pages, reviews, and social media content, substituting human mechanical analysis with automated processing that can handle large-scale data efficiently
Solution Approach 2:
The system enables self-service sentiment analysis by automatically crawling, collecting, and analyzing web content without requiring manual intervention. The opinion mining process autonomously identifies sentiment-bearing text, extracts opinions about entities, and generates summaries, allowing the system to serve itself in the data collection and analysis pipeline
2Productivity
If manual linguistic analysis is performed on web content to understand user sentiment, then some sentiment information can be extracted, but the process is inefficient and cannot handle large-scale web data
Solution Approach 1:
The patent segments the sentiment analysis process into distinct modular components: web crawling module, text preprocessing module, opinion extraction module, and sentiment analysis module. Each module handles a specific task independently, allowing parallel processing and efficient pipeline execution that dramatically increases throughput while reducing overall processing time
Solution Approach 2:
The system performs preliminary actions by pre-processing web content during crawling, including text cleaning, normalization, and initial filtering of non-relevant content. This preliminary processing prepares data in advance for the opinion mining stage, reducing the computational burden during actual sentiment analysis and accelerating the overall process
3Extent of automation
If review websites display user reviews on entities, then users can read individual opinions, but the system cannot automatically summarize or compare sentiment across multiple entities
Solution Approach 1:
The patent implements a universal opinion mining framework that can analyze sentiment across multiple entity types (products, services, brands, persons) using the same core algorithms. The system extracts opinions about any target entity mentioned in web content, enabling automated comparison between entities without requiring entity-specific processing logic, thus managing complexity while maintaining versatility
Data Source
AI summary
An entity summarization system is described herein that mines the Internet and other data source to provide answers to questions such as the relative sentiment of users towards various brands. The system uses a controlled vocabulary list describing a specific aspect of entities of interest. Given an entity name, the system scans the whole content corpus to collect statistics on the words that occur most frequently in the context of the entity name, taking into account proximity information, to produce a weighted list of vocabulary terms describing the entity. Two entities can be compared by normalizing and comparing their weighted term lists. In some embodiments, the system performs these procedures efficiently by leveraging an N-gram web model. Thus, the system provides an automated way to compare two entities to derive information about how users feel about the entities at any given time.


