Sentiment Extraction System for Web Indexing Pipelines
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional sentiment extraction methods are unsuitable for web-based document-indexing pipelines due to high processing resource requirements and inefficiencies, making them time-consuming and incomplete for forming informed opinions about research subjects.
Innovation Solution
An automated system for extraction and summarization of sentiment information that accesses and processes sentiment data from various sources, identifies opinion categories, and generates a summarized graphical presentation for users, utilizing a network setting with a server, offline training system, and components like information accessor, sentiment extractor, opinion category identifier, and summarization generator.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If conventional sentiment extraction is performed as a batch process against multiple documents, then comprehensive sentiment information can be obtained, but processing resources are severely taxed and the process becomes unsuitable for integration into a document-level indexing pipeline
Solution Approach 1:
The patent segments the sentiment extraction task from the batch processing approach to a document-level processing approach. Instead of extracting sentiment information from multiple documents simultaneously in a batch process, the system performs sentiment extraction on each document individually as it is processed through the indexing pipeline. This segmentation allows the system to maintain comprehensive sentiment information extraction while avoiding the resource exhaustion problems of batch processing multiple documents at once.
2Loss of information
If conventional sentiment extraction processes a large corpus of documents, then sufficient knowledge about the research subject can be obtained, but the time required becomes very long and the process is time-consuming
Solution Approach 1:
The patent applies preliminary action by pre-processing and identifying sentiment-bearing portions of documents before performing full sentiment extraction. The system uses techniques such as identifying sentiment expressions, opinions, and relevant text segments in advance, then focuses the computationally intensive extraction process only on these pre-identified portions rather than processing entire documents. This preliminary identification step significantly reduces the time required while maintaining completeness of sentiment information.
3Loss of information
If manual extraction and collation of information is performed from identified information sources, then thorough research can be conducted, but the process requires significant manual effort and is very time-consuming
Solution Approach 1:
The patent implements self-service by enabling the system to automatically perform sentiment extraction and summarization without requiring manual intervention. The system autonomously identifies sentiment information, extracts relevant opinions, categorizes them, and generates summaries of sentiment regarding research subjects. This automated self-service approach maintains the thoroughness of research that would otherwise require manual extraction while completely eliminating the need for manual effort in the extraction and collation processes.
Data Source
AI summary
Methods and systems for extraction and summarization of sentiment information related to a particular research subject are disclosed. A method includes accessing sources of information that contain sentiment information that is related to the research subject and extracting the sentiment information from the sources of information as opinions related to the research subject. Opinion categories related to features of the research subject are identified. From this information a summarization of the sentiment information that is related to the particular research subject that includes the identified opinion categories is generated. Subsequently, access is provided to the summarization for graphical presentation.


