Text Data Sentiment Analysis Using Predefined Word Lists
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems fail to effectively extract meaningful insights from unstructured text-based data, leading to obsolete customer and employee opinions, and are limited by the need for large training datasets, making them unsuitable for niche areas and biased towards click and like metrics.
Innovation Solution
A computer-implemented method using natural language processing algorithms to categorize and analyze text-based data from various sources, focusing on category and sentiment words, and expressions to generate trends and insights with statistical significance, avoiding click and like metrics, and enabling analysis of smaller datasets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If standard machine-learning algorithms are applied to build dictionaries, then narrative analytics can process data, but huge amounts of training data are needed which limits usage to enormous data flows only
Solution Approach 1:
The patent changes the parameters of the analysis method by using predefined category words and sentiment words instead of building dictionaries through standard machine-learning algorithms. This allows the system to work with smaller datasets by changing the fundamental approach from data-intensive training to rule-based categorization and sentiment analysis.
Solution Approach 2:
The patent uses simple, predefined word lists for categories and sentiments rather than investing in complex, data-intensive dictionary building. These lightweight word lists can be quickly created and modified without requiring huge training datasets, enabling the system to work efficiently with smaller data volumes.
2Ease of operation
If surveys with specific questions or predefined scales are used, then data collection is structured, but the results are misleading and biased without revealing why someone clicked or liked content
Solution Approach 1:
The patent segments the analysis into two distinct parts: category classification (what the user is talking about) and sentiment analysis (how the user feels). This segmentation allows the system to process unstructured text data while extracting both the topic and the emotional context, revealing the reasons behind user behavior that structured surveys miss.
Solution Approach 2:
The patent uses category words and sentiment words as intermediaries to bridge the gap between unstructured user feedback and actionable insights. These intermediaries allow the system to interpret the underlying reasons for user behavior without requiring predefined survey questions or scales.
3Loss of time
If unstructured text-based data is analyzed in real-time, then actionable insights can be extracted, but complex data processing is required which makes results harder to understand
Solution Approach 1:
The patent extracts only the essential elements from unstructured text data: category words, sentiment words, and their relationships. By taking out only these key components rather than processing the entire complex text corpus, the system achieves real-time analysis while producing simple, understandable results in the form of categorized sentiments.
Data Source
AI summary
The disclosure relates to a computer implemented method of extracting typical quotes from text-based data. The method includes obtaining a plurality of text-based data from a plurality of users or sources, wherein each text-based data of the plurality of text-based data is provided by a user or a source from the plurality of users or sources, the text-based data is unstructured; categorizing the plurality of text-based data by at least determining, a typical category and a typical sentiment; wherein the typical sentiment is connected to the typical category; and determining a typical quote representing the plurality of text-based data based on the typical category and the typical sentiment.


