Temporal Analysis of Social Media Corpora for Sentiment Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for tracking customer satisfaction through social media data lack the ability to effectively analyze changes in consumer opinion over time, providing only aggregate or net changes without tools for understanding sentiment at the lexical unit or phrase level.

Innovation Solution

The development of a system that performs temporal analysis of social media data by identifying and comparing lexical units and phrases across different time intervals, using a Volume Monitor and Periodic Analysis to visualize and filter changes, allowing for the identification of significant trends and topics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If aggregate or net change analysis is used to track customer satisfaction, then the analysis is simple and fast, but the ability to understand sentiment at the lexical unit or phrase level is lost

Engineering Contradiction:
Improveanalysis speedVSAvoidsentiment detail
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent segments the analysis process into multiple levels: aggregate-level temporal analysis for overall trends, and lexical-unit-level analysis for detailed sentiment understanding. This segmentation allows the system to provide both high-level productivity and detailed information retention by analyzing data at different granularities simultaneously.

Inventive Principle:
Principle #1Segmentation

2Loss of information

If temporal analysis at the lexical unit level is performed, then sentiment understanding is improved, but the complexity of the analysis system increases

Engineering Contradiction:
Improvesentiment detailVSAvoidanalysis system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent introduces a temporal dimension to the analysis by comparing corpora across different time periods. This allows the system to identify trends and changes over time without significantly increasing computational complexity, as the temporal comparison can be performed on aggregated statistics rather than individual lexical units.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The system applies different analysis depths to different parts of the data: aggregate-level analysis for overall sentiment trends and lexical-unit-level analysis only when detailed investigation is needed. This local quality approach optimizes resource allocation by performing detailed analysis only where necessary.

Inventive Principle:
Principle #3Local quality

3Loss of information

If detailed lexical unit analysis is performed on social media data, then consumer opinion changes are better understood, but the time required for analysis increases

Engineering Contradiction:
Improveopinion detailVSAvoidanalysis time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent performs preliminary temporal analysis at the aggregate level to identify significant changes and trends before conducting detailed lexical unit analysis. This preliminary action filters the data to focus only on relevant time periods and topics that require detailed examination, significantly reducing the overall analysis time while maintaining detailed sentiment understanding where needed.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10847144B1Methods and apparatus for identification and analysis of temporally differing corpora
Publication Date: 2020.11.24 NETBASE SOLUTIONS INC
  • US10847144B1 patent drawing
  • US10847144B1 patent drawing
  • US10847144B1 patent drawing

AI summary

Differences are identified, at the lexical unit and/or phrase level, between time-varying corpora. A corpus for a time period of interest is compared with a reference corpus. N-grams are generated for both the corpus of interest and reference corpus. Numbers of occurrences are counted. An average number of occurrences, for each n-gram of the reference corpus, is determined. A difference value, between number of occurrences in corpus of interest and average number of occurrences, is determined. Each difference value is normalized. N-grams can be selected for display, or for further processing, on the basis of the normalized difference value. Further processing can include selecting a sample period. A plurality of reference corpora are produced, where a begin time, for each sub-corpus of the plurality of reference corpora, differs, from a begin time for the corpus of interest, by an integer multiple of the sample period. Word Cloud visualization is shown.