Social Streaming Topic Tracking With Online NMF for Disaster Footprints
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for tracking disaster-related topics in social media data are inefficient due to the massive amount of noisy content and computational overhead, failing to effectively identify common and distinct topics over time.
Innovation Solution
An online Nonnegative Matrix Factorization (NMF) technique combined with a joint NMF framework to efficiently update latent factors and balance reconstruction error, enabling the identification of common and distinct topics in social streaming data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If social media data is collected during disasters, then disaster-related topics can be tracked, but the amount of noisy and unwanted data increases
Solution Approach 1:
The patent extracts relevant disaster-related topics from the vast sea of social media data by using topic modeling techniques that identify and separate meaningful patterns from noise. The system extracts only the useful information (disaster-related discussions) while filtering out unwanted content such as spam and daily chatter, thereby resolving the contradiction between tracking topics and handling noisy data.
Solution Approach 2:
The patent introduces an intermediary processing layer using topic models as mediators between raw social media data and disaster response decisions. These topic models act as filters that transform the raw noisy data into structured topic representations, enabling effective topic tracking without being overwhelmed by the quantity of noisy data.
2Loss of time
If topic tracking is performed in real-time, then timely disaster response is enabled, but computational resources are consumed
Solution Approach 1:
The patent employs dynamic topic modeling that adapts to changing disaster scenarios in real-time. The system dynamically updates topic representations as new data arrives, allowing real-time tracking of evolving disaster topics without requiring exhaustive computational resources. This dynamic approach enables timely response while managing computational consumption through efficient incremental updates.
Solution Approach 2:
The patent performs preliminary topic identification and classification before full analysis is required. By pre-processing data to identify potential disaster-related topics and pre-establishing topic frameworks, the system reduces the computational burden during critical real-time response moments, thereby enabling timely response with lower computational resource consumption.
3Loss of information
If all social media data is processed, then complete topic coverage is achieved, but storage requirements increase
Solution Approach 1:
The patent extracts and stores only the essential topic representations rather than retaining all raw social media data. By extracting key topic patterns and discarding redundant raw data, the system achieves complete topic coverage while significantly reducing storage requirements. This extraction approach allows the system to maintain comprehensive topic information without the storage burden of processing all original data.
Solution Approach 2:
The patent discards raw social media data after extracting useful topic information, and recovers only the essential topic representations for storage and analysis. This selective retention strategy ensures that complete topic coverage is maintained through preserved topic models, while storage requirements are reduced by discarding the bulk of raw data that does not contribute to topic understanding.
Data Source
AI summary
Various embodiments for systems and methods of tracking disaster footprints using social streaming media using nonnegative matrix factorization are disclosed herein. The system extracts a summarization output from historical data and compares the summarization output with incoming data to identify differing or similar topics within the data. The summarization output is projected to adjust a time-dependency of the summarization output to enable a more direct comparison. The system additionally uses the summarization output to encode topic data within historical data to reduce computational overhead.


