Geolocating Social Media via Knowledge Base Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Social media data often lacks location coordinates, making it difficult to determine the source location of posts, which hinders analysis and pattern recognition in data mining applications.
Innovation Solution
A knowledge base is created using geolocated social media data to establish location-based clusters with representative keywords, allowing for the assignment of approximate geolocation to non-geolocated data by ranking and weighting keywords based on frequency, trustworthiness, and reliability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If geographic coordinates are required for social media data analysis, then location-based pattern recognition can be achieved, but the majority of social media content without location coordinates cannot be analyzed
Solution Approach 1:
The patent introduces an intermediary system that uses keyword analysis and clustering to bridge the gap between geolocated and non-geolocated social media data. By extracting keywords from non-geolocated posts and matching them with clusters of geolocated posts, the system indirectly assigns location information without requiring direct GPS coordinates in the original data.
Solution Approach 2:
The system performs preliminary clustering of geolocated social media data by geographic location and topic keywords before attempting to assign locations to non-geolocated data. This pre-processing creates a reference framework that enables subsequent location inference for data without explicit coordinates.
2Loss of information
If keyword-based location inference is used for non-geolocated data, then data coverage is improved, but measurement precision of location decreases
Solution Approach 1:
The patent applies local quality by creating location clusters with specific geographic boundaries and assigning confidence levels to different regions within each cluster. Posts are assigned locations with varying degrees of precision based on their keyword match quality and the density of geolocated posts in the inferred region.
Solution Approach 2:
The system changes parameters by introducing confidence scores and probability distributions for location estimates rather than providing single precise coordinates. This allows the system to represent uncertainty in location inference while still enabling data analysis.
3Adaptability or versatility
If a knowledge base with clustered geolocated data is created, then location inference for non-geolocated data becomes possible, but system complexity increases
Solution Approach 1:
The patent segments the social media data space into discrete location clusters, each characterized by geographic boundaries and topic keywords. This segmentation transforms the complex problem of continuous location inference into a manageable set of discrete clusters that can be efficiently queried and matched.
Data Source
AI summary
Techniques for geolocating social media are described. According to an embodiment, information from textual content of a non-geolocated social media data item stored in a database is extracted. A knowledge database is then searched for a cluster of geo-located social media data items to which the information most closely relates, and an estimated location is assigned to the non-geolocated social media data item according to the cluster to which the information most closely relates. Each cluster comprises one or more representative tags for a spatio-temporal region. The knowledge database is created from geolocated social media data by grouping data according to location and extracting representative tags from the location's grouping of data according to textual content as well as information related to reliability and truthfulness of the textual content.


