Geolocating Social Media via Knowledge Base Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Social media data often lacks location coordinates, making it difficult to determine the source location of posts, which hinders analysis and pattern recognition in data mining applications.

Innovation Solution

A knowledge base is created using geolocated social media data to establish location-based clusters with representative keywords, allowing for the assignment of approximate geolocation to non-geolocated data by ranking and weighting keywords based on frequency, trustworthiness, and reliability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If geographic coordinates are required for social media data analysis, then location-based pattern recognition can be achieved, but the majority of social media content without location coordinates cannot be analyzed

Engineering Contradiction:
Improvelocation accuracyVSAvoidsocial media data coverage
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent introduces an intermediary system that uses keyword analysis and clustering to bridge the gap between geolocated and non-geolocated social media data. By extracting keywords from non-geolocated posts and matching them with clusters of geolocated posts, the system indirectly assigns location information without requiring direct GPS coordinates in the original data.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary clustering of geolocated social media data by geographic location and topic keywords before attempting to assign locations to non-geolocated data. This pre-processing creates a reference framework that enables subsequent location inference for data without explicit coordinates.

Inventive Principle:
Principle #10Preliminary action

2Loss of information

If keyword-based location inference is used for non-geolocated data, then data coverage is improved, but measurement precision of location decreases

Engineering Contradiction:
Improvesocial media data coverageVSAvoidlocation estimation accuracy
Core Design Contradiction:
Loss of informationVSMeasurement precision

Solution Approach 1:

The patent applies local quality by creating location clusters with specific geographic boundaries and assigning confidence levels to different regions within each cluster. Posts are assigned locations with varying degrees of precision based on their keyword match quality and the density of geolocated posts in the inferred region.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system changes parameters by introducing confidence scores and probability distributions for location estimates rather than providing single precise coordinates. This allows the system to represent uncertainty in location inference while still enabling data analysis.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If a knowledge base with clustered geolocated data is created, then location inference for non-geolocated data becomes possible, but system complexity increases

Engineering Contradiction:
Improvelocation inference capabilityVSAvoidknowledge base structure
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the social media data space into discrete location clusters, each characterized by geographic boundaries and topic keywords. This segmentation transforms the complex problem of continuous location inference into a manageable set of discrete clusters that can be efficiently queried and matched.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10191945B2Geolocating social media
Publication Date: 2019.01.29 FLORIDA INTERNATIONAL UNIVERSITY
  • US10191945B2 patent drawing
  • US10191945B2 patent drawing
  • US10191945B2 patent drawing

AI summary

Techniques for geolocating social media are described. According to an embodiment, information from textual content of a non-geolocated social media data item stored in a database is extracted. A knowledge database is then searched for a cluster of geo-located social media data items to which the information most closely relates, and an estimated location is assigned to the non-geolocated social media data item according to the cluster to which the information most closely relates. Each cluster comprises one or more representative tags for a spatio-temporal region. The knowledge database is created from geolocated social media data by grouping data according to location and extracting representative tags from the location's grouping of data according to textual content as well as information related to reliability and truthfulness of the textual content.