Merging Location Data Sets via Machine Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional approaches to merging location-based data sets are time-consuming and require manual intervention due to inconsistencies in data sources and standards, making it difficult to combine disparate data sets effectively.

Innovation Solution

A computer system transforms data sets into a standardized schema, uses a machine learning model to join and filter data sets based on geographical distances and match scores, enabling automated merging of location-based data sets across various types, such as company, place of interest, user position, and weather data sets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If manual ad hoc interventions are used to merge data sets, then data merging can be performed with flexibility, but the process becomes time-consuming and labor-intensive

Engineering Contradiction:
Improvemanual intervention flexibilityVSAvoiddata merging time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system performs automated data merging using machine learning models that self-learn from data patterns and automatically match records across different data sets without requiring manual intervention. The ML models autonomously handle schema alignment, record matching, and data fusion tasks.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Manual mechanical operations (human operators performing ad hoc merging) are replaced with automated computational systems using machine learning algorithms. The system substitutes human cognitive and manual labor with ML-based automated processes for data matching and integration.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If conventional approaches are used to merge data sets from disparate sources, then data can be combined, but inconsistencies in standards and schemas make merging difficult

Engineering Contradiction:
Improvedata source compatibilityVSAvoidmerging process complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system transforms data from different sources by changing their schema parameters to a unified standardized format. Machine learning models learn the mapping between various data schemas and automatically transform attributes, data types, and structures to ensure compatibility across disparate data sources.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

A standardized schema acts as an intermediary layer between disparate data sources. The system introduces this intermediate representation that mediates the integration process, allowing data from different sources with different standards to be merged through the common standardized interface.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If automated processing is implemented to reduce manual effort, then processing speed increases, but the complexity of the system increases

Engineering Contradiction:
Improvedata processing speedVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by pre-processing data to extract features and pre-align schemas before the actual merging operation. Machine learning models are pre-trained on sample data to learn matching patterns, enabling faster automated processing during production without requiring complex real-time decision-making.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11200239B2Processing multiple data sets to generate a merged location-based data set
Publication Date: 2021.12.14 THE WEATHER CO LLC
  • US11200239B2 patent drawing
  • US11200239B2 patent drawing
  • US11200239B2 patent drawing

AI summary

A computer system merges location-based data sets. Each of a plurality of data sets are transformed into a standardized schema, including at least two data sets including information indicating a geographic location. The schemas of the plurality of data sets are combined by data set type to produce a resulting data set for each data set type. The schemas of a first and second data sets are joined to produce a merged data set using a machine learning model to identify corresponding rows of the schemas. The schema of the merged data set is joined with the schemas of the resulting data sets for the data set types to produce a new data set. A resulting merged data set in the standardized schema is produced. Embodiments of the present invention further include a method and program product for merging location-based data sets in substantially the same manner described above.