Automated Geospatial Hypothesis Generation for Feature Engineering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Machine learning and big data analysis face challenges in effectively utilizing geospatial data due to the complexity and resource-intensiveness of feature engineering, particularly in automatically generating hypotheses and features that are relevant and non-redundant for predictive models.

Innovation Solution

A method for automatically generating hypotheses based on geospatial data by selecting auxiliary instances from a dataset based on geospatial relations, computing new attributes, and enriching labeled instances with these attributes to improve feature engineering and predictive model performance without requiring extensive domain knowledge.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual feature engineering is performed to create relevant features for machine learning models, then the quality and effectiveness of predictive models is improved, but the process becomes difficult, expensive, and time-consuming

Engineering Contradiction:
Improvepredictive model qualityVSAvoidfeature engineering time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs automated feature engineering by having the computer automatically select auxiliary instances, compute new attributes, and generate hypotheses based on geospatial relationships, eliminating the need for manual domain expert intervention in the feature creation process

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the manual mechanical process of feature engineering with an automated computational system that uses algorithms to select auxiliary instances, compute attributes, and generate hypotheses, substituting human effort with machine-based automation

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If many various features are created to prevent overfitting and improve model quality, then the model's predictive capability is improved, but the complexity of feature selection and processing increases

Engineering Contradiction:
Improvemodel predictive capabilityVSAvoidfeature selection complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system automatically evaluates the relevance of generated features through hypothesis validation, using feedback from the labeled dataset to determine which features are most relevant and should be included in the final model, thereby managing feature selection complexity

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent dynamically adjusts feature selection based on computed relevance metrics and hypothesis validation results, changing which features are included in the model based on their demonstrated predictive value rather than using a fixed feature set

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If domain knowledge is extensively used to create relevant features, then the relevance and accuracy of features is improved, but the cost and difficulty of feature engineering increases

Engineering Contradiction:
Improvefeature relevanceVSAvoidfeature engineering ease
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The system automatically discovers relevant features through automated hypothesis generation and validation, eliminating the need for extensive manual domain knowledge by having the computer system independently identify and validate feature relevance

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent introduces an automated hypothesis generation system as an intermediary between raw data and final features, using algorithmic processes to bridge the gap that previously required domain expert knowledge, thereby simplifying the feature engineering process

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11977987B2Automatic hypothesis generation using geospatial data
Publication Date: 2024.05.07 SPARKBEYOND
  • US11977987B2 patent drawing
  • US11977987B2 patent drawing
  • US11977987B2 patent drawing

AI summary

A method, apparatus and product for automatic hypothesis generation using geospatial data. A labeled dataset and an auxiliary dataset are obtained. Instances comprise geospatial attributes. Hypothesis generation is performed automatically based on the labeled dataset. For each labeled instance, one or more auxiliary instances are selected from the auxiliary dataset based on a geospatial relation between the geospatial attribute of the labeled instance and the geospatial attribute of the auxiliary instance. Based on the selected auxiliary instances, one or more new attributes are computed and added to the labeled instance.