Automated Data Model Generation for Heterogeneous Sources

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Combining data sets from different data file formats and fields stored in central and local sources is challenging due to complex data types and concatenations, requiring improved systems and methods for data modeling, especially for users with limited technical background in query and database management.

Innovation Solution

A computer-implemented method that receives a query, determines associated data elements, calculates ratings based on data set combinations, and generates a model for combining selected data sets, using a controller to weight and normalize ratings for effective data mapping and analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If data sets from different data file formats and fields are combined, then data completeness and analytical capability are improved, but data modeling complexity increases

Engineering Contradiction:
Improvedata completenessVSAvoiddata modeling complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system performs automated data analysis and self-evaluation without requiring user technical background in query and database management. The controller automatically determines data element associations, calculates ratings for data set combinations, and generates model proposals, enabling the system to serve itself in the data modeling process.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The controller acts as an intermediary between raw data sets and the final combined data model. It introduces a rating mechanism that evaluates potential data set combinations based on multiple properties, serving as a mediator that simplifies the complex process of data integration for users.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If automated data analysis is performed with minimal user interaction, then user effort is reduced, but system complexity increases

Engineering Contradiction:
Improveuser effortVSAvoidsystem complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system performs automated data analysis and self-evaluation without requiring user technical background in query and database management. The controller automatically determines data element associations, calculates ratings for data set combinations, and generates model proposals, enabling the system to serve itself in the data modeling process.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system changes the parameter of user interaction from manual technical operations to minimal query submission. By transforming the interaction model and automating the complex parameters of data element matching and combination evaluation, the system reduces user effort while managing its own increased complexity internally.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If multiple data sets with different fields and formats are integrated, then data utility is improved, but processing time increases

Engineering Contradiction:
Improvedata utilityVSAvoidprocessing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs preliminary analysis by evaluating multiple properties of data sets before final combination. The controller pre-calculates ratings for potential data set pairings based on their compatibility across different fields and formats, preparing the groundwork for efficient data integration without requiring exhaustive processing during the actual combination phase.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system segments the data integration process into distinct evaluation phases, assessing each data set's properties independently before determining combinations. By dividing the complex integration task into manageable rating calculations for individual data set pairs, the system processes diverse data formats more efficiently than attempting holistic analysis.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS9098550B2Systems and methods for performing data analysis for model proposals
Publication Date: 2015.08.04 SAP SE
  • US9098550B2 patent drawing
  • US9098550B2 patent drawing
  • US9098550B2 patent drawing

AI summary

In one embodiment, a computer-implemented method comprises receiving a query. A data store stores data as a plurality of data sets. Each data set comprises a plurality of fields and a plurality of data elements. Each field is associated with a portion of data elements. The query identifies selected data sets and selected properties of the selected data sets. For each selected property, the data elements of the selected data sets associated with each selected property are determined. A first rating of the determined data elements of the selected data sets is determined based on a type of combination of a pair of selected data sets. For the selected data set pairs, a second rating of the pair is determined based on the first ratings for the selected properties. A model of a combination of the selected data sets is generated based on the second rating.