Automated Data Exploration Using Knowledge Graph Inference
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data exploration and validation methods in computing systems are inefficient in extracting relevant information from heterogeneous data sources, leading to idle 'cold data' in databases that are not utilized for improving industrial operations, especially in domains like energy utilities where estimating signals such as electrical load or traffic flow is complex and often remains unused.
Innovation Solution
The implementation of automated data exploration and validation using a processor that generates optimal data flows based on an inference model, knowledge graph, and ontology of concepts, allowing for the abstraction of complexity, derivation of functional relations, and spatio-temporal alignment of raw data sources, with user feedback for updating the model and ranking data flows by confidence scores.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If automated data exploration using inference models and knowledge graphs is implemented, then data extraction efficiency from heterogeneous sources is improved, but system complexity increases
Solution Approach 1:
The patent introduces a knowledge graph as an intermediary structure that captures domain knowledge and relationships between heterogeneous data sources. This knowledge graph acts as a mediator between the inference model and raw data sources, enabling efficient data extraction without directly managing the complexity of individual data sources. The knowledge graph pre-structures relationships and constraints, allowing the system to query and extract data efficiently while the complexity is encapsulated in the knowledge graph construction phase rather than the query execution phase.
2Quantity of substance
If multiple heterogeneous data sources are integrated, then data completeness is improved, but data validation difficulty increases
Solution Approach 1:
The patent transforms data validation from a source-specific process to a unified parameter-based validation approach. By representing data from heterogeneous sources using common data models and schemas defined in the knowledge graph, the system changes the parameters of data representation to enable consistent validation rules. This allows data completeness to be improved by integrating multiple sources while validation difficulty is reduced through standardized parameter checking rather than source-specific validation logic.
Solution Approach 2:
The system implements feedback mechanisms where validation results from heterogeneous data sources are fed back into the inference model and knowledge graph. This feedback loop allows the system to learn from validation outcomes, refine data quality assessments, and improve future data integration. The feedback mechanism enables the system to handle increasing data completeness by systematically processing and learning from validation results rather than manually managing each data source's validation complexity.
3Loss of information
If domain knowledge is incorporated through ontology, then data relevance is improved, but processing time increases
Solution Approach 1:
The patent applies preliminary action by pre-processing and structuring domain knowledge into the knowledge graph before data extraction operations. The ontology and domain relationships are established in advance, creating a pre-computed framework that guides data extraction. This preliminary structuring of domain knowledge allows the system to quickly query and filter relevant data during execution without repeatedly processing the entire ontology, thus improving data relevance while minimizing processing time during actual data extraction operations.
Data Source
AI summary
Embodiments for automated data exploration and validation by a processor. One or more optimal data flows are provided in response to a query for one or more heterogeneous data sources according to an inference model based on a knowledge graph of heterogeneous data source relationships, a plurality of data flows between one or more heterogeneous data sources relating to the query, and an ontology of concepts and representing a domain knowledge of the one or more heterogeneous data sources.


