Knowledge Graph Generation from Siloed Data Using Transfer-Learned Ontologies
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large organizations face challenges in creating a holistic knowledge graph from siloed data sources due to unstructured data, different data formats, and disjoint schemas, which hinders effective decision-making in applications like servicing and controlling physical devices.
Innovation Solution
A computer-implemented method involving semantic analysis, natural language processing, and transfer learning to generate a knowledge graph by correcting and ranking ontologies based on data quality, using a centralized repository to integrate and analyze data from multiple sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If data from multiple isolated data sources are integrated to create a holistic knowledge graph, then the completeness and comprehensiveness of the knowledge graph is improved, but the complexity of data processing and integration increases
Solution Approach 1:
The patent introduces an intermediary system comprising a data lake, centralized repository, and ontology generation system that mediates between isolated data sources and the final knowledge graph. This intermediary layer standardizes and harmonizes data from multiple sources before integration, reducing the complexity of direct multi-source integration while preserving information completeness.
Solution Approach 2:
The integration process is segmented into distinct modular stages: data extraction from isolated sources, data cleaning and quality assessment, ontology generation and transfer learning, knowledge graph construction, and validation. Each stage handles specific tasks independently, making the overall complex process manageable and maintainable.
2Measurement precision
If semantic analysis and natural language processing are applied to analyze data, then the accuracy of knowledge extraction is improved, but the computational time and processing resources increase
Solution Approach 1:
The system performs preliminary data cleaning, quality assessment, and standardization before applying semantic analysis and NLP. By preparing and structuring data in advance, the subsequent semantic processing becomes more efficient and accurate, reducing overall computational time while maintaining extraction accuracy.
Solution Approach 2:
The system employs automated ontology generation through transfer learning that self-adapts to the specific data domain without requiring manual configuration. The model automatically learns from the data characteristics and adjusts its processing accordingly, optimizing the balance between accuracy and processing time for semantic analysis.
3Adaptability or versatility
If transfer learning is applied to generate candidate ontologies, then the adaptability to different data domains is improved, but the computational complexity of ontology generation increases
Solution Approach 1:
The system changes key parameters of the ontology generation process by using transfer learning with domain-specific pre-trained models. Instead of generating ontologies from scratch for each domain, the system adapts existing models by adjusting their parameters and knowledge representations to fit new domains, reducing computational complexity while maintaining high adaptability.
Solution Approach 2:
The system creates candidate ontologies by copying and adapting from pre-existing ontology structures and knowledge representations. Rather than inventing new ontological frameworks for each domain, it replicates and customizes proven ontology patterns, simplifying the generation process while ensuring domain appropriateness.
4Reliability
If data quality correction steps are applied to low quality data, then the reliability of the knowledge graph is improved, but the processing time and computational resources increase
Solution Approach 1:
The system implements a feedback mechanism where data quality is continuously assessed and corrected iteratively. Quality metrics are computed, and correction steps are applied selectively to data falling below thresholds. The process monitors improvement and adjusts processing intensity accordingly, ensuring reliable output while minimizing unnecessary processing time for already-high-quality data.
Data Source
AI summary
Some embodiments relate to a computer-implemented method and system, wherein the method includes generating a knowledge graph from a plurality of isolated data sources. The method includes reading data from the plurality of isolated data sources; analysing the data using semantic analysis and natural language processing vectorisation and obtaining a first knowledge graph ontology. An output from the analysis can be updated data processed based on data quality. A category of low quality data can include data having a data quality score below a predefined threshold. The method can include accessing the first knowledge graph ontology; obtaining information related to one or more previously completed knowledge graphs and ontologies; applying transfer learning to generate new candidate ontologies; utilising ranking scores to select a final ontology from the candidate ontologies; and generating a knowledge graph using the selected final ontology.


