Semantic Model for Automatic Data Annotation and Storage Optimization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Enterprise customers face challenges in accessing and leveraging data to its full potential due to cumbersome processes and the need for manual integration and annotation of new data sources, leading to suboptimal utilization, especially for unforeseen or long-tail use cases.
Innovation Solution
A holistic contextualization approach using semantic technologies and AI to analyze and integrate data sources and user behavior, generating a semantic model that automatically identifies relevant data and optimizes data integration strategies, including storage recommendations, for wide accessibility across the organization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual integration and annotation of new data sources are performed, then data accuracy and relevance are improved, but processing time and manual effort increase
Solution Approach 1:
The system performs automatic self-annotation of data sources using AI and semantic technologies. The contextualization engine automatically analyzes new data sources, extracts entities, relationships, and metadata without requiring manual intervention, thereby maintaining data accuracy while eliminating manual processing time
Solution Approach 2:
Manual annotation processes are replaced with automated AI-based semantic analysis. The system uses machine learning models to perform entity recognition, relationship extraction, and metadata generation, substituting human manual work with automated computational processes that maintain precision while reducing time consumption
2Adaptability or versatility
If comprehensive data integration is implemented, then data accessibility is improved, but system complexity increases
Solution Approach 1:
The patent introduces a contextualization engine as an intermediary layer between data sources and users. This engine automatically generates semantic models, entities, and relationships that simplify complex data integration, allowing comprehensive data accessibility without exposing users to underlying system complexity
Solution Approach 2:
The system segments data integration into modular components: data source connection, semantic analysis, entity extraction, relationship mapping, and contextualization. Each component handles specific tasks independently, reducing overall system complexity while enabling comprehensive data integration through coordinated modular operations
3Productivity
If automated semantic analysis is used, then processing speed is improved, but annotation precision may worsen
Solution Approach 1:
The system incorporates feedback mechanisms where annotation results are continuously evaluated and used to refine AI models. User interactions, correction inputs, and performance metrics feed back into the semantic analysis process, improving annotation precision over time while maintaining high processing speeds through automated operations
Solution Approach 2:
The system performs preliminary semantic analysis and pre-annotation before final data integration. This preliminary action allows the AI to prepare initial annotations at high speed, which are then refined through validation processes, achieving both fast processing and high precision through staged operations
Data Source
AI summary
The present disclosure involves systems, software, and computer implemented methods for contextualizing data to augment processes using semantic technologies and artificial intelligence. One example method includes identifying one or more data sources for semantic analysis. The data sources can include a data warehouse, a database, or a data lake. User behaviors of one or more users are identified for semantic analysis. The user behaviors include behaviors of how the users consume data in the one or more data sources. A semantic model is generated, using a knowledge graph, for the user behaviors. Nodes of the knowledge graph correspond to a class of entities in the data sources and are annotated with user behaviors and data source information. One or more queries from a user are monitored. Data from at least one node of the knowledge graph is recommended to the user, based on the semantic model and the queries.


