Data Integration System Using POLE Categorization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Integrating and harmonizing data from various sources with different formats and terminologies poses challenges, leading to fragmented insights and missed relationships when users access each source individually.
Innovation Solution
A data collection and integration system (DCIS) that maps and categorizes data using POLE categories (Person, Object, Location, Event) and stores it in a graph database, enabling unified access and visualization of relationships across multiple data sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If users access each data source individually to avoid integration complexity, then the complexity of data integration is reduced, but the user cannot see a full picture of the data and may miss valuable insights about relationships amongst the data
Solution Approach 1:
The patent introduces a data integration layer that acts as an intermediary between multiple data sources and the user. This layer includes components for data collection, normalization, storage, and presentation that mediate the interaction between diverse data sources with different formats and terminologies, allowing users to access integrated data without directly managing the complexity of each source
Solution Approach 2:
The patent creates a universal data access interface that can handle multiple types of data sources (databases, files, web services, etc.) through a single unified system. The data integration layer provides multi-functional capabilities to collect, normalize, store, and present data from various sources through consistent methods, eliminating the need for separate access mechanisms for each data source
2Loss of information
If data from various sources are integrated into a single location, then the user can see a full picture of the data and identify relationships, but the differences in format and terminology among data sources create technical challenges
Solution Approach 1:
The patent applies parameter changes by transforming data from various sources into a standardized format. The normalization component modifies data parameters (formats, terminologies, structures) to conform to a common schema, enabling integration while managing the complexity of format and terminology differences through systematic transformation rules
Solution Approach 2:
The patent segments the data integration process into distinct functional components: data collection from multiple sources, data normalization to standardize formats and terminologies, data storage in a unified repository, and data presentation to users. This segmentation allows each component to handle specific aspects of integration complexity independently
3Productivity
If a data integration system is implemented to provide unified access, then comprehensive data analysis and relationship identification are enabled, but the system complexity increases
Solution Approach 1:
The data integration layer serves as an intermediary that shields users from system complexity. It provides a simplified unified interface for data access while handling the complex operations of collection, normalization, storage, and presentation in the background, allowing users to benefit from comprehensive data analysis capabilities without directly managing system complexity
Solution Approach 2:
The patent creates a virtual copy or representation of data from multiple sources in a standardized format within the integration layer. This copied data structure maintains the relationships and insights from source systems while presenting them in a unified, manageable form that enables comprehensive analysis without requiring users to access or manage the original complex source systems
Data Source
AI summary
Disclosed herein are system, method, and computer program product embodiments for providing a data collection and integration system. An embodiment operates by determining that first data and second data retrieved from a first and second data sources are stored in a database. Both the first data and the second data are each categorized, and at least a portion of the first one of the categories includes identical information for both the first data and the second data. On a visual interface, a visual representation of the categorized first data is displayed simultaneously with the categorized second data, including the categorized identical information, and input indicating whether the identical information refers to a same entity is received. The database and the visual interface are updated based on the input.


