Unified Data Repository for CMDB via Automated CI Merging
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional Configuration Management Databases (CMDBs) face challenges with data accuracy and completeness due to duplicate entries, outdated information, and human errors, leading to unreliable CI relationship maps, which can impact IT Service Management processes.
Innovation Solution
A method and system that identify and merge Configuration Items (CIs) with the same attribute values from multiple normalized datasets using pre-defined prioritization rules to create a unified, accurate data repository, known as a golden dataset.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If manual updates from multiple data sources are used to populate CMDB, then data coverage is improved, but data accuracy deteriorates due to duplicate entries and human errors
Solution Approach 1:
The system performs self-service through automated data matching and merging processes. The processor automatically identifies duplicate CIs across multiple data sources using attribute comparison and merges them according to prioritization rules, eliminating the need for manual intervention while maintaining high data accuracy.
Solution Approach 2:
The system implements feedback mechanisms by continuously comparing CI attributes from multiple sources, identifying duplicates, and refining the unified dataset. The prioritization rules provide feedback on which data source to trust when conflicts arise, enabling iterative improvement of data quality.
2Loss of information
If manual population of CMDB is performed, then data completeness may be achieved, but reliability deteriorates due to omitted CIs and human errors
Solution Approach 1:
The system automatically discovers and imports CIs from multiple data sources through automated processes. The processor fetches CI data, normalizes attributes, and identifies duplicates without human intervention, ensuring consistent and reliable data population while maintaining completeness.
Solution Approach 2:
The system merges data from multiple sources by identifying CIs with the same attribute values across different datasets and consolidating them into a unified record. This merging process, governed by prioritization rules, ensures that no CI is omitted while eliminating duplicates, thereby improving both completeness and reliability.
3Measurement precision
If automated data processing is implemented to improve accuracy, then data quality improves, but system complexity increases
Solution Approach 1:
The automated data processing system is segmented into distinct functional modules: data fetching from multiple sources, normalization of CI attributes, duplicate identification through attribute comparison, and merging based on prioritization rules. This segmentation makes the complex process manageable and maintainable while achieving high data quality.
Solution Approach 2:
The system introduces an intermediary normalization layer that standardizes CI attributes from different data sources before comparison. This intermediary step simplifies the matching process by providing a common format for attribute comparison, reducing the overall system complexity while improving data quality.
Data Source
AI summary
A method for creating a unified data repository of clean and accurate data records is disclosed. In some embodiments, the method includes identifying one or more Configuration Items (CIs) with same attribute value from at least two of a set of normalized dataset. Each of the set of normalized dataset is generated from a plurality of CIs fetched from a plurality of data sources. The method further includes merging the one or more CIs identified with same attribute value from at least two of the set of normalized dataset, based on a set of pre-defined prioritization rules, to create a golden dataset of the clean and accurate data records.


