Unified Data Collection Pipeline for Transactional and Dimensional Metrics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data warehouse systems face challenges in maintaining optimal performance and collecting timely transactional and dimensional data due to the use of disparate sources and multiple pipelines, leading to out-of-date information and high processing costs for data joins.
Innovation Solution
A system that collects metadata and metric data from monitored servers using a single pipeline, transmitting them to a message bus for storage in a distributed database, where machine learning models analyze the data to generate performance indicators and predict potential issues, reducing the need for expensive joins and improving data accessibility.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple pipelines are used to collect data from disparate sources, then data collection coverage is improved, but system complexity and processing costs increase
Solution Approach 1:
The patent consolidates multiple data collection pipelines into a single unified pipeline that collects both transactional data and dimensional data simultaneously. This merging approach maintains comprehensive data collection coverage while reducing system complexity by eliminating redundant infrastructure and simplifying data integration processes.
Solution Approach 2:
The unified pipeline is designed to perform multiple functions: collecting transactional data, collecting dimensional data, and transmitting both types of data through a single communication channel. This multi-functional approach replaces multiple specialized pipelines, reducing overall system complexity while maintaining adaptability to various data sources.
2Adaptability or versatility
If multiple pipelines are used to collect data from disparate sources, then data collection coverage is improved, but processing costs increase
Solution Approach 1:
By merging multiple data collection operations into a single pipeline, the system eliminates redundant processing overhead associated with maintaining separate pipelines. This reduces computational resources consumed on data transmission and initial processing, thereby lowering overall processing costs while maintaining comprehensive data collection.
Solution Approach 2:
The system performs preliminary data preparation and packaging within the unified pipeline before transmission, ensuring that both transactional and dimensional data are ready for immediate use. This preliminary action reduces the need for expensive post-processing operations and data joins that would otherwise be required to integrate data from multiple separate sources.
3Quantity of substance
If data is collected from disparate sources using multiple pipelines, then data comprehensiveness is improved, but data timeliness deteriorates
Solution Approach 1:
The unified pipeline collects transactional and dimensional data simultaneously in real-time, eliminating the time delay associated with separate collection cycles. This merging approach ensures that both data types are captured together at the same moment, improving data timeliness while maintaining comprehensiveness.
Solution Approach 2:
The single pipeline operates continuously to collect both transactional and dimensional data without interruption or batching delays. This continuous operation ensures that data is captured in real-time as it becomes available, maintaining high timeliness while comprehensively covering all data sources through the unified collection mechanism.
4Measurement precision
If expensive joins are performed to integrate data, then data accuracy is improved, but processing costs increase
Solution Approach 1:
The system performs preliminary data preparation within the unified pipeline, organizing and tagging both transactional and dimensional data with appropriate metadata during the collection phase. This preliminary action reduces the complexity and cost of subsequent data integration operations, as data is already structured for efficient joining and correlation without requiring expensive post-processing operations.
Data Source
AI summary
Disclosed are methods, apparatuses and systems for collecting transactional and dimensional data. One implementation includes a configuration management database, a first data collection agent operating on a first monitored server to collect metadata associated with the first monitored server from the configuration management database using the first data collection agent, collect metric data from the first monitored server regarding operation of the first monitored server, wherein the metric data includes at least one metric collected from at least one application executing on the first monitored server; and assemble the collected metric data and at least part of the collected metadata into a packet for transmission from the first monitored server to a message bus; and a distributed database configured to receive the packet from the message bus and to store the collected metric data and metadata included in the packet.


