Job Analytics Aggregation Tool for Parallel Data Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems fail to efficiently process and upload large datasets within time constraints, and they lack the ability to gather comprehensive analytics for data processing and uploading jobs, leading to errors and failures in network node performance.
Innovation Solution
A data analytics tool that aggregates network node data to determine resource usage and identify malfunctioning nodes, generating job analytics by processing and uploading data in parallel batches across multiple network nodes, and a reporting tool that reduces processing power by creating single reports for multiple users.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data is transmitted within a network environment to perform ETL jobs, then data processing capability is improved, but computing resources such as memory, storage, network bandwidth, and CPU are consumed
Solution Approach 1:
The patent segments large datasets into smaller batches and distributes them across multiple network nodes for parallel processing. This allows the system to process larger volumes of data by dividing the workload, improving overall data processing capability while optimizing resource utilization across the network infrastructure
Solution Approach 2:
The patent creates a multi-functional analytics tool that can aggregate various types of network node data (CPU usage, memory usage, storage usage, network bandwidth usage) and perform multiple analysis functions. This universal tool improves productivity by providing comprehensive monitoring and analysis capabilities across different resource types without requiring separate systems for each function
2Reliability
If comprehensive analytics are gathered for data processing jobs, then error identification capability is improved, but processing time and resource usage increase
Solution Approach 1:
The patent implements preliminary action by proactively gathering and analyzing network node data before errors manifest. The analytics tool continuously monitors CPU usage, memory usage, storage usage, and network bandwidth usage, enabling the system to identify potential issues and take corrective action before they result in job failures or data loss
Solution Approach 2:
The patent implements feedback mechanisms where the analytics tool aggregates data from network nodes and provides insights back to the system operators. This feedback loop enables continuous improvement of system reliability by identifying patterns, detecting anomalies, and allowing operators to make informed decisions to prevent errors before they occur
3Reliability
If network node data is aggregated to identify malfunctioning nodes, then system reliability is improved, but data processing overhead increases
Solution Approach 1:
The patent introduces an intermediary analytics tool that acts as a mediator between network nodes and system operators. This tool aggregates data from multiple network nodes, processes the information centrally, and presents unified insights. The intermediary approach improves system reliability by providing comprehensive monitoring while reducing the complexity burden on individual network nodes and operators
4Speed
If parallel processing is used to upload data batches, then data upload speed is improved, but coordination complexity between network nodes increases
Solution Approach 1:
The patent segments data into multiple batches and assigns them to different network nodes for parallel upload. This segmentation strategy improves data upload speed by utilizing multiple network paths and nodes simultaneously. The analytics tool tracks and coordinates the progress of each batch, managing the complexity of parallel operations through systematic batch management and centralized monitoring
Data Source
AI summary
An analytics tool includes a network interface and an analytics engine. The network interface receives a request for job analytics of a job. The job comprises uploading a plurality of batches, each of the plurality batches comprising a subset of information of a data table. A network node of a plurality of network nodes uploads a batch of the plurality of batches. The analytics engine configured to determines the plurality of network nodes used to complete the job. The analytics engine retrieves network node data for each of the plurality of network nodes. The analytics engine generates the job analytics by aggregating the network node data for each of the plurality of network nodes.


