Edge Data Profiling for Distributed Network Bandwidth Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Industrial production processes face challenges in efficiently processing large amounts of data across distributed devices, particularly in real-time analytics, due to communication failures and the management of heterogeneous data types, which leads to increased costs and model management overhead.
Innovation Solution
A method for efficient data processing in distributed networks involves collecting data from client devices, profiling it to determine characteristic information, and implementing data processing strategies such as aggregation, compression, and selection to prioritize and transmit data effectively to master devices, using a hierarchical or peer-to-peer topology, and employing data cohorting to reduce machine learning model complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If data is transmitted from all plant devices to master device without filtering, then data completeness is improved, but network bandwidth consumption increases and processing time increases
Solution Approach 1:
The patent applies preliminary action by performing data profiling and characteristic analysis at the edge devices (plant devices) before data transmission to the master device. This preprocessing identifies and filters data based on characteristics such as recency, variability, and significance, so that only relevant data is transmitted. This resolves the contradiction by maintaining data completeness for important information while reducing overall transmission volume and processing time.
2Loss of time
If data is filtered and processed at edge devices before transmission, then transmission time is reduced, but data quality may be compromised
Solution Approach 1:
The patent applies local quality by implementing data profiling and filtering rules that are specific to each data type and each plant device's local requirements. Different data characteristics (recency, variability, significance) are evaluated with different weights and thresholds tailored to local needs. This ensures that data quality is maintained according to specific requirements while still achieving transmission time reduction through selective filtering.
3Measurement precision
If machine learning models are trained for each heterogeneous data type, then data analysis accuracy is improved, but model management overhead increases
Solution Approach 1:
The patent applies merging by consolidating data from multiple heterogeneous plant devices into unified data cohorts based on shared characteristics identified through profiling. Instead of training separate machine learning models for each data type, the system groups similar data patterns together and applies unified modeling approaches. This reduces the number of models required while maintaining analysis accuracy through characteristic-based segmentation.
4Loss of information
If all collected data is transmitted to master device, then data availability is improved, but network bandwidth consumption increases
Solution Approach 1:
The patent applies parameter changes by dynamically adjusting data transmission parameters based on profiling results. Characteristics such as data recency, variability, and significance are evaluated to determine transmission priority and format. This allows the system to maintain data availability for critical parameters while reducing bandwidth consumption by filtering or compressing less critical data, effectively changing transmission parameters based on data characteristics.
Data Source
AI summary
A system and method for enabling an efficient data processing in a distributed network of devices includes collecting first data from at least one client device sent by at least one plant device via a first communication interface; storing the first data in a data storage; profiling the first data to obtain characteristic information of the collected first data; determining a data processing strategy upon a result of the characteristic information of the first data; processing the first data according to the data processing strategy and deciding which data of the first data need to be sent to the master device via a second communication interface.

