Manufacturing Data Platform for Sequence-Dependent Aggregation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Aggregating and analyzing data across multiple manufacturing facilities in a network is challenging due to differences in security protocols, data sampling, formatting, and the sheer volume of data, as well as variations in processing equipment and configurations, which complicates the generation of secondary aggregated data values in a distributed computing environment.
Innovation Solution
A method and system for managing sequence-dependent data sets in a distributed computing environment, involving the pre-calculation of secondary aggregated data values based on data from multiple partitions, stored in a manufacturing data lake, allowing for centralized storage and analysis, and enabling the generation of tertiary aggregated data values through further processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If data is aggregated across multiple manufacturing facilities for unified analysis, then knowledge can be shared throughout the network, but differences in security protocols, data sampling, formatting, and equipment configurations complicate the aggregation process
Solution Approach 1:
The patent segments data into standardized formats at the source facilities before aggregation. Each facility maintains its own data collection systems but transforms data into a common standardized structure, allowing aggregation without requiring complex integration of heterogeneous systems. This segmentation approach resolves the contradiction by enabling data aggregation while avoiding the complexity of unifying diverse source systems.
Solution Approach 2:
The patent introduces an intermediary data standardization layer that sits between diverse manufacturing facilities and the centralized analysis system. This intermediary layer handles format conversion, sampling rate normalization, and security protocol mediation, allowing data aggregation across facilities without directly connecting complex heterogeneous systems. The intermediary resolves the contradiction by absorbing the complexity of integration while providing a clean interface for aggregation.
2Loss of information
If vast quantities of data from multiple facilities are aggregated for analysis, then comprehensive insights can be obtained, but the sheer volume of data increases processing time and computational resources required
Solution Approach 1:
The patent extracts and pre-processes data at the source facilities before aggregation, filtering out redundant information and extracting only the most relevant features for centralized analysis. This extraction approach maintains data completeness for insights while reducing the volume of data that requires centralized processing, thereby resolving the contradiction between comprehensive data retention and processing efficiency.
Solution Approach 2:
The patent performs preliminary data processing, aggregation, and feature extraction at distributed edge locations before sending processed data to the centralized system. This preliminary action reduces the data volume requiring centralized processing while preserving essential information, resolving the contradiction by advancing processing steps to earlier stages in the data pipeline.
3Productivity
If data is processed in a distributed computing environment across multiple nodes, then processing capacity is increased, but sequence-dependent data requires coordination between nodes which complicates the processing architecture
Solution Approach 1:
The patent segments sequence-dependent data processing into independent time-windowed batches that can be processed in parallel across distributed nodes. Each node processes a specific time segment independently, eliminating the need for complex inter-node coordination while maintaining the sequential nature of the data within each segment. This segmentation resolves the contradiction by enabling parallel processing of sequence-dependent data without requiring sophisticated coordination mechanisms.
Data Source
AI summary
Systems and methods are provided for handling sequence-dependent data as part of processing and/or analyzing large data sets in a distributed data processing environment. The distributed data processing environment can be suitable for handling data generated at a plurality of sites within a network of manufacturing sites. The systems and methods can allow for pre-processing of some values for sequence-dependent data. This can allow secondary aggregated values and/or secondary aggregated data sets to be generated from sequence-dependent data that can span multiple blocks or partitions. Pre-calculation of secondary aggregated values and/or secondary aggregated data sets for sequence-dependent data can allow the efficiencies of parallel or distributed computation to be at least partially retained while also allowing for desired processing of the sequence-dependent data.

