Edge Machine Learning Inference for Distributed Data Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Analyzing and searching massive quantities of machine data generated from diverse sources in IT environments is challenging due to the complexity and heterogeneity of the data, requiring efficient processing and storage solutions that allow for real-time insights and flexible data analysis.
Innovation Solution
A data intake and query system that uses a flexible schema to process and store machine data as events with timestamps, enabling field-searchability and late-binding schema for extraction rules, allowing for real-time processing and analysis across disparate data sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If all relevant data is transmitted to centralized devices for complex processing and analysis, then comprehensive data analysis capability is improved, but data transmission time and network bandwidth consumption increase
Solution Approach 1:
The patent segments the centralized data processing architecture into distributed edge computing nodes. Each edge device performs local machine learning inference, dividing the processing workload across multiple locations rather than concentrating all data transmission to a single centralized server. This reduces network bandwidth consumption and transmission time while maintaining comprehensive analysis capability.
Solution Approach 2:
The patent introduces a new dimension of processing by deploying machine learning models at the edge devices themselves, rather than only at centralized servers. This spatial dimensionality change allows simultaneous local inference and centralized training, reducing the need for continuous data transmission while preserving analytical depth.
2Measurement precision
If machine learning models are updated frequently to improve accuracy, then model performance is improved, but computational overhead and energy consumption increase
Solution Approach 1:
The patent performs model training and updates in advance at centralized servers using accumulated data, then deploys the trained models to edge devices. This preliminary action separates the computationally intensive training phase from the inference phase, allowing frequent model improvements without requiring continuous high-energy computation at resource-constrained edge devices.
Solution Approach 2:
The patent introduces a model management intermediary layer that handles model training, validation, and deployment. This intermediary coordinates between data collection, model training at centralized servers, and model distribution to edge devices, optimizing the balance between model accuracy improvements and computational energy consumption across the distributed system.
Data Source
AI summary
Various implementations of the present application set forth a computer-implemented method comprising obtaining, by a low-power hub device, a first set of data published by an edge device, where the low-power hub device subscribes to at least a subset of data published by the edge device, generating, by the low-power hub device, a second set of data from the first set of data by inputting the first set of data into a machine learning (ML) model executing on the low-power hub device, and transmitting the second set of data to a remote server computer system.


