Distributed ML Training via Fog Computing Edge Coordination

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems face challenges in distributing machine learning model training to network edge devices due to computational resource constraints and security concerns related to telemetry data exposure, while also struggling to effectively monitor and control model performance across distributed environments.

Innovation Solution

A system and method for generating and deploying machine learning model architectures to network edge devices, allowing them to instantiate and train models locally, while centrally monitoring performance and controlling deployment through performance reports, without exposing telemetry data, using a machine learning structure controller that determines model updates based on performance reports.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If machine learning model training is distributed to network edge devices, then computational resource utilization is improved and data security is enhanced, but system complexity increases and monitoring difficulty arises

Engineering Contradiction:
Improvedata securityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system segments the machine learning training process into distributed execution at edge devices while maintaining centralized coordination through the fog platform. This allows computational tasks to be distributed across multiple edge devices, enhancing data security by keeping sensitive data local, while the segmented architecture manages complexity through modular task distribution and centralized oversight.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The fog computing platform acts as an intermediary between centralized cloud systems and distributed edge devices. It coordinates model training tasks, manages resource allocation, and facilitates communication without requiring direct complex interactions between all components, thereby reducing overall system complexity while enabling distributed training.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If machine learning model training is distributed to network edge devices, then computational resource utilization is improved, but monitoring and control of model performance becomes difficult

Engineering Contradiction:
Improvecomputational resource utilizationVSAvoidmodel performance monitoring
Core Design Contradiction:
ProductivityVSDifficulty of detecting and measuring

Solution Approach 1:

The system implements feedback mechanisms where edge devices report model training status and performance metrics back to the fog computing platform. This continuous feedback loop enables centralized monitoring and control of distributed training processes, allowing the system to track computational resource utilization and model performance across multiple edge devices effectively.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The fog computing platform provides universal monitoring and control capabilities that work across diverse edge devices and model types. It implements multi-functional services including task distribution, performance tracking, resource management, and model deployment, enabling comprehensive monitoring of distributed training regardless of specific device or model variations.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If telemetry data is sent from computing nodes to centralized locations for model training, then model training capability is improved, but data exposure security risks increase

Engineering Contradiction:
Improvemodel training capabilityVSAvoiddata exposure risk
Core Design Contradiction:
ProductivityVSObject-affected harmful factors

Solution Approach 1:

The system inverts the traditional centralized training approach by distributing model training execution to edge devices where the telemetry data resides. Instead of sending data to centralized locations for training, the training computations are brought to the data source, eliminating data transmission and exposure risks while maintaining model training capability through distributed computation.

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The system implements local model training execution at edge devices where telemetry data is generated and stored. By performing training computations locally rather than centrally, the system maintains high model training capability while ensuring data remains local, thus eliminating data exposure security risks associated with transmitting sensitive telemetry data to centralized locations.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11562176B2IoT fog as distributed machine learning structure search platform
Publication Date: 2023.01.24 CISCO TECHNOLOGY INC
  • US11562176B2 patent drawing
  • US11562176B2 patent drawing
  • US11562176B2 patent drawing

AI summary

Systems, methods, and computer-readable mediums for distributing machine learning model training to network edge devices, while centrally monitoring training of the models and controlling deployment of the models. A machine learning model architecture can be generated at a machine learning structure controller. The machine learning model architecture can be deployed to network edge devices in a network environment to instantiate and train a machine learning model at the network edge devices. Performance reports indicating performance of the machine learning model at the network edge devices can be received by the machine learning structure controller from the network edge devices. The machine learning structure controller can determine whether to deploy another machine learning model architecture to the network edge devices based on the performance reports and subsequently deploy the another architecture to the network edge devices if it is determined to deploy the architecture based on the performance reports.