Edge Cluster Redundancy Across Racks for Self-Healing Resilience
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing cloud computing architectures face challenges in latency, availability, bandwidth usage, data privacy, network security, and the capacity to process large volumes of data in real-time, particularly for edge computing applications that require immediate processing and minimal latency.
Innovation Solution
Implementing resiliency and redundancy in edge computing devices through automated and redundant provisioning of management and workload clusters, utilizing machine learning (ML) and artificial intelligence (AI) models for self-healing capabilities to detect and remediate faults.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If centralized processing architecture is used, then data processing capacity is improved, but latency and bandwidth usage deteriorate
Solution Approach 1:
The patent segments the centralized data center into multiple distributed edge computing nodes deployed at different network locations. Each node independently processes data locally, eliminating the need to transmit all data to a central location. This segmentation resolves the contradiction by maintaining processing capacity through distribution while minimizing latency through local execution.
Solution Approach 2:
The patent transitions from a single-dimensional centralized architecture to a multi-dimensional distributed architecture by deploying computing nodes across multiple network locations and layers (edge, fog, cloud). This dimensional expansion allows data processing to occur closer to sources and consumers simultaneously, reducing latency while preserving processing capacity.
2Power
If centralized data center architecture is used, then processing power is improved, but network bandwidth usage and data privacy deteriorate
Solution Approach 1:
The patent segments the centralized processing power into distributed edge nodes, allowing data to be processed locally rather than transmitted across the network. This eliminates unnecessary bandwidth consumption while maintaining aggregate processing power across the distributed system.
Solution Approach 2:
The patent introduces edge computing nodes as intermediaries between data sources and centralized cloud infrastructure. These intermediaries process data locally, filtering and preprocessing information before selective transmission to the cloud, thereby reducing overall bandwidth usage while preserving processing capabilities.
3Loss of time
If edge computing nodes are deployed, then latency is improved, but system reliability and fault tolerance deteriorate
Solution Approach 1:
The patent implements local quality by equipping each edge computing node with redundant components and self-healing capabilities specific to its location. Each node maintains local backups and can independently detect and remediate faults, ensuring that latency-sensitive local operations remain reliable even if individual nodes experience issues.
Solution Approach 2:
The patent implements feedback mechanisms where edge nodes continuously monitor their own operational status and automatically trigger remediation actions when faults are detected. This closed-loop feedback system ensures that reliability is maintained at each distributed node while preserving the low-latency benefits of edge computing.
4Device complexity
If manual fault remediation is used, then system complexity is reduced, but downtime and productivity deteriorate
Solution Approach 1:
The patent implements self-service by enabling edge computing nodes to automatically detect, diagnose, and remediate their own faults without human intervention. The nodes execute self-healing routines that restore functionality, minimizing downtime while maintaining manageable system complexity through automated rather than manual processes.
Data Source
AI summary
Systems and techniques are provided for resiliency and redundancy for provisioning and/or configuring an edge compute unit. Configuration information can be obtained for provisioning an edge device with a plurality of nodes each associated with a respective rack of a plurality of racks. A first subset of the plurality of nodes can be provisioned, based on the configuration information, as a management cluster for workloads deployed to the edge device, the management cluster provisioned to include multiple redundant management control plane nodes distributed across different racks of the plurality of racks. A workload cluster can be provisioned on a remaining portion of the plurality of nodes, the workload cluster provisioned to include: multiple redundant workload control plane nodes distributed across different racks of the plurality of racks, and a respective plurality of worker nodes provisioned on each rack of the plurality of racks.


