Network Interface Device Failover for Time-Bounded Service Execution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data center architectures face challenges in managing complex, distributed applications with short-lived microservices due to the lack of effective tracking of component failures and resource utilization, leading to issues like CPU hangs and memory errors, especially in large-scale cloud environments.
Innovation Solution
Implementing a network interface device that monitors service execution progress and selectively reroutes tasks to optimize resource utilization across distributed nodes, using devices like NIC, SmartNIC, router, or DPU to manage workload distribution and ensure timely completion of services.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If application software handles distribution of work between multiple servers, then workload distribution is achieved, but tracking of component failures and resource utilization becomes complex and ineffective at large scale
Solution Approach 1:
The patent introduces a network interface device as an intermediary between the application software and the distributed servers. This intermediary automatically tracks service execution progress, monitors resource utilization, and manages failover decisions, eliminating the need for application software to directly handle complex tracking of component failures across millions of servers.
Solution Approach 2:
The network interface device autonomously monitors service execution progress and resource utilization without requiring external intervention. It automatically detects when services are stuck or resources are misallocated and triggers failover to alternative servers, enabling the system to self-manage workload distribution and failure tracking.
2Productivity
If microservice-based architectures are used, then service distribution and scalability are improved, but tracking execution progress and detecting failures becomes difficult
Solution Approach 1:
The network interface device continuously receives feedback about service execution progress from distributed microservices and uses this information to detect failures and trigger failover. The system monitors resource utilization metrics and execution status in real-time, creating a closed-loop control mechanism that automatically responds to service anomalies.
Solution Approach 2:
The network interface device serves as a mediator between microservices and the management system, centralizing the tracking of execution progress across distributed services. This intermediary aggregates status information from numerous short-lived microservices, making failure detection and progress monitoring feasible despite the distributed nature of the architecture.
3Productivity
If resources are allocated to executing services, then service execution is enabled, but resource misallocation and waste occur when services are stuck or delayed
Solution Approach 1:
The patent implements dynamic resource allocation where the network interface device continuously monitors service execution progress and automatically reassigns resources from stuck or delayed services to alternative services. This dynamic reallocation ensures that computing resources are always directed toward productive work, eliminating waste from services that cannot complete their tasks.
Solution Approach 2:
The system automatically detects when services are stuck or delayed and triggers failover to alternative servers without external intervention. This self-service mechanism ensures that resources are promptly reallocated from non-performing services, minimizing energy waste and maximizing overall system productivity.
4Productivity
If large-scale data centers are deployed, then processing capacity increases, but issues of CPU hangs, memory errors, and service failures become more frequent
Solution Approach 1:
The network interface device proactively monitors service execution progress and resource utilization to detect early signs of failures such as CPU hangs or memory errors. By identifying problematic services before they completely fail, the system can trigger failover in advance, cushioning against the impact of failures and maintaining service stability in large-scale data centers.
Solution Approach 2:
The system continuously monitors resource utilization and service execution status, using this feedback to detect patterns indicative of hardware or software failures. When anomalies are detected, the feedback loop triggers automatic failover to healthy servers, maintaining reliability despite the increased failure probability inherent in large-scale deployments.
Data Source
AI summary
Examples described herein relate to a network interface device that comprises circuitry, when operational, to select a platform to execute a function and based on load of the platform, selectively cause the function to execute on one or more other platforms to attempt to achieve or finish before the time-to-completion. In some examples, the circuitry is to detect progress of function execution to determine whether completion of execution of the function is predicted to not finish within the time-to-completion and cause the function to execute on one or more other platforms based on completion of execution of the function predicted to not finish within the time-to-completion. In some examples, the circuitry is to select the one or more other platforms to execute the function based on one or more of: processor computing utilization, available memory capacity, available cache capacity, network availability, or malfunction of a processor, memory, and/or cache.


