Infrastructure Processing Unit Local Service Failover
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In large-scale cloud native deployments, service failover techniques face challenges in handling transient failures of compute platforms, leading to inconsistent states or duplicate responses, which increase application latency and ownership costs.
Innovation Solution
The implementation of an infrastructure processing unit (IPU) and a switch that perform local service failover on compute platforms, utilizing service request replication and stale response discard techniques to ensure service level objectives are met, reducing the need for frequent failover and maintaining application reliability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If service failover techniques are implemented in large-scale cloud native deployments, then service reliability is improved, but application latency and ownership costs increase due to inconsistent states and duplicate responses
Solution Approach 1:
The patent applies preliminary action by having the infrastructure processing unit proactively monitor service requests before failures occur. The IPU tracks service request states in advance, maintains state information, and prepares for potential failures by pre-establishing monitoring mechanisms. This allows the system to detect failures early and respond quickly, reducing the latency impact of failover operations.
Solution Approach 2:
The patent implements feedback through the IPU's continuous monitoring of service request states and compute platform health. The system receives feedback about service request progression, detects when requests are lost or stuck, and uses this feedback to trigger appropriate failover actions. The feedback mechanism also helps identify duplicate responses, allowing the system to filter them out and maintain consistent state.
2Reliability
If service failover techniques are implemented in large-scale cloud native deployments, then service reliability is improved, but ownership costs increase due to inconsistent states and duplicate responses
Solution Approach 1:
The patent applies self-service by enabling the infrastructure processing unit to autonomously monitor service requests, detect failures, and execute failover decisions without requiring extensive external orchestration. The IPU independently tracks service request states, identifies when compute platforms fail, and redirects requests to healthy platforms. This self-service capability reduces the need for complex centralized management systems and lowers overall ownership costs.
Solution Approach 2:
The patent uses copying by maintaining state information about service requests in the IPU. Rather than requiring complete state replication across all compute platforms, the system copies and stores essential state information locally in the IPU, allowing it to make informed failover decisions and detect duplicate responses efficiently, thereby reducing the computational overhead and costs associated with full state synchronization.
3Productivity
If compute platforms operate under extreme conditions, then processing capacity is maintained, but service failure occurs due to processor and memory clocking oscillations
Solution Approach 1:
The patent applies beforehand cushioning by implementing the IPU as a protective layer between compute platforms operating under extreme conditions and the service request flow. The IPU monitors service request states and detects when compute platforms become unresponsive due to thermal throttling or other extreme condition effects. By cushioning against these failures through proactive monitoring and automatic failover, the system maintains service stability while allowing compute platforms to operate at high capacity under extreme conditions.
Data Source
AI summary
Example apparatus to perform service failover as disclosed herein are to detect a failure condition associated with execution of a service by a first compute platform, the execution of the service responsive to a first request. Disclosed example apparatus are also to send a second request to a second compute platform to execute the service. Disclosed example apparatus are further to monitor a queue of the first compute platform for a response to the first request, the response to indicate execution of the service by the first compute platform has completed, and when the response is detected in the queue, discard the response from the queue.


