Network I/O Device Latency Timestamping for Data Center Troubleshooting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In large data centers, identifying the source of performance slowdowns is challenging due to complex mappings of services to underlying server hardware, making traditional troubleshooting methods time-consuming and inefficient, especially in virtualized cloud computing environments where no obvious correlation exists between Service Level Agreements (SLAs) and physical/virtual infrastructure.
Innovation Solution
Implementing a network I/O device with circuitry capable of time-stamping ingress/egress packets and capturing portions of these packets to determine transaction latency values, which are then reported to management logic for analysis, allowing for faster identification of performance issues across multiple servers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional troubleshooting methods are used to identify performance issues in data centers, then thorough analysis of application logs and server resource metrics is performed, but the time required to identify problematic servers increases significantly (up to several days)
Solution Approach 1:
The patent implements preliminary action by proactively monitoring and recording transaction latency metrics between servers before performance degradation occurs. The system continuously collects baseline latency data and establishes normal performance patterns, enabling rapid anomaly detection when issues arise. This preliminary monitoring infrastructure eliminates the need for time-consuming retrospective analysis of application logs and server metrics, as the problematic servers are already identified through pre-captured latency measurements.
2Loss of information
If complex mapping of services to underlying virtual and physical infrastructure is analyzed to identify the offending server, then comprehensive service level agreement correlation is achieved, but the device complexity and analysis difficulty increase significantly
Solution Approach 1:
The patent introduces an intermediary approach by implementing a dedicated monitoring system that sits between the complex service-infrastructure mapping and the performance analysis process. This intermediary monitoring infrastructure directly measures transaction latencies between server pairs, capturing performance data independent of the complex virtualization mapping. The system correlates SLAs with physical infrastructure through these direct latency measurements rather than through complex service mapping analysis, simplifying the identification process while maintaining comprehensive correlation capability.
3Measurement precision
If additional levels of logging are enabled at application, middleware or infrastructure layers to identify root-cause, then detailed performance data is collected, but the device complexity and data processing requirements increase
Solution Approach 1:
The patent extracts the performance monitoring function from the complex multi-layer logging infrastructure and implements it as a separate, dedicated monitoring system. Instead of enabling additional logging at application, middleware, and infrastructure layers, the system directly measures transaction latencies between server pairs through specialized monitoring agents. This extraction approach captures detailed performance data at the network transaction level, eliminating the need for multiple logging layers and their associated complexity while maintaining comprehensive performance visibility.
Data Source
AI summary
Examples are disclosed for determining or using server transaction latency information. In some examples, a network input/output device coupled to a server may be capable of time stamping information related to ingress request and egress response packets for a transaction. For these examples, elements of the server may be capable of determining transaction latency values based on the time stamped information. The determined transaction latency values may be used to monitor or manage operating characteristics of the server to include an amount of power provided to the server or an ability of the server to support one or more virtual servers. Other examples are described and claimed.


