Hardware Link Failure Prediction via Signal Error Rate Monitoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems for predicting hardware component failures in hardware-based systems often react after a failure occurs, leading to inefficiencies and resource wastage, as they lack proactive methods to anticipate link failures, especially in large data centers where link instability can cause network instability and performance issues.

Innovation Solution

A method that determines whether a signal error rate for a given hardware link exceeds a threshold by comparing corrected and uncorrected error rates, using the formula abs(log(effective BER)−log(raw BER))<t, where t is a given value, and also considers the received signal-to-noise ratio, to generate an error indication predicting hardware component failure, and subsequently a recovery indication when the error rate normalizes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If existing systems react after failure occurs, then system simplicity is maintained, but system reliability deteriorates due to lack of proactive failure prediction

Engineering Contradiction:
Improvesystem reliabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by continuously monitoring signal error rates and comparing corrected versus uncorrected error rates before actual hardware failure occurs. This proactive detection mechanism identifies deteriorating link quality early, enabling preventive maintenance and avoiding complete link failures, thus improving reliability without requiring complex redundant systems

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements feedback mechanisms by continuously measuring signal error rates, comparing corrected and uncorrected error rates, and using this information to generate failure predictions. This closed-loop monitoring and comparison process provides ongoing feedback about link health, enabling the system to adapt and predict failures before they occur

Inventive Principle:
Principle #23Feedback

2Loss of time

If proactive failure prediction is implemented, then downtime is reduced, but device complexity increases due to additional monitoring and analysis mechanisms

Engineering Contradiction:
ImprovedowntimeVSAvoidmonitoring system complexity
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The system employs self-service mechanisms by utilizing existing corrected and uncorrected error rate data that are already generated by the hardware link. Instead of requiring external monitoring equipment, the system uses internally available performance metrics to predict failures, minimizing additional complexity while reducing downtime through proactive detection

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system applies partial monitoring by focusing specifically on the comparison between corrected and uncorrected error rates rather than monitoring all possible link parameters. This selective approach to monitoring captures the most critical failure indicators without requiring comprehensive surveillance of every system parameter, thus reducing complexity while effectively predicting failures

Inventive Principle:
Principle #16Partial or excessive action

3Loss of energy

If resource allocation for backups is reduced, then resource efficiency improves, but system reliability worsens due to lack of backup resources

Engineering Contradiction:
Improveresource allocation efficiencyVSAvoidsystem reliability
Core Design Contradiction:
Loss of energyVSReliability

Solution Approach 1:

The system performs preliminary failure prediction using corrected and uncorrected error rate comparisons, enabling proactive identification of deteriorating links before they fail. This allows resource allocation to be optimized based on actual predicted needs rather than maintaining constant backup resources, improving resource efficiency while maintaining reliability through targeted preventive actions

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses parameter changes in error rate metrics (comparing corrected versus uncorrected rates) to detect link deterioration. By monitoring changes in these parameters over time, the system can dynamically adjust resource allocation based on actual link health status, maintaining reliability only when and where needed rather than continuously allocating backup resources

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10997007B2Failure prediction system and method
Publication Date: 2021.05.04 MELLANOX TECHNOLOGIES LTD(IL)
  • US10997007B2 patent drawing
  • US10997007B2 patent drawing
  • US10997007B2 patent drawing

AI summary

A method including determining, for a given hardware link, whether a signal error rate for signals sent over the given hardware link is beyond a given threshold, when the signal error rate is beyond the given threshold, generating an error indication for the given hardware link, the error indication including a prediction that a hardware component associated with the given hardware link is likely to fail. Related apparatus and methods are also provided.