HPSU Thermal Control Using Equivalent Reliability Time

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing thermal management methods for semiconductor devices in network switches fail to accurately estimate the equivalent reliability time (ERT) of hardware processing sub-units (HPSUs), leading to underperformance or overperformance, and do not efficiently adjust data traffic rates and cooling capacities based on real-time temperature profiles and ambient conditions.

Innovation Solution

A method that obtains and weights the operating-temperature profile of HPSUs over time, using the dependence of ERT on operating temperature to estimate an effective ERT, which is used to modify operating conditions such as data traffic rates and cooling capacities, ensuring optimal performance and reliability by comparing the effective ERT to a prespecified value and adjusting accordingly.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If existing thermal management methods are used to manage HPSU temperature, then cooling capacity is increased to ensure reliability, but HPSU performance is reduced due to overly conservative temperature limits

Engineering Contradiction:
ImproveHPSU reliabilityVSAvoidHPSU performance
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent changes the parameter from fixed temperature limits to dynamic equivalent reliability time (ERT) thresholds. By calculating effective ERT based on actual operating temperature profiles and weighting them according to reliability dependence, the system adjusts performance limits dynamically rather than using conservative static thresholds, thus improving productivity while maintaining reliability

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent implements feedback by continuously monitoring operating temperature profiles, calculating effective ERT, and comparing it to prespecified reliability thresholds. Based on this feedback loop, the system dynamically adjusts data traffic rates and cooling capacities to optimize both performance and reliability

Inventive Principle:
Principle #23Feedback

2Productivity

If data traffic rates are increased to improve network throughput, then productivity is improved, but HPSU temperature increases reducing reliability

Engineering Contradiction:
Improvenetwork throughputVSAvoidHPSU reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies dynamics by making data traffic rates adjustable rather than fixed. The system dynamically modifies data traffic rates assigned to HPSUs based on real-time effective ERT calculations, allowing maximum throughput when temperatures are low and automatically reducing traffic when reliability thresholds are approached, thus resolving the contradiction between productivity and reliability

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the operating parameter of data traffic rate based on calculated effective ERT values. By modulating traffic rates dynamically according to thermal conditions and reliability requirements, the system achieves optimal network throughput while maintaining HPSU reliability within prespecified limits

Inventive Principle:
Principle #35Parameter changes

3Reliability

If cooling capacity is increased to maintain HPSU reliability, then reliability is improved, but energy consumption increases

Engineering Contradiction:
ImproveHPSU reliabilityVSAvoidcooling energy consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent changes cooling capacity from a fixed high level to a dynamically adjusted parameter. By modulating cooling intensity based on effective ERT calculations and actual thermal conditions, the system provides adequate cooling only when necessary to maintain reliability, thus reducing unnecessary energy consumption while preserving HPSU reliability

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent implements feedback control for cooling capacity by continuously monitoring effective ERT and adjusting cooling intensity accordingly. The system increases cooling only when effective ERT approaches prespecified reliability thresholds and reduces cooling when reliability is adequately maintained, optimizing the balance between reliability and energy consumption

Inventive Principle:
Principle #23Feedback

4Reliability

If conservative temperature limits are applied to ensure HPSU reliability, then reliability is improved, but the system operates below optimal performance

Engineering Contradiction:
ImproveHPSU reliabilityVSAvoidsystem performance
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent changes from fixed conservative temperature parameters to dynamic effective ERT parameters. By calculating effective ERT based on actual operating profiles and reliability dependencies, the system determines optimal performance limits rather than using overly conservative defaults, thus improving productivity while maintaining reliability

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent applies partial action by allowing HPSUs to operate at higher performance levels when thermal conditions and effective ERT calculations indicate it is safe to do so. Rather than consistently applying conservative limits, the system applies performance optimization partially based on actual conditions, thus improving overall productivity while maintaining reliability

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20210041927A1Raising Maximal Silicon Die Temperature Using Reliability Model
Publication Date: 2021.02.11 MELLANOX TECHNOLOGIES LTD(IL)
  • US20210041927A1 patent drawing
  • US20210041927A1 patent drawing
  • US20210041927A1 patent drawing

AI summary

A method includes obtaining (i) an operating-temperature profile of a hardware processing sub-unit (HPSU) of a network element as a function of time, and (ii) a dependence of an Equivalent Reliability Time (ERT) of the HPSU on operating temperature. The operating-temperature profile is weighted using the dependence of the ERT on operating temperature, to estimate an effective ERT of the HPSU. An operating condition of the HPSU in the network element is modified, depending on the effective ERT.