System for dynamic power management in signal conductors
By dynamically adjusting the power state transition time of the PCIe link, the problem of energy efficiency and performance imbalance in traditional power management methods is solved, achieving more efficient power management, reducing energy consumption and latency, and improving device responsiveness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MELLANOX TECHNOLOGIES LTD(IL)
- Filing Date
- 2025-11-10
- Publication Date
- 2026-05-12
AI Technical Summary
Traditional PCIe link power management methods struggle to balance performance and energy efficiency, leading to unnecessary energy consumption and latency that negatively impact device performance.
By utilizing historical link usage information and predicted future link usage information, the power state transition time of PCIe links can be dynamically adjusted to achieve adaptive power management.
The energy efficiency and performance of the PCIe link have been optimized, reducing unnecessary energy consumption and latency, and improving device responsiveness and resource utilization.
Smart Images

Figure CN122018661A_ABST
Abstract
Description
Technical Field
[0001] The exemplary embodiments of this disclosure relate to a system for dynamic power management in signal conductors. Background Technology
[0002] The implementation of conventional active-state power management (ASPM) for power management in peripheral component interconnect high-speed channel (PCIe) links can present some challenges that may affect the performance and energy efficiency of PCIe connected devices.
[0003] The applicant has identified several defects and problems associated with power management in signal conductors (e.g., PCIe). Many of these identified problems have been addressed by developing solutions included in the embodiments of this disclosure, many of which are described in detail herein. Summary of the Invention
[0004] Therefore, systems, methods, and computer program products for dynamic power management in signal conductors are provided.
[0005] In one aspect, a system for dynamic power management in a signal conductor is proposed. The system includes: a signal conductor; and a control unit operatively coupled to the signal conductor, wherein the control unit is configured to: receive historical link usage information and predicted future link usage information associated with the signal conductor; determine an impending link usage change based on the received historical link usage information and the predicted future link usage information; and trigger a transition of the signal conductor from a current operating state to a subsequent operating state in response to determining the impending change.
[0006] In some embodiments, the control unit is further configured to: dynamically adjust the time period associated with transitioning the signal conductor from the current operating state to the subsequent operating state based on the impending link usage change, wherein the current operating state and the subsequent operating state are defined in an operating power state sequence; and trigger the signal conductor to transition from the current operating state to the subsequent operating state after the adjusted time period.
[0007] In some embodiments, the impending link usage change is a decrease in link utilization, wherein the control unit is configured to dynamically adjust the time period by shortening the time period, and wherein the subsequent operating state is a low-power state immediately following the current operating state in a defined sequence of operating power states.
[0008] In some embodiments, the impending link usage change is an increase in link utilization, wherein the control unit is configured to dynamically adjust the time period by shortening the time period, and wherein the subsequent operating state is a high-power state immediately preceding the current operating state in a defined sequence of operating power states.
[0009] In some embodiments, where the historical link usage information and the predicted link usage information provide conflicting indications about the impending link usage change, the control unit is configured to dynamically adjust the time period by extending the time period.
[0010] In some embodiments, the control unit is configured to reset the time period associated with transitioning the signal conductor from the current operating state to the subsequent operating state to a default value.
[0011] In some embodiments, the control unit is configured to determine the duration of impending inactivity in the signal conductor based on the historical link usage information and the predicted link usage information, wherein, when the duration of impending inactivity is determined, the control unit is configured to dynamically adjust the time period by shortening the time period, and wherein the subsequent operating state is a low-power operating state.
[0012] In some embodiments, the low-power operating state is the final state in the defined sequence of operating power states.
[0013] In some embodiments, the low-power operating state immediately follows the current operating state in the defined sequence of operating power states.
[0014] In some embodiments, the current operating state is a low-power operating state; the impending link usage change is an increase in link utilization; wherein the control unit is further configured to dynamically adjust the time period by shortening the time period; the subsequent operating state is an active power operating state in a defined sequence of operating power states; and the control unit is configured to trigger the signal conductor to transition from the current operating state to the subsequent operating state, such that the signal conductor is in the subsequent operating state before the increase in link utilization.
[0015] In some embodiments, the active power operating state is the initial state in the defined sequence of operating power states.
[0016] In some embodiments, the signal conductor is a peripheral component fast interconnect (PCIe) link, and the operating power state sequence defined therein includes L0, L0p, L1, L2, and L3 power states.
[0017] On the other hand, a method for dynamic power management in a signal conductor is proposed. The method includes: receiving historical link usage information and predicted future link usage information associated with the signal conductor; determining an impending link usage change based on the historical link usage information and the predicted future link usage information; and triggering the signal conductor to transition from a current operating state to a subsequent operating state in response to determining the impending change.
[0018] On the other hand, a computer program product for dynamic power management in a signal conductor is proposed. This computer program product includes a non-transitory computer-readable medium containing code configured to cause a device to perform the following operations: receiving historical link usage information and predicted future link usage information associated with the signal conductor; determining an impending link usage change based on the received historical link usage information and the predicted future link usage information; and triggering a transition of the signal conductor from a current operating state to a subsequent operating state in response to determining the impending change.
[0019] In another aspect, a control unit is proposed. This control unit includes: a processor; and a non-transitory storage device containing instructions that, when executed by the processor, cause the processor to: receive historical link usage information and predicted future link usage information associated with a signal conductor; determine an impending link usage change based on the received historical link usage information and the predicted future link usage information; and, in response to determining the impending change, trigger the signal conductor to transition from a current operating state to a subsequent operating state.
[0020] The above overview is intended to summarize some exemplary embodiments to provide a basic understanding of certain aspects of this disclosure. Therefore, it should be understood that the above embodiments are merely examples and should not be construed as limiting the scope or spirit of this disclosure in any way. It should be understood that the scope of this disclosure covers many potential embodiments in addition to those outlined herein, some of which will be further described below. Attached Figure Description
[0021] Some exemplary embodiments of this disclosure have been described in general terms above; reference will now be made to the accompanying drawings. The components shown in the drawings may or may not be present in some of the embodiments described herein. Some embodiments may contain fewer (or more) components than those shown in the drawings.
[0022] Figure 1 An example of a PCIe-based interconnect architecture for dynamic power management in signal conductors according to embodiments of the present disclosure is shown;
[0023] Figure 2 An example of a control unit circuitry for dynamic power management in a signal conductor, according to an embodiment of the present disclosure, is shown.
[0024] Figure 3 This is a block diagram illustrating a computing system (e.g., a data center or high-performance computing (HPC) cluster) according to embodiments described herein; and
[0025] Figure 4 An exemplary method for dynamic power management in a signal conductor is shown according to embodiments of the present disclosure. Detailed Implementation
[0026] Overview
[0027] In deep learning (DL) systems, addressing the challenge of efficiently managing computing resources, particularly graphics processing units (GPUs), is crucial. These systems are designed to partition and distribute large workloads across multiple different GPUs to achieve scalable processing power. During operation, GPUs synchronize the data they process with the data of their corresponding GPUs, forming a collective synchronization mechanism. This process ensures that each GPU completes the transfer and reception of all relevant data before moving on to subsequent processing tasks. This synchronization may result in the interconnect structure connecting the GPUs (e.g., A unique "waveform" traffic pattern has emerged in communication networks, characterized by short duty cycles and relatively long periods of inactivity. As part of the interconnect architecture, Network Interface Controllers (NICs) and GPU-Direct technology can be used to offload GPU synchronization tasks. These NICs are connected via PCI-Express (PCIe). Connections such as GRS (Ground Reference Signaling), LPI (Low Power Interface), LLI (Low Latency Interface), and / or similar interfaces are used to interface with the GPU.
[0028] The PCIe specification includes various low-power standby states, L1, L2, and L3, designed to regulate the power consumption of PCIe-connected devices, including GPUs. These states represent progressive levels of power-saving modes compared to the higher-power-usage active state, L0. The L1 state is a moderate power-saving mode where both the PCIe-connected device and the PCIe link significantly reduce their power consumption but can quickly resume full operation. Therefore, the L1 state is configured to provide a balance between reduced power consumption and minimized wake-up latency. Compared to L1, the L2 state is a deeper power-saving state involving even greater power reductions in both the PCIe-connected device and the PCIe link. Therefore, the L2 state is in the middle of this spectrum, achieving greater power savings than L1, but with correspondingly longer entry and exit latencies. The L3 state represents the deepest level of power saving, where the device is essentially off and the PCIe link is completely de-energized. Therefore, the L3 state is used for situations where immediate device use is not anticipated, and power conservation takes precedence over rapid resumption of activity. The recovery time from L3 state is the longest of all standby states because it requires a complete reinitialization of the PCIe link and the PCIe connected devices. Therefore, each of these standby states requires a trade-off between power saving and operational responsiveness, with L1 prioritizing wake-up speed, L3 prioritizing power conservation, and L2 balancing both.
[0029] The PCIe specification also includes an enhanced power-saving mode, L0p state, which operates in an active L0 state and is designed to reduce power consumption without significantly impacting the device's readiness or performance. Unlike deep sleep states (such as L1, L2, and L3), which save more power but at the cost of higher wake-up latency, L0p state keeps at least one channel of the PCIe link active while reducing power, ensuring uninterrupted data transmission. The design of L0p state allows PCIe connected devices to significantly reduce power consumption in near-idle conditions without fully entering sleep mode. Switching into and out of L0p state is very rapid, ensuring that the device can quickly resume full data transmission with minimal latency.
[0030] PCIe configurations vary in the number of channels, which are the basic pathways for devices to send and receive data. PCIe configurations are often referred to as ×1, ×4, ×8, ×16, etc., representing the number of channels and thus the bandwidth capacity of the connection. The more channels a particular signal conductor has, the higher the data transfer rate it can support. During normal operation in the active state (L0), all channels in the PCIe link operate fully, achieving maximum data throughput. In contrast, low-power standby states (L1, L2, and L3) are designed to reduce power consumption when the link is not actively transmitting data. For example, in the L1 state, some circuitry may remain powered for quick recovery, but data transfer is paused. Unlike the paused or significantly reduced data transfer in the L1, L2, and L3 states, the L0p state allows the PCIe link to enter a power-saving mode without completely stopping data transfer. This is achieved by maintaining at least one active channel for data transfer to ensure uninterrupted traffic flow.
[0031] In PCIe technology, Active State Power Management (ASPM) aims to reduce the power consumption of devices using PCIe connections. ASPM achieves energy efficiency by dynamically adjusting the power state of the PCIe link based on active data transmission over the link. This dynamic adjustment supports various power-saving modes, ranging from slightly reducing power by lowering the clock speed (L0 state) to more significant power cuts by completely shutting down the link (transitioning to L1, L2, or L3 states). This is particularly important in environments where PCIe-connected devices may not be continuously used, resulting in significant power savings during inactive or low-activity periods. However, the transition between active and low-power states introduces latency. This latency occurs during the wake-up phase when a PCIe device returns from a low-power state to a fully active state (L0), impacting the performance of latency-sensitive applications. DL applications, known for their intensive data processing and transfer demands, are particularly vulnerable to these latency issues. Due to the nature of DL workloads, which typically require rapid, high-volume data exchange between components, any increased latency can potentially hinder application performance.
[0032] Therefore, traditional Active State Power Management (ASPM) implementations for power management in PCIe links can present several challenges that can impact the performance and energy efficiency of PCIe connected devices. For example, transitioning to a standby state (e.g., L1) typically depends on the historical idle state of the PCIe link, often requiring a wait time of 100μs to 400μs or longer. This process causes the PCIe link to remain active for longer than necessary, resulting in unnecessary energy consumption. The process of exiting standby to active state is initiated by a real-time signal. Therefore, the latency associated with this transition can adversely affect the performance of PCIe connected devices because of the delay in regaining full operational capability. The decision to transition to L0p state is based on a historical analysis of load reduction in the PCIe link, typically lasting from 50μs to 100μs or longer. This approach leads to higher-than-needed power consumption because the link remains active even without significant load. Transition strategies between standby states (specifically, from L1 to L2, and then to L3) aim to improve energy efficiency by relying on the historical idle history of the PCIe link. However, this approach may cause the PCIe link to remain in L3 state until it is reactivated, resulting in significant latency during recovery and potentially affecting the responsiveness of PCIe-connected devices. Adjusting the PCIe link width based on historical usage can lead to scenarios where the operating width exceeds the actual required width. However, this practice may result in inefficient use of energy resources, as wider PCIe links consume more power.
[0033] Embodiments of this disclosure can dynamically adjust the L0p timer using historical link usage information associated with PCIe links and predicted future link usage information. For example, if a prediction indicates that link utilization is about to decrease, the L0p timer can be shortened to enable a faster transition to an energy-saving state. Similarly, if utilization is expected to increase, the L0p timer can be shortened to enable a faster transition to a high-performance state, thereby meeting anticipated demands. In the event of a conflict between historical link usage information and predicted future link usage information, the L0p timer can maintain its default value or be adjusted upwards, thereby extending the transition latency.
[0034] Furthermore, embodiments of this disclosure can utilize historical link usage information and predicted future link usage information associated with the PCIe link to adjust stand-by entrance timers. If predictions indicate a prolonged period of low or no traffic in the PCIe link, the stand-by entrance timers can be dynamically shortened to transition to a lower state. This allows the PCIe link to quickly enter a lower power state (e.g., L1, L2, L3), thereby saving energy during inactive periods. In some cases, the stand-by entrance timers can be set based on the predicted duration of inactivity to transition the PCIe link to a lower power state other than L1 (e.g., L2 or L3), ensuring that the PCIe link does not remain in the L1 state for longer than necessary. Therefore, intermediate power states can be skipped, and the transition to the most suitable standby state can be initiated directly. Conversely, if predictions indicate that traffic will recover quickly, the current stand-by timer setting can be maintained or extended to ensure that the PCIe link remains in a ready state.
[0035] Furthermore, embodiments of this disclosure can utilize historical link usage information associated with the PCIe link and predicted future link usage information to implement an adaptive wake-up mechanism, thereby pre-waking the PCIe link from deep sleep or idle states (e.g., L3 state). In this regard, the earlier wake-up time of the PCIe link can be scheduled based on predictions indicating the duration of inactivity. In other words, embodiments of this disclosure can calculate an optimal wake-up time earlier than the actual demand of the PCIe link based on the predicted idle duration. The goal is to find a middle ground for wake-up timing, avoiding both premature wake-up (which negates energy-saving effects) and wake-up upon signal arrival (which introduces latency). A wake-up slightly earlier than the signal arrival time ensures PCIe link responsiveness without significantly impacting power consumption.
[0036] The principles and methods described in this article are primarily explained within the context of power state management for PCIe links, but are not limited to this specific interface. It is clearly understood that the concept of dynamic power management (including predictive transitions between various power states based on historical link usage information and predicted future link usage information) is applicable to and can be extended to various signal conductor technologies and communication interfaces, such as... GRS, LPI, LLI, etc. This adaptability stems from the nature of their underlying strategies, which focus on optimizing energy efficiency and performance through intelligent state transitions. Therefore, the application of these principles in other link technologies should be considered within the scope of this discussion, with necessary modifications and adjustments made according to the specific characteristics and requirements of each technology. Furthermore, the same principles of dynamic power management and predictive transitions between power states also apply to network devices using such signal conductor technologies, such as NICs, switches, and other network hardware. The adaptability of these concepts enables their effective implementation on a wide range of network devices, improving energy efficiency and performance through intelligent state transitions.
[0037] Furthermore, embodiments of this disclosure are applicable to any multi-channel interface exhibiting non-negligible idle power consumption and wake-up latency sensitivity. Specifically, the link state prediction mechanism is applicable to a variety of interfaces. In one embodiment, a multi-channel inter-die serializer-deserializer (SerDes) interface may be a suitable candidate for implementing the disclosed method. In contrast, multi-channel inter-chip designs (such as the C2C-LPI / LLI implementation used by NVIDIA) exhibit significantly lower idle power consumption, making them less suitable for this application. Additionally, in some embodiments, the NVIDIA CHI protocol running on the inter-die interface can benefit from the link state prediction mechanism. Furthermore, embodiments of this disclosure are applicable to multi-channel networking protocols, such as… And Ethernet. In such protocols, prediction algorithms can be integrated with channel width adjustment capabilities to optimize performance and energy efficiency.
[0038] Embodiments of this disclosure will be described more fully below with reference to the accompanying drawings, which illustrate some, but not all, embodiments of this disclosure. In fact, this disclosure can be embodied in many different forms and should not be construed as limited to the embodiments described herein; rather, these embodiments are provided to enable this disclosure to meet applicable legal requirements. Therefore, it should be understood that each block in the block diagrams and flowcharts can be implemented as: a computer program product; a purely hardware embodiment; a purely firmware embodiment; a combination of hardware, a computer program product, and / or firmware; and / or an apparatus, system, computing device, computing entity, etc., that performs instructions, operations, steps, and interchangeable similar terms (e.g., executable instructions, instructions for execution, program code, etc.) on a computer-readable storage medium for execution. For example, code retrieval, loading, and execution can be performed sequentially, such that one instruction is retrieved, loaded, and executed at a time. In some exemplary embodiments, retrieval, loading, and / or execution can be performed in parallel, such that multiple instructions can be retrieved, loaded, and / or executed simultaneously. Thus, such embodiments can produce machines that perform specific configurations of the steps or operations specified in the illustrations of the block diagrams and flowcharts. Accordingly, the block diagrams and flowcharts illustrate combinations of various embodiments for performing specified instructions, operations, or steps.
[0039] Where possible, any term expressed in the singular form herein also implies the inclusion of the plural form, and vice versa, unless otherwise expressly stated. Furthermore, the terms “a” and / or “one” as used herein shall mean “one or more”, even if the phrase “one or more” is also used herein. Additionally, when something is said herein to be “based on” other things, it may also be based on one or more other things. In other words, unless otherwise expressly stated, “based on” as used herein means “at least partially based on” or “at least partially based on”. The same number always refers to the same element.
[0040] As used herein, "operationally coupled" can refer to components being coupled electronically or optically to each other and / or communicating electrically or optically. Furthermore, "operationally coupled" can mean that components can be integrally formed or can be separately formed and coupled together. Additionally, "operationally coupled" can mean that components can be directly connected to each other, or connected through one or more components (e.g., connectors) located between the components and operably coupled together. Moreover, "operationally coupled" can mean that components can be detached from each other, or that components are permanently coupled together.
[0041] As used in this article, “interconnection” can refer to each component being directly or indirectly linked to all other components or switches in the network, thereby allowing seamless data transmission and communication between all components.
[0042] The term "determine" as used in this article can encompass a variety of actions. For example, "determine" can include calculation, operation, processing, derivation, investigation, and ascertainment. Furthermore, "determine" can also include receiving (e.g., receiving information), accessing (e.g., accessing data in memory), etc. Similarly, "determine" can also include parsing, selecting, choosing, calculating, and establishing. Determining can also include ascertaining whether parameters meet predetermined criteria, including whether thresholds have been reached, exceeded, exceeded, or satisfied.
[0043] As used in this article, "signal conductor" can refer to any medium or mechanism that facilitates signal transmission between components within a digital system. The term "link" is used interchangeably with "signal conductor" and encompasses various physical and logical connections used to enable information exchange between different hardware components.
[0044] It should be understood that the term "exemplary" as used herein means "serving as an example, instance, or illustration." Any "exemplary" implementation described herein is not necessarily to be construed as superior to other implementations.
[0045] Furthermore, those skilled in the art will understand from this disclosure that terms such as “basic” and “approximate” indicate that the referenced elements or related descriptions are accurate within applicable engineering tolerances.
[0046] PCIe-based interconnect architecture example
[0047] Figure 1 An example of a PCIe-based interconnect architecture 100 for dynamic power management in signal conductors according to an embodiment of this disclosure is shown. In the context of this disclosure, the PCIe-based interconnect architecture 100 can be configured to facilitate secure data transmission over a high-speed interconnect framework or structure. Specifically, the PCIe-based interconnect architecture 100 can utilize the PCIe architecture (a serial computer extended bus standard) to establish a network of devices (e.g., endpoint devices, which will be described in more detail herein) for data communication. The architecture of the PCIe architecture provides a scalable point-to-point topology, enabling direct, secure signal conductors between endpoint devices. Furthermore, the PCIe-based interconnect architecture 100 can be used for the operation of modern data centers, facilitating high-speed, low-latency communication between critical components such as CPUs, GPUs, and storage devices. The PCIe-based interconnect architecture 100 can support the infrastructure required for HPC tasks (e.g., DL tasks), where large amounts of data transfer may be required between GPUs and other hardware components.
[0048] Figure 1This is just one example of an embodiment of the PCIe-based interconnect architecture 100. It should be understood that in other embodiments, one or more systems, units, and / or devices may be combined into a single system, unit, or device, or may consist of multiple systems, units, or devices. Furthermore, it should be understood that references to the PCIe architecture are illustrative and not prescriptive. The described PCIe-based interconnect architecture 100 is not necessarily limited to the PCIe architecture and can be adapted to alternative high-speed interconnect technologies. In such cases, the configuration of the PCIe-based interconnect architecture 100 can be modified to meet the specific requirements of any chosen high-speed interconnect technology.
[0049] like Figure 1 As shown, the PCIe-based interconnect architecture 100 may include a switching circuit system 107, multiple endpoint devices 104A, 104B, 104C, a root complex 106, a root port 106A, a CPU 110, and a memory 108.
[0050] Multiple endpoint devices 104A, 104B, and 104C can refer to devices that act as communication endpoints within a PCIe architecture. Therefore, multiple endpoint devices 104A, 104B, and 104C can be peripheral devices connected to the system, such as network interface cards (NICs), GPUs, NICs, memory controllers / adapters, field-programmable gate array (FPGA) accelerators, sound cards / audio interfaces, artificial intelligence (AI) / machine learning (ML) accelerators, etc.
[0051] Root complex 106 may be the primary interface within a PCIe architecture, used to connect peripheral devices (such as multiple endpoint devices 104A, 104B, 104C and switching circuitry 107) to the PCIe fabric. Root complex 106 may be configured to manage communication between external components (such as CPU 110, memory 108, etc.) and peripheral devices. These external components may be operatively coupled to root complex 106 via bus 105. Root complex 106 may include root ports (such as root port 106A) for interfacing with peripheral devices, managing configurations for discovering and configuring peripheral devices, and managing data paths for routing data between external components and peripheral devices.
[0052] Switching circuit system 107 can refer to a switching device within a PCIe architecture, which can route data between different points in a PCIe-based interconnect architecture 100, including routing data between multiple endpoint devices 104A, 104B, 104C, and root complex 106. In this regard, the multiple endpoint devices 104A, 104B, 104C, switching circuit system 107, and root complex 106 can be operatively coupled via signal conductor 103. Signal conductor 103 can be a high-speed interconnect facilitating data transmission and communication between endpoint devices 104A, 104B, 104C, switching circuit system 107, and root complex 106. Specifically, signal conductor 103 can be implemented as a PCIe link, a standard for serial expansion buses and devices. As described herein, the PCIe specification may include various low-power standby states—L1, L2, and L3—designed to regulate power consumption in endpoint devices (e.g., endpoint devices 104A, 104B, 104C) operatively coupled via a PCIe link (e.g., signal conductor 103). These states may represent progressive levels of power-saving modes compared to the more power-efficient active state L0.
[0053] Furthermore, a control unit 102 is integrated within the switching circuit system 107. The control unit 102 can be configured to capture and subsequently utilize bandwidth usage pattern information on the PCIe link and predictions about future bandwidth usage to dynamically adjust timers associated with power state transitions. By dynamically adjusting the timers, the control unit 102 can achieve faster transitions to energy-saving states or maintenance at higher performance states based on anticipated changes in bandwidth usage. The control unit 102 can also be configured to adjust standby entry timers based on similar predictions and historical bandwidth usage data, thereby enabling more efficient transitions to or out of low-power states (e.g., L1, L2, or L3) based on anticipated periods of inactivity. Additionally, the control unit 102 can be configured to implement an adaptive wake-up mechanism for the signal conductor 103 from idle states (e.g., L3 state), scheduling wake-ups based on predicted inactivity durations to balance responsiveness and energy-saving requirements.
[0054] It is important to note that the implementation of control unit 102 is not limited to switching circuit system 107. The PCIe-based interconnect architecture 100 allows control unit 102, or portions thereof, to be placed within other components of the PCIe-based interconnect architecture 100 (e.g., endpoint devices 104A, 104B, 104C, root complex 106, etc.), thereby distributing the various functions of control unit 102 among the system components. Each component (e.g., endpoint devices 104A, 104B, 104C) may contain a portion of control unit 102, or its hardware may be configured to perform specific operations of control unit 102. This distributed capability of control unit functionality ensures that dynamic power management can be effectively implemented throughout the PCIe-based interconnect architecture 100, thereby optimizing energy efficiency and system performance.
[0055] It should be understood that the structure, components, connections, relationships, and functions of the PCIe-based interconnect architecture 100 are intended to be exemplary only and are not intended to limit the implementation of the disclosures described and / or claimed in this document. In one example, the PCIe-based interconnect architecture 100 may include more components, fewer components, or different components. In another example, some or all of the PCIe-based interconnect architecture 100 may be combined into a single part, or all parts of environment 100 may be divided into two or more distinct parts.
[0056] Example Endpoint Device Circuit System
[0057] Figure 2 An example control unit circuit system 102 for dynamic power management in a signal conductor is shown according to an embodiment of this disclosure. Figure 2 As shown, the control unit 102 may include a processor 112, a memory 114, an input / output circuit system 116, a communication circuit system 118, and a power management circuit system 120.
[0058] Although the term "circuit system" used herein with respect to components 112-120 is described in some cases using functional language, it should be understood that a particular implementation necessarily involves the use of specific hardware configured to perform functions associated with the respective circuit systems described herein. It should also be understood that some components of these components 112-120 may include similar or common hardware. For example, two sets of circuit systems may both utilize the same processor, network interface, storage medium, etc., to perform their associated functions, thus eliminating the need for duplicate hardware in each set of circuit systems. In this regard, it should be understood that some components associated with control unit 102 may be packaged together, while other components may be packaged separately (e.g., a controller communicating with control unit 102). While the term "circuit system" should be broadly understood to include hardware, in some embodiments, "circuit system" may also include software for configuring the hardware. For example, in some embodiments, "circuit system" may include processing circuit systems, storage media, network interfaces, input / output devices, etc. In some embodiments, other elements of control unit 102 may provide or supplement the functionality of a particular circuit system. For example, processor 112 can provide processing functions, memory 114 can provide storage functions, communication circuit system 118 can provide network interface functions, and so on.
[0059] In some embodiments, processor 112 (and / or coprocessor or any other auxiliary processor or processing circuitry system otherwise associated with the processor) may communicate with memory 114 via a bus to transfer information between components such as control unit 102. Memory 114 may be non-transitory, for example, it may include one or more volatile and / or non-volatile memories, or some combination thereof. In other words, memory 114 may be, for example, an electronic storage device (e.g., a non-transitory computer-readable storage medium). Memory 114 may be configured to store information, data, content, applications, instructions, etc., to enable a device (e.g., control unit 102) to perform various functions according to exemplary embodiments of this disclosure.
[0060] Despite Figure 2While shown as a single memory, memory 114 may include multiple memory components. These multiple memory components may be embodied on a single computing device or distributed across multiple computing devices. In various embodiments, memory 114 may include, for example, a hard disk, random access memory, cache memory, flash memory, read-only optical disc (CD-ROM), read-only digital versatile optical disc (DVD-ROM), optical disc, circuitry configured to store information, or some combination thereof. Memory 114 may be configured to store information, data, applications, instructions, etc., enabling control unit 102 to perform various functions according to the example embodiments discussed herein. For example, in at least some embodiments, memory 114 may be configured to buffer data for processing by processor 112. Furthermore, or alternatively, in at least some embodiments, memory 114 may be configured to store program instructions for execution by processor 112. Memory 114 may store information in the form of static and / or dynamic information. Control unit 102 may store and / or use this stored information in the course of performing its functions.
[0061] Processor 112 can be implemented in a variety of different ways, for example, it may include one or more processing devices configured to execute independently. Furthermore, or alternatively, processor 112 may include one or more processors configured in series via a bus to achieve independent execution of instructions, pipelining, and / or multithreaded processing. Processor 112 can be implemented in various ways, including one or more microprocessors with accompanying digital signal processors, one or more processors without accompanying digital signal processors, one or more coprocessors, one or more multi-core processors, one or more controllers, processing circuitry systems, one or more computers, various other processing elements (including integrated circuits, such as application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs)) or some combination thereof. The use of the term "processing circuitry system" can be understood to include single-core processors, multi-core processors, multiple processors within a device, and / or remote or "cloud" processors. Therefore, although... Figure 2 This is referred to herein as a single processor, but in some embodiments, processor 112 may include multiple processors. Multiple processors may be embodied on a single computing device or distributed across multiple such devices, which are collectively configured as control unit 102. Multiple processors may operatively communicate with each other and may be collectively configured to perform one or more functions of control unit 102 as described herein.
[0062] In one example embodiment, processor 112 may be configured to execute instructions stored in memory 114 or otherwise accessible to processor 112. Alternatively, or additionally, processor 112 may be configured to execute hard-coded functions. Thus, whether configured by hardware or software methods or a combination of both, processor 112 may represent an entity (e.g., physically embodied in a circuit system) capable of performing the operations of embodiments of this disclosure after appropriate configuration. Alternatively, as another example, when processor 112 is embodied as an executor of software instructions, these instructions may specifically configure processor 112 to perform one or more algorithms and / or operations described herein when the instructions are executed. For example, when processor 112 executes these instructions, control unit 102 may be able to perform one or more functions described herein.
[0063] In some embodiments, the control unit 102 may further include an input / output circuitry 116 that can communicate with the processor 112 to provide auditory, visual, mechanical, or other outputs, and / or, in some embodiments, to receive input instructions from a user or other source. In this sense, the input / output circuitry 116 may include devices for performing analog-to-digital and / or digital-to-analog data conversion. The input / output circuitry 116 may include, for example, support for displays, touchscreens, keyboards, mice, image capture devices (e.g., cameras), microphones, and / or other input / output mechanisms. The input / output circuitry 116 may include a user interface and may include a web user interface, mobile applications, self-service terminals, etc.
[0064] Processor 112 and / or the user interface circuitry including processor 112 may be configured to control one or more functions of the display or one or more user interface elements via computer program instructions (e.g., software and / or firmware) stored in memory accessible to processor 112 (e.g., memory 114, etc.). In some embodiments, aspects of input / output circuitry 116 may be reduced compared to embodiments where control unit 102 may be implemented as an end-user machine or other type of device designed for complex user interactions. In some embodiments (similar to other components discussed herein), input / output circuitry 116 may be removed from control unit 102. Input / output circuitry 116 may communicate with memory 114, communication circuitry 118, and / or any other components, e.g., via a bus. Although control unit 102 may include more than one input / output circuitry and / or other components, Figure 2 Only one is shown in this document to avoid making this disclosure overly complex (e.g., as with other components discussed herein).
[0065] In some embodiments, the communication circuit system 118 includes any device, such as a device or circuit system embodied in hardware, software, firmware, or a combination of hardware, software, and / or firmware, configured to receive data from and / or to a network and / or any other device, circuit system, or module associated therewith. In this regard, the communication circuit system 118 may include, for example, a network interface for enabling communication with a wired or wireless communication network. For example, in some embodiments, the communication circuit system 118 may be configured to receive and / or transmit any data that can be stored in the memory 114 using any protocol that can be used for communication between computing devices. For example, the communication circuit system 118 may include one or more communication ports, a NIC, an antenna, a transmitter, a receiver, a bus, a switch, a router, a modem, and supporting hardware and / or software, firmware / software, or any other device suitable for enabling communication via a network. Furthermore, or alternatively, in some embodiments, the communication circuit system 118 may include circuitry for interacting with an antenna to enable the transmission of signals via the antenna or the processing of reception of signals received via the antenna. These signals can be transmitted by the control unit 102 using any of a variety of wireless personal area network (PAN) technologies, such as... Bluetooth Low Energy (BLE), infrared wireless (e.g., IrDA), ultra-wideband (UWB), inductive wireless transmission, etc., are supported. Furthermore, it should be understood that these signals can be transmitted using Wi-Fi, Near Field Communication (NFC), Global Microwave Access Interoperability (WiMAX), or other proximity-based communication protocols. The communication circuitry 118 may additionally or alternatively communicate with the memory 114, the input / output circuitry 116, and / or any other component of the control unit 102, for example, via a bus. The communication circuitry 118 of the control unit 102 may also be configured to receive and transmit information between its associated components.
[0066] The power management circuitry 120 can be configured to perform dynamic power management operations within the control unit 102. In this regard, the power management circuitry 120 can be configured to adjust the transition time associated with the power state of the signal conductor 103 based on real-time and predictive data analysis of bandwidth usage. In one example embodiment, the power management circuitry 120 can be configured to adjust the L0p timer to facilitate rapid transitions between various power states, thereby improving energy efficiency during periods of fluctuating demand. For example, if a forecast indicates an impending decrease in bandwidth usage, the power management circuitry 120 can be configured to shorten the L0p timer to facilitate a faster transition to an energy-saving state. Furthermore, the power management circuitry 120 can be configured to modify the standby entry timer to optimize the timing of transitions to low-power states using bandwidth usage predictions. Additionally, the power management circuitry 120 can be configured to manage an adaptive wake-up mechanism that schedules preemptive activation of the signal conductor 103 from a deep sleep state based on the predicted duration of inactivity.
[0067] In some embodiments, control unit 102 may include hardware, software, firmware, and / or combinations of these components configured to support various aspects of the power management circuitry system described herein. It should be understood that in some embodiments, power management circuitry system 120 may be combined with other circuitry systems of control unit 102 (e.g., memory 114, processor 112, input / output circuitry system 116, and / or communication circuitry system 118) to perform one or more such example operations. For example, in some embodiments, power management circuitry system 120 may utilize processing circuitry systems (e.g., processor 112, etc.) to form a separate subsystem to perform one or more of its respective operations. In another example, and in some embodiments, some or all of the functions of power management circuitry system 120 may be performed by processor 112. In this regard, some or all of the example processes and algorithms discussed herein may be performed by at least one of processor 112 and power management circuitry system 120. It should also be understood that in some embodiments, power management circuitry system 120 may include a separate processor, a specially configured FPGA, or an ASIC to perform its respective functions.
[0068] Furthermore, or alternatively, in some embodiments, the power management circuitry 120 may use the memory 114 to store the collected information. For example, in some implementations, the power management circuitry 120 may include hardware, software, firmware, and / or combinations thereof that interact with the memory 114 to send, retrieve, update, and / or store data.
[0069] Therefore, a non-transitory computer-readable storage medium (e.g., memory 114) can be configured to store firmware, one or more application programs and / or other software, including instruction and / or other computer-readable program code portions that can be executed to direct the operation of control unit 102, thereby implementing various operations, including the examples described herein. Thus, a series of computer-readable program code portions can be embodied in one or more computer program products and can be used with devices, control unit 102, databases and / or other programmable means to produce the machine-implemented processes described herein. It should also be noted that all or part of the information described herein may be based on data received, generated and / or maintained by one or more components of control unit 102. In some embodiments, one or more external systems (e.g., remote cloud computing and / or data storage systems) may also be utilized to provide at least some of the functionality described herein.
[0070] It should be recognized that the structure of the control unit 102 detailed herein represents only one embodiment among many potential configurations. This particular structure of the control unit 102 is described to illustrate the specific arrangement and interaction of its components, including the data processing unit, network interface, and power management circuitry, which together constitute its comprehensive network functionality. However, this outlined configuration is not deterministic or limiting. The structure of the control unit 102 and its integrated components can vary to accommodate different network paradigms, technological advancements, and specific application requirements. Alternative embodiments of the control unit 102 may employ different types of processors, such as advanced multi-core CPUs or dedicated GPUs, unique network interfaces (e.g., SmartNICs or advanced wireless modules), and various methods for managing network events and communications. Furthermore, the scalability, data processing technologies, and network integration methods of the control unit 102 may vary considerably depending on the target operating environment and functional requirements. Therefore, while this disclosure describes one potential structure of the control unit 102, it should be understood that this represents only one example in the broader field of network-enabled devices. Consequently, the scope of this disclosure is not limited to this single form but can be extended to a variety of other forms, technologies, and configurations.
[0071] Example computing system
[0072] Figure 3This is a block diagram illustrating a computing system 300 (e.g., a data center environment or HPC cluster) according to embodiments described herein. According to at least one embodiment, system 300 may include multiple subsystems, such as processing devices, networking devices, and interconnect networks. Computing system 300 may include multiple integrated circuits (referred to as processing devices), wherein each integrated circuit may include one or more CPUs and GPUs, thereby providing a flexible architecture.
[0073] The processing equipment can be Interconnection can be achieved via other high-speed interconnects, allowing communication between subsystems. Furthermore, processing devices can be connected to the system network via a NIC or DPU to facilitate data transfer across computing system 300 and one or more external networks 330, 336. System 300 may include a packet switch 348 connecting NIC / DPU 328 to network 330, and a packet switch 350 connecting NIC / DPU 332 to network 336.
[0074] have The configuration of interconnected processing devices can support parallel processing, thereby improving computational efficiency. Processing devices can be connected to multiple networks via one or more NICs or DPUs, enabling the system to manage complex multi-network tasks with high bandwidth and low latency. This configuration may be suitable for applications requiring significant processing resources, such as artificial intelligence (AI), machine learning (ML), and data-intensive computing, while maintaining connectivity and scalability across network environments. The integrated circuits in the computing system 300 may include one or more CPUs and one or more GPUs.
[0075] Figure 3 A multi-GPU architecture is also demonstrated. In this embodiment, the computing system 300 may include a processing device 302 with a multi-GPU architecture. The processing device 302 may be a system-on-a-chip (SoC) having subsystems such as a CPU 306, a GPU 308, and a GPU 310. The CPU 306 can be connected to the GPU 308 via a die-to-die (D2D) or chip-to-chip (C2C) interconnect 312 (e.g., a GRS interconnect), and connected to the GPU 310 via a D2D or C2C interconnect 314. The CPU 306 can also be connected to the GPU 308 and GPU 310 via a PCIe interconnect.
[0076] The CPU 306 can connect to one or more NICs or DPUs, which can also connect to a network. For example, as Figure 3As shown, CPU 306 can connect to NIC / DPU 326, NIC / DPU 326 can connect to network 330, and can also connect to NIC / DPU 328. NIC / DPU 328 can also connect to network 330 via switch 348. NIC / DPU 326 and NIC / DPU 328 can communicate via Ethernet (ETH). or (IB) connects to network 330.
[0077] The computing system 300 may also include a processing device 304 with a multi-GPU architecture. The processing device 304 may include subsystems such as a CPU 316, a GPU 318, and a GPU 320. The CPU 316 may be connected to the GPU 318 via a D2D or C2C interconnect 322 and to the GPU 320 via a D2D or C2C interconnect 324. The CPU 316 may also be connected to the GPUs 318 and 320 via a PCIe interconnect. The CPU 316 may be connected to one or more NICs or DPUs, which may be coupled to a network. As shown, the CPU 316 may be connected to NIC / DPU 332 (which may be linked to network 336) and NIC / DPU 334 (which may be connected to network 336 via switch 350). NIC / DPU 332 and NIC / DPU 334 may be connected via Ethernet (ETH). or (IB) connects to network 336.
[0078] In at least one embodiment, processing devices 302 and 304 may be interconnected via NIC / DPU 338 (e.g., via PCIe interconnect) or via high-bandwidth communication 340 (e.g., via PCIe interconnect). (Interconnection) for communication. Figure 3 The packet switches in the data may include Quantum-2 switches, and the NIC / DPU can include DPU.
[0079] In this data center or HPC cluster, embodiments of this disclosure achieve dynamic power management of PCIe links by optimizing power state transitions using historical link usage data and predicted link usage data. In this regard, a control unit (e.g., control unit 102) can be integrated into a switching circuit system (e.g., packet switches 348 and 350) or directly embedded in one or more processing devices (e.g., processing devices 302 or 304). The control unit can be configured to manage dynamic power adjustments across PCIe links and coordinate the transitions of signal conductors between power states based on real-time usage data and predictive modeling. The control unit can capture and analyze historical link usage patterns and forecast data, enabling rapid adjustment of transition times associated with each power state (e.g., L0, L0p, L1, L2, and L3).
[0080] The control unit 102 can also be distributed across different components within the system 300, including NICs, DPUs, and even a single GPU, to ensure efficient real-time power management operations. This distributed configuration allows the control unit to dynamically adjust the power state of inter-device PCIe links (e.g., the links connecting processing devices 302 and 304 to networks 330 and 336 via NICs / DPUs 326, 328, 332, and 334) based on workload demands, thereby minimizing latency and energy consumption in data centers or HPC clusters. Through this strategic positioning within the system 300, the control unit can be configured to effectively balance performance and power efficiency, especially in applications with fluctuating computational demands, such as deep learning and data-intensive processing.
[0081] As described in this paper, the control unit can analyze real-time and forecast data on bandwidth usage patterns to adjust the power state of PCIe links, thereby optimizing energy efficiency without impacting data transmission performance. For example, when data flow forecasts indicate reduced activity between devices, the control unit can dynamically reduce the power state of the PCIe links, thereby reducing power consumption during periods of low utilization. Furthermore, the control unit can facilitate transitions into and out of standby states (e.g., L1, L2, or L3) based on forecasted traffic demands for each pair of devices connected via the PCIe links. By implementing adaptive wake-up mechanisms, the control unit can preemptively wake links from deep sleep states (e.g., L3) to ensure devices are ready for data transmission on demand, thereby reducing latency for data-intensive applications such as deep learning (DL). This approach is particularly beneficial in data center environments, as rapid data exchange and power efficiency are critical for scalable high-performance computing.
[0082] It should be understood that the embodiments and configurations described herein are for illustrative purposes only and are not intended to limit the scope of this disclosure. Various modifications, adjustments, and alternative implementations will be apparent to those skilled in the art without departing from the broader scope and spirit of this disclosure as set forth in the claims. Specific components, configurations, or functions mentioned herein, such as specific processing devices, interconnect technologies, or power states, may be replaced or modified according to different applications, technological advancements, or system requirements. The principles and techniques described herein are intended for broad application in various signal conductor technologies, network configurations, and data center architectures, including but not limited to those specifically exemplified herein.
[0083] letter Example method of dynamic power management in conductor No.
[0084] Figure 4 An example method 400 for dynamic power management in a signal conductor according to an embodiment of this disclosure is illustrated. As shown in block 402, historical link usage information and predicted future link usage information associated with the signal conductor are received. The historical link usage information may include a dataset containing records of past activity on the signal conductor (e.g., a PCIe link). This dataset may include metrics such as data transfer rates, usage frequency, activity duration, and data flow patterns observed within defined historical time periods. The purpose of analyzing historical data is to identify trends and patterns in link usage, which can inform the development of performance optimization strategies. The predicted future link usage information may include a link activity forecast based on a combination of historical data. The predicted future link usage information can predict the future state of link usage, including anticipated periods of high demand, low activity, or idle time. The prediction method may combine algorithms, machine learning models, or statistical analysis to predict future usage patterns. Details of the prediction methods and algorithms used to forecast future link usage are described in the related patent application entitled “Predicting Inactivity Patterns for a Signal Conductor”, U.S. Patent Application No. 17 / 893,692, the contents of which are incorporated herein by reference.
[0085] As shown in block 404, upcoming changes in link usage are determined based on received historical link usage information and predicted future link usage information. These upcoming changes may include anticipated increases or decreases in near-term link activity. As described herein, historical link usage information may include records of past link activity. Analyzing this information allows for the identification of established trends and patterns in link usage. On the other hand, predicted future link usage information aims to plan for the future state of link activity, identifying potential increases or decreases in demand. Ideally, the analysis of historical link usage information should independently indicate upcoming changes in link usage as indicated by predicted future link usage information. For example, if historical trends show cyclical increases in link usage at certain times, and a predictive model considering other variables also forecasts an increase in activity in the upcoming period, then both sources confirm the expectation of increased link activity. However, inconsistencies between historical link usage information and predicted future link usage information in indicating the same changes in activity are not uncommon. This can be due to many unforeseen reasons, such as sudden changes in demand, data anomalies in past events, evolution of usage patterns, model sensitivity, etc.
[0086] As shown in block 406, the time period associated with the transition of signal conductors from the current operating state to the subsequent operating state is dynamically adjusted based on impending changes in link usage. As described herein, operating states (e.g., current operating state, subsequent operating state, etc.) can be defined within a power management state framework for the PCIe link, designed to optimize the balance between energy efficiency and performance. These states, referred to as L-states, can be defined as a sequence of operating power states ranging from high performance to low power consumption. For example, in the fully operational L0 state, the PCIe link is active and capable of maximizing performance, facilitating data transfer at the highest hardware-supported rates. Due to the high activity level, power consumption in the L0 state reaches its peak. In the low-power states L1, L2, and L3, the PCIe link is inactive, providing varying degrees of energy savings. The transition from any of these low-power states back to the fully active L0 state requires a recovery time of varying durations, depending on the specific low-power state from which the transition originates. Furthermore, the L0p state, slightly less powerful than the L0 state, is a low-power idle state that allows for rapid recovery back to L0. The transition time to and from L0 is very short, making L0p suitable for short periods of inactivity.
[0087] A time period can refer to the duration required for a PCIe link to transition between specified power states, such as from a low-power state (e.g., L1, L2, L3) to an active state (L0, L0p) and vice versa. This duration can affect the readiness of the PCIe link to start, resume, reduce, or stop data transmission. Optimizing these transition times based on actual and anticipated link usage patterns allows the system to dynamically align its operating efficiency and performance capabilities with current and anticipated demands, thereby avoiding unnecessary power consumption or performance degradation.
[0088] In the context of data centers, PCIe connectivity devices such as GPUs and NICs typically strike a balance between high performance and energy efficiency requirements. The power-saving states (L1, L2, L3, and L0p) outlined in the PCIe specification provide a framework for reducing power consumption. For example, the L1 state reduces power usage while maintaining the ability to quickly resume full operation, making it ideal for workloads experiencing intermittent bursts of high demand. While the deeper L3 state offers maximum energy savings, it can introduce latency, impacting the performance of latency-sensitive deep learning applications.
[0089] As shown in block 408, after the adjusted time period has elapsed, the trigger signal conductor transitions from the current operating state to the subsequent operating state. In one example, the PCIe link may be in a semi-active L0p state, where the PCIe link remains in an operational but low-power mode. If an impending link utilization change occurs while the PCIe link is in the L0p state, resulting in a decrease in link utilization, the PCIe link may transition to a low-power state, such as the L1 state. However, in traditional systems, transitioning to the L1 state typically requires a wait of 100μs to 400μs or longer, causing the PCIe link to remain active for longer than necessary, thus increasing unnecessary energy consumption. Therefore, to avoid unnecessary power consumption, system embodiments (e.g., Figure 2 The control unit 102 can be configured to shorten the transition period (e.g., to about 10 μs to 50 μs or less) and trigger the PCIe link to switch after the shortened transition period has elapsed. Alternatively, if the impending link usage change while the PCIe link is in the L0p state is an increase in link utilization, the PCIe link can be switched to the active L0 state. Similarly, the system can shorten the transition (e.g., to about 10 μs to 50 μs or less) and trigger the PCIe link to switch after the shortened transition period has elapsed.
[0090] As described herein, in the best-case scenario, the analysis of historical link usage information aligns with the indications of predicted future link usage information, both suggesting an impending change in link usage (e.g., increase, decrease, no change, etc.). However, when historical and predicted link usage information provide conflicting indications of impending changes in link usage, the system can dynamically extend the transition period. Extending the transition period serves as a safeguard, providing an opportunity for natural resolution of any sudden changes in demand, data anomalies, evolution of usage patterns, model sensitivities, etc., without prematurely triggering PCIe link transitions. Alternatively, or additionally, when historical and predicted link usage information provide conflicting indications of impending changes in link usage, the system can reset the transition period to a predetermined default value.
[0091] In an example embodiment, an impending change in link usage might be an upcoming period of inactivity in the PCIe link. In this case, the system can transition the PCIe link from an active operating state (e.g., L0) or a semi-active operating state (e.g., L0p) to a low-power operating state (e.g., L1, L2, L3). However, as described herein, transitioning to a low-power operating state typically requires a specific transition period (e.g., 100 μs to 400 μs or longer), which can cause the PCIe link to remain in an active or semi-active operating state for longer than necessary. If the impending change is a duration of inactivity in the PCIe link (e.g., a period of low or no traffic), the system can shorten the transition period (e.g., shorten it to 10 μs to 50 μs or less) and trigger the PCIe link to transition after the shortened transition period has elapsed. In one example, if the PCIe link is currently in an active operating state (e.g., L0), and the inactivity duration is a relatively short period of no traffic, the system can shorten the transition period so that after the shortened period elapses, the PCIe link can transition from L0 to a subsequent semi-active operating state (e.g., L0p). In another example, if the PCIe link is currently in a semi-active operating state (e.g., L0p), and the inactivity duration is a relatively short period of no traffic, the system can shorten the transition period so that after the shortened period elapses, the PCIe link can transition from L0p to a subsequent low-power operating state (e.g., L1). In yet another example, if the PCIe link is currently in an active operating state (e.g., L0), and the inactivity duration is an extended period of no traffic, the system can shorten the transition period so that after the shortened period elapses, the PCIe link can be directly transitioned from L0 to a low-power state L3 (e.g., the end state in the PCIe link's operating power state sequence) without a gradual transition to an available low-power state.
[0092] In an example embodiment, the impending link usage change might be an increase in link utilization. In this case, the system can transition the PCIe link from a low-power operating state (e.g., L1, L2, L3) or a semi-active operating state (e.g., L0p) to an active operating state (e.g., L0). However, the standard transition time required for a PCIe link to adjust from a low-power operating state to an active operating state can lead to operational inefficiency. Specifically, peripheral devices (e.g., GPUs) that rely on PCIe links for data acquisition may experience latency. This latency occurs because peripheral devices typically need to wait for the duration of the transition time before accessing the necessary data via the PCIe link. Therefore, this waiting period can cause additional delays in task processing, potentially impacting overall system performance and the timely execution of functions that rely on fast data access and processing capabilities. Therefore, when the impending link usage change is an increase in link utilization, the system can shorten the transition time of the PCIe link from a low-power operating state (e.g., L2) to an active operating state (e.g., L0) and trigger the PCIe link to transition to an active operating state after the shortened transition time, before the link utilization increases.
[0093] Alternatively or additionally, in some embodiments, the control unit may be configured to directly trigger a signal conductor to transition from its current operating state to a subsequent operating state upon detecting an impending change in link usage. This approach eliminates the need for timer-based adjustments, allowing for instantaneous state transitions in response to dynamic usage conditions. This approach can potentially enhance the responsiveness of PCIe links to fluctuating demands while simplifying the control architecture by eliminating reliance on pre-configured timer settings.
[0094] By leveraging historical link usage data and predictive analytics, embodiments of this disclosure can dynamically adjust the transition timers for PCIe power states in a data center environment. This predictive mechanism allows the system to anticipate when high workloads are expected and maintain the system in a high-performance state (e.g., L0 or L0p) while transitioning to lower power states (e.g., L1, L2, or L3) during periods of inactivity. This approach not only saves energy but also minimizes latency by pre-awakening the system from deeper power states based on predicted demand.
[0095] The dynamic power management principles outlined in the various embodiments of this disclosure are not limited to PCIe-connected devices, but are also widely applicable to various signal conductor technologies and networking components within data centers. This may include NICs, switches, and other network devices in modern data centers. By intelligently switching between power states based on anticipated workloads, embodiments of this disclosure can improve the energy efficiency and performance of the entire data center infrastructure.
[0096] Those skilled in the art, upon benefiting from the teachings presented in the foregoing description and the accompanying drawings, will conceive of numerous modifications and other embodiments of the present disclosure set forth herein. Although the drawings illustrate only certain components of the methods and systems described herein, it should be understood that various other components may also be part of this disclosure. Furthermore, in some cases, the described method may contain fewer steps, while in others it may contain more steps. In some cases, the steps of the described method and modifications to the steps may be performed in any order and in any combination. Therefore, it should be understood that this disclosure is not limited to the specific embodiments disclosed, and that modifications and other embodiments are intended to be included within the scope of the appended claims. Although specific terminology is used herein, it is used only in a general and descriptive sense and not for limiting purposes.
Claims
1. A system for dynamic power management in a signal conductor, the system comprising: Signal conductor; A control unit, operatively coupled to the signal conductor, wherein the control unit is configured to: Receive historical link usage information and predicted future link usage information associated with the signal conductor; Based on the received historical link usage information and the predicted future link usage information, determine the upcoming changes in link usage; as well as In response to the determination of the impending change, the signal conductor is triggered to transition from the current operating state to a subsequent operating state.
2. The system according to claim 1, wherein, The control unit is also configured to: Based on the impending link usage change, the time period associated with transitioning the signal conductor from the current operating state to the subsequent operating state is dynamically adjusted, wherein the current operating state and the subsequent operating state are defined in an operating power state sequence; as well as After the adjusted time period, the signal conductor is triggered to switch from the current operating state to the subsequent operating state.
3. The system according to claim 2, wherein, The impending link usage change is a decrease in link utilization, wherein the control unit is configured to dynamically adjust the time period by shortening the time period, and wherein the subsequent operating state is a low-power state immediately following the current operating state in the defined sequence of operating power states.
4. The system according to claim 2, wherein, The impending link usage change is an increase in link utilization, wherein the control unit is configured to dynamically adjust the time period by shortening the time period, and wherein the subsequent operating state is a high-power state immediately preceding the current operating state in the defined sequence of operating power states.
5. The system according to claim 2, wherein, When the historical link usage information and the predicted link usage information provide conflicting indications about the impending link usage change, the control unit is configured to dynamically adjust the time period by extending the time period.
6. The system according to claim 5, wherein, The control unit is configured to: The time period associated with transitioning the signal conductor from the current operating state to the subsequent operating state is reset to a default value.
7. The system according to claim 2, wherein, The control unit is configured to determine the duration of impending inactivity in the signal conductor based on the historical link usage information and the predicted link usage information. Wherein, if the duration of the impending inactivity is determined, the control unit is configured to dynamically adjust the time period by shortening the time period; and The subsequent operation state is a low-power operation state.
8. The system according to claim 7, wherein, The low-power operating state is the final state in the defined sequence of operating power states.
9. The system according to claim 7, wherein, The low-power operating state immediately follows the current operating state in the defined sequence of operating power states.
10. The system according to claim 2, wherein: The current operating state is a low-power operating state; The impending change in link usage is an increase in link utilization. The control unit is further configured to dynamically adjust the time period by shortening the time period. The subsequent operation state is the active power operation state in the defined operation power state sequence; and The control unit is configured to trigger the signal conductor to transition from the current operating state to the subsequent operating state, such that the signal conductor is in the subsequent operating state before the link utilization increases.
11. The system according to claim 10, wherein, The active power operation state is the initial state in the defined sequence of operation power states.
12. The system according to claim 2, wherein, The signal conductor is a peripheral component fast interconnect PCIe link, and the operating power state sequence defined therein includes L0, L0p, L1, L2 and L3 power states.
13. A method for dynamic power management in a signal conductor, the method comprising: Receive historical link usage information and predicted future link usage information associated with signal conductors; Based on the historical link usage information and the predicted future link usage information, determine the upcoming changes in link usage; as well as In response to the determination of the impending change, the signal conductor is triggered to transition from the current operating state to a subsequent operating state.
14. The method of claim 13, further comprising: Based on the impending link usage change, the time period associated with transitioning the signal conductor from the current operating state to the subsequent operating state is dynamically adjusted, wherein the current operating state and the subsequent operating state are defined in an operating power state sequence; as well as After the adjusted time period, the signal conductor is triggered to switch from the current operating state to the subsequent operating state.
15. The method according to claim 14, wherein, The impending link usage change is a decrease in link utilization, wherein dynamically adjusting the time period includes shortening the time period, and wherein the subsequent operating state is a low-power state immediately following the current operating state in the defined sequence of operating power states.
16. The method of claim 14, wherein, The impending link usage change is an increase in link utilization, wherein dynamically adjusting the time period includes shortening the time period, and wherein the subsequent operating state is a high-power state immediately preceding the current operating state in the defined sequence of operating power states.
17. The method of claim 14, wherein, In cases where the historical link usage information and the predicted link usage information provide conflicting indications about the impending changes in link usage, dynamically adjusting the time period may include extending the time period.
18. The method of claim 17, further comprising: The time period associated with transitioning the signal conductor from the current operating state to the subsequent operating state is reset to a default value.
19. A computer program product for dynamic power management in a signal conductor, the computer program product comprising a non-transitory computer-readable medium containing code configured to cause a device to perform the following operations: Receive historical link usage information and predicted future link usage information associated with the signal conductor; Based on the received historical link usage information and the predicted future link usage information, determine the upcoming changes in link usage; as well as In response to the determination of the impending change, the signal conductor is triggered to transition from the current operating state to a subsequent operating state.
20. The computer program product according to claim 19, wherein, The code is also configured to cause the device to perform the following operations: Based on the impending link usage change, the time period associated with transitioning the signal conductor from the current operating state to the subsequent operating state is dynamically adjusted, wherein the current operating state and the subsequent operating state are defined in an operating power state sequence; as well as After the adjusted time period, the signal conductor is triggered to switch from the current operating state to the subsequent operating state.
21. A control unit, comprising: processor; A non-transitory storage device containing instructions that, when executed by the processor, cause the processor to: Receive historical link usage information and predicted future link usage information associated with signal conductors; Based on the received historical link usage information and the predicted future link usage information, determine the upcoming changes in link usage; as well as In response to the determination of the impending change, the signal conductor is triggered to transition from the current operating state to a subsequent operating state.