Adaptive trigger operation management in network interface controller

By implementing an adaptive window management mechanism in NIC in a high-performance computing environment, dynamically adjusting the allocation and utilization of TODS resources, the problems of resource competition and hunger among multiple processes are solved, and system performance is improved.

CN120066695APending Publication Date: 2025-05-30HEWLETT PACKARD ENTERPRISE DEV LP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410754043.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-30
Filing Date
2024-06-12
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In a high-performance computing environment, the hardware resources of the NIC are limited, resulting in that when sharing the triggered operational data structure (TODS) between multiple processes, some processes may oversubscribe resources, while others cannot effectively utilize TODS due to resource exhaustion, affecting the overall performance of the system.

Method used

By implementing an adaptive window management mechanism in NIC, the available entries of TODS are distributed uniformly or non-uniformly to each process, and the window size is dynamically adjusted according to the process's availability to ensure that each process can effectively utilize its allocated resources.

Benefits of technology

This solution effectively prevents resource competition and hunger between processes, improves the efficiency of TODS resource utilization between multiple processes, and improves the overall performance of the computing system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066695A_ABST
    Figure CN120066695A_ABST
Patent Text Reader

Abstract

The invention relates to adaptive trigger operation management in a network interface controller. A system for managing trigger operations in a computing system is provided. The computing system may include a storage medium to store descriptors identifying trigger operations to be performed based on respective trigger conditions. A network interface controller of the computing system may store a data structure. During operation, the system may determine a window size for a process, the window size indicating a number of available entries in the data structure. If the window size indicates that an entry is available, the system may insert a descriptor of a trigger operation generated by the process into a corresponding work queue. The system may determine, at the NIC, a location of the descriptor in the work queue. The system may then transmit the descriptor from the location to the data structure. The system may then decrement the window size to indicate a number of updates in the data structure available for entries of the first process.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technical Field

[0002] High-performance computing (HPC) generally enables efficient computing on nodes running applications. HPC can facilitate high-speed data transfer between a sender device and a receiver device. Brief Description of the Drawings

[0003] Figure 1 An example of adaptive trigger operation management in a network interface controller (NIC) according to one aspect of the present application is illustrated.

[0004] Figure 2 An example of facilitating inter-component communication for adaptive trigger operation management in a computing system according to one aspect of the present application is illustrated.

[0005] Figure 3A An example of partitioning a trigger operation data structure (TODS) in a NIC among multiple processes according to one aspect of the present application is illustrated.

[0006] Figure 3B An example of decrementing a corresponding window size indicating availability in a TODS according to one aspect of the present application is illustrated.

[0007] Figure 3C An example of incrementing a corresponding window size indicating availability in a TODS according to one aspect of the present application is illustrated.

[0008] Figure 4A A flowchart according to one aspect of the present application is presented, which illustrates an example of a process by which a computing system facilitates adaptive trigger operation management.

[0009] Figure 4B A flowchart according to one aspect of the present application is presented, which illustrates an example of a process by which a NIC performs a trigger operation from one process based on a trigger descriptor in a local TODS.

[0010] Figure 5 A flowchart according to one aspect of the present application is presented, which illustrates an example of a process by which a NIC performs a trigger operation from another process based on a trigger descriptor in a local TODS.

[0011] Figure 6 An example of a computing system having a NIC that facilitates adaptive trigger operation management according to one aspect of the present application is illustrated.

[0012] Figure 7 An example of a computer-readable storage medium that facilitates adaptive trigger operation management according to one aspect of the present application is illustrated.

[0013] In these figures, like reference numerals refer to like elements. DETAILED DESCRIPTION

[0014] As applications become increasingly more distributed, HPC can facilitate efficient computing on the nodes running the applications. An HPC environment can include compute nodes (e.g., computing systems), storage nodes, and high-capacity network devices coupling these nodes. Thus, an HPC environment can include a high-bandwidth, low-latency network formed by these network devices. Compute nodes can be coupled via the network to storage nodes. Compute nodes can run one or more application processes (or processes) in parallel. Storage nodes can record the output of computations performed on compute nodes. Additionally, data from one compute node can be used by another compute node for computations. Thus, compute nodes and storage nodes can cooperate with each other to facilitate high-performance computing.

[0015] One or more processes can perform computations on the processing resources (such as processors and accelerators) of a compute node. Data generated by the computations can be transferred to another node using the NIC of the compute node. Such a transfer can include a Remote Direct Memory Access (RDMA) operation. To transfer data, the process can enqueue a descriptor into a command queue in the memory of the compute node and set register values. Based on the register values, the NIC can determine the presence of the descriptor and dequeue the descriptor from the command queue.

[0016] The NIC can then obtain information associated with the RDMA operation from the descriptor, such as information about the source buffer (e.g., the location of the data to be transferred), the destination buffer (e.g., the location where the data is to be transferred), the size of the data transfer, memory registration, and destination process details. The data can be generated by the execution of the process and stored in the source buffer (e.g., stored in the storage medium of the computing system). Thus, the descriptor can be an identifier of the operation's identifier. Typically, after the descriptor is dequeued, the NIC can perform a data transfer operation (e.g., transfer a packet).

[0017] Additionally, the NIC can also support triggering operations, which can allow a process to enqueue an operation for deferred execution. For example, a process can deploy parallel loop computations that are executed in a nested and repetitive manner. Such computations are typically executed on different compute nodes and can depend on the computation outputs of each other. These computations are typically offloaded to attached hardware elements (such as accelerators) for execution. The corresponding communication operations can be deferred until after the computations are completed. Thus, the corresponding communication operations can be represented as triggering operations. When a trigger condition is met, the NIC can execute the triggering operation.

[0018] The NIC can store descriptors of trigger operations and corresponding trigger conditions in a Trigger Operation Data Structure (TODS). The descriptors of trigger operations can be referred to as trigger descriptors. When the execution of a computation is complete, a trigger event can be executed. The execution of the trigger event can then satisfy the trigger condition. For example, the trigger condition can be that a counter value reaches a threshold, and the trigger event can be incrementing the counter value. When the trigger condition is satisfied, the NIC can obtain a trigger operation based on the information in the trigger descriptor stored in the TODS. The NIC can then execute the trigger operation, which can include sending a packet that includes the output of the computation.

[0019] Aspects described herein solve the problem of efficiently distributing entries of a TODS between processes in a non-blocking manner by: (i) distributing entries of the TODS between processes that generate trigger operations; (ii) maintaining a window that indicates available entries for a process; and (iii) decrementing and incrementing a window size (WIN) in response to enqueuing and executing a trigger operation, respectively. Here, the size of the window can be referred to as the window size. The window size associated with a process can indicate the number of TODS entries allocated to that process. Since the window size can indicate the entries currently available to a process, a process can enqueue a descriptor of a trigger operation into the TODS when the window size has a non-zero value. In this way, the TODS can support trigger operations from multiple processes without overwhelming the TODS.

[0020] Unlike regular operations executed on the NIC, trigger operations provide deferred execution, where the execution of a trigger operation can be triggered later. A process that generates a trigger operation can incorporate information associated with the trigger operation into a trigger descriptor and enqueue it into a Deferred Work Queue (DWQ). The process can also set a register value to indicate the presence of the descriptor in the DWQ. In addition to a regular descriptor, a trigger descriptor can incorporate three additional parameters - a trigger counter, a completion counter, and a trigger threshold. The trigger threshold can be a predetermined value. Since the trigger descriptor includes identification information associated with the trigger operation, the trigger descriptor can also be referred to as an identifier of the trigger operation. If the trigger counter is incremented up to the threshold, the NIC can determine the location of the trigger operation based on the triggered descriptor and execute the trigger operation. Since a trigger operation can be repeated frequently (e.g., in a loop), the completion counter can indicate the number of times the trigger operation has been executed.

[0021] TODS can be deployed in the NIC to support trigger operations. The TODS can be a hardware entity, such as a storage medium. The NIC can enqueue descriptors from the DWQ into the available entries of the TODS. Descriptors for trigger operations can be referred to as trigger descriptors. When a trigger operation is executed, the entry can be released for reuse. Since the hardware resources of the NIC are limited, the number of entries in the TODS may be limited. If the computing system hosting the NIC executes multiple processes, the TODS can be shared among these processes. Since the availability of hardware resources in the NIC is limited and the TODS is shared among multiple processes, some processes may overbook the TODS, while some other processes may not be able to utilize the TODS due to resource exhaustion. Therefore, the functions and performance of underutilized processes may be adversely affected.

[0022] To address this issue, the NIC can allocate the available entries of the TODS to the respective processes that generate trigger operations and transmit the trigger descriptors issued by the processes if the processes have corresponding available entries. Here, the NIC can distribute the entries of the TODS evenly, where the NIC can allocate an equal number of TODS entries to the corresponding processes. The entries can also be distributed non-uniformly (e.g., based on the respective workloads of the processes). The corresponding processes can maintain windows to indicate the number of available entries in the TODS allocated to the processes. When a new trigger operation is generated, the process can check the window size associated with the process to determine whether an entry in the TODS is available for the process. If an entry is available, the process can enqueue the corresponding trigger descriptor into the DWQ. The NIC can then determine the presence of the trigger descriptor based on a register value. For example, the process can set a predetermined value into a register to notify the NIC that the trigger descriptor has been enqueued.

[0023] The NIC can obtain the trigger descriptor from the DWQ based on a read pointer (RP). The read pointer can point to a memory location of the computing system that stores the DWQ. The read pointer can be controlled by the NIC. On the other hand, the write pointer (WP) of the DWQ can be controlled by the corresponding process. Based on the read pointer, the NIC can determine the location of the trigger descriptor and transmit the trigger descriptor to the TODS. Transmitting the trigger descriptor can include reading from the location indicated by the read pointer, enqueuing the trigger descriptor into a segment, and updating the read pointer to indicate a subsequent location in the DWQ.

[0024] When the trigger condition indicated in the trigger descriptor is satisfied, the NIC can obtain the trigger operation from a source buffer (which can be specified by the trigger descriptor) and execute the operation. When executing the trigger operation, the process can increment the window size and allow another trigger descriptor to be queued. If the window of the process is exhausted (i.e., the window size becomes zero), the process is prohibited from inserting or queuing subsequent trigger descriptors into the DWQ. When the window size is incremented to a non-zero value, the process can insert the next triggered descriptor into the DWQ. In this way, the process is prevented from overwhelming the TODS, and the generation rate of the trigger operations of one process does not affect another independent process. Additionally, since the corresponding process can use a subset of the TODS entries allocated to that process, the TODS can support lock-free sharing, where the TODS can be shared among processes without locks.

[0025] Figure 1 Illustrated is an example of adaptive trigger operation management in a NIC according to one aspect of the present application. The computing system 100 (which can be an HPC computing node) can include multiple processing resources 102, a storage medium 104 (e.g., a memory device or a non-volatile persistent storage device), and a NIC 110. Multiple processes (such as processes 112 and 114) can perform computations on the processing resources 102. Examples of processing resources can include, but are not limited to, processors (e.g., central processing units (CPUs), CPU cores) and accelerators (such as graphics processing units (GPUs) or tensor processing units (TPUs)). The data generated by the computations performed by processes 112 and 114 can be used by corresponding processes on other computing nodes. For example, if the computation performed by process 112 includes a distributed summation operation, the output or result of the computation can be sent to the computing node that aggregates the sum.

[0026] The NIC 110 can then use remote access (such as RDMA) to send the data to another computing node. Since process 112 can know that an RDMA operation will be performed by the NIC 110 after the computation is complete, process 112 can determine that sending the data can be a trigger operation, which can be deferred for later execution. Thus, to send the data, process 112 can queue the trigger descriptor associated with the RDMA at the location indicated by the write pointer 124 into the DWQ 120 and set a predetermined value to the register 128. The DWQ 120 can be stored in the storage medium 104. Based on the value in the register 128, the NIC 110 can determine the presence of the descriptor and dequeue the descriptor from the DWQ 120 from the location indicated by the read pointer 122. The trigger descriptor can include information associated with the RDMA operation, such as information about the source buffer, the destination buffer, the size of the data transfer, memory registration, destination process details, trigger counters, completion counters, and trigger thresholds.

[0027] Similarly, to send data, process 114 can enqueue a trigger descriptor associated with RDMA at the location indicated by write pointer 134 into DWQ 130 and set a predetermined value to register 138. DWQ 130 can also be stored in storage medium 104. Based on the value in register 138, NIC 110 can determine the presence of the descriptor and dequeue the descriptor from DWQ 130 from the location indicated by read pointer 132. Here, read pointers 122 and 132 can be controlled by pointer manager 140 of NIC 110. After obtaining the corresponding trigger descriptors from DWQ 120 and 130, pointer manager 140 can update read pointers 122 and 132 to point to the next entry, respectively. Pointer manager 140 can operate based on the Heterogeneous System Architecture (HSA) specification to communicate with other components such as processing resource 102 and storage medium 104. Thus, pointer manager 140 can use HSA to access DWQ 120 and 130 and update read pointers 122 and 132.

[0028] Typically, after a regular descriptor is dequeued, NIC 110 can perform the corresponding data transfer operation without waiting for an event. In contrast, the trigger operations associated with the trigger descriptors in DWQ 120 and 130 can be deferred for later execution. To facilitate deferred execution, NIC 110 can store the trigger descriptors obtained from DWQ 120 and 130 and the corresponding trigger conditions in TODS 150. When the trigger condition is met, NIC 110 can obtain the trigger operation based on the corresponding trigger descriptor stored in TODS 150. NIC 110 can then perform the trigger operation, which can include sending a packet.

[0029] TODS 150 can be deployed in NIC 110 to support trigger operations. TODS 150 can be a hardware entity such as a storage medium. NIC 110 can enqueue the trigger descriptors from DWQ 120 and 130 into the available entries of TODS 150. When the trigger operation is performed, the entries of TODS 150 can be released for reuse. Due to the limitations of the hardware resources of NIC 110, the number of entries in TODS 150 may be limited. Since computing system 100 executes multiple processes 112 and 114, TODS 150 can be shared among processes 112 and 114. Due to the limited availability of hardware resources in the NIC and the sharing of TODS 150 among processes 112 and 114, processes may overbook TODS 150, and some other processes may not be able to utilize TODS 150 due to resource exhaustion. Therefore, the performance of the underutilized processes of computing system 100 may be adversely affected.

[0030] To solve this problem, entries of the TODS 150 can be allocated to processes 112 and 114 (e.g., during library startup). The entries of the TODS 150 can be distributed evenly or unevenly between processes 112 and 114. For example, if there are sixteen entries in the TODS 150, each of processes 112 and 114 can enqueue up to eight entries into the TODS 150 based on an even distribution. On the other hand, if the workload of process 114 is expected to be higher than the workload of process 112, more entries can be allocated to process 114. Processes 112 and 114 can then determine window sizes 152 and 154, respectively. The window size associated with a process can indicate the number of entries allowed for that process to enqueue into the TODS 150. When process 112 generates a new trigger operation, process 112 can check the window size 152 to determine whether an entry in the TODS 150 is available for process 112. If an entry is available, process 112 can enqueue the corresponding trigger descriptor into the DWQ 120 and set a predetermined value in register 128. The NIC 110 can then determine the presence of the trigger descriptor based on the predetermined value in register 128.

[0031] Subsequently, the NIC 110 can read from the location indicated by the read pointer 122 and enqueue the descriptor of the trigger into the TODS 150. When the trigger condition indicated in the trigger descriptor is satisfied, the NIC 110 can obtain the trigger operation from the source buffer (which can be specified by the trigger descriptor) and execute the operation. After the execution of the trigger operation is completed, process 112 can increment the window size 152 and allow another trigger descriptor to be enqueued into the DWQ 120. If the window size 152 is exhausted, process 112 is prohibited from inserting subsequent trigger descriptors into the DWQ 120. When the window size 152 is incremented to a non-zero value, process 112 can insert the next trigger descriptor into the DWQ 120. In this way, processes 112 and 114 are prevented from overwhelming the TODS 150. In addition, since the transfer of the trigger descriptor to the TODS 150 is controlled by the window sizes 152 and 154, the TODS 150 can support lock-free sharing, where the TODS 150 can be shared between processes 112 and 114 without a lock.

[0032] Figure 2Illustrated is an example of facilitating inter-component communication for adaptive trigger operation management in a computing system according to an aspect of the present application. The computing system 200 (which may be an HPC computing node) may include multiple processing resources such as a processor 202 and an accelerator 206 (e.g., a GPU or TPU), a storage medium 204 (e.g., a memory device or a non-volatile persistent storage device), and a NIC 210. Multiple processes (such as processes 212 and 214) may perform computations on the processor 202. Data generated by the computations performed by processes 212 and 214 may be used by corresponding processes on other computing nodes. The NIC 210 may then send the data to another computing node using remote access (such as RDMA). To send data, process 212 may enqueue a trigger descriptor associated with RDMA into DWQ 272. Similarly, to send data, process 214 may enqueue a trigger descriptor associated with RDMA into DWQ 274. DWQ 272 and 274 may be stored in the storage medium 204.

[0033] The NIC 210 may maintain a TODS 250 in the local storage medium for storing trigger descriptors from DWQ 272 and 274. The NIC TODS 250 may be shared between processes 212 and 214 based on window sizes 252 and 254, respectively. The NIC 210 may transfer trigger descriptors from DWQ 272 and 274 to TODS 250. Processes 212 and 214 may check window sizes 252 and 254, respectively, to determine the number of their available entries. Based on window sizes 252 and 254, processes 212 and 214 may enqueue trigger descriptors into DWQ 272 and 274, respectively. The NIC 210 may then transfer the trigger descriptors to TODS 250.

[0034] Processes 212 and 214 may be deployed to perform parallel loop computations executed in a nested and repetitive manner. Such computations are typically performed on different computing nodes and may depend on the computational outputs of each other. Processes 212 and 214 may offload computations from the processor 202 to the accelerator 206 for execution. During operation, when executing on the processor 202, process 212 may enqueue local computations (e.g., computations of distributed operations such as summation) into the execution stream of the accelerator 206 (operation 220). The execution stream may indicate a sequence of operations to be performed by the accelerator 206. Thus, the accelerator 206 may start executing the computations (operation 222). The computations may include collective operations such as barrier, bitwise AND operation, bitwise OR operation, bitwise XOR operation, minimum operation, maximum operation, indexed minimum / maximum operation, or summation operation.

[0035] Since the data generated by the computation will be shared with another compute node at a later time, process 212 can generate a triggering operation that includes a data transfer operation (e.g., sending a packet) based on an RDMA transaction. Process 212 can then enqueue the triggering operation into the execution stream of NIC 210 (operation 224). Enqueuing the triggering operation can include generating a trigger descriptor 260 for the triggering operation and enqueuing it into DWQ 272 if the window size 252 has a non-zero value. The trigger descriptor 260 can include a trigger counter 262, a completion counter 264, and a trigger threshold 266. The threshold 266 can be a predetermined value. The trigger counter 262 facilitates a trigger event. The trigger event can increment the trigger counter 262. When the trigger counter 262 reaches the value of the threshold 266, NIC 210 can determine the location of the triggering operation based on the trigger descriptor 260 and execute the triggering operation. Since the triggering operation can be repeated frequently (e.g., in a loop), the completion counter 264 can indicate the number of times the triggering operation has been executed.

[0036] Thus, process 212 can enqueue a trigger event into the execution stream of accelerator 206 (operation 226). Initially, the values of counters 262 and 264 can be 0, and the value of threshold 266 can be 1. Execution of the trigger event can increment the value of counter 262 to 1, which can then match the threshold 266 and initiate execution of the trigger event. NIC 210 can detect the presence of the trigger descriptor 260 in DWQ 272 and transfer the trigger descriptor 260 from DWQ 272 to an entry in TODS 250 (operation 228). Here, the triggering operation is delayed until the computation of process 212 is complete.

[0037] In addition, process 214 can execute on processor 208 concurrently with process 212. When executing on processor 208, process 214 can enqueue local computations into the execution stream of accelerator 206 (operation 230). If accelerator 206 has not completed the computation of process 212, the computation of process 214 can remain queued in the execution stream. Process 214 can then enqueue a triggering operation into the execution stream of NIC 210 (operation 232). Enqueuing the triggering operation can include generating a trigger descriptor for the triggering operation and enqueuing it into DWQ 274 if the window size 254 has a non-zero value. Process 214 can also enqueue a trigger event into the execution stream of accelerator 206 (operation 234). If NIC 210 detects the presence of the trigger descriptor in DWQ 274, NIC 210 can transfer the trigger descriptor from DWQ 274 to an entry in TODS 250 (operation 236).

[0038] When the computation is complete (operation 238), accelerator 230 may perform subsequent operations in the execution flow, which triggers an event for process 212 (operation 240). Thus, accelerator 230 may increment the value of counter 262 to 1 (e.g., in trigger descriptor 260). Accordingly, NIC 210 may determine that counter 262 has reached threshold 266 and perform a trigger operation (e.g., send a packet including the result of the computation) (operation 242). NIC 210 may send the packet from the egress buffer. To reuse the buffer for subsequent data transfers associated with the next computation, accelerator 206 may wait for the data transfer operation to complete. Accelerator 206 may then determine from NIC 210 that the trigger operation is complete (operation 244). When the trigger operation is complete, accelerator 230 may perform subsequent operations in the execution flow and initiate a computation associated with process 214 (operation 246). Thus, accelerator 206 may begin to perform the computation for process 214 (operation 248). In this way, TODS 250 may incorporate trigger operations from processes 212 and 214 based on window sizes 252 and 254, respectively, without using locks.

[0039] Figure 3A FIG. illustrates an example of partitioning TODS in a NIC among multiple processes in accordance with an aspect of the present application. A computing system 300 (which may be an HPC computing node) may include multiple processing resources 302, such as processors, GPUs, and TPUs, a storage medium 304 (e.g., a memory device or a non-volatile persistent storage device), and a NIC 310. Multiple processes, such as processes 312 and 314, may perform computations on processing resources 302. Data generated by the computations performed by processes 312 and 314 may be used by corresponding processes on other computing nodes. NIC 310 may then send the data to another computing node using remote access, such as RDMA. To send data, process 312 may enqueue a trigger descriptor associated with RDMA into DWQ 320. Similarly, to send data, process 314 may enqueue a trigger descriptor associated with RDMA into DWQ 330. DWQs 320 and 330 may be stored in storage medium 304.

[0040] The NIC 310 can maintain the TODS 350 in a local storage medium for storing trigger descriptors from the DWQs 320 and 330. The NIC 310 can allocate equal portions of the TODS 350 to the processes 312 and 314. The NIC 310 can transfer the triggered descriptors from the DWQs 320 and 330 to the TODS 350, which can operate as a circular queue. The processes 312 and 314 respectively maintain window sizes 352 and 354 to indicate the number of available entries. When the processes 312 and 314 enqueue the triggered descriptors into the DWQs 320 and 330, the NIC 310 can transfer the triggered descriptors to the TODS 350.

[0041] If the TODS 350 includes 16 entries, the window sizes 352 and 354 can each indicate 8 entries. Thus, the window size W associated with the processes 312 and 314 can be 8. Before any trigger operations are issued by the processes 312 and 324, the TODS 350 can be idle and capable of receiving 8 trigger descriptors from each of the processes 312 and 314. Thus, the maximum window size for the processes 312 and 324 can be 8. The window sizes 352 and 354 can be updated and adjusted respectively during the run time of the processes 312 and 314. However, during the execution of the process 312 or 314, the values of the window sizes 352 and 354 do not exceed the maximum window size of 8.

[0042] The processes 312 and 314 can be deployed to perform parallel loop computations executed in a nested and repetitive manner. For example, the processes 312 and 314 can repeatedly perform a summation operation. Assume that an iteration of the computation includes two trigger operations. Thus, an iteration 322 of the process 312 can enqueue two triggered descriptors into the DWQ 320. Similarly, an iteration 332 of the process 314 can enqueue two triggered descriptors 330. The process 312 can update the window size 352 when the iteration 322 is completed. In other words, the window sizes 352 and 354 can be updated at the iteration boundary.

[0043] Figure 3BIllustrated is an example of corresponding window sizes that decrease the availability in the TODS according to one aspect of the present application. Since the execution of processes 312 and 314 can continue, some triggered descriptors can be queued into the TODS 350. The size of the adaptive window can be updated by processes 312 and 314. The new window size can be equal to the previous window size minus the number of entries currently used by the process. For example, if two triggered descriptors generated by process 312 are queued into the DWQ 320, the window size 352 can be decreased by two. Similarly, if four triggered descriptors generated by process 314 are queued into the DWQ 330, the window size 354 can be decreased by four. Thus, the new values of the window sizes 352 and 354 can be six and four, respectively.

[0044] Figure 3C Illustrated is an example of corresponding window sizes that increase the availability in the TODS according to one aspect of the present application. If the corresponding trigger conditions for two trigger operations of process 314 are satisfied, the NIC 310 can execute the trigger operations. The execution of the trigger operations can release the corresponding entries in the TODS 350. Thus, process 314 can identify the completion of the executed trigger operations and increase its window size 354 by two. If the previous value of the window size is 4, the new window size can be 6. In this way, the window size can be adaptive and represent the number of entries currently available to a particular process.

[0045] Figure 4A Presented is a flowchart according to one aspect of the present application, which illustrates an example of a process by which a computing system facilitates the management of adaptive trigger operations. During operation, the computing system can store corresponding descriptors in a first storage medium of the computing system, the descriptors identifying corresponding trigger operations to be executed based on corresponding trigger conditions (operation 402). The trigger conditions can facilitate the delayed execution of the trigger operations. When the trigger conditions are satisfied, the trigger operations can be executed. The computing system can also store the TODS in a second storage medium of the NIC (operation 404). Here, the TODS can include multiple entries, and each entry can store a triggered descriptor. The descriptor can include identification information of trigger information, such as source buffer and target information.

[0046] To facilitate lock - free sharing of the TODS among processes that generate trigger operations, the computing system can determine a first window size for a first process, where the first window size indicates the number of available entries in the TODS (operation 406). The window size can be determined by distributing the entries in the TODS among the processes that generate trigger operations. For example, if there are sixteen entries and two processes, eight entries can be allocated to each process. Thus, the corresponding processes become associated with a predetermined number of entries in the TODS. If the first process and the second process generate trigger operations, the computing system can allocate the first window size and the second window size to the first process and the second process, respectively.

[0047] The computing system can determine whether the first window size indicates availability in the TODS (operation 408). Availability indicates that the number of entries in the TODS allocated to the first process can accommodate another descriptor. Thus, a non - zero value of the first window size can indicate the availability of entries. If the window size indicates availability, the computing system can insert the first descriptor of the first trigger operation generated by the first process into a first work queue (such as a DWQ) (operation 412). The work queue can be in the storage medium (e.g., memory) of the computing system. The value can indicate that a new descriptor has been enqueued. Then, the computing system can determine the presence of the first descriptor in the first work queue at the NIC based on a register value set by the first process (operation 414).

[0048] The computing system can determine the position of the first descriptor in the first work queue at the NIC based on a read pointer. The NIC can control the read pointer of the first work queue and indicate the position of the next descriptor in the DWQ. The read pointer can indicate the next descriptor to be read from the work queue. Thus, the NIC can determine the position based on the read pointer. The computing system can then read from the position in the work queue (operation 416). Since the window size already indicates the availability of entries in the TODS, the computing system can then transfer the first descriptor from the determined position to the TODS (operation 418). Transferring the first descriptor can include reading the first descriptor from that position and storing it in the next available entry in the TODS.

[0049] The NIC can then update the read pointer to indicate a subsequent position in the first work queue (operation 420). When the first descriptor is transferred to the first segment, the entry storing the first descriptor becomes unavailable. Accordingly, the number of available entries in the first segment can be reduced. Since the first window size indicates the number of available entries for the first process, the computing system can decrement the first window size to indicate the updated number of entries available for the first process in the TODS (operation 422). If the first window size indicates unavailability of an entry (e.g., the window size is zero), the computing system can determine that the first segment cannot accommodate another descriptor. Accordingly, the computing system can inhibit inserting the first descriptor into the first work queue (operation 410).

[0050] Figure 4B A flowchart in accordance with an aspect of the present application is presented, which illustrates an example of a process in which a NIC performs a trigger operation from one process based on a trigger descriptor in a local TODS. During the operation, the NIC can detect satisfaction of a trigger condition for a first trigger operation based on the execution of a first process on the processing resources of the computing system (operation 432). As described in connection with Figure 2 the first process can offload computations to processing resources (such as an accelerator) that can generate data to be transferred by the trigger operation. Here, the computation can be part of the execution of the first process. When the execution of the computation is complete, the trigger condition can be satisfied. When the trigger condition is satisfied, the NIC can initiate the trigger operation.

[0051] To initiate the trigger operation, the NIC can obtain a first descriptor from the TODS (operation 434). The first descriptor can include identification information associated with the first trigger operation, such as the location of a source buffer storing the data to be transferred by the trigger operation. The data can be generated by a computation performed by the processing resources and stored in the source buffer (e.g., stored in the storage medium of the computing system). Accordingly, the NIC can obtain the data associated with the first trigger operation based on the information in the first descriptor (operation 436). The NIC can then perform the trigger operation, which can include sending the data generated by the processing resources (e.g., a processor or an accelerator) executing the first process (operation 438). For example, the NIC can send the data to another process via a packet. The acquisition of the descriptor and the subsequent execution of the trigger operation can free the entry storing the descriptor. Accordingly, to reflect the availability of the entry, the NIC can increment the window size (operation 440).

[0052] Figure 5A flowchart in accordance with an aspect of the present application is presented, which illustrates an example of the process by which a NIC performs a trigger operation from another process based on a trigger descriptor in a local TODS. Generally, a set of processes, which may include a first process and a second process, can generate trigger operations and contend for entries in the TODS. The NIC can allocate multiple entries to the corresponding processes in the set of processes. During operation, the NIC can determine a second window size for the second process in the set of processes, where the second window size indicates the number of available entries in the TODS (operation 502). The window size can be determined by distributing the entries in the TODS among the processes that generate the trigger operations. For example, if there are sixteen entries and two processes, eight entries can be allocated to each process. Accordingly, the corresponding processes become associated with a predetermined number of entries in the TODS.

[0053] Each work queue can be associated with a register for notifying the NIC. Accordingly, the NIC can determine the presence of a second descriptor based on the value of the register associated with the second work queue. Thus, when the second process places a descriptor in the second work queue, the second process can set a predetermined value in the register. The NIC can then determine the presence of a second descriptor in the second work queue associated with the second process that identifies the second trigger operation (operation 504). Since the second window size indicates availability, the NIC can transfer the second descriptor from the second work queue to the TODS (operation 506). Due to this transfer, the entry storing the second descriptor may become unavailable. To reflect the unavailability, the NIC can decrement the second window size, which can then indicate the current number of available entries in the TODS for the second process (i.e., the reduced number of entries) (operation 508).

[0054] Figure 6 An example of a computing system with a NIC that facilitates adaptive trigger operation management in accordance with an aspect of the present application is illustrated. The computing system 600 can include a set of processors 602, a memory unit 604, a NIC 606, and a storage medium 608. The memory unit 604 can include a set of volatile memory devices (e.g., dual in-line memory modules (DIMMs)). Additionally, if needed, the computing system 600 can be coupled to a display device 612, a keyboard 614, and a pointing device 616. The storage medium 608 can store an operating system 618. The trigger operation management system 620 and data 636 associated with the trigger operation management system 620 can be maintained and executed from the storage medium 608 and / or the NIC 606. The NIC 606 can also include a storage medium 660, which can store a TODS 662 for storing trigger descriptors.

[0055] The trigger operation management system 620 may include instructions that, when executed by the computing system 600, may cause the computing system 600 (or the NIC 606) to perform the methods and / or processes described in this disclosure. The trigger operation management system 620 may include instructions for assigning entries of the TODS to the process that generates the trigger operation (partitioning subsystem 622), as described in operation 406 in conjunction with Figure 4A as described. The trigger operation management system 620 may also include instructions for determining the presence of a trigger descriptor for a trigger operation in a work queue (e.g., in the memory unit 604) (presence subsystem 624), as described in operation 414 in conjunction with Figure 4A as described. The trigger operation management system 620 may include instructions for determining the availability of an entry for a trigger descriptor based on a window size associated with a process (availability subsystem 626), as described in operation 408 in conjunction with Figure 4A as described.

[0056] The trigger operation management system 620 may also include instructions for transmitting a trigger descriptor to the TODS when an entry is available (transmission subsystem 628), as described in operations 416 and 418 in conjunction with Figure 4A as described. The trigger operation management system 620 may then include instructions for determining a trigger condition that satisfies the trigger operation (execution subsystem 630), as described in operation 432 in conjunction with Figure 4B as described. Additionally, the trigger operation management system 620 may include instructions for performing the trigger operation when the trigger condition is satisfied (execution subsystem 630), as described in operation 438 in conjunction with Figure 4B as described.

[0057] Furthermore, the trigger operation management system 620 may include instructions for adjusting the window size based on the transmission of the trigger descriptor to the TODS and the execution of the trigger operation (window size subsystem 632), as described in operation 422 in conjunction with Figure 4A and operation 440 in Figure 4B as described. The trigger operation management system 620 may also include instructions for sending and receiving data associated with the computations performed by the process (communication subsystem 634), as described in operation 438 in conjunction with Figure 4B as described. The trigger operation management system 620 may also be operated by the control circuit 664 of the NIC 606. The data 636 may include any data that may facilitate the operation of the trigger operation management system 620. The data 636 may include, but is not limited to, data generated by computations performed by a process running on the processor 602.

[0058] Figure 7Illustrated is an example of a computer-readable storage medium that facilitates the management of adaptive trigger operations according to an aspect of the present application. The computer-readable storage medium 700 may include one or more integrated circuits and may store fewer or more instruction sets than those shown in Figure 7 . Further, the storage medium 700 may be integrated with a computer system or integrated in a device capable of communicating with other computer systems and / or devices. For example, the storage medium 700 may be located in the NIC of a computer system.

[0059] The storage medium 700 may include instruction sets 702 to 714 that, when executed, may perform functions or operations similar to those of subsystems 622 to 634 of a trigger operation management system 620, respectively. Here, the storage medium 700 may include a partition instruction set 702; a presence instruction set 704, an availability instruction set 706; a transfer instruction set 708; an execution instruction set 710; a window size instruction set 712; and a communication instruction set 714. Figure 6

[0060] The description herein is presented to enable any person skilled in the art to make and use the invention and is provided in the context of a particular application and its requirements. Various modifications to the disclosed examples will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other examples and applications without departing from the spirit and scope of the invention. Thus, the invention is not limited to the examples shown, but is intended to be accorded the widest scope consistent with the claims.

[0061] One aspect of the present technology may provide a system for managing trigger operations in a computing system. The computing system may include a first storage medium for storing descriptors that identify trigger operations to be performed based on corresponding trigger conditions. The NIC of the computing system may include a second storage medium that stores a data structure. During operation, the system may determine a first window size for a first process, the first window size indicating the number of available entries in the data structure. If the first window size indicates available entries in the data structure, the system may insert a first descriptor of a first trigger operation generated by the first process into a first work queue associated with the first process. The system may determine the position of the first descriptor in the first work queue at the NIC. The system may then transfer the first descriptor from the determined position to the data structure. Subsequently, the system may decrement the first window size, thereby indicating an updated number of entries available in the data structure for the first process. These operations of the system are described in conjunction with Figure 4A .

[0062] In a variation in this regard, the system can detect the satisfaction of the trigger condition of the first trigger operation at the NIC and obtain the first descriptor from the data structure. The system can then perform the first trigger operation based on the information in the first descriptor and increment the first window size. Combined with Figure 4B These operations of the system are described.

[0063] In a further variation, the first trigger operation can be generated based on the execution of the first process on the processor of the computing system. The computing system can also include an accelerator that can execute a trigger event that satisfies the trigger condition and causes the NIC to perform the first trigger operation. Combined with Figure 2 These features of the system are described.

[0064] In a further variation, performing the first trigger operation can include sending a packet including payload data generated by the first process. Combined with Figure 2 This operation of the system is described.

[0065] In a further variation, the trigger condition can be satisfied in response to the completion of the execution of a portion of the first process that generates payload data. Combined with Figure 2 This operation of the system is described.

[0066] In a variation in this regard, the system can decrement the first window size in response to the completion of an iteration of the first process. Here, the amount by which the first window size is decremented can indicate the number of trigger operations in the iteration. Combined with Figure 3A 、 Figure 3B and Figure 3C These features of the system are described.

[0067] In a variation in this regard, the system can determine a second window size for a second process, where the second window size indicates the number of available entries in the data structure. The system can transfer a second descriptor of the second trigger operation from a second work queue associated with the second process to the data structure. The NIC can then decrement the second window size, thereby indicating an updated number of entries available in the data structure for the second process. Combined with Figure 5 These operations of the system are described.

[0068] In a variation in this regard, the system can determine the unavailability of an entry in the data structure based on the first window size. The system can then inhibit inserting a descriptor into the first work queue. Combined with Figure 4A These operations of the system are described.

[0069] In a variation in this regard, the system can determine the presence of a first descriptor in a first work queue at the NIC based on register values set by a first process. The system can then read from a position in the first work queue based on a pointer controlled by the NIC. Subsequently, the NIC can update the pointer to indicate a subsequent position in the first work queue. In combination with Figure 4A These operations of the system are described.

[0070] In the present disclosure, the term "switch" is used in a general sense and can refer to any stand-alone network device or fabric device operating at any network layer. The "switch" should not be construed as limiting the examples of the present invention to Layer 2 networks. Any device that can forward traffic to an external device or another switch can be referred to as a "switch". A switch can also be virtualized.

[0071] In addition, if a network device facilitates communication between networks, the network device can be referred to as a gateway device. Any physical or virtual device (e.g., a virtual machine or a switch operating on a computing device) that can forward traffic to a terminal device can be referred to as a "network device". Examples of "network devices" include, but are not limited to, Layer 2 switches, Layer 3 routers, routing switches, or fabric switches comprising multiple similar or heterogeneous smaller physical and / or virtual switches.

[0072] The term "packet" refers to a group of bits that can be transmitted together over a network. The "packet" should not be construed as limiting the examples of the present invention to a specific layer of the network protocol stack. The "packet" can be replaced by other terms involving a group of bits, such as "message", "frame", "cell", "datagram", or "transaction". In addition, the term "port" can refer to a port that can receive or transmit data. The "port" can also refer to the hardware, software, and / or firmware logic that can facilitate the operation of the port.

[0073] The data structures and code described in this detailed description are typically stored on a computer-readable storage medium, which can be any device or medium that can store code and / or data for use by a computer system. Computer-readable storage media can include, but are not limited to, volatile memory, non-volatile memory, magnetic storage devices, and optical storage devices (such as disks, tapes, CDs (compact discs), DVDs (digital versatile discs or digital video discs)), or other media capable of storing computer-readable media known now or developed later.

[0074] The methods and processes described in the detailed implementation section can be embodied as code and / or data, which can be stored in the computer-readable storage medium as described above. When a computer system reads and executes the code and / or data stored on the computer-readable storage medium, the computer system executes the methods and processes embodied as data structures and code and stored within the computer-readable storage medium.

[0075] The methods and processes described herein can be executed by and / or included in hardware logic blocks or devices. These logic blocks or devices can include, but are not limited to, application-specific integrated circuit (ASIC) chips, field-programmable gate arrays (FPGA), dedicated or shared processors that execute specific software logic blocks and code at a particular time, and / or other programmable logic devices known now or developed later. When the hardware logic blocks or devices are activated, they execute the methods and processes included therein.

[0076] The foregoing description of the examples of the present invention is presented for purposes of illustration and description only. The description is not intended to be exhaustive or to limit the disclosure. Accordingly, many modifications and variations will be apparent to those of ordinary skill in the art. The scope of the present invention is defined by the appended claims.

Claims

1. A method executable on a computing system, the method comprising: storing a descriptor in a first storage medium of the computing system, the descriptor identifying a corresponding trigger operation to be performed based on a corresponding trigger condition; storing a data structure in a second storage medium of a network interface controller (NIC) of the computing system; determining a first window size for a first process, the first window size indicating a number of available entries in the data structure; responsive to the first window size indicating an available entry in the data structure, inserting a first descriptor of a first triggered operation generated by the first process into a first work queue associated with the first process; determining the presence of the first descriptor in the first work queue; transferring the first descriptor from the first work queue to the data structure; as well as The first window size is decremented to indicate an updated number of entries in the data structure available for the first process.

2. The method of claim 1, further comprising: Detecting whether a trigger condition of the first trigger operation is satisfied; obtaining the first descriptor from the data structure; performing the first trigger operation based on the information in the first descriptor; as well as Increment the first window size.

3. The method of claim 2, further comprising: generating, by a processor of the computing system, the first triggering operation based on execution of the first process; as well as A trigger event that satisfies the trigger condition and causes the NIC to perform the first trigger operation is executed by an accelerator of the computing system.

4. The method of claim 3, wherein: Performing the first trigger operation further includes sending a packet including payload data generated by the first process.

5. The method of claim 4, wherein: The trigger condition is satisfied in response to completion of execution of the portion of the first process that generates the payload data.

6. The method of claim 1, further comprising decrementing the first window size in response to completion of an iteration of the first process, and wherein, The amount by which the first window size is decremented indicates the number of trigger operations in the iteration.

7. The method of claim 1, further comprising: determining a second window size for a second process, the second window size indicating a number of available entries in the data structure; transferring a second descriptor of a second triggered operation from a second work queue associated with the second process to the data structure; as well as The second window size is decremented to indicate an updated number of entries in the data structure available for the second process.

8. The method of claim 1, further comprising: determining unavailability of entries in the data structure based on the first window size; as well as Suppress the insertion of triggered descriptors into the first work queue.

9. The method of claim 1, further comprising: determining the presence of the first descriptor in the first work queue based on a register value set by the first process; reading from the first work queue based on a pointer controlled by the NIC; as well as The pointer is updated to indicate a subsequent position in the first work queue.

10. A computing system comprising: processor; a first storage medium storing a descriptor identifying a trigger operation to be performed based on a corresponding trigger condition; a network interface controller (NIC), the network interface controller comprising a second storage medium storing the data structure; A non-transitory computer-readable storage medium storing instructions that, when executed by the processor, cause the computer system to perform the following operations: determining a first window size for a first process, the first window size indicating a number of available entries in the data structure; responsive to the first window size indicating an available entry in the data structure, inserting a first descriptor of a first triggered operation generated by the first process into a first work queue associated with the first process; determining a position of the first descriptor in the first work queue; transferring the first descriptor from the determined location to the data structure; as well as The first window size is decremented to indicate an updated number of entries in the data structure available for the first process.

11. The computing system of claim 10, wherein: The control circuit is further configured to: Detecting whether a trigger condition of the first trigger operation is satisfied; obtaining the first descriptor from the data structure; performing the first trigger operation based on the information in the first descriptor; as well as Increment the first window size.

12. The computing system of claim 11, wherein: The first triggering operation is generated based on execution of the first process on the processor of the computing system; and The computing system further includes an accelerator, and the accelerator is used to execute a trigger event that satisfies the trigger condition and causes the network interface controller to perform the first trigger operation.

13. The computing system of claim 12, wherein: Performing the first trigger operation further includes sending a packet including payload data generated by the first process.

14. The computing system of claim 13, wherein: The trigger condition is satisfied in response to completion of execution of the portion of the first process that generates the payload data.

15. The computing system of claim 10, wherein: The first process is configured to decrement the first window size in response to completion of an iteration of the first process, and wherein an amount by which the first window size is decremented indicates a number of trigger operations in the iteration.

16. The computing system of claim 10, wherein: The control circuit is further configured to: determining a second window size for a second process, the second window size indicating a number of available entries in the data structure; transferring a second descriptor of a second triggered operation from a second work queue associated with the second process to the data structure; as well as The second window size is decremented to indicate an updated number of entries in the data structure available for the second process.

17. The computing system of claim 10, wherein: The first process is further used to: determining unavailability of entries in the data structure based on the first window size; and Suppress the insertion of triggered descriptors into the first work queue.

18. The computing system of claim 10, wherein: The NIC is further configured to: determining the presence of the first descriptor in the first work queue based on a register value set by the first process; reading from the location in the first work queue based on a pointer controlled by the network interface controller; as well as The pointer is updated to indicate a subsequent position in the first work queue.

19. A non-transitory computer-readable storage medium comprising instructions that, when executed by a processor of a computing system, cause the computing system to: storing a descriptor in a data structure of a network interface controller (NIC), the descriptor identifying a trigger operation to be performed based on a corresponding trigger condition; determining a first window size for a first process, the first window size indicating a number of available entries in the data structure; responsive to the first window size indicating an available entry in the data structure, inserting a first descriptor of a first triggered operation generated by the first process into a first work queue associated with the first process; determining a position of the first descriptor in the first work queue based on a register value set by the first process; transferring the first descriptor from the determined location to the data structure; as well as The first window size is decremented to indicate an updated number of entries in the data structure available for the first process.

20. The non-transitory computer readable storage medium of claim 19, wherein: When executed by the processor, the instructions cause the computing system to perform the following operations: Detecting whether a trigger condition of the first trigger operation is satisfied; obtaining the first descriptor from the data structure; performing the first trigger operation based on the information in the first descriptor; as well as Increment the first window size.