Time-aware network data transmission

By using precise time in remote data transmission to control RDMA transmission, the problem of insufficient bandwidth and memory utilization in the prior art is solved, and more efficient and deterministic data transmission is achieved.

CN120128355APending Publication Date: 2025-06-10INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411573632.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-07
Filing Date
2024-11-06
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing remote data transmission protocols are difficult to fully utilize available bandwidth and memory and do not support deterministic data transmission.

Method used

Define time windows and data consumption rates to manage transmission operations by controlling network data transmission using precise time, especially in remote direct memory access (RDMA) transmission.

Benefits of technology

It enables more efficient utilization of memory and bandwidth and provides better certainty and real-time capabilities, suitable for real-time applications in data centers and other computing environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128355A_ABST
    Figure CN120128355A_ABST
Patent Text Reader

Abstract

The name of the invention is time-aware network data transmission. Techniques for time-aware remote data transmission. In a translation protection table (TPT), time may be associated with remote direct memory access (RDMA) operations. The RDMA operation may be granted or limited based on the time in the TPT.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Remote data transfer protocols permit a device to directly access the memory of other devices via a network. However, conventional solutions may not fully utilize available bandwidth, may not fully utilize available memory, and may not support determinism. Brief Description of the Drawings

[0002] To simply identify the discussion of any particular element or action, the most significant digit or digits in the reference numerals refer to the drawing number in which the element was first introduced.

[0003] Figure 1 Illustrates one aspect of a computing architecture according to one embodiment.

[0004] Figure 2 Illustrates an example remote data transfer operation according to one embodiment.

[0005] Figure 3 Illustrates an example remote data transfer operation according to one embodiment.

[0006] Figure 4 Illustrates an example remote data transfer operation according to one embodiment.

[0007] Figure 5 Illustrates an example remote data transfer operation according to one embodiment.

[0008] Figure 6 Illustrates an example of rate-limited data according to one embodiment.

[0009] Figure 7 Illustrates an example data structure according to one embodiment.

[0010] Figure 8 Illustrates a logic flow 800 according to one embodiment.

[0011] Figure 9 Illustrates one aspect of a computing system according to one embodiment. Detailed Description

[0012] The embodiments disclosed herein utilize precise time for network data transmission, including but not limited to remote direct memory access (RDMA) transmission. Applications in a data center or other computing environment may need to run in real time. The embodiments disclosed herein provide a time-aware transmission mechanism for moving data. For example, the embodiments disclosed herein can use precise time to control RDMA transmission (e.g., to define a time window within which data can be transmitted). In some examples, the embodiments disclosed herein can use precise time and the rate of data consumption (e.g., by a processor or other computing component) to control RDMA transmission. In some examples, the embodiments disclosed herein use precise time to control RDMA fences and / or invalidate data. In some embodiments, the RDMA transmission is based on RFC 5040: A Remote Direct Memory Access Protocol Specification, or any other specification defined by the RDMA Consortium. In some embodiments, the RDMA transmission is based on iWARP or InfiniBand technology. The embodiments are not limited to these contexts.

[0013] In some embodiments, a translation protection table (TPT) can be extended to include time values. The time values can be used for any suitable purpose. For example, the time values can indicate when a queue pair assigned to an application can perform transmit and / or receive data operations. In some embodiments, a key can be associated with the time value in the TPT, where the time value indicates when the key is valid or invalid. Similarly, a key can have an associated data rate (e.g., bytes per second, etc.). In some embodiments, memory operations have an associated precise time in the TPT (e.g., loading data into memory, reading data from memory, invalidating data in memory, etc., according to the precise time in the TPT). As another example, the precise time entry in the TPT can be used to load data from a network interface controller (NIC) into cache memory. The embodiments are not limited to these contexts.

[0014] As used herein, precise time and accurate time can be used interchangeably because the systems and techniques discussed herein provide both accurate time management and precise time management for network data transmission.

[0015] By leveraging precise time, the embodiments disclosed herein may allow a computing system to utilize memory more efficiently than systems that do not use precise time. Additionally and / or alternatively, by leveraging precise time, the embodiments disclosed herein may provide better determinism than systems that do not use precise time. Additionally and / or alternatively, by leveraging precise time, the embodiments disclosed herein may allow a computing system to utilize bandwidth more efficiently than systems that do not use precise time. The embodiments are not limited in these contexts.

[0016] Reference is now made to the drawings, in which like reference numerals are always used to refer to like elements. In the following description, numerous specific details are set forth for purposes of explanation in order to provide a thorough understanding thereof. However, the novel embodiments may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form in order to facilitate the description thereof. It is intended to cover all modifications, equivalents, and alternatives consistent with the claimed subject matter.

[0017] In the drawings and the accompanying description, the reference numerals “a” and “b” and “c” (and similar reference numerals) are intended to represent variables of any positive integer. Thus, for example, if an implementation sets the value a = 5, the complete set of components 121 illustrated as components 121-1 through 121-a may include components 121-1, 121-2, 121-3, 121-4, and 121-5. The embodiments are not limited in this context.

[0018] The operation of the disclosed embodiments may be further described with reference to the following drawings. Some of the drawings herein may include a logical flow. Although such a logical flow presented herein may include a particular logical flow, it can be appreciated that the logical flow merely provides an example of how the general functionality described herein can be implemented. Additionally, unless otherwise indicated, a given logical flow need not necessarily be executed in the order presented. Additionally, in some embodiments, not all operations illustrated in the logical flow may be required. Additionally, the logical flow may be implemented by hardware elements, software elements executed by a processor, or any combination thereof. The embodiments are not limited in this context.

[0019] Figure 1FIG. 0 is a schematic diagram illustrating an example system 100 for time-aware network data transmission according to one embodiment. System 100 includes computing systems 102a and 102b communicatively coupled via network 118. As shown, computing systems 102a, 102b include respective processors 104a and 104b, respective memories 106a and 106b, respective network interface controllers (NICs) 108a and 108b, respective processor caches 110a and 110b, and respective devices 112a and 112b, each of which may be implemented at least in part in a circuit module. Computing systems 102a, 102b represent any type of physical and / or virtualized computing system. For example, computing systems 102a, 102b can be computing nodes, servers, infrastructure processing units (IPUs), data processing units (DPUs), networking devices (e.g., switches, routers, etc.), graphics processing units (GPUs), field programmable gate arrays (FPGAs), general purpose GPUs (GPGPUs), accelerator devices, artificial intelligence (AI) processors, vector processors, video processors, or any other type of computing system. In some embodiments, NIC 108a and / or NIC 108b are examples of computing systems 102a, 102b. For example, NIC 108a and / or NIC 108b can be an IPU, a die connected to a processor, and / or an IP core. Embodiments are not limited in these contexts. Devices 112a, 112b represent any type of device, such as a graphics processing unit (GPU), an accelerator device, a storage device, an FPGA, a GPGPU, a tensor flow processor, an accelerator, a peripheral device, or any other type of computing device.

[0020] Although depicted as components of computing systems 102a, 102b, memories 106a, 106b may be located external to computing systems 102a and 102b. For example, memories 106a, 106b can be part of a memory pool. The memory pool can be according to various architectures, such as Compute Express (Compute Express CXL) architecture. Embodiments are not limited in these contexts.

[0021] As shown, application 114a can be executed on processor 104a, and application 114b can be executed on processor 104b. Applications 114a, 114b represent any type of application. For example, applications 114a, 114b can be one or more of the following applications: AI applications, storage applications, networking applications, machine learning (ML) applications, video processing applications, gaming applications, applications that perform mathematical operations, TensorFlow applications, database applications, computer vision applications, compression / decompression applications, encryption / decryption applications, or quantum computing applications. Although depicted as applications, applications 114a, 114b can be implemented as any type of executable code, such as a process, a thread, or a microservice. The embodiments are not limited to these contexts.

[0022] As shown, computing systems 102a, 102b can be coupled to time source 120. The time source can be any type of time source, such as a clock, a network-based clock, an Institute of Electrical and Electronic Engineers (IEEE) 1588 time source, a precise time measurement (PTM) time source (e.g., via a link such as peripheral component interconnect express (PCIe) or Compute Express Link (CXL)), a pulse per second (PPS) (e.g., a pulse generated by a circuit module or other hardware component)-based time source, or any combination thereof. In some embodiments, time source 120 is a component of computing system 102a and computing system 102b. Time source 120 can generally provide precise time data, such as timestamps, to computing systems 102a, 102b. The timestamps can be any time granularity, such as microsecond-level, nanosecond-level, and so on. In some embodiments, synchronization with time source 120 can use a combination of techniques. For example, the synchronization can be based on IEEE 1588 over Ethernet, PTM over PCIe, and PPS between dielets.

[0023] As stated, system 100 is configured to use RDMA to facilitate direct access to resources across network 118. For example, using RDMA, application 114a can access memory 106b via RDMA-enabled NICs 108a, 108b. Similarly, using RDMA, application 114b can access memory 106a via RDMA-enabled NICs 108b, 108a. Although RDMA is used as an example herein, the present disclosure is not limited to RDMA. The present disclosure is equally applicable to other techniques for direct resource access. For example, any data transfer protocol that provides direct data placement functionality and kernel bypass functionality can be used. The direct data placement functionality can provide the source with the information necessary to place data on the target. The kernel bypass functionality can allow a user space process to directly perform fast path operations (issue work requests and retrieve work completions) with hardware without involving the kernel (which can reduce the overhead associated with system calls).

[0024] In some embodiments, RDMA runs on top of a transport protocol, such as RDMA over Converged Ethernet (RoCE), InfiniBand, iWARP, Hyper Ethernet, or another networking protocol. The embodiments are not limited to these contexts.

[0025] RDMA generally enables systems such as computing systems 102a and 102b to communicate across network 118 using high-performance, low-latency, and zero-copy direct memory access (DMA) semantics. Compared to traditional networking stacks, RDMA can reduce host processor utilization, reduce host memory bandwidth associated with the network, and reduce network latency. The RDMA resources (queues, doorbells, etc.) allocated to applications 114a, 114b can be directly mapped into the user or kernel application address space, enabling operating system bypass.

[0026] As shown, NIC 108a includes a Translation Protection Table (TPT) 116a, and NIC 108b includes a TPT 116b. Generally, the TPT can be similar to a processor's memory management unit (MMU). Some entries in the TPT may contain translations from virtual memory pages (or addresses) to physical memory pages (or addresses). Some entries in the TPT can define memory regions or memory windows in memory such as memory 106a or memory 106b. A memory region can be a virtual or logically contiguous region (e.g., a range of memory addresses) of an application address space that can be registered with an operating system (OS). Memory regions allow RDMANICs such as NIC 108a, 108b to perform DMA accesses (for both local and remote requests). Memory regions also enable user space applications such as applications 114a, 114b to process buffers in a virtual address space. Memory windows can be used to assign Steering Tags (STags) to a portion (or window) of a memory region. Some entries in the TPT can contain indications of RDMA keys. Embodiments are not limited in these contexts.

[0027] Entries in the TPTs 116a, 116b can include time information. The time information can include timestamps (or time values) generated by the time source 120. The time information can be a single timestamp or a series of two or more timestamps. Generally, the time information allows the NICs 108a, 108b to enable precise time for RDMA operations. Example RDMA operations that can use the time information in the TPTs 116a, 116b include RDMA write operations, RDMA read operations, RDMA send operations, RDM memory operations (MEMOPS) (such as fast memory registration, invalidation for invalidating a memory region or a memory window, binding operations for creating a memory window bound to an underlying memory region, and the like). In some embodiments, the time information in the TPTs 116a, 116b can include timing information for time-based access to memory regions having different CPU, cache, socket, and / or memory channel affinities. Still further, the applications 114a, 114b can include timing information to provide control over how individual messages go out onto the wire (e.g., send, write, read, etc.) via the time information inserted in the message work descriptors (e.g., work queue entries or WQEs). In such embodiments, the applications 114a, 114b publish these descriptors to the NICs 108a, 108b, thereby requesting data transfer. The NIC 108b or NIC 108b can include timing information in the message work descriptors in one or more entries of the TPTs 116a, 116b.

[0028] For example, the TPTs 116a, 116b can define the times when different hardware and / or software entities are allowed to transfer or receive data. For example, the TPTs 116a, 116b can define the times when the applications 114a, 114b can transfer or receive data. Similarly, the TPTs 116a, 116b can define the times when the queue pairs allocated to the applications 114a, 114b (e.g., see Figure 2 ) can transfer or receive data. Doing so can reduce the number of entries in the TPTs 116a, 116b at a given time. Additionally, the TPTs 116a, 116b can be updated at precise times using the queue pairs that are allowed to transfer or receive data at that time. The TPTs 116a, 116b can also define the times when the RDMA keys are valid (or invalid). In some embodiments, the entries in the TPTs 116a, 116b can include data rates (e.g., bits per time period).

[0029] In some embodiments, access to TPTs 116a, 116b can be restricted to a predetermined time. For example, in some embodiments, entries can be added to TPTs 116a, 116b at a predetermined time. Similarly, in some embodiments, entries in TPTs 116a, 116b can be accessed at a predetermined time. In some embodiments, the entries in TPTs 116a, 116b can be used to provide Quality of Service (QoS). The embodiments are not limited in these contexts.

[0030] Still further, TPTs 116a, 116b can associate time with RDMA MEMOPS. For example, using the time in TPTs 116a, 116b, fast memory registration can occur at an exact time. Similarly, using the time in TPTs 116a, 116b, local invalidation can occur at an exact time, thus freeing up memory for the next application or the next data set of the same application. For example, consider an AI cluster where images are ping-ponged into two memory regions. Using the time in TPTs 116a, 116b, the first memory region can be loaded at an exact time, and processing will stop at an exact time while the second region is being loaded. The processing of the first region has a known exact time as specified by TPTs 116a, 116b. As the AI moves to another buffer, this time can be used to invalidate the region. Since the region has been invalidated, a new RDMA transfer can fill the region, which can save time and memory regions.

[0031] RDMA can support fence operations, e.g., to block an RDMA operation from executing until one or more other RDMA operations have completed. In some embodiments, the time in TPTs 116a, 116b can be used as a fence. Returning to the ping-pong buffer example, there may be a buffer transfer every 16.7 microseconds. An RDMA command can establish a fence (e.g., using TPTs 116a, 116b): the second buffer transfer should not start until 16.7 microseconds (or 16.7 - x microseconds after the first time, where x can be the time to transfer the first byte of data over the network) have passed. Similarly, when the first buffer is "completed", an invalidate command may free up the memory, and the fence can be transferred (knowing that new data will arrive after this time-based invalidation).

[0032] In some such embodiments, images can be received from multiple sources (e.g., cameras, storage devices, etc.). In such embodiments, there may be multiple ping-pong operations in parallel, and the ping-pong operations can be interleaved at different precise times. The parallel operations can use precise times in such a way that they do not overwhelm / overflow / collide with other "jobs or threads" accessing shared resources. The embodiments are not limited to these contexts.

[0033] Figure 2 FIG. 200 is a schematic diagram illustrating an example RDMA send operation. In Figure 2 the example depicted in, application 114a can send data to an unlabeled sink buffer 204 in the memory 106b of computing system 102b for consumption by application 114b.

[0034] Regardless of the type of RDMA operation being performed, applications such as application 114a or application 114b communicate with an RDMA-enabled NIC such as NIC 108a or NIC 108b using one or more queue pairs (QP). A given queue pair can include a send queue (SQ), a receive queue (RQ), and a completion queue (CQ).

[0035] The send queue (also referred to as the "submit queue") is used by the application to issue work requests (WR) to transfer data (or read requests) to a remote system. The work request can have one or more operation codes, which can include: send, send with solicited event, RDMA write, RDMA read, etc. The receive queue is used by the application to issue work requests where buffers are used to place unlabeled messages from a remote system. Elements in the completion queue indicate to the application that the request has been completed. The application can poll the completion queue to identify any completed operations. The CQ can be associated with one or more send queues and / or one or more receive queues.

[0036] Thus, as shown, in the input / output (IO) library 206a, the application 114a is allocated one or more send queues 208a, one or more receive queues 210a, and one or more completion queues 212a. Similarly, as shown, the application 114b is allocated one or more send queues 208b, one or more receive queues 210b, and one or more completion queues 212b in the IO library 206b. Additionally, the applications 114a, 114b may be allocated corresponding sets of RDMA keys (not shown) for each queue pair. For each queue pair, the RDMA key may include P_Key, Q_Key, local key (also known as "L_Key") (STag), remote key (also known as "R_Key") (STag), or S_Key. The P_Key is carried in each transmission packet because the QP is required to be configured for the same partition to communicate. The Q_Key enforces access rights for reliable and unreliable datagram services. During the communication establishment for the datagram service, the nodes exchange the Q_Key of each QP, and the node uses the value it is passed for that remote QP in all packets it sends to the remote QP. When a consumer (e.g., an application) registers a region of memory, the consumer receives the L_Key. The consumer uses the L_Key in a work request to describe local memory to the QP. When a consumer (e.g., an application) registers a region of memory, the consumer receives the R_Key. The consumer passes the R_Key to the remote consumer for use in RDMA operations.

[0037] Thus, to send data to the memory 106b, the application 114a may post an indication of the send operation to the send queue 208a. The data for the send operation may be provided to the source buffer 202 allocated to the application 114a. The source buffer 202 may be referenced as a Scatter Gather List (SGL) in the entry in the send queue 208a. Each element in the SGL may be a triple of [Stag / L-Key, tagged offset (TO), and length]. Similarly, the application 114b may post an indication of the send operation to the receive queue 210b. The entry in the receive queue 210b may be referenced as an SGL, where the SGL element contains a triple of [STag / L-Key, TO, and length].

[0038] As stated, TPT 116a and TPT 116b can store time information. In some embodiments, the time information can be associated with a QP. In some embodiments, the time information is associated with an SQ, an RQ, and / or a CQ. For example, NIC 108a can assign a timestamp in TPT 116a that indicates the time (or range of times) at which transmit queue 208a can transmit data via NIC 108a. Similarly, NIC 108b can assign a timestamp in TPT 116b that indicates the time (or range of times) at which receive queue 210b can receive data.

[0039] Thereafter, NIC 108a can process transmit operations in transmit queue 208a based on the current time. For example, if the timestamp associated with the current time is before the timestamp in TPT 116a, NIC 108a can refrain from processing the transmit operation until the current time is greater than that timestamp. Once the current time is greater than the timestamp in TPT 116a, NIC 108a can process the transmit operation.

[0040] As another example, NIC 108a can assign two or more timestamps in TPT 116a that form one or more time windows during which transmit queue 208a can transmit data via NIC 108a. If the current time is within one of the permitted time windows, NIC 108a can permit data transmission. Otherwise, if the current time is not within one of the permitted time windows, NIC 108a can pause or otherwise restrict data transmission until the current time is within one of the permitted time windows.

[0041] Once the transmit operation is granted, NIC 108a can read data from the source buffer 202 and transmit data 214 for the transmit operation to the computing system 102b via one or more packets. NIC 108b can receive the data 214 and determine whether the receive queue 210b is granted to receive the data 214. For example, NIC 108b can refer to the TPT 116b and determine the time associated with the entry of the receive queue 210b. If the time associated with the entry of the receive queue 210b is a single timestamp, NIC 108b can determine whether the current time is after that timestamp. If the time is after that timestamp, NIC 108b can grant further processing of the data 214. Similarly, if the time in the TPT 116b for the receive queue 210b is two or more timestamps that form one or more time windows during which the receive queue 210b can receive data via NIC 108b, NIC 108b can determine whether the current time is within one of the granted time windows. If the current time is within one of the granted time windows, NIC 108b can grant the incoming data transfer. Otherwise, if the current time is not within one of the granted time windows, NIC 108a can pause or otherwise restrict the data transfer until the current time is within one of the granted time windows.

[0042] Once granted, NIC 108b can determine the location in the sink buffer 204 based on the entry of the operation in the receive queue 210b and store the data 214 at the determined location. NIC 108b can then generate an entry in the completion queue 212b for that transfer, and the entry can be read by the application 114b. The application 114b can then access the data 214 in the sink buffer 204 (or another memory location). However, in some embodiments, the data is stored in the sink buffer 204 before the time when the receive queue 210b is granted to receive the data 214. In such embodiments, the data 214 is stored in the sink buffer 204, and the time in the TPT 116b is used to determine when to generate an entry in the completion queue 212b. The sink buffer 204 can be any type of buffer. For example, the sink buffer 204 can be a low-latency memory such as a cache or a first-in first-out (FIFO) memory structure. In some embodiments, the sink buffer 204 can be part of a cache controlled by NIC 108a or NIC 108b, or a similar memory mechanism. In some embodiments, the sink buffer 204 is a locked region of the cache, which can be reserved by the QoS mechanism. The embodiments are not limited in these contexts.

[0043] NIC 108b can then generate an acknowledgement 216, which is sent to NIC 108a. NIC 108a can receive the acknowledgement 216 and generate an entry in completion queue 212a, the entry indicating successful completion of the send operation.

[0044] As stated, the entries in TPTs 116a, 116b can include data rate information. In some embodiments, the data rate is used to manage application access to RDMA operations without considering timestamps. In some embodiments, the data rate is used in combination with timestamps to manage application access to RDMA operations. For example, in a send operation, the send request can include an exact start time, a data transfer rate, and / or a final transfer time, where the final transfer time is the time by which all data should be transferred.

[0045] Figure 3 is a schematic diagram 300 illustrating an example RDMA write operation. In Figure 3 the example depicted, application 114a can write data 302 to a tagged sink buffer 204 in the memory 106b of computing system 102b.

[0046] To write data 302 to sink buffer 204, application 114a can issue an indication of the write operation to send queue 208a. The data 302 for the write operation can be provided to source buffer 202 allocated to application 114a. Source buffer 202 can be referenced as a scatter-gather list (SGL) in send queue 208a. Each element in the SGL for source buffer in send queue 208a can be a triple of [Stag / L-Key, TO, and length].

[0047] As stated, to write data 302 to sink buffer 204, NIC 108a can determine, based on TPT 116a, whether send queue 208a is permitted to perform the write operation. For example, NIC 108a can determine whether the timestamp of the current time is greater than the timestamp associated with send queue 208a in TPT 116a. If the current timestamp is not greater than the timestamp associated with send queue 208a, then NIC 108a can refrain from writing data 302 until the timestamp is greater than the timestamp associated with send queue 208a. As another example, NIC 108a can determine whether the timestamp of the current time is within one or more time windows associated with send queue 208a in TPT 116a. If the current timestamp is not within one or more time windows associated with send queue 208a, then NIC 108a can refrain from writing data 302 until the timestamp is within one or more time windows associated with send queue 208a.

[0048] Once NIC 108a determines that transmit queue 208a is permitted to write data 302, NIC 108a transmits one or more tagged packets containing data 302 to NIC 108b. NIC 108b then writes data 302 to sink buffer 204 upon receipt. Then, NIC 108b generates and sends an acknowledgement 304 to NIC 108a. When NIC 108a receives acknowledgement 304, NIC 108a generates an entry in completion queue 212a, which indicates successful completion of the write operation.

[0049] In some embodiments, NICs 108a, 108b may use TPTs 116a, 116b to regulate the extraction speed of data payloads. In some embodiments, NICs 108a, 108b may include desired data, data rate, destination time, QoS metrics, and / or SLA metrics in the extraction. In some embodiments, default read / write rates may be implemented. Acknowledgement 304 may use Enhanced Transmission Selection (ETS) to return updates regarding speed change regulation. ETS may be used as a weighted round robin algorithm, weighted queue algorithm, and / or arbiter. Generally, in ETS, each traffic class may have a minimum QoS. However, in some embodiments, all available bandwidth may be used for RDMA operations.

[0050] As stated, entries in TPTs 116a, 116b may contain data rate information. In some embodiments, the data rate is used to manage application access to RDMA operations without considering timestamps. In some embodiments, the data rate is used in conjunction with timestamps to manage application access to RDMA operations. For example, in a write operation, the write request may include an exact start time, data transfer rate, and / or final transfer time, where the final transfer time is the time by which all data should be transferred.

[0051] In some embodiments, the storage for RDMA reads and / or RDMA writes can be decomposed. In such embodiments, the rate can be defined in the TPTs 116a, 116b at memory registration to allocate the target bounce buffer. Thereafter, data can be read and / or written at the defined rate. In some embodiments, the rate is a consumption rate, a time-division multiplexing (TDM) time slot, or any other communication window. In some embodiments, the write confirmation can include details about when a certain data is requested (e.g., at a certain data rate, once, etc.). In some embodiments, the consumption rate can be dynamic. In some embodiments, the consumption rate can be based on hints from the memory, such as PCIe or CXL hints.

[0052] Figure 4 FIG. 400 is a schematic diagram illustrating an example RDMA read operation. In Figure 4 the example depicted, the application 114a can read data 406 from a source buffer 402 in the memory 106b of the computing system 102b to a sink buffer 404 in the memory 106a of the computing system 102a.

[0053] To read the data 406, the application 114a can issue an indication of the read operation to the receive queue 210a. The indication of the read operation in the receive queue 210a can include an indication of the source buffer 402, which can be referenced in a scatter-gather list (SGL). The SGL for the indication of the read operation in the receive queue 210a can be a triple of [STag / R-Key, TO, and length]. Additionally, the indication of the read operation in the receive queue 210a can include an indication of the sink buffer 404, which can be referenced as an element of the SGL, where the SGL includes a triple of [STag / L-Key, TO, and length].

[0054] As stated, to read data 406, NIC 108a can determine, based on TPT 116a, whether receive queue 210a is permitted to perform a read operation. For example, NIC 108a can determine whether the timestamp of the current time is greater than the timestamp associated with receive queue 210a in TPT 116a. If the current timestamp is not greater than the timestamp associated with receive queue 210a, then NIC 108a can refrain from reading data 302 until the timestamp is greater than the timestamp associated with transmit queue 208a. As another example, NIC 108a can determine whether the timestamp of the current time is within one or more time windows associated with receive queue 210a in TPT 116a. If the current timestamp is not within one or more time windows associated with receive queue 210a, then NIC 108a can refrain from reading data 406 until the timestamp is within one or more time windows associated with receive queue 210a.

[0055] Once NIC 108a determines that receive queue 210a is permitted to receive data 406, NIC 108a transmits a request 408 to NIC 108b. Request 408 can be an untagged message in a single Ethernet packet. Request 408 can include an indication that the operation is an RDMA read operation, a remote STag, and a TO.

[0056] NIC 108b then reads data 406 from source buffer 402 in response to receiving request 408. Then, NIC 108b generates and sends one or more packets containing data 406 to NIC 108a. In response to receiving the packets of data 406, NIC108a writes data 406 to sink buffer 404. NIC 108a can also generate an entry in completion queue 212a that indicates successful completion of the read operation.

[0057] In some embodiments, the entries in TPT 116a, 116b can be used by NICs 108a, 108b to regulate the extraction speed of elements in transmit queues 208a, 208b. In some embodiments, NICs 108a, 108b can submit a read to indicate to the remote host how to regulate the read data speed (e.g., based on one or more of data rate, destination time, QOS metrics, and / or service level agreement (SLA) metrics). In some embodiments, NICs 108a, 108b can break a read operation into smaller sub-operations to prevent memory overflow.

[0058] As stated, the entries in TPTs 116a, 116b may include data rate information. In some embodiments, the data rate is used to manage an application's access to RDMA operations without considering timestamps. In some embodiments, the data rate is used in combination with timestamps to manage an application's access to RDMA operations. For example, in a read operation, the read request may include an exact start time, a data transfer rate, and / or a final transfer time, where the final transfer time is the time by which all data should be transferred.

[0059] Figure 5 FIG. 500 is a schematic diagram illustrating an example unreliable datagram (UD) send operation. In Figure 5 the example depicted, application 114a may send data 506 to sink buffer 504 in memory 106b of computing system 102b for consumption by application 114b.

[0060] To send data to memory 106b, application 114a may post an indication of the send operation to send queue 208a. The data for the send operation may be provided to source buffer 502 allocated to application 114a. Source buffer 502 may be referenced as a scatter-gather list (SGL) in an entry in send queue 208a. Each element in the SGL may be a triple of [Stag / L-Key, TO, and length]. Similarly, application 114b may post an indication of the send operation to receive queue 210b, which may include an entry for sink buffer 504. The entry in receive queue 210b may be referenced as an SGL, where the SGL element contains a triple of [STag / L-Key, TO, and length].

[0061] As stated, to send data 506 to sink buffer 504, NIC 108a may determine, based on TPT 116a, whether send queue 208a is permitted to perform the send operation. For example, NIC 108a may determine whether the timestamp of the current time is greater than the timestamp associated with send queue 208a in TPT 116a. If the current timestamp is not greater than the timestamp associated with send queue 208a, then NIC 108a may refrain from sending data until the timestamp is greater than the timestamp associated with send queue 208a. As another example, NIC 108a may determine whether the timestamp of the current time is within one or more time windows associated with send queue 208a in TPT 116a. If the current timestamp is not within one or more time windows associated with send queue 208a, then NIC 108a may refrain from sending data until the timestamp is within one or more time windows associated with send queue 208a.

[0062] Once the transmit operation is granted, NIC 108a can read data from source buffer 502 and transmit data 506 for the transmit operation to computing system 102b via one or more packets. Since NIC 108b will not send an acknowledgment for the UD transmit operation, NIC 108a generates an entry for the transmit operation in completion queue 212a after transmitting data 506.

[0063] NIC 108b can receive data 506 and determine whether receive queue 210b is granted to receive data 214. For example, NIC 108b can refer to TPT 116b and determine the time associated with an entry of receive queue 210b. If the time associated with an entry of receive queue 210b is a single timestamp, NIC 108b can determine whether the current time is after the timestamp. If the time is after the timestamp, NIC 108b can grant further processing of data 214. Similarly, if the time in TPT 116b for receive queue 210b is two or more timestamps forming one or more time windows during which receive queue 210b can receive data via NIC 108b, NIC 108b can determine whether the current time is within one of the granted time windows. If the current time is within one of the granted time windows, NIC 108b can grant an incoming data transfer. Otherwise, if the current time is not within one of the granted time windows, NIC 108a can pause or otherwise limit the data transfer until the current time is within one of the granted time windows.

[0064] Once granted, NIC 108b can determine a location in sink buffer 504 based on the entry of the operation in receive queue 210b and store data 506 at the determined location. NIC 108b can then generate an entry in completion queue 212b for the transfer, which can be read by application 114b. Then, application 114b can access data 506 in sink buffer 204 (or another memory location). However, in some embodiments, data 506 is stored in sink buffer 204 before the time when receive queue 210b is granted to receive data 506. In such embodiments, data 506 is stored in sink buffer 204 and the time in TPT 116b is used to determine when to generate an entry in completion queue 212b.

[0065] In some embodiments, the information in TPTs 116a, 116b can be used for one-way transmission time determination. Conventional RDMA transmissions can use two-way latency measurements and divide it by two to determine the one-way latency. However, if there is congestion in one direction and not in the other, the divide-by-two method does not show the accurate congestion in one direction. Thus, if an acknowledgement (ACK) or other RDMA response includes the one-way latency, the RDMA protocol can respond appropriately. In other words, one may want rate control in one direction, but if all the variations and delays are in the other path, the rate control occurs in the wrong direction.

[0066] In addition, NICs 108a, 108b can use the exact time in TPTs 116a, TPTs 116b for the arriving data stream, and the exact time at the device or host, to place the data into the associated memories 106a, 106b in a timely manner. For example, the RDMA blocks of NICs 108a, 108b can save and merge the incoming RDMA data so that it delivers the data to the CPU and / or memory in a timely and accurate manner for using the data.

[0067] Figure 6 FIG. 600 is a schematic diagram illustrating an example of using exact time to fill data in a cache according to one embodiment. As shown, schematic diagram 600 includes a computing system 102a, which further includes an L2 cache 602, a shared interconnect 604, and an L3 cache 606. The NIC-accessible portion 608 of the L3 cache 606 can be accessible to NIC 108a for specified functions. For example, RDMA queue entries and pointers (send queue entries, completion queue entries, head / tail pointers, etc.) can be stored in the NIC-accessible portion 608 for quick checks by the processor 104a by locking the locations of these variables / parameters / entries to the L3 cache 606. Similarly, NIC 108a can reserve a portion of the NIC-accessible portion 608 of the L3 cache. Doing so allows NIC 108a to place data directly into the L3 cache in a timely manner, rather than storing the data in the memory 106a before the data is brought up to the L3 cache 606. Doing so can save memory, power, latency, and bandwidth (PCIe, memory, cache).

[0068] For example, NIC 108a can receive data (not shown). NIC 108a can refer to TPT 116a to determine the rate at which the processor 104a consumes data from the NIC-accessible portion 608 of the L3 cache 606. NIC 108a can precisely determine when to write data to the NIC-accessible portion 608 of the L3 cache 606 based on the consumed rate and the current time. For example, if the processor 104a is consuming data from the NIC-accessible portion 608 at a rate of 1 gigabit per second, NIC 108a can cause the NIC-accessible portion 608 to be filled with 1 gigabit of data at precise time intervals per second. In some embodiments, NIC 108a can save and merge incoming RDMA data such that NIC 108a delivers the data to the L3 cache 606 in a timely and precise manner for use by the processor 104a.

[0069] Although the NIC-accessible portion 608 is used as an example mechanism to reserve the L3 cache 606, embodiments are not limited to these contexts. For example, other types of caches can be accessible to NICs 108a, 108b, such as L1 cache or L2 cache 602, dedicated caches, memory structures (e.g., FIFO, scratchpad memory), and / or memory structures within the cache (e.g., FIFO, scratchpad memory). The NIC-accessible portion 608 can be accessed by NICs 108a, 108b using any suitable technique. Examples of techniques for accessing the NIC-accessible portion 608 include direct data I / O techniques ( DDIO) and cache stashing. Embodiments are not limited to these contexts.

[0070] Figure 7 A data structure 702 is illustrated. The data structure 702 can represent some or all of the entries in TPT 116a or TPT 116b. As shown, the data structure 702 includes a timestamp field 704 for one or more timestamps. Similarly, one or more timestamps 706 can be stored in the timestamp field 704 of the entries in TPT 116a or TPT 116b. In some embodiments, the timestamp 706 is stored in one bit or a reserved bit of the data structure 702. Embodiments are not limited to these contexts.

[0071] In some embodiments, the timestamp 706 includes 64 bits. However, in some embodiments, the timestamp 706 can include fewer or more than 64 bits. The number of bits can be based on fractions of nanoseconds, tens of picoseconds, or microseconds. The number of bits can be based on a known time period and / or epoch. Embodiments are not limited to this context.

[0072] Figure 8 An embodiment of a logic flow 800 is illustrated. The logic flow 800 may represent some or all of the operations performed by one or more embodiments described herein. For example, the logic flow 800 may include some or all of the operations for providing time-aware network data transmission. The embodiments are not limited in this context.

[0073] In block 802, the logic flow 800 associates time with a Remote Direct Memory Access (RDMA) operation via circuitry and in a Translation Protection Table (TPT). For example, NIC 108a may associate a timestamp with application 114a in TPT 116a. In block 804, the logic flow 800 permits or restricts an RDMA operation via circuitry based on the time in the TPT. For example, if NIC 108a determines that an RDMA is permitted based on the time in TPT 116a and the current time, then NIC 108a processes the RDMA operation. Otherwise, NIC 108a may restrict or otherwise inhibit processing of the RDMA operation until it is permitted.

[0074] More generally, embodiments disclosed herein provide RDMA TPTs, such as TPT116a, 116b, which use precise time as an attribute. An RDMA translation protection table that uses precise time as an attribute. Precise time can allow certain queue pairs to have the time at which they are allowed to transmit data. Precise time can allow certain queue pairs to have the time at which they are allowed to receive data. The use of precise time reduces the number of entries in TPT 116a, 116b at a given time. TPT 116a, 116b can be updated at precise times with queue pairs that should occur during that time. In some embodiments, an RDMA cache memory, such as L3 cache 606, is loaded based on a known time specified in TPT 116a, 116b. In some embodiments, an RDMA cache memory, such as L3 cache 606, is evicted based on a known time specified in TPT 116a, 116b. Doing so can provide a perceived performance improvement due to the precise time of updating the memory. In some embodiments, coalescing occurs based on the precise time in TPT 116a, 116b. In some embodiments, coalescing occurs in NIC 108a, NIC108b, computing system 102a, computing system 102b, or any of their components. In some embodiments, invalidation occurs based on the precise time in TPT 116a, 116b. In some embodiments, access to TPT116a, 116b may be allowed only when there is sufficient credit. In some embodiments, credit is assigned at precise times (e.g., assigned to application 114a, application 114b, other hardware elements, or other software elements). In some embodiments, credit is updated at precise times. In some embodiments, credit accumulates based on a rate at precise times.

[0075] In some embodiments, in TPT 116a, 116b, an RDMA key is associated with precise time. The key can be a P_Key. The key can be a Q_Key. The key can be an L_key. The key can be an R_Key. The key can be an S_Key. In some embodiments, the key association can occur during a precise time specified in TPT 116a, 116b. Thus, the key association may not occur during other precise times. In some embodiments, precise time can indicate when the key is valid. In some embodiments, the time is a time window. In some embodiments, precise time indicates when the key is invalid. In some embodiments, the time is a time window. In some embodiments, precise time indicates a data rate. In some embodiments, the rate indicates the number of bytes or bits per time period. In some embodiments, the rate is in bytes per nanosecond. In some embodiments, the rate is in kilobits (KB) per microsecond. In some embodiments, the rate is in megabits (MB) per microsecond.

[0076] In some embodiments, the processing of the RDMA key can occur in software. In some embodiments, the processing of the RDMA key can occur in hardware. In some embodiments, the hardware is an IPU. In some embodiments, the hardware is a NIC. In some embodiments, the hardware is a processor. In some embodiments, the hardware is a GPU. In some embodiments, precise time is used as part of an RDMA QoS scheme. In some embodiments, the QoS scheme provides an SLA and / or a service level objective (SLO).

[0077] In some embodiments, the QoS scheme uses time slots. In some embodiments, precise time is arranged to provide a class of processing or a class of traffic among RDMA consumers. In some embodiments, the QoS scheme allows for proper multi-host operation. In some embodiments, the QoS scheme provides the time required by the QoS scheme for RDMA operations for any-sized process. In some embodiments, the QoS scheme provides the bandwidth required by the QoS scheme for RDMA operations for any-sized process. In some embodiments, the QoS scheme provides the latency required by the QoS scheme for RDMA operations for any-sized process.

[0078] In some embodiments, one or more RDMA parameters are assigned at a time. In some embodiments, the RDMA parameters are part of a connection profile. In some embodiments, the connection profile is determined at precise time or within a precise time window. In some embodiments, the connection profile cannot be changed after precise time. In some embodiments, the RDMA parameters are part of the resource allocation to the host system. In some embodiments, the resources are available during precise time or within a precise time window. In some embodiments, the resource allocation is credit. In some embodiments, the resource allocation is credit per time. In some embodiments, the resource allocation is memory allocation. In some embodiments, the resource allocation is bandwidth allocation. In some embodiments, the resource allocation is based on the queue depth. In some embodiments, the queue depth is the send queue depth. In some embodiments, the queue depth is the completion queue depth. In some embodiments, the queue depth is the receive queue depth. In some embodiments, the queue depth is related to the depth of the queue pair.

[0079] In some embodiments, the remote data transfer protocol allows for precise-time-based isolation. The time can be specified in TPTs 116a, 116b. In some embodiments, the precise time indicates a transmission window to a specific host. In some embodiments, this time window is used for isolation. In some embodiments, the isolation is between two or more hosts (host isolation). In some embodiments, the isolation is between two or more network ports. In some embodiments, the isolation is between two or more memory regions. In some embodiments, the memory regions are in the cache. In some embodiments, the memory regions are in the host memory. In some embodiments, the memory regions span a CXL interface. In some embodiments, the memory regions span a PCIe interface. In some embodiments, the memory regions span an Ethernet interface. In some embodiments, the isolation is between two or more virtual machines (virtual machine isolation). In some embodiments, the isolation is between two or more devices. In some embodiments, the device includes at least one processor. In some embodiments, the device includes at least one GPU. In some embodiments, the device includes at least one accelerator. In some embodiments, the accelerator is an AI accelerator. In some embodiments, the accelerator is a Math accelerator.

[0080] In some embodiments, the precise time (e.g., in TPTs 116a, 116b) indicates one or more "no-transfer" windows to a specific host. In some embodiments, this time window is used for isolation. In some embodiments, the isolation is between two or more hosts. In some embodiments, the isolation is between two or more ports. In some embodiments, the isolation is between two or more virtual machines. In some embodiments, the isolation is between two or more devices. In some embodiments, the device includes at least one processor. In some embodiments, the device includes at least one GPU. In some embodiments, the device includes at least one accelerator. In some embodiments, the accelerator is an AI accelerator. In some embodiments, the accelerator is a Math accelerator.

[0081] In some embodiments, RDMA applications (such as applications 114a, 114b) use time to coordinate their computation and communication phases into chunks. In some embodiments, the computation phase performs mathematical operations. In some embodiments, the mathematical operations are computationally intensive. In some embodiments, the operation is an artificial intelligence operation. In some embodiments, the operation is a machine learning operation. In some embodiments, the operation is a vector processing operation. In some embodiments, the operation is a video processing operation. In some embodiments, the operation is a tensor flow operation. In some embodiments, the operation is a compression or decompression operation. In some embodiments, the operation is an encryption or decryption operation. In some embodiments, the operation is a quantum computing operation.

[0082] In some embodiments, the communication phase performs a transmission. In some embodiments, the transmission contains data. In some embodiments, the data is part of an RDMA read. In some embodiments, the data is part of an RDMA write. In some embodiments, the data is part of an RDMA send. In some embodiments, the transmission contains control information. In some embodiments, the control information includes setting up a transmission. In some embodiments, the control information includes tearing down a transmission. In some embodiments, the control information is an ACK message. In some embodiments, the control information is a send queue entry.

[0083] In some embodiments, the phase is specified by an exact time, e.g., based on TPTs 116a, 116b. In some embodiments, the phase is seen at the processor. In some embodiments, the phase is seen at the memory. In some embodiments, the memory is a cache memory. In some embodiments, the memory is a first-in first-out (FIFO) memory. In some embodiments, the memory is a structure within the cache memory. In some embodiments, the memory is a host memory. In some embodiments, the memory is on a different die. In some embodiments, the memory is connected via CXL. In some embodiments, the phase is seen at the memory controller. In some embodiments, the memory controller is a cache controller. In some embodiments, the phase is seen at the server. In some embodiments, the phase is seen at the client. In some embodiments, the phase is seen at the NIC. In some embodiments, the phase is seen at the IPU. In some embodiments, the phase is seen at the GPU. In some embodiments, the phase is seen at the accelerator. In some embodiments, the phase is seen at a common point. In some embodiments, the phase is seen at a network device. In some embodiments, the network is an Ethernet network. In some embodiments, the network is a Super Ethernet network. In some embodiments, the network is PCI or PCIe. In some embodiments, the network is CXL. In some embodiments, the device is used for storing data.

[0084] In some embodiments, the precise time in TPTs 116a, 116b includes an adjustment for latency. In some embodiments, the latency is the one-way latency from the client to the server. In some embodiments, the latency is the one-way latency from the server to the client. In some embodiments, the latency is the one-way latency between a point associated with the client and a point associated with the server. In some embodiments, the point is one or more of the following: NIC, IPU, cache management unit, memory management unit, CPU, GPU, video processing unit (VPU), video transcoding unit (VCU), tensor processing unit (TPU), switch, network device, CXL device, memory, cache, a portion of the cache (e.g., NIC-accessible portion 608), or die.

[0085] In some embodiments, the latency is the one-way latency between a point associated with the server and a point associated with the client. In some embodiments, the point is one or more of the following: NIC, IPU, cache management unit, memory management unit, CPU, GPU, VPU, VCU, TPU, switch, network device, CXL device, memory, cache, a portion of the cache (NIC-accessible portion 608), or die.

[0086] In some embodiments, the data block improves latency. In some embodiments, the data block improves performance. In some embodiments, the data block improves area. In some embodiments, the data block is all or a part of the memory. In some embodiments, this part of the memory is defined by the network interface controller. In some embodiments, this part of the memory is defined by an algorithm to improve the performance of the application. In some embodiments, the application is RDMA. In some embodiments, this part of the memory is part of a buffering scheme. In some embodiments, the buffering scheme is a ping-pong scheme. In some embodiments, one set of memory is being used while another set of memory is loading data. In some embodiments, the buffering scheme has more than two memory regions. In some embodiments, at least one region is used for computing. In some embodiments, at least one region is being loaded with new data. In some embodiments, at least one region is being invalidated, evicted, dumped, cleared, etc. In some embodiments, at least one region is being fenced off.

[0087] In some embodiments, the time is precise time. In some embodiments, the precise time is delivered via IEEE 1588. In some embodiments, PTM is used to deliver the precise time. In some embodiments, PPS is used to deliver the precise time. In some embodiments, the precise time is delivered via a clock. In some embodiments, proprietary methods are used to deliver the precise time. In some embodiments, the precise time is made effective by one or more of IEEE 1588, PTM, PPS, or proprietary methods. In some embodiments, the precise time is within the guaranteed limits for the device. In some embodiments, the device is a warehouse computer. In some embodiments, the device is in a single data center. In some embodiments, the device is in multiple data centers. In some embodiments, the device is distributed. In some embodiments, the device is distributed across geographical regions.

[0088] In some embodiments, the time is accurate time. In some embodiments, the time is both precise and accurate. In some embodiments, an RDMA application delivers precise time between RDMA nodes. In some embodiments, the nodes are servers. In some embodiments, the nodes are clients. In some embodiments, the nodes are NICs or IPUs. In some embodiments, the nodes are CPUs. In some embodiments, the nodes are GPUs. In some embodiments, the RDMA application receives precise time from a precise time source. In some embodiments, the precise time source is IEEE 1588, PTM, PPS, a proprietary method, any combination thereof.

[0089] In some embodiments, the RDMA application coordinates with the transport layer. In some embodiments, the transport layer regulates the traffic speed. In some embodiments, the transport layer delivers time. In some embodiments, the transport layer is directed by the RDMA application. In some embodiments, the transport layer is time-aware. In some embodiments, the transport layer is implemented in a device such as computing system 102a or computing system 102b. In some embodiments, the device is a server. In some embodiments, the device is a client system. In some embodiments, the device is a NIC or IPU. In some embodiments, the device is a CPU. In some embodiments, the device is a GPU. In some embodiments, the device is a switch. In some embodiments, the device is an accelerator.

[0090] In some embodiments, the RDMA application uses precise time instead of interrupts. In some embodiments, the RDMA application uses precise time to avoid large context switches. In some embodiments, the RDMA application uses precise time to avoid thrashing. In some embodiments, the thrashing is cache thrashing. In some embodiments, the transmission is throttled. In some embodiments, the throttling speed for communication is different from the computing throttling speed. In some embodiments, the communication is throttled. In some embodiments, the data used for computing is throttled.

[0091] In some embodiments, the RDMA memory operations (MEMOPS) are associated with time, such as in TPT 116a, TPT116b. In some embodiments, the time is precise time. In some embodiments, the time is accurate time. In some embodiments, the time is both precise and accurate. In some embodiments, the MEMOPS is fast memory registration. In some embodiments, the MEMOPS is bound memory window operation. In some embodiments, the binding operation is type A binding operation. In some embodiments, the binding operation is type B binding operation. In some embodiments, the MEMOPS is locally invalidated. In some embodiments, the local invalidation releases memory for the next application. In some embodiments, the memory is part of a cache. In some embodiments, the cache is an L1 cache. In some embodiments, the cache is an L2 cache. In some embodiments, the cache is a last-level cache. In some embodiments, the cache is designed for data transfer. In some embodiments, the memory is a FIFO. In some embodiments, the transfer involves a hint. In some embodiments, the transfer involves a steering tag. In some embodiments, the transfer occurs over PCIe. In some embodiments, the transfer occurs over CXL. In some embodiments, the local invalidation releases memory for data. In some embodiments, the memory is part of a cache. In some embodiments, the cache is an L1 cache. In some embodiments, the cache is an L2 cache. In some embodiments, the cache is a last-level cache. In some embodiments, the cache is designed for data transfer. In some embodiments, the memory is a FIFO. In some embodiments, the transfer involves a hint. In some embodiments, the transfer involves a steering tag. In some embodiments, the transfer occurs over PCIe. In some embodiments, the transfer occurs over CXL. In some embodiments, the memory is part of a buffering scheme. In some embodiments, the buffering scheme is a ping-pong scheme. In some embodiments, the buffering scheme has more than two buffers. In some embodiments, the buffering scheme allows for low latency. In some embodiments, the local invalidation reduces power. In some embodiments, the local invalidation reduces latency. In some embodiments, the local invalidation improves performance.

[0092] In some embodiments, the RDMA operation uses time as part of the operation. In some embodiments, the RDMA operation is a read operation. In some embodiments, the RDMA operation is a write operation. In some embodiments, the RDMA operation is a send operation. In some embodiments, the RDMA operation is an atomic operation. In some embodiments, the operation sets a lock. In some embodiments, the lock occurs at an exact time as specified by TPT 116a or TPT 116b. In some embodiments, the lock occurs within an exact time window as specified by TPT 116a or TPT 116b. In some embodiments, the RDMA operation removes the lock. In some embodiments, the removal of the lock occurs at an exact time as specified by TPT 116a or TPT 116b. In some embodiments, the removal occurs within an exact time window as specified by TPT 116a or TPT 116b. In some embodiments, the RDMA operation is a dump clear operation. In some embodiments, the dump clear time is given as part of another operation. In some embodiments, the other operation is a write. In some embodiments, the other operation is a read. In some embodiments, the other operation is a send. In some embodiments, the other operation is an atomic operation. In some embodiments, the time is an exact time. In some embodiments, the time is an accurate time. In some embodiments, the time indicates the time of the operation. In some embodiments, the time of the operation is a start time. In some embodiments, the time of the operation is an end time. In some embodiments, the time of the operation is a time window. In some embodiments, the time of the operation is associated with a rate. In some embodiments, the time of the operation is considered at the client. In some embodiments, the time of the operation is considered by the client. In some embodiments, the time is considered by the server. In some embodiments, the time of the operation is considered at the server or at the client. In some embodiments, the latency is estimated. The latency can be a best case latency. The latency can be a worst case latency.

[0093] In some embodiments, the latency is one-way latency. In some embodiments, the latency is from the server to the client. In some embodiments, the latency is from the client to the server. In some embodiments, the latency considers more than one-way latency. In some embodiments, the latency is used for calculation to determine the time for performing RDMA operations. In some embodiments, this time is the start time. In some embodiments, the start time is at the transmitter. In some embodiments, the start time is at the receiver. In some embodiments, the start time is at the device. In some embodiments, the device is a server. In some embodiments, the device is a client. In some embodiments, the device is a CPU. In some embodiments, the device is a GPU. In some embodiments, the device is an accelerator. In some embodiments, the device is an AI accelerator. In some embodiments, the device is a device. In some embodiments, the device is a storage device. In some embodiments, this time is the end time or the completion time or the final transmission time. In some embodiments, this time is used together with the data size (e.g., transmission rate). In some embodiments, there is a rate (e.g., data rate). In some embodiments, the rate is in bytes per nanosecond. In some embodiments, the rate is in KB per microsecond. In some embodiments, the rate is in MB per microsecond. In some embodiments, the rate is used to adjust the speed. In some embodiments, the operation includes an update to the completion queue. In some embodiments, the operation includes an update to the send queue. In some embodiments, the operation includes an update to the receive queue. In some embodiments, the operation is in an area defined by the network interface controller. In some embodiments, the operation is managed by a QoS algorithm. In some embodiments, the RDMA operation uses precise time instead of interrupts. In some embodiments, the RDMA operation uses precise time to avoid large context switches. In some embodiments, the RDMA operation uses precise time to avoid jitter. In some embodiments, the jitter is cache jitter.

[0094] In some embodiments, the barrier operation uses time. In some embodiments, the barrier operation is part of an RDMA operation. In some embodiments, the time is an exact time. In some embodiments, the time is an accurate time. In some embodiments, the time indicates the time of an operation. In some embodiments, the time of the operation is a start time. In some embodiments, the time of the operation is an end time. In some embodiments, the time of the operation is an invalid time. In some embodiments, the barrier operation occurs in association with a second command. In some embodiments, the second command is an invalid command. In some embodiments, the second command is an erase command. In some embodiments, the second command is an eviction command. In some embodiments, the second command is used to store new data. In some embodiments, the second command is a read. In some embodiments, the second command is a write. In some embodiments, the second command is a send. In some embodiments, the second command is atomic. In some embodiments, the barrier operation occurs after a previous command. In some embodiments, the second command is a read. In some embodiments, the second command is a write. In some embodiments, the second command is a send. In some embodiments, the second command is atomic. In some embodiments, the barrier operation occurs in a device between a server and a client, including the server and the client. In some embodiments, the device is an IPU. In some embodiments, the device is a NIC. In some embodiments, the device is a CPU. In some embodiments, the device is a GPU. In some embodiments, the device communicates with PCIe. In some embodiments, the device communicates with CXL. In some embodiments, the device communicates with Universal Chiplet Interconnect Express (UCIe). In some embodiments, the device is a memory device. In some embodiments, the device is a chiplet. In some embodiments, the device is an Ethernet device. In some embodiments, the device is a Super Ethernet device. In some embodiments, the device is a client. In some embodiments, the device is a server. In some embodiments, the device contains a memory for storage. In some embodiments, the storage is a cache. In some embodiments, the storage is a FIFO. In some embodiments, the storage has multiple levels. In some embodiments, the barrier operation uses exact time instead of an interrupt. In some embodiments, the barrier operation uses exact time to avoid large context switches. In some embodiments, the barrier uses exact time to avoid thrashing. In some embodiments, the thrashing is cache thrashing.

[0095] In some embodiments, the RDMA transfer uses a one-way transfer time. In some embodiments, the time is from a first device to a second device. In some embodiments, the first device is a server. In some embodiments, the first device is a client. In some embodiments, the second device is a server. In some embodiments, the second device is a client. In some embodiments, at least one of the devices is a virtual machine. In some embodiments, at least one of the devices is a CPU. In some embodiments, at least one of the devices is a GPU. In some embodiments, at least one of the devices is on a dielet. In some embodiments, at least one of the devices is an accelerator. In some embodiments, at least one of the devices is a memory device. In some embodiments, the memory device is a cache. In some embodiments, the transfer time is based on a probe. In some embodiments, the direction of the probe is in the direction of data transfer. In some embodiments, the time is from the client to the server. In some embodiments, the time is from the client NIC / IPU to the server or the server's NIC / IPU. In some embodiments, the time is from the server NIC / IPU to the client or the client's NIC / IPU. In some embodiments, the time is from the client memory / cache to the server or the server's NIC / IPU / memory / cache. In some embodiments, the time is from the server memory / cache to the client or the client's NIC / IPU / memory / cache. In some embodiments, the one-way transfer time is included in the RDMA message. In some embodiments, the message is an acknowledgment. In some embodiments, the time includes the latency through PCIe. In some embodiments, the time includes the latency through CXL. In some embodiments, the time includes the latency through a proprietary connection. In some embodiments, the time includes the latency through UCIe. In some embodiments, the time includes the latency between one or more dielets. In some embodiments, the one-way transfer time is used to adjust the RDMA transfer. In some embodiments, the one-way transfer time is used to adjust one or more RDMA data rates. In some embodiments, the one-way transfer time is consumed by the RDMA application. Doing so can improve performance, improve latency, and / or reduce power. In some embodiments, doing so can indicate the transmission start time. In some embodiments, doing so can indicate the transmission end time. In some embodiments, doing so can indicate "must be received before a specified time". In some embodiments, doing so indicates a transmission window.

[0096] In some embodiments, the apparatus implements a data transfer protocol to place data into a timely memory before an exact time. In some embodiments, the protocol is RDMA. In some embodiments, the data is from an RDMA write. In some embodiments, the data is from an RDMA read. In some embodiments, the data is from an RDMA send. In some embodiments, the data is from an RDMA atomic. In some embodiments, the apparatus is a NIC. In some embodiments, the apparatus is an IPU. In some embodiments, the apparatus is a CPU. In some embodiments, the apparatus is a GPU. In some embodiments, the apparatus is an FPGA. In some embodiments, the apparatus is an accelerator. In some embodiments, the accelerator performs artificial intelligence operations. In some embodiments, the accelerator performs machine learning operations. In some embodiments, the accelerator processes vectors. In some embodiments, the accelerator processes video. In some embodiments, the apparatus is a memory. In some embodiments, the memory is a cache associated with a CPU, GPU, IPU, and / or NIC. In some embodiments, the timely memory is part of the cache. The cache can be an L1 cache. The cache can be an L2 cache. The cache can be a last-level cache. The cache can be a specialized cache. The cache can be a cache for an AI accelerator. The cache can be used for a networking accelerator. In some embodiments, a portion of the cache is reserved. In some embodiments, the cache reservation mechanism is via a NIC. In some embodiments, the timely memory is a FIFO. In some embodiments, the timely memory is a scratchpad. In some embodiments, the placed data is data. In some embodiments, the placed data is information. In some embodiments, the information is related to a queue. In some embodiments, the information is related to a queue pair. In some embodiments, the information is related to a send queue. In some embodiments, the information is related to a completion queue. In some embodiments, the information is related to a receive queue. In some embodiments, the NIC / IPU / DPU stores the RDMA data until the time specified in TPT 116a or TPT 116b, and then transfers it to the apparatus. In some embodiments, the storing is combining multiple incoming data segments. In some embodiments, the apparatus is a CPU. In some embodiments, the apparatus is a memory. In some embodiments, the apparatus is a cache. In some embodiments, the transfer is delayed. In some embodiments, the time is accurate. In some embodiments, the time is precise.

[0097] In some embodiments, the processing device stores RDMA control information in a low-latency memory. In some embodiments, the low-latency memory is a cache. In some embodiments, the cache is an L1 cache. In some embodiments, the cache is an L2 cache. In some embodiments, the cache is a last-level cache. In some embodiments, the cache is a dedicated cache. In some embodiments, the low-latency memory implements a reservation algorithm. In some embodiments, the RDMA control information is one or more of the following: queue entries, send queues, completion queues, pointers, head pointers, tail pointers, buffer lists, physical memory locations, parameters, and / or variables. In some embodiments, the storage saves power. In some embodiments, the storage reduces latency. In some embodiments, the storage reduces cache thrashing. In some embodiments, the storage increases bandwidth. In some embodiments, the bandwidth is Ethernet bandwidth. In some embodiments, the bandwidth is Ultra Ethernet bandwidth. In some embodiments, the bandwidth is PCIe bandwidth. In some embodiments, the bandwidth is CXL bandwidth. In some embodiments, the bandwidth is memory bandwidth. In some embodiments, the bandwidth is die-to-die interface bandwidth.

[0098] In some embodiments, time-aware RDMA interacts with a time-aware transport protocol. In some embodiments, the protocol is the Transmission Control Protocol (TCP). In some embodiments, TCP is time-aware. In some embodiments, the protocol is the User Datagram Protocol (UDP). In some embodiments, UDP is time-aware. In some embodiments, the protocol is the reliable transport RT protocol. In some embodiments, RT is time-aware. In some embodiments, the protocol is a proprietary protocol. In some embodiments, the protocol is NVLink. In some embodiments, the NVLink protocol is time-aware. In some embodiments, the protocol is Ethernet. In some embodiments, the protocol is RoCE. In some embodiments, Ethernet or RoCE is time-aware. In some embodiments, the transport protocol consumes precise time. In some embodiments, the transport protocol consumes accurate time. In some embodiments, the transport protocol indicates a start time. In some embodiments, the transport protocol indicates an end time. In some embodiments, the transport protocol specifies a data rate. In some embodiments, the transport protocol specifies the precise time for performing the transmission. In some embodiments, the transport protocol limits when RDMA can transmit data. In some embodiments, the transport protocol updates the window for transmission. In some embodiments, RDMA can preempt packets based on precise time. In some embodiments, high-QoS data can be injected into the flow of lower-QoS data during transmission. In some embodiments, preemption occurs in a scheduler. In some embodiments, packets are preempted.

[0099] In some embodiments, an RDMA device or application uses precise time to handle faults. In some embodiments, the fault is an error. In some embodiments, the fault is reported. In some embodiments, the fault causes an interruption. In some embodiments, the fault causes packet drops. In some embodiments, one or more faults accumulate until a precise time. In some embodiments, multiple faults are reported at a precise time. In some embodiments, all faults are reported at a precise time or at a precise rate. In some embodiments, faults are handled during a precise time window. In some embodiments, faults are considered to be after a precise time. In some embodiments, faults occur at a precise time or within a precise time window. In some embodiments, faults are ignored during a precise time window. In some embodiments, the error handling of faults occurs at a precise time or during a precise time window.

[0100] In some embodiments, an RDMA device or application uses time to reduce multi-flow jitter. In some embodiments, the multi-flow contains different types of flows. In some embodiments, the multi-flow contains different destination devices. In some embodiments, the destination device is a host in a multi-host system. In some embodiments, the destination device is one or more of the following devices: CPU, GPU, NIC, IPU, DPU, TPU, accelerator, VCU, CPU, device, memory device, cache, and / or a part of the cache. In some embodiments, the time is precise time. In some embodiments, the precise time is delivered via IEEE 1588. In some embodiments, the precise time is delivered using PTM. In some embodiments, the precise time is delivered using PPS. In some embodiments, the precise time is delivered using a proprietary method. In some embodiments, the precise time is made effective via IEEE 1588, PTM, PPS, or a proprietary method. In some embodiments, the precise time is within the guaranteed limit for the device. In some embodiments, the device is a warehouse computer. In some embodiments, the device is in a single data center. In some embodiments, the device is in multiple data centers. In some embodiments, the devices are dispersed. In some embodiments, they are dispersed across geographical regions. In some embodiments, the time is accurate time. In some embodiments, the time is both precise and accurate.

[0101] In some embodiments, a first device uses time to control incast of a second device. In some embodiments, the first device runs an RDMA application. In some embodiments, the device is a NIC, an IPU, or a GPU. In some embodiments, the control method is to adjust the speed. In some embodiments, the control method uses time slots or TDM type operations. In some embodiments, the second device is a networking device. In some embodiments, it is a switch. In some embodiments, it is a router. In some embodiments, it is a network device. In some embodiments, it is an IPU. In some embodiments, it is a NIC. In some embodiments, the time is precise time. In some embodiments, the precise time is delivered via IEEE 1588. In some embodiments, PTM is used to deliver the precise time. In some embodiments, PPS is used to deliver the precise time. In some embodiments, a proprietary method is used to deliver the precise time. In some embodiments, the precise time is made effective via IEEE 1588, PTM, PPS, or a proprietary method. In some embodiments, the precise time is within guaranteed limits for the device. In some embodiments, the device is a warehouse computer. In some embodiments, the device is in a single data center. In some embodiments, the device is in multiple data centers. In some embodiments, the devices are distributed. In some embodiments, they are distributed across geographical regions. In some embodiments, the time is accurate time. In some embodiments, the time is both precise and accurate. In some embodiments, one of the devices runs Map Reduce. In some embodiments, one of the devices runs Hadoop. In some embodiments, one of the devices is a collection device. In some embodiments, the collection device sorts the results of multiple devices.

[0102] In some embodiments, one or more computing elements can use time for dynamic rerouting of flows. The dynamic rerouting can be an equal cost multi-path (ECMP) routing scheme. In some embodiments, the routing scheme uses a 5-tuple to determine which port to transmit data on. In some embodiments, the same flow exits from the same switch port. In some embodiments, the dynamic rerouting statically sprays packets (load balances). In some embodiments, AI and high performance computing (HPC) have a small number of endpoints. In some embodiments, an endpoint is a peer. In some embodiments, gigabits per second are provided to the peer.

[0103] Some ECMP collisions may occur. In some embodiments, some uplinks are pathologically congested while other uplinks are empty. In some embodiments, network utilization is insufficient, especially in the case of small connection counts. In some embodiments, dynamic routing can help make a selection to switch to a different egress port for higher utilization. Doing so may cause packets to become out-of-order and / or make the path delays different. In some embodiments, the routing scheme includes a dynamic routing algorithm. In some embodiments, the time is precise time. In some embodiments, the time is accurate time. In some embodiments, rerouting of a flow occurs at precise time or during a precise time window. In some embodiments, rerouting of a flow uses one-way delay measurements. In some embodiments, the lowest expected one-way delay is targeted for routing.

[0104] In some embodiments, PCC (programmable congestion control) may be implemented. In some embodiments, PCC may use a time control window to schedule traffic. In some embodiments, PCC may measure the time it takes for a packet to cross the network and then adjust the traffic accordingly. PCC may use round-trip time (in the absence of precise time) to make a decision. As the data transmitter adds its own timestamp and its own time base to the receiving endpoint, it returns it to the transmitter in an ACK or other packet. Using round-trip time over time for many packets, the control algorithm reacts accordingly.

[0105] In some embodiments, one-way time measurements are used with PCC. If PTP is used with two endpoints that are synchronized (e.g., to 100 ns), then the one-way delay can be measured and congestion in the path to the data transmitter only is seen in the ACK response (regardless of congestion / jitter in the ACK direction). 100 ns is in the noise of network jitter. The ACK message can be given a dedicated high priority.

[0106] Figure 9FIG. illustrates an embodiment of system 900. System 900 is a computer system with multiple processor cores, such as a distributed computing system, a supercomputer, a high-performance computing system, a computing cluster, a mainframe computer, an infrastructure processing unit (IPU), a data processing unit (DPU), a microcomputer, a client-server system, a personal computer (PC), a workstation, a server, a portable computer, a laptop computer, a tablet computer, a handheld device (such as a personal digital assistant (PDA)) or other devices for processing, displaying, or transmitting information. Similar embodiments may include, for example, entertainment devices such as a portable music player or a portable video player, a smartphone or other cellular phone, a telephone, a digital camera, a digital still camera, an external storage device, or the like. Further embodiments implement larger-scale server configurations. Examples of IPUs include Pensando IPU. Examples of DPUs include Fungible DPU, OCTEON and ARMADA DPU, NVIDIA DPU, Neoverse N2 DPU, and Pensando DPU. In other embodiments, system 900 may have a single processor (with one core) or more than one processor. Note that the term "processor" refers to a processor with a single core or a processor package with multiple processor cores. In at least one embodiment, computing system 900 represents a component of system 100. More generally, computing system 900 is configured to implement all of the logic, systems, logical flows, methods, devices, and functionality described herein with reference to the previous figures.

[0107] As used in this application, the terms "system" and "component" and "module" are intended to refer to computer-related entities (hardware, combinations of hardware and software, software, or software in execution), examples of such entities being provided by exemplary system 900. For example, a component can be, but is not limited to, a process running on a processor, a processor, a hard disk drive, multiple storage drives (of optical and / or magnetic storage media), an object, an executable, an executing thread, a program, and / or a computer. By way of illustration, both an application running on a server and the server can be components. One or more components can reside within a process and / or an executing thread, and a component can be localized on one computer and / or distributed between two or more computers. Moreover, components can be communicatively coupled to each other via various types of communication media to coordinate operations. Coordination may involve the one-way or two-way exchange of information. For example, components can transfer information in the form of signals passed through a communication medium. The information can be implemented as signals assigned to various signal lines. In such an assignment, each message is a signal. However, alternative embodiments may alternatively employ data messages. Such data messages can be sent across various connections. Exemplary connections include parallel interfaces, serial interfaces, and bus interfaces.

[0108] As Figure 9As shown in, system 900 includes a system-on-chip (SoC) 902 for mounting platform components. The system-on-chip (SoC) 902 is a point-to-point (P2P) interconnect platform that includes a first processor 904 and a second processor 906 coupled via a point-to-point interconnect 970 (such as an Ultra Path Interconnect (UPI)). In other embodiments, system 900 can be another bus architecture (such as a multi-drop bus). Additionally, each of processor 904 and processor 906 can be a processor package with multiple processor cores, respectively including (one or more) cores 908 and (one or more) cores 910. Although system 900 is an example of a dual-socket (2S) platform, other embodiments can include more than two sockets or one socket. For example, some embodiments can include a quad-socket (4S) platform or an octa-socket (8S) platform. Each socket is a mount for a processor and can have a socket identifier. Note that the term platform can refer to a motherboard that mounts certain components (such as processor 904 and chipset 932). Some platforms can include additional components, and some platforms can include sockets for mounting processors and / or chipsets. Additionally, some platforms can not have sockets (such as an SoC or the like). Although depicted as SoC 902, one or more of the components of SoC 902 can also be included in a single die package, a multi-chip module (MCM), a multi-die package, a chiplet, a bridge, and / or an interposer. Thus, embodiments are not limited to an SoC. SoC 902 is an example of a computing system 102a and a computing system 102b.

[0109] Processor 904 and processor 906 can be any of a variety of commercially available processors, including but not limited to and processors; application, embedded, and security processors; and and processors; IBM and Cell processors; and similar processors. Dual microprocessors, multi-core processors, and other multi-processor architectures can also be used as processor 904 and / or processor 906. Additionally, processor 904 does not need to be identical to processor 906.

[0110] The processor 904 includes an integrated memory controller (IMC) 920, a point-to-point (P2P) interface 924, and a P2P interface 928. Similarly, the processor 906 includes an IMC 922, a P2P interface 926, and a P2P interface 930. The IMC 920 and the IMC 922 couple the processor 904 and the processor 906 to corresponding memories (e.g., memories 916 and 918), respectively. The memories 916 and 918 may be part of the main memory of the platform (e.g., dynamic random-access memory (DRAM)), such as double data rate type 4 (DDR4) or type 5 (DDR5) synchronous DRAM (SDRAM). In this embodiment, the memories 916 and 918 are locally attached to the corresponding processors (e.g., processor 904 and processor 906). In other embodiments, the main memory may be coupled to the processors via a bus and a shared memory hub. The processor 904 includes registers 912, and the processor 906 includes registers 914.

[0111] The system 900 includes a chipset 932 coupled to the processor 904 and the processor 906. Additionally, the chipset 932 may be coupled to a storage device 950, for example, via an interface (I / F) 938. The I / F 938 may be, for example, a Peripheral Component Interconnect Express (PCIe) interface, a Compute Express Link (CXL) interface, or a Universal Chiplet Interconnect Express (UCIe) interface. The storage device 950 may store instructions executable by circuit modules of the system 900 (e.g., processor 904, processor 906, GPU 948, accelerator 954, visual processing unit 956, or the like).

[0112] The processor 904 is coupled to the chipset 932 via the P2P interface 928 and P2P 934, while the processor 906 is coupled to the chipset 932 via the P2P interface 930 and P2P 936. Direct Media Interface (DMI) 976 and DMI 978 may couple the P2P interface 928 and P2P 934, and the P2P interface 930 and P2P 936, respectively. The DMI 976 and the DMI 978 may be high-speed interconnects that facilitate, for example, eight gigatransfers per second (GT / s), such as DMI 3.0. In other embodiments, the processor 904 and the processor 906 may be interconnected via a bus.

[0113] The chipset 932 may include a controller hub, such as a platform controller hub (PCH). The chipset 932 may include a system clock for performing clock functions and an interface for an I / O bus, such as a universal serial bus (USB), a peripheral component interconnect (PCI), a CXL interconnect, a UCIe interconnect, a serial peripheral interconnect (SPI), an integrated interconnect (I2C), and the like, to facilitate the connection of peripheral devices on the platform. In other embodiments, the chipset 932 may include more than one controller hub, such as a chipset with a memory controller hub, a graphics controller hub, and an input / output (I / O) controller hub.

[0114] In the depicted example, the chipset 932 is coupled to a trusted platform module (TPM) 944 and a UEFI, BIOS, FLASH circuit module 946 via an I / F 942. The TPM 944 is a dedicated microcontroller designed to protect the hardware by integrating encryption keys into the device. The UEFI, BIOS, FLASH circuit module 946 may provide pre-boot code.

[0115] In addition, the chipset 932 includes an I / F 938 to couple the chipset 932 to a high-performance graphics engine, such as a graphics processing circuit module or a graphics processing unit (GPU) 948. In other embodiments, the system 900 may include a flexible display interface (FDI) (not shown) between the processor 904 and / or the processor 906 and the chipset 932. The FDI interconnects the graphics processor cores in one or more of the processor 904 and / or the processor 906 with the chipset 932.

[0116] The system 900 is operable to communicate with wired and wireless devices or entities via a network interface controller (NIC) 980 using the IEEE 802 standard family, such as a wireless device operatively positioned in wireless communication (e.g., IEEE 802.11 air modulation technology). This includes at least Wi-Fi (or wireless fidelity), WiMax, and Bluetooth TMWireless technologies, 3G, 4G, LTE, 5G, 6G wireless technologies, etc. Thus, the communication can be a predefined structure like a conventional network or just an ad hoc communication between at least two devices. Wi-Fi networks use radio technologies called IEEE 802.11x (a, b, g, n, ac, ax, etc.) to provide secure, reliable, and fast wireless connectivity. Wi-Fi networks can be used to connect computers to each other, to the Internet, and to wired networks (which use IEEE 802.3-related media and functions).

[0117] Additionally, accelerator 954 and / or vision processing unit 956 can be coupled to chipset 932 via I / F 938. Accelerator 954 represents any type of accelerator device (e.g., data streaming accelerator, cryptographic accelerator, crypto co-processor, offload engine, etc.). Examples of accelerator 954 include AMD or accelerators, HGX and SCX accelerators, and ARM Ethos-U NPU.

[0118] Accelerator 954 can be a device that includes circuit modules for accelerating copy operations, data encryption, hash value calculation, data comparison operations (including comparison of data in memory 916 and / or memory 918), and / or data compression. For example, accelerator 954 can be a USB device, a PCI device, a PCIe device, a CXL device, a UCIe device, and / or an SPI device. Accelerator 954 can also include circuit modules arranged to perform machine learning (ML)-related operations (e.g., training, inference, etc.) for an ML model. Generally, accelerator 954 can be specifically designed to perform compute-intensive operations, such as hash value calculation, comparison operations, cryptographic operations, and / or compression operations, in a more efficient manner than when executed by processor 904 or processor 906. Since the load on system 900 can include hash value calculation, comparison operations, cryptographic operations, and / or compression operations, accelerator 954 can greatly improve the performance of system 900 for these operations.

[0119] The accelerator 954 can be implemented as any type of device, such as a coprocessor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a functional block, an IP core, a graphics processing unit (GPU), a processor with a specific instruction set for accelerating one or more operations, or other hardware accelerators capable of performing the functions described herein. In some embodiments, the accelerator 954 can be encapsulated in a discrete package, an add-in card, a chipset, a multi-chip module (e.g., dielets, tubs, etc.), and / or a SoC. The embodiments are not limited to these contexts.

[0120] The accelerator 954 can include one or more dedicated work queues and one or more shared work queues (both not shown). Generally, the shared work queue is configured to store descriptors submitted by multiple software entities. The software can be any type of executable code, such as a process, a thread, an application, a virtual machine, a container, a microservice, etc., that share the accelerator 954. For example, the accelerator 954 can be shared according to the Single Root I / O virtualization (SR-IOV) architecture and / or the Scalable I / O virtualization (S-IOV) architecture. The embodiments are not limited to these contexts. In some embodiments, the software atomically submits a descriptor to the accelerator 954 using an instruction via a non-posted write (e.g., a deferred memory write (DMWr)). An example of an instruction for atomically submitting a work descriptor to the shared work queue of the accelerator 954 is the ENQCMD command or instruction supported by the Instruction Set Architecture (ISA) (which may be referred to herein as "ENQCMD"). However, any instruction with a descriptor (which includes an indication of the operation to be performed, the source virtual address for the descriptor, the destination virtual address for the device-specific register of the shared work queue, the virtual address of the parameters, the virtual address of the completion record, and the identifier of the address space of the submitting process) represents an instruction for atomically submitting a work descriptor to the shared work queue of the accelerator 954. The dedicated work queue can accept job submissions via a command such as the movdir64b instruction.

[0121] A variety of I / O devices 960 and a display 952 are coupled to a bus 972, along with a bus bridge 958 that couples bus 972 to a second bus 974, and an I / F 940 that connects bus 972 to chipset 932. In one embodiment, the second bus 974 may be a low pin count (LPC) bus. A variety of devices may be coupled to the second bus 974, including, for example, a keyboard 962, a mouse 964, and a communication device 966.

[0122] In addition, an audio I / O 968 may be coupled to the second bus 974. Many of the communication device 966 and the I / O devices 960 may reside on a system-on-chip (SoC) 902, while the keyboard 962 and the mouse 964 may be additional peripheral devices. In other embodiments, some or all of the I / O devices 960 and the communication device 966 are additional peripheral devices and do not reside on the system-on-chip (SoC) 902.

[0123] The components and features of the devices described above may be implemented using any combination of discrete circuit modules, application specific integrated circuits (ASICs), logic gates, and / or single-chip architectures. Additionally, where appropriate, the features of the devices may be implemented using a microcontroller, a programmable logic array, and / or a microprocessor, or any combination of the foregoing. Note that hardware, firmware, and / or software elements may be collectively or individually referred to herein as "logic" or "circuitry".

[0124] It will be appreciated that the exemplary devices shown in the block diagrams above may represent one functionally described example of many potential implementations. Thus, the partitioning, omission, or inclusion of the functions depicted in the blocks in the figures does not imply that the hardware components, circuitry, software, and / or elements for implementing these functions will necessarily be partitioned, omitted, or included in an embodiment.

[0125] At least one computer-readable storage medium may include instructions that, when executed, cause a system to perform any of the computer-implemented methods described herein.

[0126] Some embodiments may be described using the phrase "one embodiment" or "an embodiment" and derivatives thereof. These terms mean that the particular features, structures, or characteristics described in connection with the embodiment are included in at least one embodiment. The appearance of the phrase "in one embodiment" in various places in the specification does not necessarily all refer to the same embodiment. Additionally, unless otherwise indicated, the features described above are considered to be usable in any combination with each other. Thus, any features discussed separately may be employed in combination with each other, unless it is stated that the features are incompatible with each other.

[0127] Regarding the notations and nomenclatures used in this general reference, the detailed descriptions herein can be presented in accordance with program processes executed on a computer or a network of computers. These program descriptions and representations are used by those skilled in the art to most effectively convey the essence of their work to other technicians in the field.

[0128] A process is herein and generally conceived to be a self-consistent sequence of operations leading to a desired result. These operations are those that require physical manipulation of physical quantities. Ordinarily, though not necessarily, these quantities take the form of electrical, magnetic, or optical signals capable of being stored, transmitted, combined, compared, and otherwise manipulated. Sometimes (mainly for reasons of common usage) it proves convenient to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like. However, it should be noted that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to those quantities.

[0129] In addition, the manipulations performed are often referred to in terms such as addition or comparison, which are commonly associated with intellectual operations performed by a human operator. In any of the operations described herein that form part of one or more embodiments, in most cases, no such ability of a human operator is necessary or desirable. On the contrary, the operations are machine operations. Useful machines for performing the operations of the various embodiments include a general-purpose digital computer or similar devices.

[0130] The terms “coupled” and “connected” along with their derivatives can be used to describe some embodiments. These terms are not necessarily intended as synonyms for each other. For example, the terms “connected” and / or “coupled” can be used to describe some embodiments to indicate that two or more elements are in direct physical or electrical contact with each other. However, the term “coupled” can also mean that two or more elements are not in direct contact with each other but still cooperate or interact with each other.

[0131] The various embodiments also relate to devices or systems for performing these operations. This device can be specifically constructed for the required purpose, or it can include a general-purpose computer (such as selectively activated or reconfigured by a computer program stored in the computer). The processes presented herein are not inherently related to a particular computer or other device. Various general-purpose machines can be used with programs written in accordance with the teachings herein, or it may prove convenient to construct more specialized devices to perform the required methods. From the description given, the required structure for the various such machines will become apparent.

[0132] The foregoing description includes examples of the disclosed architecture. Of course, it is not possible to describe every conceivable combination of methodologies and / or components, but one of ordinary skill in the art will recognize that many additional combinations and permutations are possible. Accordingly, the novel architecture is intended to embrace all such alterations, modifications, and variations that fall within the spirit and scope of the appended claims.

[0133] The various elements of the apparatus described previously with reference to the drawings may comprise a variety of hardware elements, software elements, or a combination of both. Examples of hardware elements may include devices, logic devices, components, processors, microprocessors, circuits, processors, circuit elements (e.g., transistors, resistors, capacitors, inductors, etc.), integrated circuits, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), digital signal processors (DSPs), field-programmable gate arrays (FPGAs), memory units, logic gates, registers, semiconductor devices, chips, microchips, chip sets, and the like. Examples of software elements may include software components, programs, applications, computer programs, application programs, system programs, software development programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, procedures, software interfaces, application program interfaces (APIs), instruction sets, computing code, computer code, code segments, computer code segments, words, values, symbols, or any combination thereof. However, determining whether an embodiment is implemented using hardware elements and / or software elements may vary according to any number of factors, such as the desired computing rate, power level, heat tolerance, processing cycle budget, input data rate, output data rate, memory resources, data bus speed, and other design or performance constraints (as desired for a given implementation).

[0134] One or more aspects of at least one embodiment can be implemented by representative instructions stored on a machine-readable medium that represent various logic within a processor. When read by a machine, the instructions cause the machine to fabricate the logic to perform the techniques described herein. Such representations (referred to as “IP cores”) can be stored on a tangible machine-readable medium and supplied to various customer or production facilities to be loaded into the manufacturing machines that make the logic or processor. For example, some embodiments can be implemented using a machine-readable medium or article that can store instructions or a set of instructions that, if executed by a machine, can cause the machine to perform the methods and / or operations according to the embodiments. Such a machine can include, for example, any suitable processing platform, computing platform, computing device, processing device, computing system, processing system, computer, processor, or the like, and can be implemented using any suitable combination of hardware and / or software. The machine-readable medium or article can include, for example, any suitable type of memory unit, memory device, memory article, memory medium, storage device, storage article, storage medium, and / or storage unit, such as memory, removable or non-removable media, erasable or non-erasable media, writable or rewritable media, digital or analog media, hard disk, floppy disk, Compact Disk Read Only Memory (CD-ROM), Recordable Compact Disk (CD-R), Rewriteable Compact Disk (CD-RW), optical disk, magnetic media, magneto-optical media, removable memory card or disk, various types of Digital Versatile Disk (DVD), magnetic tape, cassette tape, or the like. The instructions can include any suitable type of code implemented using any suitable high-level, low-level, object-oriented, visual, compiled, and / or interpreted programming language, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, encrypted code, and the like.

[0135] It will be appreciated that the exemplary devices shown in the block diagrams described above can represent one functionally described example of many potential implementations. Thus, the partitioning, omission, or inclusion of the functions depicted in the blocks in the figures does not imply that the hardware components, circuits, software, and / or elements for implementing these functions will necessarily be partitioned, omitted, or included in the embodiments.

[0136] At least one computer-readable storage medium can include instructions that, when executed, cause a system to perform any of the computer-implemented methods described herein.

[0137] The phrases "one embodiment" or "an embodiment" and derivatives thereof may be used to describe some embodiments. These terms mean that the particular features, structures, or characteristics described in connection with the embodiments are included in at least one embodiment. The appearances of the phrase "in one embodiment" in various places in the specification are not necessarily all referring to the same embodiment. Additionally, unless otherwise indicated, the features described above are considered to be usable in any combination together. Thus, any features discussed separately can be employed in combination with each other, unless it is stated that the features are incompatible with each other.

[0138] The following examples relate to additional embodiments, from which numerous permutations and configurations will be apparent.

[0139] Example 1 includes a device that includes: an interface to a processor; and a circuit module that is configured to: associate a time with a Remote Direct Memory Access (RDMA) operation in a Translation Protection Table (TPT); and permit or restrict the RDMA operation based on the time in the TPT.

[0140] Example 2 includes the subject matter of Example 1, wherein the circuit module is configured to: associate a time with an RDMA key associated with the RDMA operation in the TPT; and permit or restrict the use of the RDMA key based on the time associated with the RDMA key.

[0141] Example 3 includes the subject matter of Example 1, wherein the RDMA operation is associated with an application to be executed on the processor, and the time is to be associated with a queue pair associated with the application.

[0142] Example 4 includes the subject matter of Example 1, wherein the device is to include one or more of the following: a Network Interface Controller (NIC), an Infrastructure Processing Unit (IPU), a Field Programmable Gate Array (FPGA), an accelerator device, a networking device, or a Data Processing Unit (DPU).

[0143] Example 5 includes the subject matter of Example 1, wherein the RDMA operation is to include an RDMA read operation of data from a remote device, and the circuit module is configured to: write the data to a cache of the processor based on the time.

[0144] Example 6 includes the subject matter of Example 1, wherein the time is to be based on an Institute of Electrical and Electronics Engineers (IEEE) 1588 time source.

[0145] Example 7 includes the subject matter of Example 1, wherein the circuit module is configured to: permit or restrict access to the TPT based on the amount of credit assigned to an application associated with the RDMA operation.

[0146] Example 8 includes a method that includes: associating time with a Remote Direct Memory Access (RDMA) operation through a circuit module and in a Translation Protection Table (TPT); and permitting or restricting the RDMA operation based on the time in the TPT by the circuit module.

[0147] Example 9 includes the subject matter as described in Example 8, further including: associating time with an RDMA key associated with the RDMA operation through the circuit module and in the TPT; and permitting or restricting the use of the RDMA key based on the time associated with the RDMA key by the circuit module.

[0148] Example 10 includes the subject matter as described in Example 8, where the RDMA operation is associated with an application to be executed on a processor coupled to the circuit module, and the time is to be associated with a queue pair associated with the application.

[0149] Example 11 includes the subject matter as described in Example 8, where the circuit module is to be included in one or more of the following devices: a Network Interface Controller (NIC), an Infrastructure Processing Unit (IPU), a Field Programmable Gate Array (FPGA), an accelerator device, a networking device, or a Data Processing Unit (DPU).

[0150] Example 12 includes the subject matter as described in Example 8, where the RDMA operation is to include an RDMA read operation of data from a remote device, and the method further includes: writing the data to a cache of a processor based on the time by the circuit module.

[0151] Example 13 includes the subject matter as described in Example 8, where the time is to be based on an Institute of Electrical and Electronics Engineers (IEEE) 1588 time source.

[0152] Example 14 includes the subject matter as described in Example 8, further including: permitting or restricting access to the TPT based on the number of credits assigned to an application associated with the RDMA operation by the circuit module.

[0153] Example 15 includes a non-transitory computer-readable storage medium that includes instructions that, when executed by a processor, cause the processor to: associate time with a Remote Direct Memory Access (RDMA) operation in a Translation Protection Table (TPT); and permit or restrict the RDMA operation based on the time in the TPT.

[0154] Example 16 includes the subject matter as described in Example 15, wherein the instructions further cause the processor to: associate time with an RDMA key associated with the RDMA operation in the TPT; and permit or restrict the use of the RDMA key based on the time associated with the RDMA key.

[0155] Example 17 includes the subject matter as described in Example 15, the RDMA operation being associated with an application, and the time to be associated with a queue pair associated with the application.

[0156] Example 18 includes the subject matter as described in Example 15, wherein the processor is to be included in one or more of the following devices: a network interface controller (NIC), an infrastructure processing unit (IPU), a field programmable gate array (FPGA), an accelerator device, a networking device, or a data processing unit (DPU).

[0157] Example 19 includes the subject matter as described in Example 15, the RDMA operation to include an RDMA read operation of data from a remote device, wherein the instructions further cause the processor to: write the data to a cache of the processor based on the time.

[0158] Example 20 includes the subject matter as described in Example 15, the time to be based on an Institute of Electrical and Electronics Engineers (IEEE) 1588 time source.

[0159] Example 21 includes the subject matter as described in Example 15, wherein the instructions further cause the processor to: permit or restrict access to the TPT based on the amount of credit assigned to an application associated with the RDMA operation.

[0160] Example 22 includes a device, the device including: means for associating time with a remote direct memory access (RDMA) operation; and means for permitting or restricting the RDMA operation based on the time.

[0161] Example 23 includes the subject matter as described in Example 22, further including: means for associating time with an RDMA key associated with the RDMA operation; and means for permitting or restricting the use of the RDMA key based on the time associated with the RDMA key.

[0162] Example 24 includes the subject matter as described in Example 22, the RDMA operation being associated with an application, and the time to be associated with a queue pair associated with the application.

[0163] Example 25 includes the subject matter as described in Example 22, wherein the device includes one or more of the following devices: a network interface controller (NIC), an infrastructure processing unit (IPU), a field programmable gate array (FPGA), an accelerator device, a networking device, or a data processing unit (DPU).

[0164] Example 26 includes the subject matter as described in Example 22, and the RDMA operation is to include an RDMA read operation of data from a remote device, and the device further includes: components for writing the data into a cache of a processor based on the time.

[0165] Example 27 includes the subject matter as described in Example 22, and the time is to be based on an Institute of Electrical and Electronics Engineers (IEEE) 1588 time source.

[0166] Example 28 includes the subject matter as described in Example 22, and further includes: components for permitting or restricting access to the time associated with the RDMA operation based on the number of credits assigned to an application associated with the RDMA operation.

[0167] It is emphasized that the abstract of the present disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. The abstract of the present disclosure is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. Further, in the foregoing detailed description, it can be seen that for the purposes of simplifying the present disclosure, various features are grouped together in a single embodiment. This method of disclosure should not be interpreted as reflecting an intention that the embodiments claimed require more features than are expressly recited in each claim. On the contrary, as the appended claims reflect, the subject matter of the present invention lies in less than all of the features of a single disclosed embodiment. Accordingly, the appended claims are hereby incorporated into the detailed description, where each claim stands on its own as a separate embodiment. In the appended claims, the terms "comprising" and "in" are used as the plain English equivalents of the respective terms "including" and "wherein". Further, the terms "first", "second", "third", etc. are used merely as labels and are not intended to impose numerical requirements on their objects.

[0168] For purposes of illustration and description, the foregoing description of the example embodiments has been presented. It is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. Many modifications and variations are possible in light of the disclosure. It is intended that the scope of the disclosure not be limited by this detailed description, but rather by the claims appended hereto. Future filed applications claiming priority to this application may claim the subject matter disclosed in different ways and may generally include any combination of one or more limitations differently disclosed herein or otherwise evidenced. The present disclosure thus provides the following technical solutions: Technical solution 1. A device, comprising: An interface to a processor; and A circuit module, the circuit module being configured to: Associate time with a Remote Direct Memory Access (RDMA) operation in a Translation Protection Table (TPT); and Permit or restrict the RDMA operation based on the time in the TPT. Technical solution 2. The device according to technical solution 1, wherein the circuit module is configured to: Associate time with an RDMA key associated with the RDMA operation in the TPT; and Permit or restrict the use of the RDMA key based on the time associated with the RDMA key. Technical solution 3. The device according to technical solution 1, wherein the RDMA operation is associated with an application to be executed on the processor, and the time is to be associated with a queue pair associated with the application. Technical solution 4. The device according to technical solution 1, wherein the device is to include one or more of the following devices: a Network Interface Controller (NIC), an Infrastructure Processing Unit (IPU), a Field Programmable Gate Array (FPGA), an accelerator device, a networking device, or a Data Processing Unit (DPU). Technical solution 5. The device according to technical solution 1, wherein the RDMA operation is to include an RDMA read operation of data from a remote device, and the circuit module is configured to: Write the data to a cache of the processor based on the time. Technical solution 6. The device according to technical solution 1, wherein the time is to be based on an Institute of Electrical and Electronics Engineers (IEEE) 1588 time source. Technical solution 7. The device according to technical solution 1, wherein the circuit module is configured to: Permit or restrict access to the TPT based on the number of credits assigned to an application associated with the RDMA operation. Technical solution 8. A method, comprising: Associating, by a circuit module and in a Translation Protection Table (TPT), time with a Remote Direct Memory Access (RDMA) operation; and Permitting or restricting, by the circuit module, the RDMA operation based on the time in the TPT. Technical solution 9. The method according to technical solution 8, further comprising: Associating, by the circuit module and in the TPT, time with an RDMA key associated with the RDMA operation; and The circuit module permits or restricts the use of the RDMA key based on the time associated with the RDMA key. Technical solution 10. The method according to technical solution 8, wherein the RDMA operation is associated with an application to be executed on a processor coupled to the circuit module, and the time is to be associated with a queue pair associated with the application. Technical solution 11. The method according to technical solution 8, wherein the circuit module is to be included in one or more of the following devices: a network interface controller (NIC), an infrastructure processing unit (IPU), a field programmable gate array (FPGA), an accelerator device, a networking device, or a data processing unit (DPU). Technical solution 12. The method according to technical solution 8, wherein the RDMA operation is to include an RDMA read operation of data from a remote device, and the method further includes: writing the data into a cache of the processor by the circuit module based on the time. Technical solution 13. The method according to technical solution 8, wherein the time is to be based on an Institute of Electrical and Electronics Engineers (IEEE) 1588 time source. Technical solution 14. The method according to technical solution 8, further includes: permitting or restricting access to the TPT by the circuit module based on the number of credits assigned to an application associated with the RDMA operation. Technical solution 15. A non-transitory computer-readable storage medium, the computer-readable storage medium includes instructions that, when executed by a processor, cause the processor to: associate a time with a remote direct memory access (RDMA) operation in a translation protection table (TPT); and permit or restrict the RDMA operation based on the time in the TPT. Technical solution 16. The computer-readable storage medium according to technical solution 15, wherein the instructions further cause the processor to: associate a time with an RDMA key associated with the RDMA operation in the TPT; and permit or restrict the use of the RDMA key based on the time associated with the RDMA key. Technical solution 17. The computer-readable storage medium according to technical solution 15, wherein the RDMA operation is associated with an application, and the time is to be associated with a queue pair associated with the application. Technical solution 18. The computer-readable storage medium as described in technical solution 15, wherein the processor is to be included in one or more of the following devices: network interface controller (NIC), infrastructure processing unit (IPU), field programmable gate array (FPGA), accelerator device, networking device, or data processing unit (DPU). Technical solution 19. The computer-readable storage medium as described in technical solution 15, wherein the RDMA operation is to include an RDMA read operation of data from a remote device, and wherein the instructions further cause the processor to: Write the data to a cache of the processor based on the time. Technical solution 20. The computer-readable storage medium as described in technical solution 15, wherein the time is to be based on an Institute of Electrical and Electronics Engineers (IEEE) 1588 time source.

Claims

1. A device comprising: Interface to the processor; as well as A circuit module, wherein the circuit module is used for: Associating a time with a remote direct memory access (RDMA) operation in a translation protection table (TPT); and The RDMA operation is permitted or restricted based on the time in the TPT.

2. The device according to claim 1, wherein the circuit module is used for: associating in the TPT a time with an RDMA key associated with the RDMA operation; and Use of the RDMA key is permitted or restricted based on the time associated with the RDMA key.

3. The apparatus of claim 1 or 2, wherein the RDMA operation is associated with an application to be executed on the processor, and the time is associated with a queue pair associated with the application.

4. The device according to any one of claims 1 to 3, wherein: The apparatus may include one or more of the following devices: a network interface controller (NIC), an infrastructure processing unit (IPU), a field programmable gate array (FPGA), an accelerator device, a networking device, or a data processing unit (DPU).

5. The device according to any one of claims 1 to 4, wherein the RDMA operation is to include an RDMA read operation of data from a remote device, and the circuit module is configured to: The data is written to a cache of the processor based on the time.

6. The apparatus of any one of claims 1 to 5, wherein the time is to be based on an Institute of Electrical and Electronics Engineers (IEEE) 1588 time source.

7. The device according to any one of claims 1 to 6, wherein the circuit module is used for: Access to the TPT is granted or restricted based on a number of credits allocated to an application associated with the RDMA operation.

8. A method comprising: Associating time with a remote direct memory access (RDMA) operation through a circuit module and in a translation protection table (TPT); as well as The RDMA operation is permitted or restricted, by the circuit module, based on the time in the TPT.

9. The method of claim 8, further comprising: Associating, by the circuit module and in the TPT, a time with an RDMA key associated with the RDMA operation; as well as Use of the RDMA key is permitted or restricted by the circuit module based on the time associated with the RDMA key.

10. The method of claim 8 or 9, wherein the RDMA operation is associated with an application to be executed on a processor coupled to the circuit module, and the time is to be associated with a queue pair associated with the application.

11. The method according to any one of claims 8 to 10, wherein: The circuit module is to be included in one or more of the following devices: a network interface controller (NIC), an infrastructure processing unit (IPU), a field programmable gate array (FPGA), an accelerator device, a networking device, or a data processing unit (DPU).

12. The method of any one of claims 8 to 11, wherein the RDMA operation is to include an RDMA read operation of data from a remote device, the method further comprising: The data is written into a cache of a processor based on the time by the circuit module.

13. The method of any one of claims 8 to 12, wherein the time is to be based on an Institute of Electrical and Electronics Engineers (IEEE) 1588 time source.

14. The method of any one of claims 8 to 13, further comprising: Access to the TPT is granted or restricted by the circuit module based on a number of credits allocated to an application associated with the RDMA operation.

15. A non-transitory computer-readable storage medium comprising instructions that, when executed by a processor, cause the processor to: Associating a time with a remote direct memory access (RDMA) operation in a translation protection table (TPT); and The RDMA operation is permitted or restricted based on the time in the TPT.

16. The computer-readable storage medium of claim 15, wherein: The instructions further cause the processor to: associating in the TPT a time with an RDMA key associated with the RDMA operation; and Use of the RDMA key is permitted or restricted based on the time associated with the RDMA key. 17 . The computer-readable storage medium of claim 15 , wherein the RDMA operation is associated with an application, and the time is to be associated with a queue pair associated with the application.

18. The computer-readable storage medium of any one of claims 15 to 17, wherein: The processor is to be included in one or more of the following devices: a network interface controller (NIC), an infrastructure processing unit (IPU), a field programmable gate array (FPGA), an accelerator device, a networking device, or a data processing unit (DPU).

19. The computer-readable storage medium of any one of claims 15 to 18, wherein the RDMA operation is to include an RDMA read operation of data from a remote device, wherein: The instructions further cause the processor to: The data is written to a cache of the processor based on the time.

20. The computer-readable storage medium of any one of claims 15 to 19, the time being based on an Institute of Electrical and Electronics Engineers (IEEE) 1588 time source.

21. The computer-readable storage medium of any one of claims 15 to 20, wherein: The instructions further cause the processor to: Access to the TPT is granted or restricted based on a number of credits allocated to an application associated with the RDMA operation.

22. A device comprising: A component for associating time with remote direct memory access (RDMA) operations; as well as Means for permitting or restricting the RDMA operation based on the time.

23. The apparatus of claim 22, further comprising: means for associating a time with an RDMA key associated with said RDMA operation; as well as Means for permitting or restricting use of the RDMA key based on the time associated with the RDMA key.

24. The apparatus of claim 22 or 23, wherein the RDMA operation is associated with an application, and the time is to be associated with a queue pair associated with the application.

25. The apparatus of any one of claims 22 to 24, wherein: The apparatus includes one or more of the following devices: a network interface controller (NIC), an infrastructure processing unit (IPU), a field programmable gate array (FPGA), an accelerator device, a networking device, or a data processing unit (DPU).