A smart network card and a distributed object access method based on the smart network card
By optimizing data access through read/write merging and adaptive prefetching modules of smart network interface cards (NICs), the problem of underutilization of smart NICs in distributed memory systems is solved, achieving low-latency and high-efficiency data processing.
Patent Information
- Application Number
- CN202510156343.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-02-12
AI Technical Summary
In existing technologies, the network processing capabilities of smart network interface cards (NICs) are not fully utilized, data copying overhead is high, and latency issues in distributed memory systems have not been effectively resolved, leading to a decline in system performance.
By using a smart network interface card as a proxy, the data access path is optimized through a read-write merging module and an adaptive prefetching module, reducing redundant operations and dynamically adjusting the prefetching strategy to improve system performance.
Significantly reduces network latency and redundant operations, improves overall system performance and resource utilization, and adapts to different loads and access patterns.
Smart Images

Figure CN120111107B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a smart network interface card (NIC) and a distributed object access method based on the smart NIC. Background Technology
[0002] In traditional large server architectures, computing and storage resources are tightly coupled, leading to low resource utilization, difficulty in elastically scaling CPUs (Central Processing Units) and memory, and poor fault tolerance. Distributed Memory technology decomposes the computing and storage resources of large servers into independent computing pools and memory pools. Compute nodes in the computing pool contain abundant CPU resources and a small amount of DRAM (Dynamic Random Access Memory) as an access cache; memory nodes in the memory pool contain abundant memory resources and relatively weak CPUs, which are used only for memory management and network communication. Since the weaker CPUs in memory nodes primarily utilize high-bandwidth, low-latency interconnect protocols such as CXL (Compute Fast Interconnect) or RDMA (Remote Direct Memory Access), compute nodes can directly read and write data on memory nodes, thus reducing the CPU load on memory nodes. Compute nodes can allocate and release memory blocks of arbitrary size using the alloc and free interfaces.
[0003] On the one hand, CXL (Compute Fast Interconnect) cache coherency protocol, based on PCIe 5.0 (Peripheral Component Interconnect Fast Line 5.0), supports direct memory access between devices using memory semantics, building a unified memory pool system. However, due to protocol limitations, it cannot be extended outside the rack. On the other hand, with the development of network technology, 400Gbps RDMA networks are gradually becoming widespread. CPU processing power is struggling to keep up with the rapid growth of network bandwidth. The emergence of smart NICs (Smart Network Interface Cards) offers a possible solution to this problem. They alleviate CPU pressure by offloading some CPU functions to the smart NIC, thus meeting the system's requirements for bandwidth and latency.
[0004] In distributed systems based on smart network interface cards (NICs), to alleviate CPU bottlenecks, the on-chip memory and computing resources of the smart NICs, including FPGAs (Field-Programmable Gate Arrays), ARM (Advanced RISC Machines), and ASICs (Application-Specific Integrated Circuits), can be utilized to offload processing-intensive tasks to the smart NICs and cache data in the smart NICs' on-chip memory to reduce data movement overhead. Currently, offloading primarily targets computational and management tasks. Computational task offloading mainly targets computationally intensive tasks such as compression, encryption, regular expression matching, and memory page deduplication. It leverages the pipelined and data parallel capabilities of FPGAs to copy data to the smart NICs, perform computation, and then copy it back to host memory. The drawbacks of computational task offloading are: data copying between the smart NIC and the host incurs overhead; data processing tasks may impact the network communication performance of the smart NICs; it fails to fully utilize the advantages of the smart NICs being located on the critical path of network processing; and it is typically optimized only for specific applications.
[0005] Specifically, Smart NICs have significant potential in the critical path of network data processing. They can not only offload some CPU workloads but also directly process network data streams to achieve low-latency network data forwarding and processing. However, current implementations may not fully leverage these advantages, primarily due to limitations in the software ecosystem, inadequate task partitioning, and integration challenges. Many existing software and protocol stacks are not optimized for the hardware characteristics of Smart NICs, resulting in underutilization of their network processing capabilities. Furthermore, failure to accurately identify the tasks best suited for processing on Smart NICs during task partitioning and offloading can prevent achieving optimal performance. Deeper integration of Smart NICs into existing system architectures may also present challenges, particularly in ensuring overall system stability and manageability. To fully utilize the advantages of Smart NICs, access mechanisms need to be optimized to support more efficient data paths, and applications that can fully leverage the processing capabilities of Smart NICs need to be developed.
[0006] Furthermore, optimizations for smart NICs are typically focused on specific application areas, such as network acceleration, encryption / decryption, and data compression. This application-specific optimization can introduce limitations, such as a lack of versatility; most optimization schemes focus only on specific computing patterns or operations, making them unsuitable for a wide range of applications. Developing dedicated optimization schemes for each application can also increase development costs and time, which is inflexible for rapidly iterating and changing business needs. Specific optimization schemes may also increase maintenance complexity, especially when cross-version or platform requirements are present. To overcome these limitations, it is worth considering developing more general optimization frameworks that can automatically adapt to the needs of different applications or dynamically optimize task allocation and resource usage through intelligent methods such as machine learning, thereby improving the applicability and performance of smart NICs in a wider range of scenarios.
[0007] Management task offloading primarily targets memory management, network communication, distributed file systems, and key-value (KV) operations, offloading the control and data planes of network services to the smart network interface card (NIC). The challenge of management task offloading is that network services are on the critical path of system performance, and the processing power of the smart NIC is relatively limited, potentially leading to a decline in overall system performance. Therefore, it is essential to fully leverage the high parallelism of the hardware and hide data access latency. In other words, it is necessary to fully utilize the hardware's ability to handle multiple tasks simultaneously and minimize data read latency.
[0008] CN115499157A discloses a method for accessing a Ceph distributed storage system. This method includes: receiving an access request from a client's smart network interface card (NIC) to access the Ceph distributed storage system, wherein the access request carries an identifier of a block storage device in the Ceph distributed storage system; determining the target block storage device based on the block storage device identifier carried in the access request; establishing a network connection between the Ceph distributed storage system and a gateway server to enable the client to access the target block storage device. This offloads librbd access transactions to the gateway server, allowing for sufficient CPU computing resources and significantly improving storage access performance. The smart NIC no longer directly processes librbd transactions, freeing up its CPU resources for other transactions, thereby improving the processing efficiency of other transactions. However, this technical solution, by offloading librbd access transactions to the gateway server, releases the smart NIC's resources. However, this does not utilize the smart NIC's inherent parallel processing capabilities. Smart NICs have the ability to accelerate network data stream processing and can achieve parallel processing of data packets through hardware acceleration. If these capabilities are not utilized, the hardware potential of the smart NIC is not fully exploited. Furthermore, this technical solution fails to address the latency issue caused by long data access paths. Data needs to be forwarded through a gateway server, adding an extra layer of intermediate processing, which may actually introduce new latency. The key to hiding data access latency lies in reducing the length of data request paths, optimizing data transmission speed, and employing techniques such as caching and prefetching in smart network interface cards to shorten waiting times. This technical solution does not offer significant optimizations in these areas.
[0009] As mentioned above, how to change the data processing method of the message buffer to reduce latency is a technical problem that has not yet been solved.
[0010] Furthermore, on the one hand, there are differences in understanding among those skilled in the art; on the other hand, the applicant studied a large number of documents and patents when making this invention, but due to space limitations, not all details and contents were listed in detail. However, this does not mean that the present invention does not possess the features of these prior art. On the contrary, the present invention already possesses all the features of the prior art, and the applicant reserves the right to add relevant prior art to the background art. Summary of the Invention
[0011] The current drawbacks of distributed memory systems are as follows: Accessing remote memory primarily relies on the CXL consensus protocol or the RDMA protocol. While the CXL protocol offers memory-semantic access and ultra-low latency, its PCIe architecture limits it to rack-based systems and cannot be scaled to larger networks. The RDMA network protocol supports cross-cluster expansion, enabling efficient data access and sharing, but suffers from high latency, exceeding 20 times the latency of the CXL memory-semantic access method. Furthermore, in CPU-centric systems, accelerator management tasks such as GPUs and XPUs, as well as network packet processing tasks, consume CPU resources, impacting the operation of other tasks.
[0012] To address the shortcomings of existing technologies, this invention provides a smart network interface card (NIC) and a distributed object access method based on the smart NIC. The purpose is to address the increasingly widespread use of discrete memory by using the smart NIC as a coordinator for different computing resources within a computing node, enabling unified access and prefetching of on-chip resources to reduce redundancy and improve system performance.
[0013] This invention provides a smart network interface card (NIC) from a first aspect. The smart NIC connects to and communicates with compute nodes and memory nodes. As a proxy, the smart NIC accesses remote memory on the memory nodes and returns responses to read / write requests to the corresponding devices on the compute nodes that initiated the requests. The smart NIC includes a read / write merging module and an adaptive prefetching module. This communication mechanism forms a separate memory pool centered on the smart NIC. This architecture utilizes the smart NIC as a proxy to centrally process read / write requests from compute nodes and return responses to the corresponding devices. Based on a message queue-based message passing mechanism, the smart NIC can efficiently manage communication between multiple devices, reducing the complexity of network requests. Furthermore, by introducing the read / write merging module and the adaptive prefetching module, the smart NIC not only optimizes data access paths but also significantly reduces redundant operations, thereby improving the overall system performance.
[0014] The read-write merging module merges read and write requests from compute nodes accessing the same destination address within a time window, reducing redundant read and write requests. This module fully leverages the characteristic of smart network interface cards (NICs) located on critical paths, avoiding duplicate network requests and thus reducing network load and latency. Compared to traditional methods involving simple task offloading or copy-based approaches, the read-write merging module further improves system efficiency and resource utilization through intelligent request analysis and merging strategies.
[0015] The adaptive prefetch module is used to check whether the on-chip memory of the smart network card is hit when the compute node initiates an access request to the remote memory node. If the memory is not hit, the entry in the read / write request is added to the pseudo prefetch queue in the message buffer. The prefetch module corresponding to the read / write request is scored according to the access cache hit rate. The access priority is adjusted based on the score of each prefetch module, thereby dynamically adjusting the prefetch strategy.
[0016] The adaptive prefetching module significantly improves the adaptability of smart network interface cards (NICs) to different loads by dynamically adjusting the prefetching strategy. Specifically, the adaptive prefetching module scores the prefetching module based on cache hit rate and selects the optimal prefetching strategy according to the score results. This approach not only effectively utilizes limited on-chip resources but also flexibly adjusts prefetching behavior according to actual access patterns, thereby reducing unnecessary prefetching operations. Simultaneously, through a negative feedback scoring mechanism, the smart NIC can quickly respond to changes in access patterns, further optimizing prefetching performance and improving overall performance.
[0017] According to a preferred embodiment, the read / write merging module performs the following steps: parsing the destination address of a message in the message queue, recording the access pattern, and merging accesses with the same destination address into a single access request within a time window; when the remote memory node access is completed, returning the response of the read / write request to the corresponding device of the compute node; caching the remote data of the memory node on the on-chip memory of the smart network card, accessing the data in the on-chip memory while the compute node device sends information to the smart network card, and returning directly if a local hit occurs, discarding the remote access message; and prefetching the read / write requests uniformly according to the recorded memory access pattern of the compute node device, inserting the remote data prefetched from the memory node into the on-chip memory of the smart network card.
[0018] The read-write merging module enhances its role in reducing redundant operations. By parsing the destination address in the message queue and recording access patterns, the read-write merging module can accurately identify and merge duplicate requests within a time window. This mechanism not only reduces the number of network requests but also lowers the frequency of remote memory access. Simultaneously, by caching remote data in the smart NIC's on-chip memory, the read-write merging module can directly return data upon local hit, significantly reducing access latency. Furthermore, the unified prefetching function allows the smart NIC to obtain potentially needed data in advance, further optimizing access efficiency.
[0019] According to a preferred embodiment, the steps of dynamically adjusting the prefetch strategy executed by the adaptive prefetch module include: checking whether read / write request information is hit in the pseudo prefetch queue based on the miss address information; adjusting the scores of different prefetch modules based on the hit results; selecting the optimal prefetch module to perform the prefetch operation based on the scores; updating the pseudo prefetch queue; inserting the address of the prefetch request to be issued by the prefetch module into the pseudo prefetch queue; and shutting down the prefetch operation when the values of each prefetch module are below the threshold.
[0020] The adaptive prefetch module's dynamic adjustment strategy achieves fine-grained control over prefetching behavior through a negative feedback scoring mechanism. Specifically, the adaptive prefetch module checks the hit status in the pseudo-prefetch queue based on miss address information and adjusts the scores of different prefetch modules accordingly. Modules with higher scores are prioritized for prefetching operations, while modules with lower scores are gradually eliminated or shut down. This mechanism ensures that prefetching operations always focus on the data most likely to be accessed, thereby maximizing on-chip memory utilization. Simultaneously, when the scores of all prefetch modules fall below a threshold, the smart NIC automatically disables prefetching operations to avoid unnecessary resource consumption.
[0021] According to a preferred embodiment, the step of dynamically adjusting the prefetch strategy performed by the adaptive prefetch module further includes: eviction of on-chip memory based on the LRU algorithm, wherein the priority information of the computing node's device is recorded in the cache entry, and the cache eviction priority of read and write requests is reduced for computing nodes with high latency requirements.
[0022] By employing an LRU-based cache eviction strategy, smart network interface cards (NICs) can effectively manage data in on-chip memory, ensuring that high-priority compute nodes always have sufficient cache resources. Specifically, the adaptive prefetch module records the priority information of compute node devices in cache entries and adjusts their cache eviction priorities according to the device's latency requirements. This mechanism enables smart NICs to better meet the needs of different devices, especially under high load conditions, ensuring that data access for critical tasks is not affected, thereby improving the overall stability and performance of the system.
[0023] According to a preferred embodiment, the adaptive prefetch module includes a metadata module, a prefetch data cache module, and an execution module. The metadata module stores a pseudo-prefetch queue, a prefetch score table, and / or historical information about different prefetch strategies; the prefetch data cache module temporarily stores data that is about to be accessed; the execution module checks for a hit in the pseudo-prefetch queue based on the address destination information of the read / write request; dynamically adjusts the prefetch module's score based on the hit results of the pseudo-prefetch queue; and selects the prefetch module with the highest score to perform the prefetch operation.
[0024] The composition and working principle of the adaptive prefetch module provide a solid foundation for the efficient operation of smart network interface cards (NICs). The metadata module provides comprehensive data support for prefetching decisions by storing a pseudo-prefetch queue, a prefetch score table, and historical information. The prefetch data cache module is responsible for temporarily storing data that will be accessed soon, ensuring timely data availability. The execution module dynamically selects the optimal prefetching strategy by precisely matching read / write request addresses and adjusting scores. This modular design not only enhances the flexibility and scalability of smart NICs but also ensures the accuracy and efficiency of prefetching operations.
[0025] According to a preferred embodiment, the steps of the execution module in the adaptive prefetch module to determine the prefetch module include: traversing the pseudo prefetch queue in the metadata module; comparing whether the access request address and the queue entry are equal; and determining a hit if the access request address is the same as one of the queue entries.
[0026] The specific implementation steps of the execution module clarify the selection logic of the prefetch module, ensuring the accuracy and efficiency of the prefetch operation. By traversing the pseudo-prefetch queue and comparing the access request addresses one by one, the execution module can quickly determine whether a match has occurred. This precise matching mechanism not only reduces the possibility of false positives but also improves the response speed of the smart network interface card (NIC). Furthermore, dynamically adjusting the prefetch module's score based on the match results allows the smart NIC to continuously optimize its prefetch strategy, thereby better adapting to different access patterns.
[0027] According to a preferred embodiment, the step of the execution module in the adaptive prefetching module dynamically calculating the score of the prefetching module based on the prefetching process includes: increasing the score of the prefetching module corresponding to the hit result when there is a hit; decreasing the score of the prefetching module corresponding to the hit result when there is a miss; wherein, the prefetching module score = attenuation coefficient * original prefetching module score + X; X = 1 when there is a hit, and X = -1 when there is a miss.
[0028] The prefetching module's scoring calculation method utilizes a negative feedback mechanism to achieve real-time evaluation and optimization of prefetching performance. Specifically, the adaptive prefetching module dynamically adjusts the score based on the hit result, increasing the score when a hit occurs and decreasing it when a hit occurs. This mechanism ensures that the smart NIC can quickly respond to changes in access patterns and gradually phase out ineffective prefetching strategies. Furthermore, by introducing a decay coefficient, the scoring calculation comprehensively considers historical performance and current results, thus providing more robust guidance for prefetching decisions.
[0029] The present invention provides a distributed object access method based on a smart network interface card (NIC) from a second aspect. The method includes: merging the access request counts of read and write requests from computing nodes to the same destination address within a time window to reduce redundant read and write request counts; while the computing node initiates an access request to a remote memory node, checking whether the on-chip memory of the smart NIC is hit; if not, adding the entries in the read and write request to a pseudo prefetch queue in the message buffer; scoring the prefetch module corresponding to the read and write request based on the access cache hit rate; and adjusting the access priority based on the scores of each prefetch module, thereby dynamically adjusting the prefetch strategy.
[0030] This invention constructs a lightweight multi-device communication mechanism centered on a smart network interface card (SNIC) through a message queue-based message passing mechanism. This mechanism significantly reduces request latency because it avoids network transmission bottlenecks and unnecessary duplicate requests that may exist in traditional methods. When a compute node initiates an access request to a remote memory node, the smart NIC first checks whether its on-chip memory is hit. If hit, it directly returns the data without sending another request to the remote memory node, thus greatly shortening the response time. Furthermore, by dynamically adjusting the prefetch strategy, the smart NIC can preload potentially needed data into the on-chip memory, further reducing latency caused by remote access.
[0031] According to a preferred embodiment, the step of merging the number of access requests for read / write requests in the method of the present invention includes: parsing the destination address of a message in the message queue, recording the access pattern, and merging accesses with the same destination address into one access request within a time window; when the remote memory node access is completed, returning the response of the read / write request to the corresponding device of the computing node; caching the remote data of the memory node on the on-chip memory of the smart network card, accessing the data in the on-chip memory while the device of the computing node sends information to the smart network card, and returning directly if a local hit occurs, discarding the remote access message; and uniformly prefetching the read / write requests according to the recorded memory access pattern of the device of the computing node, and inserting the remote data prefetched from the memory node into the on-chip memory of the smart network card.
[0032] By parsing the destination address in the message queue and recording the access pattern, the smart NIC can merge access requests with the same destination address into a single request within a time window. This merging mechanism effectively reduces the number of requests sent to remote memory nodes, thereby reducing network load and request latency. Once the remote memory node completes its access, the smart NIC immediately returns the read / write request response to the corresponding device on the compute node, ensuring rapid data interaction. Simultaneously, by caching remote data in the smart NIC's on-chip memory, the compute node can directly retrieve data upon local hit, without needing to initiate another remote access, further optimizing request latency. Furthermore, the unified prefetch function allows the smart NIC to preload potentially needed data, thereby reducing potential future latency.
[0033] According to a preferred embodiment, the step of dynamically adjusting the prefetch strategy in the method of the present invention includes: checking whether read / write request information is hit in the pseudo prefetch queue based on the miss address information; adjusting the scores of different prefetch modules based on the hit results; selecting the optimal prefetch module to perform the prefetch operation based on the scores; updating the pseudo prefetch queue; inserting the address of the prefetch request to be issued by the prefetch module into the pseudo prefetch queue; and shutting down the prefetch operation when each prefetch module is below a threshold.
[0034] Smart NICs can accurately determine which data is more likely to be accessed frequently by checking the hit status in the pseudo-prefetch queue based on miss address information. Based on the hit results, the scores of different prefetch modules are adjusted, and the optimal prefetch module is selected to perform the prefetch operation. This ensures that prefetching behavior always focuses on the data most likely to be needed, thereby minimizing potential latency for future remote access. Furthermore, by updating the pseudo-prefetch queue and inserting prefetch request addresses into the queue, smart NICs can continuously optimize the prefetch strategy to adapt to changing access patterns. When the scores of all prefetch modules fall below a threshold, the smart NIC automatically disables prefetching operations to avoid additional latency or resource consumption due to invalid prefetching. Attached Figure Description
[0035] Figure 1 This is a schematic diagram of the structure of the distributed object access system based on a smart network interface card provided by the present invention;
[0036] Figure 2 This is a structural diagram of the communication mechanism of the distributed object access system based on a smart network interface card provided by the present invention.
[0037] Figure 3 This is another structural diagram of the distributed object access system based on a smart network card provided by the present invention;
[0038] Figure 4 This is a logical diagram of the redundant read / write request merging strategy provided by the present invention;
[0039] Figure 5 This is a schematic diagram illustrating the principle of the adaptive prefetching operation of the smart network card provided by the present invention;
[0040] Figure 6 This is a flowchart illustrating the distributed object access method based on a smart network interface card provided by the present invention.
[0041] Figure 7 This is a flowchart illustrating the dynamic adjustment of the prefetching strategy executed by the adaptive prefetching module 205 provided by the present invention.
[0042] Figure 8 This is a flowchart illustrating the process of the execution module 208 determining the prefetch module provided by the present invention;
[0043] Figure 9 This is a flowchart illustrating an implementation scenario of the distributed object access method based on a smart network interface card provided by the present invention.
[0044] List of reference numerals
[0045] 100: Compute node; 200: Smart NIC; 201: CXL switch; 202: Stride-based prefetch module; 203: Temporal-based prefetch module; 204: Spatial-based prefetch module; 205: Adaptive prefetch module; 206: Metadata module; 207: Prefetch data cache module; 208: Execution module; 209: Read-write merging module; 300: Memory node. Detailed Implementation
[0046] The following is a detailed explanation with reference to the accompanying drawings.
[0047] Computing Node 100: A hardware unit responsible for executing specific computing tasks in a distributed computing environment. Typically, a computing node 100 includes one or more processors (CPUs), memory, storage devices, and network interfaces. The computing node 100 of this invention includes several processors or dedicated integrated chips such as CPUs, GPUs (graphics processing units), and XPUs (a general term for processing units other than CPUs, such as TPUs, FPGAs, etc.), which can efficiently handle complex computing tasks.
[0048] Smart NIC 200: A network card that performs the prefetching operation provided by this invention.
[0049] Memory node 300: This is a hardware unit specifically designed for storing data. In this invention, memory nodes 300 constitute part of a memory pool. Each memory node 300 has its own physical memory resources, which can be accessed by the compute node 100 through the smart network interface card 200.
[0050] Distributed Object Access (DOA) refers to a feature in a distributed computing environment that allows a computer program to invoke methods or services of another object located in a different address space (potentially on different computers) over a network. This approach enables remote objects to be used as if they were local objects, providing transparency; users or developers do not need to know whether the object is located locally or remotely. In this invention, resources are divided into a computing pool and a memory pool.
[0051] RNIC (RDMA Network Interface Card): A Remote Direct Memory Access network interface card, it is a network adapter specifically designed to support RDMA (Remote Direct Memory Access) technology. An RNIC can be installed on the smart network card 200, and its main function is to allow the smart network card 200 to directly access the memory of another computer without the need for the operating system or CPU of the remote computing node 100.
[0052] In this invention, multiple computing nodes 100 form a computing pool, and multiple memory nodes 300 form a memory pool. A smart network interface card 200 connects to and communicates with the computing nodes 100 and the memory nodes 300.
[0053] In this invention, various resources on compute node 100, including CPU, GPU, and XPU, share the same virtual address space. The smart network interface card (NIC) 200 accesses remote memory via the alloc, free, write, and read object access interfaces. "Remote memory" refers to memory resources not native to the current compute node 100. The smart NIC 200 acts as a proxy to access the remote memory of memory node 300. If a cache miss occurs in the smart NIC 200's local cache, the smart NIC 200, acting as a proxy, parses the access request and accesses memory node 300 via the RDMA protocol in a (address, size) manner. The response to the read / write request is returned to the corresponding device on the compute node 100 that initiated the read / write request.
[0054] RDMA communication process: The sending driver converts the request into a WQE and adds it to the sending queue. The smart network card 200 parses the WQE and copies the data to be sent from the memory of the compute node 100 to the internal buffer of the RDMA network card through DMA (Direct Memory Access) operation, and encapsulates it into a data packet according to the protocol for transmission. After receiving the data packet, the responding RDMA network card unpacks it according to the protocol to obtain the data and the destination virtual address, looks up the address translation table of the local compute node 100 to obtain the physical memory address, and finally writes the data to the destination address.
[0055] Example 1
[0056] like Figure 1 and Figure 2 As shown, the smart network interface card 200 and the computing node 100 are connected via the PCIe interface of the CXL switch 201, which performs CXL protocol conversion. The CXL switch 201 parses information such as the address and operation type in the request and sends the request to the message buffer of the smart network interface card 200 based on this information. The CXL high-speed interconnect protocol is built on PCIe 5.0 / CXL, and its logical layers are similar to the PCIe protocol, including the transaction layer, link layer, and physical layer. During data transmission, it includes mechanisms such as flow control, error detection, and correction to ensure reliable data transmission.
[0057] Preferably, such as Figure 3 As shown, the CXL switch 201 can also be integrated into the smart network card 200.
[0058] The smart network interface card 200 is equipped with a corresponding InfiniBand interface. The InfiniBand interface and the memory node 300 are connected via the InfiniBand interface, enabling remote communication between the smart network interface card 200 and the memory node 300, thereby achieving network communication between the memory node 100 and the memory node 300. Preferably, the remote communication can be RDMA communication.
[0059] like Figure 1 As shown, the Smart NIC 200 integrates an RDMA network interface controller (RNIC) and an ARM (Advanced RISC Machine) processor. The RNIC is responsible for data transmission and reception at the physical and data link layers of RDMA. It is the core component of the Smart NIC 200, ensuring that data packets can be transmitted over the network medium. The ARM processor is used for data processing. The RNIC and the ARM processor are connected via an internal bus.
[0060] The ARM processor communicates with compute node 100 via CXL switch 201. The ARM processor communicates with memory node 300 via Remote Direct Memory Access (RDMA) through a Remote Access RDMA network controller (RNIC).
[0061] Preferably, the smart network interface card 200 includes CXL memory. The CXL memory contains a circular buffer. The circular buffer message queue is used for communication with the daemon module. Specifically, devices such as the CPU and GPU in the compute node 100 use the CXL protocol to add requests to the circular buffer.
[0062] The specific communication process of the CXL protocol is as follows:
[0063] When the sender sends a memory access request to the CXL controller, the CXL hardware encapsulates the data and header information into a 68-bit FLIT packet according to the rules described in the CXL specification, and sends it to the receiver via the physical link. After receiving the data, the receiver decapsulates it according to the CXL protocol, extracts the useful data and control information, and performs the corresponding processing.
[0064] Preferably, such as Figure 2 As shown, the daemon module in the smart network interface card 200 manages access to the remote memory of the memory node 300 by the CPU or GPU in the local or compute node 100. The daemon module polls the message queue in the message buffer to obtain requests, which are then further processed by the adaptive prefetch module 205. After processing, the requests are sent to the RNIC and encapsulated into RDMA network packets to access the remote memory of the memory node 300.
[0065] In this invention, a communication mechanism centered on the smart network card 200 is constructed, which utilizes the advantage of the smart network card 200 being located on the critical path of the network to optimize the device communication process within the memory node 300 and accelerate remote object access.
[0066] like Figure 1 As shown, the ARM processor in the smart network card 200 of the present invention includes a read-write merging module 209 and an adaptive prefetch module 205.
[0067] The read / write merging module 209 is used to merge read / write requests from compute node 100 to the same destination address within a time window, thereby reducing redundant read / write requests.
[0068] The adaptive prefetch module 205 checks whether the on-chip memory of the smart network interface card 200 is hit while the compute node 100 initiates an access request to the remote memory node 300. In the case of a miss, the entry in the read / write request is added to the pseudo-prefetch queue in the message buffer, and the prefetch module corresponding to the read / write request is scored according to the access cache hit rate. The access priority is adjusted based on the scores of each prefetch module, thereby dynamically adjusting the prefetch strategy.
[0069] According to a preferred embodiment, such as Figure 6 As shown, the steps performed by the read / write merging module 209 include:
[0070] S110: Parse the destination address of messages in the message queue, record the access pattern, and merge accesses with the same destination address into one access request within the time window.
[0071] Preferably, the smart network interface card 200 extracts access requests from the message queue of the message buffer, parses the remote memory address of the memory node 300 in the access requests R from different computing nodes 100 (CPU, GPU, and XPU), such as... Figure 4 As shown.
[0072] Specifically, the read-write merging module 209 periodically polls its internal message buffer to check for new access requests. Once an unprocessed request is found, the read-write merging module 209 parses it to determine the specific requirements of the request, including the type of the computing node 100 that initiated the request (CPU, GPU, or XPU), the remote memory address, the request type (read or write), and the request size.
[0073] like Figure 4 As shown, within a preset time window, the read-write merging module 209 merges all access requests pointing to the same remote memory address 300 into a single access request, i.e., a common R. The read-write merging module 209 also divides related read-write requests with the same destination address into different time windows to avoid read-after-write and write-after-read issues. For read requests, this means that multiple threads request the same remote memory address within the same time period.
[0074] The read-write merging module 209 merges these requests into a single read operation, thereby reducing network communication and improving efficiency. For read-write requests, the message buffer accumulates these write operations, and then sends them all at once to the memory node 300 when a certain number has accumulated or the time window limit has been reached. This not only reduces network traffic but also improves performance through batch processing. During this period, the device thread that originally initiated the request (such as the CPU, GPU, or XPU thread) is added to the waiting queue, waiting for the smart network card 200 to complete the merged access operation and return the result.
[0075] S120: After the remote memory node 300 completes the access, the read-write merging module 209 returns the response to the read-write request to the corresponding device of the compute node 100.
[0076] After the smart network interface card 200 obtains the required data from the memory node 300, the read-write merging module 209 accurately distributes the data back to the corresponding compute node 100 device based on the information in the original request. This means that each device thread can receive the portion of data it needs. Simultaneously, the read-write merging module 209 sends a completion notification to the compute node 100 through the RX buffer in the message buffer, informing the relevant thread that the access operation has ended and subsequent tasks can continue. This notification mechanism ensures that the compute node 100 can promptly learn the operation status, thereby quickly resuming execution.
[0077] In addition, the read-write merging module 209 will also clean up related resources, such as releasing temporary space used to store merge requests, to ensure the effective use of system resources.
[0078] The read / write merging module 209 effectively reduces redundant requests and optimizes data transmission paths, thereby significantly reducing network latency. This strategy also enables the smart network interface card (NIC) 200 to process more requests within the same timeframe by strategically merging access requests to the same destination address, greatly improving the overall throughput of the NIC. Furthermore, the read / write merging module 209 centrally processes these merged requests and quickly returns results, allowing the NIC to respond more rapidly to the needs of various devices and significantly shorten waiting times.
[0079] As described above, when the read / write merging module 209 receives multiple access requests for the same remote memory address, it does not immediately execute each individual request. Instead, it merges them into a single, more efficient request for processing. This process not only reduces unnecessary network communication but also avoids performance bottlenecks caused by frequent small-scale operations. After the read / write merging module 209 completes the processing of the merged request, it immediately distributes the result back to the message buffers of each initiating device and notifies the relevant threads that the operation has been completed. This mechanism ensures that even under high concurrency, the smart NIC can maintain a high response speed and low latency, thereby improving overall performance. Therefore, the proxy remote access read / write merging strategy of the smart NIC 200 not only optimizes data transmission efficiency but also enhances the responsiveness and processing capabilities of the smart NIC. It is particularly suitable for application scenarios that require frequent small-scale read / write operations, such as distributed memory storage systems or high-performance computing environments.
[0080] After the read-write merging module 209 completes the merging of redundant requests, it adds the processed access request to the message buffer and simultaneously sends it to the adaptive prefetching module 205 for further processing. The message buffer is used to cache the merged read-write requests.
[0081] S130: Cache the remote data from memory node 300 on the on-chip memory of smart network interface card 200. Access the data in on-chip memory while the device in compute node 100 sends information to smart network interface card 200. If a local hit occurs, return directly and discard the remote access message.
[0082] S140: If a local memory access miss occurs, the read / write request will be uniformly prefetched by the adaptive prefetch module 205 according to the recorded memory access pattern of the compute node 100. The adaptive prefetch module 205 will insert the remote data prefetched from the memory node 300 into the on-chip memory of the smart network card 200.
[0083] Preferably, the message buffer of the present invention uses a hash table structure to maintain cache entries. Each hash slot contains information such as fingerprint, size, pointer to the actual location of the object, access frequency, and cache priority. The cache priority reflects the data access latency requirements of different computing nodes 100 and determines which data should be preferentially retained in the cache.
[0084] Preferably, the message buffer performs cache eviction of on-chip memory based on the LRU algorithm, wherein the priority information of the computing node 100 device is recorded in the cache entry, and the cache eviction priority of the computing node 100 device with high latency requirements is reduced.
[0085] When a cache entry needs to be replaced, lower priority entries should be given priority. In addition, dirty data (i.e., data that has been modified but has not yet been written back to the remote server) must be written back to the memory of memory node 300 before being evicted, while clean data can be discarded directly to simplify the cache management process.
[0086] The workloads of different computing resources (CPU, GPU, XPU) vary. Traditional methods using a single prefetch module fail to perform optimally across all workloads. For example, for irregular access patterns, methods based on access history and intra-block access pattern similarity perform better; while for sequential access patterns, step-size-based methods are more efficient. Methods based on access history and intra-block access pattern similarity require recording multiple accesses to capture new access patterns. Furthermore, a single prefetch strategy may pollute the prefetch cache by evicting useful prefetches and introducing inaccurate prefetches, thus limiting performance gains. By analyzing the access patterns of different computing resources in compute node 100, execution module 208 can dynamically select the optimal prefetch module based on changes in computing resource load and promptly shut it down when the prefetch hit rate is low.
[0087] Preferably, the execution module 208 prioritizes the processing of latency-sensitive computing resource demands, such as the GPU, to ensure these critical resources can achieve faster data response. Furthermore, by allocating larger cache spaces and higher prefetch priorities to these resources, the smart network interface card 200 not only improves their response speed but also further enhances the overall performance of the entire system.
[0088] According to a preferred embodiment, the steps of dynamically adjusting the prefetch strategy performed by the adaptive prefetch module 205 are as follows: Figure 7 As shown.
[0089] S210: Check whether the read / write request information is hit in the pseudo prefetch queue based on the miss address information, adjust the scores of different prefetch modules based on the hit results, and select the optimal prefetch module to perform the prefetch operation based on the scores.
[0090] S220: Update the pseudo prefetch queue.
[0091] S230: Insert the address of the prefetch request to be issued by the prefetch module into the pseudo prefetch queue.
[0092] S240: If the values of each prefetch module are below the threshold, disable the prefetch operation.
[0093] According to a preferred embodiment, the adaptive prefetch module 205 includes a metadata module 206, a prefetch data cache module 207, and an execution module 208. The execution module 208 uses the metadata information in the metadata module 206 to specifically perform the prefetch operation and caches the data in the prefetch data cache module 207.
[0094] Metadata module 206 is used to store pseudo prefetch queues, prefetch score tables, and / or historical information about different prefetch strategies.
[0095] Metadata module 206 records the pseudo-prefetch queues and prefetch score tables for each prefetch module. In addition, it stores historical information for different prefetch strategies, including the step size history table for stride-based prefetch module 202, the access flow history table for temporal-based prefetch module 203, and the pattern history table for spatial-based prefetch module 204. This information supports the prefetch decision-making process, ensuring the selection of the most appropriate prefetch strategy.
[0096] The prefetch data caching module 207 is used to temporarily store data that is about to be accessed. The execution module 208 is used to check whether there is a hit in the pseudo prefetch queue based on the address destination information of the read / write request; dynamically adjust the score of the prefetch module based on the hit result of the pseudo prefetch queue; and select the prefetch module with the highest score to perform the prefetch operation.
[0097] The workloads of different computing resources (CPU, GPU, XPU) vary. Traditional methods using a single prefetch module fail to perform optimally across all workloads. For example, for irregular access patterns, methods based on access history and intra-block access pattern similarity perform better; while for sequential access patterns, step-size-based methods are more efficient. Methods based on access history and intra-block access pattern similarity require recording multiple accesses to capture new access patterns. Furthermore, a single prefetch strategy may pollute the prefetch cache by evicting useful prefetches and introducing inaccurate prefetches, thus limiting performance gains. By analyzing the access patterns of different computing resources in compute node 100, execution module 208 can dynamically select the optimal prefetch module based on changes in computing resource load and promptly shut it down when the prefetch hit rate is low.
[0098] Preferably, the execution module 208 prioritizes the processing of latency-sensitive computing resource demands, such as the GPU, to ensure these critical resources can achieve faster data response. Furthermore, by allocating larger cache spaces and higher prefetch priorities to these resources, the smart network interface card 200 not only improves their response speed but also further enhances the overall performance of the entire system.
[0099] Preferably, such as Figure 3 As shown, execution module 208 includes a stride-based prefetch module 202, a temporal-based prefetch module 203, and a spatial-based prefetch module 204. Each prefetch module is responsible for predicting future memory accesses based on different access patterns and loading data from memory node 300 into the local cache in advance to reduce latency during actual access.
[0100] Stride-based prefetch module 202 performs data prefetching by analyzing the step size information in the access requests of memory node 300. Stride-based prefetch module 202 is suitable for situations where the program access pattern exhibits periodic characteristics. It can identify the fixed offset (i.e., stride) between consecutive memory accesses and predict and prefetch the memory pages that may be accessed subsequently based on this.
[0101] The temporal-based prefetch module 203 utilizes the principle of temporal locality to record access flow information in a history table and maps addresses to corresponding entries in the history table using an index table. When the first two prefetches of an address are both successful, the temporal-based prefetch module 203 considers this access to have a continuous temporal locality trend and performs further data prefetching based on the recorded address sequence flow.
[0102] Spatial-based prefetch module 204 is used for access similarity within the same memory region. It records the starting address of the memory region involved when accessing memory node 300 and its internal access pattern. Within the same memory region, if the access patterns of pages are similar, the corresponding access pattern is looked up in the pattern history table using the access address and the access offset within the region as the key, thereby realizing the prediction and data prefetching of future access patterns within that region.
[0103] In the above description, the index table is a data structure used to store the mapping relationship between addresses and their related information. It allows the smart network interface card (NIC) to quickly locate the record in the history table associated with a specific address, thus enabling rapid access to the historical access information of that address. The history table records access flow information, specifically the access history to memory node 300. In the Temporal-based prefetch module 203, the history table records the temporal locality of access information, i.e., which data is frequently accessed at similar times. Each entry in the history table may contain information such as the access timestamp, the accessed address, and the access frequency. This information is used to identify access patterns, thereby predicting future access trends.
[0104] The pattern history table focuses on recording and identifying spatial locality, i.e., the frequent grouping of data items that are close in the physical address space. In the Spatial-based prefetch module 204, the pattern history table records access patterns within memory regions, including the starting address and internal access patterns. Each entry in the pattern history table may contain the starting address of a memory region, the access offset, and the access pattern within that region. This information is used to predict future access patterns within the same memory region and to perform data prefetching accordingly.
[0105] According to a preferred embodiment, such as Figure 8 As shown, the steps by which execution module 208 determines the prefetch module include:
[0106] S310: Traverse the pseudo prefetch queue in metadata module 206.
[0107] S320: Compare the access request address with the queue entry to see if they are equal; if the access request address is the same as one of the queue entries, it is considered a hit.
[0108] If a cache hit occurs, execution module 208 reads data directly from the on-chip memory of its smart network interface card 200 and returns the requested object data to compute node 100 via the RX buffer in the message buffer. This approach significantly reduces round-trip time (RTT) for remote access and improves response speed. In this case, execution module 208 effectively accelerates the access process as part of a fast response mechanism.
[0109] If a cache miss occurs, execution module 208 needs to initiate a data read request to the remote memory node 300. In this case, execution module 208 acts as a communication medium, selecting a prefetch module and interacting with memory node 300 using message passing. Specifically, execution module 208 retrieves the request from the message buffer and then transmits it to memory node 300 via the RDMA network.
[0110] Specifically, execution module 208 checks for a hit in the pseudo-prefetch queue based on the address information of the access request. Each pseudo-prefetch queue is a FIFO queue containing 16 queue entries, and each queue entry stores a 36-bit pseudo-prefetch address. Execution module 208 iterates through the queue entries, comparing the current access request address with the queue entry to confirm whether a hit has occurred. This process helps reduce unnecessary data transmission and lower host latency.
[0111] According to a preferred embodiment, the step of the execution module 208 dynamically calculating the score of the prefetch module based on the prefetch process includes: increasing the score of the prefetch module corresponding to the hit result when there is a hit; and decreasing the score of the prefetch module corresponding to the hit result when there is a miss.
[0112] Specifically, based on the hit results of the pseudo-prefetch queue, execution module 208 modifies the score counters of the three prefetch modules (C0, C1, C2). If a hit occurs, the score of the corresponding prefetch module is increased accordingly; otherwise, its score is decreased.
[0113] To prevent pre-fetching modules with historically high scores from maintaining a dominant position for an extended period, this invention introduces a decay coefficient. The scoring formula is as follows:
[0114] Prefetch module score = attenuation coefficient * original prefetch module score + X; X = 1 in case of a hit, and X = -1 in case of a miss.
[0115] This dynamic adjustment mechanism allows the Smart NIC 200 to continuously optimize its prefetching strategy, reducing invalid prefetching and network traffic.
[0116] Preferably, the prefetch module with the highest score is selected to perform the prefetch operation.
[0117] The execution module 208 selects the prefetching module with the highest score as the optimal prefetching module to perform the data prefetching operation. The prefetching module triggers the prefetching operation if at least one prefetching module's pseudo-prefetching queue is hit twice consecutively or if the prefetching module's score is greater than a threshold (e.g., 0); otherwise, the execution module 208 does not perform the prefetching operation.
[0118] In other words, prefetching will not be performed only when the scores of all three prefetching modules are below the preset threshold of 0. However, if the pseudo-prefetch queue of a prefetching module hits twice in a row or its score exceeds the threshold of 0, the prefetching module will resume prefetching to achieve continuous and accurate prefetching.
[0119] The process of updating the pseudo-prefetch queue is described below.
[0120] Each prefetch module generates a new prefetch address based on the access request address and records it in the corresponding access history table. For example, the stride-based prefetch module 202 updates the step size history table, calculates the difference between the current address and the last access address as the new step size, and the prefetch address is the current address plus this step size. The temporal-based prefetch module 203 records the addresses in the current stream using an LRU method and prefetches the next address in the sequential stream. The spatial-based prefetch module 204 predicts and records the next address based on the relative address within the block. These newly generated prefetch addresses are placed in a pseudo-prefetch queue for subsequent use in determining the effectiveness of each prefetch module.
[0121] Preferably, the data cache in the prefetch data cache module 207 is divided into two levels. Cache entries contain device priority information. According to a preset cache eviction policy, lower-priority cache entries are evicted first. For latency-sensitive devices such as GPUs, their cached data is preferentially placed in the Level 1 cache, while other data is placed in the Level 2 cache. During eviction, data is preferentially removed from the Level 2 cache, and data at the end of the Level 1 cache is periodically moved to the Level 2 cache to avoid low-frequency data residing in the Level 1 cache for extended periods. This method ensures that frequently accessed data remains in the cache, thereby reducing overall access latency and meeting the needs of low-latency devices.
[0122] Example 2
[0123] This embodiment is a further example of Embodiment 1, and repeated content will not be repeated.
[0124] This invention provides an implementation scenario. Considering the need for a GPU computing power cloud service provider to offer computing power leasing services, and to meet the GPU's higher memory requirements, this invention leverages the characteristic of the smart network interface card 200 being located on the network critical path to achieve efficient GPU memory expansion through RDMA proxy access. Simultaneously, read-write merging and adaptive prefetching techniques are used to accelerate access to remote memory objects, reducing memory access latency and improving system performance. Specific steps are as follows... Figure 5 and Figure 9 As shown.
[0125] S410: Construct an efficient communication mechanism.
[0126] In compute node 100, multiple GPUs are connected to smart NIC 200 via CXL switch 201, and a lightweight communication model between them and smart NIC 200 is established using message passing. Smart NIC 200 is responsible for proxying requests from GPUs to access object data in memory node 300 via the RDMA protocol.
[0127] S420: Performs read-write merging operations based on time windows.
[0128] The read / write merging module 209 in the smart network interface card 200 divides the time into multiple non-overlapping 6-millisecond (ms) time windows and merges redundant read / write requests based on these windows. For example, if GPU A sends a read access request A to the smart network interface card 200 containing the address, size, request type, and device type of the accessed object, and GPU B also sends a read access request B for the same destination address within the same time window, then the read / write merging module 209 will merge these two requests into a single access request C and add it to the message buffer of the smart network interface card 200.
[0129] S430: Perform adaptive prefetch operation.
[0130] Access request C will be forwarded to the adaptive prefetch module 205. For example... Figure 5 As shown, the pseudo-prefetch queues include queues P0, P1, and P2. Execution module 208 checks whether the access address of request C hits one of the pseudo-prefetch queues corresponding to the three prefetch modules. Based on the hit result, the score of each prefetch module is adjusted using the following formula: Prefetch module score = attenuation coefficient 0.8 * original prefetch module score + X (X = +1 for a hit, X = -1 for a miss).
[0131] Next, the Stride-based prefetch module 202, Temporal-based prefetch module 203, and Spatial-based prefetch module 204 generate new prefetch addresses based on the access addresses; the Stride-based prefetch module 202 calculates the current address plus the step size; the Temporal-based prefetch module 203 searches for the next address in the sequential flow; and the Spatial-based prefetch module predicts the next address 204 based on the pattern history table entries.
[0132] Then, as Figure 5 As shown, the execution module 208 adds these newly generated prefetch addresses to the pseudo prefetch queue, and selects the prefetch module P0 with the highest score (assuming its score is 3.1) as the optimal prefetch module, encapsulates its prefetch address into a prefetch request D and adds it to the message buffer of the smart network card 200 to perform the actual prefetch operation.
[0133] S440: Agent accesses remote memory of memory node 300.
[0134] The adaptive prefetch module 205 of the smart network interface card 200 includes a proxy access module for handling read access requests C and prefetch requests D obtained from the message buffer.
[0135] First, the adaptive prefetch module 205 checks the local prefetch data cache module 207 on the smart network interface card 200. If the required data is not found, the request is encapsulated into a network packet and sent to the memory node 300 via the RDMA network interface card (RNIC). Once the data corresponding to the access request is returned, the adaptive prefetch module 205 notifies GPU A and GPU B that the data reading is complete and caches the returned data in the local prefetch data cache module 207 for fast access later.
[0136] It should be noted that the specific embodiments described above are exemplary. Those skilled in the art can devise various solutions inspired by the disclosure of this invention, and these solutions all fall within the scope of this invention and its protection. Those skilled in the art should understand that this specification and its accompanying drawings are illustrative and not intended to limit the scope of the claims. The scope of protection of this invention is defined by the claims and their equivalents. This specification contains multiple inventive concepts; phrases such as "preferredly" or "according to a preferred embodiment" indicate that the corresponding paragraph discloses an independent concept. The applicant reserves the right to file divisional applications based on each inventive concept.
Claims
1. A smart network interface card (NIC), characterized in that, The smart network card (200) is connected to and communicates with the computing node (100) and the memory node (300). The smart network card (200) acts as a proxy to access the remote memory of the memory node (300) and returns the response of the read / write request to the corresponding device on the computing node (100) that initiated the read / write request. The smart network interface card (200) includes: The read-write merging module (209) merges the access request counts of read and write requests from the computing node (100) to the same destination address within a time window, so as to reduce the number of redundant read and write requests. The adaptive prefetch module (205) checks whether the on-chip memory of the smart network card (200) is hit while the computing node (100) initiates an access request to the remote memory node (300). In the event of a cache miss, the entries in the read / write request are added to the pseudo-prefetch queue in the message buffer, and the prefetch modules corresponding to the read / write request are scored based on the access cache hit rate. The access priority is adjusted based on the scores of each prefetch module, thereby dynamically adjusting the prefetch strategy.
2. The smart network interface card according to claim 1, characterized in that, The steps performed by the read-write merging module (209) include: Parse the destination address of the message in the message queue, record the access pattern, and merge accesses with the same destination address into one access request within the time window; when the access is completed at the remote memory node (300), return the response of the read / write request to the corresponding device of the computing node (100); The remote data of the memory node (300) is cached on the on-chip memory of the smart network card (200). While the device of the computing node (100) sends information to the smart network card (200), it accesses the data in the on-chip memory. If a local hit is found, it returns directly and discards the remote access message. Based on the recorded memory access patterns of the computing node (100), read and write requests are prefetched uniformly, and remote data prefetched from the memory node (300) is inserted into the on-chip memory of the smart network card (200).
3. The smart network interface card according to claim 1 or 2, characterized in that, The steps of dynamically adjusting the prefetch strategy performed by the adaptive prefetch module (205) include: Based on the missed address information, check whether the read / write request information has been hit in the pseudo prefetch queue. The scores of different prefetching modules are adjusted based on the hit results, and the optimal prefetching module is selected based on the scores to perform the prefetching operation; Update the pseudo-prefetch queue; Insert the address of the prefetch request to be issued by the prefetch module into the pseudo prefetch queue; If the values of each of the prefetch modules are below the threshold, the prefetch operation is turned off.
4. The smart network interface card according to any one of claims 1 to 3, characterized in that, The steps of dynamically adjusting the prefetching strategy performed by the adaptive prefetching module (205) further include: On-chip memory is cached and evicted based on the LRU algorithm, wherein the priority information of the computing node (100) is recorded in the cache entry. For devices of the compute node (100) with high latency requirements, the cache eviction priority of the read and write requests is reduced.
5. The smart network interface card according to any one of claims 1 to 4, characterized in that, The adaptive prefetch module (205) includes: Metadata module (206) is used to store historical information about pseudo prefetch queues, prefetch score tables and / or different prefetch strategies; A prefetch data cache module (207) is used to temporarily store data that is about to be accessed; The execution module (208) checks whether there is a hit in the pseudo prefetch queue based on the address destination information of the read / write request; dynamically adjusts the score of the prefetch module based on the hit result of the pseudo prefetch queue; and selects the prefetch module with the highest score to perform the prefetch operation.
6. The smart network interface card according to any one of claims 1 to 5, characterized in that, The steps by which the execution module (208) in the adaptive prefetch module (205) determines the prefetch module include: Traverse the pseudo prefetch queue in the metadata module (206); Compare the access request address with the queue entry to see if they are equal; If the access request address is the same as one of the entries in the queue, it is considered a hit.
7. The smart network interface card according to any one of claims 1 to 6, characterized in that, The steps of the execution module (208) in the adaptive prefetch module (205) dynamically calculating the score of the prefetch module based on the prefetch process include: Increase the score of the prefetch module corresponding to the hit result if a hit is achieved; In the event of a miss, reduce the score of the prefetch module corresponding to the hit result; Wherein, the prefetch module score = attenuation coefficient * original prefetch module score + X; If the hit occurs, X = 1; if the hit does not occur, X = -1.
8. A distributed object access method based on a smart network interface card, characterized in that, The method includes: Within the time window, read and write requests from compute nodes (100) that access the same destination address are merged to reduce redundant read and write requests. While the computing node (100) initiates an access request to the remote memory node (300), it checks whether the on-chip memory of the smart network card (200) is hit. In the event of a cache miss, the entries in the read / write request are added to the pseudo-prefetch queue in the message buffer, and the prefetch modules corresponding to the read / write request are scored based on the access cache hit rate. The access priority is adjusted based on the scores of each prefetch module, thereby dynamically adjusting the prefetch strategy.
9. The method according to claim 8, characterized in that, The steps for merging read and write requests into access request counts include: Parse the destination address of the message in the message queue, record the access pattern, and merge accesses with the same destination address into one access request within the time window; when the access is completed at the remote memory node (300), return the response of the read / write request to the corresponding device of the computing node (100); The remote data of the memory node (300) is cached on the on-chip memory of the smart network card (200). While the device of the computing node (100) sends information to the smart network card (200), it accesses the data in the on-chip memory. If a local hit is found, it returns directly and discards the remote access message. Based on the recorded memory access patterns of the computing node (100), read and write requests are prefetched uniformly, and remote data prefetched from the memory node (300) is inserted into the on-chip memory of the smart network card (200).
10. The method according to claim 8 or 9, characterized in that, The steps for dynamically adjusting the prefetching strategy include: Based on the missed address information, check whether the read / write request information has been hit in the pseudo prefetch queue. The scores of different prefetching modules are adjusted based on the hit results, and the optimal prefetching module is selected based on the scores to perform the prefetching operation; Update the pseudo-prefetch queue; Insert the address of the prefetch request to be issued by the prefetch module into the pseudo prefetch queue; If the values of each of the prefetch modules are below the threshold, the prefetch operation is turned off.
Citation Information
Patent Citations
Ceph distributed storage system access method and device
CN115499157A
Metadata processing method in storage equipment and related equipment
CN113805789A
Data pre-reading method and device, electronic equipment and storage medium
CN116795736A