Hybrid hardware-software consistency framework
By caching data in the accelerator device and utilizing a hardware-software consistency framework, the problem of inefficient data interaction between the accelerator and the host computing system in the traditional I/O model is solved, efficient data access and processing is achieved, and the computing performance of the accelerator device is improved.
Patent Information
- Application Number
- CN202080039641.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-29
- Filing Date
- 2020-05-08
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2040-05-08
AI Technical Summary
In the traditional I/O model, data exchange between the accelerator and the host computing system is inefficient, especially when accessing large block memories, where there are problems with latency, bandwidth, and protocol message transmission overhead. In addition, the accelerator device cannot cache memory and cannot fully utilize the performance advantages of the hardware-coherent shared memory multiprocessor paradigm.
A collaborative hardware-software consistency framework is adopted to transfer the ownership of the data set from the host to the accelerator device, cache the data in the accelerator device, use hardware consistency technology to maintain data consistency between the host and accelerator device, and manage the transmission and processing of large block data through software consistency methods.
It reduces data transmission overhead, improves the computing efficiency of the accelerator device, reduces the data transmission delay between the host and the accelerator device, achieves efficient data access and processing, and maintains the performance advantages of the hardware consistent shared memory multi-processor paradigm.
Smart Images

Figure CN113924557B_ABST
Abstract
Description
Technical Field
[0001] Examples of the present disclosure generally relate to accelerators that use a cooperative hardware-software coherence framework to transfer data between a host and an accelerator device. Background Art
[0002] In the traditional I / O model, a host computing system interfaces with peripheral I / O devices (e.g., accelerators) while using custom I / O device drivers specific to the peripheral I / O devices to perform accelerator tasks or functions. A disadvantage of this model is that all accelerator-memory interactions must pass through a home agent node in the main CPU-memory complex (sometimes referred to as the server host). This model uses hardware to enforce consistency, so regardless of whether the accelerator-memory data footprint is orders of magnitude larger than a cache line (MB, GB, or even TB), all accelerator-memory interactions are serviced and tracked by the home agent node, even though the processor (generator) on the server host may ultimately only access a small portion of the data footprint for which the accelerator has performed computations. Another disadvantage is that all memory accessed and cached by the accelerator is remote, server-host attached memory. Yet another disadvantage of this model is that PCIe I / O-based accelerators cannot cache memory. Accelerator-memory interactions occur only using server-host attached memory, and this approach suffers from latency, bandwidth, and protocol message transmission overhead disadvantages.
[0003] At the same time, the hardware cache coherent shared memory multiprocessor paradigm utilizes a common instruction set architecture (ISA)-independent model of interfacing in executing tasks or functions on a multiprocessor CPU. The common ISA-independent (e.g., C code) model of interfacing scales with both the number of processing units and the amount of shared memory available to those processing units. Traditionally, peripheral I / O devices have not been able to benefit from the coherence paradigm used by the CPU executing on the host computing system. Summary of the Invention
[0004] Techniques for transferring ownership of a dataset to an accelerator device are described. One example is a computing system comprising a host and an accelerator device, the host comprising a processing unit, a home agent (HA), and an accelerator application, the accelerator device being communicatively coupled to the host, wherein the accelerator device comprises: a request agent (RA) configured to perform at least one accelerator function; a slave agent (SA); and a local memory, wherein the local memory is part of the same coherence domain as the processing unit in the host. The accelerator application is configured to transfer ownership of the dataset from the HA to the accelerator device, and the accelerator device is configured to: store a latest copy of the dataset in the local memory, use the SA to service memory requests from the RA, and transfer ownership of the dataset back to the HA.
[0005] In some embodiments, the accelerator application is configured to identify a data set, where the data set is a sub-portion of a memory block or a memory page.
[0006] In some embodiments, after ownership of a data set is transferred from the HA to the accelerator device, the remainder of the memory block or memory page continues to belong to the HA.
[0007] In some embodiments, servicing a memory request from the RA using the SA is performed without receiving permission from the HA.
[0008] In some embodiments, the processing unit includes a cache, wherein the cache stores the modified version of the data set.
[0009] In some embodiments, the accelerator device is configured to transmit a flush command to the HA after ownership of the data set has been transferred to the accelerator device, the flush command instructing the processing unit to flush the cache so that the modified version of the data set is moved to local memory.
[0010] In some embodiments, the processing unit is configured to, in response to the flush command, one of invalidate the modified version of the data set in the cache or maintain a cache-clean copy of the modified version of the data set for future low-latency re-reference.
[0011] In some embodiments, the accelerator application is configured to transmit a flush command to the HA before ownership of the data set has been transferred to the accelerator device, the flush command instructing the processing unit to flush the cache so that the modified version of the data set is moved to local memory.
[0012] In some embodiments, the accelerator application is configured to transmit an invalidate command to the HA before ownership of the data set has been transferred to the accelerator device, the invalidate command instructing the processing unit to invalidate the cache.
[0013] An example described herein is an accelerator device comprising an RA, an SA, and a memory. The RA includes a compute engine configured to perform at least one accelerator function, and the SA includes a memory controller. The RA is configured to receive ownership of a dataset from the HA in a host coupled to the accelerator device to the SA, where ownership is transferred via a software application. The memory controller is configured to service a request transmitted by the compute engine to access the dataset stored in the memory once ownership has been transferred to the SA. The RA is configured to transfer ownership from the SA back to the HA.
[0014] In some embodiments, the data set is a sub-portion of a memory block or memory page, wherein after ownership of the data set is transferred from the HA to the SA, the remainder of the memory block or memory page continues to belong to the HA, wherein the service request is executed without receiving permission from the HA.
[0015] In some embodiments, the accelerator device is configured to communicate with the host using a coherent interconnect protocol to extend the host's coherency domain to include the memory and SA in the accelerator device.
[0016] One example described herein is a method that includes: transferring ownership of a data set from an HA in a host to an accelerator device, wherein the accelerator device is communicatively coupled to the host and a coherence domain of the host is extended into the accelerator device; moving the data set to local memory in the accelerator device; servicing a memory request from a RA in the accelerator device using an SA in the accelerator device; and transferring ownership from the accelerator device back to the HA.
[0017] In some embodiments, the method includes: identifying, using an accelerator application executing on a host, a data set, wherein the data set is a sub-portion of a memory block or a memory page, wherein after transferring ownership of the data set from the HA to the accelerator device, a remainder of the memory block or memory page continues to belong to the HA.
[0018] In some embodiments, the method includes invalidating the modified data set in the cache in response to a cache maintenance operation (CMO) issued by the accelerator device after transferring ownership to the accelerator device. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order that the features described above may be understood in detail, a more particular description, briefly summarized above, may be made by reference to example implementations, some of which are illustrated in the accompanying drawings. However, it should be noted that the drawings illustrate only typical example implementations and are therefore not to be considered as limiting the scope thereof.
[0020] Figure 1 is a block diagram of a host coupled to an accelerator device according to one example.
[0021] Figure 2 is a block diagram of a host coupled to an accelerator device according to one example.
[0022] Figure 3 is a flow diagram for using a cooperative hardware-software coherence framework according to one example.
[0023] Figure 4 is a flow diagram for using a cooperative hardware-software coherence framework according to one example.
[0024] Figure 5 A field programmable gate array implementation of a programmable IC according to one example is illustrated. DETAILED DESCRIPTION
[0025] Various features are described below with reference to the accompanying drawings. It should be noted that the drawings may or may not be drawn to scale, and that elements of similar structure or function are represented by the same reference numerals throughout the drawings. It should be noted that the drawings are intended only to facilitate the description of the features. They are not intended to be an exhaustive description of the specification or to limit the scope of the claims. In addition, the illustrated examples do not necessarily have all the aspects or advantages shown. Aspects or advantages described in conjunction with a particular example are not necessarily limited to that example and can be practiced in any other example, even if not illustrated or explicitly described as such.
[0026] The examples herein describe an accelerator device (e.g., a peripheral I / O device or accelerator-attached memory) that shares the same coherency domain as hardware elements in a host computing device. As a result, computing resources in the coherency domain of the accelerator device can communicate with the host in a manner similar to CPU-to-CPU communication in the host. This means that computing resources can take advantage of coherency-type features such as direct communication (no address translation), more efficient memory usage, non-uniform memory access (NUMA) awareness, and the like. Coherent memory relies on fine-grained memory management, where the home agent node manages memory at the cache line level.
[0027] However, accelerator tasks typically handle large block memories, which can be better managed using software consistency rather than hardware consistency, where the home agent in the host manages the data handled by the accelerator device. For example, the disadvantage of the hardware consistency model is as follows: all accelerator-memory interactions still pass through the server host first. Regardless of whether the accelerator-memory data footprint is orders of magnitude larger than a cache line (MB, GB or even TB), and regardless of whether the memory is accelerator-attached memory, this model uses hardware to enforce consistency, and even if the processor (generator) may ultimately only access a small portion of the data footprint that the accelerator has performed calculations on, all accelerator-memory interactions are still first serviced and tracked by the host. In contrast, software consistency is a situation where the accelerator device manages the data stored in its memory, similar to the traditional host-I / O model, where the host transfers large blocks of data to the I / O device, the device processes the data, and the I / O device notifies the host when the processed data is ready. At the same time, the software consistency model guarantees that the host will not access the data until the data processed by the I / O accelerator is ready.
[0028] Embodiments herein describe a hybrid of hardware and software coherence to create a hybrid hardware-software framework that reduces the overhead of managing data when large blocks of data are moved from the host to the accelerator device when the host and accelerator are in the same coherence domain. In one embodiment, an accelerator application executing on the host identifies the dataset it wishes to transfer to the accelerator device for processing. The accelerator application notifies the accelerator device via software coherence methods by modifying metadata about the data or by modifying control data structures accessed by the accelerator device. In response, the accelerator device transmits a request to the host using hardware coherence techniques, requesting the most recent version of the dataset and requesting that no cached copy of the dataset be retained in the host. The host then uses hardware coherence techniques to ensure that the most recent version of the data is stored in the memory of the accelerator device and invalidates all copies of the dataset in the CPU cache. A request agent in the accelerator device processes the dataset in local memory based on the acceleration application via the slave agent. In one embodiment, the request agent may cache a subset of the local memory and process the cached copy. Software consistency already ensures that any memory operation request received from a request agent in the accelerator device can access the data set in the local memory via the slave agent without the slave agent obtaining permission from the host. For example, the slave agent does not need to use the home agent to check whether there is a more recent version of the data set, or to check whether the cache copy retains the data to be processed by the accelerator device. When the accelerator device completes processing of the data set, the accelerator device clears the cache subset, updates the local memory, and transfers ownership back to the home agent in the host via software consistency methods, for example by modifying the metadata or control data structures accessed by the host processor by the accelerator device. This re-establishes the host as the enforcer of hardware consistency for the data set with the rest of the system, which is similar to the situation when completing tasks in the traditional host-I / O model, where the accelerator device is in a different domain from the host. Therefore, by being in the same consistency domain and relying on hardware consistency enforcement, the accelerator device and the host avoid the high overhead that occurs when transferring ownership of large blocks of data to the local memory in the accelerator device via software methods. By being in the same consistency domain, the accelerator device and the host can also avoid the high overhead of moving processed data from the local memory in the accelerator device back to the host.
[0029] Figure 11 is a block diagram of a computing system 100 including a host 105 coupled to an accelerator device 140 according to one example. The accelerator device 140 can be any device (e.g., a peripheral I / O device) that performs tasks issued by the host 105. In one embodiment, the host 105 is communicatively coupled to the accelerator device 140 using a PCIe connection. In another embodiment, the accelerator device 140 is part of the host-attached memory rather than a separate I / O device. The host 105 can represent a single computer (e.g., a server) or multiple interconnected physical computing systems. In any case, the host 105 includes multiple CPUs 110, memory 115, and a home agent 135.
[0030] Memory 115 includes an operating system (OS) 120, which can be any OS capable of performing the functions described herein. In one embodiment, OS 120 (or a hypervisor or kernel) establishes a cache-coherent shared memory multiprocessor paradigm for CPU 110 and memory 115. In one embodiment, CPU 110 and memory 115 are managed by the OS (or managed by the kernel / hypervisor) to form a coherency domain that follows the cache-coherent shared memory multiprocessor paradigm.
[0031] In the embodiments herein, the shared memory multiprocessor paradigm, along with all of the performance advantages, software flexibility, and overhead reduction of that paradigm, is available for the accelerator device 140. Furthermore, adding the computing resources in the accelerator device 140 to the same coherency domain as the CPU 110 and memory 115 enables a universal ISA-independent development environment.
[0032] In one embodiment, the accelerator device 140 and the host 105 use a coherent interconnect protocol to extend the coherence domain of the host 105 into the accelerator device 140. For example, the accelerator device 140 can use the Cache Coherent Interconnect for Accelerators (CCIX) to extend the coherence domain within the device 140. CCIX is a high-performance chip-to-chip interconnect architecture that provides a cache coherence framework for heterogeneous system architectures. CCIX brings kernel-managed semantics to the accelerator device 140. Cache coherence is automatically maintained between the CPU(s) 110 on the host 105 and various other accelerators in the system, which can be located on any number of peripheral I / O devices.
[0033] However, in addition to CCIX, other coherent interconnect protocols such as Quick Path Interconnect (QPI), Omni Path, Infinity Fabric, NVLink, or OpenCAPI may be used to extend the coherency domain in the host 105 to include the computing resources in the accelerator device 140. That is, the accelerator device 140 may be customized to support any type of coherent interconnect protocol that facilitates the formation of a coherency domain that includes the computing resources in the accelerator device 140.
[0034] In one embodiment, HA 135 performs consistency actions for the memory in system 100 (both the memory 115 in the host 105 and the memory 160 in the accelerator device 140). In a consistency system, a memory address range is attributed to a designated HA, which provides the advantage of fast consistency resolution of cache lines accessed by the CPU 110 in the host 105 and the request agent 145 in the accelerator device 140. HA 135 ensures that when a request to read or modify data is received, the most recent version of the data that can be stored in different locations (or caches) is retrieved. Those requests can originate from the request agent (e.g., request agent 145) in the CPU 110 or the attached accelerator device 140. Although one HA 135 is shown, the computing system 100 can have any number of HAs, each of which serves as the home of a different memory address range.
[0035] The accelerator device 140 can be many different types of peripheral devices, such as a pluggable card (which plugs into an expansion slot in the host 105), a system on a chip (SoC), a graphics processing unit (GPU), a field programmable gate array (FPGA), etc. For example, the accelerator device 140 may include programmable logic (e.g., a programmable logic array), or may not include any programmable logic, but instead contain only hardened circuitry (which may be software programmable).
[0036] The accelerator device 140 includes a request agent (RA) 145, a slave agent (SA) 155, and a memory 160 (also referred to as local memory). In one embodiment, the RA 145 is a computing engine that is optimized to perform specific computations. In one embodiment, the RA 145 is hardware, firmware, or a combination thereof. For example, the RA 145 may be one or more computing engines implemented in programmable logic. In another example, the RA 145 may be a computing engine formed by a hardened circuit device (e.g., a core in a GPU or other ASIC). In any case, the RA 145 executes an accelerator function 150 of the accelerator device. These functions 150 may be part of a machine learning accelerator, a cryptographic accelerator, a graphics accelerator, a search accelerator, a decompression / compression accelerator, etc. In one embodiment, the accelerator function 150 is implemented in programmable logic (e.g., a programmable logic array) in the accelerator device 140. In another embodiment, the accelerator function 150 is implemented in hardened logic (which may be software configurable) such as a processing core or engine.
[0037] Typically, one or more compute engines in the RA 145 execute the accelerator functions 150 faster or more efficiently than executing those same functions using one or more of the CPUs 110 in the host 105. Although one accelerator function 150 is shown and one accelerator application 125 is shown, the computing system 100 may have any number of accelerator functions, each processing a different data set in memory.
[0038] In one embodiment, SA 155 is a memory controller (which may be a hardware circuit) that services requests to read data from or write data to memory 160. RA 145 ensures that consistency is maintained when performing data memory operations from RA 145 to memory 160. For example, if RA 145 wants to read data from memory 160, RA 145 may first send a request to HA 135 to determine whether a more recent version of the data is available elsewhere in system 100. Further, although memory 160 is shown within accelerator device 140, in another embodiment, memory 160 may be on a separate chip (e.g., attached memory) from the integrated circuit forming accelerator device 140.
[0039] As described in more detail below, SA 155 has different functionality depending on the current phase or configuration of system 100. In one phase, HA 135 is the sole agent for data stored in memory 160, and SA 155 functions similarly to a typical memory controller in a coherent system. However, in another phase, software coherence has given SA 155 ownership of the data stored in local memory 160, and RA 145 does not need to ask HA 135 for permission when processing data stored in memory 160. For example, when RA 145 wants to read data from memory 160 to execute accelerator function 150, in the second phase, SA 155 can provide the data without RA 145 first checking with HA 135 to see if a more recent version of the data exists elsewhere in system 100 (e.g., stored in memory 115 or in an on-chip cache in one of CPUs 110).
[0040] Although RA 145 is illustrated as being separate from memory 160, in one embodiment, RA 145 may be a compute engine embedded within memory 160 and performing in-memory computations. In that case, the accelerator location may be in-memory, where RA 145 and SA 155 are logically distinct and perform in-memory computations, rather than as separate components. Figure 1 The accelerator device 140 is shown having physically distinct RA 145 and SA 155 .
[0041] Figure 2 is a block diagram of a host 105 coupled to an accelerator device 140 according to one example. Unlike traditional I / O models, the memory 160 and processing elements (e.g., RA 145 and SA 155) in the accelerator device 140 are in the same coherency domain as the CPU 110 and memory in the host 105. In this way, the HA 135 in the host 105 ensures that data stored in the host 105 and the accelerator device 140 is stored in a coherent manner so that requests for memory operations, whether originating from the host 105 or the accelerator device 140, receive the most recent version of the data, regardless of whether the data is stored in memory in the host 105 or in the accelerator device 140.
[0042] In some applications, the bottleneck in processing data occurs in moving the data to the computing unit that processes the data, rather than the time it takes to process the data. In other words, moving data to and from the processor or computing engine limits the speed at which operations can be performed, rather than the processor or computing engine being the bottleneck. This situation often occurs when accelerators are used because accelerators have dedicated computing engines (e.g., accelerator engines) that can perform certain types of functions extremely quickly. Typically, it is not the computing engine that limits the efficiency of using an accelerator, but the system's ability to move data to the accelerator for processing. The situation where moving data to the computing engine limits the time required to complete a task is called a computational memory bottleneck. Moving the computing engine closer to the data can help alleviate this problem.
[0043] In a first embodiment, a traditional I / O data sharing model can be used, in which large blocks of data are moved from the host 105 to the accelerator device 140. However, the host 105 and the accelerator device 140 are no longer in the same consistency domain. As a result, the I / O-based accelerator device cannot cache memory, which also separates the computing engine (e.g., RA 145) from the data it needs to process. Instead of using hardware consistency, software consistency is used, in which a software application (e.g., accelerator application 125) uses HA 135 to establish what data it wants to manage and what data is transferred to the accelerator device 140 (e.g., data set 130). Once the accelerator device 140 has completed processing the data, the accelerator application 125 can once again take ownership of the data and use HA 135 to maintain (or re-establish) hardware consistency for the data.
[0044] In a second embodiment, the host 105 and the accelerator device 140 can share the same coherency domain, allowing data to be cached in the accelerator device 140 to help alleviate computational memory bottlenecks. For example, memory can be attached to the host 105 (e.g., host-attached memory), while data can be stored in cache in the RA 145. However, a disadvantage of this model is that all accelerator-memory interactions pass through the HA 135 in the main CPU-memory complex in the host 105. This embodiment also uses hardware to enforce coherency, so regardless of whether the accelerator-memory data footprint is orders of magnitude larger than a cache line (MB, GB, or even TB), all accelerator-memory interactions are serviced and tracked by the HA 135, even though the CPU 110 (the generator) may ultimately only access a small portion of the data footprint for which the accelerator device 140 has performed computations. In other words, even though the CPU 110A may only access a small portion of the data, the HA 135 must track all data sent to the cache in the RA 145. Another disadvantage is that all memory accessed and cached by the accelerator device 140 is remote, host-attached memory.
[0045] In a third embodiment, the memory storing data processed by the accelerator device 140 (e.g., memory 160) is attached to (or located within) the accelerator device 140 but is owned by the HA 135 in the host 105. In this embodiment, the shortcomings of the second embodiment still apply, as all accelerator-memory interactions still pass through the HA 135 in the host 105. However, unlike the previous embodiment, once the accelerator device 140 has acquired ownership of a cache line, updates to the data set 130 can occur locally in the memory 160. This embodiment also uses hardware to enforce coherency, so another shortcoming of the second embodiment still applies, as regardless of whether the accelerator-memory data footprint is orders of magnitude larger than a cache line (MB, GB, or even TB) and regardless of whether the memory is accelerator-attached memory, all accelerator-memory interactions are still first serviced and tracked by the HA 135, even though the CPU 110A may ultimately only access a small portion of the data set 130 processed by the accelerator device 140.
[0046] Unlike the three embodiments described above, a collaborative framework can be used in which the benefits of the first, second, and third embodiments described above are incorporated into the system while avoiding their disadvantages. In addition to the benefits described above, the collaborative framework has additional advantages. In one embodiment, memory continues to be owned by the HA 135 in the host 105, thereby continuing to provide the advantages of fast consistency resolution to cache lines accessed by the CPU 110 in the host 105. Moreover, in contrast to requiring all accesses of the accelerator to memory to be tracked at a coarse level through software consistency as in the first embodiment, or in contrast to requiring all accesses of the accelerator to memory to have fine-grained tracking as in the second and third embodiments, in the collaborative framework, only a fine-grained subset of accelerator-memory interactions first passes through the HA 135 to claim ownership.
[0047] Another advantage of the cooperative framework is that, unlike software coherence where all memory pages or entire memory blocks can be flushed from a CPU cache (e.g., cache 210), for a subset of accelerator-memory interactions to HA 135, only the portion cached in the host CPU 110 is snooped from the CPU cache 210. Accelerator-memory interactions that do not touch the contents of the CPU cache 210 can continue to be cached, providing low-latency access to that data to the CPU 110 (and its cores 205).
[0048] Regarding another advantage of the collaborative framework, unlike the second and third embodiments where accelerator-memory interactions occur only using host-attached memory and this approach suffers from latency, bandwidth, and protocol message transmission overhead disadvantages, using the collaborative framework, accelerator-memory interactions can occur on the accelerator-attached memory 160 where the accelerator is located close to or in the memory, thereby providing a low-latency, high-bandwidth dedicated path and also minimizing protocol message transmission overhead due to the subsequent private accelerator-memory interactions between RA 145 and SA 155.
[0049] Regarding another advantage of the collaborative framework, unlike the third embodiment (in which interactions with memory require home node ownership even if the accelerator is located near or in the memory), in the collaborative framework, once the accelerator device 140 claims ownership of the dataset 130 on which computations are to be performed, all subsequent computation memory actions occur without interaction with the HA 135 until ownership is returned to the HA 135.
[0050] The following examples describe different techniques for using the collaboration framework to achieve the advantages described above. However, the collaboration framework is not limited to these advantages. Some implementations of the framework may have fewer advantages than those listed above, while other implementations may have different advantages not listed.
[0051] Figure 3 3 is a flow chart of a method 300 for using a collaborative hardware-software coherence framework, according to one example. At block 305, an accelerator application identifies a data set to be sent to an accelerator for processing. In one embodiment, an accelerator application is a software application that performs tasks that can benefit from an accelerator device. For example, an accelerator application can be: a machine learning application, in which the accelerator device includes one or more machine learning engines; a video application, in which the accelerator device performs graphics processing; a search engine, in which the accelerator device searches for specific words or phrases; or a security application, in which the accelerator device performs encryption / decryption algorithms on data.
[0052] In any case, the accelerator device can be aware of a collaborative framework that allows ownership of data to be transferred from the HA on the host to the SA on the accelerator device, while the host and accelerator device share the same coherence domain. In this way, the accelerator device can identify the data set it wishes to transfer to the accelerator device as part of the collaborative framework.
[0053] At block 310, the accelerator application (e.g., or some other software application) transfers ownership of the dataset from the server host to the accelerator device. In one embodiment, the accelerator application updates metadata or a flag monitored by the accelerator device. Updating this metadata or flag transfers ownership of the dataset from the HA to the accelerator device (e.g., the RA in the accelerator device). This transfer of ownership lets the hardware in the accelerator device know that it can work on the dataset without worrying about the HA in the host (because the accelerator device now owns the dataset).
[0054] In one embodiment, the accelerator application transfers ownership only for a subset of the data on which the compute memory action will be performed. In other words, the accelerator application can only identify the data that the accelerator device will process as part of its accelerator function. In this way, the accelerator application in the host can transfer ownership only for the data that the accelerator device will process. This is superior to traditional I / O models and software consistency, in which all pages or entire memory blocks are cleared from the CPU cache, rather than just the required data sets (which can be subsets of those pages and memory blocks).
[0055] At block 315, the accelerator device (or more specifically, the RA) exercises its ownership of the dataset by requesting the latest copy of the dataset from the server host. Doing so moves the dataset to local memory in the accelerator device. In one embodiment, the dataset may be ensured to be in local memory in the accelerator device as part of maintaining hardware consistency between the host and the accelerator device. In that case, the dataset identified at block 305 may have been stored in local memory in the accelerator before the accelerator application transfers ownership at block 310. However, in another embodiment, the RA (or accelerator device) may use clear commands and / or invalidate commands to ensure that the dataset is moved to local memory in the accelerator device and / or that no cached copies remain. Later in Figure 4 Various techniques for ensuring that local memory in an accelerator device contains the most recent version of a dataset are described in more detail in .
[0056] At block 320, the RA uses the SA to service the memory request. When the SA services the memory request in the accelerator device, the SA can service those requests without first using the HA to check, for example, whether a more recent version of the data is available elsewhere in the system. In other words, the data set now belongs to the accelerator device.
[0057] In one embodiment, when a dataset is owned by an accelerator device, the HA does not track any modifications to the dataset. This reduces overhead in the system because the RA can read and modify the dataset without the SA first needing permission from the HA. As discussed above, the HA tracks data on a per-cache line basis, meaning that each cache line in the dataset will need to be checked by the HA before servicing memory operations. However, since the dataset is owned by the SA, this function is performed solely by the SA without assistance from the HA in the host.
[0058] Furthermore, software consistency ensures that the CPU in the host can only read the dataset, not write to it, while the RA is processing it and until the RA indicates it is complete. Software consistency also ensures that if the CPU were to read data that has not been updated, it would not derive any meaning from reading a "stale" value when performing that read. Only after the accelerator has handed ownership back to the software / HA can the CPU derive meaning from the value in the updated data.
[0059] At block 325, the RA determines whether it has completed the dataset. If the RA has not yet completed executing its accelerator function on the dataset, method 300 repeats to block 320, where the RA continues to transmit additional memory requests for portions of the dataset to the SA. However, once the RA has completed processing the dataset, method 300 proceeds to block 330, where the RA transfers ownership of the dataset back to the HA. Thus, rather than software transferring ownership as performed at block 310, at block 330, it is the hardware in the accelerator device that sets software-accessible metadata indicating to the accelerator application that the accelerator device has completed the dataset. In one embodiment, when the RA has completed executing its accelerator function on the dataset, the RA updates the metadata, as required by the producer-consumer model for notifying the CPU. The HA can ensure that hardware consistency of the data (which has been modified) is maintained (or re-established) using entities within the host, such as host-attached memory and cache memory in the CPU. In other words, software (e.g., an accelerator application) implicitly exercises ownership of the dataset using the HA. The accelerator application only knows HA as a consistency enforcer.
[0060] In one embodiment, transferring ownership of a data set from HA to SA and back to HA as shown in method 400 can be considered a type of software coherence, although different from software coherence used for traditional I / O models, because the target data set is identified and there is no need to clear / invalidate caches in the host (e.g., CPU cache).
[0061] Figure 4 is a flow chart of a method 400 for using a cooperative hardware-software coherence scheme according to one example. Generally, method 400 describes different techniques for moving data from various memory elements in a host to local memory in an accelerator device so that the HA can transfer ownership of the data to the SA, as described in method 300.
[0062] At box 405, the CPU caches at least a portion of the data set. That is, before the accelerator application transfers ownership of the data set, some data of the data set may be stored in the cache in the CPU. At box 410, the CPU modifies the portion of the data set stored in its cache. That is, before the accelerator application transfers ownership at box 310 of method 300, some data of the data set may be stored in various caches and memories in the host. The data sets stored in the various caches may be duplicate copies of the memories in the host, or updates to the data sets due to software (e.g., accelerator application) using HA to implicitly exercise ownership of the data set. In this way, before the ownership of the data set is transferred to the SA in the accelerator device, the system requires a technology for ensuring that the latest version of the data set is stored in the accelerator attached memory (i.e., local memory).
[0063] At block 415, the method 400 is split based on whether software or hardware is used to send the modified data set to the accelerator device. That is, the task of ensuring that the accelerator device has the most recent data set can be performed by software (e.g., the accelerator application) or hardware in the accelerator device (e.g., the RA).
[0064] Assuming the software has been assigned this task, the method 400 proceeds to block 420, where the accelerator application in the host ensures that the modified data is sent to the accelerator device. In one embodiment, block 420 is performed before the accelerator device has transferred ownership of the data set from the HA in the host to the accelerator device (e.g., before block 310 in method 300).
[0065] There are at least two different options for moving the modified data in the server to the accelerator device, which are illustrated as alternative sub-boxes 425 and 430. At box 425, the HA only receives a purge command from the accelerator application. In one embodiment, the purge command is a persistent purge or a purge to target to ensure that all modified data (e.g., modified data in the CPU cache) has been updated in the persistent (i.e., non-volatile) memory or target memory in the accelerator device. In response to the purge command, the HA ensures that the modified data is cleared from the various caches in the host and copies the modified data to the local memory in the accelerator device.
[0066] The CPU in the host invalidates the cache, or maintains a cache-clean copy of the contents corresponding to the data set for future low-latency re-reference. In this example, the HA may not receive an explicit invalidation command from the accelerator application, but the CPU may go ahead and invalidate the portion of its data set stored in its cache. In another example, when the data belongs to the SA in the accelerator device, the CPU may maintain its own local copy of the data set (which it can modify).
[0067] In an alternative embodiment, the HA receives a flush command and an invalidate command from the accelerator application at block 430. In response, the CPU in the host flushes its cache and invalidates its cache.
[0068] In one embodiment, the flush command and the invalidate command described in blocks 425 and 430 are cache maintenance operations (CMOs) that automatically ensure that a modified data set (e.g., a most recent data set) is stored in local memory. In one embodiment, the accelerator application can issue two types of CMOs: one type of CMO that only ensures that the most recent data set is stored in local memory but permits the CPU to maintain a cache copy (e.g., a flush command only), and another type of CMO that not only ensures that the most recent data set is stored in local memory but also ensures that no cache copy remains (e.g., a flush command and an invalidate command).
[0069] Returning to block 415, if the hardware (rather than software) in the accelerator device is tasked with ensuring that it receives the modified data set, the method proceeds to block 435, where the hardware waits until the accelerator application transfers ownership of the data set to the accelerator device (as discussed in block 310 of method 300). Once this transfer of ownership occurs, method 400 proceeds to block 435, where the hardware in the accelerator device (e.g., RA) issues one or more CMOs to ensure that the modified data is sent to the accelerator device. Similar to block 420, there are at least two different options for moving the modified data in the host to the accelerator device, which are illustrated as alternative sub-blocks 445 and 450. At block 445, the HA only receives a clear command from the RA. The actions that the CPU in the host can take can be the same as those described in block 425, so they are not repeated here.
[0070] Alternatively, the HA receives a flush command and an invalidate command from the RA at block 450. In response, the CPU in the host flushes its cache and invalidates its cache as described at block 430.
[0071] In one embodiment, the HA may have a last-level cache. When transferring ownership, the HA can ensure that the newly acquired data set is not stuck in the last-level cache, but is instead copied to the local memory of the accelerator device. The final state in the last-level cache is the same as the state of the CPU cache and is consistent with the actions in the CPU cache generated by block 420 or block 440.
[0072] Figure 3 and Figure 4 Some other non-limiting advantages of the described techniques are that, while the memory remains owned by the HA in the host (when the data is not owned by the SA on the accelerator), all CPU cache and coherency actions to the memory are quickly resolved by the local HA, thereby maintaining high-performance CPU-memory access even when the collaboration framework exists in parallel. Further, the collaboration framework can request ownership of only a subset of the accelerator-memory interactions to exercise separate or combined CMO and ownership requests to the HA. By transferring ownership only for the relevant subset, the accelerator has a performance advantage because it does not have to request ownership of the entire data set as required by a pure software consistency model in the traditional I / O model.
[0073] Another advantage is as follows: CPU-memory interactions can also gain performance advantages because the subset that does not overlap with the accelerator-memory interaction continues to be served by the local HA that maintains high-performance CPU-memory access. Moreover, in one embodiment, only the subset of data shared between the CPU and the accelerator is cleared from the CPU cache, and the content that is not shared with the accelerator can continue to be cached in the CPU, thereby providing the CPU with low-latency access to the data. In addition, independent of the CPU-memory interaction, the accelerator-memory interaction in the specific architecture of the execution compute memory domain occurs on a low-latency, high-bandwidth dedicated path. The software consistency subset of the collaboration framework can ensure that metadata switching between the host and the accelerator allows independent CPU-memory execution and accelerator-memory execution, thereby improving the overall cross-sectional bandwidth, latency and program execution performance of both the CPU and accelerator devices.
[0074] Furthermore, although the above embodiments specifically describe a hybrid hardware-software consistency framework for computing memory bottlenecks, similar techniques can be applied to related CPU-accelerator shared memory applications.
[0075] Figure 5The FPGA 500 implementation of the accelerator device 140 is illustrated, more specifically, where the FPGA has a PL array that includes a large number of different programmable tiles, including transceivers 37, CLBs 33, BRAMs 34, input / output blocks ("IOBs") 36, configuration and clock logic ("CONFIG / CLOCKS") 42, DSP blocks 35, specialized input / output blocks ("IO") 41 (e.g., configuration ports and clock ports), and other programmable logic 39 (such as digital clock managers, analog-to-digital converters, system monitoring logic, etc.). The FPGA may also include a PCIe interface 40, an analog-to-digital converter (ADC) 38, etc.
[0076] In some FPGAs, each programmable tile may include at least one programmable interconnect element ("INT") 43 having connections to input and output terminals 48 of programmable logic elements within the same tile, such as Figure 5 . Each programmable interconnect element 43 may also include a connection to an interconnect segment 49 of an adjacent programmable interconnect element in the same tile or in (one or more) other tiles. Each programmable interconnect element 43 may also include a connection to an interconnect segment 50 of a general routing resource between logic blocks (not shown). The general routing resources may include routing channels between logic blocks (not shown), including tracks of interconnect segments (e.g., interconnect segments 50) and switch blocks (not shown) for connecting the interconnect segments. The interconnect segments (e.g., interconnect segments 50) of the general routing resources may span one or more logic blocks. The programmable interconnect elements 43, together with the general routing resources, implement a programmable interconnect structure ("programmable interconnect") for the illustrated FPGA.
[0077] In one example implementation, the CLB 33 may include a configurable logic element ("CLE") 44 that can be programmed to implement user logic plus a single programmable interconnect element ("INT") 43. In addition to one or more programmable interconnect elements, the BRAM 34 may also include a BRAM logic element ("BRL") 45. Typically, the number of interconnect elements included in a tile depends on the height of the tile. In the example shown, the BRAM tile has the same height as five CLBs, but other numbers (e.g., four) may also be used. In addition to an appropriate number of programmable interconnect elements, the DSP block 35 may also include a DSP logic element ("DSPL") 46. In addition to one instance of the programmable interconnect element 43, the IOB 36 may also include, for example, two instances of an input / output logic element ("IOL") 47. It will be clear to those skilled in the art that, for example, the actual IO pads connected to the IO logic element 47 are generally not limited to the area of the input / output logic element 47.
[0078] In the example shown, a horizontal region near the center of the die (e.g. Figure 5 ) is used for configuration, clock and other control logic. Extending from this horizontal area or column are vertical columns 51 that are used to distribute clock and configuration signals across the width of the FPGA.
[0079] use Figure 5 Some FPGAs of the illustrated architecture include additional logic blocks that disrupt the regular columnar structure that makes up the majority of the FPGA. The additional logic blocks can be programmable blocks and / or dedicated logic.
[0080] Notice, Figure 5 It is intended to illustrate only an exemplary FPGA architecture. For example, the number of logic blocks in a row, the relative widths of the rows, the number and order of the rows, the types of logic blocks included in the rows, the relative sizes of the logic blocks, and Figure 5 The interconnect / logic implementation shown at the top is for illustrative purposes only. For example, in a real FPGA, any location where a CLB appears typically includes more than one adjacent CLB row to facilitate efficient implementation of user logic, but the number of adjacent CLB rows varies with the overall size of the FPGA.
[0081] In the foregoing, reference is made to the embodiments presented in this disclosure. However, the scope of the present disclosure is not limited to the specific embodiments described. On the contrary, any combination of the described features and elements (whether or not relating to different embodiments) is contemplated as implementing and practicing the contemplated embodiments. Furthermore, although the embodiments disclosed herein may achieve advantages over other possible solutions or over the prior art, whether or not a given embodiment achieves a particular advantage does not limit the scope of the present disclosure. Therefore, unless expressly set forth in claim(s), the foregoing aspects, features, embodiments and advantages are merely illustrative and are not considered to be elements or limitations of the appended claims.
[0082] It will be appreciated by those skilled in the art that the embodiments disclosed herein may be embodied as systems, methods, or computer program products. Thus, various aspects may take the form of entirely hardware embodiments, entirely software embodiments (including firmware, resident software, microcode, etc.), or embodiments combining software and hardware aspects, all of which may be collectively referred to herein as "circuits," "modules," or "systems." Further, various aspects may take the form of computer program products embodied in one or more computer-readable media having computer-readable program code embodied thereon.
[0083] Any combination of one or more computer-readable media may be utilized. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include the following: an electrical connection with one or more wires, a portable computer floppy disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium is any tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0084] A computer-readable signal medium may include a propagated data signal having computer-readable program code embodied therein, for example, in baseband or as part of a carrier wave. Such a propagated signal may take any of a variety of forms, including but not limited to electromagnetic, optical, or any suitable combination thereof. A computer-readable signal medium may be any computer-readable medium that is not a computer-readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0085] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0086] Computer program code for performing operations of aspects of the present disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, Smalltalk, C++, and the like, and conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN); or can be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0087] Aspects of the present disclosure are described below with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments presented in the present disclosure. It should be understood that each box of the flowchart and / or block diagram and the combination of boxes in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine so that the instructions executed by the processor of the computer or other programmable data processing device create components for implementing the functions / actions specified in one or more boxes of the flowchart and / or block diagram.
[0088] These computer program instructions may also be stored in a computer-readable medium, which may direct a computer, other programmable data processing apparatus, or other device to operate in a specific manner so that the instructions stored in the computer-readable medium produce an article of manufacture including instructions that implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0089] The computer program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, such that the instructions executed on the computer or other programmable apparatus provide a process for implementing the functions / actions specified in one or more boxes of the flowchart and / or block diagram.
[0090] The flow chart and block diagram in the figure illustrate the architecture, function and operation of the possible implementation of the system, method and computer program product according to various examples of the present invention.In this regard, each frame in the flow chart or block diagram can represent a module, segment or part of an instruction, and the module, segment or part of an instruction include one or more executable instructions for realizing (one or more) specified logical functions.In some alternative implementations, the function pointed out in the frame may not occur in the order pointed out in the figure.For example, according to the function involved, the two frames shown in succession can actually be performed substantially simultaneously, or the frame can sometimes be performed in reverse order.It should also be noted that the combination of each frame of the block diagram and / or flow chart and the frame in the block diagram and / or flow chart can be realized by a system based on special-purpose hardware, and this system based on special-purpose hardware performs a specified function or action or performs a combination of special-purpose hardware and computer instructions.
[0091] While the foregoing relates to particular examples, other and further examples may be devised without departing from the basic scope thereof, and the scope of the same is to be determined by the claims that follow.
Claims
1. A computing system comprising: Host, including processing unit, home agent HA and accelerator application; as well as an accelerator device communicatively coupled to the host, wherein the accelerator device comprises a request agent RA, a slave agent SA, and a local memory, the RA being configured to perform at least one accelerator function, wherein the local memory is part of the same coherency domain as the processing unit in the host, wherein the accelerator application is configured to transfer ownership of a data set from the HA to the accelerator device to permit the RA to access the data set without first receiving permission from the HA, The accelerator device is configured as follows: storing the latest copy of the data set in the local memory, using the SA to service storage requests from the RA, and Transfer ownership of the dataset back to the HA. 2 . The computing system of claim 1 , wherein the accelerator application is configured to identify the data set, wherein the data set is a sub-portion of a memory block or a memory page. 3 . The computing system of claim 2 , wherein after the ownership of the data set is transferred from the HA to the accelerator device, a remaining portion of the memory block or the memory page continues to belong to the HA. 4 . The computing system of claim 1 , wherein the memory request from the RA is performed using the SA service without receiving permission from the HA. 5 . The computing system of claim 1 , wherein the processing unit comprises a cache, wherein the cache stores a modified version of the data set.
6. The computing system of claim 5, wherein the accelerator device is configured to: After ownership of the data set has been transferred to the accelerator device, a flush command is transmitted to the HA, the flush command instructing the processing unit to flush the cache so that the modified version of the data set is moved to the local memory.
7. The computing system of claim 6, wherein the processing unit is configured to, in response to the clear command, do one of the following: invalidating the modified version of the data set in the cache; and A cache-clean copy of the modified version of the data set is maintained for future low-latency re-reference.
8. The computing system of claim 5, wherein the accelerator application is configured to: Before ownership of the data set has been transferred to the accelerator device, a flush command is transmitted to the HA, the flush command instructing the processing unit to flush the cache so that the modified version of the data set is moved to the local memory.
9. The computing system of claim 8, wherein the accelerator application is configured to: Before ownership of the data set has been transferred to the accelerator device, an invalidate command is transmitted to the HA, the invalidate command instructing the processing unit to invalidate the cache.
10. An accelerator device comprising: a request agent RA comprising a computing engine configured to perform at least one accelerator function; From the agent SA, including the memory controller; as well as Memory, wherein the SA is configured to receive ownership of a data set from a home agent HA in a host to grant the RA access to the data set without first receiving permission from the HA, the host being coupled to the accelerator device, wherein the ownership is transferred via a software accelerator application, wherein the memory controller is configured to: service a request transmitted by the compute engine to access the data set stored in the memory once ownership has been transferred to the SA; and The RA is configured to transfer ownership from the SA back to the HA. 11 . The accelerator device of claim 10 , wherein the data set is a sub-portion of a memory block or a memory page, wherein after the ownership of the data set is transferred from the HA to the SA, a remaining portion of the memory block or the memory page continues to belong to the HA. 12 . The accelerator device of claim 10 , wherein the accelerator device is configured to communicate with the host using a coherent interconnect protocol to extend a coherence domain of the host to include the memory and the SA in the accelerator device.
13. A method comprising: transferring, by an accelerator application executed on a host, ownership of a data set from a home agent HA in the host to an accelerator device to permit the accelerator device to access the data set without first receiving permission from the HA, wherein the accelerator device is communicatively coupled to the host and a coherence domain of the host is extended into the accelerator device; Moving the data set to a local memory in the accelerator device; servicing a memory request from a requesting agent RA in the accelerator device using a slave agent SA in the accelerator device; and Transfer ownership from the accelerator device back to the HA.
14. The method according to claim 13, further comprising: The data set is identified using the accelerator application executing on the host, wherein the data set is a sub-portion of a memory block or a memory page, wherein after ownership of the data set is transferred from the HA to the accelerator device, a remainder of the memory block or the memory page continues to belong to the HA.
15. The method according to claim 13, further comprising: In response to a cache maintenance operation (CMO) issued by the accelerator device after transferring ownership to the accelerator device, the modified data set in the cache is invalidated.
Citation Information
Patent Citations
Systems and methods for protocol termination in a host system driver in a virtualized software defined storage architecture
US20180314540A1
Modular Virtual Assistant Platform
US20190012311A1