Adaptive caching of memory request streams
By identifying the relevant data streams in the computing task and allocating cache partitions using the stream id, the problem of cache resource competition is solved, the performance and hit rate of the cache is improved, and the effect of power consumption reduction and battery life is achieved.
Patent Information
- Application Number
- CN202280099539.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-26
- Publication Date
- 2025-05-30
AI Technical Summary
When caches are processed with multiple computing tasks, it may lead to resource competition and jitter, reducing cache usage efficiency, thereby affecting system performance and increasing power consumption.
By identifying the relevant data flows in the computing task, using the stream id to allocate different partitions of the cache, the partition allocation is adaptively adjusted to maximize cache performance and hit rate.
Effectively reduces competition for cache resources by different computing tasks, improves cache hit rate, reduces power consumption, and extends battery life.
Smart Images

Figure CN120077367A_ABST
Abstract
Description
Background Art
[0001] This specification relates to systems that include integrated circuit devices.
[0002] A cache is an auxiliary device that manages data traffic to a memory. The cache interacts with one or more hardware devices in the system to store data retrieved from the memory, or data to be written to the memory, or both. The hardware devices can be various components of an integrated circuit and can be implemented into a system-on-chip (SOC). A device that provides read and write requests either through the cache or directly to the memory will be referred to as a client device.
[0003] Caches are often utilized to reduce power consumption by limiting the total number of requests to the main memory. Further power savings can be achieved by placing the main memory and the data path to the main memory in a low-power state. Due to the inverse relationship between cache usage and power consumption, maximizing cache usage results in an overall decrease in the power consumed. By increasing the cache usage of integrated client devices, the power capacity of battery-powered devices such as mobile computing devices can be spent more efficiently. Additionally, accessing the cache is generally faster than accessing the main memory, thus improving the performance of integrated client devices.
[0004] Caches are generally organized into partitions to increase cache usage. A partition represents a portion of the cache that is allocated for a specific purpose or to a specific entity such as a specific client device. However, due to the limited size of the partitions relative to the working data set of a system such as a mobile computing device, and the way the memory needs to change over time, effective cache partitioning and stream allocation can be challenging. Cache thrashing can occur when the memory requests of client devices compete for the same resources in their respective partitions, and thus reduce cache usage. In these cases, system operation may become unavailable, thereby degrading system performance and increasing power consumption. Therefore, maximizing cache usage depends on optimizing the stream allocation to the cache partitions. Summary of the Invention
[0005] This specification describes techniques for implementing cache policies driven by related data streams, referred to herein as "computational tasks", in a cache. In this specification, a computational task can be associated with multiple memory requests that are related to each other in software. For example, a computational task can include all requests for data or all requests for instructions of a client device (or software driver). Depending on the specific workload of the client device, the device can execute multiple computational tasks sequentially or in parallel, each computational task including multiple related memory requests.
[0006] The cache can identify a computing task by examining a stream id common to different memory requests. The cache can then allocate different partitions of the cache memory to the different tasks by referencing the respective stream ids of the different tasks. Thus, for example, requests for instructions can be allocated to different partitions of the cache compared to requests for data. Additionally, the cache can adaptively allocate computing tasks based on the corresponding hit metrics for each partition. This ability allows the cache to self-tune to an optimal allocation of computing tasks that maximizes cache performance.
[0007] Specific embodiments of the subject matter described in this specification can be implemented so as to achieve one or more of the following advantages. The cache can improve cache performance and utilization by using a stream id to determine related computing tasks. Thus, the cache can reduce competition for cache resources among different computing tasks, which improves the cache hit rate. In mobile devices that rely on battery power, improving the cache hit rate reduces power consumption and extends battery life. Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the specification, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 is a diagram of an example system.
[0009] Figure 2 is a diagram of an example subsystem.
[0010] Figure 3 is a flowchart of an example process for allocating partitions in a cache.
[0011] Figure 4 is a flowchart of an example process for servicing a memory request using a partition of a cache dedicated to a computing task.
[0012] Figure 5 is a flowchart of an example process for adaptively allocating computing tasks to partitions of a cache.
[0013] Like reference numerals and names in the various drawings indicate like elements. DETAILED DESCRIPTION
[0014] Figure 1is a diagram of an example system 100. System 100 includes a plurality of client devices 110a, 110b through 110n that provide memory requests for locations in memory device 140. The foregoing components may be integrated onto a single system-on-chip (SOC) 102. Memory controller 130 may handle data requests to and from memory device 140 of system 100. Cache 120 caches data requests of multiple client devices on SOC 102, and thus, cache 120 may be referred to as a system-level cache (SLC). However, the techniques described below may be used with various types of devices that perform caching of memory requests. For example, a cache that caches memory requests for only a single client device or software driver, or a cache that caches memory requests for client devices not integrated on the same SOC 102 as cache 120.
[0015] SOC 102 is an example of a device that may be installed on or integrated into any suitable computing device, which may be referred to as a host device. Since the techniques described in this specification are particularly suitable for reducing power consumption of the host device and improving performance of the host device, SOC 102 may be particularly beneficial when installed on a battery-powered mobile host device such as a smart phone, a smart watch or other wearable computing device, a tablet computer, or a laptop computer, to name a few.
[0016] SLC 120 is an example of a cache that may be partitioned. Partitions 112a-n are portions of the cache that are allocated to memory requests having one or more attributes, such as ways or groups.
[0017] A plurality of client devices 110a-n are integrated on SOC 102. Each of client devices 110a-n may be a suitable module, device, or functional component configured to communicate memory requests to cache 120 and memory controller 130 via SOC fabric 150. For example, client devices 110a-n or SOC 102 itself may be a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), an ambient computing module, an image processor, a sensor processing module, an application specific integrated circuit (ASIC), or other lower-level components of SOC 102 itself that are capable of issuing memory requests to memory controller 130 via SOC fabric 150.
[0018] Client devices 110a-n provide memory requests through the SOC fabric 150 during the implementation of a workload, where each workload is executed by one or more compute tasks 112a-n. Each compute task may have one or more threads executing on the client device. Thus, a compute task and / or a thread may be associated with a particular memory request provided by the client devices 110a-n.
[0019] Figure 1 The example schematics in [reference] illustrate the workload of each client device with compute tasks 112a-n having different flow IDs. In this case, the flow ID uniquely identifies a specific compute task. The flow ID can be any identifier that differentiates compute tasks, e.g., a Universally Unique Identifier (UUID). Although Figure 1 not depicted in [reference], the threads of each compute task may also have flow IDs. Thus, one or more flow IDs identifying the compute task (or thread) to which the memory request belongs may be included in the memory request.
[0020] The SOC 102 or the client devices 110a-n on the SOC 102 may pre-assign a flow ID to a specific task 112a-n such that the flow ID is included in the memory requests from these tasks. For example, for a GPU client device executing a workload with multiple threads, each thread may have a separate pre-assigned flow ID. In some implementations, the flow ID assigned to a memory request of the SOC 102 is based on the type of task being executed. As an example, a thread responsible for texture mapping in a GPU may be pre-assigned a different flow ID than a thread responsible for rendering polygons. As another example, all threads operating on the same layer of a neural network in a TPU may share data with each other. These threads may be associated with a common compute task and be pre-assigned a corresponding flow ID. Other threads operating on different layers may be associated with their own compute tasks and corresponding flow IDs. In this scenario, pre-assigning a flow ID means that the flow ID has been assigned to the compute task before the SLC 120 starts processing the memory request of the compute task.
[0021] More generally, any relevant set of memory requests may have a corresponding flow ID. The flow ID of a memory request may be programmed, pre-loaded onto the client devices 110a-n, specified by the cache 120 or the SOC fabric 150, or created dynamically when servicing the memory request. In some cases, a set of memory requests may include memory requests from multiple client devices associated with corresponding flow IDs.
[0022] The SOC architecture 150 is the communication subsystem of the SOC 102. The SOC architecture 150 includes communication paths that allow the client devices 110a-n to communicate with each other and issue requests to read and write data using the cache 120 and the memory controller 130. The SOC architecture 150 can include any suitable combination of communication hardware such as, for example, a bus or a dedicated interconnect circuitry.
[0023] The system 100 also includes communication paths that allow communication between the cache 120 and the memory controller 130, communication between the SOC architecture 150 and the memory controller 130, and an inter-chip communication path that allows communication between the memory controller 130 and the memory device 140. In some implementations, the SOC 102 can save power by powering down one or more of the communication paths. Alternatively or additionally, the SOC 102 can power down the memory device 140 to further save power. As another example, the SOC 102 can enter a clock-off mode in which the respective clock circuits of one or more devices are powered down.
[0024] The cache 120 is positioned in one of the data paths in the data path between the SOC architecture 150 and the memory controller 130. Thus, requests from the client devices 110a-n to read from or write to the memory device 140 are passed through the cache 120, or directly to the memory controller 130. For example, the client 110a can issue a request to read from the memory device 140, and the request is passed through the SOC architecture 150 to the cache 120. The cache 120 can process the request before forwarding the request to the memory controller 130 for the memory device 140. Alternatively or additionally, the client 110a can issue a request to read from the memory device 140, and the request is passed directly through the SOC architecture 150 to the memory controller 130, thus bypassing the cache 120.
[0025] The cache 120 can cache read requests, write requests, or both from the client devices 110a-n. The cache 120 can cache read requests from the client devices 110a-n by responding to the requests with data stored in the cache data instead of fetching data from the memory device 140. Similarly, the cache 120 can cache write requests from the client devices 110a-n by writing new data to the cache instead of writing new data to the memory device 140. The cache 120 can perform a write-back at a later time to write the updated data to the memory device 140.
[0026] Cache 120 may have a dedicated cache memory, which may be implemented using dedicated registers or high-speed random access memory. Cache 120 may implement a caching policy that allocates different partitions (e.g., portions, ways) of the cache memory to different corresponding computing tasks. Thus, the same allocated portion of the cache memory can be used to handle memory requests belonging to the same task. For example, Figure 1 SOC 102 of FIG. Figure 1 shows cache 120 as having a plurality of allocated partitions 122a-n. Generally, the size (i.e., the space in memory) and the number of cache partitions 122a-n can be predefined and / or adjusted dynamically when servicing memory requests. In some cases, a subset of the partitions is predefined to reserve space in the cache memory, while the remaining memory is dynamically adjustable.
[0027] In some implementations, the same partition of the cache memory can be allocated to multiple tasks. To allocate one or more tasks to a partition, cache 120 can examine the stream id of the memory request to determine which memory requests belong to the same task.
[0028] An example of these techniques includes allocating different partitions of the cache to different computing tasks executed on the same client device. For example, cache 120 can examine the stream id of incoming memory requests to determine that some of the requests in the request are related to a process owned by a first task 112a and some other requests are related to a process owned by a second task 112b. Thus, to prevent these two tasks from competing with each other for cache resources, cache 120 can allocate a first partition 122a of the cache to the first task 112a executed on client device 110a, and can allocate a second partition 122b of the cache to the second task 112b executed on the same client device 110a. Alternatively or additionally, the first task 112a and / or the second task 112b executed on client device 110a may not be allocated a partition and thus bypass cache 120.
[0029] Cache 120 can also deallocate a computing task from a partition or swap a task from that partition to a different partition.
[0030] Another example includes allocating different partitions 112a-n of cache 120 to different buffers. For example, when SOC 102 is a GPU, each client device may perform a different function in a graphics processing pipeline. Thus, different data streams can be identified for, among other things, a render buffer, a texture buffer, and a vertex buffer.
[0031] The cache 120 can use a controller pipeline to handle memory requests from the SOC architecture 150. The controller pipeline implements cache logic for determining whether data exists in the cache 120 or whether data needs to be fetched from or written to memory. Thus, when memory access is needed—e.g., a cache miss—the controller pipeline can also provide a transaction to the memory controller 130.
[0032] Figure 2 FIG. is a diagram of an example subsystem 200. The subsystem 200 includes a candidate pool 210 that specifies a set of computational tasks 212a-n based on their respective flow ids. All tasks 222a-n that are not in the candidate pool 210 are non-candidates 220, which may or may not have a flow id. Although not depicted in Figure 2 the candidate pool 210 may also include flow ids of different threads of different computational tasks. The candidate pool 210 can include any number of flow ids corresponding to any number of tasks (e.g., no flow id, one flow id, two flow ids, etc.). Generally, the candidate pool 210 specifies a subset of computational tasks from the total set of computational tasks 112a-n executed on the client devices 110a-n, and the non-candidates 220 correspond to any computational tasks or other memory requests provided by the client devices 110a-n that are not in the candidate pool 210.
[0033] Figure 2 FIG. shows the candidate pool 210 as specified by the SOC architecture 150. However, any suitable thread or processing device can specify the candidate pool 210.
[0034] The candidate pool 210 can be modified in various ways based on various criteria. For example, the SOC architecture 150 can be configured to populate or depopulate the candidate pool 210 with flow ids based on metrics of the cache 120 or the main memory. The SOC architecture 150 can also populate the candidate pool 210 with pre-specified flow ids at startup. In some implementations, the candidate pool 210 can be modified by a suitable algorithm that is executed periodically or after a certain condition is met.
[0035] Subsystem 200 shows candidate pool 210 communicating with cache 120 and memory controller 130. As previously mentioned, cache 120 can be partitioned into multiple partitions 122a - n. In this case, only client devices that execute computational tasks 212a - n specified by candidate pool 210 can provide memory requests to cache 120, while all other tasks 222a - n directly provide memory requests to memory controller 130. Thus, SOC architecture 150 can use candidate pool 210 to isolate certain computational tasks 212a - n to optimize the use of cache 120. For example, tasks that are frequently executed but have predictable and / or limited data usage can be ideal for cache 120 allocation. SOC architecture 150 can populate candidate pool 210 with these types of tasks, although generally, candidate pool 210 can include any computational task.
[0036] The allocation engine of cache 120 can be configured to allocate computational tasks to partitions 122a - n of cache 120 using a flow id. For example, the allocation engine can allocate a first partition 112a of the cache for memory requests having a first flow id and a second partition 112b of the cache for memory requests having a second flow id. Thus, cache 120 can identify different computational tasks and allocate different partitions of cache memory to the corresponding tasks based on the corresponding flow id of the corresponding tasks. The allocation engine can use dedicated hardware circuitry of cache 120 to perform the allocation techniques described below.
[0037] Optionally or additionally, the allocation process can be implemented in software, and the allocation engine can cause the CPU of the host device to execute an allocation algorithm. In some implementations, the allocation process can be executed by a dedicated thread or processing device. The processing device can be integrated into SOC 102 or integrated into SOC architecture 150 along with multiple client devices 110a - n.
[0038] Figure 3 is a flowchart of an example process 300 for allocating partitions of a cache. Example process 300 can be executed by one or more components of the cache. Example process 300 will be described as being executed by the allocation engine of the cache on the SOC, which is appropriately programmed according to this specification.
[0039] The allocation engine identifies a flow id from the candidate pool corresponding to the flow id of the computational task (310). As described above, memory requests belonging to a particular task are assigned an associated flow id. In example process 300, the allocation engine only identifies flow ids from the candidate pool for allocation.
[0040] Multiple different events can trigger the cache to start the allocation process 300 by identifying the stream id of the memory request. For example, the cache can start allocation at startup. As another example, the SOC can be configured to automatically generate a repartition trigger event when the SOC detects a change in execution or usage. The trigger event can be a signal or data received through the system indicating that the candidate pool has been modified and that the cache partitions need to be reallocated. Alternatively or additionally, the cache can identify the stream id of the memory request by monitoring the memory traffic. For example, the cache can maintain hit metric statistics on all partitions and assign partitions to stream ids that meet specific criteria. The Figure 5 is an example process for adaptively allocating stream ids by monitoring memory traffic and modifying the candidate pool.
[0041] Memory requests can be associated with more than one stream id corresponding to more than one computing task. In the case of multiple stream ids, the cache can (at least partially) repeat the example process for each of the identified stream ids.
[0042] The allocation engine assigns partitions of the cache to memory requests with stream ids (320). The allocation engine can assign any appropriate partition of the cache, such as one or more rows, groups, ways, or some combination thereof. In some implementations, the partitions are assigned exclusively so that only memory requests with a specified stream id can use the allocated cache resources.
[0043] The allocation process can distinguish different types of computing tasks based on the stream ids of different types of computing tasks. For example, the allocation engine can distinguish tasks representing instructions and tasks representing data, and can assign one partition of the cache to instructions and another part of the cache to data. Additionally, the allocation engine can distinguish a first computing task executed by a client device from a second computing task executed by the same client device or a different client device, and can assign different partitions of the cache to the different computing tasks. Taking the GPU as an example, the allocation process 300 can identify the stream ids of texture mapping threads and polygon rendering threads and assign different partitions to each corresponding thread.
[0044] In some implementations, the allocation engine may assign special priorities to tasks with stream IDs storing data structures of a specific type and may allocate different amounts of cache resources to each task. For example, one data buffer that has a significant impact on cache utilization is the page table. Thus, the allocation engine may treat the data buffer storing page table data differently from the buffers storing other types of data. For example, the allocation engine may allocate 1 MB of cache memory for page table pages and 4 kb of cache memory for other types of data buffers.
[0045] The cache then services memory requests (330) from client devices on the SOC based on the requested stream ID. In this way, the cache can effectively dedicate partitions of the cache to different computing tasks.
[0046] Figure 4 is a flowchart of an example process 400 for servicing memory requests using partitions of the cache dedicated to computing tasks. The example process may be executed by one or more components of the cache. Example process 400 will be described as being executed by a cache, such as Figure 1 cache 120 on the SOC.
[0047] The cache receives a memory request (410). The memory request may be generated by a specific client device performing a specific computing task.
[0048] The cache identifies the stream ID of the task associated with the memory request (420). The stream ID may belong to a candidate pool and may thus have a dedicated partition.
[0049] The cache determines whether the stream ID has a dedicated cache partition (430). In response to determining that the stream ID has a dedicated cache partition, the cache services the memory request by using the dedicated cache partition (440). Otherwise, the cache services the memory request by using a default cache policy (450).
[0050] For example, a memory request from a GPU texture mapping thread may be received by the cache (410), and the stream ID may be identified (420). After determining that the stream ID has a dedicated cache partition (430), the cache may use that partition to service the memory request (440). Similarly, the cache may receive (410) a memory request from a polygon rendering thread of the GPU, and the cache may identify (420) the corresponding stream ID. The cache may determine that the stream ID does not have a dedicated partition (430) and thus uses the default cache policy to service the memory request (450).
[0051] Figure 5 is a flowchart of an example process 500 for adaptively allocating computing tasks to partitions of a cache. The example process 500 may be performed by one or more components of the cache and a dedicated thread or processing device configured to perform operations by executing instructions. The processing device may be a client device on a SOC or may be integrated separately into the SOC. For example, the processing device may be a CPU. In some implementations, the operations of the processing device are performed by the cache.
[0052] The cache allocates computing tasks from a candidate pool to cache partitions (510). As previously mentioned with respect to Figure 3 the cache may use the corresponding stream id of the tasks from the candidate pool to allocate the tasks. Generally, a subset of the tasks from the candidate pool is allocated to the partitions. The cache may allocate tasks in various ways. The cache may also be instructed by the processing device to allocate partitions. For example, the cache may be instructed by the processing unit to allocate partitions randomly, based on priority, algorithms, etc. Multiple computing tasks may also be allocated to the same partition. The cache partitions may be predefined according to a desired memory configuration, which is typically based on the total available memory in the cache. In principle, the partitions may also be adjusted dynamically during the example process 500, although this requires careful cache invalidation.
[0053] The processing device monitors the hit rate of all partitions of the cache (520). In the example process 500, the processing device performs operations based on the hit metrics of the cache. However, the processing device may also monitor other performance metrics of the cache, such as cache size, associativity, replacement policy, etc.
[0054] For each partition, the processing device identifies the hit rate of the partition (521). The processing device may continuously retrieve or calculate the per-partition hit rate across all partitions to identify a specific hit rate. The per-partition hit rate is the total number of cache hits for a specific partition divided by the sum of cache hits and cache misses. Generally, the process 500 aims to maximize the hit rate of all partitions.
[0055] The processing device determines whether the hit rate of the partition is below the eviction threshold of the partition (522). The eviction threshold defines a threshold for the hit rate and may serve as a measure of the desired baseline performance of the partition. Note that the eviction threshold may be different for different partitions. The eviction threshold may also vary depending on a particular implementation to adjust the performance of the cache. For example, the eviction threshold may be programmed, specified by the cache 120 or the SOC architecture 150, or dynamically created by the cache policy.
[0056] If the hit rate is greater than the eviction threshold, the processing device continues to monitor the hit rate on all partitions (branches to 520). Thus, the cache maintenance task assignment is maintained until the processing device determines that the hit rate is less than the eviction threshold on any particular partition.
[0057] If the processing device determines that the hit rate is less than the eviction threshold of a partition, the cache deallocates the underperforming tasks from the partition (523). In some cases, for example, when multiple tasks are assigned to a partition, the cache may deallocate more than one task from that partition.
[0058] The processing device determines whether the candidate pool is empty (524). Generally, process 500 will continue indefinitely until the candidate pool is empty. At this point, the processing device can perform further operations.
[0059] In some implementations, if the processing device determines that the candidate pool is empty, the processing device issues an interrupt and / or warning message (530). The warning message can alert the user for further instructions or can be issued silently. Alternatively or additionally, the candidate pool can then be refilled, for example, by the processing device based on user instructions or an appropriate algorithm.
[0060] On the other hand, if the candidate pool is not empty, the cache assigns new tasks from the candidate pool to the partition (525). The cache can select new tasks from the candidate pool in various ways. For example, the processing device can instruct the cache to select new tasks randomly or based on an algorithm. The algorithm can be polling, FIFO (first in first out), priority-based, etc.
[0061] After determining that the hit rate of a partition is less than the eviction threshold, the processing device determines whether the hit rate is also less than the resurrection threshold of the partition (526). The resurrection threshold specifies the minimum value of the hit rate of a partition. Again, the resurrection threshold can be different for different partitions. Similar to the eviction threshold, the resurrection threshold can also vary depending on a particular implementation to adjust the performance of the cache.
[0062] If the processing device determines that the hit rate is less than the resurrection threshold of the partition, the processing device removes the deallocated computational tasks from the candidate pool (527). Thus, when process 500 repeats, the removed tasks cannot be assigned to a partition in the cache. That is, after removing the deallocated tasks from the candidate pool, the processing device continues to monitor the hit rate on all partitions (520).
[0063] On the other hand, if the hit rate is greater than the eviction threshold, the deallocated tasks remain in the candidate pool, and the processing device continues to monitor the hit rate on all partitions (branches to 520). Thus, the deallocated tasks can be reassigned at a subsequent iteration in process 500.
[0064] Taking the threads of the GPU as an example, the candidate pool has stream IDs for texture mapping threads and polygon rendering threads. At the start of the adaptive allocation process 500, a partition of the cache can be allocated to the texture mapping thread, but not to the polygon rendering thread (510). When monitoring the hit metric (520), the process 500 can identify the hit rate of the partition to which the texture mapping is allocated (521) and determine that the hit rate is less than the eviction threshold (522). Then, the process 500 can deallocate the texture mapping thread from that partition (523). After determining that the candidate pool is not empty (524), the process 500 can allocate the polygon rendering thread to that partition (525). Although the hit rate of the partition to which the texture mapping is allocated is less than the eviction threshold, the hit rate can be greater than the resurrection threshold (526). In this case, the texture mapping thread remains in the candidate pool, and the process 500 repeats by continuing to monitor the hit rate (520). Thus, the overall effect of this iteration of the process 500 is to swap the texture mapping thread with the polygon mapping thread for underperforming partitions.
[0065] In general, the adaptive allocation process 500 will self - adjust to the optimal configuration of the computational tasks assigned to the cache partitions while removing tasks that tend to perform poorly from the candidate pool.
[0066] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non - transitory storage medium for execution by, or to control the operation of, a data processing apparatus. The computer storage medium can be a machine - readable storage device, a machine - readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagated signal (e.g., a machine - generated electrical, optical, or electromagnetic signal) that is generated to encode information for transmission to a suitable receiver device for execution by the data processing apparatus.
[0067] The term "data processing device" refers to data processing hardware and includes all kinds of devices, apparatuses, and machines for processing data, including, for example, programmable processors, computers, or multiple processors or computers. The device may also be or further include dedicated logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). In addition to the hardware, the device may optionally include code for creating an execution environment for computer programs, such as code that constitutes processor firmware, protocol stacks, database management systems, operating systems, or a combination of one or more of them.
[0068] A computer program, which may also be referred to or described as a program, software, software application, app, module, software module, script, or code, can be written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages); and it can be deployed in any form, which any form includes as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may or may not correspond to a file in a file system. A program can be stored in a part of a file that holds other programs or data (such as one or more scripts in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (such as files that hold one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a data communication network.
[0069] For a system of one or more computers that is configured to perform particular operations or actions, it means that the system has installed on it software, firmware, hardware, or a combination of them that, in operation, cause the system to perform those operations or actions. For one or more computer programs that are configured to perform particular operations or actions, it means that the one or more programs include instructions that, when executed by a data processing device, cause the device to perform the said operations or actions.
[0070] As used in this specification, an "engine" or "software engine" refers to a hardware-implemented or software-implemented input / output system that provides an output different from the input. An engine can be implemented in dedicated digital circuitry or as computer-readable instructions to be executed by a computing device. Each engine can be implemented within any suitable type of computing device (such as a server, mobile phone, tablet computer, notebook computer, music player, e-book reader, laptop or desktop computer, PDA, smart phone, or other fixed or portable device) that includes one or more processing modules and a computer-readable medium. Additionally, two or more of the engines can be implemented on the same computing device or on different computing devices.
[0071] The processes and logical flows described in this specification can be performed by one or more programmable computers that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by, for example, special logic circuitry such as an FPGA or ASIC, or by a combination of special logic circuitry and one or more programmed computers.
[0072] Computers suitable for executing computer programs can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, the central processing unit will receive instructions and data from a read only memory or a random access memory or both. The basic elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special logic circuitry. Generally, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or will be operatively coupled to receive data from or transfer data to one or more mass storage devices, or both. However, a computer need not have such devices. In addition, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name just a few.
[0073] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices (such as EPROM, EEPROM, and flash memory devices); magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0074] To provide interaction with a user, embodiments of the subject matter described in this specification may be implemented on a host device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and a pointing device (e.g., a mouse, a trackball, or a presence-sensitive display or other surface) through which the user can provide input to the computer. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including sound, voice, or tactile input. In addition, the computer may interact with the user by sending documents to and receiving documents from the devices used by the user; for example, by sending a web page to a web browser on the user device in response to a request received from the web browser. In addition, the computer may interact with the user by sending a text message or other form of message to a personal device (e.g., a smart phone running a messaging application), and receiving a responsive message from the user in response.
[0075] In addition to the above embodiments, the following embodiments are also innovative:
[0076] Embodiment 1 is a system, including: A plurality of integrated client devices, each client device being configured to generate a memory request, each memory request having a corresponding pre-assigned flow id, the flow id indicating the type of computing task to which the memory request belongs; and A cache, the cache being configured to cache memory requests for memory of each of the plurality of integrated client devices, wherein the cache has a plurality of partitions, and wherein the cache is configured to allocate different partitions to corresponding memory requests according to the flow id of the memory requests.
[0077] Embodiment 2 is the system according to Embodiment 1, wherein memory requests belonging to different types of computing tasks have different flow ids.
[0078] Embodiment 3 is the system according to any one of Embodiments 1 to 2, wherein the cache is configured not to allocate a partition for a specific flow id.
[0079] Embodiment 4 is the system according to any one of Embodiments 1 to 3, wherein the cache is configured to exchange the flow id from using a first partition to using a second partition.
[0080] Embodiment 5 is the system according to any one of Embodiments 1 to 4, wherein the cache is configured to allocate a plurality of different flow ids to use the same partition.
[0081] Example 6 is a system as described in any one of Examples 1 to 5, further comprising a processing device configured to execute instructions to perform operations including the following: Providing instructions to the cache to allocate partitions to stream IDs from a candidate pool of stream IDs; Calculating a per-partition cache hit metric for each partition; and Providing instructions to the cache to change the partition allocation of one or more stream IDs.
[0082] Example 7 is a system as described in Example 6, wherein calculating the per-partition cache hit metric includes calculating a hit rate, and wherein the operations further include: Determining that the hit rate of the partition is less than an eviction threshold; and In response, deallocating one or more stream IDs from the partition and allocating new stream IDs from the candidate pool to the partition.
[0083] Example 8 is a system as described in Example 7, wherein the operations further include: Determining that the hit rate of the partition is less than a resurrection threshold; and In response, removing the deallocated one or more stream IDs from the candidate pool.
[0084] Example 9 is a system as described in any one of Examples 6 to 8, wherein the system is configured to use a selection algorithm based on any one of the following to allocate new stream IDs from the candidate pool to the partition: Randomly, Polling, First in first out, or Priority.
[0085] Example 10 is a system as described in any one of Examples 7 to 9, wherein the eviction thresholds of at least some of the partitions are different.
[0086] Example 11 is a system as described in any one of Examples 8 to 10, wherein the resurrection thresholds of at least some of the partitions are different.
[0087] Example 12 is a method performed by a device, the device comprising: A plurality of integrated client devices, each client device configured to generate a memory request, each memory request having a corresponding pre-assigned stream ID, the stream ID representing the type of computing task to which the memory request belongs; and a cache having a plurality of partitions; the method comprising: Cache memory requests to the memory for each of the multiple integrated client devices; and Allocate different partitions to the corresponding memory requests by the cache according to the stream id of the memory requests.
[0088] Embodiment 13 is the method as described in Embodiment 12, wherein memory requests belonging to different types of computing tasks have different stream ids.
[0089] Embodiment 14 is the method as described in any one of Embodiments 12 to 13, wherein the cache is configured not to allocate partitions to a specific stream id.
[0090] Embodiment 15 is the method as described in any one of Embodiments 12 to 14, further comprising swapping the stream id from using the first partition to using the second partition.
[0091] Embodiment 16 is the method as described in any one of Embodiments 12 to 15, further comprising allocating multiple different stream ids to use the same partition.
[0092] Embodiment 17 is the method as described in any one of Embodiments 12 to 16, further comprising: Providing instructions to the cache to allocate partitions to stream ids from a candidate pool of stream ids; Calculating a per-partition cache hit metric for each partition; and Providing instructions to the cache to change the partition allocation of one or more stream ids.
[0093] Embodiment 18 is the method as described in Embodiment 17, wherein calculating the per-partition cache hit metric includes calculating a hit rate, and wherein the operation further comprises: Determining that the hit rate of the partition is less than an eviction threshold; and In response, deallocating one or more stream ids from the partition and allocating new stream ids from the candidate pool to the partition.
[0094] Embodiment 19 is the method as described in Embodiment 18, further comprising: Determining that the hit rate of the partition is less than a resurrection threshold; and In response, removing the deallocated one or more stream ids from the candidate pool.
[0095] Embodiment 20 is the method as described in any one of Embodiments 17 to 19, further comprising using a selection algorithm based on any one of the following to allocate new stream ids from the candidate pool to the partition: Randomly, Polling, First in, first out, or priority.
[0096] Example 21 is the method according to any one of Examples 18 to 20, wherein the eviction thresholds of at least some of the partitions are different.
[0097] Example 22 is the method according to any one of Examples 19 to 21, wherein the resurrection thresholds of at least some of the partitions are different.
[0098] Example 23 is a computer storage medium encoded with a computer program, the program including instructions that, when executed by a data processing device, are operative to cause the data processing device to perform the method according to any one of Examples 12 to 22.
[0099] Although this specification contains many specific implementation details, these details should not be construed as limitations on the scope of any invention or of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented separately in multiple embodiments or in any suitable sub-combination. Moreover, although the features may be described above as acting in certain combinations and even initially claimed as such, in some cases one or more features from a claimed combination can be deleted from the combination, and the claimed combination may cover a sub-combination or a variant of a sub-combination.
[0100] Similarly, although the operations are depicted in the drawings in a particular order, this should not be understood as requiring that the operations be performed in the particular order shown or in a sequential order, or that all illustrated operations be performed, to achieve a desired result. In certain circumstances, multitasking and parallel processing may be advantageous. In addition, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0101] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. For example, the acts recited in the claims can be performed in a different order and still achieve a desired result. As one example, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve a desired result. In some cases, multitasking and parallel processing may be advantageous.
Claims
1. A system, comprising: a plurality of integrated client devices, each client device being configured to generate a memory request, each memory request having a corresponding pre-assigned flow id, the flow id indicating the type of computing task to which the memory request belongs; and a cache, the cache being configured to cache memory requests for memory of each of the plurality of integrated client devices, wherein the cache has a plurality of partitions, and wherein the cache is configured to allocate different partitions to corresponding memory requests according to the flow id of the memory requests.
2. The system according to claim 1, wherein memory requests belonging to different types of computing tasks have different flow ids.
3. The system according to any one of claims 1 to 2, wherein the cache is configured not to allocate a partition for a specific flow id.
4. The system according to any one of claims 1 to 3, wherein the cache is configured to swap the flow id from using a first partition to using a second partition.
5. The system according to any one of claims 1 to 4, wherein the cache is configured to allocate a plurality of different flow ids to use the same partition.
6. The system according to any one of claims 1 to 5, further comprising a processing device, the processing device being configured to execute instructions to perform operations including: providing instructions to the cache to allocate a partition to a flow id from a candidate pool of flow ids; calculating a per-partition cache hit metric for each partition; and providing instructions to the cache to change the partition allocation of one or more flow ids.
7. The system according to claim 6, wherein calculating the per-partition cache hit metric includes calculating a hit rate, and wherein the operations further comprise: determining that the hit rate of the partition is less than an eviction threshold; and in response, deallocating one or more flow ids from the partition and allocating new flow ids from the candidate pool to the partition.
8. The system according to claim 7, wherein the operations further comprise: determining that the hit rate of the partition is less than a resurrection threshold; and in response, removing the deallocated one or more flow ids from the candidate pool.
9. The system according to any one of claims 6 to 8, wherein the system is configured to use a selection algorithm based on any one of the following to allocate new flow ids from the candidate pool to the partition: randomly, polling, first in first out, or priority.
10. The system according to any one of claims 7 to 9, wherein the eviction thresholds of at least some of the partitions are different.
11. The system according to any one of claims 8 to 10, wherein the resurrection thresholds of at least some of the partitions are different.
12. A method performed by a device, the device comprising: a plurality of integrated client devices, each client device being configured to generate a memory request, each memory request having a corresponding pre-assigned flow id, the flow id indicating the type of computing task to which the memory request belongs; and A cache, the cache having a plurality of partitions; The method includes: Caching, by the cache, memory requests for memory of each of the plurality of integrated client devices; And Allocating, by the cache, different partitions to corresponding memory requests according to the flow id of the memory requests.
13. The method according to claim 12, wherein memory requests belonging to different types of computing tasks have different flow ids.
14. The method according to any one of claims 12 to 13, wherein the cache is configured not to allocate partitions to a specific flow id.
15. The method according to any one of claims 12 to 14, further comprising swapping the flow id from using a first partition to using a second partition.
16. The method according to any one of claims 12 to 15, further comprising allocating a plurality of different flow ids to use the same partition.
17. The method according to any one of claims 12 to 16, further comprising: Providing an instruction to the cache to allocate a partition to a flow id from a candidate pool of flow ids; Calculating a per-partition cache hit metric for each partition; And Providing an instruction to the cache to change the partition allocation of one or more flow ids.
18. The method according to claim 17, wherein calculating the per-partition cache hit metric includes calculating a hit rate, and wherein the operation further comprises: Determining that the hit rate of the partition is less than an eviction threshold; And In response, deallocating one or more flow ids from the partition and allocating new flow ids from the candidate pool to the partition.
19. The method according to claim 18, further comprising: Determining that the hit rate of the partition is less than a resurrection threshold; And In response, removing the deallocated one or more flow ids from the candidate pool.
20. The method according to any one of claims 17 to 19, further comprising using a selection algorithm based on any one of the following to allocate new flow ids from the candidate pool to the partition: Randomly, Polling, First in first out, or Priority.
21. The method according to any one of claims 18 to 20, wherein the eviction thresholds of at least some of the partitions are different.
22. The method according to any one of claims 19 to 21, wherein the resurrection thresholds of at least some of the partitions are different.