METHOD AND SYSTEM FOR MEMORY BANDWIDTH CONTROL - Patent application
Patent Information
- Application Number
- JP2024510657
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-02-07
- Filing Date
- 2022-07-20
- Publication Date
- 2025-06-30
AI Technical Summary
Existing systems struggle to efficiently manage cache prefetching and regular memory accesses for multiple clients in a multi-core microprocessor, leading to inefficiencies in memory bandwidth usage and resource partitioning.
Implementing a method to track and adjust memory bandwidth usage for each resource portion assigned to a client, using credit counts and thresholds to control data access requests, and dynamically manage cache prefetching based on congestion levels.
Enhances the efficiency of memory access by optimizing bandwidth usage and reducing congestion, ensuring fair and timely access for multiple clients while minimizing resource contention.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Related Applications
[0001]
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 239,702, entitled "Methods and Systems for Memory Bandwidth Control," filed on September 1, 2021, U.S. Provisional Patent Application No. 63 / 251,517, entitled "Methods and Systems for Memory Bandwidth Control," filed on October 1, 2021, and U.S. Provisional Patent Application No. 63 / 251,518, entitled "Methods and Systems for Memory Bandwidth Control," filed on October 1, 2021, each of which is incorporated by reference in its entirety into this specification.
[0002]
[0002] This application also claims priority to U.S. patent application Ser. No. 17 / 666,438, entitled "Methods and Systems for Memory Bandwidth Control," filed Feb. 7, 2022, the entire contents of which are incorporated by reference herein. [Technical field]
[0003]
[0003] This application relates generally to microprocessor technology, including, but not limited to, methods, systems, and devices for controlling memory access to memory external to one or more processing clusters of a microprocessor that provides computational and storage resources to multiple clients. [Background technology]
[0004]
[0004] A large amount of traffic is often present in the microprocessor of a computer system to facilitate both cache prefetching from slower memory or cache to faster local cache and normal memory accesses required by the operations of the individual processor units of the microprocessor. In the context of a processor cluster (i.e., a multi-core microprocessor), the computational and storage resources of the microprocessor may be partitioned to sponsor multiple tenants or clients with different portions of these resources. It would be highly desirable to provide an electronic device or system that efficiently manages the cache prefetching and normal memory accesses associated with different clients for each processor cluster of a multi-core microprocessor. Summary of the Invention
[0005]
[0005] Various implementations of systems, methods and devices within the scope of the appended claims each have several aspects, no single aspect of which is solely responsible for the attributes described herein. Without limiting the scope of the appended claims, after considering this disclosure, and in particular the section entitled "Description of the Preferred Embodiments," it will be understood how aspects of several implementations are used to manage memory request accesses to memory blocks (e.g., double data rate synchronous dynamic random access memory (DDR SDRAM)) outside a processing cluster based on memory bandwidth usage status of different clients of an electronic device. The resources of the electronic device are partitioned into resource portions utilized by different clients. A memory bandwidth usage status is tracked for each resource portion to monitor in real time how much of the memory access bandwidth allocated to each resource portion is used to access the memory block. A usage level is derived from the memory bandwidth usage status of the resource portion to control whether to issue the next data access request associated with the respective resource portion in the memory access request queue. In some implementations, for each resource portion, a lower usage level of the memory block and / or a longer duration of remaining at a low usage level leads to a higher likelihood of issuing the next data access request. By these measures, data access requests associated with different clients can be efficiently and individually managed based on the existing usage levels of the memory blocks of these clients.
[0006] In one aspect, a method is implemented in an electronic device for managing memory access. The electronic device includes one or more processing clusters and a plurality of memory blocks, each processing cluster including one or more respective processors and coupled to at least one of the memory blocks. The method includes partitioning resources of the electronic device into a plurality of resource portions to be utilized by a plurality of clients. Each resource portion is assigned to a respective client and has a respective partition identifier (ID). The method further includes receiving a plurality of data access requests associated with the plurality of clients for the plurality of memory blocks. The method further includes, for each resource portion having a respective partition ID, tracking a plurality of memory bandwidth usage states corresponding to the memory block and determining a usage level associated with the respective partition ID from the plurality of memory bandwidth usage states. Each memory bandwidth usage state is associated with a respective memory block and indicates at least how much of the memory access bandwidth assigned to the respective partition ID is used to access the respective memory block. The method further includes, for each resource portion having a respective partition ID, adjusting a credit count based on a usage level, comparing the adjusted credit count to a request issuance threshold, and issuing a next data access request associated with the respective partition ID in the memory access request queue in accordance with a determination that the credit count is greater than the request issuance threshold.
[0007]
[0007] In some circumstances, the method further includes, for each resource portion having a respective partition ID, pursuant to a determination that the credit count is less than the request issuance threshold, suspending issuance of any data access requests from the memory access request queue of the respective partition ID until the credit count is adjusted to be greater than the request issuance threshold.
[0008] In another aspect, a method is implemented in a first memory for managing memory access. The first memory is coupled to one or more processing clusters and a plurality of memory blocks in an electronic device. The method includes forwarding a plurality of data access requests associated with a plurality of clients to a plurality of memory blocks. A resource of the electronic device is partitioned into a plurality of resource portions to be utilized by a plurality of clients, each resource portion being assigned to a respective client and having a respective partition ID. The method further includes, for each resource portion having a respective partition ID, identifying a subset of data access requests associated with the respective ID to access the memory block, and tracking a plurality of memory bandwidth usage states corresponding to the memory block. Each memory bandwidth usage state is associated with the respective memory block and indicates at least how much of the memory access bandwidth assigned to the respective partition ID to access the respective memory block is used. The method further includes, for each resource portion having a respective partition ID, determining in response to each of the subset of data access requests that the respective data access request is to access a corresponding memory block, receiving a memory bandwidth usage state of the corresponding memory block, and reporting the memory bandwidth usage state of the corresponding memory block to the one or more processing clusters.
[0009] In yet another aspect, a method is implemented in a memory system for tracking memory usage. The memory system is coupled to one or more processing clusters via a first memory in an electronic device and includes a memory block. The method includes receiving a set of data access requests associated with a plurality of clients for the memory block. The resource is partitioned into a plurality of resource portions to be utilized by the plurality of clients, each resource portion being assigned to a respective client and having a respective partition ID. The method includes identifying, for each resource portion having a respective partition ID, a subset of data access requests associated with the respective ID to access the memory block, and tracking a memory bandwidth usage status associated with the respective partition ID. The memory bandwidth usage status indicates at least how much of the memory access bandwidth assigned to the respective partition ID to access the memory block is used. The method further includes reporting the memory bandwidth usage status to the one or more processing clusters in response to each of the set of data access requests.
[0010]
[0010] Other implementations and advantages will be apparent to those skilled in the art in light of the description and drawings herein. [Brief description of the drawings]
[0011] [Figure 1]
[0011] A block diagram of exemplary system modules in a typical electronic device, according to some implementations. [Diagram 2]
[0012] 1 is a block diagram of an example electronic device having one or more processing clusters, according to some implementations. [Figure 3A]
[0013] 1 is a block diagram of an exemplary electronic device that controls and tracks requests to access data stored in memory blocks external to a processing cluster, according to some implementations. [Figure 3B]1 is a block diagram of an exemplary electronic device that controls and tracks requests to access data stored in memory blocks external to a processing cluster, according to some implementations. [Figure 4]
[0014] FIG. 1 illustrates an example process implemented by a controller of a processing cluster to control requests of resource partitions to access data stored in memory blocks based on memory bandwidth usage, according to some implementations. [Figure 5A]
[0015] 1 illustrates an example process implemented by a memory to track memory bandwidth usage and current congestion levels of individual memory blocks of the memory, according to some implementations. [Figure 5B] 1 illustrates an example process implemented by a memory to track memory bandwidth usage and current congestion levels of individual memory blocks of the memory, according to some implementations. [Figure 6A]
[0016] 1 illustrates an exemplary process implemented by a cache to track memory bandwidth usage and current congestion levels for each memory block, according to some implementations. [Figure 6B]
[0017] 1 illustrates an exemplary process implemented by a cache to track its current congestion level, according to some implementations. [Figure 6C]
[0018] FIG. 1 illustrates another exemplary process implemented by a cache to track memory bandwidth usage status of each memory block, the current congestion level, and the current congestion level of the cache itself, according to some implementations. [Figure 7A]
[0019] 1 is a diagram of an example data structure for data stored in a processing cluster for managing data access requests of multiple resource partitions, according to some implementations. [Figure 7B]1 is a diagram of an example data structure of data stored in a cache to manage data access requests for multiple resource partitions, according to some implementations. [Figure 7C] 1 is a diagram of an example data structure of data stored in a memory block for managing data access requests of multiple resource partitions, according to some implementations. [Figure 8]
[0020] 1 illustrates an example method for determining a congestion level of a processing cluster for controlling cache prefetching in the processing cluster, according to some implementations. [Figure 9]
[0021] 1 illustrates an example method for determining system congestion levels for controlling cache prefetching in individual processing clusters, according to some implementations. [Figure 10]
[0022] 4 is a flowchart of a method for managing memory access to memory 104 by an electronic device according to some implementations. [Figure 11]
[0023] 13 is a flowchart of a method for tracking memory bandwidth usage in a first memory (eg, a cache) coupled to one or more processing clusters and to a plurality of memory blocks, according to some implementations. [Figure 12]
[0024] 4 is a flowchart of a method for tracking memory bandwidth usage of memory blocks of a memory system, according to some implementations. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012]
[0025] These exemplary embodiments and implementations are mentioned not to limit or define the disclosure, but to provide examples to aid in its understanding. Additional embodiments are discussed in the detailed description and further explanation is provided therein. Other implementations and advantages will be apparent to those skilled in the art in light of the description and drawings herein.
[0013]
[0026] Reference will now be made in detail to certain embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth to aid in understanding the subject matter presented herein. However, it will be apparent to those skilled in the art that various alternatives may be used without departing from the scope of the claims, and that the subject matter may be practiced without these specific details.
[0014]
[0027] FIG. 1 is a block diagram of an exemplary system module 100 in a typical electronic device according to some implementations. The system module 100 in this electronic device includes at least a system on chip (SoC) 102, a memory module 104 for storing programs, instructions and data, an input / output (I / O) controller 106, one or more communication interfaces such as a network interface 108, and one or more communication buses 140 for interconnecting these components. In some implementations, the I / O controller 106 enables the SoC 102 to communicate with I / O devices (e.g., a keyboard, a mouse or a trackpad) via a Universal Serial Bus interface. In some implementations, the network interface 108 includes one or more interfaces for Wi-Fi, Ethernet and Bluetooth networks, each of which allows the electronic device to exchange data with an external source, e.g., a server or another electronic device. In some implementations, the communication bus 140 includes circuitry (sometimes referred to as a chipset) that interconnects and controls communication between various system components included in the system module 100.
[0015]
[0028] In some implementations, the memory module 104 (e.g., memory 104 of FIGS. 2-11, memory system of FIG. 12) includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some implementations, the memory module 104 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. In some implementations, the memory module 104, or alternatively the non-volatile memory devices within the memory module 104, include a non-transitory computer-readable storage medium. In some implementations, a memory slot is reserved on the system module 100 to accept the memory module 104. When inserted into the memory slot, the memory module 104 is integrated into the system module 100.
[0016]
[0029] In some implementations, the system module 100 further includes one or more components selected from the following:
[0017] A memory controller 110 that controls communication between the SoC 102 and memory components, including the memory module 104, in the electronic device; A solid-state drive (SSD) 112, which employs integrated circuit assemblies for storing data in an electronic device, and in many implementations is based on NAND or NOR memory configurations; · a hard drive 114, which is a conventional data storage device based on electromechanical magnetic disks used to store and retrieve digital information; a power connector 116 electrically coupled to receive an external power source; A power management integrated circuit (PMIC) 118 that modulates the received external power supply to other desired DC voltage levels, e.g., 5V, 3.3V, or 1.8V, as required by various components or circuits within the electronic device (e.g., the SoC 102); a graphics module 120 that generates an output image feed to one or more display devices according to a desired image / video format; and · A sound module 122 that facilitates the input and output of audio signals to and from electronic devices under the control of a computer program.
[0018]
[0030] Also, note that communication bus 140 interconnects and controls communications between various system components, including components 110-122.
[0019]
[0031] Furthermore, those skilled in the art will appreciate that other non-transitory computer readable storage media may be used as new data storage technologies are developed for storing information on the non-transitory computer readable storage media in memory module 104 and in SSD 112. These new non-transitory computer readable storage media include, but are not limited to, those made from biological materials, nanowires, carbon nanotubes, and individual molecules, although each of these data storage technologies is currently under development and has not yet been commercialized.
[0020]
[0032] In some implementations, the SoC 102 is implemented on an integrated circuit incorporating one or more microprocessors or central processing units, memory, input / output ports, and secondary storage on a single substrate. The SoC 102 is configured to receive one or more internal power supply voltages provided by the PMIC 118. In some implementations, both the SoC 102 and the PMIC 118 are mounted on a main logic board, e.g., on two separate areas of the main logic board, and are electrically coupled to each other via conductors formed in the main logic board. As explained above, this arrangement results in parasitic effects and electrical noise that may impair the performance of the SoC, e.g., cause voltage drops in the internal voltage supplies. Alternatively, in some implementations, the SoC 102 and the PMIC 118 are vertically arranged in an integrated semiconductor device such that they are electrically coupled to each other via electrical connections that are not formed in the main logic board. Such a vertical arrangement of the SoC 102 and the PMIC 118 can reduce the length of the electrical connections between the SoC 102 and the PMIC 118 and can avoid performance degradation caused by conductors on the main logic board. In some implementations, the vertical arrangement of the SoC 102 and the PMIC 118 is facilitated, in part, by the incorporation of thin-film inductors in the limited space between the SoC 102 and the PMIC 118.
[0021]
[0033] 2 is a block diagram of an example electronic device 200 having one or more processing clusters 202 (e.g., a first processing cluster 202-1, an Mth processing cluster 202-M) according to some implementations. In addition to the processing clusters 202, the electronic device 200 further includes a cache 220 and a memory 104. The cache 220 is coupled to the processing clusters 202 on the SOC 102, which is further coupled to a memory 104 that is external to the SOC 102. The memory 104 includes a number of memory blocks 222, which may be dynamic random access memory (DRAM). Each processing cluster 202 includes one or more processors 204, a cluster cache 212, and a controller 216. The cluster cache 212 is coupled to the one or more processors 204 and maintains one or more request queues 214 for the one or more processors 204. Each processor 204 further includes a respective prefetcher 208 coupled to a controller 216 of the respective processing cluster 202 to control cache prefetching associated with the respective processor 204. In some implementations, each processor 204 further includes a core cache 218, possibly split into an instruction cache and a data cache, which stores instructions and data that may be immediately executed by the respective processor 204.
[0022]
[0034] In one example, the first processing cluster 202-1 includes a first processor 204-1, ..., Nth processor 204-N, a first cluster cache 212-1, and a first controller 216-1, where N is an integer greater than 1. The first cluster cache 212-1 has one or more first request queues 214-1, each of which includes a queue of demand requests and prefetch requests received from a subset of the processors 204 of the first processing cluster 202-1. In some implementations, the SOC 102 includes only a single processing cluster 202-1. Alternatively, in some implementations, the SOC 102 includes at least an additional processing cluster 202, e.g., an Mth processing cluster 202-M. The Mth processing cluster 202-M includes a first processor 206-1, ..., an N'th processor 206-N', an Mth cluster cache 212-M, and an Mth controller 216-M, where N' is an integer greater than 1, and the Mth cluster cache 212-M has one or more Mth request queues 214-M.
[0023]
[0035] In some implementations, the one or more processing clusters 202 are configured to provide a central processing unit (CPU) for the electronic device and are associated with a hierarchy of caches. For example, the hierarchy of caches includes three levels that are differentiated based on their different operating speeds and sizes. In this application, references to the "speed" of a memory (including cache memory) relate to the time required to write data to or read data from the memory (e.g., faster memories have shorter write and / or read times than slower memories), and references to the "size" of a memory relate to the storage capacity of the memory (e.g., smaller memories provide less storage space than larger memories). The core cache 218, the cluster cache 212, and the cache 220 correspond to a first level (L1), second level (L2), and third level (L3) cache, respectively. Each core cache 218 holds instructions and data to be directly executed by the respective processor 204, and has the fastest operating speed and smallest size of the three levels of memory. For each processing cluster 202, the cluster cache 212 operates slower than the core cache 218, is larger in size, and holds data that is less likely to be accessed by the processors 204 of the respective processing cluster 202 than the data held in the core cache 218. The cache 220 is shared by multiple processing clusters 202 and is larger in size and slower than each of the core caches 218 and the cluster cache 212. In each processing cluster 202, the respective controller 216 monitors system congestion levels associated with memory accesses to the cache 220 and memory 104, and local cluster congestion levels associated with the cluster cache 212, and controls the prefetching of instructions and data to the core cache 218 and / or cluster cache 212 based on the system and / or cluster congestion levels.Each individual processor 204 further monitors processor congestion levels to control the prefetching of instructions and data from the respective cluster cache 212 to the respective individual core cache 218 .
[0024]
[0036] In some implementations, the first cluster cache 212-1 of the first processing cluster 202-1 is coupled to a single processor 204-1 in the same processing cluster, but not to any other processors (e.g., 204-N). In some implementations, the first cluster cache 212-1 of the first processing cluster 202-1 is coupled to multiple processors 204-1 and 204-N in the same processing cluster. In some implementations, the first cluster cache 212-1 of the first processing cluster 202-1 is coupled to one or more processors 204 in the same processing cluster 202-1, but not to processors in any cluster other than the first processing cluster 202-1 (e.g., processor 206 in cluster 202-M). In such cases, the first cluster cache 212-1 of the first processing cluster 202-1 may be referred to as a second level cache.
[0025]
[0037] In each processing cluster 202, each request queue 214 includes a queue of demand requests and prefetch requests possibly received from a subset of the processors 204 of the respective processing cluster 202. Each data access request received from each processor 204 is distributed to one of the request queues 214. In some implementations, the request queue 214 receives requests only received from a particular processor 204. In some implementations, the request queue 214 receives requests from more than one processor 204 in the processing cluster 202, allowing the request load to be distributed among the multiple request queues 214. In particular, in some situations, the request queue 214 receives only one type of data access request (e.g., a prefetch request) from different processors 204 in the same processing cluster 202. Each data access request in the request queue 214 is issued under the control of the controller 216-1 to access the cache 220 and / or the memory 104 to implement a memory read or write operation. In some implementations, only data access requests that are not satisfied by the cache 220 are sent further to the memory 104 , and each such data access request may be satisfied by a respective memory block 222 of the memory 104 .
[0026]
[0038] In each processing cluster 202, a controller 216 is coupled to the output of the cluster cache 212, a request queue 214 in the cluster cache 212, and one or more processors 204 of the processing cluster 202. In particular, the controller 216 is coupled to both the cache 220 and the memory 104 via the output of the cluster cache 212. The computational and storage resources of the electronic device 200 are partitioned into a number of resource portions to be utilized by a number of clients 224. Each resource portion is assigned to a respective client 224 and has a respective partition identifier (ID). The request queue 214 includes a number of data access requests associated with the multiple clients 224 to request memory access to a number of memory blocks 222 in the cache 220 or the memory 104. For each resource portion (i.e., each client 224) with a respective partition ID, the controller 216 tracks a number of memory bandwidth usage states (i.e., 402 in FIG. 4) corresponding to different memory blocks 222 of the memory 104. Each memory bandwidth usage state is associated with a respective memory block 222 of memory 104 and indicates at least how much of the memory access bandwidth allocated to each partition ID is used to access each memory block 222 of memory 104. Controller 216 determines a usage level (i.e., 406 in FIG. 4) associated with each partition ID from the plurality of memory bandwidth usage states, adjusts a credit count (i.e., 408 in FIG. 4) based on the overall usage level, and issues a next data access request (i.e., 412 in FIG. 4) associated with each partition ID (i.e., each client 224) in request queue 214 based on the credit count.
[0027]
[0039] In some implementations, at the cluster level, the controller 216 monitors a local cluster congestion level of a corresponding processing cluster 202 based on signals received from the request queue 214. In particular, the controller 216 determines a congestion level of the processing cluster 202 based on the extent to which a plurality of data access requests sent to the cluster cache 212 from one or more processors 204 in the processing cluster 202 are not satisfied by the cluster cache 212. Pursuant to a determination that the congestion level of the processing cluster 202 satisfies a first congestion criterion requiring that the congestion level of the processing cluster 202 be above a first cluster congestion threshold, the controller 216 causes a first respective processor (e.g., processor 204-1) of the one or more processors 204 to limit prefetch requests to the cluster cache 212 to prefetch requests of at least a first threshold quality (i.e., limit prefetch requests to high quality prefetches). In particular, in one example, controller 216 sends a signal or other information to processor 204 (e.g., prefetcher 208-1 in processor 204-1) to enable prefetch throttling such that only prefetch requests of at least a first threshold quality are sent to cluster cache 212. This corresponds, in some cases, to a second prefetch throttling mode M2 that is different from the first prefetch throttle mode and limits prefetching by processor 204 from cluster cache 212 to prefetch requests of at least a first threshold quality 804 of FIG.
[0028]
[0040] Alternatively, pursuant to a determination that the congestion level of the processing cluster 202 does not satisfy the first congestion criterion (e.g., the congestion level of the processing cluster 202 falls below a first cluster congestion threshold), the controller 216 refrains from causing one or more processors to limit prefetch requests to the cluster cache 212 to prefetch requests of at least the first threshold quality. For example, the controller 216 refrains from causing the processor 204 to fully limit prefetch requests to the cluster cache 212 such that prefetch requests of any quality are not limited. This potentially corresponds to a first prefetch throttling mode M1, in which prefetching of the processor 204 from the cluster cache 212 is not limited by the controller 216 described with reference to FIG.
[0029]
[0041] In some implementations, a congestion level below the first cluster congestion threshold indicates a low degree of congestion in the cluster cache 212, and a congestion level above the first cluster congestion threshold indicates one or more higher degrees of congestion. If the one or more higher degrees of congestion correspond to a single high degree of congestion, a congestion level above the first cluster congestion threshold indicates this high degree of congestion. In contrast, if the one or more higher degrees of congestion correspond to a set of congestion degrees (e.g., medium, high, and very high), a congestion level above the first cluster congestion threshold is associated with any degree in the set of congestion degrees.
[0030]
[0042] Further, in some implementations, at a system level, the controller 216 monitors the system congestion level of the memory system coupled to the processing cluster 202 based on a system busy level signal (i.e., current congestion level 504 or 604) received from an output of the cluster cache 212. The system busy level signal includes information of outstanding in-flight requests received and not satisfied by the cache 220 or the memory 104. In particular, the controller 216 obtains the current congestion level 604 of the cache 220 (e.g., HN[2] in FIG. 6B) based on the number of outstanding in-flight requests received by the cache 220, and maintains a first congestion level history (e.g., history 902 in FIG. 9) including the obtained current congestion level 604 of the cache 220. The controller 216 also obtains a current congestion level 504 of the memory 104 (e.g., SN[2] in FIG. 5B ) based on the number of outstanding in-flight requests received by the memory 104, and maintains a second congestion level history (e.g., history 904 in FIG. 9 ) that includes the current congestion level 504 of the memory 104. In some circumstances, data access requests that are not satisfied by the cache 220 are sent further to the memory 104, and the number of outstanding in-flight requests received by the memory 104 (i.e., the current congestion level 504) is thus determined based on the extent to which the data access requests sent to the cache 220 are not satisfied by the cache 220.
[0031]
[0043] The controller 216 causes the processing cluster 202 to limit prefetch requests from the processing cluster 202 based on at least one of the current congestion level 604 of the cache 220 and the current congestion level 504 of the memory 104. In some implementations, the prefetch requests from the processing cluster 202 are limited based on the first congestion level history and / or the second congestion level history. In some implementations, the controller 216 is configured to determine a first congestion level of the cache 220 (which is a composite congestion level) based on the first congestion level history or to determine a second congestion level of the memory 104 (which is a composite congestion level) based on the second congestion level history. The prefetch requests from the processing cluster 202 may be disabled from the request queue 214 that the processing cluster 202 is joining based on the first congestion level and / or the second congestion level. In some implementations, the history of the first congestion level and / or the history of the second congestion level are maintained by the controller 216 itself. Furthermore, a cluster congestion threshold applied to control prefetch quality is indicated based on the first and / or second congestion level history of the cache 220 and the memory 104. Further details regarding the application of the system congestion levels of the cache 220 and the memory 104 are described below with reference to Figures 8 and 9.
[0032]
[0044] 3A and 3B are block diagrams of example electronic devices 300 and 350 for controlling and tracking requests to access data stored in memory blocks 222 external to a processing cluster 202, according to some implementations. In each of the electronic devices 300 and 350, one or more processing clusters 202 are coupled to a cache 220, which is further coupled to a memory 104 that includes a plurality of memory blocks 222. Each processing cluster 202 includes one or more processors 204 and a cluster cache 212 coupled to the one or more processors 204. Each processor 204 further includes a core cache 218 and a prefetcher 208, and the cluster cache 212 further includes one or more request queues 214 and a controller 216. The core cache 218, the cluster cache 212, and the cache 220 form a hierarchy of caches to provide instructions and data to the processors 204. The core cache 218 is configured to store instructions and data to be executed directly by each processor 204, and the cluster cache 212 is configured to provide instructions and data that are less likely to be executed by the processor 204 and that will be loaded into the core cache 218 when needed. The cache 220 is configured to provide instructions and data that are less likely to be executed by the processor 204 than those in the cluster cache 212 and that will be loaded into the cluster cache 212 when needed. The cluster cache 212 of the processing cluster 202 includes one or more data access request queues 214, which further include a plurality of data access requests, i.e., including all demand requests and all prefetch requests, sent from one or more processors 204 to the cache 220 within a predefined time period. In some implementations, if a data access request is not satisfied by the cache 220, it is further sent to one of the plurality of memory blocks 222 of the memory 104 (e.g., the first memory block 222A).
[0033]
[0045] 3A, in some implementations, the plurality of data access requests in the one or more data access request queues 214 includes a read request 302 configured to request retrieval of a data item from a first memory block 222A in the memory 104. The read request 302 is associated with one of the plurality of clients 224 (e.g., the first client 224A) and is performed by the processing cluster 202 for the one of the plurality of clients 224. The controller 216 controls the processing cluster 202 to issue the read request 302 to the cache 220. The cache 220 forwards the read request 302 to the first memory block 222A. Upon receiving the read request 302, the first memory block 222A retrieves the data item requested by the read request, determines that the read request is associated with one of the plurality of clients 224, and obtains a memory bandwidth usage status MBUS tracked locally for the one of the plurality of clients 224. The memory bandwidth usage status MBUS indicates at least how much of the memory access bandwidth allocated to one of the plurality of clients 224 to access the first memory block 222A is used. In response to the read request 302, the first memory block 222A sends the requested data item directly to the processing cluster 202. In some implementations, the memory bandwidth usage status MBUS of one of the plurality of clients 224 is sent directly to the processing cluster 202 with the requested data item. Alternatively, in some implementations, the memory bandwidth usage status MBUS of one of the plurality of clients 224 is sent to the cache 220, which forwards the memory bandwidth usage status MBUS to the processing cluster 202. Furthermore, in some implementations, in response to a single read request 302, the memory bandwidth usage status MBUS of one of the plurality of clients 224 is reported twice by the first memory block 222A, i.e., directly from the first memory block 222A to the processing cluster 202 and indirectly via the cache 220.
[0034]
[0046] In some implementations, the multiple data access requests in the one or more data access request queues 214 of each processing cluster 202 include multiple read requests 302, each read request 302 configured to request retrieval of a respective data item from a respective memory block 222 in the memory 104. Each read request 302 is associated with a respective client 224 and is performed by the processing cluster 202 on behalf of the respective client. In response to each read request 302, the memory block 222 corresponding to the respective read request 302 reports the memory bandwidth usage status MBUS of the respective client 224 to the processing cluster 202 directly or indirectly via the cache 220, thereby enabling the processing cluster 202 to track multiple memory bandwidth usage status MBUSs for the multiple clients 224. Each client 224 corresponds to a subset of the memory bandwidth usage status MBUSs each associated with a respective one of the memory blocks 222 of the memory 104. By these means, for each client 224 , the memory bandwidth usage status MBUS associated with a memory block 222 of memory 104 is updated in response to a read request 302 issued by the processing cluster 202 on behalf of the respective client 224 .
[0035]
[0047] 3B, in some implementations, the plurality of data access requests in the one or more data access request queues 214 includes a write request 304 configured to request storage of a data item in the first memory block 222A in the memory 104. The write request 304 is associated with one of the plurality of clients 224 and is made by the processing cluster 202 for the one of the plurality of clients 224. Thus, the write request 304 is implemented using storage resources allocated to the one of the plurality of clients 224. The controller 216 controls the processing cluster 202 to issue the write request 304 to the cache 220. The cache 220 forwards the write request 304 to the first memory block 222A. Upon receiving the write request 304, the first memory block 222A possibly writes or does not write the data item included in the write request 304 to the memory unit depending on whether the one of the plurality of clients 224 has remaining memory access bandwidth allocated to it. Further, the first memory block 222A determines that the write request 304 is associated with one of the multiple clients 224 and obtains a memory bandwidth usage status MBUS tracked for the one of the multiple clients 224. In response to the write request 304, the first memory block 222A sends a write confirmation message to the cache 220 indicating whether the data item was written to the first memory block 222A. The write confirmation message further includes the memory bandwidth usage status MBUS of the one of the multiple clients 224. The cache 220 forwards the write confirmation message including the memory bandwidth usage status MBUS of the first memory block 222A by the one of the multiple clients 224 to the processing cluster 202.
[0036]
[0048] In some implementations, the data access requests in the one or more data access request queues 214 include a plurality of read requests 304, with each write request 304 configured to request storage of a respective data item in a respective memory block 222 in the memory 104. Each write request 304 is associated with a respective client 224 and is performed by the processing cluster 202 for the respective client 224. In response to each write request 304, the memory block 222 corresponding to the respective write request 304 reports the memory bandwidth usage status MBUS of the respective client 224 to the processing cluster 202 indirectly via the cache 220. By these means, for each client 224, the memory bandwidth usage status associated with the memory block 222 of the memory 104 is updated in response to the write request 304 issued by the processing cluster 202 for the respective client 224.
[0037]
[0049] FIG. 4 illustrates an exemplary process 400 implemented by the controller 216 of the processing cluster 202 to control requests for resource partitions to access data stored in the memory block 222 based on memory bandwidth usage state 402, according to some implementations. As described above, the electronic device includes one or more processing clusters 202, a cache 220, and a memory 104. Such computational and storage resources are shared among multiple clients 224 and are therefore partitioned into multiple resource portions to be utilized by the multiple clients 224. Each resource portion is assigned to a respective client 224 and has a respective partition identifier (ID) that represents the respective resource portion and the respective client 224. The process 400 is implemented by the controller 216 of the processing cluster 202 for a first client 224A that is assigned a resource partition associated with the respective partition ID. Each client 224 is, as the case may be, a private person or a business entity that subscribes to computer services provided by the electronic device. The controller 216 of the processing cluster 202 stores a memory block usage table 401 for each of a number of clients, including the first client 224A.
[0038]
[0050] In particular, the memory block usage table 401 includes a number of rows. Each row corresponds to a respective one of the memory blocks 222 of the memory 104, and is configured to store and track a number of memory bandwidth usage states 402 corresponding to the memory blocks 222. Each memory bandwidth usage state 402 is associated with a respective memory block 222 and indicates at least how much (e.g., 75%) of the memory access bandwidth allocated to a respective partition ID of the first client 224A is used to access the respective memory block 222. For example, referring to FIG. 4, the memory block usage table 401 includes 32 rows corresponding to the 32 memory blocks 222 of the memory 104. For each row, a first column includes an integer representing a memory block identification of each memory block 222, and a second column includes a flag representing the respective memory bandwidth usage state 402, for example, whether more than 75% of the memory access bandwidth allocated to the first client 224A is used to access the respective memory block 222. At least memory blocks 0 and 31 are using more than 75% of the memory access bandwidth allocated to the first client 224A, and at least memory block 1 is not using more than 75% of the memory access bandwidth allocated to the first client 224A.
[0039]
[0051] A plurality of data access requests are waiting in one or more request queues 214 of the processing cluster 202. The controller 216 is configured to operate according to a clock frequency and manage the issuance of the plurality of data access requests based on the memory bandwidth usage state 402 of the memory block 222. In some circumstances, the plurality of data access requests are generated by two or more resource partitions of two or more clients 224 and include a subset of data access requests for the resource partition of a first client 224A. The subset of data access requests further includes a first request 404A and a second request 404B following the first request 404A. Each request 404 is possibly a read request (e.g., read request 302) to read a data item from the respective memory block 222 or a write request (e.g., write request 304) to store a data item in the respective memory block 222. The controller 216 issues a subset of data access requests associated with the resource partition of the first client 224A to access the different memory blocks 222 based on the memory bandwidth usage states 402 stored in the memory block usage table 401 in association with the different memory blocks 222.
[0040]
[0052] In some implementations, the controller 216 generates a usage level 406 associated with the partition ID of the first client 224A from the plurality of memory bandwidth usage states 402 stored in the memory block usage table 401. For example, the usage level 406 is equal to the number of memory blocks 222 that have used more than 75% of the memory access bandwidth allocated to the partition ID of the first client 224A, i.e., the number of "Y"s in the second column of the memory block usage table 401. More specifically, in one example, the usage level 406 is equal to 11, where 11 of the 32 memory blocks 222 are using more than 75% of the memory access bandwidth allocated to the partition ID of the first client 224A.
[0041]
[0053] The controller 216 adjusts (e.g., accumulates) the credit count 408 based on the usage level 406 and compares the credit count 408 to a request issuance threshold 410 to determine whether a next data access request 412 associated with the partition ID of the first client 224A needs to be issued. If the credit count 408 accumulates above the request issuance threshold 410, the next data access request 412 associated with the partition ID of the first client 224A is issued. The credit count 408 is possibly reset (414) to zero or reduced by a predefined value (e.g., by 1, the request issuance threshold 410). Conversely, if the credit count 408 is less than the request issuance threshold 410, the controller 216 suspends (416) one or more request queues 214 from issuing any data access requests for the respective partition ID until the credit count 408 is adjusted to be greater than the request issuance threshold 410.
[0042]
[0054] In some implementations, the controller 216 adjusts the credit count 408 based on the usage level 406, at least in part according to the clock frequency. After a first request 404A is issued to access a respective memory block 222 for a partition ID associated with the first client 224A, one or more of the plurality of memory bandwidth usage states 402 stored in the memory block usage table 401 associated with the first client 224A are updated. After a predefined number of clock cycles following this update of the memory bandwidth usage states 402, the usage level 406 is determined from the plurality of memory bandwidth usage states 402 stored in the memory block usage table 401. Furthermore, after a predefined number of clock cycles following the update of the memory bandwidth usage states 402, the credit count 408 is adjusted and compared to a request issuance threshold periodically, for example, once during each subsequent clock cycle or once every five clock cycles, until a next data access request 412 (e.g., the second request 404B) is issued.
[0043]
[0055] In some implementations, after determining the usage level 406 associated with the respective partition ID of the first client 224A from the memory bandwidth usage state 402, the controller 216 compares the usage level 406 to one or more usage thresholds associated with the partition ID (e.g., a high SN of a high usage threshold and a low SN of a low usage threshold). In some implementations, the high SN or low SN of the usage threshold is different for different clients 224. Alternatively, in some implementations, the high SN or low SN of the usage threshold is the same for different clients 224. In accordance with a determination that the usage level 406 is equal to or greater than the high SN of the high usage threshold of the first client 224A (418), the controller 216 reduces the credit count 408 by the respective credit unit CU corresponding to the respective partition ID of the first client 224A (420). In some implementations, the credit count 408 is periodically decreased by a respective credit unit CU every one or more clock cycles until the next data access request 412 (e.g., the second request 404B) is issued (422). Conversely, in accordance with a determination that the usage level 406 is equal to or less than the low SN of the low usage threshold of the first client (424), the controller 216 increases the credit count 408 by a respective credit unit corresponding to the partition ID (426). In some implementations, the credit count 408 is periodically increased by a respective credit unit every one or more clock cycles until the next data access request 412 (e.g., the second request 404B) is issued (428). Furthermore, in accordance with a determination that the usage level 406 is between the high SN of the high usage threshold and the low SN of the low usage threshold, the controller 216 maintains the credit count 408.
[0044]
[0056] For each partition ID of each client 224 (e.g., the first client 224A), the credit count 408 indicates a priority level for issuing a data access request of the first client 224A. In some implementations, the usage level 406 of the first client is high (i.e., quite close to its memory access bandwidth for accessing the memory block 222), and a fairly high credit count 408 can still result in a relatively high priority level for issuing the next data access request 412 associated with the first client 224A. Despite the high usage level 406 of the first client, the next data access 412 is still issued for the partition ID of the first client 224A because of this fairly high credit count 408. Conversely, in some implementations, the usage level 406 of the first client is low (i.e., quite far from its memory access bandwidth for accessing the memory block 222), and a fairly low credit count 408 can still result in a relatively low priority level for issuing the next data access request 412 associated with the first client 224A. Despite the low usage level 406 of the first client, the next data access request 412 still cannot be issued for the partition ID of the first client 224A because of this rather low credit unit 408. However, under some circumstances, despite the low usage level 406 of the first client, this rather low credit unit 408 gradually increases over time until the next data access request 412 is issued for the partition ID of the first client 224A, and the relatively low priority level for issuing the data access request of the first client 224A also gradually increases over time. In the worst case scenario, the usage level 406 of the first client is high (i.e., quite close to its memory access bandwidth for accessing the memory block 222), and the rather low credit unit 408 results in a relatively low priority level for issuing the next data access request 412 associated with the first client 224A.The controller 216 waits for the fairly low credit count 408 to increase incrementally over time until the next data access request 412 is issued for the partition ID of the first client 224A. Thus, a lower usage level 406 of the memory block and / or a longer duration of remaining at the low usage level translates to a higher likelihood of issuing the next data access request 412.
[0045]
[0057] After the controller 216 issues each request 404, the respective request 404 is received by the cache 220 and forwarded to the corresponding memory block 222 of the memory 104. In some implementations, in response to a read request 404 issued from the respective partition ID of the first client 224A for the respective memory block 222, the respective memory block 222 directly updates (430) the respective memory bandwidth usage state 402 of the respective memory block 222 to the processing cluster 202 in parallel with providing the data item requested by the read request. Alternatively, in some implementations, in response to a read request 404 issued from the respective partition ID of the first client 224A, the respective memory block 222 indirectly updates (432A) the respective memory bandwidth usage state 402 of the respective memory block 222 via the cache 220. Additionally, in some implementations, each memory bandwidth usage state 402 of each memory block 222 is updated twice in the memory block usage table 401, directly (430) from the memory 104 and indirectly (432A) via the cache 220. Further details regarding updating the memory bandwidth usage state 402 associated with the first client 224A in response to a read request were discussed above with reference to FIG.
[0046]
[0058] Further, in some implementations, in response to each write request 404 issued from the respective partition ID of the first client 224A for the respective memory block 222, the respective memory block 222 updates (432B) the respective memory bandwidth usage state 402 associated with the respective memory block 222 via the cache 220. There is no direct update of the respective memory bandwidth usage state 402 for the write request 404. In some implementations, the plurality of memory blocks 222 are configured to receive data access requests sent from one or more processing clusters 202 for the cache 220 that are not satisfied by the cache 220. Further details regarding updating the memory bandwidth usage state 402 associated with the first client 224A in response to a write request are discussed above with reference to FIG. 3B.
[0047]
[0059] In some implementations, each of the memory bandwidth usage states 402 associated with a memory block 222 is provided by the respective memory block 222 as a multi-bit state number. The usage level 406 is determined by determining how many of the respective multi-bit state numbers of the memory bandwidth usage states 402 are equal to a predefined value. For example, each memory bandwidth usage state 402 of the respective memory block 222 has two bits, and the usage level 406 is determined based on how many of the memory bandwidth usage states 402 of the memory block 222 are equal to "11". In some implementations, each of the memory bandwidth usage states 402 associated with a memory block 222 is a flag having one of two predefined values (e.g., "Y", "N").
[0048]
[0060] 5A and 5B illustrate exemplary processes 500 and 550 implemented by the memory 104 to track memory bandwidth usage states 402 and current congestion levels 504 of individual memory blocks 222 of the memory 104 according to some implementations. The memory 104 includes a plurality of memory blocks 222. The memory controller 110 is coupled to the memory blocks 222 to manage data access requests received by the memory 104 and to track memory bandwidth usage states 402 and current congestion levels 504 of the memory 104. The memory bandwidth usage states 402 (i.e., SN[0:1]) are associated with respective partition IDs of the first client 224A and indicate at least how much of the memory access bandwidth allocated to the respective partition IDs is used to access the first memory block 222A, i.e., the average data access level of the partition ID of the first client 224A to the first memory block 222A. The current congestion level 504 (i.e., SN[2]) indicates whether a second total number MCQ of data access requests waiting in the second request queue 510 of the memory 104 exceeds a second predefined portion (e.g., 75%) of the external memory capacity.
[0049]
[0061] The memory controller 110 determines that a set of data access requests issued by the processing cluster 202 is associated with the first memory block 222A, and the first memory block 222A receives the set of data access requests. The set of data access requests is associated with a plurality of clients 224, where a resource including a storage capacity of the first memory block 222A is partitioned into a plurality of resource portions to be utilized by the plurality of clients 224. Each resource portion is assigned to a respective client 224 and has a respective partition ID of the respective client 224. For the first client 224A, a subset of data access requests is identified as being associated with the respective ID of the first client 224A to access the first memory block 222A. For the respective partition ID of the first client 224A, one of the memory bandwidth usage states 402 of the first memory block is tracked. In response to each set of data access requests, the memory controller 110 reports to one or more processing clusters 202, for the first memory block 222A, a memory bandwidth usage status 402 associated with the respective partition ID of the first client 224A.
[0050]
[0062] The memory controller 110 maintains a memory block usage window 506 for each partition ID, including that of the first client 224A, where the memory block usage window 506 corresponds to a number of immediately consecutive clock cycles. In the memory block usage window 506, a third number of bytes in the first memory block 222A are accessed by the respective partition ID of the first client 224A during a second number of clock cycles. Upon receiving each data access request associated with the respective partition ID of the first client 224A, the memory controller 110 determines a total number of bytes (i.e., window bytes) processed for the first client 224A in the first memory block 222A during the memory block usage window 506. The window 506 includes a historical number of clock cycles, for example, equal to 16×128 clock cycles. This total number of bytes (i.e., window bytes) represents the average data access level of the partition ID of the first client 224A to the memory block 222 within the window 506 and is compared to the memory access bandwidth allocated to each partition ID to access the memory block 222 to determine a memory bandwidth usage state 402 indicative of how much of the memory access bandwidth allocated to each partition ID to access the memory block 222 is used, i.e., the average data access level of the first client 224A to the first memory block 222A.
[0051]
[0063] In some implementations, the memory bandwidth usage state 402 associated with each partition ID of the first client 224A is represented by a second multi-bit state number SN, e.g., two bits of the 3b state number SN[0:2] or the 2b state number SN[0:1]. If a portion of the memory access bandwidth allocated to each partition ID of the first client 224A to access the first memory block 222A is used and satisfies the first usage condition UC1, the 2b state number SN[0:1] is equal to "00". If a portion of the memory access bandwidth allocated to each partition ID of the first client 224A to access the first memory block 222A is used and satisfies the second usage condition UC2, the 2b state number SN[0:1] is equal to "01". If a portion of the memory access bandwidth allocated to each partition ID of the first client 224A to access the first memory block 222A is used and satisfies the third usage condition UC3, the 2b state number SN[0:1] is equal to "10". If the used portion of the memory access bandwidth allocated to the respective partition ID of the first client 224A to access the memory block 222 satisfies the fourth usage condition UC4 (e.g., the used portion is greater than 75% of the allocated memory access bandwidth), the 2b state number SN[0:1] is equal to "11". Thus, the magnitude of the second multi-bit state number SN[0:1] increases with how much of the memory access bandwidth allocated to the respective partition ID is used to access the memory block, and the memory bandwidth usage state 402 and the average data access level of the first client 224A to the first memory block 222A also increase therewith. In some embodiments, the usage conditions UC1, UC2, UC3, and UC4 are mutually exclusive.
[0052]
[0064] Alternatively, in some implementations, the memory bandwidth usage state (e.g., 2b state number SN[0:1]) associated with each partition ID of the first client 224A is also tracked based on an alternative current congestion level of the memory block 222 and / or based on whether a predefined memory access bandwidth is enforced (i.e., whether HardLimit=1). The memory controller 110 monitors a second total number of data access requests MCQ waiting in the second request queue 510 of the memory 104 and an alternative current congestion level indicative of whether the second total number of data access MCQ requests exceeds an alternative predefined portion of the external memory capacity.
[0053]
[0065] In some implementations, the 2b state number SN[0:1] of the first memory block 222A is equal to "11" under two conditions. In particular, in the first condition, the allocation of the first memory block 222A to the first client 224A is substantially used and the memory 104 is too busy overall. The 2b state number SN[0:1] is equal to "11" when (a) more than 75% of the memory access bandwidth allocated to the respective partition ID of the first client 224A to access the memory block 222 is used, and (b) an alternative current congestion level indicates that the second total number of data access requests MCQ exceeds an alternative predefined portion (e.g., x%, where x is possibly equal to 85) of the external memory capacity. In the second condition, the allocation of the first memory block 222A to the first client 224A is substantially used and the allocation is strictly enforced. 2b state number SN[0:1] is equal to "11" when (a) more than 75% of the memory access bandwidth allocated to the respective partition ID of the first client 224A to access the memory block 222 is used, i.e., when the average data access level to the memory block exceeds a predefined threshold portion (100%), and (b) the predefined memory access bandwidth is enforced (i.e., HardLimit=1). In other words, the memory bandwidth usage state 402 is set to a predefined value associated with a high usage state pursuant to (a) a determination that the average data access level of the first client to the first memory block 222A has exceeded a predefined threshold portion of the predefined memory access bandwidth and (b) a determination that the predefined memory access bandwidth is enforced or that the alternative current congestion level of the memory 104 is high.
[0054]
[0066] In some implementations, the memory controller 110 monitors a second total number MCQ of data access requests waiting in a second request queue 510 of the memory 104, which possibly includes requests for other partition IDs associated with other clients 224. The current congestion level 504 of the memory 104 indicates whether the second total number MCQ of data access requests exceeds a second predefined portion (e.g., 75%) of the external memory capacity of the memory 104 including the memory block 222. In some implementations, the current congestion level 504 of the memory 104 is represented by bit SN[2] of the second multi-bit status number. In some implementations, the second current congestion level 504 of the memory 104 is used to control throttling of prefetch requests. In some implementations, the second current congestion level 504 of the memory 104 including multiple memory blocks 222 is used to control the quality of prefetch requests of one or more processing clusters. Further details regarding application of the current congestion level 504 of the memory 104 are discussed below with reference to FIGS.
[0055]
[0067] FIG. 6A illustrates an exemplary process 600 implemented by the cache 220 to track memory bandwidth usage status 402 and current congestion level 504 of each memory block 222 of the memory 104 according to some implementations. The cache 220 is coupled to one or more processing clusters 202 and the memory 104 including multiple memory blocks 222. The cache 220 forwards multiple data access requests associated with multiple clients 224 from the processing clusters 202 to multiple memory blocks 222 of the memory 104. Given that the resource is partitioned into multiple resource portions to be utilized by the multiple clients 224, each resource portion is assigned to a respective client 224 and has a respective partition identifier (ID). The cache 220 tracks the memory bandwidth usage status 402 of all clients 224 for accessing all memory blocks 222 of the memory 104 and the current congestion level of the memory 104. For convenience, the description of the process 600 focuses on a first client 224A associated with a respective resource portion having a respective partition ID.
[0056]
[0068] In response to each of the subset of data access requests (e.g., all write requests and a subset of read requests associated with the first client 224A), the cache 220 receives a memory bandwidth usage state 402 (e.g., SN[0:1]) and a current congestion level 504 (e.g., SN[2]) of the memory 104. The cache 220 thereby tracks a plurality of memory bandwidth usage states 402 corresponding to the memory blocks 222 for the first client 224A. Each memory bandwidth usage state is associated with a respective memory block 222 and indicates at least how much (e.g., 75%) of the memory access bandwidth allocated to the respective partition ID is used to access the respective memory block 222. In some implementations, each memory bandwidth usage state includes a second multi-bit state number SN (e.g., “11”, “00”, “10”, and “01”) received from the respective memory block 222 and is translated into a flag stored in a first single bit (e.g., HN[0]) of the first multi-bit state number HN. For example, in some implementations, for each memory block 222, if the respective memory bandwidth usage associated with the first client 224A is equal to "11", then HN[0] is equal to "1", otherwise HN[0] is equal to "0". The cache 220 also tracks the current congestion level 504 of the memory 104 (e.g., SN[2]) and is translated into a second single bit (e.g., HN[1]) of the first multi-bit state number HN. In some implementations, the cache 220 maintains a record 602 of the most recently updated memory bandwidth usage state 402 of each memory block 222 (e.g., at HN[0]) and the current congestion level 504 of the memory 104 (e.g., at HN[1]) in association with the first client 224A.
[0057]
[0069] In response to each of the subset of data access requests forwarded by the cache 220 to the memory block 222 for the first client 224A, the cache 220 receives updates of the record 602 regarding the memory bandwidth usage state 402 and / or the current congestion level 504 of the memory 104 and reports them to the processing cluster 202 that made the respective data request. In some implementations, the cache 220 receives updates of the memory bandwidth usage state 402 and / or the current congestion level 504 of the memory block 222 of the memory 104 and reports them to the processing cluster 202 in response to each data access request and regardless of whether the data access request is a read request or a write request. In some implementations, the cache 220 receives updates of the memory bandwidth usage state 402 and / or the current congestion level 504 of the memory 104 and reports them to the processing cluster 202 in response to each write request only.
[0058]
[0070] 6B illustrates an exemplary process 650 implemented by the cache 220 to track its own current congestion level 604 according to some implementations. The cache 220 monitors a first total number HNQ of data access requests waiting in a first request queue 610 associated with the cache 220, which may possibly include requests for other partition IDs than the respective partition ID of the first client 224A. The current congestion level 604 of the cache 220 indicates whether the first total number HNQ of data access requests exceeds a first predefined portion (e.g., c%, where c may possibly be equal to 75) of the system cache capacity of the cache 220. In some implementations, the current congestion level 604 is represented by bit HN[2] of a first multi-bit state number HN. In some implementations, the first current congestion level 604 of the cache 220 is used to control throttling of prefetch requests. In some implementations, the first current congestion level 604 of the cache 220 is used to control the quality of prefetch requests of one or more processing clusters 202 .
[0059]
[0071] In response to each of the subset of data access requests forwarded by the cache 220 to the memory block 222, the cache 220 reports a first current congestion level 604 to the processing cluster 202 that made the respective data request together with the memory bandwidth usage state 402 and / or the current congestion level 504 of the respective memory block 222 of the memory 104. In some implementations, the processing cluster 202 determines whether the first current congestion level 604 satisfies a throttling condition. In accordance with the determination that the first current congestion level 604 satisfies a throttling condition, the processing cluster 202 throttles the prefetch requests from the resource portions, i.e., the prefetch requests are not entered into the one or more request queues 214 of the processing cluster 202. In some implementations, in accordance with a determination that the first and second current congestion levels 604 and 504 satisfy the prefetch control condition, the controller 216 of the processing cluster 202 selects a first subset of prefetch requests having a quality that exceeds a threshold quality corresponding to the prefetch control condition, includes the subset of prefetch requests in the memory access request queue 214, and excludes a second subset of prefetch requests having a quality that does not exceed the threshold quality from the one or more request queues 214. Further details regarding application of the current congestion level 604 of the cache 220 are discussed below with reference to Figures 8 and 9.
[0060]
[0072] 6C illustrates another example process 680 implemented by the cache 220 to track the memory bandwidth usage state 402 of each memory block, the current congestion level 504, and the current congestion level 604 of the cache 220 itself, according to some implementations. The cache 220 tracks the memory bandwidth usage state 402 of all clients 224 for accessing all memory blocks 222 of the memory 104, and the current congestion level of the memory 104. For convenience, the description of the process 680 focuses on a first client 224A associated with a respective resource portion having a respective partition ID.
[0061]
[0073] In response to each of the subset of data access requests (e.g., all write requests and a subset of read requests associated with the first client 224A), the cache 220 receives a memory bandwidth usage state 402 (e.g., SN[0:1]) and a current congestion level 504 (e.g., SN[2]) of the memory 104. The cache 220 thereby tracks a plurality of memory bandwidth usage states 402 corresponding to the memory blocks 222 for the first client 224A. Each memory bandwidth usage state is associated with a respective memory block 222 and indicates at least how much (e.g., 75%) of the memory access bandwidth allocated to the respective partition ID is used to access the respective memory block 222. In some implementations, each memory bandwidth usage state 402 (e.g., SN[0:1]) includes a second multi-bit state number SN (e.g., “11”, “00”, “10”, and “01”) received from the respective memory block 222 and is converted to a flag stored in a first single bit (e.g., HN[2]) of the first multi-bit state number HN. For example, in some implementations, for each memory block 222, if the respective memory bandwidth usage 402 associated with the first client 224A is equal to “11”, then HN[2] is equal to “1”, regardless of whether the current congestion level 504 of the memory 104 (e.g., SN[2]) is “0” or “1”. Conversely, if the respective memory bandwidth usage 402 associated with the first client 224A is not equal to “11”, then HN[2] is equal to “0”. For the first client 402, the memory bandwidth usage state 402 of the memory block 222 is provided to the controller 216 via a first single bit (e.g., HN[2]) of the first multi-bit state number HN, which is further applied by the controller 216 to control requests of the first client 402 to access data stored in the memory block 222.
[0062]
[0074] In some implementations, the first multi-bit status number HN further includes two additional bits HN[0] and HN[1]. The cache 220 monitors a first total number HNQ of data access requests waiting in a first request queue 610 associated with the cache 220, which may possibly include requests for other partition IDs than the respective partition ID of the first client 224A. A current congestion level 604 of the cache 220 is generated based on the first total number HNQ of data access requests and indicates whether the first total number HNQ of data access requests exceeds a first predefined portion (e.g., c%, where c may possibly be equal to 75) of the system cache capacity of the cache 220. In some implementations, this current congestion level 604 of the cache 220 and the current congestion level 504 of the memory 104 (e.g., SN[2]) are represented by two additional bits HN[0] and HN[1] of the first multi-bit state number HN. In some implementations, the first current congestion level 604 of the cache 220 and / or the second current congestion level 504 of the memory 104 (e.g., SN[2]) are used to control the throttling of prefetch requests. In some implementations, the first current congestion level 604 of the cache 220 and / or the second current congestion level 504 of the memory 104 (e.g., SN[2]) are used to control the quality of prefetch requests of one or more processing clusters 202. In other words, the cache 220 returns a first multi-bit state number HN, including HN[0:1], to the controller 216, which uses HN[0:1] to control the throttling and / or quality of prefetch requests. Further details regarding the application of the current congestion level 604 of the cache 220 are discussed below with reference to Figures 8 and 9.
[0063]
[0075] 7A, 7B, and 7C are exemplary data structures of data stored in the processing cluster 202, the cache 220, and the memory block 222 for managing data access requests of multiple resource partitions, respectively, according to some implementations. An electronic device (e.g., a server or server system) is configured to provide services to multiple clients 224, and thus the computational and storage resources of the electronic device are partitioned into multiple resource portions to be utilized by the multiple clients 224. Each resource portion is assigned to a respective client 224 and has a respective partition ID associated with the respective client 224. The processing cluster 202 has one or more request queues 214 that store multiple data access requests associated with the multiple clients 224 for the multiple memory blocks 222 of the memory 104. Data structures 700, 740, and 780 are applied to manage the data access requests stored in the one or more request queues 214 of each processing cluster 202.
[0064]
[0076] 7A, for each resource portion having a respective partition ID of a respective client 224 (e.g., a first client 224A), the processing cluster 202 applies a memory block usage table 401 including a plurality of memory bandwidth usage states 402 for a plurality of memory blocks 222 of the memory 104. Each memory bandwidth usage state 402 is uniquely associated with a respective memory block 222 and indicates at least how much of the memory access bandwidth allocated to the respective partition ID is used to access the respective memory block. The processing cluster 202 applies a usage level 406, a credit count 408, and a request issuance threshold 410 to dynamically control data access requests stored in one or more request queues 214 based on the memory bandwidth usage states 402 of the memory blocks 222. In particular, for each resource portion, the usage level 406 is a combination of the memory bandwidth usage states 402, and the credit count 408 is adjusted (e.g., increased or decreased by credit units CU) based on the usage level 406. In response to a determination that the credit count 408 is greater than the request issuance threshold 410, the next data access request 412 associated with the respective partition ID is issued. Conversely, in response to a determination that the credit count 408 is less than or equal to the request issuance threshold 410, the credit count 408 continues to be adjusted until the next data access request 412 can be issued.
[0065]
[0077] A predefined number of clock cycles and one or more usage thresholds (e.g., a high SN for a high usage threshold and a low SN for a low usage threshold) associated with each client 224 are applied to control adjustment of the credit count 408. While a subset of the memory bandwidth usage states 402 are updated after each data request associated with each client 224 is issued, the usage level 406 of each client 224 is not updated until a predefined number of clock cycles have passed. The usage level 406 is compared to one or more usage thresholds to determine whether the credit count 408 is increased by credit units CU, decreased by credit units CU, or remains the same. Such adjustments are implemented periodically every one or more clock cycles until the magnitude of the credit count 408 triggers the issuance of the next data access request 412.
[0066]
[0078] In some implementations, the processing cluster 202 also tracks the current congestion level 504 of the memory 104 and the current congestion level 604 of the cache 220. The controller 216 of the processing cluster maintains a first congestion level history (e.g., history 902 of FIG. 9 ) including the obtained current congestion level 604 of the cache 220 and maintains a second congestion level history (e.g., history 904 of FIG. 9 ) including the current congestion level 504 of the memory 104. In some circumstances, data access requests that are not satisfied by the cache 220 are further sent to the memory 104, and the number of outstanding in-flight requests received by the memory 104 is thus determined based on the extent to which the data access requests sent to the cache 220 are not satisfied by the cache 220. The controller 216 causes the processing cluster 202 to limit prefetch requests from the processing cluster 202 based on at least one of the current congestion level 604 of the cache 220 and the current congestion level 504 of the memory 104. In some implementations, prefetch requests from the processing clusters 202 are limited based on the first congestion level history and / or the second congestion level history. Further details regarding application of system congestion levels to the cache 220 and memory 104 are described below with reference to FIG.
[0067]
[0079] 7B, the cache 220 is coupled between the processing cluster 202 and the memory block 222 of the memory 104. The cache 220 keeps a record 602 of the most recent updated memory bandwidth usage state 402 of every client (e.g., in HN[0]) of each memory block 222 and the current congestion level 504 of the memory 104 (e.g., in HN[1]). The cache 220 stores its own current congestion level 604. The cache 220 has a first request queue 610 and monitors a first total number HNQ of data access requests waiting in the first request queue 610. The current congestion level 604 of the cache 220 indicates whether the first total number HNQ of data access requests exceeds a first predefined portion (e.g., c%) of the system cache capacity of the cache 220.
[0068]
[0080] 7C, the memory block 222 is coupled to both the processing cluster 202 and the cache 220, and receives data access requests of the different clients 224 from the processing cluster 202 via the cache 220. A memory block usage window 506 is tracked for each client 224 in the memory block 222. The total number of bytes (i.e., window bytes) processed during the window 506 for each client 224 (e.g., the first client 224A) is determined and applied to derive an average data access level of each client 224's partition ID to the memory block 222. This average data access level is used to determine each client's memory bandwidth usage state 402, i.e., how much of the memory access bandwidth allocated to the respective partition ID is used to access the memory block 222.
[0069]
[0081] The memory block 222 also tracks the second request queue 510, a second total number MCQ of data access requests waiting in the queue 510, a second predefined portion of the external memory capacity, an alternative predefined portion (e.g., x%) of the external memory capacity, and a current congestion level 504 of the memory 104. The current congestion level 504 indicates whether the second total number MCQ of data access requests waiting in the second request queue 510 exceeds the second predefined portion (e.g., 75%) of the external memory capacity. The throttling of prefetch requests in the processing cluster 202 is controlled in part by the current congestion level 504 of the memory 104. Furthermore, in some implementations, the memory bandwidth usage state 402 of each client is determined in part based on whether the second total number MCQ of data access requests waiting in the second request queue 510 exceeds the alternative predefined portion (e.g., 75%) of the external memory capacity. For example, when the average data access level to this particular memory block 222 and the second total number MCQ of data access requests waiting in the second request queue 510 are both high (e.g., when the average data access level to this particular memory block 222 exceeds a predefined threshold portion (e.g., 100%) of the predefined memory access bandwidth and the second total number MCQ of data access requests exceeds an alternative predefined portion (e.g., 75%) of the external memory capacity), the memory bandwidth usage state 402 is equal to "11".
[0070]
[0082] 8 illustrates an example method 800 for determining a congestion level for controlling cache prefetching in a processing cluster 202 (e.g., the first processing cluster 202-1 of FIG. 2) according to some implementations. In this processing cluster 202, a controller 216 of a cluster cache 212 determines a congestion level of the processing cluster 202 based on the extent to which data access requests sent to the cluster cache 212 from the processors 204 in the processing cluster 202 are not satisfied by the cluster cache 212, and controls prefetch requests from a prefetcher 208 associated with a first respective processor 204-1 in the processing cluster 202. In particular, pursuant to a determination that the congestion level of the processing cluster 202 satisfies a first congestion criterion requiring that the congestion level of the processing cluster 202 exceed a first cluster congestion threshold 802, the controller 216 causes the first respective processor 204-1 of the one or more processors 204 to limit prefetch requests to the cluster cache 212 to prefetch requests of at least a first threshold quality 804. Conversely, pursuant to a determination that the congestion level of the processing cluster 202 does not satisfy the first congestion criterion, the controller 216 refrains (806) from causing one or more processors 204 (including the first respective processor 204-1) to limit prefetch requests to the cluster cache 212 to prefetch requests of at least the first threshold quality 804. In other words, when the congestion level of the processing cluster 202 is below the first cluster congestion threshold 802, the controller 216 does not limit prefetch requests for the processing cluster 202 in the first prefetch throttling mode M1, and when the congestion level of the processing cluster 202 exceeds the cluster congestion threshold 802, the controller 216 causes the first respective processor 204-1 to limit prefetch requests to prefetch requests of at least the first threshold quality 804, i.e., limit prefetch requests to high quality prefetches in the second prefetch throttling mode M2.
[0071]
[0083] In some implementations, following a determination that the congestion level of the processing cluster 202 satisfies a second congestion criterion different from the first congestion criterion, which requires the congestion level of the processing cluster 202 to exceed a second cluster congestion threshold 808 that is above the first cluster congestion threshold 802, the controller 216 causes the first respective processor 204-1 to limit prefetch requests to prefetch requests of at least a second threshold quality 810 that is higher than the first threshold quality 804. In some implementations, when the congestion level of the processing cluster 202 is above a second cluster congestion threshold 808 (e.g., indicating high congestion as opposed to low or moderate congestion), the controller 216 operates at least each processor 204 (e.g., the first each processor 204-1) of the processing cluster 202 in a third prefetch throttling mode M3 in which prefetching is limited to prefetches of at least a second threshold quality 810 (e.g., enabling only prefetches that are at least very high quality prefetches). In contrast, in the first prefetch throttling mode M1, prefetching is not limited, and in the second prefetch throttling mode M2, prefetching is limited to prefetches having a quality between the first threshold quality 804 and the second threshold quality 810 (e.g., enabling prefetches that are at least high quality prefetches).
[0072]
[0084] In some implementations, in accordance with the determination that the congestion level of the processing cluster 202 satisfies the third congestion criterion, the controller 216 causes the first respective processor 204-1 to refrain from sending a prefetch request to the cache altogether (812), e.g., without regard to the quality of the requested prefetch. In some implementations, the third congestion criterion includes (1) a first requirement that the congestion level of the processing cluster 202 is above the cluster congestion threshold 808, and (2) a second requirement that the system congestion level history 822 of the electronic device 200 satisfies the first system congestion condition 816 (e.g., 75% of the system congestion level history is high). The system congestion level history 822 is monitored by the controller 216 based on a system busy level signal (i.e., current congestion level 604) received from the cache 220, thereby indicating the congestion level of the cache 220. For example, the system congestion level history 822 is filled with "H" or "L" based on multiple sampled values of the system busy level signal. The first system congestion condition 816 requires that 75% or more of the system congestion level history 822 be filled with "H" to enable the fourth prefetch throttling mode M4 (i.e., all throttled mode). Conversely, in some implementations, the controller 216 disables and resets the fourth prefetch throttling mode M4 when a second system congestion condition is met, e.g., when 25% or less of the system congestion level history 822 is filled with "H".
[0073]
[0085] In some implementations, the extent to which data access requests sent from the processors 204 in the processing cluster 202 to the cluster cache 212 are not satisfied by the cluster cache 212 is represented by one or more historical congestion levels for the processing cluster 202. The one or more historical congestion levels are maintained in a congestion level history 818 for the processing cluster 202. The congestion level of the processing cluster 202 is determined based on some or all of the one or more historical congestion levels in the congestion level history 818. In one example, each historical congestion level in the congestion level history 818 corresponds to a distinct respective time period and represents the extent to which data access requests were not satisfied by the cache during the respective time period. The historical congestion levels of the processing cluster 202 may be periodically sampled and stored in the congestion level history 818. In some implementations, the respective historical congestion level (or each respective historical congestion level) has a value selected from a predetermined set of congestion level values. For example, if two congestion levels are used, each historical congestion level has a first congestion level value (e.g., “low”) or a second congestion level value (e.g., “high”), for example, defined based on the first cluster congestion threshold 802. In another example, if three congestion levels are used, each historical congestion level has a first congestion level value (e.g., “low”), or a second congestion level value (e.g., “medium”), or a third congestion level value (e.g., “high”), for example, defined based on the cluster congestion thresholds 802 and 808. Those skilled in the art will recognize that any number of congestion levels may be used, and any number of distinct congestion level values may be used accordingly.
[0074]
[0086] In some implementations, the current cluster congestion level 818A of the processing cluster 202 is determined based on a comparison with the congestion thresholds 802 and 808, and is stored in the congestion level history 818, for example, instead of the oldest stored historical congestion level. The congestion level of the processing cluster 202 is determined based on some or all of the congestion level history 818, including the current cluster congestion level 818A of the processing cluster 202. For example, following a determination that the current cluster congestion level 818A (e.g., equal to "high") is greater than the congestion level of the processing cluster 202 (e.g., equal to "medium"), the congestion level of the processing cluster 202 is increased by one level or to the current cluster congestion level 818A. Following a determination that all existing historical congestion levels in the history 818 (e.g., equal to "medium" or "low") are lower than the congestion level of the processing cluster 202 (e.g., equal to "high"), the congestion level of the processing cluster 202 is reduced by one level. Otherwise, the congestion level of the processing cluster 202 does not change. The current cluster congestion level 818 is the most recent cluster congestion level measured based on the cluster congestion thresholds 802 and 808. Alternatively, in some implementations, the first and second cluster congestion thresholds 802 and 808 are applied together with a historical congestion threshold (e.g., 10% of the congestion level history 818). For example, if a portion (e.g., 75%) of the congestion level history 818 is above the first cluster congestion threshold 802 (i.e., has a value of "medium" or "high") and exceeds the historical congestion threshold (e.g., 10%), the congestion level of the processing cluster 202 meets the first congestion criterion.
[0075]
[0087] It should be noted that in some implementations, the congestion level of a processing cluster 202 is determined based on the extent to which data access requests sent to the cluster cache 212 from one or more processors 204 in the processing cluster 202 are not satisfied by the cache 212, without regard to which of the one or more processors 204 sent the data access requests. That said, the congestion level of a processing cluster 202 is determined without regard to the extent to which data access requests from a particular one of the one or more processors 204 are not satisfied by the cluster cache 212.
[0076]
[0088] In some implementations, determining the congestion level of the processing cluster 202 includes comparing the number of data access requests sent to the cluster cache 212 from one or more processors 204 in the processing cluster 202 that are not satisfied by the cluster cache 212 (e.g., also referred to as cache misses) to one or more cache miss thresholds. Each cluster congestion threshold 802 and 808 includes a respective cache miss threshold 802′ or 808′. In some implementations, the number of cache misses by the processing cluster 202 is compared to the one or more cache miss thresholds 802′ or 808′ to determine a cache miss value (e.g., low, medium, high, etc.) to be taken into account when determining the congestion level of the processing cluster 202. For example, if the number of cache misses by the processing cluster 202 is below a first cache miss threshold 802′, then the first cache miss value (e.g., a low value) is taken into account when determining the congestion level of the processing cluster 202. In another example, if the number of cache misses by a processing cluster 202 is above the first cache miss threshold 802′, a second cache miss value (e.g., a medium or high value) is taken into account when determining the congestion level of the processing cluster 202. In yet another example, if the number of cache misses by a processing cluster 202 is above the second cache miss threshold 808′, a third cache miss value (e.g., a high value) is taken into account when determining the congestion level of the processing cluster 202. In some implementations, the cache miss value is taken into account in the context of one or more historical congestion levels in the congestion level history 818 of the processing cluster 202. In one example, the cache miss value defines the historical congestion level stored in the congestion level history 818 of the processing cluster 202.
[0077]
[0089] Further, in some implementations, one or more cache miss thresholds (i.e., cache miss thresholds 802′ and 808′) are determined based on a system congestion level (e.g., 910 in FIG. 9 ) of the electronic device 200. In some implementations, a first set 820 of one or more cache miss thresholds is used according to a determination that the system congestion level is a first congestion value 826, and a different second set 820′ of one or more cache miss thresholds is used according to a determination that the system congestion level is a different second congestion value 828. Additional different sets of one or more cache miss thresholds may be used for any number of different system congestion values, as desired. In some implementations, the second congestion value 828 is lower than the first congestion value 826, and each cache miss threshold 802′ or 808′ is adjusted to a higher value in relation to the second congestion value 828, since a higher amount of cluster congestion may be tolerated when system congestion is low. For example, the first cache miss threshold 802′ is adjusted from 30% to 50% when the system congestion level drops from the first congestion value 826 to the second congestion value 828. On the other hand, the higher the system congestion level, the lower the one or more cache miss thresholds in the set 820 will be, since if the system congestion is already high, a lower amount of cluster congestion (e.g., of a processing cluster 202) may justify throttling than if the system congestion was lower.
[0078]
[0090] In some implementations, multiple data access requests are sent from one or more processors 204 to the cluster cache 212 within a predefined period of time, i.e., all data access requests, including all demand requests and all prefetch requests.
[0079]
[0091] In some implementations, the controller 216 determines that the congestion level of the respective processor 204-1 or 204-N is below a processor congestion threshold 836 that is different from the congestion threshold 802 or 808 used for the cluster cache 212, regardless of the congestion level of the processing cluster 202, and refrains from limiting prefetch requests from the respective processor 204-1 or 204-N to the cluster cache 212. That said, in these implementations, the prefetch requests from the respective processor 204-1 or 204-N are not limited based on the cluster congestion level and the system congestion level when the congestion level of the respective processor is below the processor congestion threshold 836 (e.g., equal to "L"). Conversely, if the congestion level of the respective processor 204-1 or 204-N exceeds the processor congestion threshold 836 (e.g., equal to “H”), prefetch requests from the respective processor 204-1 or 204-N to the cluster cache 212 are limited or throttled based on the congestion level of the processing cluster and the system. The congestion level of the respective processor 204-1 or 204-N is determined, for example, based on the extent to which data access requests sent to the cluster cache 212 from the respective processor 204-1 or 204-N are not satisfied by the cluster cache 212, regardless of whether data access requests sent to the cluster cache 212 from any processor other than the respective processor 204-1 or 204-N are satisfied by the cluster cache 212.
[0080]
[0092] In other words, in some implementations, the first congestion criterion further requires that the congestion level of each processor 204 exceeds the processor congestion threshold 836 in order for the controller 216 to qualify the prefetch request from the respective processor. In some implementations, a decision of whether to qualify a prefetch request from a respective processor based on whether the congestion level of the respective processor exceeds the processor congestion threshold 836 takes precedence over other decisions as to whether to qualify a prefetch request (e.g., with respect to the first congestion criterion, the second congestion criterion, and / or the third congestion criterion regarding the congestion level of the processing cluster 202).
[0081]
[0093] In some implementations, the controller 216 maintains a processor congestion level history 834 to store a historical congestion level of each processor 204. A prefetch request from each processor is limited based on the congestion level of the processor 204 determined based at least in part on the congestion level history 834 of the processor 204. The current congestion level of the processor 204 is recorded and compared to a processor congestion threshold 836, and one of a plurality of values (e.g., “L” and “H”) is determined based on the comparison result and stored as a current congestion level 834A in the congestion level history 834 of the processor 204 (e.g., instead of the oldest cache miss level in the history 834). Following a determination that the current congestion level 834A of the processor 204 indicates a higher congestion level than the congestion level of the processor 202, the congestion level of the processor 202 is increased by one level or to the current congestion level 834A. Following a determination that the overall congestion level history 834 of the processor 204 is lower than the congestion level of the processor 202, the congestion level of the processor 202 is reduced by one level or to a lower congestion level, for example from "H" to "L".
[0082]
[0094] Further, in some implementations, the processor congestion threshold 836 includes a processor cache miss threshold 836'. Determining the congestion level of the processor 204 includes comparing the number of data access requests sent from the respective processor 204 to the cluster cache 212 that are not satisfied by the cluster cache 212 (i.e., cache misses) to the processor cache miss threshold 836. For example, if the number of cache misses for the processor 204 is below the cache miss threshold 836', a first cache miss value (e.g., a low value) is taken into account when determining the congestion level of the processor 204, and if the number of cache misses for the processor 204 is above the cache miss threshold 836', a second cache miss value (e.g., a medium or high value) is taken into account when determining the congestion level of the processor 204. In particular, in some implementations, the current cache miss is determined for the current number of data access requests not satisfied by the cluster cache 212 during the sample duration. The current cache miss is compared with a cache miss threshold 836, and one of a plurality of cache miss values (e.g., “L” and “H”) is determined based on the comparison result and stored as a current cache miss level 834A in the congestion level history 834 of this processor 204 (e.g., in place of the oldest cache miss level in the history 834). In accordance with a determination that the current cache miss level 834A of the processor 204 indicates a higher congestion level than the congestion level of the processor 202, the congestion level of the processor 202 is increased by one level or to the current cache miss level 834A. In accordance with a determination that the congestion level history 834 of the processor 204 indicates a lower congestion level than the congestion level of the processor 202 (e.g., all cache miss levels in the congestion level history 834 are lower than the congestion level of the processor 202), the congestion level of the processor 202 is reduced by one level or to a lower congestion level, for example, from “H” to “L”.
[0083]
[0095] In some implementations, the electronic device 200 includes a second processing cluster 202-M having one or more second processors 206 different from the one or more processors 204 of the processing cluster 202-1. The controller 216-1 limits prefetch requests by the processing cluster 202-1 independent of whether the prefetch requests from the one or more second processors 206 of the second processing cluster 202-M are limited. In some implementations, prefetching by the second processing cluster 202-M is controlled according to any of the methods for controlling prefetching described herein with respect to the processing cluster 202-1. In some implementations, prefetching by the second processing cluster 202-M may indirectly affect prefetching by processing cluster 202-1 by indirectly affecting system congestion, but the prefetching or prefetch throttling of the second processing cluster 202-M is not directly taken into account when determining whether to limit prefetching by processing cluster 202-1.
[0084]
[0096] FIG. 9 illustrates an exemplary method 900 for determining a system congestion level for controlling cache prefetching in an individual processing cluster 202 (e.g., a first processing cluster 202-1) according to some implementations. A data access request of a processor 204 of a processing cluster 202 is sent to a cluster cache 212. If the data access request is not satisfied by the cluster cache 212, it continues to be sent by the processing cluster 202 to a cache 220 shared with one or more other processing clusters. If the data access request is not satisfied by the cache 220, it is further sent to the memory 104. The system congestion level indicates how many data access requests from the processor 204 are sent to the cache 220 or the memory 104. In particular, a first congestion level history 902 and a second congestion level history 904 are maintained by the controller 216. A current congestion level 604 of the cache 220 is obtained based on the number of outstanding in-flight requests received by the cache 220 and stored in the first congestion level history 902. The current congestion level 504 of the memory 104 is obtained based on the number of outstanding in-flight requests received by the memory 104 and stored in a second congestion level history 904. In some implementations, the information of outstanding in-flight requests not satisfied by the cache 220 or the memory 104 is determined based on system busy level signals (i.e., current congestion levels 504 and 604) received from the cache 220 and the memory 104 in response to data access requests sent to the cache 220 and the memory 104, respectively.
[0085]
[0097] The current congestion levels 504 and 604 of the memory 104 and cache 220 are monitored at respective sampling rates that may be equal or different from one another. The first and second congestion level histories 902 and 904 may store up to a respective limited number of historical congestion levels, where the respective limited numbers may be equal or different from one another. In one example, the first and second congestion level histories 902 and 904 track a first integer number of historical congestion levels of the cache 220 and a second integer number of historical congestion levels of the memory 104. The first and second integer numbers may be equal or different from one another, where the first and second integer numbers are equal or different from one another.
[0086]
[0098] In some implementations, the controller 216 is configured to cause the processing cluster 202 to limit prefetch requests from the processing cluster 202 according to a highest throttling level 920 based on the first congestion level history 902 of the cache 220, including the obtained current congestion level 604 of the cache 220. In some circumstances, the highest throttling level 920 is determined without regard to the obtained current congestion level 504 of the memory 104. In some implementations, whether the prefetch requests from the processing cluster 202 are limited according to the highest throttling level 920 is based on the obtained current congestion level 604 of the cache 220, based on the first congestion level history 902 of the cache 220, and / or based on the first congestion level of the cache 220 determined based at least in part on the first congestion level history 902 of the cache 220. For example, the highest throttling level 920 may be determined with respect to the first system congestion condition 816 (e.g., at least a predefined percentage of the first congestion level history 902 is equal to “H”). In some implementations, congestion of the cache 220, rather than congestion of the memory 104, determines whether a prefetch request from the processing cluster 202 is limited according to the highest throttling level 920. Further, in some implementations, the controller 216 is configured to cause the processing cluster 202 to limit the prefetch request according to the highest throttling level 920 based on the congestion levels of both the processing cluster 202 and the cache 220. For example, when the congestion level of the processing cluster 202 exceeds the cluster congestion threshold 808 and the first congestion level history 902 of the cache 220 satisfies the first system congestion condition 816, the highest throttling level 920 is applied to limit prefetching. In some implementations, the highest throttling level 920 corresponds to an all-throttling mode M4 (812) where prefetching is not enabled.
[0087]
[0099] Further, in some implementations, the controller 216 is configured to cause the processing cluster 202 to limit prefetch requests from the processing cluster 202 according to a highest throttling level 920 based on the first congestion level history 902 of the cache 220, e.g., based on a subset of the first congestion level history 902 and / or the second congestion level history 904. The subset of the first congestion level history 902 includes history 902 with fewer than all or all congestion levels stored. In one example, the controller 216 causes the processing cluster 202 to limit prefetch requests from the processing cluster 202 based on one or more most recently determined and recorded congestion levels of the cache 220. In some implementations, the subset of the first congestion level history 902 has the same number of recorded historical congestion levels (e.g., the same number of samples or entries) as the second congestion level history 904.
[0088]
[0100] In some implementations, the controller 216 is configured to cause the processing cluster 202 to limit prefetch requests from the processing cluster 202 according to the highest throttling level 920 to activate the highest throttling level 920, e.g., based on a determination that the first congestion level history 902 includes more than a first threshold number of determined congestion levels indicating a respective congestion level of the cache 220 (e.g., a high congestion level "H" above a system congestion threshold). For example, if the first congestion level history 902 (or a subset of the first congestion level history 902) includes more than a first threshold number (or alternatively, a first threshold percentage) of instances in which a high congestion level (e.g., "H") was recorded for the cache 220, the highest throttling level 920 is activated.
[0089]
[0101] In some implementations, the controller 216 is configured to refrain from causing the processing cluster 202 to qualify prefetch requests from the processing cluster 202 according to the highest throttling level 920 to deactivate the highest throttling level 920, for example, based on a determination that the first congestion level history 902 includes determined congestion levels less than a second threshold number indicating a respective congestion level of the cache 220 (e.g., a high congestion level "H" above a system congestion threshold). For example, if the first congestion level history 902 (or a subset of the first congestion level history 902) includes instances where a high congestion level (e.g., "H") was recorded for the cache 220 less than the second threshold number (or alternatively, a second threshold percentage), the highest throttling level 920 is deactivated. In some implementations, the first threshold number is the same as the second threshold number (or alternatively, the first threshold percentage is the same as the second threshold percentage). In some implementations, the first threshold number is different (e.g., greater than) the second threshold number (or alternatively, the first threshold percentage is different from the second threshold percentage). In one example, the first threshold percentage and the second threshold percentage are both 50%. In another example, the first threshold percentage is 75% and the second threshold percentage is 25%.
[0090]
[0102] In some implementations, throttling prefetch requests from a processing cluster 202 according to the highest throttling level 920 includes throttling all prefetch requests from a processing cluster 202, for example, all in throttling mode M4. According to the highest throttling level 920, prefetch requests from a processing cluster 202 are not enabled.
[0091]
[0103] In some implementations, the controller 216 determines a first congestion level of the cache 220 and a second congestion level of the memory 104. In accordance with a determination that the obtained current congestion level 604 of the cache 220 indicates a congestion level higher than the first congestion level, the controller 216 increases the first congestion level, for example, to the next higher level in the set of possible congestion levels. Conversely, in accordance with a determination that the first congestion level history 902 indicates a congestion level lower than the first congestion level (e.g., the entire first congestion level history 902 is lower than the first congestion level), the controller 216 decreases the first congestion level. For example, in accordance with a determination that the entries in the first congestion level history 902 do not indicate a congestion level higher than the current value of the first congestion level, the controller 216 decreases the first congestion level, for example, to the next lower level in the set of possible congestion levels. Similarly, in some implementations, in response to a determination that the retrieved current congestion level 504 of the memory 104 indicates a higher congestion level than the second congestion level (e.g., its current value), the controller 216 increases the second congestion level, e.g., to the next higher level in the set of possible congestion levels. In response to a determination that the second congestion level history 904 indicates a lower congestion level than the second congestion level (e.g., the entire second congestion level history 904 is lower than the second congestion level), the controller 216 decreases the second congestion level. For example, in some implementations, in response to a determination that the entries in the second congestion level history 904 do not indicate a higher congestion level than the current value of the second congestion level, the controller 216 decreases the second congestion level, e.g., to the next lower level in the set of possible congestion levels. Thus, the controller 216 causes the processing cluster 202 to limit a prefetch request from the processing cluster 202 based on the first congestion level and the second congestion level, which are taken into account when determining whether to limit the prefetch request according to a respective throttling level below the highest throttling level.
[0092]
[0104] In some implementations, the first system congestion level 906 is determined based on the obtained current congestion level 604 of the cache 220, determined based on the first congestion level history 902 of the cache 220, and / or determined based on a first congestion level of the cache 220 determined based at least in part on the first congestion level history 902 of the cache 220. The second system congestion level 908 is determined based on the obtained current congestion level 504 of the memory 104, determined based on the second congestion level history 904 of the memory 104, and / or determined based on a second congestion level of the memory 104 determined based at least in part on the second congestion level history 904 of the memory 104. The congestion levels 906 and 908 are combined to generate a composite system congestion level 910 having two or more congestion values, such as a first congestion value 826 and a second congestion value 828, that are applied to determine different cache miss thresholds (i.e., cache miss thresholds 802′ and 808′). In some implementations, the composite system congestion level 910 is equal to the greater of the congestion level 906 of the cache 220 and the congestion level 908 of the memory 104. For example, if the congestion level 906 is “L” and the congestion level 908 is “H”, then the composite system congestion level 910 is “H”. If the congestion level 906 is “H” and the congestion level 908 is “L”, then the composite system congestion level 910 is also “H”.
[0093]
[0105] It should be understood that the particular order described for the operations in Figures 8 and 9 is merely exemplary and is not intended to indicate that the described order is the only order in which the operations may be performed. Those skilled in the art will recognize various ways to reorder the operations described herein. Additionally, it should be noted that other process details described herein with respect to methods 800 and 900 (e.g., Figures 8 and 9) may also be applied in an interchangeable manner. For the sake of brevity, these details will not be repeated here.
[0094]
[0106] 10 is a flowchart of a method 1000 of managing memory access to a memory 104 by an electronic device, according to some implementations. The electronic device includes one or more processing clusters 202 and a plurality of memory blocks 222 of the memory 104 (1002). Each processing cluster 202 includes one or more respective processors 204 and is coupled to at least one of the memory blocks 222. In some embodiments, each processing cluster 202 has a controller 216 configured to implement the method 1000. In some embodiments, the electronic device includes a non-transitory computer-readable medium having instructions stored thereon that, when executed by the controller 216 of the electronic device, cause the controller to implement the method 1000.
[0095]
[0107] According to the method 1000, the electronic device partitions (1004) the resources of the electronic device into a plurality of resource portions to be utilized by a plurality of clients. Each resource portion is assigned to a respective client and has a respective ID. The electronic device receives (1006) a plurality of data access requests associated with a plurality of clients 224 for a plurality of memory blocks 222. In some implementations, the data access requests include both demand requests and prefetch requests. For each resource portion having a respective partition ID (1008), each processing cluster 202 tracks (1010) a plurality of memory bandwidth usage states 402 corresponding to the memory block 222. Each memory bandwidth usage state 402 is associated (1012) with a respective memory block and indicates at least how much of the memory access bandwidth assigned to the respective partition ID is used to access the respective memory block 222. The processing cluster 202 determines (1014) a usage level 406 associated with each partition ID from the plurality of memory bandwidth usage states 402, adjusts (1016) a credit count 408 based on the usage level 406, compares (1018) the adjusted credit count 408 to a request issuance threshold 410, and issues (1020) a next data access request 412 associated with the respective partition ID in the memory access request queue 214 in accordance with a determination that the credit count is greater than the request issuance threshold. In some circumstances, for each resource portion having a respective partition ID, in accordance with a determination that the credit count 408 is less than the request issuance threshold 410, the processing cluster 202 suspends issuing any data access requests from the memory access request queue 214 of the respective partition ID until the credit count 408 is adjusted to be greater than the request issuance threshold 410.
[0096]
[0108] In some implementations, for each resource portion having a respective partition ID, the processing cluster 202 updates one or more of the plurality of memory bandwidth usage states 402 in response to a previous data access request (e.g., request 404A) issued immediately prior to the next data access request 412. After a predefined number of clock cycles following the update of one or more of the plurality of memory bandwidth usage states, a usage level 406 is determined from the plurality of memory bandwidth usage states 402. The credit count 408 is periodically adjusted and compared to a request issuance threshold 410, for example, within each subsequent clock cycle, until a next data access request is issued after a predefined number of clock cycles following the update of one or more of the plurality of memory bandwidth usage states.
[0097]
[0109] In some implementations, after each of the multiple data access requests is issued, the processing cluster 202 directly or indirectly receives a respective response from each memory block associated with the issued data access request and updates each memory bandwidth usage state 502 corresponding to each memory block 222 associated with the issued data access request.
[0098]
[0110] In some implementations, pursuant to a determination that the usage level is equal to or greater than the high usage threshold, the processing cluster 202 reduces the credit count 408 by the respective credit unit CU corresponding to the respective partition ID. Pursuant to a determination that the usage level is equal to or less than the low usage threshold, the processing cluster 202 increases the credit count 408 by the respective credit unit CU. Pursuant to a determination that the usage level is between the high and low usage thresholds, the processing cluster 202 maintains the credit count 408.
[0099]
[0111] In some implementations, for each resource portion having a respective partition ID, each of the plurality of memory bandwidth usage states 402 includes a respective multi-bit state number. The processing cluster 202 determines how many of the respective multi-bit state numbers of the memory bandwidth usage states are equal to a predefined value (e.g., “11”).
[0100]
[0112] In some implementations, for each resource portion having a respective partition ID, each of the plurality of memory bandwidth usage states 402 is represented by a flag indicating whether the average data access level of the respective memory block exceeds a predefined threshold portion of the predefined memory access bandwidth allocated to the respective partition ID for accessing the respective memory block. Further, in some implementations, for each resource portion having a respective partition ID, the usage level 406 is represented by a total number of memory blocks for which the flag has a first value (e.g., “Y”). Further, in some implementations, for the first memory block 222A, the flag has a first value. For the first memory block 222A, the processing cluster 202 monitors a second total number of data access requests waiting in the second request queue 510 of the plurality of memory blocks. In accordance with a determination that (a) the first average data access level exceeds a first predefined threshold portion of a first predefined memory access bandwidth assigned to the respective partition ID for accessing the first memory block, and (b) the second total number of data access requests MCQ exceeds an alternative predefined portion of the external memory capacity, the processing cluster 202 determines that a flag representing a first memory bandwidth usage state of the first memory block has a first value.
[0101]
[0113] Further, in some implementations, for the first memory block 222A, the flag has a first value (e.g., “Y”). In accordance with the determination for the first memory block 222A that (a) the first average data access level exceeds a first predefined threshold portion of a first predefined memory access bandwidth assigned to a respective partition ID for accessing the first memory block, and (b) the first predefined memory access bandwidth is enforced, the processing cluster 202 determines that the flag representing the first memory bandwidth usage state 402 of the first memory block 222A has a first value (e.g., “Y”).
[0102]
[0114] In some implementations, for each resource portion having a respective partition ID, the processing cluster 202 sends each read or write request of the plurality of data access requests to the respective memory block 222 separately from the memory block 222 via a first memory (e.g., cache 220) associated with one or more processing clusters 202. In response to each read request issued from the respective partition ID for the respective memory block 222, the processing cluster 202 updates the respective memory bandwidth usage state 402 of the respective memory block 222 from the respective memory block 222 with the data item requested by the read request directly or indirectly via the first memory. In response to each write request issued from the respective partition ID for the respective memory block, the processing cluster 202 updates the respective memory bandwidth usage state 402 associated with the respective memory block 222 from the first memory. The plurality of memory blocks are configured to receive data access requests sent from one or more processing clusters 202 to the first memory that are not satisfied by the first memory.
[0103]
[0115] In some implementations, the electronic device further includes a first memory (e.g., cache 220) configured to receive the plurality of data access requests and pass a subset of the unsatisfied data access requests to the memory block 222. The processing cluster 202 obtains a first current congestion level 604 of the first memory indicating whether a first total number of data access requests waiting in a first request queue 610 of the first memory exceeds a first predefined portion of a system cache capacity, and a second current congestion level 504 of the plurality of memory blocks indicating whether a second total number of data access requests waiting in a second request queue 510 of the plurality of memory blocks exceeds a second predefined portion of an external memory capacity. Furthermore, in some implementations, the plurality of data access requests includes a plurality of prefetch requests. According to the determination that the first current congestion level 604 satisfies a throttling condition, the plurality of prefetch requests are throttled from the plurality of resource portions. Furthermore, in some implementations, the plurality of data access requests includes a plurality of prefetch requests. In accordance with a determination that the first and second current congestion levels satisfy the prefetch control condition, the processing cluster 202 selects a first subset of prefetch requests having a quality that exceeds a threshold quality corresponding to the prefetch control condition, includes the subset of prefetch requests in the memory access request queue, and excludes from the memory access request queue 214 a second subset of prefetch requests having a quality that does not exceed the threshold quality.
[0104]
[0116] In some implementations, the electronic device further includes a first memory (e.g., cache 220), and a plurality of memory bandwidth usage states 402 corresponding to the memory blocks 222 are tracked in one or more processing clusters 202. For each resource portion having a respective partition ID, an average data access level of the respective partition ID to the respective memory block 222 is tracked in real time in each memory block 222, and a respective memory bandwidth usage state 402 associated with the respective memory block 222 is determined based on the average data access level. The respective memory bandwidth usage states 402 are reported to the first memory and to the one or more processing clusters 202 in response to data access requests received from the one or more processing clusters 202. The first memory receives the respective memory bandwidth usage states 402 reported by the plurality of memory blocks 222 in response to the plurality of data access requests received from the one or more processing clusters 202.
[0105]
[0117] 11 is a flowchart of a method 1100 for tracking memory bandwidth usage in a first memory (e.g., cache 220) coupled to one or more processing clusters 202 and a plurality of memory blocks 222 according to some implementations. The method 1100 is implemented in the first memory (1102). The first memory is coupled to one or more processing clusters 202 and a plurality of memory blocks 222 in an electronic device. The first memory forwards a plurality of data access requests associated with a plurality of clients 224 to the plurality of memory blocks 222 (1104). A resource of the electronic device is partitioned (1106) into a plurality of resource portions to be utilized by the plurality of clients, each resource portion being assigned to a respective client and having a respective partition ID. For each resource portion having a respective partition ID (1108), the first memory identifies (1110) a subset of data access requests associated with the respective partition ID to access the memory block 222 and tracks (1112) a plurality of memory bandwidth usage states 402 corresponding to the memory block 222. Each memory bandwidth usage status 402 is associated with a respective memory block (1114) and indicates at least how much of the memory access bandwidth allocated to the respective partition ID is used to access the respective memory block. In response to each of the subset of data access requests (1116), the first memory determines (1118) that the respective data access request is to access a corresponding memory block, receives (1120) a memory bandwidth usage status for the corresponding memory block, and reports (1122) the memory bandwidth usage status for the corresponding memory block to one or more processing clusters.
[0106]
[0118] In some implementations, the first memory monitors a first total number of data access HNQ requests waiting in the first request queue 610 of the first memory and determines a first current congestion level 604 (i.e., HN[2]) indicative of whether the first total number of data access requests exceeds a first predefined portion of the system cache capacity. In response to each of the subset of data access requests, the first memory reports the first current congestion level 604 (i.e., HN[2]) to one or more processing clusters 202 along with the memory bandwidth usage status 502 of the corresponding memory block. Furthermore, in some implementations, in one or more processing clusters 202, multiple prefetch requests from multiple resource portions are throttled according to the determination that the first current congestion level 604 (i.e., HN[2]) satisfies a throttling condition.
[0107]
[0119] In some implementations, in response to each of the subset of data access requests, the first memory updates a second current congestion level 504 (i.e., SN[2]) from the corresponding memory block, indicating whether a second total number of data access requests waiting in the second request queue of the plurality of memory blocks exceeds a second predefined portion of the external memory capacity. The first memory reports the second current congestion level 504 (i.e., SN[2]) to one or more processing clusters along with the memory bandwidth usage state 402 and the first current congestion level 604 (i.e., HN[2]) of the corresponding memory block. Further, in some implementations, in accordance with a determination that the first and second current congestion levels 604 and 504 satisfy the prefetch control condition, the one or more processing clusters 202 select a first subset of prefetch requests having a quality that exceeds a threshold quality corresponding to the prefetch control condition and include the subset of prefetch requests in the memory access request queue 214, and exclude a second subset of prefetch requests having a quality that does not exceed the threshold quality from the memory access request queue 214.
[0108]
[0120] In some implementations, each memory bandwidth usage state 402 associated with a respective memory block 222 includes a respective flag configured to enable the respective memory block 222 in accordance with (a) a determination that the average data access level to the respective memory block 222 exceeds a predefined threshold portion of a predefined memory access bandwidth and (b) a determination that the predefined memory access bandwidth is enforced or an alternative congestion level of the memory block is high.
[0109]
[0121] FIG. 12 is a flowchart of a method 1200 for tracking memory bandwidth usage of a memory block 222 of a memory system according to some implementations. The memory system includes a memory controller (e.g., memory controller 110) and a memory block 222. The memory block 222 is coupled to one or more processing clusters 202 via a first memory (e.g., cache 220) in an electronic device. The method is implemented in the memory system (1202). The memory system receives (1204) a set of data access requests associated with a plurality of clients 224 for the memory block 222. The resource is partitioned (1206) into a plurality of resource portions to be utilized by the plurality of clients 224, each resource portion being assigned to a respective client and having a respective partition ID. For each resource portion having a respective partition ID (1208), the memory system (particularly the memory controller 110) identifies (1210) a subset of data access requests associated with the respective IDs to access the memory block 222, and tracks (1212) the memory bandwidth usage state 402 associated with the respective partition IDs. The memory bandwidth usage status 402 indicates at least how much of the memory access bandwidth allocated to each partition ID is used to access the memory blocks (1214). In response to each of the sets of data access requests, the memory system reports the memory bandwidth usage status to one or more processing clusters 202 (1216).
[0110]
[0122] In some implementations, in response to receiving a read request, the memory system reports the memory bandwidth usage state 402, either directly on the data item requested by the read request or indirectly via the first memory (e.g., cache 220), to one or more processing clusters 202. In response to receiving a write request, the memory system reports the memory bandwidth usage state 402 of the memory block 222 to one or more processing clusters 202 indirectly via the first memory.
[0111]
[0123] In some implementations, the memory bandwidth usage state 402 associated with each partition ID is also tracked based on an alternative current congestion level of the memory block 222 and / or based on whether a predefined memory access bandwidth is enforced. The alternative current congestion level of the memory block 222 indicates whether the second total number of data access requests MCQ exceeds an alternative predefined portion of the external memory capacity.
[0112]
[0124] In some implementations, for each partition ID, the memory system determines whether the average data access level to the memory block 222 has exceeded a predefined threshold portion of a predefined memory access bandwidth allocated to each partition ID for accessing the memory block 222. Additionally, in some implementations, the memory system monitors a second total number of data access requests waiting in a second request queue 510 of the memory system and determines an alternative current congestion level indicative of whether the second total number of data access requests exceeds an alternative predefined portion (e.g., x%) of the external memory capacity. Additionally, in some implementations, the memory system determines a second current congestion level 504 of the memory system indicative of whether the second total number of data access requests MCQ exceeds a second predefined portion of the external memory capacity. The second current congestion level 504 is used to control throttling or quality of prefetch requests of one or more processing clusters. In some cases, the second and alternative predefined portions are distinct or equal to each other. Also, in some embodiments, the memory bandwidth usage status 402 includes a flag configured to indicate a heavy memory bandwidth usage status. The memory system enables a flag in accordance with (a) a determination that the average data access level to memory block 222 has exceeded a predefined threshold portion of a predefined memory access bandwidth and (b) a determination that the predefined memory access bandwidth is enforced or alternatively that a current congestion level of memory block 222 is high.
[0113]
[0125] In some implementations, for each partition ID, the memory bandwidth usage state 402 associated with the respective partition ID includes a multi-bit state number (e.g., SN[0:1]), the magnitude of the multi-bit state number (e.g., SN[0:1]) increasing with how much of the memory access bandwidth allocated to the respective partition ID is used to access memory blocks 222.
[0114]
[0126] It should be understood that the particular order described for the operations in Figures 10-12 is merely exemplary and is not intended to indicate that the described order is the only order in which the operations may be performed. Those skilled in the art will recognize various ways to reorder the operations described herein. Additionally, it should be noted that other process details described herein with respect to methods 1000, 1100, and 1200 (e.g., Figures 10-12) may also be applied in an interchangeable manner. For the sake of brevity, these details will not be repeated here.
[0115]
[0127] Example implementations are described in at least the following numbered clauses:
[0116]
[0128] Clause 1. A method for managing memory access in an electronic device including one or more processing clusters and a plurality of memory blocks, each processing cluster including one or more respective processors and coupled to at least one of the memory blocks, comprising: partitioning resources of the electronic device into a plurality of resource portions to be utilized by a plurality of clients, receiving a plurality of data access requests associated with the plurality of clients for a plurality of memory blocks, each resource portion assigned to a respective client and having a respective partition identifier (ID), for each resource portion having a respective partition ID, tracking a plurality of memory bandwidth usage states corresponding to the memory block, determining a usage level associated with each partition ID from the plurality of memory bandwidth usage states, wherein each memory bandwidth usage state is associated with the respective memory block and indicates at least how much of the memory access bandwidth assigned to the respective partition ID is used to access the respective memory block, adjusting a credit count based on the usage level, comparing the adjusted credit count to a request issuance threshold, and issuing a next data access request associated with the respective partition ID in a memory access request queue in accordance with a determination that the credit count is greater than the request issuance threshold.
[0117]
[0129] Clause 2. The method of clause 1, further comprising, for each resource portion having a respective partition ID, pursuant to a determination that the credit count is less than a request issuance threshold, suspending issuance of any data access requests from the memory access request queue of the respective partition ID until the credit count is adjusted to be greater than the request issuance threshold.
[0118]
[0130] Clause 3. The method of clauses 1 or 2, further comprising: for each resource portion having a respective partition ID, updating one or more of a plurality of memory bandwidth usage states in response to a previous data access request issued immediately before a next data access request, wherein a usage level is determined from the plurality of memory bandwidth usage states after a predefined number of clock cycles following the update of the one or more of the plurality of memory bandwidth usage states, and wherein a credit count is periodically adjusted and compared to a request issuance threshold until a next data access request is issued after a predefined number of clock cycles following the update of the one or more of the plurality of memory bandwidth usage states.
[0119]
[0131] Clause 4. The method of any preceding clause, further comprising after each of the plurality of data access requests is issued, directly or indirectly receiving a respective response from a respective memory block associated with the issued data access request, and updating a respective memory bandwidth usage state corresponding to each memory block associated with the issued data access request.
[0120]
[0132] Clause 5. The method of any of the preceding clauses, wherein adjusting the credit count based on the usage level further comprises reducing the credit count by a respective credit unit corresponding to the respective classification ID in accordance with a determination that the usage level is equal to or greater than a high usage threshold, increasing the credit count by a respective credit unit in accordance with a determination that the usage level is equal to or less than a low usage threshold, and maintaining the credit count in accordance with a determination that the usage level is between the high usage threshold and the low usage threshold.
[0121]
[0133] Clause 6. The method of any preceding clause, wherein for each resource portion having a respective partition ID, each of a plurality of memory bandwidth usage states includes a respective multi-bit state number, and determining the level of usage includes determining how many of the respective multi-bit state numbers of the memory bandwidth usage states are equal to a predefined value.
[0122]
[0134] Clause 7. The method of any of the preceding clauses, wherein for each resource portion having a respective partition ID, each of the plurality of memory bandwidth usage states is represented by a flag indicating whether an average data access level for the respective memory block has exceeded a predefined threshold portion of the predefined memory access bandwidth allocated to the respective partition ID for accessing the respective memory block.
[0123]
[0135] Clause 8. The method of clause 7, wherein for each resource portion having a respective partition ID, the usage level is represented by a total number of memory blocks whose flags each have a first value.
[0124]
[0136] Clause 9. The method of clause 8, further comprising: for the first memory block, the flag has a first value; monitoring for the first memory block a second total number of data access requests waiting in a second request queue of the plurality of memory blocks; and determining that the flag representing a first memory bandwidth usage state of the first memory block has the first value in accordance with a determination that (a) the first average data access level exceeds a first predefined threshold portion of a first predefined memory access bandwidth assigned to a respective partition ID for accessing the first memory block, and (b) the second total number of data access requests exceeds an alternative predefined portion of an external memory capacity.
[0125]
[0137] Clause 10. The method of clause 8, further comprising determining, for the first memory block, that the flag has a first value and, for the first memory block, in accordance with a determination that (a) the first average data access level exceeds a first predefined threshold portion of a first predefined memory access bandwidth assigned to a respective partition ID for accessing the first memory block, and (b) the first predefined memory access bandwidth is enforced, the flag representing a first memory bandwidth usage state of the first memory block having a first value.
[0126]
[0138] Clause 11. The method of any preceding clause, wherein for each resource portion having a respective partition ID, tracking a plurality of memory bandwidth usage states further comprises: sending each read or write request of the plurality of data access requests to the respective memory block separately from the memory block via a first memory associated with the one or more processing clusters; updating a respective memory bandwidth usage state of the respective memory block from the respective memory block, either directly with the data item requested by the read request or indirectly via the first memory, in response to each read request issued from the respective partition ID for the respective memory block; and updating a respective memory bandwidth usage state associated with the respective memory block from the first memory in response to each write request issued from the respective partition ID for the respective memory block, wherein the plurality of memory blocks are configured to receive data access requests sent from the one or more processing clusters for the first memory that are not satisfied by the first memory.
[0127]
[0139] Clause 12. The method of any of the preceding clauses, wherein the electronic device further includes a first memory configured to receive a plurality of data access requests and pass a subset of the unsatisfied data access requests to the memory block, the method further comprising obtaining a first current congestion level of the first memory indicative of whether a first total number of data access requests waiting in a first request queue of the first memory exceeds a first predefined portion of a system cache capacity, and obtaining a second current congestion level of the plurality of memory blocks indicative of whether a second total number of data access requests waiting in a second request queue of the plurality of memory blocks exceeds a second predefined portion of an external memory capacity.
[0128]
[0140] Clause 13. The method of clause 12, wherein the plurality of data access requests includes a plurality of prefetch requests, the method further comprising throttling the plurality of prefetch requests from the plurality of resource portions in accordance with a determination that a first current congestion level satisfies a throttling condition.
[0129]
[0141] Clause 14. The method of clause 12, wherein the plurality of data access requests includes a plurality of prefetch requests, the method further comprising: selecting, in accordance with a determination that the first and second current congestion levels satisfy a prefetch control condition, a first subset of the prefetch requests having a quality that exceeds a threshold quality corresponding to the prefetch control condition, including the subset of the prefetch requests in the memory access request queue, and excluding from the memory access request queue a second subset of the prefetch requests having a quality that does not exceed the threshold quality.
[0130]
[0142] Clause 15. The method of any preceding clause, wherein the electronic device further includes a first memory, and wherein a plurality of memory bandwidth usage states corresponding to the memory blocks are tracked in one or more processing clusters, the method further comprising: for each resource portion having a respective partition ID, tracking, in each memory block, an average data access level of the respective partition ID for the respective memory block in real time, determining a respective memory bandwidth usage state associated with the respective memory block based on the average data access level, reporting the respective memory bandwidth usage states to the first memory and to the one or more processing clusters in response to data access requests received from the one or more processing clusters, and receiving, in the first memory, the respective memory bandwidth usage states reported by the plurality of memory blocks in response to the plurality of data access requests received from the one or more processing clusters.
[0131]
[0143] Clause 16. A method for managing memory access, comprising: in a first memory coupled to one or more processing clusters and a plurality of memory blocks in an electronic device, forwarding a plurality of data access requests associated with a plurality of clients to the plurality of memory blocks, wherein a resource of the electronic device is partitioned into a plurality of resource portions to be utilized by the plurality of clients, each resource portion being assigned to a respective client and having a respective partition identifier (ID), for each resource portion having a respective partition ID, identifying a subset of the data access requests associated with the respective partition ID to access the memory block, tracking a plurality of memory bandwidth usage states corresponding to the memory blocks, wherein each memory bandwidth usage state is associated with the respective memory block and indicates at least how much of the memory access bandwidth assigned to the respective partition ID is used to access the respective memory block, in response to each of the subset of data access requests, determining that the respective data access request is to access a corresponding memory block, receiving the memory bandwidth usage state of the corresponding memory block, and reporting the memory bandwidth usage state of the corresponding memory block to the one or more processing clusters.
[0132]
[0144] Clause 17. The method of clause 16, further comprising: monitoring a first total number of data access requests waiting in a first request queue of the first memory to determine a first current congestion level indicative of whether the first total number of data access requests exceeds a first predefined portion of a system cache capacity; and in response to each of the subset of the data access requests, reporting the first current congestion level along with a memory bandwidth usage status of a corresponding memory block to one or more processing clusters.
[0133]
[0145] Clause 18. The method of clause 17, further comprising throttling, in one or more processing clusters, the plurality of prefetch requests from the plurality of resource portions in accordance with a determination that a first current congestion level satisfies a throttling condition.
[0134]
[0146] Clause 19. The method of clause 17 or 18, further comprising: updating, in response to each of the subset of data access requests from a corresponding memory block, a second current congestion level indicative of whether a second total number of data access requests waiting in a second request queue of the plurality of memory blocks exceeds a second predefined portion of the external memory capacity, and reporting the second current congestion level together with a memory bandwidth usage status and the first current congestion level of the corresponding memory block to one or more processing clusters.
[0135]
[0147] Clause 20. The method of clause 19, further comprising: in one or more processing clusters, in accordance with a determination that the first and second current congestion levels satisfy the prefetch control condition, selecting a first subset of prefetch requests having a quality that exceeds a threshold quality corresponding to the prefetch control condition, including the subset of prefetch requests in the memory access request queue, and excluding a second subset of prefetch requests having a quality that does not exceed the threshold quality from the memory access request queue.
[0136]
[0148] Clause 21. The method of any of clauses 16-20, wherein each memory bandwidth usage state associated with a respective memory block includes a respective flag configured to be enabled for the respective memory block pursuant to (a) a determination that an average data access level to the respective memory block has exceeded a predefined threshold portion of the predefined memory access bandwidth, and (b) a determination that the predefined memory access bandwidth is enforced or an alternative congestion level for the memory block is high.
[0137]
[0149] Clause 22. A method for tracking memory usage, in a memory system coupled to one or more processing clusters via a first memory in an electronic device and including a memory block, comprising: receiving a set of data access requests associated with a plurality of clients for the memory block, wherein a resource is partitioned into a plurality of resource portions to be utilized by the plurality of clients, each resource portion being assigned to a respective client and having a respective partition identifier (ID), for each resource portion having a respective partition ID, identifying a subset of data access requests associated with the respective ID to access the memory block, tracking a memory bandwidth usage state associated with the respective partition ID, wherein the memory bandwidth usage state indicates at least how much of the memory access bandwidth assigned to the respective partition ID is used to access the memory block, and reporting the memory bandwidth usage state to the one or more processing clusters in response to each of the set of data access requests.
[0138]
[0150] Clause 23. The method of clause 22, wherein reporting the memory bandwidth usage to the one or more processing clusters further comprises: in response to receiving a read request, reporting the memory bandwidth usage status to the one or more processing clusters either directly on a data item requested by the read request or indirectly via the first memory; and in response to receiving a write request, reporting the memory bandwidth usage status to the one or more processing clusters indirectly via the first memory.
[0139]
[0151] Clause 24. The method of clause 22 or 23, wherein a memory bandwidth usage state associated with each partition ID is also tracked based on an alternative current congestion level of the memory block and / or whether a predefined memory access bandwidth is enforced.
[0140]
[0152] Clause 25. The method of any of clauses 22-24, wherein tracking memory bandwidth usage associated with each partition ID further comprises determining, for each partition ID, whether an average data access level to the memory block exceeds a predefined threshold portion of a predefined memory access bandwidth allocated to the respective partition ID for accessing the memory block.
[0141]
[0153] Clause 26. The method of clause 25, wherein tracking memory bandwidth usage associated with each partition ID further comprises monitoring a second total number of data access requests waiting in a second request queue of the memory system and determining an alternative current congestion level indicative of whether the second total number of data access requests exceeds an alternative predefined portion of the external memory capacity.
[0142]
[0154] Clause 27. The method of clause 26, further comprising determining a second current congestion level indicative of whether a second total number of data access requests exceeds a second predefined portion of the external memory capacity, wherein the second current congestion level is used to control throttling or quality of prefetch requests of one or more processing clusters.
[0143]
[0155] Clause 28. The method of clause 26, wherein the memory bandwidth usage state includes a flag configured to indicate a heavy memory bandwidth usage state, further comprising enabling the flag in accordance with (a) a determination that an average data access level to the memory block has exceeded a predefined threshold portion of the predefined memory access bandwidth and (b) a determination that the predefined memory access bandwidth is enforced or that an alternative current congestion level of the memory block is high.
[0144]
[0156] Clause 29. The method of any of clauses 16-28, wherein for each partition ID, the memory bandwidth usage state associated with the respective partition ID includes a multi-bit state number, the magnitude of the multi-bit state number increasing with how much of the memory access bandwidth allocated to the respective partition ID is used to access memory blocks.
[0145]
[0157] Clause 30. An electronic device comprising one or more processing clusters and a plurality of memory blocks coupled to each processing cluster, wherein each processing cluster includes one or more respective processors and a controller, the controller configured to implement the method of any of clauses 1-29.
[0146]
[0158] Clause 31. A non-transitory computer readable medium having instructions stored thereon that, when executed by a controller of an electronic device, cause the controller to perform the method of any of clauses 1 to 29.
[0147]
[0159] Clause 32. Apparatus for managing memory access in an electronic device comprising one or more processing clusters and a plurality of memory blocks, each processing cluster comprising one or more respective processors and coupled to at least one of the memory blocks, the apparatus comprising means for performing operations of the methods of any of clauses 1 to 15.
[0148]
[0160] Clause 33. Apparatus for managing memory access in a first memory coupled to one or more processing clusters and to a plurality of memory blocks in an electronic device, said apparatus comprising means for performing the operations of the methods of any of clauses 16 to 21.
[0149]
[0161] Clause 34. An apparatus for tracking memory usage in a memory system including a memory block coupled to one or more processing clusters via a first memory in an electronic device, said apparatus comprising means for performing the operations of the method of any of clauses 22 to 29.
[0150]
[0162] The above description has been provided with respect to specific implementations. However, the above exemplary description is not intended to be exhaustive or to be limited to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The implementations have been chosen and described to best explain the disclosed principles and their practical applications, thereby enabling others to best utilize the present disclosure and various implementations with various modifications suited to the particular use contemplated.
[0151]
[0163] The terms used in the description of the various described implementations herein are merely for the purpose of describing the particular implementations and are not intended to be limiting. The singular forms "a," "an," and "the," used in the description of the various described implementations and in the appended claims, are intended to include the plurals unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will further be understood that the terms "includes," "including," "comprises," and / or "comprising," as used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It will further be understood that terms such as "first," "second," and the like, may be used herein to describe various elements, but these elements should not be limited by these terms. These terms are only used to distinguish one element from another.
[0152]
[0164] The term "if" as used herein is to be interpreted to mean "when" or "upon" or "in response to determining" or "in response to detecting" or "following a determination that," depending on the context, as the case may be. Similarly, the phrase "if determined to" or "if a 'stated condition or event' is detected" is to be interpreted to mean "upon determining" or "in response to determining" or "upon detecting the 'stated condition or event'" or "in response to detecting the 'stated condition or event'" or "following a determination that a 'stated condition or event' is detected," depending on the context, as the case may be.
[0153]
[0165] The above description has been set forth in terms of specific embodiments for purposes of explanation. However, the above exemplary description is not intended to be exhaustive or to limit the scope of the claims to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments have been chosen and described in order to best explain the principles of operation and practical applications, thereby enabling others skilled in the art.
[0154] While various figures show some logical steps in a particular order, steps that are not order dependent may be rearranged and other steps may be combined or broken. While some rearrangements or other groupings are specifically mentioned, others will be apparent to those of ordinary skill in the art, and thus the ordering and groupings presented herein are not an exhaustive list of alternatives. Moreover, it should be recognized that the steps may be implemented in hardware, firmware, software, or any combination thereof.
Claims
1. In an electronic device including one or more processing clusters and a plurality of memory blocks, a method for managing memory access, wherein each processing cluster includes one or more respective processors and is coupled to at least one of the memory blocks, dividing the resources of the electronic device into a plurality of resource portions to be utilized by a plurality of clients, wherein each resource portion is assigned to a respective client and has a respective division identifier (ID), receiving a plurality of data access requests associated with the plurality of clients for the plurality of memory blocks, for each resource portion having a respective division ID, tracking a plurality of memory bandwidth usage states corresponding to the memory blocks, wherein each memory bandwidth usage state is associated with a respective memory block and indicates at least how much of the memory access bandwidth assigned to the respective division ID for accessing the respective memory block is used, determining a usage level associated with the respective division ID from the plurality of memory bandwidth usage states, adjusting a credit count based on the usage level, comparing the adjusted credit count with a request issuance threshold, issuing the next data access request associated with the respective division ID in a memory access request queue according to a determination that the credit count is greater than the request issuance threshold, comprising, for each resource portion having a respective division ID, updating one or more of the plurality of memory bandwidth usage states in response to a previous data access request issued immediately before the next data access request, further comprising, after a predefined number of clock cycles following the update of the one or more of the plurality of memory bandwidth usage states, the usage level is determined from the plurality of memory bandwidth usage states, the credit count is periodically adjusted and compared with the request issuance threshold until the next data access request is issued after the predefined number of clock cycles following the update of the one or more of the plurality of memory bandwidth usage states.
2. For each resource portion having the respective segment ID, in accordance with the determination that the credit count is less than the request issuance threshold, interrupting the issuance of any data access requests from the memory access request queue of the respective segment ID until the credit count is adjusted to be greater than the request issuance threshold; The method according to claim 1, further comprising.
3. After each of the plurality of data access requests is issued, directly or indirectly receiving each response from each memory block associated with the issued data access request; and updating the respective memory bandwidth usage status corresponding to each memory block associated with the issued data access request; The method according to claim 1, further comprising.
4. Adjusting the credit count based on the usage level includes: reducing the credit count by each credit unit corresponding to the respective segment ID in accordance with the determination that the usage level is equal to or higher than a high usage threshold; increasing the credit count by each credit unit in accordance with the determination that the usage level is equal to or lower than a low usage threshold; maintaining the credit count in accordance with the determination that the usage level is between the high usage threshold and the low usage threshold; The method according to claim 1, further comprising.
5. For each resource portion having the respective segment ID, each of the plurality of memory bandwidth usage statuses includes a respective multi-bit status number, Determining the usage level includes determining which of the respective multi-bit status numbers of the memory bandwidth usage status are equal to a predefined value. The method according to claim 1.
6. For each resource portion having the respective segment ID, each of the plurality of memory bandwidth usage statuses is represented by a flag indicating whether the average data access level of the respective memory block exceeds a predefined threshold portion of the predefined memory access bandwidth assigned to the respective segment ID for accessing the respective memory block. The method according to claim 1.
7. For each resource portion having the respective segment ID, the usage level is represented by the total number of memory blocks for which the flag has a first value, For a first memory block, the flag has the first value, for the first memory block, monitoring a second total number of data access requests waiting in a second request queue of the plurality of memory blocks; determining that the flag representing the first memory bandwidth usage state of the first memory block has the first value according to the determination that (a) a first average data access level exceeds a first predefined threshold portion of a first predefined memory access bandwidth assigned to the respective segment ID for accessing the first memory block, and (b) the second total number of data access requests exceeds an alternative predefined portion of the external memory capacity; further comprising, or, For a first memory block, the flag has the first value, for the first memory block, determining that the flag representing the first memory bandwidth usage state of the first memory block has the first value according to the determination that (a) a first average data access level exceeds a first predefined threshold portion of a first predefined memory access bandwidth assigned to the respective segment ID for accessing the first memory block, and (b) the first predefined memory access bandwidth is enforced; further comprising, the method according to claim 6.
8. For each resource portion having the respective segment ID, tracking the plurality of memory bandwidth usage states comprises sending each read or write request of the plurality of data access requests for each memory block separately from the memory block via a first memory associated with the one or more processing clusters; In response to each read request issued from the respective section ID for each memory block, directly with the data item requested by the read request or indirectly via the first memory, updating the respective memory bandwidth usage status of the respective memory block from the respective memory block; In response to each write request issued from the respective section ID for each memory block, updating the respective memory bandwidth usage status associated with the respective memory block from the first memory; further comprising; the plurality of memory blocks are configured to receive data access requests sent from the one or more processing clusters to the first memory that are not satisfied by the first memory; The method according to claim 1.
9. The electronic device further includes a first memory configured to receive the plurality of data access requests and pass a subset of the unsatisfied data access requests to the memory block, and the method includes: obtaining a first current congestion level of the first memory indicating whether a first total number of data access requests waiting in a first request queue of the first memory exceeds a first predefined portion of a system cache capacity; obtaining a second current congestion level of the plurality of memory blocks indicating whether a second total number of data access requests waiting in a second request queue of the plurality of memory blocks exceeds a second predefined portion of an external memory capacity; The method according to claim 1, further comprising.
10. The plurality of data access requests include a plurality of prefetch requests, and the method includes: throttling the plurality of prefetch requests from the plurality of resource portions according to a determination that the first current congestion level meets a throttling condition; The method according to claim 9, further comprising.
11. The plurality of data access requests include a plurality of prefetch requests, and the method includes: In accordance with the determination that the first and second current convergence levels satisfy the prefetch control condition, select a first subset of prefetch requests having a quality exceeding a threshold quality corresponding to the prefetch control condition, include the subset of prefetch requests in the memory access request queue, and exclude from the memory access request queue a second subset of prefetch requests having a quality not exceeding the threshold quality. The method according to claim 9, further comprising. **Claim 12** The electronic device further includes a first memory, the plurality of memory bandwidth usage states corresponding to the memory blocks are tracked in the one or more processing clusters, and the method includes, for each resource portion having the respective section ID, In each memory block, track the average data access level of the respective section ID for the respective memory block in real time, determine the respective memory bandwidth usage state associated with the respective memory block based on the average data access level, and report the respective memory bandwidth usage state to the first memory and the one or more processing clusters in response to the data access request received from the one or more processing clusters. In the first memory, receive the respective memory bandwidth usage states reported by the plurality of memory blocks in response to the plurality of data access requests received from the one or more processing clusters. The method according to claim 1, further comprising. **Claim 13** An electronic device, One or more processing clusters, A plurality of memory blocks coupled to each processing cluster, Comprising, Each processing cluster includes one or more respective processors and a controller, and the controller Divides the resources of the electronic device into a plurality of resource portions to be utilized by a plurality of clients, each resource portion being assigned to a respective client and having a respective section identifier (ID). Receive a plurality of data access requests associated with the plurality of clients for the plurality of memory blocks. For each resource portion having the respective section ID, Tracking a plurality of memory bandwidth usage states corresponding to the memory blocks, where each memory bandwidth usage state is associated with a respective memory block and indicates at least how much of the memory access bandwidth assigned to the respective partition ID is used to access the respective memory block. Determining a usage level associated with each of the respective partition IDs from the plurality of memory bandwidth usage states. Adjusting a credit count based on the usage level. Comparing the adjusted credit count with a request issuance threshold. Issuing the next data access request associated with each of the respective partition IDs in a memory access request queue according to a determination that the credit count is greater than the request issuance threshold. Configured to perform. For each resource portion having the respective partition ID, the controller Updating one or more of the plurality of memory bandwidth usage states in response to the previous data access request issued immediately before the next data access request. Further configured to perform. After a predefined number of clock cycles following the update of the one or more of the plurality of memory bandwidth usage states, the usage level is determined from the plurality of memory bandwidth usage states. An electronic device in which the credit count is periodically adjusted and compared with the request issuance threshold until the next data access request is issued after a predefined number of clock cycles following the update of the one or more of the plurality of memory bandwidth usage states.
14. For each resource portion having the respective partition ID, the controller Interrupting issuing any data access requests from the memory access request queue of the respective partition ID until the credit count is adjusted to be greater than the request issuance threshold according to a determination that the credit count is less than the request issuance threshold. The electronic device according to claim 13, further configured to perform.
15. A non-transitory computer-readable medium storing instructions that, when executed by a controller of an electronic device, cause the controller to In an electronic device having one or more processing clusters and a plurality of memory blocks coupled to each processing cluster, where each processing cluster includes one or more respective processors and the controller, Partition the resources of the electronic device into a plurality of resource portions to be utilized by a plurality of clients, where each resource portion is assigned to a respective client and has a respective partition identifier (ID), Receive a plurality of data access requests associated with the plurality of clients for the plurality of memory blocks, For each resource portion having a respective partition ID, Track a plurality of memory bandwidth usage states corresponding to the memory blocks, where each memory bandwidth usage state is associated with a respective memory block and indicates at least how much of the memory access bandwidth assigned to the respective partition ID for accessing the respective memory block is being used, Determine a usage level associated with the respective partition ID from the plurality of memory bandwidth usage states, Adjust a credit count based on the usage level, Compare the adjusted credit count with a request issuance threshold, Issue the next data access request associated with the respective partition ID in a memory access request queue in accordance with a determination that the credit count is greater than the request issuance threshold, For each resource portion having a respective partition ID, further Update one or more of the plurality of memory bandwidth usage states in response to a previous data access request issued immediately before the next data access request, After a predefined number of clock cycles following the update of the one or more of the plurality of memory bandwidth usage states, the usage level is determined from the plurality of memory bandwidth usage states, After the defined number of clock cycles following the one or more of the updates among the plurality of memory bandwidth usage states, until the next data access request is issued, the credit count is periodically adjusted and compared to the request issuance threshold, updating; A non-transitory computer-readable medium that causes an operation comprising.