Arbitration Scheme for Coherent and Non-Coherent Memory Requests
By implementing separate buffers and watermark thresholds for coherent and non-coherent memory traffic, the solution addresses the increased power consumption and efficiency issues in virtualization-based security systems, enhancing processing efficiency and memory performance.
Patent Information
- Application Number
- JP2022535837
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-12-20
- Filing Date
- 2020-12-11
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2040-12-11
AI Technical Summary
Virtualization-based security mechanisms in processing systems increase memory traffic and power consumption due to address translation requirements, leading to reduced processing efficiency.
Implement separate buffers and watermark thresholds for coherent and non-coherent memory traffic, managing memory requests based on the power state of the processor core, and using different buffers and watermarks for each type of traffic to optimize power consumption and performance.
Reduces power consumption and improves memory performance by efficiently processing memory traffic, maintaining coherence, and optimizing power states of processor cores.
Smart Images

Figure 0007713450000001 
Figure 0007713450000002 
Figure 0007713450000003
Abstract
Description
Background Art
[0001] To efficiently use computer resources, a server or other processing system can implement a virtual computing environment in which the processing system runs multiple virtual machines or guests simultaneously. The resources of the processing system are provided to the guests in a time-division multiplexed or other arbitrated manner, and the resources are presented as a set of dedicated hardware resources for each guest. However, when running guests simultaneously, each guest may be vulnerable to unauthorized access. To protect private guest information, virtualization-based security (VBS) and other virtualization security mechanisms impose security requirements, including address translation requirements, on the processing system. For example, in VBS, in addition to the address translation performed by the guest operating system, the hypervisor or other guest manager needs to provide an additional layer of address translation for virtual address specification. However, these address translation requirements can significantly increase the amount of memory traffic, such as by increasing the amount of coherence probe traffic throughout the processing system, thereby undesirably increasing the power consumption of the system while reducing the overall processing efficiency.
[0002] The present disclosure can be better understood by reference to the accompanying drawings, and many of its features and advantages will become apparent to those of ordinary skill in the art. Where the same reference numerals are used in different drawings, they indicate similar or identical items.
Brief Description of the Drawings
[0003]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
[0004] FIGS. 1 - 8 disclose techniques for managing memory traffic through the interconnect of a processing system by implementing separate buffers and watermark thresholds for coherent and non - coherent memory traffic. Thereby, the processing system can provide various types of traffic (i.e., coherent and non - coherent requests) at various rates depending on various states of the processing system, such as whether the processor core receiving the traffic is in a low - power state. By doing so, the overall power consumption of the system is reduced and the memory performance is improved.
[0005] In some embodiments to be described, the processing system uses an interconnect to carry memory traffic, including memory access requests and coherence probes, between different system modules such as a central processing unit (CPU), a graphics processing unit (GPU), and system memory within the system. The GPU is an example of a client module or client device that generates memory requests in the system. The “requests” and “probes” of memory are used synonymously herein, and the use of “probe” and “coherence probe” generally refers to a memory request that includes determining the coherence state of a particular memory value (data at a memory location), or executing the coherence state along with the data of read or write memory activity. Memory traffic is broadly classified into two categories: coherent memory traffic that requires maintaining coherence between different system modules, and non-coherent traffic that does not require maintaining coherence. For example, a memory request targeting a page table shared between the CPU and the GPU is coherent traffic, and coherence must be maintained between the local caches of the CPU and the GPU to ensure that the running program functions properly. In contrast, a memory access to target frame buffer data that is accessed only by the GPU so that the frame buffer data is not accessed by other system modules is non-coherent traffic. Coherent memory traffic tends to generate a large number of coherence probes, and before accessing a given data set, the memory controller of a system module determines the coherence state of the data of other system modules according to a specified coherence protocol. However, servicing these coherence probes consumes system power, such as by requiring the CPU (or one or more other modules) to exit a low-power state to provide the service of the coherence probe.
[0006] To reduce power consumption, the processing system implements a communication protocol in which memory traffic is grouped or provided to the CPU in a "stuttered" fashion depending on the power state of the CPU or a portion thereof. For example, when the CPU is in a low power mode, regardless of whether it recognizes the low power mode, the processing system holds the memory traffic in a buffer until a threshold amount of traffic is pending. The threshold amount is referred to herein as the "watermark" or "watermark threshold". By holding the memory traffic until the watermark is reached, the traffic is processed more efficiently by the CPU, cache, interconnect, etc. However, since non-coherent memory traffic does not generate coherence probes, non-coherent memory traffic does not have the same impact on the CPU with respect to power consumption over time, and thus the use of a low watermark level for non-coherent memory traffic is more efficient as further described herein. Thus, using the techniques described herein, the processing system uses different buffers and corresponding different watermarks for coherent and non-coherent memory traffic, thereby improving processing efficiency. Using a conventional shared structure for all memory traffic degrades performance between client modules and between coherent and non-coherent memory traffic, so gating of individual types of traffic is performed separately for each client module.
[0007] FIG. 1 is a block diagram of a processing system 100 that executes a series of instructions (e.g., instructions of a computer program, instructions of an operating system process) to perform tasks in place of an electronic device. Thus, in different embodiments, the processing system 100 is incorporated into any of a variety of different types of electronic devices such as a desktop computer, a laptop computer, a server, a smartphone, a tablet, a game console, an e-book reader, etc. To support the execution of a set of instructions, the processing system 100 includes a central processing unit (CPU) 110, a system memory 130, a first client module 140, and a second client module 150. The CPU 110 is a processing unit that executes general-purpose instructions according to a specific instruction architecture such as an x86 architecture or an ARM-based architecture. The client modules 140, 150 are modules such as additional processing units, engines, etc. that execute specific operations in place of the CPU 110. For example, in some embodiments, the client module 140 is a graphics processing unit (GPU) that executes graphics and vector processing operations, and the client module 150 is an input / output (I / O) engine that executes input and output operations in place of the CPU 110 itself or one or more other devices. The instructions in the system memory 130 include one or more CPU instructions 133 and GPU instructions 134.
[0008] To further support the execution of instructions, the processing system 100 includes a memory hierarchy having several levels, where each level corresponds to a different set of memory modules. In the illustrated example, the memory hierarchy includes the system memory 130 and a level 3 (L3) cache 112 in the CPU 110. In some embodiments, the memory hierarchy includes additional levels such as level 1 (L1) and level 2 (L2) caches in one or more of the CPU 110 and the client modules 140, 150.
[0009] To access the memory hierarchy, a device (either the CPU 110 or one of the client modules 140, 150) generates a memory access request, such as a write request to write data to a memory location or a read request to retrieve data from a memory location. Each memory access request includes a memory address that indicates the memory location of the data targeted by the request. In some embodiments, the memory request is generated with a virtual address that indicates a memory location within the virtual address space used by the running computer program. The processing system 100, as further described herein, converts the virtual address to a physical address to identify and access the physical memory location storing the data targeted by the request. To process the memory access request, the CPU 110 includes a memory controller 113 and the system memory 130 includes a memory controller 131. Similarly, the client modules 140, 150 include memory managers 145, 155 that operate to process memory access requests.
[0010] To ensure that different devices (the CPU 110 and the client modules 140, 150) operate on a shared set of data, the memory controllers 113, 131 and the memory managers 145, 155 (collectively referred to as the memory managers) implement a particular coherence scheme, such as the MESI scheme or the MOESI scheme. Thus, the memory managers monitor the coherence state of each memory location and adjust the coherence state of the memory location based on rules defined by the particular coherence scheme.
[0011] To maintain coherence, each memory manager identifies whether a given memory access request targets data that requires maintaining coherence across the system 100 (referred to herein as "coherent data") or data that does not require maintaining coherence across the system 100, referred to herein as "non-coherent data". Examples of coherent data are sets of data accessed via page tables (e.g., page table 132 of system memory 130) used to translate virtual memory addresses to physical memory addresses, and these tables are typically shared between the CPU 110 and the client modules 140, 150. An example of non-coherent data is frame buffer data generated by the GPU (client module 140), whose operation is based on one or more CPU instructions 133 and GPU instructions 134. Non-coherent data is not accessed by other devices of the processing system 100. As an example of coherent data, in response to identifying that a memory access request is a coherent access request, the corresponding memory managers 145, 155 generate a set of coherence probes to determine the coherence state of the targeted coherent data at different levels of the memory hierarchy. The probes communicate with the various memory managers, which identify the coherence state of the memory locations storing the target data and respond to the probes using the coherence state information from each location. Next, the originating memory manager executes appropriate actions defined by a particular coherence scheme based on the response to the coherence probe. The data is operated on by one or more instructions within the system memory 130 that include one or more of the CPU instructions 133 and GPU instructions 134.
[0012] In system 100, coherence probes and memory access requests are communicated between different devices and system memory 130 via central data fabric 120. Central data fabric 120 communicates with an input / output memory management unit (IOMMU) 123 that processes at least a portion of memory traffic 142 to and from system memory 130 and memory traffic to and from client modules 140, 150. For example, IOMMU 123 functions as a table walker and performs a table walk across page table 132 as part of a coherent memory request that causes a probe of the CPU cache (e.g., L3 cache 112) and other caches within system 100 to obtain the current coherent state. Memory traffic to and from client modules 140, 150 includes traffic channel 146 that updates various coherence states via IOMMU 123 and memory controllers 113, 131. Memory traffic to and from client modules 140, 150 includes direct memory access (DMA) traffic 147 related to the contents of system memory 130 and other memory caches within system 100 when the DMA traffic 147 is coherent or a coherent memory access request is generated otherwise. Some of the DMA traffic 147 is coherent traffic and some of the DMA traffic 147 is non-coherent traffic.
[0013] As described above, in some cases, coherence probes and other memory traffic increase the power consumption of processing devices such as CPU 110. In some embodiments to be described, processing system 100 uses a power manager (not shown in FIG. 1) to place one or more modules of CPU 110 (e.g., one or more of core 111 and L3 cache 112) in a low-power state in response to defined system conditions such as detecting that no operations are being performed by CPU 110 on the one or more modules within a threshold time. In response to a coherence probe, the processing system returns one or more modules of CPU 110 to an active high-power mode to process the probe. Thus, a large number of coherence probes that spread over time keep the modules of CPU 110 in the active mode for a relatively long time and consume power. Thus, as further described below, client modules 140, 150 group, batch, or throttle coherence probes and other memory traffic before providing the traffic to central data fabric 120. This grouping enables CPU 110 to process the probes and other memory traffic in a concentrated manner in a relatively short time. This in turn increases the time that the modules of CPU 110 remain in the low-power mode, thereby saving power. This grouping is performed on all memory traffic from client modules 140, 150, all DMA traffic from client modules 140, 150, or all coherent DMA traffic from client modules 140, 150.
[0014] However, as described above, since the non - coherent memory traffic does not result in coherence probes, the CPU 110 is never placed in the active mode. For example, the first client module (GPU) 140 reads from and writes directly to the system memory 130 to and from a dedicated graphics frame buffer within the system memory 130. Thus, as further described below, the processing system 100 manages the communication of coherent and non - coherent memory traffic in different ways (e.g., groups and gates), provides different types of memory traffic to the central data fabric 120 at different rates and different times, and can do so in batches of the same or different sizes for coherent and non - coherent memory traffic. For the sake of explanation, coherent memory traffic is a particular type of traffic, specifically, a memory address translation request, and is sometimes called, or otherwise includes, a page table walk request or one or more page table walks that generate coherent traffic. Some address translation traffic is non - coherent traffic. It will be understood that the techniques described herein apply to any type of coherent memory traffic that is different from other types of memory traffic.
[0015] Each memory manager 145, 155 executes functions such as address translation, walks the address translation table, and operates as a kind of memory management unit (MMU). Some embodiments of the client modules 140, 150 use the memory managers 145, 155 as a virtual manager (VM) and one or more corresponding translation lookaside buffers (TLBs) to perform virtual-to-physical address translation (not shown for clarity). For example, in some cases, the VM and one or more TLBs are implemented as part of a first address translation layer that generates a domain physical address from a virtual address included in a memory access request received at each client module 140, 150 or generated by the client module 140, 150 itself. Some embodiments of the VM and TLB are implemented external to the client modules 140, 150. Although not shown for convenience, the client modules 140, 150 include and use an address translation buffer for holding various memory requests. The client module translation buffer may often not be able to hold all layers of the page table hierarchy.
[0016] Also, the client modules 140, 150 perform a certain amount of prefetching of the page table 132 if physically possible to meet one or more client module service quality (QoS) requirements. For example, the QoS requirement is to normally process a minimum number of DMA requests within a given time. The prefetching is performed by the prefetch memory requests of the client modules 140, 150, and these requests can be non-coherent and coherent type prefetch memory requests. The prefetching is performed according to a prefetch policy that targets the size of the prefetch buffer of each memory manager 145, 155. The prefetch buffer is for data that depends on the properties of the accessible data (e.g., whether the data is in a contiguous range or an adjacent range within a particular memory).
[0017] During operation, client modules 140, 150 use different buffers to buffer memory accesses. Coherent memory accesses, including memory coherence states, are buffered in coherent buffers 141, 151. Non-coherent memory accesses, including some types of DMA instructions related to system memory 130 and not including coherence states, are buffered in non-coherent buffers 142, 152. Both prefetch memory requests and demand memory requests are buffered in their respective buffers 141, 151, 142, 152 according to their coherence / non-coherence characteristics, and these buffers are within respective client module memories or other structures not shown for clarity. As an example, a demand memory request is a request that is generated and released substantially in real time and without delay during operation of the system. Memory instructions or "requests" in client modules 140, 150 are buffered and released for execution based on the operations of non-coherent and coherent monitors for non-coherent and coherent memory operations, respectively, and the monitors are part of non-coherent batch controllers 144, 154 and coherent batch controllers 143, 153, respectively, as further described herein. Memory requests are sent or released as batches. Each batch released from client devices 140, 150 may include a set of non-coherent memory requests and a set of coherent memory requests, or may be of only one type or the other. Generally, coherent batch controllers 143, 153 track the states of the CPU 110 and its components, and other components within system 100 as needed. Address translation is performed on specific memory address tables, such as a set of client memory address tables within client module memory, system memory 130, or a combination of memories within system 100. An address translation table walk within shared system memory 130 may be coherent and may be cached in the CPU (if present there), thus requiring probing of a cache such as L3 cache 112.
[0018] In the case of non-shared memory locations and some direct access memory locations, the non-coherent batch controllers 144, 154 gate or throttle non-coherent requests, including any non-coherent address translation requests queued or buffered in the non-coherent buffers 142, 152, against their respective non-coherent thresholds. In a similar manner, the coherent batch controllers 143, 153 gate or throttle memory probe requests queued or buffered in the coherent buffers 141, 151 generated by each client module 140, 150. The coherent batch controllers 143, 153 gate or throttle memory probe requests targeted at one or more of the other components of the system 100. Coherent probes respond to one or more states of the state of the CPU 110 and its components, and the states of other copies within the system 100, as compared to their respective coherent thresholds.
[0019] Figure 2 is a timing diagram showing a method 200 for an arbitration scheme for coherent translation memory requests, according to some embodiments. In a system (e.g., system 100), the steps of method 200 are executed according to the particular state of components such as the CPU 110 and CPU caches such as the L3 cache 112. The states of the various components change over time, with time marked along the horizontal axis 221. The memory requests involved in method 200 are both coherent and request address translation. For example, some DMA requests are classified into this category.
[0020] For reference, the conventional timing of coherent conversion requests, including several DMA type memory requests, is shown in the upper part 210 of the figure, titled "Without Batch Processing of Coherent Conversion Requests". In the first state 201 of the conventional timing, a core or core complex such as one core 111 remains powered on and is labeled as full power in method 200. The L3 cache associated with the core also remains powered on in a power-saving state. The core and L3 cache in the first state 201 of the conventional scheme receive and continue to process coherent and non-coherent memory requests without batch processing. In the second state 202, the system 100 places the core in a low-power state, and the L3 cache is placed in a hold-only mode without changing the cache coherence state and content by the core 111. The core and L3 cache in the second state 202 of the conventional scheme receive and continue to process coherent and non-coherent memory requests. In the third state 203 and the fourth state 204, the system 100 maintains the core in a low-power state. However, in the third state 203 and the fourth state 204, the L3 cache ends the hold mode to receive and process receivable coherent requests (e.g., coherence probes). As shown by the fifth state 205, after the core 111 is placed in a low-power state, for example, during a specific time when the relative CPU is inactive, the L3 cache is flushed, and the relevant part of the central data fabric (DF) 120 is power gated (registered trademark) to further maintain power consumption by the system while the core 111 and its support structure are in a low-power state.
[0021] Instead of the conventional scheme, a batch processing scheme for coherent transformation requests is implemented and shown in the lower part 220 of the figure. The title of this scheme is "Batch Processing of Coherent Transformation Requests". Regarding timing, in the first batch processing state 211, the core or core complex is at full power until it is placed in a low-power state as shown in the second batch processing state 212. In this second state 212, the components receive and continue to process coherent and non-coherent memory requests. For example, both coherent transformation requests are processed and the video data request is directly processed by the system 100. However, the system 100 continues to batch process or initiate memory address coherent transformation requests based on the operation of one or more of the buffers 141, 142 and batch controllers 143, 444 for the first client module 140. The same is executed for the second client module 150 by each of the components 151~155.
[0022] When a specific time or a specific number of thresholds for coherent transformation requests is reached, in the third batch processing state 213, the system releases the coherent transformation requests and the L3 cache 112 terminates its holding mode to process the coherent transformation requests (shown as "transform" in method 200). In the fourth state 214, the L3 cache 112 remains in the holding mode or returns to the holding mode to maintain power, and the receivable requests (e.g., coherence probes) continue to be batch processed and released in subsequent states. Therefore, in one or more of the second state 212 and the fourth state 214, the system 100 saves power consumption by at least batch processing the transformation requests. As shown by the fifth state 215, at a specific point in time after the core 111 is put into the low-power state, the L3 cache is flushed and the relevant part of the central DF120 is power gated (registered trademark), thereby further increasing power savings.
[0023] Figure 3 is a block diagram 300 of individual coherent and non - coherent batch controllers 130, 320, according to some embodiments. Each of the batch controllers 310, 320 is provided as a pair of monitors to each client module 140, 150 within system 100. According to some embodiments, as further described herein, a pair of batch controllers 310, 320 is provided for each data processing or data generation unit of each client module 140, 150. The batch controllers 310, 320 track the number of memory requests for each client module within system 100.
[0024] According to some embodiments, each batch controller 310, 320 tracks or monitors both pre - fetch requests and demand requests 310, 311. The non - coherent batch controller 310 batches (groups) requests 301 and releases requests 301 as an emergency 302 at an emergency threshold 303 that is the same as or different from the coherent emergency threshold 313 of the coherent batch controller 320. The coherent batch controller 320 batches requests 311 and releases the requests as an emergency 312 when the coherent emergency threshold 313 is reached or exceeded. In some embodiments, the batching is performed for all conversion traffic or all DMA conversion traffic.
[0025] The specific release timing occurs at the end of the time window indicated by the state of method 200 or upon detection of a specific monitoring event such as meeting or exceeding each threshold. Batch controllers 310, 310 assert the urgency or release of batch processing of requests on each request channel. Each batch controller 310, 320 has its own respective coherent and non - coherent buffer watermarks, which are the number of each actual request 310, 311 in buffers 310, 320. When the coherent and non - coherent watermarks reach or exceed each emergency threshold 303, 313, the batch of requests is sent and completed. That is, when the batch controller 310, 320 detects the number of buffered coherent or non - coherent memory requests in its memory request buffer, the batch controller 310, 320 executes one or more additional actions.
[0026] Figure 4 is a flow diagram showing a method 400 for an arbitration scheme for coherent and non - coherent memory requests according to some embodiments, which is executed in a system such as system 100. Method 400 shows the operation of buffers 141, 142, batch controllers 143, 144, and DMA clients such as a first DMA client of a client module (e.g., client module 140) within the system. Memory access requests are generated by the DMA client or client modules 140, 150. In block 401, the system receives a memory access request from the client module.
[0027] In block 402, the system determines whether the processor core and its processor cache are in their respective active power states (low - power states). If so, in block 403, system 100 does not batch - process the memory request.
[0028] With respect to blocks 402, 403, the system 100 determines the power state of one or more of the processor core 111 and its cache 112 in one of a plurality of ways. For example, with respect to the client module 140, each of the batch controllers 143, 144 of the client module 140 receives a first signal that the processor core 111 is in a particular power state (e.g., an active power state, a low power state). Based on this signal, the system proceeds to either block 403 or 404. If the signal refers to an active or full power state, the batch controllers 143, 144 and the buffers 141, 142 operate in cooperation with the memory manager 145 (in the case of the client module 140), and in the absence of an active power state signal, cause batching and alignment as further described herein.
[0029] Regarding alignment, the client module 140 releases at least some of the currently buffered non - coherent memory requests (e.g., DMA memory requests) and at least some coherent memory requests (e.g., cache coherence probes) within the same batch. The release is based on, for example, a release signal generated by the client module 140 or its components (for threshold events regarding one or more thresholds 303, 313), or by its components such as the CPU 110 or core 111 or core complex (for state events or state changes in method 200). As an example, the release releases coherent requests along with non - coherent requests if there are non - coherent requests exceeding the emergency threshold 313, or occurs when a component power state change such as that of the CPU 110 is detected or signaled. In some embodiments, the alignment of requests is the alignment of non - coherent memory requests and coherent memory requests at the edge of the time window of operation of the client module 140. As referred to herein, "alignment" does not, unless otherwise indicated, align or adjust the timing of coherent memory requests with non - coherent memory requests. The time window of operation (e.g., represented by any of states 211 - 215) for buffering and releasing memory requests includes multiple clock cycles of operation of the system 100, the central data fabric 120, the client modules 140, 150, or the CPU 110.
[0030] In block 404, when the processor core and its cache are in a low-power state or are scheduled to be in a low-power state, batch processing is performed and specific memory requests are aligned with requests from other clients within system 100. For example, the system batches coherent requests from clients, batches non-coherent requests, and aligns non-coherent requests from a client with coherent requests from the same client. Batch processing includes buffering each request and releasing requests within a group. For example, buffering includes buffering non-coherent requests and coherent requests in an initial buffer. In another example, buffering includes buffering non-coherent requests in a first buffer (e.g., non-coherent buffer 142) and buffering coherent requests in a second buffer (e.g., coherent buffer 141).
[0031] Following block 404, system 100 performs further actions when the processor core and its cache are in a low-power state. Starting from block 405 and up to block 407, specific actions are performed based on the operation of batch controllers 143, 144 as described in connection with block diagram 300. In block 405, the system or a client notifies the urgency of non-coherent requests when the non-coherent watermark (current level) of non-coherent client memory requests in the non-coherent buffer falls below an emergency non-coherent watermark threshold or exceeds some other threshold. As an example, the non-coherent watermark is the non-coherent address translation watermark threshold for non-coherent DMA memory operations or non-coherent shared memory operations. That is, if a sufficient number of non-coherent requests are not completed within a unit of time, non-coherent requests are sent as completed with an urgency indication so that the system can timely complete non-coherent requests from clients 140, 150 via central data fabric 120.
[0032] In block 406, the system or client notifies the urgency of the coherent request when the coherent watermark (current level) of the coherent client memory requests in the coherent buffer exceeds or touches the emergency non-coherent watermark threshold. For example, the emergency coherent watermark threshold is the probe watermark threshold. The emergency non-coherent watermark threshold takes the same or different value as the emergency coherent watermark threshold, and the thresholds operate independently of each other for the coherent and non-coherent buffers. The urgency is asserted on each request channel for both the coherent and non-coherent requests. The thresholds are stored in respective registers in any of the components of the system 100 such as one or more of the client modules 140, 150, the central data fabric 120, the CPU 110, and the system memory 130.
[0033] In block 407, for specific systems or software applications such as those running in a state where the probe memory bandwidth is limited and the ratio of coherent memory requests is high or relatively high, a specific bandwidth upper limit or ceiling per unit time or per group of coherent memory requests is set. This cap is executed in response to detecting the amount of coherent memory request activity within system 100 with respect to specific client modules 140, 150. This cap helps to spread out the probe requests within system 100 from the specific client modules 140, 150 over time, thereby consuming the coherent memory bandwidth more conservatively. In block 408, system 100 determines whether the processor cache 112 is flushed. In that case, buffering and release in batches, as in blocks 404 - 407, are not executed. Instead, in block 409, coherent and non - coherent requests are batched together to maximize the fabric stutter efficiency. Since these CPU caches are flushed and there is no context information for providing coherent memory requests, the CPU caches are not considered. In some systems, when the processor cache 112 is flushed, batching no longer provides substantial power savings. When the processor cache 112 is not flushed, activities such as those in blocks 404 - 407 continue until a scheduled cache flush or until the processor core 112 and its L3 cache 112 are released from the low - power mode (e.g., return to the full - power mode).
[0034] Method 400 is applicable to any device that generates coherent and non-coherent memory requests. Although not shown, one state of system 100 occurs when processor core 111 is in a low-power state and its L3 cache 112 is in an active power state. Such a state is a common state where various memory caches process memory requests. The operation continues in system 100 as in block 404, and over time, L3 cache 112 has an opportunity to be temporarily placed in a low-power or non-active state to conserve power, as described with reference to FIG. 2. Generally, it is understood that the various components of a system that executes method 400 include circuitry that performs various activities of blocks 401 - 409, including its variations.
[0035] FIG. 5 is a block diagram of CPU 110 of system 100 according to some embodiments. In addition to L3 memory cache 112 and memory controller 113, CPU 110 includes a power manager 517 and first and second core complexes 521, 522. The first core complex 521 includes a first processing core 511, a first level (label L1) memory cache 513, and a second level (label L2) memory cache 515. The second core complex 522 includes a second processing core 512, a second L1 memory cache 514, and a second L2 memory cache 516. L3 memory cache 112 is shared across core complexes 521, 522. At least L3 cache 112 is part of the CPU coherent context used by client modules 140, 150 for their respective memory requests. In some embodiments, at least L2 and L3 caches are part of the CPU coherent context. Each of the first and second core complexes 521, 522 is independently controlled in terms of powering at least between a powered-up state and a low-power state. Power to core complexes 521, 522 is controlled, for example, by power manager 517. In other embodiments, power is managed by respective power managers provided to each core complex 521, 522. When CPU 110 is not under a heavy workload, one or both of core complexes 521, 522 are placed in a low-power state to reduce power consumption, such as by the operation of power manager 517.
[0036] When one or both of the complexes 521, 522 are in a low power state, the client modules 140, 150 can still operate to read from and write to the system memory 130, thereby still being able to benefit from having the current coherence state of the entries in the cache of the CPU 110. After a predetermined time has elapsed, one or more cache entries associated with the low power core complex are flushed because the data and instructions therein are old and are assumed to need updating. Occasionally, the central data fabric 120 communicates the CPU state to the client modules 140, 150. The CPU state includes the active or low power state of the CPU 110 (or the active or low power state of each of the core complexes 521, 522) and the equivalent - also includes the cache state with respect to whether one or more of the caches of the CPU 110 are flushed. For example, the CPU state is communicated to the memory managers 145, 155. Based on the state of the CPU or processor, each of the controllers 143, 144, 153, 154 induces the alignment and release of coherent and non - coherent memory requests of the client modules 140, 150 that conform to the scheme or schedule further described herein. The state includes the state information used during processing. For example, when in a low power state (cores 511, 512 or core complexes 521, 522 are offline) and the CPU cache (one or more of caches 112, 515, 516) has not been flushed, a coherent memory probe sent to the CPU 110 during this state removes one or more low power caches from their holding state, and the probe induces the central data fabric 120 to communicate with the CPU cache. During this process, certain memory requests are batched together to reduce the time it takes for the CPU cache to exit the cache - holding (low power) state and process memory address translation requests.
[0037] FIG. 6 is a block diagram of non - coherent and coherent batch controllers 144, 143 in client module 610 according to some embodiments. The non - coherent batch controller 144 includes a set of non - coherent (memory address) watermarks 601, a set of non - coherent watermark thresholds 602, and (non - coherent) batch control logic 603 that includes monitor logic for a client module within system 100 where non - coherent (e.g., non - coherent conversion) requests are monitored.
[0038] The coherent batch controller 143 includes a set of coherent watermarks 611, a set of coherent watermark thresholds 612, and (coherent) batch control logic 613 that includes monitor logic for a client module within system 100. The batch control logics 603, 613 compare each watermark 601, 611 (number of memory requests) with the matching watermark thresholds 602, 612 from a particular client module. For example, the batch control logic is provided to each of client modules 140, 150 within system 100. Although the term watermark is used herein, a watermark refers to the current number of memory (e.g., coherent, non - coherent, conversion, probe) requests being monitored by each of the batch control logics 603, 613.
[0039] The IOMMU 123 includes an address translator 621 and one or more translation lookaside buffers (TLBs) 622. The address translator 621 is generally referred to as a table walker (not shown), or otherwise includes a table walker that, as understood by those skilled in the art, converts a (module or device) virtual memory address to an address in physical memory, such as by walking a particular page table, such as the page table 132 of the system memory 130. When using virtual addresses, each process executing instructions (e.g., CPU instructions 133, GPU instructions 134) in the processing system 100 has a corresponding page table 132. The page table 132 of the process converts the device-generated (e.g., virtual) addresses used by the process to physical addresses within the system memory 130. For example, the IOMMU 123 executes a table walk of the page table 132 to determine the translation of the address in the memory access request. Translations frequently used by the IOMMU 123 are stored in the TLB 622 or a TLB within the system memory 130 and are used to cache frequently requested address translations. An entry is removed from the TLB 622 to make space for a new entry according to the TLB replacement policy. The TLB 622 is shown as an integrated part of the IOMMU 123. However, in other embodiments, the TLB 622 is implemented in another structure and components accessible by the IOMMU 123.
[0040] In the case of client modules 140, 150, the central data fabric 120 and the IOMMU 123 provide a translation interface to other parts of the system 100. In some embodiments, the client modules 140, 150 are high-bandwidth types of devices that generate a significant amount of memory traffic, including a significant amount of IOMMU traffic. Although not shown for clarity, the central data fabric 120 may include one or more bus controllers as peripheral controllers that include a system controller and a PCIe controller for communicating with the client modules 140, 150 and other components of the system 100. The client bus controller is bi-directionally connected to the input / output (I / O) hub and the bus, facilitating communication between various components. Through the I / O hub, various components can directly transmit and receive data, for example, to the batch controllers 143, 144 of the first client module 140 and to the registers and memory locations of various components within the system 100.
[0041] FIG. 7 is a block diagram of a client device that supports a plurality of DMA clients according to some embodiments. Device 710 is another embodiment of client modules 140, 150 of system 100. Device 710 includes at least one client virtual manager (VM) 701, at least one memory address translation cache 702, and a plurality of DMA clients 714, 724. The first DMA client 714 includes its own coherent prefetch buffer 711, non-coherent prefetch buffer 712, and MMU 713. The second DMA client 724 includes a second coherent prefetch buffer 721, a second non-coherent prefetch buffer 722, and a second MMU 723. Each of the DMA clients 714, 724 reads and writes independently to system memory 130. Each of the DMA clients 714, 724 is tracked by respective translation monitors 121 and probe monitors 122 of the central data fabric 120. The first DMA client 714 is coupled to and associated with a first external device such as display 720. The second DMA client 724 is coupled to and associated with a second external device such as camera 730. The first MMU 713 reads and / or writes to one or both of a dedicated portion of system memory 130 and a coherent shared portion of system memory 130 for display 720. The second MMU 723 reads and / or writes to one or both of a dedicated portion of system memory 130 and a coherent shared portion of system memory 130 for camera 730.
[0042] FIG. 8 is a block diagram of an address translation system 800 according to some embodiments. System 800 uses the components of system 100 to translate client module virtual address 801 from a first client module 140 to a client module physical address 802, and then to a system physical address 803 within system memory 130. A memory access request includes a device-generated address such as a client virtual address 801 that is executed on a first client module 140 or used by an application associated with the first client module 140. In the illustrated embodiment, the VBS compliance mechanism uses a two-level translation process that includes (1) a first-level translation 815 managed by the OS or device driver 810, and (2) a second-level translation 825 managed by the hypervisor 820 to provide memory protection (e.g., against kernel-mode malware). The VBS address translation traffic and the probe memory bandwidth consumed by this traffic can result in increased CPU cache power requirements and, when VBS is used to access a coherent memory such as system memory 130, can degrade the memory performance of data access. The first-level translation 815 translates an address generated by a device such as a virtual address within a memory access request to a client physical address 802. The client physical address is also referred to as a domain physical address such as a GPU physical address. In some embodiments, the first-level translation 815 is typically executed by a client module VM and associated TLB that are associated with a guest VM as described herein.
[0043] The client physical address 802 is passed to a second-level translation 825 that converts the client physical address 802 to a system physical address 803 indicating a location within the system memory 130. As described herein, the second-level translation 825 verifies that the device is permitted to access a particular region of the system memory 130 indicated by the system physical address 803, for example, using permission information encoded in entries in a related page table and a translation lookaside buffer (TLB) used to perform the second-level translation 825. In some embodiments, this translation system 800 is supported or mediated by an IOMMU such as the IOMMU 123. Based on settings 804 associated with the OS or device driver 810, one or both of the watermark thresholds 602, 612 are set or adjusted relative to the starting value of the first client module 140. Although illustrated for convenience in the OS or device driver 810, the settings 804 may alternatively be within or associated with the hypervisor 820. In some embodiments, the settings 804 are stored in a register of the system memory 130 or in its own dedicated memory register within a client module (e.g., GPU 140, I / O engine 150). Although a single setting 804 is shown, alternatively, the settings 804 may be multiple values related to one or more QoS values of a particular client module 140, 150 (e.g., GPU, I / O engine) in a particular embodiment. The settings 804 are set by the user or programmatically obtained through reading or configuration detection of hardware, firmware, or software values from external devices such as the display 720 and camera 730 coupled to the client module 150. The settings 804 can be determined at the initialization of communication with the external device and updated during operation of the system 100, for example, to adjust the operation of the controllers 143, 144 of the first client module 140.
[0044] In some embodiments, the apparatus and techniques described above are implemented in a system that includes one or more integrated circuit (IC) devices (also referred to as integrated circuit packages or microchips), such as the CPU 110, system memory 130, and client devices or modules 140, 150, 710, etc. described with reference to FIGS. 1-8. The buffers are described for storing, holding, or tracking memory requests, but can be replaced by other specific types of structures, devices, or mechanisms as would be understood by one of ordinary skill in the art. For example, instead of buffers such as buffers 141, 142, 151, 152 of client modules 140, 150, tables, memory blocks, memory registers, linked lists, etc. are used. The same applies to circuits, components, modules, devices, etc. other than those described above.
[0045] Electronic design automation (EDA) and computer-aided design (CAD) software tools can be used for the design and manufacture of the described IC devices. These design tools are typically represented as one or more software programs. The one or more software programs include computer-executable code that operates a computer system to represent the circuits of one or more IC devices so as to execute at least a portion of the processes for designing or adapting a manufacturing system for fabricating the circuits. This code can include instructions, data, or a combination of instructions and data. The software instructions representing the design tools or manufacturing tools are typically stored on a computer-readable storage medium accessible to a computing system. Similarly, the code representing one or more phases of the design or manufacture of an IC device may be stored on the same computer-readable storage medium or a different computer-readable storage medium, and may be accessed from the same computer-readable storage medium or a different computer-readable storage medium.
[0046] A computer-readable storage medium includes any non-transitory storage medium or combination of non-transitory storage media that can be accessed by a computer system during use to provide instructions and / or data to the computer system. Such storage media include, but are not limited to, optical media (e.g., compact disc (CD), digital versatile disc (DVD), Blu-ray (registered trademark) disc), magnetic media (e.g., floppy (registered trademark) disc, magnetic tape, magnetic hard drive), volatile memory (e.g., random access memory (RAM) or cache), non-volatile memory (e.g., read-only memory (ROM) or flash memory), or microelectromechanical systems (MEMS)-based storage media. A computer-readable storage medium (e.g., system RAM or ROM) may be built into a computing system, a computer-readable storage medium (e.g., magnetic hard drive) may be fixedly attached to a computing system, a computer-readable storage medium (e.g., optical disc or universal serial bus (USB)-based flash memory) may be removably attached to a computing system, or a computer-readable storage medium (e.g., network-accessible storage (NAS)) may be coupled to a computer system via a wired or wireless network.
[0047] In some embodiments, some aspects of the above technologies may be implemented by one or more processors of a processing system that executes software. The software is stored in a non-transitory computer-readable storage medium or includes one or more sets of executable instructions tangibly embodied on a non-transitory computer-readable storage medium. When executed by one or more processors, the software can include instructions and specific data that operate the one or more processors to execute one or more aspects of the above technologies. The non-transitory computer-readable storage medium can include, for example, magnetic or optical disk storage devices, solid-state storage devices such as flash memory, cache, random access memory (RAM), or other one or more non-volatile memory devices. The executable instructions stored in the non-transitory computer-readable storage medium can be in source code, assembly language code, object code, or other instruction formats interpretable or executable by one or more processors.
[0048] In addition to the above, it should be noted that not all activities or elements described in the general description are required, some activities or parts of a particular device may not be required, one or more additional activities may be performed, and one or more additional elements may be included. Further, the order in which activities are listed is not necessarily the order in which they are performed. Also, the concepts have been described with reference to particular embodiments. However, those skilled in the art will understand that various changes and modifications can be made without departing from the scope of the invention as set forth in the claims. Accordingly, the specification and drawings are to be considered in an illustrative rather than a limiting sense, and all such modifications are intended to be included within the scope of the invention.
[0049] Advantages, other advantages, and solutions to problems have been described above with respect to specific embodiments. However, advantages, benefits, solutions to problems, and features that may give rise to or manifest any advantage, benefit, or solution are not to be construed as important, essential, or indispensable features of any or all of the claims. Further, since the disclosed invention can be modified and practiced in different but similar ways that will be apparent to those skilled in the art having the benefit of the teachings herein, the specific embodiments described above are merely illustrative. There is no limitation as to the details of the construction or design shown herein other than as set forth in the appended claims. Accordingly, it is evident that the specific embodiments described above may be varied or modified and that all such variations are considered to be within the scope of the disclosed invention. Accordingly, the protection sought herein is set forth in the appended claims.
Claims
1. A processor, comprising: a coherent memory request buffer having a plurality of entries for storing coherent memory requests from a client module; a non-coherent memory request buffer having a plurality of entries for storing non-coherent memory requests from the client module; and the client module, wherein the client module buffers coherent memory requests in the coherent memory request buffer, and based on the amount of the buffered coherent memory requests exceeding a first threshold, releases the buffered coherent requests for processing; buffers non-coherent memory requests in the non-coherent memory request buffer, and based on the amount of the buffered non-coherent memory requests exceeding a second threshold different from the first threshold, releases the buffered non-coherent requests for processing; and is configured to perform the above. Processor.
2. The client module receives an indicator that the processor core is in a specific power state; and based on the received indicator, starts buffering the coherent memory requests and buffering the non-coherent memory requests; and is configured to perform the above. The processor according to Claim 1.
3. A client module memory; a set of memory address tables; and a memory manager that operates as a table walker to convert at least one client module virtual memory address of the coherent memory request and the non-coherent memory request into a physical client module address of the client module memory using the set of memory address tables. The processor according to Claim 1 or 2.
4. The non-coherent memory request includes a non-coherent prefetch memory request. The processor according to any one of Claims 1 to 3.
5. Each of the non-coherent memory requests is converted based on a first memory address conversion and a second memory address conversion related to a virtualization-based security (VBS) mechanism. The processor according to any one of Claims 1 to 4.
6. The client module After receiving a signal that the processor core is in a low-power state, when the amount of the buffered non-coherent memory requests in the non-coherent memory request buffer exceeds the second threshold, it is configured to release the buffered non-coherent requests. The processor according to any one of claims 1 to 5. **Claim 7** The client module is configured to limit the number of coherent memory requests released from the coherent memory request buffer based on the amount of activity of the coherent memory requests. The processor according to any one of claims 1 to 6. **Claim 8** The client module monitors by comparing the amount of the buffered coherent memory requests in the coherent memory request buffer with the first threshold, monitors by comparing the amount of the buffered non-coherent memory requests in the non-coherent memory request buffer with the second threshold, and based on the amount of the buffered coherent memory requests in the coherent memory request buffer and the amount of the buffered non-coherent memory requests in the non-coherent memory request buffer, releases a first set of non-coherent memory requests from the non-coherent memory request buffer and a second set of coherent memory requests from the coherent memory request buffer in the same batch. is configured to perform. The processor according to any one of claims 1 to 7. **Claim 9** The release of the non-coherent memory requests and the coherent memory requests is performed in response to receiving a release signal, The release signal is generated based on at least one of that the first number of the buffered coherent memory requests in the coherent memory request buffer exceeds the first threshold, and that after receiving a signal that the processor core is in a low-power state, the second number of the buffered non-coherent memory requests in the non-coherent memory request buffer exceeds the second threshold. The processor of claim 8. **Claim 10** A method for arbitrating coherent and non-coherent memory requests generated by a device, The method Detecting the power states of a processor core and a processor shared cache; Buffering coherent memory requests in a coherent memory request buffer in response to detecting a low power state of the processor core; Buffering non - coherent memory requests in a non - coherent memory request buffer; Releasing one or more coherent memory requests in the coherent memory request buffer based on the amount of the buffered coherent memory requests exceeding a first threshold; Releasing one or more non - coherent memory requests in the non - coherent memory request buffer based on the amount of the buffered non - coherent memory requests exceeding a second threshold different from the first threshold, the method comprising: A method.
11. Releasing the one or more non - coherent memory requests and the one or more coherent memory requests includes releasing the one or more non - coherent memory requests and the one or more coherent memory requests as a batch based on the low power state, the method of claim 10.
12. Each of the non - coherent memory requests includes at least one of a non - coherent prefetch memory request and a demand non - coherent memory request, the method of claim 10 or 11.
13. After releasing the one or more non - coherent memory requests, for each of the non - coherent memory requests, performing a first memory address translation and a second memory address translation that match a virtualization - based security (VBS) mechanism, the method according to any one of claims 10 to 12.
14. Releasing the one or more non - coherent memory requests and the one or more coherent memory requests is based on at least one of: Detecting that a first number of the buffered coherent memory requests in the coherent memory request buffer exceeds the first threshold; After receiving a signal that the processor core is in the low power state, detecting that a second number of the buffered non - coherent memory requests in the non - coherent memory request buffer exceeds the second threshold, the method according to any one of claims 10 to 13.
Citation Information
Patent Citations
Computer system power management method and apparatus
JP2007517332A
Hierarchical memory arbitration technology for heterogeneous sources
JP2012525642A
Access control to pages in memory of a computing device - Patent Application 20070122997
JP2019522298A
Method for dynamic arbitration of real-time streams in the multi-client systems
WO2019023068A1