Preserve the processor core's cache entries during a powered-down state
By storing cache entries in the reserved area when the processor core is powered off and restoring cache on power on, the problem of degradation in performance and difficult to maintain cache consistency when the processor core exits the powered off state is solved, and performance acceleration and cache consistency maintenance is achieved.
Patent Information
- Application Number
- CN201880071347.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-11-01
- Filing Date
- 2018-09-14
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2038-09-14
AI Technical Summary
Existing processing systems have degraded performance when the processor core exits the power outage state and are difficult to maintain cache consistency.
Store cache entries associated with the processor core in a reserved area that receives reserved voltage when the processor core is in a powered off state and uses this information to restore the cache when the processor core is powered on.
In this way, performance recovery can be accelerated when the processor core exits the power-off state and cache consistency can be maintained during the power-off state.
Smart Images

Figure CN111344685B_ABST
Abstract
Description
Background Art
[0001] Conventional processing systems include processing units, such as central processing units (CPUs) and graphics processing units (GPUs), which typically include multiple processor cores for executing instructions simultaneously or in parallel. Information representing the state of the processor cores is stored in caches. The information stored in the caches is used to speed up the operation of the processor cores. For example, a translation lookaside buffer (TLB) is used to cache the conversion of virtual addresses to physical addresses so that the TLB can perform virtual to physical address conversion without time-consuming page table traversal. For another example, a cache hierarchy stores instructions executed by the processor cores and data used by the instructions when executed by the processor cores, so that the instructions or data do not have to be fetched from an external memory each time an instruction or data is needed. The cache hierarchy includes: an L1 cache for caching information for each processor core, an L2 cache for caching information for a subset of processor cores in a processing unit, and an L3 cache for caching information for all processor cores in a processing unit. An inclusive cache hierarchy implements an L3 cache that includes information stored in an L2 cache that includes information stored in an L1 cache.
[0002] When the processor core is not actively performing an operation (such as executing an instruction), the processor core is placed in a power-off state (which may be referred to as a C6 or CC6 state) to reduce leakage current. Before placing the corresponding processor core in the power-off state, a cache storing state information of the processor core is flushed and then the cache is powered off. For example, when the processor core is powered off, entries in the TLB of the processor core are lost. For another example, cache entries in an L1 cache or an L2 cache associated with the processor core are flushed to an L3 cache or an external memory, such as a dynamic random access memory (DRAM) or a disk drive. Then when power is removed from the processor core and the L1 or L2 cache, the cache entries are lost from the L1 or L2 cache. The lack of up-to-date information in the cache degrades the performance of the processor core when the processor core exits the power-off state. BRIEF DESCRIPTION OF THE DRAWINGS
[0003] The present disclosure may be better understood, and its numerous features and advantages made apparent to those skilled in the art by referencing the accompanying drawings.The use of the same reference numbers in different drawings indicates similar or identical items.
[0004] Figure 1 is a block diagram of a processing system according to some embodiments.
[0005] Figure 2 is a block diagram of a portion of a processing system including a translation lookaside buffer (TLB) and an external memory hierarchy according to some embodiments.
[0006] Figure 3 is a flow chart of a method of storing information representing entries in a TLB in a reserved area before powering off a processor core associated with the TLB, according to some embodiments.
[0007] Figure 4 is a flow chart of a method of restoring entries in a TLB using information stored in a reserved area before powering off a processor core associated with the TLB, according to some embodiments.
[0008] Figure 5 is a block diagram of a cache hierarchy according to some embodiments.
[0009] Figure 6 is a block diagram of a portion of a cache hierarchy including an L2 cache and an L3 cache according to some embodiments.
[0010] Figure 7 is a flow chart of a method of storing information representing entries in a lower level cache in a reserved area before powering off a processor core associated with the lower level cache, according to some embodiments.
[0011] Figure 8 is a flow chart of a method of repopulating entries in a lower level cache using information stored in a reserved area in response to a processor core initiating an exit from a powered-down state according to some embodiments. DETAILED DESCRIPTION
[0012] At least in part to accelerate the performance of the processor core when the processor core exits the power-off state and to maintain cache coherence during the power-off state, entries in the cache associated with the powered-off processor core are stored in a reserved area that receives a reserved voltage when the processor core is in the power-off state. After the copy of the entry is stored in the reserved area, the processor core enters the power-off state. Information indicating the invalidation of one or more entries in the cache is stored when the processor core is in the power-off state. The cache is restored based on the stored copy of the entry and the stored invalidation information. The restoration is performed in response to the processor core starting to exit the power-off state. The reserved area is implemented in a portion of the memory hierarchy that remains powered on and running while the processor core is in the power-off state. The reserved area may include a higher-level cache in the cache hierarchy, an external memory such as a dynamic random access memory (DRAM), or a storage element in the processor core that is powered by a power supply that continues to be powered when the processor core is powered off. For example, the reserved area may use a storage element of the cache itself, that is, the data in the cache can be retained in situ when the cache has a power supply that continues to be powered when the rest of the core is in the power-off state.
[0013] In some embodiments, the cache is a translation lookaside buffer (TLB) that caches virtual to physical address translations for use by a processor core. The entries in the TLB are stored in a reserved area in response to the processor core entering a power-off state. When the corresponding processor core is powered off, information indicating a request to invalidate the entries in the TLB is stored, for example, individual requests are stored in a queue or a bit is set to a certain value to indicate that one or more entries in the TLB have been invalidated. The stored TLB entries and invalidation information are used to restore the TLB by refilling the entries in response to the processor core being powered on. For example, after refilling the TLB entries from the reserved area, the TLB invalidation request is replayed to invalidate the entries in the TLB. In the case of a queue overflow, the entire TLB is invalidated because all the information required to restore the TLB is no longer available in the queue. For another example, if the value of the bit indicates that one or more of the TLB entries have been invalidated when the processor core is powered off, the entire TLB is invalidated. In some embodiments, instead of storing the entries in the TLB in response to powering off the processor core, a list of virtual addresses in the TLB is stored in a reserved area. Virtual addresses are pre-fetched when the processor core is powered on to preemptively start the page table walk that populates entries in the TLB.
[0014] In some embodiments, the cache is a lower level cache in an inclusive cache hierarchy, such as an L1 cache or an L2 cache in a cache hierarchy. The reserved area may include a higher level cache (such as an L3 cache) or an external memory (such as a DRAM) that receives a reserved voltage when the processor core is powered off. The reserved area may also be implemented in the cache itself by providing a reserved power supply to the storage element of the cache. In response to the processor core entering a power-off state, for example, the modified value or dirty value in the cache is written to a higher level cache or an external memory by flushing the cache to write out the modified value of the cache entry or by refreshing all entries in the cache. Information indicating the invalidation of the cache entry in the cache is stored when the processor core is powered off. Some embodiments of the higher level cache include a shadow tag storing the physical address of the cache entry associated with the lower level cache of the processor core that is powered off and information indicating whether the entry is valid and clean. The shadow tag may also store information indicating whether the entry is invalid when the processor core is powered off. If the reserved area is implemented in the cache itself, there is no need to refill the cache for recovery. If the reserved area is not implemented in the cache itself, the cache is refilled using the information in the shadow tag in response to the processor core powering up. For example, valid entries of the lower level cache are pre-fetched based on the corresponding physical addresses stored in the shadow tag.
[0015] Some embodiments implement a probe queue that stores information indicating probes received when the processor core is powered off. If the processor core includes a shadow tag, the probe queue only records probes that hit the address in the shadow tag. The probe queue can also be implemented by adding a field to each entry of the shadow tag, which indicates that the corresponding lower-level cache line will be invalidated in response to powering on the processor core. The probes stored in the probe queue are sent to the cache in response to powering on the processor core. In the case of a probe queue overflow, the processor core is powered on to serve the probes in the probe queue or the entire cache is invalidated when the processor core is powered on. Other methods for maintaining cache consistency when the processor core is powered off can also be used. In some cases, the cache is powered on in response to receiving a probe, which allows the cache to invalidate the entry indicated by the probe and then power off again. This method consumes a lot of overhead. Each cache level can also be equipped with a "clock-on" clock that allows the cache to be clocked up and down to serve the probe. In some embodiments, other mechanisms such as Bloom filters are used to identify probes that may hit the lower-level cache and therefore must be recorded in the probe queue.
[0016] Figure 1is a block diagram of a processing system 100 according to some embodiments. Processing system 100 includes or can access memory 105 or other storage components implemented using non-transitory computer-readable media such as dynamic random access memory (DRAM). However, memory 105 can also be implemented using other types of memory, including static random access memory (SRAM), non-volatile RAM, etc. Because memory 105 is implemented external to the processing units implemented in processing system 100, it is referred to as external memory. Processing system 100 also includes bus 110 to support communication between entities implemented in processing system 100 (such as memory 105). Some embodiments of processing system 100 include a bus 110 that is not shown for clarity. Figure 1 Other buses, bridges, switches, routers, etc. not shown.
[0017] The processing system 100 includes a graphics processing unit (GPU) 115 configured to render images for presentation on a display 120. For example, the GPU 115 renders objects to generate pixel values that are provided to the display 120, which uses the pixel values to display an image representing the rendered object. Some embodiments of the GPU 115 are used for general-purpose computing. In the illustrated embodiment, the GPU 115 implements multiple processor cores 116, 117, 118 (collectively referred to herein as "processor cores 116-118") that are configured to execute instructions simultaneously or in parallel. The processor cores 116-118 are also referred to as shader engines.
[0018] The GPU 115 also includes a memory management unit (MMU) 121 for supporting communication with the memory 105. In the illustrated embodiment, the MMU 121 communicates with the memory 105 via the bus 110. However, some embodiments of the MMU 121 communicate with the memory 105 via a direct connection or via other buses, bridges, switches, routers, etc. The GPU 115 executes instructions stored in the memory 105 and the GPU 115 stores information such as the results of the executed instructions in the memory 105. For example, the memory 105 stores a copy 125 of instructions from the program code to be executed by the GPU 115. The MMU 121 includes a translation lookaside buffer (TLB) 123, which is a cache that stores virtual to physical address translations used by the processor cores 116-118. For example, the processor core 116 transmits a memory access request containing a virtual address to the MMU 121, and the MMU 121 uses the corresponding entry in the TLB 123 to translate the virtual address into a physical address. MMU 121 may then transmit a memory request (eg, to memory 105) using the physical address.
[0019] The GPU 115 includes a cache hierarchy 130 that includes one or more levels of cache for caching instructions or data for relatively low latency access by the processor cores 116-118. Instructions scheduled to the processor cores 116-118 include one or more prefetch instructions for prefetching information (such as instructions or data) into the cache hierarchy 130. For example, a prefetch instruction executed on the processor core 116 prefetches the instruction from the replica 125 so that the instruction is available in the cache hierarchy 130 before the processor core 116 executes the instruction. Although the cache hierarchy 130 is shown as being external to the processor cores 116-118, some embodiments of the processor cores 116-118 incorporate corresponding caches (such as an L1 cache) interconnected to the cache hierarchy 130.
[0020] The processing system 100 also includes a central processing unit (CPU) 140, which implements multiple processor cores 141, 142, 143, which are collectively referred to herein as "processor cores 141-143." The processor cores 141-143 are configured to execute instructions simultaneously or in parallel. The CPU 140 is connected to the bus 110 and thus communicates with the GPU 115 and the memory 105 via the bus 110. The CPU 140 includes an MMU 145 to support communication with the memory 105. The MMU 145 includes a TLB 150, which stores virtual to physical address translations used by the processor cores 141-143. The CPU 140 executes instructions such as program code 155 stored in the memory 105 and the CPU 140 stores information such as the results of the executed instructions in the memory 105. The CPU 140 is also capable of starting graphics processing by issuing a draw call to the GPU 115.
[0021] Some embodiments of CPU 140 include a cache hierarchy 160 that includes one or more levels of cache for caching instructions or data for relatively low latency access by processor cores 141-143. Although cache hierarchy 160 is shown as being external to processor cores 141-143, some embodiments of processor cores 141-143 incorporate corresponding caches interconnected to cache hierarchy 160. In some embodiments, instructions dispatched to processor cores 141-143 include one or more prefetch instructions for prefetching information, such as instructions or data, into cache hierarchy 160. For example, prefetch instructions executed by waves on processor core 141 may prefetch instructions from program code 155 so that the instructions are available in cache hierarchy 160 before processor core 141 executes the instructions.
[0022] Input / output (I / O) engine 165 handles input or output operations associated with display 120 and other elements of processing system 100 (such as keyboard, mouse, printer, external disk, etc.). I / O engine 165 is connected to bus 110, so that I / O engine 165 can communicate with memory 105, GPU 115 or CPU 140. In the illustrated embodiment, I / O engine 165 is configured to read information stored on external storage component 170, which is implemented using non-transitory computer-readable media such as compact disc (CD), digital video disc (DVD), etc. I / O engine 165 can also write information such as processing results of GPU 115 or CPU 140 to external storage component 170.
[0023] As discussed herein, conventional processing systems do not provide a mechanism for maintaining cache coherence between a powered-on cache and information stored in a cache that is powered-off in conjunction with a corresponding processor that enters a power-off state. As used herein, the term "powered-off" refers to a state in which the power supplied to a processor core and associated entities such as a cache drops below a level required to maintain the functionality of the processor core or other entities. For example, a powered-off processor core cannot execute instructions. For another example, a powered-off cache is not supplied with sufficient power to maintain stored bit values, such as by maintaining the states of transistors of bit storage elements used to construct the cache.
[0024] Conventional processing systems fail to account for the invalidation of cache entries in a powered-off cache when a processor core is in a powered-off state. At least in part to address this shortcoming in conventional practice, the processing system 100 stores information representing entries in the TLB 123, 150 or cache hierarchy 130, 160 in a reserved area in response to the corresponding processor core of the processor cores 116-118, 141-143 beginning to enter a powered-off state. The reserved area receives a reserved voltage while the processor core is in the powered-off state. The processing system 100 monitors for invalidation requests or cache probes issued while the processor core is in the powered-off state, and then selectively refills the TLB 123, 150 or cache hierarchy 130, 160 with entries that were not invalidated while the processor core was in the powered-off state.
[0025] In some embodiments of the processing system 100, multiple options are available for storing and restoring entries in the TLB 123, 150 or cache hierarchy 130, 160. When the processor core 116-118, 141-143 is in a power-off state, a probe (or other failure request) is recorded. The probe is checked against a shadow tag, a Bloom filter, or other information that identifies a potential hit of cached information in a reserved area held at a reserved voltage. Missed probes have no further impact. Probes that hit cached information, probes that are potential hits indicated by a Bloom filter, or probes for which a hit status cannot be determined are recorded in a probe queue or a corresponding shadow tag. In either case, if a probe hit (or potential probe hit) cannot be recorded in a probe queue or a shadow tag, an overflow bit is set. Based on whether the lower level cache is held at a reserved voltage, a selective refill is performed in response to the processor core 116-118, 141-143 powering on. If so, the lower level cache retains the previous entry. The recorded probe hits (or potential hits) are replayed to the lower level cache to invalidate the corresponding entries and restore the cache. If the overflow bit is set, the entire lower level cache is invalidated. If the lower level cache is not held at the retention voltage, the physical address in the shadow tag is used to prefetch (and thereby restore) the lower level cache. Physical addresses in shadow tags that were hit by a probe during the power-off state are not prefetched.
[0026] Figure 2 is a block diagram of a portion 200 of a processing system including a translation lookaside buffer (TLB) 205 and an external memory hierarchy 210 according to some embodiments. The portion 200 is used to implement Figure 1 Some embodiments of the processing system 100 are shown. For example, the external memory hierarchy 210 may include Figure 1105, the L3 cache in the cache hierarchy 130, 160, and the external storage component 170 shown. TLB 205 is used to cache virtual to physical address translations utilized by the processor core 215. Each virtual to physical address translation is stored in an entry 220 (for clarity, the reference symbol indicates only one entry). TLB 205 and processor core 215 are in the same power domain 225. TLB 205 and processor core 215 therefore receive power using the same power system. In some embodiments, TLB 205 and processor core 215 also receive clock signals from the same clock network. Therefore, when the processor core 215 is in a power-off state, TLB 205 is powered off.
[0027] At least a portion of the external memory hierarchy 210 receives power independently of the power supplied to the power domain 225. The independently powered portion of the external memory hierarchy 210 thus receives a reserved voltage when the processor core 215 is in a powered-off state and is used to implement a reserved area for storing information representing an entry 220 in the TLB 205. In the illustrated embodiment, the reserved area stores a copy 230 of an entry 220 in the TLB 205. However, in other embodiments, the reserved area stores other information representing the entry 220, such as a virtual address associated with the entry 220. In response to the processor core 215 beginning to enter a powered-off state, the information representing the entry 220 is stored in the external memory hierarchy 210. For example, in response to a signal indicating that the processor core 215 is about to be powered off, the copy 230 of the entry 220 is written to the external memory hierarchy 210.
[0028] When the processor core 215 is in a powered-off state, entries 220 in the TLB 205 are invalidated. Therefore, the processing system monitors for invalidation requests, such as TLB shootdowns that invalidate entries 220 in the TLB 205. Some embodiments of the external memory hierarchy 210 implement a queue 235 to store invalidation requests received while the processor core 215 is in a powered-off state. The queue 235 has a finite length and overflows when the number of invalidation requests received while the processor core 215 is in a powered-off state exceeds the number of available slots in the queue 235. Some embodiments of the external memory hierarchy 210 store information representing invalidation requests in other formats. For example, the external memory hierarchy 210 stores a single bit that is set to a first value (e.g., 0) to indicate that no invalidation requests have been received for the TLB 205 and a second value (e.g., 1) to indicate that one or more invalidation requests have been received for the TLB 205.
[0029] In response to the processor core 215 beginning to exit from the powered-off state, such as in response to the processor core 215 being powered on, the TLB 205 is populated using information representing the entry 220 stored in the reserved area of the external memory hierarchy 210. For example, the entry in the TLB copy 230 is written back to the TLB 205 to repopulate the entry 220. The state of the TLB 205 is then updated based on any invalidation requests received while the processor core 215 was in the powered-off state. For example, the invalidation requests in the queue 235 are replayed to invalidate the corresponding entry 220 and produce the correct state of the TLB 205. For another example, if the external memory hierarchy 210 stores a virtual address associated with the entry 220, the virtual address is pre-fetched to trigger a page table walk that repopulates the entry 220 in the TLB 205. In this example, a new page table walk is performed to populate the TLB 205, and therefore no entries in the TLB 205 need to be invalidated to keep the TLB 205 consistent with the rest of the system. For yet another example, if the external memory hierarchy 210 stores a single bit indicating whether any invalidation requests have been received, all entries 220 in the TLB 205 are invalidated if the bit indicates that one or more invalidation requests have been received.
[0030] In some embodiments, instead of refilling and then invalidating, entries 220 of TLB 205 are restored by conditionally refilling the entries based on invalidation requests received while processor core 215 is in a powered-off state. For example, only entries in TLB copy 230 that are not invalidated (as indicated by information in queue 235) are written back to TLB 205 in response to processor core 215 being powered on. For another example, if external memory hierarchy 210 stores a single bit that is set to a value indicating that one or more invalidation requests have been received, TLB copy 230 that is invalid based on the bit value is not written back to TLB 205.
[0031] Figure 3 is a flow chart of a method 300 for storing information representing entries in a TLB in a reserved area before powering off a processor core associated with the TLB, according to some embodiments. The method 300 Figure 1 The processing system 100 and Figure 2 The illustrated processing system is implemented in some embodiments as a portion 200 .
[0032] The processing system initiates powering off of the processor core at block 305. For example, in response to the absence of instructions scheduled for execution by the processor core or a prediction that no instructions will be scheduled for execution by the processor core in a subsequent time interval exceeding a power-off threshold, the processor core begins entering a power-off state.
[0033] At block 310, information representing entries in the TLB is stored to an external memory implementing a reserved area that retains power when the processor core is in a powered-off state. Some embodiments of the external memory are implemented using L3 cache, DRAM, and external storage devices such as disk drives. The information includes a copy of the entry in the TLB or a virtual address of the entry in the TLB.
[0034] The processor core is powered off at block 315. Powering off the processor core occurs after information representing entries in the TLB has been stored to external memory to prevent this information from being lost when the TLB is powered off.
[0035] At block 320, information representing the invalidation requests received for the TLB is stored in a reserved area. For example, the reserved area may implement a queue that stores invalidation requests when the processor core is in a powered-off state. For another example, the reserved area may implement a bit that is set to a first value (e.g., 0) to indicate that no invalidation requests have been received for the TLB and a second value (e.g., 1) to indicate that one or more invalidation requests have been received for the TLB.
[0036] Figure 4 is a flow chart of a method 400 for repopulating entries in a TLB using information stored in a reserved area before powering off a processor core associated with the TLB, according to some embodiments. Figure 1 The processing system 100 and Figure 2 The illustrated processing system is implemented in some embodiments as a portion 200 .
[0037] At block 405, the processing system initiates powering up of the processor core. For example, exiting from a powered-down state is initiated in response to a scheduler in the processing system scheduling an instruction for execution on the processor core.
[0038] At decision block 410, the processing system determines whether a queue storing invalidation requests has overflowed in response to the number of invalidation requests received exceeding the number of available slots in the queue. If so, the method 400 proceeds to block 415 and invalidates all entries in the TLB because the queue cannot hold all the information necessary to reconstruct the state of the TLB. If the queue has not overflowed, the method 400 proceeds to block 420.
[0039] At block 420, the TLB is refilled with information representing the entry in the TLB. For example, a copy of the entry is written from a reserved area into the TLB. For another example, the address of the entry stored in the reserved area is pre-fetched to trigger a page table walk that fills the entry in the TLB.
[0040] At block 425, the state of the TLB is modified based on the invalidation request received while the processor core is in the power-off state. For example, the invalidation request stored in the queue is replayed to invalidate the entries in the TLB. For another example, if the reserved area stores only a single bit to indicate whether any invalidation request is received while the processor core is in the power-off state, if the value of the bit indicates that one or more invalidation requests are received, all entries in the TLB are invalidated.
[0041] Figure 5 is a block diagram of a cache hierarchy 500 according to some embodiments. The cache hierarchy 500 is used to implement Figure 1 1 and 140. Cache hierarchy 500 caches information (such as instructions or data) for processor cores 501, 502, 503, 504 (collectively referred to herein as "processor cores 501-504"). Processor cores 501-504 are used to implement Figure 1 Some implementations of processor cores 116-118, 141-143 are shown.
[0042] The cache hierarchy 500 includes three cache levels: a first level including L1 caches 511, 512, 513, 514 (collectively referred to herein as "L1 caches 511-514"), a second level including L2 caches 515, 516, 517, 518 (collectively referred to herein as "L2 caches 515-518"), and a third level including L3 cache 520. However, some embodiments of the cache hierarchy 500 include more or fewer cache levels. Although the L1 caches 511-514 are shown as separate hardware structures interconnected to the corresponding processor cores 501-504, some embodiments of the L1 caches 511-514 are incorporated into the hardware structure implementing the processor cores 501-504.
[0043] The L1 caches 511-514 are used to cache information for access by the corresponding processor cores 501-504. For example, the L1 cache 511 is configured to cache information for the processor core 501. Therefore, the processor core 501 issues a memory access request to the L1 cache 511. If the memory access request hits the L1 cache 511, the requested information is returned. If the memory access request misses the L1 cache 511, the L1 cache 511 forwards the memory access request to the next higher cache level (e.g., the L2 cache 515). The information cached in the L1 cache 511 is generally not accessible by other processor cores 502-504.
[0044] The L2 caches 515-518 are also configured to cache information for the processor cores 501-504. In the illustrated embodiment, the L2 caches 515-518 include corresponding L1 caches 511-514. For example, the L2 cache 515 caches information including information cached in the L1 cache 511. However, the L2 caches 515-518 are usually larger than the L1 caches 511-514 and therefore the L2 caches 515-518 also store other information that is not stored in the corresponding L1 caches 511-514. As discussed above, if one of the processor cores 501-504 issues a memory access request that misses the corresponding L1 cache 511-514, the memory access request is forwarded to the corresponding L2 cache 515-518. If the memory access request hits the L2 cache 515-518, the requested information is returned to the processor core 501-504 that issued the request. If the memory access request misses the L2 cache 515-518, the L2 cache 515-518 forwards the memory access request to the next higher level of cache (e.g., L3 cache 520). In some embodiments, the L2 cache 515-518 is shared between multiple L1 caches 511-514 and corresponding processor cores 501-504.
[0045] The L3 cache 520 is configured as a global cache for the processor cores 501-504. Memory access requests from the processor cores 501-504 that miss the L2 caches 515, 520 are forwarded to the L3 cache 520. If the memory access request hits the L3 cache 520, the requested information is returned to the requesting processor core 501-504. If the memory access request misses the L3 cache 520, the L3 cache 520 forwards the memory access request to a memory system, such as DRAM 525.
[0046] In the illustrated embodiment, processor cores 501-504, L1 caches 511-514, and L2 caches 515-518 are implemented in power domains 530, 531, 532, 533 (which are collectively referred to herein as "power domains 530-533"). Power is independently supplied to the power domains 530-533 and therefore entities in the power domains 530-533 are independently or individually powered on or off. For example, processor core 501 may be placed in a powered-off state while processor cores 502-504 remain in a powered-on state. However, removing power from processor cores 501-504 also removes power from the corresponding L1 caches 511-514 and L2 caches 515-518, so the corresponding L1 caches 511-514 and L2 caches 515-518 lose any stored information when the corresponding processor cores 501-504 enter the powered-off state.
[0047] In response to the corresponding processor core 501-504 starting to enter the power-off state, information representing the entries in the L1 cache 511-514 or the L2 cache 515-518 is stored in the reserved area. The reserved area continues to receive the reserved voltage during the power-off state of one or more of the processor cores 501-504. The reserved area can be implemented in the DRAM 525 or the L3 cache 520. If the reserved voltage is supplied when the processor core 501-504 is in the power-off state, the reserved area can also be implemented in the L1 cache 511-514 or the L2 cache 515-518. The information representing the entry may include a copy of the entry or the physical address of the information stored in the entry. For example, the information representing the entry in the L2 cache 515 is flushed by writing the modified (or dirty) entry from the L2 cache 515 to the reserved area implemented in the L3 cache 520. For another example, information representing entries in the L2 cache 515 is flushed by writing all entries in the L2 cache 515 to a reserved area implemented in the L3 cache 520 .
[0048] The reserved area also stores information representing a failure signal (such as a cache probe) received when one or more of the processor cores 501-504 is in a power-off state. In some embodiments, the information is stored in a shadow tag associated with an entry in the cache. For example, the L3 cache 520 stores shadow tags for entries in the L2 cache 515-518. The shadow tag contains information indicating whether the corresponding entry contains clean data or whether the entry is for a cache that is powered off together with one of the processor cores 501-504 that enters a power-off state. The shadow tag also contains one or more bits indicating whether the entry is valid, for example, whether a cache probe has been received for the corresponding entry. For another example, the reserved area implements a bit, which is set to a first value (e.g., 0) to indicate that a cache probe has not been received for the entry and a second value (e.g., 1) to indicate that one or more cache probes have been received for the entry. Some embodiments of the shadow tag include the physical address of the information stored in the entry. Some embodiments of the reserved area implement a queue to save the cache probe for subsequent replay. In addition to or instead of implementing cache probe bits in the shaded tags, queues are implemented.
[0049] The information in the reserved area is used to restore the cache in response to the corresponding processor core 501-504 starting to exit from the power-off state. For example, in response to the processor core 501 starting to exit from the power-off state, a copy of the entry in the L2 cache 515 is written back from the L3 cache 520. For another example, the physical address in the shadow tag stored in the L3 cache 520 is used to prefetch the value of the entry in the L2 cache 515. Information representing an invalidation signal is used to modify the cache entry. For example, if the value of a bit in the shadow tag for an entry in the L2 cache 515 indicates that a cache probe has been received for the entry, the bit value is used to invalidate the entry. The physical address of the invalid entry in the shadow tag is not prefetched. For another example, the cache probe is replayed from the queue to modify the entry in the L2 cache 515.
[0050] Figure 6 600 is a block diagram of a portion of a cache hierarchy including an L2 cache 605 and an L3 cache 610 according to some embodiments. The L2 cache 605 is a cache memory for a processor core such as Figure 5L3 cache 610 is used to implement a reserved area for L2 cache 605 because L3 cache 610 continues to receive power when the processor core associated with L2 cache 605 is in a powered-off state. In some embodiments, a copy of the entries in L2 cache 605 is written back to L3 cache 610 by flushing or refreshing L2 cache 605.
[0051] A reserved area in the L3 cache 610 stores shadow tags 615 associated with the L2 cache 605. The shadow tag 615 contains a physical address 620 of a value cached in an entry of the L2 cache 605. A bit value 625 indicates whether the entry contains unmodified (clean) data (a value of 1 indicates clean data), a bit value 630 indicates whether the entry is associated with a cache in a powered-off state (a value of 1 indicates association with a powered-off cache), and a bit value 635 indicates whether the entry is valid (a value of 1 indicates invalidation). The bit value 635 is modified in response to a probe hitting the corresponding entry, for example, the bit value 635 is set to a value of 1 to indicate that the entry has been invalidated by a probe hit. Figure 6 The shaded label 615 shown indicates that all entries contain clean data associated with the cache being in a powered-off state and that the cache probe has invalidated the cache entry associated with physical address P_ADDR_2.
[0052] In response to the processor core associated with L2 cache 605 beginning to exit from the powered-off state, the entries in L2 cache 605 are repopulated using information stored in L3 cache 610. For example, the physical address of the valid entry stored in shadow tag 615 is prefetched into the entry of L2 cache 605. For another example, if L3 cache 610 stores a copy of the value stored in the entry of L2 cache 605, the value is written back to L2 cache 605. Then, bit value 635 is used to invalidate the entry that received the cache probe while the processor core was in the powered-off state.
[0053] Some embodiments of the reserved area include a probe queue 640 that stores information indicating probes received when the processor core is in a powered-off state. After the L2 cache 605 has been refilled with information representing cache entries stored in the L3 cache 610, in response to the processor core being powered on, the probes stored in the probe queue 640 are sent to the L2 cache 605. Replaying the probes stored in the probe queue 640 invalidates the entries indicated by the probes to place the L2 cache 605 in an appropriate state. In the event that the probe queue 640 overflows, the processor core is powered on to service the probes in the probe queue 640 or the entire L2 cache 605 is invalidated when the processor core is powered on.
[0054] Figure 7 7 is a flow chart of a method 700 for storing information representing entries in a lower level cache in a reserved area before powering off a processor core associated with the lower level cache, according to some embodiments. Figure 1 The processing system 100 shown, Figure 5 The cache hierarchy 500 and Figure 6 As shown, a portion of a processing system 600 is implemented in some embodiments.
[0055] At block 705, the processor core is powered on and in normal operating mode. In the illustrated embodiment, shadow tags in the higher level cache are used as probe filters for cache lines in the lower level cache used by the processor core. The probe filter prevents probes for cache lines that are not stored in the lower level cache associated with the processor core from being sent to the processor core. Shadow tags are thus maintained when the processor core is operating in the power-on mode.
[0056] At block 710, the processing system initiates powering off of the processor core. For example, in response to the absence of instructions scheduled for execution by the processor core or a prediction that no instructions will be scheduled for execution by the processor core in a subsequent time interval exceeding a power-off threshold, the processor core begins entering a power-off state.
[0057] At block 715, the modified or dirty entries in the lower level cache are written back to a higher level cache implementing a reserved region that receives a reserved voltage when the processor core is in a powered-off state. For example, dirty entries in an L2 cache may be written back to an L3 cache. Shadow tags in the L3 cache associated with the entries in the L2 cache are also updated. In other examples, information representing entries in other caches, such as an L1 cache, may also be stored in the reserved region. Furthermore, the reserved region may be implemented in other entities including external memory, such as DRAM, other caches, or lower level caches if a reserved voltage is provided to the lower level caches when the multiprocessor core is in a powered-off state.
[0058] The processor core is powered off at block 720. Powering off the processor core occurs after information representing entries in the lower level cache has been stored to external memory to prevent this information from being lost when the lower level cache is powered off.
[0059] At block 725, a shadow tag in the reserved region is modified in response to a cache probe received while the processor core is in a powered-off state. For example, a bit in the shadow tag of the entry indicating whether the entry is valid is set to a value indicating that the entry is invalid in response to receiving a cache probe for the entry. For another example, the cache probe is added to a probe queue implemented in the reserved region.
[0060] Figure 8 is a flow chart of a method 800 for restoring entries in a lower level cache using information stored in a retention area in response to a processor core initiating an exit from a powered-down state according to some embodiments. Figure 1 The processing system 100 shown, Figure 5 The cache hierarchy 510 and Figure 6 As shown, a portion of a processing system 600 is implemented in some embodiments.
[0061] At block 805, the processing system initiates powering up of the processor core. For example, exiting from a powered-down state is initiated in response to a scheduler in the processing system scheduling an instruction for execution on the processor core.
[0062] At box 810, the lower level cache associated with the processor core is restored. In some embodiments, the lower level cache is refilled using the physical address in the shadow tag stored in the reserved area. For example, the L2 cache is refilled by prefetching the physical address of the valid entry in the shadow tag of the L3 cache. In this case, there is no need to modify the information in the L2 cache subsequently, because only valid entries are prefetched and the invalid entries in the shadow tag of the L3 cache are not prefetched into the L2 cache. For another example, the L2 cache is refilled by writing a copy of the entry from the L3 cache to the L2 cache. When the processor core is in power-off mode, a retention voltage is provided to some embodiments of the lower level cache. In this case, it is not necessary to refill the entries in the lower level cache in order to restore the lower level cache. For example, if the storage element of the L2 cache receives a retention voltage when the core is powered off, the L2 cache entry is retained in situ in the L2 cache and the L2 cache is restored based on the invalidation information.
[0063] In an embodiment implementing a reserved region of a probe queue, information provided to a lower level cache in response to a processor core beginning to exit from a power-off state is modified (at block 815) based on cache probes stored in the probe queue. For example, cache probes stored in the probe queue are replayed to invalidate corresponding entries in a refilled L2 cache. Block 815 is therefore optional (as indicated by the dashed line) and is not performed in some embodiments of method 800.
[0064] At block 820 , the processor core powers up and begins executing instructions based on the repopulated lower level cache.
[0065] In some embodiments, the apparatus and techniques described above are implemented in a system (such as the one or more integrated circuit (IC) devices (also referred to as integrated circuit packages or microchips) described above. Figures 1 to 8The electronic design automation (EDA) and computer-aided design (CAD) software tools can be used for the design and manufacture of these IC devices. These design tools are usually represented as one or more software programs. The one or more software programs contain codes that can be executed by a computer system to manipulate the computer system to operate the code representing the circuit of one or more IC devices so as to execute at least a part of the process for designing or adapting the manufacturing system to manufacture the circuit. This code may contain instructions, data, or a combination of instructions and data. Software instructions representing design tools or manufacturing tools are usually stored in a computer-readable storage medium that can be accessed by a computing system. Similarly, codes representing one or more stages of the design or manufacture of an IC device can be stored in the same computer-readable storage medium or different computer-readable storage media, and accessed from the same computer-readable storage medium or different computer-readable storage media.
[0066] Computer-readable storage media may include any non-transitory storage media or combination of non-transitory storage media that can be accessed by a computer system during use to provide instructions and / or data to the computer system. Such storage media may include, but are not limited to, optical media (e.g., compact discs (CDs), digital versatile discs (DVDs), Blu-ray discs), magnetic media (e.g., floppy disks, tapes, or magnetic hard drives), volatile memory (e.g., random access memory (RAM) or cache), non-volatile memory (e.g., read-only memory (ROM) or flash memory), or micro-electromechanical system (MEMS)-based storage media. Computer-readable storage media may be embedded in a computing system (e.g., system RAM or ROM), fixedly attached to a computing system (e.g., a magnetic hard drive), removably attached to a computing system (e.g., an optical disc or flash memory based on a universal serial bus (USB)), or connected to a computer system via a wired or wireless network (e.g., a network accessible storage (NAS)).
[0067] In some embodiments, certain aspects of the techniques described above may be implemented by one or more processors of a processing system that executes software. The software includes one or more executable instruction sets stored or otherwise tangibly embodied on a non-transitory computer-readable storage medium. The software may include instructions and certain data that manipulate one or more processors to perform one or more aspects of the above-mentioned techniques when executed by one or more processors. A non-transitory computer-readable storage medium may include, for example, a magnetic or optical disk storage device, a solid-state storage device such as a flash memory, a cache, a random access memory (RAM), or other one or more non-volatile memory devices, etc. The executable instructions stored on a non-transitory computer-readable storage medium may be source code, assembly language code, object code, or other instruction formats that are interpreted or otherwise executable by one or more processors.
[0068] It should be noted that not all activities or elements described above in the general description are necessary, a part of a particular activity or device may not be required, and one or more other activities may be performed in addition to those described, or one or more other elements may also be included. In addition, the order in which the activities are listed is not necessarily the order in which they are performed. Moreover, the concepts have been described with reference to specific embodiments. However, it will be appreciated by those skilled in the art that various modifications and changes may be made without departing from the scope of the present disclosure as set forth in the appended claims. Therefore, the description and the drawings should be viewed in an illustrative rather than a restrictive sense, and all such modifications are intended to be included within the scope of the present disclosure.
[0069] Benefits, other advantages and solutions to problems have been described above with respect to specific embodiments. However, benefits, advantages, solutions to problems, and any one or more features that may cause any benefit, advantage or solution to occur or become more obvious should not be interpreted as key, necessary or essential features of any or all claims. In addition, the specific embodiments disclosed above are merely illustrative, because the disclosed subject matter can be modified and practiced in different but equivalent ways that are understood by those skilled in the art who benefit from the teachings herein. Except as described in the appended claims, it is not intended to be limited to the details of the construction or design shown herein. Therefore, it is obvious that the specific embodiments disclosed above can be changed or modified, and all such changes are considered to be within the scope of the disclosed subject matter. Therefore, the protection sought herein is as set forth in the appended claims.
Claims
1. A method, comprising: In response to powering off a processor core associated with a first cache, storing information representing a set of entries of the first cache in a retention area of a processing system, the retention area receiving a retention voltage while the processor core is in a powered-off state; storing in the reserved area information indicating that at least one entry in the set of entries of the first cache has been invalidated while the processor core is in a powered-off state; implementing a probe queue in the reserved area to store cache probes received while the processor core is powered off; as well as In response to the processor core initiating an exit from the powered-down state, repopulating the first cache with entries that were not invalidated while the processor core was in the powered-down state using the stored information representing the set of entries, the cache probes stored in the probe queue, and the stored information indicating that the at least one entry has been invalidated.
2. The method of claim 1 , wherein the retention area comprises at least one of: a second cache including the first cache, an external memory storing information cached in the first cache, and a portion of the first cache that receives the retention voltage when the processor core is in the power-off state.
3. The method of any one of claims 1 or 2, wherein the first cache is a translation lookaside buffer (TLB) that caches virtual to physical address translations for the processor core.
4. The method of claim 3, wherein storing the information representing the entry of the TLB comprises storing the entry in the reserved area, and wherein restoring the entry of the TLB comprises providing the stored entry to the TLB.
5. The method of claim 3, wherein: Storing the information representing the entry of the TLB includes storing a virtual address of the entry in the TLB in the reserved area; and Repopulating the entry of the TLB includes prefetching the virtual address to start a page table walk that populates the entry in the TLB.
6. The method of claim 3, wherein storing information indicating that the at least one entry has been invalidated while the processor core is in a powered-off state comprises storing the information in a queue in response to receiving a signal to invalidate the at least one entry while the processor core is in the powered-off state.
7. The method of claim 6, further comprising: The TLB is invalidated in response to the queue overflow invalidation request.
8. The method of claim 1, wherein the first cache is a lower level cache in a cache hierarchy, the cache hierarchy containing a second cache, the second cache containing the first cache.
9. The method of claim 8, wherein storing the information representing the entry in the first cache comprises at least one of: flushing the first cache to write the modified value of the entry to the second cache or an external memory, and refreshing the first cache to write all values of the entry to the second cache or the external memory.
10. The method of claim 9, wherein: Storing the information representing the entry in the first cache includes storing a physical address of the entry in a shadow tag associated with the entry in the second cache or the external memory; and Storing information indicating that the at least one entry has been invalidated while the processor core is in a powered-off state includes storing the information in the shadow tag.
11. The method of claim 10, wherein refilling the first cache comprises prefetching valid entries in the first cache based on the physical address of the entry in the shadow tag and information indicating that at least one of the entries was invalidated while the processor core was in a powered-off state.
12. A device, comprising: a processor core configured to access information from a first cache; as well as a retention area that receives a retention voltage while the processor core is in a powered-off state and implements a probe queue to store cache probes received while the processor core is powered-off, wherein in response to the processor core entering the powered-off state, information representing entries of the first cache is stored in the retention area, wherein the retention area stores information indicating that at least one of the entries of the first cache was invalidated while the processor core was in the powered-off state; and wherein in response to the processor core initiating an exit from the power-off state, the first cache is repopulated with entries that were not invalidated while the processor core was in the power-off state using stored information representing a set of entries of the first cache, the cache probes stored in the probe queue, and stored information indicating that at least one entry was invalidated while the processor core was in the power-off state.
13. The apparatus of claim 12, wherein the retention area comprises at least one of: a second cache containing the first cache, an external memory storing information cached in the first cache, and a portion of the first cache that receives the retention voltage when the processor core is in the power-off state.
14. The apparatus of claim 12, wherein: The first cache is a translation lookaside buffer (TLB) that caches virtual to physical address translations for the processor core; and The device also includes: A queue is configured to store the information indicating that the at least one invalid entry has been invalidated while the processor core is in the powered-off state in response to receiving a signal to invalidate the at least one entry while the processor core is in the powered-off state.
15. The apparatus of claim 12, wherein the first cache is a lower level cache in a cache hierarchy, the cache hierarchy containing a second cache, the second cache containing the first cache.
Citation Information
Patent Citations
Flush Engine
US20140195737A1
Method and apparatus for storing a processor architectural state in cache memory
US20150081980A1
High performance shared memory for a bridge router supporting cache coherency
US6018763A