Prefetch Termination and Recovery in Instruction Cache
By speculatively determining the hit status of the virtual address and configuring the status in the storage controller subsystem, the access of multi-level cache is optimized, and the performance bottleneck problem caused by frequent access to multi-level caches in the prior art is solved, and more efficient data access is achieved.
Patent Information
- Application Number
- CN201980067439.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-08-14
- Filing Date
- 2019-08-14
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2039-09-27
AI Technical Summary
Existing memory systems require frequent access to multi-level caches when requesting data by the processor core, resulting in performance bottlenecks and additional access overhead.
Cache access is optimized by speculatively determining the hit or miss status of the virtual address in the storage controller subsystem and configuring the status as a valid or invalid state, avoiding additional access to TAGRAM or address translation logic.
Improves processor performance and efficiency, reduces access overhead to multi-level caches, and improves the speed of data access.
Smart Images

Figure CN112840330B_ABST
Abstract
Description
Background Art
[0001] Some memory systems include a multi-level cache system. When a memory controller receives a request for a specific memory address from a processor core, the memory controller determines whether the data associated with the memory address exists in a first-level cache (L1). If the data exists in the L1 cache, the data is returned from the L1 cache. If the data associated with the memory address does not exist in the L1 cache, the memory controller accesses a second-level cache (L2), which may be larger and thus store more data than the L1 cache. If the data exists in the L2 cache, the data is returned from the L2 cache to the processor core, and a copy is also stored in the L1 cache in the case where the same data is requested again. Additional levels of storage in the hierarchy are possible. Summary of the Invention
[0002] In one example, a system includes a processor having a CPU core, a first memory cache, a second memory cache, and a memory controller subsystem. The memory controller subsystem speculatively determines a hit or miss condition of a virtual address in the first memory cache and speculatively translates the virtual address to a physical address. Associated with the hit or miss condition and the physical address, the memory controller subsystem configures the state to a valid state. In response to receiving a first indication from the CPU core that a program instruction associated with the virtual address is not needed, the memory controller subsystem reconfigures the state to an invalid state, and in response to receiving a second indication from the CPU core that a program instruction associated with the first virtual address is needed, the memory controller subsystem reconfigures the state back to the valid state without additional access to a TAGRAM or address translation logic. Brief Description of the Drawings
[0003] Figure 1 Illustrates a processor according to an example.
[0004] Figure 2 Illustrates promoting an L1 memory cache access to a full L2 cache line access according to an example.
[0005] Figure 3 Is a flowchart illustrating a performance improvement according to an example.
[0006] Figure 4 Is another flowchart illustrating another performance improvement according to an example.
[0007] Figure 5 Shows a system including Figure 1 a processor. Detailed Description
[0008] Figure 1 An example of a processor 100 including a hierarchical cache subsystem is shown. In this example, the processor 100 includes a central processing unit (CPU) core 102, a memory controller subsystem 101, an L1 data cache (L1D) 115, an L1 program cache (L1P) 130, and an L2 memory cache 155. In this example, the memory controller subsystem 101 includes a data memory controller (DMC) 110, a program memory controller (PMC) 120, and a unified memory controller (UMC) 150. In this example, at the L1 cache level, data and program instructions are separated into separate caches. Instructions to be executed by the CPU core 102 are stored in the L1P 130 and then provided to the CPU core 102 for execution. On the other hand, data is stored in the L1D 115. The CPU core 102 can read data from or write data to the L1D 115, but has read access (no write access) to the L1P 130. The L2 memory cache 155 can store both data instructions and program instructions.
[0009] Although the sizes of the L1D 115, L1P 130, and L2 memory cache 155 can vary depending on the implementation, in one example, the size of the L2 memory cache 155 is larger than the size of the L1D 115 or L1P 130. For example, the size of the L1D 115 is 32 KB and the size of the L1P is also 32 KB, while the size of the L2 memory cache can be between 64 KB and 4 MB. In addition, the cache line size of the L1D 115 is the same as the cache line size of the L2 memory cache 155 (e.g., 128 bytes), and the cache line size of the L1P 130 is smaller (e.g., 64 bytes).
[0010] When the CPU core 102 needs data, the DMC 110 receives an access request for the target data from the CPU core 102. The access request may include an address (e.g., a virtual address) from the CPU core 102. The DMC 110 determines whether the target data exists in the L1D 115. If the data exists in the L1D 115, the data is returned to the CPU core 102. However, if the data requested by the CPU core 102 does not exist in the L1D 115, the DMC 110 provides the access request to the UMC 150. The access request may include a physical address generated by the DMC 110 based on the virtual address (VA) provided by the CPU core 102. The UMC 150 determines whether the physical address provided by the DMC 110 exists in the L2 memory cache 155. If the data exists in the L2 memory cache 155, the data is returned from the L2 memory cache 155 to the CPU core 102, where a copy is stored in the L1D 115. Additional hierarchies of the cache subsystem may also exist. For example, an L3 memory cache or system memory may be accessible. Thus, if the data requested by the CPU core 102 does not exist in either the L1D 115 or the L2 memory cache 155, the data can be accessed at an additional cache level.
[0011] Regarding program instructions, when the CPU core 102 needs additional instructions to execute, the CPU core 102 provides the VA 103 to the PMC 120. The PMC responds to the VA 103 provided by the CPU core 102 by initiating a workflow to return a prefetch data packet 105 of program instructions to the CPU core 102 for execution. Although the size of the prefetch data packet may vary depending on the implementation, in one example, the size of the prefetch data packet is equal to the size of a cache line of the L1P 130. If the L1P cache line size is, for example, 64 bytes, the prefetch data packet returned to the CPU core 102 will also contain 64 bytes of program instructions.
[0012] CPU core 102 also provides a prefetch count 104 to PMC 120. In some embodiments, the prefetch count 104 is provided to PMC 120 after CPU core 102 provides VA 103. The prefetch count 104 indicates the number of prefetch units of program instructions after the prefetch unit starting at VA 103. For example, CPU core 102 may provide a VA of 200h. This VA is associated with a 64-byte prefetch unit starting at virtual address 200h. If CPU core 102 desires that the memory controller subsystem 101 send additional instructions for execution after the prefetch unit associated with virtual address 200h, CPU core 102 submits a prefetch count with a value greater than 0. A prefetch count of 0 indicates that CPU core 102 no longer requires any prefetch units. For example, a prefetch count of 6 indicates that CPU core 102 requests an additional 6 prefetch units worth of fetch instructions to be sent back to CPU core 102 for execution. The returned prefetch units are shown as prefetch packets 105 in Figure 1 as shown in
[0013] Still referring to Figure 1 the example of, PMC 120 includes a TAGRAM (tag random access memory) 121, an address converter 122, and a register 123. TAGRAM 121 includes a list of virtual addresses whose contents (program instructions) have been cached to L1P 130. Address converter 122 converts the virtual address to a physical address (PA). In one example, address converter 122 directly generates the physical address from the virtual address. For example, the lower 12 bits of the VA can be used as the least significant 12 bits of the PA, where the most significant bits of the PA (above the lower 12 bits) are generated based on a set of tables configured in the main memory before the program is executed. In this example, the L2 memory cache 155 can be addressed using the physical address instead of the virtual address. Register 123 stores a hit / miss indicator 124 looked up from TAGRAM 121, a physical address 125 generated by address converter 122, and a valid bit 126 (also referred to herein as a status bit) to indicate whether the corresponding hit / miss indicator 124 and physical address 125 are valid or invalid.
[0014] After receiving VA 103 from CPU 102, PMC 120 performs a TAGRAM 121 lookup to determine whether L1P 130 includes program instructions associated with that virtual address. The result of the TAGRAM lookup is a hit or miss indicator 124. A hit means that the VA exists in L1P 130, and a miss means that the VA does not exist in L1P 130. For an L1P 130 hit, PMC 120 retrieves the target prefetch unit from L1P 130 and returns it to CPU core 102 as a prefetch packet 105.
[0015] For an L1P 130 miss, the PA (generated based on the VA) is provided by the PMC 120 to the UMC 150, as shown at 142. The byte count 140 is also provided from the PMC 120 to the UMC 150. The byte count indicates the number of bytes of the L2 memory cache 155 to be retrieved (if present) starting from the PA 142. In one example, the byte count 140 is a multi-bit signal that encodes the number of bytes required from the L2 memory cache 155. In the example, the line size of the L2 memory cache is 128 bytes, and each line is divided into an upper half (64 bytes) and a lower half (64 bytes). The byte count 140 can thus encode the number 64 (if only the upper or lower half 64 bytes of a given L2 memory cache line are needed) or 128 (if the entire L2 memory cache line is needed). In another example, the byte count can be a single-bit signal, where one state (e.g., 1) implicitly encodes the entire L2 memory cache line, and the other state (e.g., 0) implicitly encodes half of the L2 memory cache line.
[0016] The UMC 150 also includes a TAGRAM 152. The PA 142 received by the UMC 150 from the PMC 120 is used to perform a lookup in the TAGRAM 152 to determine whether the target PA is a hit or a miss in the L2 memory cache 155. If there is a hit in the L2 memory cache 155, depending on the byte count 140, the target information can be either half or the entire cache line, and this target information is returned to the CPU core 102 and a copy is stored in the L1P 130. The next time the CPU core 102 attempts to fetch the same program instruction, the same program instruction will be provided to the CPU core 102 from the L1P 130.
[0017] In Figure 1 the example, the CPU core 102 provides the VA 103 and the prefetch count 104 to the PMC 120. As described above, the PMC 120 initiates a workflow to retrieve the prefetch packets from the L1P 130 or the L2 memory cache 155. Using the prefetch count 104 and the original VA 103, the PMC 120 calculates additional virtual addresses and continues to retrieve the prefetch packets corresponding to those calculated VAs from the L1P 130 or the L2 memory cache 155. For example, if the prefetch count is 2 and the VA 103 from the CPU core 102 is 200h, the PMC 120 calculates the next two VAs as 240h and 280h, rather than the CPU core 102 providing each such VA to the PMC 120.
[0018] Figure 2Illustrates a specific example where optimization results in performance improvement of the processor 100. As described above, the line width of the L2 memory cache 155 is greater than that of the L1P. In one example, as Figure 2 shown, the width of the L1P is 64 bytes, and the line width of the L2 memory cache 155 is 128 bytes. The L2 memory cache 155 is organized into an upper half 220 and a lower half 225. The UMC 150 can read an entire 128 - byte cache line from the L2 memory cache 155, or only read half of the L2 memory cache line (the upper half 220 or the lower half 225).
[0019] A given VA can be translated to a specific PA. If the specific PA exists in the L2 memory cache 155, it is mapped to the lower half 225 or the upper half 220 of a given line of the L2 memory cache. Based on the addressing scheme used to represent the VA and PA, the PMC 120 can determine whether a given VA will be mapped to the lower half 225 or the upper half 220. For example, a specific bit (e.g., bit 6) within the VA can be used to determine whether the corresponding PA will be mapped to the upper or lower half of the L2 memory cache line. For example, bit 6 being 0 can indicate the lower half, while bit 6 being 1 can indicate the upper half.
[0020] Reference numeral 202 shows an example where the VA of 200h provided by the CPU core 102 to the PMC 120 and the corresponding pre - fetch count is 6. Reference numeral 210 illustrates that the list of VAs passing through the cache pipeline described above includes 200h (received from the CPU core 102) and the next 6 consecutive virtual addresses 240h, 280h, 2c0h, 300h, 340h, and 380h (calculated by the PMC 120).
[0021] As described above, each address from 200h to 380h is processed. Any or all of the VAs may be misses in the L1P 130. The PMC 120 can pack two consecutive VAs that miss in the L1P 130 into a single L2 cache line access attempt. Thus, if both 200h and 240h miss in the L1P 130, and the physical address corresponding to 200h corresponds to the lower half 225 of a specific cache line of the L2 memory cache 155, and the physical address corresponding to 240h corresponds to the upper half 225 in the same cache line of the L2 memory cache, then the PMC 120 can issue a single PA 142 and a byte count 140 to the UMC 150, where the byte count 140 specifies the entire cache line from the L2 memory cache. Thus, two consecutive VA misses in the L1P 130 can be promoted to a single full - line L2 memory cache lookup.
[0022] If the last VA in a series of VAs initiated by CPU core 102 (e.g., VA 380h in a series of VAs 210) maps to the lower half 225 of a cache line in L2 memory cache 155, then according to the example described, the entire cache line in L2 memory cache 155 is retrieved even if only the lower half 225 is needed. The same response occurs if the CPU supplies VA 103 to PMC 120 with a prefetch count of 0, which means the CPU 102 only needs a single prefetch unit. Any additional overhead, time, or power consumption (if any) incurred in retrieving the entire cache line and supplying it to L1P 130 is very small. Since program instructions are typically executed in a linear order, the probability that the program instructions in the upper half 220 will be executed anyway after the instructions in the lower half 225 is usually high. Therefore, the next instruction set is received at a very low cost and would likely be needed anyway.
[0023] Figure 2 The mapping of VA 380h to the lower half 225 of cache line 260 in L2 memory cache 155 is illustrated by arrow 213. PMC 120 determines this mapping by, for example, examining one or more bits of the VA or its corresponding physical address after translation by address converter 122. PMC 120 elevates the lookup process of UMC 150 to a full cache line read by submitting the PA associated with VA 380h and a byte count 104 specifying the entire cache line. The entire 128 - byte cache line (if present in L2 memory cache 155) is then retrieved and written to L1P 130 in two separate 64 - byte cache lines, as shown at 265.
[0024] However, if the last VA in a series of VAs (or if there is only one VA in the case of a prefetch count of 0) maps to the upper half 220 of a cache line in L2 memory cache 155, then PMC 120 requests UMC 150 to look up in its TAGRAM 152 and returns only the upper half of the cache line to CPU core 102 and L1P 130. The next PA will be in the lower half 225 of the next cache line in L2 memory cache 155, and it will take additional time, overhead, and power consumption to speculatively retrieve the next cache line, and it is uncertain whether CPU core 102 will need to execute those instructions.
[0025] Figure 3An example of a flowchart 300 for the above method is shown. The operations may be performed in the order shown or in a different order. Additionally, the operations may be performed sequentially, or two or more operations may be performed simultaneously.
[0026] At 302, the method includes receiving, by the storage controller subsystem 101, an access request for N prefetch units of program instructions. In one embodiment, this operation is performed by the CPU core 102 providing an address and a count value to the PMC 120. The address may be a virtual address or a physical address, and the count value may indicate the number of additional prefetch units required by the CPU core 102.
[0027] At 304, an index value I is initialized to the value 1. This index value is used to determine when the PMC 120 is to process the last virtual address in a series of consecutive virtual addresses. At 306, the method determines whether the prefetch unit I is a hit or a miss in the L1P 130. In some examples, this determination is made by determining whether the virtual address exists in the TAGRAM 121 of the PMC. The determination 306 can have two results - a hit or a miss.
[0028] If the virtual address is a hit in the L1P 130, then at 308, the corresponding line of the L1P 130 containing the desired prefetch unit is returned from the L1P 130 and provided to the CPU core 102 as a prefetch packet 105. Then the index is incremented at 310 (I = I + 1). If I has not reached N + 1 (as determined at the decision operation 312), then the last VA of the prefetch unit that has not been evaluated is used for the hit / miss determination, and the control loops back to 306 to evaluate the I-th prefetch unit below for a hit or a miss in the L1P 130. If I has reached N + 1, then all N prefetch units have been evaluated, and the corresponding program instructions have been provided to the CPU core 102, and the process stops.
[0029] For a given I-th prefetch unit, if the PMC 120 determines at 306 that there is a miss in the L1P 130, then at 314, it is determined whether I has reached the value of N. If I is not equal to N (indicating that the last VA in a series of VAs has not been reached), then at 316, the method includes the storage controller subsystem 101 obtaining the program instructions from the L2 memory cache 155 (if present) or from the third-level cache or system memory if not present. Then the index value I is incremented at 318, and the control loops back to the determination 306.
[0030] If, at 314, I has reached N (indicating that the last VA in a series of VAs has been reached), then at 320, the method includes determining whether the VA of the Ith prefetch unit is mapped to the lower half or the upper half of a cache line in the L2 memory cache 155. Examples of how this determination is made were described above. If the VA of the Ith prefetch unit is mapped to the upper half, then at 322, the method includes obtaining program instructions only from the upper half of the cache line in the L2 memory cache.
[0031] However, if the VA of the Ith prefetch unit is mapped to the lower half, then at 324, the method includes: promoting the L2 memory cache access to a full cache line access, and at 326, obtaining program instructions from the full cache line in the L2 memory cache.
[0032] Return reference Figure 1 , as described above, after submitting to the PMC 120 from the CPU core 102 to the VA 103, the CPU core 102 can also provide a prefetch count 104 to the PMC 120. The prefetch count can be 0, meaning that the CPU core 102 does not need any other instructions except those included in the prefetch units starting from VA103. However, between receiving the VA 103 and the subsequent prefetch count, the PMC 120 does some work as described below.
[0033] Upon receiving the VA 103, the PMC 120 performs a lookup in the TAGRAM 121 to determine whether the first VA (provided by the CPU core 102) is a hit or a miss in the L1P, and also performs a VA to PA conversion using the address converter 122. Before receiving the prefetch count 104, the PMC 120 also calculates a second VA (the next consecutive VA after the VA provided by the CPU core). The PMC 120 speculatively accesses the TAGRAM 121 and uses the address converter 122 to determine the hit / miss status of the second VA, and fills the register 123 with the hit / miss indication 124 and the PA 125. The valid bit 126 in the register 123 is set to the valid state, thus allowing further processing of the second VA, as described above (e.g., retrieving the corresponding cache line from the L1P 130 if it exists, or retrieving the corresponding cache line from the L2 memory cache 155 if necessary).
[0034] However, before any further processing of the second VA occurs, CPU core 102 may send a prefetch count of 0 to PMC 120, which means that the CPU core does not require any prefetch units other than the prefetch unit starting from the original VA 103. At this time, a prefetch count of 0 is provided to PMC 120, so no prefetch unit associated with the second VA is required. However, PMC has also determined the hit / miss status of the second VA and has generated the corresponding PA. When PMC 120 has received a zero prefetch count, both the hit / miss indicator 124 and the PA 125 have been stored in register 123. PMC 120 changes the status of the valid bit 126 to indicate an invalid state, thus precluding any further processing of the second VA. This condition (setting the valid bit to the invalid state) is referred to as "kill", so PMC 120 terminates the processing of the second VA.
[0035] However, in some cases, despite a previous termination, CPU core 102 may determine that a prefetch unit associated with the second VA should indeed be obtained from L1P130 or L2 memory cache 155, as described above. For example, if CPU core 102 has no further internal prediction information to inform the next required instruction address, CPU core 102 will signal to PMC120 that it should continue linearly prefetching starting from the last requested address. For example, this condition may occur due to a misprediction in the branch prediction logic in CPU core 102. Therefore, CPU core 102 issues a resume signal 106 to PMC 120 for this purpose. PMC 120 responds to the resume signal by changing the valid bit 126 back to the valid state, thus allowing the continued processing of the second VA through the memory subsystem pipeline, as described above. In this way, CPU 102 does not need to directly submit the second VA to PMC 120. Instead, PMC 120 retains the second VA in, for example, register 123 and its hit / miss indicator 124, thus avoiding the power consumption and time spent on re-determining the hit / miss status of the second VA and converting the second VA to a PA.
[0036] Figure 4 An example of flowchart 400 is shown, which is used to initiate, then terminate, and then resume a memory address lookup. The operations may be performed in the order shown or in a different order. Additionally, the operations may be performed sequentially, or two or more operations may be performed simultaneously.
[0037] At 402, the method includes receiving an access request at a first VA by the storage controller subsystem 101. In one embodiment, this operation is performed by the CPU core 102 that provides the first VA to the PMC 120. At 404, the method includes determining whether the first VA is a hit or a miss in the L1P 30. In one example, this operation is performed by accessing the TAGRAM 121 of the PMC to determine the hit / miss status of the first VA. At 406, the first VA is translated to a first PA by using, for example, the address translator 122.
[0038] At 408, the method includes calculating a second VA based on the first VA. The second VA can be calculated by incrementing the first VA by a value to generate the address of the byte 64 bytes after the byte associated with the first VA. At 410, the method includes determining whether the second VA is a hit or a miss in the L1P 30. In one example, this operation is performed by accessing the TAGRAM 121 of the PMC to determine the hit / miss status of the second VA. As described above, at 412, the second VA is translated to a second PA by using the address translator 122. At 414, the method includes updating a register (e.g., register 123) by using the hit / miss indicator 124 and the second PA. In addition, the valid bit 126 is configured to an active state.
[0039] Then at 416, the PMC 120 receives a prefetch count. Then at 418, if the prefetch count is greater than zero, then at 420, program instructions are retrieved from the L1P 130 or the L2 memory cache 155 (or one or more additional levels), as described above. However, if the prefetch count is zero, then at 422, the valid bit 126 is changed to an inactive state. Although a prefetch count of zero has been provided to the PMC 120, the CPU core 102 can then provide a resume indication to the PMC 120 (at 424). At 426, the PMC 120 changes the valid bit 126 back to the active state, and the storage controller subsystem 101 then appropriately obtains the program instructions associated with the second PA from the L1P, the L2 memory cache, etc.
[0040] Figure 5An example use of the processor 100 described herein is shown. In this example, the processor 100 is part of a system-on-chip (SoC) 500 that includes the processor 100 and one or more peripheral ports or devices. In this example, the peripheral devices include a universal asynchronous receiver transmitter (UART) 502, a universal serial bus (USB) port 504, and an Ethernet controller 506. The SoC 500 can perform any of a variety of functions, such as those implemented by program instructions executed by the processor 100. More than one processor 100 can be provided, and within a given processor 100, more than one CPU core 102 can be included.
[0041] In this specification, the term "coupled" means an indirect or direct wired or wireless connection. Thus, if a first device is coupled to a second device, the connection can be through a direct connection or through an indirect connection via other devices and connections. Additionally, in this specification, the phrase "based on" means "at least in part based on". Thus, if X is based on Y, X can be a function of Y and any number of other factors.
[0042] Within the scope of the claims, modifications may be made in the described embodiments and also in other embodiments.
Claims
1. A device, which comprises: a central processing unit core, namely a CPU core; a first memory cache that stores instructions executed by the CPU core; a second memory cache that stores instructions executed by the CPU core, and the second memory cache is accessible in response to a miss in the first memory cache; a memory controller subsystem coupled to the CPU core and the first and second memory caches, the memory controller subsystem being configured to: determine a miss or a hit in the first memory cache for a first virtual address received from the CPU core; generate a second virtual address based on the first virtual address; determine a miss or a hit in the first memory cache for the second virtual address; convert the second virtual address to a physical address; set status bits associated with the physical address and the determination of a miss or a hit of the second virtual address to a valid state; change the status bits to an invalid state in response to receiving a count value of zero from the CPU core; and change the status bits back to the valid state in response to subsequently receiving a resume indication from the CPU core.
2. The device according to claim 1, wherein the memory controller subsystem is configured to retrieve program instructions from the second memory cache using the physical address converted from the second virtual address.
3. The device according to claim 1, wherein the receipt of the count value will occur after the second virtual address is converted to the physical address.
4. The device according to claim 1, wherein the receipt of the count value from the CPU core will occur before the receipt of the resume indication from the CPU core.
5. The device according to claim 1, which further comprises a register, and the physical address and the status bits are stored in the register.
6. The device according to claim 5, wherein an indication of a hit or a miss of the second virtual address in the first memory cache is stored in the register together with the physical address and the status bits.
7. The device according to claim 1, wherein the first memory cache will store program instructions instead of data.
8. A device, which comprises: a central processing unit core, namely a CPU core; a first memory cache that stores instructions executed by the CPU core; a second memory cache that stores instructions executed by the CPU core, and the second memory cache retrieves instructions in response to a miss in the first memory cache; and a memory controller subsystem coupled to the CPU core and the first and second memory caches, the memory controller subsystem being configured to: receive a first virtual address; determine a hit or a miss condition of a second virtual address in the first memory cache based on the first virtual address; convert the second virtual address to a physical address; Associate with the hit or miss condition and the physical address, and configure the state to a valid state; In response to receiving a first indication from the CPU core that a program instruction not associated with the second virtual address is not needed, reconfigure the state to an invalid state; And In response to receiving a second indication from the CPU core that a program instruction associated with the second virtual address is needed, reconfigure the state back to a valid state.
9. The apparatus according to claim 8, wherein, The memory controller subsystem is configured to generate the second virtual address based on the first virtual address.
10. The apparatus according to claim 8, wherein, The first indication of a program instruction from the CPU core that is not associated with the second virtual address includes a count value, and the count value has a zero value.
11. The apparatus according to claim 8, wherein, The second indication of a program instruction from the CPU core that is associated with the second virtual address includes a signal instructing the memory controller subsystem to continue retrieving program instructions starting from the second virtual address.
12. The apparatus according to claim 11, wherein, Upon receiving the second indication, the memory controller subsystem is configured to continue retrieving program instructions starting from the second virtual address without determining the hit or miss condition of the second virtual address in the first memory cache again.
13. The apparatus according to claim 12, wherein, Upon receiving the second indication, the memory controller subsystem is configured to continue retrieving program instructions starting from the second virtual address without converting the second virtual address to the physical address again.
14. The apparatus according to claim 8, wherein, The CPU core is configured to provide the second indication without providing the second virtual address to the memory controller subsystem.
15. The apparatus according to claim 8, wherein, Receiving the first indication will occur after determining the hit or miss condition and converting the second virtual address to the physical address.
16. A system-on-chip (SoC), which comprises: Input / output devices; And A processor coupled to the input / output devices and including a central processing unit core (CPU core), a first memory cache storing instructions executed by the CPU core, a second memory cache, and a memory controller subsystem coupled to the CPU core, the first memory cache, and the second memory cache, the memory controller subsystem being configured to: Before receiving an indication to retrieve program instructions from a first virtual address: Determine the hit or miss condition of the first virtual address in the first memory cache; and Convert the first virtual address to a physical address; Associate with the hit or miss condition and the physical address, and configure the state to a valid state; In response to receiving a first indication from the CPU core that a program instruction not associated with the first virtual address is not required, reconfigure the state to an invalid state; and In response to receiving a second indication from the CPU core that a program instruction associated with the first virtual address is required, reconfigure the state back to a valid state.
17. The SoC according to claim 16, wherein, the memory controller subsystem is configured to generate the first virtual address based on a second virtual address transmitted from the CPU core to the memory controller subsystem.
18. The SoC according to claim 16, wherein, the first indication of a program instruction from the CPU core that is not associated with the first virtual address includes a count value having a zero value.
19. The SoC according to claim 16, wherein, the CPU core is configured to provide the second indication without also providing the first virtual address to the memory controller subsystem.
20. The SoC according to claim 16, wherein, upon receiving the second indication, the memory controller subsystem is configured to continue retrieving program instructions starting from the first virtual address without again determining the hit or miss condition of the first virtual address in the first memory cache and without again translating the first virtual address to the physical address.
Citation Information
Patent Citations
Power Management in Federated / Distributed Shared Memory Architecture
US20090193270A1
Intelligent cache memory and prefetch method based on CPU data fetching characteristics
US5361391A