Prefetch management in hierarchical cache systems

By designing two-level memory caches in the memory system and retrieving the second-level entire row of data when the first level misses, the low performance problem caused by multiple access to different levels of caches in the prior art is solved, and more efficient data access and processor performance improvements are achieved.

CN112840331BActive Publication Date: 2025-05-16TEXAS INSTRUMENTS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201980067440.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-08-14
Filing Date
2019-08-14
Publication Date
2025-05-16
Estimated Expiration
2039-08-14

AI Technical Summary

Technical Problem

When the processor core requests data, existing memory systems need to access different levels of caches multiple times, resulting in poor performance and increased latency.

Method used

An apparatus is designed including a central processing unit (CPU) core and a two-level memory cache, the first memory cache having a fixed row size, the second memory cache having a row size greater than the first level, and each row is divided into an upper and lower half. When the first level cache misses, the storage controller subsystem determines that the target address maps to the lower half of the second level cache, and retrieves the entire row of data and returns it to the first level cache.

Benefits of technology

By improving access to the second-level cache, it can improve data access speed and processor performance, reduce latency and improve system efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112840331B_ABST
    Figure CN112840331B_ABST
Patent Text Reader

Abstract

A device includes a CPU core (102), a first memory cache (130) having a first row size, and a second memory cache (155) having a second row size greater than the first row size. Each row of the second memory cache (155) includes an upper half and a lower half. A memory controller subsystem (101) is coupled to the CPU core (102) and the first memory cache (130) and the second memory cache (155). When a first target address is missed in the first memory cache (130), the memory controller subsystem (101) determines that the first target address that caused the miss is mapped to the lower half of the row in the second memory cache (155), retrieves the entire row from the second memory cache (155), and returns the entire row from the second memory cache (155) to the first memory cache (130).
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Some memory systems include a multi-level cache system. When a memory controller receives a request for a particular memory address from a processor core, the memory controller determines whether the data associated with the memory address exists in a first level cache (L1). If the data exists in the L1 cache, the data is returned from the L1 cache. If the data associated with the memory address does not exist in the L1 cache, the memory controller accesses a second level cache (L2), which may be larger and therefore hold more data than the L1 cache. If the data exists in the L2 cache, the data is returned from the L2 cache to the processor core, and in the event that the same data is requested again, a copy is also stored in the L1 cache. Additional storage levels of the hierarchy are possible. Summary of the invention

[0002] In at least one example, an apparatus includes a central processing unit (CPU) core and a first memory cache that stores instructions executed by the CPU core. The first memory cache is configured to have a first row size. A second memory cache stores instructions executed by the CPU core. The second memory cache has a second row size that is larger than the first row size, and each row of the second memory cache includes an upper half and a lower half. A storage controller subsystem is coupled to the CPU core and the first memory cache and the second memory cache. When a first target address in the first memory cache is missed, the storage controller subsystem determines that the first target address that caused the miss is mapped to the lower half of a row in the second memory cache, retrieves the entire row from the second memory cache, and returns the entire row from the second memory cache to the first memory cache. BRIEF DESCRIPTION OF THE DRAWINGS

[0003] Figure 1 A processor according to an example is described.

[0004] Figure 2 Promoting an L1 memory cache access to a full L2 cache line access according to an example is described.

[0005] Figure 3 is a flowchart illustrating performance improvements based on examples.

[0006] Figure 4 is another flow chart illustrating another performance improvement according to an example.

[0007] Figure 5 Shown include Figure 1 processor system. DETAILED DESCRIPTION

[0008] Figure 1 An example of a processor 100 including a hierarchical cache subsystem is shown. The processor 100 in this example includes a central processing unit (CPU) core 102, a memory controller subsystem 101, an L1 data cache (L1D) 115, an L1 program cache (L1P) 130, and an L2 memory cache 155. In this example, the memory controller subsystem 101 includes a data storage controller (DMC) 110, a program memory controller (PMC) 120, and a unified memory controller (UMC) 150. In this example, at the L1 cache level, data and program instructions are divided into separate caches. Instructions executed by the CPU core 102 are stored in the L1P 130, which are then provided to the CPU core 102 for execution. On the other hand, data is stored in the L1D 115. The CPU core 102 can read data from or write data to the L1D 115, but has read access rights (no write access rights) to the L1P 130. L2 memory cache 155 may store both data instructions and program instructions.

[0009] Although the sizes of L1D 115, L1P 130, and L2 memory cache 155 may vary depending on the implementation, in one example, the size of L2 memory cache 155 is larger than the size of L1D 115 or L1P 130. For example, the size of L1D 115 is 32KB and the size of L1P is also 32KB, while the size of L2 memory cache may be between 64KB and 4MB. In addition, the cache line size of L1D 115 is the same as the cache line size of L2 memory cache 155 (e.g., 128 bytes), and the cache line size of L1P 130 is smaller (e.g., 64 bytes).

[0010] When the CPU core 102 needs data, the DMC 110 receives an access request for the target data from the CPU core 102. The access request may include an address (e.g., a virtual address) from the CPU core 102. The DMC 110 determines whether the target data exists in the L1D 115. If the data exists in the L1D 115, the data is returned to the CPU core 102. However, if the data requested by the CPU core 102 does not exist in the L1D 115, the DMC 110 provides an access request to the UMC 150. The access request may include a physical address generated by the DMC 110 based on the virtual address (VA) provided by the CPU core 102. The UMC 150 determines whether the physical address provided by the DMC 110 exists in the L2 memory cache 155. If the data exists in the L2 memory cache 155, the data is returned from the L2 memory cache 155 to the CPU core 102, where a copy is stored in the L1D 115. Additional hierarchies of cache subsystems may also exist. For example, L3 memory cache or system memory may be accessible. Thus, if the data requested by CPU core 102 is not present in L1D 115 or L2 memory cache 155, the data may be accessed in an additional cache level.

[0011] With respect to program instructions, when CPU core 102 requires additional instructions to execute, CPU core 102 provides VA 103 to PMC 120. PMC responds to VA 103 provided by CPU core 102 by initiating a workflow to return a prefetch data packet 105 of program instructions to CPU core 102 for execution. Although the size of the prefetch data packet may vary depending on the implementation, in one example, the size of the prefetch data packet is equal to the size of a cache line of L1P 130. If the L1P cache line size is, for example, 64 bytes, the prefetch data packet returned to CPU core 102 will also contain 64 bytes of program instructions.

[0012] The CPU core 102 also provides a prefetch count 104 to the PMC 120. In some embodiments, the prefetch count 104 is provided to the PMC 120 after the CPU core 102 provides the VA 103. The prefetch count 104 indicates the number of prefetch units for program instructions following the prefetch unit starting at VA 103. For example, the CPU core 102 may provide a VA of 200h. This VA is associated with a 64-byte prefetch unit starting at virtual address 200h. If the CPU core 102 wants the storage controller subsystem 101 to send additional instructions for execution after the prefetch unit associated with virtual address 200h, the CPU core 102 submits a prefetch count with a value greater than 0. A prefetch count of 0 indicates that the CPU core 102 no longer requires any prefetch units. For example, a prefetch count of 6 indicates that the CPU core 102 requests an additional 6 prefetch units worth of instructions to be fetched and sent back to the CPU core 102 for execution. The prefetch units are returned at Figure 1 Shown in FIG. 1 is a pre-fetched data packet 105 .

[0013] Still reference Figure 1 In the example of , PMC 120 includes TAGRAM (Tagged Random Access Memory) 121, address converter 122 and register 123. TAGRAM 121 includes a list of virtual addresses, and the content of the virtual address (program instructions) has been cached to L1P 130. Address converter 122 converts the virtual address into a physical address (PA). In one example, address converter 122 generates a physical address directly from the virtual address. For example, the lower 12 bits of VA can be used as the least significant 12 bits of PA, wherein the most significant bit of PA (located above the lower 12 bits) is generated based on a set of tables configured in the main memory before executing the program. In this example, physical addresses can be used instead of virtual addresses to address L2 memory cache 155. Register 123 stores a hit / miss indicator 124 from TAGRAM 121 lookup, a physical address 125 generated by address converter 122, and a valid bit 126 (also referred to as a status bit herein) to indicate whether the corresponding hit / miss indicator 124 and physical address 125 are valid or invalid.

[0014] After receiving VA 103 from CPU 102, PMC 120 performs a TAGRAM 121 lookup to determine whether L1P 130 includes a program instruction associated with the virtual address. The result of the TAGRAM lookup is a hit or miss indicator 124. A hit means that the VA is present in L1P 130, while a miss means that the VA is not present in L1P 130. For an L1P 130 hit, PMC 120 retrieves the target prefetch unit from L1P 130 and returns it to CPU core 102 as a prefetch packet 105.

[0015] For an L1P 130 miss, the PA (generated based on the VA) is provided by the PMC 120 to the UMC 150 as shown at 142. A byte count 140 is also provided from the PMC 120 to the UMC 150. The byte count indicates the number of bytes from the L2 memory cache 155 to be retrieved (if present) starting from the PA 142. In one example, the byte count 140 is a multi-bit signal that encodes the number of bytes required from the L2 memory cache 155. In an example, the line size of the L2 memory cache is 128 bytes, and each line is divided into an upper half (64 bytes) and a lower half (64 bytes). The byte count 140 may therefore encode the number 64 (if only the upper half or lower half 64 bytes are required for a given L2 memory cache line) or 128 (if the entire L2 memory cache line is required). In another example, the byte count may be a single bit signal where one state (eg, 1) implicitly encodes an entire L2 memory cache line, while another state (eg, 0) implicitly encodes half of an L2 memory cache line.

[0016] The UMC 150 also includes a TAGRAM 152. The PA 142 received by the UMC 150 from the PMC 120 is used to perform a lookup on the TAGRAM 152 to determine whether the target PA is a hit or a miss in the L2 memory cache 155. If there is a hit in the L2 memory cache 155, the target information may be half of a cache line or a full cache line depending on the byte count 140, which is returned to the CPU core 102, a copy is stored in the L1P 130, and the same program instruction will be provided to the CPU core 102 from the L1P 130 the next time the CPU core 102 attempts to fetch the same program instruction.

[0017] exist Figure 1 In the example of , the CPU core 102 provides the VA 103 and the prefetch count 104 to the PMC 120. As described above, the PMC 120 initiates a workflow to retrieve prefetch packets from the L1P 130 or the L2 memory cache 155. Using the prefetch count 104 and the original VA 103, the PMC 120 calculates additional virtual addresses and continues to retrieve prefetch packets corresponding to those calculated VAs from the L1P 130 or the L2 memory cache 155. For example, if the prefetch count is 2, and the VA 103 from the CPU core 102 is 200h, the PMC 120 calculates the next two VAs as 240h and 280h, rather than the CPU core 102 providing each such VA to the PMC 120.

[0018] Figure 2Specific examples are described where the optimization results in improved performance of the processor 100. As described above, the line width of the L2 memory cache 155 is greater than the line width of the L1P. In one example, Figure 2 As shown, the width of the L1P is 64 bytes and the line width of the L2 memory cache 155 is 128 bytes. The L2 memory cache 155 is organized into an upper half 220 and a lower half 225. The UMC 150 can read an entire 128-byte cache line from the L2 memory cache 155, or only half of the L2 memory cache line (the upper half 220 or the lower half 225).

[0019] A given VA may be translated into a specific PA, and if the specific PA is present in the L2 memory cache 155, mapped to the lower half 225 or mapped to the upper half 220 of a given line of the L2 memory cache. Based on the addressing scheme used to represent the VA and the PA, the PMC 120 may determine whether a given VA will be mapped to the lower half 225 or the upper half 220. For example, a specific bit (e.g., bit 6) within the VA may be used to determine whether the corresponding PA will be mapped to the upper half or the lower half of the L2 memory cache line. For example, a bit 6 of 0 may indicate the lower half, while a bit 6 of 1 may indicate the upper half.

[0020] Reference numeral 202 shows an example where the VA provided by the CPU core 102 to the PMC 120 is 200h and the corresponding prefetch count is 6. Reference numeral 210 illustrates that the VA list through the cache pipeline includes 200h (received from the CPU core 102) and the next 6 consecutive virtual addresses 240h, 280h, 2c0h, 300h, 340h and 380h (calculated by the PMC 120).

[0021] As described above, each address from 200h to 380h is processed. Any or all VAs may be misses in the L1P 130. The PMC 120 may pack two consecutive VAs that miss in the L1P 130 into a single L2 cache line access attempt. Thus, if both 200h and 240h miss in the L1P 130, and the physical address corresponding to 200h corresponds to the lower half 225 of a particular cache line of the L2 memory cache 155, and the physical address corresponding to 240h corresponds to the upper half 225 in the same cache line of the L2 memory cache, then the PMC 120 may issue a single PA 142 to the UMC 150 along with a byte count 140 that specifies an entire cache line from the L2 memory cache. Thus, two consecutive VA misses in the L1P 130 may be promoted to a single full-line L2 memory cache lookup.

[0022] If the last VA in a series of VAs initiated by the CPU core 102 (e.g., VA 380h in a series of VAs 210) maps to the lower half 225 of a cache line of the L2 memory cache 155, then according to the described example, even if only the lower half 225 is needed, the entire cache line of the L2 memory cache 155 is retrieved. If the CPU provides VA 103 to the PMC 120 with a prefetch count of 0, which means that the CPU 102 only needs a single prefetch unit, the same response occurs. Any additional overhead, time, or power consumption (if any) spent on retrieving the entire cache line and providing the entire cache line to the L1P 130 is very small. Since program instructions are usually executed in a linear order, the probability that the program instructions in the upper half 220 will be executed after the instructions in the lower half 225 are executed is usually high. Therefore, the next instruction set is received at very little cost, and such instructions may be needed anyway.

[0023] Figure 2 The mapping of VA 380h to the lower half 225 of cache line 260 in L2 memory cache 155 is illustrated by arrow 213. The PMC 120 determines this mapping by, for example, examining one or more bits of the VA or its corresponding physical address after translation by address translator 122. The PMC 120 promotes the lookup process of the UMC 150 to a full cache line read by submitting the PA associated with VA 380h and the byte count 104 specifying the entire cache line. The entire 128-byte cache line is then retrieved (if present in L2 memory cache 155) and written to the L1P 130 in two separate 64-byte cache lines, as shown at 265.

[0024] However, if the last VA in the series of VAs (or only one VA if the prefetch count is 0) maps to the upper half 220 of a cache line in the L2 memory cache 155, the PMC 120 requests the UMC 150 to look in its TAGRAM 152 and return only the upper half of the cache line to the CPU core 102 and L1P 130. The next PA will be located in the lower half 225 of the next cache line in the L2 memory cache 155, and additional time, overhead, and power consumption will be spent speculatively retrieving the next cache line, and it is uncertain whether the CPU core 102 will need to execute those instructions.

[0025] Figure 3An example of a flow chart 300 for the above method is shown. The operations may be performed in the order shown or in a different order. In addition, the operations may be performed sequentially, or two or more operations may be performed simultaneously.

[0026] At 302, the method includes receiving, by the storage controller subsystem 101, a request for access to N prefetch units for program instructions. In one embodiment, this operation is performed by the CPU core 102 providing an address and a count value to the PMC 120. The address may be a virtual address or a physical address, and the count value may indicate the number of additional prefetch units that the CPU core 102 requires.

[0027] At 304, an index value I is initialized to a value of 1. This index value is used to determine when the PMC 120 is about to process the last virtual address in a series of consecutive virtual addresses. At 306, the method determines whether prefetch unit I is a hit or a miss for L1P 130. In some examples, this determination is made by determining whether the virtual address is present in TAGRAM 121 of the PMC. Determination 306 may have two results—a hit or a miss.

[0028] If the virtual address is a hit into the L1P 130, then at 308, the corresponding line of the L1P 130 containing the desired prefetch unit is returned from the L1P 130 and provided to the CPU core 102 as a prefetch data packet 105. The index is then incremented (I=I+1) at 310. If I has not yet reached N+1 (as determined at decision operation 312), then the last VA of the prefetch unit has not been evaluated for a hit / miss determination, and control loops back to 306 to evaluate the next I-th prefetch unit for a hit or miss in the L1P 130. If I has reached N+1, then all N prefetch units have been evaluated and the corresponding program instructions have been provided to the CPU core 102, and the process stops.

[0029] For a given I-th prefetch unit, if the PMC 120 determines at 306 that there is a miss in the L1P 130, then at 314 it is determined whether I has reached the value of N. If I is not equal to N (indicating that the last VA in the series of VAs has not been reached), then at 316 the method includes the memory controller subsystem 101 obtaining the program instruction from the L2 memory cache 155 (if present or from the third level cache or system memory if not present). The index value I is then incremented at 318 and control loops back to determination 306.

[0030] If at 314, I has reached N (indicating that the last VA in the series of VAs has been reached), then at 320, the method includes determining whether the VA of the I-th prefetch unit is mapped to the lower half or the upper half of a cache line of the L2 memory cache 155. An example of how this determination may be made is described above. If the VA of the I-th prefetch unit is mapped to the upper half, then at 322, the method includes obtaining program instructions only from the upper half of a cache line of the L2 memory cache.

[0031] However, if the VA of the Ith prefetch unit is mapped to the lower half, then at 324, the method includes: promoting the L2 memory cache access to a full cache line access, and at 326, obtaining the program instruction from the full cache line of the L2 memory cache.

[0032] Return to reference Figure 1 As described above, after submission from CPU core 102 to PMC 120 of VA 103, CPU core 102 may also provide prefetch count 104 to PMC 120. The prefetch count may be 0, meaning that CPU core 102 does not require any instructions other than those contained in the prefetch unit starting from VA 103. However, between receipt of VA 103 and the subsequent prefetch count, PMC 120 does some work as described below.

[0033] Upon receiving VA 103, PMC 120 performs a lookup in TAGRAM 121 to determine whether the first VA (provided by CPU core 102) is a hit or miss in L1P, and also performs a VA to PA translation using address translator 122. PMC 120 also calculates a second VA (the next consecutive VA after the VA provided by the CPU core) before receiving prefetch count 104. PMC 120 speculatively accesses TAGRAM 121 and uses address translator 122 to determine the hit / miss status of the second VA and populates register 123 with hit / miss indication 124 and PA 125. Valid bit 126 in register 123 is set to a valid state, thereby allowing further processing of the second VA as described above (e.g., retrieving a corresponding cache line from L1P 130 if present, or retrieving a corresponding cache line from L2 memory cache 155 if necessary).

[0034] However, before any further processing of the second VA occurs, the CPU core 102 may send a prefetch count of 0 to the PMC 120, meaning that the CPU core does not require any prefetch units other than the prefetch units starting from the original VA 103. At this point, the prefetch count of 0 is provided to the PMC 120, so no prefetch units associated with the second VA are required. However, the PMC has also determined the hit / miss status of the second VA and has generated a corresponding PA. When the PMC 120 has received the zero prefetch count, both the hit / miss indicator 124 and the PA 125 have been stored in the register 123. The PMC 120 changes the state of the valid bit 126 to indicate an invalid state, thereby precluding any further processing of the second VA. This condition (setting the valid bit to an invalid state) is called a "kill", so the PMC 120 terminates processing of the second VA.

[0035] However, in some cases, despite the previous termination, the CPU core 102 may determine that the prefetch unit associated with the second VA should indeed be obtained from the L1P 130 or L2 memory cache 155, as described above. For example, if the CPU core 102 has no further internal prediction information to inform the next required instruction address, the CPU core 102 will signal the PMC 120 that it should continue prefetching linearly from the last requested address. For example, this situation may occur due to a misprediction in the branch prediction logic in the CPU core 102. Therefore, the CPU core 102 issues a resume signal 106 to the PMC 120 for this purpose. The PMC 120 responds to the resume signal by changing the valid bit 126 back to a valid state, thereby allowing continued processing of the second VA through the memory subsystem pipeline, as described above. In this way, the CPU 102 does not need to submit the second VA directly to the PMC 120. Instead, PMC 120 retains the second VA in, for example, register 123 and its hit / miss indicator 124, thereby avoiding power consumption and time spent in again determining the hit / miss status of the second VA and converting the second VA to a PA.

[0036] Figure 4 An example of a flow chart 400 is shown for initiating, then terminating, and then resuming a memory address lookup. The operations may be performed in the order shown or in a different order. In addition, the operations may be performed sequentially, or two or more operations may be performed simultaneously.

[0037] At 402, the method includes receiving, by the storage controller subsystem 101, an access request at a first VA. In one embodiment, the operation is performed by the CPU core 102 providing the first VA to the PMC 120. At 404, the method includes determining whether the first VA is a hit or a miss in the L1P 30. In one example, the operation is performed by accessing the TAGRAM 121 of the PMC to determine the hit / miss status of the first VA. The first VA is converted to a first PA at 406 by using, for example, the address translator 122.

[0038] At 408, the method includes calculating a second VA based on the first VA. The second VA may be calculated by incrementing the first VA by a value to generate an address of a byte 64 bytes after the byte associated with the first VA. At 410, the method includes determining whether the second VA is a hit or a miss in L1P 30. In one example, the operation is performed by accessing TAGRAM 121 of the PMC to determine the hit / miss status of the second VA. As described above, at 412, the second VA is converted to a second PA using address converter 122. At 414, the method includes updating a register (e.g., register 123) using hit / miss indicator 124 and the second PA. In addition, valid bit 126 is configured to a valid state.

[0039] The PMC 120 then receives the prefetch count at 416. Then, at 418, if the prefetch count is greater than zero, then at 420, program instructions are retrieved from the L1P 130 or L2 memory cache 155 (or one or more additional levels), as described above. However, if the prefetch count is zero, then at 422, the valid bit 126 is changed to an invalid state. Although the prefetch count of zero has been provided to the PMC 120, the CPU core 102 may then provide a resume indication (at 424) to the PMC 120. At 426, the PMC 120 changes the valid bit 126 back to a valid state, and the storage controller subsystem 101 then obtains the program instructions associated with the second PA from the L1P, L2 memory cache, etc., as appropriate.

[0040] Figure 5An example use of the processor 100 described herein is shown. In this example, the processor 100 is part of a system on a chip (SoC) 500, which includes the processor 100 and one or more peripheral ports or devices. In this example, the peripherals include a universal asynchronous receiver transmitter (UART) 502, a universal serial bus (USB) port 504, and an Ethernet controller 506. The SoC 500 may perform any of a variety of functions, such as implemented by program instructions executed by the processor 100. More than one processor 100 may be provided, and within a given processor 100, more than one CPU core 102 may be included.

[0041] In this specification, the term "coupled" refers to an indirect or direct wired or wireless connection. Thus, if a first device is coupled to a second device, the connection may be through a direct connection or through an indirect connection via other devices and connections. Additionally, in this specification, the expression "based on" means "based at least in part on." Thus, if X is based on Y, then X may be a function of Y and any number of other factors.

[0042] Modifications may be made in the described embodiments, as well as in other embodiments, within the scope of the claims.

Claims

1. A device comprising: The central processing unit core is the CPU core; a first memory cache storing instructions executed by the CPU core, the first memory cache having a first line size; a second memory cache storing instructions executed by the CPU core, the second memory cache having a second line size, the second line size being larger than the first line size, each line of the second memory cache comprising an upper half and a lower half; as well as a memory controller subsystem coupled to the CPU core and the first memory cache and the second memory cache, the memory controller subsystem being configured to: Upon determining a first miss in the first memory cache for a first target address: determining that the first target address maps to the lower half of a first line in the second memory cache; retrieve the entire first line from the second memory cache, and returning the entire first line from the second memory cache to the first memory cache; as well as Upon determining a second miss in the first memory cache for a second target address: determining that the second target address maps to the upper half of a second line in the second memory cache; as well as The upper half of the second line from the second memory cache is returned to the first memory cache instead of the lower half of the second line from the second memory cache.

2. The device according to claim 1, wherein: The second row size is twice the size of the first row size.

3. The device according to claim 1, wherein: The storage controller subsystem includes: a first memory controller that determines that the first target address maps to the lower half of the first line in the second memory cache and generates a request for the entire first line from the second memory cache; and A second memory controller receives the request and accesses the second memory cache to retrieve the entire first line.

4. The device according to claim 3, wherein: The first target address is a virtual address, and the request for the entire first line from the second memory cache includes a physical address generated based on the virtual address, and the request for the entire first line also includes an indicator, which indicates that the entire first line from the second memory cache is to be retrieved from the second memory cache.

5. The device according to claim 1, wherein: The first target address is provided by the CPU core to the storage controller subsystem to retrieve a prefetch unit including a set of program instructions starting from the first target address, and the CPU core also asserts a signal to the storage controller subsystem that the program instructions within the additional prefetch unit will not be retrieved and provided back to the CPU core.

6. The device according to claim 1, wherein: A third target address is provided by the CPU core to the storage controller subsystem to retrieve a first prefetch unit including a set of program instructions starting from the third target address, and wherein the CPU core will also provide a first prefetch count to the storage controller subsystem, the first prefetch count indicating the number of prefetch units of program instructions after the first prefetch unit.

7. The device according to claim 6, wherein: The storage controller subsystem is to calculate a first series of target addresses based on the first prefetch count and the second target address, and wherein the first target address is a last target address in the first series.

8. The device according to claim 1, wherein: Determining that the first target address maps to the lower half of the first line in the second memory cache includes determining a logic state of at least one bit in the first target address.

9. The apparatus of claim 7, wherein the storage controller subsystem is further configured to: responsive to the first target address being a last target address in a first series defined by the first prefetch unit, returning the entire first line to the first memory cache; and In response to the second target address being a last target address in a second series defined by a second prefetch unit and a second prefetch count, returning the upper half of the second line to the first memory cache.

10. The device according to claim 7, wherein: The first target address and the third target address are the same; The first prefetch count is 0; and The memory controller subsystem is further configured to, in response to the first prefetch count being zero, mark the upper half of the first row as invalid.

11. The apparatus of claim 10, wherein the storage controller subsystem is further configured to change the upper half of the first row from invalid to valid in response to a restore instruction.

12. A system comprising: Input / output devices; and a processor coupled to the input / output device and comprising a central processing unit core, i.e., a CPU core, a first memory cache, a second memory cache, and a memory controller subsystem, wherein: The first memory cache is to store instructions for execution by the CPU core, the first memory cache having a first line size; The second memory cache is to store instructions for execution by the CPU core, the second memory cache having a second line size, the second line size being larger than the first line size, each line of the second memory cache comprising an upper half and a lower half; The memory controller subsystem is coupled to the CPU core and the first memory cache and the second memory cache; and The storage controller subsystem is configured to: Upon a first miss of a first target address in the first memory cache: determining that the first target address maps to the lower half of a first line in the second memory cache; retrieving the entire first line from the second memory cache; and returning the entire first line from the second memory cache to the first memory cache; and When a second target address misses a second time in the first memory cache: determining that the second target address maps to the upper half of a second line in the second memory cache; and Only the upper half of the second line is returned from the second memory cache to the first memory cache.

13. The system according to claim 12, wherein: The second row size is twice the size of the first row size.

14. The system according to claim 12, wherein: The storage controller subsystem includes: a first memory controller that determines that the first target address maps to the lower half of the first line in the second memory cache and generates a request for the entire first line from the second memory cache; and A second memory controller receives the request and accesses the second memory cache to retrieve the entire first line.

15. The system of claim 12, wherein: The first target address is to be provided by the CPU core to the storage controller subsystem to retrieve a pre-fetch unit including a set of program instructions starting from the first target address, and the CPU core also asserts to the storage controller subsystem a signal that program instructions within an additional pre-fetch unit are not to be retrieved and provided back to the CPU core.

16. The system of claim 12, wherein: A third target address will be provided by the CPU core to the storage controller subsystem to retrieve a first prefetch unit including a set of program instructions starting from the third target address, and wherein the CPU core will also provide a prefetch count to the storage controller subsystem, the prefetch count indicating the number of prefetch units of program instructions after the first prefetch unit.

17. The system of claim 16, wherein: The storage controller subsystem is to calculate a series of target addresses based on the prefetch count and the third target address, and wherein the first target address is a last target address in the series.

18. An apparatus comprising: The central processing unit core is the CPU core; an L1 program cache storing instructions executed by the CPU core, the L1 program cache having a first line size; an L2 memory cache storing data and executable instructions, the L2 memory cache having a second line size, the second line size being twice the first line size, each line of the L2 memory cache comprising an upper half and a lower half; as well as a storage controller subsystem coupled to the CPU core, the L1 program cache, and the L2 memory cache, the storage controller subsystem configured to: receiving a first address from the CPU core, the first address corresponding to a prefetch unit including a set of instructions to be executed by the CPU core; receiving a prefetch count from the CPU core, the prefetch count indicating a number of additional prefetch units for instructions; In response to the prefetch count being zero and the first address being determined as a first miss in the L1 program cache: determining that the first address maps to the lower half of a first line in the L2 memory cache; as well as storing an entire first line from the L2 memory cache into the L1 program cache; as well as In response to the prefetch count being greater than 0: calculating a series of addresses based on the first address and the prefetch count, the series of addresses including an initial address and a final address; determining that the last address in the series is a second miss in the L1 program cache; as well as determining that the last address maps to the lower half of a second line in the L2 memory cache; as well as Responsive to determining that the last address in the series missed the second time in the L1 program cache and the last address maps to the lower half of the second line in the L2 memory cache, storing the entire second line from the L2 memory cache into the L1 program cache.

19. The device according to claim 18, wherein: The storage controller subsystem is used for: receiving a second address from the CPU core, the second address corresponding to a second prefetch unit; receiving a second prefetch count from the CPU core, the second prefetch count indicating a number of additional prefetch units for instructions; as well as In response to the second prefetch count being zero, the second address being a third miss in the L1 program cache, and the second address mapping to the upper half of a third line in the L2 memory cache, storing the upper half of the third line from the L2 memory cache into the L1 program cache instead of the lower half.

20. The device according to claim 18, wherein The storage controller subsystem includes: a first memory controller that determines that the first address maps to the lower half of the first line in the L2 memory cache and generates a request for the entire first line from the L2 memory cache; a second memory controller that receives the request from the first memory controller and accesses the L2 memory cache to retrieve the entire first line; and wherein the first address is a virtual address, and the request for the entire first line from the L2 memory cache includes a physical address generated based on the virtual address, and the request for the entire first line also includes an indicator, the indicator indicating that the entire first line from the L2 memory cache is to be retrieved from the L2 memory cache.

Citation Information

Patent Citations

  • Intelligent cache memory and prefetch method based on CPU data fetching characteristics

    US5361391A

  • Slave cache having sub-line valid bits updated by a master cache

    US5784590A