Cache set replacement order based on time-related set recording

By implementing a cache set replacement order based on time-related record information, the cache management system optimizes cache usage in set associative caches, addressing resource sharing issues in simultaneous multi-threaded processors, thereby enhancing performance by minimizing cache contamination and improving efficiency across threads.

DE102013200508B4Active Publication Date: 2025-07-10INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
DE102013200508
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2012-01-20
Filing Date
2013-01-15
Publication Date
2025-07-10
Estimated Expiration
2033-01-15

AI Technical Summary

Technical Problem

Existing cache management systems in set associative caches do not optimally handle resource sharing issues in simultaneous multi-threaded processors, leading to suboptimal performance due to one thread flushing the cache of another, especially when certain instructions consistently cause cache misses or hits.

Method used

Implement a cache set replacement order based on time-related record information, using an error count and hit position field to determine the hierarchical placement of data elements in the cache, ensuring that instructions causing consistent cache misses are placed closer to the LRU position and hits are placed closer to the MRU position, thereby optimizing cache usage.

Benefits of technology

This approach enhances cache performance by minimizing cache contamination and improving overall system efficiency by ensuring that data elements that consistently cause cache misses are not prematurely flushed, allowing more efficient use of cache resources across multiple threads.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Computer system (100) for cache management, the system (100) comprising: a processing circuit and a cache (15), wherein the system (100) is configured to perform a method comprising: Tracking (605) by the processing circuitry, for an instruction requesting access to a data item in the cache, a miss count and a hit position field (29), the miss count and the hit position field being generated by a previous execution of the instruction and being stored in a tracking table; and wherein the trace table includes the miss count and hit position field associated with an instruction address for the instruction requesting access to the data item; and Placing (610) the data element in a hierarchical replacement order based on at least one of the error count and the hit position field, wherein the hit position field (29) has a hierarchical position associated with the instruction associated with the data element.
Need to check novelty before this filing date? Find Prior Art

Description

BackgroundThe present invention relates to data processing and, more particularly, to the cache set replacement order of data elements in a set associative cache based on time-related record information.A cache is a component that transparently conserves data items (or simply data) so that future requests for any preserved data can be handled more quickly. A data element stored within a cache corresponds to a predetermined memory location within a computer system. Such a data element could be a value recently calculated or a second copy of the same location also stored elsewhere. If requested data is included in the cache, this means a cache hit, and this request can be processed by simply reading the cache, which is relatively faster because the cache is usually set close to its requester. On the other hand, if the data is not included in the cache, it is a cache miss, and the data needs to be fetched from another storage medium that is not necessarily close to the requester, and thus is relatively slower. Generally, the greater the number of requests that can be handled using the cache, the faster the overall system performance becomes.If a document is already moving, namely US 7603 522 B1. This document describes blocking aggressive neighbors in a cache subsystem. Here, the system includes a plurality of processing elements, a cache memory shared by the plurality of processing elements, and a circuit that controls allocation of data in the cache memory. In this case, the cache control circuit is configured such that data from a processing element is stored in cache memory at less attractive positions when the associated processing unit exhibits a relatively poor cache behavior compared to other processing units.Despite these advances made, there continues to be a need to further optimize cache management.SummaryThis object is achieved by the subject matters of the independent claims. Further embodiments are described by the respective dependent patent claims. According to example embodiments, a computer system, method, and computer program product for determining a cache set replacement order of data elements in a set associative cache based on time-related set record information are provided. An error count and hit position field of an instruction requesting the storage of a data item in a cache are determined. The error count and hit position field are generated based on previous executions of the instruction. In the event of a cache miss, the requested data element is then placed (installed) in a cache set memory location using a hierarchical positioning scheme based on the miss count or / and hit position field. In a cache hit, the set replacement order of the data element is changed according to the (same) error count and the (same) hit position field. The hit position field defines a hierarchical position with respect to the data element within a cache congruence class.Additional features and advantages are realized by the techniques of the present invention. Other embodiments and aspects of the invention are described in detail herein and are considered part of the claimed invention. For a better understanding of the invention, together with its advantages and features, reference is made to the description and the drawings.Brief Description of the Several Views of the DrawingsThe subject matter which is considered to be the invention is particularly pointed out and clearly claimed in the claims at the conclusion of the specification. The foregoing and other features and advantages of the invention will become apparent from the following detailed description when taken in conjunction with the accompanying drawings, in which: FIG. 1 shows a block diagram of a system according to an embodiment of the present invention; FIG. 2 shows a block diagram of hierarchical locations stored in a cache directory for cached data items per congruence class, according to an embodiment of the present invention; FIG. 3 shows a block diagram of data elements cached in a cache corresponding to hierarchical locations for a congruence class, according to an embodiment of the present invention; FIG. 4 shows a flow chart updating a hit position field and an error counter according to an embodiment of the present invention; FIG. 5 shows a flow chart that further updates a hit position field and an error counter according to an embodiment of the present invention; FIG. 6 illustrates a method for cache management of a cache by circuitry according to an embodiment of the present invention; FIG. 7 shows an example of a computer with functions that may be used in accordance with embodiments of the present invention; FIG. 8 shows an example of a computer program product on a computer readable medium according to an embodiment of the present invention.DETAILED DESCRIPTIONA microprocessor may include a level one (L1) data cache (D-cache). The L1 data cache is used for storing data elements of a subset of the system memory locations so that instructions for executing load and store operations may be processed in close proximity to the processor core. In the event of a cache miss, data elements corresponding to a requested memory location are installed into the cache. Each entry in the cache represents a cache line corresponding to a portion of a memory. One such typical installation algorithm for a cache is based on least recently used (LRU). For example, a data cache may contain 1024 lines called congruence classes, and the data cache may be a 4-way set associative cache that would have a total of 4,000 entries. For each congruence class, a sequence from the most recently used set to the least recently used set may be logged as a set of hierarchical positions used for substitutions. When installing a new entry into the data cache for a particular congruence class among one of the four sets, it is decided that the LRU entry is replaced.A simultaneous multi-threaded (SMT) processor allows multiple threads to be executed in parallel on a single processor core. The LRU scheme may remain operational for the D-cache, but the SMT function introduces new resource sharing issues, which may result in suboptimal performance. For example, one thread may stream through a large data space that flushes a cache level (such as the L1 cache), while another thread may well utilize the cache because it mostly accesses a small amount of data. The streaming process inhibits the other thread from continually flushing it. In other words, each thread should not take more cache for optimal performance than it can effectively use compared to the other threads that also use the cache.Links may be established with respect to data access instructions that never have cache misses and their corresponding instruction addresses. An example of this is register overflow and fill. A particular instruction (e.g., store) updates or installs the cache with general register content, and a subsequent instruction (e.g., load) accesses the content that was stored in the cache. The load close to the store never leads to a cache miss. Moreover, links may be established with respect to cache accesses of instructions that always result in an error. In other words, steady hits or errors for future predicted behavior may be correlated with an instruction address. However, a data cache accessed instruction that is a combination of hits and misses cannot be simply associated with its instruction address in view of predicting its future behavior.According to an example embodiment, by identifying instructions that always result in hits or errors in the data cache, these installations (in the event of an error) to the D cache and accesses (loads / stores) from the D cache may not be set as a MRU (most recently used) position (state) and / or close to the MRU position (state) in the D cache directory. In this way (for example), the processor may prevent one thread from completely flushing the data cache of another thread in the D cache. A thread flushing the cache has accesses that always result in a cache miss. Such errors should not be installed as an MRU location, as otherwise the data would have to find its way out of the cache until it eventually reaches the LRU location and is then removed. If such data is initially installed as LRU and LRU remains (if installed at all), then in an exemplary embodiment, this data is constrained to the amount of cache contamination they cause in the LRU position. For example, if a load access (or memory access) always results in a cache hit, designating the hit set as MRU may be unnecessary if marking the hit set in a position closer to the LRU position would also result in a cache hit. If the data element is marked too close to the MRU position, the no longer required data (i.e., data element) remains in the data cache longer than required. Placing a data element that always results in a hit in a position closer to LRU as determined according to example embodiments allows another data element to remain in the cache that needs to be cached for a longer period of time before accessing the data element again. Thus, if an entry is identified as MRU, the previous MRU entry becomes the second (2) MRU position shifted. The second MRU entry is now one place closer to the LRU location, and the LRU location is the entry that is removed when a new entry is installed into the particular congruence class of the cache. Example embodiments are configured to determine at which location (e.g., from the MRU location to the LRU location and / or no installation at all) an entry should be placed when an entry is installed and / or updated to the data cache. This method may be extended to be executed in the instruction cache and / or other caching structures in a microprocessor.Referring now to FIG. 1, shown is a block diagram of a system 100 generally in accordance with an embodiment. The system 100 includes a processor 105. Processor 105 has one or more processor cores, and the processor core may be referred to as circuit 10. The processor 105 may include a level one (L1) cache 15. While an L1 cache is shown, example embodiments may be implemented in an L1 cache or L2 cache, as desired. The L1 cache 15 includes an L1 data cache 20 (D cache) and an L1 instruction cache 22 (I cache). The data cache 20 is a (hardware) memory provided on the processor for buffering (i.e., holding) data on the processor 105. Data retrieved from the memory 110 may be cached in the data cache 20, while program code instructions 115 retrieved from the memory 110 may be cached in the instruction cache 22 (e.g., (hardware) memory provided on the processor).The L1 cache 15 may be an X congruence class set associative N-way cache, as will be understood by those skilled in the art. The D cache 20 contains a D cache directory 24 which may contain the LRU replacement order (the respective hierarchical positions) for each entry (data element) of the congruence classes. The replacement order of the data elements in a single congruence class can range from MRU position, second MRU position, third MRU position to LRU position per congruence class, as shown in FIG. 2. The MRU position is the highest position for a data element (an entry) in the D cache 20, the second MRU position is the second highest, the third MRU position is the third highest, and the LRU position is the lowest position for a data element (starting from a 4-way set associative cache). A data item (entry) is / denotes latched data in the D cache 20 of the L1 cache 15.FIG. 2 illustrates an example replacement order (e.g., four hierarchical positions) per set of cached data elements stored in the D cache directory 24 per congruence class in the D cache 20, according to one embodiment. In FIG. 2, congruence class 1 through congruence class X (which would correspond to congruence class 0 through congruence class X-1, as will be understood by those skilled in the art) and four sets in each congruence class are shown by way of illustration. For example, assume that data items A, B, C, and D are entries latched in each set of D cache 20 according to the hierarchical positions for congruence class 1 shown in FIG. 3. When a new data element E is installed in congruence class 1 of D cache 20, this incoming data element E is usually set to the MRU location and data element D is removed from the LRU location in the D cache. Thereafter, each data element shifts down by one position in the hierarchical order, so that the new hierarchical order from the MRU position to the LRU position is E, A, B, and C, respectively.However, the (hardware) circuits 12 of the circuit 10 are designed to install all data items (e.g., the incoming data item E) in a hierarchical position according to a log table 26 in the I-cache 22. The circuitry 12 is configured to determine the hierarchical location and install (and / or not install) an incoming data item based on the history of the associated instruction (e.g., a fetch instruction) requesting that particular data item, such that when that particular instruction is executed (i.e., viewed) from the program code 115 by the processor circuitry 10 of the processor core, the circuitry 12 stores the corresponding data item in the logging table 26 according to a function of the stored location. The corresponding instruction address is maintained (marked) in the log table 26 of the I-cache 22 along with the specifications for the processing of the corresponding data item, as discussed further below.The circuits 12 may be application specific integrated circuits (ASIC), field programmable gate arrays (FPGA), etc. In addition, the logic of circuits 12 may be implemented as software code, which is shown when implemented as software application 14. All references to the functions, logic, and features of circuits 12 relate to software application 14, as will be understood by those skilled in the art. The logging table 26 may alternatively and / or additionally be stored in the memory 110 as the logging table 26 a, and all references to the functions, logic, and features of the logging table 26 relate to the logging table 26 a, as will be understood by those skilled in the art.Under the control of circuits 12, log table 26 is configured to log data elements stored in D cache 20 as a link to the instruction performing the data access, and when the instruction is thereafter processed by circuit 10, its corresponding data element is placed in a hierarchical position in D cache directory 24 as a function of error counter 27 and hit position field 29 previously stored for that particular instruction. The circuits 12 use the logging table 26 to log the instruction address of each instruction (e.g., from program code 115 and / or any other program) accessing a data element in the D-cache 20.For example, assume that data element E has an associated instruction address E (or simply instruction E). When the data element E is first cached in the D cache 20 (and the instruction E is cached in the I cache 22), the circuits 12 are configured to log the instruction E and the data element E in the log table 26. In this first pass, data element E is queued (as usual) in the MRU position of D cache directory 24. The error counter 27 corresponding to the data accesses by the instruction E starts at zero (0) and is, for example, saturated / ended to three (3) at three (3), where 3 is the threshold error count in this example. Any cache miss above 3 exceeds the threshold, as discussed herein. For each cache error with respect to the D-cache 20 for data access by instruction E, the error counter 27 increments by one (1) until it reaches 3. At this point, the error counter 27 no longer increments by 1. If another cache hit occurs with respect to the D-cache 20, the circuits 12 decrement the error counter 27 by 1 until 0 is reached and does not go below 0. If other incoming data elements are stored in the same congruence class of the D-cache 20 after data has been stored, the instruction E has been accessed, the data element in the LRU location is removed from the D-cache 20 and the remaining data elements move down by one location in the hierarchical locations.Typically, for each cache hit, data element E is moved to the MRU position (if data element E was previously moved down the hierarchy) with respect to D cache 20 (a data element E therein), but data element E may be moved to the MRU position at maximum. Assume that the error counter 27 reaches its maximum value, e.g. threshold 3, in the first scenario of the first processing of the instruction E by the circuit 10. When circuits 12 determine that error counter 27 has reached its maximum value 3, error counter 27 does not go beyond its maximum value 3, but is saturated to 3. Circuits 12 are configured to recognize that instruction E (always) installs and maintains its corresponding data elements in the LRU position in response to the cache fault for the data accessed by instruction E (reaching maximum value 3). The subsequent processing of instruction E refers to processing instruction E at a later time when instruction E accesses a data element E to be cached again in D-cache 20. Maintaining data element E in the LRU position means that circuits 12 are configured to determine and maintain a tag (in log table 26 for instruction E) indicating that data element E accessed by instruction E is to be subsequently placed in the LRU position (when instruction E is processed at a later time) and not moved up to the MRU position of D cache directory 24, not even when cache hits occur for data element E. Unlike conventional cache management, circuits 12 disable data element E for instruction E from being up-shifted by hierarchical positions (not even at a cache hit) until data element E is eventually removed from D-cache 20 by an incoming data element. Optionally, in such a case, circuits 12 may be configured so that they do not install data element E in D-cache 20 at all with respect to instruction E when error counter 27 is saturated (exceeds its threshold). It should be noted that cache management, as discussed herein, is based on the error counter 27 and hit position field 29 stored in the log table 26 for each previously processed instruction (e.g., instruction E discussed in the various scenarios).In a second scenario for accessing a data element E by instruction E, the error counter 27 has not reached its maximum value (e.g., the error counter 27 is not set), and the data element E may eventually be removed from the D cache 20. Circuits 12 are configured to mark / mark (e.g., set corresponding bits) in hit position field 29 the hierarchical position closest to the LRU position (including the LRU position itself) at which data element E reached a cache hit before it has been removed. Four example cases are given below.In a first case, circuits 12 store / identify the MRU location in hit location field 29 for instruction E when the lowest hierarchical cache hit location for data element E has occurred in D cache 20 in the MRU location. When circuits 12 are later again processing instruction E, they check hit position field 29 for instruction E. If the tag in hit position field 29 is the MRU position, the new entry (the same or different data element E) is to be installed in the LRU position in the event of a cache miss by circuits 12 following processing of instruction E because data element E requires (only) one set in D cache 20. Assume that a cache hit occurs for data element E in D cache 20. Usually, the data element E would be set to the MRU position at a cache hit without this disclosed feature. However, with this disclosed feature, data element E would be set to the MRU position at a cache hit as close as hit position field 29 permits, and in the above case, data element E would be set to the LRU position at a cache hit. The entry is never moved up further than to the hierarchical (original) installation position according to the hierarchical hit position field 29 for instruction E. This prevents the L1 cache 15 from being contaminated by having to wait longer than necessary until the data item E is removed from the L1 cache.In a second case, circuits 12 store / tag the second MRU position in hit position field 29 for instruction E when the lowest hierarchical cache hit position for data element E in D-cache 20 has occurred in the second MRU position of a 4-way set associative cache. When circuits 12 are later processing instruction E again, they check hit position field 29 for instruction E. If the tag in hit position field 29 is the second MRU position, the new entry (the same or different data element E) is to be installed in the third MRU position (also known as the second LRU position) in the event of a cache miss from circuits 12 because the entry requires two sets (only at all times). Assume that a cache hit occurs for data element E in D cache 20. Usually, the data element E would be set to the MRU position at a cache hit. However, for a cache hit, data element E is set to the MRU position as close as the hit position field 29 permits, and in the above case, data element E is set to the third MRU position for a cache hit (by circuits 12).In a third case, circuits 12 store / tag the third MRU location in hit location field 29 for instruction E when the lowest hierarchical cache hit location for data element E in D cache 20 has occurred in the third MRU location (there may be previous cache hits in the MRU and second MRU locations). When circuits 12 are later again processing instruction E, they check hit position field 29 for instruction E. If the tag in hit position field 29 is the third MRU position, then the new entry (the same or different data element E) is to be installed from circuits 12 to the second MRU position in the event of a cache miss because the entry requires three sets. Assume that a cache hit occurs for data element E in D cache 20. Usually, the data element would be set to the MRU position at a cache hit. However, for a cache hit, data element E is set to the MRU position as close as hit position field 29 permits, and in the above case, data element E is set to the second MRU position for a cache hit.In a fourth case, circuits 12 store / tag the LRU location in the hit location field for instruction E when the lowest hierarchical cache hit location for data element E has occurred in the D-cache 20 in the LRU location (in other accesses, additional cache hits may be present in higher locations). When circuits 12 are later processing instruction E again, they check hit position field 29 for instruction E. If the tag in hit position field 29 is the LRU position, the entry from circuits 12 is to be installed into the MRU position because the new entry (the same or different data element E) uses all four sets in the D cache. Assume that a cache hit occurs for data element E in D cache 20. At the cache hit, data element E is set to the MRU position as close as hit position field 29 permits, and in the above case, data element E is set to the MRU position at a cache hit.With respect to hit position field 29, circuits 12 count the tagged position (self) back to the MRU position after the data element reaches the tagged position (as determined in the first processing of the instruction) at which the lowest hierarchical cache hit position occurred. The counted number (X) corresponds to the number of sets (hierarchical positions) required by that particular data item. To determine the hierarchical installation position, this number (X) (by the circuits 12) is used for counting from the LRU position (even as 1 set) to the MRU position, and the hierarchical position at which the counting is completed (number X) is the installation position for the data item. There is an inverse relationship between a designated position in the hierarchical hit position field 29 and the hierarchical installation position.To further explain the difference between an action for the error counter 27 and an action for the hit position field 29: the circuits 12 are configured to operate the D cache 20 such that the error counter 27 takes priority over the hit position field 29. Assume that the error counter reaches the threshold (for instruction E) in the first processing of instruction E and that the lowest cache hit position for data element E is marked in hit position field 29 (e.g., MRU position, second MRU position, third MRU position, or LRU position). In subsequent processing of instruction E, circuits 12 are configured to determine whether error counter 27 is at the threshold and hit position field 29 is set (by checking logging table 26 for markers / bits corresponding to instruction E). Circuits 12 are configured to determine that the action for error counter 27 overrides the action for hit position field 29 such that, while the action corresponding to error counter 27 (i.e., install and remain in the LRU position until each data element accessed by instruction E has been removed from D-cache 20 or no installation in D-cache 20) is applied, the action corresponding to hit position field 29 is not applied.Although only a single error counter 27 and a single hit position field 29 are shown in FIG. 1, the error counter 27 and the hit position field 29 represent a plurality of error counters 27 and hit position fields 29, both of which correspond to individual instructions (instruction addresses of data access information), for example, the program code 115 and / or other executable program code. While the logging table 26 is shown in the I-cache 22, it may be maintained elsewhere (per processor) as will be understood by those skilled in the art. Moreover, the logging table 26 may be maintained and constructed in various ways, including maintaining the instruction address as a function of a predetermined hash, maintaining the logging table as a set associative lookup table, etc., as will be understood by those skilled in the art.Moreover, it should be noted that each instruction processed by circuits 12 is marked / stored with its own respective error counter 27 and hit position field 29 in logging table 26. In this regard, the circuits 12 may later check the log table 26 for an instruction fault or hit for each respective instruction in the I-cache 22, and if an instruction hit is present, the respective data element for that instruction is processed based on the fault counter 27 (when the threshold is reached) and the hit position field 29 (which could be empty if no cache hit occurs). Each previously processed instruction is associated with its own error counter 27 and hit position field 29 to be applied in subsequent processing of that instruction.In one implementation, the error counter 27 may be two bits (or more if desired) and the hit position field 29 may be two bits. The bits of the error counter 27 may be set to count the D cache errors for the corresponding instruction. The circuits 12 are configured to detect when the error count threshold has been reached for each of the instructions. If the error count threshold of the error counter 27 for the instruction has not been reached, no corresponding action is taken with respect to the error counter 27 in the subsequent processing of the instruction from the circuits 12. When the error count has reached the error count threshold, the circuitry 12 detects this and performs the appropriate actions for the error counter 27 (e.g., maintaining the data item in the LRU position and / or not installing the data item), as discussed herein. The bits of hit position field 29 may also be set to represent each hierarchical order (from the MRU position to the LRU position) in D cache directory 24, and the determined hierarchical order is stored according to the hit position field for respective instructions. In accordance with the discussions herein, it will be appreciated that various techniques may be utilized to represent features and functions of the error counter 27 and hit position array 29, as will be understood by those skilled in the art.Various example scenarios are used for illustration with reference to instruction E and its data item E, but are not intended to be limiting. Circuits 12 are configured to process numerous instructions and their respective data elements simultaneously or nearly simultaneously in accordance with the present disclosure, as discussed herein.Processor 105 may be a simultaneous multi-threaded (SMT) processor that enables multiple threads to be executed in parallel on a single processor core (e.g., circuit 10). In logging (by circuitry 12) which hierarchical position is closest to the LRU position accessed by a particular instruction, the instruction may be logged to contain all threads of a single hierarchical order, or the logging may be performed per thread. When a larger area is used, the advantage of per-thread logging is obtaining knowledge of the required number of sets for a particular thread, thereby minimizing learning collisions from the other threads. For example, referring to FIG. 3, assume that thread 1 has data elements A and B and thread 2 has data elements C and D. Rather than considering the set as a whole, circuits 12 may be configured to involve only the hierarchical order in terms of positions per thread and thereafter apply error counter 27 and hit position field 29 as discussed herein. For example, in a per-thread basis, circuits 12 are configured to consider data element C as the MRU position and data element D as the second MRU position for thread 2, and thereafter analogously apply fault counter 27 and hit position field 29 for the thread 2 instruction. Similarly, circuits 12 are configured to include data element A in the MRU position and data element B in the second MRU, and thereafter analogously apply fault counter 27 and hit position field 29 for a thread 1 instruction, as described herein.For example, by applying SMT without altering the algorithm, thread 1 can always hit as the MRU and second MRU positions (of all positions within the hierarchy of a particular congruence class) in this pass. In a future pass of code, the data of thread 1 may be unavailable (installed as a third MRU because it initially has only occupied the first two slots) depending on what thread 2 makes, because one of the two most recently used entries in the congruence class is based on what thread 2 makes or accesses.At the first installation (no information is available in the logging table 26), the entry is installed into the D cache 20 and placed as an MRU as described above. The worst position calculation (closest to the LRU position) is logged differently. If this were the only entry from that thread in the particular congruence class (e.g., all other entries originate from another thread), that entry could be considered to be both MRU and LRU (only one entry in the thread, so that only one space is allocated, and therefore MRU= LRU). It would then be referred to as achieving a hit in the LRU space (for the particular thread) so that it would be installed in the MRU space (the entire congruence class) at future installations. If there are two entries for this thread (we call entries A and B) and "A" was always MRU between entries "A" and "B", then "A" MRU and the worst access position of "B" is the LRU space (with respect to the ordering of this thread). As such, "A" is a space away from the LRU (the string of that thread). Entry "A" (for this thread) is logged in the logging table as a "third MRU" for this reason (which is one place away from the LRU of the entire congruence class), so that "A" is installed in the second MRU position at future installations. This provides security with respect to the other thread because the other thread could position the entry "A" of that thread two slots closer to the LRU, and "A" would remain in the D cache as the LRU position of the hierarchy.Referring now to FIG. 4, a flow chart 400 is shown for updating hit position field 29 and error counter 27, according to an example embodiment. Further, the flow chart 400 shows how the D cache 20 behaves when the hit position field 29 is applied.Assume that instruction E requesting data item E is processed by circuit 10 of processor 105 to request data item E. The circuits 12 of the processor 105 determine whether a D cache hit for data element E is present in the D cache 20 at block 405. If not, there is no D cache hit for data element E and flow continues to block 505 on the next page of FIG. 5. if yes, a D cache hit is present for the data element 12, and the circuitry 12 determines whether an instruction hit is present in the log table 26 of the I-cache 22 that yields (sets) the D cache request at block 410. If not, there is no instruction hit in the log table 26 of I-cache 22 (i.e., no miss counter 27 or hit position field 29 is marked for that particular instruction (e.g., instruction E)), the circuits 12 are configured to set the hit position field 29 at block 415 at the time of the D-cache hit corresponding to the hierarchical positions (e.g., MRU position to LRU position), and set the miss counter 27 to 0. If no instruction hit is reached in the log table 26, this means that that particular instruction (e.g., instruction E) has not been processed so far and thus has no error counter 27 and hit position field 29, or that the particular instruction (e.g., instruction E) has been removed from the log table 26.If so, an instruction hit is present in the log table for that particular instruction (e.g., instruction E), the circuits 12 are configured to place the D cache entry, data element E, as close to the MRU position as possible in block 420 without exceeding the function of the hit position field 29. For example, when hit position field 29 is set to the second MRU position for instruction E, circuits 12 are configured to place the corresponding data element E (retrieved from the instruction address of instruction E) on the third MRU position in D cache directory 24 of D cache 20 (as discussed above).The circuits 12 are configured to check whether the corresponding data element (e.g., data element E) is set closer to the LRU position when installed as the MRU position than was recorded in the log table 26 for the corresponding instruction (e.g., instruction E) in block 425. If so, the circuits 12 are configured to decrease the error count of the instruction in the error counter 27 in block 430 and update the hit position field 29 to the current hierarchical position. For example, if data element E was initially referenced / marked as an MRU position and later as a second MRU position, then data element E will be installed as a third MRU location (also known as a second LRU location) upon future installation of data element E accessed by instruction E. If not, the circuits 12 are configured to decrease the error count of the instruction in the error counter 27 in block 435 and allow the hit position field 29 for the instruction (e.g., instruction E) to remain as it is.Referring now to FIG. 5, a flowchart 500 is shown that is a continuation of the flowchart 400 of FIG. 4, according to one embodiment. In the event of a D cache error (i.e., if no D cache hit has occurred in block 405), the circuits 12 are configured to determine whether there is an instruction hit for instruction E in the log table 26 of the I-cache 22 that gives (sets) the D cache request in block 505. If not, there is no instruction hit for instruction E making the D cache request, and circuits 12 are configured to install that particular instruction (e.g., instruction E) into logging table 26 in block 510, by setting the hit position to the MRU position and setting the error count of error counter 27 for that instruction E to 1.If so, an instruction hit is present in the log table 26 for instruction E, the circuits 12 are configured to check whether epsilon handling should occur for instruction E in block 515. If the answer to block 515 is "no", then no epsilon handling is made for the instruction and circuits 12 are configured to increment the error count of the instruction in error counter 27 in block 520 and install the corresponding data element (e.g., data element E) in D-cache 20 as a function of the hit position in hit position field 29.In order to enable process variability, an epsilon factor is present for epsilon handling. The epsilon defines a small percentage of the time period for which the installation space descriptor (error counter 27 and hit position field 29) stored in the log table 26 is ignored and the particular entry is installed into the MRU position by the circuits 12. By allowing such an entry to be installed in the MRU position, if the behavior of program code 115 (including instruction E) has changed over time and no potential hit is achieved because the installation is too far from the MRU position, circuits 12 of processor 105 can make corrections / adjustments to the new code behavior. If the answer to block 515 is "yes", epsilon handling is initiated per se for that instruction (e.g., based on processing the instruction a predetermined number of times and / or after a predetermined period of time has elapsed), the circuitry 12 is configured to install the corresponding data element (e.g., data element E) in block 525 into the MRU position of the D-cache 20, set the hit position field 29 to the MRU position, and increase the error count of the instruction in the error counter 27 (thereby simultaneously ignoring the previous data / actions for the error counter 27 and the hit position field 29).FIG. 6 illustrates a method 600 for cache management of the D-cache 20 by the circuits 12 according to an embodiment. It should be noted that while circuitry 12 is illustratively identified as executing certain functions (e.g., wired to hardware components for execution as discussed herein), it is also part of circuitry 10 (the hardware that constitutes the processor core). Each discussion of the circuits 12 relates to the circuit 10, both of which form the processor 105 (the processing circuit).In block 605, circuits 12 are configured to determine an error count in error counter 27 and hit position field 29 during prior processing of an instruction (e.g., instruction E) that requested storage of a data element (e.g., data element E) in a cache (e.g., D-cache 20). The error count of the error counter 27 and hit position field 29 are stored for the instruction requesting the data item. The error counter 27 and / or hit position field 29 may or may not be set when the other is set. Prior processing refers to an earlier time at which the instruction was processed by the circuits 12 to store / update the data of the error counter 27 and / or the data of the hit position field 29 in the instruction logging table 26.Circuits 12 are configured to place / install the data item based on the error count (error counter 27) or / and hit position field 29 during subsequent processing of the instruction in a hierarchical order (D cache directory 24). Hit position field 29 contains, in block 610, a hierarchical position that relates to the data element (e.g., data element E) that is within a congruence class in the cache (e.g., D-cache 20).If the error count is not set, the circuits 12 are configured to install / place the data item in accordance with the hierarchical position stored in the hit position field 29 in the hierarchical order in response thereto in block 615. Circuits 12 place the data item according to an inverse relationship of the hierarchical position stored in hit position field 29, in the hierarchical order, and in response to a cache hit, circuits 12 prevent the data item from shifting upward in the hierarchical order further than the inverse relationship to the hierarchical position.Further, circuits 12 are configured to process the instruction according to the error count (e.g., bits set to indicate reaching the threshold for error counter 27) in response to a conflict between the error count of error counter 27 and hit position field 29. Processing the instruction according to the error count includes placing the data element in the least recently used (LRU) position in the cache (when designated as the mode of operation) and / or not placing the data element in the cache (when designated as the mode of operation).The hierarchical position in hit position field 29 indicates a closest position to the least recently used (LRU) position at which a cache hit for the data item occurred during the previous processing of the instruction. The circuits 12 determine the error count of the error counter 27 and the hit position field 29 during the previous processing of the instruction that requested the data item stored in the cache by determining a lowest position in the hierarchical order for which the cache hit occurs for the data item in the D-cache 20; as such, the circuits 12 store the determined lowest hierarchical position for which the cache hit occurs as a hierarchical position in the hit position field 29 (e.g., set the corresponding bits).The instruction fetch is performed for an instruction address for detecting an instruction (e.g., instruction E). This instruction may trigger a data access (for data element E) to the D cache 20.FIG. 7 shows an example of a computer 700 having functions that may be included in the example embodiments. Various methods, operations, modules, flowcharts, tools, applications, circuits, elements, and techniques discussed herein may also include and / or utilize the functions of the computer 700. Moreover, the functions of computer 700 may be utilized to implement features of example embodiments discussed herein. One or more of the functions of computer 700 may be utilized to implement, connect to, and / or support any element discussed with respect to FIGS. 1-6 and 8 (as will be understood by those skilled in the art).With regard to the hardware architecture, the computer 700 may generally include one or more processors 710, a computer readable mass storage 720, and one or more input and / or output (I / O) devices 770 communicatively coupled to one another via a local interface (not shown). As is known in the art, the local interface may be, for example, one or more buses or other wired or wireless connections, but is not limited thereto. The local interface may include additional elements, e.g., controllers, buffers (caches), drivers, retry elements, and receivers, to enable data exchange. In addition, the local interface may include address, control and / or data connections to enable appropriate data exchange among the above-mentioned components.The processor 710 is a hardware unit for executing software that can be stored in the memory 720. Processor 710 may be virtually any custom or commercially available processor, a central processing unit (CPU), a data signal processor (DSP), or an auxiliary processor among multiple processors associated with computer 700, and processor 710 may be a semiconductor-based microprocessor (in the form of a microchip) or a macroprocessor.The computer readable memory 720 may include volatile memory elements (e.g., random access memory (RAM), such as dynamic random access memory (DRAM), static random access memory (SRAM), etc.), and nonvolatile memory elements (e.g., read only memory (ROM), erasable programmable read only memory (EPROM), electronically erasable programmable read only memory (EEPROM), electronically erasable programmable read only memory (PROM), Examples include programmable read only memory), tape, compact disk read only memory (CD-ROM), data storage medium, floppy disk, cartridge, cartridge, or the like, etc.), or a combination thereof. Moreover, the memory 720 may include electronic, magnetic, optical, and / or other types of storage media. It should be noted that memory 720 may have a distributed architecture, where various components, while spaced apart from each other, are accessible by processor 710.The software in computer readable memory 720 may include one or more separate programs, each of which includes a queued listing of executable instructions for implementing logical functions. The software in memory 720 includes a suitable operating system (OS) 750, a compiler 740, source code 730, and one or more applications 760 of the example embodiments. As shown, this application 700 includes numerous functional components for implementing the features, processes, methods, functions, and operations of the example embodiments. The application 760 of the computer 700 may represent numerous applications, agents, software components, modules, interfaces, controllers, etc., as discussed herein, but is not intended to be limiting.Operating system 750 may control the execution of the other computer programs and provides scheduling, input-output control, file and data management, memory management and data transfer control, as well as similar services.The one or more applications 760 may employ a service-oriented architecture, which may be a collection of services that communicate with each other. Moreover, the service-oriented architecture enables two or more services to coordinate and / or perform activities (e.g., on behalf of the other). Each interaction between services may be self-contained and flexibly linked, such that each interaction is independent of any other interaction.Further, application 760 may be a source program, an executable program (object code), a script, or other entity that includes a set of instructions to be executed. When in the form of a source program, the program is typically translated using a compiler (e.g., compiler 740), an assembler, an interpreter, or the like, which may be included in memory 720, to operate properly in conjunction with OS 750. Moreover, application 760 may be written as (a) an object oriented programming language having data and method classes, or (b) a procedural programming language having routines, subroutines, and / or functions.I / O devices 770 may include, but are not limited to, input devices (or peripherals) such as a mouse, keyboard, scanner, microphone, camera, etc. In addition, I / O devices 770 may also include, but are not limited to, output devices (or peripherals) such as a printer, display, etc. Finally, I / O devices 770 may also include devices capable of transmitting both input and output, such as, but not limited to, a network interface card or modulator / demodulator (for accessing remote devices, other files, devices, systems, or a network), a radio frequency (RF) or other transceiver, a telephony interface, a bridge, a router, etc. I / O devices 770 may also include components for communicating over various networks, such as the Internet or an intranet. I / O devices 770 may be connected to and / or communicate with processor 710 via Bluetooth connections and cables (e.g., Universal Serial Bus (USB) ports, serial ports, parallel ports, FireWire, HDMI (High-Definition Multimedia Interface), etc.).The processor 710 is configured to, when the computer 700 is in operation, execute software stored in the memory 720, transmit data to and from the memory 720, and generally control the operation of the computer 700 in accordance with the software. Application 760 and OS 750 are read completely or partially by processor 710, optionally buffered within processor 710, and then executed.It should be noted that the application 760, when implemented in software, may be stored on virtually any computer readable storage medium for use by or in connection with any computer-related system or method. In the context of this document, a computer readable storage medium may be an / e electronic / s, magnetic / s, optical / s, or other / s physical / s device or means that may contain or store a computer program for use by or in connection with a computer-related system or method.The application 760 may be embodied by any computer readable medium for use by or with an / r instruction execution system, device, server, or device, e.g., a computer-based system, system with processor, or other system that can fetch and execute the instructions from the instruction execution system, device, or device. In the context of this document, a "computer readable storage medium" may be any means that can store, read, write, transmit, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. The computer readable medium may be, for example, but is not limited to, an / e electronic / s, magnetic / s, optical / s, or semiconductor system, device, or device.More specific examples (an incomplete list) for the computer readable medium 720 would include an electrical connection (electronic) having one or more wires, a portable computer diskette (magnetic or optical), a random access memory (RAM), a read only memory (ROM), an erasable programmable read only memory (EPROM), an electronically erasable programmable read only memory (EEPROM), or flash memory (electronic), an optical fiber, a portable compact disc read only memory (CD-ROM), a rewritable compact disc (CD-R / W, Compact Disc Read / Write).In example embodiments where application 760 is implemented in hardware, application 760 may be implemented with any one or a combination of the following technologies, which are well known in the art: one or more discrete logic circuits having logic gates for implementing logic functions on data signals, an application specific integrated circuit (ASIC) having suitable combinatorial logic gates, one or more programmable gate arrays (PGA), a field programmable gate array (FPGA), etc.It should be appreciated that the computer 700 includes non-limiting examples of software and hardware components that may be included in various units, servers, and systems discussed herein, and it should be understood that additional software and hardware components may be included in the various units and systems discussed in the example embodiments.As described above, embodiments may be embodied in the form of computer implemented processes and devices for performing these processes. An embodiment may include a computer program product 800 shown in FIG. 8 on a computer readable medium 802 having computer program code logic 803 including instructions embodied in the form of tangible media as an article of manufacture. Example articles of manufacture for the computer readable medium 802 may include floppy disks, CD-ROMs, hard disk drives, universal serial bus (USB) flash drivers, or any other computer readable storage medium, where a computer, when the computer program code logic 804 is loaded into and executed by the computer, becomes an apparatus for practicing the invention. Embodiments include, for example, computer program code logic 804 whether stored in a storage medium, loaded into and / or executed by a computer, or transmitted over a transmission medium such as electrical wires or cabling, via optical fibers, or electromagnetic radiation, where a computer, when the computer program code logic 804 is loaded into and executed by the computer, becomes an apparatus for practicing the invention. When implemented on a general purpose microprocessor, the segments of computer program code logic 804 configure the microprocessor to form specific logic circuits.As will be understood by those skilled in the art, aspects of the present invention may be embodied in the form of a system, method, or computer program product. Accordingly, aspects of the present invention may take the form of a hardware-only embodiment, a software-only embodiment (including firmware, resident software, microcode, etc.), or an embodiment that combines software and hardware aspects, all of which may be generally referred to herein as a "circuit," "module," or "system.". Further, aspects of the present invention may take the form of a computer program product embodied in one or more computer readable media containing computer readable program code.Any combination of one or more computer readable media may be used. The computer readable medium may be a computer readable signal medium or a computer readable storage medium. A computer readable storage medium may be, for example, but is not limited to, an / e electronic / s, magnetic / s, optical / s, electromagnetic / s, infrared or semiconductor system, device, or device, or any suitable combination of the foregoing. More specific examples (incomplete list) of the computer readable storage medium include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM) or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage unit, a magnetic storage unit, or any suitable combination of the foregoing. In the context of this document, a computer readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.A computer readable signal medium may include a propagated data signal containing computer readable program code, for example, in baseband or as part of a carrier wave. Such a propagated signal may take a variety of forms, such as, but not limited to, electromagnetic, optical, or any suitable combination thereof. A computer readable signal medium may be any computer readable medium that is not a computer readable storage medium and that can transmit, propagate, or transport a program for use by or in connection with an / e instruction execution system, apparatus, or device.Program code embodied in a computer readable medium may be transmitted using any suitable medium, such as, but not limited to, wireless, wired, fiber optic cable, RF, etc., or a combination of the foregoing.Computer program code for carrying out operations for aspects of the present invention may be written in any combination of one or more programming languages, for example, an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, such as a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, may be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce such a machine that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a means for implementing the functions / acts specified in the one or more flowchart and / or block diagram blocks.These computer program instructions may also be stored in a computer readable medium that can direct a computer, other programmable data processing apparatus, or other devices to operate in a particular manner, such that the instructions stored in the computer readable medium produce an article of manufacture including instructions that implement the function / action specified in the one or more flowchart and / or block diagram blocks.The computer program instructions may also be loaded into a computer, other programmable data processing apparatus, or other devices to cause a series of operations to be performed in the computer, on the other programmable apparatus, or on other devices to produce a process executed by the computer such that the instructions executed on the computer or on the other programmable apparatus provide processes for implementing the functions / acts specified in the one or more flowchart and / or block diagram blocks.The flowchart and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block of the flowchart or block diagrams may represent a module, segment, or portion of code having one or more executable instructions for implementing the one or more specified logical functions. It should also be noted that the functions identified in the block may occur in a different order than that identified in the figures in some alternative implementations. For example, two successive blocks may actually be executed substantially simultaneously, or the blocks may sometimes be executed in the reverse order depending on the functionality. It is further noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, may be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.The flowcharts illustrated here are merely an example. This plan or steps (or acts) described herein may be varied in a variety of ways without departing from the spirit of the invention. For example, the steps may be performed in a different order or steps may be added, omitted, or changed. All of these variations are considered part of the claimed invention.

Claims

A computer system (100) for cache management, the system (100) comprising: processing circuitry and a cache (15), the system (100) being configured to execute a method comprising: tracking (605), by the processing circuitry, an error count and a hit position field (29) for an instruction requesting access to a data element in the cache, wherein the error count and the hit position field are generated by a prior execution of the instruction and are stored in a tracking table; and wherein the tracking table comprises the error count and the hit position field associated with an instruction address for the instruction requesting access to the data element; and placing (610) the data item in a hierarchical replacement order based on at least one of the error count and the hit position field, wherein the hit position field (29) has a hierarchical position associated with the instruction associated with the data item.The computer system (100) of claim 1, wherein the processing circuit, based on the non-set error count, places the data item in the hierarchical replacement order according to a reverse relationship of the hierarchical position stored in the hit position field (29); and wherein the processing circuit, based on a cache hit, prevents the data item from being shifted further upward in the hierarchical replacement order than the reverse relationship to the hierarchical position.The computer system (100) of claim 1, further comprising determining the hierarchical replacement order of the data element according to the error count based on a conflict between the error count and the hit position field (29), and wherein optionally processing the instruction according to the error count comprises: placing the data element in a least recently used replacement position in the cache based on the error count reaching a maximum value for the commanded instruction; or / and not installing the data element in the cache (15) based on the maximum value for instruction saturation.The computer system (100) of claim 1, wherein the error count and hit position field are retained for the instruction requesting the data element.The computer system (100) of claim 1, wherein the hierarchical position in the hit position field indicates a closest position to the least recently used replacement position at which a cache hit for the data element occurred during the previous execution of the instruction.The computer system (the same as claim 1, wherein determining the error count and hit position field (29) during the previous execution of the instruction comprises determining a lowest hierarchical position as a new hierarchical replacement order for which a cache hit occurs for the data item in the cache; and storing the determined lowest hierarchical position for which the cache hit occurs as a hierarchical position in the hit position field (29), and optionally wherein the instruction is a data fetch for an instruction address corresponding to the data item for the data item to be stored in the cache (15).A method (600) for cache management, the method (600) comprising: tracking (605), by a processing circuit, an error count and a hit position field (29) of a data element for an instruction requesting access to a data element in a cache (15), wherein the error count and the hit position field (29) are generated by a prior execution of the instruction and stored in a tracking table; and wherein the tracking table comprises the error count and the hit position field associated with an instruction address for the instruction requesting access to the data element; and placing (610) the data item in a hierarchical replacement order based on at least one of the error count and the hit position field, wherein the hit position field (29) has a hierarchical position associated with the instruction associated with the data item.The method (600) of claim 7, wherein placing (610) the data element in the hierarchical replacement order is based on a cache hit according to an inverse relationship of the hierarchical position stored in the hit position field (29); and wherein preventing the data element from shifting further upward in the hierarchical replacement order than the inverse relationship to the hierarchical position.The method (600) of claim 7, further comprising determining the hierarchical replacement order of the data element according to the error count based on a conflict between the error count and the hit position field (29), and wherein optionally processing the instruction according to the error count comprises: placing the data element in a least recently used replacement position in the cache (15); or / and not installing the data element in the cache (15).The method (600) of claim 7, wherein the error count and hit position field (29) are retained for the instruction requesting the data item.The method (600) of claim 7, wherein the hierarchical position in the hit position field indicates a closest position to the least recently used replacement position at which a cache hit for the data element occurred during the previous processing of the instruction.The method (600) of claim 7, wherein determining the error count and hit position field (29) during the previous execution of the instruction comprises determining a lowest hierarchical position as a new hierarchical replacement order for which a cache hit occurs for the data item in the cache (15); and storing the determined lowest hierarchical position for which the cache hit occurs as a hierarchical position in the hit position field, and optionally wherein the instruction is a data fetch for an instruction address corresponding to the data item for storing the data item in the cache (15).A computer program product for cache management, the computer program product comprising: a tangible storage medium readable by a processing circuit and storing instructions for execution by the processing circuit for performing a method according to any of claims 7-12.

Citation Information

Patent Citations

  • Blocking aggressive neighbors in a cache subsystem

    US7603522B1