Hardware prefetching system
Through the prefetch management unit in the hardware prefetch system dynamically selects stride, Markov and optimal offset value prefetch units, the problems of long training time and inflexible switching in the prior art are solved, and the cache hit rate and accuracy of the prefetch system under different address sequences are improved.
Patent Information
- Application Number
- CN202510805731.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-17
AI Technical Summary
In the prior art, the optimal offset value prefetcher is too long to switch flexibly, resulting in inaccurate prefetching results in multi-core address sequences, and it is difficult for the existing prefetcher to efficiently switch between a single address sequence and multiple address sequences, affecting the prefetching efficiency and accuracy of the cache.
A hardware prefetch system is designed, including a prefetch management unit, which can dynamically select stride prefetch, Markov prefetch and optimal offset value prefetch units, and perform prefetch requests based on the current access address and historical access address through different prefetch units to improve prediction efficiency and accuracy, and manage prefetch requests through the output buffer unit.
It improves the prefetch accuracy of the prefetch system under different address sequences, reduces the waste of broadband resources, and enhances the cache hit rate under single and multi-core address sequences.
Smart Images

Figure CN120336212A_ABST
Abstract
Description
Technical Field
[0001] The present invention is applicable to the field of processor technology, and particularly relates to a hardware prefetching system. Background Art
[0002] Hardware prefetching is an important feature of modern high-performance processors. Through prefetching algorithms, data can be fetched from the next-level cache or memory into the cache in advance, thereby reducing the cache miss rate. Modern high-performance processors generally have a three-level cache hierarchy. For the simple and regular memory access patterns of the first-level cache, many hardware and software prefetching methods have been integrated into modern microprocessors to prefetch instructions and data. However, when it comes to the secondary or even higher-level caches, the sources of addresses become multiple. When multiple address sequences are sent intertwined, the regularity of a single address sequence is disrupted. Therefore, the processor often idles while requesting memory blocks due to misses in higher cache levels.
[0003] Conventional stride prefetching and Markov prefetchers have a good prediction rate for the address sequences of a single core, but it is difficult to predict the regularity of multi-core address sequences. The optimal offset prefetching finds the optimal offset value by training a long historical address sequence, thereby achieving higher coverage and accuracy in multi-core address sequences and enabling timely prefetching.
[0004] As Figure 1 shown, Figure 1 is a flowchart of the offset prefetching algorithm in the related art. Its core is to determine an address offset value {D}, add this offset value {D} to the current memory access address {X}, and generate a prefetch address {X + D}. The offset value {D} is not a fixed value and can adaptively change according to the scenario. The prefetch address is generated by adding the offset value to the demand access address, and the optimal offset value {D} is tried to be automatically found by training the offset value set {d}.
[0005] However, the training time of the current optimal offset prefetching is too long compared to the training time of other prefetchers. As Figure 2 shown, Figure 2 is a flowchart of the offset value training of the offset prefetching algorithm in the related art. Each offset value has a corresponding score. At the beginning of the training phase, the scores of all offset values are reset to 0. During each eligible secondary cache read access (miss or prefetch hit), the offset values {d i} in the test list are tested. If {X−d i} exists in the recent request table, then the offset value {d iThe score of {} increases. In one round, each offset value in the list is tested once: d1 is tested at the first access in this round, d2 at the next access, then d3, and so on. When all offset values in the list have been tested, the current round ends and a new round starts from offset value d1. When the score of a certain offset value reaches the set threshold, the training ends, and this offset value is the optimal offset value, and its training duration is the product of the number of offset values and the threshold.
[0006] Although the existing optimal offset value prefetch can predict multiple intertwined address sequences, its training time is too long. When there is only a single address sequence, its long training time may cause the pattern of the address sequence to change after training ends, making the prefetch result inaccurate and causing cache pollution. When there is only a single address sequence, stride prefetch and Markov prefetch can more accurately predict the change pattern of the address sequence. Therefore, a prefetching device needs to be able to flexibly switch dynamically. The training duration of the existing optimal offset value prefetch is dozens of times that of stride prefetch and Markov prefetch, and it is impossible to train the three prefetching devices simultaneously, making it difficult to achieve the purpose of flexible switching of the prefetching device.
[0007] Therefore, there is an urgent need for a new hardware prefetch system to solve the above technical problems. Summary of the Invention
[0008] The present invention provides a hardware prefetch system that can flexibly select one of stride prefetch, Markov prefetch, and optimal offset value prefetch for prefetching, thereby improving prediction efficiency and prediction accuracy and reducing broadband resources.
[0009] The present invention provides a hardware prefetch system, and the hardware prefetch system includes a processor, a level-1 cache, a level-2 cache, a level-3 cache, a prefetch module, and a memory; The processor is used to read and execute instructions in the level-1 cache, the level-2 cache, the level-3 cache, and the memory; The level-1 cache is used to cache instructions; The level-2 cache is used to cache instructions and, when the processor fails to hit the level-1 cache during reading, serves as the object for the processor to read instructions; The level-3 cache is used to cache instructions and, when the processor fails to hit the instructions in the level-1 cache and the level-2 cache during reading, serves as the object for the processor to read instructions; The memory is used to store instructions and, when the processor fails to hit the instructions in the level-1 cache, the level-2 cache, and the level-3 cache during reading, serves as the object for the processor to read instructions; The prefetch module is used to send prefetch requests to the level-2 cache, the level-3 cache, and the memory respectively; The prefetch module includes a prefetch management unit, a stride prefetch unit, an optimal offset value prefetch unit, and a Markov prefetch unit; The prefetch management unit is used to dynamically select one of the stride prefetch unit, the optimal offset value prefetch unit, and the Markov prefetch unit, and send a prefetch request to the secondary cache, the tertiary cache, and the memory through the selected unit; The stride prefetch unit is used to send a prefetch request to the secondary cache, the tertiary cache, and the memory according to the stride between the current access address of the processor and the historical access address of the processor; The optimal offset value prefetch unit is used to send a prefetch request to the secondary cache, the tertiary cache, and the memory according to the optimal offset value between the current access address and the historical access address; The Markov prefetch unit is used to send a prefetch request to the secondary cache, the tertiary cache, and the memory according to the difference between the current access address and the historical access address.
[0010] Preferably, the optimal offset value prefetch unit includes an optimal offset value table, a recent request table, and a first page boundary check unit; The optimal offset value table is used to store a plurality of preset offset values, and select the optimal offset value from the preset offset values according to the current access address and a preset rule and send it to the first page boundary check unit; The recent request table is used to record the historical access address; The first page boundary check unit is used to calculate an offset value prefetch address according to the received optimal offset value and the current access address, and perform a page boundary check on the offset value prefetch address to determine whether the offset value prefetch address exceeds the address access boundary: if not, the first page boundary check unit sends a prefetch request according to the offset value prefetch address.
[0011] Preferably, the preset rule is: When the processor read misses: the optimal offset value table performs parallel training on each preset offset value according to the current access address based on an offset value training algorithm to obtain a training result corresponding to the preset offset value; The recent request table matches each training result with the historical access address to obtain a training score corresponding to the training result; When the training score corresponding to one of the training results is equal to a first preset threshold, the preset offset value corresponding to the training result is used as the optimal offset value.
[0012] Preferably, the stride prefetch unit includes a first reference prediction table, a tag scanning unit, a stride comparison unit, a first prefetch request generation unit, and a second page boundary check unit; The first reference prediction table is used to store the historical access addresses, a plurality of stride values, and the stride confidence levels corresponding to the stride values; The tag scanning unit is used to check and update the tags of the first reference prediction table; The stride comparison unit is used to calculate a current stride value according to the current access address, and compare the current stride value with all the stride values in the first reference prediction table. If the current stride value matches one of the stride values in the first reference prediction table and the corresponding stride confidence level is greater than or equal to a second preset threshold, a matching signal is sent to the first prefetch request generation unit; The first prefetch request generation unit is used to judge whether there is one of the stride values in the first reference prediction table according to the received matching signal, and the corresponding stride confidence level is greater than or equal to the second preset threshold: if so, a stride prefetch address is generated and the stride prefetch address is sent to the second page boundary check unit; The second page boundary check unit is used to perform a page boundary check on the received stride prefetch address to judge whether the stride prefetch address exceeds the address access boundary: if not, the second page boundary check unit issues a prefetch request according to the stride prefetch address.
[0013] Preferably, the Markov prefetch unit includes a difference calculation unit, a difference shift register, a difference index scanning unit, a second reference prediction table, a confidence level check unit, a prediction difference table, a second prefetch request generation unit, and a third page boundary check unit; The second reference prediction table is used to store difference indexes, a plurality of prediction differences, and the difference confidence levels corresponding to the prediction differences; The difference calculation unit is used to calculate a current difference according to the current access address and send the current difference to the difference shift register; The difference shift register is used to store the received current difference and generate a current difference index, and at the same time send the current difference index and the current difference to the difference index scanning unit and the prediction difference table respectively; The difference index scanning unit is used to check and update the difference indexes of the second reference prediction table according to the received current difference index; The confidence check unit is used to update the difference confidence in the second reference prediction table. If one of the difference confidences is greater than or equal to a third preset threshold, a prediction difference verification signal is sent to the prediction difference table. The prediction difference table is used to update the prediction difference in the second reference prediction table according to the received current difference, and to read the corresponding prediction difference from the second reference prediction table according to the received prediction difference verification signal, and output it to the second prefetch request generation unit. The second prefetch request generation unit is used to generate a Markov prefetch address according to the current access address and the received prediction difference, and send the Markov prefetch address to the third page boundary check unit. The third page boundary check unit is used to perform a page boundary check on the received Markov prefetch address to determine whether the Markov prefetch address exceeds the address access boundary. If not, the third page boundary check unit issues a prefetch request according to the Markov prefetch address.
[0014] Preferably, the prefetch management unit includes a stride prefetch counter, a Markov prefetch counter, an optimal offset value prefetch counter, and a prefetch status unit. The stride prefetch counter is used to record the confidence of the stride prefetch unit in the training mode. The Markov prefetch counter is used to record the confidence of the Markov prefetch unit in the training mode. The optimal offset value prefetch counter is used to record the confidence of the optimal offset value prefetch unit in the training mode. The prefetch status unit is used to adjust the working states of the stride prefetch unit, the Markov prefetch unit, and the optimal offset value prefetch unit respectively according to the count values of the stride prefetch counter, the Markov prefetch counter, and the optimal offset value prefetch counter.
[0015] Preferably, when the count value corresponding to any one of the stride prefetch counter, the Markov prefetch counter, and the optimal offset value prefetch counter is greater than or equal to a preset count threshold, the corresponding stride prefetch unit or Markov prefetch unit or optimal offset value prefetch unit is enabled to issue a prefetch request, and the remaining units stop training.
[0016] Preferably, when the count values corresponding to at least two of the stride prefetch counter, the Markov prefetch counter, and the optimal offset value prefetch counter are simultaneously greater than or equal to a preset count threshold, the selection priorities of the prefetch management unit from first to last are: the stride prefetch unit, the Markov prefetch unit, and the optimal offset value prefetch unit.
[0017] Preferably, the prefetch module further includes an output buffer unit; the output buffer unit is used to temporarily store the prefetch requests issued by the prefetch module.
[0018] Compared with the prior art, the prefetch module proposed by the present invention is used for prefetching of each level of cache. For the secondary and tertiary caches, if there is only a single address sequence, the stride prefetch unit and the Markov prefetch unit can be selected according to a fixed address pattern. If there are multiple address sequences at the same time, the optimal offset value prefetch unit can find the optimal offset value for prefetching. Through the prefetch manager module, different prefetch units can be selected according to the current access address of the processor, thereby improving the accuracy of prefetching; at the same time, the output buffer unit can cancel the prefetch request when the prefetch request times out and the prefetch output queue overflows, so as to reduce bandwidth resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The present invention will be described in detail below with reference to the accompanying drawings. Through the detailed description in combination with the following drawings, the above or other aspects of the present invention will become clearer and easier to understand. In the drawings: Figure 1 is a flowchart of the offset value prefetch algorithm in the related art; Figure 2 is a flowchart of the offset value training of the offset value prefetch algorithm in the related art; Figure 3 is a schematic structural diagram of the hardware prefetch system provided by an embodiment of the present invention; Figure 4 is a schematic structural diagram of the optimal offset value prefetch unit of the hardware prefetch system provided by an embodiment of the present invention; Figure 5 is a flowchart of the offset value training algorithm of the hardware prefetch system provided by an embodiment of the present invention; Figure 6 is a schematic structural diagram of the stride prefetch unit of the hardware prefetch system provided by an embodiment of the present invention; Figure 7 is a schematic structural diagram of the Markov prefetch unit of the hardware prefetch system provided by an embodiment of the present invention; Figure 8 is a schematic structural diagram of the prefetch management unit of the hardware prefetch system provided by an embodiment of the present invention; Figure 9It is a logical schematic diagram of the prefetcher status unit of the hardware prefetch system provided by an embodiment of the present invention; Figure 10 It is a structural schematic diagram of the output buffer unit of the hardware prefetch system provided by an embodiment of the present invention.
[0020] In the figure, 100 is the hardware prefetch system; 1 is the processor; 2 is the first-level cache; 3 is the second-level cache; 4 is the third-level cache; 5 is the prefetch module; 51 is the prefetch management unit; 511 is the stride prefetch counter; 512 is the Markov prefetch counter; 513 is the best offset value prefetch counter; 514 is the prefetcher status unit; 52 is the stride prefetch unit; 521 is the first reference prediction table; 522 is the tag scanning unit; 523 is the stride comparison unit; 524 is the first prefetch request generation unit; 525 is the second page boundary check unit; 53 is the best offset value prefetch unit; 531 is the best offset value table; 532 is the recent request table; 533 is the first page boundary check unit; 54 is the Markov prefetch unit; 541 is the difference calculation unit; 542 is the difference shift register; 543 is the second reference prediction table; 544 is the confidence check unit; 545 is the prediction difference table; 546 is the second prefetch request generation unit; 547 is the third page boundary check unit; 548 is the difference index scanning unit; 55 is the output buffer unit; 6 is the memory. Detailed implementation manners
[0021] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0022] Please refer to Figure 3 , the present invention provides a hardware prefetch system 100, and the hardware prefetch system 100 includes a processor 1, a first-level cache 2, a second-level cache 3, a third-level cache 4, a prefetch module 5, and a memory 6.
[0023] The processor 1 is used to read and execute instructions in the first-level cache 2, the second-level cache 3, the third-level cache 4, and the memory 6.
[0024] The first-level cache 2 is used to cache instructions.
[0025] The second-level cache 3 is used to cache instructions and, when the processor 1 fails to hit the first-level cache 2, serve as the object for the processor 1 to read instructions.
[0026] The third-level cache 4 is used to cache instructions and, when the processor 1 fails to hit the instructions in the first-level cache 2 and the second-level cache 3, serve as the object for the processor 1 to read instructions.
[0027] The memory 6 is used to store instructions and serves as the object for the processor 1 to read instructions when the instructions read from the first-level cache 2, the second-level cache 3, and the third-level cache 4 all miss.
[0028] The prefetch module 5 is used to send prefetch requests to the second-level cache 3, the third-level cache 4, and the memory 6 respectively. Specifically, the prefetch module 5 includes a plurality of units respectively arranged between the first-level cache 2, the second-level cache 3, the third-level cache 4, and the memory 6.
[0029] The prefetch module 5 includes a prefetch management unit 51, a stride prefetch unit 52, an optimal offset value prefetch unit 53, and a Markov prefetch unit 54.
[0030] The prefetch management unit 51 is used to dynamically select one of the stride prefetch unit 52, the optimal offset value prefetch unit 53, and the Markov prefetch unit 54, and send prefetch requests to the second-level cache 3, the third-level cache 4, and the memory 6 through the selected unit.
[0031] The stride prefetch unit 52 is used to send prefetch requests to the second-level cache 3, the third-level cache 4, and the memory 6 according to the stride between the current access address of the processor 1 and the historical access address of the processor 1.
[0032] The optimal offset value prefetch unit 53 is used to send prefetch requests to the second-level cache 3, the third-level cache 4, and the memory 6 according to the optimal offset value between the current access address and the historical access address.
[0033] The Markov prefetch unit 54 is used to send prefetch requests to the second-level cache 3, the third-level cache 4, and the memory 6 according to the difference between the current access address and the historical access address.
[0034] In the embodiment of the present invention, please refer to Figure 4 , Figure 4 FIG. is a schematic structural diagram of the optimal offset value prefetch unit 53 of the hardware prefetch system 100 provided by the embodiment of the present invention. The optimal offset value prefetch unit 53 includes an optimal offset value table 531, a recent request table 532, and a first page boundary check unit 533.
[0035] The optimal offset value table 531 is used to store a plurality of preset offset values, and select the optimal offset value from the preset offset values according to the current access address and a preset rule and send it to the first page boundary check unit 533. The preset rule is: When the processor 1 has a read miss, the optimal offset value table 531 parallelly trains each of the preset offset values according to the offset value training algorithm based on the current access address, and obtains the training result corresponding to the preset offset value; The recent request table 532 matches each of the training results with the historical access address to obtain the training score corresponding to the training result; When the training score corresponding to one of the training results is equal to the first preset threshold, the preset offset value corresponding to the training result is used as the optimal offset value.
[0036] The recent request table 532 is used to record the historical access address; The first page boundary check unit 533 is used to calculate the offset value prefetch address according to the received optimal offset value and the current access address, and perform a page boundary check on the offset value prefetch address to determine whether the offset value prefetch address exceeds the address access boundary: if not, the first page boundary check unit 533 issues a prefetch request according to the offset value prefetch address.
[0037] Specifically, the steps of issuing a prefetch request through the optimal offset value prefetch unit 53 are as follows: When there is a cache miss or a prefetch hit in the processor 1, the current access address of the processor 1 is subtracted from each preset offset value in the optimal offset value table 531, and a test tag is sent to the recent request table 532 according to each subtraction result to determine whether it hits the historical access address in the recent request table 532. If there is a hit, the training score is incremented by one. When the training score corresponding to one of the training scores is equal to the first preset threshold, the corresponding preset offset value is used as the optimal offset value and sent to the first page boundary check unit 533. The first page boundary check unit 533 calculates the offset value prefetch address according to the optimal offset value and the current access address, and performs a page boundary check on the offset value prefetch address. When the offset value prefetch address is within the address access boundary range, a prefetch request is issued. The optimal offset value prefetch unit 53 is used for prefetching of the tertiary cache 4 and the memory 6, and degenerates into the next line prefetcher (i.e., the offset value is fixed at 1) when prefetching the secondary cache 3.
[0038] Please refer to Figure 5 , Figure 5It is a flowchart of the offset value training algorithm of the hardware prefetching system 100 provided by an embodiment of the present invention. Since the optimal offset value prefetching unit 53 proposed in the present invention compares all preset offset values each time during training, it will occupy more combinational logic. Therefore, the offset value list is not easy to set too large. Therefore, the optimal offset value table 531 includes twelve offset value lists (i.e., 1, 2, 3, 4, 5, 6, 8, 9, 10, 12, 15, 16). If the training score corresponding to one of the preset offset values reaches the first preset threshold, the preset offset value D will be updated to the offset value with the highest score (i.e., the optimal offset value), and all training scores will be reset and the training sequence will be restarted again. If the number of missing ones is equal to the predefined missing threshold, it means that this round of training fails and the training will start again. Among them, the training time of the optimal offset value prefetching unit 53 can be customized according to the actual situation.
[0039] In an embodiment of the present invention, the stride prefetching unit 52 generates a prefetch request by detecting the stride of a continuous address sequence and adding the detected stride to the last observed access address.
[0040] Please refer to Figure 6 , Figure 6 It is a schematic structural diagram of the stride prefetching unit 52 of the hardware prefetching system 100 provided by an embodiment of the present invention. The stride prefetching unit 52 includes a first reference prediction table 521, a tag scanning unit 522, a stride comparison unit 523, a first prefetch request generation unit 524, and a second page boundary check unit 525.
[0041] The first reference prediction table 521 is used to store the historical access address, multiple stride values, and the stride confidence corresponding to the stride values. Among them, each entry in the first reference prediction table 521 is assigned a valid bit to indicate whether the content of each entry is filled. Among them, the entries in the first reference prediction table 521 are replaced by the PLRU (Pseudo - LRU) algorithm; The tag scanning unit 522 is used to check and update the tags of the first reference prediction table 521. The tag scanning unit 522 scans the first reference prediction table 521 and compares the tags with the current access address. For the address (PC)-based mode, the tag will be the program counter pointer, and for the memory 6 address-based mode, the tag will be the address within the 4KB range.
[0042] The stride comparison unit 523 is configured to calculate a current stride value according to the current access address (obtained by subtracting the current access address from the last access address), and compare the current stride value with all the stride values in the first reference prediction table 521. If the current stride value matches one of the stride values in the first reference prediction table 521 and the corresponding stride confidence is greater than or equal to a second preset threshold, a matching signal is sent to the first prefetch request generation unit 524.
[0043] The first prefetch request generation unit 524 is configured to determine whether there is a stride value in the first reference prediction table 521 whose corresponding stride confidence is greater than or equal to the second preset threshold according to the received matching signal. If so, a stride prefetch address is generated and sent to the second page boundary check unit 525.
[0044] The second page boundary check unit 525 is configured to perform a page boundary check on the received stride prefetch address to determine whether the stride prefetch address exceeds the address access boundary: if not, the second page boundary check unit 525 issues a prefetch request according to the stride prefetch address.
[0045] In an embodiment of the present invention, please refer to Figure 7 , Figure 7 FIG. is a schematic structural diagram of a Markov prefetch unit 54 of a hardware prefetch system 100 provided by an embodiment of the present invention. The Markov prefetch unit 54 includes a difference calculation unit 541, a difference shift register 542, a second reference prediction table 543, a confidence check unit 544, a prediction difference table 545, a second prefetch request generation unit 546, a third page boundary check unit 547, and a difference index scan unit 548; The second reference prediction table 543 is configured to store difference indexes, a plurality of prediction differences, and difference confidences corresponding to the prediction differences. Each entry in the second reference prediction table 543 is assigned a valid bit to indicate whether the content of each entry is filled. Among them, the replacement method for the second reference prediction table 543 and the second prediction difference table 545 entries is the PLRU algorithm.
[0046] The difference calculation unit 541 is configured to calculate a current difference according to the current access address and send the current difference to the difference shift register 542.
[0047] The difference shift register 542 is configured to store the received current difference and generate a current difference index, and send the current difference index and the current difference to the difference index scan unit 548 and the prediction difference table 545 respectively.
[0048] The difference index scanning unit 548 is configured to check and update the difference index of the second reference prediction table 543 according to the received current difference index. When the difference index scanning unit 548 scans the second reference prediction table 543, the difference index from the second reference prediction table 543 is compared with the current difference index shifted out from the difference shift register 542. If there is no matching difference index, a new difference index will be updated to the second reference prediction table 543.
[0049] The confidence check unit 544 is configured to update the difference confidence in the second reference prediction table 543. If one of the difference confidences is greater than or equal to a third preset threshold, a prediction difference verification signal is sent to the prediction difference table 545. Specifically, if the current difference index and the predicted difference both match, the corresponding difference confidence value is incremented by 1. If the difference confidence value is greater than or equal to the third preset threshold, a prediction difference verification signal is asserted to indicate that the selected predicted difference can be used to generate a prefetch request address.
[0050] The prediction difference table 545 is configured to update the predicted difference in the second reference prediction table 543 according to the received current difference, and read the corresponding predicted difference from the second reference prediction table 543 according to the received prediction difference verification signal, and output it to the second prefetch request generation unit 546.
[0051] Specifically, the prediction difference table 545 performs 2 different operations for checking, prediction difference check and prediction difference update. For the prediction difference check, if a prediction difference verification signal is received, the corresponding predicted difference is read out from the second reference prediction table 543 and output to the second prefetch request generation unit 546 to calculate the prefetch address. For the prediction difference update, if the current difference (new predicted difference) does not match the predicted difference (old predicted difference) in the second reference prediction table 543, the current difference is written into the second reference prediction table 543, and a table entry in the second reference prediction table 543 is replaced using the PLRU algorithm.
[0052] The second prefetch request generation unit 546 is configured to generate a Markov prefetch address according to the current access address and the received predicted difference, and send the Markov prefetch address to the third page boundary check unit 547.
[0053] The third page boundary check unit 547 is configured to perform a page boundary check on the received Markov prefetch address to determine whether the Markov prefetch address exceeds the address access boundary. If not, the third page boundary check unit 547 issues a prefetch request according to the Markov prefetch address.
[0054] Specifically, the Markov prefetch unit 54 observes the difference between the current access address and the last miss address, and updates the second reference prediction table 543. If the difference index corresponding to the current access address is not recorded in the second reference prediction table 543, it is recorded in the second reference prediction table 543, and the next miss address difference corresponding to the difference index of the current access address is recorded in the prediction difference table 545. Multiple prediction difference table 545 entries will be allocated for each address difference index. Once the address difference recorded in the prediction difference table 545 appears again, followed by the indexed address difference, the corresponding difference confidence value is incremented by 1. When the difference confidence value is greater than or equal to the third preset threshold, a predicted difference verification signal is asserted and sent to the second prefetch request generation unit 546 to issue a prefetch request.
[0055] In an embodiment of the present invention, the prefetch management unit 51 can dynamically manage the operations of the stride prefetch unit 52, the Markov prefetch unit 54, and the optimal offset value prefetch unit 53 in different time slots according to the access pattern or the historical prefetch accuracy result.
[0056] Please refer to Figure 8 , Figure 8 FIG. 11 is a schematic structural diagram of the prefetch management unit 51 of the hardware prefetch system 100 provided by an embodiment of the present invention. The prefetch management unit 51 includes a stride prefetch counter 511, a Markov prefetch counter 512, an optimal offset value prefetch counter 513, and a prefetch status unit 514; The stride prefetch counter 511 is used to record the confidence of the stride prefetch unit 52 in the training mode; The Markov prefetch counter 512 is used to record the confidence of the Markov prefetch unit 54 in the training mode; The optimal offset value prefetch counter 513 is used to record the confidence of the optimal offset value prefetch unit 53 in the training mode; The prefetch status unit 514 is used to adjust the working status of the stride prefetch unit 52, the Markov prefetch unit 54, and the optimal offset value prefetch unit 53 respectively according to the count values of the stride prefetch counter 511, the Markov prefetch counter 512, and the optimal offset value prefetch counter 513. Among them, the working status includes an initial state, a preparatory state, and a selection state. The initial state and the preparatory state both represent the states of the three prefetch units (the stride prefetch unit 52, the Markov prefetch unit 54, and the optimal offset value prefetch unit 53) during the training process, and the selection state represents the state where the prefetch unit starts to issue a prefetch request.
[0057] In the embodiment of the present invention, the prefetch management unit 51 can select a corresponding prefetch unit according to different address patterns. Only the selected prefetch unit will be activated, thereby reducing power consumption and improving the accuracy of prefetching.
[0058] Please refer to Figure 9 , Figure 9 which is a logical schematic diagram of the prefetch status unit 514 of the hardware prefetch system 100 provided by the embodiment of the present invention. Define the initial line, confidence value line, nearly full line, and full line in ascending order. The prefetch status unit 514 adjusts the working states of the stride prefetch unit 52, the Markov prefetch unit 54, and the optimal offset value prefetch unit 53 according to the count values of the three prefetch units within the intervals where the initial line, confidence value line, nearly full line, and full line belong. And there is a certain damping value every time a line is crossed to prevent the counter from frequently switching between various states. Among them, the interval at the initial line is the initial state, the interval from the initial line to the nearly full line is the preparatory state, and the interval from the nearly full line to the full line is the selection state.
[0059] For example, when the three prefetch units are reset, the counters corresponding to all prefetch units are at the initial line. At this time, the next line prefetch starts by default, and the other three prefetch modules 5 enter the pattern recognition state.
[0060] When the count value corresponding to any one of the stride prefetch counter 511, the Markov prefetch counter 512, and the optimal offset value prefetch counter 513 is greater than or equal to the pre-designed count threshold (nearly full line), the corresponding stride prefetch unit 52 or Markov prefetch unit 54 or optimal offset value prefetch unit 53 is enabled to issue a prefetch request, and the remaining units stop training.
[0061] When the count values corresponding to at least two of the stride prefetch counter 511, the Markov prefetch counter 512, and the optimal offset value prefetch counter 513 are simultaneously greater than or equal to the pre-designed count threshold (nearly full line), the selection priorities of the prefetch management unit 51 from first to last are: the stride prefetch unit 52, the Markov prefetch unit 54, and the optimal offset value prefetch unit 53.
[0062] In the embodiment of the present invention, the prefetch module 5 further includes an output buffer unit 55; the output buffer unit 55 is used to temporarily store the prefetch requests issued by the prefetch module 5.
[0063] Specifically, please refer to Figure 10 , Figure 10It is a schematic structural diagram of the output buffer unit 55 of the hardware prefetching system 100 provided by an embodiment of the present invention. Each prefetching unit is connected to the output buffer unit 55. The generated prefetching requests will be temporarily stored in the output buffer unit 55 first. A counter is set in the output buffer unit 55, and the counter is initially set to 0. If the output buffer unit 55 is not empty, it will increase by 1 at the start of each clock cycle. If data is read from the output buffer unit 55, the counter will be reset to 0. If the counter value exceeds a threshold (no data is sent to the downstream sub-module for a period of time), all uncompleted prefetching requests in the output buffer unit 55 will be flushed and discarded.
[0064] However, in some cases, the downstream module may be unable to receive prefetching requests for a period of time, and the output buffer unit 55 may eventually be full. If the upstream module continuously generates cache miss requests, the output buffer unit 55 may overflow. To handle this problem, once the output buffer unit 55 is full and a new prefetching request is generated, the oldest prefetching request in the output buffer unit 55 will be popped and deleted, and spare space will be reversed for the incoming new prefetching request.
[0065] Compared with the prior art, the prefetching module proposed by the present invention is used for prefetching of each-level caches. For the secondary and tertiary caches, if there is only a single address sequence, the stride prefetching unit and the Markov prefetching unit can be selected according to the fixed address pattern. If there are multiple address sequences at the same time, the optimal offset value can be found through the optimal offset value prefetching unit for prefetching. The prefetching manager module can select different prefetching units according to the current access address of the processor, thereby improving the accuracy of prefetching; at the same time, it can cancel prefetching requests when the prefetching requests time out and the prefetching output queue overflows through the output buffer unit 55 to reduce bandwidth resources.
[0066] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitations, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0067] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to enable a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0068] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. What is disclosed is only the preferred embodiments of the present invention. However, the present invention is not limited to the above specific implementation manners. The above specific implementation manners are only illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many equivalent changes in form without departing from the purpose of the present invention and the scope protected by the claims, and all of them belong to the protection scope of the present invention.
Claims
1. A hardware prefetching system, characterized in that The hardware prefetching system includes a processor, a level-1 cache, a level-2 cache, a level-3 cache, a prefetching module, and a memory; The processor is used to read and execute instructions in the level-1 cache, the level-2 cache, the level-3 cache, and the memory; The level-1 cache is used to cache instructions; The level-2 cache is used to cache instructions and, when the processor fails to hit when reading the level-1 cache, serves as the object for the processor to read instructions; The level-3 cache is used to cache instructions and, when the processor fails to hit when reading instructions in the level-1 cache and the level-2 cache, serves as the object for the processor to read instructions; The memory is used to store instructions and, when the processor fails to hit when reading instructions in the level-1 cache, the level-2 cache, and the level-3 cache, serves as the object for the processor to read instructions; The prefetching module is used to send prefetch requests to the level-2 cache, the level-3 cache, and the memory respectively; The prefetching module includes a prefetch management unit, a stride prefetch unit, an optimal offset value prefetch unit, and a Markov prefetch unit; The prefetch management unit is used to dynamically select one of the stride prefetch unit, the optimal offset value prefetch unit, and the Markov prefetch unit, and send prefetch requests to the level-2 cache, the level-3 cache, and the memory through the selected unit; The stride prefetch unit is used to send prefetch requests to the level-2 cache, the level-3 cache, and the memory according to the stride between the current access address of the processor and the historical access address of the processor; The optimal offset value prefetch unit is used to send prefetch requests to the level-2 cache, the level-3 cache, and the memory according to the optimal offset value between the current access address and the historical access address; The Markov prefetch unit is used to send prefetch requests to the level-2 cache, the level-3 cache, and the memory according to the difference between the current access address and the historical access address.
2. The hardware prefetching system according to claim 1, wherein The optimal offset value prefetch unit includes an optimal offset value table, a recent request table, and a first page boundary check unit; The optimal offset value table is used to store multiple preset offset values, select the optimal offset value from the preset offset values according to the current access address and a preset rule, and send it to the first page boundary check unit; The recent request table is used to record the historical access address; The first page boundary check unit is used to calculate an offset value prefetch address based on the received optimal offset value and the current access address, and perform a page boundary check on the offset value prefetch address to determine whether the offset value prefetch address exceeds the address access boundary: if not, the first page boundary check unit sends a prefetch request according to the offset value prefetch address.
3. The hardware prefetching system according to claim 2, wherein The preset rule is: When the processor fails to hit when reading: The optimal offset value table performs parallel training on each of the preset offset values according to the current access address based on an offset value training algorithm, and obtains a training result corresponding to the preset offset value; The recent request table matches each of the training results with the historical access address to obtain a training score corresponding to the training result; When the training score corresponding to one of the training results is equal to a first preset threshold, the preset offset value corresponding to the training result is used as the optimal offset value.
4. The hardware prefetching system according to claim 1, characterized in that, The stride prefetch unit includes a first reference prediction table, a tag scanning unit, a stride comparison unit, a first prefetch request generation unit, and a second page boundary check unit; The first reference prediction table is used to store the historical access address, a plurality of stride values, and the stride confidence corresponding to the stride values; The tag scanning unit is used to check and update the tags of the first reference prediction table; The stride comparison unit is used to calculate a current stride value according to the current access address, and compare the current stride value with all the stride values in the first reference prediction table. If the current stride value matches one of the stride values in the first reference prediction table, and the corresponding stride confidence is greater than or equal to a second preset threshold, a matching signal is sent to the first prefetch request generation unit; The first prefetch request generation unit is used to judge whether there is one of the stride values in the first reference prediction table whose corresponding stride confidence is greater than or equal to the second preset threshold according to the received matching signal; if so, a stride prefetch address is generated and sent to the second page boundary check unit; The second page boundary check unit is used to perform a page boundary check on the received stride prefetch address to judge whether the stride prefetch address exceeds the address access boundary: if not, the second page boundary check unit issues a prefetch request according to the stride prefetch address.
5. The hardware prefetching system according to claim 1, wherein, The Markov prefetch unit includes a difference calculation unit, a difference shift register, a difference index scanning unit, a second reference prediction table, a confidence check unit, a prediction difference table, a second prefetch request generation unit, and a third page boundary check unit; The second reference prediction table is used to store difference indexes, a plurality of prediction differences, and the difference confidence corresponding to the prediction differences; The difference calculation unit is used to calculate a current difference according to the current access address and send the current difference to the difference shift register; The difference shift register is used to store the received current difference and generate a current difference index, and send the current difference index and the current difference to the difference index scanning unit and the prediction difference table respectively; The difference index scanning unit is used to check and update the difference indexes of the second reference prediction table according to the received current difference index; The confidence check unit is used to update the difference confidence in the second reference prediction table. If one of the difference confidences is greater than or equal to a third preset threshold, a prediction difference verification signal is sent to the prediction difference table; The predicted difference table is used to update the predicted differences in the second reference prediction table according to the received current difference, and read the corresponding predicted differences from the second reference prediction table according to the received predicted difference verification signal, and output them to the second prefetch request generation unit; The second prefetch request generation unit is used to generate a Markov prefetch address according to the current access address and the received predicted difference, and send the Markov prefetch address to the third page boundary check unit; The third page boundary check unit is used to perform a page boundary check on the received Markov prefetch address to determine whether the Markov prefetch address exceeds the address access boundary: if not, the third page boundary check unit issues a prefetch request according to the Markov prefetch address.
6. The hardware prefetching system according to claim 1, wherein The prefetch management unit includes a stride prefetch counter, a Markov prefetch counter, an optimal offset value prefetch counter, and a prefetch status unit; The stride prefetch counter is used to record the confidence level of the stride prefetch unit in the training mode; The Markov prefetch counter is used to record the confidence level of the Markov prefetch unit in the training mode; The optimal offset value prefetch counter is used to record the confidence level of the optimal offset value prefetch unit in the training mode; The prefetch status unit is used to adjust the working states of the stride prefetch unit, the Markov prefetch unit, and the optimal offset value prefetch unit respectively according to the count values of the stride prefetch counter, the Markov prefetch counter, and the optimal offset value prefetch counter.
7. The hardware prefetching system according to claim 6, wherein When the count value corresponding to any one of the stride prefetch counter, the Markov prefetch counter, and the optimal offset value prefetch counter is greater than or equal to a preset count threshold, the corresponding stride prefetch unit or Markov prefetch unit or optimal offset value prefetch unit is enabled to issue a prefetch request, and the remaining units stop training.
8. The hardware prefetching system according to claim 6, wherein When the count values corresponding to at least two of the stride prefetch counter, the Markov prefetch counter, and the optimal offset value prefetch counter are simultaneously greater than or equal to the preset count threshold, the selection priorities of the prefetch management unit from first to last are: the stride prefetch unit, the Markov prefetch unit, the optimal offset value prefetch unit.
9. The hardware prefetching system according to claim 1, wherein The prefetch module further includes an output buffer unit; the output buffer unit is used to temporarily store the prefetch requests issued by the prefetch module.
Citation Information
Patent Citations
Offset prefetching method, device for executing offset prefetching, computing equipment and medium
CN113778520A
Data prefetching method and device, computer equipment and storage medium
CN119576809A
System, apparatus and method for performing look-ahead lookup on predictive information in a cache memory
US20060041722A1
Method and apparatus for prefetching data to a lower level cache memory
US20060047915A1
Processor performance by dynamically re-adjusting the hardware stream prefetcher stride
US20150356015A1
Cited By
Prefetching method and system of CPU prefetcher based on multi-granularity access mode recognition
CN121116395A