Perceptron-based off-chip predictor

The perceptron-based off-chip predictor with a selective delay mechanism addresses inefficiencies in existing predictors by reducing useless DRAM transactions, resulting in significant performance improvements for memory-intensive workloads.

WO2025132923A1PCT designated stage expired Publication Date: 2025-06-26BARCELONA SUPERCOMPUTING CENT CENT NAT DE SUPERCOMPUTACION

Patent Information

Application Number
PCT/EP2024/087604
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-21
Filing Date
2024-12-19
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing off-chip predictors in CPUs are inefficient due to high numbers of useless DRAM transactions, leading to performance and energy overheads, especially in memory-intensive workloads.

Method used

A perceptron-based off-chip predictor with a selective delay mechanism, utilizing two thresholds (Thigh and rtow) to determine when to issue speculative DRAM requests, thereby reducing unnecessary DRAM transactions.

Benefits of technology

The proposed solution achieves a 11.5% speedup over baseline systems without off-chip prediction, improving prior art results by at least 37% over two-step predictors and 396% over conventional FLP predictors without selective delay.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024087604_26062025_PF_FP_ABST
    Figure EP2024087604_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a first-level perceptron (FLP) off-chip predictor communicatively connectable to a computing core and to a DRAM, wherein the computing core and the DRAM are communicatively connected through a multi-level cache hierarchy of levels L1D, L2C, …, LLC. The FLP is advantageously adapted with an FLP off-chip prediction mechanism comprising two thresholds, τ low and τ high . The invention also relates to a two-level perceptron (TLP) off-chip predictor comprising a first-level perceptron (FLP) off-chip predictor according to any of the preceding claims; and a second-level perceptron (SLP) off-chip predictor communicatively connectable to a multi-level cache hierarchy of levels L1D, L2C, …, LLC through a L1D prefetcher, wherein the multi-level cache hierarchy is communicatively connected to a computing core and to a DRAM.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] DESCRIPTION

[0002] PERCEPTRON-BASED OFF-CHIP PREDICTOR

[0003] FIELD OF THE INVENTION

[0004] The present invention relates to the field of computer science and, particularly, to the microarchitecture of central processing units (CPUs), cache hierarchy and prefetching. More specifically, the invention relates to a perceptron-based predictor which improves the performance of memory-intensive workloads, by reducing the triggering of accesses to dynamic random-access memories (DRAMs).

[0005] BACKGROUND OF THE INVENTION

[0006] Emerging workloads from various computing domains have large data footprints, which are orders of magnitude larger than the capacity of current cache hierarchies. These workloads frequently trigger DRAM accesses, spending a substantial portion of their execution time waiting for data transfers to / from DRAM to complete, with a detrimental effect on performance and energy. In the prior art, various techniques to mitigate the performance and energy overheads of these applications have been proposed, and can be broadly classified into four categories: i) Off-chip prediction schemes that predict whether a memory access will result in a DRAM access or hit in the cache hierarchy. ii) Data prefetching with adaptive filters, configured to ensure that only correct prefetches will be issued. iii) Cache bypassing configured to avoid caching blocks that will not be referenced in a given period of time. iv) Specific cache designs and optimizations for specific workload types.

[0007] Despite their potential for determining the location of requested data in the memory hierarchy, the known off-chip predictors have important drawbacks that not only limit the performance of the memory subsystem, but also hinder their implementation in real-world designs. For example, some prior-art off-chip predictors trigger two memory accesses, one to DRAM and a second regular request to the cache hierarchy, when it is predicted that the corresponding demand load access will be served from DRAM. While this approach can potentially reduce the latency of a load request being served from DRAM, it may also significantly increase the number of DRAM transactions. This issue is a critical aspect of bandwidth-constrained scenarios since a large fraction of the inaccurate off-chip predictions is served by the first level data cache (L1 D). However, constantly delaying the off-chip predictions until the L1 D lookup is completed would result in suboptimal performance gains since a substantial portion of the off-chip predictions are accurate. Therefore, finding a microarchitectural scheme that selectively delays the off-chip predictions with modest confidence, until the L1 D lookup is resolved, has the potential to significantly reduce the number of useless DRAM transactions and deliver higher performance.

[0008] For instance, the research thesis study “Evaluation of L1 residence for perceptron filter enhanced signature path prefetcher” (A. Staggs, Undergraduate Research Scholars program by the Texas A&M University, May 2020) describes various integrated circuit technologies and multi-level caches. In that context, the study proposes evaluating the effect of an L1 D residence for a Perceptron-Filtered Signature Path Prefetcher (PPF) concluding that, while an unoptimized movement of the PPF from the level-two cache (L2C) to the L1 D shows performance degradation, optimizations such as using the L1 D data stream to prefetch to all cache levels, and updating table sizes and lengths, can match a scenario where PPF is located alongside the L2C.

[0009] The article “Hermes: Accelerating Long-Latency Load Requests via Perceptron-Based Off- Chip Load Prediction” (R. Bera et al., arXiv:2209.00188v3 [cs.AR], 30 September 2023) describes a lightweight, perceptron-based off-chip load predictor, that learns to identify off- chip load requests using multiple program features (e.g., sequence of program counters, byte off set of a load request). For every load request generated by the processor, the predictor observes a set of program features to predict whether the load would go off chip. If the load is predicted to go off chip, then the predictor issues a speculative load request, directly to the main memory controller once the load’s physical address is generated. If the prediction is correct, the load eventually misses the cache hierarchy and waits for the ongoing speculative load request to finish. Thus, the predictor completely hides the on-chip cache hierarchy access latency from the critical path of the correctly predicted off chip load. The perceptron learning algorithm starts by initializing the weight of each neuron and iteratively trains the weights using each input vector from the training dataset in two steps. First, for an input vector, the perceptron network computes a binary output and the current weight values of its neurons. Second, if the computed output differs from the desired output for that input vector provided by the dataset, the weight of each neuron is updated. This iterative process is repeated until the error between the computed and desired output falls below a user specified threshold.

[0010] Finally, the article “Perceptron Based Prefetch Filtering’’, (E. Bhatia et al., ISCA '19: Proceedings of the 46th International Symposium on Computer Architecture, June 2019) refers to a perceptron-based prefetch filtering (PPF) technique to increase the coverage of the prefetches generated by an underlying prefetcher without negatively impacting accuracy. PPF enables more aggressive tuning of the underlying prefetcher, leading to increased coverage by filtering out the growing numbers of inaccurate prefetches such aggressive tuning implies. This document also explores a range of features used to train PPF’s perceptron layer to identify inaccurate prefetches.

[0011] Other known approaches in the field of the invention are disclosed, for instance, in the articles: J. Kim et al., “Kill the Program Counter” XP058338263, II. David et al., “Perceptronbased filtering of futile prefetches in embedded VLIB DSPs” XP086318812, and M. Ferdman et al., “Last-touch Correlated Data Streaming”, XP031091893, or in patent application US 2016 / 019155 A1 (A. Radhakrishnan et al.).

[0012] Some of the previous approaches have successfully applied prefetch filtering at the lower- level caches. However, they are neither agile nor responsive enough, since they are typically optimized on top of specific prefetch engines, incurring significant area overheads, and do not produce fully accurate predictions. The present invention is aimed at solving this limitation, by proposing a novel perceptron-based off-chip predictor that combines the benefits of off-chip prediction schemes and data prefetching with adaptive filters, in a synergistic and cost-effective manner.

[0013] SUMMARY OF THE INVENTION

[0014] To solve the problems described in the preceding section, a first object of the invention relates to a first-level perceptron (FLP) off-chip predictor communicatively connectable to a computing core and to a DRAM, wherein the computing core and the DRAM are communicatively connected through a multi-level cache hierarchy of (L1 D, L2C, ... , LLC) levels.

[0015] Advantageously in the invention, the FLP off-chip predictor is adapted with an FLP off-chip prediction mechanism comprising two thresholds, r / owand Thigh wherein, preferably, Thigh indicates a probability threshold for the corresponding demand load request to miss in all cache levels, and TIOWindicates a probability threshold for the corresponding demand load request not to miss in all or in any of the cache levels. Under said prediction mechanism, when the FLP off-chip predictor is connected to the computing core and the DRAM, and the computing core receives a demand load request, the FLP off-chip prediction mechanism is configured to perform the following steps:

[0016] - the FLP off-chip predictor is consulted by the computing core;

[0017] - the FLP off-chip predictor produces a confidence value used to drive the FLP off- chip prediction mechanism;

[0018] - the confidence value is compared with Thigh,

[0019] - if the confidence value greater than Thigh, the FLP issues a speculative DRAM request from the computing core;

[0020] - if the confidence value does not exceed Thigh, but does exceed rtow, the demand load request is tagged as predicted off-chip, and is sent to a L1 D cache;

[0021] - if the predicted off-chip request results in a miss in the L1 D, the tag is read, and the speculative DRAM request is issued from the L1 D cache;

[0022] - if the confidence value does not exceed rtow, the FLP off-chip predictor does not issue the speculative DRAM request from the computing core.

[0023] Compared to the known alternatives of the prior art, namely: i) a conventional FLP predictor (i.e., without a selective delay mechanism); ii) a second-level perceptron (SLP) predictor alone; iii) a two-Step Predictor (TSP) consisting of an FLP off-chip predictor without the selective delay mechanism, in combination with an SLP off-chip predictor but without being based on FLP output, the present invention obtains a 11.5% speedup over a baseline system without off-chip prediction the baseline, improving the prior art results in at least a 37% yield ratio over a TSP, and a 396% over a conventional FLP off-chip predictor without selective delay mechanism.

[0024] In a preferred embodiment of the invention, the FLP off-chip prediction mechanism is further configured to correlate a probability of a demand load request going off-chip with a history of program counters, PC, and accessed memory blocks, based on a first set of legacy and / or leveling features are associated with a first weight table which is composed of confidence counters. More preferably, the first set of legacy features comprise at least one feature selected from: PC and cacheline offset, PC and byte offset, PC and first access, cacheline offset and first access, last-4 load PC. In a preferred embodiment of the invention, the FLP off-chip predictor is further configured with a first training algorithm to be trained upon completing a memory access, and the memory accessed memory block is returned to the computing core from the multi-level cache hierarchy. More preferably, the first training algorithm comprises the following steps:

[0025] - when the demand load request comes back to the computing core, the FLP off- chip predictor checks if the demand load request was a true off-chip demand load request, and the request required a DRAM access;

[0026] - if the demand load request was a true off-chip demand load request, the FLP off- chip predictor’s corresponding confidence counters of the first weight table are trained positively;

[0027] - if the demand load request was not a true off-chip demand load request, the FLP off-chip predictor’s corresponding confidence counters of the first weight table are trained negatively.

[0028] In a preferred embodiment of the invention, the off-chip predictor further comprises an SLP off-chip predictor, communicatively connectable to a multi-level cache hierarchy of (L1 D, L2C, ... , LLC) levels through a L1 D prefetcher, wherein the multi-level cache hierarchy is communicatively connected to a computing core and to a DRAM. Advantageously, the SLP off-chip predictor is adapted with an SLP off-chip prediction mechanism comprising a prefetching threshold, Tpref, such that, when the SLP off-chip predictor is connected to the multi-level cache hierarchy, and the L1 D prefetcher issues a prefetch request, the SLP off- chip prediction mechanism is configured to perform the following steps:

[0029] - the SLP off-chip predictor produces an output value used to drive the SLP off-chip prediction mechanism;

[0030] - the output value is compared with Tpref,

[0031] - if the output value exceeds Tpref, the prefetch request is discarded;

[0032] - if the output value does not exceed Tpref, the prefetch request is processed by the multi-level cache hierarchy.

[0033] As a result, the off-chip predictor of the invention can be also configured as a two-level perceptron (TLP) predictor, communicatively connectable to a computing core, to a DRAM, and to a multi-level cache hierarchy of (L1 D, L2C, ... , LLC) levels communicatively connecting the computing core and the DRAM, comprising:

[0034] - a first-level perceptron (FLP) off-chip predictor according to any of the embodiments described in the present description; and - a second-level perceptron (SLP) off-chip predictor according to any of the embodiments described in the present description.

[0035] In a preferred embodiment of the invention, the SLP off-chip prediction mechanism is further configured to correlate a probability of a demand load request going off-chip with a history of PC and accessed memory blocks, based on a second set of legacy or leveling features associated with a second weight table which is composed of confidence counters. More preferably, the second set of legacy features comprise at least one feature selected from: PC and cacheline offset, PC and byte offset, PC and first access, cacheline offset and first access, last-4 load PC, and the second set of leveling features comprise at least FLP prediction and cacheline offset.

[0036] In a preferred embodiment of the invention, the SLP off-chip prediction mechanism is further configured with a second training algorithm to be trained upon completing an L1 D prefetch request, and the L1 D prefetch is served. More preferably, the second training algorithm of the SLP off-chip prediction mechanism comprises the following steps:

[0037] - when the prefetch request comes back to the computing core, the SLP off-chip predictor checks if the prefetch request was a true off-chip prefetch request, implying that the prefetch request required a DRAM access;

[0038] - if the prefetch request was a true off-chip prefetch request, the SLP off-chip predictor’s corresponding confidence counters of the second weight table are trained positively;

[0039] - if the prefetch request was not a true off-chip prefetch request, the SLP off-chip predictor’s corresponding confidence counter of the second weight table are trained negatively.

[0040] Specific objects and preferred embodiments of the invention also refer to the claims submitted with the present document.

[0041] BRIEF DESCRIPTION OF DRAWINGS

[0042] Figure 1 shows a flowchart of a first-level perceptron (FLP) according to a preferred embodiment of the invention. Diamonds indicate decision points.

[0043] Figure 2 shows a flowchart of a second-level perceptron (SLP) according to a preferred embodiment of the invention. Diamonds indicate decision points. Figure 3 shows a diagram representing the organization and operation of a two-level perceptron (TLP) prediction mechanism according to a preferred embodiment of the invention.

[0044] DETAILED DESCRIPTION OF THE INVENTION

[0045] A set of three preferred embodiments of the invention, shown in Figures 1-3, will be now described for illustrative and not limiting purposes.

[0046] To remedy the deficiencies of the prior-art off-chip predictors, the present invention is advantageously designed to leverage off-chip prediction by forming effective prefetch filters for L1 D, thereby improving the performance of memory-intensive workloads. With this aim, the invention proposes a two-level perceptron (TLP) predictor approach. TLP constitutes, to the knowledge of the inventors, the first hardware design using a multi-level perceptron hardware approach, and preferably comprises two connected microarchitectural perceptron predictors: a first-level predictor (FLP) and a second level predictor (SLP). FLP is a perceptron hardware predictor that employs a novel selective delay mechanism, while SLP is a perceptron that leverages off-chip prediction to drive L1 D prefetch filtering using physical addresses as well as the FLP prediction as features, as described below:

[0047] First-Level Perceptron (FLP) Predictor

[0048] FLP is a microarchitectural hashed perceptron predictor that dynamically decides whether to consume its prediction, immediately or after a certain event has taken place based on two threshold values: Thigh and rtow. In a preferred embodiment, the FLP is proposed in the context of an off-chip prediction, although it can be applied to any perceptron approach using a selective delay method, based on said two thresholds. In the off-chip prediction context, an FLP dynamically decides whether to consume the prediction in the computing core, i.e., in parallel with the L1 D lookup, since L1 D caches are typically implemented as virtually indexed physically tagged (VI PT) structures, or upon an L1 D miss. This delayed decision mechanism is driven by the two threshold values: Thigh and rtow. Perceptron confidence values greater than Thigh indicate a high probability for the corresponding demand load request to miss in all cache levels, while values lower than r / owindicate the opposite, and intermediate values indicate the need for delaying the decision. In different embodiments of the invention, the FLP can consider one or more program features (see an example in Table 1 of the present document) to predict whether a demand load request will miss in the cache hierarchy. In the context of off-chip predictions, these features correlate the probability of a demand load request going off-chip with a history of program counters (PC) and accessed memory blocks. Each FLP feature is preferably associated with a weight table which is composed of confidence counters.

[0049] Table 1. List of example features used by the FLP and the SLP.

[0050] According to Table 1 , the following legacy and leveling features are described:

[0051] PC © cacheline offset: to compute this feature, an XOR operation is performed between the program counter of the load request (or the load that triggers the prefetch), and the cache line offset of the corresponding address within its virtual page (or physical frame). This feature estimates the probability of a load (or prefetch) request to be served off-chip when a load instruction with a given PC accesses (or triggers a prefetch to) a particular cache line offset within a page. By considering this probability, the predictor accurately predicts when and where off-chip requests are expected to occur. To enable PC-based features, the predictor passes the PC of a certain demand request to the corresponding prefetch request that is generated from it, thus providing prefetch requests with PC.

[0052] PC © byte offset: to compute this feature, an XOR operation is performed between the program counter of the load that requests (or triggers a prefetch to) the address and the byte offset of this address within the cache line. This feature is particularly useful in predicting off-chip load requests when a program has a streaming access pattern over a linearly allocated data structure. PC © first access: to compute this feature, the program counter is shifted one bit to the left, adding the first access hint to the most significant bit position. The first access bit indicates whether a cache line has been recently accessed by the program or not. This feature is particularly effective to predict off-chip requests when workloads display cyclic access patterns.

[0053] Cache line offset © first access: this feature is similar to the PC © first access feature, with the key difference being that it determines the probability of a request going off-chip when a specific cache line offset within a page has been recently accessed by the program.

[0054] Last-4 load PC: this feature value is computed using a shifted-XOR operation concerning the last four program counters. It is designed to capture the execution path of the program and correlate it with the probability of observing an off-chip request whenever the program follows the same execution path. By leveraging this feature, the predictor of the invention can make more accurate predictions on when off-chip requests are likely to occur, considering the current execution path of the program. This allows the predictor to accurately predict off-chip accesses of workloads that exhibit repeated execution paths.

[0055] FLP prediction © cacheline offset: this feature combines the FLP output bit of the cache block from which the prefetch request originates with the offset of the prefetched cache block in its physical memory page. The rationale of this feature is to correlate the probability of an L1 D prefetch request going off-chip when a certain cache line offset is touched with the off- chip prediction decision related to the block that triggered the prefetch request. The FLP prediction © cacheline offset feature is particularly important for workloads with high correlation between off-chip load demand requests and off-chip L1 D prefetch requests.

[0056] Figure 1 shows a flowchart of the FLP’s operation and illustrates how the confidence value produced by the FLP is used to drive the off-chip prediction mechanism. Upon a demand load request, FLP is consulted by the computing core. FLP uses the selected program features to index its weight tables, then read out and sum the corresponding weights to produce a confidence value. Then the confidence value is compared to the Thigh threshold. A confidence value greater than Thigh indicates a high probability for the corresponding demand load request to miss in all caches. In this case, the FLP issues a speculative DRAM request from the computing core without waiting for the L1 D lookup to resolve. However, if the confidence value does not exceed Thigh but does exceed the TIOWthreshold, the load request is tagged as predicted off-chip and is sent to the L1 D cache. If this request results in a miss in the L1 D, the tag is read, and a speculative DRAM request is issued from the L1 D. Thus, FLP avoids sending useless DRAM requests for loads that might hit in the on- chip caches. Finally, if the confidence value exceeds none of the two thresholds, the demand load request continues like a normal request, without triggering speculative DRAM access.

[0057] FLP is trained upon completing a memory access, i.e. , when the memory block is returned to the computing core from the cache hierarchy. When the request comes back to the computing core, the FLP checks if the request was a true off-chip load request, i.e., if this request required DRAM access. If the request was a true off-chip load request, the predictor’s corresponding weights are trained positively. Conversely, if the request was not a true off-chip load request, the predictor’s corresponding weights are trained negatively.

[0058] Second Level Perceptron (SLP) Predictor

[0059] The SLP is a perceptron-based off-chip predictor conceived to be used in the context of L1 D prefetch filtering. The SLP design is motivated by the observation that off-chip prediction can be leveraged to design effective L1 D prefetch filters. SLP can be used to improve the performance of any L1 D generic prefetcher since it makes no assumption regarding the L1 D prefetcher design. SLP uses one or more program features to perform effective prefetch filtering at L1 D.

[0060] SLP also uses program features, but these features are adapted to use physical addresses in place of virtual addresses as SLP is placed after the L1 D cache. Additionally, SLP may use a new feature denoted as FLP prediction + offset. This feature combines the FLP output bit of the cache block from which the prefetch request originated with the offset of the prefetched cache block in its physical memory page. The rationale of this feature is to correlate the probability of an L1 D prefetch request going off-chip when a certain cache line offset is touched with the off-chip prediction concerning the block that triggered the prefetch request. The SLP produces a binary off-chip prediction when an L1 D prefetch request is issued.

[0061] Figure 2 shows a flowchart of the SLP operation. SLP is consulted when the L1 D prefetcher issues a prefetch request. The confidence value is built similarly to the FLP. The output value is compared to the Tpref threshold. If it exceeds Tpref, the prefetch is considered as eventually requiring DRAM access and, therefore, being useless with high probability. In this situation, the prefetch request is discarded. Conversely, if the confidence value does not exceed Tpref, the prefetch request is processed as usual by the cache hierarchy.

[0062] The SLP is trained analogously as the FLP. Upon the completion of an L1 D prefetch request, the predictor’s weights are trained positively or negatively depending on whether the prefetch request was served off-chip.

[0063] Two Level Perceptron (TLP) Predictor

[0064] This section describes the Two-Level Perceptron (TLP) predictor, a hardware design using a multi-level perceptron approach. Figure 3 shows the design and the operation of TLP when used to combine off-chip prediction and prefetch filtering. In this example, TLP uses FLP and SLP as its fundamental building blocks.

[0065] Upon a load demand access, the computing core consults FLP to obtain a confidence value (referred to as ‘Conf’) driving the off-chip prediction (1). This prediction can give one of the three following outcomes: i) the load request is predicted to be off-chip with high confidence (Conf > Thigh) , thus a speculative DRAM request is thrown from the computing core (2) besides the regular load demand access; ii) the load request is predicted to be off-chip with low confidence (rtowConf < Thigh , thus the speculative DRAM request will be thrown only if the load misses in the L1 D (3); iii) the load request is predicted to be on-chip; therefore no additional action is taken (4) besides triggering the regular demand access. Metadata relative to the prediction (hashed PC, history of last load PC, and perceptron confidence value) are stored in the matching Load Queue entry for later training and an off-chip prediction tag is set in the load request thrown to the cache hierarchy depending on the FLP prediction.

[0066] SLP is consulted upon L1 D prefetch requests (5). To make a prediction, SLP takes as input the metadata attached to the prefetch request and the off-chip prediction tag attached to the demand load request from which the prefetch request originates. This information is used to produce an off-chip Conf prediction specific to L1 D prefetch request. This prediction can result in two possible outcomes: i) the prefetch request is predicted to be off-chip (Conf TPref) and the prefetch request is discarded (6), and ii) the prefetch request is predicted to be on-chip (Conf > Tpref) and the prefetch request is processed as usual by the cache hierarchy (7). Analogously to FLP, SLP stores metadata relative to its prediction in the L1 D miss status / handler register (MSHR) entries for later training.

[0067] The training routines of the FLP and the SLP are triggered upon completion of the corresponding requests, i.e. , for FLP when the load request returns to the computing core and SLP when the prefetch request is served.

[0068] It can be shown that the known alternatives of the prior art, namely: i) a conventional FLP predictor (i.e., without a selective delay mechanism); ii) an SLP predictor alone; iii) a two- Step Predictor (TSP) consisting of an FLP without the selective delay mechanism, in combination with an SLP but without being based on FLP output, obtain, respectively, 2.9%, 6.9%, 8.4% geometric mean speedups over a baseline system without off-chip prediction. Compared to these known alternatives, the claimed proposal provides a 11.5% speedup over the baseline, thus improving the prior art results in at least a 37% yield ratio over the TSP, and a 396% over the conventional FLP without selective delay mechanism.

Claims

CLAIMS1.- A first-level perceptron, FLP, off-chip predictor communicatively connectable to a computing core and to a DRAM, wherein the computing core and the DRAM are communicatively connected through a multi-level cache hierarchy comprising first-level data cache, L1 D, second-level cache, L2C, ... , and last-level cache, LLC, and characterized in that the FLP off-chip predictor is adapted with an FLP off-chip prediction mechanism comprising two thresholds, r / owand Thigh, such that, when the FLP off- chip predictor is connected to the computing core and the DRAM, and the computing core receives a demand load request, the FLP off-chip prediction mechanism is configured to perform the following steps:- the FLP off-chip predictor is consulted by the computing core;- the FLP off-chip predictor produces a confidence value used to drive the FLP off- chip prediction mechanism;- the confidence value is compared with Thigh',- if the confidence value greater than Thigh, the FLP off-chip predictor issues a speculative DRAM request from the computing core;- if the confidence value does not exceed Thigh, but does exceed rtow, the demand load request is tagged as predicted off-chip, and is sent to a L1 D cache;- if the predicted off-chip request results in a miss in the L1 D, the tag is read, and the speculative DRAM request is issued from the L1 D cache;- if the confidence value does not exceed rtow, the FLP off-chip predictor does not issue the speculative DRAM request from the computing core.2.- FLP off-chip predictor according to the preceding claim, wherein Thigh indicates a probability threshold for the corresponding demand load request to miss in all cache levels, and Tiow indicates a probability threshold for the corresponding demand load request not to miss in all or in any of the cache levels.3.- FLP off-chip predictor according to any of the preceding claims, wherein the FLP off-chip prediction mechanism is further configured to correlate a probability of a demand load request going off-chip with a history of program counters, PC, and accessed memory blocks, based on a first set of legacy or leveling features associated with a first weight table which is composed of confidence counters.4.- FLP off-chip predictor according to the preceding claim, wherein the first set of legacy features comprise at least one feature selected from: PC and cacheline offset, PC and byte offset, PC and first access, cacheline offset and first access, last-4 load PC.5.- FLP off-chip predictor according to any of claims 3-4, wherein the FLP off-chip prediction mechanism is further configured with a first training algorithm to be trained upon completing a memory access, and the accessed memory block is returned to the computing core from the multi-level cache hierarchy.6.- FLP off-chip predictor according to the preceding claim, wherein the first training algorithm comprises the following steps:- when the demand load request comes back to the computing core, the FLP off- chip predictor checks if the demand load request was a true off-chip demand load request, and the request requires a DRAM access;- if the demand load request is a true off-chip demand load request, the FLP off-chip predictor’s corresponding confidence counters of the first weight table are trained positively;- if the demand load request was not a true off-chip demand load request, the FLP off-chip predictor’s corresponding confidence counters of the first weight table are trained negatively.7.- A two-level perceptron, TLP, off-chip predictor, communicatively connectable to a computing core, to a DRAM, wherein the computing core and the DRAM are communicatively connected through a multi-level cache hierarchy comprising first-level data cache, L1 D, second-level cache, L2C, ... , and last-level cache, LLC, and wherein the TLP off-chip predictor comprises:- an FLP off-chip predictor according to any of the preceding claims; and- a second-level perceptron, SLP, off-chip predictor communicatively connectable to the multi-level cache hierarchy through a L1 D prefetcher, and wherein the SLP off-chip predictor is adapted with an SLP off-chip prediction mechanism comprising a prefetching threshold, Tpref, such that, when the SLP off-chip predictor is connected to the multi-level cache hierarchy, and the L1 D prefetcher issues a prefetch request, the SLP off-chip prediction mechanism is configured to perform the following steps: a) the SLP off-chip predictor produces an output value used to drive the SLP off- chip prediction mechanism; b) the output value is compared with Tpref,c) if the output value exceeds Tpref, the prefetch request is discarded; d) if the output value does not exceed Tpref, the prefetch request is processed by the multi-level cache hierarchy.8.- TLP off-chip predictor according to the preceding claim, wherein the SLP off-chip prediction mechanism is further configured to correlate a probability of a demand load request going off-chip with a history of PC and accessed memory blocks, based on a second set of legacy or leveling features associated with a second weight table which is composed of confidence counters.9.- TLP off-chip predictor according to the preceding claim, wherein the second set of legacy features comprise at least one feature selected from: PC and cacheline offset, PC and byte offset, PC and first access, cacheline offset and first access, last-4 load PC.10.- TLP off-chip predictor according to any of claims 8-9, wherein the second set of leveling features comprise at least FLP prediction and cacheline offset.11.- TLP off-chip predictor according to any of claims 7-10, wherein the SLP off-chip prediction mechanism is further configured with a second training algorithm to be trained upon completing an L1 D prefetch request, and the L1 D prefetch request is served.12.- TLP off-chip predictor according to the preceding claim, wherein the second training algorithm of the SLP off-chip prediction mechanism comprises the following steps:- when the prefetch request comes back to the computing core, the SLP off-chip predictor checks if the prefetch request was a true off-chip prefetch request, implying that the prefetch request requires a DRAM access;- if the prefetch request is a true off-chip prefetch request, the SLP off-chip predictor’s corresponding confidence counters of the second weight table are trained positively;- if the prefetch request is not a true off-chip prefetch request, the SLP off-chip predictor’s corresponding confidence counters of the second weight table are trained negatively.

Citation Information

Patent Citations

  • Adaptive mechanism to tune the degree of pre-fetches streams

    US20160019155A1

Cited By

  • Request target prediction method and device of memory access request

    CN121255666A