REDUCING A GROUP OF POLICIES WITH TWO OR MORE HYPERDIMENSIONS

By allocating CRPVs based on prefetcher engines and event types using machine learning agents, the system optimizes cache efficiency and processing performance by reducing the number of policies through iterative and greedy algorithms.

DE102025112709A1Pending Publication Date: 2025-10-02MELLANOX TECHNOLOGIES LTD(IL)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102025112709
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-01
Filing Date
2025-04-01
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing cache systems face inefficiencies due to the limited memory space and the need for effective cache replacement policies, which are not optimally managed by current methods, affecting processing efficiency.

Method used

A system and method are introduced to allocate Cache Replacement Priority Values (CRPVs) based on prefetcher engines and event types, utilizing machine learning agents to reduce the number of policies through iterative processes and greedy algorithms, combining multi-armed bandit and spatially distributed learning agents to optimize cache efficiency.

Benefits of technology

This approach enhances cache efficiency by dynamically selecting the best cache replacement policies, improving processing performance and reducing the number of policies required for effective cache management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

In one embodiment, a system includes a processor to reduce a number of cache replacement priority value policies available for selection by a machine learning agent, each cache replacement priority value policy including an array of predefined cache replacement priority values ​​for corresponding different combinations of (a) prefetcher engines that have loaded cache lines and (b) event types of events performed on the cache lines, and a memory to store data used by the processor.
Need to check novelty before this filing date? Find Prior Art

Description

AREA OF REVELATION

[0001] The present invention relates to computer systems and, in particular, but not exclusively, to the reduction of a set of policies. BACKGROUND

[0002] By definition, a cache has only a limited amount of storage space. Therefore, when a new entry is loaded into the cache, an old entry is generally evicted according to a cache replacement policy. Each cache entry or cache line can have a cache replacement precedence value (CRPV), which defines the importance of that cache entry for eviction. The cache entry with the worst (i.e., highest) CRPV is usually evicted from the cache first. There are rules for allocating and updating CRPVs based on what happens to a cache entry. SUMMARY

[0003] The invention is defined by the claims. To illustrate the invention, aspects and embodiments are described herein, which may or may not be within the scope of the claims.

[0004] According to one embodiment of the present disclosure, a system is provided including a processor to reduce a number of cache replacement priority value policies available for selection by a machine learning agent, each cache replacement priority value policy including an array of predefined cache replacement priority values ​​for corresponding different combinations of (a) prefetcher engines that have loaded cache lines and (b) event types of events performed on the cache lines, and a memory to store data used by the processor.

[0005] According to another embodiment of the present disclosure, a system is also provided that includes a processor for determining reduced sets of policies for respective hyperdimensions, comprising determining a first reduced set of policies for a first hyperdimension and determining a second reduced set of policies for a second hyperdimension, combining the reduced sets of policies in various combinations, resulting in a third set of policies that includes combining the first reduced set of policies for the first hyperdimension with the second reduced set of policies for the second hyperdimension, and reducing the third set of policies to obtain a fourth reduced set of policies available for selection by a machine learning agent, wherein each policy in the fourth reduced set includes an array of values,associated with each of the hyperdimensions, including an array of first values ​​associated with the first hyperdimension and an array of second values ​​associated with the second hyperdimension, and a memory for storing data used by the processor.

[0006] Furthermore, according to an embodiment of the present disclosure, the policies are cache replacement priority value policies, wherein the first hyperdimension is a first event type, the second hyperdimension is a second event type, and each of the cache replacement priority value policies includes an array of predefined cache replacement priority values ​​for corresponding different combinations of (a) prefetcher engines that have loaded cache lines and (b) event types of events performed on the cache lines.

[0007] Furthermore, according to an embodiment of the present disclosure, the processor is configured to determine an impact on the performance of a device resulting from changing target values ​​of the first hyperdimension and the second hyperdimension during the execution of benchmark applications, to define groups of the first hyperdimension and groups of the second hyperdimension in response to the impact on the performance of the device, and to determine preliminary groups of policies for the first hyperdimension and the second hyperdimension based on the defined groups such that an equal value is to be applied to each of the defined groups.

[0008] Additionally, according to an embodiment of the present disclosure, the processor is further configured to apply a greedy algorithm to reduce a size of the preliminary sets of policies to obtain the first reduced set of policies and the second reduced set of policies.

[0009] Furthermore, according to an embodiment of the present disclosure, the processor is configured to receive measurements of the performance of the device executing the benchmark applications, define the preliminary sets of policies for the first hyperdimension and the second hyperdimension based on the defined sets, define orderings of at least some of the policies of the preliminary sets of policies, and provide the first reduced set of policies and the second reduced set of policies based on the defined orderings.

[0010] Furthermore, according to an embodiment of the present disclosure, the processor is configured to reduce the third set of policies to obtain the fourth reduced set of policies by applying a greedy algorithm.

[0011] Furthermore, according to an embodiment of the present disclosure, the processor is configured to receive measurements of the performance of the device executing the benchmark applications, define the third set of policies, an ordering of at least some of the policies of the third set of policies, and provide the fourth reduced set of policies based on the defined ordering.

[0012] Additionally, according to an embodiment of the present disclosure, the processor is configured to determine the second reduced set of policies for the second hyperdimension based on the first reduced set of policies of the first hyperdimension.

[0013] Furthermore, according to an embodiment of the present disclosure, the processor is configured to determine the second reduced set of policies for the second hyperdimension based on a best policy per benchmark application of the first reduced set of policies of the first hyperdimension.

[0014] Furthermore, according to another embodiment of the present disclosure, the processor is configured to iteratively determine the first reduced set of policies and the second reduced set of policies by performing an iterative process in which (a) intermediate results from the process for determining the first reduced set of policies are used in the process for determining the second reduced set of policies, and (b) intermediate results from the process for determining the second reduced set of policies are used in the process for determining the first reduced set of policies.

[0015] According to yet another embodiment of the present disclosure, a method is also provided, comprising: determining reduced sets of policies for respective hyperdimensions, including determining a first reduced set of policies for a first hyperdimension and determining a second reduced set of policies for a second hyperdimension, combining the reduced sets of policies in various combinations, resulting in a third set of policies comprising combining the first reduced set of policies for the first hyperdimension with the second reduced set of policies for the second hyperdimension, and reducing the third set of policies to obtain a fourth reduced set of policies available for selection by a machine learning agent, wherein each policy in the fourth reduced set includes an array of values,associated with each of the hyperdimensions, including an array of first values ​​associated with the first hyperdimension and an array of second values ​​associated with the second hyperdimension.

[0016] Furthermore, according to an embodiment of the present disclosure, the policies are cache replacement priority value policies, wherein the first hyperdimension is a first event type, the second hyperdimension is a second event type, and each of the cache replacement priority value policies includes an array of predefined cache replacement priority values ​​for corresponding different combinations of (a) prefetcher engines that have loaded cache lines and (b) event types of events performed on the cache lines.

[0017] Additionally, according to an embodiment of the present disclosure, the method includes determining an impact on the performance of a device resulting from changing target values ​​of the first hyperdimension and the second hyperdimension during execution of benchmark applications, defining groups of the first hyperdimension and groups of the second hyperdimension in response to the impact on the performance of the device, and determining preliminary groups of policies for the first hyperdimension and the second hyperdimension based on the defined groups such that an equal value is to be applied to each of the defined groups.

[0018] Moreover, according to an embodiment of the present disclosure, the method further comprises applying a greedy algorithm to reduce a size of the preliminary sets of policies to obtain the first reduced set of policies and the second reduced set of policies.

[0019] Furthermore, according to an embodiment of the present disclosure, the method comprises receiving measurements of the performance of the device executing the benchmark applications, the preliminary sets of policies for the first hyperdimension and the second hyperdimension based on the defined sets, defining orders of at least some of the policies of the preliminary sets of policies, and providing the first reduced set of policies and the second reduced set of policies based on the defined orders.

[0020] Furthermore, according to an embodiment of the present disclosure, the method comprises reducing the third set of policies to obtain the fourth reduced set of policies by applying a greedy algorithm.

[0021] Additionally, according to an embodiment of the present disclosure, the method includes receiving measurements of the performance of the device executing the benchmark applications, defining the third set of policies, defining an ordering of at least some of the policies of the third set of policies, and providing the fourth reduced set of policies based on the defined ordering.

[0022] Furthermore, according to an embodiment of the present disclosure, determining comprises determining the second reduced set of policies for the second hyperdimension based on the first reduced set of policies of the first hyperdimension.

[0023] Furthermore, according to an embodiment of the present disclosure, determining comprises determining the second reduced set of policies for the second hyperdimension based on a best policy per benchmark application of the first reduced set of policies of the first hyperdimension.

[0024] Furthermore, according to an embodiment of the present disclosure, determining comprises iteratively determining the first reduced set of policies and the second reduced set of policies by performing an iterative process in which (a) intermediate results from the process for determining the first reduced set of policies are used in the process for determining the second reduced set of policies, and (b) intermediate results from the process for determining the second reduced set of policies are used in the process for determining the first reduced set of policies.

[0025] Any feature of one aspect or embodiment may be applied to other aspects or embodiments in any suitable combination. In particular, any feature of a method aspect or embodiment may be applied to a device aspect or embodiment, and vice versa. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The present invention will become more fully understood from the following detailed description taken in conjunction with the drawings in which: Fig. 1 is a block diagram view of an apparatus constructed and operative in accordance with an embodiment of the present disclosure; Fig. Figure 2 is a flowchart showing steps in a method for selecting configurations of prefetcher engines in exploration and exploitation phases; Fig. 3 is a flowchart showing steps in a method of operating a multi-armed bandit learning agent for use with the apparatus of Fig. 1 is; Fig. 4 is a flowchart showing steps in a method for reducing a number of configurations for use in the device of Fig. 1 is; Fig. Figure 5 is a view illustrating the calculation of a vector for use in the method of Fig. 4 represents; Fig. 6 and Fig. 7 tables are the statistical calculations used in the procedure of Fig. 4 are carried out; Fig. 8-10 views are those that define groups in the procedure of Fig. 4 represent; Fig. 11 and Fig. 12 views are illustrating exemplary configurations for use in the method of Fig. 2 represent; Fig. 13 is a flowchart showing steps in a further method for reducing a number of configurations for use in the device of Fig. 1 is; Fig. 14 a flow chart with sub-steps in the method of Fig. 13 to determine a best configuration; Fig. 15 is a view showing vector operations in the process of Fig. 13 for determining a second-best configuration; Fig. 16 a flow chart with sub-steps in the method of Fig. 13 to determine a second-best configuration; Fig. 17 is a view showing vector operations in the process of Fig. 13 to determine a third-best configuration; Fig. 18 a flow chart with sub-steps in the method of Fig. 13 to determine a third-best configuration; Fig. 19 a flow chart with sub-steps in the method of Fig. 13 to determine an Nth best configuration; Fig. 20 is a flowchart showing steps in a method for use in the device of Fig. 1 is; Fig. 21 is a flowchart showing steps in a method for training regression models for use in the apparatus of Fig. 1 is: Fig. 22 is a view of an example illustrating the use of trained regression models in the apparatus of Fig. 1 shows; Fig. 23 is a view of an example illustrating the training of regression models in the apparatus of Fig. 1 shows; Fig. 24 is a flowchart showing steps in a method for using trained regression models in the device of Fig. 1 is; Fig. 25 is a block diagram view of the device of Fig. 1, which represents a cache eviction control; Fig. 26 is a flowchart showing steps in a method for controlling cache displacement in the device of Fig. 25 is; Fig. 27 and Fig. 28 schematic views of parts of the device of Fig. 25 and exemplary data used therein to illustrate cache eviction control; Fig. 29 is a flowchart showing steps in a method for selecting one or more policies for use in the method for controlling cache eviction of Fig. 26 is: Fig. 30 is a schematic view of a cache of the device of Fig. 25, which shows an example use of the cache by a spatially distributed machine learning agent; Fig. 31 a flowchart with steps in a first optional sub-method of the method of Fig. 29 is; Fig. 32 is a flowchart with steps in a second optional sub-method of the method of Fig. 29 is; Fig. 33 is a flowchart showing steps in a method for reducing a number of cache replacement priority value policies for use in the apparatus of Fig. 25; and Fig. 34 a flowchart with steps in a sub-process of the process of Fig. 33 is. DESCRIPTION OF EMBODIMENTS OVERVIEW

[0027] As mentioned previously, each cache entry or cache line can have a cache replacement precedence value (CRPV), which defines the importance of that cache entry for eviction. Additionally, there are rules for assigning and updating CRPVs based on what happens to a cache entry. The choice of rules for assigning and updating CRPVs is important because the rules affect cache efficiency and, therefore, processing efficiency.

[0028] Therefore, embodiments of the present disclosure address at least some of the above-noted disadvantages by providing a system in which CRPVs are assigned according to the prefetcher engines that loaded the cache lines and the event types performed on the cache lines. The event types may be any suitable event type, such as loading data into the cache or using a cache line (e.g., via a central processing unit (CPU)).

[0029] The use of prefetcher engines as input for determining CRPVs is based on the goal of prefetcher engines to improve memory access performance by applying different access patterns to memory sequences. For example, one prefetcher can track large structures and one prefetcher can track small structures. Prefetching predicts which data or instructions will be needed in the near future and proactively fetches them from slower memory (e.g., RAM) to faster cache memory. This reduces the latency associated with retrieving data when needed. Prefetching can be performed either by hardware (special mechanisms in the processor) or by software (instructions generated by the compiler).

[0030] In some embodiments, the system may include a single policy that includes an array of values ​​for each combination of prefetcher engine and event type. For example, if there are 15 prefetcher engines and 2 event types, then the policy includes 30 values, for example, one value for prefetcher engine 1 and event type A, another value for prefetcher engine 1 and event type B, and so on. For example, when data is loaded into the cache by prefetcher engine 8 (i.e., event type A), the policy value for prefetcher engine 8 and event type A is used for the CRPV for the cache line associated with the loaded data. When the cache line is then used by the CPU (i.e., event type B), the policy value for prefetcher engine 8 and event type B is used for the CRPV for the used cache line.

[0031] In some embodiments, the system may include multiple different policies, with each policy providing an array of values ​​for each combination of prefetcher engine and event type. The system then selects the best policy to use, for example, for a current software program being executed by the CPU. In some embodiments, the best policy(ies) for determining CRPVs may be selected using appropriate machine learning techniques.

[0032] In some embodiments, a multi-armed bandit machine learning agent performs exploration to determine the best policy and performs exploitation using the best policy. The best policy may change over time due to ongoing exploration. The best policy may be defined by any suitable metric, such as maximizing instructions per cycle (IPC) or cache efficiency.

[0033] In some embodiments, a spatially distributed machine learning agent (e.g., a machine learning model similar to a species distribution model (SDM)) may be used to perform exploration with different policies in different parts of the cache simultaneously and perform exploitation with the best policy in another part of the cache concurrently with exploration. A sliding window may be used to select exploration data to use to determine the current best policy.

[0034] In some embodiments, a combination of multi-armed bandits and spatially distributed machine learning agents may be used. Several examples of such a combination are described below.

[0035] In one example, a spatially distributed machine learning agent performs exploration with different policies in different areas of the cache, and a multi-armed bandit machine learning agent performs the exploitation at a different time than the exploration with the currently best policy. The system alternates between exploration and exploitation.

[0036] In another example, a spatially distributed machine learning agent performs exploration with different policies in different areas of the cache and determines a preliminary best policy according to a first metric (e.g., cache efficiency). A multi-armed bandit machine learning agent uses the preliminary best policy determined by the spatially distributed machine learning agent to bias exploration toward the preliminary best policy to accelerate the discovery of the best policy according to a second metric (e.g., IPC). The multi-armed bandit machine learning agent also performs exploitation with the best policy determined according to the second metric (e.g., IPC).

[0037] In some embodiments, the spatially distributed machine learning agent may perform exploration to identify changes in the first metric (e.g., cache efficiency) above a target threshold and notify the multi-armed bandit machine learning agent to increase exploration (e.g., for all policies) for a specified period of time to more quickly converge toward a new best policy determined according to the second metric (e.g., IPC).

[0038] In some embodiments, the system may use a conversion function to convert the cache efficiency measured by the spatially distributed machine learning into IPC for use by the multi-armed bandit machine learning agent, which is used by the multi-armed bandit machine learning agent to add to an accumulated reward value used by the multi-armed bandit machine learning agent. The conversion function may be static, observation-based, or based on a trained predictive function. The predictive function may be trained to provide the conversion function using examples where both cache efficiency and IPC have been measured and may be policy-dependent.

[0039] As previously mentioned, the best policy may be selected from a number of cache replacement priority value policies. The number of policies available for selection can be a very large group. For example, if there are 15 prefetcher engines in the system and two event types, and each CRPV can be assigned one of four values ​​(e.g., 0-3), there could be 4 to the power of 2 to the power of 15 policies. Therefore, embodiments of the present disclosure address at least some of the above-noted disadvantages by providing a system and method for reducing the number of cache replacement priority value policies available for selection by a machine learning agent, as described in the disclosed embodiments.

[0040] As previously mentioned, each policy contains CRPVs for different combinations of prefetcher engines and event types. In general, each policy can be considered as containing values ​​for different combinations of dimensions (e.g., prefetcher engines) and hyperdimensions (e.g., event types). For example, each policy may contain values ​​for each dimension (e.g., each prefetcher engine) in a first hyperdimension (e.g., event type A) and values ​​for each dimension (e.g., each prefetcher engine) in a second hyperdimension (e.g., event type B), and so on. Therefore, in some embodiments, a system and method are provided to reduce the number of policies that contain values ​​in two or more hyperdimensions.In a first step, the number of policies in each hyperdimension is reduced, for example, by defining groups of dimensions in each hyperdimension such that all dimensions in each group must be assigned the same value. By restricting the assignment of values ​​by groups, the maximum number of possible policies is reduced. In some embodiments, in the first step, the results from the reduced group of one of the hyperdimensions are used to reduce the group of the other hyperdimension, as described in the disclosed embodiments. The results from each hyperdimension can be used in the other hyperdimension as an iterative process to further refine the calculation of the reduced groups.In a second step, the number of policies can then be further reduced, for example, by using a greedy algorithm (or another suitable algorithm) to determine a set of the best policies for each hyperdimension. In a third step, the reduced sets of policies can then be combined in various combinations. For example, if there are 12 policies for the first hyperdimension and 12 policies for the second hyperdimension, the combined set of policies can comprise 144 policies. In a fourth step, the combined set of policies (e.g., 144 policies) can then be reduced, e.g., using a greedy algorithm, to obtain a final reduced set of policies (e.g., 12 policies) for selection by the machine learning agent. SYSTEM DESCRIPTION

[0041] It will now Fig. 1, which is a block diagram view of an apparatus 10 constructed and operative in accordance with an embodiment of the present disclosure. The apparatus 10 includes a processor 12, prefetcher engines 14, a memory 16, and at least one cache 18. The prefetcher engines 14 are configured to predict the next memory access addresses of the memory 16 from which data is to be loaded into the cache 18 during execution of a software application 20 by the processor 12. The prefetcher engines 14 may operate in different ways to predict the next memory access addresses, even if all of the prefetcher engines 14 are set to the same aggressiveness level. The prefetcher engines 14 are configured to load the data (e.g., cache lines) from the predicted next memory access addresses into the cache 18 during the execution of the software application 20.

[0042] Each of the prefetcher engines 14 is configured to selectively provide different aggressiveness levels (e.g., levels 0 to 3). One of the aggressiveness levels (e.g., level 0) of one of the prefetcher engines 14 may indicate that the prefetcher engine 14 is turned off, i.e., the prefetcher engine 14 is not predicting next memory access addresses and is not loading data into the cache. Different configurations of the prefetcher engines 14 enable different aggressiveness configurations of the prefetcher engines 14. The device 10 may be implemented with any number of prefetcher engines 14 having any number of aggressiveness levels. The example in Fig. Figure 1 assumes fourteen prefetcher engines 14, and each prefetcher engine 14 can be assigned any aggressiveness level selected from four aggressiveness levels. For illustration, the prefetcher engines 14 can have configuration A shown in inset 22, where prefetcher engine 1 has an aggressiveness level of 0, prefetcher engine 2 has an aggressiveness level of 1, prefetcher engine 3 has an aggressiveness level of 1, prefetcher engine 4 has an aggressiveness level of 3, and so on. Inset 22 also shows configuration Z, in which prefetcher engine 1 has an aggression level of 1, prefetcher engine 2 has an aggression level of 2, prefetcher engine 3 has an aggression level of 0, prefetcher engine 4 has an aggression level of 0, and so on.

[0043] If each prefetcher engine 14 can be assigned any aggressiveness level, regardless of the assigned aggressiveness level of the other prefetcher engines 14, then there would be 4 14 (i.e., 268,435,456) possible configurations of the prefetcher engines 14. The number of possible configurations is generally restricted (e.g., reduced) to enable more efficient machine learning, as described in more detail below. The number of prefetcher engine configurations may be reduced using any suitable method. Two different methods for reducing the number of prefetcher engine configurations are described in the disclosed embodiments with reference to Fig. 4-20 described.

[0044] In practice, some or all of the functions of the prefetcher engines 14 may be combined in a single physical component or, alternatively, implemented with multiple physical components. These physical components may be hard-wired or programmable devices, or a combination of both. In some embodiments, at least some of the functions of the prefetcher engines 14 may be implemented by a programmable processor under the control of suitable software. This software may be downloaded to a device in electronic form, for example, over a network. Additionally or alternatively, the software may be stored in physical, non-transitory, computer-readable storage media, such as optical, magnetic, or electronic memory.

[0045] The processor 12 is configured to execute the software application 20. The processor 12 is configured to select one of the prefetcher engine configurations to control the prefetcher engines 14 in each processing interval with the aid of a machine learning agent 24, which is described in more detail below. The processing intervals are described below with reference to Fig. 2 described in more detail.

[0046] The processor 12 is configured to control the prefetcher engines 14 according to the configurations of the prefetcher engines 14 selected by the machine learning agent 24 during exploration phases and exploitation phases during execution of the software application 20. The machine learning agent 24 can determine whether a next phase should be an exploration phase or an exploitation phase according to any suitable scheme, such as an epsilon-greedy scheme. In some embodiments, the machine learning agent 24 is a multi-armed bandit machine learning agent.

[0047] The processor 12 is configured to execute the machine learning agent 24 to select from the prefetcher engine configurations (e.g., a curated prefetcher engine configuration set) to control the prefetcher engines 14 during the exploration phases. The machine learning agent 24 may select the prefetcher configurations to use during the exploration phases according to any suitable scheme. For example, the machine learning agent 24 may randomly select the prefetcher configuration to use during an exploration phase.For example, the multi-armed bandit machine learning agent may select the prefetcher configuration to be used during an exploration phase based on different exploration probabilities of the different configurations, where the exploration probabilities may be based on prior knowledge gained from controlling benchmark applications according to the prefetcher configurations, as described in the disclosed embodiments with reference to FIG. Fig. 3 is described.

[0048] Processor 12 is configured to perform measurements on the system (e.g., device 10) during execution of machine learning agent 24. The measurements performed may include one or more of the following: instructions executed per cycle; memory transactions per cycle; power cost per memory transaction; average core frequency; average core power; and measured temperature (e.g., of the processor). In some embodiments, the measurements performed may include two or more of the following: instructions executed per cycle, memory transactions per cycle, power cost per memory transaction, average core frequency, average core power, and measured temperature (e.g., of the processor).

[0049] The processor 12 is configured to execute the machine learning agent 24 to select from the configurations (e.g., the cleaned configuration set) to maximize potential rewards by controlling the prefetcher engines 14 (and optionally by controlling other system parameters) in the exploitation phases based on the measurements taken. The calculation of values ​​based on the measurements taken and the maximization of potential rewards are described with reference to Fig. 2 described in more detail.

[0050] In some embodiments, the processor 12 is configured to execute the machine learning agent 24 to select from the configurations to determine potential rewards (which may be based on performance values ​​determined with reference to Fig. 2) by controlling the prefetcher engines 14 in the exploitation phases and minimizing potential reward losses by controlling the prefetcher engines 14 in the exploration phases, based on the measurements performed. Minimizing potential reward losses by controlling the prefetcher engines 14 in the exploration phases can be achieved by selecting the prefetcher engine configurations according to an exploration probability, as described in the disclosed embodiments with reference to Fig. 3 described.

[0051] It will now Fig. 2, which is a flowchart 200 illustrating steps in a method for selecting configurations of prefetcher engines 14 in exploration and exploitation phases. Following a previous processing interval, which may have been an exploration or an exploitation phase, the machine learning agent 24 is configured to determine whether a next processing interval should be an exploration or an exploitation phase (decision block 202). The machine learning agent 24 may decide whether the next processing interval should be an exploration or an exploitation phase using any suitable scheme, for example, but not limited to, an epsilon-greedy scheme.In some embodiments, the machine learning agent may perform 24 exploration and exploitation phases without explicitly determining whether the next phase should be exploration or exploitation, for example, a UCB algorithm or Thompson's Multi-Armed Bandit algorithm.

[0052] If the machine learning agent 24 determines that the next processing interval should be an exploration phase, the machine learning agent 24 is configured to select a prefetcher configuration from a prefetcher configuration set to control the prefetcher engines 14 in the exploration phase of the next processing interval (block 204). The machine learning agent 24 may select the prefetcher configuration to control the prefetcher engines 14 in the exploration phase using any suitable scheme, for example, randomly.In some embodiments, the machine learning agent 24 is configured to select a prefetcher configuration from the prefetcher configuration set to control the prefetcher engines 14 in the exploration phase of the next processing interval to minimize potential losses, for example, based on probabilities associated with each of the prefetcher configurations, as described with respect to FIG. Fig. 3 described in more detail.

[0053] If the machine learning agent 24 determines that the next processing interval should be an exploitation phase, the machine learning agent 24 is configured to select a prefetcher configuration from the prefetcher configuration set to control the prefetcher engines 14 in the exploitation phase of the next processing interval to maximize potential rewards by controlling the prefetcher engines 14 (and optionally by controlling other system parameters) in the exploitation phase (block 206). The machine learning agent 24 is configured to maximize potential rewards by selecting the prefetcher configuration that provides maximum potential (e.g., estimated) rewards compared to potential rewards of other prefetcher configurations.The potential rewards of each prefetcher configuration can be calculated based on past values ​​of system performance from the operation of the prefetcher engines 14 according to that prefetcher configuration in previous exploration and optional exploitation phases, as described in more detail below.

[0054] The processor 12 is configured to control the prefetcher engines 14 according to the prefetcher configuration selected either in the step of block 204 or in the step of block 206 (block 208) in the next processing interval. The processor 12 is configured to perform measurements on the system (e.g., at the end of the next processing interval, which becomes the previous processing interval) (block 210). The measurements taken may include any one or more of the following: instructions executed per cycle, memory transactions per cycle, power cost per memory transaction, average core frequency, average core power, and measured temperature (e.g., of the processor). The processor 12 is configured to calculate a value based on the measurements taken (block 212) and other factors such as the power budget, which is described in more detail below.Therefore, the value provides an indication of the system performance achieved by controlling the prefetcher engines 14 according to the selected prefetcher configuration during the just-completed previous processing interval. The value can be calculated according to any suitable function based on one or more of the measurements taken.

[0055] In some embodiments, the value may be calculated as follows: where: PwrratioMPC=PCoeff∗MPC∗CorefCorePower and Core power the average core performance (e.g. per application) is Core f is the average core frequency, P coeffwhere μ is the power cost in joules for a single memory transaction, IPC is instructions per cycle, and MPC is memory transactions per cycle. The above value calculation assumes that significant power savings are achieved by reducing memory transactions and that the saved power can be transferred to accelerate the CPU clock speed, thus achieving higher performance. The above value provides an example of how to maximize IPC and minimize MPC.

[0056] Processor 12 is configured to calculate a potential reward for selecting a particular configuration (e.g., the configuration used in the previous processing interval) of prefetcher engines 14 (block 214). In some embodiments, processor 12 is configured to calculate the potential reward for selecting the particular configuration (e.g., the configuration used in the previous processing interval) of prefetcher engines 14 based on past performance values ​​achieved by previously controlling prefetcher engines 14 according to the particular configuration.For example, if the configuration used in the previous processing interval is configuration A, then in the step of block 212, processor 12 calculates the value for configuration A, and then processor 12 calculates the potential reward for selecting configuration A in the future based on the calculated value for configuration A and the previously calculated values ​​for configuration A calculated in other previous processing intervals. In some embodiments, processor 12 is configured to calculate the potential reward for selecting the particular configuration of prefetcher engines 14 based on an average of past performance values ​​by previously controlling the prefetcher engines (in previous processing intervals) according to the particular configuration.In some embodiments, processor 12 is configured to calculate the potential reward for selecting the particular configuration of prefetcher engines based on a weakened average of past performance values, as described in more detail below.

[0057] The processor 12 can execute multiple software applications serially or in parallel. When different applications are executed serially, the machine learning agent 24 ideally detects a new application and adapts accordingly. However, the application boundary is typically unknown to the processor 12. In such a case, the environment is said to be non-stationary, in which the best configuration changes over time. A naive solution is to periodically reset the states of the machine learning agent 24, for example, to delete old performance values. While this is a good solution, it assumes that all applications have the same duration.In some embodiments, the machine learning agent 24 is therefore configured to calculate the potential reward for selecting a particular configuration based on past performance values ​​using a time-based exponential moving average over all past values. The time-based exponential moving average attenuates the historical values ​​while giving greater weight to new values. A state of the machine learning agent 24 (e.g., the potential reward for selecting a particular configuration) may be updated (with attenuation) when that state is assigned a value. Therefore, an attenuation factor (alpha) is a function of the time interval between the current value and the last value for a particular configuration, and the attenuated average potential reward for a particular configuration may be calculated as follows: [(1−alpha(dt))×previous attenuated average reward potential]+ [alpha(dt)×current value for a given configuration], where alpha(dt) is defined based on the expected duration of an application.

[0058] Any suitable method may be used to calculate the potential reward for selecting a particular configuration. In some embodiments, processor 12 is configured to calculate the potential reward for selecting a particular configuration of prefetcher engines 14 based on maximizing executed instructions per cycle. Additionally or alternatively, processor 12 is configured to calculate the potential reward for selecting the particular configuration of prefetcher engines 14 based on minimizing memory transactions per cycle.In general, the processor 12 may be configured to calculate a potential reward for selecting a particular configuration of the prefetcher engines 14 based on past performance values ​​by previously controlling the prefetcher engines 14 according to the particular configuration, where the past performance values ​​are based on one or more of the following: instructions executed per cycle; memory transactions per cycle; power cost per memory transaction; average core frequency; average core power; power budget; and measured temperature (e.g., of the processor). The power budget may be the amount of power that may be used based on other (e.g., higher-order) considerations.

[0059] In some embodiments, when multiple applications run in parallel on processor 12, an instance of machine learning agent 24 (e.g., a multi-armed Bandit context) is attached to each application, and a context switch is configured to switch a machine learning agent context (e.g., Bandit) when other context information is exchanged.

[0060] In some embodiments, different multi-armed bandit types or configurations can be optimized for different applications. A further layer of machine learning can perform additional exploration and use the best bandit type for a running application.

[0061] The size of the time window for each processing interval is carefully selected. If the time window is too small, the data provided to the machine learning agent 24 may be too noisy. If the time window is too large, it may take significantly longer for the machine learning agent 24 to adjust to the best configuration, reducing overall performance. In some embodiments, the size of the time window may be set to a value in the range of 10 microseconds to 10 milliseconds. The size of the time window may be application-dependent. An application may be divided into checkpoints, which represent large sequences of instructions that may be executed repeatedly during application execution. The size of the checkpoints (or the average size of the checkpoints) may be an indicator of an optimal time window for a particular application.In other words, the time window can be of the same order of magnitude as the average size of the checkpoints. Each checkpoint can be assigned a weight based on the number of repetitions of that checkpoint in the application. The checkpoint weights can also be used to calculate the average size of the checkpoints.

[0062] In practice, some or all of the processors 12 may be combined into a single physical component, or alternatively, implemented with multiple physical components. These physical components may be hard-wired or programmable devices, or a combination of both. In some embodiments, at least some of the functions of the processor 12 may be implemented by a programmable processor under the control of suitable software. This software may be downloaded to a device in electronic form, for example, over a network. Additionally or alternatively, the software may be stored in physical, non-transitory, computer-readable storage media, such as optical, magnetic, or electronic memory.

[0063] It will now Fig. 3, which is a flowchart 300 including steps in a method of operating a multi-armed bandit learning agent 24 for use with the apparatus 10 of Fig. 1. The machine learning agent 24 is, as previously described with reference to Fig. 1 and Fig. 2, to select configurations of the prefetcher engines 14. It should be noted that the Fig. 3 represents a general improvement to the Multi-Armed Bandit (MAB) and may be applied to any suitable MAB used to select policies for controlling a resource. In some embodiments, the policies may be configurations, e.g., the configurations of the prefetcher engines 14 or any suitable preprocessor engines. In some embodiments, the resource may include the prefetcher engines 14 or any suitable preprocessor engines.

[0064] As already mentioned with reference to Fig. As described in Figure 1, the machine learning agent 24 selects the configurations or policies to be applied in the exploration phases according to an appropriate scheme. Some schemes treat all policies as equally important by randomly selecting the policies to be applied in the exploration phases. For example, given an exploration probability of epsilon and there are m policies, the probability that a particular policy will be explored is epsilon divided by m and is evenly distributed across all policies.

[0065] In embodiments of the present disclosure, not all policies are treated equally; the machine learning agent 24 can be configured to bias exploration toward policies that, based on prior knowledge, are more advantageous than other policies and have a lower risk of harm. Although the benefits and losses of selecting different policies are unknown for a current scenario, prior knowledge based on a large number of examples (e.g., benchmark applications) provides data about the benefits and losses of selecting different policies, including the negative impacts in a worst-case scenario. The prior knowledge can be used to determine a probability for each of the policies to be explored.The machine learning agent 24 can then be configured to select policies for exploration based on the corresponding probabilities of the policies, as described in more detail below.

[0066] There is a probability trade-off. If the probability of a particular policy is high and the system is already in a steady state with another policy, then exploring that policy is disadvantageous. If a particular policy has a low probability and eventually becomes the policy in the steady state, the low probability of the policy means that it takes a long time to reach the steady state because the policy is only explored at a very low rate. The fact that the probability is low in this case is disadvantageous. Therefore, a policy with high probability can be disadvantageous in the steady state, and a low probability can be disadvantageous in the transition state (before the steady state is reached).

[0067] Embodiments of the present disclosure calculate exploration probabilities taking into account the disadvantages described above, such that policies that have a high probability of significantly degrading one or more system performance metrics (e.g., instructions per cycle (IPC)) should receive a lower probability of exploration - and this must be considered relative to the probability that these policies could significantly improve the performance metric(s) in question.

[0068] The following description with reference to Fig. 3 gives an overview of the calculation of exploration probabilities and use of the probabilities in the device 10 of Fig. 1 or any suitable device or system. Following this overview, a detailed description of two methods for calculating the probabilities follows.

[0069] In some embodiments, processor 12 is configured to control a resource (e.g., prefetcher engines 14) according to different policies (e.g., different configurations of the resource) while benchmark applications are executed (block 302). For example, processor 12 may be configured to control the resource according to Policy A for benchmark applications 1-10, and also control the resource according to Policy B for benchmark applications 1-10, and so on. Processor 12 is configured to measure system performance during execution of the benchmark applications while different policies are applied to the resource (block 304). The measurements provide prior knowledge that, for each policy, includes measurements of system performance for different ones of the benchmark applications or different parts (e.g., checkpoints) of the benchmark applications. For example, the prior knowledge mayIPC for Policy A for Benchmark Application 1, IPC for Policy A for Benchmark Application 2, etc.) and IPC for Policy B for Benchmark Application 1, IPC for Policy B for Benchmark Application 2, and so on for all policies and all benchmark applications.

[0070] The following description describes the calculation of exploration probabilities using the processing circuitry 26 of the device 10. In some embodiments, the processing circuitry 26 may be integrated into the processor 12. In other embodiments, the processing circuitry 26 may be an additional element in the device 10 (as in Fig. 1) or in another device. The processing circuit 26 is configured to calculate the probabilities with which the corresponding policies are explored based on the prior knowledge gained by controlling the resource (e.g., the prefetcher engines 14) according to the policies during the execution of benchmark applications (block 306). The processing circuit 26 is configured to calculate the probability with which any particular policy of the policies is explored based on the prior knowledge of the degree of success and failure of the particular policy during the execution of the benchmark applications. The processing circuit 26 is configured to calculate the probability with which the particular policy is explored such that the probability with which the particular policy is explored optimizes a loss function (e.g.,maximized or minimized) representing a system parameter to be optimized (block 308). The loss function may include a parameter that compares: (a) a first value of a quality metric when the particular policy is applied during execution of a particular benchmark application; with (b) a second value of the quality metric of the best policy, which is one of the policies that provides a highest value of the quality metric among the policies applied during execution of the particular benchmark application. See Equation 2 below for an example of a suitable loss function.The loss function may include any one or more of the following parameters: an empirical parameter that measures the average or distribution of the frequency with which certain policies need to be explored to identify the relative quality of the certain policies compared to other policies for the particular benchmark application; a weight of the particular benchmark application with respect to other benchmark applications; and a number of checkpoints (or a similar measure of the application length) in the particular benchmark application.

[0071] In some embodiments, processing circuitry 26 is configured to calculate the various probabilities with which the corresponding policies are explored such that the exploration probabilities optimize (e.g., maximize or minimize) a loss function representing a system parameter to be optimized. See Equation 10 below for an example of a suitable loss function representing various policies (and not treating each policy separately as described in Equation 2). The loss function of Equation 10 includes a first term representing the utility loss due to the time required by the Multi-Armed Bandit machine learning agent during exploration (during execution of the benchmark applications) to determine which policies are better than others.The loss function of Equation 10 includes a second term that represents the utility loss due to the exploration (during the execution of the benchmark applications) of suboptimal policies. Suboptimal policies are policies that are not the best policy for a given benchmark application.

[0072] The processing circuit 26 is configured to calculate the various probabilities with which the corresponding policies are explored based on an expression formed by comparing a derivative of the loss function to zero. See, for example, equations 3, 4, 11, and 12 below.

[0073] The following overview describes the use of the calculated exploration probabilities by the machine learning agent 24 or another suitable multi-armed bandit machine learning agent. The processor 12 is configured to control a resource according to the policies selected by the multi-armed bandit machine learning agent 24 during exploration phases and exploitation phases (block 310).

[0074] In some embodiments, the resource may include prefetcher engines 14 configured to predict the next memory access addresses of the memory 16. Each of the prefetcher engines 14 may be configured to provide a certain level of aggressiveness, as described above with reference to Fig. 1. The policies can correspond to different configurations of the prefetcher engines 14.

[0075] The processor 12 is configured to execute the multi-armed bandit machine learning agent 24 to select from the policies to explore the resource during the exploration phases (e.g., during execution of a software application such as the software application 20) according to the probabilities corresponding to one of the policies, where the probabilities include different probabilities (block 312). The processor 12 is configured to perform measurements on the system (e.g., on the device 10) during the exploration phases and, optionally, during the exploitation phases while the multi-armed bandit machine learning agent 24 is executed (block 314).The measurements taken may include any one or more of the following: instructions executed per cycle, memory transactions per cycle, power cost per memory transaction, average core frequency, average core power, and measured temperature (e.g., of processor 12). Processor 12 is configured to execute the multi-armed bandit machine learning agent 24 to select policies based on the measurements taken to maximize the potential rewards from controlling the resource during the exploitation phases (block 316).

[0076] The following detailed description describes an initial method for calculating the probabilities to be explored. Note that the following examples assume that the metric to be maximized is IPC. However, IPC can be replaced by any other suitable metric or metrics. For example, the natural logarithm (ln) of IPC can be maximized.

[0077] The following are definitions of the terms used in the first procedure.

[0078] P p is the probability of choosing policy p.

[0079] i is an index of the benchmark application i, which provides prior knowledge for the policy p.

[0080] W i is the weight of application i relative to the other benchmark applications. This indicates how important the benchmark application i is for optimization and can be specified by the system administrator.

[0081] N i is the total number of control points (not just the different number of control points) in the application i.

[0082] IPC p i is the composite IPC of policy p in application i based on prior knowledge. The composite IPC can include any suitable factor that influences the IPC. For example, the composite IPC can include the effect of the number of memory transactions. If there are too many memory transactions, power consumption increases, causing the system to respond by reducing the clock frequency, resulting in a lower rate of instructions per second.

[0083] IPC B i is the composite IPC of the best policy (except for B=p) for application i. The best policy for application i is the policy that provides the highest IPC for application i.

[0084] Δ i= abs(ln(IPC p i ) - ln(IPC B i )) is Equation 1, which compares the IPC for a given policy p with the IPC for the best policy B for an application i. Δ i represents the loss caused by choosing policy p over the best policy B for application i. A suitable metric is the natural logarithm (ln) of the IPC, which provides accurate results.

[0085] i ∈ B represents the sum over all applications i, where for each i the IPC p i is greater than IPC B i , i.e., the sum over all applications i for which policy p is the best policy for this application i (i.e., if for application i the IPC with policy p is greater than the IPC of the other policies). It should be noted that, since IPC B i is defined as the composite IPC of the best policy (except for B = p) for application i, IPCp i larger than IPC B i can be if IPC B i is the second-best policy for application i.

[0086] i ∈ NB represents the sum over all applications i for which IPC p i is smaller than IPC B i , i.e. the sum of all applications i for which policy p is not the best policy for this application i.

[0087] Therefore, the benefit of including policy p in a previous set of policies where policy p was not previously used for exploration is as follows: Bp=∑i∈B(1−1Nipp)wiΔi−∑i∈NBppwiΔi

[0088] In the first term, the summation is performed for the different applications i for policy p. However, the summation is only performed for terms where policy p is the "best" policy for application i based on prior knowledge. Since the IPC depends on the application, policy p could be the best in none, in one or more applications, and not in others, and so on.

[0089] The "1" in the first term represents the proportion of time in which policy p, as the best policy, delivers a benefit. The "minus" portion in the first term represents the proportion of time in which policy p delivers no benefit as the best policy. This occurs during the transition phase, in which policy p is not used in the exploitation phase because the system needs time during the exploration phases to recognize that policy p is the best policy. Each benchmark application is divided into Ni checkpoints, and at the beginning of each checkpoint, the system can decide which policy, for example, to use for exploration. The value Ni in the first term takes into account how many checkpoints must be "wasted" on average to recognize that policy p is the best.

[0090] In the second term, the summation is performed for the applications i where the current policy is worse than the best policy. Therefore, if policy p is not the best policy, selecting policy p for exploration results in a loss in the steady state. The fraction of time lost by exploring policy p is the probability P. p , with which policy p is explored, and the amount lost is given by Δ i If policy p is the best policy, then there is no loss from exploring this policy, and therefore it is not included in the second term in these cases.

[0091] The optimal probability, P p , for the choice of policy p for exploration can be determined by calculating the derivative of equation 2.

[0092] The derivative of equation 2 is given by: ∑i∈B1Nipp2wiΔi−∑i∈NBwiΔi=0

[0093] The optimal probability, P p , for the choice of policy p for exploration based on prior knowledge is therefore given by: ∑i∈BwiΔiNi∑i∈NBwiΔi

[0094] Therefore, the probability of choosing policy p is based on the prior knowledge of the success and failure of policy p in controlling the resource (e.g., the prefetcher engines 14) while the benchmark applications i are executed.

[0095] Based on the above, the prior knowledge acquired during the execution of the benchmark applications i is used to calculate the probability of each policy (using Equation 4). During the execution of another application Q (not one of the benchmark applications), the machine learning agent selects 24 policies for exploration based on the calculated probabilities. The knowledge acquired during the execution of the current application Q does not affect the calculated probabilities. The knowledge acquired during the execution of the current application Q affects which policy is selected during the exploitation phase based on the multi-armed bandit algorithm, which searches for the best policies.

[0096] The decision of the machine learning agent 24 whether to explore or exploit can be based on an epsilon, which can be defined as the sum of the calculated probabilities as follows: ε=∑pj

[0097] In other words, the probabilities define epsilon.

[0098] Another approach predefines epsilon and normalizes the probabilities under the predefined epsilon: pp˜=pp∑pjε

[0099] In some embodiments, the processor 12 or the processing circuit 26 may be configured to perform an offline process that selects another policy from a selection of policies to be added to a current set of policies. The process includes calculating p p or pp˜ for each policy candidate to be added to the set of policies using Equation 4, and calculating B p for each candidate (taking into account the performance of the policies for the benchmark applications i) using Equation 2. The process then selects the candidate with the best B p and adds the best candidate to the group.

[0100] Due to noise, the best policy will not be identified the first time that policy is explored. A given policy for a given application may not behave deterministically every time it is applied and may not behave as the best policy, even if it is the best on average. Therefore, a new factor m can be introduced to account for this factor. "m" is a heuristic parameter that specifies the average frequency with which a policy must be explored to be identified as the best.

[0101] The introduction of the parameter “m” results in: Bp=∑i∈B(1−mNipp)wiΔi−∑i∈NBppwiΔi pp=m∑i∈BwiΔiNi∑i∈NBwiΔi ε=m∑p∑i∈BwiΔiNi∑i∈NBwiΔi

[0102] The following detailed description describes a second method for calculating exploration probabilities. Definitions of the terms used in the second method follow. w i is the weight of application i relative to the other applications. N i is the total number of checkpoints in application i. N is the number of policies. d is the default policy (i.e., the policy initially chosen for exploitation). This can be based on the policy with the highest average IPC or chosen arbitrarily. IPC p i is the (composite) IPC of policy p for application i. P p is the probability of choosing policy p. p i (k) is the k-th best policy for application i. For example, p i (1) the best policy for application i, pi (2) is the second best and p i (N) is the worst. n i p is the rank of policy p in application i. For example, n i p =1 for the best policy, n i p =2 for the second best and n i p =N for the worst. m i (n) is an empirical parameter. m i (n) is the (average) frequency with which one of the n best policies of application i must be executed to determine that one of the n best policies is the best policy for application i. For example, m i (1) the (average) frequency with which the best policy of application i must be explored in order to identify the best policy as the best one for application i. m i(2) is the (average) frequency with which the best or second-best policy for application i must be explored to identify the best or second-best policy as the best policy. This is also prior knowledge. m i (1) can be determined by running the machine learning agent 24 with application i to see how many times the best policy needs to be explored to identify the best policy as the best policy. m i (2) can be determined by running the machine learning agent 24 with application i to see how many times the best policy or the second-best policy needs to be explored to identify the best policy or the second-best policy as the best policy, respectively. This process can be repeated, and an average value can be determined. A similar analysis can be performed for other values ​​of m. i (n) are carried out. Δ i (n) = ln(IPC pi(n) i) - ln(IPCpi(n+1) i ) and returns the advantage of the policy with rank n over the policy with the next best rank. Δ^ip=ln(IPCpi(1)i)−ln(IPCpi) represents an undesirable effect of exploring an undesirable policy compared to the best policy. IPC pi(1) is the composite IPC with the best policy for the application i.

[0103] The loss function, which represents the combined effects of all policies, is given by: (L=Potential−∑i∑n=1nid−1wimi(n)Δi(n)Ni∑k=1nPpi(k)−∑i∑pwiPpΔ^ip)

[0104] Equation 10 provides the effects of all policies, while Equation 2 considers the policies individually. Equation 10 is more precise. Roughly speaking, the three terms in Equation 10 correspond to the terms in Equation 2. For example, the "potential" corresponds to the "1" in the first term of Equation 2, the second term corresponds to the second half of the first term of Equation 2, and the third term corresponds to the second term of Equation 2.

[0105] The potential contained in Equation 10 is the contribution of all the best policies for the respective applications, assuming they are chosen correctly. It is defined as: Potential=∑iwiln(IPCipi(1))

[0106] It becomes clear below that the potential plays no role, since it vanishes (e.g., equals zero) when differentiating Equation 10, as shown in Equation 11. It is the weighted sum of the natural logarithms of the IPCs of the best policies for the respective applications.

[0107] The middle term of equation 10: ∑∑n=1nid−1wiNimi(n)Δi(n)∑k=1nPpi(k) describes the effect of the transition phase and represents the proportion of time in which better policies (or the best policy) do not provide any benefit - due to the transition phase in which a better policy (or the best policy) is not used in the exploitation phase because the system needs time during the exploration phase to recognize that the better policy(ies) is (are) better.

[0108] As mentioned above, n i p the rank of policy p in application i (for example, ni p =1 for the best policy, n i p =2 for the second best and n i p =N for the worst) and d is the default policy. Therefore, nid the rank of the default policy for application i.

[0109] If the default policy is the best policy, then nid equals 1. In this case (nid−1) is zero and summation does not occur for this application because there is no adverse effect of the transition phase for this application.

[0110] If the default policy is the second best policy, then nid equals 2. In this case (nid−1) equal to 1, since there is only one policy that is important in the transition phase, and this is the best policy.

[0111] If n=1 (i.e. the best policy for this application), then the middle term is equal to: wiNimi(1)Δi(1)∑k=11Ppi(k) and Δ i (1) = ln(IPC pi(1) i ) - ln(IPC pi(2) i ), which represents the advantage of the best policy over the next best policy.

[0112] Therefore, if k=1, the middle term is equal to: wiNimi(1)Δi(1)Ppi(1)

[0113] Where P pi(1) the probability of the best policy for application i.

[0114] Therefore, the above term indicates the effect in the transition phase when switching from the second-best policy to the best policy.

[0115] If the default policy is the third best policy, then nid equals 3. In this case (nid−1) equals 2, since there are two better policies than the default policy in the transition phase. Therefore, the mean term from n = 1 to n = 2 is summed for application i. The term for n = 1 was shown above.

[0116] If n=2, the middle term is equal to: wiNimi(2)Δi(2)∑k=12Ppi(k) where Δ i (2) = ln(IPC pi(2) i ) - ln(IPC pi(3) i ) and ∑k=12Ppi(k) P pi(k) equal to P pi(1) + P pi(2) where P pi(2) is the probability of the second-best policy for application i. Therefore, P pi(1) + P pi(2) the probability of choosing Policy 1 or Policy 2 for application i.

[0117] Therefore, wiNimi(2)Δi(2)∑k=12Ppi(k) a term for the effect in the transition phase when switching from the third-best policy to the second-best or best policy.

[0118] The last term of equation 10: ∑i∑pwiPpΔ^ip performs a summation for each application i and each policy p in each application i, where the policies lead to undesirable results because they are not the best policies for the applications in question. The last term sums, for a given application i, the product of: a weight w i the application, a probability P P that a policy is being explored for this application, and an undesirable effect of exploring an undesirable policy is equal to Δ^ip, which can be zero if it is the best policy and therefore does not contribute to the last term.

[0119] Differentiating Equation 10 and comparing the derivative as a function of P p with zero results in: ∑i∑n=nipnid−1wiNimi(n)Δi(n)(∑k=1nPpi(k))2−∑iwiΔ^ip=0

[0120] The above equation defines N equations with N variables. The solution to Equation 11 appears non-analytical and seems to point to a solver that determines the probability for each policy. Equation 11 can be "solved" by defining a standard policy and knowing for each application whether this policy is the best, second-best, and so on.

[0121] Alternatively, it can be assumed that the default policy is the worst choice and it can be assumed that the summation runs to N-1 instead of (nid−1), so that equation 11 is as follows: ∑i∑n=nipN−1wiNimi(n)Δi(n)(∑k=1nPpi(k))2−∑iwiΔ^ip=0

[0122] The following description with reference to Fig. 4-20 refers to prefetcher engines 14, aggressiveness levels, and prefetcher engine configurations. In some embodiments, the prefetcher engines 14 can be replaced with any suitable preprocessor engines, the aggressiveness levels can be replaced with any suitable settings, and the prefetcher engine configurations can be replaced with any suitable configurations or policies.

[0123] It will now Fig. 4, which is a flowchart 400 with steps in a method for reducing the number of configurations for use in the device 10 of Fig. 1. The processor 12 is configured to determine an impact on the performance of the device 10 resulting from changing settings (e.g., aggressiveness levels) of prefetcher engines 14 applied to benchmark applications executed by the device 10 (block 402). In some embodiments, the processor 12 is configured to determine the impact on performance based on the measured performance of the device 10. The step of block 402 may include one or more of the steps of blocks 404-416 described with reference to Fig. 5-7 below.

[0124] It will now Fig. 5, which is a view illustrating the calculation of a vector 50 for use in the method of Fig. 4. In addition, Fig. 4 referred to.

[0125] The processor 12 is configured to execute the benchmark applications and, in the process, change the settings of the prefetcher engines 14 (block 404). The processor 12 is configured to measure the performance of the device 10 during the execution of the benchmark applications (block 406). The steps of blocks 404 and 406 will now be described in more detail. The processor 12 executes the benchmark applications (e.g., sequentially, one after the other) with an initial configuration of the prefetcher engines 14 in which the aggressiveness levels of the prefetcher engines 14 are set to a particular level (e.g., the highest aggressiveness level, level 3). The processor 12 measures the performance of the device 10 during the execution of each benchmark application. The performance may be measured using one or more suitable measurements and may include calculating a performance value, as described above with reference to Fig. 2 described. Fig. 5 shows a vector 52 with performance measurements for 10 corresponding benchmark applications 1-10. For example, the value for benchmark application 1 is 78, the value for benchmark application 2 is 50, and so on. The aggressiveness level of one of the prefetcher engines 14 (e.g., prefetcher 1) is then changed (e.g., from level 3 to level 2), resulting in a changed configuration of the prefetcher engines 14. The processor 12 then executes the benchmark applications (e.g., in order, one after the other) with the changed configuration of the prefetcher engines 14. The processor 12 measures the performance of the device 10 during the execution of each of the benchmark applications. Fig. Figure 5 shows a vector 54 with performance measurements for the 10 corresponding benchmark applications based on the changed configuration. For example, the value for benchmark application 1 is 75, the value for benchmark application 2 is 50, and so on.

[0126] Processor 12 is configured to generate vector 50 from vector 52 and vector 54. For example, vector 50 is generated by an element-wise subtraction of vector 54 from vector 52, where each element is a performance measurement of one of the benchmark applications. Vector 50 describes the effect of changing the settings of one of the prefetcher engines 14 (e.g., Prefetcher 1) on the performance of device 10, as applied to the benchmark applications executed by device 10. Processor 12 then calculates an average 56 of the elements in vector 50 and a measure of the dispersion 58 (e.g., standard deviation or variance) of the elements in vector 50.

[0127] The above steps are performed such that the processor 12 is configured to generate the vectors 50 for the corresponding prefetcher engines 14 that describe the impact on the performance of the device 10 when the settings of the prefetcher engines 14 are changed for the benchmark applications executed by the device (block 408). For example, based on the initial configuration, the aggressiveness level of another of the prefetcher engines 14 (e.g., Prefetcher 2) is then changed (e.g., from level 3 to level 2), resulting in a second changed configuration of the prefetcher engines 14. The processor 12 then executes the benchmark applications (e.g., sequentially, one after the other) with the second changed configuration of the prefetcher engines 14.Processor 12 measures the performance of device 10 during the execution of each of the benchmark applications, resulting in a second vector 54, and then generates vector 50 for prefetcher 2 based on second vector 54 and vector 52, and so on for all prefetchers. The vectors 50 described above contain one value per benchmark application. In some embodiments, vector 50 may contain one value per part of an application, e.g., per checkpoint.

[0128] It will now Fig. 6 and Fig. 7, which are tables containing statistical calculations in the context of the procedure of Fig. 4. In addition, Fig. 4. Processor 12 is configured to calculate statistical measures based on vectors 50 for respective prefetcher engines 14 that describe the impact on the performance of device 10 when changing settings of prefetcher engines 14 applied to benchmark applications executed by device 10 (block 410). The step of block 410 includes the steps of blocks 412-416, which are described in more detail below.

[0129] The processor 12 is configured to calculate a dispersion measure 58 for each of the vectors 50 (block 412). For example, a dispersion measure is calculated for the elements of the vector 50 of prefetcher 1, and a dispersion measure is calculated for the elements of the vector 50 of prefetcher 2, and so on. The dispersion measure may be, for example, a variance or a standard deviation. The processor 12 is configured to calculate an average 56 of each of the vectors 50 (block 414). For example, an average of the elements of the vector 50 of prefetcher 1 is calculated, and an average of the elements of the vector 50 of prefetcher 2 is calculated, and so on. The average values ​​represent the average impact of the corresponding prefetcher engine 14 on the performance of the device 10. Fig. 6 shows a table 60 with the average values ​​56 and the dispersion measures 58 for the corresponding prefetcher engines 14.

[0130] Processor 12 is configured to calculate similarity values ​​between pairs of vectors 50 (block 416). For example, processor 12 may calculate a similarity measure between vector 50 of prefetcher 1 and vector 50 of prefetcher 2, another similarity measure between vector 50 of prefetcher 1 and vector 50 of prefetcher 3, yet another similarity measure between vector 50 of prefetcher 2 and vector 50 of prefetcher 3, and so on for all pairs of prefetcher engines 14. Each of the similarity measures may be, for example, a cosine similarity or correlation. Fig. Figure 7 shows a table 70 containing similarity measures (e.g., cosine similarity) for pairs of prefetcher engines 14. For example, the similarity measure between prefetcher engines 1 and 2 is equal to 0.6.

[0131] If one or more of the vectors have a high absolute average value, it may be necessary in some embodiments to change the initial configuration and repeat the above process of steps 404-416 based on the new initial configuration. A high average can be determined by comparing the average with the inherent noise of the measurement and the granularity of the configuration. A vector with a high degree of dispersion (and a low absolute average value) may be of interest for machine learning. If a pair of vectors 50 has a high degree of similarity, then the prefetcher engines 14 associated with that pair of vectors 50 move in a similar manner and are good candidates for the same group of prefetcher engines 14. The above observations can be used in selecting the prefetcher engines 14 to be grouped, as in Fig. 8-10 described in more detail.

[0132] The average values, dispersion measures and similarity measures calculated with reference to Fig. 4-12 may be any suitable averages, dispersion measures, and similarity measures. In some embodiments, the averages and / or dispersion measures and / or similarity measures may be weighted to obtain weighted averages, weighted dispersion measures, and / or weighted similarity measures that are weighted according to the different scenarios (e.g., different benchmark applications or parts thereof) in the vectors. The weights may be assigned according to the estimated importance of each scenario. The weights may be set by a system administrator or by any suitable method, for example, by the method described in Fig. 21 described procedures.

[0133] It will now Fig. 8-10, which are views that support the definition of groups 80 in the method of Fig. 4. In addition, Fig. 4 referred to.

[0134] The prefetcher engines 14 whose associated vectors 50 have the highest dispersion measures 58 (e.g., the highest standard deviations) are generally selected as central members of various ones of the groups 80, with one central member per group. Fig. Figure 8 shows that prefetcher engines 4, 5, 9, and 13 have the highest dispersion measures 58 and form the central members of groups A, B, C, and D, respectively. If a pair of vectors 50 with the highest dispersion measures 58 has a very high similarity measure, the prefetcher engines 14 of this pair of vectors 50 can be assigned to the same group, and another prefetcher engine 14 with a next higher dispersion measure 58 is assigned as the central member of another group 80.

[0135] The prefetcher engines 14 that are not yet assigned to groups are then assigned to groups 80 based on the highest similarity measures of the associated vectors (of the unassigned prefetcher engines 14) with the vector(s) of one or more existing members of the groups 80 (e.g., with the vectors of the central members of the respective groups 80). Fig. Figure 9 illustrates that prefetcher engines 3 and 10 are added to group A because the vectors 50 of prefetcher engines 3 and 10 are most similar to the vector 50 of prefetcher engine 4 (the central member of group A). Fig. Figure 10 shows the final distribution of the prefetcher engines 14 to the groups 80.

[0136] Therefore, the processor 12 is generally configured to define the groups 80 of the prefetcher engines 14 in response to the impact of changing the settings of the prefetcher engines on device performance (block 418). In some embodiments, the processor 12 is configured to define the groups 80 based on the calculated statistical measures. In some embodiments, the processor 12 is configured to define the groups 80 based on the calculated dispersion measure of at least some of the vectors 50 (e.g., the vector 50 with the highest dispersion measures) and the calculated similarity measures between at least some of the pairs of vectors 50.

[0137] In some embodiments, the processor 12 is configured to select the central members of the groups 80 such that a different prefetcher engine 14 is selected as the central member of each group 80 based on the dispersion measure of the corresponding vectors 50 (block 420). In some embodiments, the processor 12 is configured to select the central members based on the highest dispersion measures of the corresponding vectors 50. In some embodiments, the processor 12 is configured to select the central members based on the dispersion measure of the corresponding vectors 50, thereby minimizing the similarity measures between the corresponding vectors 50 of the central members. In other words, if the similarity measure between a pair of vectors is very high (e.g., close to 1), the prefetcher engines 14 may be assigned to the same group 80.

[0138] In some embodiments, the processor 12 is configured to, for each group 80, add other prefetcher engines 14 to that group based on the measures of similarity of vectors to the vector of the central member of that group (block 422).

[0139] In some embodiments, the addition of prefetcher engines 14 may be compared to more than one existing member of a group and / or even to a member that is not a central member of that group. Therefore, in some embodiments, the processor 12 is configured to add one or more prefetcher engines 14 to that group for each group 80 based on one or more measures of the similarity of one or more corresponding vectors (of the prefetcher engine(s) 14 being considered for addition to that group) to one or more vectors of one or more existing members of that group.

[0140] It will now Fig. 11 and Fig. 12, which are views illustrating exemplary configurations for use in the method of Fig. 2. In addition, Fig. 4 referred to. Fig. 11 shows a first exemplary configuration 82 in which all prefetcher engines 14 of group A are assigned aggressiveness level 2, all prefetcher engines 14 of group B are assigned aggressiveness level 1, all prefetcher engines 14 of group C are assigned aggressiveness level 0, and all prefetcher engines 14 of group D are assigned aggressiveness level 3. Fig. 12 shows a second exemplary configuration 84 in which all prefetcher engines 14 of group A are assigned aggressiveness level 1, all prefetcher engines 14 of group B are assigned aggressiveness level 3, all prefetcher engines 14 of group C are assigned aggressiveness level 1, and all prefetcher engines 14 of group D are assigned aggressiveness level 2.

[0141] The processor 12 is configured to provide different prefetcher engine configurations based on the settings to be applied to the prefetcher engines 14 (block 424), such that for each defined group 80, a corresponding setting is to be applied equally to the prefetcher engines 14 of that group 80, thereby reducing the number of prefetcher engine configurations available to the machine learning agent 24 (or another suitable machine learning agent) for selection. In general, if the number of prefetcher engines 14 is X, the number of settings (e.g., aggressiveness levels) is Y, and the number of defined groups 80 is Z, then the number of prefetcher engine configurations of Y X on Y Z reduced.

[0142] In use, the reduced configuration set may be used by the machine learning agent 24 to select from this reduced configuration set to control the prefetcher engines 14. In use, the reduced configuration set may be used by the machine learning agent 24 to select from this reduced configuration set to control the prefetcher engines 14. In some embodiments, the steps of block 204 and / or block 206 and block 208 of Fig. 2 as follows. The processor 12 may be configured to execute the machine learning agent 24 to select from various prefetcher engine configurations (e.g., from the reduced configuration set) to control the prefetcher engines 14, wherein the various prefetcher engine configurations are based on settings to be applied to the prefetcher engines 14, such that the groups 80 of prefetcher engines 14 are defined, and for each defined group 80, a corresponding setting (e.g., aggressiveness level) is to be applied equally to the prefetcher engines 14 of that group. The processor 12 may be configured to control the prefetcher engines 14 according to the various prefetcher engine configurations selected by the machine learning agent 24 during execution of the software application 20.

[0143] It will now Fig. 13, which is a flowchart 1300 including steps in another method for reducing a number of configurations for use in the device 10 of Fig. 1. The processor 12 is configured to receive data of a configuration set from prefetcher engines 14 (block 1302). The configuration set may be an already reduced configuration set, for example, using the method described above with reference to Fig. 4-12. In some embodiments, the method described in Fig. 13-20 described methods can be used to reduce an original configuration set.

[0144] Processor 12 is configured to execute benchmark applications, changing the aggressiveness levels of prefetcher engines 14 (block 1304), so that each benchmark application executes with each configuration. For example, if there are four groups 80 of prefetcher engines 14 with four different aggressiveness levels to be set for each group 80, then there are 256 possible configurations of prefetcher engines 14.

[0145] Processor 12 is configured to measure the performance of device 10 for each of the various configurations during execution of the benchmark applications (block 1306). For example, if there are 100 benchmark applications and 256 possible configurations, then there are 100 x 256 measurements of the performance of device 10, assuming that device performance can be measured using a metric (e.g., IPC). If device performance is measured using Q metrics, then there are 100 x 256 x Q measurements of device performance.

[0146] The processor 12 is configured to receive measurements of the performance of the device 10 that executed the benchmark applications while changing the configuration of the prefetcher engines 14 selected from the configuration set of prefetcher engines 14 (block 1308).

[0147] The processor 12 is configured to generate vectors, one vector for each configuration of the prefetcher engines 14, that indicate device performance per benchmark application (block 1310). Each vector contains elements, where each element indicates device performance when executing a corresponding benchmark application. For example, a vector for configuration A may contain one element indicating device performance when executing benchmark application 1 and another element indicating device performance when executing benchmark application 2, and so on. For example, if there are 100 benchmark applications, the vector may contain 100 elements describing device performance for the 100 corresponding benchmark applications. In some embodiments, performance may be measured for portions of the benchmark applications (e.g., checkpoints).Therefore, the vector may contain elements indicating device performance for corresponding parts of the benchmark applications (e.g., checkpoints). Similar vectors exist for configurations B and C, and so on. If there are 256 configurations, there are 256 corresponding vectors. Device performance may be expressed as IPC or ln(IPC), or any other suitable performance value, as discussed above with reference to . Fig. 2 described in more detail.

[0148] The processor 12 is configured to define an order for at least some of the configurations based on the measurements (e.g., based on the vectors) (block 1312). The processor 12 is configured to determine the best configuration of the configurations based on the measurements (block 1314) (e.g., based on the vectors), as described with reference to Fig. 14. The processor 12 is configured to determine a second-best configuration of the configurations based on the best configuration that has been found to be the best configuration and based on the measurements (block 1316) (e.g., based on the vectors), as described with reference to Fig. 15 and Fig. 16. The processor 12 is configured to determine a third-best configuration of the configurations based on the best configuration and the second-best configuration that have been found to be the best configuration and the second-best configuration, respectively, and based on the measurements (block 1318) (e.g., based on the vectors), as described with reference to Fig. 17 and Fig. 18. In general, the processor 12 is configured to determine an N-th best configuration of the configurations based on the best to N-1 best configurations of the configurations that were found to be the best to N-1 best configurations, respectively, and based on the measurements (block 1320) (e.g., based on the vectors), as described with reference to Fig. 19 is described in more detail.

[0149] Processor 12 is configured to provide a cleaned configuration set based on the defined order of at least some configurations (block 1322). For example, processor 12 may determine the tenth best configurations (or any number of best configurations) from a configuration set comprising, for example, 256 configurations, and the tenth best configurations form the cleaned configuration set.

[0150] It will now Fig. 14, which shows a flowchart 1400 with substeps in the method of Fig. 13 for determining the best configuration. As previously mentioned, each vector contains elements, each element indicative of device performance when executing a corresponding one of the benchmark applications (or a portion of the benchmark application). The processor 12 is configured to calculate averages (e.g., geometric averages across scenarios, e.g., benchmark applications or portions of the benchmark applications) of the elements of the vectors indicative of the performance of the device 10 (which executed the benchmark applications for corresponding configurations of the prefetcher engines 14) (block 1402). In other words, the processor 12 is configured to calculate an average of the elements of each of the vectors, thus averaging the device performance indications for each of the benchmark applications or portions of the benchmark applications in each vector.For example, if there are 256 vectors, processor 12 calculates 256 average values. Processor 12 is configured to: Fig. 15) finds the vectors with the highest average of the average values ​​calculated in the step of block 1402; and defines the best configuration as the configuration corresponding to the vector 86 (associated with the highest average) (block 1404).

[0151] It will now Fig. 15 and Fig. 16 referred to. Fig. 15 is a view showing vector operations in the method of Fig. 13 to determine a second-best configuration. Fig. 16 is a flowchart 1600 with substeps in the method of Fig. 13 to determine a second-best configuration.

[0152] The processor 12 is configured to select a candidate vector 88 from the vectors (block 1602) other than the vector 86 of the best configuration. Fig. For example, the vectors 86 and 88 shown in Figure 15 each contain 10 elements corresponding to the device performance when executing 10 corresponding benchmark applications (1-10). The vectors 86, 88 may contain any number of elements corresponding to any suitable number of benchmark applications or portions of the benchmark applications (e.g., checkpoints). The processor 12 is configured to perform element-wise max operations (arrow 90) on the candidate vector 88 and the vector 86, resulting in a resultant vector 92 (block 1604). The element-wise max operations select the "best" or highest element value for each of the elements from the vectors 86, 88. For example, element 1 corresponding to benchmark application 1 has an element value of 78 in vector 86 and a value of 60 in candidate vector 88, which is why the value 78 is used for element 1 in the resulting vector 92.Likewise, element 2 corresponding to benchmark application 2 has an element value of 50 in vector 86 and a value of 72 in candidate vector 88, therefore the value 72 is used for element 2 in the resulting vector 92, and so on. The processor 12 is configured to calculate an average (e.g., a geometric average) over all elements of the resulting vector 92 (block 1606), which yields an average value (reference numeral 94).

[0153] The processor 12 is configured to repeat steps of blocks 1602-1606 for other candidate vectors of the vectors, resulting in a plurality of resulting vectors 92 with a plurality of corresponding average values ​​94 (block 1608). For example, if there are 256 configurations in the configuration set, the steps are repeated for an additional 254 candidate vectors 88. The processor 12 is configured to determine the second-best configuration corresponding to the candidate vector 88 that results in the resulting vector 92 (hereinafter referred to as "the second-best configuration resulting vector 92") with the highest average value of the average values ​​94 (block 1610). In other words, the candidate vector 88 that results in the resulting vector 92 with the highest average value of the average values ​​94 is the second-best configuration.

[0154] It will now Fig. 17 and Fig. 18 referred to. Fig. 17 is a view showing vector operations in the method of Fig. 13 to determine a third-best configuration. Fig. 18 is a flowchart 1800 with substeps in the method of Fig. 13 to determine a third-best configuration.

[0155] The processor 12 is configured to select from the vectors a different candidate vector 88 (block 1802) as the vector 86 of the best configuration and a vector 96 of the second best configuration. Fig. For example, for simplicity, the vectors 86, 96, and 88 shown in Figure 17 each contain 10 elements corresponding to the device performance when executing 10 corresponding benchmark applications (1-10). The vectors 86, 96, 88 may contain any suitable number of elements corresponding to any suitable number of benchmark applications or portions of the benchmark applications (e.g., checkpoints). The processor 12 is configured to perform element-wise maximum operations (arrow 90) on the candidate vector 88 and the vectors 86, 96 (or on the candidate vector 88 and the "second-best configuration resultant vector 92," which is already an element maximum of the vectors 86, 96), yielding a resultant vector 98 (block 1804). The element-wise maximum operation selects the “best” or highest element value for each of the elements in the vectors 86, 96, 88.For example, element 1 corresponding to benchmark application 1 has an element value of 78 in vector 86, a value of 60 in vector 96, and a value of 80 in candidate vector 88, therefore the value 80 is used for element 1 in the resulting vector 98. Likewise, element 2 corresponding to benchmark application 2 has an element value of 50 in vector 86, a value of 72 in vector 96, and a value of 55 in candidate vector 88, therefore the value 72 is used for element 2 in the resulting vector 98, etc. The processor 12 is configured to calculate an average (e.g., a geometric or weighted average) over all elements of the resulting vector 98 (block 1806), resulting in an average value (reference numeral 99).

[0156] The above with reference to Fig. 13-20 calculated averages can be any suitable average, e.g., a geometric average or a weighted average, where different weights are assigned to the different scenarios (e.g., different benchmark applications or parts thereof) in the vectors. The weights are generally assigned according to an estimated importance of each scenario. The weights can be set by a system administrator or by any suitable method, for example, by the Fig. 21 described procedures.

[0157] The processor 12 is configured to repeat steps of blocks 1802-1806 for other candidate vectors of the vectors, resulting in multiple resulting vectors 98 with multiple corresponding average values ​​99 (block 1808). For example, if there are 256 configurations in the configuration set, then the steps are repeated for an additional 253 candidate vectors 88, resulting in 253 resulting vectors and 253 corresponding average values. The processor 12 is configured to find the third-best configuration corresponding to the candidate vector 88 that results in the resulting vector 98 with the highest average value of the average values ​​99 (block 1810). In other words, the candidate vector 88 that results in the resulting vector 98 with the highest average value of the average values ​​99 is for the third-best configuration.

[0158] It will now Fig. 19, which shows a flowchart 1900 with substeps in the method of Fig. 13 for determining an N-th best configuration. Based on the above description with reference to Fig. 15-18, which describes the determination of the second and third best configurations, describes Fig. 19 generally involves determining the Nth best configuration. The processor 12 is configured to select a candidate vector from the vectors (block 1902). The processor 12 is configured to perform element-wise maximum operations on the candidate vector and at least one vector (for example, on the vector(s) of (all) configurations better than the Nth configuration, or on a vector representing the element-wise maximum of the vectors of all configurations better than the Nth configuration) indicating the configurations from the best configuration to an (N-1)th best configuration, yielding a resulting vector (block 1904). The processor 12 is configured to calculate an average of the resulting vector (block 1906).Processor 12 is configured to repeat steps of blocks 1902-1906 for other candidate vectors of the vectors, resulting in multiple resulting vectors with multiple corresponding average values ​​(block 1908). Processor 12 is configured to determine the Nth best configuration corresponding to the candidate vector that results in the resulting vector with the highest average value of the multiple corresponding average values ​​(block 1910).

[0159] It will now Fig. 20, which is a flowchart 2000 with steps in a method for use in the device 10 of Fig. 1. The processor 12 is configured to reduce a collection of configurations of the prefetcher engines 14 to obtain a reduced configuration set of the prefetcher engines 14 by using a first configuration reduction method, for example, the method described with reference to Fig. 4-12. The processor 12 is configured to reduce the reduced configuration set of the prefetcher engines 14 to obtain a cleaned configuration set using a second configuration reduction method (different from the first configuration reduction method), for example, the one described with reference to Fig. Configuration reduction methods described in 13-19.

[0160] It will now Fig. 21, which shows a flowchart 2100 with steps in a method for training regression models for use in the device 10 of Fig. 1. As mentioned previously, supervised learning can be used to select prefetcher configurations (e.g., using multiple regression models for corresponding prefetcher engine configurations), as described in more detail below.

[0161] As an overview, machine learning training data is used to train configuration-specific machine learning regression models to produce configuration-specific predictions of device performance based on particular device hardware states of device 10. For example, a regression model for prefetcher configuration A is trained to provide a prediction of device performance (i.e., predict how device 10 will behave) when prefetcher engines 14 are configured according to configuration A based on a particular hardware state (e.g., represented by the counter values ​​of counters 28 ( Fig. 1) of device 10), a regression model is trained for prefetcher configuration B to provide a prediction of device performance (i.e., predict how device 10 will behave) when prefetcher engines 14 are set according to configuration B based on a particular hardware state (e.g., represented by counter values), and so on for all different configurations. In some embodiments, each regression module is also trained to provide a prediction based on the particular hardware state and the configuration used to control the prefetcher engines in the previous processing interval. The prediction of device performance may be an indication of instructions executed per cycle (IPC) or any suitable metric or combination of suitable metrics.Hardware health can be indicated by central processing unit (CPU) statistical counters, such as performance counters in ARM cores.

[0162] In use, before each processing interval, the hardware state of device 10 is retrieved, and the retrieved hardware state (and optionally the previous configuration used to control the prefetcher engines) is input to each of the trained regression models, resulting in a prediction for each trained regression model. For example, the retrieved hardware state is input to the trained regression model for prefetcher engine configuration A, which yields a prediction; the same retrieved hardware state is input to the trained regression model for prefetcher engine configuration B, which yields a prediction; and so on. The regression model that yields the best prediction (e.g., the highest or lowest, depending on whether the performance indicator is maximized or minimized, respectively) is selected.minimized) specifies the prefetcher engine configuration to be used to control the prefetcher engines in the next processing interval. This way, the hardware state (and optionally the configuration used to control the prefetcher engines) of the previous processing interval is used to find the prefetcher engine configuration to be used in the next processing interval.

[0163] The training of the regression models is now described in more detail below.

[0164] The processor 12 is configured to execute benchmark applications with different configurations of the prefetcher engines 14 (block 2102). For example, the processor 12 is configured to execute the benchmark applications with configuration A, then the same benchmark applications with configuration B, and so on. The processor 12 is configured to record the hardware states during the execution of the benchmark applications (e.g., recording the hardware state at each processing interval) (block 2104) and to perform measurements of device performance (e.g., at each processing interval) (block 2106). In some embodiments, the processor 12 is configured to execute each benchmark application (or a portion thereof, e.g., each checkpoint) with each prefetcher engine configuration (while the prefetcher engines 14 are executed according to configuration A, B, etc.).are set) and performs device performance measurements (e.g., IPC, memory transactions per cycle, power cost per memory transaction, average core frequency, average core power, and temperature) and collects hardware states (e.g., counter values ​​of counters 28) for each processing interval (while the prefetcher engines 14 are set according to configuration A, B, etc.).

[0165] The processor 12 is configured to calculate each of a plurality of label scores (block 2108) based on one or more of the following: executed instructions per cycle; memory transactions per cycle; power cost per memory transaction; average core frequency; average core power; power budget; and measured temperature. The label scores may be calculated according to the method described above with reference to Fig. 2 described value can be calculated.

[0166] The processor 12 is configured to receive (e.g., from memory or cache or from another device) machine learning training data, including label values ​​based on device performance measurements during execution of the benchmark applications for various prefetcher engine configurations and corresponding device hardware states (block 2110), where each value corresponds to a hardware state, as described with reference to Fig. 23. In some embodiments, the machine learning training data includes the label values, the corresponding device hardware states, and the corresponding previous prefetcher engine configurations, as described with reference to Fig. 23. In some embodiments, the hardware states are indicated at least in part by counter values, e.g., counter 28.

[0167] The regression models are trained to associate the detected hardware states with labels (e.g., values ​​or measurements of device performance). The processor 12 is configured to train the configuration-specific machine learning regression models based on the received machine learning training data to provide corresponding configuration-specific predictions of device performance based on device hardware states (block 2112) and optionally previous prefetcher engine configurations, as described in Fig. 23. For example, a regression model for prefetcher configuration A is trained to provide a prediction of device performance when the prefetcher engines 14 are set according to configuration A based on a particular hardware state of the device 10, a regression model for prefetcher configuration B is trained to provide a prediction of device performance when the prefetcher engines 14 are set according to configuration B based on a particular hardware state, and so on for all different configurations. For example, the regression model for prefetcher A can be trained based on a dataset containing pairs, each containing the hardware state and the label value (and optionally the configuration used in the previous processing interval) for corresponding different scenarios (e.g.different benchmark application or parts thereof) when the prefetcher engines 14 are set according to configuration A. The regression models may be implemented using any suitable model type, such as linear models, decision trees, or neural networks.

[0168] The processor is configured to train the configuration-specific machine learning regression models using corresponding configuration-specific cost (or loss) functions (based on optimization, e.g., minimizing or maximizing the cost functions) based on data indicative of various scenarios while executing the benchmark applications. In some embodiments, the data of the various scenarios are weighted in the cost functions, as described in more detail below.

[0169] For a regression model, a loss function can be defined as the squared error (or mean square error) between the prediction (i.e., the output of the regression model based on a given input, e.g., the hardware state of device 10) and the label (e.g., the value calculated from the measured device performance). Since the supervised machine learning model has limited resources, it may be more efficient to focus on scenarios that are more important. For example, accuracy may be less important for configurations (hereafter referred to as policies) that are far from the optimal policy for a given scenario. Therefore, assigning weights to the different scenarios can make the training of the supervised learning model more efficient.

[0170] An example loss function for a particular regression model might look like this: ∑iWi(Pi−Li)2, where Pi is the prediction for scenario i, Li is the label for scenario i, and Wi is the weight for scenario i. The loss function is summed over all scenarios i (e.g., checkpoints or applications) in the dataset.

[0171] Given a particular scenario (e.g., a benchmark application or a part of it, e.g., a checkpoint), there is a label for each of the policies when the scenario is executed if the prefetcher engines 14 are configured accordingly for each of the policies. The policy with the best reward (e.g., the highest value) for that scenario is called the best policy, and the one with the second-best reward is called the nearest challenger. For a policy that is not the best policy, that policy is more important if the prediction of the non-best policy is close to the prediction of the best policy, and less important if it is far from the best policy. If the prediction for the best policy is far from the nearest challenger, the prediction is generally unimportant because the best policy will be chosen anyway.If the prediction is close to the challenger, there is a possibility that the challenger will be chosen instead of the best policy due to an error.

[0172] The following factors are assumed when calculating the weights of the scenarios in the loss function. The weight decreases depending on the distance of the result from a reference value: (a) for the best policy, the distance is the distance to the nearest challenger, and (b) for any other policy, the distance is the distance to the best policy. In addition to the above heuristic, two measures are taken to protect the machine learning from anomalous behavior in edge cases. A worst-case scenario is assumed between the labeling and the actual prediction. The significance weight does not decrease upon inversion, i.e., when the best policy has a lower value than another policy.

[0173] The following terms are defined. Li k is the label of policy k for checkpoint i. P i k is the label of policy k for checkpoint i. B i = argmax k (L i k ) and is the best policy for checkpoint i. L i b is the identification of the best policy for checkpoint i. P i b is the prediction of the best policy for checkpoint i. CL i = argmax k≠b (L i k ) and is the challenger's policy for control point i. L i ci is the challenger's identification for checkpoint i. CP i = argmax k≠b (P i k ) and is the predicted challenger for checkpoint i. Note that CP i by Cli i may vary. P i cp is the prediction of the predicted challenger for checkpoint i. W i k is a significance weight of policy k for control point i, as defined below. W l b is a significance weight of the best policy for checkpoint i, as defined below. α is a hyperparameter and can have any value. A suggested initial value is α=0.01. Wib={α2α2+(GBi−GCi)2GBi>GCi1GBi≤GCi

[0174] The above formula shows that the weight is small when GB i and GC i are far apart. If GB i and GC i are close to each other, the weight is close to 1. Wik≠b={α2α2+(GBi−Gik)2GBi>Gik1GBi≤Gik

[0175] The above formula provides the weight of a non-best policy. If GB i near Gi k , the weight is close to 1. If GBi is far from G i k away, the weight is lower.

[0176] In some embodiments, the processor 12 is therefore configured to calculate a weight of a best configuration of the various prefetcher engine configurations for a particular scenario i (block 2114), such that the calculated weight is a function of a comparison between GB i and GC i and where GB i is a function of a label and / or a prediction of the best configuration for the particular scenario i, and GC i is a function of a label and / or a prediction of a second-best configuration for the particular scenario i. The weights of the best configurations for each of the different scenarios i are calculated in a similar manner.

[0177] In some embodiments, the processor 12 calculates a weight of a non-best configuration of the various prefetcher engine configurations for a particular scenario i (block 2116), such that the weight of the non-best configuration is a function of a comparison between GB i and G i k and where G i k is a function of a label and / or a prediction of the non-best configuration for the particular scenario. The weights of the non-best configurations for each of the different scenarios i are calculated in a similar manner. In some embodiments, the sum of the weights for a particular policy may be required to equal 1, and the weights may need to be normalized so that the sum of the weights for a particular policy equals 1.

[0178] It should be noted that the weights are a function of the prediction of the models, but the weights are also used to train the predictions of the regression models. Therefore, for each of the regression models, the processor 12 is configured to calculate the weights based on a default value of the prediction, create the cost function based on the calculated weights (block 2118), calculate predictions based on the scenarios and the cost function (block 2120), and perform an iteration (block 2122) of calculating the weights based on the new predictions, update the cost function with the updated weights, and calculate the predictions based on the scenarios and the updated cost function. The iterations can be performed until a steady state is reached. Suggested default values ​​are P i b = ∞ and P i k≠b =-∞.

[0179] Therefore, the processor 12 is configured to perform multiple iterations to calculate the weights of the best configurations and the non-best configurations for the different scenarios, using the configuration-specific device performance predictions as input for calculating the weights.

[0180] In some embodiments, each configuration-specific machine learning regression model includes different submodels for different previous prefetcher engine configurations of the previous time window. Therefore, processor 12 is configured to train the different submodels for previous prefetcher engine configurations (block 2124). For example, the regression model for configuration A may include a submodel for configuration A (referred to as "submodel AA"), a submodel for configuration B (referred to as "submodel AB"), and so on. Likewise, the regression model for configuration D may include a submodel for configuration A (referred to as "submodel DA"), a submodel for configuration B (referred to as "submodel DB"), and so on.For example, submodel DB is a regression model trained to predict device performance when the prefetcher engines are to be tuned according to configuration D based on a specific hardware state (e.g., represented by counter values) and the configuration of the previous processing interval is configuration B. Therefore, if there are N configurations, there can be N. 2 provide partial models.

[0181] In use, for a given previous prefetcher engine configuration, a corresponding submodel of each of the configuration-specific machine learning regression models is selected to generate a prediction based on the device hardware state and the given previous prefetcher engine configurations. In other words, in use, the relevant submodels are selected and applied to generate predictions. The best prediction (e.g., the highest or lowest) specifies the configuration to be used in the next processing interval. For example, if the configuration used in the previous processing interval is configuration G, submodel AG, submodel BG, submodel CG, etc. are selected and applied to generate predictions based on the given hardware state of the previous processing interval.

[0182] If submodel BG provides the best prediction, configuration B is selected to configure the prefetcher engines 14 in the next processing interval.

[0183] In some embodiments, instead of N regression models being trained to provide predictions for N configurations, a single classification engine may be trained to predict the best configuration for the next processing interval. In other embodiments, instead of a single classification engine, N classification engines (e.g., N regression models or other suitable classification engines) may be trained to provide predictions for N corresponding previous prefetcher engine configurations as input to the respective classification engines.For example, classification system A may provide a prediction of the best configuration for the next processing interval based on configuration A of the previous processing interval, and classification system B may provide a prediction of the best configuration for the next processing interval based on configuration B of the previous processing interval, and so on.

[0184] The size of the time window for each processing interval is carefully selected. Shorter time windows allow for faster response to changes in the hardware state. However, if the time window is too small, the collected statistics may be too short-term to be stable enough to provide a good estimate of the configuration for the next processing interval. Any time interval can be used, for example, 10s of milliseconds. An application can be divided into checkpoints, which represent large sequences of instructions that can be executed repeatedly during application execution. The size of the checkpoints (or the average size of the checkpoints) can be an indicator of an optimal time window for a particular application. In other words, the time window can be on the same order of magnitude as the average size of the checkpoints.

[0185] Each checkpoint can be assigned a weight based on the number of repetitions of that checkpoint in the application. The checkpoint weights can also be used to calculate the average size of the checkpoints. The checkpoint weights can be used in the training process to bias the model toward higher-frequency samples.

[0186] It will now Fig. 22, which is a view of an example illustrating the use of trained regression models 2200 (for simplicity, only 3 are shown) in the device 10 of Fig. 1 shows. Fig. 22 shows that the prefetcher engines 14 were controlled in a previous processing interval 2204 using configuration D (block 2202). Fig. 22 also shows a hardware state 2208 in (or at the end of) the previous processing interval(s) 2204. The hardware state and optionally the configuration D of the previous processing interval 2204 are used to predict the configuration to be used for a next processing interval 2206. The trained regression models 2200 are applied with the hardware state and optionally the configuration D of the previous processing interval 2204 as input to obtain predictions. The trained regression model 2200 that provides the best (e.g., the highest or lowest) prediction indicates the configuration to be selected in the next processing interval 2206 (block 2210).

[0187] It will now Fig. 23, which is a view of an example illustrating the training of a regression model 2312 in the device 10 of Fig. 1 shows. In the example of Fig. 23, training is performed while the benchmark application 3 is running (block 2300). In a previous processing interval 2302, configuration C (block 2304) is applied to the prefetcher engines 14, and a hardware state 2306 is captured (e.g., at the end of the previous processing interval 2302). In the next processing interval 2308, configuration C (block 2310) is also applied to the prefetcher engines 14, and the device performance 2312 is measured. In general, the same configuration is applied to the prefetcher engines 14 while all benchmark applications are running, and measurements are taken and the hardware states are captured per processing interval. The same is repeated for each configuration, resulting in a data set for each configuration.For example, the prefetcher engines 14 can be set to configuration C while all benchmark applications are running, and measurements are taken and the hardware states are captured per processing interval, resulting in a data set for configuration C, and so on.

[0188] The hardware state 2306 from the previous processing interval 2302 and the device performance 2312 measured in the next processing interval 2308 (along with similar data from running the benchmark applications with configuration C) can be used to train the regression model 2314, which is used to make a prediction for configuration C based on a hardware state as input.

[0189] If the regression model 2314 is trained based on the configuration used in the previous processing interval as input, more data is used. For example, if the regression model 2314 for configuration C is to be trained based on the previous processing interval with configuration D, the data set for configurations C and D is used as described below. The corresponding hardware states determined in the data set for configuration D (based on running the benchmark applications while the prefetcher engines 14 are set to configuration D) are used with the device performance 2312 determined in the data set for configuration C (based on running the benchmark applications while the prefetcher engines 14 are set to configuration C).

[0190] In some embodiments, the corresponding hardware states from the configuration D dataset are used along with the device performance 2312 of the configuration C dataset to train a regression submodel (e.g., submodel CD) that is used to make a prediction for configuration C when the configuration of the previous processing interval 2302 is equal to configuration D.

[0191] It will now Fig. 24, which is a flowchart 2400 illustrating steps in a method for using trained regression models in the apparatus 10 of Fig. 1. In addition, Fig. 1. The processor 12 is configured to execute the software application 20 (block 2402). The processor 12 is configured to execute the machine learning agent 24 to select from various prefetcher engine configurations to control the prefetcher engines 14 (block 2404) while the software application 20 is executing. In some embodiments, the machine learning agent 24 is configured to select from a cured set of prefetcher engine configurations that has been cured using one or more suitable methods, for example, using one or more of the methods described above with reference to Fig. Procedures described in 4 to 20.

[0192] The machine learning agent 24 is configured to receive a device hardware state (e.g., (at the end of) the previous processing interval) (block 2406). If submodels are used, the machine learning agent is configured to select, for the previous prefetcher engine configuration (i.e., the prefetcher engine configuration used in the previous processing interval), a corresponding submodel of each configuration-specific machine learning regression model to create a prediction based on the received device hardware state and the previous prefetcher engine configuration (block 2408). For example, if the configuration used in the previous processing interval is configuration G, submodel AG, submodel BG, submodel CG, and so on are selected.

[0193] If no submodels are used, the machine learning agent 24 is configured to apply the configuration-specific machine learning regression models to generate corresponding configuration-specific performance predictions based on the received device hardware state and, optionally, the prefetcher engine configuration of the prefetcher engines 14 (block 2410). If submodels are used, the machine learning agent 24 is configured to apply the selected submodels (selected in the step of block 2408) for the previous prefetcher engine configuration to generate a prediction based on the received device hardware state and the determined previous prefetcher engine configuration.

[0194] The machine learning agent 24 is configured to select the prefetcher configuration to control the prefetcher engines 14 in the next processing interval based on the best prediction (e.g., the highest or lowest prediction) of the corresponding configuration-specific performance predictions (block 2412). The processor 12 is configured to control the prefetcher engines 14 according to the various prefetcher engine configurations selected by the machine learning agent 24 during execution of the software application. The steps of blocks 2406 through 2414 are repeated at each processing interval.

[0195] It will now Fig. 25, which is a block diagram view of the device 10 of Fig. 1, which represents a cache eviction control. As previously described, the prefetcher engines 14 are configured to: predict the next memory access addresses of the main memory 16 from which data is to be loaded into the cache 18 during execution of the software application 20; and load the data from the predicted next memory access addresses into the cache 18 during execution of the software application 20. The main memory 16 is configured to store a cache eviction table 30 that provides a corresponding cache replacement priority value 32 for each cache line 34. The cache replacement priority values ​​32 contained in the cache eviction table 30 are used to determine which of the cache lines 34 should be evicted first and replaced with new data, as in Fig. 26. The memory 16 is also configured to store one or more cache replacement value policies 36. If there is only one cache replacement priority value policy 36, the cache replacement priority values ​​32 in the cache replacement priority value policy 36 are used to assign the cache replacement priority values ​​32 in the cache displacement table 30, as described with reference to Fig. 26-28. If there is more than one cache replacement priority value policy 36, the processor 12 may be configured to select the best policy to assign the cache replacement priority values ​​32 from the best policy to the cache displacement table 30, as shown in Fig. 29-32 described in more detail.

[0196] It will now Fig. 26, which is a flowchart 2600 with steps in a method for controlling cache eviction in the device 10 of Fig. 25. The processor 12 is configured to assign cache replacement priority values ​​32 to the cache lines 34 based on the prefetcher engines 14 that loaded the cache lines 34 (block 2602). In some embodiments, the processor 12 is configured to assign the cache replacement priority values ​​32 to the cache lines 34, e.g., in the cache eviction table 30, based on the prefetcher engines 14 that loaded the cache lines 34 and the event types of events performed on the cache lines 34, as shown in Fig. 27 and Fig. 28. The event types may include: (i) loading into the cache; and (ii) returning a corresponding one of the cache lines 34 to the cache 18 after initial use, e.g., by the processor 12 or a higher level in a hierarchy of the cache 18. For example, a cache line 34 may be assigned a first cache replacement priority value 32 when it is loaded from memory 16 into the cache 18 and then assigned a different cache replacement priority value 32 when the cache line 34 is returned to the cache 18 after initial use, e.g., by the processor 12 or a higher level in the cache hierarchy.Therefore, in response to an event performed on one of the cache lines 34, the processor 12 is configured to select a relevant cache replacement priority value 32 from the predefined or current cache replacement priority value policy 36 (block 2604) and update the cache eviction table 30 with the selected cache replacement priority value 32 for that cache line 34 (block 2606), as described with reference to FIG. Fig. 27 and Fig. 28. Therefore, the processor 12 is generally configured to select the cache replacement priority values ​​32 to be assigned to the cache lines 34 in the cache eviction table 30 from an array of predefined cache replacement priority values ​​32 (e.g., in the predefined or current cache replacement priority value policy 36). The array may include the cache replacement priority value 32 for correspondingly different combinations of the prefetcher engines 14 and event types. The processor 12 is configured to evict the cache lines 34 from the cache 18 based on the cache replacement priority values ​​32 of the cache lines 34 (block 2608).

[0197] It will now Fig. 27 and Fig. 28, which are schematic views of parts of the device 10 of Fig. 25 and exemplary data 38 used therein to illustrate cache eviction control. Fig. Figure 27 shows that one of the prefetcher engines 14 (e.g., prefetcher engine 3) has selected data 38 from the memory 16 to be loaded into the cache 18 as one of the cache lines 34 (e.g., cache line 1). Loading one of the cache lines 34 is defined in this example as event type A. Therefore, the processor 12 ( Fig. 25) performs a search in the cache replacement priority value policy 36 to find the cache replacement priority value 32 for prefetcher 3 and event type A. The processor 12 determines that the relevant cache replacement priority value 32 is equal to 3. The processor 12 writes (block 40) the found cache replacement priority value 32 (e.g., 3) to the relevant entry (e.g., for cache line 1) of the cache eviction table 30.

[0198] Fig. Figure 28 shows that data 38 (e.g., cache line 1) has been read by processor 12 and is being returned to cache 18. Returning a cache line 34 to cache 18 is defined as event type B in this example. Therefore, processor 12 performs a search of cache replacement priority value policy 36 to find the cache replacement priority value 32 for prefetcher 3 and event type A. Processor 12 determines that the relevant cache replacement priority value 32 is equal to 3. Processor 12 writes (block 42) the found cache replacement priority value 32 (e.g., 1) to the relevant entry (e.g., for cache line 1) of cache eviction table 30.

[0199] Reference is now made to Fig. 29, which shows a flowchart 2900 with steps in a method for selecting one or more policies 36 for use in the method for controlling cache eviction of Fig. 26. Although a best policy is selected for use at any one time, the best policy may change from one policy to another over time due to system conditions as well as the software application or portion thereof currently being executed by processor 12.

[0200] The processor 12 is configured to select, using the machine learning agent 24, one or more cache replacement priority value policies 36 from a group of different cache replacement priority value policies 36 to optimize the operation of the cache according to a metric (block 2902). Each cache replacement priority value policy includes a corresponding array of predefined cache replacement priority values ​​32 for corresponding different combinations of the prefetcher engines 14 and the event types. The metric may be any suitable metric, for example, one or more of the following: IPC; based on the IPC of the software application; or cache efficiency. In some embodiments, the metric may be the metric described above with reference to Fig. 2. The processor 12 is configured to select the cache replacement priority values ​​32 to be assigned to the cache lines 34 in the cache displacement table 30 from the predefined cache replacement priority values ​​32 of the currently selected cache replacement priority value policy 36 (block 2904), as described above with reference to Fig. 27 and Fig. 28 described in more detail.

[0201] In some embodiments, the processor 12 is configured to execute the machine learning agent 24 during execution of the software application 20 in exploration and exploitation to select particular policies from the various cache replacement priority value policies (block 2906) and assign the cache replacement priority values ​​according to the particular policies selected by the machine learning agent. The machine learning agent 24 may explore various cache replacement priority value policies 36 and select a best cache replacement priority value policy 36 to utilize. Best may be defined by the metric measured during exploitation of the cache with the respective cache replacement priority value policies 36.

[0202] In some embodiments, the machine learning agent 24 includes a multi-armed bandit machine learning agent to select from the various cache replacement priority value policies 36 during exploration and exploitation (block 2908).

[0203] Reference is now made to Fig. 30, which is a schematic view of the cache 18 of the device 10 of Fig. 25, which illustrates an exemplary use of the cache 18 by a spatially distributed machine learning agent. In some embodiments, the machine learning agent 24 comprises a spatially distributed machine learning agent configured to perform exploration in different portions 44 of the cache 18 concurrently with certain cache replacement priority value policies 36 (block 2910 of Fig. 29). The spatially distributed machine learning agent can explore a suitable number of cache replacement priority value policies 36 (e.g., all cache replacement priority value policies 36 or a subset of cache replacement priority value policies 36) simultaneously in different parts 44 of the cache 18, so that many policies 36 can be explored simultaneously. When exploring different parts of the cache 18, a suitable metric such as cache efficiency is used, as it can be measured for the respective parts 44 of the cache 18, while the IPC represents a measure of the operation of the entire cache 18. The spatially distributed machine learning agent may also be configured to perform an exploitation in another portion 46 of the cache 18 with one of the policies 36 that was identified as the best policy based on the exploration (block 2912 of Fig. 29).

[0204] It will be repeated again Fig. 29 is referred to.

[0205] In some embodiments, the machine learning agent 24 includes: a multi-armed bandit machine learning agent for selecting from the various cache replacement priority value policies 36; and a spatially distributed machine learning agent for simultaneously exploring different portions of the cache 18 using specific policies 36. Some examples of combining the multi-armed bandit machine learning agent with the spatially distributed machine learning agent are described below.

[0206] In a first example, the machine learning agent 24 is configured to perform exploration and exploitation in different time periods (i.e., exploration and exploitation at different times), the spatially distributed machine learning agent is configured to simultaneously explore different portions 44 of the cache 18 using particular policies 36 (block 2914), and the multi-armed bandit machine learning agent is configured to perform exploitation of the particular policy 36 identified as the best policy (block 2916). It should be noted that the policy 36 identified as the best policy may change over time due to the spatially distributed machine learning agent's exploration of other policies.

[0207] In a second example, the spatially distributed machine learning agent is configured to simultaneously explore different portions 44 of the cache 18 using certain policies to determine a preliminary best policy of the policies 36 that optimizes a first metric (block 2918). The multi-armed bandit machine learning agent is configured to perform a portion of the exploration toward the preliminary best policy to determine a certain best policy of the policies 36 that optimizes a second metric (block 2920). The multi-armed bandit machine learning agent is configured to perform the exploitation with the determined best policy (block 2922). In some embodiments, the first metric is cache efficiency or is based on cache efficiency, and the second metric is IPC or is based on the IPC of the software application.

[0208] It will now Fig. 31, which shows a flowchart 3100 with steps in a first optional sub-method of the method of Fig. 29. When the spatially distributed machine learning agent is deployed in conjunction with the multi-armed bandit machine learning agent, the spatially distributed machine learning agent may perform exploration (e.g., intermittently or continuously) to determine if there is a significant change in cache efficiency by exploring different cache replacement priority value policies 36 in different portions 44 of the cache 18. In response to detecting a significant change in cache efficiency, the spatially distributed machine learning agent may alert the multi-armed bandit machine learning agent to perform further exploration to determine a new best policy 36, if appropriate. Therefore, in some embodiments, the spatially distributed machine learning agent is configured to detect a particular change (e.g.,exceeding a threshold) of cache efficiency and delivers a notification to the multi-armed bandit machine learning agent to increase exploration (block 3102); and the multi-armed bandit machine learning agent is configured to increase exploration for a specified period of time in response to the notification (block 3104).

[0209] It will now Fig. 32, which shows a flowchart 3200 with steps in a second optional sub-method of the method of Fig. 29. The exploration data (and optionally the exploitation data) measured with respect to cache efficiency may be converted to IPC. The converted IPC data may then be added to an accumulated gain used by the machine learning agent 24 (e.g., the Multi-Armed Bandit machine learning agent). In some embodiments, the processor 12 is configured to determine or train a prediction function to predict a measure of IPC from a measure of cache efficiency (block 3202). The training may be based on examples where both cache efficiency and IPC have been measured and may be policy-specific.

[0210] The spatially distributed machine learning agent is configured to measure cache efficiency (block 3204), e.g., during exploration (and exploitation). The processor 12 is configured to convert the cache efficiency measure into an instructions per cycle (IPC) measure (block 3206), e.g., based on the prediction function. The multi-armed bandit machine learning agent is configured to add the converted instructions per cycle measure to the accumulated gain maintained by the multi-armed bandit machine learning agent (block 3208). The steps of blocks 3204-3208 may be repeated intermittently.

[0211] It will now Fig. 33, which is a flowchart 3300 including steps in a method for reducing a number of cache replacement priority value policies 36 for use in the device 10 of Fig. 25. As previously mentioned, the best policy may be selected from a number of cache replacement priority value policies 36. The number of policies available for selection may be a very large group. For example, if there are 15 prefetcher engines in the system and two event types, and each CRPV can be assigned one of four values ​​(e.g., 0-3), there could be 4 to the power of 2 to the power of 15 policies. The processor 12 may be configured to reduce the number of cache replacement priority value policies 36 available for selection by the machine learning agent 24, as described in more detail below. Each cache replacement priority value policy may have a set of predefined cache replacement priority values ​​for corresponding different combinations of: (a) prefetcher engines 14 that have loaded cache lines 34; and (b) event types of events executed on the cache lines 34.

[0212] In general, each policy can be considered to contain values ​​for various combinations of dimensions (e.g., prefetcher engines) and hyperdimensions (e.g., event types). For example, each policy may contain values ​​for each dimension (e.g., each prefetcher engine) in a first hyperdimension (e.g., event type A) and values ​​for each dimension (e.g., each prefetcher engine) in a second hyperdimension (e.g., event type B), and so on. For simplicity, the following description assumes that there are 15 dimensions and 2 hyperdimensions, and therefore each policy includes 15 elements for the first hyperdimension and subsequently 15 elements for the second hyperdimension. The following description also assumes that each element can be set to one of four values, e.g., 0-3.The use of 15 dimensions, two hyperdimensions, and four values ​​in the description below is for example purposes only. The reference to . Fig. 33 and Fig. 34 can include any number of dimensions, hyperdimensions and values.

[0213] The processor 12 is configured to determine reduced sets of policies for corresponding hyperdimensions (block 3302), including determining a first reduced set of policies for the first hyperdimension (block 3304) and determining a second reduced set of policies for the second hyperdimension (block 3306). The steps of blocks 3304 and 3306 are described with reference to Fig. 33 and Fig. 34 described in more detail.

[0214] The processor 12 is configured to combine the reduced sets of policies (i.e., the first reduced set of policies and the second reduced set of policies) in various combinations, resulting in a third set of policies (block 3308), including the combination of the first reduced set of policies for the first hyperdimension with the second reduced set of policies for the second hyperdimension. The step of block 3308 will be described below with reference to Fig. 34 described in more detail below.

[0215] The processor 12 is then configured to reduce the third set of policies to obtain a fourth reduced set of policies available for selection by the machine learning agent 24 (block 3310). The step of block 3310 will be described below with reference to Fig. 34 is described in more detail below. Each policy in the fourth reduced set of policies includes an array of values ​​associated with each of the hyperdimensions, including an array of first values ​​associated with the first hyperdimension and an array of second values ​​associated with the second hyperdimension. In some embodiments, the policies are cache replacement priority value policies 36, the first hyperdimension is a first event type, the second hyperdimension is a second event type, and each cache replacement priority value policy includes an array of predefined cache replacement priority values ​​for corresponding different combinations of: (a) prefetcher engines 14 that have loaded cache lines 34; and (b) event types of events that have been executed on the cache lines 34.

[0216] It will now Fig. 34, which shows a flowchart 3400 with steps in a sub-method of the method of Fig. 33. The steps of blocks 3402-3414 are first briefly described below and then explained in more detail after the brief description.

[0217] The processor 12 is configured to detect an impact on the performance of the device 10 when changing certain values ​​of the first hyperdimension and the second hyperdimension while executing benchmark applications (block 3402). The processor 12 is configured to define groups of dimensions for the first hyperdimension and groups of dimensions for the second hyperdimension in response to the impact on the performance of the device 10 (block 3404). The processor 12 is configured to determine preliminary groups of policies for the first hyperdimension and the second hyperdimension based on the defined groups, such that for each of the defined groups, the same value must be applied to all associated dimensions (block 3406).The processor 12 is configured to apply a greedy algorithm (or other suitable algorithm) to reduce the size of the preliminary policy sets to obtain the first reduced policy set and the second reduced policy set (block 3408). The step of block 3408 may include the substeps of blocks 3410-3414, which are described in more detail below. The processor 12 is configured to receive measurements of the performance of the device 10 executing the benchmark applications according to the preliminary policy sets for the first hyperdimension and the second hyperdimension based on the defined sets (block 3410).The processor 12 is configured to define orderings of at least some of the policies of the preliminary sets of policies (block 3412) and to provide the first reduced set of policies and the second reduced set of policies based on the defined orderings (block 3414).

[0218] The steps of blocks 3402-3414 will now be described in more detail with respect to the first hyperdimension. As previously mentioned, for simplicity, the following description assumes that there are 15 dimensions and 2 hyperdimensions, and therefore each policy includes 15 elements for the first hyperdimension and then 15 elements for the second hyperdimension.

[0219] The step of block 3402 will now be described in more detail. The processor 12 executes the benchmark applications (e.g., 100 benchmark applications) with a default policy in which all values ​​for both hyperdimensions are set to a value such as 2. This provides an output vector such as the one described above with reference to Fig. 5. Then, one of the values ​​of the first hyperdimension in the default policy for one of the dimensions is changed to another value, e.g., to 3, to obtain an updated policy. The processor 12 then executes the benchmark applications with the updated policy to obtain a change vector, such as the one described above with reference to Fig. 5. The processor 12 then calculates an impact vector (such as the one described with reference to Fig. 5) by performing an element-wise subtraction of the vector 54 from the vector 52. The processor 12 then calculates an average and a variance (e.g., standard deviation or variance) of the impact vector, as described above with reference to the average 56 and variance 58 of Fig. 5. An updated policy is generated from the standard vector, and a change vector is calculated for each of the dimensions of the first hyperdimension based on the steps described above. The processor 12 is configured to calculate an average and a variance for each of the change vectors. The processor 12 calculates the similarity values ​​(e.g., cosine similarity) between the different change vectors, as shown in Fig. 7. Based on the similarity and dispersion values ​​58, the processor 12 finds groups of dimensions (e.g., prefetcher engines 14) in the first hyperdimension (e.g., event type A) (block 3404), as described with reference to Fig. 8-10 described in more detail.

[0220] In the step of block 3406, the processor 12 is configured to determine the preliminary set of policies for the first hyperdimension based on the groups determined in the step of block 3404. For example, each policy contains values ​​for the first hyperdimension by assigning values ​​to the elements of the first hyperdimension based on the groups according to the Fig. 11 and Fig. 12, while the values ​​of the elements of the second hyperdimension are assigned a default value such as 2. For example, in a policy, all values ​​for the second hyperdimension can be assigned the value 2, while the elements of the dimensions in group A of the first hyperdimension can each be assigned the value 0, and the elements of the dimensions in group B of the first hyperdimension can each be assigned the value 3, and so on.

[0221] The preliminary set of policies for the first hyperdimension may be reduced using a suitable reduction method, for example a greedy algorithm, which is described in more detail below with reference to the steps of blocks 3410-3414.

[0222] Processor 12 is configured to execute the benchmark applications for each preliminary set of policies for the first hyperdimension and obtain the measurement of the performance of device 10 (e.g., IPC) based on the execution of the benchmark applications according to the preliminary set of policies (block 3410). For example, if there are 100 benchmark applications and 256 possible preliminary policies, then there are 100 x 256 measurements of the performance of device 10, assuming that device performance can be measured with one metric (e.g., IPC). If device performance is measured with Q metrics, then there are 100 x 256 x Q measurements of device performance.

[0223] The step of block 3412 will now be described in more detail below. For each policy, there is a vector that indicates, for example, the performance for each benchmark application. Each vector contains elements, where each element indicates the device performance when executing a corresponding benchmark application (or part of the benchmark application), e.g., element 1 indicates the device performance for application 1, element 2 indicates the device performance for application 2, etc. The processor 12 is configured to calculate an average (e.g., a geometric mean) of the elements of each vector. In other words, the processor 12 is configured to calculate an average of the elements of each of the vectors, thus averaging the device performance indications for each of the benchmark applications or parts of the benchmark applications in each vector. For example, if there are 256 vectors, the processor 12 calculates 256 averages.The processor 12 is configured to generate a vector (e.g., similar to vector 86 in . Fig. ) of the vectors with the highest average and defines the best policy as the policy corresponding to the vector with the highest average. The order of the other policies can also be determined to find the second-best policy, the third-best policy, and so on according to the greedy algorithm (or any other suitable algorithm) in a similar manner as described with reference to Fig. 17-19. In the above procedure, the preliminary set of policies can be reduced to a smaller number, e.g., 12 policies, thereby defining the first reduced set of policies for the first hyperdimension (block 3414).

[0224] The processor 12 may be configured to determine the second reduced set of policies for the second hyperdimension (block 3306). In some embodiments, the processor 12 is configured to determine the second reduced set of policies for the second hyperdimension using the method described above with reference to finding the first reduced set of policies for the first hyperdimension, with appropriate modifications, without taking into account the values ​​of the first reduced set of policies. As previously mentioned above, in the step of block 3402, the processor 12 executes the benchmark applications (e.g., 100 benchmark applications) with a default policy where all values ​​for both hyperdimensions are set to a value such as 2. This provides an output vector such as the one described above with reference to Fig. 5. Then, one of the values ​​of the second hyperdimension in the default policy for one of the dimensions is changed to another value, for example, to 3, to obtain an updated policy.

[0225] However, in other embodiments, the processor 12 is configured to determine the second reduced set of policies for the second hyperdimension based on the first reduced set of policies of the first hyperdimension (arrow 3320). The processor 12 may be configured to determine the second reduced set of policies for the second hyperdimension based on the best policy per benchmark application of the first reduced set of policies of the first hyperdimension. For example, the policies used to determine the second reduced set of policies for the second hyperdimension may include values ​​for the first hyperdimension from the best policy per benchmark application previously determined in the steps of blocks 3402-3414 to determine the first reduced set of policies.For example, for a default policy used in the execution of Application 1, the default policy may include default values ​​(e.g., all set to 2) for the second hyperdimension, and for the values ​​of the first hyperdimension, processor 12 selects the values ​​of the first hyperdimension from the best policy for Application 1. The best policy for Application 1 may be determined from the 12 policies of the first reduced set of policies determined in the step of block 3414, wherein the policy of these 12 policies that results in the highest value (e.g., IPC) when executing Benchmark Application 1 is determined. Then, one of the values ​​of the default policy for one of the dimensions is changed to a different value, e.g., 3, to provide an updated policy for that dimension to be used in the execution of Application 1, and the method described above with respect to the steps of block 3402 is performed.This is repeated for each of the applications and each of the dimensions. In the step of block 3406 and the policies used in the step of block 3314 of the greedy algorithm, the policies used may be per-application and may include first hyperdimension values ​​from the best policy for each application, as described above. In the step of block 3316, the per-application values ​​from the first hyperdimension are ignored, and step 3318 results in only 12 policies for the second hyperdimension.

[0226] It will be repeated again Fig. 33 referred to.

[0227] In the step of block 3308, the processor 12 is configured to combine the reduced sets of policies in various combinations, resulting in the third set of policies. For example, the processor 12 takes the 12 best policies for the first hyperdimension (i.e., the values ​​of the first hyperdimension in the 12 best policies of the first reduced set of policies) and the 12 best policies for the second hyperdimension (i.e., the values ​​of the second hyperdimension in the 12 best policies of the second reduced set of policies) and combines the policies in various combinations. For example, the best policy of the first hyperdimension is combined with the third best policy of the second hyperdimension, and so on. For example, if there are 12 best policies for the first hyperdimension and 12 best policies for the second hyperdimension, there could be 144 policies in the third set of policies.

[0228] In the step of block 3310, the processor 12 is configured to reduce the number of the third set of policies, resulting in a fourth set of policies. In some embodiments, the processor 12 is configured to reduce the third set of policies to obtain the fourth reduced set of policies by applying a greedy algorithm (block 3312), for example, using substantially the same method as described above with reference to the steps of blocks 3408-3414.

[0229] In some embodiments, the processor 12 is configured to: provide measurements of the performance of the device 10 executing the benchmark applications according to the third set of policies (block 3314); define an ordering of at least some of the policies of the third set of policies (block 3316); and provide the fourth reduced set of policies based on the defined ordering (block 3318) determined in the step of block 3314. In some embodiments, the greedy algorithm may be replaced by any suitable algorithm for selecting a reduced set of policies, such as selecting the policies based on the average value of the vectors representing the device's performance for each of the policies, or by exhaustively trying different combinations of 2 of the vectors and 3 of the vectors, and so on.

[0230] In some embodiments, the results of the first reduced set of policies can be used in the second reduced set of policies, so that the results of the second reduced set of policies can be used in the first reduced set of policies (arrow 3322), for example, to re-perform the various calculations to refine the first reduced set of policies. Likewise, the results of the refined first reduced set of policies can be used to recalculate the second reduced set of policies (arrow 3320), and so on. This process can be repeated a certain number of times or until the incremental improvement in the results is below a certain threshold.In some embodiments, therefore, the processor 12 is configured to iteratively determine the first reduced set of policies and the second reduced set of policies by performing an iterative process in which (a) intermediate results from the process of determining the first reduced set of policies are used in the process of determining the second reduced set of policies, and (b) intermediate results from the process of determining the second reduced set of policies are used in the process of determining the first reduced set of policies.

[0231] In some embodiments, processor 12 may be configured to reduce the number of policies if each policy comprises three or more hyperdimensions. The above reduction method may be generalized to more than two hyperdimensions. For example, the method described above may be performed for the first and second hyperdimensions, while the values ​​of higher hyperdimensions have default values, e.g., set to 2, resulting in the fourth reduced set of policies. The steps of blocks 3402-3414 may then be performed for the third hyperdimension, using the best policy per application of the combined first and second hyperdimensions, with default values, e.g., 2, set for hyperdimensions higher than three (if relevant).Processor 12 is then configured to combine the reduced combined set of policies for the first and second hyperdimensions with the vectors found for the third hyperdimension in various combinations. The combined vectors are then reduced using an algorithm, e.g., a greedy algorithm. This is repeated for the fourth hyperdimension, and so on, until all hyperdimensions have been processed and a final reduced set of policies is obtained.

[0232] Various features of the disclosure that are described in conjunction with separate embodiments for clarity may also be provided in combination in a single embodiment. Conversely, various features of the disclosure that are described in conjunction with a single embodiment for brevity may also be provided separately or in any suitable subcombination.

[0233] The embodiments described above are provided as examples, and the present disclosure is not limited to what has been specifically shown and described above. Rather, the scope of the disclosure includes combinations and subcombinations of the various features described herein, as well as variations and modifications thereof that would occur to those skilled in the art upon reading the above description and that are not disclosed in the prior art.

[0234] It is to be understood that aspects and embodiments described above are purely exemplary and that modifications of details may be made within the scope of the claims.

[0235] Each device, method, and feature disclosed in the description, and (where appropriate) the claims and drawings may be provided independently or in any suitable combination.

[0236] Reference signs appearing in the claims are for illustrative purposes only and do not limit the scope of the claims.

Claims

[1] System comprising: a processor to reduce a number of cache replacement priority value policies available for selection by a machine learning agent, each cache replacement priority value policy containing an array of predefined cache replacement priority values ​​for corresponding different combinations of: (a) prefetcher engines that have loaded cache lines; and (b) event types of events that have been executed on the cache lines; and a memory to store data used by the processor. [2] System comprising: a processor to: determine reduced sets of policies for respective hyperdimensions, comprising determining a first reduced set of policies for a first hyperdimension and determining a second reduced set of policies for a second hyperdimension; combining the reduced sets of policies in various combinations, resulting in a third set of policies comprising combining the first reduced set of policies for the first hyperdimension with the second reduced set of policies for the second hyperdimension; and reduce the third set of policies to obtain a fourth reduced set of policies available for selection by a machine learning agent, each policy in the fourth reduced set containing an array of values ​​associated with each of the hyperdimensions, including an array of first values ​​associated with the first hyperdimension and an array of second values ​​associated with the second hyperdimension; and a memory to store data used by the processor. [3] The system of claim 2, wherein: the policies are cache replacement priority value policies; the first hyperdimension is a first event type; the second hyperdimension is a second event type; and each of the cache replacement priority value policies contains an array of predefined cache replacement priority values ​​for corresponding different combinations of: (a) prefetcher engines that have loaded cache lines; and (b) event types of events that have been performed on the cache lines. [4] A system according to claim 2 or 3, wherein the processor is arranged to: to determine an impact on the performance of a device resulting from changing setpoints of the first hyperdimension and the second hyperdimension during the execution of benchmark applications; to define groups of the first hyperdimension and groups of the second hyperdimension in response to the effect on the performance of the device; and to determine preliminary groups of policies for the first hyperdimension and the second hyperdimension based on the defined groups, so that an equal value is to be applied to each of the defined groups. [5] The system of claim 4, wherein the processor is configured to apply a greedy algorithm to reduce a size of the preliminary sets of policies to obtain the first reduced set of policies and the second reduced set of policies. [6] The system of claim 5, wherein the processor is arranged to: receive measurements of the performance of the device executing the benchmark applications according to the preliminary sets of policies for the first hyperdimension and the second hyperdimension based on the defined sets; define orderings of at least some of the policies of the preliminary groups of policies; and to provide the first reduced set of policies and the second reduced set of policies based on the defined sequences. [7] The system of any of claims 2-6, wherein the processor is arranged to reduce the third set of policies to obtain the fourth reduced set of policies by applying a greedy algorithm. [8] The system of claim 7, wherein the processor is arranged to: receive measurements of the performance of the device executing the benchmark applications according to the third group of policies; define an order of at least some of the policies of the third group of policies; and and provide the fourth reduced set of policies based on the defined order. [9] The system of any of claims 2-8, wherein the processor is arranged to determine the second reduced set of policies for the second hyperdimension based on the first reduced set of policies of the first hyperdimension. [10] The system of claim 9, wherein the processor is configured to determine the second reduced set of policies for the second hyperdimension based on a best policy per benchmark application of the first reduced set of policies of the first hyperdimension. [11] The system of claim 9 or 10, wherein the processor is arranged to iteratively determine the first reduced set of policies and the second reduced set of policies by performing an iterative process in which: (a) intermediate results from the process for determining the first reduced set of policies are used in the process for determining the second reduced set of policies, and (b) intermediate results from the process for determining the second reduced set of policies are used in the process for determining the first reduced set of policies. [12] Procedure comprising: Determining reduced sets of policies for respective hyperdimensions, comprising determining a first reduced set of policies for a first hyperdimension and determining a second reduced set of policies for a second hyperdimension; Combining the reduced sets of policies in various combinations, resulting in a third set of policies comprising combining the first reduced set of policies for the first hyperdimension with the second reduced set of policies for the second hyperdimension; and Reducing the third set of policies to obtain a fourth reduced set of policies available for selection by a machine learning agent, wherein each policy in the fourth reduced set contains an array of values ​​associated with each of the hyperdimensions, including an array of first values ​​associated with the first hyperdimension and an array of second values ​​associated with the second hyperdimension. [13] The method of claim 12, wherein: the policies are cache replacement priority value policies; the first hyperdimension is a first event type; the second hyperdimension is a second event type; and each of the cache replacement priority value policies contains an array of predefined cache replacement priority values ​​for corresponding different combinations of: (a) prefetcher engines that have loaded cache lines; and (b) event types of events that have been performed on the cache lines. [14] A method according to claim 12 or 13, further comprising: Determining an impact on the performance of a device resulting from changing setpoints of the first hyperdimension and the second hyperdimension during the execution of benchmark applications; Defining groups of the first hyperdimension and groups of the second hyperdimension in response to the effect on the performance of the device; and Determine preliminary groups of policies for the first hyperdimension and the second hyperdimension based on the defined groups, so that an equal value is to be applied to each of the defined groups. [15] The method of claim 14, further comprising applying a greedy algorithm to reduce a size of the preliminary sets of policies to obtain the first reduced set of policies and the second reduced set of policies. [16] The method of claim 15, further comprising: Receiving measurements of the performance of the device executing the benchmark applications according to the preliminary sets of policies for the first hyperdimension and the second hyperdimension based on the defined sets; Defining orderings of at least some of the policies of the preliminary groups of policies; and Deploy the first reduced set of policies and the second reduced set of policies based on the defined orders. [17] The method of any of claims 12-16, further comprising reducing the third set of policies to obtain the fourth reduced set of policies by applying a greedy algorithm. [18] The method of claim 17, further comprising: Receiving measurements of the performance of the device executing the benchmark applications according to the third set of policies; Defining an order of at least some of the policies of the third group of policies; and Deploy the fourth reduced set of policies based on the defined order. [19] The method of any of claims 12-18, wherein determining comprises determining the second reduced set of policies for the second hyperdimension based on the first reduced set of policies of the first hyperdimension. [20] The method of claim 19, wherein determining comprises determining the second reduced set of policies for the second hyperdimension based on a best policy per benchmark application of the first reduced set of policies of the first hyperdimension. [21] The method of claim 19 or 20, wherein determining comprises iteratively determining the first reduced set of policies and the second reduced set of policies by performing an iterative process in which: (a) intermediate results from the process of determining the first reduced set of policies are used in the process of determining the second reduced set of policies, and (b) intermediate results from the process of determining the second reduced set of policies are used in the process of determining the first reduced set of policies.