A Prefetch Control Strategy Based on Improved Hill Climbing Method under Asymmetric Multi-Core Architecture
By adopting improved mountain climbing method and simulated annealing algorithm under the asymmetric multi-core architecture, the degree of prefetching radicalization is dynamically adjusted, and the system performance degradation caused by the performance differences between large and small cores is solved, and higher core IPC and system performance is achieved.
Patent Information
- Application Number
- CN202210750282.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-28
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-06-28
AI Technical Summary
Under the asymmetric multi-core architecture, the performance differences between large and small cores make it difficult for existing prefetch control strategies to effectively adjust the degree of prefetching aggressiveness, resulting in a degradation of system performance.
The improved mountain climbing method combined with simulated annealing algorithm is used to dynamically adjust the degree of prefetching radicality of different cores, and improve the core IPC by searching for the best combination of prefetch distance and prefetching degree.
Through the improved prefetch control strategy, the degree of prefetching can be dynamically adjusted at different cores and stages, improving the overall performance of the system, avoiding local optimal solutions, and improving search efficiency.
Smart Images

Figure CN115114189B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of cache prefetching and asymmetric multi-core, and particularly relates to a prefetch control strategy under an asymmetric multi-core architecture based on an improved hill climbing method. Background Art
[0002] As a buffer component between the CPU and the main memory, the cache plays an important role in balancing the relationship between their frequencies and capacities. In order to give full play to the role of the cache, cache prefetching, as a mature technology, can reduce cache misses to a certain extent and improve the cache hit rate. The basic idea of cache prefetching is that it can predict the cache blocks that may be missing in the future through a certain strategy, so as to request the data of the cache block from the memory in advance. If the missing cache block is successfully predicted, the latency of cache misses will be reduced, thereby improving the overall performance of the system.
[0003] Cache prefetching is usually closely related to two metrics: prefetch distance and prefetch degree. Taking the commonly used streaming prefetching as an example, it will record the missing addresses that appear regularly at a certain span. If the number of occurrences of such regular missing addresses reaches a certain threshold, the prefetching device will determine that the addresses that may be missing according to this rule will appear subsequently, and this threshold is called the prefetch distance. When the prefetching device determines that this rule takes effect, it will send requests for multiple subsequent addresses at a time, and this request number is called the prefetch degree.
[0004] In the entire system, there will be two types of memory access requests. One is the memory access request when the CPU has a data miss, and this type of request is called a demand request. This type of request will cause the CPU to pause and affect the system performance. The other type of request is called a prefetch request. When this type of request successfully predicts the missing address, it will reduce the CPU pause time, thereby improving the system. For a single-core processor, prefetching can achieve better results in a relatively aggressive case. However, for a multi-core processor, aggressive prefetching may have the opposite effect, because too many prefetch requests will cause the bandwidth of the last-level shared cache and the main memory to saturate, resulting in a longer response time for demand requests and a correspondingly longer CPU pause time, thereby causing a decline in system performance.
[0005] Aggressive prefetching strategies usually bring higher performance improvements to applications with accurate prefetch requests, but they may also impose significant pressure on the bandwidth of the last-level cache and main memory. When the bandwidth of the last-level cache and main memory saturates, the response time of demand requests will increase, thus affecting the overall system performance. Many scholars have conducted research in this area. Among them, Ebrahimi et al. demonstrated the performance degradation and memory bandwidth waste caused by aggressive prefetching in multi-core systems and proposed a multi-level aggressive prefetching control strategy (HPAC), which can dynamically adjust the aggressiveness of the prefetcher for multiple cores to control the inter-core interference caused by the prefetcher. Panda proposed that HPAC did not consider the impact between the throttling strategies of multiple prefetcher, and on this basis, proposed a cooperative prefetching control strategy (SPAC), which explored the interaction between the limiting decisions of the prefetcher and controlled the aggressiveness of the prefetcher based on the fairness speedup improvement of the multi-core system. At the same time, Panda et al. pointed out some drawbacks of the prefetching control strategy. Commonly used prefetching control strategies usually set fixed thresholds, which change according to the system configuration requirements, and no single threshold can be well-suited for all workloads. To overcome these shortcomings, the authors proposed the CAFFEE prefetching control strategy. Ridharan et al. showed that CAFFEE makes biased decisions due to the approximate estimation of the loss caused by last-level cache misses. Based on two observations, the authors proposed the Band-Pass prefetching control strategy. Heirman et al. proposed the idea of proximal throttling based on the principle of detecting the latency of prefetching and adjusting the prefetch distance to closely track the point where most prefetching is not delayed. Previous prefetching control strategies usually rely on homogeneous architectures, where the core configurations are the same, and these strategies have the same control requirements for the cores. Such control strategies often set corresponding thresholds based on metrics such as prefetch accuracy, prefetch coverage, and cache pollution caused by prefetching. When the evaluation metrics reach the thresholds, aggressive prefetching is controlled. However, such prefetching strategies will not be applicable to asymmetric multi-core architectures, such as big.LITTLE architectures.
[0006] As an asymmetric multi-processor architecture, the big.LITTLE architecture has been widely applied to modern embedded platforms. Among them, the big cores have higher clock frequencies and are suitable for computing-intensive tasks, while the little cores have lower clock frequencies and perform better in memory-intensive tasks. In the current architecture, the big and little cores can work simultaneously to adapt to more flexible scenarios, but this also poses greater challenges to system load balancing. Figure 1 It is the system architecture diagram of the big.LITTLE architecture. This architecture consists of a big core DIE and a little core DIE. Each core has a private L1 cache and a private L2 cache. Each L2 cache, last-level cache, and Mem are interconnected through a network. There is a prefetcher beside each L1 cache.
[0007] Use the multi-threaded benchmark PARSEC to test the memory bandwidth requirements of different cores. Since PARSEC is multi-threaded during operation, the same program can run on different cores simultaneously, enabling us to observe the differences between big and small cores in running multi-threaded programs. The memory bandwidth results required by big and small cores during program execution are as Figure 2 shown.
[0008] To test the single-threaded scenario, use SPEC for testing. Bind the same benchmark to different big and small cores. The IPC results of big and small cores are different under different prefetch distances and prefetch degrees, and the IPC trend changes reflected by big and small cores on some benchmarks are also different. The results are as Figure 3 、 4 shown, indicating that different prefetch aggressiveness levels should be selected for big and small cores. If the same throttling strategy is adopted for big and small cores, some cores will not achieve their best performance. Summary of the Invention
[0009] To solve the above problems, the present invention proposes a prefetch control strategy based on an improved hill-climbing method for an asymmetric multi-core architecture. This research considers the situation where big and small cores have different performances and uses the hill-climbing method to control the prefetch aggressiveness level. By combining simulated annealing algorithm optimization, it avoids falling into local optimal solutions and aims to improve the core IPC at different stages of program execution for different cores. To achieve this goal, a prefetch distance set, a prefetch degree set, and a candidate node set are first set. We sequentially traverse the values in the prefetch degree set. When the prefetch degree is fixed, use the hill-climbing method combined with simulated annealing to search for the prefetch distance. In each round of search, we record the prefetch distance and prefetch degree corresponding to the maximum IPC that appears during the current round of search, and after the search is completed, add the optimal prefetch degree and prefetch distance of the current round to the candidate node set. After searching all prefetch degree nodes, select the optimal value from the candidate node set for each subsequent adjustment and continuously update the optimal value in the candidate node set. The specific technical solutions are as follows:
[0010] A prefetch control strategy based on an improved hill-climbing method for an asymmetric multi-core architecture. This strategy takes into account the fact that the performance differences between big and small cores will lead to different feedback on the prefetch aggressiveness level, and adjusts the aggressiveness level of different cores through the improved hill-climbing method to improve the core IPC. It mainly includes the following steps:
[0011] Step 1), set a prefetch distance set PDIS and a prefetch degree set PDEG;
[0012] Step 2), in a way of controlling variables, select an un-searched prefetch degree from the prefetch degree set. In this case of the prefetch degree, randomly select a prefetch distance from the prefetch distance set. The selected two are used as the initial values of the prefetch degree and the prefetch distance. Without changing the prefetch degree, continuously change the prefetch distance to find the prefetch distance corresponding to the maximum IPC obtained by each core. The prefetch degree and the prefetch distance corresponding to the maximum IPC become the candidate best prefetch aggressiveness;
[0013] Step 3), traverse all the prefetch degrees in the prefetch degree set. Each core selects the candidate best prefetch aggressiveness that maximizes its own IPC from its corresponding candidate best prefetch aggressiveness. This candidate best prefetch aggressiveness is the best prefetch aggressiveness.
[0014] Beneficial effects
[0015] In view of the different performances among cores in the asymmetric multi-core architecture, the present invention controls the prefetch aggressiveness for different cores by combining the hill climbing method and the simulated annealing strategy to improve the IPC of the cores. At the same time, considering that the hill climbing method is not applicable to solving two-dimensional search problems, this method adopts the way of controlling variables and adjusts the prefetch strategy through the IPC feedback in the sampling stage to enhance the IPC of the cores.
[0016] The present invention is designed based on the improved hill climbing method, and avoids falling into the local optimum through the simulated annealing algorithm. At different stages of program operation, the prefetch aggressiveness of different cores is adjusted according to the historical results, which can reduce the search of the global space and improve the search efficiency. Through this strategy, the optimal prefetch distance and prefetch degree are selected to improve the IPC of the cores. Description of the drawings
[0017] Figure 1 It is the architecture diagram of the big.LITTLE system
[0018] Figure 2 It is the memory bandwidth required by the big.LITTLE during program operation
[0019] Figure 3 It is the IPC result of the big core under different prefetch distances and prefetch degrees
[0020] Figure 4 It is the IPC result of the small core under different prefetch distances and prefetch degrees
[0021] Figure 5 It is the process composition of the program operation
[0022] Figure 6(a) is the experimental comparison result of the big core 0
[0023] Figure 6(b) is the experimental comparison result of the big core 1
[0024] Figure 6(c) shows the experimental comparison results of small core 0
[0025] Figure 6(d) shows the experimental comparison results of small core 1
[0026] Figure 7 This is the flowchart of the method of the present invention Detailed implementation manners
[0027] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the implementation of the present invention will be described in detail below with reference to the accompanying drawings
[0028] The present invention relates to a prefetch control strategy under an asymmetric multi-core architecture based on an improved hill climbing method. This method considers the different performances among cores under the asymmetric multi-core architecture. Through the improved hill climbing method, it jumps out of the local optimum by combining with the simulated annealing strategy, controls the prefetch aggressiveness for different cores, and improves the IPC of the cores. At the same time, considering that the hill climbing method is not applicable to solving two-dimensional search problems, this method adopts the method of controlling variables, and adjusts the prefetch strategy through the IPC feedback in the sampling stage to improve the IPC of the cores. This method uses the hill climbing method to adjust the prefetch aggressiveness of different cores according to different stages of program operation, and realizes the IPC growth of different cores
[0029] The specific steps are as follows
[0030] Step 1: Set the prefetch distance set PDIS, the prefetch degree set PDEG, the candidate node set C, initialize the temperature T, the cooling rate R, and the search round S
[0031] At the beginning of the system, the prefetch distance set PDIS and the prefetch degree set PDEG are set. Thereafter, the prefetch aggressiveness will consist of two parts: the prefetch distance dis and the prefetch degree deg. The entire optimization process consists of two parts: the sampling stage and the adjustment stage. The sampling stage will be simulated by the initial values set by the system or the prefetch degree and prefetch distance adjusted in the adjustment stage, and the best combination of the optimal prefetch distance and prefetch degree is continuously searched through the adjustment stage to achieve the best IPC of the core. The process is as Figure 5 shown
[0032] Step 2: A prefetch control strategy under an asymmetric multi-core architecture based on an improved hill climbing method, the specific steps are as follows
[0033] S1. Select an unselected prefetch degree deg from the prefetch degree set PDEG, and randomly select a prefetch distance dis from the prefetch distance set PDIS under the prefetch degree deg. The selected two are used as the initial values of the prefetch degree and the prefetch distance. Set the prefetch distance modification amplitude ΔD, set the initial round s = 0, and initialize the best prefetch distance disbest , Optimal prefetch degree deg best , Highest IPC best . In each round of the search process, the control variable method is adopted. Here, the prefetch degree is kept constant, and the hill-climbing method is used to search for the prefetch distance.
[0034] S2. For each core, make adjustments according to its own IPC change. Sample the IPC during the running process under the current prefetch degree deg and prefetch distance dis, and count the sampled IPC of the current segment, denoted as C ipc , if s > 1, then retain the IPC recorded in the previous sampling process (i.e., the (s - 1)th round), denoted as L ipc , select a random value rand, and set the probability p = exp(-(C ipc - L ipc )) / T).
[0035] S3. In the adjustment stage, if C ipc > L ipc , and {dis + ΔD} ∈ PDIS, then dis` = dis + ΔD, otherwise dis` = dis, where dis` represents the updated prefetch distance;
[0036] If C ipc ≤ L ipc and p ≥ rand and {dis + ΔD} ∈ PDIS, then dis` = dis + ΔD, otherwise dis` = dis, where dis` represents the updated prefetch distance;
[0037] If C ipc ≤ L ipc and p < rand and {dis + ΔD} ∈ PDIS, then ΔD` = -ΔD, dis` = dis + ΔD`, otherwise ΔD` = -ΔD, dis` = dis, where dis` represents the updated prefetch distance, and ΔD` represents the updated prefetch distance modification amplitude;
[0038] S4. T = T * R, s = s + 1. If C ipc > ipc best , then dis best = dis, deg best = deg. If s < S, jump to S2; if s = S and there are still unselected deg in PEDG, {deg best , dis best , ipc best} is added to C, and jump to S1; if s = S and there are no candidate deg in PEDG, {deg best , dis best , ipc best} is added to C, and jump to Step Three.
[0039] Step 3: Select deg and dis with the highest IPC from the candidate set C as the prefetch degree and prefetch distance.
[0040] Sample the IPC during the running process under the current prefetch degree deg and prefetch distance dis, update {deg best , dis best , ipc best} and add it to C. Repeat Step 3 until the program ends.
[0041] It can be seen from the experimental result graphs 6(a)-(d) that the improved hill climbing method can timely adjust the prefetch aggressiveness of different cores. Between prefetch conservatism and prefetch aggressiveness, it can always make a better choice as the prefetch strategy during the program running and can improve the IPC of the cores.
Claims
1. A prefetch control strategy under an asymmetric multi-core architecture based on an improved hill climbing method, characterized in that it includes the following steps: Step 1), set a prefetch distance set PDIS and a prefetch degree set PDEG; Step 2), in a way of controlling variables, select an un-searched prefetch degree from the prefetch degree set. In this case of the prefetch degree, randomly select a prefetch distance from the prefetch distance set. The selected two are used as the initial values of the prefetch degree and the prefetch distance. Without changing the prefetch degree, continuously change the prefetch distance to find the prefetch distance corresponding to the maximum IPC obtained by each core. The prefetch degree and the prefetch distance corresponding to the maximum IPC become the candidate best prefetch aggressiveness; Step 3), traverse all the prefetch degrees in the prefetch degree set. Each core selects the candidate best prefetch aggressiveness that maximizes its own IPC from its corresponding candidate best prefetch aggressiveness. This candidate best prefetch aggressiveness is the best prefetch aggressiveness.
2. The prefetch control strategy under an asymmetric multi-core architecture based on an improved hill climbing method according to claim 1, characterized in that: The method for changing the prefetch distance is specifically as follows, For each core, adjust the prefetch distance according to the change of its own IPC, specifically: Sample the IPC during the running process under the current prefetch degree deg and prefetch distance dis, and count the sampled IPC of the current segment, denoted as C ipc , if the search round s > 1, then retain the record of the previous sampling process, that is, the IPC of round s - 1, denoted as L ipc , select a random value rand, and set the probability p = exp(–(C ipc –L ipc ) / T, where T represents the temperature; If C ipc > L ipc , and {dis + ΔD} ∈ PDIS, then dis` = dis + ΔD; otherwise, dis` = dis, where ΔD represents the prefetch distance modification amplitude and dis` represents the updated prefetch distance; If C ipc ≤L ipc and p ≥ rand and {dis + ΔD} ∈ PDIS, then dis` = dis + ΔD, otherwise dis` = dis, where dis` represents the updated prefetch distance; If C ipc ≤L ipc and p < rand and {dis + ΔD} ∈ PDIS, then ΔD` = -θD, dis` = dis + ΔD`, otherwise ΔD` = -ΔD, dis` = dis, where dis` represents the updated prefetch distance and ΔD` represents the updated prefetch distance modification magnitude; The updated prefetch distance dis` is used as the prefetch distance during the (s + 1)-th round of IPC sampling; Update the temperature T = T * R, and update the round number s = s + 1, where R represents the cooling rate.