Apparatus and method for optimizing application execution based on memory-level parallelism (MLP) metrics

Through a performance analyzer based on memory-level parallel metrics, analyzing the missed state processing register queue of the hardware processor and generating optimization recommendation data, solving the problem of modern processor performance bottleneck identification and optimization, improving processor performance and simplifying the process.

CN117009261BActive Publication Date: 2025-07-08HEWLETT PACKARD ENTERPRISE DEV LP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211309682.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-05-06
Filing Date
2022-10-25
Publication Date
2025-07-08
Estimated Expiration
2042-10-25

AI Technical Summary

Technical Problem

Due to the architectural complexity of modern hardware processors and the difficulty of interpreting performance counters, the reasons for performance bottlenecks are difficult to accurately identify and optimize, and existing performance tools cannot effectively guide performance improvements.

Method used

Using a performance analyzer based on memory-level parallel (MLP) metrics, it generates optimized recommended data to guide the performance improvement of hardware processors by analyzing the average occupancy and capacity of the missed state processing register (MSHR) queue.

Benefits of technology

Simplifies performance bottleneck identification and optimization process, improves processor performance, is suitable for a wide range of hardware processor vendors, and reduces expertise requirements for microarchitecture details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117009261B_ABST
    Figure CN117009261B_ABST
Patent Text Reader

Abstract

The present disclosure relates to optimizing application execution based on memory-level parallel (MLP) metrics. A process includes: determining a memory bandwidth of a processor subsystem corresponding to execution of an application by the processor subsystem. The process includes: determining an average memory latency corresponding to execution of the application; and determining an average occupancy of a miss status handling register queue associated with execution of the application based on the memory bandwidth and the average memory latency. The process includes: generating data representing a recommendation for an optimization to be applied to the application based on the average occupancy of the miss status handling register queue and a capacity of the miss status handling register queue.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an apparatus and method for optimizing application execution based on memory - level parallelism (MLP) metrics. Background Art

[0002] When executing a particular application, a hardware processor (e.g., a central processing unit (CPU) package or “socket”) may experience one or more performance bottlenecks. Due to the increasing complexity of modern hardware processor architectures (which have features such as multiple - instruction issue and out - of - order execution), there can be many potential causes of performance bottlenecks. By way of example, performance bottlenecks can include: instruction issue stalls due to the scheduler or re - order buffer (ROB) being fully filled; instruction issue stalls due to long - latency loads; fetch - related front - end pipeline stalls; issue - related back - end pipeline stalls; memory access problems; and so on. Summary of the Invention

[0003] An apparatus for optimizing application execution includes: a hardware processor; and a memory for storing instructions that, when executed by the hardware processor, cause the hardware processor to: determine a memory bandwidth of the processor subsystem corresponding to the execution of an application by the processor subsystem; determine an average memory latency corresponding to the execution of the application by the processor subsystem; determine a metric characterizing the memory - level parallelism associated with the execution of the application by the processor subsystem based on the memory bandwidth and the average memory latency; and generate data representing a recommendation for an optimization to be applied to the application based on the metric.

[0004] A method for optimizing application execution includes: determining, by a hardware processor, a memory bandwidth of a processor subsystem corresponding to the execution of an application by the processor subsystem; determining, by the hardware processor, an average memory latency corresponding to the execution of the application by the processor subsystem; determining, by the hardware processor, an average occupancy rate of an un - hit status handling register queue associated with the execution of the application by the processor subsystem; and generating, by the hardware processor, data representing a recommendation for an optimization to be applied to the application based on the average occupancy rate and the capacity of the un - hit status handling register queue.

[0005] A non-transitory storage medium for storing machine-readable instructions that, when executed by a machine, cause the machine to perform the following operations: determine an average miss status handling register (MSHR) queue occupancy rate associated with executing an application; based on a primary memory access type associated with executing the application, designate a given MSHR queue among a plurality of MSHR queues as limiting execution performance; determine the capacity of the given MSHR queue; and generate data for a graphical user interface (GUI) representing a selection of an optimization for the application from among a plurality of candidate optimizations based on a comparison of the average occupancy rate with the capacity of the given MSHR queue. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Figure 1 is a block diagram of a computer system including a performance analyzer based on memory-level parallelism (MLP) metrics according to an example embodiment.

[0007] Figure 2 is according to an example embodiment Figure 1 of a performance analyzer.

[0008] Figure 3 is according to an example embodiment depicting a process performed by Figure 1 the performance analyzer to provide data recommending optimizations for an application.

[0009] Figure 4 is a block diagram of an apparatus for generating data representing recommendations for optimizations to be applied to an application based on metrics characterizing memory-level parallelism associated with executing the application according to an example embodiment.

[0010] Figure 5 is a flowchart of a process according to an example embodiment depicting generating data representing recommendations for optimizations to be applied to an application based on a determined average occupancy rate of a miss status handling register queue (MSHR) queue and the capacity of the MSHR queue.

[0011] Figure 6 is an illustration of machine-readable instructions stored on a non-transitory storage medium that, when executed by a machine, cause the machine to generate data for a graphical user interface (GUI) representing a selection of an optimization for an application based on a comparison of the average occupancy rate of an MSHR queue with the capacity of the MSHR queue. DETAILED DESCRIPTION

[0012] Because there are so many potential causes of the performance bottleneck of a hardware processor, determining the cause of a specific performance bottleneck can be a daunting task. In this context, a "hardware processor" refers to the actual physical components that include one or more processing cores (e.g., CPU cores) that execute machine-readable instructions (i.e., "software"). According to an example embodiment, the hardware processor can be a multi-core CPU semiconductor package (or "socket"). A "performance bottleneck" generally refers to a condition (e.g., the average queue occupancy reaches or approaches full capacity) associated with a component (e.g., a queue) of the hardware processor, which imposes a limit or constraint on the processor's ability to execute at a higher level. For the purpose of resolving a specific performance bottleneck to enhance the performance of the processor, the application executed on the hardware processor can be changed or optimized.

[0013] For the purpose of determining the cause(s) of a performance bottleneck, a user (e.g., a software developer) can use one or more performance evaluation tools for the purpose of visualizing the architecture of the hardware processor and more specifically the way the components of the hardware processor execute when executing a specific application. If the performance evaluation tool(s) happen to expose the correct set of performance counters of the hardware processor, the cause(s) of the performance bottleneck can be revealed to the user. A "performance counter" generally refers to a hardware counter built into the hardware processor that counts the occurrence of specific hardware events of the hardware processor (e.g., cache misses, cycles per instruction, stalls, etc.). However, due to the disconnection between the performance tool and the processor architecture and / or the disconnection between the performance tool and the user, the performance tool may not be able to reveal the cause(s) of the performance bottleneck to the user.

[0014] The disconnection between the performance tool and the processor architecture can be at least partially attributable to the complexity of modern processor architectures. Out-of-order execution in modern hardware processors is achieved at the cost of complex interactions between various processor structures, which complicate the interpretation of the performance counters of the processor and may not clearly describe the performance of the processor when executing a specific application. These challenges can be further exacerbated if a specific hardware processor does not expose the appropriate performance counter(s) to the performance evaluation tool to allow tracking of the cause(s) of the performance bottleneck.

[0015] The disconnection between the performance tool and the user can be at least partially attributable to the user's level of expertise. For non-expert users who do not have a sufficient understanding of the microarchitecture details of the hardware processor, the values reported by the processor's performance counters may not make sense. For expert users who have a sufficient understanding of the microarchitecture details of the processor, the performance counters may still be rather useless because the performance counters cannot guide the user to take specific actionable steps (e.g., optimization) to improve (or at least attempt to improve) the performance of the hardware processor.

[0016] According to an example embodiment described herein, when executing a specific sub - part of an application, a profiler (which may also be referred to as a "memory - level parallel (MLP) metric - based profiler" or an "MLP - metric - based profiler") can be used to analyze the performance of a hardware processor. Here, a "sub - part" of an application corresponds to a unit of machine - readable instructions that correspond to a selected portion of the application (such as a routine, sub - routine, or loop of the application). The MLP metric is a measure of the ability of the hardware processor to perform multiple memory operations simultaneously when executing a sub - part of the application. According to the example embodiment, the MLP metric represents the average occupancy rate of the miss - status handling register (MSHR) queue of the hardware processor when the application sub - part is executed (e.g., the average number of occupied registers in the MSHR queue).

[0017] The hardware processor may include multiple MSHR queues, and the multiple MSHR queues are respectively associated with corresponding caches of the hardware processor (e.g., a level - 1 (L1) cache and a level - 2 (L2) cache). Generally, an MSHR queue contains a set of registers, and when occupied, the registers represent outstanding memory requests due to cache misses in the cache associated with the MSHR queue. As an example, the registers in the MSHR queue can correspond to associated outstanding memory requests and contain information about the requests (such as the address of the requested block, whether the requested block corresponds to a read or a write, and the cycle when the requested block is ready).

[0018] According to an example embodiment, the MLP metric is generic in nature. In this way, the MLP metric is not theoretically associated with a particular MSHR queue of a hardware processor. However, based on the type of main memory access (e.g., streaming or random) that occurs when the hardware processor executes an application sub-part, the performance analyzer associates the MLP metric with a particular MSHR queue. According to an example embodiment, the performance analyzer associates the MLP metric with the MSHR queue corresponding to the L1 cache or the MSHR queue corresponding to the L2 cache based on the type of main memory access. For example, when the type of main memory access is random access, the performance analyzer associates the MLP metric with the MSHR queue of the L1 cache (since this MSHR queue is more likely to be a performance bottleneck). When the type of main memory access is streaming access, the performance analyzer associates the MLP metric with the MSHR queue of the L2 cache (since this MSHR queue is more likely to be a performance bottleneck). According to an example embodiment, the performance analyzer compares the average MSHR queue occupancy rate (represented by the MLP metric) with the total capacity of the MSHR queue with which the performance analyzer associates the MLP metric. According to an example embodiment, the performance analyzer generates data (e.g., data for a graphical user interface (GUI)) representing one or more recommended optimizations that can be applied to the application to enhance the processing performance of the processor based on the comparison.

[0019] Among the potential advantages of the performance analyzer, the performance analyzer can determine the MLP metric using relatively few performance counter values. The performance counter values can correspond to performance counters that are widely available for processors provided by different vendors. The performance analyzer can also be beneficial for users with limited knowledge of the microarchitecture details of the hardware processor, since the MLP metric abstracts the out-of-order execution details from the user. In this way, the MLP metric can be directly related to a particular MSHR queue associated with a particular cache, and thus, the user can deal with a single understandable structure of the hardware processor.

[0020] Reference Figure 1 , as a more specific example, according to some embodiments, computer system 100 includes one or more nodes 101 ( Figure 1 the N example nodes 101-1 to 101-N depicted) that can be interconnected by a network fabric 148. According to some embodiments, a given node 101 can correspond to computer platform 100. According to some embodiments, computer system 100 can be a cluster computer system, and nodes 101 can include the compute nodes of the cluster and possibly other nodes (such as management nodes, storage nodes, etc.). According to further embodiments, computer system 100 can not be a cluster computer system. Figure 1 Details of a particular node 101-1 described herein are depicted according to an example embodiment.

[0021] According to an example embodiment, node 101 can be a modular unit including a frame or chassis. Additionally, the modular unit can include hardware mounted on the chassis and capable of executing machine-executable instructions. According to an example embodiment, a blade server is an example of node 101. However, according to a further embodiment, node 101 can be any number of different platforms other than a blade server, such as a rack server, a stand-alone server, a client, a desktop computer, a smart phone, a wearable computer, a network component, a gateway, a network switch, a storage array, a portable electronic device, a portable computer, a tablet computer, a thin client, a laptop computer, a television, a modular switch, a consumer electronic device, an appliance, an edge processing system, a sensor system, a watch, a removable peripheral device card, etc.

[0022] It should be noted that, according to one of many possible embodiments, Figure 1 the depicted architecture of node 101-1 is one of many possible architectures of node 101. Additionally, according to a further example embodiment, node 101-1 can be a stand-alone node (i.e., not part of a computer system 100 having multiple nodes 101 as depicted in Figure 1 ). Other nodes 101 of computer system 100 may or may not have an architecture similar to that of node 101-1. Thus, many embodiments within the scope of the appended claims are contemplated.

[0023] The network fabric 148 can be associated with one or more types of communication networks, such as (by way of example) a Fibre Channel network, a Gen-Z fabric, a dedicated management network, a local area network (LAN), a wide area network (WAN), a global network (e.g., the Internet), a wireless network, or any combination thereof.

[0024] According to an example embodiment, node 101-1 can include one or more hardware processors 104. In this case, a "hardware processor" refers to an actual physical device or component having one or more processing cores 120 that execute machine-readable instructions (or "software"). As a specific example, according to some embodiments, the hardware processor 104 can be a multi-core CPU semiconductor package or "socket" containing multiple CPU processing cores 120.

[0025] The hardware processor 104 can include one or more level 1 (L1) caches 114. According to an example embodiment, each processing core 120 can have its own dedicated L1 cache 114, and according to a further example embodiment, multiple processing cores 120 (e.g., two or more processing cores 120) can share the L1 cache 114. Additionally, as also in Figure 1Depicted, according to an example embodiment, the hardware processor 104 may include one or more level 2 (L2) caches 118. According to some embodiments, each processing core 120 may have its own dedicated L2 cache 118, and according to further embodiments, multiple processing cores 120 (e.g., two or more processing cores 120) may share an L2 cache. It should be noted that, according to an example embodiment, the hardware processor 104 may include higher-level caches, such as one or more level 3 (L3) caches 119.

[0026] According to some embodiments, the L1 cache 114 may have a relatively small size (in terms of memory capacity) and may be formed of a memory having an associated, relatively fast response time. For example, according to some embodiments, the L1 cache 114 may be formed of static random access memory (SRAM) devices. According to an example embodiment, the L2 cache 118 may be a relatively large memory (as compared to the capacity of the L1 cache 114) that may be formed, for example, of dynamic random access memory (DRAM) devices. The L1 cache 114, the L2 cache 118, and other memories described herein are generally non-transitory storage media, which may be formed of non-transitory memory devices, such as semiconductor storage devices, flash memory devices, memristors, phase change memory devices, combinations of devices formed of one or more of the foregoing storage technologies, and the like. Additionally, unless otherwise specified herein, a memory device may be a volatile memory device (e.g., DRAM devices, SRAM devices, etc.) or a non-volatile memory device (e.g., flash memory devices, read-only memory (ROM) devices, etc.).

[0027] According to an example embodiment, the hardware processor 104 includes a miss status handling register (MSHR) queue 115 dedicated to each L1 cache 114 and an MSHR queue 115 dedicated to each L2 cache 118 (i.e., according to an example embodiment, each MSHR queue 115 has a corresponding associated L1 cache 114 or L2 cache 118). Additionally, as Figure 1 Depicted, according to an example embodiment, the hardware processor 104 may include one or more performance counters 116. The performance counters 116 count different events occurring in the hardware processor 104. As an example, a particular performance counter 116 may reveal an L3 miss count, which may be used, as further described herein, for the purpose of evaluating memory bandwidth utilization.

[0028] According to an example embodiment, node 101-1 includes a performance analyzer 170 based on MLP metrics (referred to herein as "performance analyzer 170"), which can generally be used to analyze the execution of application 130 (or a selected sub-part of application 130) by the processor subsystem 102 of node 101-1. According to an example embodiment, processor subsystem 102 includes one or more hardware processors 104 and the system memory 140 of node 101-1. According to an example embodiment, application 130 (or a selected sub-part of application 130) can be executed simultaneously on one or more processing cores 120 of a specific hardware processor 104 of node 101-1. Additionally, according to an example embodiment, application 130 (or a selected sub-part of application 130) can be executed simultaneously on multiple hardware processors 104 of node 101-1. According to an example embodiment, performance analyzer 170 can be used to target the execution of a specific sub-part of application 130. In this way, the target sub-part can be machine-executable instructions (i.e., program code or "software") that correspond to a specific routine, subroutine, or loop of application 130, and the specific routine, subroutine, or loop has been specified by the user of performance analyzer 170 for analysis by performance analyzer 170.

[0029] As further described herein, performance analyzer 170 calculates an MLP metric that represents a measure of the MLP of processor subsystem 102 when executing a selected sub-part of application 130. According to an example embodiment, the MLP also represents the calculated average MSHR queue occupancy. According to an example embodiment, performance analyzer 170 selects the MSHR queue 115 associated with L1 cache 114 or the MSHR queue 115 associated with L2 cache 118 based on the type of main memory access (e.g., streaming or random) that occurs during the execution of the sub-part of application 130. Through this selection, performance analyzer 170 designates the selected MSHR queue 115 as a potential performance bottleneck, i.e., performance analyzer 170 determines that the selected MSHR queue 115 is most likely to affect the performance of processor subsystem 102 when executing the sub-part of application 130. It should be noted that, according to an example embodiment, selecting the MSHR queue 115 is selecting the MSHR queue type, e.g., selecting the MSHR queue 115 associated with L1 cache or the MSHR queue 115 associated with L2 cache. According to an example embodiment, all MSHR queues 115 associated with L1 cache have the same size or capacity (i.e., the same number of registers), and all MSHR queues 115 associated with L2 cache have the same capacity (i.e., the same number of registers).

[0030] According to an example embodiment, the performance analyzer 170 compares the average MSHR queue occupancy rate (represented by the MLP metric) with the size or capacity (e.g., the number of registers) of the selected MSHR queue 115. According to an example embodiment, based on this comparison, the performance analyzer 170 selects one or more optimizations for the application 130. Generally, an "optimization" of an application is a change to be applied to the application 130 for the purpose of improving the processor execution performance for the subpart of the application being analyzed.

[0031] According to an example embodiment, the performance analyzer 170 provides data to the graphical user interface (GUI) 172, which causes the GUI 172 to display the recommended optimization(s). The performance analyzer 170 can further provide data to the GUI 172, which causes the GUI 172 to display an analysis that characterizes the execution of a subpart of the application 130. These analyses can include one or more values of the performance counters 116, MLP metric values, a determination of the cache type associated with the selected MSHR queue 115, the capacity of the selected MSHR queue 115, one or more performance metrics derived from the value(s) of the performance counters 116, and so on. According to an example embodiment, the GUI 172 can receive user input. For example, according to some embodiments, the user can provide to the GUI 172 via one or more input / output devices (e.g., keyboard, touch screen, mouse, touchpad, etc.): an input representing the selection of a subpart of the application 130 for analysis; an input representing parameters for controlling the analysis of the performance analyzer 170; an input representing control buttons and options of the GUI 172; an input used by the performance analyzer to determine the MLP metric (e.g., inputs such as memory bandwidth, cache line size, average latency, type of main memory access associated with the application subpart, bandwidth-versus-latency graph of the hardware processor 104, performance counter values, etc.); and so on.

[0032] According to an example embodiment, the performance analyzer 170 is a software entity hosted on the node 101-1 and provided by one or more processing cores 120 of one or more hardware processors 104 of the node 101-1 that execute machine-readable instructions, while one or more processing cores 120 of one or more hardware processors 104 of the node 101-1 execute the subpart of the application being analyzed. According to an example embodiment, the machine-readable instructions 142 corresponding to the performance analyzer 170 can be stored in the system memory 140. In addition, the machine-readable instructions corresponding to the application 130 can be stored in the system memory 140. Similarly, as Figure 1Depicted, according to some embodiments, memory 140 may further store data 144. Data 144 may include data associated with performance analyzer 170 and / or GUI 172, such as inputs for performance analyzer 170, inputs for GUI 172, control parameters for performance analyzer 170, outputs for performance analyzer 170, outputs for GUI 172, intermediate values obtained by performance analyzer 170 as part of its analysis and recommendation process, and the like. System memory 140 may further store data related to application 130. Although Figure 1 Performance analyzer 170 is depicted as being located on the same node 101-1 as the application 130 being evaluated, according to further embodiments, performance analyzer 170 may be located on another node 101 other than the node 101 executing application 130. In a similar manner, according to further embodiments, GUI 172 may not be located on the same node 101 as application 130. Additionally, according to further embodiments, GUI 172 and performance analyzer 170 may be located on different nodes 101.

[0033] According to further embodiments, all or a portion of performance analyzer 170 may be formed of special-purpose hardware that does not execute machine-readable instructions. For example, according to further embodiments, all or a portion of performance analyzer 170 may be formed of an application specific integrated circuit (ASIC), complex logic device (CLD), field programmable gate array (FPGA), and the like.

[0034] Also as Figure 1 Depicted, according to some embodiments, node 101 may include one or more performance evaluation tools 117. As an example, according to some embodiments, a particular performance evaluation tool 117 may provide a latency versus bandwidth utilization graph. As another example, according to some embodiments, a particular performance evaluation tool 117 may provide an average memory latency based on the bandwidth utilization provided as an input to the performance evaluation tool 117. According to some embodiments, as another example, a particular performance evaluation tool 117 may expose particular performance counters 116 for the purpose of determining bandwidth utilization. According to further embodiments, other performance evaluation tools 117 may be used in conjunction with performance analyzer 170. According to some embodiments, performance analyzer 170 may interface directly with one or more performance evaluation tools 117. Additionally, according to some embodiments, GUI 172 may interface directly with one or more performance evaluation tools 117.

[0035] According to an example embodiment, out-of-order execution performed by the hardware processor 104 relies on parallel execution of multiple operations and memory requests. At any given time, the hardware processor 104 uses the MSHR queues 115 associated with the caches 114, 118 to track all unique memory requests that miss the L1 cache 114 or the L2 cache 118 (at the cache line granularity). This tracking thus avoids duplicate memory requests. According to an example embodiment, the hardware processor 104 includes one or more hardware prefetchers (not shown). As an example, the hardware processor 104 may include hardware prefetchers for the L1 cache 114 and the L2 cache 118, which, when triggered, issue prefetch requests at their respective caches 114 and 118.

[0036] Depending on the main memory access type associated with the execution of the application subpart, the MSHR queue 115 corresponding to the L1 cache 114 or the MSHR queue 115 corresponding to the L2 cache 118 may create a performance bottleneck. Whether the MSHR queue 115 causes a performance bottleneck may depend on two factors: 1. the size of the MSHR queue 115; and 2. the nature of the application subpart. According to an example embodiment, for the purpose of meeting the L1 cache access timing constraints, the size of the MSHR queue 115 associated with the L1 cache 114 is kept relatively small (e.g., compared to the size of the MSHR queue 115 corresponding to the L2 cache 118). In this way, the L1 cache access timing constraints may specify that all entities in the MSHR queue 115 are searched simultaneously for each memory request. According to an example embodiment, the size of the MSHR queue 115 corresponding to the L2 cache 118 may be much larger than the size of the MSHR queue 115 corresponding to the L1 cache 114.

[0037] According to an example embodiment, the performance analyzer 170 uses the type of main memory access associated with the execution of a particular application sub - part as an indicator of which type of MSHR queue 115 (e.g., the L1 - cache - associated MSHR queue 115 or the L2 - cache - associated MSHR queue 115) might be a potential cause of a performance bottleneck. In this case, the "type" of memory access refers to whether the memory access is a streaming access or a random access. A "streaming memory access" refers to a memory access to a predictable address in memory (e.g., a virtual address) (e.g., an access that coincides with the same cache line or the same set of cache lines, an access to the same memory page or the same set of memory pages), such that the hardware processor 104 can predict future memory accesses based on a particular pattern of previous memory accesses. A "random memory access" refers to a memory access that does not follow a particular pattern, and as such, the hardware processor 104 may not be able to accurately predict future memory accesses based on previous memory accesses. The "dominant" type of memory access refers to the type of memory access that is more prevalent or greater in quantity than the other type of memory access. Thus, if the execution of a given application sub - part results in more random accesses to memory than streaming accesses to memory, the execution of the application sub - part is predominantly associated with random accesses. Conversely, if the execution of a given application sub - part results in more streaming accesses to memory than random accesses to memory, the execution of the application sub - part is predominantly associated with streaming accesses.

[0038] If the execution of an application sub - part does not trigger the hardware prefetcher of the L2 cache (as in the case of random memory accesses), the average occupancy of the MSHR queue 115 corresponding to the L2 cache 114 may not be greater than the average occupancy of the MSHR queue 115 corresponding to the L1 cache 118. Thus, according to an example embodiment, the performance analyzer 170 concludes that for an application sub - part predominantly associated with random memory accesses, the MSHR queue 115 corresponding to the L1 cache 114 is a potential cause of MLP limitation. Additionally, according to an example embodiment, the performance analyzer 170 concludes that for an application sub - part predominantly associated with streaming memory accesses that benefit from the L2 - cache hardware prefetcher, the MSHR queue 115 associated with the L2 cache is a potential cause of MLP limitation.

[0039] According to an example embodiment, the performance analyzer 170 can determine the average MSHR queue occupancy (referred to herein as "n" 平均 ") or the MLP metric based on Little's Law. Little's Law states that the average number of customers in a steady - state system is equal to the long - term average effective arrival rate multiplied by the average time a customer spends in the system. Since Little's Law assumes a steady - state system, thus according to an example embodiment, the average MSHR queue occupancy n 平均is determined for an application sub - part (e.g., a separate routine, sub - routine, loop, etc. of application 130). Applying Little's Law, the average MSHR queue occupancy rate n of a given application sub - part (e.g., a routine, sub - routine, or loop of application 130) 平均 can be described as the long - term average memory request arrival rate (i.e., the rate at which requests enter MSHR queue 115) multiplied by the average memory latency (i.e., the average time a request stays in MSHR queue 115). The long - term average memory request arrival rate is the total number of memory requests (referred to herein as "R") during the execution of the application sub - part divided by the total time of execution of the application sub - part (referred to herein as "T"). Thus, the average occupancy rate n 平均 or MLP can be described as follows:

[0040]

[0041] where "lat" 平均 represents the average memory latency. The memory bandwidth utilization or observed memory bandwidth (referred to herein as "BW") during the execution of the application sub - part can be described as follows:

[0042]

[0043] where "cls" represents the cache line size. Using Equation 2, Equation 1 can be rewritten as follows:

[0044]

[0045] It should be noted that the average memory latency lat 平均 refers to the memory latency observed in the hardware processor 104 at a specific BW memory bandwidth (and not, for example, the idle latency). Typically, the observed latency increases as the bandwidth utilization increases and can be two times or more the idle latency at peak bandwidth utilization. According to an example embodiment, the performance analyzer 170 can obtain the BW memory bandwidth indirectly (e.g., via the L3 cache miss count provided by the performance counter 116 of the x86 - based processing core 120) or directly (e.g., via the memory read / write count provided by the performance counter 116 of the ARM - based processing core 120). Using, for example, a bandwidth - latency relationship graph of the hardware processor 104, the performance analyzer 170 can use the determined BW memory bandwidth to determine the average memory latency lat 平均 . The bandwidth - latency relationship graph of the hardware processor 104 can be calculated once using, for example, the performance evaluation tool 117.

[0046] Figure 2 depicts a block diagram of the performance analyzer 170 according to an example embodiment. In combination with Figure 1 referenceFigure 2 , according to the example embodiment, the performance analyzer 170 includes an MLP metric determination engine 220 and a recommendation engine 230. As Figure 2 depicted, according to the example embodiment, the MLP metric determination engine 220 may receive data representing the average memory latency 206, memory bandwidth 208, cache line size 212, and core frequency 214 as inputs. Then, the MLP metric determination engine 220 may generate data representing the MLP metric 224 based on these inputs.

[0047] The recommendation engine 230 of the performance analyzer 170 may provide recommendation data 250 representing recommendations for one or more optimizations for the application 130 that are specific selections of the recommendation engine 230 based on data 234 identifying the primary memory access type and the MLP metric 224. According to the example embodiment, the recommendation data 250 may be configured to cause the GUI 172 to display the recommended optimization(s).

[0048] According to the example embodiment, the recommendation engine 230 may consider any of a plurality of candidate optimizations. For example, one candidate optimization is vectorization, in which a single operation is applied to multiple operands. In addition to thread-level parallelism, vectorization provides another level of parallelism and can therefore be quite effective in improving the MLP. Vectorization can be particularly helpful in improving the MLP on processors with high-bandwidth memory (HBM). The degree of parallelism (vector width) and coverage (using gather / scatter, prediction, etc.) obtained through vectorization are also increasing in more modern processors, making it more widely applicable than before. Since vectorization improves the MLP, it also increases the average MSHR queue occupancy. Therefore, if the average occupancy of the MSHR queue 115 of the application is close to the capacity of the MSHR queue 115, the application 130 may not benefit from vectorization. Otherwise, according to the example embodiment, the recommendation engine 230 may recommend the vectorization optimization.

[0049] Software prefetching is another example of a candidate optimization. In this optimization, the user or compiler inserts software prefetch instructions in the source code for the purpose of prefetching data into a cache at a specific level. Prefetching can be particularly useful for certain irregular access patterns because the hardware prefetcher may not recognize these patterns, or the hardware prefetcher may not recognize these patterns in a timely manner. Each software prefetch request occupies the MSHR queue 115, and the MSHR queue rejects another demand load request or rejects the hardware prefetcher from obtaining the MSHR queue 115. Therefore, when the average occupancy rate of the MSHR queue 115 of a program code unit of an application is relatively high, the program code unit may not benefit from software prefetch optimization. When a program code unit is associated with a relatively high degree of random access to memory, the recommendation engine 230 may recommend software prefetch optimization for the program code unit. For random access, software prefetch optimization may result in the use of the MSHR queue 115 associated with the L2 cache, otherwise, when the hardware prefetcher for the L2 cache is ineffective, the queue is not used.

[0050] Loop tiling is another example of a candidate optimization. Loop tiling divides the iteration space of an application loop into smaller chunks or tiles such that the data accessed within these smaller tiles remains in the cache until it is reused. Loop tiling can target reuse in different levels of the memory hierarchy. According to an example embodiment, the recommendation engine 230 may recommend loop tiling in response to a relatively high average occupancy rate of the MSHR queue 115 experienced by a sub-part of the application 130 because loop tiling reduces the number of memory requests and thus reduces the occupancy rate of the MSHR queue 115.

[0051] Register tiling (or "unroll and jam" optimization) is another example of a candidate optimization. Register tiling is similar to loop tiling, except that register tiling targets data reuse in registers (as opposed to cache reuse). Since there are few memory accesses (i.e., most data fits in a higher-level cache), register tiling can be particularly beneficial when memory accesses have already experienced short latency. A low MSHR queue 115 occupancy rate can be used to infer short latency and, therefore, can be used as a metric for the recommendation engine 230 to recommend register tiling.

[0052] Another candidate optimization is loop fusion optimization. Loop fusion combines the bodies of different loops or loop nests, and therefore, loop fusion can significantly shorten the reuse distance for certain memory accesses. Like loop tiling, loop fusion is particularly useful in reducing the occupancy rate of the MSHR queue 115 because loop fusion promotes data reuse. Therefore, according to an example embodiment, the recommendation engine 230 may recommend loop fusion optimization for a relatively high MSHR queue occupancy rate.

[0053] Another candidate optimization is loop distribution optimization. Loop distribution is the exact opposite of loop fusion. Loop distribution is an optimization that enables loop fusion or vectorization (such as loop interchange). Loop distribution is expected to improve performance when used alone when the distributed loops can reduce the number of active streams or memory bandwidth contention. Thus, according to an example embodiment, the recommendation engine 230 may recommend loop distribution optimization for relatively high MLP metrics and corresponding relatively high average occupancy rates of the MSHR queue 115.

[0054] According to an example embodiment, the performance analyzer 170 may recommend simultaneous multithreading (SMT) or hyperthreading (HT). These are not optimizations but different ways of executing the application 130, which involve using the simultaneous multithreading or hyperthreading capabilities of the hardware processor 104. SMT can be very beneficial for a hardware processor 104 with HBM because SMT can significantly improve the MLP. Threads on the processing core 120 participating in SMT share most of the core resources (including the MSHR queue 115), and the occupancy rate of the MSHR queue 115 directly contributes to understanding the benefits of SMT. A nearly full MSHR queue 115 means there are not enough resources in the processing core 120 to accommodate more threads. Thus, according to an example embodiment, the recommendation engine 230 recommends SMT for all applications 130, except for applications 130 with a high MSHR queue occupancy rate and except for special cases such as cache residency contention between threads.

[0055] Figure 3 Depicts an example process 300 that may be performed by the performance analyzer 170 according to an example embodiment. In conjunction with Figure 1 and Figure 2 reference Figure 3 According to an example embodiment, blocks 304, 308, and 312 may be performed by the MLP metric determination engine 220, and blocks 316 to 348 may be performed by the recommendation engine 230.

[0056] According to block 304, the performance analyzer 170 determines the memory bandwidth. As an example, the performance analyzer 170 may make this determination based on the appropriate performance counter(s) 116, may obtain the memory bandwidth via data provided through the GUI 172, may use the output of the performance evaluation tool 117 to obtain the memory bandwidth, etc. Next, according to an example embodiment, the performance analyzer 170 determines (block 308) the average memory latency. According to an example embodiment, the performance analyzer 170 may infer the average memory latency from the observed bandwidth based on the number of observed load latencies of the processor 104. For this purpose, one or more performance evaluation tools 117 may be used, a bandwidth-versus-latency graph of the processor 104 may be used, or the user may provide an input specifying the average memory latency via the GUI 172, etc.

[0057] According to an example embodiment, process 300 then includes determining (block 312) an MLP metric using equation 3 above. Contention found in the MSHR queue 115 can be associated with the L1 cache 114 or the L2 cache 118. According to an example embodiment, determining a particular MSHR queue type (e.g., L1 cache-associated or L2 cache-associated) lies in the application sub-part under discussion. In this way, if the execution of the application sub-part is dominated by random memory accesses (e.g., the hardware prefetcher is largely ineffective), then the MSHR queue 115 associated with the L1 cache 114 is a source of potential bottlenecks. Otherwise, the MSHR queue 115 associated with the L2 cache 118 is the source of the bottleneck.

[0058] Determining the main memory access type is performed in decision block 316 of process 300 and involves transitioning to block 320 for a random access-dominated case or to block 340 for a stream access-dominated case. According to an example embodiment, the decision in decision block 316 can be made as a result of an input to the performance analyzer 170 (e.g., an input provided by the user via the GUI 172). According to a further embodiment, the performance analyzer 170 can perform decision block 316 by observing the ratio of memory requests generated by the hardware prefetcher to the demand load. For example, this data can be exposed through one or more performance counters 116, or alternatively, the memory access type can be exposed by the user disabling the hardware prefetcher. In the case of a mixed sequential and random memory access, such as in a sparse matrix-vector multiplication operation, the data structure generating the random memory accesses typically tends to dominate the memory traffic because each reference typically targets a different cache line, as opposed to different words on the same cache line.

[0059] With knowledge of the average occupancy of the MSHR queue 115 and the particular MSHR queue type as a potential bottleneck, the performance analyzer 170 can then proceed to block 340 (for a stream access-dominated case) or block 320 (for a random access-dominated case).

[0060] For a random access-dominated case, the performance analyzer 170 compares the average occupancy of the MSHR queue 115 (represented by the MLP metric) with the size or capacity of the MSHR queue 115 associated with the L1 cache (block 320). If the occupancy is less than the size, then according to block 324, the performance analyzer 170 can recommend vectorization, SMT, or L1 software prefetching. If the occupancy is nearly equal to the size of the MSHR queue, then according to block 328, the performance analyzer 170 can recommend L2 cache software prefetching, loop fusion, or loop tiling.

[0061] As an example, according to some embodiments, "substantially the same" or "substantially equal in size" may mean that the average MSHR queue occupancy rate is greater than or equal to a threshold representing a certain percentage (e.g., 90%) of the MSHR queue capacity. According to further embodiments, thresholds and / or techniques other than the capacity-based percentage threshold may be used to evaluate whether the average MSHR queue occupancy rate is "substantially equal to" the capacity of the MSHR queue 115. Regardless of how this is determined, according to example embodiments, if the average occupancy rate of the MSHR queue 115 is substantially the same as the capacity or size of the MSHR queue 115, the performance analyzer 170 recommends optimizations for reducing rather than increasing the average occupancy rate of the MSHR queue 115. If the performance analyzer 170 determines that the occupancy rate of the MSHR queue 115 is less than the size of the MSHR queue 115 (e.g., the occupancy rate is less than 90% of the size of the MSHR queue 115), then according to example embodiments, the performance analyzer 170 may consider all optimizations, including optimizations for increasing the MSHR queue occupancy rate or optimizations to which the MLP can be applied.

[0062] If the performance analyzer 170 determines that the average occupancy rate of the MSHR queue is greater than the size of the MSHR queue, the performance bottleneck may be the MSHR queue associated with the L2 cache. In this case, control transfers to block 340.

[0063] For predominantly streaming access (according to decision block 316), the performance analyzer 170 compares the average occupancy rate of the MSHR queue 115 with the size of the L2 cache MSHR queue 115 (block 340). If the occupancy rate is less than the size, then according to block 348, the performance analyzer 170 may recommend vectorization, SMT, or L1 cache software prefetching. If the occupancy rate is substantially equal to the size, then according to block 344, the performance analyzer 170 may recommend loop fusion or loop tiling.

[0064] It should be noted that the process 300 may be repeated to consider other optimizations, depending on the changes in the average MSHR queue occupancy rate and the observed performance due to the recommended optimizations being applied.

[0065] Reference Figure 4 , according to example embodiments, the apparatus 400 includes a memory 404 and a processor 414. The memory stores instructions 410. The processor 414 is configured to determine the memory bandwidth of the processor subsystem corresponding to the execution of an application by the processor subsystem and determine the average memory latency corresponding to the execution of the application by the processor subsystem. The processor 414 is configured to determine a metric characterizing the memory-level parallelism associated with the execution of the application by the processor subsystem based on the memory bandwidth and the average memory latency. Based on the metric, the processor 414 generates data representing recommendations for optimizations to be applied to the application.

[0066] Reference Figure 5 According to an example embodiment, process 500 includes: a hardware processor determining (block 504) the memory bandwidth of a processor subsystem corresponding to the execution of an application by the processor subsystem. According to block 508, process 500 includes: a hardware processor determining the average memory latency corresponding to the execution of an application by the processor subsystem. According to block 512, process 500 includes: a hardware processor determining the average occupancy of a miss status handling register queue associated with the execution of an application by the processor subsystem. According to block 516, process 500 includes: based on the average occupancy of the miss status handling register queue and the capacity of the miss status handling register queue, a hardware processor generating data representing a recommendation for an optimization to be applied to the application.

[0067] Reference Figure 6 According to an example embodiment, non-transitory storage medium 600 stores machine-readable instructions 604 that, when executed by a machine, cause the machine to: determine an average miss status handling register (MSHR) queue occupancy associated with the execution of an application; and based on a primary memory access type associated with the execution of the application, designate a given MSHR queue as limiting execution performance. Instructions 604, when executed by the machine, may cause the machine to determine the capacity of the given MSHR queue and generate data for a graphical user interface (GUI) representing a selection of an optimization for the application based on a comparison of the average MSHR queue occupancy with the capacity of the given MSHR queue.

[0068] According to an example embodiment, the instructions, when executed by a hardware processor, further cause the hardware processor to process data provided by at least one performance counter of the processor subsystem to determine the memory bandwidth. In certain advantages, metrics can be determined using a relatively small number of performance counter values; metrics can be determined for a wide range of hardware processors corresponding to a wide range of hardware processor vendors; and vectorization of a processor architecture corresponding to a performance bottleneck can be simplified.

[0069] According to an example embodiment, the instructions, when executed by a hardware processor, further cause the hardware processor to access data representing the memory bandwidth provided by a performance tool. In certain advantages, metrics can be determined using a relatively small number of performance counter values; metrics can be determined for a wide range of hardware processors corresponding to a wide range of hardware processor vendors; and vectorization of a processor architecture corresponding to a performance bottleneck can be simplified.

[0070] According to an example embodiment, the instructions, when executed by a hardware processor, further cause the hardware processor to determine an average memory latency based on the memory bandwidth and the bandwidth-to-latency relationship of the hardware processor. In certain advantages, a metric can be determined using a relatively small number of performance counter values; the metric can be determined for a wide range of hardware processors corresponding to a wide range of hardware processor vendors; and the vectorization of the processor architecture corresponding to a performance bottleneck can be simplified.

[0071] According to an example embodiment, the metric represents an average occupancy rate of an unhit status handling register queue associated with a cache of a processor subsystem. In certain advantages, a metric can be determined using a relatively small number of performance counter values; the metric can be determined for a wide range of hardware processors corresponding to a wide range of hardware processor vendors; and the vectorization of the processor architecture corresponding to a performance bottleneck can be simplified.

[0072] According to an example embodiment, the processor subsystem includes a level 1 (L1) cache, a level 2 (L2) cache, a first unhit status handling register (MSHR) queue associated with the L1 cache, and a second unhit status handling register (MSHR) queue associated with the L2 cache. The instructions, when executed by a hardware processor, further cause the hardware processor to: associate the metric with one of the first MSHR queue or the second MSHR queue; use the metric as an indication of the occupancy rate of the associated MSHR queue; compare the occupancy rate with the capacity of the associated MSHR queue; and select an optimization in response to the result of the comparison. In certain advantages, a metric can be determined using a relatively small number of performance counter values; the metric can be determined for a wide range of hardware processors corresponding to a wide range of hardware processor vendors; and the vectorization of the processor architecture corresponding to a performance bottleneck can be simplified.

[0073] According to an example embodiment, the instructions, when executed by a hardware processor, further cause the hardware processor to: determine whether the memory requests associated with the execution of an application are mainly stream accesses or mainly random accesses; and select an optimization in response to determining whether the memory requests are mainly stream accesses or mainly random accesses. In certain advantages, a metric can be determined using a relatively small number of performance counter values; the metric can be determined for a wide range of hardware processors corresponding to a wide range of hardware processor vendors; and the vectorization of the processor architecture corresponding to a performance bottleneck can be simplified.

[0074] According to an example embodiment, the instructions, when executed by a hardware processor, further cause the hardware processor to determine that a memory request associated with executing an application is dominated by streaming access. According to an example embodiment, the instructions, when executed by a hardware processor, further cause the hardware processor to, in response to determining that the memory request is dominated by streaming access: use a metric as an indication of an average occupancy rate of a miss status handling register (MSHR) queue associated with a level 2 (L2) cache of the processor subsystem; compare the capacity of the MSHR queue with the average occupancy rate; and select an optimization based on the result of the comparison. In a particular advantage, the metric can be determined using a relatively small number of performance counter values; the metric can be determined for a wide range of hardware processors corresponding to a wide range of hardware processor vendors; and the vectorization of a processor structure corresponding to a performance bottleneck can be simplified.

[0075] According to an example embodiment, the instructions, when executed by a hardware processor, further cause the hardware processor to: compare the average occupancy rate with a threshold derived from the capacity, where the threshold includes a boundary between when the MSHR queue is considered almost full and when the MSHR queue is considered not full; and select an optimization in response to the comparison of the average occupancy rate with the threshold. In a particular advantage, the metric can be determined using a relatively small number of performance counter values; the metric can be determined for a wide range of hardware processors corresponding to a wide range of hardware processor vendors; and the vectorization of a processor structure corresponding to a performance bottleneck can be simplified.

[0076] According to an example embodiment, the instructions, when executed by a hardware processor, further cause the hardware processor to generate data for display as a recommendation on a graphical user interface. In a particular advantage, the metric can be determined using a relatively small number of performance counter values; the metric can be determined for a wide range of hardware processors corresponding to a wide range of hardware processor vendors; and the vectorization of a processor structure corresponding to a performance bottleneck can be simplified.

[0077] According to an example embodiment, the instructions, when executed by a hardware processor, further cause the hardware processor to determine that a memory request associated with executing an application is dominated by random access. The instructions, when executed by a hardware processor, cause the hardware processor to, in response to determining that the memory request is dominated by random access: use a metric as an indication of a first average occupancy rate of a miss status handling register (MSHR) queue associated with a level 1 (L1) cache of the processor subsystem; compare the capacity of the MSHR queue with the first average occupancy rate; and select an optimization based on the result of the comparison. In a particular advantage, the metric can be determined using a relatively small number of performance counter values; the metric can be determined for a wide range of hardware processors corresponding to a wide range of hardware processor vendors; and the vectorization of a processor structure corresponding to a performance bottleneck can be simplified.

[0078] According to an example embodiment, the instruction, when executed by a hardware processor, further causes the hardware processor to compare the first average occupancy with a threshold derived from the capacity. The threshold defines a boundary between when the MSHR queue is considered to be almost full and when the MSHR queue is considered to be not full. The instruction, when executed by a hardware processor, further causes the hardware processor to select an optimization in response to the comparison of the first average occupancy with the threshold. In a particular advantage, the metric can be determined using a relatively small number of performance counter values; the metric can be determined for a wide range of hardware processors corresponding to a wide range of hardware processor vendors; and the vectorization of the processor architecture corresponding to the performance bottleneck can be simplified.

[0079] According to an example embodiment, the instruction, when executed by a hardware processor, further causes the hardware processor to, in response to the first average occupancy being greater than the capacity: use the metric as an indication of a second average occupancy of an MSHR queue associated with a level 2 (L2) cache of the processor subsystem; compare the capacity of another MSHR queue with the second average occupancy; and select an optimization based on the result of the comparison of the capacity of the another MSHR queue with the second average occupancy. In a particular advantage, the metric can be determined using a relatively small number of performance counter values; the metric can be determined for a wide range of hardware processors corresponding to a wide range of hardware processor vendors; and the vectorization of the processor architecture corresponding to the performance bottleneck can be simplified.

[0080] Although the present disclosure has been described with respect to a limited number of embodiments, those skilled in the art having the benefit of this disclosure will appreciate many modifications and variations of the present disclosure. The appended claims are intended to cover all such modifications and variations.

Claims

1. An apparatus for optimizing application execution, comprising: A hardware processor; And A memory for storing instructions that, when executed by the hardware processor, cause the hardware processor to perform the following operations: Determine the memory bandwidth of the processor subsystem corresponding to the execution of an application by the processor subsystem; Determine the average memory latency corresponding to the execution of the application by the processor subsystem; Based on the memory bandwidth and the average memory latency, determine a metric characterizing memory-level parallelism associated with the execution of the application by the processor subsystem; Based on the main memory access type associated with the execution of the application, designate a given MSHR queue among a plurality of MSHR queues as limiting execution performance; Determine the capacity of the given MSHR queue; And Generate data representing a recommendation for an optimization to be applied to the application based on the metric and the capacity of the given MSHR queue, Wherein the instructions, when executed by the hardware processor, further cause the hardware processor to perform the following operations: Determine whether the memory requests associated with the execution of the application are mainly stream accesses or mainly random accesses; and In response to determining whether the memory requests are mainly stream accesses or mainly random accesses, select the optimization.

2. The device according to claim 1, wherein, The instructions, when executed by the hardware processor, further cause the hardware processor to process data provided by at least one performance counter of the processor subsystem to determine the memory bandwidth.

3. The device according to claim 1, wherein The instructions, when executed by the hardware processor, further cause the hardware processor to access data representing the memory bandwidth provided by a performance tool.

4. The device according to claim 1, wherein, The instructions, when executed by the hardware processor, further cause the hardware processor to determine the average memory latency based on the memory bandwidth and the bandwidth-to-latency relationship of the hardware processor.

5. The apparatus according to claim 1, wherein The metric represents the average occupancy rate of the miss status handling register queue associated with the cache of the processor subsystem.

6. The apparatus according to claim 1, wherein: The processor subsystem includes a first-level L1 cache, a second-level L2 cache, a first miss status handling register (MSHR) queue associated with the L1 cache, and a second miss status handling register (MSHR) queue associated with the L2 cache; and The instructions, when executed by the hardware processor, further cause the hardware processor to perform the following operations: Associate the metric with one of the first MSHR queue or the second MSHR queue; Use the metric as an indication of the occupancy rate of the associated MSHR queue in the first MSHR queue or the second MSHR queue; Compare the occupancy rate with the capacity of the associated MSHR in the first MSHR or the second MSHR; And In response to the result of the comparison, select the optimization.

7. The device according to claim 1, wherein, The instructions, when executed by the hardware processor, further cause the hardware processor to perform the following operations: Determine that the memory requests associated with the execution of the application are mainly stream accesses; and In response to determining that the memory requests are mainly stream accesses: Use the metric as an indication of the average occupancy of a miss status handling register (MSHR) queue associated with a level 2 (L2) cache of the processor subsystem; Compare the capacity of the MSHR queue with the average occupancy; And Select the optimization based on the result of the comparison.

8. The device according to claim 7, wherein, When executed by the hardware processor, the instruction further causes the hardware processor to perform the following operations: Compare the average occupancy with a threshold obtained from the capacity, where the threshold includes a boundary between when the MSHR queue is considered almost full and when the MSHR queue is considered not full; and Select the optimization in response to the comparison of the average occupancy with the threshold.

9. The device according to claim 1, wherein When executed by the hardware processor, the instruction further causes the hardware processor to generate the data for displaying the recommendation on a graphical user interface.

10. The device according to claim 1, wherein, When executed by the hardware processor, the instruction further causes the hardware processor to perform the following operations: Determine that memory requests associated with executing the application are mainly random access; and In response to determining that the memory requests are mainly random access: Use the metric as an indication of a first average occupancy of a miss status handling register (MSHR) queue associated with a level 1 (L1) cache of the processor subsystem; Compare the capacity of the MSHR queue with the first average occupancy; And Select the optimization based on the result of the comparison.

11. The device according to claim 10, wherein, When executed by the hardware processor, the instruction further causes the hardware processor to perform the following operations: Compare the first average occupancy with a threshold obtained from the capacity, where the threshold defines a boundary between when the MSHR queue is considered almost full and when the MSHR queue is considered not full; and Select the optimization in response to the comparison of the first average occupancy with the threshold.

12. The device according to claim 10, wherein, When executed by the hardware processor, the instruction further causes the hardware processor to perform the following operations in response to the first average occupancy being greater than the capacity: Use the metric as an indication of a second average occupancy of an MSHR queue associated with a level 2 (L2) cache of the processor subsystem; Compare the capacity of another MSHR with the second average occupancy; And Select the optimization based on the result of the comparison of the capacity of the another MSHR queue with the second average occupancy.

13. A method for optimizing application execution, comprising: Determine, by a hardware processor, the memory bandwidth of the processor subsystem corresponding to the processor subsystem executing an application; Determine, by the hardware processor, the average memory latency corresponding to the processor subsystem executing the application; Determine, by the hardware processor, the average occupancy of a miss status handling register queue associated with the processor subsystem executing the application; Designate a given MSHR queue among a plurality of MSHR queues as limiting execution performance based on the main memory access type associated with executing the application; Determine the capacity of the given MSHR queue; And Based on the average occupancy of the miss status handling register queue and the capacity of the given MSHR queue, the hardware processor generates data representing a recommendation for an optimization to be applied to the application. Wherein, the method further includes: Determining whether the memory requests associated with executing the application are mainly stream accesses or mainly random accesses; and Selecting the optimization in response to determining whether the memory requests are mainly stream accesses or mainly random accesses.

14. The method according to claim 13, wherein, Determining the memory bandwidth includes determining the memory bandwidth corresponding to a sub - part of the application.

15. The method according to claim 14, wherein, The sub - part includes routines, sub - routines, or loops of the application.

16. A non - transitory storage medium for storing machine - readable instructions that, when executed by a machine, cause the machine to perform the following operations: Determine the average occupancy of the miss status handling register (MSHR) queue associated with executing an application; Based on the main memory access type associated with executing the application, designate a given MSHR queue among a plurality of MSHR queues as limiting execution performance; Determine the capacity of the given MSHR queue; and Generate data for a graphical user interface (GUI) representing the selection of an optimization for the application from a plurality of candidate optimizations based on a comparison of the average occupancy with the capacity of the given MSHR queue. Among them, When executed by the machine, the instructions further cause the machine to perform the following operations: Determine whether the memory requests associated with executing the application are mainly stream accesses or mainly random accesses; and Select the optimization in response to determining whether the memory requests are mainly stream accesses or mainly random accesses.

17. The non-transitory storage medium according to claim 16, wherein, When executed by the machine, the instructions further cause the machine to perform the following operations: In response to the given MSHR queue being associated with a level 1 (L1) cache and the main memory access type associated with the application being random access, designate the given MSHR queue.

18. The non-transitory storage medium according to claim 16, wherein, When executed by the machine, the instructions further cause the machine to perform the following operations: In response to the given MSHR queue being associated with a level 2 (L2) cache and the main memory access type associated with the application being stream access, designate the given MSHR queue.

19. The non-transitory storage medium according to claim 16, wherein, When executed by the machine, the instructions further cause the machine to perform the following operations: Based on the average memory latency associated with executing the application, the bandwidth associated with executing the application, and the cache line size, determine the average occupancy of the miss status handling register (MSHR) queue.

Citation Information

Patent Citations

  • Methods and apparatus for reducing memory latency in a software application

    CN1890635A