Resource provisioning

By using resource models to predict GPU performance constraints and determine resource allocation, the problem of poor resource allocation in the existing technology is solved, and more efficient GPU resource utilization and query performance improvement is achieved.

CN119948461APending Publication Date: 2025-05-06MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202380068159.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-30
Filing Date
2023-08-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively select the correct GPU resource allocation for workloads, resulting in poor query performance and lack of simple methods for resource optimization.

Method used

Using a resource model, such as a roofline model, the processing unit predicts performance constraints for the workload, determines resource allocation based on this, and instructs the processing unit to allocate resources to optimize the execution of the workload.

Benefits of technology

Through automated resource allocation, overall GPU utilization and workload performance are improved, such as a significant reduction in query execution time and throughput improvements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948461A_ABST
    Figure CN119948461A_ABST
Patent Text Reader

Abstract

A system for provisioning resources of a processing unit. The system predicts a performance impact on the workload attributable to performance constraints of the processing unit for the workload according to a resource model, wherein the workload comprises a query and the resource model characterizes the available computing bandwidth, available memory bandwidth, and computing strength based on a peak computing bandwidth and a peak memory bandwidth of the processing unit. The system determines a resource allocation to the processing unit based on the predicted performance impact, and instructs the processing unit to allocate resources for processing the workload based on the determined resource allocation.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Graphics processing units (GPUs), which have the potential for massively parallel computing, high-bandwidth memory access capabilities, and are relatively easy to program as accelerators, have seen increasing interest in their use for accelerating data analytics, with several GPU database systems developed in recent years in both academic and industrial settings. Summary of the invention

[0002] In some aspects, the technology described herein relates to a method of provisioning resources for a processing unit, the method comprising: predicting, based on a resource model, a performance impact on a workload attributable to performance constraints of the processing unit for the workload, wherein the workload comprises queries and the resource model characterizes available compute bandwidth, available memory bandwidth, and operational intensity based on a peak compute bandwidth and a peak memory bandwidth of the processing unit; determining a resource allocation to the processing unit based on the predicted performance impact; and instructing the processing unit to allocate resources for processing the workload based on the determined resource allocation.

[0003] In some aspects, the technology described herein relates to a system for provisioning resources for a processing unit, the system comprising: one or more hardware processors; a performance analyzer executable by the one or more hardware processors and configured to predict a performance impact on a workload attributable to a performance constraint of the processing unit for the workload based on a resource model, wherein the workload comprises a query and the resource model characterizes available compute bandwidth, available memory bandwidth, and operational intensity based on a peak compute bandwidth and a peak memory bandwidth of the processing unit; and a resource manager executable by the one or more hardware processors and configured to determine a resource allocation to the processing unit based on the predicted performance impact, and instruct the processing unit to allocate resources for processing the workload based on the determined resource allocation.

[0004] In some aspects, the technology described herein relates to one or more tangible processor-readable storage media embodied with instructions for executing on one or more processors and circuits of a computing device for provisioning resources for a processing unit, the process comprising: predicting a performance impact on a workload attributable to performance constraints of the processing unit for the workload based on a resource model, wherein the workload includes queries and the resource model characterizes available compute bandwidth, available memory bandwidth, and operational intensity based on a peak compute bandwidth and a peak memory bandwidth of the processing unit; determining a resource allocation to the processing unit based on the predicted performance impact; and instructing the processing unit to allocate resources for processing the workload based on the determined resource allocation.

[0005] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

[0006] Other implementations are also described and described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 An example resource provisioning system in computing device 102 is shown.

[0008] Figure 2 An example roofline model applied to resource allocation in a processing unit is shown.

[0009] Figure 3 Performance modeling results from the DRAM roofline model in four example computing systems are shown.

[0010] Figure 4 Performance modeling results in the L2 roofline model for four example computing systems are shown.

[0011] Figure 5 An example memory hierarchy, an example multi-instance GPU, and an example multi-process service that may be employed in executing a workload are shown.

[0012] Figure 6 A scenario using representative queries from Crystal (Q11, Q31) and a full-DRAM to half-DRAM allocation change is shown.

[0013] Figure 7 A scenario showing the bandwidth (ie, throughput) and AI (2886.74 Gops / sec and 27.15 ops / byte, respectively) achievable for query Q34.

[0014] Figure 8 Example operations for provisioning resources of a processing unit are shown.

[0015] Fig. 9 Various example comparisons of query execution times for different queries across multiple scale factors and example comparisons between actual throughput and estimated throughput are shown.

[0016] Fig.10 An example computing device is shown for implementing the features and operations of the described techniques. DETAILED DESCRIPTION

[0017] Advances in interconnect protocols and architectural enhancements have made GPUs more attractive as accelerators for data analytics. It is expected that the popularity of GPU database systems and the surge in research on their efficient design will continue.

[0018] Query performance depends on query characteristics and input data size, and the same query can have different performance on different GPU database systems and different GPUs. A good understanding of GPU resource utilization and bottlenecks encountered can help design systems that use available hardware resources efficiently.

[0019] Over the years, GPUs have become more powerful as resources for computation, memory capacity, and bandwidth have increased. The ability to support concurrent kernel execution through partial or full partitioning of GPU resources provides users with the opportunity to schedule their workloads and allocate resources to balance performance costs according to their business needs. The relative cost is a fraction of the allocated full GPU resources, and the performance is also relative to the performance with full allocation. The performance impact varies with resource allocation and a scaling factor (SF) related to the input data size. For example, running with half GPU resources causes a 54.5% performance loss for SF=64, compared to only 5.4% for SF=16. For other queries, the trade-off will be different.

[0020] However, while GPUs allow for multiple possible resource allocations, there is no simple way to select the right allocation for a workload. In a naive approach, the user must run multiple representative workloads with different allocations and then select the most appropriate configuration, which can present an impractical, time-consuming, and expensive process. In contrast, the described techniques utilize and adapt models, such as a popular bottleneck analysis and visualization framework known as the "roofline" model, to automatically present estimates to the user about cost-to-performance tradeoffs that can lead to informed decisions about resource allocation. For example, the roofline model reveals that database queries often underutilize GPU resources. Therefore, there are previously unrealized technical benefits to improving overall GPU utilization and workload performance (e.g., speeding up to two times (and more than two times) performance) by supporting concurrent query execution with model-based resource allocation in the GPU. In one implementation, results have been obtained using 7-way concurrency, showing a 6.43x performance improvement.

[0021] The described techniques can be applied to various GPU database systems, including but not limited to: Crystal, a highly optimized academic prototype but with limited coverage (i.e., only supporting certain queries); HeavyDB, a GPU database that supports a few classes of queries; BlazingSQL, another GPU database system that supports a few classes of queries; TQP, a general-purpose GPU database system that uses the PyTorch framework as its backend for performing relational operations; and PG-Strom, an execution of the existing PostgreSQL database system to support offloading operations to the GPU.

[0022] Some, but not all, GPU-based database systems can benefit from a query optimization and query compilation phase. After query optimization, in one implementation, the query plan is compiled and the physical execution of the GPU code is optimized. Specifically, such an implementation can take advantage of a GPU cost model, allowing the GPU execution mode to be part of the query optimization phase. Based on the relative GPU performance cost, the optimizer can determine whether an operator should be offloaded to the GPU. This approach depends on the system having a good estimate of the cost of each operator in the query plan.

[0023] Figure 1 An example resource provisioning system 100 is shown in a computing device 102. Computing device 102 may be or include a simple workstation or laptop computer, a mobile computing device, a computing device within a data center, a robust IoT (Internet of Things) device, or any other type of computing device.

[0024] A query 104 (such as a database query) is received at a computing device 102 via a communication interface 106 (such as a graphical user interface, a network interface, or an I / O port), which passes the query 104 to a resource provider 108, which may include a performance analyzer 118, a resource manager 120, an optimizer, a resource scheduler, and other components executable by a hardware processor. The query may be received in a workload that includes one or more queries for execution by the computing device 102. In an implementation, the resource provider may be executed by one or more processing units (such as (multiple) central processing units 110) to determine resource allocation in another processing unit (such as a graphics processing unit 112) selected to execute the query 104. Example resources allocated in such a processing unit may include, but are not limited to, processor cores, dynamic random access memory (DRAM), L1 cache memory, L2 cache memory, integer / floating point arithmetic logic units (ALUs), and tensor cores.

[0025] The resource provider 108 may include or work in conjunction with a resource scheduler (not shown) that determines whether the query 104 should be assigned to the central processing unit (s) 110 or the graphics processing unit 112 for execution. In some implementations, the resource scheduler is referred to as a "query optimizer". In one implementation, the query optimizer may employ a rule-based approach. In another implementation, the query optimizer may employ a cost-based approach. The resource modeling of the described techniques may be used in a cost-based approach, although it may also be used in other approaches.

[0026] If the resource scheduler determines that the query 104 should be executed by the graphics processing unit 112, the resource provider 108 consults the model repository 114 to determine the resource allocation instructions 116 to configure the graphics processing unit 112 when processing the query 104. In one implementation, the performance analyzer 118 can be executed by one or more hardware processors and is configured to predict the performance impact of the workload on the performance constraints of the processing unit due to the workload according to the resource model. The resource model characterizes the available computing bandwidth, the available memory bandwidth, and the operation intensity based on the peak computing bandwidth and the peak memory bandwidth of the processing unit. The resource manager 120 can be executed by one or more hardware processors and is configured to determine the resource allocation of the processing unit based on the predicted performance impact, and instruct the processing unit to allocate resources for processing the workload based on the determined resource allocation.

[0027] In one implementation, the resource provider 108 sends the resource allocation instruction 116 and the query 104 to the graphics processing unit 112 for execution according to the resource allocation instruction 116. In another implementation, the query 104 may be passed to the graphics processing unit 112 by another component of the computing device 102, such as the communication interface 106. The graphics processing unit 112 configures its resources according to the resource allocation instruction 116 and executes the query 104.

[0028] Figure 2 An example roofline model 200 applied to resource allocation in a processing unit is shown. Typically, a topline model is a type of resource model (also referred to as a performance model) used to estimate the system performance of a given computing kernel, application, or process running on a multi-core, many-core, or accelerator processor architecture by displaying inherent hardware limitations. Such a resource model can also predict the potential benefits and priorities of optimization. In one implementation, the topline model can be visualized by plotting floating-point performance as a function of machine peak performance, machine peak bandwidth, and computational intensity. The resulting curve is actually the performance boundary where the kernel, application, or process performance resides, and includes two platform-specific performance ceilings: an upper limit derived from memory bandwidth and another upper limit derived from the peak computational performance of the processor. One or more axes tend to a logarithmic scale.

[0029] As described herein, the performance of a processing unit (such as a graphics processing unit 112) for different queries and / or different query classifications may be modeled. In one implementation, a roofline model 200 is employed, but other models may be used, such as a machine learning (ML) model (e.g., a tree-based machine learning model). Example features of the ML model may include, but are not limited to, obtaining GPU resource allocations, workload attributes, GPU resource utilization, and resource performance from GPU performance counters. The topline model 200 assumes that any execution in a particular hardware is bounded by its memory (e.g., DRAM, L1 cache memory, L2 cache memory), its compute resources, or some other resources. For ML model implementations, the ML model may be trained on profiled data from previous query executions of a training workload. In order to select GPU resource allocations for a training query or workload, the ML model may be trained based on known runtime labels of the training query or workload based on different GPU resource allocations. Accordingly, the ML model may predict runtime labels for unlabeled queries / workloads.

[0030] Visually, Figure 2 As seen in FIG. 2 , the top line model 200 includes two lines to indicate its peak memory bandwidth (sloping line α) and its peak compute bandwidth (flat line β). Query execution on that particular processing unit will correspond to points within the space defined by these lines, so these two lines are considered to be the performance ceiling for that processing unit. The X-axis represents the operational intensity (AI), which is calculated as the total number of operations (e.g., integer or floating point operations) divided by the total number of bytes used to perform reads. The Y-axis indicates the achieved throughput, calculated as operations executed per second.

[0031] Query execution can be memory-bound (e.g. ) or computationally bounded. Bounded execution is subject to changes in the allocation of corresponding resources (e.g., Figure 2 Algorithmic or compiler inefficiencies will increase the AI ​​and make other memory bound execution unbounded where queries take longer to complete. However, as will be shown, it is relevant to consider L2 cache bandwidth as an additional constraint on the GPU for query execution, and compute bound execution can be affected by changes in compute resource allocation.

[0032] Figure 3Performance modeling results 300 in the DRAM roofline model in 4 example computing systems (Crystal, HeavyDB, BlazingSQL, and TQP) are shown. AI and the throughput achieved are derived from the metrics described in Table 1 below. In Table 1, the peak integer operation performance essentially indicates the upper limit of the computing bandwidth. The theoretical GPU DRAM bandwidth is used for the upper limit of the DRAM bandwidth. It should be noted that the GPU also has other functional units, such as floating-point operation units. However, because most OLAP (online analytical processing) queries may only require integer operations (although other OLAP queries can be modeled as floating-point operations), describing integer operations is sufficient to construct the topline model. However, the described techniques can be applied to floating-point-based queries and a mixture of queries of different types. It should also be noted that each query can consist of multiple kernels, which are addressed by aggregate scalar values ​​(such as execution duration, total bytes, and total integer operation instructions). The aggregate values ​​are then used to obtain the relevant values ​​for constructing the roofline model.

[0033]

[0034] Table 1 Metrics for the roofline model (DRAM)

[0035] The Star Schema Benchmark (SSB) is a collection of performance tests in lightweight data warehouse scenarios. Based on TPC-H, SSB provides a simplified version of the star schema dataset, which is mainly used to test the performance of multi-table join queries under the star schema. In computing, the star schema is the simplest style of data mart architecture and is a widely used approach for developing data warehouses and dimensional data marts. A star schema consists of one or more fact tables that reference any number of dimension tables. The star schema is an important special case of the snowflake schema and is more efficient for processing simpler queries.

[0036] The SSB benchmark has a total of 13 queries. Figure 3 Each data point in corresponds to the execution of a query from one of four systems. Figure 3As shown, the AIs of Crystal, HeavyDB, and BlazingSQL are all relatively low. In particular, for Crystal, 3 queries have saturated the peak GPU DRAM bandwidth. Crystal implements hash joins as filters for these 3 queries. The filter operation simply runs a fast data scan, which is easier to saturate the GPU DRAM bandwidth than a hash join. For the other queries, they all have hash joins, which cause multiple random memory accesses to the GPU DRAM, so their performance is still far from the peak GPU DRAM bandwidth. In contrast, compared to the other 3 systems, BlazingSQL is very compute-intensive. It proves that even simple OLAP queries can be very compute-intensive depending on the query implementation from the system. Interestingly, BlazingSQL has the highest AI as well as operations / second (e.g., example throughput metrics). However, BlazingSQL and TQP load much more data than Crystal and HeavyDB. Therefore, they have more instructions to operate, so the final runtimes of these 2 systems are higher.

[0037] Figure 4 Performance modeling results 400 in the L2 roofline model in 4 example computing systems (Crystal, HeavyDB, BlazingSQL, and TQP) are shown. GPU DRAM bandwidth is not the only resource constraint for query execution. Especially for very optimized systems, such as Crystal and HeavyDB, such systems can be limited by other resources (e.g., L2 cache bandwidth). The AI ​​for different resources is very different. For example, when a query has a very good L2 hit rate, most memory requests will be satisfied by the L2. Therefore, the number of bytes loaded from the L2 cache will be higher. Therefore, the AI ​​relative to the L2 cache is lower. On the other hand, because fewer bytes are loaded from the GPU DRAM, the AI ​​is high in the case of a fixed number of integer operation instructions. Therefore, a separate roofline model can be used to characterize the same query with respect to different memory resources.

[0038] Most of the metrics profiled in Table 1 can be used to build a roofline model of the L2 cache, except for the total bytes that may be read from DRAM. In addition, to estimate the bytes read from L2, the number of L2 requests loaded by the core (shown in Table 2 below) is profiled and multiplied by the cache line size per request (128 bytes).

[0039] Metric Name describe lts__t_requests_srcunit_tex_op_read.sum Total requests to L2 cache

[0040] Table 2 Metrics for the roofline model (L2 cache)

[0041] Figure 4 Indicates that two systems (Crystal and HeavyDB) are very optimized: their performance is sometimes limited by the L2 cache bandwidth. However, for other systems such as BlazingSQL and TQP, their implementations are still far from saturating the L2 cache bandwidth. In our experiments, we profile the same set of queries with the same SF=16 for both DRAM and L2 cache top-line models.

[0042] Back to Figure 3 , Figure 3 shows that very few queries saturate the peak DRAM bandwidth. Figure 4 In , the same queries for the L2 cache topline model show AI degradation. This is reasonable, especially in cases where the queries have good utilization of the L2 cache bandwidth because most memory requests are done at the L2 cache level and therefore the GPU has less data to process in DRAM. The second observation is that several of these queries are actually limited by the peak L2 cache bandwidth. Accordingly, queries with hash joins are more likely to saturate the L2 cache bandwidth. The reason proposed is that the SSB benchmark has a relatively small hash table that may fit in the L2 cache of a high-end GPU. Even though the hash join may incur random accesses, the query can still have good L2 cache utilization due to the small working set size. On the other hand, queries with simple filtering are more likely to be limited by DRAM bandwidth. This is because there is minimal data reuse, but most data is streamed to the core and cache utilization is generally low.

[0043] Figure 5 An example memory hierarchy 500, an example multi-instance GPU (MIG 502), and an example multi-process service (MPS 504) that can be employed to execute workloads are shown. As shown in the memory hierarchy 500, kernel execution can access data stored in shared memory or L1 cache. Shared memory is managed by the user, but the L1 cache is managed by hardware. Each GPU core has a private L1 cache and a shared memory area, but the L2 cache and DRAM are shared across all GPU cores. For example, the NVIDIA A100 GPU has a 40MB L2 cache and 40GB DRAM. The memory hierarchy 500 is consistent for all currently available NVIDIA GPUs, but specific values ​​for capacity and bandwidth vary across GPUs.

[0044] MIG 502 enables physical partitioning of GPU resources (SM (streaming multiprocessors), L2 cache, DRAM capacity, and bandwidth), which can create complete isolation between concurrent users. In this example, MIG 502 shows an example of resource allocation through MIG 502 to support two concurrent clients with equal allocation (1 / 2GPU resources). MIG 502 also supports heterogeneous resource partitioning to meet the different needs of different clients. Currently, the finest resource allocation granularity supported by MIG 502 is 1 / 7GPU resources, so it can support up to 7 concurrent clients with isolated resources on the GPU. MIG 502 currently provides a total of 18 choices for resource partitioning on NVIDIA A100. These features and limitations may vary for different GPUs.

[0045] MPS 504 can provide logical resource partitioning to support concurrent execution. In the old generation GPU, MPS 504 only allows time sharing of the GPU. The new generation MPS 504 allows actual concurrent execution through lightweight resource partitioning by time-sharing SMs. In MPS 504, L2 cache and DRAM are still unified resources without any isolation. When MPS 504 starts, it creates a resource scheduler 506 for the GPU.

[0046] Virtual GPU (vGPU) is another feature of some GPU technologies that allows multiple clients to share the GPU. However, its main purpose is to ensure secure isolation between clients so that malicious clients cannot monitor the activities of others. vGPU adds another layer of protection on top of the shared GPU.

[0047] Resource models (such as a roofline model) can be used to estimate the performance impact of changes in resource allocation, which can then be used to select the best configuration to execute the workload on a processing unit (such as a GPU). Resource model use can be obtained from previous executions of loop queries or by executing runtime statistics of representative workloads.

[0048] For illustration purposes, we use DRAM as the target resource, let t denote the query time under the current allocation and let Bandwidth DRAM Denoting the new DRAM bandwidth, the new query time t' can be predicted by using the following equation:

[0049]

[0050] The equation is proposed based on the observation that the AI ​​is determined by the implementation of the query and is unlikely to change when the resource allocation changes. The resource model picks the maximum value of the two terms in the equation, the first term is used for the scenario where DRAM bandwidth is not a bottleneck (and therefore, the time remains unchanged), and the second term is used for the scenario where DRAM bandwidth is a bottleneck. For the latter case, the denominator in the fraction is the maximum throughput given the AI ​​and the allocated DRAM bandwidth. In this case, the query time is the time to perform the total integer operations at this throughput. A similar equation can be applied to L2 bandwidth.

[0051] Example resource allocation instructions may direct the processing unit to physically allocate portions of memory to execute a query, logically allocate portions of memory to execute a query, and / or allocate all or a substantial subset of processor cores in the processing unit to execute a query.

[0052] Figure 6 Scenario 600 and scenario 602 using representative queries (Q11, Q31) from Crystal and a full-DRAM to half-DRAM allocation change are shown. Q31 underutilizes DRAM bandwidth and has no performance impact, while Q11 loses throughput (geometrically, the points shift downward) because it saturates DRAM bandwidth. Finally, in scenario 600 where DRAM bandwidth is not a bottleneck (and therefore, time remains constant), and in scenario 602 where half DRAM is allocated to query execution and DRAM bandwidth is a bottleneck, the performance impact is computed as SlowdownDRAM=t' / t.

[0053] As discussed, recent GPU systems support not only DRAM bandwidth partitioning, but also L2 bandwidth partitioning. Similar to the above approach, the L2 roofline model can be used to determine the performance impact due to the changed L2 allocation. One challenge is how to combine both the DRAM and L2 roofline models into a unified model to estimate query slowdown, which can be implemented using a max function to provide a total slowdown estimate, as follows:

[0054]

[0055] Note: Can be used with Slowdown DRAM The metric is calculated similarly Metric, but using AI and bandwidth for L2Cache.

[0056] It has been found empirically that queries are rarely bottlenecked by both resources (L2 and DRAM). For example, if a query has very high utilization of L2 cache bandwidth (i.e., it is bottlenecked by L2), it generates only minimal traffic to DRAM, and thus it is not substantially affected by changes in DRAM bandwidth. Therefore, one of the estimated slowdown terms may be "1" (indicating no slowdown). Therefore, the max function produces the dominant slowdown value.

[0057] The same type of resource model can be used to estimate the performance impact for more computationally intensive queries. In this case, BlazingSQL was chosen as an example to illustrate the concept because it has the most computationally intensive implementation compared to the other systems.

[0058] Figure 7 Scenario 700 and scenario 702 show the available bandwidth (i.e., throughput) and AI (2886.74 Gops / sec and 27.15 ops / byte, respectively) for query Q34. The peak computational bandwidth for the full GPU is 18247.00 Gops / sec, and for half GPU resources, the peak computational bandwidth is close to 9123.50 Gops / sec (dashed and dotted lines in scenario 700). Clearly, even though the peak computational bandwidth for the half GPU exceeds the available bandwidth for Q34, the traditional roofline model would predict no performance slowdown. However, it is found that this is not the case for the following reasons: for each query, there is an actual achievable computational bandwidth = achievable bandwidth per SM × the number of SMs on the GPU (dashed line in scenario 700), which is an upper limit much lower than the theoretical peak computational bandwidth. When the GPU is allocated fewer computational resources, it reduces the number of allocated SMs, but it does not improve the execution efficiency (i.e., achievable computational bandwidth) of each SM. As a result, the overall achievable computational bandwidth will then decrease (scenario 702). To estimate the resulting slowdown, the ratio of resource allocation is used as follows:

[0059]

[0060] For example, if the GPU compute resources are halved (ComputeAllocationRation=1 / 2), the achievable bandwidth may be calculated as half of the original achievable bandwidth with full GPU resources.

[0061] Unified Model. Now that we have proposed two models for estimating slowdowns as allocations change: one for memory resources and one for compute resources, the last step is to decide which model to use. Heuristic evaluation is used to determine whether an application is compute-intensive or memory-intensive. As shown below, it can be determined whether an application tends to be compute-intensive based on the AI ​​and peak compute and DRAM bandwidth of the GPU.

[0062]

[0063] If the application is more compute intensive, then a resource model with reduced compute bandwidth is considered. Otherwise, a resource model with reduced DRAM or L2 cache bandwidth is applied.

[0064] The resource model can be extended to estimate the end-to-end performance impact for different degrees of concurrency, which can be used to determine the optimal concurrency for optimal performance. In one aspect, the resource model can be used to evaluate CPU and constant overhead when executing a workload (e.g., a workload including one or more queries) within a processing unit. In another aspect, not necessarily mutually exclusive with the first aspect, the resource model can be used to evaluate the end-to-end performance of the workload execution.

[0065] In order to build a resource model for evaluating CPU and constant overhead, additional overhead for query execution is considered. For CPU overhead, the resource model includes the overhead of query optimization and compilation. For some systems (e.g., HeavyDB, BlazingSQL), even if the same query has been optimized and compiled into binary, each query call still introduces some constant overhead on the CPU side. For such systems, these overheads can be considered to have an impact on query execution performance. At least two additional major overheads can be considered GPU setup overheads, which include GPU context initialization and memory allocation, and data transfer overhead. The relevant system cache tables on the GPU device are used for future query execution, so the data transfer overhead is also a one-time cost.

[0066] As discussed previously, the resource model can be used to estimate the end-to-end query execution time for one process. As the system changes resources, the resource model will adjust the query in GPU execution time. Other overheads will remain constant. Now consider the execution times from multiple concurrent processes, using the max function in one implementation because the longest running process will determine the end-to-end query execution time for concurrent executions.

[0067] ExecTime = max{ExecTime P 1.ExecTime P 2...ExecTime P n}

[0068] Query scheduling may also be used to estimate query performance and concurrency for the MPS, although interference in accesses to the shared L2 cache and DRAM may be excluded from consideration in some implementations.

[0069] Figure 8 An example operation 800 for provisioning resources of a processing unit is shown. A receiving operation 802 receives a workload that may include one or more queries (such as through a communication interface or a user interface). Another receiving operation 804 receives a resource model that characterizes available compute / memory bandwidth and operational intensity based on the peak compute bandwidth and peak memory bandwidth of the processing unit.

[0070] Predicting operation 806 predicts a performance impact on the execution of the workload attributable to the performance constraints of the processing unit for the workload based on the resource model. Determining operation 808 determines a resource allocation for the processing unit based on the predicted performance impact. Instructing operation 810 instructs the processing unit to allocate resources for processing the workload based on the determined resource allocation. Executing operation 812 executes the workload in the processing unit based on the determined workload allocation.

[0071] It should be understood that the predicted performance impact can be negative or positive relative to a baseline or some previous execution. In some cases, instructions for resource allocation can cause a slowdown in query execution (e.g., a negative performance impact). For example, reducing the amount of computing resources, DRAM, and / or L2 cache allocated to a query can slow down the query by reducing the computing and / or memory bandwidth available to the query. In other cases, instructions for resource allocation can cause an acceleration in query execution (e.g., a positive performance impact). For example, increasing the amount of computing resources, DRAM, and / or L2 cache allocated to a query can accelerate the query by increasing the computing and / or memory bandwidth available to the query. By extension, the reduction / increase of any execution resource can be predicted using an appropriate resource model corresponding to a given query.

[0072] Fig. 9 Various example comparisons of query execution times for different queries across multiple scale factors and example comparisons between actual throughput and estimated throughput are shown. Comparison 900 shows the execution time of query Q41 with different scale factors (SF) and GPU resource allocation (x-axis). Comparison 902 shows the execution time of queries Q12, Q41, and Q43 with the same scale factor and GPU resource allocation (x-axis).

[0073] Comparison 904 shows actual throughput versus concurrency (x-axis) across different scaling factors. Comparison 906 shows estimated throughput (estimated by the described techniques) versus concurrency (x-axis) across different scaling factors.

[0074] Such visualization may be presented in a graphical user interface to assist a user in setting or influencing resource application instructions sent to a processing unit, such as a GPU.

[0075] Fig.10 An example computing device 1000 for implementing the features and operations of the described techniques is shown. The computing device 1000 may be embodied as a remote control device or a physical control device, and is an example networked and / or network-capable device, and may be a client device such as a laptop, a mobile device, a desktop computer, a tablet computer, a server / cloud device, an Internet of Things device, an electronic accessory, or another electronic device. The computing device 1000 includes one or more processors 1002 and a memory 1004. The memory 1004 typically includes both volatile memory (e.g., RAM) and non-volatile memory (e.g., flash memory). An operating system 1010 resides in the memory 1004 and is executed by the processor(s) 1002.

[0076] In the example computing device 1000, as Fig.10 As shown, one or more modules or segments, such as applications 1050, user interface manager, query manager, optimizer and / or resource scheduler, resource provisioning, communication interface, and other modules, are loaded into the operating system 1010 on the memory 1004 and / or storage 1020 and executed by the processor(s) 1002. The storage 1020 may include one or more tangible storage media devices and may store queries, workloads, resource models, compute metrics, DRAM metrics, cache (L1 or L2) metrics, or other data and may be local to the computing device 1000 or may be remote and communicatively connected to the computing device 1000.

[0077] Computing device 1000 includes power supply 1016, which is powered by one or more batteries or other power sources and provides power to other components of computing device 1000. Power supply 1016 may also be connected to an external power supply that overcharges or recharges the internal batteries or other power sources.

[0078] The computing device 1000 may include one or more communication transceivers 1030, which may be connected to one or more antennas 1032 to provide a network connection (e.g., a mobile phone network, Wi-Fi, Bluetooth) to one or more other servers and / or client devices (e.g., a mobile device, a desktop computer, or a laptop computer). The computing device 1000 may also include a communication interface 1036 (e.g., a network adapter), which is a type of computing device. The computing device 1000 may use the communication interface 1036 and any other type of computing device to establish a connection through a wide area network (WAN) or a local area network (LAN). It should be understood that the network connections shown are examples, and other computing devices and means for establishing a communication link between the computing device 1000 and other devices may be used.

[0079] The computing device 1000 may include one or more input devices 1034 so that a user can enter commands and information (e.g., a keyboard or mouse). These and other input devices may be coupled to a server by one or more interfaces 1038, such as a serial port interface, a parallel port, or a universal serial bus (USB). The computing device 1000 may also include a display 1022, such as a touch screen display.

[0080] The computing device 1000 may include various tangible processor-readable storage media and intangible processor-readable communication signals. Tangible processor-readable storage may be embodied by any available medium accessed by the computing device 1000, and includes both volatile and non-volatile storage media, removable and non-removable storage media. Tangible processor-readable storage media exclude communication signals (e.g., the signal itself), and include volatile and non-volatile, removable and non-removable storage media implemented in any method or technology for storing information (such as processor-readable instructions, data structures, program modules or other data). Tangible processor-readable storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CDROM, digital versatile disk (DVD) or other optical disk storage, cassettes, tapes, magnetic disk storage or other magnetic storage devices, or any other tangible media that can be used to store desired information and can be accessed by the computing device 1000. Compared to tangible processor-readable storage media, intangible processor-readable communication signals may embody processor-readable instructions, data structures, program modules or other data residing in modulated data signals (such as carrier waves or other signal transmission mechanisms). The term "modulated data signal" means a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, intangible communication signals include signals that travel through wired media, such as a wired network or direct-wired connection, and wireless media, such as acoustic, RF, infrared and other wireless media.

[0081] Item 1: A method for provisioning resources for a processing unit, the method comprising: predicting, based on a resource model, a performance impact on a workload attributable to performance constraints of the processing unit for the workload, wherein the workload comprises queries and the resource model characterizes available compute bandwidth, available memory bandwidth, and operational intensity based on a peak compute bandwidth and a peak memory bandwidth of the processing unit; determining a resource allocation to the processing unit based on the predicted performance impact; and instructing the processing unit to allocate resources for processing the workload based on the determined resource allocation.

[0082] Clause 2: The method of clause 1, further comprising: executing the workload in the processing unit according to the determined resource allocation.

[0083] Clause 3: The method of clause 1, wherein the resource model comprises a roofline model.

[0084] Clause 4: A method according to clause 1, wherein the resource model comprises a machine learning model trained on profile data from previous query executions on a training workload, wherein the machine learning model is configured to predict runtimes on queries for different resource allocations to a processing unit.

[0085] Clause 5: The method of clause 1, wherein the resource allocation includes allocating a portion of a DRAM of a processing unit to the query and allocating a portion of a cache memory of the processing unit to the query.

[0086] Clause 6: The method of clause 1, wherein the resource allocation comprises allocating a processor core of a processing unit to the query.

[0087] Clause 7: A method according to clause 1, wherein the predicted performance impact represents an estimated slowdown in processing the query on a processing unit between two different resource allocations applied to the query in the processing unit, and determining the resource allocation includes: selecting a resource allocation for the processing unit to use when processing the query based on the predicted performance impact between the two different resource allocations.

[0088] Item 8: A system for provisioning resources for a processing unit, the system comprising: one or more hardware processors; a performance analyzer executable by the one or more hardware processors and configured to predict a performance impact on a workload attributable to performance constraints of the processing unit for the workload based on a resource model, wherein the workload comprises a query and the resource model characterizes available compute bandwidth, available memory bandwidth, and operational intensity based on a peak compute bandwidth and a peak memory bandwidth of the processing unit; and a resource manager executable by the one or more hardware processors and configured to determine a resource allocation to the processing unit based on the predicted performance impact and instruct the processing unit to allocate resources for processing the workload based on the determined resource allocation.

[0089] Clause 9: The system of clause 8, wherein the processing unit is configured to execute the workload according to the determined resource allocation.

[0090] Clause 10: The system of clause 8, wherein the resource model comprises a roofline model.

[0091] Clause 11: A system as described in clause 8, wherein the resource model comprises a machine learning model trained on profile data from previous query executions on a training workload, wherein the machine learning model is configured to predict runtimes on queries for different resource allocations for the processing unit.

[0092] Clause 12: The system of clause 8, wherein the resource allocation comprises allocating a portion of a DRAM of a processing unit to the query and allocating a portion of a cache memory of the processing unit to the query.

[0093] Clause 13: The system of clause 8, wherein the resource allocation comprises allocating a processor core of a processing unit to the query.

[0094] Clause 14: A system as described in clause 8, wherein the predicted performance impact represents an estimated slowdown in processing the query on the processing unit between two different resource allocations applied to the query in the processing unit, and the resource manager is configured to determine the resource allocation by selecting the resource allocation for use by the processing unit when processing the query based on the predicted performance impact between the two different resource allocations.

[0095] Item 15: One or more tangible processor-readable storage media embodied with instructions for executing, on one or more processors and circuits of a computing device, a process for provisioning resources for a processing unit, the process comprising: predicting, based on a resource model, a performance impact on a workload attributable to a performance constraint of the processing unit, wherein the workload comprises a query and the resource model characterizes available compute bandwidth, available memory bandwidth, and operational intensity based on a peak compute bandwidth and a peak memory bandwidth of the processing unit; determining a resource allocation to the processing unit based on the predicted performance impact; and instructing the processing unit to allocate resources for processing the workload based on the determined resource allocation.

[0096] Clause 16: The one or more tangible processor-readable storage media of clause 15, wherein the process further comprises: executing the workload in the processing unit according to the determined resource allocation.

[0097] Clause 17: The one or more tangible processor-readable storage media of clause 15, wherein the resource model comprises a roofline model.

[0098] Clause 18: One or more tangible processor-readable storage media as described in clause 15, wherein the resource model comprises a machine learning model trained on profiling data from previous query executions on a training workload, wherein the machine learning model is configured to predict runtimes on queries for different resource allocations for a processing unit.

[0099] Clause 19: One or more tangible processor-readable storage media as described in clause 15, wherein the resource allocation includes allocating a portion of a DRAM of a processing unit to the query, allocating a portion of a cache memory of a processing unit to the query, or allocating a processor core of a processing unit to the query.

[0100] Clause 20: One or more tangible processor-readable storage media as described in clause 15, wherein the predicted performance impact represents an estimated slowdown in processing the query on the processing unit between two different resource allocations applied to the query in the processing unit, and determining the resource allocation includes: selecting a resource allocation for the processing unit to use when processing the query based on the predicted performance impact between the two different resource allocations.

[0101] Item 21: A system for provisioning resources for a processing unit, the system comprising: a device for predicting a performance impact on a workload attributable to performance constraints of the processing unit for the workload based on a resource model, wherein the workload comprises a query and the resource model characterizes available compute bandwidth, available memory bandwidth, and operational intensity based on a peak compute bandwidth and a peak memory bandwidth of the processing unit; a device for determining an allocation of resources to the processing unit based on the predicted performance impact; and a device for instructing the processing unit to allocate resources for processing the workload based on the determined resource allocation.

[0102] Clause 22: The system of clause 21, further comprising: means for executing the workload in the processing unit according to the determined resource allocation.

[0103] Clause 23: The system of clause 21, wherein the resource model comprises a roofline model.

[0104] Clause 24: A system as described in clause 21, wherein the resource model comprises a machine learning model trained on profile data from previous query executions on a training workload, wherein the machine learning model is configured to predict runtimes on queries for different resource allocations for a processing unit.

[0105] Clause 25: The system of clause 21, wherein the resource allocation comprises allocating a portion of a DRAM of a processing unit to the query and allocating a portion of a cache memory of the processing unit to the query.

[0106] Clause 26: The system of clause 21, wherein the resource allocation comprises allocating a processor core of a processing unit to the query.

[0107] Clause 27: A system according to clause 21, wherein the predicted performance impact represents an estimated slowdown in processing the query on a processing unit between two different resource allocations applied to the query in the processing unit, and the device for determining the resource allocation includes: a device for selecting a resource allocation for the processing unit to use when processing the query based on the predicted performance impact between the two different resource allocations.

[0108] The various software components described herein may be executed by one or more processors, which may include a logic machine configured to execute hardware or firmware instructions. For example, a processor may be configured to execute instructions that are part of one or more applications, services, programs, routines, libraries, objects, components, data structures, or other logical constructs. Such instructions may be implemented to perform tasks, implement data types, transform the state of one or more components, achieve technical effects, or otherwise achieve desired results.

[0109] The processor and storage aspects may be integrated together into one or more hardware logic components. For example, such hardware logic components may include field programmable gate arrays (FPGAs), programmable and application specific integrated circuits (PASIC / ASIC), application specific and application specific standard products (PSS / ASSP), systems on chips (SOCs), and complex programmable logic devices (CPLDs).

[0110] The terms "module", "program", and "engine" may be used to describe aspects of a remote control device and / or a physical control device that is implemented to perform a particular function. It should be understood that different modules, programs, and / or engines may be instantiated from the same application, service, code block, object, library, routine, API, function, etc. Likewise, the same module, program, and / or engine may be instantiated by different applications, services, code blocks, objects, routines, APIs, functions, etc. The terms "module", "program", and "engine" may cover individual or groups of executable files, data files, libraries, drivers, scripts, database records, and the like.

[0111] It should be understood that a "service" as used herein is an application that can be executed across one or more user sessions. A service is an application that can be used for one or more system components, programs, and / or other services. In some implementations, a service can run on one or more server computing devices.

[0112] Although this specification contains many specific implementation details, these should not be interpreted as limitations on any technology or the scope that can be claimed, but rather as descriptions of features specific to a particular implementation of a particular described technology. Certain features described in this specification in the context of separate implementations may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented in multiple implementations individually or in any suitable sub-combination. In addition, although features may be described above as working in certain combinations and even initially claimed as such, in some cases, one or more features from the claimed combination may be removed from the combination, and the claimed combination may involve a sub-combination or a variation of the sub-combination.

[0113] Similarly, although operations are depicted in a particular order in the accompanying drawings, this should not be understood as requiring to perform such operations in the particular order shown or in order, or to perform all displayed operations to achieve desired results. In addition, it should be understood that logical operations can be performed in any order, adding or omitting operations as required, regardless of whether operations are marked or identified as optional, unless otherwise explicitly stated or the claim language inherently requires a particular order. In some cases, the actions recorded in the claims can be performed in different orders and still achieve desirable results. In some cases, multitasking and parallel processing may be advantageous. The logical operations constituting the implementation of the technology described herein may be referred to as operations, steps, objects, or modules in different ways.

[0114] In addition, the separation of various system components in the implementation described above should not be understood as requiring such separation in all implementations, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products. Therefore, a specific implementation of the subject matter has been described. Other implementations are within the scope of the attached claims. However, it should be understood that various modifications can be made without departing from the spirit and scope of the enumerated claims.

Claims

1. A method for provisioning resources of a processing unit, the method comprising: predicting, based on a resource model, a performance impact on the workload attributable to a performance constraint of the processing unit for the workload, wherein the workload comprises a query and the resource model characterizes available compute bandwidth, available memory bandwidth, and operational intensity based on a peak compute bandwidth and a peak memory bandwidth of the processing unit; determining a resource allocation to the processing unit based on the predicted performance impact; as well as The processing unit is instructed to allocate the resources for processing the workload based on the determined resource allocation.

2. The method according to claim 1, further comprising: The workload is executed in the processing unit according to the determined resource allocation. The method of claim 1 , wherein the asset model comprises a roofline model.

4. The method of claim 1 , wherein the resource model comprises a machine learning model trained on profile data from previous query executions on a training workload, wherein the machine learning model is configured to predict runtimes on the query for different resource allocations for the processing unit.

5. The method of claim 1, wherein the resource allocation comprises allocating a portion of a DRAM of the processing unit to the query and allocating a portion of a cache memory of the processing unit to the query. The method of claim 1 , wherein the resource allocation comprises allocating a processor core of the processing unit to the query.

7. The method of claim 1 , wherein the predicted performance impact represents an estimated slowdown in processing the query on the processing unit between two different resource allocations applied to the query in the processing unit, and determining the resource allocations comprises: The resource allocation for use by the processing unit in processing the query is selected based on the predicted performance impact between the two different resource allocations.

8. A system for provisioning resources of a processing unit, the system comprising: one or more hardware processors; a performance analyzer executable by the one or more hardware processors and configured to predict, based on a resource model, a performance impact on a workload attributable to a performance constraint of the processing unit for the workload, wherein the workload comprises a query and the resource model characterizes an achievable compute bandwidth, an achievable memory bandwidth, and an operational intensity based on a peak compute bandwidth and a peak memory bandwidth of the processing unit; as well as A resource manager executable by the one or more hardware processors and configured to determine a resource allocation to the processing unit based on the predicted performance impact, and to instruct the processing unit to allocate the resources for processing the workload based on the determined resource allocation.

9. The system of claim 8, wherein the processing unit is configured to execute the workload according to the determined resource allocation.

10. The system of claim 8, wherein the asset model comprises a roofline model.

11. The system of claim 8, wherein the resource model comprises a machine learning model trained on profiling data from previous query executions on a training workload, wherein the machine learning model is configured to predict runtimes on the query for different resource allocations for the processing unit.

12. The system of claim 8, wherein the resource allocation comprises allocating a portion of a DRAM of the processing unit to the query and allocating a portion of a cache memory of the processing unit to the query.

13. The system of claim 8, wherein the resource allocation comprises allocating a processor core of the processing unit to the query.

14. A system according to claim 8, wherein the predicted performance impact represents an estimated slowdown in processing the query on the processing unit between two different resource allocations applied to the query in the processing unit, and the resource manager is configured to determine the resource allocation by selecting the resource allocation for use by the processing unit when processing the query based on the predicted performance impact between the two different resource allocations.

15. One or more tangible processor-readable storage media embodied with instructions for execution on one or more processors and circuits of a computing device for provisioning resources of a processing unit, the process comprising: predicting, based on a resource model, a performance impact on the workload attributable to a performance constraint of the processing unit for the workload, wherein the workload comprises a query and the resource model characterizes available compute bandwidth, available memory bandwidth, and operational intensity based on a peak compute bandwidth and a peak memory bandwidth of the processing unit; determining a resource allocation to the processing unit based on the predicted performance impact; as well as The processing unit is instructed to allocate the resources for processing the workload based on the determined resource allocation.

16. The one or more tangible processor-readable storage media of claim 15, wherein the process further comprises: The workload is executed in the processing unit according to the determined resource allocation.

17. The one or more tangible processor-readable storage media of claim 15, wherein the resource model comprises a roofline model.

18. One or more tangible processor-readable storage media according to claim 15, wherein the resource model comprises a machine learning model trained on profiling data from previous query executions on a training workload, wherein the machine learning model is configured to predict the running time on the query for different resource allocations for the processing unit.

19. One or more tangible processor-readable storage media according to claim 15, wherein the resource allocation includes allocating a portion of the DRAM of the processing unit to the query, allocating a portion of the cache memory of the processing unit to the query, or allocating a processor core of the processing unit to the query.

20. The one or more tangible processor-readable storage media of claim 15, wherein the predicted performance impact represents an estimated slowdown in processing the query on the processing unit between two different resource allocations applied to the query in the processing unit, and determining the resource allocations comprises: The resource allocation for use by the processing unit in processing the query is selected based on the predicted performance impact between the two different resource allocations.

Citation Information

Cited By

  • Method, computing device, medium and program product for predicting performance of computing system executed by application program

    CN120631705A