A GPU-based Optimization Method for Large-scale Parallel Computing

By obtaining the load amount and number of kernel functions of the model task in real time, calculating the comprehensive evaluation score, analyzing the chain competition and frequency reduction performance, optimizing the allocation of GPU equipment resources, solving the problems of parallel computing resource competition and load imbalance of GPU equipment, and improving the computing efficiency of the GPU cluster.

CN120123107BActive Publication Date: 2025-07-18SUPER TELECOM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510614618.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-07-18
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

In multi-model deployment scenarios, the parallel computing resource competition and load imbalance of GPU devices lead to overload or paralysis of some devices, affecting computing efficiency.

Method used

By obtaining the load, batch size and number of kernel functions of the model task in real time, calculating the comprehensive evaluation score, analyzing the chain competition weight and parallel competition intensity, combining the frequency reduction performance waste, determining the priority of resource allocation, and using a multi-criteria decision-making algorithm to optimize the resource allocation of GPU devices.

Benefits of technology

Effectively reduce resource competition and frequency reduction problems, balance GPU equipment load, and improve parallel computing efficiency of large-scale GPU clusters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123107B_ABST
    Figure CN120123107B_ABST
Patent Text Reader

Abstract

This application relates to the field of parallel computing technology, and specifically relates to an optimization method for large-scale parallel computing based on GPU. The method includes: obtaining in real time the minimum GPU resource amount required for each queuing model task at the current moment, as well as the load amount, batch size, and number of kernel functions of each execution model task in each GPU device; obtaining in real time the operating power, working frequency, and available resource amount of each GPU device at each moment; calculating a comprehensive evaluation score; determining the chain contention weight, parallel contention intensity, down-frequency performance waste degree, parallel computing efficiency, and resource allocation priority of each GPU device at the current moment, and performing resource allocation for the GPU devices. This application can reduce the resource contention and down-frequency problems caused by GPU resource fragmentation, can balance the load of GPU devices, and improve the efficiency of large-scale GPU cluster parallel computing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of parallel computing technology, and particularly to an optimization method for large-scale parallel computing based on GPU. Background Art

[0002] As a coprocessor widely used in major supercomputers, the GPU (Graphics Processing Unit) has powerful floating-point computing capabilities, providing hardware support for large-scale science and engineering technologies. In complex multi-model deployment scenarios, parallel computing of model tasks on GPU devices will cause competition for computing power resources, and it is difficult to balance GPU resources, often resulting in overload or even paralysis of some GPU devices.

[0003] For job requests with multiple model tasks existing simultaneously, when allocating resources to GPU devices through the first-fit algorithm, due to the arrangement order of available GPU devices, it is easy to cause GPU devices with low available resource amounts to parallelly process more model tasks, while GPU devices with high available resource amounts are idle for a long time, thereby greatly increasing the competition for computing power resources of model tasks and reducing the efficiency of parallel computing, resulting in poor load balancing of GPU devices. Summary of the Invention

[0004] To solve the above technical problems, an optimization method for large-scale parallel computing based on GPU is provided to solve the existing problems.

[0005] The solution of this application to solve the technical problem is to provide an optimization method for large-scale parallel computing based on GPU, including the following steps:

[0006] Obtain in real time the minimum GPU resource amount required by each queued model task at the current moment, as well as the load, batch size, and number of kernel functions of each executing model task in each GPU device; obtain in real time the operating power, working frequency, and available resource amount of each GPU device at each moment;

[0007] Conduct a comprehensive evaluation and analysis of the executing model tasks through the load, batch size, and number of kernel functions of each executing model task, calculate the comprehensive evaluation score of each executing model task in each GPU device at the current moment, and obtain each high-demand task based on the comprehensive evaluation score;

[0008] Analyze the deviation of the comprehensive evaluation scores of all high-demand tasks in each GPU device, determine the chain competition weight of each GPU device at the current moment; based on the comprehensive evaluation scores of all executing model tasks in each GPU device, combine the chain competition weight to determine the parallel competition intensity of each GPU device at the current moment;

[0009] Analyze the running power and the deviation degree of the working frequency of each GPU device at different times to obtain the downclocking performance waste degree of each GPU device at the current time; fuse the parallel contention intensity and the downclocking performance waste degree to determine the parallel computing efficiency of each GPU device at the current time;

[0010] Comprehensively evaluate the GPU devices based on the available resource amount and the parallel computing efficiency, and calculate the resource allocation priority of each GPU device at the current time; according to the proportion of the minimum GPU resource amount required by all queuing model tasks in the available resource amount of all GPU devices, and in combination with the resource allocation priority, obtain the device allocation sequence; combine the first-fit algorithm to allocate resources to the GPU devices.

[0011] Preferably, calculating the comprehensive evaluation score of each execution model task in each GPU device at the current time includes:

[0012] Form the execution task vector of each execution model task in each GPU device at the current time with the load amount, batch size, and number of kernel functions of each execution model task in each GPU device at the current time;

[0013] Adopt a multi-criteria decision-making algorithm to comprehensively evaluate the execution task vectors of all execution model tasks in all GPU devices at the current time, and obtain the comprehensive evaluation score of each execution model task in each GPU device at the current time.

[0014] Preferably, obtaining each high-demand task includes:

[0015] Adopt a threshold segmentation algorithm to obtain the segmentation threshold of the comprehensive evaluation scores of all execution model tasks in all GPU devices at the current time;

[0016] Record all execution model tasks with comprehensive evaluation scores greater than the segmentation threshold as each high-demand task.

[0017] Preferably, the th GPU device at the current time The calculation formula of the chain contention weight is: , where is the comprehensive evaluation score of the th high-demand task in the th GPU device at the current time, is the segmentation threshold, is at the current time the number of all high-demand tasks in the th GPU device, is a preset weight value.

[0018] Preferably, the th GPU device at the current moment parallel contention intensity is calculated as follows: , where is the chain contention weight of the th GPU device at the current moment , is the sum of all the comprehensive evaluation scores of the model tasks executed by the th GPU devices at the current moment, is the current moment th the number of all model tasks executed by the GPU devices at the current moment.

[0019] Preferably, obtaining the down - frequency performance waste degree of each GPU device at the current moment includes:

[0020] Taking the operating power of each GPU device at multiple moments before the current moment as the input of the prediction model to obtain the power prediction value of each GPU device at the next moment;

[0021] Forming a working frequency sequence with the working frequencies of each GPU device at multiple moments before the current moment, and obtaining the minimum value points within the working frequency sequence;

[0022] Obtaining the rated power and the maximum working frequency of the GPU device; calculating the difference between the power prediction value and the rated power, denoted as the power difference; the calculation result of the exponential function with the natural constant as the base and the power difference as the exponent is denoted as the relative power difference;

[0023] Denoting the sum of the differences between the maximum working frequency and the working frequencies corresponding to all the minimum value points within the working frequency sequence as the relative frequency difference;

[0024] The down - frequency performance waste degree is the product of the relative power difference and the relative frequency difference.

[0025] Preferably, the parallel computing efficiency of each GPU device at the current moment is the normalized result of the reciprocal of the product of the parallel contention intensity and the down - frequency performance waste degree.

[0026] Preferably, the resource allocation priority of each GPU device at the current moment includes:

[0027] ​The available resource amount of each GPU device at the current moment and the parallel computing efficiency are combined to form a feature vector; the feature vectors of all GPU devices at the current moment are used as the input of the multi-criteria decision-making algorithm, and the comprehensive evaluation scores of each GPU device obtained are used as the resource allocation priorities of each GPU device at the current moment.

[0028] Preferably, the obtaining of the device allocation sequence includes:

[0029] The sum of the minimum GPU resource amounts required by all queuing model tasks at the current moment is denoted as the first sum value; the sum value of the available resource amounts of all GPU devices at the current moment is denoted as the second sum value; the ratio of the first sum value to the second sum value is used as the task resource ratio.

[0030] If the task resource ratio is greater than a preset first threshold, select the GPU devices with a parallel computing efficiency greater than a preset second threshold, and arrange them in ascending order according to the resource allocation priority to form a device allocation sequence; otherwise, arrange all GPU devices in ascending order according to the resource allocation priority to form a device allocation sequence, where the preset first threshold is greater than the preset second threshold.

[0031] Preferably, the resource allocation to the GPU devices includes:

[0032] Arrange all queuing model tasks at the current moment in descending order according to the minimum GPU resource amount required by them to form a task queuing sequence.

[0033] Based on the task queuing sequence and the device allocation sequence, use the first-fit algorithm to perform resource allocation to the GPU devices.

[0034] This application has at least the following beneficial effects:

[0035] This application comprehensively evaluates and analyzes the execution of model tasks based on the workload, batch size, and number of kernel functions of each execution model task, calculates the comprehensive evaluation scores of each execution model task in each GPU device at the current moment, and obtains each high-demand task based on the comprehensive evaluation scores. The beneficial effect is that it considers the resource requirements of each execution model task to reflect whether each execution model task in each GPU device requires high resource requirements; analyzes the deviation of the comprehensive evaluation scores of all high-demand tasks in each GPU device to determine the chain contention weight of each GPU device at the current moment. The beneficial effect is that it considers the possibility of mutual resource contention among high-demand tasks in the GPU device to reflect the chain contention phenomenon existing in the GPU device resources; based on the comprehensive evaluation scores of all execution model tasks in each GPU device, combined with the chain contention weight, determines the parallel contention intensity of each GPU device at the current moment. The beneficial effect is that it considers the significance of resource contention that occurs during the parallel computing execution of model tasks in the GPU device; analyzes the deviation degree of the operating power and working frequency of each GPU device at different moments to obtain the downclocking performance waste degree of each GPU device at the current moment. The beneficial effect is that it considers the changes in the operating power and working frequency of the GPU device to reflect the GPU device performance waste situation and illustrate the efficiency of parallel computing of the GPU device; fuses the parallel contention intensity and the downclocking performance waste degree to determine the parallel computing efficiency of each GPU device at the current moment. The beneficial effect is that it reduces the additional GPU resource requirements caused by the increase in resource contention and downclocking, facilitating the subsequent provision of a more efficient GPU device resource allocation strategy; comprehensively evaluates the GPU device through the available resource amount and the parallel computing efficiency, calculates the resource allocation priority of each GPU device at the current moment. The beneficial effect is that it considers the available resource amount situation of the GPU device and the degree of parallel computing efficiency, so as to preferentially allocate to GPU devices with a higher available resource amount and high parallel computing efficiency; based on the proportion of the minimum GPU resource amount required by all queued model tasks at the current moment in the available resource amount of all GPU devices, combined with the resource allocation priority, obtains the device allocation sequence; combines the first-fit algorithm to allocate resources to the GPU device. The beneficial effect is that it reduces the resource contention and downclocking problems caused by GPU resource fragmentation, and can also balance the load of the GPU device, avoiding the low parallel computing efficiency of GPU devices with low available resource amounts and the idle state of GPU devices with high available resource amounts during the non-peak period of parallel computing of the model deployment platform, and improving the efficiency of parallel computing of large-scale GPU clusters. Brief Description of the Drawings

[0036] The following further elaborates in detail a GPU-based large-scale parallel computing optimization method of the present application with reference to the accompanying drawings.

[0037] Figure 1 It is a flowchart of the steps of an optimization method for large-scale parallel computing based on GPU provided by an embodiment of the present application;

[0038] Figure 2 It is a flowchart of the steps of a method for obtaining the degree of wasted downclock performance of each GPU device at the current moment provided by an embodiment of the present application. Detailed implementation manners

[0039] In order to make the objectives, technical solutions and advantages of the present application clearer, the following further details an optimization method for large-scale parallel computing based on GPU proposed in the present application in conjunction with the accompanying drawings and implementation examples. It should be understood that the specific implementation examples described herein are only used to explain the present application and are not used to limit the present application.

[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs.

[0041] Please refer to Figure 1 , which shows a flowchart of the steps of an optimization method for large-scale parallel computing based on GPU provided by an embodiment of the present application. The method includes the following steps:

[0042] Step 1, obtain in real time the minimum GPU resource amount required by each queuing model task at the current moment, as well as the load, batch size, and number of kernel functions of each execution model task in each GPU device; obtain in real time the operating power, working frequency, and available resource amount of each GPU device at each moment.

[0043] With the rapid development of deep learning, the number of parameters and the amount of computation of deep network models are increasing day by day. There is a great waste of computing resources when a single model task monopolizes a GPU device. The current parallel computing deployment optimization method using a large-scale GPU cluster with multiple models greatly improves the utilization rate of GPU resources.

[0044] In a model deployment platform, model tasks include queuing model tasks and execution model tasks. Among them, a queuing model task is a model task that queues in the model deployment platform waiting for GPU resource allocation and parallel computing; an execution model task is a model task that is performing parallel computing in a GPU device.

[0045] The model deployment platform has multiple GPU devices. Each GPU device has multiple executing model tasks performing parallel computing. By using the monitoring tool of the model deployment platform, the minimum GPU resource amount required for each queuing model task at the current moment is obtained in real time, as well as the load amount, batch size, and number of kernel functions of each executing model task in each GPU device at the current moment. The operating power, working frequency, and available resource amount of each GPU device at each moment are obtained in real time.

[0046] It should be noted that the load amount is the resource amount stored by the executing model task in the GPU device; the batch size is the number of samples processed by the GPU device in a single batch; the number of kernel functions reflects the complexity of the model task type.

[0047] In this embodiment, the monitoring tool includes a GPU scheduler and the Nvidia System Management Interface (Nvidia-smi); secondly, the acquisition time interval is 20 ms. As other implementation manners, the implementer can set it according to the actual situation.

[0048] Thus, the minimum GPU resource amount required for each queuing model task at the current moment is obtained, as well as the load amount, batch size, and number of kernel functions of each executing model task in each GPU device at the current moment. The operating power, working frequency, and available resource amount of each GPU device at each moment are obtained in real time.

[0049] Step 2: Conduct a comprehensive evaluation and analysis of the executing model tasks through the load amount, batch size, and number of kernel functions of the executing model tasks, calculate the comprehensive evaluation scores of each executing model task in each GPU device at the current moment, and obtain each high-demand task based on the comprehensive evaluation scores; analyze the deviation of the comprehensive evaluation scores of all high-demand tasks in each GPU device to determine the chain contention weight of each GPU device at the current moment; based on the comprehensive evaluation scores of all executing model tasks in each GPU device, combined with the chain contention weight, determine the parallel contention intensity of each GPU device at the current moment.

[0050] The processing core unit of the GPU device is the Streaming Multiprocessor (SM). There is a warp scheduler in the Streaming Multiprocessor (SM) responsible for coordinating and distributing instructions to the CUDA (Compute Unified Device Architecture) cores, thereby realizing the parallel computing of multiple executing model tasks in a single GPU device. The parallel computing of the executing model tasks will inevitably lead to resource contention in the GPU device, and the resource contention ability of different executing model tasks mainly depends on the load amount, batch size, and number of kernel functions of the executing model tasks.

[0051] Once an execution model task is assigned to a certain GPU device, the execution model task must start processing and execution after its associated data is loaded into the exclusive memory of the GPU device. Among them, the larger the workload of the execution model task, the longer the queuing delay of memory data reading, and the greater the resource demand for the GPU memory. The execution model task has a larger batch size, and a single calculation requires more CUDA cores, so the resource demand for occupying the GUDA cores is greater; the larger the number of kernel functions of the execution model task, the greater the resource demand for the kernel function scheduling of the GPU device.

[0052] The workload, batch size, and number of kernel functions of each execution model task in each GPU device at the current moment are combined to form an execution task vector of each execution model task in each GPU device at the current moment;

[0053] Using a multi-criteria decision-making algorithm, comprehensively evaluate the execution task vectors of all execution model tasks in all GPU devices at the current moment, and obtain the comprehensive evaluation scores of each execution model task in each GPU device at the current moment;

[0054] In this embodiment, the Topsis algorithm (Technique for Order Preference by Similarity to Ideal Solution) is used for comprehensive evaluation. Among them, the Topsis algorithm is a well-known technology and will not be elaborated here. As other implementation manners, implementers can use other methods of the existing technology, for example, the Analytic Hierarchy Process (AHP), etc. This embodiment does not make special restrictions on this.

[0055] It should be noted that the comprehensive evaluation score reflects the resource demand ability of the corresponding execution model task; the larger the comprehensive evaluation score, the more likely the corresponding execution model task is a high-resource-demand task. When multiple execution model tasks with high resource demand capabilities are calculated in parallel, more GPU scheduling delays are required, and it is more likely to cause resource chain contention, which will inevitably increase the amplitude of resource contention.

[0056] Furthermore, based on the comprehensive evaluation score, calculate the chain contention weight to reflect the chain contention phenomenon of GPU device resources, specifically:

[0057] Using a threshold segmentation algorithm, obtain the segmentation threshold of the comprehensive evaluation scores of all execution model tasks in all GPU devices at the current moment;

[0058] In this embodiment, the Otsu threshold segmentation algorithm is used to obtain the segmentation threshold. Among them, the Otsu threshold segmentation algorithm is a well-known technology and will not be elaborated here.

[0059] All execution model tasks with the comprehensive evaluation score greater than the segmentation threshold are recorded as high-demand tasks respectively;

[0060] Based on the above analysis, calculate the chain contention weight of each GPU device at the current moment. Taking the chain contention weight of the th GPU device in the model deployment platform at the current moment as an example, the calculation formula is:

[0061]

[0062] where, is the chain contention weight of the th GPU device at the current moment , is the comprehensive evaluation score of the th high-demand task among all high-demand tasks of the th GPU device at the current moment, is the segmentation threshold, is the number of all high-demand tasks of the th GPU device at the current moment , is a preset weight value. In this embodiment, the preset weight value takes the value of 1. As other implementation manners, the implementer can set it according to the actual situation.

[0063] It should be noted that the function of the preset weight value is that if there are no high-demand tasks in the GPU device, the chain contention weight of its corresponding GPU device is at least 1; secondly, the larger is, the higher the GPU device resource requirement for the corresponding high-demand task at this time, and the more likely this high-demand task will affect other execution model tasks in parallel computing; the larger

[0064] is, the greater the possibility of mutual resource contention among multiple high-demand tasks in the corresponding GPU device at this time, and the larger the obtained chain contention weight , the more likely it is to cause the chain contention phenomenon of GPU device resources. ​

[0065]

[0066] Among them, is the parallel contention intensity of the th GPU device at the current moment , is the chain contention weight of the th GPU device at the current moment , is the sum of the comprehensive evaluation scores of all the execution model tasks in the th GPU devices at the current moment, is the number of all the execution model tasks in the th GPU devices at the current moment.

[0067] It should be noted that the above formula is in the form of a power-exponential function, where is the base, is the exponent. The function of the value 1 in the base is to prevent the base from being less than 1, resulting in a monotonically decreasing power-exponential function. Secondly, The larger it is, the stronger the resource demand of all the execution model tasks in the th GPU device at the current moment. The more intense the mutual resource contention is, the lower the utilization rate of the th GPU will be; the larger the chain contention weight , the more likely it is that chain contention will occur in the mutual resource contention of the execution model tasks at this time, which will increase the latency of the GPU scheduler to handle the chain phenomenon; The larger it is, the more parallel load quantities there are in the th GPU device at the current moment, the more complex the resource contention relationship is, and the larger the obtained parallel contention intensity , the more significant the resource contention caused by the parallel computing of each execution model task of the th GPU device at the current moment will be.

[0068] Thus, the parallel contention intensity of each GPU device at the current moment is obtained.

[0069] Step 3: Analyze the deviation degrees of the operating power and working frequency of each GPU device at different moments, and obtain the downclocking performance waste degree of each GPU device at the current moment; fuse the parallel contention intensity and the downclocking performance waste degree to determine the parallel computing efficiency of each GPU device at the current moment.

[0070] Furthermore, if the number of all execution model tasks that the GPU device is performing parallel computing on increases, it not only causes an increase in latency and resource contention, but also is accompanied by an increase in the load of the GPU device, requiring higher power carrying, and the power of the GPU device continues to rise. At the same time, the specification table of the GPU device has an upper limit value for power, that is, the rated power, which is used to prevent the GPU device from overheating. When the operating power of the GPU device approaches or reaches the upper limit, the GPU device will automatically reduce the clock frequency and working voltage, lower the working frequency of the GPU device, and maintain the GPU power not exceeding the upper limit. However, parallel computing causes the working frequency of the GPU device to decrease, which is a great waste of resources for the GPU device. The frequency reduction may cause greater waste of the performance of the GPU device, making the GPU device unable to utilize its maximum performance for a long time and reducing the efficiency of parallel computing of the GPU device.

[0071] First, analyze the changes in the operating power and working frequency of the GPU device, and calculate the degree of performance waste due to frequency reduction. The flowchart of the steps for obtaining the degree of performance waste due to frequency reduction of each GPU device at the current moment provided by the embodiments of the present application is as Figure 2 shown, and specifically includes:

[0072] Take the operating power of each GPU device at multiple moments before the current moment as the input of the prediction model, and obtain the power prediction value of each GPU device at the next moment;

[0073] In this embodiment, take the operating power of each GPU device at all moments within 1 second before the current moment as the input of the Autoregressive Integrated Moving Average model (ARIMA), and obtain the power prediction value of each GPU device at the next moment. The ARIMA model is a well-known technology and will not be elaborated here.

[0074] Form a working frequency sequence with the working frequencies of each GPU device at multiple moments before the current moment, and obtain the minimum value points within the working frequency sequence;

[0075] In this embodiment, form a working frequency sequence with the working frequencies of each GPU device at all moments within 1 second before the current moment. As other implementation manners, the implementer can set it by himself according to the actual situation. Secondly, use the difference method to obtain the minimum value points, where the difference method is a well-known technology and will not be elaborated here.

[0076] Obtain the rated power and the maximum working frequency of the GPU device;

[0077] It should be noted that the rated power and the maximum working frequency of the GPU device can be queried through the product specification table.

[0078] Based on the differences between the operating frequencies corresponding to the minimum points in the operating frequency sequence and the maximum operating frequency, as well as the differences between the power prediction values and the rated power, calculate the degradation performance waste degree, specifically:

[0079] Calculate the difference between the power prediction value and the rated power, denoted as the power difference;

[0080] The calculation result of the exponential function with the natural constant as the base and the power difference as the exponent is denoted as the relative power difference;

[0081] In this embodiment, calculate the difference between the power prediction value and the rated power, denoted as the power difference. Among them, the calculation formula for the relative power difference of each GPU device at the current moment is: , is the th GPU device at the current moment of the relative power difference, is the th GPU device at the moment of the power prediction value, is the rated power of the GPU device, is the exponential function with the natural constant as the base; secondly, is the power difference.

[0082] Denote the sum of the differences between the maximum operating frequency and the operating frequencies corresponding to all the minimum points in the operating frequency sequence as the relative frequency difference;

[0083] In this embodiment, denote the sum of the differences between the maximum operating frequency and the operating frequencies corresponding to all the minimum points in the operating frequency sequence as the relative frequency difference; the calculation formula for the relative frequency difference of each GPU device at the current moment is: , is the th GPU device at the current moment of the relative frequency difference, is the maximum operating frequency of the GPU device, is the th GPU device at the current moment in the operating frequency sequence at the th minimum point corresponding to the operating frequency, is the th GPU device at the current moment in the operating frequency sequence of all the minimum points.

[0084] Take the product of the relative power difference and the relative frequency difference as the degradation performance waste degree of each GPU device at the current moment;

[0085] It should be noted that the greater the relative frequency difference, the more serious the frequency reduction caused by parallel computing for the corresponding GPU device. The greater the relative power difference, the more the predicted power value of the GPU device exceeds the upper limit, and the future operating frequency of the GPU device also needs to be reduced, resulting in a decrease in the efficiency of parallel computing of the GPU device and more serious performance waste. Then, the greater the degree of frequency reduction performance waste obtained.

[0086] Furthermore, based on the parallel contention intensity and the degree of frequency reduction performance waste, the parallel computing efficiency is calculated specifically as follows:

[0087] The normalized result of the reciprocal of the product of the parallel contention intensity and the degree of frequency reduction performance waste is used as the parallel computing efficiency of each GPU device at the current moment;

[0088] In this embodiment, the sigmoid function is used for normalization processing. Among them, the sigmoid function is a well-known technology and will not be elaborated here. As other implementation manners, implementers can adopt other methods of existing technologies, such as the tanh function, etc. This embodiment does not make special restrictions on this.

[0089] It should be noted that the greater the parallel contention intensity, the more intense the resource contention caused by parallel computing of the GPU device, the more the parallel computing delay increases, and the more it affects the performance of the GPU device, and the lower the parallel computing efficiency; the greater the degree of frequency reduction performance waste, the more serious the performance waste caused by frequency reduction due to parallel computing of the GPU device, the more it affects the performance of the GPU device, and the lower the parallel computing efficiency, and the obtained parallel computing efficiency.

[0090] Thus, the parallel computing efficiency of each GPU device at the current moment is obtained.

[0091] Step 4: Comprehensively evaluate the GPU devices through the available resource amount and the parallel computing efficiency, and calculate the resource allocation priority of each GPU device at the current moment; according to the proportion of the minimum GPU resource amount required by all queuing model tasks at the current moment in the available resource amount of all GPU devices, combined with the resource allocation priority, obtain the device allocation sequence; combine the first-fit algorithm to allocate resources to the GPU devices.

[0092] There are multiple GPU devices in the model deployment platform. The traditional first-fit algorithm is restricted by the arrangement order of GPU devices. When dynamically matching resources for queued model tasks according to the available resource amounts of GPU devices, the traditional first-fit algorithm is suitable for the peak period of the model deployment platform. When the model deployment platform is not in the peak period, it is easy to cause GPU devices with low available resource amounts to parallelly process a large number of queued model tasks, greatly increasing the competition for computing power resources of model tasks and reducing the efficiency of parallel computing. At the same time, GPU devices with high available resource amounts are idle for a long time, resulting in low utilization rate and poor balance of the large-scale GPU cluster.

[0093] Based on the above analysis, analyze whether the model deployment platform at the current moment is in the peak period of parallel computing, so as to reasonably allocate GPU device resources for queued model tasks. Specifically:

[0094] Form a feature vector with the available resource amount of each GPU device at the current moment and the efficiency of parallel computing.

[0095] Take the feature vectors of all GPU devices at the current moment as the input of the multi-criteria decision-making algorithm, and take the comprehensive evaluation score of each GPU device obtained as the resource allocation priority of each GPU device at the current moment.

[0096] In this embodiment, the Topsis algorithm is used to obtain the resource allocation priority of each GPU device at the current moment. The Topsis algorithm is a well-known technology and will not be elaborated here.

[0097] Record the sum of the minimum GPU resource amounts required by all queued model tasks at the current moment as the first sum value.

[0098] Record the sum value of the available resource amounts of all GPU devices at the current moment as the second sum value.

[0099] Take the ratio of the first sum value to the second sum value as the task resource ratio.

[0100] If the task resource ratio is greater than the preset first threshold, select the GPU devices with the efficiency of parallel computing greater than the preset second threshold, and arrange them in ascending order according to the resource allocation priority to form a device allocation sequence. If the task resource ratio is less than or equal to the preset first threshold, arrange all GPU devices in ascending order according to the resource allocation priority to form a device allocation sequence, where the preset first threshold is greater than the preset second threshold.

[0101] In this embodiment, the preset first threshold is set to 0.8, and the preset second threshold is set to 0.4. As other implementation manners, the implementer can set them according to the actual situation.

[0102] It should be noted that if the proportion of the task resources is greater than a preset first threshold, it indicates that the model deployment platform is at the peak period of parallel computing at the current moment. For GPU devices with a parallel computing efficiency less than or equal to a preset second threshold, it means that the execution model tasks of these GPU devices in parallel computing already have strong competition for computing power resources and frequency reduction, and it is not suitable to allocate queuing model tasks for computing anymore. Therefore, GPU devices with a parallel computing efficiency less than or equal to the preset second threshold are excluded. If the proportion of the task resources is less than the preset first threshold, it indicates that the model deployment platform is at the non-peak period of parallel computing at the current moment, and then all GPU devices are sorted to be allocated to queuing model tasks for parallel computing.

[0103] Arrange all the queuing model tasks at the current moment in descending order according to the minimum GPU resource amount they require to form a task queuing sequence;

[0104] It should be noted that by arranging the queuing model tasks in descending order, it is ensured that queuing model tasks with larger resource requirements are given priority, reducing the generation of GPU resource fragmentation.

[0105] Based on the task queuing sequence and the device allocation sequence, the first-fit algorithm is used to allocate resources to the GPU devices;

[0106] It should be noted that the first-fit algorithm is a well-known technology and will not be elaborated here.

[0107] It should be noted that by optimizing the parallel computing of the large-scale GPU cluster in the model deployment platform, it not only reduces the resource competition and frequency reduction caused by GPU resource fragmentation, but also can balance the load of GPU devices as much as possible, avoiding the situation that GPU devices with a higher available resource amount are idle when the model deployment platform is at the non-peak period of parallel computing, and can improve the efficiency of parallel computing of the large-scale GPU cluster.

[0108] It should be understood that although Figure 1 the steps in the flowchart of Figure 1 are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,

[0109] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0110] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation to the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made. Therefore, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present application without departing from the content of the technical solution of the present application all belong to the protection scope of the technical solution of the present application.

Claims

1. An optimization method for large-scale parallel computing based on GPU, characterized in that, The method includes the following steps: Obtain in real time the minimum GPU resource amount required for each queuing model task at the current moment, as well as the load, batch size, and number of kernel functions of each execution model task in each GPU device; obtain in real time the operating power, working frequency, and available resource amount of each GPU device at each moment; Conduct a comprehensive evaluation and analysis of the execution model tasks through the load, batch size, and number of kernel functions of each execution model task, calculate the comprehensive evaluation score of each execution model task in each GPU device at the current moment, and based on the comprehensive evaluation score, obtain each high-demand task; Analyze the deviation of the comprehensive evaluation scores of all high-demand tasks in each GPU device, and determine the chain contention weight of each GPU device at the current moment. The chain contention weight of the nth GPU device at the current moment is calculated as follows: , where is the comprehensive evaluation score of the mth high-demand task in the nth GPU device at the current moment , is the segmentation threshold, is the number of all high-demand tasks in the nth GPU device at the current moment , and is the preset weight value; Based on the comprehensive evaluation scores of all execution model tasks in each GPU device, combined with the chain contention weight, determine the parallel contention intensity of each GPU device at the current moment. The parallel contention intensity of the nth GPU device at the current moment is calculated as follows: , where is the sum of the comprehensive evaluation scores of all execution model tasks in the nth GPU device at the current moment , and is the number of all execution model tasks in the nth GPU device at the current moment ; Analyze the deviation degree of the operating power and working frequency of each GPU device at different moments to obtain the down-frequency performance waste degree of each GPU device at the current moment; fuse the parallel contention intensity and the down-frequency performance waste degree to determine the parallel computing efficiency of each GPU device at the current moment; Conduct a comprehensive evaluation of the GPU devices through the available resource amount and the parallel computing efficiency, calculate the resource allocation priority of each GPU device at the current moment; according to the proportion of the minimum GPU resource amount required by all queuing model tasks in the available resource amount of all GPU devices, combine with the resource allocation priority to obtain the device allocation sequence; combine with the first-fit algorithm to allocate resources to the GPU devices.

2. The optimization method for large-scale parallel computing based on GPU according to claim 1, characterized in that, The calculation of the comprehensive evaluation score of each execution model task in each GPU device at the current moment includes: Form the execution task vector of each execution model task in each GPU device at the current moment by the load, batch size, and number of kernel functions of each execution model task in each GPU device at the current moment; Adopt a multi-criteria decision-making algorithm to conduct a comprehensive evaluation of the execution task vectors of all execution model tasks in all GPU devices at the current moment, and obtain the comprehensive evaluation score of each execution model task in each GPU device at the current moment.

3. A large-scale parallel computing optimization method based on GPU according to claim 1, characterized in that, The obtaining of each high-demand task includes: Adopt a threshold segmentation algorithm to obtain the segmentation threshold of the comprehensive evaluation scores of all execution model tasks in all GPU devices at the current moment; Record all execution model tasks with the comprehensive evaluation score greater than the segmentation threshold as each high-demand task.

4. A method for optimizing large-scale parallel computing based on GPU according to claim 1, characterized in that The obtaining of the down-frequency performance waste degree of each GPU device at the current moment includes: Use the operating power of each GPU device at multiple moments before the current moment as the input of the prediction model to obtain the power prediction value of each GPU device at the next moment; Form the working frequency sequence with the working frequencies of each GPU device at multiple moments before the current moment, and obtain the minimum value points in the working frequency sequence; Obtain the rated power and the maximum working frequency of the GPU device; calculate the difference between the power prediction value and the rated power, and record it as the power difference; the calculation result of the exponential function with the natural constant as the base and the power difference as the exponent is recorded as the relative power difference; Record the sum of the differences between the maximum working frequency and the working frequencies corresponding to all minimum value points in the working frequency sequence as the relative frequency difference; The down-frequency performance waste degree is the product of the relative power difference and the relative frequency difference.

5. The large-scale parallel computing optimization method based on GPU according to claim 1, characterized in that The parallel computing efficiency of each GPU device at the current moment is the normalized result of the reciprocal of the product of the parallel contention intensity and the downclocking performance waste degree.

6. The large-scale parallel computing optimization method based on GPU according to claim 1, characterized in that The resource allocation priority of each GPU device at the current moment includes: Combining the available resource amount and the parallel computing efficiency of each GPU device at the current moment to form a feature vector; using the feature vectors of all GPU devices at the current moment as the input of a multi-criteria decision-making algorithm, and taking the comprehensive evaluation score of each GPU device obtained as the resource allocation priority of each GPU device at the current moment.

7. An optimization method for large-scale parallel computing based on GPU according to claim 1, characterized in that The obtaining of the device allocation sequence includes: Denoting the sum of the minimum GPU resource amounts required by all queuing model tasks at the current moment as the first sum value; denoting the sum value of the available resource amounts of all GPU devices at the current moment as the second sum value; taking the ratio of the first sum value to the second sum value as the task resource occupancy ratio. If the task resource occupancy ratio is greater than a preset first threshold, select the GPU devices with a parallel computing efficiency greater than a preset second threshold, and arrange them in ascending order according to the resource allocation priority to form a device allocation sequence; otherwise, arrange all GPU devices in ascending order according to the resource allocation priority to form a device allocation sequence, where the preset first threshold is greater than the preset second threshold.

8. A method for optimizing large-scale parallel computing based on GPU according to claim 1, characterized in that The resource allocation to the GPU devices includes: Arranging all queuing model tasks at the current moment in descending order according to the minimum GPU resource amount required by them to form a task queuing sequence. Based on the task queuing sequence and the device allocation sequence, use the first-fit algorithm to allocate resources to the GPU devices.

Citation Information

Patent Citations

  • Computing power resource allocation method oriented to large model parallel processing

    CN118819858A

  • Computing power resource allocation control method and device, electronic equipment and storage medium

    CN119440808A