A control system and control method for computing unit transmission and load
By introducing a control system for computing unit transmission and load in the GPU system, the PID controller is used to adjust the resource utilization rate of the computing unit, the problem of mismatch in resource requirements of GPUs in different application scenarios is solved, and more efficient computing performance and energy efficiency are achieved.
Patent Information
- Application Number
- CN202111249142.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-29
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-01-29
AI Technical Summary
When facing different application scenarios, how can GPUs flexibly meet the transmission and computing resource needs of different scenarios based on fixed hardware configurations, especially in the context of the continuous and rapid development of data and computing volume.
A control system for the transmission and load of the computing unit is provided. According to the transmission occupancy rate and load rate transmitted by the computing unit, and combined with pre-configured scheduling rules, the scheduling strategies of each computing unit are determined, including adjusting the calculation accuracy, working parameters and the associated number of computing subunits.
Through the implementation of the regulation system, the computing speed of the data center can be accelerated, the computing quality can be improved, or the energy consumption can be reduced, the efficiency can be improved, and the system performance can be optimized while maintaining the computing speed.
Smart Images

Figure CN114138456B_ABST
Abstract
Description
Technical Field
[0001] This article belongs to the field of chip technology, and specifically relates to a control system and control method for computing unit transmission and load. Background Art
[0002] The application scenarios of GPU have been evolving since the birth of GPU. Early GPUs were used to load and render 2D graphics calculations. Today, GPUs can also be used for video processing, natural language processing, high-performance computing, massive data processing, etc. They are used in the Internet, industry, finance, government affairs, sports, medical care, scientific research, autonomous driving and other fields. In addition to the objective needs of people's production and life, the driving force behind the rapid development of GPU applications is the development of three key factors, namely algorithms, computing power and data.
[0003] Algorithms and data reflect the inherent properties of everything in the world, so GPU algorithms and models also present different characteristics. Some algorithm models are very small and have few parameters, while others are very large. For example, OpenAI released GPT-2 in 2019, which has 1.5 billion parameters. This is the first nonlinear programming model with more than 1 billion parameters. In 2020, OpenAI released GPT-3 worldwide, with 175 billion parameters. In terms of computing power, the computing power of current mainstream GPUs has reached the level of processing trillions of operations per second, or even higher.
[0004] The value of GPU is to use computing power to help CPU accelerate computing and meet the needs of scene applications. From the microscopic process of each accelerated computing, the data is transmitted from the CPU to the GPU through the PCIe protocol, and the GPU transmits the result to the CPU after the calculation is completed. This process can be simplified as transmission->computation->transmission. Transmission and calculation are the two most important steps in GPU operation. To meet the time requirements of calculation, there needs to be a good coordination between data transmission to the GPU and data processing by the GPU. The GPU must have both powerful computing power and sufficient data throughput to enable the powerful computing power to be exerted, which brings about the problem of matching transmission speed and processing speed.
[0005] GPUs will face multiple application scenarios, even at the same time. The algorithms and data volumes of different application scenarios are also different. Some algorithms have large computational volumes but small data volumes, while others have the opposite. This means that different applications have very different demand ratios for transmission and computing resources. For GPUs, how to flexibly meet the needs of different scenarios based on fixed hardware configurations is an inevitable problem that GPU chips will face as data volumes and computing volumes continue to grow at a high speed in the future. Summary of the invention
[0006] In view of the above problems in the prior art, the purpose of this article is to provide a control system and method for computing unit transmission and load, which can help chip optimization and improve system performance.
[0007] Specifically, this document provides a computing unit transmission and load control system, the system comprising: an uplink device and a plurality of computing units communicating with the uplink device, wherein:
[0008] The calculation unit is used to collect the transmission occupancy rate and load rate between the calculation unit and the uplink device, and transmit the collected transmission occupancy rate and load rate to the uplink device;
[0009] The uplink device is provided with a computing unit management module, which includes a PID controller. The PID controller determines the scheduling strategy of each computing unit according to the transmission occupancy rate and load rate of each computing unit transmission received and the pre-configured computing unit scheduling rules, and regulates each computing unit according to the scheduling strategy of each computing unit.
[0010] In an optional embodiment, the scheduling strategy includes adjusting the computing accuracy of the computing unit, adjusting the working parameters of the computing unit, and adjusting the number of computing sub-units associated with the computing unit.
[0011] In an optional embodiment, when the scheduling strategy is to adjust the calculation accuracy of the calculation unit, the PID controller is used to adjust the calculation accuracy of the target calculation unit through the local software of the target calculation unit.
[0012] In an optional embodiment, when the scheduling strategy is to adjust the working parameters of the computing unit, the PID controller is used to adjust the working parameters of the target computing unit by driving the hardware driving module of the target computing unit.
[0013] In an optional embodiment, the system further includes: a cloud management platform;
[0014] The uplink device is provided with a cloud management interface module, and the uplink device communicates with the cloud management platform through the cloud management interface module;
[0015] The uplink device is used to upload the number of computing sub-units associated with the adjustment target computing unit to the cloud management platform through the cloud management interface module;
[0016] The cloud management platform is used to adjust the number of computing sub-units associated with the target computing unit according to the received number of computing sub-units associated with the target computing unit.
[0017] In an optional embodiment, the operating parameters include: operating frequency and operating voltage.
[0018] On the other hand, the present invention provides a method for regulating computing unit transmission and load, the method being applied to the computing unit transmission and load regulation system described above, the method comprising:
[0019] Collect the transmission occupancy rate and load rate between each computing unit and the uplink device;
[0020] Determine the scheduling strategy of each computing unit according to the received transmission occupancy rate and load rate of each computing unit transmission and the pre-configured computing unit scheduling rule;
[0021] Each computing unit is regulated according to its scheduling strategy.
[0022] In an optional embodiment, determining the scheduling strategy of each computing unit according to the received transmission occupancy rate and load rate of each computing unit transmission and the pre-configured computing unit scheduling rule includes:
[0023] Determine a deviation value of each computing unit according to the received transmission occupancy rate and the corresponding load rate transmitted by each computing unit, wherein the deviation value is the difference between the transmission occupancy rate and the load rate of the same computing unit;
[0024] The deviation values of the various computing units are input into a PID controller to determine a scheduling strategy for each computing unit, wherein the scheduling strategy includes adjusting the computing accuracy of the computing unit, adjusting the working parameters of the computing unit, and adjusting the number of computing subunits associated with the computing unit.
[0025] In an optional embodiment, the PID controller includes a plurality of PID parameters, and the PID parameters include at least one of the following: a proportional adjustment coefficient, an integral adjustment coefficient, and a differential adjustment coefficient.
[0026] In an optional embodiment, the scheduling strategy of each computing unit is determined by the following formula:
[0027]
[0028] Where u(t) is the load rate of the target computing unit at time t, e(t) is the deviation value calculated at time t, and K p is the proportional adjustment coefficient, K i is the integral adjustment coefficient, K d is the differential adjustment coefficient.
[0029] Adopting the above technical scheme, a control system and control method for computing unit transmission and load described in this article, the system includes: an uplink device and multiple computing units communicating with the uplink device, wherein: the computing unit is used to collect the transmission occupancy rate and load rate between the computing unit and the uplink device, and transmit the collected transmission occupancy rate and load rate to the uplink device; the uplink device is provided with a computing unit management module, and the computing unit management module includes a PID controller, and the PID controller determines the scheduling strategy of each computing unit according to the received transmission occupancy rate, load rate and pre-configured computing unit scheduling rules transmitted by each computing unit, and regulates each computing unit according to the scheduling strategy of each computing unit. This article can speed up the overall computing speed of the data center, or improve the computing quality while the local GPU maintains the computing speed, or reduce energy consumption and improve efficiency while maintaining the computing speed.
[0030] In order to make the above and other purposes, features and advantages of this article more obvious and easy to understand, the following specifically cites preferred embodiments and describes them in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of this article or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of this article. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0032] Figure 1 A schematic diagram showing the principle of a PID controller determining a scheduling strategy provided by an embodiment of this invention;
[0033] Figure 2 A schematic diagram showing the principle of another PID controller determining a scheduling strategy provided by an embodiment of this invention;
[0034] Figure 3 A schematic diagram showing the principle of another PID controller determining a scheduling strategy provided in an embodiment of this invention is shown;
[0035] Figure 4 A schematic diagram of the steps of the control system of computing unit transmission and load in the embodiment of this article is shown;
[0036] Figure 5 A schematic diagram showing the steps of the method for regulating the transmission and load of the computing unit in the embodiment of this invention is shown;
[0037] Figure 6 Another schematic diagram of the method for controlling the transmission and load of the computing unit in the embodiment of this invention is shown;
[0038] Figure 7 A schematic diagram of the structure of a device in an embodiment of this invention is shown. DETAILED DESCRIPTION
[0039] The following will be combined with the drawings in the embodiments of this article to clearly and completely describe the technical solutions in the embodiments of this article. Obviously, the described embodiments are only part of the embodiments of this article, not all of the embodiments. Based on the embodiments of this article, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this article.
[0040] It should be noted that the terms "first", "second", etc. in the specification and claims of this article and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of this article described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, device, product or equipment that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0041] The traditional control method of GPU resource utilization is open-loop control. The system allocates application data transmission and processing requests according to the task queue, in a sequential polling or DLA (deep learning accelerator GPU, computing core) load level, without considering the matching degree of data transmission and computing requirements, such as Figure 1 As shown in the figure. When the fixed GPU hardware capability is switched in different scenarios, there will be scenarios where data transmission (occupancy rate) and data calculation (load rate) do not match. For example, when the data transmission volume is large (occupancy rate is high) and the calculation volume is small (load rate is low), the transmission occupancy rate will always be higher than the calculation load rate, and there will be a transmission bottleneck. Correspondingly, there may also be a situation where the transmission will wait for the calculation, and there will be a calculation bottleneck, which will cause resource waste and affect performance.
[0042] In order to solve the above problems, the embodiments of this specification provide a control system for computing unit transmission and load. Since the load changes of chip data transmission and data processing affect each other, they belong to the category of system changes and show similar system change laws as electronic and electrical systems. Therefore, introducing system PID control theory into chip design can help chip optimization and improve system performance.
[0043] It should be noted that the above system is suitable for efficient computing scenarios in which large data centers process massive amounts of data.
[0044] Specifically, Figure 1-4 As shown, the system includes: an uplink device and a plurality of computing units communicating with the uplink device, wherein:
[0045] The calculation unit is used to collect the transmission occupancy rate and load rate between the calculation unit and the uplink device, and transmit the collected transmission occupancy rate and load rate to the uplink device;
[0046] The uplink device is provided with a computing unit management module, which includes a PID controller. The PID controller determines the scheduling strategy of each computing unit according to the transmission occupancy rate and load rate of each computing unit transmission received and the pre-configured computing unit scheduling rules, and regulates each computing unit according to the scheduling strategy of each computing unit.
[0047] It can be understood that the uplink device can be a central processing unit (CPU, Central Processing Unit / Processor), and the computing unit can be a graphics processing unit (Graphics Processing UnitGPU, GPU), GPGPU or any other form of XPU computing unit. Correspondingly, the uplink device of the computing unit can be a CPU or other forms of network element devices. Data exchange can be achieved between the computing unit and the uplink device through a transmission protocol, which can be PCIe or other transmission interface protocols.
[0048] The uplink device can transmit the data to be calculated to the corresponding computing unit through a preset link, and the computing unit will transmit the calculation result back to the uplink device after calculating the data. A computing unit management module can be set in the uplink device, and the computing unit management module can include a PID controller. The PID controller determines the scheduling strategy of each computing unit based on the collected transmission occupancy rate of the uplink device and each computing unit, the load rate of each computing unit, and the pre-configured computing unit scheduling rules, and regulates each computing unit according to the scheduling strategy of each computing unit. Among them, the transmission occupancy rate can represent the transmission utilization rate between the corresponding computing unit and the corresponding uplink device, and the load rate can represent the ratio of the computing power used by the computing unit to calculate the data to the total computing power of the computing unit.
[0049] Specifically, the PID controller can include a proportional control unit (Proportional), an integral control unit (Integral) and a derivative control unit (Derivative). The PID controller can be understood as an algorithm. The PID controller can adjust the adjustment parameters K of these three units. p ,K i ,K d To adjust the load rate of each computing unit, the PID controller has the advantages of simple algorithm, good robustness and high reliability.
[0050] It is understandable that the pre-configured computing unit scheduling rules can be trained by intelligent models based on expert experience or historical data by technicians in this field. The computing unit scheduling rules obtained by the trained model can characterize the mapping relationship between the input transmission occupancy rate, the input load rate and the output load rate. The scheduling strategy of each computing unit can be determined by the transmission occupancy rate and load rate transmitted by each computing unit and the pre-configured computing unit scheduling rules, and the corresponding computing unit can be regulated by the corresponding computing unit's regulation strategy.
[0051] In the embodiment of the present specification, the scheduling strategy may include adjusting the computing accuracy of the computing unit, adjusting the working parameters of the computing unit, and adjusting the number of computing sub-units associated with the computing unit.
[0052] It is understandable that the data types processed by GPU mainly include integer type and floating point type. The integer types mainly include INT4, INT8, INT16, CINT32, etc., and the floating point types include FP16, FP32, FP64, etc. The data types with high precision occupy more memory space, and the multiplier and accumulation (MAC) also consumes more computing resources. The calculation accuracy adjusted in the embodiment of this specification can be within the preset threshold range of the application. For example, when the PCIe occupancy rate is less than the DLA load rate to a certain extent, the calculation accuracy can be reduced so that the DLA can process more data per unit time, thereby completing the calculation task faster and avoiding possible calculation bottlenecks. When the PCIe occupancy rate is greater than the DLA load rate to a certain extent, the calculation accuracy can be appropriately increased. Although it cannot increase the number of tasks completed by the GPU per unit time, it can improve the quality of the GPU completing the task, which is also an effective use of computing resources. It is understandable that the adjustment of the calculation accuracy of the calculation unit in the embodiment of this specification is to select the controlled object from the perspective of the quality of task completion.
[0053] Adjusting the working parameters of the computing unit may include the working frequency and the working voltage. The working parameters of the computing unit in the embodiments of the present specification may be executed by a hardware driver module in the computing unit. The load rate of the corresponding computing unit is adjusted by adjusting the working frequency or the working voltage of the computing unit to optimize the balance between the transmission occupancy rate and the load rate.
[0054] In addition, the computing speed of the GPU is related to the operating frequency, which is maintained and guaranteed by the operating voltage. The frequency affects the operating current, and the current and voltage determine the power consumption. Specifically, when the PCIe occupancy rate is less than a certain range of the DLA load rate, the GPU computing power can be increased by overclocking, thereby improving the load balance. When the PCIe occupancy rate is greater than a certain degree of the DLA load rate, the bottleneck pressure of data transmission can be alleviated by reducing the frequency, while reducing power consumption. This is to select the controlled object from the perspective of task completion energy efficiency. Therefore, the calculation accuracy of the computing unit can be adjusted by adjusting the working parameters (working frequency or working voltage) of the computing unit.
[0055] In the embodiment of the present specification, when the scheduling strategy is to adjust the calculation accuracy of the calculation unit, the PID controller can adjust the calculation accuracy of the target calculation unit through the local software of the target calculation unit.
[0056] It is understandable that the local software can be a platform software set in the target computing unit, and the platform software can be used to adjust the accuracy of the calculation data in the target computing unit. That is, when the transmission occupancy rate of the target computing unit is higher than the load rate, the transmission occupancy rate can be increased by adjusting the calculation accuracy to achieve an optimal balance between the transmission occupancy rate and the load rate.
[0057] In the embodiment of the present specification, when the scheduling strategy is to adjust the working parameters of the computing unit, the PID controller can adjust the working parameters of the target computing unit by driving the hardware driving module of the target computing unit.
[0058] It can be understood that a hardware driver module can be set in the computing unit to adjust the working parameters of the computing unit. The PID controller can communicate with the computing unit through the hardware driver module. When the scheduling strategy determined by the PID controller is to adjust the working parameters of the computing unit, the hardware driver module can adjust the working parameters of the computing unit based on the scheduling strategy determined by the PID controller. The load rate (computing power) of the computing unit can be improved by adjusting the working parameters of the computing unit. For example, the load rate of the computing unit is 50% before adjusting the working parameters, and the load rate of the computing unit is 60% after adjusting the working parameters.
[0059] In the embodiments of this specification, the system further includes: a cloud management platform;
[0060] The uplink device is provided with a cloud management interface module, and the uplink device communicates with the cloud management platform through the cloud management interface module;
[0061] The uplink device is used to upload the number of computing sub-units associated with the adjustment target computing unit to the cloud management platform through the cloud management interface module;
[0062] The cloud management platform is used to adjust the number of computing sub-units associated with the target computing unit according to the received number of computing sub-units associated with the target computing unit.
[0063] Specifically, the cloud management platform is connected to multiple upstream devices, and each upstream device communicates with the cloud management platform through a cloud management interface module. Each upstream device can communicate with multiple computing sub-units, and multiple computing sub-units together constitute a computing unit. Among them, when the computing unit is a GPU, the computing sub-unit can be a vGPU (Virtualized GPU). It can be understood that the upstream device can virtualize a physical GPU into multiple vGPU cards, each VM has an exclusive vGPU, and each vGPU is directly connected to the physical GPU.
[0064] For example, if a data center has two GPU clusters (computing units) I and II, 100 VGPUs (computing sub-units) are allocated to each GPU cluster, and 100TOPS (Tera Operations Per Second, 1TOPS means that the processor can perform one trillion (10^12) operations per second) computing power is allocated to each VGPU. The two GPU clusters process two different application scenarios respectively. The operating status of the GPU clusters is shown in the following table (the unit of VGPU number is , and the unit of real-time computing power is TOPS, the same below):
[0065] Cluster I / Scenario I PCIe transfer occupancy % DLA load factor % frequency Accuracy Number of VGPUs Real-time computing power Small load 50% 25% 1.0GHz INT8 100 25*100 Maximum load 100% 50% 1.0GHz INT8 100 50*100
[0066] Cluster II / Scene II PCIe transfer occupancy % DLA load factor % frequency Accuracy Number of VGPUs Real-time computing power Small load 25% 50% 1.0GHz INT8 100 50*100 Maximum load 50% 100% 1.0GHz INT8 100 100*100
[0067] It is understandable that different application scenarios have different algorithms and models, and the number of MACs (Multiplier and Accumulation) that each data that needs to be calculated goes through is also different. Therefore, the transmission occupancy rate and load rate of the above two GPU clusters show different rules: for GPU cluster I, there is a transmission bottleneck. When the transmission occupancy rate is 100%, the load rate (computing power) is only 50%. At this time, the real-time maximum computing power of GPU cluster I is 50% of the theoretical maximum value; for GPU cluster II, there is a computing bottleneck. When the load rate is 100%, that is, the load rate is the maximum, the computing power can reach the theoretical maximum value, but the PCIe transmission occupancy rate at this time is only 50%, and the GPU throughput capacity is not fully utilized.
[0068] When the PID controller set in the upstream device detects a positive signal of the difference between the PCIe transmission occupancy rate and the DLA load rate of GPU cluster I, that is, PCIe transmission occupancy rate - DLA load rate = 100% - 50% = 50%, after calculation by the PID controller, it can be determined that there are 50 VGPUs in GPU cluster I that can realize data calculation. The PID controller can send a request to the cloud management platform through the cloud management interface module that GPU cluster I can release 50 VGPUs.
[0069] The PID controller set in the upstream device detects a negative signal of the difference between the PCIe transmission occupancy rate and the DLA load rate of GPU cluster II, that is, PCIe transmission occupancy rate - DLA load rate = 50% - 100% = -50%. After calculation by the PID controller, it can be determined that 100 VGPUs are still needed in GPU cluster II for data transmission to achieve the maximum PCIe transmission occupancy rate and DLA load rate. The PID controller can send a request to the cloud management platform to add 100 VGPUs through the cloud management interface module.
[0070] After scheduling by the cloud management platform, GPU cluster I can release 50 VGPUs and release 50 VGPUs for GPU cluster II to use. The results are shown in the following table:
[0071] Cluster I / Scenario I PCIe occupancy % DLA load factor % frequency Accuracy Number of VGPUs Real-time computing power Small load 50% 50% 1.0GHz INT8 50 50*50 Maximum load 100% 100% 1.0GHz INT8 50 100*50
[0072] Cluster II / Scene II PCIe occupancy % DLA load factor % frequency Accuracy Number of VGPUs Real-time computing power Small load 25% 33.3% 1.0GHz INT8 150 33.3*150 Maximum load 75% 100% 1.0GHz INT8 150 100*150
[0073] It can be seen that after adjustment by the cloud management platform, the load difference between the PCIe transmission occupancy rate and the DLA load rate of GPU cluster II still exists. At this time, the working frequency of each vGPU in GPU cluster II can be adjusted by adjusting the working parameters of the computing unit, as shown in the following table:
[0074] Cluster II / Scene II PCIe occupancy % DLA load factor % frequency Accuracy Number of VGPUs Real-time computing power Small load 27.5% 33.3% 1.1GHz INT8 150 36.6*150 Maximum load 82.5% 100% 1.1GHz INT8 150 110*150
[0075] Overclocking by 10% theoretically increases computing power by 10%. At the same DLA load rate, the amount of data processed per unit time also increases by 10%. Correspondingly, the PCIe occupancy rate will increase by 10%, which can further achieve the goal of balancing transmission and computing.
[0076] After adjusting the number of computing sub-units associated with the computing unit and adjusting the working parameters of the computing unit, the load difference between the PCIe transmission occupancy rate and the DLA load rate of the GPU cluster II still exists. The P controller and the I controller will continue to have output regulation strategies. If the application scenario permits, the control strategy determined by the PID controller can be to adjust the calculation accuracy of the computing unit. By adjusting the calculation accuracy of some data, the transmission capacity is further released, so that the transmission and computing capabilities are further balanced.
[0077] After the PID controller is used to adjust each computing sub-unit in the two GPU clusters in the embodiment of this specification, the state changes of the two GPU clusters at maximum load are compared with the data before the adjustment as shown in the following table:
[0078] Cluster I / Scenario I Prior art PID Control PCIe occupancy 100% 100% DLA load factor 50% 100% VGPU Scaling 100 50 Real-time computing power 5000TOPS 5000TOPS
[0079] Cluster II / Scene II Prior art PID Control PCIe occupancy 50% 82.5% DLA load factor 100% 100% VGPU Scaling 100 150 Real-time computing power 10000TOPS 16500TOPS
[0080] When two GPU clusters form a data center, it can be seen that after the PID controller is used to adjust each computing sub-unit in the two GPU clusters, the data before and after the adjustment are compared as shown in the following table:
[0081]
[0082] It can be seen from the above table that the computing unit transmission and load control system provided by this application, compared with the prior art, not only increases the original computing power from 15000TOPS to 21500TOPS, but also improves the energy efficiency ratio by about 33%. This application uses the transmission resource occupancy rate and load rate between the computing unit and the upstream device as the control source. The PID controller can specify different control strategies based on the control source to achieve precise control of the GPU cluster scale, GPU frequency voltage, GPU computing accuracy, etc., so that idle transmission resources and computing resources flow intelligently and orderly to the corresponding GPU cluster applications with higher load rates, thereby improving the overall computing speed of the data center, and can also improve the computing quality while maintaining the computing speed of the GPU. It can also reduce system energy consumption and improve efficiency while maintaining the computing speed, and realize the optimization of each chip in the server and improve system performance.
[0083] Based on the computing unit transmission and load control system provided above, the embodiments of this specification also provide a computing unit transmission and load control method, which can realize the control of each computing unit in the computing unit transmission and load control system.
[0084] Specifically, Figure 5 The figure is a schematic diagram of the steps of the method for regulating the transmission and load of the computing unit in the embodiment of this article. This specification provides the method operation steps described in the embodiment or flowchart, but may include more or fewer operation steps based on conventional or non-creative labor. The order of steps listed in the embodiment is only one way of executing the steps among many orders, and does not represent the only order of execution. When the system or device product is executed in practice, it can be executed in the order or in parallel according to the method shown in the embodiment or the drawings. Specifically, Figure 4 As shown, the method may include:
[0085] S101: Collecting the transmission occupancy rate and load rate between each computing unit and the uplink device;
[0086] S102: determining a scheduling strategy for each computing unit according to the received transmission occupancy rate and load rate of each computing unit transmission and a pre-configured computing unit scheduling rule;
[0087] In a preferred embodiment, Figure 6 As shown, Figure 6 1 is another step schematic diagram of the method for regulating the transmission and load of the computing unit in the embodiment of this invention, wherein the scheduling strategy of each computing unit is determined according to the transmission occupancy rate and load rate of each computing unit transmission received and the pre-configured computing unit scheduling rule, including:
[0088] S201: determining a deviation value of each computing unit according to the received transmission occupancy rate and the corresponding load rate transmitted by each computing unit, wherein the deviation value is a difference between the transmission occupancy rate and the load rate of the same computing unit;
[0089] S202: Input the deviation value of each computing unit into the PID controller to determine the scheduling strategy of each computing unit, wherein the scheduling strategy includes adjusting the calculation accuracy of the computing unit, adjusting the working parameters of the computing unit, and adjusting the number of computing sub-units associated with the computing unit.
[0090] Specifically, the PID controller can collect the transmission occupancy rate and the corresponding load rate of each computing unit in real time.
[0091] Among them, the PID controller may include a proportional control unit (Proportional), an integral control unit (Integral) and a derivative control unit (Derivative). Figure 2 As shown, the P controller is a proportional controller, which is used to control the deviation value of the system in proportion. The deviation value reflects the deviation between the PCIe occupancy rate and the DLA load rate. When the deviation occurs, the P controller immediately produces a control effect to reduce the error. When the deviation disappears, the control effect also disappears. Specifically, the P controller can adjust the controlled object (computing unit) through the following formula:
[0092] u(t)=K p e(t)
[0093] Among them, P P is the occupancy rate of the PCIe channel, P D is the load rate of DLA, the deviation value of the controlled object is e, e=P P -P D , K p is the proportional adjustment coefficient of the P controller.
[0094] It can be understood that the P controller can be used to eliminate the difference between the GPU's current transmission (occupancy) and the calculated load rate. When a deviation value exists, the occupancy rate and load rate can be roughly balanced through adjustment of the P controller, that is, there will be a steady-state error between the occupancy rate and the load rate.
[0095] The I controller is an integral controller, which can be used to eliminate the deviation value of the computing unit caused by the imbalance of resource utilization in the historical time. That is, when the computing unit has a deviation value, the deviation value of the computing unit can be integrated to make the output continue to increase or decrease until the deviation value is zero, the integration stops, and the output no longer changes. The integration can reflect the accumulation of deviations over a period of time. Specifically, the I controller can adjust the controlled object (computing unit) through the following formula:
[0096]
[0097] Among them, K i is the integral adjustment coefficient of the I controller.
[0098] It is understandable that the deviation value can represent data backlog, which indicates the existence of computing bottlenecks, that is, the computing resource utilization rate is significantly higher than the transmission resource utilization rate within the historical preset cycle time. The I controller will adjust the controlled object according to the degree of data backlog to eliminate the phenomenon of data backlog. That is, the I controller can eliminate static error and improve the control accuracy of the system.
[0099] The D controller is a differential controller. The D controller can be used to predict the trend of the deviation between GPU transmission and computing resource utilization, reflect the rate of change of the deviation value, and introduce an effective correction value into the system before the deviation value becomes too large, thereby speeding up the system's action speed and reducing the adjustment time. Specifically, the D controller can adjust the controlled object (computing unit) through the following formula:
[0100]
[0101] Among them, K d is the differential adjustment coefficient of the D controller.
[0102] In practical applications, if the acceleration of the deviation value increases, it indicates that the deviation value has a tendency to increase. Even if the current deviation value is still small, the system will increase the adjustment command in advance. On the contrary, if the acceleration of the deviation value is negative, even if there is still a certain deviation, the system will weaken the adjustment command in advance. The D controller is used to reduce oscillations in the control process.
[0103] It can be understood that the D controller's regulation of the computing unit is time-sensitive. The proportional controller controls the current deviation value of the system, which is the control of the "present". The integral controller controls the history of the system deviation value, which is the control of the "past". The differential controller reflects the changing trend of the system deviation value, which is the control of the "future".
[0104] Specifically, the PID controller can be adjusted for the controlled object (computation unit) by the following formula:
[0105]
[0106] Among them, u(t) is a continuous function over time, while the PCIe occupancy rate and DLA load rate in the GPU are reported by the computing unit in the sampling period or actively read by the PID controller. Therefore, u is actually in discrete form. For the discrete form, the PID controller can adjust the controlled object (computing unit) through the following formula:
[0107]
[0108] It can be understood that the adjustment of the controlled object can be to adjust the calculation accuracy of the calculation unit, the DLA operating frequency, and the number of sub-computing units, etc. The calculation accuracy, DLA operating frequency, and the number of sub-computing units can be used as controlled objects alone or in combination.
[0109] It should be noted that the input of the PID controller can be discrete, and accordingly, the adjustment signal can also be discrete. In practical applications, u(k) in different intervals can be matched to different inputs of the GPU controlled object by using interval control according to the characteristics of the controlled object. For example, as shown in the following table, the following table is a control strategy comparison table of u(k) and the calculation accuracy of the calculation unit, the frequency voltage of the calculation unit, and the number of sub-calculation units.
[0110] U(k) DLA calculation accuracy DLA Frequency Voltage Number of sub-computing units [-0.9,-0.7) 2nd gear downshift +300MHz +4 [-0.7,-0.5) Downshift +200MHz +3 [-0.5,-0.3) Keep it the same +100MHz +2 [-0.3,-0.1) Keep it the same Keep it the same +1 [-0.1,0.1] Keep it the same Keep it the same Keep it the same (0.1,0.3] Keep it the same Keep it the same -1 (0.3,0.5] Keep it the same -100MHz -2 (0.5,0.7] Up a gear -200MHz -3 (0.7,0.9] 2nd gear -300MHz -4
[0111] It can be understood that the adjustment strategy is pre-set. When u(k) falls within different value ranges, the controlled object can be adjusted individually or in combination to achieve the purpose of balancing transmission and processing. In actual use, the adjustment combination can be set based on comprehensive considerations such as GPU product positioning, functional characteristics, software implementation, and application scenarios.
[0112] For example, when the GPU system does not require high balancing accuracy but is sensitive to transient disturbances, PD control can also be used. The control structure diagram can be found in Figure 3 :
[0113] The PD controller can be adjusted for the controlled object (computing unit) by the following formula:
[0114] u(k)=K p e(k)+K d (e(k)-e(k-1))
[0115] For example, when the GPU system is not sensitive to instantaneous disturbances but requires high control accuracy, PI control can also be used. The control structure diagram can be found in Figure 4 :
[0116] The PI controller can be adjusted for the controlled object (computation unit) by the following formula:
[0117]
[0118] Specifically, the PID controller may be a pre-trained algorithm, that is, in practical applications, the PID controller may be pre-set with PID parameters, and the PID parameters may be related to the PCIe bandwidth occupancy rate and the GPU computing unit DLA load rate.
[0119] The occupied time includes the actual data transmission time and the idle time.
[0120] The occupied time includes the actual data processing time and the idle time, and MAC stands for multiplier-accumulator.
[0121] It can be seen that the PCIe occupancy rate and the DLA load rate are linearly related to the data volume. Without considering the loss of other software and hardware in the system, there is also a linear relationship between the PCIe occupancy rate and the DLA load rate in theory, that is, the PID parameters in the PID controller can be selected according to the deviation value, and different deviation values can correspond to different PID parameters.
[0122] The PID parameters can be determined in the following way: using fuzzy rules to adaptively adjust the PID parameters of the PID controller according to the deviation value and the deviation change rate, and outputting the change amount of the PID parameters; using an expert knowledge base to obtain the initial value of the PID parameters according to the deviation value and the deviation change rate.
[0123] Preferably, the PID parameters include: proportional adjustment coefficient, integral adjustment coefficient and differential adjustment coefficient. For example, taking the deviation value e and the relevant characteristic quantity of the deviation value as input signals, a set of changes in PID parameters, i.e., the change in the proportional adjustment coefficient ΔK, is obtained through the process of fuzzification, fuzzy reasoning and defuzzification. p , the change of integral adjustment coefficient ΔK i And the change of differential adjustment coefficient ΔK d .
[0124] Preferably, the initial values of the corresponding PID parameters can be generated according to the deviation value and the deviation change rate according to the following regular mathematical model;
[0125] |e|≥ε 1 When Kp 0 =Kp 01 ,Ki 0 =Ki 01 ,Kd 0 =Kd 01 ;
[0126] ε 2 ≤|e|<ε 1 When Kp 0 =Kp 02 ,Ki 0 =Ki 02 ,Kd 0 =Kd 02
[0127] |e|<ε 2 And |ec|≥δ 1 When Kp 0 =Kp 03 ,Ki 0 =Ki 03 ,Kd 0 =Kd 03 ;
[0128] |e|<ε 2 and |ec|<δ 1 When Kp 0 =Kp 03 ,Ki 0 =Ki 03 ,Kd 0 =Kd 04
[0129] Where e is the calculated deviation value, ec is the calculated deviation change rate; ε 1 , ε 2 The first error level and the second error level, δ 1 is the first error change level value, Kp0, Ki0, Kd0 are the initial values of the proportional adjustment coefficient Kp, the integral adjustment coefficient Ki and the differential adjustment coefficient Kd respectively; Kp01, Kp02, Kp03 are the preset first proportional adjustment value, the second proportional adjustment value and the third proportional adjustment value respectively; Ki01, Ki02, Ki03 are the preset first integral adjustment value, the second integral adjustment value and the third integral adjustment value respectively; Kd01, Kd02, Kd03, Kd04 are the preset first differential adjustment value, the second differential adjustment value, the third differential adjustment value and the fourth differential adjustment value respectively; and Kp 01 ≈Kp 03 >Kp 02 ,Kd 01 ≈Kd 02 ≈Kd 04 <Kd 03 ,0=Ki 01 <Ki 02 <Ki 03 ,ε 1 >ε 2 >0,δ 1 >0.
[0130] In practical applications, when the PID controller regulates each computing unit, the PID parameter value can be obtained according to the initial value and change of the PID parameter, and the scheduling strategy of each computing unit can be calculated according to the PID parameter value.
[0131] Specifically, the scheduling strategy of each computing unit is determined by the following formula:
[0132]
[0133] Where u(t) is the load rate of the target computing unit at time t, e(t) is the deviation value calculated at time t, and K p is the proportional adjustment coefficient, K i is the integral adjustment coefficient, K d is the differential adjustment coefficient.
[0134] It can be understood that u(t) is continuously output until the difference between the load rate and the transmission occupancy rate meets a preset threshold value and then stops being output.
[0135] S103: regulating each computing unit according to its scheduling strategy.
[0136] It can be understood that the executor of the method is a PID controller, which can collect the transmission occupancy rate and load rate between the computing unit and the uplink device through the local software in the computing unit, and determine the scheduling strategy of each computing unit according to the transmission occupancy rate, load rate and pre-configured computing unit scheduling rules, wherein the scheduling strategy may include adjusting the calculation accuracy of the computing unit, adjusting the working parameters of the computing unit and adjusting the number of computing sub-units associated with the computing unit, and then controlling each computing unit through the determined control strategy to achieve balanced optimization of the transmission occupancy rate and load rate.
[0137] It should be noted that the PID controller can adjust the system behavior of the controlled object according to the deviation between the PCIe occupancy rate and the DLA load rate to improve the balance between GPU transmission resources and computing resources, so as to optimize the system performance of the computing unit. In addition to the amount of data, the factors that can affect the DLA load rate include the number of concurrent operations, calculation accuracy, and calculation frequency. Therefore, the scheduling strategy can be to adjust the number of concurrent operations, calculation accuracy, calculation frequency, and number of sub-computing units in each sub-computing unit.
[0138] For example, when the controlled object is the size of the GPU cluster, the computing speed of the data center can be increased by adjusting the number of VGPUs (sub-computing units). For example, when the difference between the PCIe occupancy rate and the DLA load rate is less than the preset deviation threshold, that is, the data received by the GPU may be waiting in line for processing, and the number of VGPUs (Virtual GPU means virtual GPU) can be dynamically increased to complete the computing task faster and avoid possible computing bottlenecks. When the difference between the PCIe occupancy rate and the DLA load rate is greater than the preset deviation threshold, that is, the DLA may be idle waiting for data transmission, and the number of VGPUs can be dynamically reduced to avoid the occurrence of transmission bottlenecks, and the released VGPUs can be placed in the GPU cluster with computing bottlenecks. Through the system's automatic adjustment of peak-shaving and valley-filling, the computing speed of the entire data center can be effectively improved. This is to select the controlled object from the perspective of task completion speed.
[0139] Further, if Figure 7As shown, a device provided in an embodiment of this invention may include the PID controller provided above. Optionally, the computer device 802 may include one or more processors 804, such as one or more central processing units (CPUs), each of which may implement one or more hardware threads. The computer device 802 may also include any memory 806 for storing any kind of information such as code, settings, data, etc. Non-limitingly, for example, the memory 806 may include any one or more combinations of the following: any type of RAM, any type of ROM, flash memory device, hard disk, optical disk, etc. More generally, any memory may use any technology to store information. Further, any memory may provide volatile or non-volatile retention of information. Further, any memory may represent a fixed or removable component of the computer device 802. In one case, when the processor 804 executes an associated instruction stored in any memory or a combination of memories, the computer device 802 may perform any operation of the associated instruction. The computer device 802 also includes one or more drive mechanisms 808 for interacting with any memory, such as a hard disk drive mechanism, an optical disk drive mechanism, etc.
[0140] The computer device 802 may also include an input / output module 810 (I / O) for receiving various inputs (via input devices 812) and for providing various outputs (via output devices 814). A specific output mechanism may include a presentation device 816 and an associated graphical user interface (GUI) 818. In other embodiments, the input / output module 810 (I / O), the input device 812, and the output device 814 may not be included, and the computer device 802 may be used as a computer device in a network. The computer device 802 may also include one or more network interfaces 820 for exchanging data with other devices via one or more communication links 822. One or more communication buses 824 couple the components described above together.
[0141] The communication link 822 may be implemented in any manner, for example, through a local area network, a wide area network (e.g., the Internet), a point-to-point connection, etc., or any combination thereof. The communication link 822 may include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc. governed by any protocol or combination of protocols.
[0142] Corresponds to Figure 5-Figure 6 The method in this article also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are executed.
[0143] The embodiment of the present invention also provides a computer readable instruction, wherein when the processor executes the instruction, the program therein causes the processor to execute the following Figures 5 and 6 The method shown.
[0144] It should be understood that in the various embodiments of this document, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this document.
[0145] It should also be understood that in the embodiments of this article, the term "and / or" is only a description of the association relationship of the associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0146] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this article.
[0147] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0148] In the several embodiments provided herein, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, or can be electrical, mechanical or other forms of connection.
[0149] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments of this article.
[0150] In addition, each functional unit in each embodiment of this invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of software functional unit.
[0151] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this article is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of this article. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0152] Specific embodiments are used in this article to illustrate the principles and implementation methods of this article. The description of the above embodiments is only used to help understand the methods and core ideas of this article. At the same time, for general technicians in this field, according to the ideas of this article, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on this article.
Claims
1. A computing unit transmission and load control system, characterized in that: The system comprises: an uplink device and a plurality of computing units communicating with the uplink device, wherein: The calculation unit is used to collect the transmission occupancy rate and load rate between the calculation unit and the uplink device, and transmit the collected transmission occupancy rate and load rate to the uplink device; The uplink device is provided with a computing unit management module, the computing unit management module includes a PID controller, the PID controller determines the scheduling strategy of each computing unit according to the transmission occupancy rate and load rate of each computing unit transmission received and the pre-configured computing unit scheduling rule, and regulates each computing unit according to the scheduling strategy of each computing unit; The scheduling strategy includes adjusting the calculation accuracy of the calculation unit, adjusting the working parameters of the calculation unit, and adjusting the number of calculation subunits associated with the calculation unit; When the scheduling strategy is to adjust the calculation accuracy of the calculation unit, the PID controller is used to adjust the calculation accuracy of the target calculation unit through the local software of the target calculation unit; When the scheduling strategy is to adjust the working parameters of the computing unit, the PID controller is used to adjust the working parameters of the target computing unit by driving the hardware driving module of the target computing unit; The system also includes: a cloud management platform; The uplink device is provided with a cloud management interface module, and the uplink device communicates with the cloud management platform through the cloud management interface module; The uplink device is used to upload the number of computing sub-units associated with the adjustment target computing unit to the cloud management platform through the cloud management interface module; The cloud management platform is used to adjust the number of computing subunits associated with the target computing unit according to the received number of computing subunits associated with the target computing unit; The above control system uses a method for controlling the transmission and load of a computing unit, the method comprising: Collect the transmission occupancy rate and load rate between each computing unit and the uplink device; Determine the scheduling strategy of each computing unit according to the received transmission occupancy rate and load rate of each computing unit transmission and the pre-configured computing unit scheduling rule; Regulate each computing unit according to its scheduling strategy; The step of determining the scheduling strategy of each computing unit according to the received transmission occupancy rate and load rate of each computing unit transmission and the pre-configured computing unit scheduling rule includes: Determine a deviation value of each computing unit according to the received transmission occupancy rate and the corresponding load rate transmitted by each computing unit, wherein the deviation value is the difference between the transmission occupancy rate and the load rate of the same computing unit; Inputting the deviation value of each computing unit into a PID controller to determine a scheduling strategy for each computing unit, wherein the scheduling strategy includes adjusting the calculation accuracy of the computing unit, adjusting the working parameters of the computing unit, and adjusting the number of computing subunits associated with the computing unit; The scheduling strategy of each computing unit is determined by the following formula: Where u(t) is the load rate of the target computing unit at time t, e(t) is the deviation value calculated at time t, and K p is the proportional adjustment coefficient, K i is the integral adjustment coefficient, K d is the differential adjustment coefficient.
2. The control system for computing unit transmission and load according to claim 1, characterized in that: The operating parameters include: operating frequency and operating voltage.
3. The control system for computing unit transmission and load according to claim 2, characterized in that: The PID controller includes a plurality of PID parameters, and the PID parameters include at least one of the following: a proportional adjustment coefficient, an integral adjustment coefficient, and a differential adjustment coefficient.
Citation Information
Patent Citations
Cloud service resource collaborative optimization scheduling method, system, medium and equipment
CN113411369A