A method, product, device, server, and medium for adjusting computing device parameters
By acquiring and analyzing the GPU's energy efficiency gradient map and adjusting the batch size and core frequency, the problem of how to balance GPU's energy efficiency and latency in the cloud computing era is solved, and the optimal operating state of the device is achieved between performance and energy efficiency.
Patent Information
- Application Number
- CN202411976965.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-12-31
AI Technical Summary
In the age of cloud computing, how to find the best batch size and core frequency configuration to balance GPU energy efficiency and latency and achieve optimal operating state between performance and energy efficiency.
By obtaining the current energy efficiency gradient map of the computing device, determining the target position point and performing adjustments, ensure that the computing device meets the delay and energy efficiency balance conditions in the new configuration.
A dynamic balance between performance and energy efficiency of computing devices is achieved, ensuring optimal operating status of devices under different workloads and operating environments.
Smart Images

Figure CN119376960B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of servers, and particularly to a method, product, device, server and medium for adjusting computing device parameters. Background Art
[0002] In the era of cloud computing, the demand for GPU (Graphics Processing Unit) - accelerated inference is increasing continuously. Optimizing the energy efficiency and performance of GPUs has become crucial. The batch size and kernel frequency are key parameters for adjusting GPU performance. Adjusting the batch size affects the memory utilization rate and parallel processing efficiency of the GPU, while adjusting the GPU kernel frequency directly affects the processing speed and energy consumption. Therefore, finding the optimal batch size and kernel frequency configuration to balance the energy efficiency and latency of the GPU is a problem that those skilled in the art need to solve currently. Summary of the Invention
[0003] The purpose of the present invention is to provide a method, product, device, server and medium for adjusting computing device parameters, which can automatically adjust the configuration according to the actual operating conditions of the computing device, adapt to different workloads and operating environments, achieve a dynamic balance between energy efficiency and latency, and ensure the optimal operating state of the computing device between performance and energy efficiency.
[0004] To solve the above - mentioned technical problems, the present invention provides a method for adjusting computing device parameters, including: obtaining the current energy - efficiency gradient map of the computing device; the position points of the current energy - efficiency gradient map are used to represent the comprehensive performance value of energy efficiency and latency of the computing device under the corresponding combination of batch size and kernel frequency; determining a target position point on the current energy - efficiency gradient map according to the current batch size and current kernel frequency of the computing device, and performing a current adjustment operation on the computing device based on the comprehensive performance value of energy efficiency and latency corresponding to the target position point; in response to the new batch size and new kernel frequency after completing the current adjustment operation satisfying the latency - energy - efficiency balance condition, controlling the computing device to continuously operate at the new batch size and the new kernel frequency.
[0005] Optionally, the process of obtaining the current energy efficiency gradient map of the computing device includes: determining a reference batch size range and a reference core frequency range corresponding to the computing device according to the specifications and / or performance limitation conditions of the computing device; selecting a plurality of batch sizes within the reference batch size range; selecting a plurality of core frequencies within the reference core frequency range; arranging and combining the plurality of batch sizes and the plurality of core frequencies to generate a plurality of combinations of batch sizes and core frequencies; calculating the comprehensive performance value of energy efficiency and latency of the computing device under each combination of batch size and core frequency; and filling the position points of a preset two-dimensional grid based on the plurality of combinations of batch sizes and core frequencies and their corresponding comprehensive performance values of energy efficiency and latency to obtain the current energy efficiency gradient map.
[0006] Optionally, the process of calculating the comprehensive performance value of energy efficiency and latency of the computing device under each combination of batch size and core frequency includes: for each combination of batch size and core frequency, obtaining the energy efficiency data and latency data of the computing device under the combination of batch size and core frequency, and calculating the comprehensive performance value of energy efficiency and latency of the combination of batch size and core frequency based on the energy efficiency data and the latency data.
[0007] Optionally, the computing device parameter adjustment method further includes: performing data fitting on the energy efficiency data of the plurality of core frequencies to obtain an energy efficiency fitting curve; performing data fitting on the latency data under the plurality of batch sizes to obtain a latency fitting curve; and filling the position points in the current energy efficiency gradient map where the comprehensive performance value of energy efficiency and latency is not filled based on the energy efficiency fitting curve and the latency fitting curve.
[0008] Optionally, the process of calculating the comprehensive performance value of energy efficiency and latency of the combination of batch size and core frequency based on the energy efficiency data and the latency data includes: calculating the comprehensive performance value of energy efficiency and latency of the combination of batch size and core frequency based on a first relational expression, and the first relational expression is ; where is the combination of batch size and core frequency, is the comprehensive performance value of energy efficiency and latency of the combination of batch size and core frequency, is the energy efficiency data of the combination of batch size and core frequency, is the latency data of the combination of batch size and core frequency, is a first preset weight value, is a second preset weight value.
[0009] Optionally, the process of performing the current adjustment operation on the computing device based on the comprehensive performance value of energy efficiency and latency corresponding to the target position point includes: obtaining the actual operating state of the computing device, where the actual operating state includes the current batch size, the current core frequency, the current performance metric, and the comprehensive performance value of energy efficiency and latency corresponding to the target position point; inputting the actual operating state into an action prediction model to obtain a predicted optimal adjustment operation;
[0010] obtaining the current adjustment operation based on the predicted optimal adjustment operation, and performing the current adjustment operation on the computing device.
[0011] Optionally, the actual operating state further includes the comprehensive performance value of energy efficiency and latency corresponding to a neighborhood position point, where the neighborhood position point is a position point within a preset range from the target position point on the current energy efficiency gradient map.
[0012] Optionally, the computing device parameter adjustment method further includes: calculating a gradient vector corresponding to each position point on the current energy efficiency gradient map; the process of obtaining the current adjustment operation based on the predicted optimal adjustment operation includes: obtaining the gradient vector of the target position point;
[0013] obtaining the current adjustment operation by using the gradient vector and the predicted optimal adjustment operation.
[0014] Optionally, the process of obtaining the current adjustment operation by using the gradient vector and the predicted optimal adjustment operation includes: obtaining the current adjustment operation based on a second relational expression, where the second relational expression is ; where A exec is the current adjustment operation, A DQN is the predicted optimal adjustment operation, W is a third preset weight value, and G is the gradient vector of the target position point.
[0015] Optionally, the computing device parameter adjustment method further includes: using the action prediction model to obtain a predicted reward function value corresponding to the current adjustment operation; obtaining an observation result after performing the current adjustment operation on the computing device, where the observation result includes the new actual operating state of the computing device and an immediate reward function value; updating the action prediction model by using the observation result, the current adjustment operation, and the actual operating state before the current adjustment operation is not performed, so that the difference between the immediate reward function value and the predicted reward function value is minimized.
[0016] Optionally, the method for adjusting computing device parameters further includes: determining the maximum value among the energy efficiency and latency comprehensive performance values corresponding to each position point on the current energy efficiency gradient map; determining the maximum value as the optimal performance value; in response to the difference between the energy efficiency and latency comprehensive performance value corresponding to the target position point and the optimal performance value not being within a preset range, obtaining a new energy efficiency gradient map as the current energy efficiency gradient map.
[0017] Optionally, after performing the current adjustment operation on the computing device based on the energy efficiency and latency comprehensive performance value corresponding to the target position point, the method for adjusting computing device parameters further includes: obtaining the measured comprehensive performance value when the computing device operates at the new batch size and the new core frequency; in response to the measured comprehensive performance value being within the high-performance range, controlling the computing device to operate at the new batch size and the new core frequency for a preset number of times, and if the measured comprehensive performance values obtained each time are all within the high-performance range, determining that the new batch size and the new core frequency meet the latency and energy efficiency balance condition.
[0018] Optionally, after obtaining the measured comprehensive performance value when the computing device operates at the new batch size and the new core frequency, the method for adjusting computing device parameters further includes: in response to the measured comprehensive performance value not being within the high-performance range, determining that the new batch size and the new core frequency do not meet the latency and energy efficiency balance condition, taking the new batch size as the current batch size and the new core frequency as the current core frequency, and performing the operation of determining the target position point on the current energy efficiency gradient map according to the current batch size and the current core frequency of the computing device.
[0019] To solve the above technical problems, the present invention also provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the method for adjusting computing device parameters as described in any one of the above are implemented.
[0020] To solve the above technical problems, the present invention also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of the method for adjusting computing device parameters as described in any one of the above when executing the computer program.
[0021] To solve the above technical problems, the present invention also provides a server, including a plurality of computing devices and the above-described electronic device.
[0022] To solve the above technical problems, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method for adjusting computing device parameters as described in any one of the above are implemented.
[0023] The present invention provides a method for adjusting computing device parameters. First, the current energy efficiency gradient map of the computing device is obtained. The current energy efficiency gradient map shows the comprehensive performance of energy efficiency and latency of the computing device under different combinations of batch sizes and core frequencies. Through the energy efficiency gradient map, all possible combinations of batch sizes and core frequencies are considered, ensuring the comprehensiveness of the solution. During the actual operation of the computing device, the current batch size and current core frequency of the computing device are monitored, a position point matching the current batch size and current core frequency is found on the energy efficiency gradient map, and the batch size and core frequency of the computing device are adjusted based on the comprehensive performance value of energy efficiency and latency at the position point. After the adjustment is completed, it is verified whether the new configuration meets the preset latency and energy efficiency balance condition. If it is satisfied, the computing device will continue to operate according to this configuration, automatically adjusting the configuration according to the actual operation situation of the computing device to adapt to different workloads and operating environments, achieving a dynamic balance between energy efficiency and latency, and ensuring the optimal operating state of the computing device between performance and energy efficiency.
[0024] The present invention also provides a computer program product, an electronic device, a server, and a computer-readable storage medium, which have the same beneficial effects as the above-mentioned method for adjusting computing device parameters. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0026] Figure 1 It is a flowchart of the steps of a method for adjusting computing device parameters provided by the present invention.
[0027] Figure 2 It is a schematic structural diagram of a system for adjusting computing device parameters provided by the present invention.
[0028] Figure 3 It is a schematic structural diagram of an electronic device provided by the present invention.
[0029] Figure 4 It is a schematic structural diagram of a computer-readable storage medium provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] The core of the present invention is to provide a method, product, device, server, and medium for adjusting computing device parameters, which can automatically adjust configurations according to the actual operating conditions of the computing device, adapt to different workloads and operating environments, achieve a dynamic balance between energy efficiency and latency, and ensure the optimal operating state of the computing device between performance and energy efficiency.
[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0032] In a first aspect, please refer to Figure 1 , the present invention provides a method for adjusting computing device parameters, including:
[0033] S101: Obtain the current energy efficiency gradient map of the computing device; the position points of the current energy efficiency gradient map are used to represent the comprehensive performance value of energy efficiency and latency of the computing device under the corresponding batch size and core frequency combination.
[0034] In this embodiment, the computing device is a computing device in a server, including but not limited to a GPU, an FPGA (Field-Programmable Gate Array), etc. Energy efficiency gradient maps corresponding to each computing device in the server are respectively created in advance. Considering that with the use of the computing device and the change of environmental conditions, the same computing device may exhibit different performances even under the same workload. Therefore, the energy efficiency gradient map of the computing device can be dynamically updated during the use process of the computing device, and the energy efficiency gradient map currently used by the computing device at the current moment is determined as the current energy efficiency gradient map of the computing device.
[0035] It can be understood that the energy efficiency gradient map can be a two-dimensional grid, where one dimension is the batch size and the other dimension is the core frequency. The batch size affects the memory usage, throughput, and parallel processing ability of the computing device, and the change of the core frequency affects the power consumption and computing speed of the device. Each position point on the energy efficiency gradient map is used to represent the comprehensive performance value R(b j , f i ) of energy efficiency and latency of the computing device under the corresponding batch size and core frequency combination (b j , f i ), where b j represents the jth batch size, and f iRepresents the i-th core frequency. The combined energy efficiency and latency performance value combines two key performance metrics, energy efficiency and latency. The larger the combined energy efficiency and latency performance value, the better the balance between the energy efficiency and latency of the computing device under this batch size and core frequency combination. This balance is crucial for optimizing the performance of the computing device, ensuring that power consumption is effectively controlled while meeting performance requirements.
[0036] S102: Determine a target position point on the current energy efficiency gradient map according to the current batch size and current core frequency of the computing device, and perform a current adjustment operation on the computing device based on the combined energy efficiency and latency performance value corresponding to the target position point.
[0037] In this embodiment, first obtain the current batch size (b 当前 ) and the current core frequency (f 当前 ) of the computing device during actual operation, and search for the position point matching b 当前 and f 当前 on the current energy efficiency gradient map. Specifically, traverse each position point (b j , f i ) on the current energy efficiency gradient map. For each position point, check whether the following conditions are met: b j (the batch size of this position point) is equal to b 当前 (the current batch size of the computing device), and f i (the core frequency of this position point) is equal to f 当前 (the current core frequency of the device). If a position point that satisfies b j = b 当前 and f i = f 当前 is found, then this position point is the target position point. If no exactly matching position point is found, interpolation or searching for the closest position point may be required, which can be set according to actual engineering needs and is not specifically limited in this embodiment. It can be understood that by traversing the energy efficiency gradient map to find the position point that exactly matches the current batch size and current core frequency, it ensures that the adjustment is based on the most accurate data, improves the accuracy of the adjustment, and avoids performance degradation caused by adjustment based on inaccurate data. When there is no exactly matching position point, interpolation or the method of finding the closest point is used to estimate performance data, enabling reasonable performance estimation and adjustment even when the data on the energy efficiency gradient map is incomplete, improving the flexibility and robustness of the system.
[0038] It can be understood that each position point in the current energy efficiency gradient map of the computing device has its corresponding comprehensive performance value of energy efficiency and latency. After determining the target position point in this embodiment, the comprehensive performance value of energy efficiency and latency corresponding to the target position point is obtained, and an adjustment operation is performed on the computing device according to the comprehensive performance value of energy efficiency and latency corresponding to the target position point. The comprehensive performance value of energy efficiency and latency is used to analyze whether an adjustment operation needs to be performed on the computing device. Exemplarily, assuming that the comprehensive performance value of energy efficiency and latency at the target position point is lower than a preset threshold, it indicates that the computing device cannot achieve the balance of energy efficiency and latency when running at the current batch size and the current adjustment direction. At this time, it is necessary to adjust the operating parameters (batch size and / or core frequency) of the computing device. The adjustment operation includes the adjustment direction and amplitude of the current batch size of the computing device, and also includes the adjustment direction and amplitude of the current core frequency of the computing device. In this embodiment, according to the situation that the comprehensive performance value is lower than the threshold, the batch size and the core frequency are adjusted, which can improve the energy efficiency and latency performance of the device, enhance the overall performance, optimize the energy efficiency, reduce unnecessary energy consumption, lower the operating cost, avoid the computing device running in a non-optimal state for a long time, reduce failures and downtime, and also avoid extreme working conditions, which helps to extend the service life of the computing device.
[0039] S103: In response to the new batch size and the new core frequency after the current adjustment operation is completed satisfying the latency and energy efficiency balance condition, control the computing device to continuously run at the new batch size and the new core frequency.
[0040] After the current adjustment operation is completed on the computing device, it is also necessary to monitor the new batch size and the new core frequency of the computing device. If the adjusted new batch size and new core frequency can make the computing device satisfy the latency and energy efficiency balance condition, then the new batch size and new core frequency are used as the optimal operating parameters of the computing device, and the computing device is controlled to run at the new batch size and new core frequency to ensure the balance of energy efficiency and latency. Through this process, the computing device can run in a continuously optimized environment, which not only improves the performance of the device, but also enhances the overall efficiency and reliability of the data center.
[0041] In response to measuring that the comprehensive performance value is not within the high-performance range, it is determined that the new batch size and the new core frequency do not satisfy the latency and energy efficiency balance condition. The new batch size is used as the current batch size, and the new core frequency is used as the current core frequency, and the operation of determining the target position point on the current energy efficiency gradient map according to the current batch size and the current core frequency of the computing device is performed.
[0042] Among them, the latency and energy efficiency balance condition is set according to actual engineering needs, and this embodiment does not make specific limitations here.
[0043] It can be seen that in this embodiment, first, the current energy efficiency gradient map of the computing device is obtained. The current energy efficiency gradient map shows the comprehensive performance of the energy efficiency and latency of the computing device under different combinations of batch sizes and core frequencies. Through the energy efficiency gradient map, all possible combinations of batch sizes and core frequencies are considered, ensuring the comprehensiveness of the solution. During the actual operation of the computing device, the current batch size and current core frequency of the computing device are monitored, and the position point matching the current batch size and current core frequency is found on the energy efficiency gradient map. Based on the comprehensive performance value of energy efficiency and latency at the position point, the batch size and core frequency of the computing device are adjusted. After the adjustment is completed, it is verified whether the new configuration meets the preset latency and energy efficiency balance condition. If it is satisfied, the computing device will continue to operate according to this configuration, automatically adjusting the configuration according to the actual operation situation of the computing device to adapt to different workloads and operating environments, achieving the dynamic balance between energy efficiency and latency, and ensuring the optimal operating state of the computing device between performance and energy efficiency.
[0044] Based on the above embodiment:
[0045] In an exemplary embodiment, the process of obtaining the current energy efficiency gradient map of the computing device includes: determining the corresponding reference batch size range and reference core frequency range of the computing device according to the specifications and / or performance limit conditions of the computing device; selecting multiple batch sizes in the reference batch size range; selecting multiple core frequencies in the reference core frequency range; arranging and combining the multiple batch sizes and multiple core frequencies to generate multiple combinations of batch sizes and core frequencies; calculating the comprehensive performance value of energy efficiency and latency of the computing device under each combination of batch sizes and core frequencies; filling the position points of the preset two-dimensional grid based on the multiple combinations of batch sizes and core frequencies and their corresponding comprehensive performance values of energy efficiency and latency to obtain the current energy efficiency gradient map.
[0046] In this embodiment, first, the reference batch size range and the reference core frequency range need to be determined. This range should include all possible values from the minimum to the maximum to ensure that the data can cover all operating points. Specifically, the possible value ranges of the batch size and the core frequency are determined based on the specifications and performance limit conditions of the computing device. For example, the batch size range can start from 1 and gradually increase to the maximum value that each computing device in the cloud data center can effectively process, and the core frequency starts from the lowest frequency supported by the computing device and increases to the highest frequency.
[0047] When determining the reference batch size range (b min to b max ), and the reference core frequency range (f min to f maxAfter that, for each increase in the batch size and the core frequency, a reasonable increment is determined. For example, the batch size can be increased by 1, 4, or 8 each time, depending on the overall range and the desired number of data points. The increment of the core frequency can be set according to the frequency step of the computing device. The increment can be fixed or dynamically changed, and it can be set according to the actual engineering needs. This embodiment does not make any limitations here.
[0048] After determining the increment, multiple batch sizes (b1, b2,..., b m ) can be determined within the reference batch size range, where b1 = b min , b m = b max . Similarly, multiple core frequencies (f1, f2,..., f n ) can be determined within the reference core frequency range, where f1 = f min , f n = f max . After determining multiple core frequencies and multiple batch sizes, the multiple batch sizes and multiple core frequencies are arranged and combined to generate multiple combinations of batch sizes and core frequencies. Specifically, one batch size can be fixed first, and then all core frequencies are traversed. After that, the batch size is adjusted, and all core frequencies are traversed again, and so on, until all batch sizes have been fixed once. Exemplarily, b1 can be fixed first, and then all core frequencies are traversed. The combinations of batch sizes and core frequencies obtained include (b1, f1), (b1, f2),..., (b1, f n ). Then b2 is fixed, and all core frequencies are traversed again. The combinations of batch sizes and core frequencies obtained include (b2, f1), (b2, f2),..., (b2, f n ), and so on, until b m is fixed and all core frequencies are traversed, and the combinations of batch sizes and core frequencies obtained include (b m , f1), (b m , f2),..., (b m , f n ). Another strategy is to fix one core frequency first, then traverse all batch sizes, and then change the core frequency, repeating this process until f max is reached. The process of obtaining the combinations of batch sizes and core frequencies is the same as above, and this embodiment will not elaborate here.
[0049] Suppose there are the following ranges and increments:
[0050] Batch size range: b min = 16, b max= 64, increment = 16, then multiple batch sizes include 16, 32, 48, 64;
[0051] Core frequency range: f min = 1000 MHz, f max = 1500 MHz, increment = 100 MHz, then multiple core frequencies include 1000 MHz, 1100 MHz, 1200 MHz, 1300 MHz, 1400 MHz, 1500 MHz.
[0052] The following batch size and core frequency combinations can be obtained: (16, 1000 MHz), (16, 1100 MHz), …, (16, 1500 MHz), (32, 1000 MHz), (32, 1100 MHz), …, (32, 1500 MHz) … (64, 1000 MHz), (64, 1100 MHz), …, (64, 1500 MHz).
[0053] These batch size and core frequency combinations will be used to test and evaluate the performance of the computing device under different conditions, so as to construct an energy efficiency gradient map.
[0054] Then let the computing device run one by one under each batch size and core frequency combination. It can be understood that considering the possible impact of the test order on the results, for example, before changing the batch size and core frequency combination, give the computing device a certain cooling time to avoid the impact of the previous batch size and core frequency combination on the test results of the next batch size and core frequency combination. Then for each batch size and core frequency combination, determine the number of tests and the duration of each test, which need to be determined according to the specific requirements of the test, as long as the data under each configuration is representative and reliable enough. Obtain the comprehensive performance values of energy efficiency and latency when the computing device runs under each batch size and core frequency combination, and then fill the comprehensive performance values of energy efficiency and latency corresponding to each batch size and core frequency combination into the position points of the two-dimensional grid, so as to create an energy efficiency gradient map of the computing device. This map provides an intuitive way to understand the performance under different configurations and can help determine the optimal batch size and core frequency combination.
[0055] In an exemplary embodiment, the process of calculating the comprehensive performance values of energy efficiency and latency for the computing device under each batch size and core frequency combination includes: for each batch size and core frequency combination, obtain the energy efficiency data and latency data of the computing device under the batch size and core frequency combination, and calculate the comprehensive performance values of energy efficiency and latency under the batch size and core frequency combination based on the energy efficiency data and latency data.
[0056] In this embodiment, after the control computing device operates at a certain combination of batch processing size and core frequency, the energy efficiency data and latency data of the computing device are obtained. Among them, the energy efficiency data is the energy consumed per batch calculation of the computing device at this batch processing size and core frequency, and the latency data is the latency time of the computing device at this combination of batch processing size and core frequency. The energy efficiency data can be obtained by reading the real-time power value of the graphics card of the computing device and integrating it with the latency time to obtain the total energy value.
[0057] To comprehensively consider energy efficiency and latency, this embodiment defines an inference revenue model, and the formula is as follows:
[0058] ;
[0059] ;
[0060] Among them, is the inference revenue model, batchsize is the batch processing size, s is seconds, is the number of pictures processed per second, p is the power consumption, t is the time, E is the total energy consumption, W is watts, is the combination of batch processing size and core frequency, is the comprehensive performance value of energy efficiency and latency under the combination of batch processing size and core frequency, is the energy efficiency data under the combination of batch processing size and core frequency, is the latency data under the combination of batch processing size and core frequency, is the first preset weight value, is the second preset weight value.
[0061] It can be understood that as the batch processing size increases, both the energy efficiency data and the latency data will increase. Therefore, the performance trade-off is the ratio of the two, f i is the frequency of the i-th core of the computing device, that is, the i-th core frequency of the computing device. Among them, is the first preset weight value, is the second preset weight value, used to balance the importance of energy efficiency and latency. In short, this model aims to optimize the "revenue" of the inference task by considering energy efficiency and latency. The first preset weight and the second preset weight are set according to actual engineering needs, allowing the model to adjust the degree of emphasis on energy efficiency and latency according to the requirements of the task. For example, if the first preset weight is greater than the second preset weight, this inference revenue model pays more attention to energy efficiency than latency. On the contrary, if the first preset weight is less than the second preset weight, the inference revenue model pays more attention to latency. The goal of this optimization problem is to find an optimal b and f i combination to maximize .
[0062] By Shows how to maximize the inference revenue model by selecting the optimal batch size and kernel frequency , which is represented as an optimization problem and can be expressed as: max b,fi , b and f i respectively need to satisfy: b ∈ N, f i ∈ C gpu .
[0063] In the above representation, max represents finding the value that maximizes the objective function , b represents the batch size, which must be an element in the set of positive integers N, and f i represents the GPU kernel frequency and must be an element in the set of kernel frequencies C gpu and, is the revenue model to be maximized, determined based on energy efficiency and latency as well as the first preset weight value and the second preset weight value .
[0064] In an exemplary embodiment, the method for adjusting computing device parameters further includes: performing data fitting on the energy efficiency data of multiple kernel frequencies to obtain an energy efficiency fitting curve; performing data fitting on the latency data under multiple batch sizes to obtain a latency fitting curve; filling the position points in the current energy efficiency gradient map that do not have filled comprehensive performance values of energy efficiency and latency based on the energy efficiency fitting curve and the latency fitting curve.
[0065] In this embodiment, considering that when deploying a model in a cloud data center, in order to effectively balance the time and resources of data sampling, a coarse-grained and large-span method is used to initially establish a two-dimensional grid map of values. Therefore, in order to further improve the accuracy, comprehensiveness, and reliability of the energy efficiency gradient map, this embodiment also performs data fitting on the basis of the initial data, which not only improves the time efficiency and cost efficiency, but also, due to its flexibility and adaptability, can easily handle changes in hardware and workloads without having to perform comprehensive data sampling every time, thus continuously optimizing and iterating the model in a dynamically changing environment.
[0066] Specifically, data fitting can be performed on the energy efficiency data of multiple kernel frequencies to obtain an energy efficiency fitting curve, including but not limited to performing data fitting on the energy efficiency data of multiple kernel frequencies through a Fourier curve. The Fourier curve can handle periodic changes well and is suitable for processing regular changes that may occur at certain frequencies.
[0067] When performing Fourier curve fitting, the general form of the Fourier curve is:
[0068] .
[0069] Among them, EE(f i ) represents the energy efficiency data of the core frequency, a0 is the average or DC component, a n and b n are Fourier coefficients, which are parameters to be determined through a fitting process. n represents the harmonic order, and N is the maximum number of harmonics selected. Parameter estimation: Use numerical methods (such as the least squares method) to estimate the Fourier coefficients a n and b n . This usually involves constructing a cost function, such as the mean square error, and then optimizing these coefficients to minimize the cost function.
[0070] Data fitting can be performed on the latency data for multiple batch sizes to obtain a latency fitting curve, including but not limited to using a rational function curve fitting. Rational functions are suitable for describing non-linear relationships and help to better understand the variation of latency with frequency and batch size.
[0071] Of course, in addition to using the above curves for data fitting, the least squares method can also be used to perform curve fitting on the energy efficiency data and latency data for each batch size, or on the energy efficiency data and latency data for each core frequency, to ensure the accuracy and reliability of the fitting.
[0072] Then, for the position points in the current energy efficiency gradient map that are not filled with the combined performance values of energy efficiency and latency, according to the batch size and core frequency combination corresponding to this position point, without controlling the computing device to run at this batch size and core frequency combination, directly determine the energy efficiency data corresponding to the batch size and core frequency combination based on the energy efficiency fitting curve, determine the latency data corresponding to the batch size and core frequency combination based on the latency fitting curve, calculate the combined performance value of energy efficiency and latency corresponding to the batch size and core frequency combination according to the energy efficiency data and latency data determined on the fitting curve, and fill it into this position point.
[0073] In this embodiment, it is not necessary to perform actual tests on each possible batch size and core frequency combination, which can significantly reduce the time and resources required for testing. Through the fitting curve, the data in the untested area can be quickly estimated, thus accelerating and improving the creation process of the energy efficiency gradient map.
[0074] In an exemplary embodiment, the process of performing a current adjustment operation on a computing device based on the comprehensive performance value of energy efficiency and latency corresponding to a target location point includes: obtaining the actual operating state of the computing device, where the actual operating state includes the current batch size, the current core frequency, the current performance metric, and the comprehensive performance value of energy efficiency and latency corresponding to the target location point; inputting the actual operating state into an action prediction model to obtain a predicted optimal adjustment operation; obtaining the current adjustment operation based on the predicted optimal adjustment operation, and performing the current adjustment operation on the computing device.
[0075] It can be understood that, for the convenience of adjusting the optimal batch size and core frequency of the computing device, in this embodiment, a reinforcement learning network for the core frequency and batch size of the computing device is pre-constructed and trained to obtain an action prediction model. Specifically, a reinforcement learning algorithm is constructed: assuming that the Deep Q-Network (DQN) is selected as the learning algorithm, DQN is used, which is a reinforcement learning algorithm that combines Q-learning and deep learning and is suitable for dealing with complex problems with continuous state spaces. DQN uses a deep neural network to estimate the reward function value of the state-action pair, and is a reinforcement learning algorithm that combines Q-learning and deep neural networks and is used to estimate the expected total return of taking a certain action in a given state. In the reinforcement learning framework, especially in DQN (Deep Q-Network), the input is the state, and the output is the best action for the given state. Here, the action refers to how to adjust f i and batchsize to maximize the long-term reward, and DQN selects an action by evaluating the Q value of each possible action.
[0076] Environment definition: In the environment of a cloud data center, the environment includes the current configuration and performance metrics of the computing device. The current configuration includes, but is not limited to, the core frequency and batch size, and the current performance metrics generated from this configuration include, but are not limited to, processing speed, energy efficiency, and latency, etc.
[0077] State definition: The state S is a multi-dimensional vector that includes the current batch size batchsize, the core frequency f i , the performance metric performance metrics and the current comprehensive performance value R of energy efficiency and latency current , that is, S = [batchsize, f i , performance metrics , R current .
[0078] Action definition: Action A is predicted by the DQN network, indicating how to adjust the current core frequency and the current batch size. .
[0079] It can be understood that after obtaining the actual operating state of the computing device, as the input of the action prediction model, that is, the current state S, the current state S is sent into the action prediction model for optimal action prediction, and the predicted optimal adjustment operation A output by the action prediction model is obtained. DQN , and send the predicted optimal adjustment operation to the external controller of the computing device to execute f i and the adjustment of batchsize.
[0080] The DQN network can automatically select the optimal action according to the current state, adapt to the changing workload and environmental conditions. Through continuous learning and adjustment, the DQN network can find the configuration that maximizes the comprehensive performance value of energy efficiency and latency. The automated adjustment process reduces manual intervention and improves the operation efficiency of the data center. In this way, the reinforcement learning network can effectively optimize the performance of the computing device and improve the overall efficiency and reliability of the data center.
[0081] In an exemplary embodiment, the actual operating state further includes the comprehensive performance value of energy efficiency and latency corresponding to the neighborhood location points, and the neighborhood location points are the location points within a preset range from the target location point on the current energy efficiency gradient map.
[0082] In this embodiment, the input of the DQN model, that is, the current state S, includes not only the comprehensive performance value of energy efficiency and latency corresponding to the target location point, but also the comprehensive performance value of energy efficiency and latency corresponding to the neighborhood location points of the target location point. This can provide more local information, enabling the DQN to consider the immediate effect of the state and its performance in neighboring states, which helps to make more comprehensive decisions when selecting actions and select better actions. That is, S = [batchsize, f i , performance metrics , R current , R local ;
[0083] where R local is a vector containing the comprehensive performance values of energy efficiency and latency within a certain range around batchsize and f i , including but not limited to , , , , etc.
[0084] In this embodiment, incorporating neighborhood information into the state space of the DQN model can improve the accuracy and robustness of the model's decision-making and accelerate the convergence speed.
[0085] In an exemplary embodiment, the method for adjusting computing device parameters further includes: for each position point on the current energy efficiency gradient map, calculating the gradient vector corresponding to the position point; the process of obtaining the current adjustment operation based on the predicted optimal adjustment operation includes: obtaining the gradient vector of the target position point; and using the gradient vector and the predicted optimal adjustment operation to obtain the current adjustment operation.
[0086] In this embodiment, for each position point on the energy efficiency gradient map, its corresponding gradient vector is calculated. The gradient vector represents the change trend of the comprehensive performance value of energy efficiency and latency when adjusting the batch size and kernel frequency at this position point. After the predicted optimal adjustment operation A output by the DQN network DQN is obtained, the gradient vector of the target position point is combined with the optimal adjustment operation A predicted by the DQN DQN to obtain the current actual adjustment operation. The gradient vector provides local information about the change trend of energy efficiency around the target position point, which can help the DQN model make more accurate decisions. Specifically, for the filled gradient map, the gradient vector of each position point is calculated, and the gradient vector can reveal in which direction adjusting the batch size and kernel frequency can maximize the comprehensive performance value of energy efficiency and latency.
[0087] The scheme for calculating the gradient vector is as follows:
[0088] S1: Select neighboring points: For each position point on the current energy efficiency gradient map, select its neighboring points in the directions of batch size and core frequency. For example, if the position point is (b, f i ), the neighboring points may be the first neighboring point (b + b, f i ) and the second neighboring point (b, f i + f i ), where b is the minimum increment of the batch size, f i is the minimum increment of the kernel frequency.
[0089] S2: Calculate the difference: For the batch size direction, calculate the difference R b = R(b + b, f i ) - R(b, f i ); R b is the first difference, R(b + b, f i ) is the actual performance value of the first neighboring point, and R(b, f i) is the actual performance value of the position point; for the core frequency direction, calculate the difference R fi =R(b,f i + f i )-R(b,f i ); R fi is the second difference, R(b,f i + f i ) is the actual performance value of the second adjacent point, and R(b,f i ) is the actual performance value of the position point.
[0090] S3: Gradient approximation. The gradient can be approximated by dividing these differences by the corresponding increments: Gradient direction b =( R b ) / b; Gradient direction fi =( R fi ) / f i 。
[0091] S4: Obtain the gradient vector. The gradient vector can be expressed as: R = (Gradient direction b , Gradient direction fi ).
[0092] In this embodiment, combining local gradient information can make up for the defect that the DQN model lacks understanding of local information, thereby improving the accuracy of decision-making. The gradient information can help the DQN model find the optimal solution faster, thereby improving the search efficiency, and thus improving the efficiency and accuracy of the computing device configuration.
[0093] In an exemplary embodiment, the process of obtaining the current adjustment operation using the gradient vector and the predicted optimal adjustment operation includes: obtaining the current adjustment operation based on the second relational expression, and the second relational expression is ; where A exec is the current adjustment operation, A DQN is the predicted optimal adjustment operation, W is the third preset weight value, and G is the gradient vector of the target position point.
[0094] In this embodiment, G is the local gradient vector calculated from R local ( R b , Rf i ), and the third preset weight value W is set according to actual engineering needs. Where R localis a vector containing the comprehensive performance values of energy efficiency and latency within a certain range around the target position point. That is, the local gradient vector G reflects the trend of energy efficiency change near the target position point. For example: R b represents the change amount of the comprehensive performance value when the batch size changes, R fi represents the change amount of the comprehensive performance value when the kernel frequency changes.
[0095] In this embodiment, by considering the local gradient information, the influence of the current state can be better understood, and the change trend of the future state can be predicted. The local gradient information can help the DQN model find the optimal solution faster, thereby improving the search efficiency and making more accurate adjustment operations.
[0096] In an exemplary embodiment, the method for adjusting the computing device parameters further includes: obtaining the predicted reward function value corresponding to the current adjustment operation by using the action prediction model; obtaining the observation result after performing the current adjustment operation on the computing device, where the observation result includes the new actual operating state of the computing device and the immediate reward function value; updating the action prediction model by using the observation result, the current adjustment operation, and the actual operating state before the current adjustment operation is not performed, so that the difference between the immediate reward function value and the predicted reward function value is minimized.
[0097] In this embodiment, the reward function of the action prediction model is described. The reward function R' is designed to measure the improvement in performance after taking a specific action. It can be defined according to the changes in energy efficiency and latency. Specifically, the reward function should encourage higher energy efficiency and lower latency, and is defined according to the following structure:
[0098] ; where b is the batch size, f i is the kernel frequency, is the energy efficiency data under the processing batch size b and the kernel frequency f i and is the latency data under the processing batch size b and the kernel frequency f i and and are weights used to balance the importance of energy efficiency and latency in the reward calculation, is the reward function value.
[0099] There is another definition of the reward function. The reward function is designed to measure the improvement in performance after taking a specific action and can be defined according to the changes in energy efficiency and latency: ; where and are weight factors used to balance the importance of energy efficiency and latency in the decision-making, is the energy efficiency data, is the delayed data, which is the reward function value under state S and action A.
[0100] It can be understood that in addition to the above reward function, other forms of reward functions aimed at measuring the improvement of performance after taking a specific action can also be adopted, which are not limited in this embodiment.
[0101] In the reinforcement learning model, the agent (i.e., the reinforcement learning algorithm) executes each action (i.e., changing b or f i ) and then calculates the new energy efficiency and latency, and then applies the reward function to obtain the predicted reward function value of this action .
[0102] After sending the current adjustment operation to the computing device, obtain the observation result after the computing device executes the current adjustment operation. The observation result includes the new state S' and the observed immediate reward function value .
[0103] Use the observation result to update the DQN model. Specifically, train the network by minimizing the difference between the predicted reward function value and the observed immediate reward function value. During the actual operation of the computing device, repeat the above steps to execute the loop process of selecting an action, executing the action, obtaining the observation result, and updating the model, so that the DQN model gradually learns which action to take in different states to obtain the maximum reward, thereby realizing the dynamic optimization of the computing device configuration.
[0104] In this embodiment, the computing device configuration can be dynamically adjusted according to different tasks and environments, so that it is always in the best state. The DQN model can learn complex non-linear relationships and gradually improve its prediction ability. In addition, this embodiment can be applied to various types of computing devices and tasks, and only need to adjust the reward function and model structure according to the actual situation.
[0105] In an exemplary embodiment, the method for adjusting computing device parameters further includes: determining the maximum value among the comprehensive performance values of energy efficiency and latency corresponding to each position point on the current energy efficiency gradient map; determining the maximum value as the optimal performance value; in response to the difference between the comprehensive performance value of energy efficiency and latency corresponding to the target position point and the optimal performance value not being within the preset range, obtaining a new energy efficiency gradient map as the current energy efficiency gradient map.
[0106] In this embodiment, after defining the above reinforcement learning parameters, start to execute the reinforcement learning training. First, each computing device in the cloud data center selects the optimal f i and batchsize value within the latency tolerance range, that is, the batch processing size and core frequency corresponding to the maximum value in the comprehensive performance value of energy efficiency and latency, and start training the DQN network from this value.
[0107] Meanwhile, determine the maximum value in the comprehensive performance value of energy efficiency and latency as the optimal R value. During the actual operation of the computing device, when it is determined that the comprehensive performance value of energy efficiency and latency corresponding to the current batch size and the current core frequency deviates from the optimal R value by more than a certain threshold, that is, when the difference between the comprehensive performance value of the energy efficiency and latency corresponding to the target position point and the optimal performance value is not within the preset range, it indicates that the current energy efficiency gradient map is no longer valid and needs to be recreated. At this time, recreate the energy efficiency gradient map of the computing device to ensure that the model always operates in the best state.
[0108] In an exemplary embodiment, after performing the current adjustment operation on the computing device based on the comprehensive performance value of energy efficiency and latency corresponding to the target position point, the computing device parameter adjustment method further includes:
[0109] Obtain the measured comprehensive performance value when the computing device runs at the new batch size and the new core frequency; in response to the measured comprehensive performance value being within the high-performance range, control the computing device to run at the new batch size and the new core frequency for a preset number of times. If the measured comprehensive performance values obtained each time are all within the high-performance range, determine that the new batch size and the new core frequency meet the latency and energy efficiency balance condition.
[0110] In this embodiment, if after changing to the new batch size and the new core frequency, the measured comprehensive performance value remains within the high-performance range of the delta threshold for a long time within h steps, fix the new batch size and the new core frequency, and the computing device reaches the highest energy efficiency operating state. If after a certain period, due to changes in the environmental state such as the latency not being satisfied or the energy consumption being found abnormal, restart the DQN model to predict and adjust the batch size and the core frequency.
[0111] Exemplarily, assume that the DQN model predicts that the optimal batch size is 128 and the core frequency is 1.5 GHz. After executing this configuration, the energy efficiency and latency of the GPU meet the requirements. At this time, the parameters of the DQN model no longer change. If the comprehensive performance values for 10 consecutive steps (h = 10) are all within the high-performance range (delta = 0.1), it can be determined that this configuration meets the latency and energy efficiency balance condition and fix it. If after a period of time, due to an increase in the server load, the latency of the GPU no longer meets the requirements, it indicates that the environmental state has changed. At this time, it is necessary to restart the DQN model, predict the new optimal configuration, and make adjustments.
[0112] As described above, the online training of the reinforcement learning network, the adaptive adjustment of the batch size and the core frequency fi are completed, so as to optimize the energy efficiency and performance of the computing device in big data processing and deep learning tasks during the long-term operation of the cloud server. By using the reinforcement learning technology, the optimal GPU configuration within the given latency tolerance range is automatically found, and an energy efficiency gradient map is established. A data-driven method is used to guide the adjustment of the GPU configuration, improving the accuracy and efficiency of decision-making. The Fourier curve fitting method is adopted to accurately fit the energy efficiency data, effectively handle the periodic changes, and enhance the adaptability and accuracy of the model. For tasks with strict running time requirements, such as real-time and interactive tasks, an optimization strategy is proposed to minimize the energy consumption while ensuring the performance.
[0113] In a second aspect, please refer to Figure 2 , the present invention provides a computing device parameter adjustment system, including:
[0114] An acquisition module 11, configured to acquire the current energy efficiency gradient map of the computing device; the position points of the current energy efficiency gradient map are used to characterize the comprehensive performance value of the energy efficiency and latency of the computing device under the corresponding batch size and core frequency combination.
[0115] A determination module 12, configured to determine a target position point on the current energy efficiency gradient map according to the current batch size and the current core frequency of the computing device, and perform a current adjustment operation on the computing device based on the comprehensive performance value of the energy efficiency and latency corresponding to the target position point.
[0116] A control module 13, configured to control the computing device to continuously operate according to the new batch size and the new core frequency in response to the new batch size and the new core frequency after the current adjustment operation is completed satisfying the latency and energy efficiency balance condition.
[0117] In an exemplary embodiment, the process of acquiring the current energy efficiency gradient map of the computing device includes: determining the corresponding reference batch size range and the reference core frequency range of the computing device according to the specifications and / or performance limitation conditions of the computing device; selecting a plurality of batch sizes in the reference batch size range; selecting a plurality of core frequencies in the reference core frequency range; arranging and combining the plurality of batch sizes and the plurality of core frequencies to generate a plurality of batch size and core frequency combinations; calculating the comprehensive performance value of the energy efficiency and latency of the computing device under each batch size and core frequency combination; filling the position points of a preset two-dimensional grid based on the plurality of batch size and core frequency combinations and their corresponding comprehensive performance values of the energy efficiency and latency to obtain the current energy efficiency gradient map.
[0118] In an exemplary embodiment, the process of calculating the comprehensive performance value of energy efficiency and latency of a computing device for each combination of batch processing size and core frequency includes: for each combination of batch processing size and core frequency, obtaining the energy efficiency data and latency data of the computing device under the combination of batch processing size and core frequency, and calculating the comprehensive performance value of energy efficiency and latency under the combination of batch processing size and core frequency based on the energy efficiency data and latency data.
[0119] In an exemplary embodiment, the computing device parameter adjustment system is further configured to: perform data fitting on the energy efficiency data of multiple core frequencies to obtain an energy efficiency fitting curve; perform data fitting on the latency data under multiple batch processing sizes to obtain a latency fitting curve; and fill the position points in the current energy efficiency gradient map where the comprehensive performance value of energy efficiency and latency is not filled based on the energy efficiency fitting curve and the latency fitting curve.
[0120] In an exemplary embodiment, the process of calculating the comprehensive performance value of energy efficiency and latency under the combination of batch processing size and core frequency based on the energy efficiency data and latency data includes: calculating the comprehensive performance value of energy efficiency and latency under the combination of batch processing size and core frequency based on a first relational expression, and the first relational expression is ; where is the combination of batch processing size and core frequency, is the comprehensive performance value of energy efficiency and latency under the combination of batch processing size and core frequency, is the energy efficiency data under the combination of batch processing size and core frequency, is the latency data under the combination of batch processing size and core frequency, is the first preset weight value, is the second preset weight value.
[0121] In an exemplary embodiment, the process of performing the current adjustment operation on the computing device based on the comprehensive performance value of energy efficiency and latency corresponding to the target position point includes: obtaining the actual operating state of the computing device, where the actual operating state includes the current batch processing size, the current core frequency, the current performance metric, and the comprehensive performance value of energy efficiency and latency corresponding to the target position point; inputting the actual operating state into an action prediction model to obtain a predicted optimal adjustment operation; and obtaining the current adjustment operation based on the predicted optimal adjustment operation and performing the current adjustment operation on the computing device.
[0122] In an exemplary embodiment, the actual operating state further includes the comprehensive performance value of energy efficiency and latency corresponding to a neighborhood position point, and the neighborhood position point is a position point within a preset range from the target position point on the current energy efficiency gradient map.
[0123] In an exemplary embodiment, the computing device parameter adjustment system is further configured to: for each position point on the current energy efficiency gradient map, calculate the gradient vector corresponding to the position point; the process of obtaining the current adjustment operation based on the predicted optimal adjustment operation includes: obtaining the gradient vector of the target position point; using the gradient vector and the predicted optimal adjustment operation to obtain the current adjustment operation.
[0124] In an exemplary embodiment, the process of using the gradient vector and the predicted optimal adjustment operation to obtain the current adjustment operation includes: obtaining the current adjustment operation based on a second relational expression, and the second relational expression is ; where A exec is the current adjustment operation, A DQN is the predicted optimal adjustment operation, W is a third preset weight value, and G is the gradient vector of the target position point.
[0125] In an exemplary embodiment, the computing device parameter adjustment system is further configured to: use an action prediction model to obtain the predicted reward function value corresponding to the current adjustment operation; obtain the observation result after performing the current adjustment operation on the computing device, and the observation result includes the new actual operating state of the computing device and the immediate reward function value; use the observation result, the current adjustment operation, and the actual operating state before the current adjustment operation is not performed to update the action prediction model, so that the difference between the immediate reward function value and the predicted reward function value is minimized.
[0126] In an exemplary embodiment, the computing device parameter adjustment system is further configured to: determine the maximum value among the energy efficiency and latency comprehensive performance values corresponding to each position point on the current energy efficiency gradient map; determine the maximum value as the optimal performance value; in response to the difference between the energy efficiency and latency comprehensive performance value corresponding to the target position point and the optimal performance value not being within the preset range, obtain a new energy efficiency gradient map as the current energy efficiency gradient map.
[0127] In an exemplary embodiment, after performing the current adjustment operation on the computing device based on the energy efficiency and latency comprehensive performance value corresponding to the target position point, the computing device parameter adjustment system is further configured to: obtain the measured comprehensive performance value when the computing device runs at the new batch size and the new core frequency; in response to the measured comprehensive performance value being within the high-performance range, control the computing device to run at the new batch size and the new core frequency for a preset number of times, and if the measured comprehensive performance values obtained each time are all within the high-performance range, determine that the new batch size and the new core frequency meet the latency and energy efficiency balance condition.
[0128] In an exemplary embodiment, after obtaining the measured comprehensive performance value when the computing device runs at a new batch size and a new core frequency, the computing device parameter adjustment system is further configured to, in response to the measured comprehensive performance value not being within the high-performance range, determine that the new batch size and the new core frequency do not meet the latency and energy efficiency balance condition, use the new batch size as the current batch size and the new core frequency as the current core frequency, and perform an operation of determining a target position point on the current energy efficiency gradient map according to the current batch size and the current core frequency of the computing device.
[0129] In a third aspect, the present invention further provides a computer program product, including a computer program / instructions, which when executed by a processor, implement the steps of the computing device parameter adjustment method described in any one of the above embodiments.
[0130] For the introduction of a computer program product provided by the present invention, please refer to the above embodiments, and the present invention will not be elaborated herein.
[0131] The computer program product provided by the present invention has the same beneficial effects as the above-described computing device parameter adjustment method.
[0132] In a fourth aspect, please refer to Figure 3 , the present invention further provides an electronic device, including:
[0133] A memory 21 for storing a computer program;
[0134] A processor 22 for implementing the steps of the computing device parameter adjustment method described in any one of the above embodiments when executing the computer program.
[0135] The electronic device further includes:
[0136] An input interface 23 connected to the processor 22 via a communication bus 26 for obtaining externally imported computer programs, parameters, and instructions, and storing them in the memory 21 under the control of the processor 22. The input interface can be connected to an input device to receive parameters or instructions manually input by a user. The input device can be a touch layer covered on a display screen, or a button, a trackball, or a touchpad provided on the terminal housing.
[0137] A display unit 24 connected to the processor 22 via a communication bus 26 for displaying data sent by the processor 22. The display unit can be a liquid crystal display screen or an electronic ink display screen, etc.
[0138] The network port 25 is connected to the processor 22 via the communication bus 26 and is used for communication connection with various external terminal devices. The communication technology adopted for this communication connection can be a wired communication technology or a wireless communication technology, such as Mobile High-Definition Link (MHL) technology, Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), Wi-Fi technology, Bluetooth communication technology, Bluetooth Low Energy (BLE) communication technology, communication technology based on IEEE 802.11s, etc.
[0139] For the introduction of an electronic device provided by the present invention, please refer to the above embodiments, and the present invention will not be elaborated herein.
[0140] The electronic device provided by the present invention has the same beneficial effects as the above-mentioned method for adjusting parameters of a computing device.
[0141] In a fifth aspect, the present invention further provides a server, which includes a plurality of computing devices and the above-mentioned electronic device.
[0142] For the introduction of a server provided by the present invention, please refer to the above embodiments, and the present invention will not be elaborated herein.
[0143] The server provided by the present invention has the same beneficial effects as the above-mentioned method for adjusting parameters of a computing device.
[0144] In a sixth aspect, please refer to Figure 4 , the present invention further provides a computer-readable storage medium 30, on which a computer program 31 is stored. When the computer program 31 is executed by a processor, it implements the steps of the method for adjusting parameters of a computing device described in any one of the above embodiments.
[0145] The computer-readable storage medium 30 may include: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.
[0146] For the introduction of a computer-readable storage medium 30 provided by the present invention, please refer to the above embodiments, and the present invention will not be elaborated herein.
[0147] The computer-readable storage medium 30 provided by the present invention has the same beneficial effects as the above-mentioned method for adjusting parameters of a computing device.
[0148] It should also be noted that in this specification, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0149] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for adjusting parameters of a computing device, characterized in that: include: Get the current energy efficiency gradient map of the computing device; The position points of the current energy efficiency gradient map are used to characterize the comprehensive performance value of energy efficiency and latency of the computing device under the corresponding batch size and kernel frequency combination; the current energy efficiency gradient map is a two-dimensional grid, the first dimension of the two-dimensional grid is the batch size, and the second dimension is the kernel frequency; Determine a target location point on the current energy efficiency gradient map according to a current batch size and a current kernel frequency of the computing device, and perform a current adjustment operation on the computing device based on a comprehensive performance value of energy efficiency and latency corresponding to the target location point; In response to the new batch size and the new kernel frequency satisfying a delay and energy efficiency balance condition after the current adjustment operation is completed, the computing device is controlled to continue to operate according to the new batch size and the new kernel frequency.
2. The method for adjusting computing device parameters according to claim 1, characterized in that: The process of obtaining the current energy efficiency gradient map of a computing device includes: Determining a reference batch size range and a reference kernel frequency range corresponding to the computing device according to the specification and / or performance constraints of the computing device; Selecting a plurality of batch sizes within the reference batch size range; Selecting a plurality of kernel frequencies in the reference kernel frequency range; Arrange and combine a plurality of the batch sizes and a plurality of the kernel frequencies to generate a plurality of combinations of the batch sizes and the kernel frequencies; Calculating the comprehensive performance value of energy efficiency and latency of the computing device under each combination of the batch size and the core frequency; Based on the plurality of batch size and kernel frequency combinations and their corresponding energy efficiency and latency comprehensive performance values, the position points of a preset two-dimensional grid are filled to obtain a current energy efficiency gradient map.
3. The method for adjusting computing device parameters according to claim 2, characterized in that: The process of calculating the comprehensive performance value of energy efficiency and latency of the computing device under each combination of the batch size and the core frequency includes: For each of the batch size and kernel frequency combinations, energy efficiency data and latency data of the computing device under the batch size and kernel frequency combination are obtained, and a comprehensive performance value of energy efficiency and latency under the batch size and kernel frequency combination is calculated based on the energy efficiency data and the latency data.
4. The method for adjusting computing device parameters according to claim 3, characterized in that: The computing device parameter adjustment method further includes: Performing data fitting on the energy efficiency data of the plurality of core frequencies to obtain an energy efficiency fitting curve; Performing data fitting on the delay data under the plurality of batch sizes to obtain a delay fitting curve; Based on the energy efficiency fitting curve and the delay fitting curve, the position points in the current energy efficiency gradient map that are not filled with the comprehensive performance value of energy efficiency and delay are filled.
5. The method for adjusting computing device parameters according to claim 3, characterized in that: The process of calculating the comprehensive performance value of energy efficiency and latency under the combination of the batch size and the kernel frequency based on the energy efficiency data and the latency data includes: The comprehensive performance value of energy efficiency and latency under the combination of the batch size and the kernel frequency is calculated based on the first relational expression, where the first relational expression is: ; in, For the batch size and kernel frequency combination, is the comprehensive performance value of energy efficiency and latency under the combination of batch size and kernel frequency, is the energy efficiency data for the batch size and kernel frequency combination, is the latency data for the batch size and kernel frequency combination, is the first preset weight value, is the second preset weight value.
6. The method for adjusting computing device parameters according to claim 1, characterized in that: The process of performing the current adjustment operation on the computing device based on the comprehensive performance value of energy efficiency and delay corresponding to the target location point includes: Acquire the actual running state of the computing device, wherein the actual running state includes a current batch size, a current kernel frequency, a current performance index, and a comprehensive performance value of energy efficiency and delay corresponding to the target location point; Inputting the actual operating state into an action prediction model to obtain a predicted optimal adjustment operation; A current adjustment operation is obtained based on the predicted optimal adjustment operation, and the current adjustment operation is performed on the computing device.
7. The method for adjusting computing device parameters according to claim 6, characterized in that: The actual operating status also includes a comprehensive performance value of energy efficiency and delay corresponding to a neighborhood position point, and the neighborhood position point is a position point within a preset range from the target position point on the current energy efficiency gradient map.
8. The method for adjusting computing device parameters according to claim 6, characterized in that: The computing device parameter adjustment method further includes: For each of the position points on the current energy efficiency gradient map, calculating a gradient vector corresponding to the position point; The process of obtaining the current adjustment operation based on the predicted optimal adjustment operation includes: Obtaining a gradient vector of the target position point; A current adjustment operation is obtained by using the gradient vector and the predicted optimal adjustment operation.
9. The method for adjusting computing device parameters according to claim 8, characterized in that: The process of obtaining the current adjustment operation by using the gradient vector and the predicted optimal adjustment operation includes: The current adjustment operation is obtained based on the second relational expression, where the second relational expression is: ; Among them, A exec For the current adjustment operation, A DQN is the predicted optimal adjustment operation, W is the third preset weight value, and G is the gradient vector of the target position point.
10. The method for adjusting computing device parameters according to claim 6, characterized in that: The computing device parameter adjustment method further includes: Using the action prediction model to obtain a predicted reward function value corresponding to the current adjustment operation; Obtaining an observation result after executing the current adjustment operation on the computing device, the observation result comprising a new actual operating state and an instant reward function value of the computing device; The action prediction model is updated using the observation result, the current adjustment operation, and the actual running state before the current adjustment operation is performed, so as to minimize the difference between the immediate reward function value and the predicted reward function value.
11. The method for adjusting computing device parameters according to claim 10, characterized in that: The computing device parameter adjustment method further includes: Determine the maximum value of the comprehensive performance values of energy efficiency and delay corresponding to each position point on the current energy efficiency gradient map; determining the maximum value as the optimal performance value; In response to the difference between the energy efficiency and delay comprehensive performance value corresponding to the target position point and the optimal performance value not being within a preset range, a new energy efficiency gradient map is acquired as the current energy efficiency gradient map.
12. The method for adjusting computing device parameters according to any one of claims 1 to 11, characterized in that: After performing the current adjustment operation on the computing device based on the comprehensive performance value of energy efficiency and delay corresponding to the target location point, the computing device parameter adjustment method further includes: obtaining a measured comprehensive performance value of the computing device when operating at the new batch size and the new kernel frequency; In response to the measured comprehensive performance value being within the high-performance range, the computing device is controlled to run a preset number of times according to the new batch size and the new core frequency; if the measured comprehensive performance value obtained each time is within the high-performance range, it is determined that the new batch size and the new core frequency satisfy the delay and energy efficiency balance condition.
13. The method for adjusting computing device parameters according to claim 12, characterized in that: After obtaining the measured comprehensive performance value of the computing device when running at the new batch size and the new kernel frequency, the computing device parameter adjustment method further includes: In response to the measured comprehensive performance value not being within the high-performance range, it is determined that the new batch size and the new kernel frequency do not satisfy the latency and energy efficiency balance condition, the new batch size is used as the current batch size, the new kernel frequency is used as the current kernel frequency, and an operation of determining a target position point on the current energy efficiency gradient map according to the current batch size and the current kernel frequency of the computing device is performed.
14. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the method for adjusting parameters of a computing device as described in any one of claims 1 to 13 are implemented.
15. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the method for adjusting parameters of a computing device as described in any one of claims 1 to 13 when executing the computer program.
16. A server, characterized in that: The invention comprises a plurality of computing devices and the electronic device as claimed in claim 15.
17. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for adjusting parameters of a computing device according to any one of claims 1 to 13 are implemented.
Citation Information
Patent Citations
Equipment parameter adjusting method and system, equipment, server, product and medium
CN119376959A