Device Parameter Adjustment Method, System, Device, Server, Product and Medium
Adjusting the batch size and core frequency of the GPU through the particle swarm optimization algorithm solves the balance of energy efficiency and latency in cloud computing, improves equipment efficiency and reduces costs, and enhances the competitiveness of cloud computing services.
Patent Information
- Application Number
- CN202411976961.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-12-31
AI Technical Summary
The prior art is difficult to effectively balance the energy efficiency and latency of GPUs in cloud computing, making it difficult to optimize the performance and operational costs of computing devices.
The particle swarm optimization algorithm is used to initialize the position of the particle swarm by obtaining multiple batch sizes and kernel frequency combinations, and iteratively update the position of the particle based on the predicted and actual performance values until the optimal solution is found to adjust the parameter configuration of the computing device.
It achieves a balance between energy efficiency and latency of computing equipment, improves work efficiency, reduces operating costs, and enhances the competitiveness of cloud computing services.
Smart Images

Figure CN119376959B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of servers, and particularly to a method, a system, a device, a server, a product and a medium for adjusting device parameters. Background Art
[0002] In the era of cloud computing, the inference demand accelerated by GPU (Graphics Processing Unit) is increasing continuously, and it becomes crucial to optimize the energy efficiency and performance of the GPU. The batch size and the core frequency are the key parameters for adjusting the GPU performance. Adjusting the batch size affects the memory utilization rate and the parallel processing efficiency of the GPU, while adjusting the GPU core frequency directly affects the processing speed and the energy consumption. Therefore, finding the optimal batch size and core frequency configuration to balance the energy efficiency and latency of the GPU is a problem that those skilled in the art need to solve currently. Summary of the Invention
[0003] The object of the present invention is to provide a method, a system, a device, a server, a product and a medium for adjusting device parameters, which can balance the energy efficiency and latency of a computing device, improve the working efficiency of the computing device, reduce the operation cost, and enhance the competitiveness of cloud computing services.
[0004] To solve the above technical problem, the present invention provides a method for adjusting device parameters, including: obtaining multiple combinations of batch sizes and core frequencies of a computing device; initializing the positions of multiple particles in a particle swarm based on the multiple combinations of batch sizes and core frequencies; for each particle in the multiple particle swarms, obtaining a predicted performance value and an actual performance value of the computing device in the current iteration according to the position of the particle in the current iteration, and updating the position of the particle in the next iteration based on the predicted performance value and the actual performance value, repeating this step until an iteration end condition is satisfied; taking the particle corresponding to the optimal actual performance value when the iteration end condition is satisfied as the optimal solution, and controlling the operation of the computing device according to the combination of the batch size and the core frequency in the optimal solution.
[0005] Optionally, the process of obtaining the actual performance value of the computing device according to the position of the particle in the current iteration includes: determining the current combination of the batch size and the core frequency in the position of the particle in the current iteration; obtaining the actual performance value of the computing device when operating under the current combination of the batch size and the core frequency.
[0006] Optionally, the process of obtaining the predicted performance value of the computing device according to the position of the particle in the current iteration includes: inputting the positions of the particle in consecutive preset iterations including the current iteration and the corresponding actual performance values into a time series network to obtain the predicted performance value of the computing device in the current iteration.
[0007] Optionally, the process of updating the position of the particle in the next iteration based on the predicted performance value and the actual performance value includes: determining the individual best position of the particle in the current iteration based on the actual performance value of the particle in the current iteration; determining the global best position in the current iteration based on the individual best positions of each particle in the particle swarm in the current iteration; updating the position of the particle in the next iteration according to the predicted performance value, the actual performance value, the individual best position, and the global best position.
[0008] Optionally, the process of determining the individual best position of the particle in the current iteration based on the actual performance value of the particle in the current iteration includes: determining whether the actual performance value of the particle in the current iteration is greater than the actual performance value corresponding to the individual best position of the particle in the previous iteration; if so, determining the position of the particle in the current iteration as the individual best position of the particle in the current iteration; if not, determining the individual best position of the particle in the previous iteration as the individual best position of the particle in the current iteration.
[0009] Optionally, the process of determining the global best position in the current iteration based on the individual best positions of each particle in the particle swarm in the current iteration includes: determining the maximum value of the actual performance values corresponding to the individual best positions of each particle in the particle swarm in the current iteration; determining the individual best position corresponding to the maximum value as the global best position in the current iteration.
[0010] Optionally, the process of updating the position of the particle in the next iteration according to the predicted performance value, the actual performance value, the individual best position, and the global best position includes: obtaining the gradient information corresponding to the position of the particle in the current iteration; updating the position of the particle in the next iteration according to the predicted performance value, the actual performance value, the individual best position, the global best position, and the gradient information.
[0011] Optionally, the process of updating the position of the particle in the next iteration according to the predicted performance value, the actual performance value, the individual best position, the global best position, and the gradient information includes: updating the velocity of the particle in the next iteration according to the velocity update relation, and the velocity update relation is Update the position of the particle in the next iteration according to the position update relation, where the position update relation is where j is the label of the particle, w is the inertia weight, c1, c2, c3, and c4 are all learning factors, and r1, r2, r3, and r4 are all random numbers, is the velocity of the particle in the next iteration, is the predicted performance value of the particle in this iteration, Gradient j is the gradient information corresponding to the position of the particle in this iteration, is the position of the particle in the next iteration, is the position of the particle in this iteration, V j is the velocity of the particle in this iteration, P best,j is the individual best position of the particle in this iteration, G best is the global best position in this iteration, f(X j ) is the actual performance value of the particle in this iteration.
[0012] Optionally, after obtaining multiple combinations of batch sizes and kernel frequencies of the computing device, the device parameter adjustment method further includes: calculating the actual performance value of the computing device under each combination of batch sizes and kernel frequencies; filling the position points of a preset two-dimensional grid based on multiple combinations of batch sizes and kernel frequencies and their corresponding actual performance values to obtain the current energy efficiency gradient map; calculating the gradient information corresponding to each position point on the current energy efficiency gradient map; the process of obtaining the gradient information corresponding to the position of the particle in this iteration includes: determining the current batch size and kernel frequency combination in the position of the particle in this iteration; determining the target position point corresponding to the current batch size and kernel frequency combination on the current energy efficiency gradient map; and using the gradient information corresponding to the target position point as the gradient information corresponding to the position of the particle in this iteration.
[0013] Optionally, for each of the position points on the current energy efficiency gradient map, the process of calculating the gradient information corresponding to the position point includes: for each of the position points on the current energy efficiency gradient map, determining a first neighboring point and a second neighboring point corresponding to the position point on the current energy efficiency gradient map; the core frequency of the first neighboring point is the same as the core frequency of the position point, the batch size of the first neighboring point is different from the batch size of the position point, the core frequency of the second neighboring point is different from the core frequency of the position point, and the batch size of the second neighboring point is the same as the batch size of the position point; calculating a first gradient value of the position point based on the actual performance value of the first neighboring point and the actual performance value of the position point; calculating a second gradient value of the position point based on the actual performance value of the second neighboring point and the actual performance value of the position point; and obtaining the gradient information corresponding to the position point by using the first gradient value and the second gradient value.
[0014] Optionally, the process of calculating the first gradient value of the position point based on the actual performance value of the first neighboring point and the actual performance value of the position point includes: obtaining a first difference between the actual performance value of the first neighboring point and the actual performance value of the position point; obtaining a first increment between the batch size of the first neighboring point and the batch size of the position point; and calculating a first gradient value based on the first difference and the first increment.
[0015] Optionally, the process of calculating the second gradient value of the position point based on the actual performance value of the second neighboring point and the actual performance value of the position point includes: obtaining a second difference between the actual performance value of the second neighboring point and the actual performance value of the position point; obtaining a second increment between the core frequency of the second neighboring point and the core frequency of the position point; and calculating a second gradient value based on the second difference and the second increment.
[0016] Optionally, the process of calculating the actual performance value of the computing device under each combination of the batch size and the core frequency includes: for each combination of the batch size and the core frequency, obtaining the energy efficiency data and the latency data of the computing device under the combination of the batch size and the core frequency, and calculating the actual performance value under the combination of the batch size and the core frequency based on the energy efficiency data and the latency data.
[0017] Optionally, the process of calculating the actual performance value under the combination of the batch size and the core frequency based on the energy efficiency data and the latency data includes: calculating the actual performance value under the combination of the batch size and the core frequency based on a first relational expression, and the first relational expression is ; where For the batch size and kernel frequency combination, is the actual performance value under the batch size and kernel frequency combination, is the energy efficiency data under the batch size and kernel frequency combination, is the latency data under the batch size and kernel frequency combination, is the first preset weight value, is the second preset weight value.
[0018] Optionally, the iteration end condition includes reaching the maximum number of iterations, or the change amount of the actual performance value corresponding to the global optimal position in a continuous preset number of iterations including the current iteration is less than the preset threshold.
[0019] To solve the above technical problems, the present invention further provides a device parameter adjustment system, including: an acquisition module, configured to acquire multiple batch size and kernel frequency combinations of a computing device; an initialization module, configured to initialize the positions of multiple particles in a particle swarm based on the multiple batch size and kernel frequency combinations; an iterative update module, configured to, for each particle in the multiple particle swarms, obtain the predicted performance value and the actual performance value of the computing device in the current iteration according to the position of the particle in the current iteration, update the position of the particle in the next iteration based on the predicted performance value and the actual performance value, and repeat this step until the iteration end condition is satisfied; a configuration module, configured to use the particle corresponding to the optimal actual performance value when the iteration end condition is satisfied as the optimal solution, and control the computing device to operate according to the batch size and kernel frequency combination in the optimal solution.
[0020] To solve the above technical problems, the present invention further provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the device parameter adjustment method described in any one of the above are implemented.
[0021] To solve the above technical problems, the present invention further provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of the device parameter adjustment method described in any one of the above when executing the computer program.
[0022] To solve the above technical problems, the present invention further provides a server, including multiple computing devices and the electronic device described above.
[0023] To solve the above technical problems, the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the device parameter adjustment method described in any one of the above are implemented.
[0024] The present invention provides a method for adjusting device parameters. First, multiple combinations of batch sizes and core frequencies of a computing device are obtained. The positions of each particle are initialized as different combinations of batch sizes and core frequencies using the particle swarm optimization algorithm, providing a diverse initial population for searching for the optimal configuration and helping to avoid local optimal solutions. In each iteration, the predicted performance value and the actual performance value of the current position of the particle (i.e., a specific combination of batch size and core frequency) are evaluated. By comparing the predicted and actual performances, the next movement of the particle can be more accurately guided, thus converging to the optimal solution faster. Through iterative updates, the particle gradually approaches the optimal combination of batch size and core frequency, realizing the dynamic optimization of the configuration of the computing device. After all iterations are completed, the particle with the best performance value is selected, and its position is the optimal configuration. The finally determined optimal combination of batch size and core frequency can balance the energy efficiency and latency of the computing device, not only improving the working efficiency of the computing device but also reducing the operating cost and enhancing the competitiveness of cloud computing services.
[0025] The present invention also provides a device parameter adjustment system, a computer program product, an electronic device, a server, and a computer-readable storage medium. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] To more clearly illustrate the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0027] Figure 1 It is a flowchart of the steps of a method for adjusting device parameters provided by the present invention.
[0028] Figure 2 It is a schematic structural diagram of a device parameter adjustment system provided by the present invention.
[0029] Figure 3 It is a schematic structural diagram of an electronic device provided by the present invention.
[0030] Figure 4 It is a schematic structural diagram of a computer-readable storage medium provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] The core of the present invention is to provide a method, a system, a device, a server, a product, and a medium for adjusting device parameters, which can balance the energy efficiency and latency of a computing device, improve the working efficiency of the computing device, reduce the operating cost, and enhance the competitiveness of cloud computing services.
[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0033] In a first aspect, please refer to Figure 1 , the present invention provides a method for adjusting device parameters, including:
[0034] S101: Obtain multiple combinations of batch sizes and core frequencies of a computing device.
[0035] In this embodiment, the computing device is a computing device in a server, including but not limited to a GPU, an FPGA (Field-Programmable Gate Array), etc. The server can be a server deployed in a cloud computing center. It can be understood that when the computing device operates under different combinations of batch sizes and core frequencies, its energy efficiency and latency will change. To determine the combination of batch size and core frequency that can balance energy efficiency and latency, this embodiment obtains multiple possible combinations of batch sizes and core frequencies of the computing device. Any combination of batch size and core frequency can be represented in the following form (batchsize, f i ), where b represents the batch size, and f i represents the core frequency. As an optional embodiment, before selecting multiple combinations of batch sizes and core frequencies, the batch size range (batchsize min , batchsize min ) and the core frequency range [f min , f max can be determined first based on reference conditions such as the capabilities of the computing device, task characteristics, hardware specifications, and / or operation limitations. Then, multiple batch sizes are determined within the batch size range, multiple core frequencies are determined within the core frequency range, and then the multiple batch sizes and multiple core frequencies are arranged and combined to obtain multiple combinations of batch sizes and core frequencies (batchsize, f i ).
[0036] Specifically, after determining the batch size range and the core frequency range, for each increase in the batch size and the core frequency, a reasonable increment is determined. For example, the batch size can be increased by 1, 4, or 8 each time, depending on the overall range and the desired number of data points. The increment of the core frequency can be set according to the frequency step of the computing device. The increment can be fixed or dynamically changed, which can be set according to actual engineering needs, and this embodiment does not make any limitation here. After determining the increment, multiple batch sizes (b1, b2,... b m ) can be determined within the batch size range, where b1 = b min , b m = b max . Similarly, multiple core frequencies (f1, f2,... f n ) can be determined within the core frequency range, where f1 = f min , f n = f max . After determining multiple core frequencies and multiple batch sizes, the multiple batch sizes and multiple core frequencies are arranged and combined to generate multiple combinations of batch sizes and core frequencies.
[0037] As an alternative embodiment, one can choose to first fix a batch size, then traverse all core frequencies, and then adjust the batch size and traverse all core frequencies again, and so on until all batch sizes have been fixed once. Exemplarily, one can first fix b1 and then traverse all core frequencies. The combinations of batch sizes and core frequencies obtained include (b1, f1), (b1, f2),..., (b1, f n ). Then fix b2 and traverse all core frequencies again. The combinations of batch sizes and core frequencies obtained include (b2, f1), (b2, f2),..., (b2, f n ), and so on until b m is fixed and all core frequencies are traversed, and the combinations of batch sizes and core frequencies obtained include (b m , f1), (b m , f2),..., (b m , f n ). Another strategy is to first fix a frequency, traverse all batch sizes, then change the core frequency, and repeat this process until f max is reached. The process of obtaining the combinations of batch sizes and core frequencies is the same as above, and this embodiment will not elaborate here.
[0038] S102: Initialize the positions of multiple particles in the particle swarm based on multiple combinations of batch sizes and core frequencies.
[0039] In this embodiment, the particle swarm algorithm is used to find the optimal combination of batch size and core frequency. First, determine the particle swarm size (N), that is, select the total number of particles in the particle swarm. Then, define the search space. The search space can be the possible ranges of batch size and core frequency. For example, the batch size range is [batchsize min ,batchsize max , and the core frequency range is [f min ,f max . First, based on multiple combinations of batch size and core frequency, initialize the position X j and velocity V j of each particle j (j = 1, 2, …, N) in the particle swarm.
[0040] That is, for each particle j, randomly select a combination of batch size and core frequency as its initial position X j_0 = [batchsize j_0 ,f i,j_0 ; where batchsize j_0 is a random value within the range of [batchsize min ,batchsize max , and f i,j_0 is a random value within the range of [f min ,f max .
[0041] For each particle j, randomly initialize its velocity V j_0 : V j_0 = [v batchsizej_0 ,v fi,j_0 ; it can be understood that the velocity of the particle can represent the movement trend of the particle in the search space. As an alternative embodiment, the velocity can be limited within a certain preset range to ensure that the particle does not move excessively in the search space. Among them, v batchsizej_0 is the change rate of the batch size when the particle is at this position X j_0 , and v fi,j_0 is the change rate of the core frequency when the particle is at this position X j_0 .
[0042] The initialization of each particle in the particle swarm is completed above. Through random initialization, each particle represents a potential solution in the search space. The particle swarm starts searching from different starting points, which ensures that the particle swarm algorithm can comprehensively cover the search space at the beginning and explore different combinations of batch size and core frequency, so that the optimal combination of batch size and core frequency can be found by iteratively updating the positions of the particles subsequently.
[0043] S103: For each particle in multiple particle swarms, based on the position of the particle in the current iteration, obtain the predicted performance value and the actual performance value of the computing device in the current iteration, update the position of the particle in the next iteration based on the predicted performance value and the actual performance value, and repeat this step until the iteration end condition is met.
[0044] In this embodiment, one particle in the particle swarm is taken as an example for illustration. The iterative update processes of the velocities and positions of other particles are the same by analogy.
[0045] First, obtain the position X of particle j in the current iteration j_c = (batchsize j_c , f i,j_c ), where c represents the current iteration, X j_c is the position of particle j in the current iteration, batchsize j_c is the batch size of particle j in the current iteration, and f i,j_c is the core frequency of particle j in the current iteration. Then, based on the combination of the batch size and the core frequency (batchsize j_c , f i,j_c ) of particle j in the current iteration, control the computing device to run, obtain the performance metrics of the computing device under the combination of the batch size and the core frequency (batchsize j_c , f i,j_c ), where the performance metrics include but are not limited to energy efficiency and latency, and calculate the actual performance value of the computing device in the current iteration based on the energy efficiency and the latency. At the same time, input the position X j_c of particle j in the current iteration into a pre-established time series network (such as LSTM (Long Short-Term Memory)), and the input of the time series network includes at least the position X j_c of particle j in the current iteration. The time series network predicts the performance metrics of the computing device in the current iteration, including but not limited to the predicted energy efficiency and the predicted latency, and then calculates and outputs the predicted performance value based on the predicted energy efficiency and the predicted latency. It can be understood that the difference between the predicted performance metrics and the actual performance metrics of the computing device in the current iteration is used to update the position of particle j in the next iteration. Repeat the above steps until the iteration end condition is met, and these conditions may include reaching the maximum number of iterations, the improvement of the performance metrics being lower than a certain threshold, or finding a satisfactory performance value. This process ensures that the particle swarm algorithm can dynamically adjust the search strategy according to the actual and predicted performance values, thereby more effectively searching for the optimal solution. By combining the actual performance and the prediction model, the particle swarm algorithm can converge to a good solution faster and reduce the dependence on actual hardware testing.
[0046] S104: Use the particle corresponding to the optimal actual performance value when the iteration end condition is met as the optimal solution, and control the operation of the computing device according to the batch size and kernel frequency combination in the optimal solution.
[0047] It can be understood that during the iteration process, each particle will be continuously updated according to its position, and at the same time, record its own best performance value in history (personal optimal solution P best ). At the same time, the entire particle swarm will also record the best personal optimal solution among all particles, that is, the global optimal solution (G best ).
[0048] The particle swarm algorithm will continue to iterate until the preset end conditions are met. These conditions may include: reaching the predetermined maximum number of iterations, the performance improvement is lower than a certain threshold, indicating that the space for further optimization is limited, and a solution that meets specific performance requirements is found. Once the iteration end condition is met, select the particle with the optimal actual performance value from all particles as the optimal solution. Extract the combination of its batch size and kernel frequency from the selected optimal particle. These parameters correspond to the optimal actual performance value. Use the combination of batch size and kernel frequency in the optimal solution to configure and control the operation of the computing device, that is, apply these parameters to the actual hardware or software environment to expect to obtain the optimal performance.
[0049] In this way, the particle swarm optimization algorithm not only helps to find the theoretically optimal parameter combination, but also transforms these theoretical results into practical applications, thereby improving the performance and efficiency of the computing device to adapt to scenarios that need to find the best configuration among multiple parameter combinations, such as cloud computing, high-performance computing, and big data processing.
[0050] It can be seen that in this embodiment, first, obtain multiple combinations of batch size and kernel frequency of the computing device, use the particle swarm optimization algorithm to initialize the position of each particle to different combinations of batch size and kernel frequency, providing a diverse initial population for searching for the optimal configuration, which helps to avoid local optimal solutions. In each iteration, evaluate the predicted performance value and the actual performance value of the current position of the particle (that is, a specific combination of batch size and kernel frequency). By comparing the predicted and actual performances, it can more accurately guide the next move of the particle, so as to converge to the optimal solution faster. Through iterative updates, the particle gradually approaches the optimal combination of batch size and kernel frequency, realizing the dynamic optimization of the configuration of the computing device. After all iterations are completed, select the particle with the best performance value, and its position is the optimal configuration. The finally determined optimal combination of batch size and kernel frequency can balance the energy efficiency and latency of the computing device, not only improving the working efficiency of the computing device, but also reducing the operating cost and enhancing the competitiveness of cloud computing services.
[0051] Based on the above embodiments:
[0052] In an exemplary embodiment, the process of obtaining the predicted performance value of the computing device according to the position of the particle in the current iteration includes: inputting the positions and corresponding actual performance values of the particle in a continuous preset number of iterations including the current iteration into a time series network to obtain the predicted performance value of the computing device in the current iteration.
[0053] In this embodiment, the preset number of iterations is first determined, that is, how many consecutive iterations need to be considered to predict the performance value of the current iteration. In addition to the position X in the current iteration, the input parameters of the time series network j_c and its corresponding actual performance value, also include the positions and corresponding actual performance values in a continuous preset number of iterations before the current iteration. For the case where the preset number is 2, the data that needs to be collected includes: the position X of particle j in the current iteration (the c-th iteration) j_c and its corresponding actual performance value, also include the position X of particle j in the previous iteration (the c-1-th iteration) j_c-1 and its corresponding actual performance value, the position X of particle j in the iteration before the previous iteration (the c-2-th iteration) j_c-2 and its corresponding actual performance value. The above collected data is formed into a historical parameter setting sequence and input into the time series network. The time series network can process these sequence data and consider the temporal dependence relationship, so as to predict the predicted performance value of the computing device in the current iteration. This prediction takes into account the historical changes in the position of particle j and the corresponding performance values, and can provide a more accurate prediction.
[0054] It can be understood that the time series network can capture the influence of the change in the position of the particle on the performance value of the computing device and predict the performance value of future iterations, which is applicable to scenarios where there is a temporal dependence between the performance value and the position of the particle, and can improve the efficiency and accuracy of the optimization process.
[0055] In an exemplary embodiment, the process of updating the position of the particle in the next iteration based on the predicted performance value and the actual performance value includes: determining the individual best position of the particle in the current iteration based on the actual performance value of the particle in the current iteration; determining the global best position in the current iteration based on the individual best positions of each particle in the particle swarm; updating the position of the particle in the next iteration according to the predicted performance value, the actual performance value, the individual best position, and the global best position.
[0056] Among them, the process of determining the individual best position of a particle in the current iteration based on the actual performance value of the particle in the current iteration includes: determining whether the actual performance value of the particle in the current iteration is greater than the actual performance value corresponding to the individual best position of the particle in the previous iteration; if so, determining the position of the particle in the current iteration as the individual best position of the particle in the current iteration; if not, determining the individual best position of the particle in the previous iteration as the individual best position of the particle in the current iteration.
[0057] Among them, the process of determining the global best position in the current iteration based on the individual best positions of the particles in the particle swarm in the current iteration includes: determining the maximum value of the actual performance values corresponding to the individual best positions of the particles in the particle swarm in the current iteration; determining the individual best position corresponding to the maximum value as the global best position in the current iteration.
[0058] In this embodiment, for each particle j, according to its actual performance value in the current iteration, the individual best position P best,j is determined. This position is the position where particle j obtains the best actual performance value up to the current iteration. In the entire particle swarm, by comparing the individual best positions of all particles, the global best position (G best ) is determined. The global best position is the position where the entire particle swarm obtains the best actual performance value in history.
[0059] According to the predicted performance value, actual performance value, individual best position (P best,j ) and global best position (G best ) of particle j, the position of particle j in the next iteration is updated. Among them, the update of the particle velocity takes into account the current velocity of the particle, the distance from the current position of the particle to the individual best position, and the distance from the current position of the particle to the global best position. The update of the particle position is based on the updated velocity.
[0060] Exemplarily, it is assumed that in the c-th iteration, the position of particle j is X j_c = [batchsize j_c , f i,j_c , and its actual performance value is p c,j . The individual best position of particle j in the (c - 1)-th iteration is P best,j = [batchsize j_best , f i,j_best , and the corresponding actual performance value is p best,j . If p c,j > p best,j , then X j_c = [batchsize j_c , f i,j_cis the new personal best position of particle j, that is, X j_c =[batchsize j_c ,f i,j_c is the personal best position P of particle j in this iteration best,j , if p c,j ≤p best,j , then keep the personal best position P of the previous iteration best,j =[batchsize best,j ,f i,best,j as the personal best position of particle j in this iteration, and the global best position G in this iteration best is updated. Similarly, in the entire particle swarm, compare the actual performance values corresponding to the personal best positions of all particles, determine the maximum value among these values, and find the corresponding personal best position. This position is determined as the global best position of this iteration.
[0061] Through this process, the particle swarm optimization algorithm can dynamically adjust the positions of particles to find the optimal combination of batch size and kernel frequency, thereby optimizing the performance of the computing device. At the same time, the particle swarm algorithm can adaptively search the solution space and gradually approach the optimal solution.
[0062] In an exemplary embodiment, the process of updating the position of a particle in the next iteration according to the predicted performance value, the actual performance value, the personal best position, and the global best position includes: obtaining the gradient information corresponding to the position of the particle in this iteration; updating the position of the particle in the next iteration according to the predicted performance value, the actual performance value, the personal best position, the global best position, and the gradient information.
[0063] In this embodiment, for the position X of particle j in this iteration j_c =[batchsize j_c ,f i,j_c calculate the gradient information of its corresponding actual performance value. The gradient information provides local information about the parameter change of the performance value with respect to the position of particle j, so as to more precisely adjust the position of the particle and make the particle move faster in the direction of increasing performance value. By introducing gradient information, the particle swarm optimization algorithm can search the solution space more effectively, especially in the case where the performance value function is complex or nonlinear. This method combines the global search ability of the particle swarm optimization algorithm and the local search ability of the gradient descent method, improving the performance and robustness of the algorithm.
[0064] In an exemplary embodiment, the process of updating the position of a particle in the next iteration according to the predicted performance value, the actual performance value, the individual best position, the global best position, and the gradient information includes: updating the velocity of the particle in the next iteration according to the velocity update relation, and the velocity update relation is ; updating the position of the particle in the next iteration according to the position update relation, and the position update relation is ; where j is the label of the particle, w is the inertia weight, c1, c2, c3, and c4 are all learning factors, r1, r2, r3, and r4 are all random numbers, is the velocity of the particle in the next iteration, is the predicted performance value of the particle in the current iteration, Gradient j is the gradient information corresponding to the position of the particle in the current iteration, is the position of the particle in the next iteration, is the position of the particle in the current iteration, V j is the velocity of the particle in the current iteration, P best,j is the individual best position of the particle in the current iteration, G best is the global best position in the current iteration, f(X j ) is the actual performance value of the particle in the current iteration.
[0065] It can be understood that the inertia weight controls the degree to which particle j retains the velocity of the current iteration, and the learning factor is used to control the degree to which particle j moves in different directions. The learning factor is a number randomly drawn from the interval [0,1], and the random number is used to increase the randomness and exploration ability of the algorithm.
[0066] In an exemplary embodiment, after obtaining multiple combinations of batch sizes and kernel frequencies of a computing device, the device parameter adjustment method further includes: calculating the actual performance value of the computing device under each combination of batch sizes and kernel frequencies; filling the position points of a preset two-dimensional grid based on the multiple combinations of batch sizes and kernel frequencies and their corresponding actual performance values to obtain the current energy efficiency gradient map; calculating the gradient information corresponding to each position point on the current energy efficiency gradient map; the process of obtaining the gradient information corresponding to the position of the particle in the current iteration includes: determining the current combination of batch size and kernel frequency in the position of the particle in the current iteration; determining the target position point corresponding to the current combination of batch size and kernel frequency on the current energy efficiency gradient map; using the gradient information corresponding to the target position point as the gradient information corresponding to the position of the particle in the current iteration.
[0067] In this embodiment, after obtaining multiple combinations of batch sizes and core frequencies of the computing device, the computing device is made to run one by one under each combination of batch sizes and core frequencies. It can be understood that considering the possible impact of the test order on the results, for example, before changing the combination of batch sizes and core frequencies, a certain cooling time is given to the computing device to avoid the influence of the previous combination of batch sizes and core frequencies on the test results of the next combination of batch sizes and core frequencies. Then, for each combination of batch sizes and core frequencies, the number of tests and the duration of each test are determined according to the specific requirements of the test, as long as the data under each configuration is representative and reliable enough. The actual performance values of the computing device running under each combination of batch sizes and core frequencies are obtained, and then the actual performance values corresponding to each combination of batch sizes and core frequencies are filled into the position points of the two-dimensional grid, thereby creating an energy efficiency gradient map of the computing device. This map provides an intuitive way to understand the performance under different configurations and can help determine the optimal combination of batch sizes and core frequencies.
[0068] In an exemplary embodiment, the process of calculating the actual performance value of the computing device under each combination of batch sizes and core frequencies includes: for each combination of batch sizes and core frequencies, obtaining the energy efficiency data and latency data of the computing device under the combination of batch sizes and core frequencies, and calculating the actual performance value under the combination of batch sizes and core frequencies based on the energy efficiency data and latency data.
[0069] Specifically, the actual performance value under the combination of batch sizes and core frequencies can be calculated according to the following relationship: ; where is the combination of batch sizes and core frequencies, is the actual performance value under the combination of batch sizes and core frequencies, is the energy efficiency data under the combination of batch sizes and core frequencies, is the latency data under the combination of batch sizes and core frequencies, is the first preset weight value, is the second preset weight value.
[0070] Specifically, in this embodiment, after controlling the computing device to run under a certain combination of batch sizes and core frequencies, the energy efficiency data and latency data of the computing device are obtained. Among them, the energy efficiency data is the energy consumed by each batch of calculations of the computing device under the batch size and core frequency, and the latency data is the latency time of the computing device under the combination of batch sizes and core frequencies. The energy efficiency data can be obtained by reading the real-time power value of the graphics card of the computing device and integrating it with the latency time to obtain the total energy value.
[0071] To comprehensively consider energy efficiency and latency, this embodiment defines an inference revenue model, and the formula is as follows: ; ; where is the inference revenue model, batchsize is the batch size, s is in seconds, is the number of images processed per second, p is the power consumption, t is the time, E is the total energy consumption, W is in watts, is the combination of batch size and core frequency, is the actual performance value under the combination of batch size and core frequency, is the energy efficiency data under the combination of batch size and core frequency, is the latency data under the combination of batch size and core frequency, is the first preset weight value, is the second preset weight value.
[0072] It can be understood that as the batch size increases, both energy efficiency and latency increase. Therefore, the performance trade-off is the ratio of the two, f i is the frequency of the i-th core of the computing device, that is, the i-th core frequency of the computing device. Among them, the first preset weight value and the second preset weight value , are used to balance the importance of energy efficiency and latency. In short, this model aims to optimize the "revenue" of the inference task by considering energy efficiency and latency. The first preset weight and the second preset weight are set according to actual engineering needs, allowing the model to adjust the degree of emphasis on energy efficiency and latency according to the requirements of the task. For example, if the first preset weight is greater than the second preset weight, this inference revenue model pays more attention to energy efficiency rather than latency. Conversely, if the first preset weight is less than the second preset weight, the inference revenue model pays more attention to latency. The goal of this optimization problem is to find an optimal b and f i value to maximize .
[0073] Through it shows how to maximize the inference revenue model by selecting the optimal batch size and core frequency , which is expressed as an optimization problem and can be expressed as: max b,fi , b and f i respectively need to satisfy: b ∈ N, f i ∈ C gpu .
[0074] In the above representation, max means to find the maximum value of the objective function where a represents a value, b represents the batch size, which must be an element of the set N of positive integers, and fi represents the GPU core frequency, which must be an element of the core frequency set C gpu in is the maximized revenue model, based on energy efficiency and latency as well as the first preset weight value and the second preset weight value to determine.
[0075] In an exemplary embodiment, the method for adjusting computing device parameters further includes: performing data fitting on the energy efficiency data of multiple core frequencies to obtain an energy efficiency fitting curve; performing data fitting on the latency data under multiple batch sizes to obtain a latency fitting curve; and filling the position points in the current energy efficiency gradient map where the actual performance values are not filled based on the energy efficiency fitting curve and the latency fitting curve.
[0076] In this embodiment, considering that when deploying a model in a cloud data center, in order to effectively balance the time and resources of data sampling, a coarse-grained and large-span method is adopted to initially establish a two-dimensional grid map of values. Therefore, in order to further improve the accuracy, comprehensiveness, and reliability of the energy efficiency gradient map, this embodiment also performs data fitting on the basis of the initial data, which not only improves the time efficiency and cost efficiency, but also, due to its flexibility and adaptability, can easily cope with changes in hardware and workloads without having to perform comprehensive data sampling every time, thereby continuously optimizing and iterating the model in a dynamically changing environment.
[0077] Specifically, data fitting can be performed on the energy efficiency data of multiple core frequencies to obtain an energy efficiency fitting curve, including but not limited to performing data fitting on the energy efficiency data of multiple core frequencies through a Fourier curve. The Fourier curve can handle periodic changes well and is suitable for dealing with regular changes that may occur at certain frequencies.
[0078] When performing Fourier curve fitting, the general form of the Fourier curve is: ; where EE(f i ) represents the energy efficiency data of the core frequency, a0 is the average or DC component, a n and b n are Fourier coefficients, which are parameters to be determined through the fitting process, and N is the maximum harmonic number selected. Parameter estimation: Use numerical methods (such as the least squares method) to estimate the Fourier coefficients a n and b n . This usually involves constructing a cost function, such as the mean square error, and then optimizing these coefficients to minimize the cost function.
[0079] Data fitting can be performed on the latency data under multiple batch sizes to obtain a latency fitting curve, including but not limited to using a rational function curve fitting. Rational functions are suitable for describing non-linear relationships and help to better understand the variation of latency with frequency and batch size.
[0080] Of course, in addition to using the above curves for data fitting, the least squares method can also be used to perform curve fitting on the energy efficiency data and latency data for each batch size, or on the energy efficiency data and latency data for each core frequency, to ensure the accuracy and reliability of the fitting.
[0081] Then, for the position points in the current energy efficiency gradient map that are not filled with actual performance values, according to the batch size and core frequency combination corresponding to this position point, without controlling the computing device to run at this batch size and core frequency combination, directly determine the energy efficiency data corresponding to the batch size and core frequency combination based on the energy efficiency fitting curve, determine the latency data corresponding to the batch size and core frequency combination based on the latency fitting curve, calculate the actual performance value corresponding to the batch size and core frequency combination according to the energy efficiency data and latency data determined on the fitting curve, and fill it into this position point.
[0082] In this embodiment, it is not necessary to perform actual tests on each possible batch size and core frequency combination, which can significantly reduce the time and resources required for testing. Through the fitting curve, the data in the untested area can be quickly estimated, thus accelerating and improving the creation process of the energy efficiency gradient map.
[0083] It can be understood that for particle j, according to the position X of particle j in this iteration j_c =[batchsize j_c ,f i,j_c match the corresponding target position point on the current energy efficiency gradient map, and use the actual performance value corresponding to this target position point as the actual performance value of the computing device in this iteration. If no exactly matching position point is found, interpolation or finding the closest position point may be required, which can be set according to actual engineering needs and is not specifically limited in this embodiment. It can be understood that by traversing the energy efficiency gradient map, finding the position point that exactly matches the position X j_c =[batchsize j_c ,f i,j_c ensures that the adjustment is based on the most accurate data, improves the accuracy of the adjustment, and avoids performance degradation caused by adjustment based on inaccurate data. When there is no exactly matching position point, interpolation or the method of finding the closest point is used to estimate the performance data. Even when the energy efficiency gradient map data is incomplete, reasonable performance estimation and adjustment can be performed, improving the flexibility and robustness of the system.
[0084] In an exemplary embodiment, for each position point on the current energy efficiency gradient map, the process of calculating the gradient information corresponding to the position point includes: for each position point on the current energy efficiency gradient map, determining a first neighboring point and a second neighboring point corresponding to the position point on the current energy efficiency gradient map; the core frequency of the first neighboring point is the same as that of the position point, and the batch size of the first neighboring point is different from that of the position point, the core frequency of the second neighboring point is different from that of the position point, and the batch size of the second neighboring point is the same as that of the position point; calculating a first gradient value of the position point based on the actual performance value of the first neighboring point and the actual performance value of the position point; calculating a second gradient value of the position point based on the actual performance value of the second neighboring point and the actual performance value of the position point; and obtaining the gradient information corresponding to the position point by using the first gradient value and the second gradient value.
[0085] In this embodiment, for each position point on the energy efficiency gradient map, its corresponding gradient information is calculated. The gradient information represents the change trend of the actual energy efficiency performance value when adjusting the batch size and core frequency at this position point. Specifically, for the filled energy efficiency gradient map, the gradient information of each position point is calculated, and the gradient information can reveal in which direction to adjust the batch size and core frequency to maximize the actual energy efficiency performance value.
[0086] First, determine the first neighboring point in the batch size direction and the second neighboring point in the core frequency direction of each position point. For example, if the position point is (b, f i ), the first neighboring point may be , that is, the core frequency of the first neighboring point is the same as that of the position point, and the second neighboring point may be , that is, the batch size of the second neighboring point is the same as that of the position point, where and are the minimum increments of the batch size and frequency.
[0087] Then calculate the first gradient value of the position point based on R(b, f i ) and , calculate the second gradient value of the position point based on R(b, f i ) and , and obtain the gradient information corresponding to the position point by using the first gradient value and the second gradient value. It can be understood that the gradient information can be represented as a vector, which contains two components: one corresponding to the gradient in the batch size direction, and the other corresponding to the gradient in the core frequency direction.
[0088] Specifically, a first difference between the actual performance value of the first neighboring point and the actual performance value of the position point can be obtained, as well as a first increment between the batch size of the first neighboring point and the batch size of the position point. Based on the first difference and the first increment, a first gradient value can be calculated. A second difference between the actual performance value of the second neighboring point and the actual performance value of the position point can be obtained, as well as a second increment between the core frequency of the second neighboring point and the core frequency of the position point. Based on the second difference and the second increment, a second gradient value can be calculated.
[0089] For the batch size direction, calculate the difference: ; is the first difference, is the actual performance value of the first neighboring point, R(b,f i ) is the actual performance value of the position point.
[0090] For the core frequency direction, calculate the difference: ; is the second difference, is the actual performance value of the second neighboring point, R(b,f i ) is the actual performance value of the position point. The gradient can be approximated by dividing these differences by the corresponding increments: gradient direction b = ; gradient direction fi = ; Obtain the gradient vector, and the gradient vector can be expressed as: = (gradient direction b , gradient direction fi ).
[0091] Through this process, the gradient information of each position point can be obtained. The gradient information characterizes in which direction adjusting the batch size and kernel frequency can maximize the actual energy efficiency performance value, facilitating the optimization of the performance of the computing device. It can be understood that the gradient information provides the rate of performance change in the directions of batch size and kernel frequency, enabling the particle swarm optimization algorithm to more accurately identify the direction of performance improvement, thereby more effectively searching the solution space. Utilizing the gradient information, the particles can move faster in the direction of increasing performance value, which helps the particle swarm optimization algorithm to accelerate the convergence to the optimal solution or approximate optimal solution. The gradient information reduces the need for random search, making the particle swarm optimization algorithm more efficient during the search process. In the case where the performance value function is complex or there are multiple local optimal solutions, the gradient information can help the algorithm avoid being trapped in local optimal solutions. Since the gradient information provides a clear search direction, it can reduce unnecessary iteration times, thereby reducing the consumption of computing resources. In this embodiment, by combining the time-series network prediction and gradient map information, the particle swarm optimization algorithm can more effectively navigate the parameter space and is expected to find the performance-optimal batch size and core frequency settings faster. In summary, the present invention proposes a systematic method to optimize the energy efficiency and performance of GPUs in cloud data centers. By precisely adjusting the batch size and GPU kernel frequency, higher energy efficiency and improved performance are achieved. By collecting and analyzing data in detail to drive the decision-making process, the accuracy and practicality of the optimization strategy are ensured. The construction and utilization of the energy efficiency gradient map provide an intuitive view of the performance under different GPU kernel frequencies and batch sizes. A comprehensive benefit model considering energy efficiency and latency is developed to provide a multi-dimensional framework for evaluating the performance of different configurations. The use of meta-heuristic learning-based algorithms, such as particle swarm optimization, combined with time-series network prediction, provides an effective method for finding the optimal configuration. It provides practical guidance for the optimization of GPUs in cloud data centers, helps improve the operating efficiency, and reduces energy consumption.
[0092] In a second aspect, please refer to Figure 2 , the present invention also provides a device parameter adjustment system, including:
[0093] An acquisition module 11, configured to acquire multiple combinations of batch sizes and kernel frequencies of a computing device.
[0094] An initialization module 12, configured to initialize the positions of multiple particles in a particle swarm based on multiple combinations of batch sizes and kernel frequencies.
[0095] An iterative update module 13, configured to, for each particle in multiple particle swarms, obtain the predicted performance value and the actual performance value of the computing device in the current iteration according to the position of the particle in the current iteration, and update the position of the particle in the next iteration based on the predicted performance value and the actual performance value. Repeat this step until the iteration end condition is met.
[0096] A configuration module 14 is used to take the particle corresponding to the optimal actual performance value when the iteration end condition is met as the optimal solution, and control the operation of the computing device according to the batch size and kernel frequency combination in the optimal solution.
[0097] In an exemplary embodiment, the process of obtaining the actual performance value of the computing device according to the position of the particle in the current iteration includes: determining the current batch size and kernel frequency combination in the position of the particle in the current iteration; obtaining the actual performance value of the computing device when operating under the current batch size and kernel frequency combination.
[0098] In an exemplary embodiment, the process of obtaining the predicted performance value of the computing device according to the position of the particle in the current iteration includes: inputting the positions and corresponding actual performance values of the particle in a continuous preset number of iterations including the current iteration into a time series network to obtain the predicted performance value of the computing device in the current iteration.
[0099] In an exemplary embodiment, the process of updating the position of the particle in the next iteration based on the predicted performance value and the actual performance value includes: determining the individual best position of the particle in the current iteration based on the actual performance value of the particle in the current iteration; determining the global best position in the current iteration based on the individual best positions of each particle in the particle swarm in the current iteration; updating the position of the particle in the next iteration according to the predicted performance value, the actual performance value, the individual best position, and the global best position.
[0100] In an exemplary embodiment, the process of determining the individual best position of the particle in the current iteration based on the actual performance value of the particle in the current iteration includes: judging whether the actual performance value of the particle in the current iteration is greater than the actual performance value corresponding to the individual best position of the particle in the previous iteration; if so, determining the position of the particle in the current iteration as the individual best position of the particle in the current iteration; if not, determining the individual best position of the particle in the previous iteration as the individual best position of the particle in the current iteration.
[0101] In an exemplary embodiment, the process of determining the global best position in the current iteration based on the individual best positions of each particle in the particle swarm in the current iteration includes: determining the maximum value of the actual performance values corresponding to the individual best positions of each particle in the particle swarm in the current iteration; determining the individual best position corresponding to the maximum value as the global best position in the current iteration.
[0102] In an exemplary embodiment, the process of updating the position of a particle in the next iteration according to the predicted performance value, the actual performance value, the individual best position, and the global best position includes: obtaining the gradient information corresponding to the position of the particle in the current iteration; and updating the position of the particle in the next iteration according to the predicted performance value, the actual performance value, the individual best position, the global best position, and the gradient information.
[0103] In an exemplary embodiment, the process of updating the position of a particle in the next iteration according to the predicted performance value, the actual performance value, the individual best position, the global best position, and the gradient information includes: updating the velocity of the particle in the next iteration according to the velocity update relation, where the velocity update relation is ; updating the position of the particle in the next iteration according to the position update relation, where the position update relation is ; where j is the label of the particle, w is the inertia weight, c1, c2, c3, and c4 are all learning factors, r1, r2, r3, and r4 are all random numbers, is the velocity of the particle in the next iteration, is the predicted performance value of the particle in the current iteration, Gradient j is the gradient information corresponding to the position of the particle in the current iteration, is the position of the particle in the next iteration, is the position of the particle in the current iteration, V j is the velocity of the particle in the current iteration, P best,j is the individual best position of the particle in the current iteration, G best is the global best position in the current iteration, f(X j ) is the actual performance value of the particle in the current iteration.
[0104] In an exemplary embodiment, after obtaining multiple combinations of batch sizes and kernel frequencies of a computing device, the device parameter adjustment system is further configured to: calculate the actual performance value of the computing device under each combination of batch sizes and kernel frequencies; fill the position points of a preset two-dimensional grid based on the multiple combinations of batch sizes and kernel frequencies and their corresponding actual performance values to obtain the current energy efficiency gradient map; calculate the gradient information corresponding to each position point on the current energy efficiency gradient map; the process of obtaining the gradient information corresponding to the position of the particle in the current iteration includes: determining the current combination of batch size and kernel frequency in the position of the particle in the current iteration; determining the target position point corresponding to the current combination of batch size and kernel frequency on the current energy efficiency gradient map; and using the gradient information corresponding to the target position point as the gradient information corresponding to the position of the particle in the current iteration.
[0105] In an exemplary embodiment, for each position point on the current energy efficiency gradient map, the process of calculating the gradient information corresponding to the position point includes: for each position point on the current energy efficiency gradient map, determining a first neighboring point and a second neighboring point corresponding to the position point on the current energy efficiency gradient map; the core frequency of the first neighboring point is the same as that of the position point, and the batch size of the first neighboring point is different from that of the position point, the core frequency of the second neighboring point is different from that of the position point, and the batch size of the second neighboring point is the same as that of the position point; calculating a first gradient value of the position point based on the actual performance value of the first neighboring point and the actual performance value of the position point; calculating a second gradient value of the position point based on the actual performance value of the second neighboring point and the actual performance value of the position point; and obtaining the gradient information corresponding to the position point by using the first gradient value and the second gradient value.
[0106] In an exemplary embodiment, the process of calculating the first gradient value of the position point based on the actual performance value of the first neighboring point and the actual performance value of the position point includes: obtaining a first difference between the actual performance value of the first neighboring point and the actual performance value of the position point; obtaining a first increment between the batch size of the first neighboring point and the batch size of the position point; and calculating the first gradient value based on the first difference and the first increment.
[0107] In an exemplary embodiment, the process of calculating the second gradient value of the position point based on the actual performance value of the second neighboring point and the actual performance value of the position point includes: obtaining a second difference between the actual performance value of the second neighboring point and the actual performance value of the position point; obtaining a second increment between the core frequency of the second neighboring point and the core frequency of the position point; and calculating the second gradient value based on the second difference and the second increment.
[0108] In an exemplary embodiment, the process of calculating the actual performance value of the computing device under each combination of batch size and core frequency includes: for each combination of batch size and core frequency, obtaining the energy efficiency data and latency data of the computing device under the combination of batch size and core frequency, and calculating the actual performance value under the combination of batch size and core frequency based on the energy efficiency data and the latency data.
[0109] In an exemplary embodiment, the process of calculating the actual performance value under the combination of batch size and core frequency based on the energy efficiency data and the latency data includes: calculating the actual performance value under the combination of batch size and core frequency based on a first relational expression, and the first relational expression is ; where is the combination of batch size and core frequency, is the actual performance value under the combination of batch size and core frequency, Is the energy efficiency data under the combination of batch size and core frequency, Is the latency data under the combination of batch size and core frequency, Is the first preset weight value, Is the second preset weight value.
[0110] In an exemplary embodiment, the iteration end condition includes reaching the maximum number of iterations, or the change amount of the actual performance value corresponding to the global optimal position in a continuous preset number of iterations including the current iteration is less than a preset threshold.
[0111] In a third aspect, the present invention further provides a computer program product, including a computer program / instructions, which when executed by a processor, implement the steps of the device parameter adjustment method described in any one of the above embodiments.
[0112] For the introduction of a computer program product provided by the present invention, please refer to the above embodiments, and the present invention will not be elaborated herein.
[0113] A computer program product provided by the present invention has the same beneficial effects as the above device parameter adjustment method.
[0114] In a fourth aspect, please refer to Figure 3 , the present invention further provides an electronic device, including:
[0115] A memory 21 for storing a computer program;
[0116] A processor 22 for implementing the steps of the device parameter adjustment method described in any one of the above embodiments when executing the computer program.
[0117] The electronic device further includes:
[0118] An input interface 23, connected to the processor 22 via a communication bus 26, for obtaining externally imported computer programs, parameters, and instructions, and storing them in the memory 21 under the control of the processor 22. The input interface can be connected to an input device to receive parameters or instructions manually input by the user. The input device can be a touch layer covered on the display screen, or a button, trackball, or touchpad provided on the terminal housing.
[0119] A display unit 24, connected to the processor 22 via a communication bus 26, for displaying the data sent by the processor 22. The display unit can be a liquid crystal display screen or an electronic ink display screen, etc.
[0120] The network port 25 is connected to the processor 22 via the communication bus 26 and is used for communication connection with various external terminal devices. The communication technology adopted for this communication connection can be a wired communication technology or a wireless communication technology, such as Mobile High-Definition Link (MHL) technology, Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), Wi-Fi technology, Bluetooth communication technology, Bluetooth Low Energy (BLE) communication technology, communication technology based on IEEE 802.11s, etc.
[0121] For the introduction of an electronic device provided by the present invention, please refer to the above embodiments, and the present invention will not be elaborated herein.
[0122] An electronic device provided by the present invention has the same beneficial effects as the above device parameter adjustment method.
[0123] Fifthly, the present invention further provides a server, which includes a plurality of computing devices and the above electronic device.
[0124] For the introduction of a server provided by the present invention, please refer to the above embodiments, and the present invention will not be elaborated herein.
[0125] A server provided by the present invention has the same beneficial effects as the above device parameter adjustment method.
[0126] Sixthly, please refer to Figure 4 , the present invention further provides a computer-readable storage medium 30, on which a computer program 31 is stored. When the computer program 31 is executed by a processor, it implements the steps of the device parameter adjustment method described in any one of the above embodiments.
[0127] The computer-readable storage medium 30 may include: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.
[0128] For the introduction of a computer-readable storage medium provided by the present invention, please refer to the above embodiments, and the present invention will not be elaborated herein.
[0129] A computer-readable storage medium provided by the present invention has the same beneficial effects as the above device parameter adjustment method.
[0130] It should also be noted that in this specification, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0131] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for adjusting device parameters, characterized in that, Including: Obtain multiple combinations of batch sizes and core frequencies of a computing device; Initialize the positions of multiple particles in a particle swarm based on the multiple combinations of batch sizes and core frequencies. For each particle, randomly select one combination of batch size and core frequency as the initial position of the particle; For each particle in multiple particle swarms, according to the position of the particle in the current iteration, obtain the predicted performance value and the actual performance value of the computing device in the current iteration, and update the position of the particle in the next iteration based on the predicted performance value and the actual performance value. Repeat this step until the iteration end condition is met; Take the particle corresponding to the optimal actual performance value when the iteration end condition is met as the optimal solution, and control the operation of the computing device according to the combination of batch size and core frequency in the optimal solution.
2. The method for adjusting device parameters according to claim 1, characterized in that The process of obtaining the actual performance value of the computing device according to the position of the particle in the current iteration includes: Determine the current combination of batch size and core frequency in the position of the particle in the current iteration; Obtain the actual performance value of the computing device when running under the current combination of batch size and core frequency.
3. The method for adjusting device parameters according to claim 1, characterized in that The process of obtaining the predicted performance value of the computing device according to the position of the particle in the current iteration includes: Input the positions and corresponding actual performance values of the particle in consecutive preset iterations including the current iteration into a time series network to obtain the predicted performance value of the computing device in the current iteration.
4. The method for adjusting device parameters according to claim 1, wherein The process of updating the position of the particle in the next iteration based on the predicted performance value and the actual performance value includes: Determine the individual best position of the particle in the current iteration based on the actual performance value of the particle in the current iteration; Determine the global best position in the current iteration based on the individual best positions of all particles in the particle swarm in the current iteration; Update the position of the particle in the next iteration according to the predicted performance value, the actual performance value, the individual best position, and the global best position.
5. The method for adjusting device parameters according to claim 4, wherein, The process of determining the individual best position of the particle in the current iteration based on the actual performance value of the particle in the current iteration includes: Judge whether the actual performance value of the particle in the current iteration is greater than the actual performance value corresponding to the individual best position of the particle in the previous iteration; If so, determine the position of the particle in the current iteration as the individual best position of the particle in the current iteration; If not, determine the individual best position of the particle in the previous iteration as the individual best position of the particle in the current iteration.
6. The method for adjusting device parameters according to claim 4, characterized in that, The process of determining the global best position in the current iteration based on the individual best positions of all particles in the particle swarm in the current iteration includes: Determine the maximum value of the actual performance values corresponding to the individual best positions of all particles in the particle swarm in the current iteration; Determine the individual best position corresponding to the maximum value as the global best position in the current iteration.
7. The method for adjusting device parameters according to claim 4, wherein The process of updating the position of the particle in the next iteration according to the predicted performance value, the actual performance value, the individual best position, and the global best position includes: Obtain the gradient information corresponding to the position of the particle in the current iteration; Update the position of the particle in the next iteration according to the predicted performance value, the actual performance value, the individual best position, the global best position, and the gradient information.
8. The method for adjusting device parameters according to claim 7, characterized in that The process of updating the position of the particle in the next iteration according to the predicted performance value, the actual performance value, the individual best position, the global best position, and the gradient information includes: Update the velocity of the particle in the next iteration according to the velocity update relationship, where the velocity update relationship is ; Update the position of the particle in the next iteration according to the position update relationship, where the position update relationship is ; Among them, j is the label of the particle, w is the inertia weight, c1, c2, c3, and c4 are all learning factors, and r1, r2, r3, and r4 are all random numbers. is the velocity of the particle in the next iteration. is the predicted performance value of the particle in this iteration, Gradient j is the gradient information corresponding to the position of the particle in this iteration. is the position of the particle in the next iteration. is the position of the particle in this iteration, V j is the velocity of the particle in this iteration, P best,j is the individual best position of the particle in this iteration, G best is the global best position in this iteration, f(X j ) is the actual performance value of the particle in this iteration.
9. The method for adjusting device parameters according to claim 7, wherein, After obtaining multiple combinations of batch sizes and kernel frequencies of the computing device, the device parameter adjustment method further includes: Calculate the actual performance value of the computing device under each combination of batch size and kernel frequency; Fill the position points of the preset two-dimensional grid based on multiple combinations of batch sizes and kernel frequencies and their corresponding actual performance values to obtain the current energy efficiency gradient map; For each position point on the current energy efficiency gradient map, calculate the gradient information corresponding to the position point; The process of obtaining the gradient information corresponding to the position of the particle in the current iteration includes: Determine the current batch size and kernel frequency combination in the position of the particle in the current iteration; Determine the target position point corresponding to the current batch size and kernel frequency combination on the current energy efficiency gradient map; Use the gradient information corresponding to the target position point as the gradient information corresponding to the position of the particle in the current iteration.
10. The method for adjusting device parameters according to claim 9, wherein, The process of calculating the gradient information corresponding to each position point on the current energy efficiency gradient map includes: For each position point on the current energy efficiency gradient map, determine the first adjacent point and the second adjacent point corresponding to the position point on the current energy efficiency gradient map; the kernel frequency of the first adjacent point is the same as the kernel frequency of the position point, the batch size of the first adjacent point is different from the batch size of the position point, the kernel frequency of the second adjacent point is different from the kernel frequency of the position point, and the batch size of the second adjacent point is the same as the batch size of the position point; Calculate the first gradient value of the position point based on the actual performance value of the first adjacent point and the actual performance value of the position point; Calculate the second gradient value of the position point based on the actual performance value of the second adjacent point and the actual performance value of the position point; Use the first gradient value and the second gradient value to obtain the gradient information corresponding to the position point.
11. The method for adjusting device parameters according to claim 10, wherein The process of calculating the first gradient value of the position point based on the actual performance value of the first adjacent point and the actual performance value of the position point includes: Obtain the first difference between the actual performance value of the first adjacent point and the actual performance value of the position point; Obtain the first increment between the batch size of the first adjacent point and the batch size of the position point; Calculate the first gradient value based on the first difference and the first increment.
12. The method for adjusting device parameters according to claim 10, wherein The process of calculating the second gradient value of the position point based on the actual performance value of the second neighboring point and the actual performance value of the position point includes: Obtaining a second difference between the actual performance value of the second neighboring point and the actual performance value of the position point; Obtaining a second increment between the core frequency of the second neighboring point and the core frequency of the position point; Calculating a second gradient value based on the second difference and the second increment.
13. The method for adjusting device parameters according to claim 9, wherein, The process of calculating the actual performance value of the computing device under each combination of the batch size and the core frequency includes: For each combination of the batch size and the core frequency, obtaining the energy efficiency data and the latency data of the computing device under the combination of the batch size and the core frequency, and calculating the actual performance value under the combination of the batch size and the core frequency based on the energy efficiency data and the latency data.
14. The method for adjusting device parameters according to claim 13, wherein The process of calculating the actual performance value under the combination of the batch size and the core frequency based on the energy efficiency data and the latency data includes: Calculate the actual performance value under the combination of the batch size and the kernel frequency based on the first relational expression, where the first relational expression is ; Among them, is the batch size and kernel frequency combination, is the actual performance value under the batch size and kernel frequency combination, is the energy efficiency data under the batch size and kernel frequency combination, is the latency data under the batch size and kernel frequency combination, is the first preset weight value, is the second preset weight value.
15. The method for adjusting device parameters according to any one of claims 1-14, characterized in that, The iteration end condition includes reaching the maximum number of iterations, or the change amount of the actual performance value corresponding to the global optimal position in a continuous preset number of iterations including the current iteration is less than a preset threshold.
16. An equipment parameter adjustment system, characterized in that, including: An obtaining module, configured to obtain multiple combinations of the batch size and the core frequency of the computing device; An initialization module, configured to initialize the positions of multiple particles in the particle swarm based on multiple combinations of the batch size and the core frequency, wherein for each particle, randomly select a combination of the batch size and the core frequency as the initial position of the particle; An iterative update module, configured to, for each particle in multiple particle swarms, obtain the predicted performance value and the actual performance value of the computing device in the current iteration according to the position of the particle in the current iteration, and update the position of the particle in the next iteration based on the predicted performance value and the actual performance value, and repeat this step until the iteration end condition is satisfied; A configuration module, configured to use the particle corresponding to the optimal actual performance value when the iteration end condition is satisfied as the optimal solution, and control the operation of the computing device according to the combination of the batch size and the core frequency in the optimal solution.
17. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, the steps of the device parameter adjustment method according to any one of claims 1-15 are implemented.
18. An electronic device, characterized in that, including: A memory, configured to store a computer program; A processor, configured to implement the steps of the device parameter adjustment method according to any one of claims 1-15 when executing the computer program.
19. A server, characterized in that, including multiple computing devices and the electronic device according to claim 18.
20. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the device parameter adjustment method according to any one of claims 1-15 are implemented.
Citation Information
Patent Citations
Micro-grid optimization scheduling system based on improved particle swarm optimization
CN118589588A