CPU (central processing unit) and GPU (graphics processing unit)-oriented computing power resource measurement method

By deploying benchmark models and data sets on smart devices, combining batch size and learning rate, determining the optimal running time, and using entropy weight method for comprehensive measurement, the problem of difficult to accurately reflect the computing power resources of CPU and GPU devices in the existing technology is solved, and the accurate measurement and evaluation of the computing power resources of smart devices is achieved.

CN120196432APending Publication Date: 2025-06-24NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510236854.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art is difficult to accurately reflect the computing power resources of CPU and GPU devices, resulting in inaccurate evaluation of computing power value of smart devices.

Method used

By deploying the benchmark model and dataset on the client device, running an exhaustive combination of batch size and learning rate, obtaining the optimal combination of batch size and learning rate for each device, and determining the optimal run time through the accuracy-convergence-time computing capability measurement method. Send computing resources and optimal running time to the server, and use the entropy weight method to comprehensively measure computing resources.

Benefits of technology

The accurate measurement of the computing power resources of smart devices is achieved, and the accuracy of the evaluation of the computing power value of smart devices is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196432A_ABST
    Figure CN120196432A_ABST
Patent Text Reader

Abstract

The invention discloses a computing power resource measurement method for a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit). The method comprises the steps that on each device of a client side, the hardware type of each device is obtained, and the hardware type of each device is a CPU chip or a GPU chip; deploying a plurality of reference models and a data set corresponding to each reference model on each piece of equipment; and when each device calls a function corresponding to each reference model. The method solves the problems that in the prior art, a weighted measurement mode based on resource attributes cannot reflect the computing power condition in real use of computing power resources although a calculation mode is simple and efficient, and a time test mode based on an artificial intelligence reference model cannot visually reflect the computing power condition although running time results are ranked. And the calculation power value evaluation of the intelligent equipment is not accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computing power resource perception, computing power resource measurement, and artificial intelligence deep learning. Specifically, it relates to a method for measuring computing power resources for CPUs and GPUs. Background Art

[0002] In recent years, the development of artificial intelligence computing services has continuously increased the demand for computing power. In response to the computing and running characteristics of its neural network models, various manufacturers have continuously developed and iterated dedicated computing acceleration chips in recent years. Among them, large-scale high-performance heterogeneous computing clusters configured with CPUs, GPUs, etc. are representative hardware architectures. The emergence of computing power network technology has connected various mature network service technologies, and further proposed to coordinate and schedule the computing, network, and storage resources distributed ubiquitously in space and time in the cloud, edge, and terminal, providing an end-to-end solution for the intelligent computing power service requirements in various industries such as biometric recognition, intelligent manufacturing, intelligent finance, intelligent healthcare, and intelligent transportation.

[0003] Computing power measurement is an important part of the underlying technology of the computing power network. The industrial community and academia have continuously explored computing measurement models and benchmark testing tools to measure the computing capabilities of various intelligent computing power devices. Currently, the most widely used computing resources are CPUs and GPUs. However, rapid iteration has resulted in a diverse range of categories, architectures, and versions. Therefore, the heterogeneity and dynamicity of devices pose a severe challenge to the unified measurement of computing power resources. In current research work, the weighted measurement method based on resource attributes, although simple and efficient in calculation, cannot reflect the actual computing power situation during the real use of computing power resources. The time test method based on artificial intelligence benchmark models, although there is a ranking of running time results, cannot intuitively reflect the computing power situation.

[0004] In this regard, there is a need to design a computing power resource measurement method that is applicable to various intelligent devices containing CPU or GPU computing resources, and can achieve unified measurement and evaluation methods, intuitive and reliable measurement results, and timely measurement results. Summary of the Invention

[0005] An embodiment of the present invention provides a method for measuring computing power resources for CPUs and GPUs, so as to at least solve the technical problem in the prior art that the weighted measurement method based on resource attributes, although simple and efficient in calculation, cannot reflect the actual computing power situation during the real use of computing power resources, and the time test method based on artificial intelligence benchmark models, although there is a ranking of running time results, cannot intuitively reflect the computing power situation, resulting in inaccurate evaluation of the computing power value of intelligent devices.

[0006] According to one aspect of the embodiments of the present invention, a computing power resource measurement method for CPU and GPU is provided. The method may include: on each device of the client, obtaining the hardware type of each device, where the hardware type of each device is a CPU chip or a GPU chip; deploying a number of benchmark models and the corresponding data sets for each benchmark model on each device, where each benchmark model includes its own number, and each benchmark model runs on the CPU chip or GPU chip corresponding to each device; when each device calls the function corresponding to each benchmark model, by exhaustively combining each batch size and learning rate of each benchmark model, running the corresponding benchmark model under each set of batch sizes and learning rates, obtaining the accuracy value and running time of each benchmark model under each set of batch sizes and learning rates; based on the accuracy value and running time of each benchmark model under each set of batch sizes and learning rates, obtaining the optimal batch size and learning rate combination of each benchmark model of each device; based on the optimal batch size and learning rate combination of each benchmark model of each device, processing each benchmark model through its corresponding data set, and during the processing, determining the optimal end round through the change amplitude of the accuracy value corresponding to the running round of each benchmark model; based on the difference between the time stamp of the optimal end round and the time stamp of the start round under the optimal batch size and learning rate combination of each benchmark model of each device, obtaining the optimal running time of each benchmark model of each device, where this optimal running time is used as the computing power measurement reference of this benchmark model on this device; sending the computing resources of each device of the client and the optimal running time of each benchmark model to the server; the server obtaining the average running time of the same-numbered benchmark models of each device of the client based on the optimal running time of the same-numbered benchmark models of each device of the client; obtaining the score of each benchmark model of each device based on the average running time and the optimal running time of the same-numbered benchmark models of each device of the client; calculating the weights of the same-numbered benchmark models of different devices using the entropy weight method based on the scores of each benchmark model of each device; and obtaining the target score of each device based on the weights of the corresponding benchmark models of each device and the scores of the corresponding benchmark models of each device, where the target score of each device is used to characterize the computing power resource measurement of each device.

[0007] Optionally, the obtaining the optimal batch size and learning rate combination of each benchmark model of each device based on the accuracy value and running time of each benchmark model under each set of batch sizes and learning rates includes: determining the set of batch size and learning rate with the maximum accuracy value and the shortest running time corresponding to each benchmark model of each device as the optimal batch size and learning rate combination of each benchmark model of each device.

[0008] Optionally, determining the optimal end round based on the change amplitude of the accuracy value corresponding to each running round of each reference model includes: when, during the running process of each reference model based on the running round, the change in the accuracy values corresponding to adjacent 5 rounds is less than the target threshold, determining the middle round of the 5 rounds as the optimal end round.

[0009] Optionally, obtaining the optimal running time of each reference model of each device based on the difference between the timestamp of the optimal end round and the timestamp of the start round under the optimal batch size and learning rate combination of each reference model of each device includes: based on the optimal batch size and learning rate combination of each reference model of each device, processing each of its reference models five times through its corresponding data set to obtain the optimal end round of each processing; based on the difference between the timestamp of the optimal end round and the timestamp of the start round of each processing, obtaining the running time of each reference model of each device during each processing; taking the average of the running times of each reference model of each device for the five processes to obtain the optimal running time of each reference model of each device.

[0010] Optionally, obtaining the score of each reference model of each device based on the average running time and the optimal running time of the reference models with the same number of each device on the client includes: when the attribute of the reference models with the same number of each device is a maximum value attribute, multiplying the quotient of the optimal running time and the average running time of the reference models with the same number of each device by 50 to obtain the score of each reference model of each device; when the attribute of the reference models with the same number of each device is a minimum value attribute, multiplying the quotient of the average running time and the optimal running time of the reference models with the same number of each device by 50 to obtain the score of each reference model of each device.

[0011] Optionally, obtaining the target score of each device based on the weight of the corresponding reference model of each device and the score of the corresponding reference model of each device includes: determining the sum of the products of the weight of the corresponding reference model of each device and the score of the corresponding reference model of each device as the target score of each device.

[0012] Advantages of the present invention: The present invention proposes a computing power resource measurement method for CPUs and GPUs. A client is deployed on an intelligent device to obtain the attribute information of each computing power resource of the device. After running a specified artificial intelligence deep learning benchmark model task, the resource information and test results are packaged and sent. Then, the server receives the data packets sent by the online devices in real time, and uses the entropy weight method for comprehensive measurement of computing power resources. To enable the intelligent device to achieve a relatively optimal test result, a grid search hyperparameter optimization method and a precision-convergence-time calculation ability measurement method are added to the benchmark model test. Based on the formed device resource status and benchmark operation data set, a benchmark operation prediction model is established to achieve real-time evaluation of the computing power value of the intelligent device. Description of the Drawings

[0013] The drawings described herein are used to provide a further understanding of the present invention and form a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 is a flowchart of a computing power resource measurement method for CPUs and GPUs according to an embodiment of the present invention; Figure 2 is a block diagram of a computing power resource measurement architecture for CPUs and GPUs according to an embodiment of the present invention. Detailed Embodiments

[0014] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0015] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects and are used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0016] Embodiment 1 According to an embodiment of the present invention, there is provided a computing power resource measurement method for a CPU and a GPU. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system including at least one set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0017] Figure 1 is a flowchart of a computing power resource measurement method for a CPU and a GPU according to an embodiment of the present invention, as Figure 1 shown, the method may include the following steps: Step S101, on each device of the client, obtain the hardware type of each device, where the hardware type of each device is a CPU chip or a GPU chip.

[0018] In the technical solution provided in step S101 of the present invention above, Figure 2 is a block diagram of a computing power resource measurement architecture for a CPU and a GPU according to an embodiment of the present invention, as Figure 2 shown, on each device of the client, detect the hardware type of each device, the hardware type of each device is a CPU chip or a GPU chip, that is, the reference model on each device runs on the CPU chip or the GPU chip, the attributes of the CPU chip are the number of cores, rated frequency, and real-time utilization rate of the CPU; the attributes of the GPU chip are the video memory, frequency, and real-time utilization rate of the GPU, etc.; the size and swap rate of the memory, etc.; the capacity and read / write speed of the disk, etc.; the storage space and read / write speed of the disk, etc.; the bandwidth, upload / download speed, and ping latency of the network, etc., and the attributes of the CPU chip and the GPU chip are Figure 2 the computing resources in the client.

[0019] Step S102, deploy a number of reference models and the data set corresponding to each reference model on each device, where each reference model includes its own number, and each reference model runs on the CPU chip or the GPU chip corresponding to each device.

[0020] In the technical solution provided in step S102 of the present invention above, as Figure 2 shown, deploy a number of reference models and the data set corresponding to each reference model on each device, where each reference model includes its own number, for example Figure 2 the reference model_1, reference model_2,..., reference model_n in, and each reference model runs on the CPU chip or the GPU chip corresponding to each device, where the reference model is a neural network model.

[0021] Step S103, when each device calls the function corresponding to each of its baseline models, by exhaustively enumerating each combination of batch size and learning rate for each baseline model, running its corresponding baseline model under each set of batch size and learning rate, and obtaining the accuracy value and running time of each baseline model under each set of batch size and learning rate.

[0022] In the technical solution provided in step S103 of the present invention, when each device calls the function corresponding to each of its baseline models, by exhaustively enumerating each combination of batch size (Batch Size) and learning rate (LearningRate) for each baseline model, such as (BS, LR), running its corresponding baseline model under each set of batch size and learning rate, and obtaining the accuracy value and running time of each baseline model under each set of batch size and learning rate.

[0023] Step S104, based on the accuracy value and running time of each baseline model of each device under each set of batch size and learning rate, obtaining the optimal batch size and learning rate combination of each baseline model of each device.

[0024] In the technical solution provided in step S104 of the present invention, sorting the accuracy value and running time of each baseline model of each device under each set of batch size and learning rate to obtain the optimal batch size and learning rate combination of each baseline model of each device , where is the maximum accuracy value of each baseline model under each set of batch size and learning rate, is the minimum running time of each baseline model under each set of batch size and learning rate.

[0025] Step S105, based on the optimal batch size and learning rate combination of each baseline model of each device, processing each baseline model through its corresponding dataset. During the processing, determining the optimal end round through the change amplitude of the accuracy value corresponding to the running rounds of each baseline model.

[0026] In the technical solution provided in step S105 of the present invention, processing each baseline model of each device through its corresponding dataset according to the optimal batch size and learning rate combination of each baseline model of each device. During the actual processing, obtaining the optimal end round through the change amplitude of the accuracy value corresponding to the running rounds of each baseline model.

[0027] Step S106, based on the difference between the time stamp of the optimal end round and the time stamp of the start round under the optimal batch size and learning rate combination of each baseline model of each device, obtaining the optimal running time of each baseline model of each device, where this optimal running time is used as a reference for measuring the computing power of this baseline model on this device.

[0028] In the technical solution provided in step S106 of the present invention, the difference between the timestamp of the optimal end round and the timestamp of the start round under the optimal batch size and learning rate combination of each reference model of each device is used as the optimal running time of each reference model of each device.

[0029] Step S107: Send the computing resources of each device of the client and the optimal running time of each reference model to the server.

[0030] In the technical solution provided in step S107 of the present invention, the computing resources of each device of the client and the optimal running time of each reference model are sent to the server.

[0031] Step S108: The server obtains the average running time of the reference models with the same number of each device of the client based on the optimal running time of the reference models with the same number of each device of the client.

[0032] In the technical solution provided in step S108 of the present invention, the server calculates the average value based on the optimal running time of the reference models with the same number of each device of the client to obtain the average running time of the reference models with the same number of each device of the client. For example, the optimal running time of device_1 corresponding to reference model_2 is 0.3s, the optimal running time of device_2 corresponding to reference model_2 is 0.6s, and the optimal running time of device_3 corresponding to reference model_2 is 0.7s. Then the average running time of device_1, device_2, and device_3 corresponding to reference model_2 is 0.53s.

[0033] Step S109: Based on the average running time and the optimal running time of the reference models with the same number of each device of the client, obtain the score of each reference model of each device.

[0034] In the technical solution provided in step S109 of the present invention, the average running time and the optimal running time of the reference models with the same number of each device of the client are calculated to obtain the score of each reference model of each device.

[0035] Step S110: Based on the scores of each reference model of each device, use the entropy weight method to calculate the weights of the reference models with the same number of different devices.

[0036] In the technical solution provided in step S110 of the present invention, the score matrix of each reference model of each device is normalized by Max - Min , to obtain , and then the probability matrix of each reference model of each device is calculated , and then the information entropy of each reference model of each device is calculated , and then, according to the redundancy of information entropy, calculate the information entropy weight of each reference model of each device , and finally calculate the comprehensive score of the computing power resources of the i-th device , represents the representation of the i-th device and the j-th reference model, which is a two-dimensional matrix.

[0037] Step S111: Obtain the target score of each device based on the weight of the corresponding reference model of each device and the score of the corresponding reference model of each device. Among them, the target score of each device is used to represent the measurement of the computing power resources of each device.

[0038] In the technical solution provided in step S111 of the present invention above, calculate the weight of the corresponding reference model of each device and the score of the corresponding reference model of each device to obtain the target score of each device.

[0039] The above method of this embodiment will be further introduced below.

[0040] As an optional embodiment, in step S104, the method for obtaining the optimal batch size and learning rate combination of each reference model of each device based on the accuracy value and running time of each reference model under each group of batch sizes and learning rates includes: determining the group of batch sizes and learning rates with the maximum accuracy value and the shortest running time corresponding to each reference model of each device as the optimal batch size and learning rate combination of each reference model of each device.

[0041] In this embodiment, determine the group of batch sizes and learning rates with the maximum accuracy value and the shortest running time corresponding to each reference model of each device as the optimal batch size and learning rate combination of each reference model of each device.

[0042] As an optional embodiment, in step S105, the method for obtaining the optimal running time of each reference model of each device based on the difference between the time stamp of the optimal end epoch and the time stamp of the start epoch under the optimal batch size and learning rate combination of each reference model of each device includes: based on the optimal batch size and learning rate combination of each reference model of each device, process each reference model through its corresponding data set five times to obtain the optimal end epoch of each process; based on the difference between the time stamp of the optimal end epoch and the time stamp of the start epoch of each process, obtain the running time of each reference model of each device at each time of processing; take the average of the running times of each reference model of each device for the five processes to obtain the optimal running time of each reference model of each device.

[0043] In this embodiment, after determining the optimal batch size and learning rate combination, the criterion for judging the end of the operation of the baseline model will be determined by the accuracy-convergence-time method, that is, the training accuracy of the baseline model reaches the threshold, and the change in the training accuracy of the baseline model thereafter is not obvious. For example, for a combination of the optimal batch size and learning rate of a baseline model of a device, the corresponding dataset is used to perform an operation on the baseline model once, and it is obtained that at the 32nd, 33rd, 34th, 35th, and 36th rounds, the training accuracy of the baseline model reaches the threshold, and the change in the training accuracy of the baseline model thereafter is , so the optimal end round for a baseline model to perform an operation once through its corresponding dataset is 34. According to the difference between the timestamp of the optimal end round and the timestamp of the start round for each processing, the running time of each baseline model of each device for each processing is obtained; the average value of the running time of each baseline model of each device for the five processes is calculated to obtain the optimal running time of each baseline model of each device.

[0044] As an alternative embodiment, step S109, obtaining the score of each baseline model of each device based on the average running time and the optimal running time of the baseline models with the same number of each device on the client includes: when the attribute of the baseline models with the same number of each device is a maximum value attribute, multiplying the quotient of the optimal running time and the average running time of the baseline models with the same number of each device by 50 to obtain the score of each baseline model of each device; when the attribute of the baseline models with the same number of each device is a minimum value attribute, multiplying the quotient of the average running time and the optimal running time of the baseline models with the same number of each device by 50 to obtain the score of each baseline model of each device.

[0045] In this embodiment, when the attribute of the baseline models with the same number of each device is a maximum value attribute, multiplying the quotient of the optimal running time and the average running time of the baseline models with the same number of each device by 50 to obtain the score of each baseline model of each device; when the attribute of the baseline models with the same number of each device is a minimum value attribute, multiplying the quotient of the average running time and the optimal running time of the baseline models with the same number of each device by 50 to obtain the score of each baseline model of each device.

[0046] As an alternative embodiment, step S111, obtaining the target score of each device based on the weight of the corresponding baseline model of each device and the score of the corresponding baseline model of each device includes: determining the sum of the products of the weight of the corresponding baseline model of each device and the score of the corresponding baseline model of each device as the target score of each device.

[0047] In this embodiment, the sum of the products of the weights of the corresponding reference models of each device and the scores of the corresponding reference models of each device is determined as the target score of each device.

[0048] After obtaining the target score of each device, as Figure 2 shown, obtain a device to be predicted, train the DNN network model according to historical data (a number of reference models are deployed on each device and the data set corresponding to each reference model), obtain a successfully trained DNN network model, and process the device to be predicted through the successfully trained DNN network model to obtain the target score of the device to be predicted.

[0049] In an embodiment of the present invention, on each device of the client, the hardware type of each device is obtained, where the hardware type of each device is a CPU chip or a GPU chip; a number of benchmark models and the data set corresponding to each benchmark model are deployed on each device, where each benchmark model includes its own number, and each benchmark model runs on the CPU chip or GPU chip corresponding to its device; when each device calls the function corresponding to its benchmark model, by exhaustively enumerating each combination of the batch size and learning rate of each benchmark model, the corresponding benchmark model is run under each set of batch sizes and learning rates, and the accuracy value and running time of each benchmark model under each set of batch sizes and learning rates are obtained; based on the accuracy value and running time of each benchmark model under each set of batch sizes and learning rates, the optimal combination of batch size and learning rate of each benchmark model of each device is obtained; based on the optimal combination of batch size and learning rate of each benchmark model of each device, the corresponding benchmark model of each device is processed through its corresponding data set. During the processing, the optimal end round is determined through the change amplitude of the accuracy value corresponding to the running round of each benchmark model; based on the difference between the time stamp of the optimal end round and the time stamp of the start round under the optimal combination of batch size and learning rate of each benchmark model of each device, the optimal running time of each benchmark model of each device is obtained, where this optimal running time is used as a reference for measuring the computing power of this benchmark model on this device; the computing resources of each device of the client and the optimal running time of each benchmark model are sent to the server; the server obtains the average running time of the same-numbered benchmark models of each device of the client based on the optimal running time of the same-numbered benchmark models of each device of the client; based on the average running time and the optimal running time of the same-numbered benchmark models of each device of the client, the score of each benchmark model of each device is obtained; based on the scores of each benchmark model of each device, the entropy weight method is used to calculate the weights of the same-numbered benchmark models of different devices; based on the weights of the corresponding benchmark models of each device and the scores of the corresponding benchmark models of each device, the target score of each device is obtained, where the target score of each device is used to represent the measurement of the computing power resources of each device, solving the technical problem in the prior art that the weighted measurement method based on resource attributes, although simple and efficient in calculation, cannot reflect the computing power situation in the actual use of computing power resources, and the time test method based on the artificial intelligence benchmark model, although there is a ranking of running time results, cannot intuitively reflect the computing power situation, resulting in inaccurate evaluation of the computing power value of intelligent devices, and achieving the technical effect of improving the accuracy of evaluating the computing power value of intelligent devices by proposing a method for measuring computing power resources for CPUs and GPUs.

[0050] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0051] In the above embodiments of the present invention, the descriptions of the various embodiments each have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0052] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of units can be a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.

[0053] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0054] In addition, the functional units in the various embodiments of the present invention can be integrated in a first processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0055] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A CPU and GPU computing resource measurement method, characterized in that: include: On each device of the client, obtain the hardware type of each device, wherein the hardware type of each device is a CPU chip or a GPU chip; Deploy several benchmark models and data sets corresponding to each benchmark model on each device, wherein each benchmark model includes its own serial number, and each benchmark model runs on a CPU chip or GPU chip corresponding to each device; When each device calls the function corresponding to each benchmark model, it runs the corresponding benchmark model under each batch size and learning rate by exhaustively enumerating each combination of batch size and learning rate of each benchmark model, and obtains the accuracy value and running time of each benchmark model under each batch size and learning rate. Based on the accuracy and running time of each benchmark model under each set of batch size and learning rate, the optimal batch size and learning rate combination of each benchmark model on each device is obtained; Based on the optimal batch size and learning rate combination of each benchmark model of each device, each benchmark model is processed through its corresponding data set. During the processing, the optimal end round is determined by the change range of the accuracy value corresponding to the running round of each benchmark model; Based on the difference between the timestamp of the optimal end round and the timestamp of the start round under the optimal batch size and learning rate combination of each benchmark model of each device, the optimal running time of each benchmark model of each device is obtained, wherein the optimal running time is used as a reference for measuring the computing power of the benchmark model on this device; Send the computing resources of each device of the client and the optimal running time of each benchmark model to the server; The server obtains the average running time of the same numbered benchmark model of each device of the client based on the optimal running time of the same numbered benchmark model of each device of the client; Based on the average running time and the best running time of the same numbered benchmark model of each device of the client, the score of each benchmark model of each device is obtained; Based on the score of each benchmark model of each device, the entropy weight method is used to calculate the weight of the benchmark models with the same number of different devices; Based on the weight of the corresponding benchmark model of each device and the score of the corresponding benchmark model of each device, a target score of each device is obtained, wherein the target score of each device is used to characterize the computing resource measurement of each device.

2. The method according to claim 1, characterized in that The method of obtaining the optimal batch size and learning rate combination of each benchmark model of each device based on the accuracy value and running time of each benchmark model under each set of batch size and learning rate includes: A set of batch sizes and learning rates with the maximum accuracy value and the shortest running time corresponding to each benchmark model of each device is determined as the optimal batch size and learning rate combination of each benchmark model of each device.

3. The method according to claim 2, characterized in that The step of determining the optimal end round by the accuracy value change range corresponding to the running round of each benchmark model includes: When the change in the accuracy value corresponding to the five adjacent rounds of each benchmark model is less than the target threshold during the running process based on the running rounds, the middle round of the five rounds is determined as the optimal ending round.

4. The method according to claim 3, characterized in that The optimal running time of each benchmark model of each device is obtained based on the difference between the timestamp of the optimal end round and the timestamp of the start round under the optimal batch size and learning rate combination of each benchmark model of each device, including: Based on the optimal batch size and learning rate combination of each benchmark model on each device, each benchmark model is processed five times through its corresponding dataset to obtain the optimal end round for each processing; Based on the difference between the timestamp of the optimal end round and the timestamp of the start round of each processing, the running time of each benchmark model of each device during each processing is obtained; The running time of each benchmark model of each device in the five processings was averaged to obtain the optimal running time of each benchmark model of each device.

5. The method according to claim 4, characterized in that The step of obtaining the score of each benchmark model of each device based on the average running time and the optimal running time of the benchmark model with the same number of each device of the client includes: When the attribute of the same numbered benchmark model of each device is a maximum value attribute, the quotient of the optimal running time and the mean running time of the same numbered benchmark model of each device is multiplied by 50 to obtain the score of each benchmark model of each device; When the attribute of the same-numbered benchmark model of each device is a minimum attribute, the quotient of the mean running time and the optimal running time of the same-numbered benchmark model of each device is multiplied by 50 to obtain the score of each benchmark model of each device.

6. The method according to claim 5, characterized in that The step of obtaining a target score for each device based on the weight of the corresponding benchmark model of each device and the score of the corresponding benchmark model of each device comprises: The sum of the products of the weight of the corresponding benchmark model of each device and the score of the corresponding benchmark model of each device is determined as the target score of each device.

7. A processor, characterized in that: The processor is used to run a program, wherein the program executes the method according to any one of claims 1 to 6 when running.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Method and device for dynamically metering computing power of intelligent computing center cloud platform based on computing power operation task

    CN121233441A