Evaluation method and device, electronic device and computer-readable storage medium
By obtaining the scoring results and hardware information of the target algorithm and conducting iterative testing, the problem of not being able to comprehensively evaluate the algorithm performance in the existing technology is solved, and a more accurate evaluation and optimization of the relationship between the algorithm and hardware performance is achieved.
Patent Information
- Application Number
- CN202110358457.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-01
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-04-01
AI Technical Summary
The prior art cannot comprehensively and efficiently evaluate algorithm performance, especially when considering the relationship between algorithm characteristics and hardware performance.
By obtaining the scoring results of the target algorithm and the physical hardware information of the running algorithm, the preset algorithm load is used for iterative testing, and the performance evaluation results are obtained based on the test data, which comprehensively reflects the relationship between algorithm characteristics and hardware performance.
A more accurate evaluation of the operating performance of the target algorithm on the current hardware is achieved, providing more comprehensive information to optimize the use of algorithms and hardware.
Smart Images

Figure CN113419941B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computing technology, and in particular to an evaluation method and device, an electronic device, and a computer-readable storage medium. Background Art
[0002] With the development of artificial intelligence technology, more and more fields have begun to use artificial intelligence technology to solve various problems in various fields. In the application, the computational efficiency and accuracy of artificial intelligence technology have become the most critical factors in the application of artificial intelligence technology. In particular, due to the widespread popularity of artificial intelligence, the research and development of artificial intelligence technology, especially algorithms, has been promoted. It can be said that a large number of new algorithms are published every day, new frameworks are upgraded every month, or new artificial intelligence chips are released every few months. These new technological developments will have different impacts on the application of artificial intelligence. Therefore, how artificial intelligence researchers choose and evaluate new technologies when applying artificial intelligence technology has become the most important issue on the road to artificial intelligence application.
[0003] Therefore, there is a need for a technical solution that can improve the evaluation efficiency of artificial intelligence technology. Summary of the invention
[0004] The embodiments of the present application provide an evaluation method and device, an electronic device, and a computer-readable storage medium to address the defect in the prior art that the algorithm performance cannot be comprehensively and efficiently evaluated.
[0005] To achieve the above objectives, the present application provides an evaluation method, including:
[0006] Obtaining a scoring result of a target algorithm and physical hardware information for running the target algorithm;
[0007] Using at least one preset algorithm load to iteratively test the target algorithm according to the scoring result and the physical hardware information;
[0008] A performance evaluation result of the target algorithm is obtained according to the test data of the test, wherein the performance evaluation result indicates a relationship between a feature of the target algorithm and a physical hardware performance.
[0009] The present application also provides an evaluation device, including:
[0010] An acquisition module, used to acquire the scoring result of the target algorithm and the physical hardware information for running the target algorithm;
[0011] A collection module, configured to use at least one preset algorithm load to iteratively test the target algorithm according to the scoring result and the physical hardware information;
[0012] An analysis module is used to obtain a performance evaluation result of the target algorithm based on the test data of the test, wherein the performance evaluation result indicates a relationship between a feature of the target algorithm and physical hardware performance.
[0013] The present application also provides an electronic device, including:
[0014] Memory, used to store programs;
[0015] The processor is used to run the program stored in the memory, and the evaluation method provided in the embodiment of the present application is executed when the program is run.
[0016] The embodiment of the present application further provides a computer-readable storage medium on which a computer program executable by a processor is stored, wherein the program, when executed by the processor, implements the evaluation method provided in the embodiment of the present application.
[0017] The evaluation method and device, electronic device, and computer-readable storage medium provided in the embodiments of the present application can use the scoring results for the target algorithm and the physical hardware information running the target algorithm as input to perform iterative testing on the target algorithm using a preset load, thereby obtaining a performance evaluation result based on the test data. Since the scoring results and hardware information reflecting the algorithm characteristics of the target algorithm are used as input for testing during the iterative test, the test results can comprehensively reflect the performance of the target algorithm running on the current hardware, thereby obtaining a more accurate performance evaluation result.
[0018] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present application. Also, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:
[0020] Figure 1 A schematic diagram of the principle architecture of the evaluation scheme provided in the embodiment of the present application;
[0021] Figure 2 A flowchart of an embodiment of the evaluation method provided by the present application;
[0022] Figure 3 A flowchart of another embodiment of the evaluation method provided by the present application;
[0023] Figure 4 A schematic diagram of the structure of an evaluation device embodiment provided in the present application;
[0024] Figure 5 A schematic diagram of the structure of an electronic device embodiment provided in this application. DETAILED DESCRIPTION
[0025] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0026] Embodiment 1
[0027] The solution provided in the embodiments of the present application can be applied to any device or system with algorithm execution, such as a distributed server, etc. Figure 1 A schematic diagram of the principle architecture of the evaluation scheme provided in the embodiment of the present application, Figure 1 The illustrated architecture is merely one example of an architecture to which the technical solution of the present application can be applied.
[0028] With the development of artificial intelligence technology, more and more fields have begun to use artificial intelligence technology to solve various problems in various fields. In the application, the computational efficiency and accuracy of artificial intelligence technology have become the most critical factors in the application of artificial intelligence technology. In particular, due to the widespread popularity of artificial intelligence, the research and development of artificial intelligence technology, especially algorithms, has been promoted. It can be said that a large number of new algorithms are published every day, new frameworks are upgraded every month, or new artificial intelligence chips are released every few months. These new technological developments will have different impacts on the application of artificial intelligence. Therefore, how artificial intelligence researchers choose and evaluate new technologies when applying artificial intelligence technology has become the most important issue on the road to artificial intelligence application.
[0029] In the prior art, since the use of artificial intelligence algorithms or frameworks depends on various physical computing hardware, such as CPU or GPU, the operation of tasks running on the computing hardware can only be analyzed through the hardware resource analysis tools provided by the hardware manufacturer. For example, the GPU performance analysis tool provided by NVIDIA can analyze the execution of the GPU, and can analyze the performance bottleneck of the GPU and the computing power of the operator according to the timeline; the CPU-oriented performance analysis tool provided by a company can support functions such as instruction / cache / branch prediction; there are also CPU-oriented analysis tools that can be used to analyze the performance bottleneck of the CPU architecture.
[0030] However, the above-mentioned analysis tools of the prior art are all oriented to the hardware usage of the device, such as the utilization rate of hardware resources for analysis. For example, the balance analysis of computing / data transmission / memory and other aspects can be carried out, but since these solutions are only based on the hardware level, they basically do not consider the characteristics of the load, especially the lack of consideration of the characteristics of artificial intelligence computing tasks, and do not consider the performance analysis between distributed multiple devices. In other words, the hardware-based analysis tools of the prior art are only based on the utilization rate of the current hardware or hardware units for analysis, so they cannot really consider the rationality of resource allocation from the computing task side, especially when distributed computing is currently used, that is, using multiple computing units connected by communication lines, or even using part of the hardware resources of multiple computing units to perform a computing task. Such a fade-out analysis tool based on a single hardware cannot provide comprehensive analysis and suggestions, especially it cannot provide resource allocation optimization suggestions based on task characteristics.
[0031] In the above-mentioned prior art, since only the hardware usage evaluation tool provided by the hardware used to run the artificial intelligence algorithm can be used to understand the running performance of the artificial intelligence algorithm, such an evaluation tool cannot take into account the characteristics of the artificial intelligence algorithm running on the hardware, and therefore cannot obtain a comprehensive evaluation result for the artificial intelligence algorithm. In other words, the existing solutions are all oriented to analyze the utilization of the equipment, the micro-architecture level analysis of the hardware equipment, and the balance analysis of computing / transmission / memory, etc., which can analyze the performance bottleneck of the hardware to a certain extent, but cannot provide optimization suggestions and specific shortcomings.
[0032] For example, in Figure 1 In the architecture shown in , the evaluation method according to the embodiment of the present application can use the calculation results of the target algorithm on the artificial intelligence scoring platform as the input that can reflect the algorithm characteristics of the target algorithm. For example, in the embodiment of the present application, the artificial intelligence scoring platform can be a scoring platform of a single algorithm type, or it can be a test platform covering a variety of mainstream algorithm models, and can support the current mainstream algorithm running hardware, such as GPU / third-party algorithm chips, etc. Therefore, the score for the target algorithm itself can be obtained by inputting the target algorithm into the scoring platform. In addition, Figure 1As shown in , the target algorithm is usually run on a hardware device, for example, a distributed server or a current mainstream distributed GPU system. Therefore, in an embodiment of the present application, the hardware information of the physical hardware running the target algorithm can be further obtained. For example, in an embodiment of the present application, these hardware information may include hardware performance specifications, hardware topology, and hardware networking forms. Specifically, for example, when a chip is used as a computing carrier to run the target algorithm, the model and computing power of the chip can be obtained as chip specifications. In another embodiment, the chip specifications can also be modeled according to the model and computing power, so that the chip specifications can be automatically identified. In addition, since a computing system with a distributed structure is usually used to run various algorithms recently, the connection between the various computing units constituting the computing system (for example, the connection mode, i.e., the topology, and the communication bandwidth) also has a certain influence on the operation performance of the algorithm. Therefore, in this case, the topology and communication bandwidth between the computing units of the hardware running the target algorithm can be further obtained. Here, according to the networking form of the cluster constituting the computing system, the communication bandwidth between the servers or the communication bandwidth between the GPUs can be obtained.
[0033] Therefore, by obtaining the above input, the scoring result reflecting the algorithm characteristics of the target algorithm itself and the hardware information reflecting the hardware specifications of the target algorithm can be obtained as the input of the evaluation method of the embodiment of the present application. Therefore, compared with the solution in the prior art that only considers the usage rate or specification information of the hardware to evaluate the algorithm, the embodiment of the present application can use more comprehensive information to evaluate the algorithm.
[0034] In addition, after obtaining the algorithm information and hardware information, automated testing can be performed based on these input information and test data can be collected. In the embodiment of the present application, different algorithm loads can be used to perform testing and performance collection, and the test and performance data can be stored in a structured manner to facilitate subsequent various analyses. In addition, various analysis models can be pre-stored in the storage space, so that when performing analysis, the corresponding analysis model can be adapted according to the collected test data for analysis.
[0035] In addition, in order to facilitate analysis, the target algorithm can be iteratively tested during the above test to collect sufficient performance data. Of course, the number of iterations can be selected and determined according to actual needs, as long as it can ensure that sufficient performance data is collected and the amount of data does not cause excessive pressure on storage and analysis.
[0036] Therefore, when analyzing, the embodiments of the present application can, for example, consider the iteration time and the number of layers of the neural network and the amount of data entering and leaving the computing unit, such as the GPU, to calculate the computing power of each computing unit to execute the algorithm. In particular, the computing power within each iteration cycle can be calculated and compared to determine whether the computing power within each iteration is balanced, or whether the deviation from the theoretical value of the computing power requirement of the target algorithm determined according to the analysis model is too large. Therefore, by analyzing the computing power, the calculation of the target algorithm on the current hardware can be evaluated.
[0037] In addition, it is also possible to calculate whether the computing power consumption of the computing unit in each iteration cycle is balanced, whether there is any jitter in computing power consumption, and because distributed computing units are used to execute the target algorithm, the parallelism of the operation sequence on each computing unit can also be calculated based on the test data, for example, whether there is a waiting time process for a certain computing unit, etc., and the computing power consumption and time for the additional data transmission of the computing unit can also be calculated. Similarly, these evaluation results can reflect whether the computing efficiency of the distributed computing system is reasonable.
[0038] In addition, the computing efficiency of the entire hardware system for the operation of the target algorithm can be calculated to see whether the computing power consumption of the target algorithm meets the design value or expectation. Figure 1 The display part shown in is displayed to the user. On the other hand, optimization suggestions can be calculated based on these evaluation results, such as optimization of transmission efficiency, optimization of utilization of computing units, etc., and such optimization suggestions can also be displayed to the user.
[0039] Specifically, since an evaluation result that fully reflects the performance of the target algorithm running on the current hardware can be obtained in the embodiment of the present application, an algorithm or hardware adjustment suggestion can be further generated according to the evaluation result in the embodiment of the present application. For example, the evaluation result in the embodiment of the present application can be based on the computational efficiency of the hardware system for the operation of the target algorithm, that is, it can understand the consumption of computing power by the target algorithm. Therefore, in addition to giving an evaluation result such as the low computational efficiency of the target algorithm on the current hardware system, it can be further determined that the low computational efficiency is caused by the low memory in the hardware resources, which makes the hardware resources mismatched. Therefore, in this case, a hardware resource adjustment suggestion such as increasing the memory capacity can be further generated and sent to the operator or hardware provider for reference or to guide them to adjust the hardware system. Alternatively, it can also be further confirmed based on the evaluation result of low computational efficiency that the low computational efficiency is caused by the unreasonable deployment of the operator unit of the algorithm. In this case, the adjustment scheme of the algorithm can also be further generated and output to the user or the user of the algorithm for reference.
[0040] Therefore, the evaluation scheme provided in the embodiment of the present application can use the scoring results for the target algorithm and the physical hardware information running the target algorithm as input to use a preset load to iteratively test the target algorithm, thereby obtaining a performance evaluation result based on the test data. Since the scoring results and hardware information reflecting the algorithm characteristics of the target algorithm are used as input for testing during the iterative test, the test results can comprehensively reflect the performance of the target algorithm running on the current hardware, thereby obtaining a more accurate performance evaluation result.
[0041] The above embodiments are illustrations of the technical principles and exemplary application frameworks of the embodiments of the present application. The specific technical solutions of the embodiments of the present application are further described in detail below through multiple embodiments.
[0042] Embodiment 2
[0043] Figure 2 This is a flowchart of an embodiment of the evaluation method provided by the present application. The execution subject of the method can be various terminals or server devices with algorithm evaluation capabilities, or it can be a device or chip integrated on these devices. Figure 2 As shown, the evaluation method includes the following steps:
[0044] S201, obtaining the scoring result of the target algorithm and the physical hardware information for running the target algorithm.
[0045] In an embodiment of the present application, the scoring results of the target algorithm on various artificial intelligence scoring platforms can be obtained as the algorithm feature input for extracting the target algorithm. For example, in an embodiment of the present application, the artificial intelligence scoring platform can be a scoring platform for a single algorithm type, or it can be a test platform covering a variety of mainstream algorithm models, and can support the current mainstream algorithm running hardware, such as GPU / third-party algorithm chips, etc. Therefore, the score for the target algorithm itself can be obtained by inputting the target algorithm into the scoring platform. In addition, Figure 1 As shown in , the target algorithm usually runs on a hardware device, for example, a distributed server or a current mainstream distributed GPU system. Therefore, in an embodiment of the present application, the hardware information of the physical hardware running the target algorithm can be further obtained. By obtaining the above-mentioned input, the scoring result reflecting the algorithm characteristics of the target algorithm itself and the hardware information reflecting the hardware specifications of the target algorithm can be obtained as the input of the evaluation method of the embodiment of the present application. Therefore, compared with the solution of evaluating the algorithm by only considering the utilization rate or specification information of the hardware in the prior art, the embodiment of the present application can use more comprehensive information to evaluate the algorithm.
[0046] S202: Use at least one preset algorithm load to iteratively test the target algorithm according to the scoring result and the physical hardware information.
[0047] After various input information is obtained in step S201, at least one analysis algorithm may be used in step S202 to test the target algorithm. For example, in an embodiment of the present application, different algorithm loads may be used to perform testing and performance acquisition, and the test and performance data may be structured and stored to facilitate subsequent various analyses. In addition, various analysis models may be pre-stored in a storage space, so that when performing analysis, the corresponding analysis model may be adapted according to the collected test data for analysis.
[0048] In addition, in order to facilitate analysis, the target algorithm can be iteratively tested during the above test to collect sufficient performance data. Of course, the number of iterations can be selected and determined according to actual needs, as long as it can ensure that sufficient performance data is collected and the amount of data does not cause excessive pressure on storage and analysis.
[0049] S203, obtaining a performance evaluation result of the target algorithm according to the test data.
[0050] After the test data is obtained in step S202, a performance evaluation calculation may be further performed based on the test data in step S203. The performance evaluation results may indicate the relationship between the characteristics of the target algorithm and the performance of the physical hardware, so that the user or the server system itself may optimize the algorithm or hardware based on the performance evaluation results to achieve higher performance.
[0051] Therefore, the evaluation scheme provided in the embodiment of the present application can use the scoring results for the target algorithm and the physical hardware information running the target algorithm as input to use a preset load to iteratively test the target algorithm, thereby obtaining a performance evaluation result based on the test data. Since the scoring results and hardware information reflecting the algorithm characteristics of the target algorithm are used as input for testing during the iterative test, the test results can comprehensively reflect the performance of the target algorithm running on the current hardware, thereby obtaining a more accurate performance evaluation result.
[0052] Embodiment 3
[0053] Figure 3 This is a flowchart of another embodiment of the evaluation method provided by the present application. The execution subject of the method can be various terminals or server devices with algorithm evaluation capabilities, or it can be a device or chip integrated on these devices. Figure 3 As shown, the evaluation method includes the following steps:
[0054] S301, using the artificial intelligence testing platform to calculate the scoring result of the target algorithm.
[0055] In an embodiment of the present application, the scoring results of the target algorithm on various artificial intelligence scoring platforms can be obtained as the algorithm feature input for extracting the target algorithm. For example, in an embodiment of the present application, the artificial intelligence scoring platform can be a scoring platform of a single algorithm type, or it can be a test platform covering a variety of mainstream algorithm models, and can support the current mainstream algorithm running hardware, such as GPU / third-party algorithm chips, etc. Therefore, the score for the target algorithm itself can be obtained by inputting the target algorithm into the scoring platform.
[0056] S302, obtaining physical hardware information for running the target algorithm.
[0057] like Figure 1 As shown in , the target algorithm is usually run on a hardware device, such as a distributed server or a current mainstream distributed GPU system. Therefore, in an embodiment of the present application, hardware information of the physical hardware running the target algorithm can be further obtained.
[0058] For example, in an embodiment of the present application, these hardware information may include performance specifications, hardware topology, and hardware networking forms of the hardware. Specifically, for example, when a chip is used as a computing carrier to run a target algorithm, the model and computing power of the chip can be obtained as chip specifications. In another embodiment, the chip specifications can also be modeled according to the model and computing power, so that the chip specifications can be automatically identified. In addition, since a computing system with a distributed structure is generally used to run various algorithms recently, the connection between the various computing units constituting the computing system (for example, the connection mode, i.e., the topology, and the communication bandwidth) also has a certain influence on the operation performance of the algorithm. Therefore, in this case, the topology and communication bandwidth between the computing units of the hardware running the target algorithm can also be further obtained. Here, according to the networking form of the cluster constituting the computing system, the communication bandwidth between the servers or the communication bandwidth between the GPUs can be obtained.
[0059] Therefore, the scoring result reflecting the algorithm characteristics of the target algorithm itself obtained in step S301 and the hardware information reflecting the hardware specifications of the target algorithm obtained in step S302 are used as inputs of the evaluation method of the embodiment of the present application. Therefore, compared with the solution in the prior art that only considers the usage rate or specification information of the hardware to evaluate the algorithm, the embodiment of the present application can use more comprehensive information to evaluate the algorithm.
[0060] S303: Use at least one preset algorithm load to iteratively test the target algorithm according to the scoring result and the physical hardware information.
[0061] After obtaining the scoring result in step S301 and the hardware information in step S302, at least one analysis algorithm may be used to test the target algorithm in step S303. For example, in an embodiment of the present application, different algorithm loads may be used to perform testing and performance acquisition, and the test and performance data may be structured and stored to facilitate subsequent various analyses. In addition, various analysis models may be pre-stored in a storage space, so that when performing analysis, the corresponding analysis model may be adapted according to the collected test data for analysis.
[0062] In addition, in order to facilitate analysis, the target algorithm can be iteratively tested during the above test to collect sufficient performance data. Of course, the number of iterations can be selected and determined according to actual needs, as long as it can ensure that sufficient performance data is collected and the amount of data does not cause excessive pressure on storage and analysis.
[0063] S304, using the time of a single test of the iterative test, the number of layers of the neural network of the target algorithm, and the data entering and exiting the computing unit running the target algorithm to calculate the computing power of the computing unit to execute the target algorithm.
[0064] In step S304, the embodiment of the present application can consider the time of a single test of the iterative test in step S303, the number of layers of the neural network of the target algorithm, and the data entering and exiting the computing unit running the target algorithm to calculate the computing power of each computing unit to execute the algorithm. In particular, the computing power within each iteration cycle can be calculated and compared to determine whether the computing power within each iteration is balanced, or whether the deviation from the theoretical value of the computing power requirement of the target algorithm determined according to the analysis model is too large. Therefore, by analyzing the computing power, the calculation of the target algorithm on the current hardware can be evaluated.
[0065] S305, time slicing a single test according to predetermined time units.
[0066] S306, calculating the computing time of the computing unit of the physical hardware in each time slice as the performance evaluation result.
[0067] In addition, in the embodiment of the present application, the performance of the algorithm, such as the two indicators of computing power and transmission, is calculated based on time slicing. For example, each iteration, that is, the time of a single test, can be time-sliced according to a predetermined time unit, such as 10ms, in step S305, so that the computing time ratio and transmission ratio of each 10ms period can be calculated in step S306, so as to evaluate whether the consumption of each second of the computing unit is balanced and whether there is jitter during the entire training process based on the evaluation result.
[0068] S307, calculating the computing power consumption of the physical hardware in each time slice as a performance evaluation result.
[0069] In addition, the computing power consumption in each time slice can also be calculated in step S307 to reflect the computing efficiency of the target algorithm on the hardware. For example, such an evaluation result can be used to determine whether the computing power consumption of the target algorithm meets the design.
[0070] S308, using a preset analysis model to obtain a performance evaluation result of the target algorithm according to the test data.
[0071] In addition, in step S308, a pre-stored analysis model may be used to analyze the test data obtained in step S303. In the embodiment of the present application, these analysis models may be pre-stored in a database, so that the above test may be performed offline according to the needs of the user.
[0072] In the case where the performance evaluation result of the target algorithm on the current hardware is obtained in steps S306 to S308, the evaluation result is based on, for example, the physical hardware information obtained in step S302 and the calculation results calculated by various parameters of the algorithm used in step S304, so in the embodiment of the present application, an evaluation result that fully reflects the performance of the target algorithm running on the current hardware can be obtained. In the embodiment of the present application, an algorithm or hardware adjustment suggestion can be further generated based on the performance evaluation results obtained in steps S306 to S308.
[0073] For example, in steps S306 to S308, it is determined that the target algorithm consumes too much computing power. Therefore, the cause of the evaluation result can be further analyzed based on the hardware information obtained in step S302 and the information of the algorithm used in step S304, and adjustment suggestions for the cause can be generated by looking up a table or using a model. For example, after understanding the consumption of computing power by the target algorithm in steps S306 to S308, it can be further determined that the low computing efficiency is caused by the low memory in the hardware resources, which causes the hardware resources to be mismatched. Therefore, in this case, hardware resource adjustment suggestions such as increasing memory capacity can be further generated and sent to the operator or hardware provider for reference or guidance to adjust the hardware system. Alternatively, it can also be further confirmed based on the evaluation result of low computing efficiency that the low computing efficiency is caused by the unreasonable deployment of the operator unit of the algorithm. For example, in particular, different test methods can be used to test the algorithm from different angles in steps S306 to S308, so an adjustment scheme for the target algorithm can be generated based on at least one of the calculation results of steps S306 to S308 and output to the user or the user of the algorithm for reference. For example, when the calculation time of the calculation unit calculated in step S306 is too long compared to the transmission time, it is possible to further determine whether the calculation efficiency of a certain operator in the target algorithm is low or the convergence speed is slow due to unreasonable settings of convergence parameters, etc., based on which iterative step in the calculation process takes too long or whether the number of iterations is too many. Therefore, algorithm adjustment suggestions can be generated for the user based on such analysis results, or recommended adjustment values can be further given, etc., so that the user can adjust the algorithm based on the evaluation results generated by steps S306 to S308 and the suggestions provided, so as to further improve the computing efficiency of the target algorithm on hardware resources.
[0074] S309, displaying the performance evaluation result to the user.
[0075] Since the performance evaluation results obtained in steps S304-S308 can indicate the relationship between the characteristics of the target algorithm and the performance of the physical hardware, they can be displayed to the user in step S309, and optimization suggestions can be calculated based on these evaluation results, such as optimization of transmission efficiency or optimization of the utilization rate of computing units, and such optimization suggestions can also be displayed to the user. Therefore, the user or the server system itself can optimize the algorithm or hardware according to such performance evaluation results to achieve higher performance.
[0076] Therefore, the evaluation scheme provided by the embodiment of the present application can adopt the scoring result for the target algorithm and the physical hardware information of the target algorithm as input to adopt the preset load to iteratively test the target algorithm, so as to obtain the performance evaluation result according to the test data, because the scoring result and the hardware information reflecting the algorithm characteristics of the target algorithm are adopted as input to test together during the iterative test, so the test result can fully reflect the performance of the target algorithm running on the current hardware, so as to obtain a more accurate performance evaluation result. In addition, the embodiment of the present application can also further generate the adjustment suggestion scheme of the hardware resources or the algorithm itself according to the process data of the evaluation according to the evaluation result that can reflect the performance of the target algorithm, and can further generate specific adjustment suggestions, so that the user can refer to or guide the user to execute. Even in the embodiment of the present application, the adjustment suggestion generated in this way can also be used directly as the parameter update value in the iterative cycle to participate in iterative calculation, and realize automatic adjustment.
[0077] Embodiment 4
[0078] Figure 4 The schematic diagram of the structure of the evaluation device embodiment provided in the present application can be used to perform the following steps: Figure 2 and Figure 3 The method steps shown. Figure 4 As shown, the evaluation device may include: an acquisition module 41 , a collection module 42 and an analysis module 43 .
[0079] The acquisition module 41 may be used to acquire the scoring result of the target algorithm and the physical hardware information for running the target algorithm.
[0080] In an embodiment of the present application, the scoring results of the target algorithm on various artificial intelligence scoring platforms can be obtained by the acquisition module 41 as the algorithm feature input for extracting the target algorithm. For example, in an embodiment of the present application, the artificial intelligence scoring platform can be a scoring platform of a single algorithm type, or a test platform covering a variety of mainstream algorithm models, and can support the current mainstream algorithm running hardware, such as GPU / third-party algorithm chips, etc. Therefore, the score for the target algorithm itself can be obtained by inputting the target algorithm into the scoring platform and then input into the acquisition module 41. The target algorithm usually runs on a hardware device, such as a distributed server or a current mainstream distributed GPU system. Therefore, in an embodiment of the present application, the hardware information of the physical hardware running the target algorithm can be further obtained. For example, in an embodiment of the present application, the hardware information can include the performance specifications, hardware topology, and hardware networking form of the hardware. Specifically, for example, when a chip is used as a computing carrier to run the target algorithm, the chip model and computing power can be obtained as chip specifications. In another embodiment, the chip specifications can also be modeled according to the model and computing power, so that the chip specifications can be automatically identified. In addition, since a computing system with a distributed structure is often used to run various algorithms recently, the connection between the various computing units constituting the computing system (e.g., the connection mode, i.e., the topology, and the communication bandwidth) also has a certain influence on the operation performance of the algorithm. Therefore, in this case, the topology and communication bandwidth between the computing units of the hardware running the target algorithm can be further obtained. Here, according to the networking form of the cluster constituting the computing system, the communication bandwidth between servers or the communication bandwidth between GPUs can be obtained.
[0081] After obtaining the above input, the acquisition module 41 can obtain the scoring result reflecting the algorithm characteristics of the target algorithm itself and the hardware information reflecting the hardware specifications of the running of the target algorithm as the input of the evaluation method of the embodiment of the present application.
[0082] The acquisition module 42 may be configured to use at least one preset algorithm load to iteratively test the target algorithm according to the scoring result and the physical hardware information.
[0083] After the acquisition module 41 acquires various input information, the acquisition module 42 can use at least one analysis algorithm to test the target algorithm. For example, in the embodiment of the present application, different algorithm loads can be used to perform testing and performance acquisition, and the test and performance data can be stored in a structured manner to facilitate subsequent various analyses. In addition, various analysis models can be pre-stored in the storage space, so that when performing analysis, the corresponding analysis model can be adapted according to the collected test data for analysis.
[0084] In addition, in particular, for the convenience of analysis, the acquisition module 42 can perform iterative testing on the target algorithm during the above test so as to collect sufficient performance data. Of course, the number of iterations can be selected and determined according to actual needs, as long as it can ensure that sufficient performance data is collected and the amount of data does not cause excessive pressure on storage and analysis.
[0085] The analysis module 43 may be used to obtain a performance evaluation result of the target algorithm based on the test data.
[0086] After the acquisition module 42 acquires the test data, it can further perform performance evaluation calculations based on the test data through the analysis module 43. These performance evaluation results can indicate the relationship between the characteristics of the target algorithm and the physical hardware performance, so that the user or the server system itself can optimize the algorithm or hardware based on such performance evaluation results to achieve higher performance.
[0087] In addition, the analysis module 43 may be further configured to use a preset analysis model to obtain a performance evaluation result of the target algorithm according to the test data.
[0088] Specifically, the analysis module 43 may use pre-stored analysis models to analyze the performance of the algorithm based on the test data obtained by the acquisition module 42. In the embodiment of the present application, these analysis models may be pre-stored in a database, so that the above test may be performed offline according to the needs of the user.
[0089] The analysis module 43 may further include a computing power analysis unit 431, which is used to calculate the computing power of the computing unit to execute the target algorithm using the time of a single test of the iterative test, the number of layers of the neural network of the target algorithm, and the data entering and exiting the computing unit running the target algorithm.
[0090] In the embodiment of the present application, the computing power analysis unit 431 can consider the time of a single test of the iterative test performed by the acquisition module 42, the number of layers of the neural network of the target algorithm, and the data entering and exiting the computing unit running the target algorithm to calculate the computing power of each computing unit to execute the algorithm. In particular, the computing power in each iteration cycle can be calculated and compared to determine whether the computing power in each iteration is balanced, or whether the deviation from the theoretical value of the computing power requirement of the target algorithm determined according to the analysis model is too large. Therefore, the computing power analysis unit 431 can evaluate the computing situation of the target algorithm executed on the current hardware by analyzing the computing power. In the embodiment of the present application, the evaluation result obtained by the computing power analysis unit 431 may include at least one of the following: the number of core functions executed in a single test, the proportion of the core used for calculation and the core of the operating memory relative to all cores; the execution time of the target algorithm in a single test, the total execution time of the iterative test, the calculation time and data transmission time in a single test; the calculation and transmission consumption of each operation queue in a single test.
[0091] In addition, the analysis module 43 may further include an efficiency analysis unit 432 and a computing power consumption analysis unit 433 .
[0092] The efficiency analysis unit 432 may be used to time slice a single test according to predetermined time units, and calculate the computing time of the computing unit of the physical hardware in each time slice as a performance evaluation result.
[0093] The computing power consumption analysis unit 423 can be used to time slice a single test according to predetermined time units, and calculate the computing power consumption of the physical hardware in each time slice as a performance evaluation result.
[0094] Therefore, both the efficiency analysis unit 432 and the computing power consumption analysis unit 433 can calculate the performance of the algorithm based on time slicing, such as the two indicators of computing power and transmission. For example, the efficiency analysis unit 432 can time slice each iteration, that is, the time of a single test, according to a predetermined time unit, such as 10ms, so as to calculate the computing time ratio and the transmission ratio during each 10ms period, and then evaluate whether the consumption of the computing unit per second in the entire training process is balanced and whether there is jitter based on the evaluation result.
[0095] The computing power consumption analysis unit 433 can also calculate the computing power consumption in each time slice after time slicing, so as to reflect the computing efficiency of the target algorithm on the hardware. For example, such an evaluation result can be used to see whether the computing power consumption of the target algorithm is in line with the design.
[0096] In addition, the evaluation device according to the embodiment of the present application may further include a display module 44, which is used to display the performance evaluation results and / or optimization suggestions generated according to the performance evaluation results to the user.
[0097] Since the performance evaluation results obtained by the evaluation module can indicate the relationship between the characteristics of the target algorithm and the performance of the physical hardware, the display module 44 can display them to the user, and further, based on these evaluation results, optimization suggestions can be calculated, such as optimization of transmission efficiency, or optimization of the utilization rate of the computing unit, etc., and such optimization suggestions can also be displayed to the user. Therefore, the user or the server system itself can optimize the algorithm or hardware according to such performance evaluation results to achieve higher performance.
[0098] When the analysis module 43 obtains the performance evaluation result of the target algorithm on the current hardware, the evaluation result is calculated based on, for example, the physical hardware information obtained by the acquisition module 41 and the various parameters of the algorithm used by the acquisition module 42 to perform the test. Therefore, in the embodiment of the present application, an evaluation result that fully reflects the performance of the target algorithm running on the current hardware can be obtained. In the embodiment of the present application, the analysis module 43 can also further generate algorithm or hardware adjustment suggestions based on the obtained performance evaluation results.
[0099] For example, the analysis module 43 can determine that the target algorithm consumes too much computing power. Therefore, the cause of the evaluation result can be further analyzed based on the hardware information obtained by the acquisition module 41 and the information of the algorithm used by the acquisition module 42 to perform the test, and an adjustment suggestion for the cause can be generated by looking up a table or using a model. For example, after the analysis module 43 understands the consumption of computing power by the target algorithm, it can be further determined that the low computing efficiency is caused by the low memory in the hardware resources, which causes the hardware resources to be mismatched. Therefore, in this case, a hardware resource adjustment suggestion such as increasing the memory capacity can be further generated and sent to the operator or hardware provider for reference or to guide them to adjust the hardware system. Alternatively, it can also be further confirmed based on the evaluation result of low computing efficiency that the low computing efficiency is caused by the unreasonable deployment of the operator unit of the algorithm. For example, in particular, the analysis module 43 can use different test methods to evaluate the algorithm from different angles, so an adjustment scheme for the target algorithm can be generated based on the calculation results of at least one of the computing power analysis unit 431, the efficiency analysis unit 432 and the computing power consumption analysis unit 433, and output to the user or the user of the algorithm for reference. For example, when the calculation time of the computing unit calculated by the computing power analysis unit 431 is too long compared with the transmission time, it is possible to further determine whether the computational efficiency of a certain operator in the target algorithm is low or the convergence speed is slow due to unreasonable settings of convergence parameters, etc., based on which iterative step in the calculation process takes too long or whether the number of iterations is too many. Therefore, algorithm adjustment suggestions can be generated for the user based on such analysis results, or recommended adjustment values can be further given, etc., so that the user can adjust the algorithm based on the evaluation results and suggestions generated by the computing power analysis unit 431, the efficiency analysis unit 432 and the computing power consumption analysis unit 433, so as to further improve the computing efficiency of the target algorithm on hardware resources.
[0100] Therefore, the evaluation scheme provided by the embodiment of the present application can adopt the scoring result for the target algorithm and the physical hardware information of the target algorithm as input to adopt the preset load to iteratively test the target algorithm, so as to obtain the performance evaluation result according to the test data, because the scoring result and the hardware information reflecting the algorithm characteristics of the target algorithm are adopted as input to test together during the iterative test, so the test result can fully reflect the performance of the target algorithm running on the current hardware, so as to obtain a more accurate performance evaluation result. In addition, the embodiment of the present application can also further generate the adjustment suggestion scheme of the hardware resources or the algorithm itself according to the process data of the evaluation according to the evaluation result that can reflect the performance of the target algorithm, and can further generate specific adjustment suggestions, so that the user can refer to or guide the user to execute. Even in the embodiment of the present application, the adjustment suggestion generated in this way can also be used directly as the parameter update value in the iterative cycle to participate in iterative calculation, and realize automatic adjustment.
[0101] Embodiment 5
[0102] The internal functions and structure of the evaluation device are described above. The device can be implemented as an electronic device. Figure 5 This is a schematic diagram of the structure of an electronic device embodiment provided by this application. Figure 5 As shown, the electronic device includes a memory 51 and a processor 52 .
[0103] The memory 51 is used to store programs. In addition to the above programs, the memory 51 can also be configured to store various other data to support operations on the electronic device. Examples of such data include instructions for any application or method operating on the electronic device, contact data, phone book data, messages, pictures, videos, etc.
[0104] The memory 51 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0105] The processor 51 is not limited to a central processing unit (CPU), but may also be a processing chip such as a graphics processing unit (GPU), a field programmable gate array (FPGA), an embedded neural network processor (NPU) or an artificial intelligence (AI) chip. The processor 52 is coupled to the memory 51 and executes a program stored in the memory 51. When the program is running, the evaluation method of the above-mentioned embodiments 2 and 3 is executed.
[0106] Further, if Figure 5As shown, the electronic device may also include: a communication component 53, a power component 54, an audio component 55, a display 56 and other components. Figure 5 Only some components are shown schematically, which does not mean that the electronic device only includes Figure 5 Components shown.
[0107] The communication component 53 is configured to facilitate wired or wireless communication between the electronic device and other devices. The electronic device can access a wireless network based on a communication standard, such as WiFi, 3G, 4G or 5G, or a combination thereof. In an exemplary embodiment, the communication component 53 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 53 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0108] The power supply component 54 provides power to various components of the electronic device. The power supply component 54 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device.
[0109] The audio component 55 is configured to output and / or input audio signals. For example, the audio component 55 includes a microphone (MIC), and when the electronic device is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 51 or sent via the communication component 53. In some embodiments, the audio component 55 also includes a speaker for outputting audio signals.
[0110] The display 56 includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundary of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation.
[0111] Those skilled in the art can understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk and other media that can store program codes.
[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An evaluation method comprising: Obtaining the scoring result of the target algorithm and the physical hardware information for running the target algorithm; wherein the physical hardware is a chip, and the physical hardware information includes chip specifications, and the chip specifications are modeled according to the chip model and the chip computing power; Using at least one preset algorithm load to iteratively test the target algorithm according to the scoring result and the physical hardware information; A performance evaluation result of the target algorithm is obtained according to the test data of the test, wherein the performance evaluation result indicates a relationship between a feature of the target algorithm and a physical hardware performance.
2. The evaluation method according to claim 1, wherein: The physical hardware information includes: hardware performance specifications, hardware topology and hardware networking form.
3. The evaluation method according to claim 1, wherein: The scoring result of obtaining the target algorithm includes: The artificial intelligence testing platform is used to calculate the scoring results of the target algorithm.
4. The evaluation method according to claim 2, wherein: The hardware networking form includes the communication bandwidth between physical computing units.
5. The evaluation method according to claim 1, wherein: The step of obtaining the performance evaluation result of the target algorithm according to the test data of the test includes: The computing power of the computing unit in executing the target algorithm is calculated using the time of a single test of the iterative test, the number of layers of the neural network of the target algorithm, and the data entering and exiting the computing unit running the target algorithm.
6. The evaluation method according to claim 5, wherein: The performance evaluation results further include at least one of the following: The number of core functions executed in the single test, and the ratio of cores used for calculation and cores used for operating memory to all cores; The execution time of the target algorithm in the single test, the total execution time of the iterative test, the calculation time and the data transmission time in the single test; The computation and transmission consumption of each operation queue within the single test.
7. The evaluation method according to claim 1, wherein: The step of obtaining the performance evaluation result of the target algorithm according to the test data of the test includes: Time slices are performed for a single test according to predetermined time units; The computing time of the computing unit of the physical hardware in each time slice is calculated as the performance evaluation result.
8. The evaluation method according to claim 1, wherein: The step of obtaining the performance evaluation result of the target algorithm according to the test data of the test includes: Time slices are performed for a single test according to predetermined time units; The computing power consumption of the physical hardware in each time slice is calculated as the performance evaluation result.
9. An evaluation device comprising: An acquisition module, used to acquire the scoring result of the target algorithm and the physical hardware information for running the target algorithm; wherein the physical hardware is a chip, and the physical hardware information includes chip specifications, and the chip specifications are modeled according to the chip model and the chip computing power; A collection module, configured to use at least one preset algorithm load to iteratively test the target algorithm according to the scoring result and the physical hardware information; An analysis module is used to obtain a performance evaluation result of the target algorithm based on the test data of the test, wherein the performance evaluation result indicates a relationship between a feature of the target algorithm and physical hardware performance.
10. An electronic device, wherein: include: Memory, used to store programs; A processor is used to run the program stored in the memory, and the program executes the evaluation method according to any one of claims 1 to 8 when it is run.
11. A computer-readable storage medium having stored thereon a computer program executable by a processor, wherein: When the program is executed by a processor, the evaluation method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Benchmark test method and device of supervised learning algorithm in distributed environment
CN107203467A