Method for acquiring acceleration card performance, acceleration card performance test equipment and product
By deploying a universal access platform and training framework on the accelerator card, the problem of incompatibility of tools by different manufacturers is solved, the hardware universality of the accelerator card performance test and the accuracy of the test results are achieved, reducing the difficulty of selecting an accelerator card and improving efficiency.
Patent Information
- Application Number
- CN202510780863.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-12
AI Technical Summary
The performance analysis tools of different acceleration card manufacturers are incompatible, resulting in increased difficulty and reduced efficiency when users choose acceleration cards, mainly due to differences in the test environment, usage methods and result presentation.
The general access platform is adopted to deploy the training framework and task modules, and the task modules are obtained through the platform and the task modules are scheduled to be tested to accelerated card performance, isolate the underlying hardware dependence, and ensure that the test is carried out under the same training framework.
The hardware universality of accelerator card performance test and the accuracy of test results are achieved, reducing the difficulty of users to select accelerator cards, improving efficiency, and breaking down barriers between manufacturers' special tools.
Smart Images

Figure CN120295883A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of acceleration card performance testing, and particularly relates to a method for obtaining the performance of an acceleration card, an acceleration card performance testing device, and a product. Background Art
[0002] With the rapid development of artificial intelligence technology, the requirements for the computing performance of acceleration cards are getting higher and higher. In order to understand the performance of acceleration cards, it is necessary to analyze the performance of acceleration cards.
[0003] In the related art, the performance of acceleration cards is analyzed by relying on the performance analysis tools provided by acceleration card manufacturers. However, the performance analysis tools provided by acceleration card manufacturers are only for their own acceleration cards, that is, the performance analysis tools of different manufacturers are not compatible; and due to the fact that different performance analysis tools often have large differences in terms of test environment, usage method, analysis object, result presentation, etc., it increases the difficulty and reduces the efficiency for users to select acceleration cards.
[0004] It can be seen that while realizing the performance analysis of different acceleration cards, how to reduce the difficulty for users to select acceleration cards and improve the efficiency for users to select acceleration cards are technical problems that need to be solved urgently by those skilled in the art. Summary of the Invention
[0005] The present invention provides a method for obtaining the performance of an acceleration card, an acceleration card performance testing device, and a product, so as to at least solve the problems in the related art that the performance analysis tools of different manufacturers are not compatible; and different performance analysis tools often have large differences in terms of test environment, usage method, analysis object, result presentation, etc., resulting in an increase in the difficulty and a decrease in the efficiency for users to select acceleration cards.
[0006] The present invention provides a method for obtaining the performance of an acceleration card, including: Obtain an acceleration card to be tested and a general access platform for performing acceleration card performance testing; wherein, the general access platform is deployed with a training framework and modules corresponding to tasks; the modules corresponding to tasks at least include a runtime module and an operator module; When it is detected that the general access platform is deployed on the acceleration card to be tested, obtain a task to be executed through the training framework of the general access platform, and schedule the modules corresponding to the task according to the task to be executed; through the modules corresponding to the task, find the interface corresponding to the task to be executed from the task library of the task to be executed located in the acceleration card to be tested, so that the acceleration card to be tested executes the task to be executed according to the interface corresponding to the task to be executed; Obtain the performance of the acceleration card to be tested when executing the task to be executed through the general access platform.
[0007] The beneficial effects of the present invention are as follows. First, in the method for obtaining the performance of an acceleration card, when performing a performance test on the acceleration card to be tested, a general access platform is deployed on the acceleration card to be tested. The training framework of the general access platform is used to obtain the task to be executed, and the module corresponding to the task is scheduled according to the task to be executed. The module corresponding to the task searches for the interface corresponding to the task to be executed from the task library of the task to be executed located in the acceleration card to be tested, so that the acceleration card to be tested executes the task to be executed according to the interface corresponding to the task to be executed. Finally, the performance of the acceleration card to be tested when executing the task to be executed is obtained through the general access platform. The performance test of the acceleration card is realized by this method. Second, the module corresponding to the task is deployed on the general access platform, so that the task to be executed is indirectly deployed on the acceleration card to be tested through the module corresponding to the task, rather than directly deploying the task to be executed on the acceleration card to be tested, realizing the decoupling of the task to be executed and the acceleration card to be tested actually used. Therefore, the performance analysis of the module corresponding to the task of the general access platform does not naturally depend on the underlying acceleration card to be tested, thus realizing the hardware generality of performance analysis. Third, in the method where the performance analysis tool provided by the acceleration card manufacturer is only for the performance test of its own acceleration card, due to the different software R & D progress of each acceleration card manufacturer, it is often difficult to align the test environments provided by them in terms of the version of the training framework. In the case where the versions of the training frameworks cannot be aligned, it is very difficult to fairly evaluate the hardware performance of different acceleration cards when executing tasks. In the method provided by the present invention, a training framework is set in the general access platform. When using this general access platform to test multiple acceleration cards to be tested, the task to be executed is obtained from the same training framework, that is, the version of the training framework is ensured to be the same. The performance of different acceleration cards to be tested is tested in the same test environment, improving the accuracy and fairness of the performance test results. It can be seen that the method provided by the present invention realizes the performance test of different acceleration cards while breaking the usage barriers between the dedicated performance analysis tools of each acceleration card and isolating the development differences of different acceleration card manufacturers at the training framework level, greatly reducing the difficulty for users to select acceleration cards and improving the efficiency for users to select acceleration cards.
[0008] The present invention also provides an acceleration card performance test device. The acceleration card performance test device is connected to the acceleration card to be tested. The acceleration card performance test device deploys a training framework and a module corresponding to the task. The module corresponding to the task at least includes a runtime module and an operator module. The training framework of the acceleration card performance test device is used to obtain the task to be executed, and schedule the module corresponding to the task according to the task to be executed. The module corresponding to the task searches for the interface corresponding to the task to be executed from the task library of the task to be executed located in the acceleration card to be tested, so that the acceleration card to be tested executes the task to be executed according to the interface corresponding to the task to be executed. The performance of the acceleration card to be tested when executing the task to be executed is obtained.
[0009] The present invention also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above methods for obtaining the performance of an acceleration card when executing the computer program.
[0010] The present invention also provides a computer-readable storage medium storing a computer program, wherein the computer program implements the steps of any of the above methods for obtaining the performance of an acceleration card when executed by a processor.
[0011] The present invention also provides a computer program product including a computer program, which implements the steps of any of the above methods for obtaining the performance of an acceleration card when executed by a processor. Description of the Drawings
[0012] To more clearly illustrate the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0013] Figure 1 It is a flowchart of a method for obtaining the performance of an acceleration card provided by an embodiment of the present invention; Figure 2 It is a schematic diagram of deploying a pre-trained language model training task to an acceleration card provided by an embodiment of the present invention; Figure 3 It is a multi-level implementation architecture diagram of model training performance analysis provided by an embodiment of the present invention; Figure 4 It is a general-purpose implementation architecture diagram of an artificial intelligence computing power execution framework and training performance analysis provided by an embodiment of the present invention; Figure 5 It is a schematic diagram of a training time-consuming extraction tool extracting the time-consuming of a model training task provided by an embodiment of the present invention. Detailed Embodiments
[0014] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.
[0015] It should be noted that in the description of the present invention, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0016] In order to perform performance testing on the acceleration card, in the related art, performance analysis is carried out based on the performance analysis tools provided by the acceleration card manufacturers. It relies on the acceleration card manufacturers to provide performance analysis tools. The performance analysis tools of different acceleration card manufacturers are not compatible, and there are often significant differences in aspects such as the test environment, usage method, analysis object, result presentation, etc., which will greatly increase the research workload of users in selecting a model training acceleration card. Specifically: In terms of the test environment, each acceleration card manufacturer provides its own test image, and users need to be familiar with the setup methods of different environments, and the workload linearly increases with the types of acceleration cards to be investigated. Secondly, due to the different software R & D progress of each acceleration card manufacturer, it is often difficult to align the test environments provided by them in terms of the versions of training software frameworks such as PyTorch, Transformers, etc. Known software will also have an obvious impact on the training performance. Therefore, it is difficult to fairly evaluate the hardware performance of different acceleration cards in the model training task when the software versions cannot be aligned.
[0017] In terms of the analysis object, the performance analysis tools provided by the acceleration card manufacturers only target their own acceleration cards. In terms of the usage method, some manufacturers' performance analysis tools are configured through the command line, while some manufacturers' performance analysis tools are configured through the graphical interface. In terms of the result presentation, the results of different performance analysis tools also vary in form. These specialized designs and differences will become obstacles for model training users to select training hardware devices, increase their workload in measuring the performance of each acceleration card, and make it more difficult for users to select acceleration cards and reduce efficiency.
[0018] Therefore, the present invention provides a general performance analysis method, which is applicable to the performance testing of multiple acceleration cards.
[0019] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Figure 1 The flowchart of a method for obtaining the performance of an acceleration card provided by an embodiment of the present invention is as Figure 1 shown, and the method includes: S10: Obtain the acceleration card to be tested and a general access platform for performing acceleration card performance testing; wherein, the general access platform is deployed with a training framework and modules corresponding to the task; the modules corresponding to the task at least include a runtime module and an operator module; S11: When it is detected that the general access platform is deployed on the acceleration card to be tested, obtain the task to be executed through the training framework of the general access platform, and schedule the modules corresponding to the task according to the task to be executed; through the modules corresponding to the task, find the interface corresponding to the task to be executed from the task library of the task to be executed located in the acceleration card to be tested, so that the acceleration card to be tested executes the task to be executed according to the interface corresponding to the task to be executed; S12: Obtain the performance of the acceleration card to be tested when executing the task to be executed through the general access platform.
[0020] There is no limitation on the acceleration card to be tested, which is determined according to the actual situation. Considering that in the method of using the performance analysis tool provided by the acceleration card manufacturer only for the performance test of its own acceleration card, due to the different software R & D progress of each acceleration card manufacturer, it is often difficult to align the provided test environment in terms of the version of the training framework. And when the versions of the training frameworks cannot be aligned, it is very difficult to fairly evaluate the hardware performance of different acceleration cards when executing tasks. Therefore, in the method provided by the present invention, a training framework is set in the general access platform. The training framework includes PyTorch, Transformers, etc. It should be noted that different types of tasks to be executed use different training frameworks. The types of tasks to be executed are such as training task types and inference task types. The training frameworks used for training task types are different from those used for inference task types.
[0021] In order to enable the general access platform to test the performance of most acceleration cards, modules corresponding to the task are also deployed on the general access platform. The modules corresponding to the task at least include a runtime module and an operator module. Runtime and operators are software very close to the underlying hardware. In the general platform, different acceleration cards have different runtime and operator implementations.
[0022] During the process of using the general access platform to perform performance testing on the acceleration card to be tested, deploy the general access platform on the acceleration card to be tested. When it is detected that the general access platform is deployed on the acceleration card to be tested, obtain the task to be executed through the training framework of the general access platform. In order to achieve refined analysis of the performance of the acceleration card to be tested, the total time consumed for model training or the total time consumed for inference is gradually refined to the operator granularity. That is, decompose the task to be executed. Specifically, before the general access platform obtains the task to be executed and then, through the modules corresponding to the task, finds the interface corresponding to the task to be executed from the task library of the task to be executed, it further includes: Decompose the task to be executed into operator calculation tasks and runtime call tasks; Deploy the decomposed operator computing tasks to the operator module, and deploy the decomposed runtime call tasks to the runtime module.
[0023] Searching for the interface corresponding to the task to be executed from the task library corresponding to the task to be executed through the module corresponding to the task includes: Search for the interface corresponding to the operator computing task from the operator library corresponding to the operator computing task through the operator module; Search for the interface corresponding to the runtime call task from the runtime library corresponding to the runtime call task through the runtime module.
[0024] In order to accurately divide the tasks to be executed, in implementation, decomposing the tasks to be executed into operator computing tasks and runtime call tasks includes: Obtain the type of the task to be executed; wherein, the type of the task to be executed includes at least a training task type and an inference task type; Determine the processing flow of the task to be executed according to the type of the task to be executed; According to the processing flow of the task to be executed and in accordance with the preset task classification rules, decompose the task to be executed step by step to obtain the operator computing task and the runtime call task; wherein, the task to be executed is the highest-level task among all levels of tasks; the operator computing task and the runtime call task are the lowest-level tasks among all levels of tasks.
[0025] If the type of the task to be executed is the training task type, the processing flow of the task to be executed sequentially includes forward calculation, backward calculation, and gradient update. From the perspective of the training process, the model training task can be divided into three parts: forward calculation, backward calculation, and gradient update. Forward calculation and backward calculation depend on the model structure, and gradient update is related to the optimizer. The bottom layer is the runtime and the operator. Figure 2 This is a schematic diagram of deploying the pre-training language model training task to the acceleration card provided by the embodiment of the present invention. Figure 2 In it, the pre-training language model training task is divided into forward calculation, backward calculation, gradient update, model structure, optimizer, operator, and runtime. The training tasks are respectively deployed on acceleration card 1, acceleration card 2,..., acceleration card N.
[0026] There is no limitation on the pre-set task grading rules, which are determined according to the actual situation. For example, the common pre-trained language model network structure mainly consists of an embedding block, multiple repeated decoding blocks, and a normalization block. The number of repetitions of the decoding block is generally dozens of times. For example, the decoding layer of the Llama2 model with 7 billion parameters has 32 layers. The decoding block can be composed of a normalization layer, a multi-head self-attention layer, and a fully connected layer, and each layer is composed of many various operator calculations, such as matrix multiplication, addition, etc. The training tasks (i.e., the tasks to be executed) are divided into the first level, the forward calculation, backward calculation, and gradient update are divided into the second level; the blocks in the network structure are divided into the third level; the layers under the blocks are divided into the fourth level, and the operator definitions are divided into the fifth level. Figure 3 This is a multi-level implementation architecture diagram for model training performance analysis provided by an embodiment of the present invention. As Figure 3 shown, the first level is the pre-trained language model training task, and the second level includes forward calculation, backward calculation, and gradient update. The third level includes: an embedding block, a decoding block, and a normalization block; taking the decoding block as an example, the fourth level includes a normalization layer, a multi-head self-attention layer, a normalization layer, and a fully connected layer; taking the multi-head self-attention layer as an example, the fifth level includes matrix multiplication, addition, and subtraction. If the task to be executed is an inference task, then the first level is the pre-trained language model inference task, and the second level includes forward calculation. The clearly defined process structure can be used for the multi-level implementation of training performance analysis: gradually refining from the total training time of the model to the operator granularity.
[0027] After refining the tasks to be executed, in order to implement the execution of each refined task. Accordingly, corresponding functions for representing the execution of each refined task are set in the training function.
[0028] Specifically, the training framework includes a function for representing the execution of the task to be executed, a function for representing the execution of subtasks, a function for representing the execution of operator calculation tasks, and a function for representing the execution of runtime call tasks; among them, the subtasks are all the tasks remaining after the tasks to be executed are gradually decomposed, except for the operator calculation tasks and runtime call tasks.
[0029] Among them, taking the task to be executed as the training task as an example, the function for representing the execution of the task to be executed is the function for representing the execution of the training task; in the multi-level implementation architecture for model training performance analysis in the present invention embodiment, Figure 3 the tasks in the second level, the tasks in the third level, and the tasks in the fourth level in the architecture are all called subtasks; the function for representing the execution of subtasks is the function for representing the execution of the tasks in the second level, the function for representing the execution of the tasks in the third level, and the function for representing the execution of the tasks in the fourth level.
[0030] The tasks to be performed through the training framework of the universal access platform include: Receiving information sent by a user representing a function for calling a task to be executed through a training framework of a universal access platform; Decomposing the tasks to be executed into operator computing tasks and runtime calling tasks includes: The control function used to characterize the calling of the task to be executed calls the functions corresponding to the tasks at each level in sequence according to the hierarchical order of the tasks to be executed, until the function used to characterize the execution of the operator calculation task and the function used to characterize the execution of the runtime calling task are called, so as to decompose the task to be executed into the operator calculation task and the runtime calling task.
[0031] Figure 4 This is a general implementation architecture diagram of an artificial intelligence computing power execution framework and training performance analysis provided by an embodiment of the present invention. Figure 4 As shown in FIG. 1 , the model training task is deployed on the general access platform, and the general training performance analysis module is deployed on the general access platform. The general training performance analysis module includes a training framework, a runtime module, and a deep learning operator module. The general access platform is deployed on accelerator cards 1 to N. The accelerator cards are respectively deployed with a computing operator library and a runtime library.
[0032] The general training performance analysis module acts on the training framework, the runtime module of the universal access platform of the accelerator, and the deep learning operator module to obtain the model training time. The universal access platform of the accelerator defines a set of virtual runtime interfaces and operator interfaces, which are decoupled from the actual accelerator card. Therefore, the performance analysis of the runtime module and operator module of the universal access platform of the accelerator naturally does not depend on the underlying accelerator card, thus realizing the hardware universality of the training performance analysis. Based on the general training performance analysis module, Figure 2 The training performance of accelerator card 1, accelerator card 2 and more accelerator cards connected to the accelerator card universal access platform are evaluated. The advantages and disadvantages of the performance of each accelerator card are clearly visible, which solves the pain point that different performance analysis tools need to be used for different accelerator cards in related technologies.
[0033] After breaking down the tasks to be executed into operator computing tasks and runtime calling tasks, in order to achieve refined performance analysis, during implementation, the general access platform obtains the performance of the accelerator card to be tested when executing the tasks to be executed, including: Obtaining tasks to be timed in the process of executing tasks to be executed; wherein the tasks to be timed include one or more of tasks to be executed, subtasks, operator calculation tasks, and runtime call tasks; Get the time taken by the task to be timed; The performance of the accelerator card to be tested when executing the task to be executed is determined according to the time consumption of the task to be timed.
[0034] Taking the forward calculation as the task to be timed, obtain the time consumption of the forward calculation task, which is used as the performance of the acceleration card to be tested when executing the forward calculation task.
[0035] In order to time the task to be timed, in implementation, obtaining the time consumption of the task to be timed includes: Receive the function inserted before the function representing the execution of the task to be timed, which is used to represent the start of timing; After detecting that the task to be timed has completed execution, send a prompt message to the user to represent the completion of the execution of the task to be timed, so that the user can insert a function representing the end of timing after the function representing the execution of the task to be timed according to the prompt message; Determine the first time point according to the function representing the start of timing, and determine the second time point according to the function representing the end of timing; Determine the time consumption of the acceleration card to be tested when executing the task to be timed according to the difference between the second time point and the first time point.
[0036] Since different users have different test requirements, the operations of inserting the function representing the start of timing (i.e., the start timing function) and inserting the function representing the end of timing (i.e., the end timing function) can both be inserted by the user. For example, in order to time the forward calculation task, the start timing function can be inserted before the function representing the execution of the forward calculation task, and the end timing function can be inserted after the function representing the execution of the forward calculation task.
[0037] In order to automatically calculate the time consumption, in addition to the start timing function and the end timing function, a function representing the result printing (i.e., the result printing function) is also set.
[0038] Obtaining the difference between the second time point and the first time point includes: Receive the function representing the result printing inserted after the function representing the end of timing; Output the difference between the second time point and the first time point through the function representing the result printing.
[0039] In this method, through the function representing the result printing, the corresponding time consumption is output for each executed timing task.
[0040] After calculating the time consumption of each level of tasks, in order to intuitively understand the time consumption of each level of tasks, after determining the time consumption of the acceleration card to be tested when executing the task to be timed according to the difference between the second time point and the first time point, it also includes: Display the time consumption of the task to be timed in sequence according to the order of the task levels.
[0041] After the task to be executed is completed, the time taken for tasks at all levels is output. After detecting that the user inserts a function for characterizing the end of timing after the function for characterizing the execution of the task to be executed according to the prompt information, it further includes: Sending information to the user for characterizing inserting a function for characterizing result printing after the function for the task to be executed; After detecting that there is a function for characterizing result printing, the time taken for the tasks to be timed is sequentially displayed in the order of the task levels.
[0042] The lowest layer of the implemented general multi-level training performance analysis method depends on the extraction of the time-consuming data of each part of the training. The time-consuming extraction tool is implemented by the timing class. Taking the extraction of the time taken for the model training task as an example, Figure 5 It is a schematic diagram of extracting the time taken for the model training task by a training time-consuming extraction tool provided by an embodiment of the present invention. As Figure 5 shown, the time-consuming extraction tool includes a start timing function, an end timing function, and a result printing function. The start timing function, the end timing function, and the result printing function are collectively referred to as the timing class. Among them, the start timing function and the end timing function are implemented by the time function (time()) of the Python time library. The start timing function is inserted before the execution of the model training task, the end timing function is inserted after the completion of the execution of the model training task, and the result printing function is inserted after the end timing function. After the model training task ends, the result printing function sequentially displays the results according to the task levels.
[0043] After obtaining the time taken for tasks at all levels, the performance of the acceleration cards can be horizontally compared according to the time taken and the performance bottleneck can be longitudinally investigated, as Figure 2 shown. First, the process of horizontal comparison will be described below. There are multiple acceleration cards to be tested. After sequentially displaying the time taken for the tasks to be timed in the order of the task levels, it further includes: Obtaining the time taken for the tasks to be timed by the acceleration card to be tested at the first level; If it is detected that there is a target acceleration card whose time taken for the task to be timed is greater than the time-consuming threshold, then obtain the time-consuming difference between the time taken for the tasks to be timed at other levels of the target acceleration card and the time taken for the tasks to be timed at the same level by other acceleration cards; where, the other acceleration cards are the acceleration cards except the target acceleration card among all the acceleration cards to be tested; Adjust the target acceleration card according to the time-consuming difference, use the target acceleration card as the new acceleration card to be tested, and return to the steps of obtaining the acceleration card to be tested and the general access platform for accelerating card performance testing.
[0044] It should be noted that the training performance analysis in the related technology acts on the run-time and operator interfaces of each acceleration card, as Figure 2As shown, there will be slight differences in the training time consumed in the general training performance analysis and statistics of the present invention, that is, the time consumed introduced by the general access platform for the acceleration card. This part of the time is mainly the time from the virtual interface of the general platform to the actual interface call, which belongs to the time consumed by the Central Processing Unit (CPU). Compared with the training time of the pre-trained language model, it can be almost ignored. In addition, this part of the time is the same for each acceleration card connected to the general platform, so it does not affect the horizontal comparison of the training performance of each acceleration card in the artificial intelligence computing power execution framework.
[0045] Through horizontal comparison, the acceleration cards with poor performance can be identified, and their optimization can be carried out with reference to the performance test results of other acceleration cards, so as to ensure that the performance of all acceleration cards is improved.
[0046] In implementation, if the time consumed is calculated for all operators and runtimes, it will lead to an increase in the amount of calculation. Therefore, in the embodiments of the present invention, in order to reduce the amount of calculation, it is determined that the tasks to be timed include: Regarding each subtask at the current level as a task to be timed respectively; wherein, the current level starts from the second level; After obtaining the time consumed by each subtask at the current level, the subtask with the most time consumed at the current level is obtained according to the time consumed by each subtask at the current level; Regarding the subtasks of the subtask with the most time consumed at the current level at the next level as each subtask at the new current level; returning to the step of regarding each subtask at the current level as a task to be timed respectively.
[0047] The horizontal comparison is described above. In this embodiment, a method for vertical comparison is provided. After sequentially displaying the time consumed by the tasks to be timed in the order of the task levels, it further includes: Comparing the time consumed by the tasks to be timed at the same level to determine the task to be timed with the most time consumed at the same level, and determining the target operator calculation task and / or the target runtime call task with the most time consumed when executing the task to be timed; Sending information to the user for characterizing the update of the operator library and / or the runtime library corresponding to the target operator calculation task in the acceleration card to be tested; Regarding the updated acceleration card to be tested as a new acceleration card to be tested, and returning to the step of obtaining the acceleration card to be tested and the general access platform for performing the performance test of the acceleration card.
[0048] First find the one with the most time consumed in the second level. If it is found that the forward time consumption is the most, then only the time consumption of the levels under the forward time consumption is counted subsequently. For example, continue to obtain the one with the most time consumed among the subtasks of the third level under the forward time consumption. Figure 4For example, compare the time consumption of the forward calculation, the reverse calculation, and the gradient update at the second level. If it is found that the forward calculation takes the most time, there is no need to calculate the time consumption of the subsequent hierarchical tasks of the reverse calculation and the time consumption of the gradient update of the subsequent hierarchical tasks. Only the time consumption of the subsequent hierarchical tasks of the forward calculation needs to be continuously analyzed. Continue to compare the time consumption of the embedding block, the decoding block, and the normalization block at the subsequent third level of the forward calculation. If it is found that the decoding block takes the most time, there is no need to analyze the time consumption of the subsequent hierarchical levels of the embedding block and the normalization block; continue to compare the time consumption of the normalization layer, the multi-head self-attention layer, the normalization layer, and the fully connected layer at the fourth level subsequent to the decoding block. If it is found that the multi-head self-attention layer takes the most time, there is no need to analyze the time consumption of the subsequent hierarchical levels of the normalization layer and the fully connected layer. Only the time consumption of the subsequent hierarchical levels of the multi-head self-attention layer needs to be analyzed, such as analyzing the time consumption of matrix multiplication, addition, and subtraction at the fifth level. For example, for Figure 2 in the acceleration card 1, perform the fourth-level performance analysis. If it is found that the self-attention layer takes a relatively large proportion of the time, the self-attention implementation algorithm can be considered for optimization. Perform the fifth-level performance analysis. If it is found that the matrix multiplication takes a relatively large proportion of the time, the matrix multiplication operator implementation can be considered for optimization.
[0049] In this method, the multi-level performance analysis helps users vertically troubleshoot the training performance bottleneck of a certain acceleration card in the artificial intelligence computing power execution framework and find the direction of performance optimization.
[0050] The general, multi-level performance analysis method for different acceleration card model training / inference tasks described above. First, deploy the large model training task on each acceleration card through the artificial intelligence computing power execution framework, and horizontally compare the training performance of each acceleration card to achieve the universality of performance analysis based on the general access platform for acceleration cards of the artificial intelligence computing power execution framework. Secondly, decompose the large model training task step by step according to the model training process, refine from the complete training granularity to the operator granularity, and realize multi-level performance analysis to vertically troubleshoot the training performance bottleneck.
[0051] In summary, this method enables users to horizontally evaluate the performance of different acceleration cards in various model training tasks with only one set of training environments, breaking the usage barriers between dedicated performance analysis tools for different acceleration cards in related technologies and isolating the development differences of different acceleration card manufacturers at the training framework level. The artificial intelligence computing power execution framework aims to facilitate the adaptation of more acceleration cards to large model training through a set of general platforms and standards, and has currently completed docking with multiple acceleration cards. The general access platform for acceleration cards based on the artificial intelligence computing power execution framework provided by the present invention realizes the hardware generality of training / inference performance analysis, and realizes the multi-level extraction of training / inference time-consuming based on the clearly structured process at the large model training / inference level, greatly facilitating users to horizontally compare the training / inference performance advantages and disadvantages of different acceleration cards, select the acceleration card most suitable for their own business, further vertically explore the training performance bottleneck of the acceleration card, and find the direction of performance optimization.
[0052] A method for obtaining the performance of an acceleration card is described above. This embodiment also provides an acceleration card performance testing device. The acceleration card performance testing device is connected to the acceleration card to be tested; a training framework and a module corresponding to a task are deployed on the acceleration card performance testing device; the module corresponding to the task at least includes a runtime module and an operator module; The training framework of the acceleration card performance testing device is used to obtain a task to be executed, schedule the module corresponding to the task according to the task to be executed; find an interface corresponding to the task to be executed from the task library of the task to be executed located in the acceleration card to be tested through the module corresponding to the task, so that the acceleration card to be tested executes the task to be executed according to the interface corresponding to the task to be executed; obtain the performance of the acceleration card to be tested when executing the task to be executed.
[0053] In practice, there will be a need to test the performance of multiple acceleration cards. To meet this need, the acceleration card performance testing device includes a main control unit and multiple independent test channels, and each channel is connected to the acceleration card test interface through an isolation circuit. Before testing, the general access platform is burned into the storage area of each acceleration card. After the acceleration card performance testing device is started, the main control unit issues a test instruction set to each channel through a time-division multiplexing bus; after receiving the test instruction, the general access platform deployed on the acceleration card performs a performance test on the acceleration card; the test task is identified based on a unique device identifier (Identifier, ID), and the test data is transmitted back to the data cache array of the device through an asynchronous interrupt mechanism, and the main control unit sorts and verifies the data according to the timestamp and ID.
[0054] In this method, before testing, the general access platform is burned into the storage area of each acceleration card, the main control unit issues a test instruction set to each channel through a time-division multiplexing bus, and uses the instruction cycle difference to realize the parallel loading of multi-acceleration card instructions. The multi-acceleration cards use the general access platform deployed on themselves to perform a performance test on the acceleration cards, and finally realize the parallel test of multiple acceleration cards by the same device.
[0055] The acceleration card performance testing device provided in this embodiment has the same or corresponding technical features as the method for obtaining the acceleration card performance described above, and the effects are the same.
[0056] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0057] An embodiment of the present invention also provides a device for obtaining the acceleration card performance, including: A first acquisition module, configured to acquire an acceleration card to be tested and a general access platform for performing acceleration card performance testing; wherein, a training framework and a module corresponding to a task are deployed on the general access platform; the module corresponding to the task at least includes a runtime module and an operator module; A second acquisition module, configured to, when detecting that the general access platform is deployed on the acceleration card to be tested, acquire a task to be executed through the training framework of the general access platform, and schedule the module corresponding to the task according to the task to be executed; through the module corresponding to the task, find an interface corresponding to the task to be executed from the task library of the task to be executed located in the acceleration card to be tested, so that the acceleration card to be tested executes the task to be executed according to the interface corresponding to the task to be executed; A third acquisition module, configured to acquire the performance of the acceleration card to be tested when executing the task to be executed through the general access platform.
[0058] In some embodiments, the device for obtaining the acceleration card performance further includes: A decomposition module, configured to decompose the task to be executed into an operator calculation task and a runtime call task; A deployment module, configured to deploy the decomposed operator calculation task to the operator module, and deploy the decomposed runtime call task to the runtime module; The second acquisition module includes a search module, configured to find an interface corresponding to the task to be executed from the task library corresponding to the task to be executed through the module corresponding to the task.
[0059] The search module specifically includes: A first search sub-module, configured to find an interface corresponding to the operator calculation task from the operator library corresponding to the operator calculation task through the operator module; A second search sub-module, configured to find an interface corresponding to the runtime call task from the runtime library corresponding to the runtime call task through the runtime module.
[0060] In some embodiments, the decomposition module includes: A fourth acquisition module, configured to acquire the type of the task to be executed; wherein, the type of the task to be executed includes at least a training task type and an inference task type; A first determination module, configured to determine the processing flow of the task to be executed according to the type of the task to be executed; A hierarchical decomposition module, configured to hierarchically decompose the task to be executed according to the processing flow of the task to be executed and in accordance with a preset task hierarchical rule, so as to obtain an operator calculation task and a runtime call task; wherein, the task to be executed is the highest-level task among all levels of tasks; the operator calculation task and the runtime call task are the lowest-level tasks among all levels of tasks.
[0061] In some embodiments, the second acquisition module includes a first sub-acquisition module, and the first sub-acquisition module is configured to acquire the task to be executed through the training framework of the general access platform.
[0062] The first sub-acquisition module includes: A first reception module, configured to receive, through the training framework of the general access platform, information sent by the user representing a function for invoking the task to be executed; The decomposition module is specifically configured to control the function representing the invocation of the task to be executed to sequentially call the functions corresponding to the tasks at each level in accordance with the hierarchical order of the task to be executed, until the function representing the execution of the operator calculation task and the function representing the execution of the runtime call task are called, so as to decompose the task to be executed into an operator calculation task and a runtime call task.
[0063] In some embodiments, the third acquisition module includes: A fifth acquisition module, configured to acquire the task to be timed during the execution of the task to be executed; wherein, the task to be timed includes one or more of the task to be executed, a subtask, an operator calculation task, and a runtime call task; A sixth acquisition module, configured to acquire the elapsed time of the task to be timed; A seventh acquisition module, configured to determine the performance of the to-be-tested acceleration card when executing the task to be executed according to the elapsed time of the task to be timed.
[0064] In some embodiments, the sixth acquisition module includes: A second reception module, configured to receive a function representing the start of timing inserted before the function representing the execution of the task to be timed; A first transmission module, configured to send a prompt message representing the completion of the execution of the task to be timed to the user after detecting that the task to be timed has completed execution, so that the user can insert a function representing the end of timing after the function representing the execution of the task to be timed according to the prompt message. A second determination module, configured to determine a first time point according to a function for characterizing the start of timing, and determine a second time point according to a function for characterizing the end of timing; A third determination module, configured to determine the time taken for the acceleration card under test to execute the task to be timed according to the difference between the second time point and the first time point; The apparatus for obtaining the performance of the acceleration card further includes: A first display module, configured to sequentially display the time taken for the task to be timed in the order of the task hierarchy.
[0065] In some embodiments, the apparatus for obtaining the performance of the acceleration card further includes: an eighth acquisition module, configured to acquire the difference between the difference of the second time point and the first time point.
[0066] The eighth acquisition module includes: A third receiving module, configured to receive a function for characterizing result printing inserted after the function for characterizing the end of timing; An output module, configured to output the difference between the second time point and the first time point through the function for characterizing result printing.
[0067] In some embodiments, the apparatus for obtaining the performance of the acceleration card further includes: A second sending module, configured to send information to the user for characterizing inserting a function for characterizing result printing after the function of the task to be executed; A second display module, configured to sequentially display the time taken for the task to be timed in the order of the task hierarchy after detecting the existence of a function for characterizing result printing.
[0068] In some embodiments, the apparatus for obtaining the performance of the acceleration card further includes: A ninth acquisition module, configured to acquire the time taken for the tasks to be timed on the first level of each acceleration card under test; A tenth acquisition module, configured to, if a target acceleration card whose time taken for the task to be timed is greater than the time threshold is detected, acquire the time difference between the time taken for the tasks to be timed on other levels of the target acceleration card and the time taken for the tasks to be timed on the same level of other acceleration cards; wherein, the other acceleration cards are acceleration cards other than the target acceleration card among all the acceleration cards under test; An adjustment module, configured to adjust the target acceleration card according to the time difference, use the target acceleration card as a new acceleration card under test, and return to trigger the first acquisition module.
[0069] In some embodiments, the apparatus for obtaining the performance of the acceleration card further includes: a fourth determination module, configured to determine the task to be timed.
[0070] The fourth determination module includes: A first acting module, configured to use each subtask at the current level as a task to be timed; wherein, the current level starts from the second level; An eleventh obtaining module, configured to, after obtaining the time consumption of each subtask at the current level, obtain the subtask with the most time consumption at the current level according to the time consumption of each subtask at the current level; A second acting module, configured to use the subtasks of the subtask with the most time consumption at the current level at the next level as each subtask at the new current level; and return to trigger the first acting module.
[0071] In some embodiments, the device for obtaining the performance of the acceleration card further includes: A comparator determination module, configured to compare the time consumption of the tasks to be timed at the same level to determine the task to be timed with the most time consumption at the same level, and determine the target operator calculation task and / or the target runtime call task with the most time consumption when executing the task to be timed; An update module, configured to send information to the user for characterizing the update of the operator library and / or the target runtime library corresponding to the target operator calculation task in the acceleration card to be tested; A third acting module, configured to use the updated acceleration card to be tested as the new acceleration card to be tested, and return to trigger the first obtaining module.
[0072] For the description of the features in the embodiments corresponding to the device for obtaining the performance of the acceleration card, reference can be made to the relevant description of the embodiments corresponding to the method for obtaining the performance of the acceleration card, which will not be elaborated here one by one.
[0073] An embodiment of the present invention further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above method embodiments for obtaining the performance of the acceleration card.
[0074] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above method embodiments for obtaining the performance of the acceleration card when running.
[0075] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: a USB flash drive, a read-only memory (ROM for short), a random access memory (RAM for short), a mobile hard disk, a magnetic disk, or an optical disc and other various media that can store computer programs.
[0076] An embodiment of the present invention further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any one of the above method embodiments for obtaining the performance of an acceleration card.
[0077] An embodiment of the present invention further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any one of the above method embodiments for obtaining the performance of an acceleration card.
[0078] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.
[0079] The above has introduced in detail a method for obtaining the performance of an acceleration card, an acceleration card performance test device, and a product provided by the present invention. Specific examples are used herein to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.
Claims
1. A method for obtaining the performance of an acceleration card, characterized in that, include: Obtain an accelerator card to be tested and a universal access platform for performing accelerator card performance testing; wherein the universal access platform is deployed with a training framework and modules corresponding to the task; the modules corresponding to the task include at least a runtime module and an operator module; In the case where it is detected that the universal access platform is deployed on the accelerator card to be tested, the task to be executed is obtained through the training framework of the universal access platform, and the module corresponding to the task is scheduled according to the task to be executed; the interface corresponding to the task to be executed is searched from the task library of the task to be executed located in the accelerator card to be tested through the module corresponding to the task, so that the accelerator card to be tested executes the task to be executed according to the interface corresponding to the task to be executed; The performance of the acceleration card to be tested when executing the task to be executed is obtained through the universal access platform.
2. The method for obtaining the performance of an acceleration card according to claim 1, wherein After acquiring the task to be executed, the universal access platform searches for the interface corresponding to the task to be executed from the task library corresponding to the task to be executed through the module corresponding to the task, and also includes: Decomposing the tasks to be executed into operator calculation tasks and runtime call tasks; Deploy the decomposed operator computing tasks in the operator module, and deploy the decomposed runtime calling tasks in the runtime module; The searching, through the task corresponding module, for the interface corresponding to the task to be executed from the task library corresponding to the task to be executed comprises: Searching, by the operator module, for an interface corresponding to the operator computing task from an operator library corresponding to the operator computing task; The runtime module searches for an interface corresponding to the runtime calling task from a runtime library corresponding to the runtime calling task.
3. The method for obtaining the performance of an acceleration card according to claim 2, wherein Decomposing the task to be executed into operator computing tasks and runtime calling tasks includes: Obtaining the type of the task to be performed; wherein the type of the task to be performed includes at least a training task type and a reasoning task type; Determining a processing flow of the task to be executed according to the type of the task to be executed; According to the processing flow of the tasks to be executed and the pre-set task classification rules, the tasks to be executed are decomposed level by level to obtain operator calculation tasks and runtime calling tasks; wherein, the tasks to be executed are the highest-level tasks among all levels of tasks; operator calculation tasks and runtime calling tasks are the lowest-level tasks among all levels of tasks.
4. The method for obtaining the performance of an acceleration card according to claim 3, wherein The types of the tasks to be executed are different, and the training frameworks are different; the training framework includes a function for characterizing the execution of the tasks to be executed, a function for characterizing the execution of subtasks, a function for characterizing the execution of operator calculation tasks, and a function for characterizing the execution of runtime call tasks; wherein the subtasks are the remaining tasks after decomposing the tasks to be executed level by level, except for the operator calculation tasks and runtime call tasks; The acquiring of tasks to be performed through the training framework of the universal access platform includes: Receiving, through the training framework of the universal access platform, information sent by a user representing a function for calling the task to be executed; Decomposing the to-be-executed task into operator calculation tasks and runtime call tasks includes: Controlling the function for characterizing the to-be-executed task to sequentially call the functions corresponding to the tasks at each level according to the hierarchical order of the to-be-executed task until the functions for characterizing the execution of the operator calculation task and the function for characterizing the execution of the runtime call task are called, so as to decompose the to-be-executed task into operator calculation tasks and runtime call tasks.
5. The method for obtaining the performance of an acceleration card according to claim 4, wherein The general access platform obtaining the performance of the to-be-tested acceleration card when executing the to-be-executed task includes: Obtaining the to-be-timed task during the execution of the to-be-executed task; wherein, the to-be-timed task includes one or more of the to-be-executed task, subtasks, operator calculation tasks, and runtime call tasks; Obtaining the time consumption of the to-be-timed task; Determining the performance of the to-be-tested acceleration card when executing the to-be-executed task according to the time consumption of the to-be-timed task.
6. The method for obtaining the performance of an acceleration card according to claim 5, wherein The obtaining the time consumption of the to-be-timed task includes: Receiving the function for characterizing the start of timing inserted before the function for characterizing the execution of the to-be-timed task; After detecting that the to-be-timed task has completed execution, sending a prompt message for characterizing the completion of execution of the to-be-timed task to the user, so that the user can insert a function for characterizing the end of timing after the function for characterizing the execution of the to-be-timed task according to the prompt message; Determining a first time point according to the function for characterizing the start of timing, and determining a second time point according to the function for characterizing the end of timing; Determining the time consumption of the to-be-tested acceleration card for executing the to-be-timed task according to the difference between the second time point and the first time point; After determining the time consumption of the to-be-tested acceleration card for executing the to-be-timed task according to the difference between the second time point and the first time point, it further includes: Sequentially displaying the time consumption of the to-be-timed task according to the order of the task levels.
7. The method for obtaining the performance of an acceleration card according to claim 6, wherein Obtaining the difference between the second time point and the first time point includes: Receiving the function for characterizing the result printing inserted after the function for characterizing the end of timing; Outputting the difference between the second time point and the first time point through the function for characterizing the result printing.
8. The method for obtaining the performance of an acceleration card according to claim 6, wherein When the to-be-timed task is the to-be-executed task, after detecting that the user inserts a function for characterizing the end of timing after the function for characterizing the execution of the to-be-executed task according to the prompt message, it further includes: Sending a message to the user for characterizing inserting a function for characterizing the result printing after the function for the to-be-executed task; After detecting the existence of the function for characterizing the result printing, sequentially displaying the time consumption of the to-be-timed task according to the order of the task levels.
9. The method for obtaining the performance of an acceleration card according to claim 8, wherein When there are multiple to-be-tested acceleration cards, after sequentially displaying the time consumption of the to-be-timed task according to the order of the task levels, it further includes: Obtaining the time consumption of each to-be-tested acceleration card for the to-be-timed task at the first level; If a target accelerator card is detected where the time consumption of the to-be-timed task is greater than the time consumption threshold, obtain the time consumption difference between the time consumption of the to-be-timed task on other levels of the target accelerator card and the time consumption of the to-be-timed task on the same level of other accelerator cards; wherein, the other accelerator cards are accelerator cards other than the target accelerator card among all the to-be-tested accelerator cards. Adjust the target accelerator card according to the time consumption difference, use the target accelerator card as a new to-be-tested accelerator card, and return to the step of obtaining the to-be-tested accelerator card and the general access platform for performing accelerator card performance testing.
10. The method for obtaining the performance of an acceleration card according to claim 6 or 7, characterized in that Determine that the to-be-timed task includes: Take each sub-task on the current level as the to-be-timed task respectively; wherein, the current level starts from the second level. After obtaining the time consumption of each sub-task on the current level, obtain the sub-task with the most time consumption on the current level according to the time consumption of each sub-task on the current level. Take the sub-tasks of the next level of the sub-task with the most time consumption on the current level as each sub-task on the new current level; return to the step of taking each sub-task on the current level as the to-be-timed task respectively.
11. The method for obtaining the performance of an acceleration card according to claim 10, wherein After sequentially displaying the time consumption of the to-be-timed tasks in the order of task levels, further include: Compare the time consumption of the to-be-timed tasks on the same level to determine the to-be-timed task with the most time consumption on the same level, and determine the target operator calculation task and / or target runtime call task with the most time consumption when executing the to-be-timed task. Send information to the user for characterizing the update of the operator library and / or target runtime library corresponding to the target operator calculation task in the to-be-tested accelerator card. Use the updated to-be-tested accelerator card as a new to-be-tested accelerator card, and return to the step of obtaining the to-be-tested accelerator card and the general access platform for performing accelerator card performance testing.
12. An acceleration card performance testing device, characterized in that, The accelerator card performance testing device is connected to the to-be-tested accelerator card; the accelerator card performance testing device is deployed with a training framework and a module corresponding to the task; the module corresponding to the task at least includes a runtime module and an operator module. The training framework of the accelerator card performance testing device is used to obtain the to-be-executed task, schedule the module corresponding to the task according to the to-be-executed task; find the interface corresponding to the to-be-executed task from the task library of the to-be-executed task located in the to-be-tested accelerator card through the module corresponding to the task, so that the to-be-tested accelerator card executes the to-be-executed task according to the interface corresponding to the to-be-executed task; obtain the performance of the to-be-tested accelerator card when executing the to-be-executed task.
13. An electronic device, characterized in that, Include: A memory for storing a computer program. A processor for implementing the steps of the method for obtaining accelerator card performance according to any one of claims 1 to 11 when executing the computer program.
14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the method for obtaining accelerator card performance according to any one of claims 1 to 11 when executed by a processor.
15. A computer program product, comprising a computer program, characterized in that, The computer program implements the steps of the method for obtaining accelerator card performance according to any one of claims 1 to 11 when executed by a processor.
Citation Information
Patent Citations
Performance test method and system for artificial intelligence accelerator card of domestic heterogeneous platform
CN117370088A
Model training method, product, equipment and computer readable storage medium
CN118395194A
Method, device and equipment for testing training performance of artificial intelligence acceleration card
CN118796632A
Accelerator card test method, device and equipment and storage medium
CN118916255A
Automatic model parallel scheduling strategy generation method and device based on heterogeneous computing power
CN118939391A
Cited By
Hierarchical modeling method based on large language model
CN121301158A