An AI processor performance evaluation method, device, equipment and medium
By splitting and prioritizing simulation tasks, and utilizing simulation execution clusters and performance analysis components, the performance of AI processors is automatically evaluated, solving the problems of high cost and low efficiency in existing technologies, and achieving fast and accurate performance evaluation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI SUIYUAN TECH CO LTD
- Filing Date
- 2026-06-18
- Publication Date
- 2026-07-21
Smart Images

Figure CN122432006A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of electronic digital data processing technology, and in particular to a method, apparatus, device, and medium for evaluating the performance of an AI processor. Background Technology
[0002] More and more enterprises are starting to use large AI models running on Artificial Intelligence (AI) servers for image processing, text processing, and speech processing. AI servers typically contain multiple AI processors for data transmission and computation in these processes. Before deploying an AI processor, a performance evaluation is required to determine if any performance issues exist.
[0003] In related technologies, a common performance evaluation scheme for AI processors involves technicians assessing various aspects of the AI processor's performance based on their experience, determining performance evaluation information to describe whether the AI processor has performance problems. With the significant increase in the complexity of AI processors, the complexity of AI processor performance evaluation work, both qualitatively and quantitatively, is also facing an exponential increase. Current AI processor performance evaluation schemes rely on manual operation, resulting in high time and labor costs, low efficiency, and difficulty in guaranteeing accuracy. Summary of the Invention
[0004] This invention provides a method, apparatus, device, and medium for evaluating the performance of AI processors, in order to solve the problems of high time and labor costs, low efficiency, and difficulty in guaranteeing accuracy in related technologies for evaluating the performance of AI processors.
[0005] According to one aspect of the present invention, a method for evaluating the performance of an AI processor is provided, comprising: Based on the performance evaluation instructions of the AI processor under test, the simulation task of the AI processor under test is determined, and the simulation task is broken down into multiple simulation sub-tasks. The simulation subtasks are arranged according to their multidimensional priority evaluation values to obtain a priority task queue. For each simulation subtask in the priority task queue, the machine nodes and resource nodes in the simulation execution cluster that match the simulation subtask are determined. Each simulation subtask in the priority task queue is submitted to the simulation execution cluster for execution as a target thread, and the task status changes of each simulation subtask are monitored. The detected completed subtasks, suspended subtasks, and timed-out subtasks are responded to. Based on the monitored effective sub-tasks in each completed sub-task, the performance data of the AI processor under test is determined and stored. The performance analysis component determines the performance evaluation information of the AI processor under test based on the performance data, and manages the performance evaluation information.
[0006] According to another aspect of the present invention, a performance evaluation apparatus for an AI processor is provided, comprising: The simulation task splitting module is used to determine the simulation task of the AI processor under test according to the performance evaluation instructions of the AI processor under test, and split the simulation task into multiple simulation sub-tasks. The simulation task arbitration module is used to arrange each simulation subtask according to the multi-dimensional priority evaluation value of each simulation subtask to obtain a priority task queue. For each simulation subtask in the priority task queue, the machine node and resource node in the simulation execution cluster that match the simulation subtask are determined. The simulation task execution module is used to submit each simulation subtask in the priority task queue to the simulation execution cluster for execution in the form of a target thread, and to monitor the task status changes of each simulation subtask, and to respond to the detected completed subtasks, suspended subtasks and timed-out subtasks. The performance data processing module is used to determine the performance data of the AI processor under test based on the simulated effective sub-tasks in each completed sub-task monitored, and to store the performance data. The performance evaluation module is used to determine the performance evaluation information of the AI processor under test based on the performance data through the performance analysis component, and to manage the performance evaluation information.
[0007] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores a computer program that is executed by the at least one processor, which enables the at least one processor to perform the performance evaluation method of the AI processor or the inference method of the hybrid expert model according to any embodiment of the present invention.
[0008] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions, the computer instructions being configured to cause a processor to execute and implement the performance evaluation method of the AI processor according to any embodiment of the present invention or the inference method of the hybrid expert model according to any embodiment of the present invention.
[0009] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the performance evaluation method of the AI processor or the inference method of the hybrid expert model according to any embodiment of the present invention.
[0010] The technical solution of this invention involves determining the simulation task of the AI processor under test according to the performance evaluation instructions, and then dividing the simulation task into multiple simulation subtasks. The simulation subtasks are then arranged according to their multi-dimensional priority evaluation values to obtain a priority task queue. For each simulation subtask in the priority task queue, the matching machine nodes and resource nodes in the simulation execution cluster are determined. Each simulation subtask in the priority task queue is submitted to the simulation execution cluster for execution as a target thread, and the task status changes of each simulation subtask are monitored. Responses are given to completed, suspended, and timed-out subtasks. Based on the monitored valid simulation subtasks among the completed subtasks, the performance data of the AI processor under test is determined and... Performance data is stored; finally, through a performance analysis component, the performance evaluation information of the AI processor under test is determined based on the performance data, and the performance evaluation information is managed. This solves the problems of high time and manpower costs, low efficiency, and difficulty in guaranteeing accuracy in the performance evaluation schemes of AI processors in related technologies. It can automatically and quickly determine the performance data of the AI processor based on simulation tasks, simulation execution clusters, and performance analysis components, and determine the performance evaluation information of the AI processor based on the performance data, and manage the performance evaluation information. This enables fast and accurate performance evaluation of AI processors, determines the performance evaluation information used to describe whether the AI processor has performance problems, reduces the time and manpower costs of the performance evaluation process, and improves the efficiency and accuracy of the performance evaluation process.
[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a flowchart of a performance evaluation method for an AI processor provided in Embodiment 1 of the present invention.
[0014] Figure 2 This is a flowchart of a performance evaluation method for an AI processor provided in Embodiment 2 of the present invention.
[0015] Figure 3 This is a schematic diagram of the structure of an AI processor performance evaluation device provided in Embodiment 3 of the present invention.
[0016] Figure 4 A schematic diagram of the structure of an electronic device for implementing the performance evaluation method of the AI processor in this embodiment of the invention. Detailed Implementation
[0017] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0018] It should be noted that the terms "target," "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising," "including," and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0019] Example 1 Figure 1This is a flowchart illustrating a performance evaluation method for an AI processor according to Embodiment 1 of the present invention. This embodiment is applicable to evaluating the performance of an AI processor and determining performance evaluation information to describe whether the AI processor has performance problems. This method can be executed by an AI processor performance evaluation device, which can be implemented in hardware and / or software and can be configured in an electronic device. This electronic device can be an electronic device installed in an enterprise for performing performance evaluation on the AI processor and determining performance evaluation information to describe whether the AI processor has performance problems. Figure 1 As shown, the method includes: Step 101: Based on the performance evaluation instructions of the AI processor under test, determine the simulation task of the AI processor under test, and break down the simulation task into multiple simulation sub-tasks.
[0020] Optionally, the AI processor under test can be an AI processor that requires performance evaluation. The performance evaluation instructions for the AI processor under test can be text used to instruct the AI processor under test to perform a performance evaluation.
[0021] Optionally, the performance evaluation instructions for the AI processor under test (UTP) include a simulation task for the UTP. The simulation task can consist of multiple test case groups required for the performance evaluation of the UTP. Each test case group consists of multiple related test cases. Simulation subtasks can be independently executable and configurable test cases separated from the simulation task. The performance characteristics of the AI processor include, but are not limited to: computing power, energy efficiency, architecture specificity, and scene adaptability. Each test case can be a text file used to instruct the input of specified simulation stimuli into the simulation model of the UTP, thereby triggering the simulation model to execute a specified performance-related operation and obtaining the operation record data of the specified performance-related operation output by the simulation model. The simulation model of the UTP can be a model pre-built in a specified simulation platform to simulate the UTP chip. The specified simulation platform can be a simulation platform selected from various simulation platforms. The simulation stimuli input for each test case are different. The simulation stimuli can be information used to instruct the simulation model of the UTP to execute a performance-related operation. The performance-related operation can be an operation related to the performance of the AI processor. Operation log data for performance-related operations can be performance-related data generated or collected during the execution of performance-related operations. By inputting simulation stimuli into the simulation model of the AI processor under test, the simulation model can be triggered to execute a specified performance-related operation, and the operation log data of the specified performance-related operation output by the simulation model can be obtained.
[0022] Optionally, the performance evaluation instructions for the AI processor under test can be input by the technician responsible for managing the AI processor under test. After obtaining the performance evaluation instructions input by the technician, the simulation task of the AI processor under test can be determined according to the performance evaluation instructions, and the simulation task can be broken down into multiple simulation sub-tasks.
[0023] Optionally, based on the performance evaluation instructions of the AI processor under test, the simulation task of the AI processor under test is determined, and the simulation task is divided into multiple simulation subtasks, including: parsing the performance evaluation instructions of the AI processor under test, and extracting the simulation task of the AI processor under test from the performance evaluation instructions; wherein, the simulation task contains multiple test case groups, and each test case group contains multiple test cases; each test case in the simulation task is determined as a simulation subtask, thereby dividing the simulation task into multiple simulation subtasks.
[0024] Optionally, the predefined performance evaluation instruction template can be a template used by technicians when writing performance evaluation instructions for the AI processor under test. The semantics of the performance evaluation instructions for the AI processor under test and the predefined performance evaluation instruction template can be parsed to extract the simulation task of the AI processor under test. Each test case in the simulation task of the AI processor under test can be identified as a simulation subtask, resulting in multiple simulation subtasks, thus breaking down the simulation task of the AI processor under test into multiple simulation subtasks. Therefore, the complete simulation task of the AI processor under test is decomposed into multiple simulation subtasks at the granularity of individual test cases. Each simulation subtask is independent and configurable.
[0025] Step 102: Arrange the simulation subtasks according to the multi-dimensional priority evaluation values of each simulation subtask to obtain a priority task queue. For each simulation subtask in the priority task queue, determine the machine node and resource node in the simulation execution cluster that match the simulation subtask.
[0026] Optionally, the multidimensional priority evaluation value of the simulation subtask can be a value used to characterize the priority of the simulation subtask, determined based on multiple dimensions of information describing the simulation subtask. A higher multidimensional priority evaluation value indicates a higher priority for the simulation subtask, while a lower value indicates a lower priority. The priority task queue can be a sequence obtained by arranging the simulation subtasks according to their multidimensional priority evaluation values.
[0027] Optionally, the simulation execution cluster can be a computer system set up within an enterprise to execute test cases in simulation tasks of an AI processor. The simulation execution cluster contains multiple machine nodes and multiple resource nodes. Machine nodes can be hardware modules used to execute test cases. Resource nodes can be hardware modules used to provide the necessary resources to the machine nodes. Resources include, but are not limited to, a Central Processing Unit (CPU) and memory. A machine node matching a simulation subtask can be a machine node used to execute the simulation subtask. A resource node matching a simulation subtask can be a resource node that provides the necessary resources to the machine node executing the simulation subtask.
[0028] Optionally, the performance evaluation instruction includes the task urgency, scenario description, and estimated simulation time for each test case. The task urgency can be a pre-set numerical value representing the urgency of executing the test case. The scenario description can be a pre-set numerical value describing the type of scenario in which the test case will be executed. The estimated simulation time can be the estimated duration of executing the test case. Multiple dimensions of information describing the simulation subtask include, but are not limited to, the task urgency, scenario description, and estimated simulation time of the simulation subtask. The task urgency, scenario description, and estimated simulation time of the simulation subtask can be obtained from the performance evaluation instruction.
[0029] Optionally, the simulation subtasks are arranged according to their multidimensional priority evaluation values to obtain a priority task queue. This includes: determining the multidimensional priority evaluation values of each simulation subtask based on its urgency, scenario description, and estimated simulation time; and arranging the simulation subtasks in descending order of their multidimensional priority evaluation values to obtain the priority task queue.
[0030] Optionally, based on the task urgency, scenario description, and estimated simulation time of each simulation subtask, the multidimensional priority evaluation value of each simulation subtask is determined, including: performing the following operation for each simulation subtask: using a preset weighted summation formula to perform a weighted summation on the task urgency, scenario description, and estimated simulation time of the simulation subtask, and determining the weighted summation result as the multidimensional priority evaluation value of the simulation subtask.
[0031] Optionally, the preset weighted summation formula can be a pre-defined calculation formula used to weight and sum the task urgency, scenario description, and estimated simulation time of the simulation subtasks. The weighted summation result of the task urgency, scenario description, and estimated simulation time of the simulation subtasks can be used to characterize the priority of the simulation subtasks. After determining the multidimensional priority evaluation values of each simulation subtask, the simulation subtasks can be arranged in descending order of their multidimensional priority evaluation values to form a sequence, thus obtaining a priority task queue.
[0032] Optionally, for each simulation subtask in the priority task queue, determine the machine nodes and resource nodes in the simulation execution cluster that match the simulation subtask, including: performing the following operations sequentially for each simulation subtask in the priority task queue: obtaining the simulation subtask from the priority queue, and selecting the machine nodes and resource nodes that match the simulation subtask from the machine nodes and resource nodes in the simulation execution cluster according to the trigger resource awareness strategy or feature information matching strategy.
[0033] Optionally, starting with the simulation subtask ranked first in the priority task queue, each simulation subtask is retrieved sequentially from the priority queue. Based on the trigger resource awareness strategy or characteristic information matching strategy, machine nodes and resource nodes matching the simulation subtask are selected from the machine nodes and resource nodes in the simulation execution cluster until the machine node and resource node matching the simulation subtask ranked last in the priority task queue are selected.
[0034] Optionally, based on the resource awareness triggering strategy, select machine nodes and resource nodes matching the simulation subtask from the machine nodes and resource nodes in the simulation execution cluster. This includes: collecting queue status information, machine configuration information, and load information of the simulation execution cluster, and storing the queue status information, machine configuration information, and load information in a real-time resource database; determining the health assessment information of each machine node and each resource node in the simulation execution cluster based on the queue status information, machine configuration information, and load information; wherein the health assessment information indicates availability or unavailability; selecting one machine node from the available machine nodes as the machine node matching the simulation subtask; and selecting one resource node from the available resource nodes as the resource node matching the simulation subtask.
[0035] Optionally, queue status information, machine configuration information, and load information of the simulation execution cluster can be collected. Queue status information can describe the status of task queues in the simulation execution cluster. Task queues in the simulation execution cluster can be queues used to store simulation subtasks received by the simulation execution cluster. Machine configuration information can be parameter values related to processor, CPU, model, stepping, clock frequency, cache, etc., for each machine node and resource node in the simulation execution cluster. Load information can describe the CPU and memory usage in the simulation execution cluster. Machine node health assessment information can characterize whether a machine node can be used to execute simulation subtasks. Machine node health assessment information is either available or unavailable. A machine node health assessment of available indicates that it can be used to execute simulation subtasks. A machine node health assessment of unavailable indicates that it cannot be used to execute simulation subtasks. Resource node health assessment information is also either available or unavailable. A resource node health assessment of available indicates that it can provide resources for machine nodes executing simulation subtasks. A resource node's health assessment information indicating it is unavailable means that the resource node cannot provide resources to machine nodes executing simulation subtasks. The real-time resource repository can be a database used to store information related to machine nodes and resource nodes in the simulation execution cluster. It can periodically collect queue status information, machine configuration information, and load information from the simulation execution cluster and update the queue status information, machine configuration information, and load information stored in the real-time resource repository.
[0036] Optionally, based on queue status information, machine configuration information, and load information, the health assessment information of each machine node and each resource node in the simulation execution cluster is determined. This includes: inputting the queue status information, machine configuration information, and load information into a resource health assessment component, and obtaining the health assessment information of each machine node and each resource node in the simulation execution cluster output by the resource health assessment component. The resource health assessment component can be a software module used to determine the health assessment information of each machine node and each resource node in the simulation execution cluster based on queue status information, machine configuration information, and load information. The queue status information, machine configuration information, and load information are input into the resource health assessment component. The resource health assessment component performs data cleaning and normalization on the queue status information, machine configuration information, and load information, then analyzes and detects the queue status information, machine configuration information, and load information to determine the health assessment information of each machine node and each resource node, and finally outputs the health assessment information of each machine node and each resource node.
[0037] Optionally, the real-time resource repository stores attribute information and model numbers for each machine node and resource node. Machine node attribute information can be text describing the machine node's independence, reversibility, and other attributes. Machine node model numbers can be text describing the type of the machine node. Similarly, resource node attribute information can be text describing the resource node's independence, reversibility, and other attributes. Resource node model numbers can be text describing the type of the resource node.
[0038] Optionally, the performance evaluation instruction includes the expected machine attributes, expected machine type, optimal resource attributes, and optimal resource type for each test case. The expected machine attributes of a test case can be attribute information of the machine node suitable for executing the test case. The expected machine type of a test case can be the model of the machine node suitable for executing the test case. The optimal resource attributes of a test case can be attribute information of the resource node suitable for executing the test case. The optimal resource type of a test case can be the model of the resource node suitable for executing the test case. The expected machine attributes, expected machine type, optimal resource attributes, and optimal resource type of the simulation subtask can be obtained from the performance evaluation instruction.
[0039] Optionally, based on the characteristic information matching strategy, select machine nodes and resource nodes that match the simulation subtask from the machine nodes and resource nodes in the simulation execution cluster, including: matching the simulation subtask and each machine node with attribute information and model based on the expected machine attributes and expected machine model of the simulation subtask to determine the machine node that matches the simulation subtask; and matching the simulation subtask and each resource node with attribute information and model based on the optimal resource attributes and optimal resource model of the simulation subtask to determine the resource node that matches the simulation subtask.
[0040] Optionally, based on the expected machine attributes and expected machine model of the simulation subtask, attribute information matching and model matching are performed on the simulation subtask and each machine node to determine the machine node that matches the simulation subtask, including: selecting a machine node as the machine node that matches the simulation subtask from each machine node whose attribute information is the same as the expected machine attributes of the simulation subtask and whose model is the same as the expected machine model of the simulation subtask.
[0041] Optionally, based on the optimal resource attributes and optimal resource model of the simulation subtask, attribute information matching and model matching are performed on the simulation subtask and each resource node to determine the resource node that matches the simulation subtask, including: selecting a resource node as the resource node that matches the simulation subtask from each resource node whose attribute information is the same as the optimal resource attribute of the simulation subtask and whose model is the same as the optimal resource model of the simulation subtask.
[0042] Step 103: Submit each simulation subtask in the priority task queue to the simulation execution cluster for execution as a target thread, and monitor the task status changes of each simulation subtask, and respond to the detected completed subtasks, suspended subtasks and timed-out subtasks.
[0043] Optionally, the target thread can be single-threaded or multi-threaded. Task flows can be used to submit each simulation subtask in the priority task queue to the simulation execution cluster, instructing the machine nodes and resource nodes in the simulation execution cluster that match each simulation subtask to execute each simulation subtask in single-threaded or multi-threaded mode.
[0044] Optionally, a completed subtask can refer to a simulation subtask that completes normally. A suspended subtask can refer to a simulation subtask that stops executing due to an exception. A timeout subtask can refer to a simulation subtask whose execution time exceeds a preset time threshold. The preset time threshold can be a pre-set time threshold.
[0045] Optionally, monitoring the task status changes of each simulation subtask includes: collecting change perception information of each simulation subtask, and monitoring whether each simulation subtask is updated to a completed subtask, a suspended subtask, or a timed-out subtask based on the collected change perception information; wherein, the change perception information includes task start time, task end time, execution machine information, and / or real-time task status.
[0046] Optionally, change awareness information for each simulation subtask can be collected periodically at preset configurable time intervals. The collection frequency can be adjusted to reduce system load when a new simulation subtask is submitted, during peak or off-peak periods of simulation subtask processing. Specifically, the collection frequency can be decreased when a new simulation subtask is submitted, increased during peak periods of simulation subtask processing, and decreased during off-peak periods.
[0047] Optionally, the task start time can be the time recorded by the machine node matching the simulation subtask to begin executing the simulation subtask. The task end time can be the time recorded by the machine node matching the simulation subtask to complete executing the simulation subtask. Execution machine information can be relevant information about the machine node matching the simulation subtask. The task real-time status can be text set by the machine node matching the simulation subtask to characterize the execution status of the simulation subtask. The task real-time status can be initial suspended status, execution status, environment resource suspended status, disconnected suspended status, system suspended status, user suspended status, normal completion status, or abnormal exit status.
[0048] Optionally, the initial suspension status can be text indicating that the machine node matching the simulation subtask has not yet started executing the simulation subtask. The execution status can be text indicating that the machine node matching the simulation subtask is executing the simulation subtask. The environment resource suspension status can be text indicating that the machine node matching the simulation subtask cannot execute the simulation subtask due to unsatisfactory environment or resources. The disconnection suspension status can be text indicating that the machine node matching the simulation subtask cannot execute the simulation subtask due to program disconnection. The user suspension status can be text indicating that the machine node matching the simulation subtask cannot execute the simulation subtask due to user operation. The system suspension status can be text indicating that the machine node matching the simulation subtask cannot execute the simulation subtask due to operation of the simulation execution cluster. The normal completion status can be text indicating that the machine node matching the simulation subtask has successfully completed the simulation subtask. The abnormal exit status can be text indicating that the machine node matching the simulation subtask has forcibly terminated the execution of the simulation subtask. For example, the initial suspension state is PEND, the execution state is RUN, the environment resource suspension state is PSUSP, the disconnection suspension state is UNKWN, the system suspension state is SSUSP, the user suspension state is USUSP, the normal completion state is DONE, and the abnormal exit state is EXIT.
[0049] Optionally, based on the collected change perception information, monitor whether each simulation subtask is updated to a completed subtask, a suspended subtask, or a timed-out subtask, including: performing the following operations for each simulation subtask: when the latest collected task real-time status is a normal completion status, determine that the simulation subtask is updated to a completed subtask; when the latest collected task real-time status is an environment resource suspended status, a disconnected suspended status, a system suspended status, or a user suspended status, determine that the simulation subtask is updated to a suspended subtask; when the time difference between the current time and the latest collected task start time is greater than a preset duration threshold, and no task end time has been collected, determine that the simulation subtask is updated to a timed-out subtask; when the latest collected execution machine information contains a timeout prompt, determine that the simulation subtask is updated to a timed-out subtask; when the latest collected execution machine information contains a suspension prompt, determine that the simulation subtask is updated to a suspended subtask.
[0050] Optionally, the timeout message can be used to indicate that the execution time of the simulation subtask by the machine node matching the simulation subtask exceeds a preset time threshold. The suspension message can be used to indicate that the execution of the simulation subtask by the machine node matching the simulation subtask has stopped due to an exception.
[0051] Optionally, responses can be made to detected completed subtasks, suspended subtasks, and timed-out subtasks, including: after detecting a completed subtask, checking whether the simulation function of the completed subtask is correct and whether the completed subtask is effective for performance analysis; determining completed subtasks with correct simulation function and effective performance analysis as valid simulation subtasks; and rerunning completed subtasks with failed simulation function or ineffective performance analysis according to the automatic rerun strategy.
[0052] Optionally, "successfully completing the simulation function of the subtask" can mean that the operation of inputting the specified simulation stimulus into the simulation model of the AI processor under test, thereby triggering the simulation model of the AI processor under test to execute the specified performance-related operation, and obtaining the operation record data of the specified performance-related operation output by the simulation model of the AI processor under test, has been correctly executed. "Failed to complete the simulation function of the subtask" can mean that the operation of inputting the specified simulation stimulus into the simulation model of the AI processor under test, thereby triggering the simulation model of the AI processor under test to execute the specified performance-related operation, and obtaining the operation record data of the specified performance-related operation output by the simulation model of the AI processor under test, has not been correctly executed. "Effective for performance analysis" means that the operation record data obtained by executing the subtask can be used to measure whether the AI processor has performance problems. "Ineffective for performance analysis" means that the operation record data obtained by executing the subtask cannot be used to measure whether the AI processor has performance problems. A valid simulation subtask is a simulation subtask in which the indicated operation has been correctly executed and the obtained operation record data can be used to measure whether the AI processor has performance problems.
[0053] Optionally, the simulation function of completing the subtask is checked for correctness, including: checking whether the operation record data obtained by executing the subtask contains integrity flag information and simulation result flag information; if the operation record data obtained by executing the subtask contains integrity flag information and simulation result flag information, the simulation function of completing the subtask is determined to be correct; if the operation record data obtained by executing the subtask does not contain integrity flag information or simulation result flag information, the simulation function of completing the subtask is determined to have failed. The integrity flag information can be text used to indicate that the operation record data is complete. The simulation result flag information can be text used to indicate that the operation to which the operation record data belongs was successfully executed.
[0054] Optionally, the determination of whether the completed subtask is effective for performance analysis includes: detecting whether the operation log data obtained by executing the completed subtask contains key log feature markers; if the operation log data obtained by executing the completed subtask contains key log feature markers, then the completion of the subtask is determined to be effective for performance analysis; if the operation log data obtained by executing the completed subtask does not contain key log feature markers, then the completion of the subtask is determined to be invalid for performance analysis. Key log feature markers can be text used to characterize that the operation log data contains key data that can be used to measure whether the AI processor has performance problems.
[0055] Optionally, according to the automatic rerun strategy, completed subtasks that fail in simulation or are invalid for performance analysis are rerunned, including: for each completed subtask that fails in simulation or is invalid for performance analysis, the following operation is performed: instructing the machine node matching the completed subtask to re-execute the completed subtask until the completed subtask is determined to be a valid simulation subtask or the number of times the completed subtask is re-executed reaches a preset threshold. The preset threshold can be a pre-set threshold. For example, the preset threshold is 3.
[0056] Optionally, after the number of times a completed subtask has been re-executed reaches a preset threshold, the re-execution of the completed subtask is stopped, and the completed subtask is designated as a forcibly terminated task. A forcibly terminated task can be a simulation subtask that cannot be successfully executed and does not need to be re-executed.
[0057] Optionally, responses can be made to detected completed subtasks, suspended subtasks, and timed-out subtasks, including: confirming the wake-up of a suspended subtask after it is detected; and adjusting resources for a timed-out subtask after it is detected.
[0058] Optionally, the process of waking up a suspended subtask includes: instructing the machine node matching the suspended subtask to re-execute the suspended subtask; if the suspended subtask is updated to a completed subtask, then the waking up of the suspended subtask is confirmed to be successful; if the suspended subtask is updated to a suspended subtask again, then the waking up of the suspended subtask is confirmed to be a failed task, and the suspended subtask is determined to be a task to be forcibly terminated.
[0059] Optionally, resource adjustments can be made to the timed-out subtask, including: re-identifying the machine node and resource node that match the timed-out subtask, and instructing the re-identified machine node that matches the timed-out subtask to execute the timed-out subtask. If the timed-out subtask is updated to a timed-out subtask again, then the timed-out subtask is determined as a task to be forcibly terminated.
[0060] Optionally, the simulation success queue can be a pre-defined queue for storing valid simulation subtasks and related information. The simulation failure queue can be a pre-defined queue for storing forced termination tasks and related information. The simulation pending queue can be a pre-defined queue for storing simulation subtasks that need to be re-executed.
[0061] Valid simulation subtasks and operation logs obtained from executing valid simulation subtasks can be stored in the simulation success queue. Forced termination tasks and their failure reasons can be stored in the simulation failure queue. The failure reason information for forced termination tasks can be collected or registered information describing why the forced termination task could not be executed normally. Simulation subtasks that need to be re-executed can be stored in the simulation pending queue.
[0062] Step 104: Based on the simulated effective sub-tasks in each completed sub-task monitored, determine the performance data of the AI processor under test and store the performance data.
[0063] Optionally, after each simulation subtask has been updated to a valid simulation subtask or the task has been forcibly terminated, the performance data of the AI processor under test can be determined and stored based on the valid simulation subtasks among the completed subtasks.
[0064] Optionally, the performance data of the AI processor under test is data that can be used to measure whether the AI processor under test has performance problems. Operation record data obtained by executing each effective sub-task of the simulation can be used to measure whether the AI processor under test has performance problems.
[0065] Optionally, based on the monitored effective sub-tasks in each completed sub-task, the performance data of the AI processor under test is determined and stored, including: determining the operation record data obtained by executing each effective sub-task as the performance data of the AI processor under test; and storing the performance data of the AI processor under test in a performance database. The performance database can be a database used to store the performance data of the AI processor.
[0066] Step 105: Using the performance analysis component, determine the performance evaluation information of the AI processor under test based on the performance data, and manage the performance evaluation information.
[0067] Optionally, the performance analysis component can be a pre-configured software system used to determine the performance evaluation information of the AI processor under test based on the performance data of the AI processor, and to manage the performance evaluation information of the AI processor. The input to the performance analysis component is the performance data of the AI processor, and the output is the performance evaluation information of the AI processor. The performance evaluation information of the AI processor can be information describing whether the AI processor has performance problems. The performance evaluation information includes performance evaluation results. The performance evaluation results are text used to characterize whether the AI processor has performance problems. The performance evaluation results are either normal or abnormal. A performance evaluation result of normal indicates that the AI processor has performance problems. A performance evaluation result of abnormal indicates that the AI processor does not have performance problems. The performance data of the AI processor under test can be input into the performance analysis component. The performance analysis component analyzes and detects the performance data of the AI processor under test, determines the performance evaluation information of the AI processor under test, and manages the performance evaluation information of the AI processor under test. The performance analysis component's management of the performance evaluation information includes making the performance evaluation information of the AI processor under test available for viewing and analysis by technical personnel.
[0068] Optionally, the performance analysis component can provide performance query and analysis services based on the performance evaluation information of the AI processor under test. Performance query and analysis services can include performance presentation, performance analysis, performance mapping, performance alerts, and performance tuning. Performance presentation refers to the performance analysis component providing the performance evaluation information of the AI processor under test to technicians based on the performance points they are interested in, allowing them to analyze the processor's performance. Performance analysis refers to the performance analysis component providing the performance evaluation information of the AI processor under test to technicians based on different dimensions specified by the technicians, allowing them to analyze the processor's performance. Performance mapping refers to the performance analysis component providing a performance map of the AI processor under test to technicians based on its performance evaluation information. Performance alerts refer to the performance analysis component viewing the performance evaluation information of the AI processor under test, broadcasting the triggering simulation stimuli via email or other means to generate new simulation stimulus configurations and initiate a new round of performance evaluation. Performance tuning refers to the performance analysis component using the performance evaluation information of the AI processor under test, through AI modeling analysis and optimization, to trigger the generation of new simulation stimulus configurations and initiate a new round of performance evaluation, improving the efficiency of performance verification and optimization.
[0069] The technical solution of this invention involves determining the simulation task of the AI processor under test according to the performance evaluation instructions, and then dividing the simulation task into multiple simulation subtasks. The simulation subtasks are then arranged according to their multi-dimensional priority evaluation values to obtain a priority task queue. For each simulation subtask in the priority task queue, the matching machine nodes and resource nodes in the simulation execution cluster are determined. Each simulation subtask in the priority task queue is submitted to the simulation execution cluster for execution as a target thread, and the task status changes of each simulation subtask are monitored. Responses are given to completed, suspended, and timed-out subtasks. Based on the monitored valid simulation subtasks among the completed subtasks, the performance data of the AI processor under test is determined and... Performance data is stored; finally, through a performance analysis component, the performance evaluation information of the AI processor under test is determined based on the performance data, and the performance evaluation information is managed. This solves the problems of high time and manpower costs, low efficiency, and difficulty in guaranteeing accuracy in the performance evaluation schemes of AI processors in related technologies. It can automatically and quickly determine the performance data of the AI processor based on simulation tasks, simulation execution clusters, and performance analysis components, and determine the performance evaluation information of the AI processor based on the performance data, and manage the performance evaluation information. This enables fast and accurate performance evaluation of AI processors, determines the performance evaluation information used to describe whether the AI processor has performance problems, reduces the time and manpower costs of the performance evaluation process, and improves the efficiency and accuracy of the performance evaluation process.
[0070] The technical solution of this invention provides a highly efficient performance evaluation method to ensure the best cost-effectiveness of manpower deployment, server resource deployment, performance problem detection, and performance scenario optimization, ending the current situation where manpower deployment needs surge with the increase of processor complexity. The technical solution of this invention can identify the most critical scenario subspaces affecting performance from a massive number of performance evaluation scenarios, and efficiently control and manage the simulation process, control and optimize computing power consumption, automatically predict future more valuable performance evaluation subspaces, automatically perform multiple iterations, and record performance data from each round to complete analysis and performance presentation.
[0071] The technical solution of this invention is a hardware architecture performance evaluation method. From an engineering implementation perspective, it meticulously designs the flow strategies between each stage, ensuring zero performance defects in multi-generational projects. This invention's solution is compatible with multiple performance evaluation simulation platforms, a unified task distribution management mechanism, a unified data collection and storage mechanism, multi-digit performance data analysis methods, performance data comparison methods, and performance data presentation methods. It avoids various inefficiencies inherent in traditional methods, as well as the lack of susceptibility and compatibility in traditional processes.
[0072] The technical solution of this invention models the performance evaluation process of processor hardware architecture, overcoming the problems of poor portability, low integration, and insufficient concurrency of traditional regression verification methods on different verification platforms. It also reduces development and maintenance costs and improves overall verification efficiency. The technical solution of this invention establishes a methodological model of the overall performance evaluation process. All key links in the model support flexible configuration to meet any adjustment needs during actual project execution, ensuring seamless implementation. The technical solution of this invention models all real platforms involved in the performance evaluation process, eliminating platform differences in the performance evaluation cycle. Addressing the limitations of task deployment between compilation tools and computing core scheduling, the technical solution of this invention can automatically decompose the compilation process, breaking the coupling between compilation and computing power scheduling. The decomposed regression simulation tasks are scheduled to the server cluster using a proprietary process, overcoming the limitation of the original process scheduling's maximum number of parallel tasks being affected by the number of server cores, greatly improving parallel execution efficiency. The technical solution of this invention creates a performance data framework, supported by a performance database, creating a dynamic automated regression process with feedback capabilities. This regression process, combined with AI methods, can continuously and automatically iterate according to performance targets to complete performance evaluation or optimization work.
[0073] Example 2 Figure 2 This is a flowchart illustrating a performance evaluation method for an AI processor according to Embodiment 2 of the present invention. Embodiments of the present invention can be combined with various optional solutions from one or more of the above embodiments. For example... Figure 2 As shown, the method includes: Step 201: Based on the performance evaluation instructions of the AI processor under test, determine the simulation task of the AI processor under test, and break down the simulation task into multiple simulation sub-tasks.
[0074] Step 202: Determine the multidimensional priority evaluation value of each simulation subtask based on its urgency, scenario description, and estimated simulation time.
[0075] Step 203: Arrange each simulation subtask in descending order of multidimensional priority evaluation values to obtain a priority task queue.
[0076] Step 204: For each simulation subtask in the priority task queue, retrieve the simulation subtask from the priority queue, and select the machine node and resource node that match the simulation subtask from the machine node and resource node in the simulation execution cluster according to the trigger resource awareness strategy or characteristic information matching strategy.
[0077] Step 205: Submit each simulation subtask in the priority task queue to the simulation execution cluster for execution as a target thread, and monitor the task status changes of each simulation subtask, and respond to the detected completed subtasks, suspended subtasks and timed-out subtasks.
[0078] Step 206: Based on the simulated effective sub-tasks in each completed sub-task monitored, determine the performance data of the AI processor under test and store the performance data.
[0079] Step 207: Using the performance analysis component, determine the performance evaluation information of the AI processor under test based on the performance data, and manage the performance evaluation information.
[0080] The technical solution of this invention can determine the multi-dimensional priority evaluation value of each simulation subtask based on the urgency of the task, the scenario description, and the expected simulation time, and arrange each simulation subtask according to the multi-dimensional priority evaluation value to obtain a priority task queue. Based on the trigger resource awareness strategy or the characteristic information matching strategy, the machine nodes and resource nodes that match the simulation subtask can be selected quickly and accurately from the machine nodes and resource nodes in the simulation execution cluster.
[0081] Example 3 Figure 3 This is a schematic diagram of a performance evaluation device for an AI processor provided in Embodiment 3 of the present invention. The device can be configured in an electronic device. Figure 3 As shown, the device includes: a simulation task splitting module 301, a simulation task arbitration module 302, a simulation task execution module 303, a performance data processing module 304, and a performance evaluation module 305.
[0082] The simulation task splitting module 301 is used to determine the simulation task of the AI processor under test according to the performance evaluation instructions of the AI processor under test, and split the simulation task into multiple simulation subtasks; the simulation task arbitration module 302 is used to arrange the simulation subtasks according to the multi-dimensional priority evaluation values of each simulation subtask to obtain a priority task queue, and for each simulation subtask in the priority task queue, determine the machine node and resource node in the simulation execution cluster that match the simulation subtask; the simulation task execution module 303 is used to submit each simulation subtask in the priority task queue to the simulation execution cluster for execution in the form of target threads, and monitor the task status changes of each simulation subtask, and respond to the monitored completed subtasks, suspended subtasks, and timed-out subtasks; the performance data processing module 304 is used to determine the performance data of the AI processor under test according to the monitored effective simulation subtasks in each completed subtask and store the performance data; the performance evaluation module 305 is used to determine the performance evaluation information of the AI processor under test according to the performance data through a performance analysis component, and manage the performance evaluation information.
[0083] The technical solution of this invention involves determining the simulation task of the AI processor under test according to the performance evaluation instructions, and then dividing the simulation task into multiple simulation subtasks. The simulation subtasks are then arranged according to their multi-dimensional priority evaluation values to obtain a priority task queue. For each simulation subtask in the priority task queue, the matching machine nodes and resource nodes in the simulation execution cluster are determined. Each simulation subtask in the priority task queue is submitted to the simulation execution cluster for execution as a target thread, and the task status changes of each simulation subtask are monitored. Responses are given to completed, suspended, and timed-out subtasks. Based on the monitored valid simulation subtasks among the completed subtasks, the performance data of the AI processor under test is determined and... Performance data is stored; finally, through a performance analysis component, the performance evaluation information of the AI processor under test is determined based on the performance data, and the performance evaluation information is managed. This solves the problems of high time and manpower costs, low efficiency, and difficulty in guaranteeing accuracy in the performance evaluation schemes of AI processors in related technologies. It can automatically and quickly determine the performance data of the AI processor based on simulation tasks, simulation execution clusters, and performance analysis components, and determine the performance evaluation information of the AI processor based on the performance data, and manage the performance evaluation information. This enables fast and accurate performance evaluation of AI processors, determines the performance evaluation information used to describe whether the AI processor has performance problems, reduces the time and manpower costs of the performance evaluation process, and improves the efficiency and accuracy of the performance evaluation process.
[0084] In an optional embodiment of the present invention, the simulation task splitting module 301 is specifically used to: parse the performance evaluation instructions of the AI processor under test, and extract the simulation task of the AI processor under test from the performance evaluation instructions; wherein, the simulation task contains multiple test case groups, and each test case group contains multiple test cases; and determine each test case in the simulation task as a simulation subtask, thereby splitting the simulation task into multiple simulation subtasks.
[0085] In an optional embodiment of the present invention, the simulation task arbitration module 302 is specifically used to: determine the multi-dimensional priority evaluation value of each simulation sub-task based on the task urgency, scenario description and estimated simulation time of each simulation sub-task; and arrange each simulation sub-task in descending order of the multi-dimensional priority evaluation value to obtain a priority task queue.
[0086] In an optional embodiment of the present invention, the simulation task arbitration module 302 is specifically configured to: sequentially perform the following operations for each simulation subtask in the priority task queue: obtain the simulation subtask from the priority queue, and select the machine node and resource node that match the simulation subtask from the machine nodes and resource nodes in the simulation execution cluster according to the trigger resource awareness strategy or characteristic information matching strategy.
[0087] In an optional embodiment of the present invention, the simulation task execution module 303 is specifically used to: monitor the task status changes of each simulation subtask, including: collecting change perception information of each simulation subtask, and monitoring whether each simulation subtask is updated to a completed subtask, a suspended subtask, or a timed-out subtask based on the collected change perception information; wherein, the change perception information includes task start time, task end time, execution machine information, and / or task real-time status.
[0088] In an optional embodiment of the present invention, the simulation task execution module 303 is specifically used to: after detecting the completion of a subtask, detect whether the simulation function of the completed subtask is correct and whether the completed subtask is effective for performance analysis, determine the completed subtask with correct simulation function and effective performance analysis as a valid simulation subtask, and rerun the completed subtask with failed simulation function or ineffective performance analysis according to the automatic rerun strategy.
[0089] In an optional embodiment of the present invention, the simulation task execution module 303 is specifically used to: confirm the wake-up of the suspended subtask after detecting the suspended subtask; and adjust the resources of the timed-out subtask after detecting the timed-out subtask.
[0090] The AI processor performance evaluation device provided in this embodiment of the invention can execute the AI processor performance evaluation method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method execution.
[0091] Example 4 Figure 4 A schematic diagram of an electronic device 10, which can be used to implement the performance evaluation method of an AI processor according to embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, electronic devices, blade electronic devices, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0092] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory 12 or a random access memory 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the read-only memory 12 or loaded from storage unit 18 into the random access memory 13. The random access memory 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, read-only memory 12, and random access memory 13 are interconnected via a bus 14. An input / output interface 15 is also connected to the bus 14.
[0093] Multiple components in electronic device 10 are connected to input / output interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of monitors, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0094] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, central processing units, graphics processing units, various special-purpose artificial intelligence computing chips, various processors running machine learning model algorithms, digital signal processors, and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as performance evaluation methods for AI processors.
[0095] In some embodiments, the AI processor performance evaluation method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as a storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed on a heterogeneous hardware accelerator via read-only memory and / or a communication unit. When the computer program is loaded into random access memory and executed by the processor, one or more steps of the AI processor performance evaluation method described above can be performed. Alternatively, in other embodiments, the processor can be configured to perform the AI processor performance evaluation method by any other suitable means (e.g., by means of firmware).
[0096] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays, application-specific integrated circuits (ASICs), application-specific standard products (ASICs), systems-on-a-chip (SoCs), payload programmable logic devices (PLCs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0097] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or electronic device.
[0098] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0099] To provide user interaction, the systems and techniques described herein can be implemented on a heterogeneous hardware accelerator, which includes: a display device (e.g., a cathode ray tube or liquid crystal display monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the heterogeneous hardware accelerator. Other types of devices can also be used to provide user interaction; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or haptic feedback); and input from the user can be received in any form (including sound input, voice input, or haptic input).
[0100] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data electronic devices), or computing systems that include middleware components (e.g., application electronic devices), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0101] A computing system can include clients and electronic devices. Clients and electronic devices are generally geographically separated and typically interact via communication networks. The client-electronic device relationship is created by computer programs running on the respective computers and establishing a client-electronic device relationship between them. Electronic devices can be cloud electronic devices, also known as cloud computing electronic devices or cloud servers, which are host products within the cloud computing service system. These address the shortcomings of traditional physical hosts and virtual private server services, such as high management difficulty and weak business scalability.
[0102] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0103] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A performance evaluation method for an AI processor, characterized in that, include: Based on the performance evaluation instructions of the AI processor under test, the simulation task of the AI processor under test is determined, and the simulation task is broken down into multiple simulation sub-tasks. The simulation subtasks are arranged according to their multidimensional priority evaluation values to obtain a priority task queue. For each simulation subtask in the priority task queue, the machine nodes and resource nodes in the simulation execution cluster that match the simulation subtask are determined. Each simulation subtask in the priority task queue is submitted to the simulation execution cluster for execution as a target thread, and the task status changes of each simulation subtask are monitored. The detected completed subtasks, suspended subtasks, and timed-out subtasks are responded to. Based on the monitored effective sub-tasks in each completed sub-task, the performance data of the AI processor under test is determined and stored. The performance analysis component determines the performance evaluation information of the AI processor under test based on the performance data, and manages the performance evaluation information.
2. The performance evaluation method for an AI processor according to claim 1, characterized in that, Based on the performance evaluation instructions for the AI processor under test, the simulation task of the AI processor under test is determined, and the simulation task is broken down into multiple simulation sub-tasks, including: The performance evaluation instructions for the AI processor under test are parsed, and the simulation task of the AI processor under test is extracted from the performance evaluation instructions; wherein, the simulation task contains multiple test case groups, and each test case group contains multiple test cases; Each test case in the simulation task is identified as a simulation subtask, thereby splitting the simulation task into multiple simulation subtasks.
3. The performance evaluation method for an AI processor according to claim 1, characterized in that, The simulation subtasks are arranged according to their multidimensional priority evaluation values to obtain a priority task queue, including: Based on the task urgency, scenario description, and estimated simulation time of each simulation subtask, determine the multidimensional priority evaluation value of each simulation subtask. The simulation subtasks are arranged in descending order of their multidimensional priority evaluation values to obtain a priority task queue.
4. The performance evaluation method for an AI processor according to claim 1, characterized in that, For each simulation subtask in the priority task queue, determine the machine nodes and resource nodes in the simulation execution cluster that match the simulation subtask, including: Perform the following operations sequentially for each simulation subtask in the priority task queue: The simulation subtask is obtained from the priority queue. Based on the trigger resource awareness strategy or feature information matching strategy, the machine node and resource node that match the simulation subtask are selected from the machine nodes and resource nodes in the simulation execution cluster.
5. The performance evaluation method for an AI processor according to claim 1, characterized in that, Monitor the task state changes of each simulation subtask, including: Collect change perception information for each simulation subtask, and monitor whether each simulation subtask is updated to a completed subtask, a suspended subtask, or a timed-out subtask based on the collected change perception information; wherein, the change perception information includes task start time, task end time, execution machine information, and / or task real-time status.
6. The performance evaluation method for an AI processor according to claim 1, characterized in that, Respond to detected completed subtasks, suspended subtasks, and timed-out subtasks, including: After detecting the completion of a subtask, it checks whether the simulation function of the completed subtask is correct and whether the completion of the subtask is effective for performance analysis. Completed subtasks with correct simulation function and effective performance analysis are identified as valid simulation subtasks. According to the automatic rerun strategy, completed subtasks with failed simulation function or ineffective performance analysis are rerun.
7. The performance evaluation method for an AI processor according to claim 1, characterized in that, Respond to detected completed subtasks, suspended subtasks, and timed-out subtasks, including: After detecting a suspended subtask, the suspended subtask is awakened and confirmed. After detecting a timeout subtask, resources are adjusted for the timeout subtask.
8. A performance evaluation device for an AI processor, characterized in that, include: The simulation task splitting module is used to determine the simulation task of the AI processor under test according to the performance evaluation instructions of the AI processor under test, and split the simulation task into multiple simulation sub-tasks. The simulation task arbitration module is used to arrange each simulation subtask according to the multi-dimensional priority evaluation value of each simulation subtask to obtain a priority task queue. For each simulation subtask in the priority task queue, the machine node and resource node in the simulation execution cluster that match the simulation subtask are determined. The simulation task execution module is used to submit each simulation subtask in the priority task queue to the simulation execution cluster for execution in the form of target threads, and to monitor the task status changes of each simulation subtask, and to respond to the detected completed subtasks, suspended subtasks and timed-out subtasks. The performance data processing module is used to determine the performance data of the AI processor under test based on the simulated effective sub-tasks in each completed sub-task monitored, and to store the performance data. The performance evaluation module is used to determine the performance evaluation information of the AI processor under test based on the performance data through the performance analysis component, and to manage the performance evaluation information.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores a computer program that is executed by the at least one processor, which enables the at least one processor to perform the performance evaluation method of the AI processor according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the performance evaluation method of the AI processor according to any one of claims 1-7.