Artificial intelligence chip performance test method and device, electronic equipment and medium
Through an artificial intelligence chip performance testing method, the problem of difficult and costly development of existing test methods is solved, and efficient and fair test results are achieved through a unified testing framework.
Patent Information
- Application Number
- CN202411937187.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-09
AI Technical Summary
The development of existing artificial intelligence chip test methods is difficult, has a large workload, has a long cycle and is costly, especially in the case of a variety of chip manufacturers and specifications.
A method for testing artificial intelligence chip performance is proposed. By obtaining the chip to be tested and the target workload, including test models, data sets, model weight accuracy, machine learning framework and batch size, data preprocessing, model adaptation, and inference processing, model accuracy and power consumption information are determined, and test results are output.
By building a unified testing software framework, the development difficulty is reduced, the testing efficiency is improved, and the fairness and comparability of the test results are ensured.
Smart Images

Figure CN119961066A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of testing, and in particular to an artificial intelligence chip performance testing method, device, electronic device and medium. Background Art
[0002] Artificial intelligence chips are a new generation of microprocessors specially designed to handle artificial intelligence tasks. They have the advantages of high performance and high energy efficiency and are widely used in the fields of autonomous driving, smart home appliances, robots, etc. With the development of unmanned and intelligent technology, artificial intelligence chips have gradually been applied to the fields of aviation, aerospace and military industry.
[0003] The current testing method is for designers to design and test programs according to requirements. When there are many optional chip manufacturers and specifications, it will lead to problems such as difficult development, heavy workload, long cycle, and high cost. Summary of the invention
[0004] In view of this, the purpose of the present disclosure is to propose an artificial intelligence chip performance testing method, device, electronic device and medium to solve or partially solve the above-mentioned problems.
[0005] Based on the above purpose, the first aspect of the present disclosure provides an artificial intelligence chip performance testing method, the method comprising:
[0006] Obtaining the artificial intelligence chip to be tested and the target workload, wherein the target workload includes a test model, a test data set, model weight numerical accuracy, a machine learning framework, and a batch size;
[0007] Preprocessing the test data set to obtain a target data set;
[0008] Determine a programming language for the performance test of the artificial intelligence chip, and use the programming language to perform model adaptation processing on the test model to obtain a target test model;
[0009] Using the target test model equipped with the artificial intelligence chip to be tested, the target data set is processed to obtain a model inference result;
[0010] Determining a model reasoning accuracy value corresponding to the target test model according to the model reasoning result and the target data set;
[0011] Determine the target test model to obtain power consumption information corresponding to the model reasoning result, use the model reasoning result, the model reasoning accuracy value and the power consumption information as the test result, and output the test result.
[0012] Based on the same inventive concept, the second aspect of the present disclosure proposes an artificial intelligence chip performance testing device, comprising:
[0013] A data acquisition module is configured to acquire the artificial intelligence chip to be tested and the target workload, wherein the target workload includes a test model, a test data set, a model weight numerical accuracy, a machine learning framework, and a batch size;
[0014] A data preprocessing module is configured to preprocess the test data set to obtain a target data set;
[0015] A model adaptation module is configured to determine a programming language for the artificial intelligence chip performance test, and use the programming language to perform model adaptation processing on the test model to obtain a target test model;
[0016] An inference module is configured to use the target test model equipped with the artificial intelligence chip to perform inference on the target data set to obtain a model inference result;
[0017] A model reasoning accuracy determination module is configured to determine a model reasoning accuracy value corresponding to the target test model according to the model reasoning result and the target data set;
[0018] The output module is configured to determine the target test model to obtain power consumption information corresponding to the model reasoning result, take the model reasoning result, the model reasoning accuracy value and the power consumption information as the test result, and output the test result.
[0019] Based on the same inventive concept, the third aspect of the present disclosure proposes an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor implements the chip performance testing method as described above when executing the computer program.
[0020] Based on the same inventive concept, the fourth aspect of the present disclosure proposes a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to execute the chip performance testing method as described above.
[0021] As can be seen from the above, the present disclosure proposes an artificial intelligence chip performance test method, device, electronic device and medium, obtains the artificial intelligence chip to be tested and the target workload, wherein the target workload includes a test model, a test data set, a model weight numerical accuracy, a machine learning framework and a batch size, and the test model is a model of the target workload equipped with the artificial intelligence chip to be tested. The test data set is preprocessed to obtain a target data set. The programming language for the artificial intelligence chip performance test is determined, and the test model is adapted to obtain a target test model using the programming language. The target test model equipped with the artificial intelligence chip to be tested is used to infer the target data set to obtain a model inference result. The performance test of the artificial intelligence chip to be tested is realized by running the model equipped with the artificial intelligence chip to be tested. According to the model inference result and the target data set, the model inference accuracy value corresponding to the target test model is determined. The target test model is determined to obtain the power consumption information corresponding to the model inference result, and the model inference result, the model inference accuracy value and the power consumption information are used as the test result, and the test result is output. By building a test software framework for different workloads, the same framework can be used to test different chips, which reduces the development difficulty and improves the efficiency of chip testing. At the same time, using the same framework for testing can also make the test results of different artificial intelligence chips to be tested more fair and comparable. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the drawings required for use in the embodiments or related technical descriptions are briefly introduced below. Obviously, the drawings described below are only embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0023] Figure 1 This is a flow chart of the artificial intelligence chip performance testing method according to an embodiment of the present disclosure;
[0024] Figure 2 This is a flow chart of an artificial intelligence chip performance testing method according to another embodiment of the present disclosure;
[0025] Figure 3 This is a flow chart of an artificial intelligence chip performance testing method according to another embodiment of the present disclosure;
[0026] Figure 4 This is a structural block diagram of the artificial intelligence chip performance test benchmark framework of the embodiment of the present disclosure;
[0027] Figure 5 It is a schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0028] In order to make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.
[0029] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should be understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. "Including" or "comprising" and similar words mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0030] The terms used in this disclosure are explained as follows:
[0031] Python: Python is a high-level scripting language that is currently used in Web and Internet development, scientific computing and statistics, education, software development, and back-end development. It has the advantages of being easy to learn, portable, extensible, and embeddable.
[0032] MobileNet: The MobileNet model is a lightweight model proposed by Google in 2017 for mobile phones or embedded systems.
[0033] LSTM: LSTM, the full name of Long Short Term Memory, is a special recursive neural network.
[0034] ResNet: Deep Residual Networks (ResNet), a standard ResNet model consists of multiple residual blocks, usually starting with a common convolutional layer and pooling layer for preliminary feature extraction.
[0035] Mobilenet-SSD: MobileNet-SSD is an object detection network model that combines MobileNet and SSD. By using deep separable convolution and feature pyramid networks, MobileNet-SSD has low computational and storage costs while maintaining high accuracy.
[0036] YOLO: YOLO (You Only Look Once) is a real-time object detection system that predicts object bounding boxes and categories through a single neural network.
[0037] FAR: False Acceptance Rate. When performing face recognition, FAR refers to the rate at which the authentication system mistakenly allows an illegal or unauthorized user to pass authentication. FAR = number of false acceptances / total number of illegal user authentications.
[0038] FRR: False Rejection Rate. When performing face recognition, it refers to the proportion of legitimate users (users registered in the system) who are falsely rejected when attempting to pass face recognition verification. FRR = number of false rejections / total number of legitimate user verifications.
[0039] Identification Rate: When performing face recognition, it refers to the proportion of users whose identities are correctly found by the system from the candidate identity database given the facial features of a query user. IR = number of successful identifications / total number of queries.
[0040] PSNR: Peak signal-to-noise ratio. In super-resolution tasks, it indicates the quality of image restoration or the degree of distortion in the image reconstruction process. It measures image quality by the ratio between the maximum possible pixel value of the image and the noise of the reconstruction error.
[0041] F1-SCORE: In the semantic segmentation task, each pixel needs to be classified. F1-score evaluates the classification correctness of each pixel, and ensures that neither false positives nor false negatives are ignored through the harmonic average of precision and recall. Precision refers to the proportion of pixels predicted by the model as positive that actually belong to the positive class. Recall refers to the proportion of all pixels that are actually positive that are correctly predicted by the model as positive.
[0042] Artificial intelligence chips are a new generation of microprocessors specially designed to handle artificial intelligence tasks. They have the advantages of high performance and high energy efficiency and are widely used in the fields of autonomous driving, smart home appliances, robots, etc. With the development of unmanned and intelligent technology, artificial intelligence chips have gradually been applied to the fields of aviation, aerospace and military industry.
[0043] The current testing method is for designers to design and test programs according to requirements. When there are many optional chip manufacturers and specifications, it will lead to problems such as difficult development, heavy workload, long cycle, and high cost.
[0044] Based on the above description, this embodiment proposes a chip performance testing method, such as Figure 1 As shown, the method includes:
[0045] Step 101, obtain the artificial intelligence chip to be tested and the target workload, wherein the target workload includes a test model, a test data set, a model weight numerical accuracy, a machine learning framework, and a batch size.
[0046] Step 102: preprocess the test data set to obtain a target data set.
[0047] Step 103, determine the programming language for the artificial intelligence chip performance test, use the programming language to perform model adaptation processing on the test model, and obtain the target test model.
[0048] Step 104: Use the target test model equipped with the artificial intelligence chip to process the target data set to obtain a model inference result.
[0049] Step 105: Determine the model inference accuracy value corresponding to the target test model based on the model inference result and the target data set.
[0050] Step 106, determining the target test model to obtain power consumption information corresponding to the model inference result, taking the model inference result, the model inference accuracy value and the power consumption information as the test result, and outputting the test result.
[0051] In the specific implementation, the artificial intelligence chip to be tested and the target workload are obtained, and the target workload includes the test model and test data set, the numerical accuracy of the model weight, the machine learning framework and the batch size. The test model is a classic neural network model or a custom neural network model, and the classic neural network model includes at least one of the following: MobileNet, LSTM, ResNet, Mobilenet-SSD, YOLO model, etc.
[0052] The test data set is a public data set or a custom data set corresponding to a real application scenario, and the numerical precision of the model weight includes at least one of the following: 32-bit floating point precision, 16-bit floating point precision, 8-bit integer precision, 4-bit integer precision, 16-bit integer precision or mixed precision.
[0053] The batch size is 2 n , n≥0 and n is an integer, the machine learning framework is a mainstream open source machine learning framework, which is a tool and library that provides support for machine learning (especially deep learning), including Tensorflow, Pytorch, etc.
[0054] The test data set is preprocessed to obtain a target data set, and a programming language for the artificial intelligence chip performance test is determined, where the programming language is Python, C, or C++, etc.
[0055] The programming language is used to perform model adaptation processing on the test model to obtain a target test model. The target test model equipped with the artificial intelligence chip to be tested is used to infer the target data set to obtain a model inference result.
[0056] According to the model reasoning result and the target data set, the model reasoning accuracy value corresponding to the target test model is determined. The target test model is determined to obtain power consumption information corresponding to the model reasoning result, and the model reasoning result, the model reasoning accuracy value and the power consumption information are used as test results, and the test results are output.
[0057] Through the above scheme, the artificial intelligence chip to be tested and the target workload are obtained, wherein the target workload includes a test model, a test data set, a model weight numerical accuracy, a machine learning framework and a batch size, and the test model is a model equipped with the artificial intelligence chip to be tested. The test data set is preprocessed to obtain a target data set. The programming language for the performance test of the artificial intelligence chip is determined, and the test model is adapted to obtain a target test model using the programming language. The target test model equipped with the artificial intelligence chip to be tested is used to infer the target data set to obtain a model inference result. The performance test of the artificial intelligence chip to be tested is realized by running the model equipped with the artificial intelligence chip to be tested. According to the model inference result and the target data set, the model inference accuracy value corresponding to the target test model is determined. The target test model is determined to obtain the power consumption information corresponding to the model inference result, and the model inference result, the model inference accuracy value and the power consumption information are used as the test result, and the test result is output. By building a test software framework for different workloads, the same framework can be used to test different chips, which reduces the development difficulty and improves the efficiency of chip testing. At the same time, using the same framework for testing can also make the test results of different artificial intelligence chips to be tested more fair and comparable.
[0058] In some embodiments, step 102 specifically includes:
[0059] Step 1021, obtaining a preset reasoning requirement, and preprocessing the test data set according to the reasoning requirement to obtain a target data set.
[0060] During specific implementation, a preset reasoning requirement is obtained, and a test data set is preprocessed according to the reasoning requirement, and a target data set is obtained after the preprocessing is completed.
[0061] The preprocessing method includes at least one of the following: adjusting image size, normalization, data enhancement, standardization, cropping, denoising, image smoothing, color space conversion, channel order conversion, etc.
[0062] Among them, adjusting the image size is to adjust the input image to a fixed size to adapt to the input requirements of the network.
[0063] Normalization helps speed up training and increase the convergence rate of the model by scaling pixel values to the range of [0, 1] or [-1, 1].
[0064] Normalization converts each channel of an image (such as RGB) into a distribution with zero mean and unit variance to reduce the scale differences between individual features.
[0065] Cropping is either center cropping or random cropping of the image.
[0066] Denoising is to remove noise from the image so that the model can focus on important features. Common methods include Gaussian filtering, median filtering, etc.
[0067] Image smoothing is to use a filter (such as Gaussian filter) to smooth the image and reduce the impact of noise.
[0068] Color space conversion is usually used when the test data set is image information. Because images are usually three-channel RGB, it is necessary to convert color images into gray images during processing, that is, change three channels into one channel.
[0069] The channel order conversion changes the RGB order of the image information.
[0070] Through the above scheme, by preprocessing the data in the test data set, when the test model is subsequently used to process the data in the test data set, the performance of the model is improved and the sensitivity to model noise and unnecessary information is reduced.
[0071] In some embodiments, step 103 specifically includes:
[0072] Step 1031, determining the adaptation model file format corresponding to the artificial intelligence chip to be tested, wherein the adaptation model file format is a model file format that allows the artificial intelligence chip to be tested to compile;
[0073] Step 1032, obtaining an initial file format of the test model corresponding to the target workload, and comparing the initial file format with the adaptation model file format;
[0074] Step 1033, in response to the initial file format being different from the adaptation model file format, the initial file format is compiled and converted using a model compiler corresponding to the programming language to obtain a target file format that is the same as the adaptation model file format, and the test model including the target file format is used as a target test model.
[0075] During specific implementation, the adaptation model file format corresponding to the artificial intelligence chip to be tested is determined, wherein the adaptation model file format is a model file format that allows the artificial intelligence chip to run, and the adaptation model file format includes the model file format, hyperparameter configuration of the chip adaptation model, weights, bias initialization and weight format.
[0076] An initial model file format of a test model corresponding to a target workload is obtained, and the initial model file format is compared with the adaptation model file format.
[0077] If the initial model file format is different from the adapted model file format, it means that the model file format of the model is not compatible with the artificial intelligence chip to be tested. A model compiler corresponding to the programming language is determined according to the programming language. The model compiler is used to convert the format of the initial model file format of the test model.
[0078] The format of the converted model file is the same as the model file format adapted for the artificial intelligence chip to be tested. At this time, the test model can be used to carry the artificial intelligence chip to be tested, and the target test model equipped with the artificial intelligence chip to be tested can be used to infer the test data set in the target workload to obtain the model inference result.
[0079] In some embodiments, step 105 specifically includes:
[0080] Step 1051, determining the model type corresponding to the target test model;
[0081] Step 1052, according to the model type, determining the model reasoning accuracy corresponding to the model type;
[0082] Step 1053: Determine a model reasoning accuracy value corresponding to the model reasoning accuracy according to the model reasoning result and the target data set.
[0083] During specific implementation, the model type corresponding to the target test model is determined, and the model type indicates the type of reasoning effect achieved by the target test model, such as image classification, image generation, big language, target detection, semantic segmentation, face recognition or super-resolution.
[0084] According to the model type, determine the model reasoning accuracy corresponding to the model type, where the model reasoning accuracy is an indicator representing the model reasoning accuracy. Exemplarily, the model reasoning accuracy includes at least one of the following: Top-1 accuracy, Top-5 accuracy, mean average precision (mAP), intersection-over-union ratio, F1-SCORE, etc.
[0085] Specifically, if the model type is image classification, the model inference accuracy is Top-1 accuracy and / or Top-5 accuracy. If the model type is object detection, the model inference accuracy is mean average precision (mAP). If the model type is semantic segmentation, the model inference accuracy is intersection-over-union ratio and / or F1-SCORE. If the model type is face recognition, the model inference accuracy is FAR, FRR, Identification Rate. If the model type is super-resolution, the model inference accuracy is PSNR.
[0086] Exemplarily, the Top-1 accuracy and Top-5 accuracy are the model inference accuracy corresponding to the classification model. Exemplarily, if the classification model determines the category corresponding to the object in the image based on the image, the Top-1 accuracy indicates that the probability of determining the category of the object using the classification model is 70% for cats, 10% for dogs, 2% for pigs, 3% for bears, 4% for mice, 2% for cows, 5% for sheep, and 4% for rabbits. Then the Top-1 is a cat. If the actual category of the object is a cat, the Top-1 accuracy is 100%.
[0087] Another example, the Top-5 accuracy rate means: the probability of determining the category of an object using the classification model is 70% for cat, 10% for dog, 2% for pig, 3% for bear, 1% for mouse, 2% for cow, 5% for sheep, and 7% for rabbit. The Top-5 are cat, dog, rabbit, sheep, and bear respectively. If the actual category of the object is cat, the Top-5 accuracy rate is 100%.
[0088] In some embodiments, step 104 specifically includes:
[0089] Step 1041, using the target test model equipped with the artificial intelligence chip to be tested, inferring the target data set to obtain a model inference result;
[0090] Step 1042, recording the inference time required to obtain the model inference result, and using the inference time as the inference delay of the artificial intelligence chip to be tested;
[0091] Step 1043, obtaining the amount of test data in the target data set;
[0092] Step 1044, performing ratio processing on the test data volume and the inference delay to obtain an average forward inference rate of the artificial intelligence chip to be tested.
[0093] During specific implementation, the target test model of the target workload equipped with the artificial intelligence chip to be tested is used to infer the target data set corresponding to the test data set in the target workload to obtain a model inference result.
[0094] Exemplarily, the test model is a classification model, and the classification model equipped with the artificial intelligence chip to be tested is used to classify the data in the target data set to obtain a classification result.
[0095] The timing starts from when the data in the target data set is input into the target test model, and stops when the model inference result is obtained. The time required for the target test model to receive the target data set and obtain the model inference result can be obtained, and the inference time is used as the inference delay of the artificial intelligence chip to be tested.
[0096] The target data volume in the target data set is obtained, where the target data volume represents the total amount of data contained in the target data set. The ratio between the target data volume and the determined inference delay is calculated, and the ratio is used as the average forward inference rate of the artificial intelligence chip to be tested, that is, the number of input samples that the model can process within a unit time at the average forward inference rate.
[0097] In some embodiments, step 106 specifically includes:
[0098] Step 1061, acquiring the target test model to obtain power consumption information corresponding to the model inference result, and using the power consumption information as the inference power consumption of the artificial intelligence chip to be tested;
[0099] Step 1062, performing ratio processing on the average forward reasoning rate and the reasoning power consumption to obtain the energy efficiency ratio of the artificial intelligence chip to be tested.
[0100] During specific implementation, the test model is obtained to process the target data set, and the power consumption information required for the model inference result is obtained, and the power consumption information is used as the inference power consumption of the artificial intelligence chip to be tested.
[0101] The average forward inference rate obtained is ratioed to the inference power consumption to obtain the energy efficiency ratio of the artificial intelligence chip to be tested. The energy efficiency ratio indicates the computing performance that can be achieved per unit energy consumption when performing artificial intelligence chip inference tasks.
[0102] Based on the same inventive concept, another embodiment of the present disclosure provides an artificial intelligence chip performance test method, wherein the programming language of the artificial intelligence chip performance test is C++ language, such as Figure 2 As shown, the method specifically includes:
[0103] Initialize the initialization module, create model description information and initialize the chip device.
[0104] Obtain the artificial intelligence chip to be tested and the target workload, wherein the target workload includes a test model, a test data set, the numerical accuracy of the model weights, a machine learning framework, and a batch size.
[0105] Get the preset reasoning requirements, use dataset_preprocessing.cpp to preprocess the test dataset according to the reasoning requirements, and obtain the target dataset.
[0106] Model_adaptation is used to perform model adaptation processing on the test model, specifically preparing a model weight file, saving it as a model compiler compilable file, designing model operation parameters, and writing a model compiler configuration file. The model compiler corresponding to the programming language is used to compile and convert the initial file format to obtain a target file format that is the same as the adaptation model file format, and the test model including the target file format is used as the target test model.
[0107] The target test model equipped with the artificial intelligence chip to be tested is used to infer the target data set using Model_inference to obtain a model inference result. Specifically:
[0108] Using the target test model equipped with the artificial intelligence chip to be tested, reasoning the target data set to obtain a model reasoning result;
[0109] The inference time required to obtain the model inference result is recorded, and the inference time is used as the inference delay of the artificial intelligence chip to be tested. The amount of test data in the target data set is obtained, and the test data amount and the inference delay are ratio-processed to obtain the average forward inference rate of the artificial intelligence chip to be tested.
[0110] Use Compute_metrics_ModelName.cpp to perform model post-processing operations, determine the model type corresponding to the target test model, and determine the model inference accuracy corresponding to the model type based on the model type. Determine the model inference accuracy value corresponding to the model inference accuracy based on the model inference result and the target data set. The model inference accuracy includes at least one of the following: Top-1 accuracy, Top-5 accuracy, mean average precision (mAP), intersection-over-union ratio, F1-SCORE, etc.
[0111] Use Power_detect to perform power consumption testing, obtain the target test model to obtain the power consumption information corresponding to the model inference result, and use the power consumption information as the inference power consumption of the artificial intelligence chip to be tested. Ratio the average forward inference rate to the inference power consumption to obtain the energy efficiency ratio of the artificial intelligence chip to be tested.
[0112] Based on the same inventive concept, another embodiment of the present disclosure provides an artificial intelligence chip performance testing method, wherein the programming language of the artificial intelligence chip performance testing is Python, such as Figure 3 As shown, the method specifically includes:
[0113] Initialize resources, create model description information, and prepare input and output data structures for reasoning.
[0114] Obtain the artificial intelligence chip to be tested and the target workload, wherein the target workload includes a test model, a test data set, the numerical accuracy of the model weights, a machine learning framework, and a batch size.
[0115] Build a test data set, perform image scaling, normalization, convert color space format, etc. Build a loader for batch loading data, perform random shuffling, multi-process loading, etc.
[0116] Perform model inference and use the model conversion tool to convert the original networks of different frameworks into offline models suitable for the AI chip to be tested based on the document parameters. Load the converted model through the interface, perform model inference, and obtain the model inference result.
[0117] The inference time required to obtain the model inference result is recorded, and the inference time is used as the inference delay of the artificial intelligence chip to be tested. The amount of test data in the target data set is obtained, and the test data amount and the inference delay are ratio-processed to obtain the average forward inference rate of the artificial intelligence chip to be tested.
[0118] Perform data post-processing to determine the model type corresponding to the target test model, and determine the model inference accuracy corresponding to the model type based on the model type. Determine the model inference accuracy value corresponding to the model inference accuracy based on the model inference result and the target data set. The model inference accuracy includes at least one of the following: Top-1 accuracy, Top-5 accuracy, mean average precision (mAP), intersection-over-union ratio, F1-SCORE, etc.
[0119] Perform a power consumption test, obtain the target test model to obtain the power consumption information corresponding to the model inference result, and use the power consumption information as the inference power consumption of the artificial intelligence chip to be tested. Ratio the average forward inference rate to the inference power consumption to obtain the energy efficiency ratio of the artificial intelligence chip to be tested.
[0120] The specific process is as follows:
[0121] Given a test scenario, a benchmark dataset, a specified model, a batch size, and a machine learning framework, such as PyTorch, TensorFlow, etc.
[0122] Use DatasetName_preprocessing.py to build the Dataset class of the dataset and preprocess the data, such as scaling the image to the size required by the model, performing mean and normalization processing, changing channel information, data format, etc.; build a data loader Dataloader for batch loading data and batch load data from the dataset.
[0123] Based on the selected network structure and model hyperparameter configuration, weight and bias initialization, and weight format, use the model conversion tool to convert the original network models of different frameworks to adapt them to the offline models of different artificial intelligence processors, and then use the interface to load the model;
[0124] Start timing. Read inference input data and perform model inference. End timing, return forward inference latency, and calculate average forward inference rate and throughput performance.
[0125] Use Compute_metrics_ModelName.py to post-process the obtained inference results and calculate the corresponding test indicators for the test task scenario, such as Top-1, Top-5 accuracy, mAP, mIOU, F1-SCORE, etc.;
[0126] Use Detect_power.c to detect the power consumption information of inference, and use the returned power consumption value and the average forward inference rate to calculate the energy efficiency ratio. Output the test results.
[0127] It should be noted that the method of the embodiment of the present disclosure can be performed by a single device, such as a computer or a server. The method of the present embodiment can also be applied in a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only perform one or more steps in the method of the embodiment of the present disclosure, and the multiple devices will interact with each other to complete the described method.
[0128] It should be noted that the above describes some embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the above embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0129] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present disclosure also provides an artificial intelligence chip performance testing benchmark framework.
[0130] refer to Figure 4 , Figure 4 The artificial intelligence chip performance test benchmark framework of the embodiment includes:
[0131] The data acquisition module 401 is configured to acquire the artificial intelligence chip to be tested and the target workload, wherein the target workload includes a test model, a test data set, a model weight numerical accuracy, a machine learning framework, and a batch size;
[0132] The data preprocessing module 402 is configured to preprocess the test data set to obtain a target data set;
[0133] The model adaptation module 403 is configured to determine a programming language for the artificial intelligence chip performance test, and use the programming language to perform model adaptation processing on the test model to obtain a target test model;
[0134] The reasoning module 404 is configured to use the target test model equipped with the artificial intelligence chip to perform reasoning on the target data set to obtain a model reasoning result;
[0135] A model reasoning accuracy determination module 405 is configured to determine a model reasoning accuracy value corresponding to the target test model according to the model reasoning result and the target data set;
[0136] The output module 406 is configured to determine the target test model to obtain power consumption information corresponding to the model reasoning result, take the model reasoning result, the model reasoning accuracy value and the power consumption information as the test result, and output the test result.
[0137] In some embodiments, the data preprocessing module 402 is configured to:
[0138] Obtain preset reasoning requirements, preprocess the test data set according to the reasoning requirements, and obtain the target data set.
[0139] In some embodiments, the model adaptation module 403 is specifically configured to:
[0140] Determine the adaptation model file format corresponding to the artificial intelligence chip to be tested, wherein the adaptation model file format is a model file format that allows the artificial intelligence chip to be tested to compile;
[0141] Obtaining an initial file format of a test model corresponding to the target workload, and comparing the initial file format with the adaptation model file format;
[0142] In response to the initial file format being different from the adaptation model file format, the initial file format is compiled and converted using a model compiler corresponding to the programming language to obtain a target file format that is the same as the adaptation model file format, and the test model including the target file format is used as the target test model.
[0143] In some embodiments, the model reasoning accuracy determination module 405 is specifically configured to:
[0144] Determine the model type corresponding to the target test model;
[0145] According to the model type, determining the model reasoning accuracy corresponding to the model type;
[0146] According to the model inference result and the target data set, a model inference accuracy value corresponding to the model inference accuracy is determined.
[0147] In some embodiments, the reasoning module 404 is specifically configured to:
[0148] Using the target test model equipped with the artificial intelligence chip to be tested, reasoning the target data set to obtain a model reasoning result;
[0149] Record the inference time required to obtain the model inference result, and use the inference time as the inference latency of the artificial intelligence chip to be tested;
[0150] Obtaining the amount of test data in the target data set;
[0151] The test data volume and the inference delay are ratio-processed to obtain the average forward inference rate of the artificial intelligence chip to be tested.
[0152] In some embodiments, the output module 406 is specifically configured to:
[0153] Acquire the target test model to obtain power consumption information corresponding to the model inference result, and use the power consumption information as the inference power consumption of the artificial intelligence chip to be tested;
[0154] The average forward reasoning rate is ratioed to the reasoning power consumption to obtain the energy efficiency ratio of the artificial intelligence chip to be tested.
[0155] In some embodiments, in the device, the test model is a classical neural network model or a custom neural network model;
[0156] The test data set is a public data set or a custom data set corresponding to a real application scenario;
[0157] The numerical precision of the model weight includes at least one of the following: 32-bit floating point precision, 16-bit floating point precision, 8-bit integer precision, 4-bit integer precision, 16-bit integer precision or mixed precision;
[0158] The batch size is 2 n , n ≥ 0 and n is an integer;
[0159] The programming language is Python, C or C++;
[0160] The machine learning framework is a mainstream open source machine learning framework.
[0161] For the convenience of description, the above device is described by dividing it into various modules according to its functions. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0162] The device of the above embodiment is used to implement the corresponding chip performance testing method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.
[0163] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the artificial intelligence chip performance testing method described in any of the above embodiments is implemented.
[0164] Figure 5 A more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment is shown, and the device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are connected to each other through the bus 1050 in the device.
[0165] The processor 1010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0166] The memory 1020 may be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0167] The input / output interface 1030 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0168] The communication interface 1040 is used to connect a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).
[0169] The bus 1050 includes a path that transmits information between the various components of the device (eg, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).
[0170] It should be noted that, although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040 and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiments of the present specification, and does not necessarily include all the components shown in the figure.
[0171] The electronic device of the above-mentioned embodiment is used to implement the corresponding artificial intelligence chip performance testing method in any of the above-mentioned embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0172] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present disclosure also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the artificial intelligence chip performance testing method described in any of the above embodiments.
[0173] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0174] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the artificial intelligence chip performance testing method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0175] It is understandable that before using the technical solutions of each embodiment of the present disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner, and the user's authorization will be obtained.
[0176] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly remind the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can independently choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.
[0177] As an optional but non-limiting implementation, in response to receiving the user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0178] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0179] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples. Based on the concept of the present disclosure, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of simplicity.
[0180] In addition, to simplify the description and discussion, and in order not to make the embodiments of the present disclosure difficult to understand, the known power / ground connections to the integrated circuit (IC) chips and other components may or may not be shown in the provided figures. In addition, the device can be shown in the form of a block diagram to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure will be implemented (that is, these details should be fully within the scope of understanding of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present disclosure, it is apparent to those skilled in the art that the embodiments of the present disclosure can be implemented without these specific details or with changes in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0181] Although the present disclosure has been described in conjunction with specific embodiments of the present disclosure, many replacements, modifications and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.
[0182] The embodiments of the present disclosure are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present disclosure should be included in the scope of protection of the present disclosure.
Claims
1. A method for testing the performance of an artificial intelligence chip, characterized in that: include: Obtaining the artificial intelligence chip to be tested and the target workload, wherein the target workload includes a test model, a test data set, model weight numerical accuracy, a machine learning framework, and a batch size; Preprocessing the test data set to obtain a target data set; Determine a programming language for the performance test of the artificial intelligence chip, and use the programming language to perform model adaptation processing on the test model to obtain a target test model; Using the target test model equipped with the artificial intelligence chip to be tested, reasoning the target data set to obtain a model reasoning result; Determining a model reasoning accuracy value corresponding to the target test model according to the model reasoning result and the target data set; Determine the target test model to obtain power consumption information corresponding to the model reasoning result, use the model reasoning result, the model reasoning accuracy value and the power consumption information as the test result, and output the test result.
2. The method according to claim 1, characterized in that The preprocessing of the test data set to obtain the target data set includes: Obtain preset reasoning requirements, preprocess the test data set according to the reasoning requirements, and obtain the target data set.
3. The method according to claim 1, characterized in that The using the programming language to perform model adaptation processing on the test model to obtain a target test model includes: Determine the adaptation model file format corresponding to the artificial intelligence chip to be tested, wherein the adaptation model file format is a model file format that allows the artificial intelligence chip to be tested to compile; Obtaining an initial file format of a test model corresponding to the target workload, and comparing the initial file format with the adaptation model file format; In response to the initial file format being different from the adaptation model file format, the initial file format is compiled and converted using a model compiler corresponding to the programming language to obtain a target file format that is the same as the adaptation model file format, and the test model including the target file format is used as the target test model.
4. The method according to claim 1, characterized in that: Determining the model inference accuracy value corresponding to the target test model according to the model inference result and the target data set includes: Determine the model type corresponding to the target test model; According to the model type, determining the model reasoning accuracy corresponding to the model type; According to the model inference result and the target data set, a model inference accuracy value corresponding to the model inference accuracy is determined.
5. The method according to claim 1, characterized in that The method of using the target test model equipped with the artificial intelligence chip to be tested to infer the target data set to obtain a model inference result includes: Using the target test model equipped with the artificial intelligence chip to be tested, reasoning the target data set to obtain a model reasoning result; Record the inference time required to obtain the model inference result, and use the inference time as the inference latency of the artificial intelligence chip to be tested; Obtaining the amount of test data in the target data set; The test data volume and the inference delay are ratio-processed to obtain the average forward inference rate of the artificial intelligence chip to be tested.
6. The method according to claim 5, characterized in that The determining the target test model to obtain power consumption information corresponding to the model inference result includes: Acquire the target test model to obtain power consumption information corresponding to the model inference result, and use the power consumption information as the inference power consumption of the artificial intelligence chip to be tested; The average forward reasoning rate is ratioed to the reasoning power consumption to obtain the energy efficiency ratio of the artificial intelligence chip to be tested.
7. The method according to claim 1, characterized in that The test model is a classic neural network model or a custom neural network model; The test data set is a public data set or a custom data set corresponding to a real application scenario; The numerical precision of the model weight includes at least one of the following: 32-bit floating point precision, 16-bit floating point precision, 8-bit integer precision, 4-bit integer precision, 16-bit integer precision or mixed precision; The batch size is 2 n , n ≥ 0 and n is an integer; The programming language is Python, C or C++; The machine learning framework is a mainstream open source machine learning framework.
8. An artificial intelligence chip performance test benchmark framework, characterized in that: include: A data acquisition module is configured to acquire the artificial intelligence chip to be tested and the target workload, wherein the target workload includes a test model, a test data set, a model weight numerical accuracy, a machine learning framework, and a batch size; A data preprocessing module is configured to preprocess the test data set to obtain a target data set; A model adaptation module is configured to determine a programming language for the artificial intelligence chip performance test, and use the programming language to perform model adaptation processing on the test model to obtain a target test model; An inference module is configured to use the target test model equipped with the artificial intelligence chip to perform inference on the target data set to obtain a model inference result; A model reasoning accuracy determination module is configured to determine a model reasoning accuracy value corresponding to the target test model according to the model reasoning result and the target data set; The output module is configured to determine the target test model to obtain power consumption information corresponding to the model reasoning result, take the model reasoning result, the model reasoning accuracy value and the power consumption information as the test result, and output the test result.
9. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 8 is implemented.
10. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute the method according to any one of claims 1 to 8.
Citation Information
Cited By
Processor test method and device
CN121524025A