AI chip-based quantitative evaluation methods, devices, equipment, media, and programs

By conducting multi-dimensional evaluation of the chip hardware features and operator-level object features of AI chips, a quantization strategy is generated, which solves the problem of insufficient refinement in the quantization evaluation of AI chips in the existing technology, improves the quantization accuracy and operating efficiency, and optimizes the chip performance and the competitiveness of terminal devices.

CN121167227BActive Publication Date: 2026-03-06SHANGHAI SUIYUAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511725468.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-24
Publication Date
2026-03-06
Estimated Expiration
2045-11-24

AI Technical Summary

Technical Problem

Existing quantization methods for AI chips lack the ability to perform refined evaluations of individual computing nodes or operators, fail to fully consider the characteristics of different AI chip architectures, resulting in significant discrepancies between quantization evaluation results and actual deployment performance. Furthermore, these methods primarily target static quantization scenarios and lack support for dynamic and adaptive quantization.

Method used

By extracting the chip hardware features of the target AI chip and the features of the target operator-level objects, a multi-dimensional evaluation is performed to generate a quantization strategy for the target operator-level objects, thereby improving the refinement and accuracy of the quantization evaluation.

Benefits of technology

It has improved the quantization accuracy and operating efficiency of AI chips, optimized the computing performance of chips, shortened the product evaluation cycle, and enhanced the competitiveness of terminal devices and the real-time performance and reliability of autonomous driving systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121167227B_ABST
    Figure CN121167227B_ABST
Patent Text Reader

Abstract

This invention discloses a quantitative evaluation method, apparatus, device, medium, and program based on AI chips. The method includes: extracting the chip hardware features of the target AI chip based on its chip configuration information; extracting target operator features from target operator-level objects in the target AI chip; performing a multi-dimensional evaluation of the quantization method of the target operator-level objects based on the chip hardware features of the target AI chip and the target operator features of the target operator-level objects, obtaining a quantitative evaluation result; and generating a target quantization strategy for the target operator-level objects based on the quantitative evaluation result. The technical solution of this invention can improve the refinement and accuracy of AI chip quantitative evaluation, thereby enhancing the quantization precision and operational efficiency of AI chips.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the fields of chip and artificial intelligence technology, and in particular to a quantitative evaluation method, apparatus, electronic device, storage medium and program based on AI chip. Background Technology

[0002] AI (Artificial Intelligence) chip quantization refers to the technology of reducing the chip's storage space and computing resource consumption by lowering the accuracy of model parameters within the chip, thereby improving the chip's operating efficiency.

[0003] In the process of AI chip quantization, the parameters in the model, originally represented by high-precision floating-point numbers (such as 32-bit FP32), are converted to low-precision formats (such as INT8 or FP16) to reduce the model size and speed up computation. For example, the parameters of a large model that originally required a lot of storage can be compressed from tens of billions of parameters to a more compact representation.

[0004] In the process of developing this invention, the inventors discovered the following shortcomings in existing technologies: Existing AI chip quantization methods primarily evaluate the entire model or network layer within the chip, lacking the ability to fine-tune the evaluation of individual computing nodes or operators. Furthermore, existing AI chip quantization methods fail to fully consider the characteristics of different AI chip architectures, including instruction sets, memory access patterns, and computing unit configurations, resulting in insufficient hardware adaptability and significant deviations between quantization evaluation results and actual deployment performance. Currently, there is a lack of effective methods to evaluate the quantization relationship between quantization results (such as accuracy and performance), making it difficult to find the optimal balance between multi-dimensional quantization results. Additionally, existing AI chip quantization methods mainly target static quantization scenarios, lacking sufficient support for advanced quantization techniques such as dynamic quantization and adaptive quantization. Summary of the Invention

[0005] This invention provides a quantitative evaluation method, apparatus, electronic device, storage medium, and program based on AI chips, which can improve the precision and accuracy of chip quantitative evaluation, thereby enhancing the quantitative accuracy and operating efficiency of AI chips.

[0006] According to one aspect of the present invention, a quantitative evaluation method based on an AI chip is provided, comprising:

[0007] Extract the chip hardware features of the target AI chip based on the chip configuration information of the target AI chip;

[0008] Extract target operator features from the target operator-level objects in the target AI chip;

[0009] The quantization method of the target operator-level object is evaluated from multiple dimensions based on the chip hardware characteristics of the target AI chip and the target operator characteristics of the target operator-level object, and the quantization evaluation result is obtained.

[0010] The target quantization strategy for the target operator-level object is generated based on the quantization evaluation results.

[0011] According to another aspect of the present invention, a quantitative evaluation device based on an AI chip is provided, comprising:

[0012] A chip hardware feature extraction module is used to extract the chip hardware features of the target AI chip based on the chip configuration information of the target AI chip.

[0013] The target operator feature extraction module is used to extract target operator features from target operator-level objects in the target AI chip;

[0014] The quantization evaluation module is used to perform multi-dimensional evaluation of the quantization method of the target operator-level object based on the chip hardware characteristics of the target AI chip and the target operator characteristics of the target operator-level object, and obtain the quantization evaluation result.

[0015] The target quantization strategy generation module is used to generate the target quantization strategy for the target operator-level object based on the quantization evaluation results.

[0016] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0017] At least one processor; and

[0018] A memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the AI ​​chip-based quantitative evaluation method according to any embodiment of the present invention.

[0020] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the AI ​​chip-based quantitative evaluation method according to any embodiment of the present invention.

[0021] According to another aspect of the present invention, a computer program product is also provided, comprising a computer program that, when executed by a processor, implements the AI ​​chip-based quantitative evaluation method described in any embodiment of the present invention.

[0022] This invention extracts the chip hardware features of a target AI chip based on its chip configuration information and extracts target operator features from target operator-level objects within the target AI chip. It then performs a multi-dimensional evaluation of the quantization method of the target operator-level objects based on both the chip hardware features and the target operator features, obtaining a quantization evaluation result. Based on this result, a target quantization strategy for the target operator-level objects is generated. This addresses the problem that existing AI chip quantization evaluation methods cannot quantize operator-level objects within AI chips, improving the precision and accuracy of chip quantization evaluation, thereby enhancing the quantization accuracy and operational efficiency of AI chips.

[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a flowchart of a quantitative evaluation method based on an AI chip provided in Embodiment 1 of the present invention;

[0026] Figure 2 This is a flowchart of a quantitative evaluation method based on an AI chip provided in Embodiment 2 of the present invention;

[0027] Figure 3 This is a schematic diagram of a process for extracting hardware features from a target AI chip according to Embodiment 2 of the present invention;

[0028] Figure 4 This is a schematic diagram of a process for extracting target operator features from a target operator-level object, provided in Embodiment 2 of the present invention;

[0029] Figure 5 This is a flowchart illustrating a multi-dimensional evaluation of the quantization method for target operator-level objects provided in Embodiment 2 of the present invention;

[0030] Figure 6 This is a flowchart illustrating a target quantization strategy for generating target operator-level objects, provided in Embodiment 2 of the present invention.

[0031] Figure 7This is a schematic diagram of the framework structure of a quantitative evaluation system based on an AI chip provided in Embodiment 2 of the present invention;

[0032] Figure 8 This is a schematic diagram of a quantitative evaluation device based on an AI chip provided in Embodiment 3 of the present invention;

[0033] Figure 9 This is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of the present invention. Detailed Implementation

[0034] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0035] It should be noted that the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such process, method, product or device.

[0036] Example 1

[0037] Figure 1 This is a flowchart of a quantitative evaluation method based on an AI chip, provided in Embodiment 1 of the present invention. This embodiment is applicable to the quantitative evaluation of operator-level objects within an AI chip. The method can be executed by an AI chip-based quantitative evaluation device, which can be implemented in software and / or hardware and is generally integrated into an electronic device. This electronic device can be a terminal device or a server device, as long as it can execute the AI ​​chip-based quantitative evaluation method. The present invention does not limit the specific type of electronic device. Correspondingly, as... Figure 1 As shown, the method includes the following operations:

[0038] S110. Extract the chip hardware features of the target AI chip based on the chip configuration information of the target AI chip.

[0039] The target AI chip can be any type of AI chip, including but not limited to CPUs (Central Processing Units), GPUs (Graphics Processing Units), TPUs (Tensor Processing Units), ASICs (Application-Specific Integrated Circuits), FPGAs (Field-Programmable Gate Arrays), and NPUs (Neural Processing Units), as long as they can run AI models and execute AI algorithms. This embodiment of the invention does not limit the type of target AI chip. Chip hardware characteristics refer to the characteristic information of the hardware modules or units of the target AI chip.

[0040] In this embodiment of the invention, for a chip requiring quantitative evaluation, it can be used as the target AI chip. The chip configuration information of the target AI chip is loaded, and this information is parsed to extract its hardware characteristics. It is understood that the hardware characteristics of a chip determine its core functions in an electronic device, including but not limited to data processing, control, storage, and communication. Therefore, by extracting the hardware characteristics of the target AI chip, its functional characteristics can be essentially obtained.

[0041] S120. Extract target operator features from the target operator-level objects in the target AI chip.

[0042] The operator-level object can be an operator-level object type in the target AI chip, such as, but not limited to, nodes, operators, or operations within operators capable of independently performing computational functions in the target AI chip. This embodiment of the invention does not limit the specific type of the operator-level object or the specific computational functions it can perform in the target AI chip. The target operator-level object can be an operator-level object in the target AI chip that needs to be quantified and evaluated. The target operator feature can be a software-type feature extracted from the target operator-level object that reflects its computational characteristics.

[0043] It's understandable that different types of AI chips have different core functionalities. For AI chips, certain important operator-level objects directly affect the overall quantization method. If the quantization method of an AI chip is considered only at the model level, ignoring the sensitivity of specific core operator-level objects to the quantization algorithm, it is highly likely to lead to unreasonable quantization, resulting in a decrease in the model's computational performance. For example, if the applicable quantization method for a certain AI chip is INT4 (Integer 4-bit Quantization) at the model level, but important core operators in this AI chip experience a sharp drop in accuracy after using INT4 quantization, and only INT8 (8-bit Integer Quantization) can maintain ideal computational accuracy, then it can be determined that the actual applicable quantization method for this AI chip is INT8, not INT4. This demonstrates that existing AI chip quantization evaluation methods lack the ability to fine-tune the evaluation of individual computing nodes or operators, resulting in low quantization evaluation accuracy.

[0044] To address the lack of refined evaluation capabilities for individual computing nodes or operators in existing AI chip quantization assessments, when quantizing a target AI chip, the first step is to determine the core operator-level objects within the target AI chip based on its type. It's understood that a target AI chip may contain one or more core operator-level objects. Each core operator-level object within the target AI chip can be treated as a target operator-level object, and the applicable quantization strategy can be determined based on these target operator-level objects. To fully leverage these target operator-level objects to determine the applicable quantization strategy, corresponding operator features can be extracted from the target operator-level objects within the target AI chip as target operator features. These extracted target operator features can then be used as the basis for evaluating and selecting the appropriate quantization strategy for the target AI chip.

[0045] S130. Based on the chip hardware characteristics of the target AI chip and the target operator characteristics of the target operator-level object, the quantization method of the target operator-level object is evaluated in multiple dimensions to obtain the quantization evaluation result.

[0046] After extracting the hardware features of the target AI chip and the target operator features of the target operator-level objects, we can analyze the quantization requirements of the target AI chip at the hardware level by combining its hardware characteristics. Simultaneously, focusing on the target operator features of the target operator-level objects, we can analyze the software characteristics of these objects within the target AI chip. Furthermore, we can identify multiple quantization methods suitable for the target operator-level objects as candidate quantization methods. Then, based on both the hardware and software characteristics of the target AI chip and the target operator-level objects, we can evaluate the quantization performance of each candidate quantization method across multiple dimensional quantization metrics to obtain the evaluation results of the quantization effects of each candidate quantization method.

[0047] S140. Generate the target quantization strategy for the target operator-level object based on the quantization evaluation results.

[0048] Among them, the target quantization strategy can be a quantization strategy for evaluating and screening target AI chips at the operator level, and is used to quantize the target operator level objects of the target AI chips.

[0049] Accordingly, the quantization method with the best quantization effect can be selected from the candidate quantization methods based on the multi-dimensional quantization performance evaluation results of each candidate quantization method for the target operator-level object, and used as the target quantization strategy for the target operator-level object.

[0050] Since the target quantization strategy applicable to target operator-level objects is a quantization strategy that is comprehensively evaluated and selected from multiple dimensions of quantization indicators based on chip hardware characteristics and operator software characteristics, the use of the target quantization strategy to quantize target operator-level objects in the target AI chip can significantly improve the overall computational efficiency of the target AI chip while ensuring the computational accuracy of the target operator-level objects, thereby improving the quantization accuracy and computational efficiency of the target AI chip.

[0051] In a specific example, taking a chip manufacturer's verification scenario, chip manufacturers often need to verify the performance of their self-developed quantization algorithms. They can use their self-developed NPU or ASIC as the target AI chip, identify the target operator-level objects in the target AI chip, and verify the performance of quantization algorithms such as INT8, INT4, and FP16 on the target operator-level objects in the self-developed NPU / ASIC. By comparing the quantization effects of different quantization methods, they can optimize the quantization parameter configuration at the operator level, evaluate the impact of operator-level quantization on the overall chip accuracy and performance, and achieve the technical value of quickly verifying the performance of quantization algorithms and shortening the product evaluation cycle.

[0052] In a specific example, taking the selection scenario of terminal manufacturers as an example, terminal equipment manufacturers need to select the most suitable AI chip and quantization scheme. They need to compare the performance of different AI chips such as ASIC, GPU, NPU, and FPGA. After identifying the target operator-level object in the target AI chip for each of the above types of AI chips, they need to evaluate the quantization effects of various quantization schemes such as INT8, INT4, FP16, mixed precision, PTQ (Post-Training Quantization), AWQ (Activation-aware Weight Quantization), AQ (Adaptive Quantization), and QAT (Quantization Aware Training) on ​​the target operator-level object corresponding to different AI chips. They need to comprehensively consider multiple dimensions of quantization evaluation indicators such as latency, throughput, power consumption, and accuracy, and select the optimal "AI chip + quantization" combination scheme to achieve the technical value of improving product competitiveness and optimizing cost-effectiveness.

[0053] In a specific example, taking the optimization scenario of autonomous driving as an illustration, autonomous driving systems require efficient AI inference. Deploying deep learning models on onboard chips to achieve real-time object detection and path planning allows for the identification of corresponding target operator-level objects by the deep learning models loaded on the AI ​​chips integrated into the autonomous driving system, and the determination of various quantization methods suitable for these target operator-level objects. Furthermore, by using multiple sensor data as model input data and considering the specific computational requirements of various sensor data types, the optimal quantization scheme is obtained by screening various quantization methods suitable for the target operator-level objects. This optimizes model accuracy and inference speed, thereby enhancing the technical value of improving the real-time performance and reliability of the autonomous driving system.

[0054] In a specific example, taking the Industrial Internet of Things (IIoT) scenario as an example, industrial equipment needs to deploy intelligent analysis functions. A target AI chip can be deployed on the industrial equipment, and the quantization strategy for adapting the target operator-level objects in the target AI chip can be determined. This enables the IIoT equipment to perform equipment status monitoring and fault prediction through the target AI chip, thereby optimizing the quantization method of the target AI chip at the operator level to adapt to the industrial environment, supporting edge computing and cloud collaboration, and realizing the technological value of improving the intelligence level of industrial equipment.

[0055] This invention extracts the chip hardware features of a target AI chip based on its chip configuration information and extracts target operator features from target operator-level objects within the target AI chip. It then performs a multi-dimensional evaluation of the quantization method of the target operator-level objects based on both the chip hardware features and the target operator features, obtaining a quantization evaluation result. Based on this result, a target quantization strategy for the target operator-level objects is generated. This addresses the problem that existing AI chip quantization evaluation methods cannot quantize operator-level objects within AI chips, improving the precision and accuracy of chip quantization evaluation, thereby enhancing the quantization accuracy and operational efficiency of AI chips.

[0056] Example 2

[0057] Figure 2 This is a flowchart of a quantitative evaluation method based on an AI chip, provided in Embodiment 2 of the present invention. This embodiment is a specific embodiment based on the above embodiment. In this embodiment, various specific optional implementation methods are given for extracting the chip hardware features of the target AI chip, extracting target operator features from the target operator-level objects in the target AI chip, performing multi-dimensional evaluation of the quantization method of the target operator-level objects, and generating the target quantization strategy of the target operator-level objects. Correspondingly, as... Figure 2 As shown, the method in this embodiment may include:

[0058] S210. Extract the hardware characteristic information of the target AI chip based on the chip configuration information.

[0059] The hardware characteristic information may include at least one of the following: computing unit association information, memory hierarchy association information, instruction set association information, and data precision support association information.

[0060] Specifically, the computing unit association information can be information associated with the computing units of the target AI chip. The memory hierarchy association information can be information associated with the memory of the target AI chip. The data precision support association information can be information associated with the data precision that the target AI chip can support.

[0061] In an optional embodiment of the present invention, the step of extracting the hardware characteristic information of the target AI chip based on the chip configuration information may include: analyzing the number of computing units, parallelism, and frequency characteristics of the target AI chip based on the chip configuration information to obtain the computing unit association information; analyzing the memory bandwidth, cache hierarchy, and memory latency of the target AI chip based on the chip configuration information to obtain the memory hierarchy association information; analyzing the instruction set characteristics, operation support, and vectorized operation characteristics of the target AI chip based on the chip configuration information to obtain the instruction set association information; and analyzing the support information of the target AI chip for various data precisions based on the chip configuration information to obtain the data precision support association information.

[0062] Figure 3 This is a flowchart illustrating the extraction of hardware features from a target AI chip according to Embodiment 2 of the present invention. In a specific example, such as... Figure 3 As shown, when extracting hardware features from a target AI chip, at the computational unit level, information such as the number of computational units, parallelism, and frequency characteristics can be extracted as computational unit association information. At the chip memory level, information such as memory bandwidth, cache hierarchy, and memory latency can be extracted as memory hierarchy association information. At the instruction set level, information such as instruction set characteristics, computational support, and vectorization capabilities (i.e., vectorized computation characteristics) can be extracted as instruction set association information. At the data precision level, information such as the target AI chip's support for various data precisions, including FP32, FP16, INT8, INT4, mixed precision, PTQ, AWQ, AQ, QAT, and other quantization types, can be extracted as data precision support association information to analyze the target AI chip's ability to support different data precisions.

[0063] S220. Generate a hardware characteristic matrix of the target AI chip based on the hardware characteristic information of the target AI chip, and determine the hardware adaptation weight of the hardware characteristic matrix based on the chip characteristics of the target AI chip.

[0064] The hardware characteristic matrix can be a full matrix composed of various hardware parameter information, reflecting the hardware characteristics of the target AI chip in matrix form. Hardware adaptation weights are the weights assigned to each element in the hardware characteristic matrix, representing the importance of the corresponding element.

[0065] S230. Generate the chip hardware features of the target AI chip based on the hardware feature matrix of the target AI chip and the hardware adaptation weights of the hardware feature matrix.

[0066] After extracting the hardware characteristic information of the target AI chip, a hardware characteristic matrix can be generated based on this information. This matrix reflects the chip's capabilities in areas such as computational unit analysis, memory hierarchy analysis, instruction set analysis, and data precision support analysis. Simultaneously, weights can be assigned to hardware-related parameters such as computational units, memory hierarchy, instruction sets, and data precision support based on the chip's characteristics. This indicates the chip's importance in computation, memory, instructions, and data, clarifying the appropriate quantization method for its adaptation. Optionally, since chips typically prioritize computational and memory performance, to simplify the evaluation process, corresponding hardware adaptation weights can be configured primarily for the computational and memory-related parameters of the target AI chip. Accordingly, the hardware characteristic matrix with configured hardware adaptation weights can be output as the chip hardware features of the target AI chip.

[0067] S240. Extract the multidimensional operator features of the target operator-level object based on the operator computation graph of the target operator-level object.

[0068] The multidimensional operator features may include at least one of computation mode, data flow mode, and computational complexity.

[0069] S250. Classify the operator types of the target operator-level object according to the multidimensional operator features of the target operator-level object to obtain the target operator type.

[0070] The target operator type can be the operator type determined by classifying the operator types of the target operator level object.

[0071] Figure 4 This is a flowchart illustrating the extraction of target operator features from a target operator-level object according to Embodiment 2 of the present invention. In a specific example, such as... Figure 4 As shown, a matching operator computation graph can be generated for a target operator-level object. An operator computation graph is a directed acyclic graph used in deep learning frameworks to abstract and optimize neural network computations. It consists of operators (operation units) and tensors (data), representing the computational flow from input to output. If the target operator-level object is a new operator, a new node, or a new operator operation, its computation graph can be defined based on its computational logic. Furthermore, a deep analysis of the operator computation graph of the target operator-level object is performed to extract multi-dimensional operator features such as computation patterns, data flow patterns, and computational complexity.

[0072] like Figure 4As shown, the computational patterns of target operator-level objects can be analyzed based on their operator computation graphs, including but not limited to convolution, matrix multiplication, activation function, pooling, and other patterns. Simultaneously, the data flow patterns of target operator-level objects can be analyzed based on their operator computation graphs, including but not limited to data dependency patterns, memory access patterns, and parallelism patterns. Furthermore, the computational complexity of target operator-level objects can be analyzed based on their operator computation graphs, including but not limited to time complexity, space complexity, and computational intensity.

[0073] In an optional embodiment of the present invention, classifying the operator type of the target operator-level object according to the multidimensional operator features of the target operator-level object to obtain the target operator type may include: calculating the computational complexity and computational data flow of the target operator-level object according to the multidimensional operator features of the target operator-level object; classifying the operator type of the target operator-level object according to the computational complexity and computational data flow of the target operator-level object to obtain the target operator type.

[0074] Understandably, for operator-level objects, computational complexity and computational data flow are the two most important factors affecting operator performance. Therefore, specifically, the computational complexity and computational data flow of the target operator-level object can be calculated based on its multidimensional operator characteristics. This allows for the classification of the target operator-level object's operator type based on its computational complexity and computational data flow. For example, if the target operator-level object has high computational complexity, it can be classified as a computational operator. If the target operator-level object has relatively simple computational complexity but a large data flow during computation, it can be classified as a dataflow operator.

[0075] Taking convolution and matrix multiplication operators as examples, the computational complexity of the convolution operator is... It can be modeled as: Where C is the number of output channels, I is the number of input channels, K is the kernel size, and H and W are the output feature map sizes. For the matrix multiplication operator, its computational complexity is... It can be modeled as: , where M and N are the dimensions of the output matrix, and P is the common dimension.

[0076] S260. Determine the feature quantization mode adapted to the target operator-level object based on the target operator type of the target operator-level object, and generate the target operator feature based on the target operator type and the adapted feature quantization mode of the target operator-level object.

[0077] The computational pattern of a target operator-level object primarily reflects its computational characteristics, including but not limited to the number of 1D and 2D computations. The data flow of a target operator-level object reflects its data transport characteristics, indicating data transport requirements, such as the required bandwidth. For example... Figure 4 As shown, after classifying the operator types of the target operator-level object to obtain the target operator type, the feature quantization mode adapted to the target operator-level object can be determined according to the target operator type of the target operator-level object. Then, the feature information of the target operator-level object in terms of operators is quantized and generated into the corresponding target operator features according to the target operator type and the adapted feature quantization mode of the target operator-level object.

[0078] The target operator features can include operator computation features, data transfer features, and post-quantization time-consuming computation features. Post-quantization time-consuming computation features can provide information such as the time consumption percentage and computational efficiency of the operator under different quantization modes. These three types of operator features enable comprehensive quantization of the operator characteristics of the target operator-level object.

[0079] S270. Based on at least one of the chip hardware characteristics of the target AI chip, the target operator characteristics of the target operator-level object, and the operator quantization index, perform accuracy evaluation, performance evaluation, and power consumption evaluation on the quantization method of the target operator-level object to obtain initial multi-dimensional evaluation results.

[0080] The quantification metrics for operators may include, but are not limited to, PSNR (Peak signal-to-noise ratio), SSIM (Structural Similarity Index Measure), and classification accuracy. The initial multi-dimensional evaluation results can be preliminary assessments of the target operator-level object from multiple dimensions.

[0081] Optionally, the chip hardware characteristics of the target AI chip, the target operator characteristics of the target operator-level object, and the operator quantization indicators can be comprehensively analyzed. The various quantization methods that can be used for the target operator-level object can be comprehensively evaluated from the dimensions of accuracy, performance, and power consumption to obtain the initial multi-dimensional evaluation results.

[0082] In an optional embodiment of the present invention, the step of performing accuracy evaluation, performance evaluation, and power consumption evaluation on the quantization method of the target operator-level object based on at least one of the chip hardware characteristics of the target AI chip, the target operator characteristics of the target operator-level object, and operator quantization indicators to obtain an initial multi-dimensional evaluation result may include: performing accuracy evaluation on the quantization method of the target operator-level object based on the operator quantization indicators to obtain an accuracy evaluation result; wherein, the operator quantization indicators include peak signal-to-noise ratio, structural similarity, and classification accuracy; performing performance evaluation on the quantization method of the target operator-level object based on the chip hardware characteristics of the target AI chip to obtain a performance evaluation result; and performing power consumption evaluation on the quantization method of the target operator-level object based on the chip hardware characteristics of the target AI chip and the target operator characteristics of the target operator-level object to obtain a power consumption evaluation result.

[0083] Figure 5 This is a flowchart illustrating a multi-dimensional evaluation of a quantization method for a target operator-level object, as provided in Embodiment 2 of the present invention. In a specific example, such as... Figure 5 As shown, when evaluating the quantization methods of target operator-level objects from multiple dimensions, we can calculate operator quantization indicators such as peak signal-to-noise ratio, structural similarity, and classification accuracy after quantization using the corresponding quantization method. Combined with the loss function used in the quantization algorithm, we can evaluate the impact of each quantization method on the quantization accuracy of the target operator-level object, obtaining the accuracy evaluation results. Simultaneously, based on the chip hardware characteristics of the target AI chip, we can evaluate the performance of the quantization methods of the target operator-level object in dimensions such as computation time, memory bandwidth, computing unit utilization, and throughput, obtaining the performance evaluation results of each quantization method. Furthermore, based on the chip hardware characteristics of the target AI chip and the target operator characteristics of the target operator-level object, we can calculate various power consumption information such as dynamic power consumption, static power consumption, energy efficiency ratio, and power consumption distribution of hardware modules or the entire chip, obtaining the power consumption evaluation results of each quantization method. Optionally, in the power consumption evaluation process, dynamic power consumption and static power consumption can be obtained by measuring the actual power consumption of the target operator-level object in the target AI chip, while the energy efficiency ratio can be calculated theoretically.

[0084] In an optional embodiment of the present invention, the performance evaluation of the quantization method of the target operator-level object based on the chip hardware characteristics of the target AI chip to obtain the performance evaluation result may include: calculating the computation time of the quantization method of the target operator-level object based on the number of operations per second and throughput of the target AI chip, such as calculating the computation time T of the quantization method of the target operator-level object based on the formula: T=OPS / Throughput, where OPS represents the number of operations per second of the target AI chip and Throughput represents the throughput. Alternatively, the memory bandwidth utilization rate of the quantization method of the target operator-level object may be calculated based on the actual memory bandwidth and test memory bandwidth of the target AI chip, such as based on the formula: Calculate the memory bandwidth utilization of the quantization method for the target operator-level object. ,in, The actual memory bandwidth of the target AI chip. This refers to the test memory bandwidth of the target AI chip. Optionally, the computational unit utilization rate of the quantization method for the target operator-level object can also be calculated based on the actual number of computational units used and the total number of computational units in the target AI chip, such as based on the formula: The utilization rate of computational units for quantization methods of target operator-level objects. ,in, This refers to the actual computing units used in the target AI chip. This indicates the total number of computing units in the target AI chip.

[0085] S280. Determine the multi-dimensional evaluation weights of the target operator-level object based on the chip application scenario of the target AI chip and the actual operating data of the target AI chip, and generate the quantization evaluation result of the quantization method of the target operator-level object based on the initial multi-dimensional evaluation result and the multi-dimensional evaluation weights.

[0086] The multi-dimensional evaluation weights can include accuracy weight, performance weight, and power consumption weight. Accuracy weight is used to characterize the importance of accuracy evaluation results, performance weight is used to characterize the importance of performance evaluation results, and power consumption weight is used to characterize the importance of power consumption evaluation results.

[0087] Understandably, different types of AI chips have different application scenarios, thus requiring different quantization methods. For example, when an AI chip focuses on computational applications, the emphasis is on the quantization accuracy corresponding to the quantization method; when an AI chip focuses on lightweight terminal applications, the emphasis is on the performance or power consumption corresponding to the quantization method. Furthermore, the specific performance of the AI ​​chip during operation, i.e., its actual operating data, will also affect its specific requirements for the quantization method. Therefore, as... Figure 5As shown, after evaluating the quantization methods of the target operator-level object from multiple dimensions such as accuracy, performance, and power consumption to obtain initial multi-dimensional evaluation results, the specific quantization requirements of the target AI chip can be evaluated by combining the chip application scenario and the actual operating data of the target AI chip. Based on the evaluation results of the quantization requirements, the weights corresponding to the quantization evaluation results of the target operator-level object in each dimension are determined, resulting in accuracy weights, performance weights, and power consumption weights. Then, the quantization evaluation results of each quantization method are comprehensively scored based on these weights. For example, the score calculated based on the accuracy evaluation result of the target operator-level object is A, the score calculated based on the performance evaluation result is B, and the score calculated based on the power consumption evaluation result is C. If the accuracy weight, performance weight, and power consumption weight are a, b, and c respectively, then the quantization evaluation result of the quantization method of the target operator-level object can be expressed as: .

[0088] S290. Generate the target quantization strategy for the target operator-level object based on the quantization evaluation results.

[0089] In an optional embodiment of the present invention, generating the target quantization strategy for the target operator-level object based on the quantization evaluation result may include: determining the constraints of the target AI chip and / or the target operator-level object; generating a multi-dimensional quantization strategy based on the quantization evaluation result and the constraints of the target operator-level object; wherein the multi-dimensional quantization strategy includes a precision allocation quantization strategy and a multi-objective optimization quantization strategy; performing strategy verification and feasibility checks on the multi-dimensional quantization strategy, and generating a quantization score for each quantization strategy in the multi-dimensional quantization strategy based on the verification and check results; and selecting the target quantization strategy for the target operator-level object from the multi-dimensional quantization strategy based on the quantization scores of each quantization strategy in the multi-dimensional quantization strategy.

[0090] Figure 6 This is a flowchart illustrating a target quantization strategy for generating target operator-level objects, provided in Embodiment 2 of the present invention. In a specific example, such as... Figure 6As shown, when generating a target quantization strategy for a target operator-level object based on the quantization evaluation results, the pre-defined constraints on the target AI chip and / or the target operator-level object can be determined first. Optionally, the constraints can limit the conditions that the quantization strategy must follow from dimensions such as accuracy loss constraints, performance constraints, and power consumption constraints. For example, from the perspective of accuracy and power consumption, it can be required that the accuracy of the quantization method for the target operator-level object is not lower than a set accuracy threshold, and the power consumption is not higher than a set power consumption threshold. The constraints can be set from at least one dimension such as accuracy loss constraints, performance constraints, and power consumption constraints, and this embodiment of the invention does not limit this.

[0091] After setting preset conditions, the quantization evaluation results of various quantization methods for the target operator-level object can be comprehensively analyzed under preset constraints. From the perspectives of precision allocation quantization strategies and multi-objective optimization quantization strategies, a suitable multi-dimensional quantization strategy can be generated for the target operator-level object. For example, regarding precision allocation quantization strategies, the quantization precision corresponding to the input and output of the target operator-level object and / or its internal operations can be determined from the perspective of precision allocation in the quantization method evaluation results, thereby selecting quantization methods that meet its quantization precision requirements. Alternatively, the corresponding quantization method can be determined from the perspective of operator importance, based on the operator importance evaluation results of the target operator-level object. Furthermore, the corresponding quantization method can be determined from the perspective of precision sensitivity, based on the specific precision sensitivity requirements of the target operator-level object. For multi-objective optimization measurement, the quantization method corresponding to the target operator-level object can be determined from aspects such as objective function construction, weight coefficient adjustment, and optimization algorithm selection. Objective function construction involves determining the corresponding quantization method based on the type of objective function corresponding to the target operator-level object and its operational requirements. Weight coefficient adjustment can be understood as adjusting the weights in the multi-objective optimization function to determine the final quantization method. Optimization algorithm selection involves determining the quantization method suitable for the target operator-level object based on the optimization algorithm chosen by the AI ​​model corresponding to the target operator-level object. For example, a multi-objective optimization function could be:

[0092]

[0093]

[0094]

[0095]

[0096] in, Represents a multi-objective optimization function. The loss of precision in representation, Indicates delay loss, Let P represent the power consumption function, and let P represent the precision configuration vector. For the precision threshold, The delay threshold, For power consumption threshold, , and These are the weighting coefficients.

[0097] Figure 7 This is a schematic diagram of the framework structure of a quantitative evaluation system based on an AI chip, provided in Embodiment 2 of the present invention. In a specific example, such as... Figure 7 As shown, the AI ​​chip-based quantitative evaluation system is configured with a hardware adapter layer, which can configure and manage adapters corresponding to various AI chips. The adapter can extract the hardware interface information of the target AI chip as hardware abstraction information, obtaining a series of chip hardware characteristics of the target AI chip. The AI ​​chip-based quantitative evaluation system also has a user interface layer, providing user-facing functions such as operator specification specification, evaluation parameter configuration, and result visualization. The core processing layer of the AI ​​chip-based quantitative evaluation system can interface with the hardware adapter layer and the user interface layer. Through the hardware feature extraction module, it connects to the hardware abstraction interface provided by the hardware abstraction layer to extract the hardware characteristics of the target AI chip from the corresponding adapter according to the type of AI chip specified by the user, thus achieving hardware abstraction. The core processing layer parses the user-specified target operator-level object through the operator type recognition module to extract the operator features of the target operator-level object. The quantitative modeling module can load various quantization methods that can be used for the target operator-level object. The performance evaluation module can perform multi-dimensional evaluation of the quantization method of the target operator-level object based on the chip hardware characteristics of the target AI chip and the target operator features of the target operator-level object. The optimization strategy generation module can generate target quantization strategies for target operator-level objects based on the quantization evaluation results output by the performance evaluation module.

[0098] In the actual deployment of AI chips, quantization is often required for certain key nodes, operators, or operations, thus creating a need for single-node quantization evaluation calculations. Furthermore, different quantization instructions used in various chip architecture and instruction set designs can significantly impact performance, power consumption, and accuracy. When evaluating the performance of single-operator quantization strategies for various chips, different quantization strategies exhibit significant differences in performance across different chips. Considering the quantization needs of these application scenarios, this invention provides a quantization evaluation scheme for single-node quantization methods based on AI chips. It achieves multi-dimensional evaluation of operator-level object quantization methods from the perspectives of the target AI chip's hardware characteristics and operator characteristics, obtaining the performance characteristics of different quantization methods. Based on the multi-dimensional evaluation results, it generates the optimal quantization strategy suitable for operator-level objects. It can select or search for quantization precision based on user or operator characteristics, enabling refined and accurate comparative evaluation of single-node quantization. It also performs fine-grained evaluation of the chip's instruction set and quantized memory access evaluation of the chip's memory access characteristics, significantly improving the refinement and accuracy of chip quantization evaluation, thereby enhancing the quantization accuracy and operating efficiency of AI chips.

[0099] It should be noted that all information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this disclosure are information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data comply with the relevant laws, regulations and standards of the relevant regions.

[0100] It should be noted that any arrangement or combination of the technical features in the above embodiments also falls within the protection scope of this invention.

[0101] Example 3

[0102] Figure 8 This is a schematic diagram of a quantitative evaluation device based on an AI chip provided in Embodiment 3 of the present invention, as shown below. Figure 8 As shown, the device includes: a chip hardware feature extraction module 810, a target operator feature extraction module 820, a quantization evaluation module 830, and a target quantization strategy generation module 840, wherein:

[0103] The chip hardware feature extraction module 810 is used to extract the chip hardware features of the target AI chip based on the chip configuration information of the target AI chip.

[0104] The target operator feature extraction module 820 is used to extract target operator features from target operator-level objects in the target AI chip;

[0105] The quantization evaluation module 830 is used to perform multi-dimensional evaluation of the quantization method of the target operator-level object based on the chip hardware characteristics of the target AI chip and the target operator characteristics of the target operator-level object, and obtain the quantization evaluation result.

[0106] The target quantization strategy generation module 840 is used to generate the target quantization strategy of the target operator-level object based on the quantization evaluation result.

[0107] This invention extracts the chip hardware features of a target AI chip based on its chip configuration information and extracts target operator features from target operator-level objects within the target AI chip. It then performs a multi-dimensional evaluation of the quantization method of the target operator-level objects based on both the chip hardware features and the target operator features, obtaining a quantization evaluation result. Based on this result, a target quantization strategy for the target operator-level objects is generated. This addresses the problem that existing AI chip quantization evaluation methods cannot quantize operator-level objects within AI chips, improving the precision and accuracy of chip quantization evaluation, thereby enhancing the quantization accuracy and operational efficiency of AI chips.

[0108] Optionally, the chip hardware feature extraction module 810 is further configured to: extract hardware characteristic information of the target AI chip based on the chip configuration information; wherein the hardware characteristic information includes at least one of computing unit association information, memory hierarchy association information, instruction set association information, and data precision support association information; generate a hardware characteristic matrix of the target AI chip based on the hardware characteristic information of the target AI chip; determine the hardware adaptation weight of the hardware characteristic matrix based on the chip characteristics of the target AI chip; and generate chip hardware features of the target AI chip based on the hardware characteristic matrix of the target AI chip and the hardware adaptation weight of the hardware characteristic matrix.

[0109] Optionally, the chip hardware feature extraction module 810 is further configured to: analyze the number of computing units, parallelism, and frequency characteristics of the target AI chip based on the chip configuration information to obtain the computing unit association information; analyze the memory bandwidth, cache hierarchy, and memory latency of the target AI chip based on the chip configuration information to obtain the memory hierarchy association information; analyze the instruction set characteristics, computational support, and vectorized computation characteristics of the target AI chip based on the chip configuration information to obtain the instruction set association information; and analyze the support information of the target AI chip for various data precisions based on the chip configuration information to obtain the data precision support association information.

[0110] Optionally, the target operator feature extraction module 820 is further configured to: extract multi-dimensional operator features of the target operator-level object based on the operator computation graph of the target operator-level object; wherein the multi-dimensional operator features include at least one of computation mode, data flow mode, and computational complexity; classify the operator type of the target operator-level object based on the multi-dimensional operator features of the target operator-level object to obtain the target operator type; determine the feature quantization mode adapted to the target operator-level object based on the target operator type of the target operator-level object; generate the target operator features based on the target operator type of the target operator-level object and the adapted feature quantization mode; wherein the target operator features include operator computation features, data transport features, and quantized time-consuming computation features.

[0111] Optionally, the target operator feature extraction module 820 is further configured to: calculate the computational complexity and computational data flow of the target operator-level object based on the multidimensional operator features of the target operator-level object; classify the operator type of the target operator-level object based on the computational complexity and computational data flow of the target operator-level object to obtain the target operator type.

[0112] Optionally, the quantization evaluation module 830 is further configured to: perform accuracy evaluation, performance evaluation, and power consumption evaluation on the quantization method of the target operator-level object based on at least one of the chip hardware characteristics of the target AI chip, the target operator characteristics of the target operator-level object, and the operator quantization index, to obtain an initial multi-dimensional evaluation result; determine the multi-dimensional evaluation weight of the target operator-level object based on the chip application scenario of the target AI chip and the actual operating data of the target AI chip; wherein the multi-dimensional evaluation weight includes accuracy weight, performance weight, and power consumption weight; and generate a quantization evaluation result of the quantization method of the target operator-level object based on the initial multi-dimensional evaluation result and the multi-dimensional evaluation weight.

[0113] Optionally, the quantization evaluation module 830 is further configured to: perform accuracy evaluation on the quantization method of the target operator-level object according to the operator quantization index, and obtain an accuracy evaluation result; wherein the operator quantization index includes peak signal-to-noise ratio, structural similarity, and classification accuracy; perform performance evaluation on the quantization method of the target operator-level object according to the chip hardware characteristics of the target AI chip, and obtain a performance evaluation result; and perform power consumption evaluation on the quantization method of the target operator-level object according to the chip hardware characteristics of the target AI chip and the target operator characteristics of the target operator-level object, and obtain a power consumption evaluation result.

[0114] Optionally, the target quantization strategy generation module 840 is further configured to: determine the constraints of the target AI chip and / or the target operator-level object; generate a multi-dimensional quantization strategy based on the quantization evaluation results and the constraints of the target operator-level object; wherein the multi-dimensional quantization strategy includes a precision allocation quantization strategy and a multi-objective optimization quantization strategy; perform strategy verification and feasibility checks on the multi-dimensional quantization strategy, and generate a quantization score for each quantization strategy in the multi-dimensional quantization strategy based on the verification and check results; and select the target quantization strategy for the target operator-level object from the multi-dimensional quantization strategy based on the quantization scores of each quantization strategy in the multi-dimensional quantization strategy.

[0115] The aforementioned AI chip-based quantitative evaluation device can execute the AI ​​chip-based quantitative evaluation method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in this embodiment can be found in the AI ​​chip-based quantitative evaluation method provided in any embodiment of the present invention.

[0116] Since the AI ​​chip-based quantitative evaluation device described above is capable of executing the AI ​​chip-based quantitative evaluation method in the embodiments of the present invention, those skilled in the art can understand the specific implementation and various variations of the AI ​​chip-based quantitative evaluation device in this embodiment based on the AI ​​chip-based quantitative evaluation method described in the embodiments of the present invention. Therefore, how the AI ​​chip-based quantitative evaluation device implements the AI ​​chip-based quantitative evaluation method in the embodiments of the present invention will not be described in detail here. Any device used by those skilled in the art to implement the AI ​​chip-based quantitative evaluation method in the embodiments of the present invention falls within the scope of protection of this application.

[0117] Example 4

[0118] Figure 9 A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0119] like Figure 9As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0120] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0121] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as quantitative evaluation methods based on AI chips.

[0122] Optionally, the quantitative evaluation method based on AI chips may include: extracting the chip hardware features of the target AI chip based on the chip configuration information of the target AI chip; extracting target operator features from the target operator-level objects in the target AI chip; performing a multi-dimensional evaluation of the quantization method of the target operator-level objects based on the chip hardware features of the target AI chip and the target operator features of the target operator-level objects to obtain a quantitative evaluation result; and generating a target quantization strategy for the target operator-level objects based on the quantitative evaluation result.

[0123] In some embodiments, the AI ​​chip-based quantitative evaluation method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the AI ​​chip-based quantitative evaluation method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the AI ​​chip-based quantitative evaluation method by any other suitable means (e.g., by means of firmware).

[0124] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0125] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0126] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0127] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0128] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0129] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0130] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0131] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A quantitative evaluation method based on AI chips, characterized in that, The method comprises the following steps: extracting chip hardware features of the target AI chip according to chip configuration information of the target AI chip; extracting target operator features of a target operator level object in the target AI chip; performing multi-dimensional evaluation on a quantization mode of the target operator level object according to the chip hardware features of the target AI chip and the target operator features of the target operator level object, to obtain a quantization evaluation result; generating a target quantization strategy of the target operator level object according to the quantization evaluation result, wherein the target quantization strategy is a multi-objective optimization quantization strategy.

2. The method of claim 1, wherein, The method of extracting the chip hardware features of the target AI chip according to the chip configuration information of the target AI chip comprises: extracting hardware characteristic information of the target AI chip according to the chip configuration information, wherein the hardware characteristic information comprises at least one of calculation unit association information, memory hierarchy association information, instruction set association information and data precision support association information; generating a hardware characteristic matrix of the target AI chip according to the hardware characteristic information of the target AI chip; determining a hardware adaptation weight of the hardware characteristic matrix according to the chip characteristics of the target AI chip; generating the chip hardware features of the target AI chip according to the hardware characteristic matrix of the target AI chip and the hardware adaptation weight of the hardware characteristic matrix.

3. The method of claim 2, wherein, The method of extracting the hardware characteristic information of the target AI chip according to the chip configuration information comprises: analyzing a calculation unit number, parallelism and frequency characteristic information of the target AI chip according to the chip configuration information, to obtain the calculation unit association information; analyzing a memory bandwidth, cache hierarchy and memory delay information of the target AI chip according to the chip configuration information, to obtain the memory hierarchy association information; analyzing instruction set characteristics, operation support and vector operation characteristic information of the target AI chip according to the chip configuration information, to obtain the instruction set association information; analyzing support information of the target AI chip for various data precisions according to the chip configuration information, to obtain the data precision support association information.

4. The method of claim 1, wherein, The method of extracting the target operator features of the target operator level object in the target AI chip comprises: extracting multi-dimensional operator features of the target operator level object according to an operator computation graph of the target operator level object, wherein the multi-dimensional operator features comprise at least one of a calculation mode, a data flow mode and a calculation complexity; classifying operator types of the target operator level object according to the multi-dimensional operator features of the target operator level object, to obtain target operator types; determining a characteristic quantization mode adapted by the target operator level object according to the target operator types of the target operator level object; generating the target operator features according to the target operator types of the target operator level object and the adapted characteristic quantization mode, wherein the target operator features comprise operator computation features, data transfer features and time-consuming computation features after quantization.

5. The method of claim 4, wherein, The method of classifying the operator types of the target operator level object according to the multi-dimensional operator features of the target operator level object, to obtain the target operator types, comprises: According to the multi-dimensional operator characteristics of the target operator level object, the calculation complexity and the calculation data flow of the target operator level object are calculated; According to the calculation complexity and the calculation data flow of the target operator level object, the operator type of the target operator level object is classified, and the target operator type is obtained.

6. The method of claim 1, wherein, According to the multi-dimensional evaluation of the quantization mode of the target operator level object according to the chip hardware characteristics of the target AI chip and the target operator characteristics of the target operator level object, a quantization evaluation result is obtained, including: According to at least one of the chip hardware characteristics of the target AI chip, the target operator characteristics of the target operator level object and the operator quantization index, the quantization mode of the target operator level object is evaluated in precision, performance and power consumption, and an initial multi-dimensional evaluation result is obtained; According to the chip application scenario of the target AI chip and the actual running data of the target AI chip, the multi-dimensional evaluation weight of the target operator level object is determined; wherein the multi-dimensional evaluation weight includes precision weight, performance weight and power consumption weight; According to the initial multi-dimensional evaluation result and the multi-dimensional evaluation weight, the quantization evaluation result of the quantization mode of the target operator level object is generated.

7. The method of claim 6, wherein, According to at least one of the chip hardware characteristics of the target AI chip, the target operator characteristics of the target operator level object and the operator quantization index, the quantization mode of the target operator level object is evaluated in precision, performance and power consumption, and an initial multi-dimensional evaluation result is obtained, including: According to the operator quantization index, the quantization mode of the target operator level object is evaluated in precision, and a precision evaluation result is obtained; wherein the operator quantization index includes peak signal-to-noise ratio, structural similarity and classification accuracy; According to the chip hardware characteristics of the target AI chip, the performance of the quantization mode of the target operator level object is evaluated, and a performance evaluation result is obtained; According to the multi-dimensional evaluation of the quantization mode of the target operator level object according to the chip hardware characteristics of the target AI chip and the target operator characteristics of the target operator level object, a quantization evaluation result is obtained, including:

8. The method of claim 1, wherein, Determine the constraint condition of the target AI chip and / or the target operator level object; According to the quantization evaluation result and the constraint condition of the target operator level object, a multi-dimensional quantization strategy is generated; wherein the multi-dimensional quantization strategy includes a multi-objective optimization quantization strategy; The multi-dimensional quantization strategy is verified and checked for feasibility, and the quantization score of each quantization strategy in the multi-dimensional quantization strategy is generated based on the verification and checking result; According to the quantization score of each quantization strategy in the multi-dimensional quantization strategy, the target quantization strategy of the target operator level object is selected from the multi-dimensional quantization strategy. Including: 9.A device for quantitatively evaluating an AI chip, comprising: The chip hardware feature extraction module is used for extracting the chip hardware features of the target AI chip according to the chip configuration information of the target AI chip; ​ The target operator feature extraction module is configured to extract target operator features of target operator level objects in the target AI chip. The quantization evaluation module is configured to perform multi-dimensional evaluation on a quantization manner of the target operator level objects according to chip hardware features of the target AI chip and the target operator features of the target operator level objects, and obtain a quantization evaluation result. The target quantization strategy generation module is configured to generate a target quantization strategy of the target operator level objects according to the quantization evaluation result, wherein the target quantization strategy is a multi-objective optimization quantization strategy.

10. An electronic device, comprising: The electronic device comprises: at least one processor; and a memory connected to the at least one processor in communication; wherein The memory stores a computer program executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the AI chip-based quantization evaluation method of any one of claims 1-8.

11. A computer readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for enabling the processor to execute the AI chip-based quantization evaluation method of any one of claims 1-8 when executed.

12. A computer program product, characterised in that, The computer program / instructions are executed by the processor to implement the AI chip-based quantization evaluation method of any one of claims 1-8.

Citation Information

Patent Citations

  • AI chip reasoning quantification method for intelligent servo driver

    CN115018076A

  • Chip evaluation method and device, electronic equipment and medium

    CN117851208A