Energy efficiency ratio tuning method of neural network processor and related equipment

By constructing Bayesian distribution model and acquisition function for Bayesian optimization processing, the problem of difficulty in determining NPU matrix parameters is solved, and the effect of improving energy efficiency ratio and reducing energy consumption is achieved.

CN119938318AActive Publication Date: 2025-05-06PENG CHENG LAB

Patent Information

Application Number
CN202411946612.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-06
Estimated Expiration
2044-12-25

AI Technical Summary

Technical Problem

In the prior art, when determining the NPU matrix parameters through traversal search methods, the huge search space makes it difficult to determine the target matrix parameters, and cannot effectively improve the energy efficiency ratio of the NPU, resulting in excessive energy consumption.

Method used

A method for tuning energy efficiency ratio of neural network processors is proposed. By obtaining the matrix parameter set, building a Bayesian distribution model, using the acquisition function to perform Bayesian optimization processing, and obtaining the target matrix parameter set, thereby optimizing the energy efficiency ratio of NPU.

Benefits of technology

Through this method, the energy efficiency ratio of the NPU can be effectively improved, energy consumption can be reduced, energy waste caused by blind sampling, and more efficient computing processing can be achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938318A_ABST
    Figure CN119938318A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an energy efficiency ratio tuning method of a neural network processor and related equipment, and belongs to the technical field of computer processing, and the method comprises the steps: obtaining a matrix parameter set of the neural network processor, and determining a working energy efficiency ratio function of the neural network processor based on the matrix parameter set; a Bayesian distribution model is constructed based on the working energy efficiency ratio function and the matrix parameter set, and the Bayesian distribution model comprises an acquisition function used for evaluating the quality of the matrix parameter set; performing Bayesian optimization processing on the matrix parameter set by using the acquisition function to obtain a target matrix parameter set; and controlling the neural network processor to enter a running state based on the target matrix parameter set. According to the invention, the energy efficiency ratio of the NPU can be improved, and the energy consumption is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer processing technology, and in particular to a method for optimizing the energy efficiency ratio of a neural network processor and related equipment. Background Art

[0002] A neural processing unit (NPU) is a processor specifically designed for neural network calculations. Generally, intelligent computing clusters use matrix multiplication calculations inside the NPU to perform deep learning parallel tasks, thereby helping the intelligent computing clusters to exert their performance advantages; furthermore, how to divide the matrix on the NPU into blocks and combine and reuse the divided matrix blocks has become the key to NPU matrix multiplication tuning. Energy efficiency ratio is an important indicator for measuring NPU energy consumption. The larger the energy efficiency ratio, the more computing work the NPU can complete under the same energy consumption.

[0003] In the related art, in order to maximize the energy efficiency of the NPU and thus enter the best operating state, a traversal search method is used to determine the target matrix parameters from multiple matrix parameters of the NPU matrix. However, the huge search space brings great difficulties to the determination of the target matrix parameters. As a result, the tuning method used in the related art cannot effectively improve the energy efficiency of the NPU, which leads to excessive energy consumption of the NPU. Summary of the invention

[0004] The main purpose of the embodiments of the present application is to propose a method for tuning the energy efficiency of a neural network processor and related equipment, aiming to improve the energy efficiency of the NPU and reduce energy consumption.

[0005] To achieve the above-mentioned purpose, a first aspect of an embodiment of the present application proposes a method for optimizing the energy efficiency ratio of a neural network processor, the method comprising:

[0006] Acquire a matrix parameter set of a neural network processor, and determine a working energy efficiency ratio function of the neural network processor based on the matrix parameter set;

[0007] Building a Bayesian distribution model based on the work energy efficiency ratio function and the matrix parameter set, wherein the Bayesian distribution model includes a collection function for evaluating the quality of the matrix parameter set;

[0008] Using the acquisition function to perform Bayesian optimization on the matrix parameter set, the target matrix parameter set is obtained;

[0009] The neural network processor is controlled to enter a running state based on a target matrix parameter set.

[0010] In some embodiments, the neural network processor includes a target matrix for performing operation processing, the target matrix including a plurality of matrix elements;

[0011] Get the matrix parameter set of the neural network processor, including:

[0012] Based on a preset first value range, determining at least one blocking factor;

[0013] Based on a preset second value range, determining a loop expansion factor that matches the blocking factor, wherein the blocking factor is used to divide the target matrix into a plurality of small sub-matrices, and the loop expansion factor represents a combined calculation of a plurality of matrix elements in the sub-matrices;

[0014] Determining a current operating frequency of the neural network processor based on a preset third value range;

[0015] A matrix parameter set is obtained according to the blocking factor, the loop expansion factor and the current running frequency.

[0016] In some embodiments, a plurality of different levels of cache space are provided in the neural network processor;

[0017] After determining the working energy efficiency ratio function of the neural network processor based on the matrix parameter set, it also includes:

[0018] Predicting the usage frequency of the sub-matrix to obtain frequency level information, and storing the sub-matrix in a cache space matching the frequency level information;

[0019] Based on the matrix parameter set, the neural network processor is controlled to enter a working state, so that the neural network processor processes the data in all cache spaces to obtain the calculation amount information and average power information per unit time;

[0020] Based on the ratio of the computational amount information and the average power information, the energy efficiency ratio information corresponding to the matrix parameter set is obtained.

[0021] In some embodiments, a Bayesian distribution model is constructed based on the work energy efficiency ratio function and the matrix parameter set, including:

[0022] Based on multiple matrix parameter sets and working energy efficiency ratio functions, determining corresponding multiple energy efficiency ratio information;

[0023] Based on multiple energy efficiency ratio information and corresponding matrix parameter sets, construct a mean function representing the correlation between the multiple energy efficiency ratio information and a covariance function representing the correlation between the matrix parameter set and the energy efficiency ratio information;

[0024] A Bayesian distribution model is constructed based on the mean function and covariance function.

[0025] In some embodiments, the acquisition function comprises a first acquisition function;

[0026] The first acquisition function is determined by the following steps, the steps comprising:

[0027] The mean function is updated based on a preset cumulative distribution function to obtain an updated mean function;

[0028] Updating the covariance function based on a preset probability density function to obtain a first updated covariance function;

[0029] The updated mean function and the first updated covariance function are superimposed to construct a first acquisition function for evaluating the quality of a matrix parameter set.

[0030] In some embodiments, the acquisition function further comprises a second acquisition function;

[0031] The second acquisition function is determined by the following steps, comprising:

[0032] The covariance function is updated based on the preset hyperparameters to obtain a second updated covariance function;

[0033] The mean function and the second updated covariance function are superimposed to construct a second acquisition function for evaluating the quality of the matrix parameter set.

[0034] In some embodiments, the acquisition function is used to perform Bayesian optimization processing on the matrix parameter set to obtain a target matrix parameter set, including:

[0035] Optimizing the acquisition function to obtain an updated matrix parameter set, and generating updated energy efficiency ratio information based on the updated matrix parameter set and the working energy efficiency ratio function;

[0036] Calculate the energy efficiency ratio difference between the updated energy efficiency ratio information and the initial optimal energy efficiency ratio information, update the acquisition function based on the updated energy efficiency ratio information to obtain an updated acquisition function, and update the initial optimal energy efficiency ratio based on the updated energy efficiency ratio information;

[0037] The updated acquisition function is used as a new acquisition function to iteratively update the matrix parameter set until the energy efficiency ratio difference is less than a preset threshold difference, and the updated matrix parameter set is used as the target matrix parameter set.

[0038] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application provides an energy efficiency ratio tuning device for a neural network processor, the device comprising:

[0039] An acquisition module, used to acquire a matrix parameter set of a neural network processor, and determine a working energy efficiency ratio function of the neural network processor based on the matrix parameter set;

[0040] A construction module, used to construct a Bayesian distribution model based on the work energy efficiency ratio function and the matrix parameter set, wherein the Bayesian distribution model includes a collection function for evaluating the quality of the matrix parameter set;

[0041] The Bayesian optimization module is used to perform Bayesian optimization processing on the matrix parameter set using the acquisition function to obtain the target matrix parameter set;

[0042] The target operation module is used to control the neural network processor to enter the operation state based on the target matrix parameter set.

[0043] To achieve the above-mentioned purpose, the third aspect of an embodiment of the present application proposes an electronic device, the electronic device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the method of the above-mentioned first aspect when executing the computer program.

[0044] To achieve the above objectives, the fourth aspect of an embodiment of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the method of the first aspect is implemented.

[0045] The energy efficiency ratio tuning method and related equipment of the neural network processor proposed in the present application include obtaining the matrix parameter set of the neural network processor, and determining the working energy efficiency ratio function of the neural network processor based on the matrix parameter set; constructing a Bayesian distribution model based on the working energy efficiency ratio function and the matrix parameter set, wherein the Bayesian distribution model includes an acquisition function for evaluating the quality of the matrix parameter set; the tuning device can clarify the acquisition direction from multiple sets of matrix parameter sets according to the acquisition function, avoiding blindly sampling the matrix parameter set and causing energy waste; then, using the acquisition function to perform Bayesian optimization processing on the matrix parameter set to obtain the target matrix parameter set, thereby achieving the purpose of selecting the most valuable set for sampling from other matrix parameter sets whose energy efficiency ratio information has not been calculated; then, controlling the neural network processor to enter the running state based on the target matrix parameter set. In this way, the neural network processor can efficiently process the current task while improving the energy efficiency ratio information and reducing energy consumption. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 This is a schematic diagram of an optional application scenario of the energy efficiency ratio tuning device for a neural network processor provided in an embodiment of the present application;

[0047] Figure 2 It is an optional flow chart of the energy efficiency ratio tuning method of the neural network processor provided in the embodiment of the present application;

[0048] Figure 3 yes Figure 2 An optional implementation flowchart of step 101 in FIG. 1 ;

[0049] Figure 4 yes Figure 2An optional implementation flowchart after step 101;

[0050] Figure 5 yes Figure 2 An optional implementation flowchart of step 102;

[0051] Figure 6 This is another optional implementation flow chart of the energy efficiency ratio tuning device for a neural network processor provided in an embodiment of the present application;

[0052] Figure 7 This is another optional implementation flow chart of the energy efficiency ratio tuning device for the neural network processor provided in the embodiment of the present application;

[0053] Figure 8 yes Figure 2 An optional implementation flowchart of step 103;

[0054] Fig. 9 It is an optional flow chart of the energy efficiency ratio tuning device of the neural network processor provided in the embodiment of the present application;

[0055] Fig.10 It is a schematic diagram of the hardware structure of the electronic device provided in the embodiment of the present application. DETAILED DESCRIPTION

[0056] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0057] It should be noted that, although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first", "second", etc. in the specification, claims and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0059] First, some nouns involved in this application are analyzed:

[0060] Artificial Neural Network (ANN) abstracts the human brain neural network from the perspective of information processing, establishes a simple model, and forms different networks according to different connection methods. In engineering and academia, it is often simply referred to as neural network or quasi-neural network. A neural network is a computational model composed of a large number of nodes (or neurons) connected to each other. Each node represents a specific output function, called an activation function; the connection between each two nodes represents a weighted value for the signal passing through the connection, called a weight, which is equivalent to the memory of an artificial neural network.

[0061] Neural network tasks: mainly include pattern recognition, data prediction, system control, feature importance identification, regression tasks, and classification tasks. Since neural network tasks involve constraints and interactions in large-scale parameters, nonlinear mapping, training data, hardware resources, etc., neural network tasks are more complex than other tasks.

[0062] Next, the background technology related to the embodiments of the present application is described.

[0063] NPU is a processor specially designed for neural network calculations. Generally, intelligent computing clusters use matrix multiplication calculations inside NPUs to perform deep learning parallel tasks, thereby helping intelligent computing clusters to exert their performance advantages; therefore, how to divide the matrix on the NPU into blocks and combine and reuse the divided matrix blocks has become the key to NPU matrix multiplication tuning.

[0064] Among them, data centers and computing clusters, as the core infrastructure for the current development of the information technology industry, bear the huge data processing needs of modern society and have become an important support for the global digital economy. Especially today when artificial intelligence technology is advancing by leaps and bounds, intelligent computing clusters (also referred to as "intelligent computing clusters") have gradually become a key support for promoting artificial intelligence research and industrial development. Intelligent computing clusters are built on advanced intelligent computing processors such as graphics processing units (GPUs) and NPUs, which can efficiently handle complex artificial intelligence tasks and large-scale data computing needs.

[0065] Moreover, with the rapid development of artificial intelligence technology, the current intelligent computing cluster architecture has changed significantly compared with traditional data centers and supercomputer clusters: in traditional clusters, computing tasks mainly rely on the general computing power provided by the central processing unit (CPU), so the CPU, as the core of the node, is responsible for processing various types of tasks. However, the CPU's versatility for a wide range of tasks makes the CPU in the intelligent computing cluster relatively limited in its ability to process parallel tasks, especially when facing highly parallel computing tasks such as deep learning and large-scale image processing.

[0066] In comparison, NPU, as a processor specially designed for artificial intelligence tasks, can process a large number of matrix operations and data flow calculations in neural networks with extremely high efficiency. Therefore, it is widely used in intelligent computing clusters and can better exert its performance advantages when performing deep learning tasks, significantly improving the intelligent computing power of intelligent computing clusters.

[0067] That is, in the current intelligent computing cluster, NPU has become the basic unit for providing intelligent computing power, supporting large-scale artificial intelligence computing needs. A large intelligent computing cluster may contain tens of thousands of NPUs. The existence of these NPUs enables the intelligent computing cluster to have super parallel computing capabilities and can efficiently handle artificial intelligence tasks, such as the training and reasoning of deep learning models. Among them, energy efficiency is an important indicator for measuring NPU energy consumption. The larger the energy efficiency, the more computing work the NPU can complete under the same energy consumption.

[0068] In the related art, in order to maximize the energy efficiency of the NPU and thus enter the best operating state, a traversal search method is used to determine the target matrix parameter set from multiple matrix parameters of the NPU matrix. However, the huge search space brings great difficulties to the determination of the target matrix parameter set. As a result, the tuning method used in the related art cannot effectively improve the energy efficiency of the NPU, which leads to excessive energy consumption of the NPU.

[0069] Specifically, taking the world's top intelligent computing clusters as an example, the peak power consumption of the US Frontier supercomputer system is as high as 22.7 megawatts (MW); the peak power consumption of the supercomputer (Aurora system) is 24.7 MW; and the peak power consumption of China's Sunway Ocean Light cluster is also at the same order of magnitude. According to estimates, the electricity cost of a cluster with a power consumption of 20 MW is more than 100 million yuan in one year, and the corresponding carbon dioxide emissions exceed 70,000 tons. These data reflect that while intelligent computing clusters provide powerful computing capabilities, they also bring huge challenges to energy and the environment.

[0070] Based on this, an embodiment of the present application provides an energy efficiency tuning method for a neural network processor and related equipment, aiming to improve the energy efficiency of the NPU and reduce energy consumption.

[0071] The energy efficiency ratio tuning method and related equipment of the neural network processor provided in the embodiments of the present application are specifically explained through the following embodiments. First of all, it should be noted that the energy efficiency ratio tuning method of the neural network processor in the embodiments of the present application relates to the field of computer processing technology. It can be applied to the terminal or to the server side, as long as the deployed system includes a neural network processor and the energy efficiency ratio of the neural network processor needs to be optimized. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or it can be configured as a server cluster or distributed system composed of multiple physical servers, and it can also be configured as a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, and basic cloud computing services such as big data and artificial intelligence platforms; this is only an example, and is not limited to the above forms.

[0072] It should also be noted that in the embodiments of the present application, when it comes to information related to user characteristics such as user basic information or user identity, the user's permission or consent will be obtained first, and the collection, use and processing of these data will comply with relevant laws, regulations and standards. In addition, when the embodiments of the present application need to obtain sensitive personal information of the user, the user's separate permission or consent will be obtained first. After the user's separate permission or consent is clearly obtained, the necessary data for the normal operation of the embodiments of the present application will be obtained. For example, the energy efficiency ratio tuning device of the neural network processor proposed in the embodiment of the present application (for the convenience of description, it may also be referred to as "tuning device" below) will obtain the consent of the personnel who manage the relevant servers before determining the target matrix parameter set of the NPU using the energy efficiency ratio tuning method of the neural network processor proposed in the embodiment of the present application (for the convenience of description, it may also be referred to as "tuning method" below). Otherwise, the matrix parameter set for determining the target matrix parameters cannot be obtained. In addition, all relevant data obtained in the embodiments of the present application are authorized data after obtaining consent, and will not be repeated here.

[0073] like Figure 1 As shown, Figure 1 is a schematic diagram of an optional application scenario of the energy efficiency ratio tuning device for a neural network processor provided in an embodiment of the present application, Figure 1 In the example, the tuning method proposed in the embodiment of the present application can be applied to a server cluster. Figure 1The server cluster includes multiple servers such as server a, server b and server c (for example only and does not limit the number of servers). Assume that in order to implement a certain intelligent computing task, an NPU is deployed in each server and the relevant parameters (matrix parameter set) that affect the operation of the NPU need to be adjusted; in order to enable the NPU to enter a working state while minimizing energy consumption, the matrix parameter set is input into the tuning device proposed in the embodiment of the present application. The tuning device uses the tuning method proposed in the embodiment of the present application. Without trying all possible matrix parameter sets, it can determine and output a target matrix parameter set that optimizes the energy efficiency ratio of the NPU, and apply the target matrix parameter set to the NPU. It can be understood that the energy consumption of the NPU that enters the working state based on the target matrix parameter set is the lowest among all feasible decision solutions of the matrix parameter set.

[0074] After understanding the example application scenarios of the tuning device proposed in the present application, the tuning method proposed in the embodiment of the present application is described next.

[0075] In the embodiments of the present application, the optimization device will be described from the perspective of: Figure 2 As shown, Figure 2 is an optional flow chart of the method for optimizing the energy efficiency ratio of a neural network processor provided in an embodiment of the present application. Figure 2 The method in the embodiment may include but is not limited to the following steps 101 to 104. When the tuning device executes the tuning method, the specific process is as follows. It should be noted that the present embodiment Figure 2 The order of step 101 to step 104 is not specifically limited, and the order of steps can be adjusted or some steps can be reduced or added according to actual needs.

[0076] Step 101, obtaining a matrix parameter set of a neural network processor, and determining a working energy efficiency ratio function of the neural network processor based on the matrix parameter set.

[0077] The following is a detailed description of step 101.

[0078] Among them, the neural network processor (NPU) is a hardware processor specially designed to accelerate neural network computing tasks. Compared with the traditional central processing unit (CPU) and graphics processing unit (GPU), the NPU can provide efficient neural network computing support for the server. In practical applications, in order to solve the neural network processing tasks, it is necessary to build an intelligent computing cluster based on multiple servers to solve more complex neural network processing tasks. That is to say, the NPU can be deployed in a single server; or, the NPU can be deployed in multiple servers of the intelligent computing cluster; or, when the intelligent computing cluster also includes smart cameras, smart homes and other smart devices, the NPU can also be deployed in these smart devices. The embodiments of the present application do not limit the actual deployment scenarios of the NPU. No matter which one, the tuning method proposed in this application can be used to tune the energy efficiency ratio of the NPU so as to improve the energy efficiency ratio of the NPU and reduce the energy consumption of the NPU while completing the neural network tasks.

[0079] Furthermore, one of the core functions of the NPU is to efficiently perform matrix multiplication operations. Matrix multiplication is the most commonly used basic operator in deep learning calculations and is widely used in operations such as forward propagation, backpropagation, and weight update of neural networks. For example, the matrix multiplication performed in the NPU is represented as C=AB, where matrix A is a matrix of dimension M×K, matrix B is a matrix of dimension K×N, and matrix C is a matrix of dimension M×N. By performing matrix multiplication operations, the NPU can complete a series of complex neural network task processing.

[0080] Furthermore, compared to the traditional NPU triple-loop matrix multiplication, when the matrix dimension is quite large, processing the matrix in blocks can greatly improve the computational efficiency. In other words, when a target neural network task C needs to be processed, how to divide the task C into multiple small sub-matrices c=ab and determine the operation mode of the sub-matrix is ​​the key to improving the energy efficiency of the NPU.

[0081] The following will explain in detail how to tune the energy efficiency ratio of the NPU, and the beneficial effects of the tuning method proposed in the embodiment of the present application will gradually become apparent during the description process.

[0082] In some embodiments, Figure 3 As shown, Figure 3 yes Figure 2 An optional implementation flowchart of step 101 in the embodiment of the present invention, obtaining a matrix parameter set of a neural network processor, includes the following steps 201 to 204:

[0083] Step 201: determine at least one blocking factor based on a preset first value range.

[0084] Step 202: determine a loop expansion factor that matches the blocking factor based on a preset second value range, wherein the blocking factor is used to divide the target matrix into multiple small sub-matrices, and the loop expansion factor represents a combined calculation of multiple matrix elements in the sub-matrix.

[0085] Step 203: Determine the current operating frequency of the neural network processor based on the preset third value range.

[0086] Step 204, obtaining a matrix parameter set according to the blocking factor, the loop expansion factor and the current operating frequency.

[0087] Steps 201 to 204 are described in detail below.

[0088] In some embodiments, the neural network processor includes a target matrix for performing arithmetic processing, and the target matrix includes multiple matrix elements. The target matrix is ​​used to perform arithmetic processing in the neural network and is the core part of the NPU for calculation. The target matrix is ​​composed of multiple matrix elements, and the matrix elements can be numbers, vectors, or other data types, depending on the specific application and requirements of the neural network. The arrangement and combination of the matrix elements in the target matrix determines the type and accuracy of calculations that the NPU can perform. Furthermore, the NPU performs arithmetic processing on the matrix elements of the target matrix by executing specific arithmetic rules (such as matrix multiplication, addition, convolution, etc.) to realize the processing and transformation of data in the neural network task.

[0089] Among them, the matrix parameter set represents the sum of the NPU related parameters that need to be adjusted, and the matrix parameter set includes the blocking factor, the loop expansion factor and the current operating frequency. The blocking factor is a parameter used to divide the target matrix into multiple small sub-matrices. For example, for the target matrix C, the blocking factor includes m, n, k. The target matrix can be divided into multiple sub-matrices c = ab based on the blocking factor, where a = m × k, b = k × n; the size of the blocking factor directly affects the size and number of the sub-matrices obtained by the division.

[0090] Furthermore, the loop expansion factor characterizes the magnitude of the combined calculation of multiple matrix elements in the sub-matrix, so as to improve the computational efficiency through loop expansion. For example, assuming that the loop expansion factor is 1, it means that in each loop iteration, the NPU will only calculate one matrix element in a sub-matrix. In this case, the computation process is serialized and the efficiency is low; assuming that the loop expansion factor is 4, then the NPU will simultaneously calculate 4 matrix elements in a sub-matrix in each loop iteration, thereby reducing the number of iterative calculations and improving the computational efficiency of the NPU.

[0091] Furthermore, the current operating frequency of the neural network processor refers to the clock rate of the processor when executing a task, which determines the number of instruction cycles that the processor can execute per second. The current operating frequency is usually expressed in Hertz (Hz). A higher current operating frequency allows the NPU to perform more operations per unit time, which means that in the same time, a high-frequency NPU can process more data, but it will also increase power consumption. In other words, the increase in NPU frequency is proportional to the increase in power consumption. Therefore, it is very important to set a suitable frequency for the NPU to ensure that the NPU can efficiently process neural network tasks while reducing energy consumption.

[0092] Furthermore, the blocking factor, loop expansion factor and current operating frequency interact with each other, and thus jointly affect the computational efficiency and energy consumption performance of the NPU matrix multiplication. Based on this, the blocking factor, loop expansion factor and current operating frequency are taken as factors affecting the NPU operating efficiency and energy efficiency ratio to obtain a matrix parameter set. For example, when the blocking factor includes m, n, k, and the loop expansion factor includes U m , U n , U k , when the current operating frequency of NPU is f, the matrix parameter set is (m, n, k, U m , U n , U k , f).

[0093] It should be noted that the number of loop expansion factors usually matches the number of blocking factors. The specific number of blocking factors can be set according to actual conditions. The embodiments of the present application are only used as examples and do not limit this to the embodiments of the present application.

[0094] Furthermore, in order to determine the matrix parameter set subsequently input to the Bayesian distribution model, first obtain the first value range of the predefined blocking factor, the second value range of the loop expansion factor, and the third value range of the current operating frequency. Generally, due to the limitation of the NPU cache size, the values ​​of m, n, and k are usually in the integer range of [1, 128]. m , U n , U k Usually in the range of [1, 16], the current operating frequency f of the NPU has multiple adjustable options, usually more than 10 optional frequencies, and the specific value range is determined by the specific NPU model. The first value range, the second value range and the third value range can all be set according to actual conditions, and the embodiment of the present application does not limit this.

[0095] In some embodiments, Figure 4 As shown, Figure 4 yes Figure 2An optional implementation flow chart after step 101 in the embodiment of the present invention further includes the following steps 301 to 303 after determining the working energy efficiency ratio function of the neural network processor based on the matrix parameter set:

[0096] Step 301 : predict the usage frequency of the sub-matrix to obtain frequency level information, and store the sub-matrix in a cache space matching the frequency level information.

[0097] Step 302, based on the matrix parameter set, controls the neural network processor to enter a working state, so that the neural network processor processes the data in all cache spaces to obtain the amount of calculation information and average power information per unit time.

[0098] Step 303: Based on the ratio of the calculation amount information and the average power information, energy efficiency ratio information corresponding to the matrix parameter set is obtained.

[0099] Steps 301 to 303 are described in detail below.

[0100] In some embodiments, the working energy efficiency ratio function of the neural network processor is determined based on the matrix parameter set, as shown in the following formula <1>:

[0101]

[0102] Among them, Performance represents the computing performance of the NPU, which is defined as the amount of matrix multiplication calculations completed per unit time, and is therefore equal to the inverse of the time it takes to complete a matrix multiplication; Power represents the average power of the NPU during the matrix multiplication calculation process. Both Performance and Power can be calculated in real time by the server where the NPU is located or other related devices.

[0103] In some embodiments, a plurality of different levels of cache space are provided in the neural network processor, and the processing efficiency and storage capacity of different levels of cache space are usually different. Specifically, the NPU usually has a multi-level cache structure, including a first-level cache (First-Level Cache, L1 cache), a second-level cache (Second-Level Cache, L2 cache) and a global shared memory. Each cache layer has its specific capacity and access speed difference. In order to improve the energy efficiency of matrix multiplication operations, when performing block matrix multiplication, the size of each block of matrix data must be reasonably planned to ensure that it can be loaded into the cache exactly, avoiding performance loss caused by cache overflow.

[0104] Furthermore, the L1 cache is usually the high-speed cache closest to the computing unit. It has extremely low latency but small capacity. Therefore, the L1 cache usually stores the data that the NPU processes most frequently, such as the matrix block currently being processed; the L2 cache has a relatively large capacity, but the access speed is slower than the L1 cache. The L2 cache is usually used to store large blocks of data that have been loaded to avoid frequent access to the global memory. Therefore, the optimization of the NPU's block matrix multiplication often stores larger data blocks in the L2 cache; the global shared memory is the slowest part of the NPU storage system and is usually used to store complete matrix data. Due to its high access cost, when optimizing matrix multiplication, frequent access to global memory should be minimized, and more reliance should be placed on data in the L1 and L2 caches.

[0105] Furthermore, after the target matrix is ​​divided into multiple sub-matrices using different blocking factors, the frequency of use of the sub-matrices is first predicted to obtain frequency level information to determine the cache space required to store the sub-matrices so that the NPU can process the sub-matrices later. The frequency of use of the sub-matrices can be achieved by analyzing historical data, matrix usage patterns, or through specific prediction algorithms.

[0106] Furthermore, the frequency level information matches different cache spaces. For example, the frequency level information includes a first-level frequency, a second-level frequency, and a third-level frequency. The first-level frequency matches the L1 cache, the second-level frequency matches the L2 cache, and the third-level frequency matches the global shared memory. Of course, the specific amount of frequency level information and its matching relationship with different cache spaces can be adaptively adjusted according to actual conditions, and the embodiments of the present application do not limit this.

[0107] Furthermore, based on the matrix parameter set, the neural network processor is controlled to enter a working state, so that the neural network processor processes the data in all cache spaces, obtains the computational information representing the specific value of Performance, and the average power information representing the specific value of Power, and then determines the energy efficiency ratio information under the current matrix parameter set.

[0108] Furthermore, in order to reduce the calculation error of the energy efficiency ratio information, the averaged calculation amount information and average power information are often obtained by repeated measurements. The number of repeated measurements can be set according to actual conditions, usually more than 5 times, but the embodiments of the present application do not limit this.

[0109] For example, when the matrix parameter set is (m1, n1, k1, U m1 , U n1 , U k1 , f1), the NPU enters the running state and obtains the energy efficiency ratio information EER1; when the matrix parameter set is (m2, n2, k2, Um2 , U n2 , U k2 , f2), the NPU enters the running state and obtains the energy efficiency ratio information EER2. In the case of different matrix parameter sets, EER1 and EER2 are usually different.

[0110] Step 102: construct a Bayesian distribution model based on the work energy efficiency ratio function and the matrix parameter set, wherein the Bayesian distribution model includes a collection function for evaluating the quality of the matrix parameter set.

[0111] The following is a detailed description of step 102.

[0112] In some embodiments, since the relationship between the energy efficiency ratio information of the NPU and different matrix parameter sets cannot be described by a simple analytical formula, the embodiment of the present application constructs a Bayesian distribution model based on Bayesian optimization to characterize the distribution relationship between the matrix parameter set and the NPU energy efficiency ratio, and the Bayesian distribution model includes an acquisition function for evaluating the quality of the matrix parameter set. The tuning device can clarify the acquisition direction from multiple sets of matrix parameter sets according to the acquisition function, avoiding blindly sampling the matrix parameter set and causing energy waste.

[0113] In some embodiments, Figure 5 As shown, Figure 5 yes Figure 2 An optional implementation flowchart of step 102 in the embodiment of the present invention is to construct a Bayesian distribution model based on the work energy efficiency ratio function and the matrix parameter set, including the following steps 401 to 403:

[0114] Step 401: determine corresponding multiple energy efficiency ratio information based on multiple matrix parameter sets and working energy efficiency ratio functions.

[0115] Step 402: Based on multiple energy efficiency ratio information and corresponding matrix parameter sets, construct a mean function characterizing the correlation between multiple energy efficiency ratio information and a covariance function characterizing the correlation between the matrix parameter set and the energy efficiency ratio information.

[0116] Step 403: construct a Bayesian distribution model based on the mean function and the covariance function.

[0117] Steps 401 to 403 are described in detail below.

[0118] In some embodiments, a Bayesian distribution model is constructed to approximate the true objective function. Gaussian Process (GP) is a commonly used proxy model in Bayesian optimization, which describes the probability distribution of the objective function through mean function μ(x) and k(x, x′), as shown in the following formula <2>:

[0119] EER(x)~GP(μ(x),k(x,x))<2>

[0120] Where x = (m, n, k, U m , U n , U k , f) is the input matrix parameter set; GP stands for Gaussian process; μ(x)=E[EER(x)] is the mean function that characterizes the correlation between multiple energy efficiency ratio information; k(x, x′)=E[(EER(x)-μ(x))(EER(x′)-μ(x′))] is the covariance function that characterizes the correlation between the matrix parameter set and the energy efficiency ratio information; E represents the expected value.

[0121] It should be noted that random forests, neural networks, etc. can also be selected as proxy models in the Bayesian optimization process. The specific selection can be made according to the actual situation, and the embodiments of the present application do not limit this. The Bayesian distribution model constructed based on the existing energy efficiency ratio information and the corresponding matrix parameter set can capture the complex interactions between multiple different types of processors, and then consider the impact of the interaction on the energy consumption of the NPU, helping to reduce the exploration of unknown matrix parameter sets in the subsequent sampling process, thereby reducing computing costs and unnecessary energy waste.

[0122] In some embodiments, Figure 6 As shown, Figure 6 This is another optional implementation flow chart of the energy efficiency ratio tuning device of the neural network processor provided in the embodiment of the present application. The acquisition function includes a first acquisition function, and the first acquisition function is determined by the following steps 501 to 503:

[0123] Step 501: Update the mean function based on a preset cumulative distribution function to obtain an updated mean function.

[0124] Step 502: Update the covariance function based on a preset probability density function to obtain a first updated covariance function.

[0125] Step 503: superimpose the updated mean function and the first updated covariance function to construct a first acquisition function for evaluating the quality of the matrix parameter set.

[0126] Steps 501 to 503 are described in detail below.

[0127] In some embodiments, the NPU can obtain corresponding different energy efficiency ratio information when entering the running state based on different matrix parameter sets, and determine the current optimal energy efficiency ratio information from the multiple existing energy efficiency ratio information. Usually, the energy efficiency ratio information with the largest energy efficiency ratio information is selected as the current optimal energy efficiency ratio information. Then, the first acquisition function shown in the following formula <3> is constructed, that is, the expected improvement (EI):

[0128] EI(x) = E[max(0, EER(x)-EER best )]

[0129] =(μ(x)-EER best )Φ(Z)+σ(x)φ(Z) <3>

[0130] Among them, EER best is the current optimal energy efficiency ratio information; E represents the expected value; Z = (μ(x)-EER best ) / σ(x); μ(x) and σ(x) are the predicted mean and standard deviation of the Gaussian process at x, respectively. The predicted mean is calculated by the mean function, and the standard deviation is calculated by the square root of the covariance function; Φ(Z) is the preset cumulative distribution function (CDF), (μ(x)-EER best )Φ(Z) is the updated mean function; φ(Z) is the preset probability density function (PDF), and σ(x)φ(Z) is the first updated covariance function.

[0131] In some embodiments, Figure 7 As shown, Figure 7 This is another optional implementation flow chart of the energy efficiency ratio tuning device of the neural network processor provided in the embodiment of the present application. The acquisition function also includes a second acquisition function, and the second acquisition function is determined by following the steps 601 to 602:

[0132] Step 601, updating the covariance function based on preset hyperparameters to obtain a second updated covariance function.

[0133] Step 602: superimpose the mean function and the second updated covariance function to construct a second acquisition function for evaluating the quality of the matrix parameter set.

[0134] Steps 601 to 602 are described in detail below.

[0135] In some embodiments, the following formula is constructed <4> The second acquisition function shown is the upper confidence limit (UpperConfidence Bound, UCB):

[0136] UCB(X)=μ(x)+k·σ(x) <4>

[0137] Wherein, μ(x) and σ(x) are the predicted mean and standard deviation of the Gaussian process at x, respectively; k·σ(x) is the second updated covariance function; k is a hyperparameter used to control the balance between exploration and utilization. The common value range of k is between 1 and 10. In practical applications, a smaller k value can be used for a smoother objective function, while for a highly uncertain or complex objective function, a larger k value is usually selected to enhance exploration. That is, the specific k value can be set according to the actual situation, and the embodiments of the present application do not limit this.

[0138] Step 103: Perform Bayesian optimization processing on the matrix parameter set using the acquisition function to obtain a target matrix parameter set.

[0139] The following is a detailed description of step 103.

[0140] Furthermore, compared with the traditional method of debugging multiple matrix parameter sets one by one, the embodiment of the present application can use the acquisition function to perform Bayesian optimization processing on the matrix parameter set, and determine the sampling direction from multiple matrix parameter sets based on the Bayesian distribution model, and then sample other matrix parameter sets, so as to achieve the purpose of selecting the most valuable sampling from other matrix parameter sets whose energy efficiency information has not been calculated, thereby reducing repeated debugging of parameters and energy consumption.

[0141] When the sampling function is the first sampling function, the matrix parameter set for the next sampling is selected by maximizing EI, that is, the next matrix parameter set is x * =argmaxEI(x). According to the EI calculation formula, when μ(x)>EER best When σ(x) is large, the EI value is large, indicating that there is a higher expectation of improvement at this point; when σ(x) is large, even if μ(x) is not high, EI may be large, which can encourage the distribution model to explore areas of high uncertainty. In this way, it balances exploration (understanding unknown areas) and utilization (fine search in known good areas) to improve sampling accuracy.

[0142] When the sampling function is the second sampling function, the next sampling point is selected by maximizing UCB(x), that is, x*=argmaxUCB(x).

[0143] Further, input the matrix parameter set into the Bayesian distribution model to obtain the corresponding energy efficiency ratio information. Select the optimal one as the target energy efficiency ratio information from the currently existing energy efficiency ratio information and the energy efficiency ratio information corresponding to the newly sampled matrix parameter set. Then, use the matrix parameter set corresponding to the target energy efficiency ratio information as the target matrix parameter set, and adjust the matrix parameters of each processor according to the target matrix parameter set, thus completing the optimization.

[0144] Step 104, control the neural network processor to enter the operating state based on the target matrix parameter set.

[0145] The following provides a detailed description of Step 104.

[0146] In some embodiments, the existing matrix parameter set and the corresponding energy efficiency ratio information are (m1, n1, k1, U m1 , U n1 , U k1 , f1, EER1), (m2, n2, k2, U m2 , U n2 , U k2 , f2, EER2), the newly sampled matrix parameter set and the corresponding energy efficiency ratio information are (m3, n3, k3, U m3 , U n3 , U k3 , f3, EER3), and EER1 < EER2 < EER3. Then, determine EER3 as the optimal energy efficiency ratio information, and adjust the relevant parameters of the NPU according to (m3, n3, k3, U m3 , U n3 , U k3 , f3). It can be understood that in this state, the NPU has the best energy efficiency ratio and the least energy consumption.

[0147] It can be understood that the optimization method proposed in the embodiments of this application can not only effectively avoid the performance bottleneck caused by improper single computing unit matrix, but also more precisely balance the relationship between the computing speed and energy consumption, thereby achieving more efficient NPU energy efficiency ratio optimization.

[0148] In some embodiments, as Figure 8 shown, Figure 8 is Figure 2 an optional implementation flowchart of Step 103 in

[0149] Step 701, perform optimization calculation on the acquisition function to obtain an updated matrix parameter set, and generate updated energy efficiency ratio information based on the updated matrix parameter set and the working energy efficiency ratio function.

[0150] Step 702, calculate the energy efficiency ratio difference between the updated energy efficiency ratio information and the initial optimal energy efficiency ratio information, update the acquisition function based on the updated energy efficiency ratio information to obtain an updated acquisition function, and update the initial optimal energy efficiency ratio based on the updated energy efficiency ratio information.

[0151] Step 703: Use the updated acquisition function as a new acquisition function to iteratively update the matrix parameter set until the energy efficiency ratio difference is less than a preset threshold difference, and then use the updated matrix parameter set as the target matrix parameter set.

[0152] Steps 701 to 703 are described in detail below.

[0153] In some embodiments, after completing the sampling of a matrix parameter set, the Bayesian distribution model can be updated according to the original matrix parameter set and the corresponding energy efficiency ratio information, as well as the newly sampled matrix parameter set and the corresponding energy efficiency ratio information, and resample other matrix parameter sets whose energy efficiency ratio information is unknown; by repeating the iterative model updating and resampling steps, the optimal output solution is gradually approached, and the accuracy of the target matrix parameter set finally determined as the output is ensured.

[0154] Among them, the initial optimal energy efficiency ratio refers to the best one among multiple energy efficiency ratio information obtained when the NPU enters the working state according to different matrix parameter sets. As the number of iterations increases, the initial optimal energy efficiency ratio information will also change accordingly. Based on the initial optimal efficiency ratio, the acquisition function is optimized to resample other matrix parameter sets to obtain an updated matrix parameter set; then, the updated energy efficiency ratio information corresponding to the updated matrix parameter set is determined according to the working energy efficiency ratio function; the initial optimal energy efficiency ratio information is updated according to the current updated energy efficiency ratio information and the original energy efficiency ratio information, and then a new acquisition function is obtained; then, the new acquisition function is used to sample the matrix parameter set to obtain an updated matrix parameter set again; thereafter, the iterative processing process of the matrix parameter set is repeated until the energy efficiency ratio difference is less than the preset threshold difference, or the preset iteration termination condition is met, and the updated matrix parameter set is used as the target matrix parameter set.

[0155] Among them, after multiple iterations, when the new energy efficiency ratio information is much lower than the existing energy efficiency ratio information value and the improvement of the target energy efficiency ratio information is no longer significant, the optimization process converges and returns the target matrix parameter set. Alternatively, the condition for ending repeated iterations may be that the change between the newly sampled energy efficiency ratio information and the last sampled energy efficiency ratio information is less than a threshold, the maximum number of iterations is reached, the maximum iteration time is reached, etc. The condition for ending repeated iterations can be adaptively adjusted according to actual conditions, and the embodiments of the present application do not limit this. Compared with related technologies, the Bayesian distribution model used in the embodiments of the present application usually only needs to be performed dozens of times or even less times to determine the optimal matrix parameter set of the NPU. By effectively exploring the search space, the computing resources and time costs required to find the global optimal point are greatly reduced, providing strong support for the green and efficient operation of the NPU.

[0156] It is understandable that the tuning method proposed in the embodiment of the present application intelligently samples the matrix parameter space by establishing a Bayesian distribution model. After each sampling, the model is updated and the search range is narrowed, gradually approaching the optimal solution. Compared with traditional grid search or random search, Bayesian optimization can find the global optimal point in a smaller number of evaluations, so it is suitable for optimization problems in high-dimensional and large-scale search spaces.

[0157] Furthermore, since the energy efficiency ratio information needs to be obtained through experiments on real servers, the cost is high and it is difficult to quickly obtain a large number of data points, the fitting methods that rely on a large number of data points in traditional methods, such as neural network methods, cannot be used. In comparison, the tuning method proposed in the embodiment of the present application can reduce the number of adjustments to NPU related parameters while determining the optimal target matrix parameter set, thereby reducing energy consumption.

[0158] In addition, the iterative process of the Bayesian distribution model is adaptive. It can dynamically adjust the next search direction according to the existing sample data, which makes the optimization process more intelligent and can gradually approach the global optimal solution in a complex multivariable environment. At the same time, the Bayesian distribution model supports online optimization, which enables it to maintain high-efficiency energy-efficiency tuning under the ever-changing computing task load, and is therefore suitable for real-time application scenarios.

[0159] In order to enable readers to further understand the beneficial effects of the tuning method proposed in the embodiments of the present application, a summary description is given below.

[0160] The energy efficiency ratio tuning method of the neural network processor proposed in the embodiment of the present application seeks the global optimal solution of energy efficiency ratio among the matrix blocking factor, the loop expansion factor and the current operating frequency of the NPU by Bayesian optimization to achieve a balance between performance and energy efficiency, and has the following expected effects:

[0161] (1) Improve energy efficiency ratio:

[0162] The tuning method proposed in the embodiment of the present application can significantly reduce unnecessary energy consumption by optimizing matrix blocking, loop expansion and frequency adjustment. Unlike the traditional method that simply pursues the highest performance, this tuning method can reduce total energy consumption while ensuring performance, thereby achieving the best energy efficiency ratio. Especially in large-scale deep learning tasks, the optimized NPU nodes can complete the same or even more computing work at lower power consumption. Experiments on artificial intelligence processors show that the energy efficiency ratio of the optimal matrix blocking can be improved by 5% compared with the default matrix blocking; the energy efficiency ratio of the optimal loop expansion can be improved by 13% compared with the default loop expansion; the energy efficiency ratio of the optimal frequency setting can be improved by 3% compared with the default frequency setting. The overall NPU energy efficiency ratio can be improved by more than 20%.

[0163] (2) Reduce memory access overhead:

[0164] The tuning method proposed in the embodiment of the present application uses block matrix multiplication to decompose a large matrix into small blocks for calculation. These small blocks can better match the multi-level cache of the NPU (such as L1 and L2 cache), thereby reducing the number of global memory accesses and reducing the pressure and latency of the memory bandwidth. Reducing global memory accesses not only improves the calculation speed, but also reduces the energy consumption during data transmission, further optimizing the overall energy efficiency.

[0165] (3) Improve computing efficiency:

[0166] The tuning method proposed in the embodiment of the present application further improves the utilization of computing resources through the loop expansion factor. Loop expansion not only reduces the overhead of loop control, but also improves the register reuse rate, making the data flow more efficient in the processor pipeline. This process reduces waiting time and improves computing efficiency, allowing more operations to be completed in the same time, and improving the overall computing power of the NPU node.

[0167] (4)Flexible adaptation to different computing tasks:

[0168] The tuning method proposed in the embodiment of the present application can automatically adjust parameters such as the block factor, loop expansion factor, and current operating frequency according to the characteristics of different computing tasks, so that the optimization method can flexibly adapt to tasks of different scales and types. This means that the NPU node can dynamically adjust its computing mode when facing different loads, achieving dual optimization of computing performance and energy efficiency without manual adjustment and complex pre-configuration.

[0169] (5) Reduce system operating costs:

[0170] The tuning method proposed in the embodiment of the present application improves the energy efficiency ratio, and the NPU nodes in the intelligent computing cluster will consume less electricity, significantly reducing energy consumption when performing large-scale computing tasks. For large-scale data centers and computing clusters, this not only reduces direct power consumption, but also reduces indirect costs related to energy consumption such as cooling and maintenance. Therefore, the application of this method will help data centers achieve energy conservation and consumption reduction while maintaining high-performance computing capabilities, thereby reducing long-term operating costs.

[0171] (6) Supporting Sustainable Development Goals:

[0172] The tuning method proposed in the embodiment of the present application can significantly reduce the carbon emissions per unit computing task by improving the energy efficiency of NPU nodes, provide technical support for the green transformation of data centers and computing clusters, and help promote the development of related industries in a more environmentally friendly direction.

[0173] like Fig. 9 As shown, Fig. 9 It is an optional flow chart of the energy efficiency ratio tuning device of the neural network processor provided in the embodiment of the present application. The tuning device includes the following modules 801 to 804:

[0174] An acquisition module 801 is used to acquire a matrix parameter set of a neural network processor, and determine a working energy efficiency ratio function of the neural network processor based on the matrix parameter set;

[0175] A construction module 802 is used to construct a Bayesian distribution model based on the work energy efficiency ratio function and the matrix parameter set, wherein the Bayesian distribution model includes a collection function for evaluating the quality of the matrix parameter set;

[0176] The Bayesian optimization module 803 is used to perform Bayesian optimization processing on the matrix parameter set using the acquisition function to obtain a target matrix parameter set;

[0177] The target operation module 804 is used to control the neural network processor to enter the operation state based on the target matrix parameter set.

[0178] The energy efficiency ratio tuning method and related equipment of the neural network processor proposed in the present application include obtaining the matrix parameter set of the neural network processor, and determining the working energy efficiency ratio function of the neural network processor based on the matrix parameter set; constructing a Bayesian distribution model based on the working energy efficiency ratio function and the matrix parameter set, wherein the Bayesian distribution model includes an acquisition function for evaluating the quality of the matrix parameter set; the tuning device can clarify the acquisition direction from multiple sets of matrix parameter sets according to the acquisition function, avoiding blindly sampling the matrix parameter set and causing energy waste; then, using the acquisition function to perform Bayesian optimization processing on the matrix parameter set to obtain the target matrix parameter set, thereby achieving the purpose of selecting the most valuable set for sampling from other matrix parameter sets whose energy efficiency ratio information has not been calculated; then, controlling the neural network processor to enter the running state based on the target matrix parameter set. In this way, the neural network processor can efficiently process the current task while improving the energy efficiency ratio information and reducing energy consumption.

[0179] The specific implementation of the tuning device is basically the same as the specific implementation of the above-mentioned tuning method, and will not be repeated here.

[0180] The embodiment of the present application also provides an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the above tuning method when executing the computer program. The electronic device can be any intelligent terminal including a tablet computer, a car computer, etc.

[0181] like Fig.10 As shown, Fig.10 : is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application, the electronic device includes:

[0182] The processor 901 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0183] The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When the technical solution provided in the embodiment of this specification is implemented by software or firmware, the relevant program code is stored in the memory 902, and the processor 901 calls and executes the tuning method of the embodiment of this application;

[0184] Input / output interface 903, used to implement information input and output;

[0185] Communication interface 904, used to realize communication interaction between the device and other devices, which can be realized by wired mode (such as USB, network cable, etc.) or wireless mode (such as mobile network, WIFI, Bluetooth, etc.);

[0186] A bus 905 that transmits information between various components of the device (e.g., the processor 901, the memory 902, the input / output interface 903, and the communication interface 904);

[0187] The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .

[0188] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program, and the computer program implements the above-mentioned tuning method when executed by a processor.

[0189] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0190] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0191] Those skilled in the art will appreciate that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0192] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0193] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.

[0194] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0195] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0196] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0197] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0198] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0199] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store programs.

[0200] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.

Claims

1. A method for optimizing the energy efficiency ratio of a neural network processor, characterized in that: The method comprises: Acquire a matrix parameter set of a neural network processor, and determine a working energy efficiency ratio function of the neural network processor based on the matrix parameter set; Constructing a Bayesian distribution model based on the work energy efficiency ratio function and the matrix parameter set, wherein the Bayesian distribution model includes an acquisition function for evaluating the quality of the matrix parameter set; Using the acquisition function to perform Bayesian optimization processing on the matrix parameter set to obtain a target matrix parameter set; The neural network processor is controlled to enter a running state based on the target matrix parameter set.

2. The method according to claim 1, characterized in that The neural network processor includes a target matrix for performing operation processing, and the target matrix includes a plurality of matrix elements; The step of obtaining a matrix parameter set of a neural network processor includes: Based on a preset first value range, determining at least one blocking factor; Based on a preset second value range, determining a loop expansion factor that matches the blocking factor, wherein the blocking factor is used to divide the target matrix into a plurality of small sub-matrices, and the loop expansion factor represents a merged calculation of a plurality of the matrix elements in the sub-matrices; Determining a current operating frequency of the neural network processor based on a preset third value range; The matrix parameter set is obtained according to the blocking factor, the loop expansion factor and the current operating frequency.

3. The method according to claim 2, characterized in that The neural network processor is provided with a plurality of cache spaces of different levels; After determining the working energy efficiency ratio function of the neural network processor based on the matrix parameter set, the method further includes: Predicting the usage frequency of the sub-matrix to obtain frequency level information, and storing the sub-matrix in the cache space matching the frequency level information; Based on the matrix parameter set, the neural network processor is controlled to enter a working state, so that the neural network processor processes the data in all the cache spaces to obtain the amount of calculation information and the average power information per unit time; Based on the ratio of the calculation amount information to the average power information, energy efficiency ratio information corresponding to the matrix parameter set is obtained.

4. The method according to claim 3, characterized in that The constructing of the Bayesian distribution model based on the working energy efficiency ratio function and the matrix parameter set includes: Based on the plurality of matrix parameter sets and the working energy efficiency ratio function, determining a corresponding plurality of energy efficiency ratio information; Based on the plurality of energy efficiency ratio information and the corresponding matrix parameter sets, constructing a mean function characterizing the correlation between the plurality of energy efficiency ratio information and a covariance function characterizing the correlation between the matrix parameter set and the energy efficiency ratio information; The Bayesian distribution model is constructed according to the mean function and the covariance function.

5. The method according to claim 4, characterized in that The acquisition function comprises a first acquisition function; The first acquisition function is determined by the following steps, which include: Updating the mean function based on a preset cumulative distribution function to obtain an updated mean function; Updating the covariance function based on a preset probability density function to obtain a first updated covariance function; The updated mean function and the first updated covariance function are superimposed to construct the first acquisition function for evaluating the quality of the matrix parameter set.

6. The method according to claim 4, characterized in that The acquisition function also includes a second acquisition function; The second acquisition function is determined by the following steps, which include: Updating the covariance function based on preset hyperparameters to obtain a second updated covariance function; The mean function and the second updated covariance function are superimposed to construct the second acquisition function for evaluating the quality of the matrix parameter set.

7. The method according to claim 1, characterized in that The method of performing Bayesian optimization processing on the matrix parameter set using the acquisition function to obtain a target matrix parameter set includes: Optimizing the acquisition function to obtain an updated matrix parameter set, and generating updated energy efficiency ratio information based on the updated matrix parameter set and the working energy efficiency ratio function; Calculating an energy efficiency ratio difference between the updated energy efficiency ratio information and the initial optimal energy efficiency ratio information, updating the acquisition function based on the updated energy efficiency ratio information to obtain an updated acquisition function, and updating the initial optimal energy efficiency ratio based on the updated energy efficiency ratio information; The updated acquisition function is used as a new acquisition function to iteratively update the matrix parameter set until the energy efficiency ratio difference is less than a preset threshold difference, and the updated matrix parameter set is used as a target matrix parameter set.

8. A device for optimizing energy efficiency ratio of a neural network processor, characterized in that: The device comprises: An acquisition module, used to acquire a matrix parameter set of a neural network processor, and determine a working energy efficiency ratio function of the neural network processor based on the matrix parameter set; A construction module, configured to construct a Bayesian distribution model based on the work energy efficiency ratio function and the matrix parameter set, wherein the Bayesian distribution model includes an acquisition function for evaluating the quality of the matrix parameter set; A Bayesian optimization module, used to perform Bayesian optimization processing on the matrix parameter set using an acquisition function to obtain a target matrix parameter set; A target operation module is used to control the neural network processor to enter an operation state based on the target matrix parameter set.

9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the energy efficiency tuning method of the neural network processor according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the energy efficiency tuning method of the neural network processor according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Efficiency optimization method of processor and mainboard using same

    CN111381663A

  • Block parameter space optimization method of matrix multiplication based on array packaging

    CN112765552A

  • NPU power consumption optimization system and method based on neural network structure

    CN114217688A

  • Multi-intelligent body cooperation routing method and device, and computer storage medium

    CN115955685A

  • NPU power consumption optimization method and system based on neural network structure

    CN116502685A

Cited By

  • Variable mode decomposition method for step frequency continuous wave signal noise reduction

    CN120508751A