Training method of power consumption prediction model, power consumption prediction method and device
Patent Information
- Application Number
- CN202210804066.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-07
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2042-07-07
AI Technical Summary
在NPU上使用不同的神经网络(Neural-Networks,NN)算子,计算不同尺寸的输入矩阵会消耗不同的功耗
[0029] The power prediction model training method of this application embodiment can train the NPU power prediction model using NPU circuit signal flipping data, thereby improving the accuracy of NPU power prediction. Furthermore, the power prediction method of this application embodiment can automatically search for customized network structures under NPU power constraints and select the required sub-network structure configuration based on the prediction results of the NPU power prediction model.
Smart Images

Figure CN117436497B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence (AI) technology, and in particular to a training method, power prediction method and apparatus for a power prediction model. Background Technology
[0002] In various AI scenarios, massive amounts of computation are required, and ordinary chips may be inefficient. Neural Network Processing Units (NPUs) are particularly adept at processing massive amounts of multimedia data such as video and images. Using an NPU can accelerate neural network computation and improve computational efficiency. Using different Neural Networks (NN) operators on an NPU to compute input matrices of different sizes consumes varying amounts of power. In some scenarios, such as Neural Architecture Search (NAS), it may be necessary to predict the power consumption of NPUs for different network structures in advance. Summary of the Invention
[0003] This application provides a method, apparatus, device, and storage medium for training a power consumption prediction model.
[0004] In a first aspect, embodiments of this application provide a method for training a power consumption prediction model, comprising:
[0005] The training sample set is selected from the circuit signal flipping data of the neural network processor (NPU).
[0006] The initial NPU power consumption prediction model was trained using this training sample set to obtain the trained NPU power consumption prediction model.
[0007] Secondly, embodiments of this application provide a power consumption prediction method, including:
[0008] Searching for neural networks under NPU power consumption constraints yields sub-network structure configurations that satisfy these constraints.
[0009] The power consumption of the NPU configured in this sub-network structure is predicted using an NPU power consumption prediction model.
[0010] The NPU power prediction model is obtained by using the training method of the power prediction model provided in the first aspect above.
[0011] Thirdly, embodiments of this application provide a network search method, including:
[0012] The power consumption of the NPU in the sub-network structure configuration of the neural network is predicted by the power consumption prediction method of the second aspect described above, as implemented in this application.
[0013] The accuracy of the sub-network structure configuration is predicted using an accuracy prediction model.
[0014] The target sub-network structure configuration is determined based on the NPU power consumption and accuracy of each sub-network structure configuration.
[0015] Fourthly, embodiments of this application provide a training apparatus for a power consumption prediction model, comprising:
[0016] The sample selection unit is used to select a training sample set from the circuit signal flip data of the neural network processor (NPU).
[0017] The training unit is used to train the initial NPU power prediction model using the training sample set to obtain the trained NPU power prediction model.
[0018] Fifthly, embodiments of this application provide a power consumption prediction device, comprising:
[0019] The search unit is used to search for neural networks under NPU power consumption constraints to obtain sub-network structure configurations that satisfy the NPU power consumption constraints.
[0020] A power consumption prediction unit is used to predict the NPU power consumption of the sub-network structure configuration using an NPU power consumption prediction model.
[0021] The NPU power consumption prediction model is obtained by using the training device of the power consumption prediction model provided in the fourth aspect above.
[0022] Sixthly, embodiments of this application provide a network search device, including:
[0023] The power prediction device provided in the fifth aspect above;
[0024] The accuracy prediction unit is used to predict the accuracy of the sub-network structure configuration using an accuracy prediction model, and obtain the accuracy of the sub-network structure configuration.
[0025] The sub-network determination unit is used to determine the target sub-network structure configuration based on the NPU power consumption and accuracy of each sub-network structure configuration.
[0026] In a seventh aspect, embodiments of this application provide an electronic device, including: a processor, a memory, a display, and one or more programs, the one or more programs being stored in the memory and configured to be executed by the processor, the programs including instructions for performing the steps in the methods provided in the first, second, or third aspects described above.
[0027] Eighthly, embodiments of this application provide a computer-readable storage medium storing at least one instruction that is executed by a processor to implement the method provided in the first, second, or third aspects described above.
[0028] Ninthly, embodiments of this application provide a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the first, second, or third aspects described above.
[0029] The power prediction model training method of this application embodiment can train the NPU power prediction model using NPU circuit signal flipping data, thereby improving the accuracy of NPU power prediction. Furthermore, the power prediction method of this application embodiment can automatically search for customized network structures under NPU power constraints and select the required sub-network structure configuration based on the prediction results of the NPU power prediction model. Attached Figure Description
[0030] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this application. Wherein:
[0031] Figure 1 This is a flowchart of a power consumption prediction model training method according to an embodiment of this application;
[0032] Figure 2 This is a flowchart of a training method for a power consumption prediction model according to another embodiment of this application;
[0033] Figure 3 This is a flowchart of a power consumption prediction method according to an embodiment of this application;
[0034] Figure 4 This is a flowchart of a power consumption prediction method according to another embodiment of this application;
[0035] Figure 5 This is a flowchart of a power consumption prediction method according to another embodiment of this application;
[0036] Figure 6 This is a structural diagram of a training device for a power consumption prediction model according to an embodiment of this application;
[0037] Figure 7 This is a structural diagram of a training apparatus for a power consumption prediction model according to another embodiment of this application;
[0038] Figure 8This is a structural diagram of a power consumption prediction device according to an embodiment of this application;
[0039] Figure 9 This is a structural diagram of a network search device according to an embodiment of this application;
[0040] Figure 10 This is a structural diagram of a network search device according to another embodiment of this application;
[0041] Figures 11a-11c This is a schematic diagram of an example of Once-For-All;
[0042] Figure 12 This is a diagram illustrating the Once-For-All steps;
[0043] Figure 13 This is a schematic diagram of the training process of the Once-For-All hypernetwork in this application;
[0044] Figure 14 This is an overall framework diagram of this application;
[0045] Figure 15 This is a schematic diagram of the training method for the power consumption prediction model of this application;
[0046] Figure 16 This is a structural block diagram of an exemplary electronic device of this application. Detailed Implementation
[0047] The embodiments of this application are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.
[0048] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0049] As described above in the background section, wearable devices need to display at least some information from third-party applications on the terminal in real time. Here, at least some information includes user interface prompts or user interface content.
[0050] Figure 1This is a flowchart of a training method for a power consumption prediction model according to an embodiment of this application. The training method may include:
[0051] S101. Select a training sample set from the circuit signal flipping data of the NPU.
[0052] S102. The initial NPU power consumption prediction model is trained using the training sample set to obtain the trained NPU power consumption prediction model.
[0053] In this embodiment, circuit signal toggle data, such as the number of circuit signal toggles, can be collected for a specific NPU under one or more neural networks. NPU circuit signal toggle data can also be simply referred to as NPU circuit toggle data or NPU toggle data. A portion of the NPU circuit signal toggle data can be selected as a training sample set for an NPU power consumption prediction model. The NPU power consumption prediction model can also be called an NPU power consumption predictor model or NPU power consumption predictor. An initial structure for the NPU power consumption prediction model can be pre-constructed. Then, the initial NPU power consumption prediction model is trained using the aforementioned training sample set, and the model parameters are updated until the model converges, resulting in the trained NPU power consumption prediction model. The trained NPU power consumption prediction model for a specific NPU can be used to predict the power consumption of various types of neural networks using that NPU.
[0054] In one possible implementation, the NPU power prediction model is a Gradient Boosting Decision Tree (GBDT) model or a Multilayer Perceptron (MLP) model. For example, an initial NPU power prediction model is constructed using a GBDT model; the initial GBDT model is then trained using a training sample set, resulting in the trained NPU power prediction model. Similarly, an initial NPU power prediction model is constructed using an MLP model; the initial MLP model is then trained using a training sample set, resulting in the trained NPU power prediction model.
[0055] For example, the GBDT model is an additive model that sequentially trains a set of Classification and Regression Trees (CART) regression trees, then sums the predictions of all regression trees to obtain a strong learner. Each new tree can fit the negative gradient direction of the current loss function. Therefore, the NPU power prediction model can also be called a machine learning regression model, machine learning regressor, etc.
[0056] For example, an MLP model is formed by connecting multiple perceptron models. In two adjacent perceptron models, the output of the first perceptron model becomes the input of the second. In a multilayer perceptron model, besides the input and output layers of the overall model, the intermediate layers can be called hidden layers. Parameters that need to be calculated in a multilayer perceptron model include the connection weights between all adjacent layers, and the thresholds for all hidden and output layers.
[0057] In one possible implementation, such as Figure 2 As shown, the training sample set is selected from the circuit signal flipping data of the NPU, including:
[0058] S201. The circuit signal toggling data of the NPU is grouped according to each sub-module of the NPU, and each group includes the circuit signal toggling characteristics of the sub-module.
[0059] S202. Sort the circuit signal switching characteristics of the submodule within the group.
[0060] S203. Select the features within each group from the features after sorting within the group.
[0061] In this embodiment, an NPU typically needs to process multiple neural network (NN) operators, with different modules within the NPU responsible for processing different NN operators. Based on the function and calculation method of the NN operators, they are assigned to different sub-modules within the NPU for computation. If the circuit signal inversion data collected from the NPU under multiple neural networks includes m circuit signal inversion features, and there are s important sub-modules within the NPU, then the m circuit signal inversion features can be divided into s groups. For example, if the m features are divided equally into s groups, each group can have m / s features. If m / s is not an integer, it can be rounded up, with the last group having fewer features; or it can be rounded down, with the last group having more features. Alternatively, the m features can be divided into s groups according to the relationship between the features and the sub-modules, resulting in a different number of features in each group. When collecting the circuit signal inversion data from the NPU, if the relationship between the features in the circuit signal inversion data and the NPU's sub-modules can be determined, the features in the circuit signal inversion data can be grouped according to this relationship.
[0062] After grouping, features within each group can be sorted by importance, and features with high importance can be selected. For example, the top 100 features from each group can be selected. This process of filtering features within a group is a feature selection process from coarse-grained to fine-grained, which is beneficial for comprehensively acquiring features related to each sub-module of the NPU. The selected features can be directly used as the training sample set for the NPU power prediction model to train the initial NPU power prediction model. The selected features can also serve as the basis for further filtering to obtain a new training sample set.
[0063] In one possible implementation, the NPU's submodules include at least one of the following: a matrix operation submodule, a vector operation submodule, an accumulation operation submodule, a data storage submodule, and an instruction set storage submodule. For example, the NPU may include a Scheduler Sequence Processor (SSP), a Vector Scalar Processor (VSP), a Matrix Multiply Processor (MMP), and a Partial Sum Accumulator (PSUM). The MMP may include multiple mmp_cores (multiplication matrix processor cores), the PSUM may include multiple psum_cores (partial sum accumulator cores), the VSP may include vsp_core_logic (vector processor core logic) and vsp_core_mem (vector processor core memory), and the SSP may include ssp_inst (scheduler instruction set). The mmp_core is primarily responsible for matrix operations and is a matrix operation submodule. The vsp core logic is primarily responsible for vector and scalar multiplication and accumulation operations and is a vector operation submodule. `psum_core` is primarily responsible for accumulation operations and is a submodule for accumulation operations. `vsp core mem` is primarily used for data storage and is a submodule for data storage. `ssp inst` is primarily used for storing instruction sets and is a submodule for storing instruction sets.
[0064] In one possible implementation, such as Figure 2 As shown, the training sample set is selected from the circuit signal flipping data of the NPU, and also includes:
[0065] S204. Perform unified sorting of the features within multiple groups;
[0066] S205. Select the features required for training from the uniformly sorted features and use them as the training sample set.
[0067] For example, if each of the s groups yields f features, there are a total of s×f features. These s×f features can be reordered according to importance, and then the top-ranked features (e.g., 500 features) can be selected. These reordered features can then be used as the training sample set for the NPU power prediction model to train the initial model. Reordering and selecting features within a group represents a fine-grained to coarse-grained feature selection process, which is beneficial for distinguishing the importance of features in different sub-modules of the NPU.
[0068] Figure 3 This is a flowchart of a power consumption prediction method according to an embodiment of this application. The prediction method may include:
[0069] S301. Search the neural network under the NPU power consumption constraint to obtain the sub-network structure configuration that satisfies the NPU power consumption constraint.
[0070] S302. Use the NPU power consumption prediction model to predict the NPU power consumption of the sub-network structure configuration;
[0071] The NPU power consumption prediction model is obtained by training any of the power consumption prediction models in the above embodiments of this application.
[0072] In this embodiment, NPU power consumption constraints can be pre-defined according to actual needs. Under these constraints, a neural network, such as a supernetwork (or super-network, etc.) trained once and deployed multiple times, is searched. The subnetwork structure configurations within this supernetwork that satisfy the given NPU power consumption constraints may vary. An NPU power consumption prediction model can be used to predict the NPU power consumption of this subnetwork structure configuration.
[0073] In one possible implementation, the NPU power consumption of the sub-network structure configuration is predicted using an NPU power consumption prediction model, which includes: converting the sub-network structure configuration into circuit signal inversion data of the NPU; and inputting the converted circuit signal inversion data into the NPU power consumption prediction model to obtain the NPU power consumption of the sub-network structure configuration.
[0074] In this embodiment, a converter module can be used to convert the sub-network structure configuration into NPU circuit signal inversion data (i.e., the hardware pattern corresponding to the sub-network). Then, the NPU circuit signal inversion data obtained from the sub-network structure configuration conversion is input into an NPU power consumption prediction model to obtain the NPU power consumption of that sub-network structure configuration. For example, for each sub-network structure configuration of a supernetwork found, the NPU circuit signal inversion data obtained from that sub-network structure configuration can be input into the NPU power consumption prediction model to predict the NPU power consumption of that sub-network structure configuration. Alternatively, after finding multiple sub-network structure configurations of a supernetwork, the NPU circuit signal inversion data obtained from each sub-network structure configuration can be input into the NPU power consumption prediction model to predict the NPU power consumption of each sub-network structure configuration. Then, based on the NPU power consumption of each sub-network structure configuration, the sub-network structure configuration with the lowest power consumption can be selected as the target sub-network structure configuration.
[0075] Figure 4 This is a flowchart of a power consumption prediction method according to another embodiment of this application. The prediction method may include: predicting the NPU power consumption of the sub-network structure configuration of the neural network using S301 and S302 in the above-described power consumption prediction method embodiment.
[0076] In one possible implementation, such as Figure 4 As shown, the method also includes:
[0077] S401. Use the accuracy prediction model to predict the accuracy of the sub-network structure configuration and obtain the accuracy of the sub-network structure configuration.
[0078] S402. Determine the target sub-network structure configuration based on the NPU power consumption and accuracy of each sub-network structure configuration.
[0079] In this embodiment, power consumption prediction and accuracy prediction can be combined to select a sub-network structure configuration with lower NPU power consumption and higher accuracy from among various sub-network structure configurations as the target sub-network structure configuration. For example, if there is a sub-network structure configuration with the lowest power consumption and highest accuracy among multiple sub-network structure configurations, this configuration can be selected as the target sub-network structure configuration. Alternatively, if there is no sub-network structure configuration with both the lowest power consumption and highest accuracy among multiple sub-network structure configurations, the configuration with the highest accuracy can be selected from several sub-network structure configurations with lower power consumption; or, the configuration with the lowest power consumption can be selected from several sub-network structure configurations with higher accuracy.
[0080] In one possible implementation, the neural network is a supernetwork, such as... Figure 5 As shown, the method also includes:
[0081] S501. Determine the dynamic search variables based on the static network;
[0082] S502. Construct a dynamic network based on the dynamic search variable;
[0083] S503. The supernetwork is obtained through dynamic network training;
[0084] S504. Determine the sampling method for the subnetworks of the supernetwork;
[0085] S505. Based on the sampling method of the sub-network, the supernetwork is sampled to obtain the training data of the accuracy prediction model.
[0086] In one possible implementation, the dynamic search variables include at least one of the following: the number of channels in each layer, the kernel size, the network depth, and the input image resolution.
[0087] For example, a dynamic network D1 can be constructed using the number of channels t1, kernel size c1, network depth d1, and input image resolution r1. Similarly, a dynamic network D2 can be constructed using the number of channels t2, kernel size c2, network depth d2, and input image resolution r2.
[0088] A supernetwork can be obtained by training a dynamic network using training data from a specific function or scenario. This supernetwork can have information such as the maximum number of channels and the kernel size. After the supernetwork is trained, subnetwork structure configurations and their corresponding accuracy are obtained by randomly sampling from the supernetwork. Different subnetwork structure configurations include one or more different configurations of the number of channels, kernel size, network depth, and input image resolution for each layer.
[0089] In the embodiments of this application, the sub-network sampling method may include random uniform sampling and random independent sampling. For example, the random uniform sampling method may include: randomly generating a number between [0, 1.0] as a product factor for the number of network channels (the product factor is the same for each channel), multiplying this product factor by the maximum number of channels in the supernetwork and taking the integer part, as the number of channels in the sub-network. As another example, in the random independent sampling method, the product factor for each channel is different. It is necessary to randomly generate a random number between [0, 1.0] for each channel, and then multiply the maximum number of channels in each layer of the network by the generated random number and take the integer part, as the number of channels in the sub-network.
[0090] Furthermore, a precision prediction model can be trained using training data. The precision prediction model training process and the process of acquiring training data can be performed on the same device or on different devices. The precision prediction model training process and the aforementioned NPU power consumption prediction model training process can be performed on the same device or on different devices, and there are no restrictions on their execution timing.
[0091] In one possible implementation, the training data includes multiple data pairs, which include subnetwork structure configuration and test accuracy.
[0092] Then, using the data pairs from this training data, an initial accuracy prediction model, such as an MLP model, can be trained to obtain a trained accuracy prediction model. The trained accuracy prediction model can predict the test accuracy of each sub-network configuration.
[0093] Figure 6 This is a structural diagram of a training apparatus for a power consumption prediction model according to an embodiment of this application. The apparatus may include:
[0094] The sample filtering unit 601 is used to filter out the training sample set from the circuit signal flip data of the neural network processor NPU;
[0095] The training unit 602 is used to train the initial NPU power prediction model using the training sample set to obtain the trained NPU power prediction model.
[0096] In one possible implementation, such as Figure 7 As shown, the sample screening unit 601 includes:
[0097] Grouping subunit 701 is used to group the circuit signal toggling data of the NPU according to each submodule of the NPU, and each group includes the circuit signal toggling features of the submodule.
[0098] The first sorting subunit 702 is used to sort the circuit signal flipping characteristics of the submodule within a group;
[0099] The first filtering subunit 703 is used to filter out the features within each group from the features sorted within the group.
[0100] In one possible implementation, the sample screening unit 801 is further configured to:
[0101] The second sorting subunit 704 is used to uniformly sort the features within multiple groups.
[0102] The second screening subunit 705 is used to select the features required for training from the uniformly sorted features, and use them as the training sample set.
[0103] In one possible implementation, the NPU's submodules include at least one of the following: a matrix operation submodule, a vector operation submodule, an accumulation operation submodule, a data storage submodule, and an instruction set storage submodule.
[0104] In one possible implementation, the NPU power consumption prediction model is a gradient boosting decision tree (GBDT) model or a multilayer perceptron (MLP) model.
[0105] Figure 8 This is a structural diagram of a power consumption prediction device according to an embodiment of this application. The device may include:
[0106] Search unit 801 is used to search for neural networks under NPU power consumption constraints to obtain sub-network structure configurations that satisfy the NPU power consumption constraints.
[0107] The power consumption prediction unit 802 is used to predict the NPU power consumption of the sub-network structure configuration using an NPU power consumption prediction model.
[0108] The NPU power consumption prediction model is a model obtained by training a power consumption prediction model using any one of the power consumption prediction models in the embodiments of this application.
[0109] In one possible implementation, the power consumption prediction unit is used to convert the sub-network structure configuration into circuit signal inversion data of the NPU; the converted circuit signal inversion data is input into the NPU power consumption prediction model to obtain the NPU power consumption of the sub-network structure configuration.
[0110] Figure 9 This is a structural diagram of a network search device according to an embodiment of this application. Figure 9 As shown, the device may include: Figure 8 The power prediction device shown includes a search unit 801 and a power prediction unit 802.
[0111] In one possible implementation, the device further includes: an accuracy prediction unit 901, used to perform accuracy prediction on the sub-network structure configuration using an accuracy prediction model, to obtain the accuracy of the sub-network structure configuration.
[0112] In one possible implementation, the device further includes:
[0113] The subnetwork determination unit 902 is used to determine the target subnetwork structure configuration based on the NPU power consumption and accuracy of each subnetwork structure configuration.
[0114] In one possible implementation, the neural network is a supernetwork, such as... Figure 10 As shown, the device also includes:
[0115] The variable determination unit 1001 is used to determine dynamic search variables based on the static network.
[0116] Network building unit 1002 is used to build a dynamic network based on the dynamic search variable;
[0117] The supernetwork training unit 1003 is used to obtain the supernetwork based on the dynamic network training.
[0118] The sampling method determination unit 1004 is used to determine the sampling method of the subnetworks of the supernetwork;
[0119] The sampling unit 1005 is used to sample the supernetwork according to the sampling method of the subnetwork to obtain the training data of the accuracy prediction model.
[0120] In one possible implementation, the dynamic search variables include at least one of the following: the number of channels in each layer, the kernel size, the network depth, and the input image resolution.
[0121] In one possible implementation, the training data includes multiple data pairs, which include subnetwork structure configuration and test accuracy.
[0122] In one possible implementation, the accuracy prediction model is an MLP model.
[0123] The specific functions and examples of each unit and subunit of the apparatus in this disclosure embodiment can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.
[0124] Most Neural Architecture Search (NAS) methods are limited to specific devices or resource-constrained platforms. For different devices, retraining is required on each device. This approach has poor scalability and high computational cost. From the perspective of Train Once and Deploy Many Times (OFA), the goal is to decouple the training and search processes, training a SuperNet that supports different architecture configurations. By selecting a subnetwork from the SuperNet, a specialized subnetwork can be obtained without additional training.
[0125] Once-For-All (OFA) is a novel neural network search solution proposed from the perspective of ease of neural network deployment. This solution designs an Once-For-All network (also known as an Once-For-All supernetwork or OFA supernetwork) that can be directly deployed across different architectures, amortizing training costs. Inference can be performed using only a subset of subnetworks within the Once-For-All supernetwork. The Once-For-All supernetwork can flexibly support different depths, widths, kernel sizes, and resolutions without requiring retraining. A simple example of Once-For-All (OFA) is shown below. Figure 11a As shown: After training an OFA network, multiple specialized subnets can be deployed. Examples include specialized subnets for cloud AI, mobile AI, and micro-AI. These specialized subnets can be deployed directly without retraining. Figure 11b As shown, the design cost of using OFA remains largely unchanged regardless of the number of deployment scenarios. Compared to conventional training deployment schemes, the design cost is significantly reduced. Figure 11c As shown, the horizontal axis represents the measured latency of the searched network on a certain NPU, and the vertical axis represents the classification accuracy on the ImageNet dataset. The closer the performance of the searched network is to the top left corner, the better its performance. OFA allows training once to obtain multiple networks. Conventional training methods, such as MobileNetV3, have the same number of training iterations as the number of networks; for example, training four times yields four networks. Comparatively, OFA offers significantly better performance.
[0126] An exemplary OFA process step is as follows: Figure 12As shown. First, a corresponding dynamic network needs to be constructed based on the original static network (S1201). The constructed dynamic network can include the number of channels in each layer, the size of the convolutional kernel, the network depth, the resolution of the input image, etc. Then, a supernetwork (SuperNet) is trained by inputting training data from the training set or the ground truth (GT, the ground truth value corresponding to the data in the training set) (S1202). This supernetwork has information such as the maximum number of channels and the size of the convolutional kernel. After the supernetwork is trained, subnetworks are randomly sampled from the supernetwork, and the subnetworks are encoded to obtain the subnetwork structure encoding (i.e., the subnetwork structure configuration), and the accuracy value corresponding to the subnetwork is obtained, which is the sample generated for the accuracy predictor (S1203). Based on the training data (i.e., samples) of the accuracy predictor obtained in the above steps, a simple accuracy predictor model (or accuracy prediction model), such as a multi-layer perceptron machine (MLP) (S1204), is constructed and trained until convergence (S1205). Then, given constraints such as floating-point operations (FLOPs), the system searches for subnetwork structure configurations that satisfy the constraints (S1206). Finally, based on the searched subnetwork structure configurations, the corresponding subnetwork structure and weights are obtained (S1207), i.e., the optimal configuration is output.
[0127] In NPU power consumption prediction, the power consumption and time consumption of the neural network operator can be calculated based on the number of signal toggles within the NPU's internal circuitry. Since different neural network operators used on the NPU consume different amounts of power when operating on input matrices of different sizes, it is possible to combine neural network operators to design high-precision and low-power network structures for NPU design.
[0128] In the process of training a supernetwork in a one-for-all manner, if a subnetwork structure is randomly searched for and trained each time, it is quite challenging to train all subnetworks to achieve the same performance as the main network within the same large network model framework. Figure 13The diagram illustrates the training flow of an Once-For-All supernetwork. Typically, the performance of smaller networks decreases to varying degrees compared to larger networks. The training flow can include: first, setting the iteration count to 1 (S1301). Second, obtaining a batch of training data (S1302). Third, randomly sampling an active subnetwork (S1303) and training the active subnetwork using the dataset (S1304). Fourth, determining if all data (the training dataset) is covered (S1305). Covering all data includes: each training epoch (generation training) needs to traverse all data in the training set; that is, during training, the network needs to undergo at least one forward and backward propagation using all data. If all data is not covered, return to step S1301. If all data is covered, check if the iteration count is less than the maximum value (S1306). If the iteration count is less than the maximum value, increment the iteration count by 1; otherwise, terminate the process. Calculating power consumption based on the number of signal flips in the NPU's internal circuitry requires collecting a large number of signal flip counts for different network structures under different input sizes, which limits the design efficiency of software-side AI networks.
[0129] This application presents a customized network optimization method for predicting NPU power consumption based on a regression model. This method is implemented on the Once-For-All framework through an external power prediction interface. By attaching the NPU power prediction method of this application to the Once-For-All computing power interface, it can achieve optimal sub-network search for NPU power consumption optimization on a specific platform.
[0130] Figure 14The diagram shows the overall framework of an embodiment of this application. The upper dashed line represents the Once-For-All subnetwork search method, and the lower modules represent the communication between the regression model and the network search module. The subnetwork search method of the regression model can include acquiring training data, building a baseline model, training a once-for-all supernet for multiple deployments, searching under constraints, retraining the searched model, and inferring the final model. In the search under constraints step, the network structure flipping data of each subnetwork of the supernet (e.g., the number of NPU circuit signal flips corresponding to the subnetwork structure encoding) can be input into the NPU power prediction model through the ARCH model (Architecture Model). The NPU power prediction model can output the corresponding NPU power consumption. The ARCH model includes the NPU hardware response to the AI model, such as latency, power consumption, and memory hierarchy.
[0131] In one example, the training method for the power consumption prediction model is as follows: Figure 15 As shown, the method includes: collecting NPU circuit flip data, such as circuit signal flip data or the number of circuit signal flips (S1501). Dividing the training dataset into blocks according to NPU sub-module instances (S1502), for example, grouping the collected circuit flip data according to one or more important NPU modules such as mmp_core, psum_core, vsp_core_logic, vsp_core_mem, and ssp_inst. Performing feature selection from coarse to fine (S1503), for example, ranking features by importance within each group and selecting highly important samples, such as the top 100 samples within the group. Then, performing further feature selection from fine to coarse (S1504), for example, uniformly ranking the features selected from each sub-module group and further selecting highly important samples. Finally, feeding the selected features into a machine learning regressor to train the NPU power prediction model (or NPU power prediction model). The NPU power prediction model can be a prediction model such as GBDT or MLP. For example, an NPU power predictor model (S1505) was obtained through GBDT supervised training.
[0132] The specific implementation of this application embodiment can be accomplished through the following steps:
[0133] (1) Train the NPU power consumption prediction model.
[0134] (2) Determine dynamic search variables based on static networks, such as the baseline model mentioned above, such as the number of channels in each layer of the network, the size of the convolution kernel, the network depth, and the resolution of the input image.
[0135] (3) Construct a dynamic network based on dynamic search variables.
[0136] (4) Train the supernetwork based on the dynamic network and training data related to the model function.
[0137] (5) Determine the sampling method for subnetworks of the supernetwork.
[0138] (6) Data required for training the sampling precision predictor. For example, select 5000 data pairs according to the sub-network sampling method. For example, the data pair can be represented as [sub-network structure encoding, test precision].
[0139] (7) Train the accuracy predictor. For example, the accuracy predictor model can be an MLP.
[0140] (8) Search for the optimal network architecture configuration under the given NPU power consumption constraints. The search method can be found in the search method described above in Once-For-All. The power consumption of the sub-network architecture configuration can be predicted by the power consumption prediction model output in step (1).
[0141] (9) Generate the optimal subnetwork structure configuration. Furthermore, the weights of the subnetwork structure configuration can be extracted.
[0142] Figure 16 A structural block diagram of an electronic device provided in an exemplary embodiment of this application is shown. It can be implemented as a terminal device in the above embodiments. The electronic device in this application may include one or more of the following components: a processor 1610 and a memory 1620.
[0143] Processor 1610 may include one or more processing cores. Processor 1610 connects to various parts of the electronic device using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 1620, and by calling data stored in memory 1620. Optionally, processor 1610 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). Processor 1610 may integrate one or more of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), Neural-network Processing Unit (NPU), and modem. Specifically, the CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required to be displayed on the touch screen; the NPU is used to implement Artificial Intelligence (AI) functions; and the modem is used to handle wireless communication. It is understandable that the aforementioned modem may not be integrated into the processor 1610, but may be implemented as a separate chip.
[0144] The memory 1620 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 1620 may include a non-transitory computer-readable storage medium. The memory 1620 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 1620 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the various method embodiments described above, etc.; the data storage area may store data created according to the use of the electronic device (such as audio data, phone book, etc.).
[0145] In addition, the electronic device may also include a display component 1630, which may include a display screen for displaying images.
[0146] In addition, those skilled in the art will understand that the structure of the electronic device shown in the above figures does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements. For example, the electronic device may also include radio frequency circuits, input units, sensors, audio circuits, speakers, microphones, power supplies, etc., which will not be described in detail here.
[0147] This application also provides a computer-readable storage medium storing at least one instruction, which is executed by a processor to implement the power prediction model training method or power prediction method as described in the above embodiments.
[0148] This application provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the power prediction model training method or power prediction method provided in the above embodiments.
[0149] Those skilled in the art will recognize that the functions described in the embodiments of this application in one or more of the above examples can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of a computer program from one place to another. Storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0150] The above are merely optional embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A training method for a power consumption prediction model, comprising: Selecting a training sample set from the circuit signal flipping data of a neural network processor (NPU) includes: grouping the circuit signal flipping data of the NPU according to each sub-module of the NPU, wherein the circuit signal flipping data includes multiple circuit signal flipping features, each group includes the circuit signal flipping features of one sub-module of the NPU, and the circuit signal flipping features included in each group are determined according to the relationship between the circuit signal flipping features of the NPU and the sub-module of the NPU at the time of acquisition, and the neural network (NN) operators that the sub-module of the NPU is responsible for processing are allocated according to the function and calculation method of each NN operator; sorting the circuit signal flipping features of the sub-module within each group; and selecting the features within each group from the sorted features. The initial NPU power consumption prediction model is trained using the training sample set to obtain the trained NPU power consumption prediction model.
2. The method according to claim 1, further comprising selecting a training sample set from the circuit signal inversion data of the NPU: Perform unified sorting of features within multiple groups; The features required for training are selected from the uniformly sorted features and used as the training sample set.
3. The method according to claim 1 or 2, wherein the NPU submodules include at least one of the following: a matrix operation submodule, a vector operation submodule, an accumulation operation submodule, a data storage submodule, and an instruction set storage submodule.
4. The method according to claim 1 or 2, wherein, The NPU power consumption prediction model is either the Gradient Boosting Decision Tree (GBDT) model or the Multilayer Perceptron (MLP) model.
5. A power consumption prediction method, comprising: Searching for neural networks under NPU power consumption constraints yields sub-network structure configurations that satisfy the NPU power consumption constraints. The NPU power consumption of the sub-network structure configuration is predicted using an NPU power consumption prediction model. The NPU power prediction model is a model obtained by using the training method of the power prediction model according to any one of claims 1 to 4.
6. The method according to claim 5, wherein the NPU power consumption of the sub-network structure configuration is predicted using an NPU power consumption prediction model, comprising: The sub-network structure configuration is converted into NPU circuit signal inversion data; The converted circuit signal inversion data is input into the NPU power consumption prediction model to obtain the NPU power consumption of the sub-network structure configuration.
7. A web search method, comprising: The power consumption of the NPU in the sub-network structure configuration of the neural network is predicted using the power consumption prediction method of claim 5 or 6. The accuracy of the sub-network structure configuration is obtained by using an accuracy prediction model to predict the accuracy of the sub-network structure configuration. The target sub-network structure configuration is determined based on the NPU power consumption and accuracy of each sub-network structure configuration.
8. The method according to claim 7, wherein the neural network is a supernetwork, and the method further comprises: Determine dynamic search variables based on the static network; Construct a dynamic network based on the dynamic search variables; The supernetwork is obtained based on dynamic network training; Determine the sampling method for the subnetworks of the supernetwork; The supernetwork is sampled according to the subnetwork sampling method to obtain the training data of the accuracy prediction model.
9. The method according to claim 8, wherein, The dynamic search variables include at least one of the following: the number of channels in each layer of the supernetwork, the kernel size, the network depth, and the input image resolution.
10. The method according to claim 8 or 9, wherein, The training data for the accuracy prediction model includes multiple data pairs, which include sub-network structure configuration and test accuracy.
11. The method according to claim 8 or 9, wherein the accuracy prediction model is an MLP model.
12. A training device for a power consumption prediction model, comprising: The sample selection unit is used to select a training sample set from the circuit signal flip data of the neural network processor (NPU). The training unit is used to train the initial NPU power consumption prediction model using the training sample set to obtain the trained NPU power consumption prediction model. The sample screening unit includes: The grouping subunit is used to group the circuit signal toggling data of the NPU according to each submodule of the NPU. Each group includes the circuit signal toggling features of the submodule. The circuit signal toggling data includes multiple circuit signal toggling features. Each group includes the circuit signal toggling features of one submodule of the NPU. The circuit signal toggling features included in each group are determined according to the relationship between the circuit signal toggling features of the NPU and the submodule of the NPU at the time of acquisition. The neural network (NN) operators that the submodule of the NPU is responsible for processing are allocated according to the function and calculation method of each NN operator. The first sorting subunit is used to sort the circuit signal flipping characteristics of the submodule within a group; The first filtering subunit is used to filter out the features within each group from the features sorted within the group.
13. The apparatus according to claim 12, wherein the sample screening unit further comprises: The second sorting subunit is used to uniformly sort the features within multiple groups; The second filtering subunit is used to filter out the features required for training from the uniformly sorted features, as the training sample set.
14. The apparatus according to claim 12 or 13, wherein the submodule of the NPU includes at least one of the following: a matrix operation submodule, a vector operation submodule, an accumulation operation submodule, a data storage submodule, and an instruction set storage submodule.
15. The apparatus according to claim 12 or 13, wherein, The NPU power consumption prediction model is either the Gradient Boosting Decision Tree (GBDT) model or the Multilayer Perceptron (MLP) model.
16. A power consumption prediction device, comprising: The search unit is used to search for neural networks under NPU power consumption constraints to obtain sub-network structure configurations that satisfy the NPU power consumption constraints. A power consumption prediction unit is used to predict the NPU power consumption of the sub-network structure configuration using an NPU power consumption prediction model. The NPU power prediction model is a model obtained by training the power prediction model according to any one of claims 12 to 15.
17. The apparatus according to claim 16, wherein the power consumption prediction unit is configured to convert the sub-network structure configuration into NPU circuit signal inversion data; and input the converted circuit signal inversion data into the NPU power consumption prediction model to obtain the NPU power consumption of the sub-network structure configuration.
18. A network search device, comprising: The power consumption prediction device as claimed in claim 16 or 17; The accuracy prediction unit is used to perform accuracy prediction on the sub-network structure configuration using an accuracy prediction model to obtain the accuracy of the sub-network structure configuration. The sub-network determination unit is used to determine the target sub-network structure configuration based on the NPU power consumption and accuracy of each sub-network structure configuration.
19. The apparatus of claim 18, wherein the neural network is a supernetwork, and the apparatus further comprises: The variable determination unit is used to determine dynamic search variables based on the static network. A network construction unit is used to construct a dynamic network based on the dynamic search variables; A hypernetwork training unit is used to obtain the hypernetwork based on dynamic network training. A sampling method determination unit is used to determine the sampling method of the sub-networks of the supernetwork; The sampling unit is used to sample the supernetwork according to the subnetwork sampling method to obtain training data for the accuracy prediction model.
20. The apparatus according to claim 19, wherein, The dynamic search variables include at least one of the following: the number of channels in each layer of the supernetwork, the kernel size, the network depth, and the input image resolution.
21. The apparatus according to claim 19 or 20, wherein, The training data for the accuracy prediction model includes multiple data pairs, which include sub-network structure configuration and test accuracy.
22. The apparatus according to claim 19 or 20, wherein the accuracy prediction model is an MLP model.
23. An electronic device, characterized in that, The device includes a processor, a memory, a display, and one or more programs, said one or more programs being stored in the memory and configured to be executed by the processor, said programs including instructions for performing the steps of the method as described in any one of claims 1-11.
24. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program that, when executed by a processor, implements the method according to any one of claims 1-11.
25. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions to cause the computer device to perform the method according to any one of claims 1-11.
Citation Information
Patent Citations
Operational circuit of neural network
CN111738427A
Recommendation model processing method and device, electronic equipment and storage medium
CN114117206A
Selecting a subset of training data from a data pool for a power prediction model
US20220138388A1