Method and apparatus for training

By screening the multi-group subtraining data sets of NPU and further screening the total feature data sets, the problem of missing selection of important features in feature screening in the prior art is solved, and the accuracy and efficiency of the NPU power consumption prediction model are improved.

CN116523016BActive Publication Date: 2025-05-27GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210062693.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-19
Publication Date
2025-05-27
Estimated Expiration
2042-01-19

AI Technical Summary

Technical Problem

In the prior art, when training the power consumption prediction model of NPU, the feature screening process may miss the selection of important features, resulting in low model accuracy.

Method used

By performing feature screening on multiple sets of subtraining data sets, multiple sets of feature data sets are obtained, and the total feature data set formed by these feature data sets are further screened to obtain the target feature data set, and the power consumption prediction model is finally trained based on the target feature data set.

Benefits of technology

The accuracy of selected feature data and model accuracy are improved, the probability of missing important features is reduced, and the efficiency of power consumption prediction is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116523016B_ABST
    Figure CN116523016B_ABST
Patent Text Reader

Abstract

The present application provides a training method and apparatus. The method is used to train a power consumption prediction model of an NPU. The NPU includes multiple hardware modules, and the training data set of the power consumption prediction model includes multiple groups of sub-training data sets corresponding to the multiple hardware modules one by one. The method includes: respectively performing feature screening on the multiple groups of sub-training data sets to obtain multiple groups of feature data sets; performing feature screening on the total feature data set formed by the multiple groups of feature data sets to obtain a target feature data set; and training the power consumption prediction model according to the target feature data set. The training method provided by the present application improves the accuracy of the selected feature data and the accuracy of the power consumption prediction model trained using the feature data by performing local screening from coarse to fine and overall screening from fine to coarse on the training data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and specifically relates to a training method and device. Background Art

[0002] In the related art, when training a power consumption prediction model of an NPU, usually all training data is directly screened for features, and then a power consumption model is trained according to the screened feature data set.

[0003] This method may miss important feature data during the process of screening features, resulting in a low accuracy of the finally trained power consumption prediction model. Summary of the Invention

[0004] This application provides a training method and device to improve the accuracy of the trained NPU power prediction model.

[0005] In a first aspect, a prediction method is provided. This method is used to train a power consumption prediction model of an NPU. The NPU includes multiple hardware modules, and the training data set of the power consumption prediction model includes multiple groups of sub-training data sets corresponding one-to-one to the multiple hardware modules. The method includes: respectively screening features for the multiple groups of sub-training data sets to obtain multiple groups of feature data sets; screening features for the total feature data set formed by the multiple groups of feature data sets to obtain a target feature data set; training a power consumption prediction model according to the target feature data set.

[0006] Optionally, in some embodiments, the training data set includes the number of flips of the electrical signal inside the NPU and the power consumption corresponding to the number of flips of the electrical signal.

[0007] Optionally, in some embodiments, the multiple hardware modules include some or all of the following modules: matrix multiplication processor module, partial accumulation processor module, vector data processor calculation unit module, vector data processor storage unit module, and instruction issuing processor module.

[0008] Optionally, in some embodiments, the feature screening adopts a base model selection method and / or a variance selection method.

[0009] Optionally, in some embodiments, the base model of the base model selection method includes the GBDT regression method.

[0010] Second aspect, a training device is provided. The device is used to train a power consumption prediction model of an NPU. The NPU includes multiple hardware modules, and the training data set of the power consumption prediction model includes multiple groups of sub-training data sets corresponding one-to-one to the multiple hardware modules. The device includes: an acquisition unit configured to perform feature screening on the multiple groups of sub-training data sets respectively to obtain multiple groups of feature data sets; a screening unit configured to perform feature screening on the total feature data set formed by the multiple groups of feature data sets to obtain a target feature data set; a training unit configured to train the power consumption prediction model according to the target feature data set.

[0011] Optionally, in some embodiments, the training data set includes the number of flips of the electrical signals inside the NPU and the power consumption corresponding to the number of flips of the electrical signals.

[0012] Optionally, in some embodiments, the multiple hardware modules include some or all of the following modules: matrix multiplication processor module, partial accumulation processor module, vector data processor calculation unit module, vector data processor storage unit module, and instruction issue processor module.

[0013] Optionally, in some embodiments, the feature screening adopts a base model selection method and / or a variance selection method.

[0014] Optionally, in some embodiments, the base model of the base model selection method includes the GBDT regression method.

[0015] Third aspect, a training device is provided, including a memory and a processor. An executable code is stored in the memory, and the processor is configured to execute the executable code to implement the method described in the first aspect.

[0016] Fourth aspect, a computer-readable storage medium is provided, on which an executable code is stored. When the executable code is executed, the method described in the first aspect can be implemented.

[0017] Fifth aspect, a computer program product is provided, including an executable code. When the executable code is executed, the method described in the first aspect can be implemented.

[0018] The embodiment of the present application provides a training method. By performing local screening from coarse to fine and overall screening from fine to coarse on the training data, the accuracy of the selected feature data is improved, and the accuracy of the power consumption prediction model trained using the feature data is also improved. Description of the Drawings

[0019] Figure 1 It is a schematic diagram of the electrical signal flip inside the NPU provided by an embodiment of the present application.

[0020] Figure 2It is a schematic flowchart of a training method provided by an embodiment of the present application.

[0021] Figure 3 It is a schematic flowchart of a training method provided by another embodiment of the present application.

[0022] Figure 4 It is a schematic structural diagram of a training device provided by an embodiment of the present application.

[0023] Figure 5 It is a schematic structural diagram of a training device provided by another embodiment of the present application. Detailed implementation manners

[0024] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments.

[0025] In recent years, artificial intelligence research represented by neural networks has achieved very great results in many fields, and it will also play an important role in people's production and life in the future for a long time.

[0026] A neural-network processing unit (NPU) is a network processor that can be used to implement various neural network-related functions. For example, an NPU can generate a neural network model. Another example is that an NPU can train (or learn) a neural network model. Another example is that an NPU can perform calculations based on the received data to be processed and generate an information signal based on the calculation result, etc.

[0027] A neural network model can be composed of one or more neural network (NN) operators. Neural network operators can include, for example, fully connected operators, convolutional operators, pooling operators, etc.

[0028] Using different neural network operators on a neural network processor to operate matrices of different sizes will result in different power consumptions.

[0029] In the actual design process of a neural network, for each designed network structure, it is necessary to know how much power the network structure consumes on the neural network processor for further evaluation.

[0030] How to combine neural network operators to design a network structure with high accuracy and low power consumption is a crucial step for the design of a neural network processor.

[0031] Currently, the industrial software commonly used in the industry calculates the power consumption of neural network operators based on the number of internal electrical signal flips in the NPU. Multiple gate circuits can be included inside the NPU. The number of internal electrical signal flips in the NPU can refer to the number of flips of the gate circuits inside the NPU.

[0032] However, when using industrial software to calculate power consumption, it is necessary to collect data for a long time to obtain the result. Therefore, the power consumption calculation process is extremely time-consuming and seriously affects the design efficiency of neural networks.

[0033] To improve the efficiency of calculating the power consumption of neural network operators, a power consumption prediction model can be trained using a training dataset.

[0034] For example, the number of flips of the internal electrical signals in the NPU and the power consumption corresponding to the number of flips of the electrical signals can be used as training data to train a power consumption prediction model. When using the trained model for power prediction, only the number of flips of a small number of internal electrical signals within a short period of time needs to be collected as input, and the predicted power consumption can be obtained. Therefore, the efficiency of power consumption prediction can be effectively improved.

[0035] Improving the efficiency of power consumption prediction can further improve the design efficiency of neural networks. That is, for each designed neural network model, power consumption prediction can be completed at a faster speed. Correspondingly, the entire neural network design can also be completed at a faster speed.

[0036] Two methods for constructing prediction models are provided in the related art.

[0037] The first method is to directly model based on the relationship between input data and output data to generate a power consumption prediction model. During the process of designing a neural network on the NPU, various data can be related to the power consumption of the NPU. Using data related to the power consumption of the NPU can generate a power consumption prediction model for the NPU.

[0038] Multiple methods can be used to implement the modeling of the power consumption prediction model. For example, a multilayer perception (MLP) can be used to model the input data to obtain a power consumption prediction model.

[0039] It can be understood that different data selections result in different accuracies of the generated power consumption prediction models. Relatively speaking, since the correspondence between the internal electrical signals in the NPU and the power consumption of the NPU is stronger, using the number of flips of the internal electrical signals in the NPU and the power consumption corresponding to the number of flips as training data results in a higher accuracy of the power consumption prediction model.

[0040] However, there are a very large number of internal circuit signals in the NPU, usually reaching the order of hundreds of millions. For a certain neural network operator (such as a convolution operator), among hundreds of millions of internal electrical signals, only some electrical signals are flipped. That is, for a certain neural network operator, the number of flips of most internal electrical signals is zero.

[0041] Figure 1 is a schematic diagram of the flipping of internal electrical signals in the NPU provided by an embodiment of the present application. Figure 1 The squares in represent multiple gate circuits in the NPU. If Figure 1 the flipping of the electrical signal of the gate circuit in is represented by 1 and the non-flipping of the electrical signal of the gate circuit is represented by 0, Figure 1 the corresponding schematic diagram of the flipping of internal electrical signals in the NPU can be represented as a matrix. Most of the data in this matrix is 0, and only a very small number of data is 1. This data set with most data being 0 can be called a data set with high sparsity.

[0042] Figure 1 is only a schematic diagram of an exemplary data set with high sparsity. In reality, the data formed by the flipping of internal electrical signals in the NPU may have higher sparsity. For example, assume that there are 100 million electrical signals inside an NPU, and in a training data, only 300 electrical signals may have flipped.

[0043] The method of directly modeling data based on the correspondence between input data and output data has high requirements for data. Since the number of flips of internal electrical signals in the NPU is a data set with high sparsity, the data distribution varies greatly, and there are many outliers and abnormal values. Therefore, when directly using this data for modeling, it is difficult to generate the power consumption prediction model and it is difficult to converge.

[0044] In addition, since the model generation method is closely related to the provided data, providing different input and output data will result in a large difference in the accuracy of the fitted power consumption prediction model. Therefore, the accuracy of the model trained using this method cannot be guaranteed.

[0045] To solve the above problems, the related art also provides a second method. That is, first perform feature screening on the training data to select important features. After selecting the important feature data, use the important feature data for model training to obtain the power consumption prediction model of the NPU.

[0046] Feature screening can reduce the dimension of the training data, can reduce the difficulty of the learning task, and at the same time improve the efficiency of the model.

[0047] There are various methods for feature screening. For example, the variance or correlation coefficient of each feature can be calculated. A certain number of data are selected as feature data based on the variance or correlation coefficient of the data. The larger the variance or correlation coefficient of a data, the more discriminative and representative the data is.

[0048] After obtaining the target feature data set through screening, various different methods can be used to train the power consumption prediction model. For example, the gradient boosting decision tree (GBDT) method can be used for supervised training to obtain the power consumption prediction model of the NPU.

[0049] The second method directly performs feature screening on all training data. For the training data used to generate the NPU power consumption prediction model, this screening process is too rough, and many important features will be missed, resulting in inaccurate final fitting. That is, the accuracy of the power consumption prediction model generated by the second method is not high.

[0050] In addition, when using the number of internal electrical signal flips of the NPU and the NPU power consumption corresponding to the number of flips as training data, the amount of data is extremely large (for example, the training data may be a 1000 * 100 million matrix). It is difficult to directly put this data into memory for training. And even if it is put into memory, it has extremely high requirements for the memory of the computing device (such as a computer, server, etc.) to screen all features of all samples at one time, which is difficult to achieve.

[0051] In view of this, the present application provides a training method and device to improve the accuracy of the trained NPU power prediction model.

[0052] Figure 2 It is a schematic flowchart of the training method provided by an embodiment of the present application. This method can be used to train the power consumption prediction model of the NPU. The NPU may include multiple hardware modules.

[0053] The specific division method and division granularity of the hardware modules inside the NPU can be selected according to the specific design of the NPU and the training requirements of the NPU power consumption model. The present application does not limit the specific division method and division granularity of the hardware modules of the NPU.

[0054] For example, the multiple hardware modules may include some or all of the following modules: matrix multiplication processor module (or mmp_core), partial sum processor module (or psum_core), vector data processor calculation unit module (or vsp_core_logic), vector data processor storage unit module (or vsp_core_mem), and instruction issuing processor module (or ssp_inst), etc.

[0055] In some embodiments, the hardware module may also select a secondary divided module. That is, on the basis of the above-mentioned divided hardware modules, each hardware module may be further divided into multiple sub-modules. In other words, each hardware module (such as a matrix multiplication processor module) may also include multiple sub-modules.

[0056] For a module on the NPU (eg, a matrix multiplication processor module), the signals on each submodule express or implement the same function, so the submodule can also be called an instance.

[0057] The training data set used to train the power consumption prediction model of the NPU may include a plurality of groups of sub-training data sets corresponding one-to-one to the plurality of hardware modules.

[0058] The specific module division of multiple hardware modules can be selected as needed. For example, the primary hardware module described above (such as the matrix multiplication processor module) can be selected, or the secondary hardware module described above (such as the submodule on the matrix multiplication processor module) can be selected.

[0059] A variety of data may be used as training data as long as a relationship can be established between the data and the power consumption of the NPU.

[0060] For example, the characteristics of the neural network and the power consumption of the NPU corresponding to the characteristics may be used as training data. The characteristics of the neural network may include the size of the convolution kernel, the number of hidden layers, etc.

[0061] For another example, the physical properties of the NPU and the power consumption of the NPU corresponding to the physical properties may be used as training data. The physical properties of the NPU may include the number of flips of the electrical signal inside the NPU.

[0062] This application does not limit the specific data format and acquisition method of the training data set. For example, when the number of NPU internal electrical signal flips and the NPU power consumption corresponding to the flip times are used as the training data set, the industrial software mentioned above can be used to obtain the data.

[0063] Figure 2 The method shown includes steps S210 to S230, and each step is introduced below.

[0064] In step S210, feature screening is performed on the multiple groups of sub-training data sets respectively to obtain multiple groups of feature data sets.

[0065] Multiple sets of sub-training data sets can be obtained in a variety of ways. For example, the training data set can be divided into blocks according to the hardware modules of the NPU to obtain multiple sets of sub-training data sets corresponding to each hardware module. For another example, the sub-training data set corresponding to each module can be directly obtained.

[0066] The feature screening of multiple groups of sub-training data sets can be carried out in various ways. For example, the feature screening can be carried out by the base model selection method and / or the variance selection method. The base model of the base model selection method can include the GBDT regression method, etc.

[0067] After selecting the feature screening method, the screening strategy of the features can be selected according to actual needs.

[0068] For example, the top K features with the highest importance can be selected as the screened feature data set. The parameters for judging the feature importance are different for different screening methods. For example, when the variance selection method is selected, since the feature with a larger variance has a higher distinctiveness, the larger the variance of the feature, the higher its importance.

[0069] Another example is that an importance threshold can be set. The threshold can be selected according to actual needs. When performing feature screening, the features greater than the threshold can be selected as the screened feature data set.

[0070] Feature screening of the training data according to the granularity of the hardware modules of the NPU reduces the requirements for computer memory in the feature selection process. In the case of the same number of features, by means of feature screening of the sub-modules, the number of samples that can be processed simultaneously is greatly increased. Correspondingly, the probability of missing important features is greatly reduced. It can be understood that the finer the division granularity of the hardware modules of the NPU, the higher the accuracy of the selected features. In actual selection, the division or selection of the sub-modules can be adjusted according to the accuracy requirements.

[0071] In step S220, feature screening is performed on the total feature data set formed by multiple groups of feature data sets to obtain the target feature data set.

[0072] The total feature data set can be obtained in various ways. For example, the multiple groups of feature data obtained in step S210 can be merged first to obtain the total feature data set.

[0073] After obtaining the total feature data set, feature screening can be performed on the total feature data set to obtain the target feature data set. For example, importance analysis can be performed on the total feature data set to select an appropriate number of features as the target feature data set. The method of feature screening can refer to the previous text and will not be elaborated here.

[0074] Through the secondary feature screening from fine to coarse, the fineness and accuracy of feature selection are improved. The finally selected target feature data set has better representativeness.

[0075] In step S230, a power consumption prediction model is trained according to the target feature data set.

[0076] The power consumption prediction model of the NPU can be trained in various ways. For example, the GBDT regression tree can be used to perform supervised training on the target feature data set to obtain the power consumption prediction model.

[0077] The training method provided in this application improves the accuracy of the selected feature data and the precision of the power consumption prediction model trained using this feature data by performing local screening from coarse to fine and overall screening from fine to coarse on the training data.

[0078] The following Figure 3 introduces the prediction method provided in this application with a specific embodiment. Figure 3 is a schematic flowchart of the training method provided in the embodiment of this application. Figure 3 It includes steps S301 to S305.

[0079] As Figure 3 shown, in step S301, the NPU electrical signal flip data is collected. There are usually many gate circuits inside the NPU. The NPU internal electrical signal flip data can refer to the flip data of the gate circuits inside the NPU.

[0080] Since using the number of NPU internal electrical signal flips and the corresponding NPU power consumption of the number of flips as training data results in a higher-precision NPU power consumption prediction model. Therefore, in this embodiment, the number of NPU internal electrical signal flips and the corresponding NPU power consumption are selected as the training data set for training the prediction model.

[0081] In step S302, the data is split by module. The hardware modules included in the NPU may have different divisions according to different designs.

[0082] Table 1 is the NPU internal module division table provided in the embodiment of this application.

[0083] Table 1

[0084]

[0085] As shown in Table 1, the NPU can internally include 5 hardware modules: matrix multiplication processor module, partial accumulation processor module, vector data processor calculation unit module, vector data processor storage unit module, and instruction issue processor module.

[0086] Each hardware module (such as the matrix multiplication processor module) can also include multiple sub-modules. In some embodiments, the signal performances of the multiple sub-modules of a hardware module are the same. Therefore, the sub-module can also be referred to as an instance.

[0087] Different hardware modules may include different numbers of sub - modules. As shown in Table 1, in this embodiment, the matrix multiplication processor module may include 64 sub - modules, the partial accumulation processor module may include 8 sub - modules, the vector data processor calculation unit module may include 8 sub - modules, the vector data processor storage unit module may include 8 sub - modules, and the instruction issuing processor module may include 1 sub - module.

[0088] In this embodiment, the training data can be divided according to the granularity of the sub - modules of each module of the NPU. That is, first, the training data is divided into five groups of sub - training data sets according to five modules. Then, each group of sub - training data sets is divided according to the sub - modules of each module, and multiple groups of sub - training data sets are obtained respectively. Or directly divide the training data set according to the granularity of the sub - modules to obtain the sub - training data sets on each sub - module.

[0089] As shown in Table 1, the NPU hardware partitioning method provided in this embodiment can divide the training data set into 89 groups of sub - training data sets.

[0090] In step S303, feature screening for each module. When performing feature screening on the sub - training data sets of each module, each sub - training data set can be separately fed into a base model with a penalty term. The base model can, for example, select the GBDT regression method.

[0091] The importance of each feature can be calculated according to the GBDT base model. In this embodiment, the feature can refer to selecting representative gate circuits from hundreds of millions of gate circuits inside the NPU and obtaining the electrical signal flip data of these representative gate circuits as feature data.

[0092] It should be understood that the higher the discriminability and representativeness of the data, the higher the accuracy of the model trained using this data. In the process of feature screening from coarse to fine and from the whole to the module, the training data is divided according to the granularity of the sub - modules, and the amount of data contained in each module decreases. In the process of feature selection, the memory requirement is reduced, and the number of samples that can be processed at one time increases. At the same time, by performing feature screening for each module, the probability of missing important features is reduced.

[0093] In step S304, merge the module features and perform re - screening of the features. The feature data sets obtained from each sub - module can be merged to obtain the total feature data set. Perform re - screening of the features on the total feature data set to obtain the target feature data set. This feature screening can still use the GBDT regression method.

[0094] An appropriate number of features can be selected as the target feature data set according to the requirements for the accuracy of the prediction model.

[0095] Through the second feature screening from fine to coarse and from module to whole, the obtained target feature dataset is more delicate, more discriminative and representative.

[0096] In step S305, GBDT supervised training is performed. The selected features can be supervised and trained using GBDT regression trees to obtain a power consumption prediction model.

[0097] The accuracy of the generated power consumption model can be adjusted by adjusting the parameters of feature screening in steps S303 and S304.

[0098] For example, by selecting different levels of sub-modules for the division of training data, feature data with different accuracies can be obtained.

[0099] Also, for example, the accuracy of the screened features can be adjusted by changing the feature screening rules. Specifically, the accuracy of the screened features can be adjusted by changing (increasing or decreasing) the number of selected feature data.

[0100] As described above in conjunction with Figures 1 to 3 , the method embodiments of the present disclosure have been described in detail. Below, in conjunction with Figures 4 to 5 , the apparatus embodiments of the present disclosure will be described in detail. It should be understood that the descriptions of the method embodiments and the apparatus embodiments correspond to each other. Therefore, for the parts not described in detail, reference can be made to the previous method embodiments.

[0101] Figure 4 It is a schematic structural diagram of a training apparatus provided by an embodiment of the present application. Figure 4 The training apparatus 400 can be used to train a power consumption prediction model for the NPU. The NPU can include multiple hardware modules. The training dataset of the power consumption prediction model can include multiple groups of sub-training datasets corresponding one-to-one to the multiple hardware modules.

[0102] As Figure 4 shown, the training apparatus 400 can include an acquisition unit 410, a screening unit 420, and a training unit 430. Each module will be introduced below.

[0103] The acquisition unit 410 can be configured to perform feature screening on multiple groups of sub-training datasets respectively to obtain multiple groups of feature datasets.

[0104] The screening unit 420 can be configured to perform feature screening on the total feature dataset formed by multiple groups of feature datasets to obtain a target feature dataset.

[0105] The training unit 430 can be configured to train a power consumption prediction model according to the target feature dataset.

[0106] The NPU power consumption prediction model trained by the training device provided in the embodiments of the present application reduces the probability of missing important features and improves the representativeness of the selected features through two feature screenings from coarse to fine and from fine to coarse. Therefore, the power consumption prediction model trained using the target feature data has high accuracy and high stability.

[0107] Optionally, in some embodiments, the training dataset may include the number of flips of the electrical signals inside the NPU and the power consumption corresponding to the number of flips of the electrical signals.

[0108] Since the relationship between the number of flips of the internal electrical signals of the NPU and the power consumption is the closest, using the number of flips of the internal electrical signals of the NPU and the power consumption corresponding to the number of flips of the electrical signals as training data to train the power consumption prediction model can further improve the accuracy of the power consumption prediction model.

[0109] Optionally, in some embodiments, the multiple hardware modules include some or all of the following modules: matrix multiplication processor module, partial accumulation processor module, vector data processor calculation unit module, vector data processor storage unit module, and instruction issuing processor module.

[0110] Dividing the data by the hardware modules of the NPU makes the margins between the data clearer and the training samples more accurate.

[0111] Optionally, in some embodiments, the feature screening adopts the base model selection method and / or the variance selection method.

[0112] The base model selection method and the variance selection method are easy to implement, can reduce the debugging steps, and improve the model generation efficiency.

[0113] Optionally, in some embodiments, the base model of the base model selection method includes the GBDT regression method.

[0114] Figure 5 It is a schematic structural diagram of the training device provided in the embodiments of the present application. The training device 500 may be a computer, a server, etc. The device 500 may include a memory 510 and a processor 520. The memory 510 may be used to store executable code. The processor 520 may be used to execute the executable code stored in the memory 510 to implement the steps in the various methods described above. In some embodiments, the device 500 may further include a network interface 530, and the data exchange between the processor 520 and external devices may be implemented through the network interface 530.

[0115] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a Digital Video Disc (DVD)), or a semiconductor medium (such as a Solid State Disk (SSD)), etc.

[0116] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments of the present application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0117] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in an electrical, mechanical, or other form.

[0118] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0119] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, may also exist physically alone for each unit, or two or more units may be integrated in one unit.

[0120] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A training method, characterized in that, the method is used to train a power consumption prediction model of an NPU, the NPU includes multiple hardware modules, and the training data set of the power consumption prediction model includes multiple groups of sub-training data sets corresponding one-to-one to the multiple hardware modules, the method includes: performing feature screening on the multiple groups of sub-training data sets respectively to obtain multiple groups of feature data sets; performing feature screening on the total feature data set formed by the multiple groups of feature data sets to obtain a target feature data set; training the power consumption prediction model according to the target feature data set.

2. The method according to claim 1, characterized in that, the training data set includes the number of flips of the electrical signal inside the NPU and the power consumption corresponding to the number of flips of the electrical signal.

3. The method according to claim 1, characterized in that, the multiple hardware modules include some or all of the following modules: matrix multiplication processor module, partial accumulation processor module, vector data processor calculation unit module, vector data processor storage unit module, and instruction issuing processor module.

4. The method according to claim 1, characterized in that, the feature screening adopts the base model selection method and / or the variance selection method.

5. The method according to claim 4, characterized in that, the base model of the base model selection method includes the GBDT regression method.

6. A training device, characterized in that, the device is used to train a power consumption prediction model of an NPU, the NPU includes multiple hardware modules, and the training data set of the power consumption prediction model includes multiple groups of sub-training data sets corresponding one-to-one to the multiple hardware modules, the device includes: an acquisition unit configured to perform feature screening on the multiple groups of sub-training data sets respectively to obtain multiple groups of feature data sets; a screening unit configured to perform feature screening on the total feature data set formed by the multiple groups of feature data sets to obtain a target feature data set; a training unit configured to train the power consumption prediction model according to the target feature data set.

7. The device according to claim 6, characterized in that, the training data set includes the number of flips of the electrical signal inside the NPU and the power consumption corresponding to the number of flips of the electrical signal.

8. The device according to claim 6, characterized in that, the multiple hardware modules include some or all of the following modules: matrix multiplication processor module, partial accumulation processor module, vector data processor calculation unit module, vector data processor storage unit module, and instruction issuing processor module.

9. The device according to claim 6, characterized in that, the feature screening adopts the base model selection method and / or the variance selection method.

10. The device according to claim 9, characterized in that, the base model of the base model selection method includes the GBDT regression method.

11. A training device includes a memory and a processor, an executable code is stored in the memory, and the processor is configured to execute the executable code to implement the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Core scheduling method and terminal

    CN109937410A

  • GPU power consumption estimation method and system of computer platform and medium

    CN111427750A