Calculation Method, Device, Terminal Device, and Storage Medium for Computing Power
By establishing a pre-training model adapted to different neural network processors on the terminal device, the problem of inaccurate computing power prediction of NPU acceleration cards in the prior art is solved, and accurate computing power prediction of user tasks and deep learning applications are achieved.
Patent Information
- Application Number
- CN202111281679.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-01
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-11-01
AI Technical Summary
The prior art is difficult to accurately predict the computing power of NPU accelerator card to user tasks, especially MLPerf is not applicable to some NPU accelerator cards, and it is unable to effectively predict the computing power of user tasks.
By establishing a pre-established pre-trained model on the terminal device, the model includes a number of target neural network models of different task types. The model is transformed to adapt to a preset neural network processor, including an NPU acceleration card or a combination of CPU and NPU acceleration card, so as to perform model reasoning on user tasks and determine the required computing power information.
It realizes accurate computing power prediction of user tasks on different types of preset neural network processors, solves the problem of inaccurate computing power prediction in the existing technology, and improves the support ability of terminal devices for deep learning applications.
Smart Images

Figure CN114239844B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a method, device, terminal device, and storage medium for calculating computing power. Background Art
[0002] With the rapid development of the field of artificial intelligence, various applications based on deep learning models have been continuously developed. How to efficiently provide intelligent services to users is a concern for IT practitioners. Hardware is one of the more critical issues. Currently, there are many NPU acceleration cards developed for intelligent computing in China, and the computing power of these acceleration cards cannot be simply calculated through hardware data. Moreover, the computing power calculated through hardware data is only an ideal value, and the actual computing power needs to be tested according to specific deep learning applications.
[0003] MLPerf is a set of general benchmarks for measuring and improving the performance of machine learning software and hardware, mainly used to measure the time required to train and infer different neural networks. However, MLPerf is not applicable to some NPU acceleration cards and cannot predict the computing power of user tasks. Summary of the Invention
[0004] The present invention aims to provide a method, device, terminal device, and storage medium for calculating computing power to solve the deficiencies in the prior art. The technical problems to be solved by the present invention are achieved through the following technical solutions.
[0005] In a first aspect, an embodiment of the present invention provides a method for calculating computing power, the method comprising:
[0006] Obtain a user task for which computing power is to be predicted, where the user task includes at least a target task type and a target task volume;
[0007] According to a pre-established pre-training model, perform model inference on the user task of the target task type to determine the computing power information corresponding to the target task volume for executing the user task, where the pre-established pre-training model includes at least multiple target neural network models of different task types, and the target neural network model is obtained by converting a preset neural network processor, where the preset neural network processor includes an NPU acceleration card or a combination of a CPU and an NPU acceleration card.
[0008] Optionally, the pre-established pre-training model is obtained by the following method:
[0009] Obtain training sample sets corresponding to different task types, where the different task types include at least: image classification task, object recognition task, recommendation task, speech recognition task, text recognition task, or reinforcement learning task;
[0010] Train different neural network models using different training sample sets to obtain different initial neural network models;
[0011] According to different types of preset neural network processors, convert the initial neural network models to determine pre-trained models corresponding to the preset neural network processors.
[0012] Optionally, the obtaining of the training sample sets corresponding to different task types includes:
[0013] Obtain the training sample sets corresponding to different task types through the ImageNet database, the COCO database, or the Wikipedia database.
[0014] Optionally, the training different neural network models using different training sample sets to obtain different initial neural network models includes:
[0015] Train the VGG19 model according to the image classification sample set to obtain an initial image classification neural network model;
[0016] Train the yolov3 module according to the object recognition sample set to obtain an initial object recognition neural network model;
[0017] Train the DLRM model according to the recommendation task sample set to obtain an initial recommendation task neural network model;
[0018] Train the RNN-T model according to the speech recognition sample set to obtain an initial speech recognition neural network model;
[0019] Train the BERT model according to the text recognition sample set to obtain an initial text recognition neural network model;
[0020] Train the MINIGO model according to the reinforcement learning sample set to obtain an initial reinforcement learning neural network model.
[0021] Optionally, the converting the initial neural network models according to different types of preset neural network processors to determine pre-trained models corresponding to the preset neural network processors includes:
[0022] Obtain a deep learning sample set;
[0023] Use a deep learning framework to establish a network architecture, where the deep learning framework includes at least one of tensorflow and pytorch;
[0024] According to the deep learning sample set, train the initial neural network models corresponding to different types of preset neural network processors to obtain training results;
[0025] If the training result meets the preset conditions, the initial neural network models corresponding to different types of preset neural network processors are determined as the pre-trained models.
[0026] In a second aspect, an embodiment of the present invention provides a computing device for computing power, and the device includes:
[0027] An acquisition module, configured to acquire a user task for which computing power is to be predicted, where the user task includes at least a target task type and a target task volume;
[0028] A calculation module, configured to perform model inference on the user task of the target task type according to a pre-established pre-trained model, and determine the computing power information corresponding to the target task volume required to execute the user task, where the pre-established pre-trained model includes at least multiple target neural network models of different task types, and the target neural network model is obtained by converting a preset neural network processor, where the preset neural network processor includes an NPU acceleration card or a combination of a CPU and an NPU acceleration card.
[0029] Optionally, the device further includes a training module, and the training module is configured to:
[0030] Acquire training sample sets corresponding to different task types, where the different task types include at least: image classification task, object recognition task, recommendation task, speech recognition task, text recognition task or reinforcement learning task;
[0031] Train different neural network models with different training sample sets to obtain different initial neural network models;
[0032] Convert the initial neural network model according to different types of preset neural network processors, and determine a pre-trained model corresponding to the preset neural network processor.
[0033] Optionally, the training module is configured to:
[0034] Acquire training sample sets corresponding to different task types through the ImageNet database, the COCO database or the Wikipedia database.
[0035] Optionally, the training module is specifically configured to:
[0036] Train the VGG19 model according to the image classification sample set to obtain an initial image classification neural network model;
[0037] Train the yolov3 module according to the object recognition sample set to obtain an initial object recognition neural network model;
[0038] Train the DLRM model according to the recommendation task sample set to obtain an initial recommendation task neural network model;
[0039] Train the RNN-T model according to the speech recognition sample set to obtain an initial speech recognition neural network model;
[0040] Train the BERT model according to the text recognition sample set to obtain an initial text recognition neural network model;
[0041] Train the MINIGO model according to the reinforcement learning sample set to obtain an initial reinforcement learning neural network model.
[0042] Optionally, the training module is specifically configured to:
[0043] Obtain a deep learning sample set;
[0044] Build a network architecture using a deep learning framework, where the deep learning framework includes at least one of tensorflow and pytorch;
[0045] Train the initial neural network models corresponding to different types of preset neural network processors according to the deep learning sample set to obtain a training result;
[0046] If the training result meets the preset conditions, determine the initial neural network models corresponding to different types of preset neural network processors as the pre-trained models.
[0047] In a third aspect, an embodiment of the present invention provides a terminal device, including: at least one processor and a memory;
[0048] The memory stores a computer program; the at least one processor executes the computer program stored in the memory to implement the computing method of the computing power provided in the first aspect.
[0049] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed, it implements the computing method of the computing power provided in the first aspect.
[0050] The embodiments of the present invention have the following advantages:
[0051] The calculation method, device, terminal device, and storage medium for computing power provided by the embodiments of the present invention obtain a user task for which computing power is to be predicted, where the user task includes at least a target task type and a target task volume; according to a pre-established pre-trained model, perform model inference on the user task of the target task type to determine the computing power information corresponding to the target task volume required for executing the user task. The pre-established pre-trained model includes at least multiple target neural network models of different task types, and the target neural network model is obtained by converting a preset neural network processor. By establishing a pre-trained model on the terminal device in the embodiments of the present invention, in this way, when a user task is input, regardless of the type of the preset neural network processor on the terminal device, the computing power of the user task can be predicted. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 is a flowchart of the steps of an embodiment of the calculation method for computing power of the present invention;
[0053] Figure 2 is a flowchart of the steps of another embodiment of the calculation method for computing power of the present invention;
[0054] Figure 3 is a flowchart of the steps of still another embodiment of the calculation method for computing power of the present invention;
[0055] Figure 4 is a flowchart of the steps for establishing the pre-trained model of the present invention;
[0056] Figure 5 is a block diagram of the structure of an embodiment of the calculation device for computing power of the present invention;
[0057] Figure 6 is a schematic diagram of the structure of a terminal device of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present invention will be described in detail below with reference to the drawings and in combination with the embodiments.
[0059] An embodiment of the present invention provides a calculation method for computing power for predicting the computing power of a user task. The execution subject of this embodiment is a calculation device for computing power, which is set on a terminal device. For example, the terminal device includes at least a mobile phone terminal, a tablet terminal, and a computer terminal, etc.
[0060] Referring to Figure 1 , a flowchart of the steps of an embodiment of the calculation method for computing power of the present invention is shown, and the method may specifically include the following steps:
[0061] S101. Obtain a user task for which the computing power is to be predicted, where the user task includes at least a target task type and a target task volume;
[0062] Specifically, when predicting the computing power of a terminal device, simply relying on the hardware device NPU acceleration card on the terminal device for calculation is inaccurate. Some auxiliary software is required for more accurate calculation. Therefore, MLPerf is a set of general benchmarks for measuring and improving the performance of machine learning software and hardware, mainly used to measure the time required for training and inferring different neural networks. The MLPerf test set contains Benchmark sub-items in different fields, mainly including image classification, object recognition, translation, recommendation, speech recognition, sentiment analysis, and reinforcement learning.
[0063] However, MLPerf is not applicable to some domestic NPU (Neural-Network Processing Unit) acceleration cards. These NPU acceleration cards do not support training and can only be used for inference applications. For the pre-trained model for inference, it needs to be converted before use. At the same time, MLPerf does not have the operation results of other types of CPUs (Central Processing Unit / Processor), and it is impossible to compare the computing power differences of different CPU and NPU combination devices. Therefore, the embodiment of the present invention provides a method for calculating computing power. Different types of CPUs and / or NPU acceleration cards are installed on the terminal device, and the terminal device obtains a user task for which the computing power is to be predicted, and the user task includes a target task type and a target task volume.
[0064] Specifically, according to different deep learning fields and commonly used applications of users, determine the target task type of the user task. The specific method is as follows:
[0065] By crawling network news and various information in the field of artificial intelligence, and at the same time conducting a demand survey on users, different task types in the field of deep learning are obtained, including: image classification, target recognition, recommendation, speech, text, and reinforcement learning.
[0066] Exemplarily, the user task is to identify target objects in 100 images.
[0067] S102. According to the pre-established pre-trained model, perform model inference on the user task of the target task type, and determine the computing power information corresponding to the target task volume required to execute the user task. The pre-established pre-trained model includes at least multiple target neural network models of different task types. The target neural network model is obtained by converting a preset neural network processor, where the preset neural network processor includes an NPU acceleration card or a combination of a CPU and an NPU acceleration card.
[0068] Specifically, a pre-trained model is pre-established on the terminal device. The pre-trained model is a target neural network model trained according to different task types. Since a preset neural network processor is installed on the terminal device, where the preset neural network processor at least includes various different types of CPUs and / or NPU acceleration cards. For example, the preset neural network processor can be an NPU acceleration card, or a combination of a CPU and an NPU acceleration card. Therefore, the target neural network model is obtained by converting different CPUs or NPU acceleration cards, and the target neural network model can be recognized by the CPU or NPU.
[0069] After the terminal device obtains the user task input by the user, through the pre-trained model on the CPU and / or NPU of the terminal device, the corresponding neural network model is selected according to the target task type, and the target task amount in the user task is calculated through the corresponding neural network model to obtain the computing power information corresponding to the user task.
[0070] Among them, during the training process of the pre-trained model, by continuously increasing the task amount, different computing power information is calculated. Finally, a pre-trained model that can maximize the utilization of the performance of the acceleration card and select the optimal computing power result in a stable operating state is determined.
[0071] The computing power calculation method provided by the embodiments of the present invention obtains a user task for which the computing power is to be predicted, where the user task at least includes a target task type and a target task amount; according to the pre-established pre-trained model, model inference is performed on the user task of the target task type to determine the computing power information corresponding to the target task amount required to execute the user task. The pre-established pre-trained model at least includes target neural network models of multiple different task types, and the target neural network model is obtained by converting the preset neural network processor. By establishing a pre-trained model on the terminal device in the embodiments of the present invention, in this way, when inputting a user task, regardless of the type of the preset neural network processor on the terminal device, the computing power of the user task can be predicted.
[0072] Another embodiment of the present invention further supplements and explains the computing power calculation method provided in the above embodiment.
[0073] Optionally, the pre-established pre-trained model is obtained through the following method:
[0074] Step A1, obtain a training sample set corresponding to different task types, where the different task types at least include: image classification task, object recognition task, recommendation task, speech recognition task, text recognition task, or reinforcement learning task;
[0075] Step A2: Train different neural network models with different training sample sets to obtain different initial neural network models;
[0076] Step A3: Convert the initial neural network models according to different types of preset neural network processors to determine pre-trained models corresponding to the preset neural network processors.
[0077] Optionally, obtaining training sample sets corresponding to different task types includes:
[0078] Obtain training sample sets corresponding to different task types through the ImageNet database, COCO database, or Wikipedia database.
[0079] Optionally, training different neural network models with different training sample sets to obtain different initial neural network models includes:
[0080] Train the VGG19 model according to the image classification sample set to obtain an initial image classification neural network model;
[0081] Train the yolov3 module according to the object recognition sample set to obtain an initial object recognition neural network model;
[0082] Train the DLRM model according to the recommendation task sample set to obtain an initial recommendation task neural network model;
[0083] Train the RNN-T model according to the speech recognition sample set to obtain an initial speech recognition neural network model;
[0084] Train the BERT model according to the text recognition sample set to obtain an initial text recognition neural network model;
[0085] Train the MINIGO model according to the reinforcement learning sample set to obtain an initial reinforcement learning neural network model.
[0086] Specifically, collect different data sets and construct neural network structures, and the specific methods are as follows:
[0087] In the field of artificial intelligence, different applications have very different requirements for data. Therefore, it is necessary to find specific data sets for each application. At the same time, it is necessary to set up deep neural networks corresponding to the different data sets, that is, sample sets, to give full play to the performance of the acceleration card. In the embodiments of the present invention, sample sets are obtained through data sets such as ImageNet, COCO, and Wikipedia and stored in the data warehouse.
[0088] In the embodiments of the present invention, it is also necessary to construct network models, and construct different deep neural network models for each field, that is, initial neural network models:
[0089] (1) Image Classification - VGG19
[0090] VGG19 (Visual Geometry Group) uses several consecutive 3x3 convolutional kernels instead of the larger convolutional kernels (11x11, 7x7, 5x5) in AlexNet, and contains 19 hidden layers (16 convolutional layers and 3 fully connected layers).
[0091] (2) Object Recognition - YOLO
[0092] Yolo uses the first 52 layers of darknet - 53. Yolov3 is a fully convolutional network that makes extensive use of residual skip connections. And to reduce the negative effect of pooling on gradients, it directly abandons POOLing and uses the stride of conv to achieve downsampling.
[0093] (3) DLRM Deep Learning Recommendation Model
[0094] The DLRM model uses embeddings to process the sparse features representing categorical data, uses an MLP to process the dense features, and then explicitly crosses these features using the statistical techniques in 24. Finally, another MLP post - processes the cross - results to find the event probability.
[0095] (4) Text - BERT (Bidirectional Encoder Representation from Transformers, text training model)
[0096] BERT is a pre - trained language representation model. It adopts a new MLM structure so as to generate deep bidirectional language representations.
[0097] (5) Speech - RNN - T Powerful End - to - End Speech Recognition Framework
[0098] RNN - T enables the model to have outstanding advantages such as end - to - end joint optimization, language modeling ability, and facilitating online speech recognition, making it more suitable for speech tasks.
[0099] (6) Reinforcement Learning - MINIGO
[0100] MINIGO uses reinforcement learning to solve the policy problem. It analyzes the current environment and, based on the existing experience, selects a behavior with higher value, and will receive feedback within a certain period of time.
[0101] As Figure 4 shown, Figure 4It is a flowchart of the steps for establishing the pre-training model of the present invention; optionally, according to different types of preset neural network processors, the initial neural network model is converted to determine the pre-training model corresponding to the preset neural network processor, including:
[0102] Step B1, obtain a deep learning sample set;
[0103] Step B2, establish a network architecture using a deep learning framework, where the deep learning framework includes at least one of tensorflow and pytorch;
[0104] Step B3, train the initial neural network model corresponding to different types of preset neural network processors according to the deep learning sample set to obtain a training result;
[0105] Step B4, if the training result meets the preset conditions, determine the initial neural network model corresponding to different types of preset neural network processors as the pre-training model.
[0106] Specifically, in the embodiments of the present invention, deep learning frameworks such as tensorflow and pytorch are used. Some NPU acceleration cards do not support training, while the inference process supports most deep learning frameworks. Therefore, common deep learning frameworks and NVIDIA graphics cards are used for training. After the training effect reaches the target quality, the pre-training model is retained.
[0107] Figure 2 It is a flowchart of the steps of another embodiment of the computing method of the computing power of the present invention. As Figure 2 shown, the embodiments of the present invention propose an NPU acceleration card computing power test method based on model inference. Different types of CPUs and NPU acceleration cards are installed on a combined device, that is, a terminal device. Among them, the CPU may include an ARM processing chip or an X86 processing chip. A pre-training model is installed on the combined device, where the pre-training model is obtained by converting the initial neural network model through the NPU acceleration technology stack.
[0108] The combined device performs model inference calculations through the obtained pre-training model. Under the condition of continuously changing the input value, the performance of the acceleration card can be maximally utilized, and finally the optimal computing power result in a stable operating state is selected.
[0109] The computing power calculation method provided by the embodiments of the present invention includes defining task types, determining the application fields of deep learning according to actual applications, such as image classification, object recognition, recommendation, speech, text, and reinforcement learning; collecting the data sets required for relevant tasks and designing corresponding network models; using NVIDIA acceleration cards for model training and storing the pre-trained models; using the acceleration stack toolkit to convert the pre-trained models, and using different combinations of CPU and NPU acceleration card devices for model inference and collecting computing power information.
[0110] Figure 3 It is a step flowchart of another embodiment of the computing power calculation method of the present invention. As Figure 3 shown, the computing power calculation method includes:
[0111] S1. Task definition: Determine the corresponding intelligent applications according to different deep learning fields and the applications commonly used by users.
[0112] S2. For different deep learning models, different data sets need to be collected and the corresponding neural network structures need to be constructed.
[0113] S3. Use NVIDIA graphics cards for deep learning model training, train according to the constructed deep neural network using the data set, reach the corresponding target quality, and store the pre-trained models.
[0114] S4. Use the NPU acceleration stack to convert the pre-trained models, and perform inference operations on different combinations of CPU and NPU acceleration card devices to collect computing power information.
[0115] Specifically, different NPU acceleration cards have different acceleration stacks to convert the initial neural network model, aiming to convert the initial neural network model into content that the acceleration card can run, that is, to obtain the pre-trained model.
[0116] First, different combinations of CPU and NPU acceleration cards are used to form specific service devices. Among them, the CPU has different chips with x86 and arm architectures respectively, and the NPU acceleration cards also have various domestic brands; then, the specific acceleration stack of the selected NPU acceleration card is used for conversion; then the model is run. During the computing power calculation process, the input of the model is continuously adjusted to make the best use of the performance of the acceleration card. Finally, the optimal computing power result under the stable operation state is selected.
[0117] It should be noted that for method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the embodiments of the present invention are not limited by the described action sequences, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present invention.
[0118] The computing method of computing power provided by the embodiments of the present invention includes obtaining a user task for which computing power is to be predicted, where the user task at least includes a target task type and a target task volume; according to a pre-established pre-trained model, performing model inference on the user task of the target task type to determine the computing power information corresponding to the target task volume required to execute the user task. The pre-established pre-trained model at least includes target neural network models of multiple different task types, and the target neural network model is obtained by converting a preset neural network processor. By establishing a pre-trained model on the terminal device in the embodiments of the present invention, in this way, when a user task is input, regardless of the type of the preset neural network processor on the terminal device, the computing power of the user task can be predicted.
[0119] Another embodiment of the present invention provides a computing device for computing power, which is used to execute the computing method of computing power provided in the above embodiments.
[0120] Referring to Figure 5 , a structural block diagram of an embodiment of a computing device for computing power according to the present invention is shown. The device may specifically include the following modules: an acquisition module 501 and a calculation module 502, where:
[0121] The acquisition module 501 is used to obtain a user task for which computing power is to be predicted, where the user task at least includes a target task type and a target task volume;
[0122] The calculation module 502 is used to perform model inference on the user task of the target task type according to a pre-established pre-trained model to determine the computing power information corresponding to the target task volume required to execute the user task. The pre-established pre-trained model at least includes target neural network models of multiple different task types, and the target neural network model is obtained by converting a preset neural network processor.
[0123] The computing device for computing power provided by the embodiments of the present invention obtains a user task for which computing power is to be predicted, where the user task at least includes a target task type and a target task volume; according to a pre-established pre-trained model, performs model inference on the user task of the target task type, and determines the computing power information corresponding to the target task volume required to execute the user task. The pre-established pre-trained model at least includes target neural network models of multiple different task types, and the target neural network model is obtained by converting a preset neural network processor. By establishing a pre-trained model on the terminal device in the embodiments of the present invention, in this way, when inputting a user task, regardless of the type of the preset neural network processor on the terminal device, the computing power of the user task can be predicted.
[0124] Another embodiment of the present invention further supplements the computing device for computing power provided in the above embodiment.
[0125] Optionally, the device further includes a training module, and the training module is used for:
[0126] Obtain training sample sets corresponding to different task types, where the different task types at least include: image classification task, object recognition task, recommendation task, speech recognition task, text recognition task, or reinforcement learning task;
[0127] Use different training sample sets to train different neural network models to obtain different initial neural network models;
[0128] According to different types of preset neural network processors, convert the initial neural network model to determine a pre-trained model corresponding to the preset neural network processor.
[0129] Optionally, the training module is used for:
[0130] Obtain training sample sets corresponding to different task types through the ImageNet database, COCO database, or Wikipedia database.
[0131] Optionally, the training module is specifically used for:
[0132] Train the VGG19 model according to the image classification sample set to obtain an initial image classification neural network model;
[0133] Train the yolov3 module according to the object recognition sample set to obtain an initial object recognition neural network model;
[0134] Train the DLRM model according to the recommendation task sample set to obtain an initial recommendation task neural network model;
[0135] Train the RNN-T model according to the speech recognition sample set to obtain an initial speech recognition neural network model;
[0136] Train the BERT model according to the text recognition sample set to obtain an initial text recognition neural network model;
[0137] Train the MINIGO model according to the reinforcement learning sample set to obtain an initial reinforcement learning neural network model.
[0138] Optionally, the training module is specifically configured to:
[0139] Obtain a deep learning sample set;
[0140] Build a network architecture using a deep learning framework, where the deep learning framework includes at least one of tensorflow and pytorch;
[0141] Train the initial neural network models corresponding to different types of preset neural network processors according to the deep learning sample set to obtain training results;
[0142] If the training results meet the preset conditions, determine the initial neural network models corresponding to different types of preset neural network processors as pre-trained models.
[0143] For the apparatus embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For related parts, refer to the partial description of the method embodiment.
[0144] The computing device for computing power provided by the embodiment of the present invention obtains a user task for which computing power is to be predicted, where the user task includes at least a target task type and a target task volume; according to a pre-established pre-trained model, perform model inference on the user task of the target task type to determine the computing power information corresponding to the target task volume required to execute the user task. The pre-established pre-trained model includes at least multiple target neural network models of different task types, and the target neural network model is obtained by converting a preset neural network processor. By establishing a pre-trained model on the terminal device in the embodiment of the present invention, in this way, when inputting a user task, regardless of the type of the preset neural network processor on the terminal device, the computing power of the user task can be predicted.
[0145] Another embodiment of the present invention provides a terminal device for executing the computing power calculation method provided in the above embodiment.
[0146] Figure 6 is a schematic structural diagram of a terminal device of the present invention, as Figure 6 shown, the terminal device includes: at least one processor 601 and a memory 602;
[0147] The memory stores a computer program; at least one processor executes the computer program stored in the memory to implement the computing method of computing power provided in the above embodiments.
[0148] The terminal device provided in this embodiment obtains a user task for which computing power is to be predicted, where the user task includes at least a target task type and a target task volume; according to a pre-established pre-trained model, model inference is performed on the user task of the target task type to determine the computing power information corresponding to the target task volume required to execute the user task. The pre-established pre-trained model includes at least target neural network models of multiple different task types, and the target neural network model is obtained by converting a preset neural network processor. By establishing a pre-trained model on the terminal device in the embodiment of the present invention, in this way, when a user task is input, regardless of the type of the preset neural network processor on the terminal device, the computing power of the user task can be predicted.
[0149] Another embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed, it implements the computing method of computing power provided in any of the above embodiments.
[0150] According to the computer-readable storage medium of this embodiment, a user task for which computing power is to be predicted is obtained, where the user task includes at least a target task type and a target task volume; according to a pre-established pre-trained model, model inference is performed on the user task of the target task type to determine the computing power information corresponding to the target task volume required to execute the user task. The pre-established pre-trained model includes at least target neural network models of multiple different task types, and the target neural network model is obtained by converting a preset neural network processor. By establishing a pre-trained model on the terminal device in the embodiment of the present invention, in this way, when a user task is input, regardless of the type of the preset neural network processor on the terminal device, the computing power of the user task can be predicted.
[0151] It should be noted that the above detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0152] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should also be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0153] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of this application described herein can be implemented in an order other than those illustrated or described herein.
[0154] In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0155] For ease of description, spatial relative terms such as "above", "over", "on the upper surface", "above" etc. may be used here to describe the spatial positional relationship of one device or feature to other devices or features as shown in the figures. It should be understood that spatial relative terms are intended to encompass different orientations in use or operation in addition to the orientation depicted in the figures. For example, if the device in the figures is inverted, a device described as "above" or "over" other devices or structures will then be positioned "below" or "under" other devices or structures. Thus, the exemplary term "above" can include both the orientations of "above" and "below". The device can also be positioned in other different ways, such as rotated 90 degrees or in other orientations, and corresponding interpretations of the spatial relative descriptions used here will be made.
[0156] In the detailed description above, reference has been made to the drawings which form a part hereof. In the drawings, like symbols typically identify like components, unless the context indicates otherwise. The illustrated embodiments described in the detailed description, drawings and claims are not meant to be limiting. Other embodiments may be used and other changes may be made without departing from the spirit or scope of the subject matter presented herein.
[0157] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for calculating computing power, characterized in that, the method includes: Obtain a user task for which computing power is to be predicted, where the user task includes at least a target task type and a target task volume; According to a pre-established pre-training model, perform model inference on the user task of the target task type to determine the computing power information corresponding to the target task volume required to execute the user task, where the pre-established pre-training model includes at least multiple target neural network models of different task types, and the target neural network model is obtained by converting a preset neural network processor, where the preset neural network processor includes an NPU acceleration card or a combination of a CPU and an NPU acceleration card; wherein, the pre-established pre-training model is obtained through the following method: Obtain training sample sets corresponding to different task types, where the different task types include at least: image classification task, object recognition task, recommendation task, speech recognition task, text recognition task, or reinforcement learning task; Use different training sample sets to train different neural network models to obtain different initial neural network models; According to different types of preset neural network processors, convert the initial neural network models to determine pre-training models corresponding to the preset neural network processors.
2. The method according to claim 1, characterized in that, the obtaining of the training sample sets corresponding to different task types includes: Obtain training sample sets corresponding to different task types through the ImageNet database, COCO database, or Wikipedia database.
3. The method according to claim 1, characterized in that, the using different training sample sets to train different neural network models to obtain different initial neural network models includes: Train the VGG19 model according to the image classification sample set to obtain an initial image classification neural network model; Train the yolov3 module according to the object recognition sample set to obtain an initial object recognition neural network model; Train the DLRM model according to the recommendation task sample set to obtain an initial recommendation task neural network model; Train the RNN-T model according to the speech recognition sample set to obtain an initial speech recognition neural network model; Train the BERT model according to the text recognition sample set to obtain an initial text recognition neural network model; Train the MINIGO model according to the reinforcement learning sample set to obtain an initial reinforcement learning neural network model.
4. The method according to claim 3, characterized in that, the converting the initial neural network models according to different types of preset neural network processors to determine pre-training models corresponding to the preset neural network processors includes: Obtain a deep learning sample set; Use a deep learning framework to establish a network architecture, where the deep learning framework includes at least one of tensorflow and pytorch; Train the initial neural network models corresponding to the different types of preset neural network processors according to the deep learning sample set to obtain a training result; If the training result meets the preset conditions, determine the initial neural network models corresponding to the different types of preset neural network processors as the pre-trained models.
5. A computing device for computing power, characterized in that, the device includes: an acquisition module for acquiring a user task for which computing power is to be predicted, where the user task includes at least a target task type and a target task volume; a calculation module for performing model inference on the user task of the target task type according to a pre-established pre-trained model to determine the computing power information corresponding to the target task volume required to execute the user task, where the pre-established pre-trained model includes at least multiple target neural network models of different task types, and the target neural network model is obtained by conversion of a preset neural network processor, where the preset neural network processor includes an NPU acceleration card or a combination of a CPU and an NPU acceleration card; wherein the device further includes a training module, and the training module is used for: acquiring training sample sets corresponding to different task types, where the different task types include at least: image classification tasks, object recognition tasks, recommendation tasks, speech recognition tasks, text recognition tasks, or reinforcement learning tasks; training different neural network models with different training sample sets to obtain different initial neural network models; converting the initial neural network models according to different types of preset neural network processors to determine pre-trained models corresponding to the preset neural network processors.
6. The device according to claim 5, characterized in that, the training module is used for: acquiring training sample sets corresponding to different task types through the ImageNet database, the COCO database, or the Wikipedia database.
7. A terminal device, characterized in that, it includes: at least one processor and a memory; the memory stores a computer program; the at least one processor executes the computer program stored in the memory to implement the computing power calculation method according to any one of claims 1-4.
8. A computer-readable storage medium, characterized in that, a computer program is stored in the computer-readable storage medium, and when the computer program is executed, the computing power calculation method according to any one of claims 1-4 is implemented.
Citation Information
Patent Citations
Neural network automatic training method and device based on cloud platform and model recommendation
CN109376844A
Neural network model determination method and device
CN111144561A