A method and device for constructing a neural network model
By obtaining performance indicators on the target chip and adjusting the model generator, the compatibility problem of neural network structure in different chip environments is solved, and chip utilization and operation efficiency are improved.
Patent Information
- Application Number
- CN202080104556.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-30
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2040-07-30
AI Technical Summary
In the prior art, neural network structures have poor compatibility in different chip environments, resulting in problems such as excessive time consumption and low chip utilization.
By obtaining the performance indicators of the neural network model running on the target chip and adjusting the model generator based on these indicators, a neural network model with better performance on the target chip is built.
Improves the hardware performance indicators of the neural network model on the target chip, and improves compatibility and efficiency.
Smart Images

Figure CN116261729B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and in particular to a method and device for constructing a neural network model. Background Art
[0002] In recent years, deep neural networks have achieved outstanding results in processing and analyzing various media signals such as images, videos, and voices. A neural network with good performance often has a sophisticated network structure, which requires human experts with superb skills and rich experience to spend a lot of effort to design.
[0003] The structural search of neural networks, that is, the construction of neural network models, has changed this manual design mode, automatically searching for neural network structures, and obtaining neural network structures with excellent performance, achieving excellent results in tasks such as image recognition, image semantic segmentation, and natural language processing.
[0004] In traditional structure search, the neural network structure search is trained in a certain chip environment according to the target indicators of the task (for example, the model test accuracy of applications such as image classification and image segmentation). When the neural network structure search trained based on a certain chip environment is used in other chip environments, due to the difference in chip parameters, compatibility issues will arise during the operation of the neural network structure search, such as excessive time consumption, low chip utilization, etc. Summary of the invention
[0005] The embodiment of the present application provides a neural network model construction method and device thereof, which is used to generate a neural network model using a model generator, by obtaining a first theoretical performance indicator of a first neural network model when it is running on a target chip, and adjusting the corresponding weights in the first model generator according to the first theoretical performance indicator, thereby constructing a neural network model with better hardware performance indicators when running on the target chip.
[0006] A first aspect of an embodiment of the present application provides a method for constructing a neural network model.
[0007] The neural network model building device builds a first neural network model through a first model generator preset in the neural network model building device, and the first neural network model is built by the first model generator based on various building units.
[0008] After the first model generator constructs the first neural network model, the neural network model construction device obtains a first performance indicator of the first neural network model when the first neural network model runs on the target chip according to the first neural network model.
[0009] The neural network model construction device adjusts the first model generator according to the first performance metric to obtain a second model generator. After obtaining the second model generator, the neural network model construction device constructs a second neural network model according to the second model generator. The second performance metric of the second neural network model is better than the first performance metric, that is, the performance metric of the second neural network model when running on the target chip is better than the performance metric of the first neural network model when running on the target chip.
[0010] In the embodiments of the present application, by obtaining the first performance metric of the first neural network model when running on the target chip and adjusting the first model generator according to the first performance metric, a second neural network model with better hardware performance metric when running on the target chip is constructed.
[0011] Based on the neural network model construction method of the first aspect of the embodiments of the present application, in a possible implementation
[0012] After the first model generator constructs the first neural network model, the neural network model construction device obtains the theoretical performance metric of the first neural network model. The first theoretical performance metric represents the theoretical value of the performance metric of the first neural network model when running on the target chip.
[0013] The neural network model construction device adjusts the first model generator according to the first theoretical performance metric to obtain a second model generator. After obtaining the second model generator, the neural network model construction device constructs a second neural network model according to the second model generator. The second neural network model has a second theoretical performance metric, that is, the theoretical value of the performance metric of the second neural network when running on the target chip is better than the first theoretical performance metric.
[0014] In the embodiments of the present application, by obtaining the first theoretical performance metric of the first neural network model when running on the target chip and adjusting the first model generator according to the first theoretical performance metric, a second neural network model with better hardware performance metric when running on the target chip is constructed.
[0015] Based on the neural network model construction method of the first aspect of the embodiments of the present application, in a possible implementation, after the neural network model construction device constructs the second neural network model through the second model generator, the neural network model construction device further obtains a first measured performance metric. The first measured performance metric represents the measured value of the performance metric of the second neural network model when running on the target chip, that is, the measured performance metric of the target chip obtained after the second neural network model runs on the target chip.
[0016] The neural network model construction device adjusts the corresponding weight factors in the second model generator according to the first measured performance index to obtain a third model generator. After obtaining the third model generator, the neural network model construction device constructs a third neural network model through the third model generator, and the second measured performance index of the third neural network model, that is, the measured value of the performance index when the third neural network model runs on the target chip, is better than the first measured performance index.
[0017] In the embodiments of the present application, by running the second neural network in the actual target chip and obtaining the corresponding measured performance index, and then adjusting the second model generator according to the measured performance index, a neural network model more suitable for the target chip can be obtained.
[0018] Based on the neural network model construction method in the first aspect of the embodiments of the present application, in a possible implementation manner, before the neural network model construction device obtains the first measured performance index, the neural network model construction device also trains the second neural network model to obtain a fourth neural network model. After obtaining the fourth neural network model, the neural network model construction device obtains the model performance of the fourth neural network model, and adjusts the second model generator according to the model performance of the fourth neural network and the first measured performance index to obtain a third model generator.
[0019] In the embodiments of the present application, by adjusting the second model generator according to the model performance of the fourth neural network model after training the second neural network model and the first measured performance index, it is possible to better improve the performance index of the neural network model when running on the target chip while ensuring the same model performance of the neural network model generated by the adjusted third model generator.
[0020] Based on the neural network model construction method in the first aspect of the embodiments of the present application, in a possible implementation manner, the neural network model construction device obtains a first theoretical performance index through a performance evaluation tool, and the theoretical performance evaluation tool includes a calculation function for calculating the first neural network model to obtain the first theoretical performance index.
[0021] In the embodiments of the present application, the neural network model construction device obtains the first theoretical performance index through the performance evaluation tool, which improves the feasibility of obtaining the first theoretical performance index.
[0022] Based on the neural network model construction method of the first aspect of the embodiments of the present application, in a possible implementation, the neural network model construction device determines the first construction unit of the first neural network through a performance evaluation tool. The first construction unit includes at least one of the following: the convolutional layer of the first neural network model, the pooling layer of the first neural network model, the activation function of the first neural network model, and the normalization layer of the first neural network model. The convolutional layer, pooling layer, activation function, and normalization layer included in the first construction unit can be one or more. The neural network model construction device performs calculations based on the first construction unit to obtain the first theoretical performance index.
[0023] In the embodiments of the present application, the neural network model construction device calculates one or more first construction units of the first neural network model through a performance evaluation tool to obtain the first theoretical performance index. The first construction unit further includes at least one convolutional layer, pooling layer, activation function, and normalization layer. Therefore, the theoretical performance index of the first neural network model can be adjusted according to the layers that make up the first neural network model, improving flexibility.
[0024] Based on the neural network model construction method of the first aspect of the embodiments of the present application, in a possible implementation, after the neural network model construction device constructs the first neural network model, the neural network model construction device further trains the first neural network model to obtain the fifth neural network model. After the neural network model construction device obtains the fifth neural network model, it acquires the model performance of the fifth neural network model and adjusts the first model generator according to the model performance of the fifth neural network and the first theoretical performance index to obtain the second model generator.
[0025] In the embodiments of the present application, adjusting the first model generator according to the model performance of the fifth neural network model after training the first neural network model and the first theoretical performance index can better improve the theoretical performance index of the neural network model running on the target chip while ensuring the same model performance of the neural network model generated by the adjusted second model generator.
[0026] Based on the neural network model construction method of the first aspect of the embodiments of the present application, in a possible implementation, the first theoretical performance index and the second theoretical performance index respectively include at least one of the following: theoretical vector module bound, theoretical memory bound, theoretical cube utilization rate, theoretical high-speed parallel multiplier-accumulator MAC utilization rate, theoretical cube module (vector cycle) operation count, and theoretical vector module (cube) operation count.
[0027] In the embodiments of the present application, the specific references of the first theoretical performance index and the second theoretical performance index are exemplarily described, improving the feasibility of the solution.
[0028] Based on the neural network model construction method of the first aspect of the embodiments of the present application, in a possible implementation, the first measured performance index and the second measured performance index respectively include at least one of the following: measured vector module limit, measured memory limit, measured cube module utilization rate, measured high-speed parallel multiplier-accumulator (MAC) utilization rate, measured cube module operation count, and measured vector module operation count.
[0029] In the embodiments of the present application, the specific references of the first measured performance index and the second measured performance index are exemplarily described, improving the feasibility of the solution.
[0030] The second aspect of the embodiments of the present application provides a neural network model construction device.
[0031] A neural network model construction device includes:
[0032] A construction unit, configured to construct a first neural network model through a first model generator;
[0033] An acquisition unit, configured to obtain a first performance index of the first neural network model when running on a target chip according to the first neural network model;
[0034] A processing unit, configured to adjust the first model generator according to the first performance index to obtain a second model generator;
[0035] The construction unit is further configured to construct a second neural network model through the second model generator, and the second performance index of the second neural network model is better than the first performance index.
[0036] Optionally, the first performance index is a first theoretical performance index, the first theoretical performance index represents the theoretical value of the performance index of the first neural network model when running on a target chip, the second performance index is a second theoretical performance index, and the second theoretical performance index is better than the first theoretical performance index.
[0037] Optionally, the acquisition unit is specifically configured to obtain a first measured performance index, and the first measured performance index represents the measured value of the performance index of the second neural network model when running on a target chip;
[0038] The processing unit is further configured to adjust the second model generator according to the first measured performance index to obtain a third model generator;
[0039] The construction unit is further configured to construct a third neural network model through the third model generator, and the second measured performance index of the third neural network model is better than the first measured performance index.
[0040] Optionally, the neural network model construction device further includes:
[0041] A training unit, configured to train the second neural network model to obtain a fourth neural network model;
[0042] The processing unit is further configured to adjust the second model generator according to the first measured performance index to obtain a third model generator, including:
[0043] The processing unit is further configured to adjust the second model generator according to the first measured performance index and the model performance of the fourth neural network model to obtain a third model generator.
[0044] Optionally, the first performance index is a first theoretical performance index, and the obtaining unit is specifically configured to obtain the first theoretical performance index through a performance evaluation tool, where the performance evaluation tool includes a calculation function, and the calculation function is used to calculate the first neural network model to obtain the first theoretical performance index.
[0045] Optionally, the neural network model construction device further includes:
[0046] A determination unit, configured to determine the first construction unit of the first neural network model through a performance evaluation tool, where the first construction unit includes at least one of the following: the convolutional layer of the first neural network model, the pooling layer of the first neural network model, the activation function of the first neural network model, the normalization layer of the first neural network model;
[0047] The processing unit is further configured to perform calculations according to the first construction unit to obtain the first theoretical performance index.
[0048] Optionally, the training unit is further configured to train the first neural network model to obtain a fifth neural network model;
[0049] The processing unit is further configured to adjust the first model generator according to the first theoretical performance index and the model performance of the fifth neural network model to obtain a second model generator.
[0050] Optionally, the first theoretical performance index and the second theoretical performance index respectively include at least one of the following: theoretical vector module limit, theoretical memory limit, theoretical cube module utilization rate, theoretical high-speed parallel multiplier-accumulator MAC utilization rate, theoretical cube module operation times, theoretical vector module operation times.
[0051] Optionally, the first measured performance index and the second measured performance index respectively include at least one of the following: measured vector module limit, measured memory limit, measured cube module utilization rate, measured high-speed parallel multiplier-accumulator MAC utilization rate, measured cube module operation times, measured vector module operation times.
[0052] In the third aspect of the embodiments of the present application, a neural network model construction device is provided, including:
[0053] A processor, a memory, and an input / output interface, where the processor, the memory, and the input / output interface are connected; the memory is used to store program codes; when the processor calls the program codes in the memory, it executes the method provided in the implementation manner of the first aspect of the present application.
[0054] In the fourth aspect of the embodiments of the present application, a storage medium is provided. It should be noted that the technical solution of the present invention can be embodied in the form of a software product, either essentially or in part contributing to the prior art, or the whole or part of the technical solution. The computer software product is stored in a storage medium and is used to store computer software instructions for the above-mentioned device, which includes a program designed for the metadata storage method in the first aspect above.
[0055] The storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (abbreviation: ROM, full name: Read-Only Memory), a random access memory (abbreviation: RAM, full name: Random Access Memory), a magnetic disk, or an optical disc.
[0056] In the fifth aspect of the embodiments of the present application, a computer program product containing instructions is provided. When it runs on a computer, it causes the computer to execute the method according to the implementation manner of the first aspect of the present application.
[0057] Among them, the processor mentioned anywhere above can be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the method for port detection in the first aspect above.
[0058] In the technical solution provided by the embodiments of the present application, by obtaining the first theoretical performance index of the first neural network running on the target chip and adjusting the first model generator according to the first theoretical performance index, the first model generator can be adjusted according to the performance index of the target chip, improving the compatibility of the second neural network model generated by the adjusted second model. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 It is a framework schematic diagram of an embodiment of the neural network model construction method in the embodiments of the present application;
[0060] Figure 2Another framework schematic diagram of the neural network model construction method embodiment in this application
[0061] Figure 3 Another framework schematic diagram of the neural network model construction method embodiment in this application
[0062] Figure 4 Another framework schematic diagram of the neural network model construction method embodiment in this application
[0063] Figure 5 Another framework schematic diagram of the neural network model construction method embodiment in this application
[0064] Figure 6 A process schematic diagram of the neural network model construction method embodiment in this application
[0065] Figure 7 Another process schematic diagram of the neural network model construction method embodiment in this application
[0066] Figure 8 Another process schematic diagram of the neural network model construction method embodiment in this application
[0067] Figure 9 Another process schematic diagram of the neural network model construction method embodiment in this application
[0068] Figure 10 A structure schematic diagram of the neural network model construction device embodiment in this application
[0069] Figure 11 Another structure schematic diagram of the neural network model construction device embodiment in this application
[0070] Figure 12 Another structure schematic diagram of the neural network model construction device embodiment in this application Detailed implementation manners
[0071] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0072] Figure 1 Show an artificial intelligence main framework schematic diagram, which describes the overall working process of the artificial intelligence system and is applicable to the general artificial intelligence field requirements.
[0073] The above artificial intelligence theme framework will be elaborated from two dimensions: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis).
[0074] The "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it can be the general processes of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution and output. In this process, data undergoes the refinement process of "data - information - knowledge - wisdom".
[0075] The "IT value chain" reflects the value brought by artificial intelligence to the information technology industry from the underlying infrastructure of artificial intelligence, information (provision and processing technology implementation) to the industrial ecological process of the system.
[0076] (1) Infrastructure:
[0077] The infrastructure provides computing power support for the artificial intelligence system, enables communication with the external world, and is supported through the basic platform. It communicates with the external world through sensors; the computing power is provided by intelligent chips (such as hardware acceleration chips like CPU, NPU, GPU, ASIC, FPGA, etc.); the basic platform includes relevant platform guarantees and supports such as distributed computing frameworks and networks, and can include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the external world to obtain data, and these data are provided to the intelligent chips in the distributed computing system provided by the basic platform for calculation.
[0078] (2) Data
[0079] The data at the upper layer of the infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voices, videos, texts, and also involves Internet of Things data of traditional devices, including business data of existing systems and perception data such as force, displacement, liquid level, temperature, humidity, etc.
[0080] (3) Data processing
[0081] Data processing usually includes data training, machine learning, deep learning, search, reasoning, decision-making, etc.
[0082] Among them, machine learning and deep learning can perform symbolic and formal intelligent information modeling, extraction, preprocessing, training, etc. on data.
[0083] Reasoning refers to the process of simulating the intelligent reasoning method of humans in a computer or intelligent system, and using formal information for machine thinking and problem-solving according to the reasoning control strategy. The typical function is search and matching.
[0084] Decision-making refers to the process of making decisions after intelligent information is inferred, and usually provides functions such as classification, ranking, prediction, etc.
[0085] (4) General capabilities
[0086] After the data is processed as mentioned above, some general capabilities can be further formed based on the results of the data processing. For example, it can be an algorithm or a general system. For example, translation, text analysis, computer vision processing (such as image recognition, object detection, etc.), speech recognition, and so on.
[0087] (5) Intelligent products and industry applications
[0088] Intelligent products and industry applications refer to the products and applications of artificial intelligence systems in various fields, which are the encapsulation of the overall artificial intelligence solution, productize intelligent information decision-making, and achieve practical applications. Its application fields mainly include: intelligent manufacturing, intelligent transportation, smart home, intelligent healthcare, intelligent security, autonomous driving, safe city, intelligent terminals, etc.
[0089] See the appendix Figure 2 , the embodiment of the present application provides a system architecture 200. The system architecture includes a database 230 and a client device 240. The data acquisition device 260 is used to collect data and store it in the database 230, and the training module 220 generates a target model / rule 201 based on the data maintained in the database 230.
[0090] The work of each layer in a deep neural network can be described by the mathematical expression y = a(W * x + b): From a physical level, the work of each layer in a deep neural network can be understood as completing the transformation from the input space (the set of input vectors) to the output space (i.e., from the row space to the column space of the matrix) through five operations on the input space. These five operations include: 1. Dimensionality increase / dimensionality reduction; 2. Enlargement / shrinkage; 3. Rotation; 4. Translation; 5. "Bending". Among them, operations 1, 2, and 3 are completed by W * x, operation 4 is completed by +b, and operation 5 is implemented by a(). The reason for using the word "space" here is that the objects to be classified are not individual things, but a class of things, and space refers to the set of all individuals of this class of things. Among them, W is a weight vector, and each value in the vector represents the weight value of a neuron in this layer of the neural network. This vector determines the space transformation from the above-mentioned input space to the output space, that is, the weight of each layer controls how to transform the space. The purpose of training a deep neural network, that is, to finally obtain the weight matrices of all layers of the trained neural network. Therefore, the training process of the neural network is essentially the process of learning the way to control space transformation, and more specifically, learning the weight matrix. In the following embodiments of the present application, this weight matrix can be refined into a set of structural parameters and a set of network parameters. For specific reference, see the following Figure 2The relevant introduction in
[0091] Since it is expected that the output of the deep neural network is as close as possible to the target value, the weight vector of each layer of the neural network can be updated by comparing the predicted value of the current network with the target value and then according to the difference between the two. (Of course, there is usually an initialization process before the first update, that is, parameters are pre-configured for each layer in the deep neural network). For example, if the predicted value of the network is too high, the weights in the weight matrix are adjusted to reduce the predicted value. After continuous adjustment, until the value output by the neural network is close to or equal to the target value. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", that is, the loss function or objective function. The loss function is an important equation for measuring the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference. The training of the neural network can be understood as a process of minimizing the loss as much as possible.
[0092] The computing module may include a training module 220, and the target model / rule obtained by the training module 220 can be applied to different systems or devices. In the appendix Figure 2 In the figure, the execution device 210 is configured with a transceiver 212. The transceiver 212 can be a wireless transceiver, an optical transceiver or a wired interface (such as an I / O interface), etc., to interact with external devices. The "user" can input data to the transceiver 212 through the client device 240. For example, in the following embodiments of the present application, the client device 240 can send a target task to the execution device 210, request the execution device to build a neural network, and send a database for training to the execution device 210.
[0093] The execution device 210 can call data, code, etc. in the data storage system 250, and can also store data, instructions, etc. in the data storage system 250.
[0094] The computing module 211 processes the input data using the target model / rule 201. Specifically, the computing module 211 is used to: build a first neural network model through a first model generator, obtain a first performance index of the first neural network model when running on a target chip according to the first neural network model, adjust the first model generator according to the first performance index to obtain a second model generator, and build a second neural network model through the second model generator. The second performance index of the second neural network model is better than the first performance index.
[0095] The association function module 21 can specifically be a module for training the model generator.
[0096] The associated function module 214 can be used to perform search construction based on the basic operations included in the search space to obtain the first model generator.
[0097] Finally, the transceiver 212 returns the constructed neural network model to the client device 240 for deploying the neural network model in the client device 240 or other devices.
[0098] More deeply, the training module 220 can obtain the corresponding target model / rule 201 based on different data for different target tasks to provide better results for users.
[0099] In the situation shown in the appendix Figure 2 The user can manually specify the data in the input execution device 210. For example, the user can operate in the interface provided by the transceiver 212. In another situation, the client device 240 can automatically input data to the transceiver 212 and obtain results. If the client device 240 needs user authorization to automatically input data, the user can set corresponding permissions in the client device 240. The user can view the results output by the execution device 210 in the client device 240, and the specific presentation forms can be display, sound, action and other specific ways. The client device 240 can also be used as a data collection end to store the collected data associated with the target task in the database 230.
[0100] It should be noted that the appendix Figure 2 is only a schematic diagram of a system architecture provided by an embodiment of the present application. The positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, in the appendix Figure 2 the data storage system 250 is an external memory relative to the execution device 210. In other scenarios, the data storage system 250 can also be placed in the execution device 210.
[0101] Referring to the appendix Figure 3 , an embodiment of the present application provides a system architecture 300. The execution device 210 is implemented by one or more servers. Optionally, it cooperates with other computing devices, such as devices for data storage, routers, load balancers, etc.; the execution device 210 can be arranged on one physical site or distributed on multiple physical sites. The execution device 210 can use the data in the data storage system 250 or call the program code in the data storage system 250 to implement the steps of the following Figures 6 - 8 corresponding neural network model construction method of the present application.
[0102] Users can operate their respective user devices (such as local device 301 and local device 302) to interact with the execution device 210. Each local device can represent any computing device, such as a personal computer, a computer workstation, a smartphone, a tablet computer, a smart camera, a smart car, or other types of cellular phones, media consumption devices, wearable devices, set-top boxes, game consoles, etc.
[0103] Each user's local device can interact with the execution device 210 through a communication network of any communication mechanism / communication standard. The communication network can be a wide area network, a local area network, a point-to-point connection, etc., or any combination thereof. Specifically, the communication network can include a wireless network, a wired network, or a combination of a wireless network and a wired network, etc. The wireless network includes, but is not limited to: a fifth-generation mobile communication technology (5th-Generation, 5G) system, a long term evolution (LTE) system, a global system for mobile communication (GSM), or a code division multiple access (CDMA) network, a wideband code division multiple access (WCDMA) network, a wireless fidelity (WiFi), a bluetooth, a Zigbee protocol, a radio frequency identification (RFID), a Long Range (Lora) wireless communication, a near field communication (NFC), or any combination of one or more of them. The wired network can include a fiber optic communication network or a network composed of coaxial cables, etc.
[0104] In another implementation, one or more aspects of the execution device 210 can be implemented by each local device. For example, the local device 301 can provide local data or feedback calculation results for the execution device 210.
[0105] It should be noted that all functions of the execution device 210 can also be implemented by the local device. For example, the local device 301 implements the functions of the execution device 210 and provides services for its own user, or provides services for the user of the local device 302.
[0106] Please refer to Figure 4 , which is a schematic diagram of a neural network model construction framework provided by an embodiment of this application.
[0107] The neural network model construction framework includes at least one controller model and a neural network model generated via the controller model. Among them, the controller model obtains the architecture of a neural network model through search, trains the architecture of the neural network model with a training set, and then evaluates the architecture of the neural network model with a validation set to obtain the accuracy. After that, the feedback result (such as accuracy) is returned to the control model, and the controller model is updated using reinforcement learning so that the controller can generate a better network structure in the next cycle. After this process is repeated multiple times, a new architecture is generated, then tested, and the feedback result is sent to the controller model for reinforcement learning again. Finally, the above-mentioned controller model will tend to design architectures that can obtain higher accuracy in the validation set.
[0108] Further, based on Figure 4 the schematic diagram of the neural network model construction framework, please refer to Figure 5 , which is another schematic diagram of the neural network model construction framework provided by the embodiments of this application.
[0109] The schematic diagram of the neural network model construction framework at least includes a network structure population and a performance evaluation tool. Among them, the network structure population contains a variety of neural network model construction units. The model generator searches for suitable neural network construction units in the network structure population and constructs a neural network model from the searched neural network construction units. It can be understood that the neural network model can include one or more neural network model construction units, and specific details are not limited here.
[0110] After constructing the neural network model, input the neural network model into the performance evaluation tool to obtain the chip performance metrics when the neural network model runs on the target chip. Optionally, the neural network model can also be trained on a chip in the cloud to obtain the network structure performance of the neural network structure. It can be understood that the neural network model can also be trained in other chip environments, and specific details are not limited here.
[0111] In the continuous iteration process, adjust the network structure population through the chip performance metrics and the network structure performance to obtain a model generator that can construct a model with higher chip performance and better network structure performance. In a possible implementation, update the network structure population through the Pareto optimization method, that is, without affecting the network structure performance, make the constructed neural network model have higher chip performance.
[0112] It should be noted that the neural network model construction device in the embodiments of this application can be a computer device with a chip such as a server, a desktop computer, a laptop computer, a computer cluster, etc., and specific details are not limited here.
[0113] Based on the foregoing application scenarios, the neural network model construction method provided by this application will be described below.
[0114] Please refer to Figure 6 , which is a schematic flowchart of a neural network model construction method provided by an embodiment of this application.
[0115] In step 601, the neural network model construction device constructs a first neural network model through a first model generator.
[0116] The neural network model construction device constructs a first neural network model through a first model generator, and the first model generator is pre-set in the neural network model construction device.
[0117] Specifically, in a possible implementation manner, the neural network model construction device constructs a first neural network through a first model generator and a target task. Before the first model generator constructs the first neural network, the neural network model construction device will obtain the target task. The target task is determined according to its own needs or can also be determined according to the user's operation. For example, the target task may include: the type of neural network, the accuracy of the neural network, etc. Among them, the type of neural network includes the output type of the neural network to be constructed. For example, the target task may be to construct a face recognition neural network for recognizing faces and outputting corresponding person information. Another example is that the target task may be initiated by a terminal to construct a neural network for vehicle recognition to identify the information of the vehicle included in the picture obtained by the tampering device.
[0118] It should be noted that the neural network in this application may be a convolutional neural network, a recurrent neural network, a perceptron neural network, etc., and can be specifically adjusted according to the actual application scenario, and this application does not make any limitations in this regard.
[0119] Simultaneously with or after obtaining the target task, corresponding training data can also be obtained. The training data is data associated with the target task. The training data may include the input data of the target task and the actual measurement data. For example, if the target task is to construct a face recognition neural network, the training data includes a large number of face pictures and the task information corresponding to each picture. Among them, in a possible implementation manner, the training data can be divided into a training set and a validation set. The training set represents the pictures for training the neural network model, and the validation set represents the pictures for verifying the accuracy of the network model.
[0120] In a possible implementation, before the first model generator constructs the first neural network, hardware constraint conditions are also input to the first model generator, and the hardware constraint conditions include various parameters of the target chip. Specifically, the hardware constraint conditions may include at least one of the following: the frequency of the chip, the size of the chip memory, the size of the chip operation module, the bandwidth between memories in the chip, and so on. It can be understood that in the actual application process, more parameters may also be included, and specific details are not limited here. After inputting the hardware constraint conditions into the first model generator, the first model generator constructs the first neural network according to the hardware constraint conditions when constructing the first neural network model.
[0121] In step 602, the neural network model construction device obtains the first theoretical performance metric.
[0122] The neural network model construction device obtains the first theoretical performance metric, and the first theoretical performance metric represents the theoretical value of the performance metric when the first neural network model runs on the target chip, that is, the theoretical performance metric is a theoretical performance metric.
[0123] The target chip may be a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or a general-purpose processor, etc. For example, the first theoretical performance metric is the theoretical performance metric when the CPU runs the first neural network.
[0124] In a possible implementation, the first theoretical performance metric includes at least one of the following: theoretical vector module bound, theoretical memory bound, theoretical cube utilization, theoretical multiplier-accumulator (MAC) utilization, number of operations of the theoretical cube module (cube cycle), number of operations of the theoretical vector module (vector cycle), L1 and L2 memory fusion effect, compute batch effect, Tiling strategy and its performance, performance effect under mixed precision, performance effect under different data flow modes, number of cycles or latency of each operator or network layer in the neural network model, total number of cycles or latency of the entire neural network model.
[0125] In a possible implementation, the neural network model construction device obtains the first theoretical performance metric through a performance evaluation tool, and the performance evaluation tool includes a calculation function for calculating the first neural network model to obtain the first theoretical performance metric. Specifically, the performance evaluation tool can be a software-based tool for obtaining the performance metric corresponding to the target chip when running the first neural network. For example, in a preferred manner, the performance evaluation tool is PTM, and the neural network model construction device obtains the theoretical performance metric through PTM. In actual application, the performance evaluation tool can also exist in other forms, such as a hardware module, and specific details are not limited here.
[0126] Specifically, in a possible implementation, the neural network model construction device divides the neural network into one or more first construction units with at least one constituent unit (which may include at least one operator or one layer as a unit) that makes up the first neural network as a first construction unit. The multiple first construction units are input into a performance evaluation tool, and the performance evaluation tool calculates based on the multiple first construction units and the parameters of the target chip to obtain the theoretical performance indicators of each first construction unit. Furthermore, the theoretical performance indicators of each first construction unit are added up, which is the first theoretical performance indicator for the target chip to run the first neural network. For example, the theoretical performance indicator is the time taken for the target chip to run the first neural network. If a construction unit is an operator in the first neural network, the performance evaluation tool calculates the time taken for each operator respectively, and then adds them up to finally obtain the time taken for all operators in the entire neural network, that is, the time taken for the target chip to run the first neural network. Among them, the time taken for the entire network can be calculated by the following formula:
[0127] Total network time = ∑ n f (Time taken for each layer operator estimated by the performance evaluation tool)
[0128] It should be noted that other performance indicators can also be implemented through similar formulas. For example, when calculating the number of cube cycles, it can be calculated by the following formula:
[0129] Total network cube cycles = ∑ n f (Cube cycles of each layer operator estimated by the performance evaluation tool)
[0130] Specifically, in a possible implementation, the neural network model construction device inputs the entire first neural network model into a performance evaluation tool, and the performance evaluation tool calculates based on the first neural network model and the parameters of the target chip to obtain the theoretical performance indicator of the first neural network model, which is the first theoretical performance indicator for the target chip to run the first neural network. For example, if the first theoretical performance indicator is the time taken for the target chip to run the first neural network, the performance evaluation tool calculates the time taken for the first neural network to obtain the time taken for the target chip to run the first neural network.
[0131] In a possible implementation, the neural network model construction device obtains a first theoretical performance metric by inputting a data stream into a performance evaluation tool. Specifically, the neural network model construction device determines one or more first construction units of the first neural network model, and the first construction unit includes at least one of the following: the convolutional layer of the first neural network model, the pooling layer of the first neural network model, the activation function of the first neural network model, and the normalization layer of the first neural network model. Then, according to the one or more first construction units, different layers in the neural network are divided into dimensions of a task. For example, one or more convolutional layers, pooling layers, activation functions, normalization layers, etc. are divided into one dimension. In a preferred manner, one convolutional layer, pooling layer, activation function, and normalization layer are divided into one dimension. And analyze the data stream according to each task. For example, when the neural network runs on a chip, different types of chip memory will process different data. For example, during the process from L2 to L1, L2 needs to transfer data to L1, then analyze the efficiency in this process under the dimension of a neural network. This process may further include: L1->L0A / L0B, UB->L1, etc. Then, according to the data stream calculation and analysis, when each dimension in the neural network performs data transmission, calculate the cycle number, cube utilization rate, mac utilization rate, vector bound, memory bound (DDR, L2), etc. of the theoretical performance metrics through which pipelines (pipes) are used for calculation and data transmission.
[0132] Memory bound: The situation of the reuse of the smallest memory unit in the chip. When the output feature map is larger than the smallest memory unit of the chip, the smallest memory unit of the chip needs to be reused multiple times. In the embodiments of the present application, the output feature map can be adjusted to be smaller by corresponding weights when constructing the neural network to adapt to the size of the smallest memory unit of the chip, thereby improving the usage efficiency of the smallest memory unit of the chip.
[0133] Vector bound: The situation of the reuse of the vector module in the chip.
[0134] Cube utilization rate: The cube module is mainly used to calculate matrices. The cube utilization rate refers to the number of times the cube module is used per unit time. When training a neural network, if you want to improve the subsequent cube utilization rate, you can set the cube utilization rate as the target task. When training the model, the closer the model is to convergence, the higher the cube utilization rate.
[0135] MAC (multiplier-accumulator / multiply-accumulate operation) utilization rate: A module in the chip that counts the number of multiplication and addition operations in a neural network. The MAC utilization rate refers to the number of times the MAC module is used per unit time. The method for improving the MAC utilization rate in the embodiments of this application is similar to the method for improving the cube utilization rate, and will not be elaborated here specifically.
[0136] Cube cycle count: The number of operations using the cube module in the chip.
[0137] Vector cycle count: The number of operations using the vector module.
[0138] L1 and L2 memory fusion: L1 and L2 are two different types of memory in the chip. By setting the corresponding weights for constructing the neural network, the compatibility of L1 and L2 in processing the corresponding data of the neural network can be improved.
[0139] Compute batch: The number of feature maps processed in batches within the same time period. By setting the corresponding weights for constructing the neural network, the compute batch can be increased as much as possible while ensuring the running performance of the chip.
[0140] In addition to the performance metrics of the chip described above, the performance metrics in the embodiments of this application may also include more parameters, as long as they can affect the running efficiency of the neural network on the chip, and no specific limitations are provided here.
[0141] In step 603, the neural network model construction device adjusts the first model generator according to the first theoretical performance metric to obtain a second model generator.
[0142] After obtaining the first theoretical performance metric, the neural network model construction device adjusts the first model generator according to the first theoretical performance metric to obtain a second model generator. The second model generator can construct a neural network with better theoretical performance metrics when running on the target chip compared to the first model generator.
[0143] In a possible implementation, the weight factor in the corresponding first model generator can be adjusted according to the first theoretical performance metric, so that the first model generator can construct a second neural network model with better performance metrics when running on the target chip, and the second theoretical performance metric of the second neural network model is better than the first theoretical performance metric. For example, when the theoretical performance metric is the theoretical elapsed time, then when adjusting the model generator, the weight factor corresponding to the elapsed time is adjusted so that the model generator constructs a neural network with a shorter elapsed time when running on the target chip. It can be understood that when there are multiple theoretical performance metrics, the corresponding multiple weight factors are adjusted so that the first model generator constructs a neural network with better performance metrics in all aspects when running on the target chip.
[0144] In a possible implementation, the first neural network is trained. When the first neural network tends to converge, a fifth neural network model is obtained, and the performance parameters of the fifth neural network are obtained. The performance parameters of the fifth neural network may include accuracy, peak signal-to-noise ratio for describing pictures, etc., and are not specifically limited here. Then, the weight factor in the first model generator is adjusted according to the performance parameters of the fifth neural network and the first theoretical performance metric. In this way, it can be ensured that the neural network generated by the model generator has excellent performance parameters and better performance metrics when running on the target chip.
[0145] It should be noted that as Figure 7 shown, steps 601 to 603 are a process of the first model generator iterating once according to the first theoretical performance metric. In actual application, it can be iterated once or multiple times, and finally a model generator that can construct a neural network with the optimal theoretical performance metric is generated, that is, the second model generator.
[0146] In step 604, the neural network model construction device obtains the measured performance metric.
[0147] After the neural network model construction device obtains the second model generator, the neural network model construction device obtains the first measured performance metric, and the first measured performance metric represents the actual performance metric when the target chip runs the second neural network.
[0148] After the neural network model construction device obtains the second model generator, it can directly use the second neural network generated by the second model generator, or can perform a second round of improvement on the second model generator. The second round of improvement is an adjustment based on the performance metrics obtained when the second neural network runs on the actual target chip.
[0149] In a possible implementation, a second neural network model is constructed from a second model generator, and the second neural network model is placed in a target chip for operation, and actual performance metrics corresponding to the operation of the second network model on the target chip are obtained.
[0150] The target chip can be a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a general-purpose processor, or the like.
[0151] In a possible implementation, the first measured performance metrics may include at least one of the following: measured vector bound, measured memory bound, measured cube utilization, measured multiplier-accumulator (MAC) utilization, measured number of cube cycles, measured L1 and L2 memory fusion effect, measured compute batch effect, measured tiling strategy and its performance, measured performance under mixed precision, measured performance under different data flow modes, measured latency of each operator or network layer cycle in the neural network model, measured total cycle count or measured latency of the entire neural network model, etc. It can be understood that in actual application processes, there may also be more chip performance metrics, such as measured number of vector cycles, as long as they can reflect the performance metrics reflected by the operation of the second network on the target chip. Specific details are not limited here.
[0152] In a possible implementation, the system where the target chip is located can obtain the performance metrics corresponding to the operation of the second neural network on the target chip in the form of software or hardware. The specific method for obtaining the performance metrics is not limited here.
[0153] In step 605, the neural network model construction device adjusts the second model generator according to the first measured performance metrics to obtain a third model generator.
[0154] After the neural network model construction device obtains the first measured performance metric, the neural network model construction device adjusts the second model generator according to the first measured performance metric to obtain a third model generator. The third model generator can construct a third neural network model, and the second measured performance metric of the third neural network model is better than the first measured performance metric.
[0155] In a possible implementation, the weight factor in the corresponding second model generator can be adjusted according to the measured performance metric, so that the second model generator can construct a neural network with better performance when running on the target chip. For example, when the first measured performance metric is the measured time taken for the second neural network to run on the target chip, then when adjusting the second model generator, the weight factor corresponding to the time taken is adjusted, so that the model generator constructs a neural network with a shorter running time on the target chip. It can be understood that when the first measured theoretical performance metric is multiple metrics, then the corresponding multiple weight factors are adjusted, so that the second model generator constructs a neural network with better performance metrics when running on the target chip.
[0156] In a possible implementation, the second neural network is trained. When the second neural network converges, a fourth neural network model is obtained, and the performance parameters of the fourth neural network are obtained. The performance parameters of the fourth neural network can include accuracy, peak signal-to-noise ratio for describing pictures, etc., and are not specifically limited here. Then, the weight factor in the first model generator is adjusted according to the performance parameters of the fourth neural network and the first measured performance metric. This can ensure that the neural network generated by the model generator has excellent performance parameters and better performance metrics when running on the target chip.
[0157] It should be noted that as Figure 8 shown, steps 604 to 605 are the process of the second model generator iterating once according to the measured performance metric. In actual application, it can be iterated once or multiple times, and finally a model generator that can construct a neural network with the optimal measured performance metric is generated, that is, the third model generator.
[0158] In the embodiments of this application, steps 604 to 605 are optional steps. When steps 604 to 605 are not executed, the second neural network model constructed by the second model generator is used as the neural network model for use on the target chip.
[0159] It should be noted that in the embodiments of this application, when obtaining the theoretical performance metric and the measured performance metric, it can be obtained in the neural network model construction device, or it can be obtained by other computer devices and then sent to the neural network model construction device, which is not specifically limited here.
[0160] In the embodiments of the present application, by obtaining the theoretical performance metrics of a target chip running a first neural network and adjusting the corresponding weights in the first model generator according to the theoretical performance metrics, a neural network model with better hardware performance metrics when running on the target chip is constructed.
[0161] The above Figure 6 The embodiment shown is an application scenario of the embodiments of the present application. The following describes another application scenario of the neural network model construction method of the embodiments of the present application.
[0162] Please refer to Figure 9 , which is another flowchart of the neural network model construction method provided by the embodiments of the present application.
[0163] In this embodiment, the model generator is described by taking the structure search space as an example.
[0164] According to the requirements of the application scenario task, before constructing the neural network model, it is also necessary to construct a search space according to the metrics. Each constituent unit in this search space is an encoding indicating the construction of the neural network structure. During the structure search process, each sampling utilizes the distribution relationship of the unevaluated constituent units in the search space and the network model constructed based on the evaluated constituent units. Specifically, as Figure 9 shown, the part in the dashed box is the step of constructing the initial search space. This search space consists of multiple network structures, and each network structure consists of multiple constituent units. After constructing the initial search space of multiple network structures, based on preset rules or application scenario tasks, unreasonable network structures are filtered to obtain a second search space. Then, through network structure clustering, a new third search space is formed. Then, based on the third search space, the cluster center structure is trained, the unevaluated structure is modeled according to the training loss value of the evaluated network structure, and several network structures are selected for training based on Bayesian optimization. Then, the unevaluated structure is modeled according to the training loss value again, and this cycle continues until the search space construction is completed.
[0165] After constructing the initial search space, then Figure 6 the neural network model construction method in the embodiment shown is used to further adjust the search space, and the details are not described here.
[0166] In the embodiments of the present application, by initially constructing and training the search space, the performance of the search space is improved.
[0167] The neural network model construction method in the embodiments of the present application has been described above. Next, the neural network model construction device in the embodiments of the present application will be described.
[0168] Please refer to Figure 10, which is a schematic structural diagram of the neural network model construction device provided by the embodiment of the present application.
[0169] A neural network model construction device includes:
[0170] A construction unit 1001, configured to construct a first neural network model through a first model generator;
[0171] An acquisition unit 1002, configured to obtain a first performance index of the first neural network model when it runs on a target chip according to the first neural network model;
[0172] A processing unit 1003, configured to adjust the first model generator according to the first performance index to obtain a second model generator;
[0173] The construction unit 1001 is further configured to construct a second neural network model through the second model generator, and the second performance index of the second neural network model is better than the first performance index.
[0174] In this embodiment, the operations performed by each unit of the neural network model construction device are similar to those described in the foregoing Figure 6 and Figure 7 illustrated embodiments, and will not be elaborated here specifically.
[0175] Please refer to Figure 11 , which is another schematic structural diagram of the neural network model construction device provided by the embodiment of the present application.
[0176] A neural network model construction device includes:
[0177] A construction unit 1101, configured to construct a first neural network model through a first model generator;
[0178] An acquisition unit 1102, configured to obtain a first performance index of the first neural network model when it runs on a target chip according to the first neural network model;
[0179] A processing unit 1103, configured to adjust the first model generator according to the first performance index to obtain a second model generator;
[0180] The construction unit 1101 is further configured to construct a second neural network model through the second model generator, and the second performance index of the second neural network model is better than the first performance index.
[0181] Optionally, the first performance index is a first theoretical performance index, and the first theoretical performance index represents the theoretical value of the performance index of the first neural network model when it runs on the target chip. The second performance index is a second theoretical performance index, and the second theoretical performance index is better than the first theoretical performance index.
[0182] Optionally, the obtaining unit 1102 is specifically configured to obtain a first measured performance metric, where the first measured performance metric represents a measured value of the performance metric when the second neural network model runs on the target chip;
[0183] The processing unit 1103 is further configured to adjust the second model generator according to the first measured performance metric to obtain a third model generator;
[0184] The constructing unit 1101 is further configured to construct a third neural network model through the third model generator, and the second measured performance metric of the third neural network model is better than the first measured performance metric.
[0185] Optionally, the neural network model constructing apparatus further includes:
[0186] A training unit 1104, configured to train the second neural network model to obtain a fourth neural network model;
[0187] The processing unit 1103 is further configured to adjust the second model generator according to the first measured performance metric to obtain a third model generator, including:
[0188] The processing unit 1103 is further configured to adjust the second model generator according to the first measured performance metric and the model performance of the fourth neural network model to obtain a third model generator.
[0189] Optionally, the first performance metric is a first theoretical performance metric, and the obtaining unit is specifically configured to obtain the first theoretical performance metric through a performance evaluation tool, where the performance evaluation tool includes a calculation function, and the calculation function is used to calculate the first neural network model to obtain the first theoretical performance metric.
[0190] Optionally, the neural network model constructing apparatus further includes:
[0191] A determining unit 1105, configured to determine a first constructing unit of the first neural network model through the performance evaluation tool, where the first constructing unit includes at least one of the following: a convolutional layer of the first neural network model, a pooling layer of the first neural network model, an activation function of the first neural network model, a normalization layer of the first neural network model;
[0192] The processing unit 1103 is further configured to perform calculations according to the first constructing unit to obtain the first theoretical performance metric.
[0193] Optionally, the training unit 1104 is further configured to train the first neural network model to obtain a fifth neural network model;
[0194] The processing unit 1103 is further configured to adjust the first model generator according to the first theoretical performance metric and the model performance of the fifth neural network model to obtain a second model generator.
[0195] Optionally, the first theoretical performance metric and the second theoretical performance metric each include at least one of the following: theoretical vector module limit, theoretical memory limit, theoretical cube module utilization rate, theoretical high-speed parallel multiplier-accumulator (MAC) utilization rate, theoretical cube module operation count, and theoretical vector module operation count.
[0196] Optionally, the first measured performance metric and the second measured performance metric each include at least one of the following: measured vector module limit, measured memory limit, measured cube module utilization rate, measured high-speed parallel multiplier-accumulator (MAC) utilization rate, measured cube module operation count, and measured vector module operation count.
[0197] In this embodiment, the operations performed by each unit of the neural network model construction device are similar to those described in the foregoing Figure 6 and Figure 7 illustrated embodiments, and will not be elaborated herein for the sake of brevity.
[0198] Please refer to Figure 12 , which is another schematic structural diagram of the neural network model construction device provided by the embodiments of the present application.
[0199] A processor 1201, a memory 1202, a bus 1205, and an interface 1204. The processor 1201 is connected to the memory 1202 and the interface 1204. The bus 1205 is respectively connected to the processor 1201, the memory 1202, and the interface 1204. The interface 1204 is used to receive or send data. The processor 1201 is a single-core or multi-core central processing unit, or a specific integrated circuit, or one or more integrated circuits configured to implement the embodiments of the present invention. The memory 1202 can be a random access memory (RAM), or a non-volatile memory, such as at least one hard disk memory. The memory 1202 is used to store computer-executable instructions. Specifically, the computer-executable instructions may include a program 1203.
[0200] In this embodiment, when the processor 1201 calls the program 1203, it can cause the Figure 12 neural network model construction device to perform the operations performed by the neural network model construction device in the foregoing Figure 6 or Figure 9 illustrated embodiments, and will not be elaborated herein for the sake of brevity.
[0201] It should be understood that the processor mentioned in the neural network model construction device in the above embodiments of the present application, or the processor provided in the above embodiments of the present application, may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0202] It should also be understood that the number of processors in the neural network model construction device in the above embodiments of the present application may be one or multiple, which can be adjusted according to the actual application scenario. This is only an exemplary illustration here and is not limited. The number of memories in the embodiments of the present application may be one or multiple, which can be adjusted according to the actual application scenario. This is only an exemplary illustration here and is not limited.
[0203] It should also be noted that when the neural network model construction device includes a processor (or processing unit) and a memory, the processor in the present application may be integrated with the memory, or the processor and the memory may be connected through an interface, which can be adjusted according to the actual application scenario and is not limited.
[0204] The embodiments of the present application also provide a computer program or a computer program product including the computer program. When the computer program is executed on a certain computer, the computer will implement the method flow related to the neural network model construction device in any of the above method embodiments.
[0205] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the computer, it implements the method flow related to the neural network model construction device in any of the above method embodiments.
[0206] In the above Figures 6 - 9 In each of the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0207] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, they generate, in whole or in part, the processes or functions described in the embodiments of the present application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0208] In the description of the present application, the terms "first", "second", etc. in the specification, claims, and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such terms may be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing the embodiments of the present application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product, or device that includes a series of units does not have to be limited to those units, but may include other units that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0209] The names of messages / frames / information, modules, or units provided in the embodiments of the present application are only examples, and other names may be used as long as the functions of the messages / frames / information, modules, or units are the same.
[0210] The terms used in the embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "the", and "said" used in the embodiments of the present application are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that in the description of the present application, unless otherwise specified, " / " indicates that the related objects before and after are in an "or" relationship. For example, A / B may represent A or B; the "and / or" in the present application is only a description of the relationship between related objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B may be singular or plural.
[0211] Depending on the context, the words "if" or "when" as used herein may be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" may be interpreted as "when determined" or "in response to determining" or "when detected (stated condition or event)" or "in response to detecting (stated condition or event)".
[0212] As described above, the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or equivalently replace some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for constructing a neural network model, characterized in that, Including: Construct a first neural network model through a first model generator; Obtain a first theoretical performance metric of the first neural network model when running on a target chip, where the first theoretical performance metric represents the theoretical value of the performance metric of the first neural network model when running on the target chip; Adjust the first model generator according to the first theoretical performance metric to obtain a second model generator; Construct a second neural network model through the second model generator, where the second theoretical performance metric of the second neural network model is better than the first theoretical performance metric; Obtain a first measured performance metric, where the first measured performance metric represents the measured value of the performance metric of the second neural network model when running on the target chip; Adjust the second model generator according to the first measured performance metric to obtain a third model generator; Construct a third neural network model through the third model generator, where the second measured performance metric of the third neural network model is better than the first measured performance metric.
2. The method according to claim 1, wherein Before obtaining the first measured performance metric, the method further includes: Train the second neural network model to obtain a fourth neural network model; Adjusting the second model generator according to the first measured performance metric to obtain a third model generator includes: Adjust the second model generator according to the first measured performance metric and the model performance of the fourth neural network model to obtain the third model generator.
3. The method according to claim 1, characterized in that Obtaining the first theoretical performance metric of the first neural network model when running on a target chip according to the first neural network model includes: Obtain the first theoretical performance metric through a performance evaluation tool, where the performance evaluation tool includes a calculation function for calculating the first neural network model to obtain the first theoretical performance metric.
4. The method according to claim 3, wherein Obtaining the first theoretical performance metric through a performance evaluation tool includes: Determine a first building unit of the first neural network model through the performance evaluation tool, where the first building unit includes at least one of the following: the convolutional layer of the first neural network model, the pooling layer of the first neural network model, the activation function of the first neural network model, the normalization layer of the first neural network model; Perform calculations according to the first building unit to obtain the first theoretical performance metric.
5. The method according to any one of claims 1 to 4, characterized in that, After constructing the first neural network model through the first model generator, the method further includes: Train the first neural network model to obtain a fifth neural network model; Adjusting the first model generator according to the first theoretical performance metric to obtain a second model generator includes: Adjust the first model generator according to the first theoretical performance metric and the model performance of the fifth neural network model to obtain the second model generator.
6. The method according to any one of claims 1 to 4, characterized in that, The first theoretical performance metric and the second theoretical performance metric respectively include at least one of the following: theoretical vector module bound, theoretical memory bound, theoretical cube module utilization rate, theoretical high-speed parallel multiply-accumulate (MAC) utilization rate, theoretical cube module operation count, theoretical vector module operation count.
7. The method according to any one of claims 1 to 4, characterized in that, The first measured performance index and the second measured performance index respectively include at least one of the following: measured vector module limit, measured memory limit, measured cube module utilization rate, measured high-speed parallel multiplier-accumulator (MAC) utilization rate, measured number of operations of the cube module, and measured number of operations of the vector module.
8. A neural network model construction device, characterized in that including: a construction unit configured to construct a first neural network model through a first model generator; an acquisition unit configured to obtain a first theoretical performance index of the first neural network model when running on a target chip according to the first neural network model, where the first theoretical performance index represents a theoretical value of the performance index of the first neural network model when running on the target chip; a processing unit configured to adjust the first model generator according to the first theoretical performance index to obtain a second model generator; the construction unit is further configured to construct a second neural network model through the second model generator, and a second theoretical performance index of the second neural network model is better than the first theoretical performance index; the acquisition unit is further configured to obtain a first measured performance index, where the first measured performance index represents a measured value of the performance index of the second neural network model when running on the target chip; the processing unit is further configured to adjust the second model generator according to the first measured performance index to obtain a third model generator; the construction unit is further configured to construct a third neural network model through the third model generator, and a second measured performance index of the third neural network model is better than the first measured performance index.
9. The neural network model construction device according to claim 8, characterized in that, The neural network model construction device further includes: a training unit configured to train the second neural network model to obtain a fourth neural network model; the processing unit is further configured to adjust the second model generator according to the first measured performance index to obtain a third model generator, including: the processing unit is further configured to adjust the second model generator according to the first measured performance index and the model performance of the fourth neural network model to obtain the third model generator.
10. The neural network model construction device according to claim 8, characterized in that, The acquisition unit is specifically configured to obtain the first theoretical performance index through a performance evaluation tool, and the performance evaluation tool includes a calculation function for calculating the first neural network model to obtain the first theoretical performance index.
11. The neural network model construction device according to claim 10, characterized in that, The neural network model construction device further includes: a determination unit configured to determine a first construction unit of the first neural network model through the performance evaluation tool, where the first construction unit includes at least one of the following: a convolutional layer of the first neural network model, a pooling layer of the first neural network model, an activation function of the first neural network model, and a normalization layer of the first neural network model; the processing unit is further configured to perform calculations according to the first construction unit to obtain the first theoretical performance index.
12. The neural network model construction device according to claim 9, characterized in that, the training unit is further configured to train the first neural network model to obtain a fifth neural network model; the processing unit is further configured to adjust the first model generator according to the first theoretical performance index and the model performance of the fifth neural network model to obtain the second model generator.
13. The neural network model construction device according to any one of claims 8 to 11, characterized in that The first theoretical performance index and the second theoretical performance index respectively include at least one of the following: theoretical vector module limit, theoretical memory limit, theoretical cube module utilization rate, theoretical high-speed parallel multiplier-accumulator (MAC) utilization rate, theoretical cube module operation times, theoretical vector module operation times.
14. The neural network model construction device according to any one of claims 8 to 11, characterized in that, The first measured performance index and the second measured performance index respectively include at least one of the following: measured vector module limit, measured memory limit, measured cube module utilization rate, measured high-speed parallel multiplier-accumulator (MAC) utilization rate, measured cube module operation times, measured vector module operation times.
15. A computer storage medium, characterized in that, Instructions are stored in the computer storage medium, and when the instructions are executed on the computer, the computer is caused to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Ai model development method and device
CN111357014A