Efficient processing of neural network models
By identifying and compiling the sequence of frequently called neural network models and storing the compiled model in SRAM, the delay and power consumption problems caused by frequent switching of neural network models on hardware accelerators are solved, achieving more efficient processing speed and lower latency.
Patent Information
- Application Number
- CN202080041079.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-03-09
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2040-03-09
AI Technical Summary
When processing neural network models on hardware accelerators, frequent switching of different neural network models results in frequent loading and cleaning of parameters in memory, resulting in increased latency and power consumption.
The compiler identifies the sequence of frequently called neural network models, compiles the model according to the sequence, and only loads when the compiled model does not exist in SRAM, avoiding unnecessary loading and cleaning.
Significantly reduces latency and power consumption, improves processing speed and efficiency, and prevents redundant cleaning and reloading of SRAM.
Smart Images

Figure CN113939830B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to efficient processing of neural network models on hardware accelerators. Background Art
[0002] Various mobile device applications (e.g., camera applications, social media applications, etc.) may all require the use of corresponding neural network models on a hardware accelerator. Typically, the architecture of a hardware accelerator allows the parameters of a single neural network model to be stored in its memory, and therefore the compiler of the hardware accelerator can only compile one neural network model at a time. When another neural network model (e.g., a model of another application) needs to be executed, the parameters of the neural network model replace the parameters of the previous neural network model. Therefore, when the previous neural network model needs to be processed by the hardware accelerator, the parameters of this model are loaded again into one or more memories of the hardware accelerator. Reloading the parameters into the memory in this way consumes a lot of memory and power, resulting in delays. Summary of the invention
[0003] A compiler for a computing device is described that identifies a sequence of neural network models that are frequently called by an application of the computing device, compiles the model according to the sequence, and loads the compiled model into a static random access memory (SRAM) of a hardware accelerator only if the same compiled model from another identical sequence of previous calls does not already exist in the SRAM.
[0004] In one aspect, a method performed by a compiler is described. The compiler is capable of identifying a first neural network model set that has been executed on a hardware accelerator of a computing device more than a threshold number of times in the past preset amount of time. The compiler is capable of identifying a sequence in which the first model set is executed on the hardware accelerator. Each neural network model in the first neural network model set is capable of being compiled for execution by the hardware accelerator. For each neural network model in the first neural network model set, the compiler is capable of outputting the compiled model to the hardware accelerator for storage in one or more memories of the hardware accelerator according to the sequence. The storage of the compiled model for each neural network model in the first neural network model set according to the sequence prevents the need to recompile and reload the compiled results of the sequence of the first neural network model set into one or more memories when the first neural network model set is to be executed again on the hardware accelerator.
[0005] In some embodiments, the method can also include one or more of the following aspects, which can be implemented separately or in any feasible combination. For each neural network model in the first neural network model set, a data structure including neural network model parameters can be received. The compiling can further include compiling the data structure of each neural network model in the first neural network model set to generate a compiled data structure for each neural network model in the first neural network model set, wherein the compiled data structure is a compiled model.
[0006] The same first hash can be assigned to each compiled model in the sequence. The first hash can be output to the hardware accelerator along with each compiled model in the sequence for storage in one or more memories of the hardware accelerator. The same second hash can be assigned to each compiled model in a second sequence of models. The second sequence can be after the first sequence. When the second sequence is the same as the first sequence, the second hash can be the same as the first hash. When the second sequence is different from the first sequence, the second hash can be different from the first hash. If the second hash is different from the first hash, the hardware accelerator is configured to replace each compiled model in the first sequence in the one or more memories with each compiled model in the second sequence in the one or more memories. If the second hash is the same as the first hash, the hardware accelerator is configured to prevent erasure of each compiled model in the first sequence from the one or more memories.
[0007] Each neural network model in the first neural network model set has been processed on the hardware accelerator more than a preset number of times (e.g., 5 times) in the past. The compiler compiles the first neural network model set, while the hardware accelerator simultaneously performs neural network calculations of one or more other neural network models. The identifier of the first model set and the identifier of the sequence can be updated after a preset time interval. The compiler can abandon the update for a preset time in response to a compilation failure of the first neural network model set. The abandonment can include abandoning the update for 7500 milliseconds in response to a compilation failure of the first neural network model set.
[0008] The compilation of the first set of neural network models can include determining that each neural network model in the first set of neural network models has a particular size compatible with the compilation. The compilation of the first set of neural network models can include compiling only a single neural network model at any time. The sequence can include a facial recognition neural network model and one or more subordinate neural network models to be processed after processing the facial recognition neural network model.
[0009] In another aspect, a system including a compiler and a hardware accelerator is described. The compiler is capable of identifying a first neural network model set that has been executed on a hardware accelerator of a computing device more than a threshold number of times in the past preset amount of time. The compiler is capable of identifying a sequence in which the first model set is executed on the hardware accelerator. The compiler is capable of compiling each neural network model in the first neural network model set for execution by the hardware accelerator. For each neural network model in the first neural network model set, the compiler is capable of outputting the compiled model to the hardware accelerator for storage in one or more memories of the hardware accelerator according to the sequence. The hardware accelerator is capable of including one or more memories for storing the compiled model of each neural network model in the first neural network model set according to the sequence. The storage of each neural network model in the first neural network model set according to the sequence of compiled models in one or more memories can prevent the need to recompile and reload the compiled results of the sequence of the first neural network model set into the one or more memories when the first neural network model set is to be executed again on the hardware accelerator.
[0010] In certain embodiments, one or more of the following can be additionally implemented alone or in any feasible combination. One or more memories can be static random access memories (SRAMs). The hardware accelerator can also include multiple computing units configured to process a first set of neural network models. Each of the multiple computing units can include at least one processor and a memory. The multiple computing units can be serially coupled via at least one bus. The first set of neural network models can include at least one facial recognition neural network model. In response to the controller receiving an instruction to execute at least one facial recognition neural network model from the computing device, the at least one facial recognition neural network model can be activated.
[0011] The computing device can include an application and an application programming interface (API). The application can generate instructions sent via the API. The application can be a camera application. The API can be a neural network API (NNAPI). The computing device can be an Android device.
[0012] The subject matter described herein provides many advantages. For example, storing multiple compiled models in a sequence simultaneously in the SRAM of the hardware accelerator prevents redundant deletion and reloading of the same, previously loaded compiled models of the same sequence that were also called previously from the SRAM. This avoidance of unnecessary cleaning of the SRAM and reloading of compiled data into the SRAM can significantly reduce latency and increase processing speed. In addition, storing parameters in the SRAM and retrieving these parameters from the SRAM is significantly faster and more energy-efficient than storing in the main memory, i.e., dynamic random access memory (DRAM) and retrieving them from the main memory. In addition, in the event of a compilation process failure that is not necessarily but may occur, the compiler can prevent repeated compilation failures by pausing the recognition of frequently called neural network models or their sequences for a period of time. In addition, the compiler will only attempt to compile them if each of the frequently occurring neural network models has a size less than a preset amount (e.g., 5 megabytes), which can increase (e.g., maximize) the number of models for which the compiled models are stored in the SRAM at the same time.
[0013] The details of one or more variations of the subject matter described herein are set forth in the drawings and the specification. Other features and advantages of the subject matter described herein will be apparent from the specification, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 A device with a hardware accelerator that processes a neural network model with reduced latency is illustrated.
[0015] Figure 2 A method performed by a compiler to prevent redundant cleanup and reloading of compiled data in a static random access memory (SRAM) of a hardware accelerator is illustrated.
[0016] Figure 3 The diagram shows two different sequences of the model for which data structures along with corresponding hashes are compiled and compared to determine if the SRAM needs to be cleaned and reloaded with the compiled data.
[0017] Figure 4 The steps performed by the hardware accelerator are illustrated to store compiled data in the SRAM and to perform actions in response to determining whether existing compiled data in the SRAM needs to be cleared and whether new compiled data needs to be reloaded in the SRAM.
[0018] Figure 5Illustrated are aspects of a hardware accelerator that includes an SRAM for storing compiled data that includes parameters for processing a neural network model.
[0019] Like reference numbers in the various drawings represent like elements. DETAILED DESCRIPTION
[0020] Figure 1 The diagram shows a computing device 102 having a hardware accelerator 104 that processes a neural network model with reduced latency (e.g., low latency). The computing device 102 can be a mobile device, such as a phone, a tablet, a tablet, a laptop, and / or any other mobile device. Although the computing device 102 is described as a mobile device, in some embodiments, the computing device 102 can be a desktop computer or a computer cluster or a computer network. The hardware accelerator 104 refers to computer hardware that is specifically manufactured to perform certain functions more efficiently than software running on a general-purpose processor. For example, the hardware accelerator 104 can use dedicated hardware, such as matrix multiplication, to perform specified operations, and the dedicated hardware allows the hardware accelerator to execute deep feedforward neural networks, such as convolutional neural networks (CNNs), more efficiently than a general-purpose processor. In order for the hardware accelerator 104 to execute the neural network model, the neural network model is compiled specifically for the accelerator 104.
[0021] The hardware accelerator 104 can be a tensor processing unit (TPU). Although a TPU is described, in other embodiments, the hardware accelerator 104 can be any other hardware accelerator 104, such as a graphics processing unit (GPU), a digital signal processor (DSP), a field programmable analog array (FPAA), a sound card, a network processor, a cryptographic accelerator, an artificial intelligence accelerator, a physical processing unit (PPU), a data compression accelerator, a network on chip, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a complex programmable logic device, and / or a system on chip.
[0022] The computing device 102 further includes "N" software applications 106, such as a camera application, a social networking application, and any other software applications. N can be any integer, such as 5, 10, 20, 100, or any other integer. In order for these applications 106 to communicate with the hardware accelerator 106 (e.g., provide input to the hardware accelerator 106 and receive output from the hardware accelerator 106), the mobile computing device 102 further employs an application programming interface (API) 108, a processor 110, and a processor 112, the API 108 outputting a specific data structure to be processed in response to the execution of the application 106, the processor 110 for performing quantization on the specific data structure (also referred to as a quantizer 110), and the processor 112 implementing a compiler configured to use the quantized data structure.
[0023] The API 108 enables communication between the application 106 and the processor 110. The API 108 can be a communication protocol that facilitates such communication (e.g., syntax, semantics, and synchronization of the communication and possible error recovery methods). The API 108 can be, for example, a neural network API (NN-API). The API 108 can allow the application 106 to generate a data structure that includes mathematical operations that constitute a neural network model to be processed in response to execution of the application 106. For example, in response to acquiring an image through the camera application 106, the API 108 can allow the camera application 106 to generate a data structure (e.g., a TensorFlow data structure) 109 that indicates mathematical operations that constitute a facial recognition model to be implemented on the accelerator. The data structure 109 can have parameter data (e.g., weights and input data of the neural network) represented as floating point numbers, and the floating point numbers have a preset number of bits (e.g., 32-bit floating point numbers).
[0024] The processor / quantizer 110 can receive the data structure 109 from the API 108 (or, in some embodiments, from the application 106) and convert it into a smaller data structure (e.g., a TensorFlowLite data structure) 111 having the same parameter data (e.g., weights and input data of the neural network model, as in the data structure 109), which is represented as a fixed-point number with a lower preset number of bits (e.g., an 8-bit fixed-point number). Converting all 32-bit floating-point numbers in the data structure 109 to the closest 8-bit fixed-point number in the data structure 111 is called quantization. Quantization advantageously makes the data structure smaller, and therefore makes operations on the data by the hardware accelerator 104 faster and less computationally intensive. In addition, although these low-bit (e.g., 8-bit) representations may not be as accurate as the corresponding high-bit (e.g., 32-bit) representations in the data structure 109, they do not significantly (i.e., not noticeably) affect the inference accuracy of the neural network. Although quantization is described herein as occurring during an API call, in some embodiments, quantization can be performed at any time before compiling the quantized data, as described below. Furthermore, while quantization is described as an automatic process in which quantizer 110 automatically receives data and performs quantization based on the data, in some implementations, at least a portion of the values in data structure 109 may already be quantized, for example, in an external system.
[0025] The processor 112 implements a compiler 114. The compiler 114 compiles the quantized data structure 111 into a compiled data structure 116 that is compatible with the hardware accelerator 104. In addition to the quantized data structure 111, the compiled data structure 116 can include machine-level code that includes low-level instructions to be executed by the hardware accelerator 104. In general, the compiler 114 can be run on any suitable operating system, such as an Android system or a Debian-based Linux system.
[0026] Furthermore, while quantization and compilation of quantized data are described, in some embodiments, quantization is not performed because quantization may not necessarily be necessary (e.g., if a hardware accelerator 104 such as a GPU or TPU is capable of floating point operations, then the compiler 114 can work directly on a floating point model without quantization).
[0027] The hardware accelerator 104 is capable of performing various neural network calculations to process a neural network model (e.g., a facial recognition model) based on the compiled data structure 116. Each time the hardware accelerator 104 processes a neural network model (e.g., a facial recognition model), the hardware accelerator 104 needs to access parameters of the neural network model. For such access, the hardware accelerator 104 includes one or more memories—specifically, static random access memories (SRAMs)—which store data structures 116, including parameters of the neural network model (e.g., a facial detection neural network model). The parameters are stored in the SRAM rather than in the main memory of the computing device 102 because the SRAM allows faster access to data stored therein by the hardware accelerator 104, thereby improving processing speed and energy efficiency and reducing latency.
[0028] The SRAM on the hardware accelerator 104 has limited memory space (e.g., up to 8 megabytes) that can store the compiled data structure 116 (including parameters) of the model. When compiling the data structure with parameters separately, the compiler 114 provides a unique hash (e.g., a 64-bit number) for each compiled data structure for unique identification. When executing the neural network model at runtime, the hardware accelerator compares the hash with the hash of the previously compiled data structure stored in the SRAM. If the token matches, the hardware accelerator uses the stored previously compiled data structure, thereby avoiding the need to reload the compiled data structure 116 into the SRAM. If the token does not match, the hardware accelerator 104 eliminates / erases the stored data structure (i.e., the previously compiled data structure) and instead writes the compiled data structure 116 to the SRAM (i.e., cleans and reloads the SRAM), thereby improving (e.g., in some embodiments, maximizing) the efficiency of using the limited memory space (e.g., up to 8 megabytes) in the SRAM. However, the hardware accelerator 104 is configured to store data (including parameters) corresponding to a single hash. Therefore, when all compiled data structures 116 have different hashes, after each separate compilation by compiler 114, the SRAM is reloaded with the new compiled data 116, which will cause delays.
[0029] The delay and power requirements due to this reloading of the SRAM after each compilation are significantly reduced by the compiler 114 by:
[0030] (1) identifying a sequence of frequently occurring neural network models to be executed—for example, each time a user clicks multiple images using camera application 106 on computing device 102, computing device 102 may invoke and process the following neural network models in a sequence: (a) first a face detection model for detecting faces in each image, (b) then an orientation detection model for detecting orientation in each image, (c) then a blur detection model for detecting blur in each image, (d) then another neural network model for suggesting the best image based on the detections of (a), (b), and (c);
[0031] (2) compiling all data structures 111 of the neural network model in the sequence together to generate a plurality of compiled data structures 116, wherein each compiled data structure 116 corresponds to a compilation result of a corresponding neural network model in the sequence; and
[0032] (3) Outputting the compiled data structure 116 for the model in the sequence to the hardware accelerator 104 for storage according to the sequence in one or more memories (more specifically—the SRAM of the hardware accelerator 104).
[0033] Whenever the same frequently occurring sequence of the model set is called up again (e.g., in response to a user clicking multiple images again using the camera application 106 on the computing device 102), the hardware accelerator 104 can quickly access the compiled data structure 116 directly from the SRAM, thereby avoiding the need to clean (i.e., eliminate / erase) the SRAM and reload the SRAM with the same compiled data structure. Avoiding cleaning and reloading the SRAM can advantageously convert processing resources and power, and increase processing speed, thereby significantly reducing latency.
[0034] The simultaneous storage of the compiled data structures 116 for all models in a sequence in the SRAM of the hardware accelerator 104 is performed as follows. Each time the compiler 114 outputs a compiled data structure to the hardware accelerator 104, the compiler 114 calculates and sends a separate unique hash (e.g., a 64-bit number) for unique identification of the compiled data structure 116. However, when multiple compiled data structures 116 are determined for the neural network models in the sequence, the compiler 114 assigns a single hash (e.g., a 64-bit number) to identify all of these compiled data structures. For example, the compiled data structures 116 for all models in the sequence are assigned the same hash, and the compiled data structure for any model not in the sequence will have a different hash. The compiler 114 calculates the same hash for the same model, thereby calculating the same hash for the same sequence of models. Therefore, if the hash is the same as the hash of the model for which the compiled data structure was previously stored in the SRAM, this indicates that the current model sequence is the same as the previous sequence for which the compiled data structure was previously stored, thereby avoiding the need to clean the SRAM and then reload the SRAM with the same compiled data structure. This prevention of clearing and then reloading the SRAM unnecessarily reduces latency significantly.
[0035] The amount of SRAM allocated to each model is fixed at compile time and is prioritized based on the order of the data structures compiled by the compiler. For example, when two models A and B in a sequence (where A is called before B) are compiled and the corresponding data structures 116 are assigned the same hash, as much SRAM space as required is first allocated to the data structure 116 of model A, and if there is still SRAM space remaining thereafter, the SRAM space is given to the data structure 116 of model B. If a portion of the model data structure 116 cannot fit in SRAM, it is instead stored in an external memory (e.g., the main memory of the computing device 102) and extracted from it at run time. If the entire model does not fit in SRAM, the compiler 114 generates appropriate instructions for the accelerator 104 to extract data from a dynamic random access memory (DRAM). Maximizing the use of SRAM in this way can advantageously improve (e.g., maximize) processing speed.
[0036] In some embodiments, if several models are compiled, some models may not be allocated space in SRAM, so these models must load all data from external memory (e.g., the main memory of the computing device 102). Loading from external memory is slower than loading from SRAM, but when the models are run in a sequence of frequently called models, this is still faster than swapping (i.e., clearing and reloading) the SRAM every time any model is run. As described above, if the entire model does not fit in SRAM, the compiler 114 generates appropriate instructions executed by the accelerator 104 to extract data from dynamic random access memory (DRAM). Note that this interaction between the compiler 114 and the accelerator 104 is different from a typical central processing unit (CPU), which typically has a hardware cache that automatically stores the most frequently used data in SRAM.
[0037] The compiler 114 can continue to compile the data structures of the neural network model while the hardware accelerator 104 simultaneously performs neural network calculations for one or more other neural network models. This synchronization function prevents the processing of the neural network model from being paused by the compiler during compilation activities, thereby advantageously increasing speed and reducing latency.
[0038] The compiler 114 can update the identification of the first model set and the identification of the sequence (i.e., re-identify the first model set and re-identify the sequence) after a preset time interval (e.g., 1 second, 30 seconds, 1 minute, 5 minutes, 10 minutes, 20 minutes, 30 minutes, 1 hour, 24 hours, 5 days, or any other suitable time). This updating ensures that the SRAM is used in an optimal manner to simultaneously store parameters of the most relevant models (i.e., the models currently or most recently determined to be most frequently called, rather than such determinations that were made a long time ago (e.g., more than a threshold time ago)). The compiler can abandon updates for a preset time (e.g., 7500 milliseconds) in response to a compilation failure of the first neural network model set. Abandoning for a preset time can provide protection against transient faults that might otherwise trigger continuous compilation cycles, resulting in a significant increase in power consumption. Pausing joint compilation for a preset time after a failure increases the likelihood that the active model set will change, thereby avoiding the recurrence of the compilation failure. In some embodiments, the preset time can have other values, such as 1 second, 2 seconds, 5 seconds, 10 seconds, 30 seconds, 1 minute, or any other value that can avoid the compilation failure from happening again. This abandonment can save compilation resources, because using these resources to compile immediately may cause another compilation failure.
[0039] Compilation of the first set of neural network models can include determining, prior to compiling, that each data structure to be compiled is compatible (e.g., in megabytes) with compiler 114. For example, compiler 114 may not compile data structures having a size greater than a preset amount (e.g., 5 megabytes), thus ensuring that parameters of the model can be stored in SRAM simultaneously with parameters of other neural network models.
[0040] Figure 2 The method performed by the compiler 114 is illustrated to prevent redundant cleaning and reloading of the SRAM. In step 202, the compiler 114 can identify a frequently occurring sequence of neural network models. The frequently occurring sequence of models can be a set of neural network models that are called together in a specific sequence by the application 106 and processed by the hardware accelerator more than a threshold number of times in the past preset amount of time. For example, a user may frequently click multiple images using the camera application 106 on a mobile phone (e.g., more than a threshold number of times), and in this case, the frequently occurring models and corresponding sequences can be: (a) first a face detection model for detecting a face in each image, (b) then an orientation detection model for detecting an orientation in each image, (c) then a blur detection model for detecting blur in each image, (d) then another neural network model for suggesting the best image based on the detection of (a), (b), and (c). In order for a sequence to qualify as a frequently occurring sequence, the threshold number of times the sequence needs to be repeated can be 5 times, and in other embodiments can be 4 times, 6 times, 10 times, 15 times, 20 times, or any other integer (greater than 1). The past preset amount of time that can be considered for this determination can be from the time when the hardware accelerator 104 was deployed to perform neural network calculations for the corresponding application 106. In another embodiment, this past preset amount of time can be 1 minute, 5 minutes, 10 minutes, 20 minutes, 30 minutes, 1 hour, 24 hours, 5 days, or any other suitable time. In some embodiments, the past preset amount of time to be used for determining frequently occurring sequences can be dynamically calculated based on the usage of users of one or more applications 106.
[0041] At step 204, for each neural network model in the frequently occurring sequence of neural network models, compiler 114 can receive a data structure including model parameters. In some examples, the data structure received by compiler 114 can have 8-bit fixed-point numbers in data structure 111, which can be obtained, for example, by quantizing data structure 109 having 32-bit floating-point numbers. The model parameters of each neural network model can include weights and data of the neural network model.
[0042] At step 206, compiler 114 can compile the data structure of each neural network model in the sequence to generate a compiled data structure 116 for each neural network model. This compilation can be performed in the order of the sequence. For example, first the first model in the sequence is compiled, then the second model in the sequence is compiled, then the third model in the sequence is compiled, and so on, until finally the last model in the sequence is compiled. The individual models can be compiled using any suitable technique (including any conventional technique) for performing compilation by generating machine-level code to be accessed and executed by hardware accelerator 104.
[0043] At step 208, for each neural network model in the set of neural network models, compiler 114 can assign a hash to compiled data structure 116. The hash can be a unique 64-bit number that uniquely identifies compiled data structure 116. Although the hash is described as a 64-bit number, in other embodiments it can have any other number of bits. Typically, a hash function receives a compiled data structure as input and outputs a hash (e.g., a 64-bit number) that can be used, for example, as an index in a hash table. In various embodiments, the hash can also be referred to as a hash value, a hash code, or a digest. The hash function can be MD5, SHA-2, CRC32, any other one or more hash functions, and / or any combination thereof.
[0044] At step 210, when the hash (e.g., the first 64-bit number) is different from another hash (e.g., the second 64-bit number) of another sequence previously identified in the SRAM, the compiler can output the compiled data structure and hash to the hardware accelerator 104 according to the sequence for storage in the SRAM (i.e., reloading the SRAM). When the two hashes are the same, such reloading of the SRAM is unnecessary, and the compiler 114 therefore prevents cleaning and reloading the SRAM in this case. In other words, if the hash of all models in the first sequence is different from another hash of all models in the previous sequence (which indicates that the two sequences are different), then cleaning and reloading of the SRAM is performed. If the two hashes are the same, this indicates that the two sequences are the same, and therefore there is no need to clean and reload the SRAM, thereby reducing delays.
[0045] Figure 3Two different sequences of illustrated models are compiled for them along with the corresponding hashes, and compared to determine whether the SRAM needs to be cleaned up and reloaded with the compiled data. Each model in a frequently occurring sequence is assigned the same hash (64-bit unique number). For example, each model in the first sequence is assigned a first hash (calculated by the compiler) and each model in the second sequence is assigned a second hash (calculated by the compiler). The compiler 114 calculates the same hash for the same model, and thereby calculates the same hash for the same sequence of models. Therefore, if the second hash is the same as the first hash, then this indicates that the second sequence is the same as the first sequence, and therefore there is no need to reload the compiled data structure 116 for the second sequence of models in the SRAM of the hardware accelerator 104, thereby reducing delays. In this case, the compiler 114 therefore prevents the cleaning and reloading of the SRAM.
[0046] Figure 4 The steps performed by the hardware accelerator 104 are illustrated to store compiled data in SRAM and to act in response to determining whether the SRAM needs to be cleaned and reloaded. In step 402, the hardware accelerator 104 can allocate SRAM space for storing Figure 3 The hardware accelerator 104 can compare the compiled data structure 116 and the first hash. If the second hash is different from the first hash, then at step 404, the hardware accelerator 104 can erase the SRAM space used for the compiled data structure 116 and replace it with the compiled result of the second sequence. If the second hash is the same as the first hash, then at step 406, the hardware accelerator 104 can avoid erasing the compiled data structure 116 from the SRAM, which advantageously reduces latency.
[0047] Figure 5 1. Various aspects of the hardware accelerator 104 are illustrated, and the hardware accelerator 104 includes an SRAM 501 for storing a compiled data structure 116 that includes parameters for processing a neural network model. The hardware accelerator 106 communicates with the application 106 via an application programming interface (API) 108. The API 108 sends data to the hardware accelerator 104 via a compiler, and can output data from the hardware accelerator 104 (e.g., via a decompiler not shown). For example, the API 108 can send a particular data structure to the compiler 114 to be processed in response to execution of the application 106. The API 108 may need to send data to the compiler via a quantizer, depending on the configuration of the hardware accelerator 104, such as Figure 1 described.
[0048] The hardware accelerator 104 is configured to perform neural network computations in response to instructions and input data received from an application running on the computing device 102. The accelerator 102 can have a controller 502 and a plurality of individual computing units 504. Although eight computing units 504 are shown, in alternative implementations, the hardware accelerator 104 can have any other number of computing units 504, such as any number between 2 and 16. Each computing unit 504 can have at least one programmable processor 506 and at least one memory 508. In some implementations, as shown in the compiled data structure 116, the parameters for processing the neural network model can be distributed across one or more (e.g., all) memories 508.
[0049] The computing units 504 are capable of accelerating the machine learning inference workload of the neural network layer. Each computing unit 504 is self-contained and can independently perform the calculations required for a given layer of the multi-layer neural network. The hardware accelerator 104 is capable of performing calculations for the neural network layer by distributing tensor calculations across multiple computing units 504. The computational processing performed within the neural network layer can include the product of an input tensor including input activations and a parameter tensor including weights. The calculation can include multiplying the input activations by the weights over one or more loops and performing the accumulation of the products over multiple loops. The term "tensor" as used herein refers to a multidimensional geometric object, which can be a matrix or a data array.
[0050] Each computing unit 504 can implement a software algorithm to perform tensor calculations by processing nested loops that traverse N-dimensional tensors (where N can be any integer). In an example computing process, each loop can be responsible for traversing a specific dimension of an N-dimensional tensor. For a given tensor structure, a computing unit 504 can require access to elements of a specific tensor to perform multiple dot product calculations associated with the tensor. The calculation occurs when the input activation is multiplied by a parameter or weight. When the multiplication result is written to the output bus, the tensor calculation ends, and the output bus connects the computing unit 504 in series, and the data is transferred between the computing units through the bus and stored in the memory.
[0051] The hardware accelerator 104 is capable of supporting certain types of data structures (e.g., structure 109 with 32-bit floating point numbers) that are quantized (e.g., to obtain structure 111 with 8-bit floating point numbers) and then compiled specifically for the hardware accelerator 104 (e.g., to obtain compiled structure 116).
[0052] The hardware accelerator 104 can perform various neural network calculations based on the compiled data structure 116 generated by the compiler 114 to process a neural network model (e.g., a facial recognition model). Each time the hardware accelerator 104 processes a neural network model (e.g., a facial recognition model), the hardware accelerator 104 needs to access the parameters in the compiled data structure 116 of the neural network model. In order to store data received from the compiler 114 and the API 108 (including the parameters in the compiled data structure 116), the hardware accelerator 104 further includes an instruction memory 510, an SRAM 501, and a data memory 512.
[0053] The SRAM 501 has a limited memory space (e.g., up to 8 megabytes) that can store the compiled data structure 116 of the model. In order to best use the SRAM 501, the compiler 114 can identify the sequence in which the frequently occurring neural network model is executed, compile all the data structures in the sequence together to generate the compiled data structure 116 assigned the same identification (e.g., hash), and output the compiled data structure 116 to the SRAM 501 in a selective manner (specifically, only when a different sequence of the model is called). This prevents redundant reloading of the SRAM 501.
[0054] The amount of SRAM 501 allocated to each model is fixed at compile time and is prioritized based on the order of the data structures compiled by the compiler. For example, for compiled models, if two models A and B are compiled using the same hash, then as much SRAM 501 space as required is first allocated to the data structures of model A, and if there is still SRAM 501 space remaining thereafter, the SRAM 501 space is given to the data structures of model B. If the data structure 116 of one of model A or model B cannot fit in SRAM 501, it is instead stored in external memory (e.g., the main memory of the computing device 102) and picked up from the external memory at runtime.
[0055] If several models are compiled, some models may not have space allocated in SRAM 501, so these models must load all data from external memory. Loading from external memory is slower than loading from SRAM 501, but when models are run in frequent sequences, this is still faster than swapping SRAM 501 every time any model is run.
[0056] The embodiments of the subject matter and functional operations described in this specification can be implemented in digital electronic circuits, in tangibly embodied computer software or firmware, computer hardware, including the structures disclosed in this specification and their structural equivalents, or a combination of one or more of them. The embodiments of the subject matter described in this specification can be implemented as one or more computer programs, that is, one or more modules of computer program instructions encoded on a tangible non-temporary program carrier for executing or controlling the operation of a data processing device by a data processing device. Alternatively or in addition, program instructions can be encoded on an artificially generated propagation signal (e.g., a machine-generated electrical, optical or electromagnetic signal), which is generated to encode information for transmission to an appropriate receiver device for execution by a data processing device. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.
[0057] The processes and logic flows described in this specification can be performed by one or more programmable computers that execute one or more computer programs to perform functions by operating on input data and generating (multiple) outputs. The processes and logic flows can also be performed by dedicated logic circuits, and the apparatus can also be implemented as dedicated logic circuits, such as FPGAs (field programmable gate arrays), ASICs (application-specific integrated circuits), or GPGPUs (general purpose graphics processing units).
[0058] As an example, a computer suitable for executing a computer program can be based on a general or special purpose microprocessor or both, or any other type of central processing unit. Typically, the central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a central processing unit for performing or executing instructions and one or more storage devices for storing instructions and data. Typically, a computer will also include receiving data from or transmitting data to one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or operatively coupled to one or more mass storage devices. However, a computer does not need to have such a device.
[0059] Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including, by way of example, semiconductor memory devices, such as EPROM, EEPROM and flash memory devices; magnetic disks, such as internal hard disks or removable disks. The processor and memory can be supplemented by or incorporated in special purpose logic circuits.
[0060] Although this specification contains many specific implementation details, it should not be interpreted as a limitation on any invention or the scope of what may be claimed, but rather as a description of features that may be specific to a particular implementation of a particular invention. In the context of separate implementations, certain features described in this specification can also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation can also be implemented in multiple implementations individually or by any appropriate sub-combination. In addition, although features are described as working in certain combinations, and even claimed as such at the beginning, in some cases, one or more features from the claimed combination can be deleted from the claimed combination, and the claimed combination may be directed to a sub-combination or a variant of the sub-combination.
[0061] Similarly, although operations are described in a particular order in the accompanying drawings, this should not be understood as requiring that the operations be performed in the particular order shown or in a sequential order, or that all illustrated operations be performed, to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated into a single software product, or packaged into multiple software products.
[0062] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve the desired results. As an example, the processes described in the accompanying drawings do not necessarily require the particular order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing may be advantageous.
Claims
1. A method performed by a compiler, the method comprising: identifying a first set of neural network models that have been executed on a hardware accelerator of a computing device more than a threshold number of times in a past preset amount of time; Identifying a first sequence in which the first neural network model set is executed on the hardware accelerator; compiling each neural network model in the first set of neural network models for execution by the hardware accelerator; as well as For each neural network model in the first neural network model set, the compiled model is output to the hardware accelerator for storage in one or more memories of the hardware accelerator according to the first sequence, and when the first neural network model set is to be executed again on the hardware accelerator, the storage of the compiled model according to the first sequence for each neural network model in the first neural network model set prevents the need to recompile and reload the compilation results of the first sequence of the first neural network model set into the one or more memories.
2. The method according to claim 1, further comprising: For each neural network model in the first set of neural network models, receiving a data structure including parameters of the neural network model, Wherein, the compilation further includes compiling the data structure of each neural network model in the first neural network model set to generate a compiled data structure of each neural network model in the first neural network model set, and the compiled data structure is the compiled model.
3. The method according to claim 1, further comprising: assigning the same first hash to each compiled model in the first sequence; outputting the first hash along with each compiled model in the first sequence to the hardware accelerator for storage in the one or more memories of the hardware accelerator; assigning the same second hash to each compiled model in a second sequence of models, the second sequence following the first sequence, the second hash being the same as the first hash when the second sequence is the same as the first sequence and being different from the first hash when the second sequence is different from the first sequence, in: If the second hash is different from the first hash, the hardware accelerator is configured to replace each compiled model in the first sequence in the one or more memories with each compiled model in the second sequence in the one or more memories; If the second hash is the same as the first hash, the hardware accelerator is configured to prevent erasure of each compiled model in the first sequence from the one or more memories.
4. The method according to claim 1, wherein: Each neural network model in the first set of neural network models has been processed more than five times on the hardware accelerator in the past.
5. The method according to claim 4, wherein: The compiler compiles the first neural network model set, while the hardware accelerator simultaneously performs neural network calculations of one or more other neural network models.
6. The method according to claim 1, further comprising: The identifier of the first neural network model set and the identifier of the first sequence are updated after a preset time interval.
7. The method according to claim 6, further comprising: In response to the compilation failure of the first neural network model set, abandoning updating for a preset time.
8. The method according to claim 7, wherein: The waiver includes: In response to the compilation failure of the first neural network model set, updating is abandoned for 7500 milliseconds.
9. The method according to claim 1, wherein: The compiling of the first neural network model set includes: Determining that each neural network model within the first set of neural network models has a specific size that is compatible with the compilation.
10. The method according to claim 1, wherein: The compiling of the first neural network model set includes: Only a single neural network model is compiled at any time.
11. The method according to claim 1, wherein: The first sequence includes a facial recognition neural network model and one or more subordinate neural network models to be processed after processing the facial recognition neural network model.
12. A system for processing a neural network model, comprising: A compiler, the compiler being configured to: identifying a first set of neural network models that have been executed on a hardware accelerator of a computing device more than a threshold number of times in a past preset amount of time; Identifying a sequence in which the first neural network model set is executed on the hardware accelerator; compiling each neural network model in the first set of neural network models for execution by the hardware accelerator; as well as For each neural network model in the first neural network model set, outputting the compiled model to the hardware accelerator for storage in one or more memories of the hardware accelerator according to the sequence; as well as The hardware accelerator includes one or more memories for storing the compiled model of each neural network model in the first neural network model set according to the sequence. When the first neural network model set is to be executed again on the hardware accelerator, the storage of the compiled model according to the sequence for each neural network model in the first neural network model set in the one or more memories prevents the need to recompile and reload the compilation results of the sequence of the first neural network model set into the one or more memories.
13. The system according to claim 12, wherein: The one or more memories are static random access memories SRAM.
14. The system according to claim 12, wherein: The hardware accelerator further includes a plurality of computing units configured to process the first set of neural network models.
15. The system of claim 14, wherein: Each of the plurality of computing units comprises at least one processor and a memory; and The plurality of computing units are coupled in series via at least one bus.
16. The system of claim 12, wherein: The first neural network model set includes at least one facial recognition neural network model.
17. The system of claim 16, wherein: The at least one facial recognition neural network model is activated in response to the controller receiving an instruction from the computing device to execute the at least one facial recognition neural network model.
18. The system of claim 17, wherein: The computing device includes an application and an application programming interface (API), the application generating instructions to be sent via the API.
19. The system of claim 18, wherein: The application is a camera application.
20. The system of claim 18, wherein: The API is a neural network API; and The computing device is an Android device.
Citation Information
Patent Citations
Debugging device and method of convolutional neural network accelerator, and storage medium
CN109858621A
Hardware accelerated neural network subgraphs
US20190286973A1