Compile code for machine learning models for execution on specialized processors
By allocating memory according to the resource constraints of the target device at compile time and avoiding dynamic memory allocation, the memory and computing resource limitations of running neural networks on low-computing dedicated processors is solved, and the computing function of the device is improved.
Patent Information
- Application Number
- CN202010465993.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-25
- Filing Date
- 2020-05-28
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-05-13
AI Technical Summary
The prior art is difficult to effectively run neural network models on dedicated processors with low computing capabilities, mainly due to the limitations of memory and computing resources.
By allocating memory according to the resource constraints of the target device at compile time and avoiding dynamic memory allocation and release technologies, the memory usage of neural networks is reduced and performance is improved.
It realizes running neural networks on dedicated processors with resource constraints, improving the computing functionality of electronic devices, especially in end-user devices with limited computing resources.
Smart Images

Figure CN112015424B_ABST
Abstract
Description
Technical Field
[0001] The present specification generally relates to compiling neural network model source code for execution on a target platform, including compiling neural network model source code for execution on a dedicated processor such as a resource-constrained processor. Background Art
[0002] Software engineers and scientists have been using computer hardware for machine learning to make improvements in different industry applications including image classification, video analysis, speech recognition, and natural language processing. Notably, neural networks are being used more frequently to create systems that can perform different computing tasks based on training on very large amounts of data. BRIEF DESCRIPTION OF THE DRAWINGS
[0003] Certain features of the subject technology are set forth in the appended claims. However, for purposes of explanation, several embodiments of the subject technology are set forth in the following drawings.
[0004] Figure 1 An exemplary network environment is shown according to one or more implementations.
[0005] Figure 2 An exemplary software architecture for generating code for a neural network for execution on a special-purpose processor according to one or more implementations is shown.
[0006] Figure 3 An exemplary model format for a neural network (NN) model documentation file and corresponding NN model code according to one or more specific implementations is shown.
[0007] Figure 4 Examples of convolutional neural networks according to one or more implementations are shown.
[0008] Figure 5 An exemplary table illustrating memory allocation according to one or more implementations is shown.
[0009] Figure 6 A flowchart is shown of an exemplary process for generating code for a neural network model according to one or more implementations.
[0010] Figure 7 An exemplary process for determining memory allocation to generate code for a convolutional neural network according to one or more implementations is shown.
[0011] Figure 8 An electronic system is shown that can be used to implement one or more implementations of the subject technology. DETAILED DESCRIPTION
[0012] CROSS-REFERENCE TO RELATED APPLICATIONS
[0013] This application claims the benefit of U.S. Provisional Patent Application Serial No. 62 / 855,840, filed on May 31, 2019, entitled “Compiling Code for a Machine Learning Model for Execution on a Specialized Processor,” which is hereby incorporated by reference in its entirety for all purposes.
[0014] The specific embodiments shown below are intended to be descriptions of various configurations of the subject technology and are not intended to represent the only configuration that the subject technology can be practiced. The accompanying drawings are incorporated herein and constitute a part of the specific embodiments. The specific embodiments include specific details intended to provide a thorough understanding of the subject technology. However, the subject technology is not limited to the specific details set forth herein, but can be practiced using one or more other specific implementations. In one or more specific implementations, structures and components are shown in block diagram form to avoid blurring the concept of the subject technology.
[0015] The popularity of machine learning has increased significantly in recent years due to the availability of large amounts of training data and the advancement of more powerful and efficient computing hardware. A common approach is to use a graphics processing unit (GPU) to train a deep neural network, and also to perform deep neural networks after training on new input data. In addition, as further described below, dedicated, customized and / or specialized hardware, such as a low-power dedicated processor that can always be powered on (e.g., to detect audio triggers, collect and process sensor data from integrated accelerometers, gyroscopes, and compasses, etc.) can be provided to perform certain operations in a more computationally efficient and / or power-efficient manner. However, when deploying a given deep neural network for execution on a target platform and / or on a target processor on a target platform, resource constraints (e.g., memory and / or computing resources) that may limit the execution of a given neural network may be encountered, depending on the available hardware. For example, in order to be able to deploy a neural network model on a dedicated processor with lower computing power than a main processor (e.g., a CPU), it may be necessary to modify the neural network model to make it compatible with the architecture of the dedicated processor. Without such modifications, the neural network model, when running on a dedicated processor, may need to use another processor (such as a CPU) in order to perform some operations of the neural network model, resulting in further consumption of power / memory / computing resources.
[0016] In addition, as further described herein, a given electronic device may include a dedicated processor that is always powered on and / or in an active mode, for example, even when the device's host processor / application processor is in a low power mode or where such electronic device does not include a host processor / application processor (e.g., a CPU and / or GPU). Such a dedicated processor may be a low-computing power processor that is designed to also utilize less energy than a CPU or GPU, and in one example is also designed to run continuously on the electronic device to collect audio and / or sensor data. In one example, such a dedicated processor may be an always-on processor (AOP), which is a small, low-power auxiliary processor implemented as an embedded motion coprocessor, such as in an electronic device such as or In existing solutions, it is not feasible to run machine learning models on such low-computing-power dedicated processors because such low-computing-power dedicated processors are incompatible with the structural and / or operational requirements for running machine learning models (e.g., they may require additional computing power and / or memory requirements of a CPU or GPU).
[0017] Specific implementations of the subject technology described herein reduce the memory footprint of a neural network by providing code that reuses portions of memory and allocates all memory at compile time (e.g., before running the neural network) based on the resource constraints of a given target device / dedicated processor. In addition, the performance of the neural network can be improved by avoiding the use of dynamic memory allocation and deallocation techniques that are typically performed during the execution of the neural network model. In addition, some processors (such as some dedicated processors provided on a given electronic device) may not allow for or be infeasible for performing dynamic memory allocation. Therefore, the subject technology described herein enables the neural network to run on such dedicated (e.g., resource-constrained) processors. Therefore, these benefits are understood to improve the computing functionality of a given electronic device, such as an end-user device, which may typically have fewer available computing resources than, for example, one or more cloud-based servers.
[0018] Figure 1 An exemplary network environment 100 is shown according to one or more implementations. However, not all of the depicted components may be used in all implementations, and one or more implementations may include additional or different components than those shown in the figures. Changes in the arrangement and types of these components may be made without departing from the spirit or scope of the claims set forth herein. Additional components, different components, or fewer components may be provided.
[0019] The network environment 100 includes a wireless audio output device 104, an electronic device 110, an electronic device 115, and a server 120. The network 106 may communicatively couple (directly or indirectly) the electronic device 110 and / or the server 120, the electronic device 115, and / or the server 120 and / or the electronic device 110 and / or the electronic device 115. In one or more specific implementations, the network 106 may be an interconnected network of devices that may include the Internet or be communicatively coupled to the Internet. Figure 1 , the wireless audio output device 104 is shown as not being directly coupled to the network 106; however, in one or more implementations, the wireless audio output device 104 may be directly coupled to the network 106. For purposes of explanation, the network environment 100 is shown as being directly coupled to the network 106. Figure 1 1. Network environment 100 is shown as including wireless audio output device 104, electronic device 110, electronic device 115, and server 120; however, network environment 100 may include any number of electronic devices and any number of servers.
[0020] The wireless audio output device 104 may be, for example, a wireless headset device, one or more wireless earbuds, a smart speaker, or generally any device that includes audio output circuitry and one or more wireless interfaces, such as a near field communication (NFC) radio, a WLAN radio, a Bluetooth radio, a Zigbee radio, and / or other radios. Figure 1 In the embodiment of the present invention, by way of example, the wireless audio output device 104 is depicted as a set of wireless earbuds. The wireless audio output device 104 may be and / or may include the following with respect to Figure 8 The wireless audio output device 104 may be paired with one or more of the electronic devices 110 and / or 115, such as via Bluetooth. In one embodiment, the wireless audio output device 104 may not include a main processor such as a CPU and / or a GPU, but may only include the following: Figure 2 A dedicated processor as further described in.
[0021] The electronic device 110 may be, for example, a desktop computer, a portable computing device such as a laptop computer, a smart phone, a peripheral device (e.g., a digital camera, a headset), a tablet device, a wearable device such as a watch, a band, etc. Figure 1 In the embodiment, the electronic device 110 is depicted as a desktop computer by way of example. The electronic device 110 may be and / or may include the following with respect to Figure 8 All or part of the electronic system described.
[0022] In one or more specific implementations, the electronic device 110 may provide a system for converting the neural network model into code (e.g., C code) in a specific programming language as described herein. Specifically, the subject system may include a neural network compiler for compiling the code. In one example, by using the compiled code, the subject system may create an executable software package for deployment on a target platform (such as the electronic device 115) with the assistance of the server 120. When executing the compiled code, the target platform may perform one or more given operations of the neural network model on a dedicated processor disposed on the target platform.
[0023] The electronic device 115 may be, for example, a portable computing device such as a laptop computer, a smart phone, a peripheral device (e.g., a digital camera, a headset), a tablet device, a wearable device such as a watch, a band, etc., or any electronic device. The electronic device may also include processors with different computing capabilities, including, for example, a CPU, a GPU, a neural processor, and / or a dedicated processor. Figure 1 In the embodiment, by way of example, the electronic device 115 is depicted as a smart phone device. In one or more specific implementations, the electronic device 115 may be and / or may include the following relative to the following relative to Figure 8 The electronic system is all or part of the electronic equipment.
[0024] In one or more specific implementations, the server 120 deploys the compiled code included in the executable software package to the target device for execution. In one or more specific implementations, the server 120 may send the executable software package to an intermediate device such as the electronic device 115 to enable deployment on a target device such as the wireless audio output device 104. In one example, the wireless audio output device 104 may be a target device for receiving a software package with compiled neural network code and for executing the compiled code in an operating environment of the wireless audio output device 104. As further described herein, the subject technology advantageously enables the wireless audio output device 104 to run the compiled neural network code without utilizing a framework. A framework may refer to a software environment that provides specific functionality as part of a larger software platform to facilitate software application development.
[0025] Figure 2 An exemplary software architecture for generating code for a neural network for execution on a dedicated processor according to one or more specific implementations is shown. For the purpose of explanation, the software architecture is described as consisting of Figure 1The software architecture may be provided by the electronic device 110, such as by a processor and / or memory of the electronic device 110; however, the software architecture may be implemented by any other electronic device. However, not all of the components depicted may be used in all implementations, and one or more implementations may include additional or different components than those shown in the figures. Changes in the arrangement and type of these components may be made without departing from the spirit or scope of the claims listed herein. Additional components, different components, or fewer components may be provided.
[0026] As shown, the software architecture includes a machine learning (ML) framework 220 and a neural network compiler 215, wherein the ML framework includes a code generator 230. The memory 240 includes a neural network model document file 244. In one example, each of the neural network model document files 244 may include at least information representing a set of operations to be performed by corresponding nodes from different layers of a given neural network model. In addition, information including descriptions of one or more input features and output features, data structures, and feature types may be included in a given neural network model document file.
[0027] The code generator 230 may obtain the NN model document file from the neural network model document file 244 and convert the NN model document file into code in a specific programming language so that it can be executed on a dedicated processor of the target device once compiled. The neural network compiler 215 obtains the generated code from the code generator 230 and compiles the code into a neural network binary executable file, which can be stored in the neural network executable file 242 and then deployed to one or more different target devices (e.g., wireless audio output device 104) for execution. Although the code generator 230 is shown as separate from the neural network compiler 215 for explanation purposes, in at least one specific implementation, the code generator 230 may be part of the neural network compiler 215, so that the neural network compiler 215 can convert a given network model file and generate code in a specific programming language, which is then compiled by the neural network compiler 215.
[0028] Despite Figure 2 In the example of FIG. 1 , the neural network compiler 215 is set on the electronic device 110, but in some specific implementations, such a compiler may be set on a specific electronic device that compiles the code for the neural network model and executes the compiled neural network model on the same device.
[0029] As described above, the neural network model may be compiled for a specific target platform and then deployed to a different device such as the wireless audio output device 104 for execution.
[0030] As shown, the wireless audio output device 104 includes a system on chip (SOC) 260. The SOC 260 includes a host processor 262 and a dedicated processor 264. The host processor 262 may include appropriate logic components, circuits, and / or codes capable of processing data and / or controlling the operation of the wireless audio output device 104. In this regard, the host processor 262 may be enabled to provide control signals to various other components of the wireless audio output device 104, respectively. In addition, the host processor 262 may enable an operating system to be implemented or otherwise execute code to manage the operation of the wireless audio output device 104. In one specific implementation, the dedicated processor 264 is a processor that is considered to be "always on" and continuously runs on the wireless audio output device 104. In this specific implementation, certain machine learning applications (such as predicting the movement of a person based on sensor data, detecting voice verbal voice triggers, and other types of machine learning applications, etc.) may be advantageously executed on the dedicated processor 264. In one example, the dedicated processor 264 may be utilized to perform operations from a compiled neural network model. In one or more implementations, the wireless audio output device 104 can communicate directly with the server 120. In one or more implementations, the wireless audio output device 104 can include only the dedicated processor 264 (eg, without including the host processor 262).
[0031] As further shown, in one implementation, the electronic device 115 includes a system on a chip (SOC) 250. The SOC 250 includes a dedicated processor 252, a CPU 254, and a GPU 255, as well as a neural processor 256 that can be used to execute operations from the compiled neural network model. In implementations where the dedicated processor 252 is a processor that is considered to be "always on" and continuously running on the electronic device 115, certain machine learning applications (such as predicting a person's movement based on sensor data, detecting voice spoken voice triggers, and other types of machine learning applications) can be advantageously executed on such a dedicated processor.
[0032] As further described herein, the code generator 230 may generate corresponding code based on a given neural network model file from the neural network model documentation file 244, which corresponding code may be compiled by the neural network compiler 215 for execution only on the dedicated processor 264 provided by the wireless audio output device 104.
[0033] As described herein, a CPU may refer to a main processor in a given electronic device that performs operations for basic arithmetic, logic, control, and input / output operations specified by instructions of a computer program or application, including some operations for neural network models. As described herein, a GPU may refer to a dedicated electronic circuit designed to perform operations for rendering graphics, which in many cases is also used to process computational workloads for machine learning operations (e.g., specified by instructions of a computer program or application). CPUs, GPUs, neural processors, and dedicated processors may each have different computing specifications and capabilities, depending on their respective specific implementations, in which each of the aforementioned components may provide different degrees of performance for certain operations compared to other components.
[0034] Recently, dedicated (e.g., specialized) hardware has been developed that is optimized for performing specific operations from a given NN. A given electronic device may include a neural processor that may be implemented as a circuit that performs various machine learning operations based on calculations including multiplication, addition, and accumulation. Such calculations may be arranged to perform, for example, convolution of input data. In one example, the neural processor is specifically configured to perform a machine learning algorithm, typically by operating on a predictive model such as a NN. In one or more specific implementations, an electronic device may include a dedicated processor and / or a neural processor in addition to a CPU and / or GPU.
[0035] Figure 3 An exemplary model format for data for an existing NN model documentation file and corresponding NN model code according to one or more implementations is shown. However, not all of the depicted components may be used in all implementations, and one or more implementations may include additional or different components than those shown in the figures. Changes in the arrangement and type of these components may be made without departing from the spirit or scope of the claims set forth herein. Additional components, different components, or fewer components may be provided.
[0036] As described herein, a neural network (NN) is a computational model that uses a collection of connected nodes to process input data based on machine learning techniques. A neural network can be represented by connecting different operations together, and is therefore referred to as a network. A model of a NN (e.g., a feedforward neural network) can be represented as a graph representing how these operations are connected together from an input layer through one or more hidden layers and finally connected to an output layer, wherein each layer includes one or more nodes, and wherein different layers perform different types of operations on corresponding inputs. However, it should be understood that the specific implementations described herein contemplate other types of neural networks. For example, a convolutional neural network (CNN) for execution on a given dedicated processor may be provided. In addition, a NN as described herein may also refer to a deep neural network corresponding to a neural network with multiple hidden layers. The number of layers and the number of nodes in each layer may be set as part of a neural network architecture. The settings for the architecture of a neural network (e.g., the number of layers, the connections between nodes of a layer, etc.) are also referred to as hyperparameters.
[0037] As described above, an existing NN model (e.g., a given NN model documentation file) can be converted into code in a programming language and compiled into a binary file for deployment on a target platform such as a wireless audio output device 104. As shown, the NN model documentation file 310 represents an existing NN model with information in a format different from a programming language. In one example, the NN model documentation file may conform to a specific model specification. The NN model documentation file 310 may include NN data types 324 (e.g., input features, output values, etc.) of NN data, and information about one or more NN layers 326. The NN data type 324 may include information about a data type or data structure (e.g., vector, matrix, array, etc.). The NN layer 326 includes information about the structure of the NN model, such as the number of layers and the number of nodes per layer, the connections between the nodes of the layer, and the functions or operations performed at each node in the nodes in the layer of the NN model. In one example, each layer in the NN layer 326 includes a name, a layer type (e.g., an input layer, a convolutional layer, a pooling layer, a linear rectifier unit layer, and a fully connected layer), an input name list, an output name list, and a parameter set specific to the layer type.
[0038] The transformed NN model code 330 includes code in a specific programming language (e.g., C language) that represents the aforementioned information from the NN model documentation file 310. For example, the transformed NN model code 330 includes operations 342, memory allocations 344, data formats 346, and data layers 350. The operations 342 correspond to corresponding operations performed at each layer of the NN. In one example, the operations 342 may include code for corresponding function calls for performing the operations for each layer of the NN and / or a set of parameters for the function calls. The data format 346 (e.g., data blobs, arrays, arrays of arrays, matrices) may correspond to code corresponding to the NN data type 324 and / or include code for specifying a compatible binary format for NN data utilized by a given dedicated processor of a target platform (e.g., the wireless audio output device 104). The data layers 350 may correspond to code for each layer of the NN, and the memory allocations 344 correspond to code for allocating memory portions based on a determined size for each layer of the NN and / or based on an amount of memory available at the target device. Figure 5 Determining the appropriate size for each layer of the NN is discussed in more detail.
[0039] When analyzing the NN model documentation file, the code generator 230 may perform various optimizations to generate code that is smaller and can run more efficiently on a dedicated processor such as a resource-constrained processor. For example, when analyzing the NN model documentation file, the code generator 230 may perform an operation fusion optimization, in which multiple operations are combined into the same code segment or function call. For example, the code generator 230 may perform a vertical fusion optimization, in which multiple operations (e.g., 2 to 3 operations) are combined. For example, a given set of operations may be represented as the following operations:
[0040] (1) ReLU with Z = X
[0041] (2) Convolution of A = Z
[0042] The code generator 230 may determine that if the results of operation (1) and / or operation (2) are not used by other operations (or layers), the code generator 230 may combine the above operations into a single combined operation, as shown below:
[0043] (3) ReLU convolution of A = X
[0044] The code generator 230 may also perform graph coloring optimization on the NN model document file. As referred to herein, graph coloring refers to the optimization of memory allocations for the layers of a neural network, which in one example involves determining which memory allocations are reused by the layers. Figure 5 One example of a memory allocation technique is described in more detail.
[0045] In one implementation, the code generator 230 may also generate code for debugging purposes (including, for example, data for a test network and / or a set of compilation flags and metadata) to indicate that the binary file is to be compiled for debugging or testing purposes.
[0046] In one specific implementation, the code generator 230 may also perform quantization of data included in the neural network based on, for example, the amount of memory (and / or other resources) available at a target device (e.g., a resource-constrained processor). In one example, such data may be in a floating point format, which in some computing architectures provides 32 bits of data precision. In some cases, the functionality of the network is not affected if the data is formatted in a different format that uses a smaller number of bits (e.g., lower precision) than the 32 bits described above for floating point values. Thus, the code generator 230 may perform quantization optimizations on floating point data and generate code for a data format that uses a smaller number of bits (e.g., 16 bits, 8 bits, 4 bits, etc.).
[0047] The following discussion relates to examples of code generated by code generator 230 of ML framework 220 from a given neural network model document.
[0048] The following example code defines a structure (e.g., a user-defined data type) for a neural network in the C programming language, including code indicating the operation of the neural network and / or the layer type of each of its layers:
[0049]
[0050]
[0051] The following code defines static storage allocations for data from a layer of a neural network, which in one example are determined based on the amount of memory of a target device (e.g., wireless audio output device 104):
[0052] static unsigned char buffer_color_0[625*4];
[0053] static unsigned char buffer_color_1[976*4];
[0054] The following code defines various binary formats for data (e.g., blob shape, blob topography, etc.) that may be the result of graph coloring optimization performed by the code generator 230:
[0055]
[0056]
[0057]
[0058] In one example, dependency information between corresponding layers of a network may be indicated in the following code:
[0059] static unsigned short topology_bin_info[]={1,3,1,2}
[0060] In the above code examples, each line of code with similar syntax corresponds to the operation of a given neural network model. By way of example, for a line in the order 1,3,1,2: 1 is the number of output blobs / tensors, 3 is the index of the output blob, 1 is the number of input blobs, and 2 is the index of the input blob.
[0061] Although the examples described herein are related to generating code using the C programming language, it should be understood that this is only one possible goal of the compiler. In a specific implementation, the compiler of the subject technology can generate LLVM IR (intermediate representation) or binary files.
[0062] By compiling a given neural network model into a binary file and pruning all unused configurations for any operation as described herein, the subject technology enables running a neural network without utilizing a deep learning or machine learning framework on an embedded processor (e.g., dedicated processor 252) with limited memory (e.g., within tens of kB) by selecting portions of a framework (e.g., ML framework 220) for inference tasks (or other machine learning tasks) for such a network.
[0063] As described herein, a convolutional neural network refers to a specific type of neural network, but uses different types of layers consisting of nodes that exist in three dimensions, which can vary between layers. In a convolutional neural network, nodes in a layer can be connected to only a subset of nodes in the previous layer. The final output layer can be fully connected and can be sized according to the number of classifiers. In an example of a convolutional neural network performing image classification for digital images representing numbers, an exemplary final output layer can have dimensions of [1×1×10]. In another example, the dimensions of the final output layer of a convolutional neural network used to recognize 500 different objects in an image (e.g., cats, dogs, people, bridges, etc.) can have dimensions of [1×1×500].
[0064] As described herein, a convolutional neural network model may include various combinations of the following types of layers, and in some cases multiples of each of the following types of layers and the order of these layers: input layers, convolutional layers, pooling layers, rectified linear unit layers (ReLU), and fully connected layers. A portion of the operations performed by a convolutional neural network includes obtaining a set of filters (or kernels) that are iterated over the input data based on one or more parameters. In one example, the depth of a convolutional layer may be equal to the number of filters used. It should be understood that given the hyperparameters of the convolutional neural network, the size of the different volumes at each layer may be mathematically determined.
[0065] In one example, a convolutional layer reads input data (e.g., a 3D input volume, a 2D image, or a 1D signal) using a kernel that reads in small segments at a time and steps across the entire input field. Each read may result in an input that is projected onto a filter graph and represents an internal interpretation of the input. Convolutional neural networks may be applied to human activity recognition data (e.g., sensor data corresponding to motion or movement), where the convolutional neural network model learns to map a given window of signal data to an activity in which the model reads across each window of data and prepares an internal representation of the window.
[0066] Convolutional neural networks are often run on cloud-based computing platforms due to the amount of data being processed. In such cases, memory management is often added as an afterthought, as cloud-based systems have no real memory issues (e.g., more computing power / larger memory is readily available). In contrast, it may be impossible or impractical to store all of the weights and resulting node values of a convolutional neural network in memory on a resource / memory-limited / constrained device (e.g., a mobile electronic device such as a smartphone).
[0067] Figure 4 An example of a convolutional neural network 400 is shown according to one or more implementations.
[0068] like Figure 4 As shown in the example of , convolutional neural network 400 shows intermediate data layers 402, 404, 406, 408 and 410. For the purpose of explanation, the intermediate data layers are shown as 2D objects, but it should be understood that the intermediate data layers may correspond to 3D input volumes. Intermediate data layers 402, 404, 406, 408 and 410 may be different types of layers, such as convolutional layers, ReLU layers, etc. Therefore, different intermediate data layers may have different dimensions. Different computing architectures may use different formats to store intermediate data layers. For example, when a convolutional neural network is processed on a dedicated processor (e.g., a motion processor), a specific binary format compatible with the architecture of the dedicated processor may be used to represent and store input data.
[0069] Convolutional neural network 400 is shown along a vertical time axis starting at t0 and ending at t3. The axis shows different and relative times of intermediate data layers that can be processed by an electronic device. For example, intermediate data layer 402 can be processed first, and then intermediate data layer 404 and intermediate data layer 406 can be processed in parallel at t1.
[0070] Convolutional neural network 400 also shows dependencies between different intermediate data layers. Thus, intermediate data layer 404 and intermediate data layer 406 both use the output of intermediate data layer 402; intermediate data layer 408 uses the output of intermediate data layer 406; and intermediate data layer 410 uses the output of intermediate data layer 408 and intermediate data layer 404. In one implementation, the hyperparameters and architecture (e.g., the number of layers and how the layers are connected) of convolutional neural network 400 may be included in the examples described above. Figure 2 and Figure 3 In various examples, convolutional neural network 400 can be executed on a dedicated processor of a single electronic device (e.g., a mobile device, a laptop, a desktop computer).
[0071] The dependencies between the layers of the convolutional neural network 400 can be used to infer the minimum number of memory allocations required to execute the convolutional neural network. Once the dependencies are known, the code generator 230 can determine at a particular execution point whether an output from a data layer will be needed in the future. If the output is needed, then memory allocations may be needed to hold the output until any intermediate data layers that need the output have used the output. In one example, the minimum number of memory allocations is based on the maximum number of memory allocations required to hold the dependent outputs during execution of the convolutional neural network.
[0072] Figure 5 A visualization of the results (eg, in a tabular format) of the inference process when executed by the code generator 230 is shown.
[0073] Figure 5 An exemplary table of memory allocation 500 according to one or more implementations is shown. The rows of table 500 correspond to times t0 to t3. The columns of table 500 represent three different memory allocations: memory allocation 502, memory allocation 504, and memory allocation 506. Labels B1, B2, B3, B4, and B5 correspond to intermediate data layer 402, intermediate data layer 404, intermediate data layer 406, intermediate data layer 408, and intermediate data layer 410.
[0074] The following discussion refers to times (e.g., t0) as if the convolutional neural network is actually running. However, based on the dependency information and the relative order of execution time of operations in the network, the code generator 230 performs the inference process without actually running the convolutional neural network. The dependency information can be generated as part of the code corresponding to the convolutional neural network. For example, the dependency information about the convolutional neural network 400 can be represented as:
[0075] B1:B2,B3
[0076] B2:B5
[0077] B3:B4
[0078] B4:B5
[0079] B5: empty value (for example, the output of B1 is used by B2 and B3, the output of B2 is used by B5, etc.)
[0080] The following discussion describes how the code generator 230 determines memory allocations for a network. In one or more specific implementations, the total amount of memory available for allocation may be determined based at least in part on the amount of available memory for a given target device (e.g., a dedicated processor provided by the wireless audio output device 104). For example, the code generator 230 may utilize information about the total amount of available memory for a target device (e.g., the wireless audio output device 104), which may be provided in a database or another source such as a table (e.g., a lookup table) that includes corresponding entries for various target devices and related information about hardware capabilities (e.g., minimum and / or maximum memory allocation sizes, etc.) and the total amount of memory for such target devices. In one specific implementation, based on previous allocations (if any) for the network, the code generator 230 may track the available amount of memory relative to the total amount of memory for the target device.
[0081] For example, starting at t0, the first memory allocation (memory allocation 502) is used to hold data about B1. Then, at t1, both the intermediate data layer 404 (B2) and the intermediate data layer 406 (B3) require memory allocation. Therefore, the code generator 230 can perform a check to determine the content stored in the memory allocation 502. As previously described, B1 is currently stored in the memory allocation 502. Then, the code generator 230 can access the dependency information to determine whether B1 is used by other intermediate data layers of the network. In this example, B1 is used by both B2 and B3. Therefore, the memory allocation 502 may not be allocated to B2 or B3. Therefore, two new memory allocations are required, namely, the memory allocation 504 and the memory allocation 506. These allocations are allocated to B2 and B3 respectively by the code generator 230.
[0082] Transitioning to t2, intermediate data layer 408 (B4) requires a memory allocation. Again, a check may be performed to see if an existing memory allocation may be reused. B1 is still in memory allocation 502, but because both B2 and B3 are now fully completed, the data from B1 is no longer needed. Therefore, memory allocation 502 may be reallocated to B4. Similarly, at t3, memory allocation 506 may be reallocated to B5 because B3 is no longer needed. Therefore, based on the dependency information, code generator 230 may infer that a minimum number of three memory allocations are required to execute convolutional neural network 400, which is the maximum number required at any point after traversing the dependency tree (e.g., performing a simulated execution of the convolutional neural network by code generator 230).
[0083] The code generator 230 may also determine the intermediate data layers allocated to the memory allocation during execution and generate code for such memory allocation. For example, both B1 and B4 are used by the memory allocation 502. The allocation information may be determined while determining how many memory allocations are needed.
[0084] Next, the code generator 230 may determine the memory storage size required for each memory allocation in the minimum number of memory allocations. Different computing architectures may allocate memory in different ways. For example, some computing architectures allow linear memory allocations, such as in some types of special-purpose processors. Similarly, different computing architectures may have different requirements for minimum or maximum memory allocation sizes. As described above, the code generator 230 may determine the total amount of available memory on the target device (e.g., the wireless audio output device 104) so as to determine the amount of available memory for the corresponding memory allocation based on previous allocations (if any) (e.g., this will likely reduce the amount of available memory).
[0085] In one implementation, the code generator 230 may iterate over each memory allocation to determine the amount of memory storage to reserve. Figure 4 , the code generator 230 may examine the underlying intermediate data layers of B1 and B4. As previously described, each intermediate data layer may be considered a 3D input volume. In one example, the code generator 230 may determine the memory storage required for the intermediate data layer based on the product of the dimensions of the intermediate data layer and the size of the data at the entry point in the volume. In addition, the code generator 230 may examine the resource constraints for the target device (e.g., the total amount of memory on the wireless audio output device 104 and the current available amount of memory) to further determine whether the required memory size for such allocation is possible, and if so, generate code for memory allocation accordingly.
[0086] Some computer architectures may allow memory allocations to be made using linear memory. In such cases, the code generator 230 may determine the size of the memory allocation based on the maximum total size of the intermediate data layers of any layer that will reuse the memory allocation. For example, this may be expressed as max(W B1 W B4 ). In other cases where textures or linear memory may not be used, the code generator 230 may determine the size based on both the maximum width and maximum height of the stored texture. For example, this may be expressed as max(W B1 H B1 ,W B4 H B4 ). The code generator 230 may determine the amount of allocated memory space required based on the depth information of the volume. For example, when determining the size of the memory allocation, the code generator 230 may process the [32×32×3] volume as three consecutive [32×32] arrays (e.g., [32×96]) volumes. In addition, the code generator 230 may check the resource constraints for the target device (e.g., the total amount of memory on the wireless audio output device 104 and the current available amount of memory) to further determine whether the required memory size for such allocation is feasible, and if so, generate code for such allocation.
[0087] Figure 6 A flowchart of an exemplary process 600 for generating code for a neural network model according to one or more specific implementations is shown. For the purpose of explanation, this article mainly refers to Figure 2 The process 600 is described by the components of the software architecture of Figure 1 The process 600 is performed by one or more processors of the electronic device 110. However, the process 600 is not limited to the electronic device 110, and one or more blocks (or operations) of the process 600 may be performed by one or more other components of other suitable devices (such as by the electronic device 115). Further for the purpose of explanation, the blocks of the process 600 are described herein as occurring sequentially or linearly. However, multiple blocks of the process 600 may occur in parallel. In addition, the blocks of the process 600 do not have to be performed in the order shown, and / or one or more blocks of the process 600 do not have to be performed and / or may be replaced by other operations.
[0088] The ML framework 220 receives a neural network model in a model format that includes information about a set of layers of the neural network model, each layer in the set of layers including a set of corresponding operations (610). In one example, the NN model includes multiple layers that include operations that can be executed on a dedicated processor of a target platform. In one example, the target platform can be a different electronic device, such as the wireless audio output device 104.
[0089] The code generator 230 generates a neural network (NN) code based on the neural network model, the NN code using a programming language different from the model format, and the NN code includes a corresponding memory allocation for each corresponding layer in a set of layers of the neural network model (612). In one example, the code includes specific code (e.g., C code) corresponding to the memory allocation for each layer in the set of layers. In addition, the corresponding memory allocation for each corresponding layer is determined based at least in part on a resource constraint (e.g., a total amount of memory and / or an amount of available memory) of a target device (e.g., a wireless audio output device 104).
[0090] The neural network compiler 215 compiles the NN code into a binary format (614). In one example, the binary format is compatible with the hardware architecture of the dedicated processor of the target platform (e.g., the wireless audio output device 104).
[0091] The neural network compiler 215 generates a package for deploying the compiled NN code on a target device (616).
[0092] Figure 7 An exemplary process 700 for determining memory allocation to generate code for a convolutional neural network according to one or more specific implementations is shown. For the purpose of explanation, this document mainly refers to Figure 2 The process 700 is described with reference to the components of the electronic device shown in FIG. Figure 1 The process 700 is performed by one or more processors of the electronic device 110. However, the process 700 is not limited to the electronic device 110, and one or more blocks (or operations) of the process 700 may be performed by one or more other components of other suitable devices. Further for the purpose of explanation, the blocks of the process 700 are described herein as occurring sequentially or linearly. However, multiple blocks of the process 700 may occur in parallel. In addition, the blocks of the process 700 do not have to be performed in the order shown, and / or one or more blocks of the process 700 do not have to be performed and / or may be replaced by other operations.
[0093] Code generator 230 determines dependencies between intermediate data layers of the neural network (710). In one example, the neural network is a convolutional neural network based on a NN documentation file (e.g., from neural network model documentation file 244). The NN documentation file may identify the number of intermediate data layers, dependencies between layers, dimensions of each layer (e.g., height, width, depth), and an execution order of the layers. In some examples, ML framework 220 is configured to analyze the NN documentation file.
[0094] Code generator 230 determines the dimensions of the neural network (712). In some examples, the size of the intermediate data layer is obtained from the metadata. In some examples, the size of the intermediate data layer is calculated based on hyperparameters of the neural network.
[0095] The code generator 230 determines a minimum number of memory allocation portions for executing the neural network based on the dependencies (714). The minimum number of memory allocation portions may be inferred based on the order of the intermediate data layers within the neural network. For example, if three later intermediate data layers use data from earlier intermediate data layers, the data in the earlier intermediate data layers may be stored at least until the three later intermediate data layers are executed. In one example, the minimum number of dependencies is stored as part of the metadata for the neural network. In addition, the code generator 230 determines a designation for assigning the intermediate data layers to the memory allocation portions. In one example, this is accomplished by traversing the architecture, as if running the neural network to determine which intermediate data layer is stored in which data storage portion when the neural network is to be run. In addition, more than one intermediate data layer may be assigned to a memory allocation portion. In some examples, different memory allocation portions are assigned to different intermediate data layers. The resulting designation may be stored as a table identifying the intermediate data layers and the memory allocation portions assigned to the intermediate data layers.
[0096] The code generator 230 determines the memory allocation size for each corresponding memory allocation portion in the memory allocation portion based on the dimensions and the dependencies (716). The code generator 230 generates the memory allocation size for each corresponding data storage portion based on the dimensions and the dependencies. For example, the dependencies may determine which intermediate data layer is allocated to the memory allocation portion as described above. Then, the dimensions of one or more intermediate data layers allocated to the corresponding memory allocation portion may be checked to determine the largest intermediate data layer by volume. The memory allocation size for the corresponding memory allocation portion may be set to at least the size of the largest intermediate data layer. The type of execution environment may affect the memory allocation size. For example, if memory may not be allocated using a texture or linear method, the memory allocation size may be greater than the size of the largest intermediate data layer.
[0097] The code generator 230 generates code for allocating memory on the target platform (eg, the wireless audio output device 104 ) for each memory allocation portion based at least in part on the corresponding determined memory allocation sizes ( 718 ).
[0098] When compiled and deployed to a target device such as the wireless audio output device 104, memory on the target device may be allocated for each memory allocation portion of the neural network according to the corresponding determined memory allocation size of each memory allocation portion. After allocation, the designated table between the intermediate data layer and the data storage portion may be updated to include the memory address of the allocated memory. In one example, the memory for the data storage portion is allocated as a continuous block, but is actually divided into the number of memory portions. During execution of the neural network, pointers may be moved around the blocks corresponding to the memory portions in the continuous block.
[0099] Figure 8 An electronic system 800 is shown that can be used to implement one or more implementations of the subject technology. The electronic system 800 can be Figure 1 The electronic device 110, the electronic device 115 and / or the server 120 are shown and / or may be a part thereof. The electronic system 800 may include various types of computer-readable media and interfaces for various other types of computer-readable media. The electronic system 800 includes a bus 808, one or more processing units 812, a system memory 804 (and / or a buffer), a ROM 810, a permanent storage device 802, an input device interface 814, an output device interface 806, and one or more network interfaces 816, or a subset and variation thereof.
[0100] The bus 808 generally represents all system buses, peripheral busses, and chipset buses that communicatively connect the many internal devices of the electronic system 800. In one or more specific implementations, the bus 808 communicatively connects one or more processing units 812 with the ROM 810, the system memory 804, and the permanent storage device 802. The one or more processing units 812 retrieve instructions to be executed and data to be processed from these various memory units in order to perform the processes disclosed in the present subject matter. In different specific implementations, the one or more processing units 812 can be a single processor or a multi-core processor.
[0101] ROM 810 stores static data and instructions required by one or more processing units 812 and other modules of the electronic system 800. On the other hand, permanent storage device 802 can be a read-write memory device. Permanent storage device 802 can be a non-volatile memory unit that stores instructions and data even when the electronic system 800 is turned off. In one or more specific implementations, a mass storage device (such as a magnetic disk or optical disk and its corresponding disk drive) can be used as permanent storage device 802.
[0102] In one or more implementations, a removable storage device (such as a floppy disk, a flash drive, and its corresponding disk drive) may be used as the permanent storage device 802. Like the permanent storage device 802, the system memory 804 may be a read-write memory device. However, unlike the permanent storage device 802, the system memory 804 may be a volatile read-write memory, such as a random access memory. The system memory 804 may store any of the instructions and data that one or more processing units 812 may need at runtime. In one or more implementations, the processes disclosed in the subject matter are stored in the system memory 804, the permanent storage device 802, and / or the ROM 810. The one or more processing units 812 retrieve instructions to be executed and data to be processed from these various memory units in order to perform the processes of one or more implementations.
[0103] The bus 808 is also connected to an input device interface 814 and an output device interface 806. The input device interface 814 enables a user to transmit information and select commands to the electronic system 800. Input devices that can be used with the input device interface 814 may include, for example, an alphanumeric keyboard and a pointing device (also referred to as a "cursor control device"). The output device interface 806 may, for example, enable the display of images generated by the electronic system 800. Output devices that can be used with the output device interface 806 may include, for example, a printer and a display device, such as a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, a flexible display, a flat panel display, a solid-state display, a projector, or any other device for outputting information. One or more specific implementations may include a device that acts as both an input device and an output device, such as a touch screen. In these specific implementations, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including acoustic, voice, or tactile input.
[0104] Finally, if Figure 8 As shown, bus 808 also couples electronic system 800 to one or more networks and / or to one or more network nodes, such as Figure 1 800. In this manner, electronic system 800 may be part of a computer network, such as a LAN, a wide area network ("WAN"), or an intranet, or may be part of a network of networks, such as the Internet. Any or all components of electronic system 800 may be used with the subject disclosure.
[0105] One aspect of the present technology may include collecting and using data obtained from specific and legitimate sources to improve the delivery of inspirational content or any other content that may be of interest to users. The present disclosure contemplates that, in some instances, the collected data may include personal information data that uniquely identifies or can be used to identify a specific person. Such personal information data may include demographic data, location-based data, online identifiers, phone numbers, email addresses, home addresses, data or records related to a user's health or fitness level (e.g., vital sign measurements, medication information, exercise information), date of birth, or any other personal information.
[0106] The present disclosure recognizes that the use of such personal information data in the present technology can be used to benefit users. For example, personal information data can be used to deliver targeted content that users may be more interested in based on their preferences. Therefore, the use of such personal information data enables users to have greater control over the content delivered. In addition, the present disclosure also anticipates other uses of personal information data that benefit users. For example, health and fitness data can be used according to the user's preferences to provide insights into their overall health, or can be used as positive feedback to individuals who use technology to pursue health goals.
[0107] The present disclosure envisions that entities responsible for collecting, analyzing, disclosing, transmitting, storing or otherwise using such personal information data will comply with established privacy policies and / or privacy practices. Specifically, such entities will be expected to implement and consistently apply privacy practices that are generally recognized as meeting or exceeding the requirements of the industry or government that maintain user privacy. Such information about the use of personal data should be prominently and easily accessible to users, and should be updated as the collection and / or use of data changes. The user's personal information should be collected only for legitimate use. In addition, such collection / sharing should only occur after receiving user consent or other legal basis specified in applicable law. In addition, such entities should consider taking any necessary steps to defend and safeguard access to such personal information data and ensure that others who have access to personal information data comply with their privacy policies and processes. In addition, such entities may subject themselves to third-party assessments to demonstrate that they comply with widely accepted privacy policies and practices. In addition, policies and practices should be adjusted for specific types of personal information data that are collected and / or accessed, and applied to applicable laws and standards, including specific considerations that are exclusive to jurisdictions that can be used to impose higher standards. For example, in the United States, collection or access to certain health data may be governed by federal and / or state laws, such as the Health Insurance Portability and Accountability Act (HIPAA); while health data in other countries may be subject to other regulations and policies and should be handled accordingly.
[0108] Regardless of the foregoing, the present disclosure also contemplates implementation schemes in which users selectively block the use or access of personal information data. That is, the present disclosure contemplates that hardware elements and / or software elements may be provided to prevent or block access to such personal information data. For example, with respect to an advertising delivery service, the technology of the present invention may be configured to allow a user to choose to "opt in" or "opt out" at any time during or after registration for the service to participate in the collection of personal information data. In another example, a user may choose not to provide emotion-related data for a target content delivery service. As another example, a user may choose to limit the length of time that emotion-related data is retained, or completely block the development of basic emotional conditions. In addition to providing "opt-in" and "opt-out" options, the present disclosure contemplates providing notifications related to access or use of personal information. For example, a user may be notified that their personal information data will be accessed when downloading an application, and then reminded again just before the personal information data is accessed by the application.
[0109] In addition, it is an object of the present disclosure that personal information data should be managed and processed to minimize the risk of unintentional or unauthorized access or use. Risks can be minimized by limiting data collection and deleting data once it is no longer needed. In addition, and when applicable, including in certain health-related applications, data de-identification can be used to protect the privacy of users. De-identification can be facilitated by removing identifiers, controlling the amount or specificity of stored data (e.g., collecting location data at a city level rather than at an address level), controlling how data is stored (e.g., aggregating data across users), and / or other methods such as differential privacy, where appropriate.
[0110] Thus, while the present disclosure broadly covers the use of personal information data to implement one or more of the various disclosed embodiments, the present disclosure also contemplates that various embodiments may also be implemented without access to such personal information data. That is, various embodiments of the present technology will not fail to function properly due to the lack of all or a portion of such personal information data. For example, content may be selected and delivered to a user based on aggregated non-personal information data or an absolute minimum amount of personal information, such as content processed only on the user's device or other non-personal information that may be used for content delivery services.
[0111] Implementations within the scope of the present disclosure may be implemented in part or in whole using a tangible computer-readable storage medium (or multiple tangible computer-readable storage media of one or more types) programmed with one or more instructions. Tangible computer-readable storage media may also be non-transitory in nature.
[0112] Computer-readable storage media can be any storage media that can be read, written, or otherwise accessed by a general or special computing device, including any processing electronics and / or processing circuitry capable of executing instructions. For example, without limitation, computer-readable media can include any volatile semiconductor memory, such as RAM, DRAM, SRAM, T-RAM, Z-RAM, and TTRAM. Computer-readable media can also include any non-volatile semiconductor memory, such as ROM, PROM, EPROM, EEPROM, NVRAM, flash memory, nvSRAM, FeRAM, FeTRAM, MRAM, PRAM, CBRAM, SONOS, RRAM, NRAM, racetrack memory, FJG, and Millipede memory.
[0113] In addition, the computer-readable storage medium may include any non-semiconductor memory, such as optical disk storage, magnetic disk storage, magnetic tape, other magnetic storage devices, or any other medium capable of storing one or more instructions. In one or more specific implementations, the tangible computer-readable storage medium may be directly coupled to the computing device, while in other specific implementations, the tangible computer-readable storage medium may be indirectly coupled to the computing device, for example, via one or more wired connections, one or more wireless connections, or any combination thereof.
[0114] Instructions can be directly executable, or can be used to develop executable instructions. For example, instructions can be implemented as executable or non-executable machine code, or can be implemented as high-level language instructions that can be compiled to produce executable or non-executable machine code. In addition, instructions can also be implemented as data, or can include data. Computer executable instructions can also be organized in any format, including routines, subroutines, programs, data structures, objects, modules, applications, applets, functions, etc. As recognized by those skilled in the art, the details including but not limited to the number, structure, sequence and organization of instructions can be significantly different without changing the underlying logic, function, processing and output.
[0115] Although the above discussion mainly involves microprocessors or multi-core processors that execute software, one or more implementations are performed by one or more integrated circuits such as ASICs or FPGAs. In one or more implementations, such integrated circuits execute instructions stored on the circuits themselves.
[0116] Those skilled in the art will recognize that the various illustrative frames, modules, elements, parts, methods and algorithms described herein can be implemented as electronic hardware, computer software or a combination of the two. In order to illustrate this interchangeability of hardware and software, various illustrative frames, modules, elements, parts, methods and algorithms have been generally described above according to functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the entire system. The technician can implement the described functionality in different ways for each specific application. Various parts and frames can be arranged differently (e.g., arranged in different orders, or divided in different ways) without departing from the scope of the subject technology.
[0117] It should be understood that the specific order or the hierarchical structure of the frame in the process disclosed by the present invention are the illustration of exemplary methods. Based on the design preferred requirements, it should be understood that the specific order or the hierarchical structure of the frame in the process can be rearranged or all the frames shown are executed. Any frame in these frames can be executed simultaneously. In one or more specific implementations, multitasking and parallel processing may be advantageous. In addition, the division of each system component in the above-mentioned specific implementation should not be understood as requiring such division in all specific implementations, and it should be understood that program components and systems can be generally integrated together in a single software product or encapsulated in multiple software products.
[0118] As used in this specification and any claims of this patent application, the terms "base station", "receiver", "computer", "server", "processor" and "memory" all refer to electronic devices or other technical devices. These terms exclude people or groups of people. For the purpose of this specification, the term "display" or "displaying" means displaying on an electronic device.
[0119] As used herein, the phrase "at least one of" following a list of items, any of which is separated by the terms "and" or "or", modifies the list as a whole, rather than each member (i.e., each item) of the list. The phrase "at least one of" does not require selection of at least one of each item listed; rather, the phrase allows for a meaning that includes at least one of any one item and / or at least one of any combination of items and / or at least one of each item. For example, the phrases "at least one of A, B, and C" or "at least one of A, B, or C" each refer to only A, only B, or only C; any combination of A, B, and C; and / or at least one of each of A, B, and C.
[0120] The predicate words "configured to," "operable to," and "programmed to" do not imply any specific tangible or intangible modification of a subject matter but are intended to be used interchangeably. In one or more specific implementations, a processor configured to monitor and control an operation or component may also mean that the processor is programmed to monitor and control the operation or that the processor is operable to monitor and control the operation. Similarly, a processor configured to execute code may be interpreted as a processor programmed to execute code or operable to execute code.
[0121] Phrases such as aspect, this aspect, on the other hand, some aspects, one or more aspects, a specific implementation, this specific implementation, another specific implementation, some specific implementations, one or more specific implementations, an embodiment, this embodiment, another embodiment, some embodiments, one or more embodiments, configuration, this configuration, other configurations, some configurations, one or more configurations, subject technology, disclosure, the present disclosure, other variations thereof, etc. are for convenience and do not mean that the disclosure involving such one or more phrases is essential to the subject technology, nor does it mean that such disclosure applies to all configurations of the subject technology. Disclosures involving such one or more phrases may apply to all configurations or one or more configurations. Disclosures involving such one or more phrases may provide one or more examples. Phrases such as aspect or some aspects may refer to one or more aspects and vice versa, and this applies similarly to the other aforementioned phrases.
[0122] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any implementation described herein as "exemplary" or as an "example" is not necessarily to be construed as preferred or advantageous over other implementations. Furthermore, to the extent that the terms "including," "having," and the like are used in the specification or claims, such terms are intended to be inclusive, in a manner similar to the way the term "comprising" is interpreted when used as a transitional word in a claim.
[0123] All structural and functional equivalents to the elements of the various aspects described throughout this disclosure that are known or later become known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be covered by the claims. In addition, nothing disclosed herein is intended to be made available to the public regardless of whether the disclosure is explicitly stated in the claims. No claim element shall be interpreted under the provisions of 35 U.S.C. §112(f) unless the element is explicitly stated using the phrase “means for…” or, in the case of a method claim, the element is stated using the phrase “step for…”
[0124] The previous description is provided to enable those skilled in the art to practice the various aspects described herein. Various modifications of these aspects are apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. Therefore, the claims are not intended to be limited to the aspects shown herein, but are intended to make the full scope consistent with the language claims, wherein reference to elements in singular values is not intended to mean "only one", but refers to "one or more", unless specifically noted. Unless otherwise specifically stated, the term "some" refers to one or more. Male pronouns (e.g., his) include female and neutral (e.g., her and its), and vice versa. Titles and subtitles (if any) are used only for convenience and do not limit the subject disclosure.
Claims
1. A method for generating a package, the method comprising: Receiving, at a device, a neural network model in a model format, the model format comprising information about a set of layers of the neural network model, each layer in the set of layers comprising a set of corresponding operations; generating, at the device, neural network (NN) code from the neural network model, the NN code being in a programming language different from a format of the model and including a respective memory allocation for each respective layer in the set of layers of the neural network model, wherein the generating comprises determining the respective memory allocation for each respective layer based at least in part on a resource constraint of a target device separate from the device; compiling the NN code into a binary format at the device, the compiling comprising pruning a set of unused configurations of operations of the neural network model, wherein the pruning reduces a size of the NN code in the binary format by performing operation fusion optimization in which multiple operations are combined into the same code segment or function call; as well as A package is generated at the device for deploying the compiled NN code on the target device for execution by a dedicated processor of the target device.
2. The method according to claim 1, wherein the model format comprises a NN model document file adopting a specific specification for a neural network, and the programming language is an imperative programming language.
3. The method of claim 1 , wherein the package is executed by the target device on the dedicated processor without utilizing a machine learning framework.
4. The method of claim 1, wherein generating the NN code further comprises: Determining dependencies between intermediate layers of the neural network model; Determining the dimension of the intermediate layer in the neural network model; determining a minimum number of memory allocation portions for executing the neural network model based on the dependencies; determining a memory allocation size for each respective one of the memory allocation portions based on the dimensions and the dependencies; as well as Generating code for allocating memory on the target device for each memory allocation portion based at least in part on the determined corresponding memory allocation size, wherein the NN code includes the code.
5. The method of claim 1, wherein generating the NN code further comprises generating a set of compiler flags or a set of test data for inclusion in the compiled NN code.
6. The method according to claim 1, wherein generating a neural network (NN) code according to the neural network model further comprises: Determining a set of operations to be performed in a sequential manner in an execution flow of the neural network model, the set of operations being determined based on a lack of dependencies between the set of operations; as well as The set of operations is combined for compilation.
7. The method of claim 1 , wherein the set of layers comprises a set of intermediate data layers, and for each respective intermediate data layer in the set of intermediate data layers: Respective code is generated to allocate respective portions of memory for the respective intermediate data layers, wherein allocating the respective portions of memory is based on which intermediate layers will be executed concurrently as the respective intermediate data layers on the target device at a particular time.
8. The method of claim 7, wherein a first portion of the memory is allocated for the first intermediate data tier and a second portion of the memory is allocated for the second intermediate data tier.
9. The method according to claim 1, wherein generating NN code according to the neural network model further comprises: Based at least in part on the resource constraints of the target device, operations having a higher precision are quantized into corresponding operations having a lower precision.
10. The method of claim 1, wherein the target device includes an operating environment utilizing the dedicated processor, the dedicated processor utilizing less power than a main processor of the target device, the dedicated processor having lower computing power than the main processor, and the dedicated processor being powered on at all times, wherein the package is loaded into a memory of the target device for execution by the dedicated processor.
11. A system for generating a package, the system comprising: processor; a memory device comprising instructions which, when executed by the processor, cause the processor to: Receiving a neural network model in a model format, the model format comprising information about a set of layers of the neural network model, each layer in the set of layers comprising a set of corresponding operations; generating neural network (NN) code based on the neural network model, the NN code being in a programming language different from a format of the model and including a respective memory allocation for each respective layer in the set of layers of the neural network model, wherein generating the NN code comprises determining the respective memory allocation for each respective layer based at least in part on a resource constraint of a target device separate from the system; compiling the NN code into a binary format, the compiling comprising pruning a set of unused configurations of operations of the neural network model, wherein the pruning reduces the size of the NN code in the binary format by performing operation fusion optimization in which multiple operations are combined into the same code segment or function call; as well as A package is generated for deploying the compiled NN code on the target device for execution by a dedicated processor of the target device.
12. The system of claim 11, wherein the model format comprises a NN model document file adopting a specific specification for a neural network, and the programming language is an imperative programming language.
13. The system of claim 11, wherein generating the NN code further causes the processor to: Determining dependencies between intermediate layers of the neural network model; Determining the dimension of the intermediate layer in the neural network model; determining a minimum number of memory allocation portions for executing the neural network model based on the dependencies; determining a memory allocation size for each respective one of the memory allocation portions based on the dimensions and the dependencies; as well as Generating code for allocating memory on the target device for each memory allocation portion based at least in part on the determined corresponding memory allocation size, wherein the NN code includes the code.
14. The system of claim 11, wherein generating the NN code further causes the processor to: A set of compiler flags or a set of test data is generated for inclusion in the compiled NN code.
15. The system of claim 11, wherein generating the NN code according to the neural network model further comprises: Operations with higher precision are quantized to corresponding operations with lower precision.
16. The system of claim 11, wherein generating the NN code further causes the processor to: Determining a set of operations to be performed in a sequential manner in an execution flow of the neural network model, the set of operations being determined based on a lack of dependencies between the set of operations; and The set of operations is combined for compilation.
17. The system of claim 11, wherein the set of layers comprises a set of intermediate layers, and for each corresponding intermediate data layer in the set of intermediate layers, the processor is further caused to: Respective code is generated to allocate respective portions of memory for the respective intermediate data layers, wherein allocating the respective portions of memory is based on which intermediate layers will be executed concurrently as the respective intermediate data layers on the target device at a particular time.
18. The system of claim 17, wherein a first portion of the memory is allocated for the first intermediate data tier and a second portion of the memory is allocated for the second intermediate data tier.
19. The system of claim 11, wherein the target device includes an operating environment utilizing the dedicated processor, the dedicated processor utilizing less power than a main processor of the target device, the dedicated processor having lower computing power than the main processor, and the dedicated processor being powered on at all times.
20. A non-transitory computer-readable medium comprising instructions that, when executed by a computing device, cause the computing device to perform operations comprising: Receiving, by the computing device, a neural network model in a model format, the model format comprising information about a set of layers of the neural network model, each layer in the set of layers comprising a set of corresponding operations; generating, by the computing device, neural network (NN) code from the neural network model, the NN code being in a programming language different from the model format and including a respective memory allocation for each respective layer in the set of layers of the neural network model, wherein the generating comprises determining the respective memory allocation for each respective layer based at least in part on a resource constraint of a target device separate from the computing device; compiling, by the computing device, the NN code into a binary format, the compiling comprising pruning a set of unused configurations of operations of the neural network model, wherein the pruning reduces a size of the NN code in the binary format by performing operation fusion optimization in which multiple operations are combined into the same code segment or function call; as well as A package is generated for deploying the compiled NN code on the target device for execution by a dedicated processor of the target device.
Citation Information
Patent Citations
Systems and Methods of Memory Allocation for Neural Networks
US20180088996A1