Controllable neural network sparsity through dynamic activation functions
By integrating a dynamic activation function layer that adjusts sparsity based on system resources, neural networks can optimize resource usage and maintain accuracy in resource-constrained environments.
Patent Information
- Application Number
- PCT/US2023/036290
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-30
- Publication Date
- 2025-05-08
AI Technical Summary
Existing neural networks lack the ability to dynamically adjust sparsity in response to changes in available system resources, leading to suboptimal trade-offs between accuracy and resource consumption.
Implementing a dynamic activation function layer in neural networks that can adjust its sparsity based on current system resources, allowing for real-time optimization of resource usage without retraining the network.
This approach enables neural networks to maintain acceptable accuracy while reducing resource consumption, particularly in resource-constrained environments like mobile devices.
Smart Images

Figure US2023036290_08052025_PF_FP_ABST
Abstract
Description
[0001] CONTROLLABLE NEURAL NETWORK SPARSITY THROUGH DYNAMIC ACTIVATION FUNCTIONS
[0002] BACKGROUND
[0003] This specification generally relates to processing inputs using neural networks.
[0004] Neural networks are machine learning models that employ one or more layers of nonlinear units to predict an output for a received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to the next layer in the network, e.g., the next hidden layer or the output layer. Each layer of the network generates an output from a received input in accordance with current values of a respective set of parameters.
[0005] Some or all of the layers of the neural network can apply an activation function as part of generating the output of the layer. An activation function is an element-wise, nonlinear function that is applied independently of each element of a given input. One example of an activation function is the rectified linear unit (ReLU) function, which, for a given input x, outputs zero if x is less than zero and outputs x if x is greater than or equal to zero.
[0006] SUMMARY
[0007] This specification generally describes techniques for dynamically controlling the sparsity of a neural network that is deployed on a computing device, e.g., on a hardware accelerator of the computing device.
[0008] As used in this specification, the "sparsity" of a neural network refers to the sparsity of the outputs of one of the layers of the neural network, i.e., the fraction of the values within the outputs of the layer of the neural network that are zero.
[0009] The subject matter described in this specification can be implemented in particular embodiments so as to realize one or more of the following advantages.
[0010] System operating conditions of mobile systems, e.g., mobile phones, wearable devices, augmented reality devices, virtual reality devices, and other edge devices, can significantly change the usable resources (e.g., battery life, thermal power, and so on) that are available for running applications.
[0011] Many of the applications that consume significant resources on a mobile system are those that require processing one or more inputs using a neural network deployed on a hardware accelerator of the device. Many hardware accelerators, e.g., application-specific integrated circuits (ASICs), for accelerating the computations of neural networks, e.g.. that include hardware for performing matrix multiplications in hardware, include hardware or are controlled by software that allows sparse matrices and higher-order tensors to be processed more efficiently, e.g., with decreased latency and while consuming less power and other resources, i .e., consuming less of the available resource budget. The "‘resource budget" of the mobile system can include the power, thermal, and and / or performance budget currently allocated to the mobile system.
[0012] How ever, increasing the sparsity of the intermediate tensors generated by a given neural network while processing a given input generally reduces the accuracy of the output generated by the neural network for the given input.
[0013] Thus, by adjusting the sparsity, a system could implement a trade-off between accuracy and resource consumption when processing inputs using the given neural network when deployed on the hardware accelerator.
[0014] However, existing techniques require that the (average) amount of sparsity’ in the outputs of a given layer of a neural network be fixed before the neural network is deployed. Thus, existing systems cannot dynamically adjust the sparsity to account for changes in available system resources. Instead, some systems turn to limiting the operating frequency of the hardware IP block on which the neural network is deployed. This results in lower FPS (frames / second) and longer latency, degrading the user experience.
[0015] This specification, on the other hand, describes techniques for dynamically adjusting the sparsity of the output of a given layer of a neural network to account for the available system resources, e.g., so that when more resources are available, sparsity is lower and accuracy is improved while, when fewer resources are available, sparsity is higher and accuracy is reduced (while still being maintained at an acceptable level).
[0016] In particular, this specification describes dynamically controlling an activation function of one of the layers of the neural netw ork to modify the amount of sparsity that is introduced into the neural network, thereby modulating how many resources are consumed by the neural network. This controllable activation function can be implemented to increase the sparsity in a given feature map without requiring any retraining or architecture changes to the neural netw ork.
[0017] If multiple neural networks are deployed on the same hardware accelerator, the different neural networks can have different target sparsity levels to reflect the priority assigned to the neural networks, e.g., based on the criticality of the accuracy of each of the neural networks. Thus, when system resources are limited, the accuracy of the outputs produced by critical neural networks can be maintained while decreasing the overall system resource usage of the neural networks deployed on the accelerator.
[0018] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.
[0019] BRIEF DESCRIPTION OF THE DRAWINGS
[0020] FIG. 1 shows an example computing device.
[0021] FIG. 2 illustrates an example process for processing a set of one or more inputs using a neural netw ork.
[0022] FIG. 3 illustrates an example of the operation of the sparsity controller.
[0023] FIG. 4 illustrates another example of the operation of the sparsity' controller.
[0024] FIG. 5 illustrates an example of the operation of the sparsity controller when there are multiple neural networks deployed on the integrated circuit.
[0025] FIG. 6 illustrates an example of the performance of the described techniques relative to a conventional scheme.
[0026] FIG. 7 illustrates another example of the performance of the described techniques relative to a conventional scheme.
[0027] DETAILED DESCRIPTION
[0028] This specification describes techniques for deploying a neural network that has a dynamic activation function layer on a computing device. This specification also describes techniques for controlling the sparsity of outputs generated by the dynamic activation function layer in order to modify the computational load imposed on the computing device by processing inputs using the neural network.
[0029] FIG. 1 shows an example computing device 100.
[0030] The computing device 100 can be any appropriate device that includes an integrated circuit 101 (also referred to as a ‘’hardware accelerator”) that performs neural network computations in hardware.
[0031] For example, the computing device 100 can be a mobile device, e.g., mobile phone, a tablet computer, a wearable device, and so on. As another example, the computing device 100 can be an augmented reality (AR) device, a virtual reality (VR) device, a mixed reality (MR) device, and so on.
[0032] The integrated circuit 101 can be any appropriate integrated circuit that includes hardware for performing neural network computations. Examples of such circuits 101 include tensor processing units (TPUs), graphics processing units (GPUs), and so on.
[0033] As a particular example, the integrated circuit 101 can be a TPU or other hardware accelerator that includes a systolic array circuit for performing multiplication in hardware. For example, the systolic array can be an array of multiply accumulate units (MACs) that performs matrix multiplication in hardware using any of a variety of computation paradigms, e.g., output stationary or input stationary computation.
[0034] As another particular example, the integrated circuit 101 can be a GPU or other hardware accelerator that includes a tensor core circuit for performing multiplication in hardware.
[0035] As part of the operation of the device 100, the device 100 uses a neural network 110 that is deployed on the integrated circuit 101 to generate predictions. For example, the device 100 can generate predictions in response to user requests or in response to requests received from other software programs running on the device 100.
[0036] The neural network 110 can be configured to generate predictions for any appropriate machine learning task.
[0037] For example, the neural network 110 can be configured to perform a computer vision task, e.g., an image classification task, an image modification or editing task, an image generation task, an object detection task, and so on. As another example, the neural network 110 can be configured to perform an audio processing task, e.g., a speech recognition task, an audio classification task, and so on. As yet another example, the neural network 110 can be configured to perform a natural language processing task, e.g., a language modeling task, a natural language understanding task, a text generation task, a computer code generation task, and so on.
[0038] Thus, in response to receiving a request that includes an input to the neural network 110, the integrated circuit 101 processes the input using the neural network 110 by performing at least some of the operations of the neural network 110 in hardware to generate a prediction for the machine learning task. The integrated circuit 101 can then provide the prediction as output, e.g., to another software program running on the device 100. Generally, the operating conditions of the computing device 101 can define the resources that are usable by the integrated circuit 101 for generating predictions using the neural network 110. That is, the operating conditions of the computing device 101 at any given time can define the resource budget, e.g., the power, thermal, and performance budget that can be allocated to the integrated circuit 101.
[0039] For example, when system load is high or the battery level of the computing device 101 is low, the resource budget that is available for use by the integrated circuit 101 can be lower than when system load is high or the computing device 101 is charging.
[0040] As another example, when the ambient temperature of the environment in which the computing device 101 is located is high, the resource budget that is available for use by the integrated circuit 101 can be lower than when the ambient temperature is lower in order to mitigate the risk of overheating.
[0041] To account for this, the computing device 101 uses a sparsity controller 120 that controls the resources consumed by the integrated circuit 101 when performing inference using the neural network 110.
[0042] For example, the sparsity controller 120 can be implemented as one or more software programs on the computing device 101 or can be implemented as one or more software programs on a server that is remote from the computing device 101.
[0043] As yet another example, the sparsity controller 120 can be implemented as specialpurpose hardware on the device 101. e.g., as part of the integrated circuit 110 or off-chip relative to the integrated circuit 1 10.
[0044] In particular, the sparsity controller 120 can dynamically modify the amount of computational resources consumed by performing inference using the neural network 110 in response to changes in the resource budget that is allocated to the neural network 110.
[0045] More specifically, the neural network 110 has an architecture that includes multiple neural network layers, one of which is a dynamic activation function layer 130.
[0046] When performing inference on a batch of one or more network inputs, each neural network layer in the architecture receives a respective layer input that has a first number of elements and performs one or more operations on the respective layer input to generate a respective layer output that has a second number of elements. Depending on the operations performed by the layer, the first number of elements can be equal to or different from the second number of elements. Examples of layers that can be included in the architecture include fully -connected layers, convolutional layers, self-attention layers, cross-attention layers, batch normalization layers, layer normalization layers, residual connection layers, and so on.
[0047] Generally, however, the layers in the neural network 1 10 include an output layer that generates the prediction of, i.e., the output of, the neural network 110 and the dynamic activation function layer 130.
[0048] The dynamic activation function layer 130 applies an elementwise, non-linear activation function to each element in the layer input to the layer 130 to generate a corresponding element of the layer output of the layer 130.
[0049] More specifically, the dynamic activation function layer 130 receives a layer input that has a plurality of input elements and generates a layer output that has a respective output element for each input element by applying, to each input element, the element-wise controllable activation function.
[0050] Additionally, the element-wise controllable activation function has one or more parameters that define a level of sparsity in the layer output of the dynamic activation function layer.
[0051] That is, the activation function has one or more parameters such that different values of the parameter(s) change the fraction of input elements that are expected to be mapped to zero values by applying the activation function.
[0052] As one example, the activation function can be a '“throttleable’7rectified linear unit (ReLU) function that satisfies, for a given input element x: y = max(0, a(x - v)), where y is a given output element y for the given input element x, a is a positive slope, and v is a cut-off parameter.
[0053] In this example, the cut-off parameter v defines the level of sparsity in the layer output of the element-wise controllable activation function because different values of v result in different values of x being mapped to 0, i.e., with each value of x that is less than or equal to v being mapped to 0 by the function.
[0054] The integrated circuit 101 generally includes hardware or is operated by software that allows the integrated circuit 101 to consume fewer resources when processing layer outputs and inputs that have a higher level of sparsity than when processing layer outputs and inputs that have relatively lower levels of sparsity.
[0055] One example of this is hardware or software that implements clock gating that can switch off circuitry that would otherwise only be used to compute a multiplication, addition, or subtraction of a given value with a zero value and instead pass the given value unmodified to the output of the circuitry. When a given layer input has a larger number of zeros, more circuitry can be switched off, reducing the amount of power consumed by performing the operations.
[0056] Another example of this is hardware or software that implements zero-skipping, where computations that operate on zeros are skipped, i.e., are bypassed without being performed, resulting in a shorter overall execution time and fewer clock cycles being necessary to complete the operations, thereby decreasing latency and power consumption.
[0057] Thus, the sparsity controller 120 can control the resources consumed by the integrated circuit 101 by obtaining data specity ing a target level of sparsity for layer outputs of the dynamic activation function layer 130 generated during the processing of the one or more network inputs.
[0058] In some implementations, the controller 120 can determine the target level of sparsity based on the current amount of resources that are available for consumption by the integrated circuit 101. In some other implementations, the controller 120 receives the target level of sparsity, e.g., from another software component executing on the device 100.
[0059] The controller 120 then determines, based on the data specifying the target level of sparsity for the layer outputs of the dynamic activation function layer 130 generated during the processing of the one or more network inputs, a respective new value for each of the one or more parameters of the element-wise controllable activation function.
[0060] That is, the controller 120 determines respective new values for each of the parameters that results in the layer outputs of the layer 130 being expected to have the target level of sparsity .
[0061] The controller 120 then causes the integrated circuit 101 to process the one or more inputs using the neural network 1 10 with the one or more parameters of the layer 130 set to the respective new values. For example, the controller 120 can overwrite the current values of the one or more parameters stored in on-chip or off-chip memory wi th the new values, which causes the circuit 101 to use the new values when processing layer inputs using the layer 130.
[0062] By repeatedly adjusting the values of the one or more parameters in this manner when the amount of available resources changes, the sparsity manager 120 can effectively trade-off between accuracy and resource consumption to maximize the accuracy of outputs generated by the neural network 110 given the current resource budget. In some cases, multiple neural networks, i.e., the neural network 120 and one or more additional neural networks, can be deployed on the integrated circuit 101. That is, the integrated circuit 110 may be able to execute multiple neural networks simultaneously. In these cases, the sparsity controller 130 can set a respective target sparsity value for each of the multiple neural networks. For example, the respective target sparsity values can be determined based in part on the priority assigned to each of the multiple neural networks, i.e., with neural networks that have higher priorities, e.g., due to a stricter quality requirement for the outputs of the neural networks, being assigned lower target sparsity values.
[0063] FIG. 2 illustrates an example process 200 for controlling the sparsity of a neural network. For convenience, the process 200 will be described as being performed by a system of one or more computers, e.g., the device 100 of FIG. 1.
[0064] In particular, the system can perform iterations of the process 200 dynamically without re-training the neural network. That is, after the neural network has been trained and deployed on a particular device, the system can perform the process 200 each time that the sparsity level of the neural network needs to be changed. For example, the system can perform the process 200 in response to receiving an indication that the state of the device 100 has changed such that a different target sparsity level is required.
[0065] The system receives a set of one or more network inputs for processing by the neural network (step 202). As described above, the neural network includes multiple neural network layers, e.g., that are connected in a directed graph in the architecture of the neural network, and the layers include a dynamic activation function layer.
[0066] The system obtains data specifying a target level of sparsity for layer outputs of the dynamic activation function layer generated during the processing of the one or more network inputs (step 204).
[0067] A “level” of sparsity for the layer outputs refers to the portion of the elements in any given layer output that are expected to be zero. The portion can be represented as, e.g., a percentage, a fraction, or a decimal.
[0068] In some implementations, the data specifies an “absolute” level of sparsity, i.e., an absolute portion of the elements.
[0069] In some other implementations, the data specifies a “relative” level of sparsity', i.e., a change in the current target level of sparsity.
[0070] As described above, in some cases the system receives the target level of sparsity. In some other implementations, the system determines the target level of sparsitybased on at least the current resource budget that is allocated to performing inference using the neural network. An example of this is described below with reference to FIG. 4.
[0071] The system determines, based on the data specifying the target level of sparsity for the layer outputs of the dynamic activation function layer generated during the processing of the one or more network inputs, a respective new value for each of the one or more parameters of the element-wise controllable activation function (step 206).
[0072] That is, the system selects a respective new value for each of the one or more parameters of the element- wise controllable activation function in order to adjust the expected sparsity of the layer outputs generated by the dynamic activation function layer.
[0073] In particular, the system maintains data that maps each of a plurality of different target sparsity levels to the corresponding value(s) of the parameter(s) of the activation function. That is, the system maintains data that defines, for each of the different target sparsity levels, the value(s) of the parameter(s) of the activation function that are required for the target sparsity level to be reached.
[0074] The system or another system can have generated this data in any of a variety of ways.
[0075] For example, the system can have processed a set of test inputs through the layers of the neural network and can have identified the distribution of values of the elements of the layer inputs to the dynamic activation function layer. Based on this distribution, the system or another system can have generated a mapping that identifies what percentage of elements in a given layer input will be set to zero in accordance with different value(s) of the parameter(s) of the dynamic activation function layer.
[0076] As another example, the layer that precedes the activation function layer can be a normalization layer, e.g., batch normalization or layer normalization, that normalizes the inputs to the layer and then applies a learned shift and scale parameters to the normalized inputs. In this example, the system can apply a statistical analysis using the normalization statistics used for normalization and the values of the shift and scale parameters to generate a mapping that identifies what percentage of elements in a given layer input will be set to zero in accordance with different value(s) of the parameter(s) of the dynamic activation function layer.
[0077] The system processes the set of one or more network inputs using the neural network with the one or more parameters of the element-wise controllable activation function set to the respective new values to generate a respective network output for each network input in the set (step 208).
[0078] By repeatedly performing the process 200 to adapt the parameter(s) of the dynamic activation function layer, the system can modify the sparsity levels of the layer outputs of the activation function layer, thereby modifying how many resources are consumed by performing inference using the neural network.
[0079] FIG. 3 shows an example 320 of the operation of the sparsity controller 120.
[0080] In the example of FIG 3, the activation function is a throttleable ReLU with slope a and cutoff value v.
[0081] FIG. 3 also shows an example 310 of the operation of a conventional ReLu layer with no slope parameter and no cutoff value v (or, equivalently, with slope a equal to one and v equal to zero).
[0082] Conventionally, the operation of the ReLu layer is fixed and all input elements x that are less than or equal to zero are mapped to zero.
[0083] When the sparsity controller 120 is used, however, the controller 120 can dynamically control the sparsity of the output elements by changing the cutoff value v.
[0084] For example, as shown in FIG. 3, the sparsity controller 120 receives an input specifying the current performance budget for the neural network, e.g., based on operating conditions of system, e.g., a system on a chip (SOC), on which the neural network is deployed.
[0085] The sparsity controller 120 determines, from the current performance budget, a new value for the cut-off threshold value v, e.g., that results in a sparsity level that causes the neural network to operate within the current performance budget.
[0086] In some implementations, to maintain the accuracy of the outputs of the neural network, the sparsity controller 120 can also update the slope value a.
[0087] For example, the system can identify a particular input element value that is not assigned to zero by the new' cut-off threshold value v and then identify' an output element value that is generated by the element-wise controllable activation function for the particular input element value in accordance with current values for the positive slope and the cut-off threshold parameter.
[0088] The system can then determine a new' value for the positive slope value a such that the element-wise controllable activation function generates an output element value for the particular input element value in accordance with the new values for the positive slope and the cut-off threshold parameter. Modifying the slope value a in this manner minimizes the impact that changing the cut-off threshold value v will have on the operation of downstream layers in the neural network, mitigating any decreases in accuracy that could result.
[0089] Alternatively, the controller 120 can maintain, for each possible value of v, the corresponding value of a that needs to be set to mitigate the accuracy decrease, and can use the mapping to determine the new value of a that corresponds to v. For example, the controller 120 or another system can have pre-computed the corresponding values of a using the above technique.
[0090] FIG. 4 is another example 400 of the operation of the sparsity controller 120. In particular, the example of FIG. 4 shows how the sparsity controller 120 can use estimates of the accuracy of the neural network for various sparsity targets to determine how to modify the sparsity target.
[0091] As shown in FIG. 4, the sparsity controller 120 receives the current resource budget for the neural network given the current system conditions.
[0092] The sparsity controller 120 then uses the current resource budget to determine the target sparsity level for the layer outputs of the dynamic activation function layers.
[0093] For example, the sparsity controller 120 can maintain data mapping different target sparsity levels to different amounts of resources consumed for computing inference. For example, the sparsity controller 120 or another system can have generated this data byprocessing multiple sets of one or more test inputs using the neural network when deployed on the integrated circuit 101 (or in a simulation of the integrated circuit 101 ) and measured for each set of inputs (i) the average level of sparsity of the network inputs in the set and (ii) the amount of resources consumed by processing the test inputs. The sparsity controller 120 or the other system can have then generated the data by determining, for each level of sparsity observed for any of the sets of test inputs, how many resources were consumed, on average, when processing inputs that had, on average, that level of sparsity.
[0094] In some implementations, the sparsity controller 120 then sets the target sparsity level to the lowest sparsity level that will result in an amount of resources being consumed that is within the current budget.
[0095] In some other implementations, however, the sparsity controller 120 also considers the accuracies of the neural network for different sparsity targets when setting the target sparsity level.
[0096] As one example, the controller 120 can determine an estimate (or a "‘proxy”) of the current accuracy of the neural network given the current sparsity level based on the softmax output values, i.e., probabilities, assigned by the neural network to the inputs that have been processed with the current sparsity level.
[0097] As one example, the controller 120 can generate this estimate by determining either (i) the overall average softmax output value assigned to the highest class in the probability distribution for the inputs processed with the current sparsity level or (ii) for each class, the individual average softmax output value assigned to the class when the class is the highest class in the probability distribution for a given input processed with the current sparsity level. For (i), the system can then set the estimate to the overall average softmax output and for (ii), the system can, e.g., set the estimate to the lowest individual average softmax output for any class.
[0098] That is, the controller 120 can use the '■confidence” of the neural network in its classification outputs as a proxy for the accuracy of the neural network, e.g., based on the fact that the neural network is more likely to be accurate when the softmax outputs indicate a higher degree of confidence.
[0099] In this example, if the current accuracy estimate is below a minimum accuracythreshold for the neural network, the controller 120 will not increase the current sparsity target (even if the current resource budget is mapped to a higher sparsity level in the data maintained by the controller 120).
[0100] As another example, the controller 120 can maintain respective accuracy estimates (computed as above) for multiple different sparsity levels (in addition to the current sparsity level). In this example, the controller 120 can set the target sparsity level to the highest sparsity- level that is (i) equal to or less than the sparsity level that the current resource budget is mapped to in the data maintained by the controller 120 and (ii) that is mapped to an accuracy estimate that is not below the accuracy threshold for the neural network.
[0101] Thus, in both of these examples, the system uses the accuracy estimates to ensure that accuracy maintains at an acceptable level throughout dynamically adjusting the sparsity of the neural netw ork.
[0102] FIG. 5 shows an example 500 of the operation of the sparsity controller 120 when there are multiple neural networks deployed on the integrated circuit 101.
[0103] As shown in FIG. 5, there are four neural networks deployed on the integrated circuit 101 and, given the current resource budget for the integrated circuit 101, the sparsity controller 120 determines respective sparsity targets for each of the four neural networks.
[0104] As described above, the sparsity targets can be different for each of the four neural networks. For example, as described above, the sparsity controller 120 can maintain respective accuracy estimates for each of the four neural networks. The sparsity controller 120 can also maintain respective minimum accuracy thresholds.
[0105] The sparsity controller 120 can then set the respective sparsity targets for each of the four neural networks so that the resource budget is not exceeded but subject to the constraint that none of the neural network falls below the respective minimum accuracy thresholds. Thus, neural networks that have higher minimum accuracy thresholds, e.g., because their predictions are more critical to the operation of the device 100, or that have accuracies that are more sensitive to sparsity changes can have lower sparsity targets than other neural networks.
[0106] FIG. 6 illustrates an example 600 of the performance of the described techniques relative to a conventional scheme.
[0107] In particular, FIG. 6 shows the layer output 610 of a given, non-dynamic activation function layer of an example convolutional neural network in a conventional scheme and the layer output 620 of a dynamic activation function layer that replaces the given nondynamic activation function layer within the the same neural network. That is. the layer output 610 and 620 are the output of the two layers for the same particular image and with the remainder of the neural network otherwise unmodified and without any additional training. More specifically, for ease of illustration, FIG. 6 shows the first ten channels of the output of the two layers for the particular image. For example, channel 612 is the third channel of the output of the non-dynamic layer and channel 622 is the third channel of the output of the dynamic layer.
[0108] As can be seen from FIG. 6, the layer outputs 620 of the dynamic activation function layer have significantly more zeros than the layer outputs 610 of the non-dynamic activation function. As described above, this leads to reduced resource usage while processing the particular image using the convolutional neural network. However, as can be seen from FIG. 6, the key features are retained across the channels. That is, each channel of the layer output 620 still retains the same key features as the corresponding channel in the layer output 610. Thus, the system still generates outputs of comparable quality’ even with the increase in sparsity.
[0109] FIG. 7 illustrates another example 700 of the performance of the described techniques relative to a conventional scheme.
[0110] In particular, FIG. 7 shows a histogram 710 of values w ithin the layer outputs (“activations”) of the given, non-dynamic activation function layer of the example convolutional neural network in the conventional scheme and a histogram 720 of values of layer outputs of a dynamic activation function layer that replaces the given non-dynamic activation function layer within the the same neural network.
[0111] As can be seen from FIG. 7, over 300,000 of the values in the layer outputs of the dynamic layer are zero, while less than 250,000 of the values in the layer outputs of the conventional outputs are zero. Moreover, apart from the zero values, the two histograms have the same shape, showing that the key features are retained in the layer outputs of the dynamic activation function layer and only low-impact, close to zero activations are set equal to zero. Thus, the system still generates outputs of comparable quality even with the increase in sparsity.
[0112] This specification uses the term "configured" in connection with systems and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions. Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry', in tangibly- embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory' device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus.
[0113] The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including byway of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0114] A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a stand alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.
[0115] In this specification, the term “database’' is used broadly to refer to any collection of data: the data does not need to be structured in any particular way, or structured at all, and it can be stored on storage devices in one or more locations. Thus, for example, the index database can include multiple collections of data, each of which may be organized and accessed differently.
[0116] Similarly, in this specification the term “engine” is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers.
[0117] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.
[0118] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read only memoiy or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memoiy' can be supplemented by, or incorporated in, special purpose logic circuitry'. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.
[0119] Computer readable media suitable for storing computer program instructions and data include all forms of non volatile memory', media and memory' devices, including by way of example semiconductor memoiy devices, e.g., EPROM, EEPROM, and flash memory’ devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM and DVD-ROM disks.
[0120] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid cry stal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory’ feedback, e.g.. visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user's device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return.
[0121] Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and compute-intensive parts of machine learning training or production, i.e., inference, workloads.
[0122] Machine learning models can be implemented and deployed using a machine learning framework, e g., a TensorFlow framework or a Jax framework.
[0123] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g.. an application server, or that includes a front end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.
[0124] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device.
[0125] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0126] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0127] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
[0128] What is claimed is:
Claims
CLAIMS1. A method comprising: receiving a set of one or more network inputs for processing by a neural network, wherein: the neural network comprises a plurality of neural network layers that comprise a dynamic activation function layer, the dynamic activation function layer receives a layer input comprising a plurality of input elements and generates a layer output comprising a respective output element for each input element by applying, to each input element, an element-wise controllable activation function, and the element- wise controllable activation function has one or more parameters that define a level of sparsity in the layer output of the controllable activation function layer; obtaining data specifying a target level of sparsity for layer outputs of the dynamic activation function layer generated during the processing of the one or more network inputs; determining, based on the data specifying the target level of sparsity for the layer outputs of the controllable activation function layer generated during the processing of the one or more network inputs, a respective new value for each of the one or more parameters of the element-wise controllable activation function; and processing the set of one or more network inputs using the neural network with the one or more parameters of the element-wise controllable activation function set to the respective new values to generate a respective network output for each network input in the set.
2. The method of claim 1, wherein the element-wise controllable activation function is a modified Rectified Linear Unit (ReLU) and satisfies, for a given input element x: y = max(0, a(x - v)). where y is a given output element y for the given input element x, a is a positive slope, and v is a cut-off threshold parameter that defines the level of sparsity in the layer output of the element-wise controllable activation function, and wherein determining a respective new value for each of the one or more parameters of the element-wise controllable activation function comprises determining a new value for the cut-off threshold parameter.
3. The method of claim 2, further comprising: identifying a particular input element value; identifying an output element value that is generated by the element-wise controllable activation function for the particular input element value in accordance with current values for the positive slope and the cut-off threshold parameter; determining a new value for the positive slope value such that the element-wise controllable activation function generates an output element value for the particular input element value in accordance with the new values for the positive slope and the cut-off threshold parameter; wherein processing the set of one or more network inputs using the neural network with the one or more parameters of the element-wise controllable activation function set to the respective new values to generate a respective network output for each network input in the set comprises: processing the set of one or more network inputs using the neural network with the positive slope and the cut-off threshold parameter set to the respective new values to generate a respective network output for each network input in the set.
4. The method of any preceding claim, wherein determining, based on the data specifying the target level of sparsity for the layer outputs of the dynamic activation function layer generated during the processing of the one or more network inputs, a respective new value for each of the one or more parameters of the element-wise controllable activation function comprises: maintaining data specifying an expected distribution of values of input elements; and determining values for each of the one or more parameters that, given the expected distribution of values of input elements, result in an expected distribution of values of output elements that have the target level of sparsity.
5. The method of claim 4, wherein the data specifying an expected distribution of values indicates, for each of a plurality of intervals, an expected proportion of input elements that have a value within the interval.
6. The method of any preceding claim, wherein obtaining data specifying a target level of sparsity for layer outputs of the dynamic activation function layer generated during the processing of the one or more network inputs comprises: maintaining resource data mapping each of a plurality of sparsity levels to a respective amount of computational resources consumed by the neural network when processing inputs with the sparsity level; receiving data specifying a target amount of computational resources to be consumed by the neural network when processing the set of one or more network inputs; and selecting the target level of sparsity based on the target amount of computational resources and the resource data.
7. The method of claim 6, wherein selecting the target level of sparsity based on the target amount of computational resources and the resource data comprises: setting the target level of sparsity to a sparsity level that is mapped to by the target level of sparsity in the resource data.
8. The method of claim 6, wherein selecting the target level of sparsity based on the target amount of computational resources and the resource data comprises: maintaining accuracy data mapping each of a plurality of sparsity levels to a respective accuracy of the neural network when processing inputs with the sparsity level; obtaining data specifying a minimum acceptable accuracy for the processing of the one or more network inputs; and when a sparsity level that is mapped to by the target level of sparsity in the resource data is mapped to a respective accuracy that is greater than or equal to the minimum acceptable accuracy, setting the target level of sparsity' to a sparsity level that is mapped to by the target level of sparsity in the resource data.
9. The method of any preceding claim, wherein the neural network is deployed on a hardware accelerator, and wherein processing the set of one or more network inputs using the neural network with the one or more parameters of the element-wise controllable activation function set to the respective new values to generate a respective network output for each network input in the set comprises: providing data specifying the new values for the one or more parameters to the hardware accelerator.
10. The method of claim 9, wherein one or more other neural networks are also deployed on the hardware accelerator, wherein the neural network and the one or more other neural networks are each associated with a respective priority, and wherein obtaining data specifying a target level of sparsity for layer outputs of the dynamic activation function layer generated during the processing of the one or more network inputs comprises: determining the target level of sparsity based on the respective priority associated with the neural network.
11. A system comprising one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform the operations of the respective method of any one of claims 1-10.
12. A computer storage medium encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the respective method of any one of claims 1-10.
Citation Information
Patent Citations
Systems and methods for accelerating sparse neural network execution
US20210004665A1
Dynamic activation sparsity in neural networks
US20220383121A1
Acceleration method and accelerator used for convolutional neural network
WO2019196223A1