Artificial Intelligence Processor Architecture for Dynamic Scaling of Neural Network Quantization
Dynamic AI quantization adjustment addresses thermal management issues in neural network processing by balancing throughput and accuracy, improving performance in computing systems.
Patent Information
- Application Number
- JP2023557775
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-03-24
- Filing Date
- 2022-02-25
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2042-02-25
AI Technical Summary
Modern computing systems face thermal management challenges in neural network processing due to heavy workloads, leading to reduced processor operating frequency, which affects performance and can cause critical issues in mission-critical systems.
Dynamically adjusting AI quantization levels for neural networks based on operating conditions such as temperature, power consumption, and utilization to manage throughput and accuracy, using a dynamic quantization controller and MAC array.
Enhances neural network processing performance by balancing throughput and accuracy under adverse conditions, reducing heat buildup and power constraints without sacrificing user experience or safety.
Smart Images

Figure 0007810717000001 
Figure 0007810717000002 
Figure 0007810717000003
Abstract
Description
[Technical Field]
[0001] Related Applications This application claims the benefit of priority to U.S. Patent Application No. 17 / 210,644, filed March 24, 2021, the entire contents of which are incorporated herein by reference. [Background technology]
[0002] Modern computing systems run multiple neural networks on a system-on-chip (SoC), leading to heavy neural network loads for the SoC's processor. Despite optimization of processor architectures for running neural networks, heat remains a limiting factor for neural network processing under heavy workloads, as thermal management is implemented by reducing the processor's operating frequency, which impacts processing performance. In mission-critical systems, reducing the operating frequency can cause critical issues that can result in degradation of user experience, product quality, operational safety, etc. Summary of the Invention [Means for solving the problem]
[0003] Various disclosed aspects may include apparatus and methods for processing a neural network by an artificial intelligence (AI) processor. Various aspects may include receiving AI processor operating condition information, dynamically adjusting AI quantization levels for a segment of the neural network in response to the operating condition information, and processing the segment of the neural network using the adjusted AI quantization levels.
[0004] In some aspects, dynamically adjusting the AI quantization level for a segment of the neural network may include increasing the AI quantization level in response to operating condition information indicating a level of operating conditions that increased the processing power constraints of the AI processor, and decreasing the AI quantization level in response to operating condition information indicating a level of operating conditions that decreased the processing power constraints of the AI processor.
[0005] In some aspects, the operating condition information may be at least one of the following group: temperature, power consumption, operating frequency, or utilization of the processing unit.
[0006] In some aspects, dynamically adjusting an AI quantization level for a segment of a neural network may include adjusting an AI quantization level for quantizing weight values to be processed by the segment of the neural network.
[0007] In some aspects, dynamically adjusting an AI quantization level for a segment of a neural network may include adjusting an AI quantization level for quantizing activation values to be processed by the segment of the neural network.
[0008] In some aspects, dynamically adjusting an AI quantization level for a segment of a neural network may include adjusting an AI quantization level for quantizing weight values and activation values to be processed by the segment of the neural network.
[0009] In some aspects, the AI quantization level may be configured to indicate a dynamic bit of a value to be processed by the neural network that is to be quantized, and processing a segment of the neural network using the adjusted AI quantization level may include bypassing a portion of a multiply-accumulate unit (MAC) associated with the dynamic bit of the value.
[0010] Some embodiments may further include determining an AI Quality of Service (QoS) value using an AI QoS factor and determining an AI quantization level to achieve the AI QoS value. In some embodiments, the AI QoS value may represent a goal for the accuracy of results generated by the AI processor and the throughput (e.g., inferences per second) of the AI processor.
[0011] Further embodiments may include an AI processor including a dynamic quantization controller and a MAC array configured to perform the operations of any of the methods summarized above. Further embodiments may include a computing device having an AI processor including a dynamic quantization controller and a MAC array configured to perform the operations of any of the methods summarized above. Further embodiments may include an AI processor including means for performing the functions of any of the methods summarized above.
[0012] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate exemplary embodiments of various embodiments and, together with the general description above and the detailed description below, serve to explain the features of the claims. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 1 is a component block diagram illustrating an exemplary computing device suitable for implementing various embodiments. [Figure 2A] FIG. 1 is a component block diagram illustrating an example artificial intelligence (AI) processor with a dynamic neural network quantization architecture suitable for implementing various embodiments. [Figure 2B] FIG. 1 is a component block diagram illustrating an example AI processor with a dynamic neural network quantization architecture suitable for implementing various embodiments. [Figure 3]FIG. 1 is a component block diagram illustrating an exemplary system-on-chip (SoC) having a dynamic neural network quantization architecture suitable for implementing various embodiments. [Figure 4A] FIG. 1 is a graph illustrating exemplary AI quality of service (QoS) relationships suitable for implementing various embodiments. [Figure 4B] FIG. 1 is a graph illustrating an example AI QoS relationship suitable for implementing various embodiments. [Figure 5] FIG. 10 is a graph illustrating an exemplary benefit in AI processor operating frequency from implementing a dynamic neural network quantization architecture in accordance with various embodiments. [Figure 6] FIG. 10 is a graphical comparison diagram illustrating exemplary benefits in AI processor operating frequency from implementing a dynamic neural network quantization architecture, according to various embodiments. [Figure 7] FIG. 1 is a component diagram illustrating an example of a bypass in a multiply-accumulate (MAC) unit in a dynamic neural network quantization architecture suitable for implementing various embodiments. [Figure 8] FIG. 1 is a process flow diagram illustrating a method for AI QoS determination, according to an embodiment. [Figure 9] FIG. 1 is a process flow diagram illustrating a method for dynamic neural network quantization architecture configuration control, according to an embodiment. [Figure 10] FIG. 1 is a process flow diagram illustrating a method for dynamic neural network quantization architecture reconfiguration, according to an embodiment. [Figure 11] FIG. 1 is a component block diagram illustrating an exemplary mobile computing device suitable for implementing an AI processor, according to various embodiments. [Figure 12] FIG. 1 is a component block diagram illustrating an exemplary mobile computing device suitable for implementing an AI processor, according to various embodiments. [Figure 13]FIG. 1 is a component block diagram illustrating an exemplary server suitable for implementing an AI processor, according to various embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0014] Various embodiments will now be described in detail with reference to the accompanying drawings. Wherever possible, the same reference numbers will be used throughout the drawings to refer to the same or like parts. References made to specific examples and implementations are for illustrative purposes only and are not intended to limit the scope of the claims.
[0015] Various embodiments may include methods for dynamically configuring a neural network quantization architecture and computing devices implementing such methods. Some embodiments may include dynamic neural network quantization logic hardware configured to modify quantization, masking, and / or neural network pruning based on operating conditions of an artificial intelligence (AI) processor, a system-on-chip (SoC) having the AI processor, memory accessed by the AI processor, and / or other peripherals of the AI processor. Some embodiments may include configuring the dynamic neural network quantization logic for quantization of activation values and weight values based on a number of dynamic bits for dynamic quantization. Some embodiments may include configuring the dynamic neural network quantization logic for masking of activation values and weight values and bypassing a portion of a MAC of a multiply-accumulate (MAC) array based on a number of dynamic bits for bypass. Some embodiments may include configuring the dynamic neural network quantization logic for masking of weight values and bypassing the entire MAC based on a threshold weight value for neural network pruning. Some embodiments may include determining whether to configure dynamic neural network quantization logic and using an AI quality of service (QoS) value incorporating accuracy of AI processor results and responsiveness of the AI processor to implement the configuration of the dynamic neural network quantization logic.
[0016] The term "dynamic bits" is used herein to refer to bits of activation and / or weight values for configuring dynamic neural network quantization logic for quantization of activation and weight values and / or for masking activation and weight values and bypassing portions of MAC. In some embodiments, the dynamic bits may be any number of low-order bits of the activation and / or weight values.
[0017] The term "AI quantization level" is described herein in relative terms, such that multiple AI quantization levels are described relative to one another. For example, a higher AI quantization level may relate to increased quantization, in which a greater number of dynamic bits for activation values and / or weight values are masked (zeroed) than a lower AI quantization level. A lower AI quantization level may relate to reduced quantization, in which a smaller number of dynamic bits for activation values and / or weight values are masked (zeroed) than a higher AI quantization level.
[0018] The terms "computing device" and "mobile computing device" are used interchangeably herein to refer to any one or all of mobile phones, smartphones, personal or mobile multimedia players, personal digital assistants (PDAs), laptop computers, tablet computers, convertible laptops / tablets (2-in-1 computers), smartbooks, ultrabooks, netbooks, palmtop computers, wireless email receivers, multimedia Internet-enabled mobile phones, mobile game consoles, wireless game controllers, and similar personal electronic devices that include memory and a programmable processor. The term "computing device" may also refer to stationary computing devices, including personal computers, desktop computers, all-in-one computers, workstations, supercomputers, mainframe computers, embedded computers (such as in vehicles and other large-scale systems), computerized vehicles (e.g., passenger vehicles, commercial vehicles, recreational vehicles, military vehicles, partially or fully autonomous land, air, and / or water vehicles such as drones), servers, multimedia computers, and game consoles.
[0019] Neural networks are implemented in an array of computing devices capable of simultaneously executing multiple neural networks. AI processors are implemented using architectures specifically designed for neural network execution, such as in neural processing units, and / or are advantageous for neural network execution, such as in digital signal processing units. AI processor architectures can offer higher processing performance in terms of latency, accuracy, power consumption, etc., compared to other processor architectures, such as central processing units and graphics processing units. However, AI processors typically have high power density, and under heavy workloads, often resulting from simultaneously executing multiple neural networks, AI processors can suffer performance degradation caused by heat buildup. An example of such an AI processor running multiple neural networks is in an automobile with an active driver assistance system, where the AI processor simultaneously runs one set of neural networks for vehicle navigation / operation and another set of neural networks for driver monitoring. Current strategies for thermal management in AI processors include reducing the operating frequency of the AI processor based on sensed temperature.
[0020] Reducing the operating frequency of an AI processor in a mission-critical system can cause fatal problems that can result in degradation of user experience, product quality, operational safety, and the like. The throughput of an AI processor is an important factor in the performance of an AI processor that is more adversely affected by reducing the operating frequency. Another important factor in the performance of an AI processor is the accuracy of the AI processor's results. This accuracy may not be affected by reducing the operating frequency because the operating frequency may affect how quickly the AI processor performs its operations, such as completing data processing using all of the data provided, rather than whether the operations are fully performed. Thus, by reducing the operating frequency in response to heat accumulation, the throughput of the AI processor may be sacrificed, but the accuracy of the AI processor's results may not be sacrificed. In some systems, such as autonomous vehicles, drones, and other self-propelled machines, throughput is critical, and as a result, sacrificing some accuracy for faster throughput is acceptable and even desirable.
[0021] Similar problems arise when the operating frequency is reduced in response to other adverse operating conditions, such as power constraints on the AI processor's power supply and / or performance constraints on the computing device having the AI processor. For clarity and simplicity of explanation, examples herein are described with respect to heat buildup, but such reference is not intended to limit the scope of the claims and the description herein.
[0022] Furthermore, the quantization applied to neural network inputs, including activation values and weight values, is static in conventional systems: neural network developers pre-configure the quantized features of a neural network in a compiler or development tool, setting the quantization for the neural network to a fixed number of significant bits.
[0023] In some embodiments described herein, dynamically configuring the neural network quantization architecture may be configured to manage the throughput of the AI processor and the accuracy of the AI processor's results under adverse operating conditions, such as heat buildup. While the accuracy of the AI processor's results is an important factor in the performance of an AI processor, a certain degradation may be acceptable in many situations. The accuracy of the AI processor's results may be affected by modifying the inputs, activation values, and weight values to the neural network running on the AI processor. By sacrificing some accuracy of the AI processor, the throughput of the AI processor may be less affected in response to heat buildup than when responding to heat buildup by only reducing the throughput of the AI processor. In some embodiments, sacrificing some accuracy of the AI processor and some throughput of the AI processor may achieve greater power and / or main memory traffic reduction than when responding to heat buildup by only reducing the throughput of the AI processor.
[0024] In some embodiments, the dynamic neural network quantization logic may be configured at runtime to modify quantization, masking, and / or neural network pruning based on operating conditions, such as temperature, power consumption, and processing unit utilization, of the AI processor, the SoC having the AI processor, memory accessed by the AI processor, and / or other peripheral devices of the AI processor. Some embodiments may include configuring the dynamic neural network quantization logic for quantization of activation values and weight values based on a number of dynamic bits for dynamic quantization. Some embodiments may include configuring the dynamic neural network quantization logic for masking of activation values and weight values and bypassing a portion of the MAC based on a number of dynamic bits for bypass. Some embodiments may include configuring the dynamic neural network quantization logic for masking of weight values and bypassing the entire MAC based on a threshold weight value for neural network pruning. In some embodiments, the dynamic neural network quantization logic may be configured to modify the pre-configured quantization of the neural network as needed based on operating conditions.
[0025] Some embodiments may include a dynamic quantization controller configured to generate a dynamic quantization signal and send it to any number of AI processors, dynamic neural network quantization logic, and MACs, and combinations thereof. The dynamic quantization controller may determine parameters for implementing quantization, masking, and / or neural network pruning by the AI processors, dynamic neural network quantization logic, and MACs. The dynamic quantization controller may determine these parameters based on an AI quantization level that incorporates the accuracy of the AI processor's results and the responsiveness of the AI processor.
[0026] Some embodiments may include an AI QoS manager configured to determine whether to perform dynamic neural network quantization reconfiguration of the AI processor, the dynamic neural network quantization logic, and / or the MAC. The AI QoS manager may receive a data signal representing an AI QoS factor. The AI QoS factor may be an operating condition, and the dynamic neural network quantization logic reconfiguration may be based on the operating condition to modify quantization, masking, and / or neural network pruning. These operating conditions may include the temperature, power consumption, processing unit utilization, etc., of the AI processor, the SoC having the AI processor, memory accessed by the AI processor, and / or other peripheral devices of the AI processor. The AI QoS manager may determine an AI QoS value that considers the throughput of the AI processor, the accuracy of the AI processor results, and / or the operating frequency of the AI processor to be achieved by the AI processor under certain operating conditions. The AI QoS value may be used to determine an AI quantization level that considers the throughput of the AI processor and the accuracy of the AI processor results as a result of configuring the dynamic neural network quantization logic, and / or the AI processor operating frequency for the operating conditions.
[0027] 1 illustrates a system including a computing device 100 suitable for use with various embodiments. Computing device 100 may include a SoC 102 with a processor 104, memory 106, a communications interface 108, a memory interface 110, and a peripheral device interface 120. Computing device 100 may further include a communications component 112, such as a wired or wireless modem, memory 114, an antenna 116 for establishing a wireless communications link, and / or peripheral devices 122. Processor 104 may include any of a variety of processing devices, e.g., several processor cores.
[0028] The term "system on a chip" or "SoC" is used herein to generally refer to a set of interconnected electronic circuits including, but not limited to, a processing device, memory, and a communication interface. A processing device may include a variety of different types of processors 104 and / or processor cores, such as a general-purpose processor, a central processing unit (CPU), a digital signal processor (DSP), a graphics processing unit (GPU), an accelerated processing unit (APU), a secure processing unit (SPU), a subsystem processor for a particular component of a computing device, such as an image processor for a camera subsystem or a display processor for a display, an auxiliary processor, a single-core processor, a multi-core processor, a controller, and / or a microcontroller. A processing device may further embody other hardware and combinations of hardware, such as a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), other programmable logic devices, discrete gate logic, transistor logic, performance monitoring hardware, watchdog hardware, and a time reference. An integrated circuit may be configured such that the components of the integrated circuit reside on a single semiconductor material, such as silicon.
[0029] The memory 106 of the SoC 102 may be volatile or nonvolatile memory configured to store data and processor-executable code for access by the processor 104 or by other components of the SoC 102, including the AI processor 124. The computing device 100 and / or the SoC 102 may include one or more memories 106 configured for various purposes. The one or more memories 106 may include volatile memory, such as random access memory (RAM) or main memory, or cache memory. These memories 106 may be configured to temporarily hold limited amounts of data received from data sensors or subsystems, data and / or processor-executable code instructions requested from and loaded into the memory 106 from nonvolatile memory, and / or intermediate processing data and / or processor-executable code instructions generated by the processor 104 and / or the AI processor 124 that are not stored in nonvolatile memory but are temporarily stored for quick future access. Memory 106 may be configured to at least temporarily store data and processor-executable code that is loaded into memory 106 from another memory device, such as another memory 106 or memory 114, for access by one or more of processors 104 or other components of SoC 102, including AI processor 124. In some embodiments, any number of memories 106 and combinations thereof may include one-time programmable memory or read-only memory.
[0030] Memory interface 110 and memory 114 may cooperate to enable computing device 100 to store and retrieve data and processor-executable code to and from volatile and / or non-volatile storage media. Memory 114 may be configured much like embodiments of memory 106, and memory 114 may store data or processor-executable code for access by one or more of processors 104 or other components of SoC 102, including AI processor 124. Memory interface 110 may control access to memory 114, enabling processor 104 or other components of SoC 12, including AI processor 124, to read data from and write data to memory 114.
[0031] The SoC 102 may also include an AI processor 124. The AI processor 124 may be the processor 104, a portion of the processor 104, and / or a standalone component of the SoC 102. The AI processor 124 may be configured to execute a neural network for processing activation values and weight values on the computing device 100. The computing device 100 may also include an AI processor 124 not associated with the SoC 102. Such an AI processor 124 may be a standalone component of the computing device 100 and / or integrated into another SoC 102.
[0032] Some or all of the components of computing device 100 and / or SoC 102 may be arranged and / or combined differently while still performing the functionality of the various embodiments. Computing device 100 may not be limited to one of each of the components, and multiple instances of each component may be included in various configurations of computing device 100.
[0033] 2A illustrates an exemplary AI processor having a dynamic neural network quantization architecture suitable for implementing various embodiments. Referring to FIGS. 1 and 2A, the AI processor 124 may include any number of MAC arrays 200, weight buffers 204, activation buffers 206, dynamic quantization controllers 208, AI QoS managers 210, and dynamic neural network quantization logic 212, 214, as well as combinations thereof. The MAC array 200 may include any number of MACs 202a-202i and combinations thereof.
[0034] The AI processor 124 may be configured to execute a neural network. The executed neural network may process activation values and weight values. The AI processor 124 may receive and store activation values in activation buffer 206 and weight values in weight buffer 204. In general, the MAC array 200 may receive activation values from activation buffer 206 and weight values from weight buffer 204 and process the activation values and weight values by multiplying and accumulating the activation values and weight values. For example, each MAC 202a-202i may receive any number of activation and weight value combinations, multiply the bits of each received activation and weight value combination, and accumulate the results of the multiplications. A transform (CVT) module (not shown) of the AI processor 124 may modify the MAC results by performing a function that uses the MAC results, such as scaling, adding a bias, and / or applying an activation function (e.g., sigmoid, ReLU, Gaussian, SoftMax, etc.). The MACs 202a-202i may receive multiple combinations of activation values and weight values by receiving each combination in turn. As described further herein, in some embodiments, the activation values and weight values may be modified before being received by the MACs 202a-202i, and as described further herein, in some embodiments, the MACs 202a-202i may be modified to process the activation values and weight values.
[0035] AI QoS manager 210 may be configured as hardware, software executed by AI processor 124, and / or a combination of hardware and software executed by processor 124. AI QoS manager 210 may be configured to determine whether to perform dynamic neural network quantization reconfiguration of AI processor 124, dynamic neural network quantization logic 212, 214, and / or MACs 202a-202i. AI QoS manager 210 may be communicatively coupled to any number and combination of sensors (not shown), such as temperature sensors, voltage sensors, current sensors, etc., as well as processor 104. AI QoS manager 210 may receive data signals representing AI QoS factors from these communicatively coupled sensors and / or processor 104. The AI QoS factors may be operating conditions, and the decision to reconfigure the dynamic neural network quantization logic may be based on the operating conditions to modify quantization, masking, and / or neural network pruning. These operating conditions may include temperature, power consumption, processing unit utilization, performance, etc., of the AI processor 124, the SoC 102 having the AI processor 124, memories 106, 114 accessed by the AI processor 124, and / or other peripheral devices 122 of the AI processor 124. For example, a temperature operating condition may be a temperature sensor value representing the temperature at a location on the AI processor 124. In a further example, a power operating condition may be a value representing the peak of a power supply rail compared to the power supply, and / or the capabilities of a power management integrated circuit, and / or the state of charge of a battery. As a further example, a performance operating condition may be a value representing utilization, time completely idle, frames per second, and / or end-to-end latency of the AI processor 124.
[0036] The AI QoS manager 210 may be configured to determine whether to perform dynamic neural network quantization reconfiguration based on operating conditions. The AI QoS manager 210 may determine to perform dynamic neural network quantization reconfiguration based on a level of operating conditions that increase the constraints on the processing power of the AI processor 124. The AI QoS manager 210 may determine to perform dynamic neural network quantization reconfiguration based on a level of operating conditions that decrease the constraints on the processing power of the AI processor 124. The constraints on the processing power of the AI processor 124 may be caused by the level of operating conditions, such as the level of heat buildup, power consumption, and processing unit utilization, which affect the ability of the AI processor 124 to maintain a level of processing power.
[0037] In some embodiments, AI QoS manager 210 may be configured with any number of algorithms, thresholds, lookup tables, and the like, and combinations thereof, for determining whether to perform dynamic neural network quantization reconfiguration from the operating conditions. For example, AI QoS manager 210 may compare the received operating conditions to the operating condition thresholds. In response to an unfavorable result of the comparison of the operating conditions to the operating condition thresholds, such as exceeding the thresholds, AI QoS manager 210 may decide to perform dynamic neural network quantization reconfiguration. Such an unfavorable comparison may indicate to AI QoS manager 210 that the operating conditions have increased the processing power constraints of AI processor 124. In response to a favorable result of the comparison of the operating conditions to the operating condition thresholds, such as falling below the thresholds, AI QoS manager 210 may decide to perform dynamic neural network quantization reconfiguration. Such a favorable comparison may indicate to AI QoS manager 210 that the operating conditions have reduced the processing power constraints of AI processor 124. In some embodiments, AI QoS manager 210 may be configured to compare the plurality of received operating conditions to a plurality of threshold operating condition values and determine to perform dynamic neural network quantization reconfiguration based on a combination of unfavorable and / or favorable comparison results. In some embodiments, AI processor 124 may be configured with an algorithm to combine the plurality of received operating conditions and compare the results of the algorithm to a threshold value. In some embodiments, the plurality of received operating conditions may be of the same type and / or different types. In some embodiments, the plurality of received operating conditions may be for a particular time and / or over a period of time.
[0038] For dynamic neural network quantization reconfiguration, the AI QoS manager 210 may determine an AI QoS value to be achieved by the AI processor 124. The AI QoS value may be configured to consider the AI processor throughput and accuracy of the AI processor results to be achieved as a result of the dynamic neural network quantization reconfiguration, and / or the AI processor operating frequency of the AI processor 124 under certain operating conditions. The AI QoS value may represent a level of latency, quality, accuracy, etc. for the AI processor 124 that is perceptible by a user and / or acceptable for mission-critical applications. In some embodiments, the AI QoS manager 210 may be configured with any number of algorithms, thresholds, lookup tables, etc., and combinations thereof, for determining the AI QoS value from the operating conditions. For example, the AI QoS manager 210 may determine an AI QoS value that considers the AI processor throughput and accuracy of the AI processor results as goals to be achieved by an AI processor 124 that exhibits a temperature above a temperature threshold. As a further example, AI QoS manager 210 may determine AI QoS values that consider the throughput of the AI processor and the accuracy of the AI processor's results as goals to be achieved by an AI processor 124 that exhibits a current (power consumption) above a current threshold. As a further example, AI QoS manager 210 may determine AI QoS values that consider the throughput of the AI processor and the accuracy of the AI processor's results as goals to be achieved by an AI processor 124 that exhibits a throughput value and / or utilization value that exceed a throughput threshold and / or utilization threshold. The foregoing examples described with respect to operating conditions exceeding thresholds are not intended to limit the scope of the claims or the specification and are equally applicable to embodiments in which operating conditions are below thresholds.
[0039] As described further herein, the dynamic quantization controller 208 may determine how to dynamically configure the AI processor 124, the dynamic neural network quantization logic 212, 214, and / or the MACs 202a-202i to achieve the AI QoS value. In some embodiments, the AI QoS manager 210 may be configured to execute an algorithm that calculates an AI quantization level to achieve the AI QoS value from values representing the accuracy of the AI processor and the throughput of the AI processor. For example, the algorithm may be an additive and / or minimum function of the accuracy of the AI processor and the throughput of the AI processor. As a further example, the value representing the accuracy of the AI processor may include an error value of the output of the neural network executed by the AI processor 124, and the value representing the throughput of the AI processor may include a value of the inferences per period produced by the AI processor 124. The algorithm may be weighted to prioritize either the accuracy of the AI processor or the throughput of the AI processor. In some embodiments, the weights may be associated with any number of operating conditions and combinations of operating conditions of the AI processor 124, the SoC 102, the memories 106, 114, and / or other peripherals 122. In some embodiments, the AI quantization level may be calculated along with the operating frequency of the AI processor to achieve an AI QoS value. The AI quantization level may vary relative to a previously calculated AI quantization level based on the impact of the operating conditions on the processing power of the AI processor 124. For example, operating conditions that indicate to the AI QoS manager 210 an increased processing power constraint for the AI processor 124 may result in an increase in the AI quantization level. As another example, operating conditions that indicate to the AI QoS manager 210 a decreased processing power constraint for the AI processor 124 may result in a decrease in the AI quantization level.
[0040] In some embodiments, AI QoS manager 210 may also determine whether to implement a conventional reduction in the AI processor operating frequency, alone or in combination with dynamic neural network quantization reconfiguration. For example, some of the thresholds for the operating conditions may be associated with a conventional reduction in the AI processor operating frequency and / or dynamic neural network quantization reconfiguration. An unfavorable result of comparing any number of received operating conditions or combinations thereof against thresholds associated with a reduction in the AI processor operating frequency and / or dynamic neural network quantization reconfiguration may cause AI QoS manager 210 to determine to implement a reduction in the AI processor operating frequency and / or dynamic neural network quantization reconfiguration. In some embodiments, AI QoS manager 210 may be adapted to control the operating frequency of MAC array 200.
[0041] The AI QoS manager 210 may generate an AI quantization level signal having the AI quantization level and send it to the dynamic quantization controller 208. The AI quantization level signal may cause the dynamic quantization controller 208 to determine parameters for implementing dynamic neural network quantization reconfiguration and provide the AI quantization level as an input for the parameter determination. In some embodiments, the AI quantization level signal may also include an operating condition that caused the AI QoS manager 210 to decide to implement dynamic neural network quantization reconfiguration. The operating condition may also be an input for determining parameters for implementing dynamic neural network quantization reconfiguration. In some embodiments, the operating condition may be represented by a value representing the value of the operating condition and / or the result of an algorithm that uses the operating condition, a comparison of the operating condition to a threshold value, a value from a lookup table for the operating condition, etc. For example, the value representing the result of the comparison may include the difference between the value of the operating condition and the value of the threshold value. In some embodiments, the AI QoS manager 210 may be adapted to vary the AI quantization level used by the MAC array 200, for example, by setting a particular AI quantization level or commanding the current level to be increased or decreased.
[0042] In some embodiments, AI QoS manager 210 may also generate and transmit an AI frequency signal to MAC array 200. The AI frequency signal may cause MAC array 200 to implement a reduction in the AI processor operating frequency. In some embodiments, MAC array 200 may be configured with means for implementing a reduction in the AI processor operating frequency. In some embodiments, AI QoS manager 210 may generate and transmit either or both of an AI quantization level signal and an AI frequency signal.
[0043] The dynamic quantization controller 208 may be configured as hardware, software executed by the AI processor 124, and / or a combination of hardware and software executed by the AI processor 124. The dynamic quantization controller 208 may be configured to determine parameters for the dynamic neural network quantization reconfiguration. In some embodiments, the dynamic quantization controller 208 may be pre-configured to determine parameters for any number of specific types of dynamic neural network quantization reconfiguration and combinations of those specific types of dynamic neural network quantization reconfiguration. In some embodiments, the dynamic quantization controller 208 may be configured to determine which parameters to determine for any number of types of dynamic neural network quantization reconfiguration and combinations of those types of dynamic neural network quantization reconfiguration.
[0044] Determining which parameters to determine for those types of dynamic neural network quantization reconfiguration may control which types of dynamic neural network quantization reconfiguration may be implemented. Types of dynamic neural network quantization reconfiguration may include configuring the dynamic neural network quantization logic 212, 214 for quantization of activation and weight values, configuring the dynamic neural network quantization logic 212, 214 for masking of activation and weight values and configuring the MAC array 200 and / or MACs 202a-202i for bypassing a portion of the MACs 202a-202i, and configuring the dynamic neural network quantization logic 212 for masking of weight values and configuring the MAC array 200 and / or MACs 202a-202i for bypassing the entire MACs 202a-202i. In some embodiments, the dynamic quantization controller 208 may be configured to determine a number of dynamic bits of parameters for configuring the dynamic neural network quantization logic 212, 214 for quantization of activation and weight values. In some embodiments, the dynamic quantization controller 208 may be configured to determine an additional parameter of a number of dynamic bits to configure the dynamic neural network quantization logic 212, 214 for masking activation values and weight values and bypassing portions of the MACs 202a-202i. In some embodiments, the dynamic quantization controller 208 may be configured to determine an additional parameter of a threshold weight value to configure the dynamic neural network quantization logic 212 for masking weight values and bypassing entire MACs 202a-202i.
[0045] The AI quantization level may differ from a previously calculated AI quantization level, which may result in differences in the determined parameters for performing the dynamic neural network quantization reconfiguration. For example, increasing the AI quantization level may cause the dynamic quantization controller 208 to determine a larger number of dynamic bits and / or lower threshold weight values for configuring the dynamic neural network quantization logic 212, 214. Increasing the number of dynamic bits and / or lowering the threshold weight values may cause fewer bits and / or fewer MACs 202a-202i to be used to perform the neural network calculations, which may reduce the accuracy of the neural network inference results. As another example, decreasing the AI quantization level may cause the dynamic quantization controller 208 to determine a smaller number of dynamic bits and / or higher threshold weight values for configuring the dynamic neural network quantization logic 212, 214. Reducing the number of dynamic bits and / or increasing the threshold weight values may result in more bits and / or more MACs 202a-202i being used to perform the neural network calculations, which may increase the accuracy of the neural network's inference results.
[0046] In some embodiments, the dynamic neural network quantization logic 212, 214 may dynamically implement the AI quantization levels using parameters determined by the dynamic quantization controller 208, which may be by masking, quantization, bypassing, or any other suitable means. The dynamic quantization controller 208 may receive an AI quantization level signal from the AI QoS manager 210. The dynamic quantization controller 208 may use the AI quantization levels received along with the AI quantization level signal to determine parameters for the dynamic neural network quantization reconfiguration. In some embodiments, the dynamic quantization controller 208 may also use operating conditions received along with the AI quantization level signal to determine parameters for the dynamic neural network quantization reconfiguration. In some embodiments, the dynamic quantization controller 208 may be configured with algorithms, thresholds, lookup tables, etc. to determine which parameters and / or parameter values of the dynamic neural network quantization reconfiguration to use based on the AI quantization levels and / or operating conditions. For example, the dynamic quantization controller 208 may use the AI quantization level and / or operating conditions as input to an algorithm that may output a number of dynamic bits to use for quantizing activation and weight values. In some embodiments, an additional algorithm may be used that may output a number of dynamic bits for masking activation and weight values and bypassing portions of the MACs 202a-202i. In some embodiments, an additional algorithm may be used that may output a threshold weight value for masking weight values and bypassing the entire MACs 202a-202i.
[0047] The dynamic quantization controller 208 may generate and send a dynamic quantization signal having parameters for the dynamic neural network quantization reconfiguration to the dynamic neural network quantization logic 212, 214. The dynamic quantization signal may cause the dynamic neural network quantization logic 212, 214 to perform the dynamic neural network quantization reconfiguration and provide parameters for performing the dynamic neural network quantization reconfiguration. In some embodiments, the dynamic quantization controller 208 may send the dynamic quantization signal to the MAC array 200. The dynamic quantization signal may cause the MAC array 200 to perform the dynamic neural network quantization reconfiguration and provide parameters for performing the dynamic neural network quantization reconfiguration. In some embodiments, the dynamic quantization signal may include an indicator of the type of dynamic neural network quantization reconfiguration to be implemented. In some embodiments, the indicator of the type of dynamic neural network quantization reconfiguration may be parameters for the dynamic neural network quantization reconfiguration.
[0048] The dynamic neural network quantization logic 212, 214 may be implemented in hardware. The dynamic neural network quantization logic 212, 214 may be configured to quantize the activation and weight values received from the activation buffer 206 and the weight buffer 204, such as by rounding the activation and weight values. Quantization of the activation and weight values may be performed using any type of rounding, such as rounding up or down to one dynamic bit, rounding up or down to one significant bit, rounding up or down to the nearest value, or rounding up or down to a specific value. For clarity and simplicity of explanation, an example of quantization is described with respect to rounding to dynamic bits, but this does not limit the scope of the claims and the description herein. The dynamic neural network quantization logic 212, 214 may provide the quantized activation and weight values to the MAC array 200. The dynamic neural network quantization logic 212, 214 may be configured to receive a dynamic quantization signal and perform dynamic neural network quantization reconfiguration.
[0049] The dynamic neural network quantization logic 212, 214 may receive a dynamic quantization signal from the dynamic quantization controller 208 and determine parameters for the dynamic neural network quantization reconfiguration. The dynamic neural network quantization logic 212, 214 may also determine a type of dynamic neural network quantization reconfiguration to implement from the dynamic quantization signal, which may include configuring the dynamic neural network quantization logic 212, 214 for a particular type of quantization. In some embodiments, the type of dynamic neural network quantization reconfiguration to implement may also include configuring the dynamic neural network quantization logic 212, 214 for masking of activation values and / or weight values. In some embodiments, masking of activation values and weight values may include replacing a number of dynamic bits with zero values. In some embodiments, masking of weight values may include replacing all of the bits with zero values.
[0050] The dynamic quantization signal may include a parameter of a certain number of dynamic bits to configure the dynamic neural network quantization logic 212, 214 for quantization of the activation and weight values. The dynamic neural network quantization logic 212, 214 may be configured to quantize the activation and weight values by rounding the bits of the activation and weight values to the number of dynamic bits indicated by the dynamic quantization signal.
[0051] The dynamic neural network quantization logic 212, 214 may include configurable logic gates that may be configured to round the bits of the activation values and weight values to the number of dynamic bits. In some embodiments, the logic gates may be configured to output zero values for the least significant bits of the activation values and weight values up to and / or including the number of dynamic bits. In some embodiments, the logic gates may be configured to output values for the most significant bits of the activation values and weight values up to and / or including the number of dynamic bits. For example, each bit of the activation values or weight values may be input to the logic gates sequentially, e.g., from least significant bit to most significant bit. The logic gates may output zero values for the least significant bits of the activation values and weight values up to and / or including the number of dynamic bits indicated by the parameter. The logic gates may output values for the most significant bits of the activation values and weight values up to and / or including the number of dynamic bits indicated by the parameter. As a further example, the weight values and activation values may be 8-bit integers, and a dynamic bit in that number may instruct the dynamic neural network quantization logic 212, 214 to round to the lower half of the 8-digit integer. The number of dynamic bits may be different from the default number of dynamic bits or the previous number of dynamic bits to round to for the default or previous configuration of the neural network quantization logic 212, 214. Thus, the configuration of the logic gates may also be different from the default or previous configuration of the logic gates.
[0052] The dynamic quantization signal may include a parameter of a number of dynamic bits for configuring the dynamic neural network quantization logic 212, 214 for masking the activation and weight values and bypassing a portion of the MACs 202a-202i. The dynamic neural network quantization logic 212, 214 may be configured to quantize the activation and weight values by masking the number of dynamic bits of the activation and weight values indicated by the dynamic quantization signal.
[0053] The dynamic neural network quantization logic 212, 214 may include configurable logic gates that may be configured to mask the number of dynamic bits of the activation and weight values. In some embodiments, the logic gates may be configured to output zero values for the lower-order bits of the activation and weight values up to and / or including the number of dynamic bits. In some embodiments, the logic gates may be configured to output values for the upper-order bits of the activation and weight values up to and / or including the number of dynamic bits. For example, each bit of the activation and weight values may be input to the logic gates sequentially, e.g., from least significant bit to most significant bit. The logic gates may output zero values for the lower-order bits of the activation and weight values up to and / or including the number of dynamic bits indicated by the parameter. The logic gates may output values for the upper-order bits of the activation and weight values up to and / or including the number of dynamic bits indicated by the parameter. The number of dynamic bits may be different from the default number of dynamic bits or the previous number of dynamic bits to be masked for the default or previous configuration of the dynamic neural network quantization logic 212, 214. Thus, the configuration of the logic gates may also be different from the default or previous configuration of the logic gates.
[0054] In some embodiments, the logic gates may be clock-gated to not receive and / or output the least significant bits of the activation and weight values up to and / or including the number of dynamic bits. Clock-gating the logic gates may effectively replace those least significant bits of the activation and weight values with values of 0, since the MAC array 200 may not receive values for those least significant bits of the activation and weight values.
[0055] In some embodiments, the dynamic neural network quantization logic 212, 214 may signal a parameter of the number of dynamic bits to the MAC array 200 for bypassing some of the MACs 202a-202i. In some embodiments, the dynamic neural network quantization logic 212, 214 may signal to the MAC array 200 which bits of the activation values and weight values are to be masked. In some embodiments, the absence of a signal for a bit of the activation values and weight values may be a signal from the dynamic neural network quantization logic 212, 214 to the MAC array 200.
[0056] In some embodiments, MAC array 200 may receive a dynamic quantization signal that includes a parameter for a number of dynamic bits for configuring dynamic neural network quantization logic 212, 214 for masking activation values and weight values and bypassing portions of MACs 202a-202i. In some embodiments, MAC array 200 may receive a signal from dynamic neural network quantization logic 212, 214 for a parameter for the number of dynamic bits and / or which dynamic bits for bypassing portions of MACs 202a-202i. MAC array 200 may be configured to bypass portions of MACs 202a-202i for dynamic bits of activation values and weight values indicated by the dynamic quantization signal and / or the signal from dynamic neural network quantization logic 212, 214. These dynamic bits may correspond to bits of activation values and weight values that are masked by dynamic neural network quantization logic 212, 214.
[0057] MACs 202a-202i may include logic gates configured to implement multiplication and accumulation functions. In some embodiments, MAC array 200 may clock gate logic gates in MACs 202a-202i configured to multiply and accumulate activation and weight value bits corresponding to the number of dynamic bits indicated by the parameters of the dynamic quantization signal. In some embodiments, MAC array 200 may clock gate logic gates in MACs 202a-202i configured to multiply and accumulate activation and weight value bits corresponding to the number of dynamic bits and / or particular dynamic bits indicated by the signals from dynamic neural network quantization logic 212, 214.
[0058] In some embodiments, MAC array 200 may power collapse logic gates in MACs 202a-202i configured to multiply and accumulate activation and weight value bits corresponding to the number of dynamic bits indicated by the parameters of the dynamic quantization signal. In some embodiments, MAC array 200 may power collapse logic gates in MACs 202a-202i configured to multiply and accumulate activation and weight value bits corresponding to the number of dynamic bits and / or particular dynamic bits indicated by the signals from dynamic neural network quantization logic 212, 214.
[0059] By clock gating and / or powering down the logic gates of MACs 202a-202i, MACs 202a-202i may not receive the activation and weight value bits corresponding to that number of dynamic bits or particular dynamic bits, effectively masking those bits. Further examples of clock gating and / or powering down the logic gates of MACs 202a-202i are described herein with reference to FIG.
[0060] The dynamic quantization signal may include a threshold weight value parameter for configuring the dynamic neural network quantization logic 212 for masking the weight values and bypassing the entire MACs 202a-202i. The dynamic neural network quantization logic 212 may be configured to quantize the weight values by masking all of the bits of the weight values based on a comparison of the weight values to the threshold weight values indicated by the dynamic quantization signal.
[0061] The dynamic neural network quantization logic 212 may include a configurable logic gate that may be configured to compare weight values received from the weight buffer 204 to a threshold weight value and mask weight values that result in an unfavorable comparison, such as being less than the threshold weight value or less than or equal to the threshold weight value. In some embodiments, the comparison may be a comparison of the absolute value of the weight value to the threshold weight value. In some embodiments, the logic gate may be configured to output a value of 0 for all of the bits of a weight value that result in an unfavorable comparison with the threshold weight value. All of the bits may be a different number of bits than the default number of bits or a different number of bits than a previous number of bits to be masked for a default or previous configuration of the dynamic neural network quantization logic 212. Thus, the configuration of the logic gate may also be different from the default or previous configuration of the logic gate.
[0062] In some embodiments, the logic gates may be clock gated to prevent receiving and / or outputting bits of weight values that result in an unfavorable comparison with the threshold weight value. Clock gating the logic gates may effectively replace the bits of the weight value with a value of 0, since the MAC array 200 may not receive the value of the bits of the weight value. In some embodiments, the dynamic neural network quantization logic 212 may signal to the MAC array 200 which bits of the weight value are masked. In some embodiments, the absence of a signal for a bit of the weight value may be a signal from the dynamic neural network quantization logic 212 to the MAC array 200.
[0063] In some embodiments, MAC array 200 may receive a signal from dynamic neural network quantization logic 212 for which bits of a weight value are masked. MAC array 200 may interpret the entire masked weight value as a signal to bypass MACs 202a-202i entirely. MAC array 200 may be configured to bypass MACs 202a-202i for the weight values indicated by the signal from dynamic neural network quantization logic 212. These weight values may correspond to the weight values that are masked by dynamic neural network quantization logic 212.
[0064] MACs 202a-202i may include logic gates configured to implement multiplication and accumulation functions. In some embodiments, MAC array 200 may clock gate the logic gates of MACs 202a-202i configured to multiply and accumulate weight value bits corresponding to the masked weight values. In some embodiments, MAC array 200 may power down the logic gates of MACs 202a-202i configured to multiply and accumulate weight value bits corresponding to the masked weight values. By clock gating and / or powering down the logic gates of MACs 202a-202i, MACs 202a-202i may not receive activation values and weight value bits corresponding to the masked weight values.
[0065] Masking weight values by dynamic neural network quantization logic 212 and / or clock gating and / or powering down MACs 202a-202i may prune the neural network executed by MAC array 200. Removing weight values and MAC operations from the neural network may effectively remove synapses and nodes from the neural network. The weight threshold may be determined based on the fact that removing weight values that compare unfavorably to the weight threshold from the neural network execution may cause only an acceptable loss in accuracy of the AI processor's results.
[0066] 2B illustrates an embodiment of the AI processor 124 shown in FIG. 2A. Referring to FIGS. 1-2B, the AI processor 124 may include dynamic neural network quantization logic 212, 214, which may be implemented as hardware circuit logic rather than as a software tool or in a compiler. The activation buffer 206 and weight buffer 204, the dynamic quantization controller 208, the hardware dynamic neural network quantization logic 212, 214, and the MAC array 200 may function and interact as described with reference to FIG. 2A.
[0067] 3 illustrates an exemplary SoC having a dynamic neural network quantization architecture suitable for implementing various embodiments. Referring to FIGS. 1-3, SoC 102 may include any number of AI processing subsystems 300 and memory 106, as well as combinations thereof. AI processing subsystem 300 may include any number of AI processors 124a-124f, input / output (I / O) interfaces 302, and memory controller / physical layer components 304a-304f, as well as combinations thereof.
[0068] As discussed herein with respect to an AI processor (e.g., 124), in some embodiments, the dynamic neural network quantization reconfiguration may be performed using an AI processor. In some embodiments, the dynamic neural network quantization reconfiguration may be performed, at least in part, before the activation values and weight values are received by AI processors 124a-124f.
[0069] The I / O interface 302 may be configured to control communications between the AI processing subsystem 300 and other components of the computing device (e.g., 100), including a processor (e.g., 104), a communications interface (e.g., communications interface (e.g., 108)), communications components (e.g., 112), peripheral device interfaces (e.g., 120), peripheral devices (e.g., 120), etc. Some such communications may include receiving activation values. In some embodiments, the I / O interface 302 may be configured to include and / or implement functionality of an AI QoS manager (e.g., 210), a dynamic quantization controller (e.g., 208), and / or dynamic neural network quantization logic (e.g., 212). In some embodiments, the I / O interface 302 may be configured to implement functionality of an AI QoS manager, a dynamic quantization controller, and / or dynamic neural network quantization logic through hardware, software executing on the I / O interface 302, and / or hardware and software executing on the I / O interface 302.
[0070] The memory controller / physical layer components 304a-304f may be configured to control communications between the AI processors 124a-124f, the memory 106, and / or memories local to the AI processing subsystem 300 and / or the AI processors 124a-124f. Some such communications may include reading and writing weights and activation values from and to the memory 106.
[0071] In some embodiments, the memory controller / physical layer components 304a-304f may be configured to include and / or implement the functionality of an AI QoS manager, a dynamic quantization controller, and / or dynamic neural network quantization logic. For example, the memory controller / physical layer components 304a-304f may quantize and / or mask activation values and / or weight values during the initial writing or reading of the weight values and / or activation values to or from memory 106. As a further example, the memory controller / physical layer components 304a-304f may quantize and / or mask weight values while writing the weight values to local memory when transferring the weight values from memory 106. As a further example, the memory controller / physical layer components 304a-304f may quantize and / or mask activation values while the activation values are being generated.
[0072] In some embodiments, memory controller / physical layer components 304a-304f may be configured to implement the functionality of the AI QoS manager, dynamic quantization controller, and / or dynamic neural network quantization logic through hardware, software executing on memory controller / physical layer components 304a-304f, and / or hardware and software executing on memory controller / physical layer components 304a-304f.
[0073] The I / O interface 302 and / or memory controller / physical layer components 304a-304f may be configured to provide quantized and / or masked weights and / or activation values to the AI processors 124a-124f. In some embodiments, the I / O interface 302 and / or memory controller / physical layer components 304a-304f may be configured not to provide fully masked weight values to the AI processors 124a-124f.
[0074] 4A and 4B illustrate example AI QoS relationships suitable for implementing various embodiments. Referring to FIGS. 1-4B, for dynamic neural network quantization reconfiguration, an AI QoS manager (e.g., 210) may determine an AI QoS value that takes into account the throughput of the AI processor and the accuracy of the AI processor's results to be achieved as a result of the dynamic neural network quantization reconfiguration under certain operating conditions.
[0075] FIG. 4A shows a graph 400a depicting a measurement of the accuracy of an AI processor's results in terms of AI QoS values on the vertical axis, relative to the bit width of weights and activation values quantized using dynamic neural network quantization reconfiguration on the horizontal axis. Curve 402a indicates that the larger the bit width of the weights and activation values, the more accurate the AI processor's results may be. However, curve 402a also indicates that the benefit of the bit width of the weights and activation values diminishes as the bit width of the weights and activation values increases as the slope of curve 402a approaches zero. Thus, for some bit widths of weights and activation values that are less than the maximum bit width, the accuracy of the AI processor's results may show only negligible change.
[0076] Curve 402a further shows that the slope of curve 402a increases at a greater rate at points where the bit width of some of the weights and activation values is smaller than the maximum bit width. Thus, for some of the bit widths of the weights and activation values that are smaller than the maximum bit width, the accuracy of the AI processor's results may show a non-negligible change. For weight and activation value bit widths that show negligible change, dynamic neural network quantization reconfiguration of the accuracy of the AI processor's results may be implemented to quantize the weights and activation values and still achieve an acceptable level of accuracy of the AI processor's results.
[0077] 4B shows a graph 400b depicting a measure of AI processor responsiveness, sometimes referred to as latency, in terms of AI processor throughput for a dynamic neural network quantization reconfiguration implementation on the horizontal axis and AI QoS values on the vertical axis. In some embodiments, throughput may include a value of inferences per period produced by the AI processor, such as inferences per second. Throughput may increase for a dynamic neural network quantization reconfiguration implementation in response to smaller bit widths of activation values and / or weight values.
[0078] Curve 402b shows that the higher the throughput of the AI processor, the more responsive the AI processor may be. However, curve 402b also shows that the benefits of the AI processor's throughput diminish as the AI processor's throughput increases as the slope of curve 402b approaches zero. Thus, the responsiveness of the AI processor may show only a negligible change for any AI processor throughput that is lower than the throughput of the highest AI processor.
[0079] Curve 402b further shows that at points where the throughput of some AI processors is even lower than that of the best AI processor, the slope of curve 402b increases at a greater rate. Thus, for some AI processors whose throughput is even lower than that of the best AI processor, the responsiveness of the AI processor may exhibit a non-negligible change. For AI processors whose throughput exhibits a negligible change, AI processor responsiveness and dynamic neural network quantization reconfiguration may be implemented to quantize the activation values and / or weight values and still achieve an acceptable level of AI processor responsiveness.
[0080] 5 illustrates an exemplary benefit in AI processor operating frequency implementing a dynamic neural network quantization architecture in various embodiments. With reference to FIGS. 1-5, for dynamic neural network quantization reconfiguration, the dynamic neural network quantization logic (e.g., 212, 214), the I / O interface (e.g., 302), and / or the memory controller / physical layer components (e.g., 304a-304f) may implement the dynamic neural network quantization reconfiguration to achieve a level of AI processor throughput and / or accuracy of AI processor results.
[0081] FIG. 5 shows a graph 500 representing the measurement of the AI processor operating frequency, which can affect the throughput of the AI processor, on the vertical axis, relative to the bit width of the weights and activation values on the horizontal axis. Graph 500 is also shaded to represent operating conditions under which the AI processor may operate. For example, the operating condition may be the temperature of the AI processor, with darker shading representing higher temperatures. Thus, the lowest temperature may be at the origin of the graph and the highest temperature may be opposite the origin. For point 502, dynamic neural network quantization reconfiguration is not performed, the weights and activation values may remain at their maximum bit width, and the only way to reduce the temperature is to reduce the operating frequency of the AI processor. Excessive reduction in the operating frequency of the AI processor can result in poor AI QoS and latency, which can cause catastrophic problems in mission-critical systems such as automotive systems. For point 504, dynamic neural network quantization reconfiguration is performed, and both the operating frequency of the AI processor may be reduced and the bit width of the weights and activation values may be quantized to be smaller than the maximum bit width to achieve a temperature reduction similar to that shown by point 502. Point 504 indicates that by reducing the bit width of the weight values and activation values and using dynamic neural network quantization reconfiguration, the AI processor operating frequency may be higher compared to the AI processor operating frequency at point 502, while the temperature operating conditions at both points 502 and 504 are similar. Therefore, dynamic neural network quantization reconfiguration may achieve better AI processor performance, such as AI processor throughput, under similar operating conditions, such as the temperature of the AI processor, compared to not using dynamic neural network quantization reconfiguration.
[0082] FIG. 6 illustrates an exemplary benefit in AI processor operating frequency implementing a dynamic neural network quantization architecture in various embodiments. Referring to FIGS. 1-6 , for dynamic neural network quantization reconfiguration, the dynamic neural network quantization logic (e.g., 212, 214), I / O interface (e.g., 302), and / or memory controller / physical layer components (e.g., 304a-304f) may implement dynamic neural network quantization reconfiguration to achieve a level of AI processor throughput and / or accuracy of AI processor results. FIG. 6 illustrates graphs 600a, 600b, 604a, 604b, and 608 representing measurements of AI processor operating conditions, which may affect the AI processor's throughput, plotted with respect to time. Graph 600a represents measurements of AI processor temperature without implementing dynamic neural network quantization reconfiguration on the vertical axis, with respect to time on the horizontal axis. Graph 600b represents measurements of AI processor temperature with implementing dynamic neural network quantization reconfiguration on the vertical axis, with respect to time on the horizontal axis. Graph 604a represents a measurement of AI processor frequency on the vertical axis without implementing dynamic neural network quantization reconfiguration, with respect to time on the horizontal axis. Graph 604b represents a measurement of AI processor frequency on the vertical axis with implementing dynamic neural network quantization reconfiguration, with respect to time on the horizontal axis. Graph 608 represents a measurement of AI processor bitwidth for activation and / or weight values on the vertical axis with implementing dynamic neural network quantization reconfiguration, with respect to time on the horizontal axis.
[0083] Before time 612, AI processor temperature 602a in graph 600a may increase, while AI processor frequency 606a in graph 604a may remain stable. Similarly, before time 612, AI processor temperature 602b in graph 600b may increase, while AI processor frequency 606b in graph 604b and AI processor bitwidth 610 in graph 608 may remain stable. A reason for an increase in AI processor temperature 602a, 602b without a change in AI processor frequency 606a, 606b and / or AI processor bitwidth 610 may be an increase in workload on the AI processors (e.g., 124, 124a-124f).
[0084] At time 612, AI processor temperature 602a may peak, and AI processor frequency 606a may decrease. The lower AI processor frequency 606a may cause the AI processor to generate less heat than before time 612, and AI processor temperature 602a may stop increasing because the lower AI processor frequency 606a consumes less power. Similarly, at time 612, AI processor temperature 602b may peak, and AI processor frequency 606b may decrease. However, at time 612, AI processor bit-width 610 may also decrease. The lower AI processor frequency 606b and smaller AI processor bit-width 610 may cause the AI processor to generate less heat than before time 612, and AI processor temperature 602b may stop increasing because the lower AI processor frequency 606b consumes less power and processes smaller bit-width data.
[0085] The difference between the AI processor frequency 614a before time 612 and at time 612 may be greater than the difference between the AI processor frequency 614b before time 612 and at time 612, relative to each other. Reducing the AI processor bit-width 610 in conjunction with reducing the AI processor operating frequency 606b may allow the reduction in the AI processor operating frequency 606b to be less than the reduction in the AI processor operating frequency 606a when reducing the AI processor operating frequency 606a alone. Reducing the AI processor bit-width 610 may provide similar benefits to the AI processor temperature 602a, 602b as reducing the AI processor operating frequency 606a alone, but may also provide the benefit of a higher AI processor operating frequency 606b, which may affect the throughput of the AI processor.
[0086] FIG. 7 illustrates an example of a detour in a MAC in a dynamic neural network quantization architecture for implementing various embodiments. Referring to FIGS. 1-7, MAC 202 may include logic circuitry including various logic components 700, 702, such as any number of AND gates, full adders (labeled "F" in FIG. 7), and / or half adders (labeled "H" in FIG. 7), and combinations thereof. The example illustrated in FIG. 7 illustrates MAC 202 having logic circuitry typically configured for 8-bit multiplication and accumulation functions. However, MAC 202 may typically be configured for multiplication and accumulation functions of any bit-width data, and the example illustrated in FIG. 7 does not limit the scope of the claims and the description herein.
[0087] In some embodiments, lines X0-X7, Y0-Y7 may provide activation and weight inputs to MAC 202. X0 and Y0 may represent the least significant bits, and X7 and Y7 may represent the most significant bits of the activation and weight values. As described herein, dynamic neural network quantization reconfiguration may include quantizing and / or masking any number of dynamic bits of the activation and / or weight values. Quantizing and / or masking bits of the activation and / or weight values may round and / or replace bits of the weight values to a value of zero. Thus, multiplying a quantized and / or masked bit of an activation and / or weight value with another bit of the activation and / or weight value may result in a value of zero. Given the known result of multiplying the quantized and / or masked activation and / or weight values, it may not be necessary to actually perform the multiplication and addition of the results. Thus, an AI processor (e.g., 124, 124a-123f) including a MAC array (e.g., 200) may clock gate to turn off logic components 702 for multiplication of quantized and / or masked activation values and / or weight values and summing the results. Clock gating logic components 702 for multiplication of masked weight values and summing the results may reduce circuit switching power dissipation, also referred to as dynamic power reduction.
[0088] 7, the lower two bits of the activation and weight values on lines X0, X1, Y0, or Y1 are masked. The shaded corresponding logic components 702 that receive as inputs X0, X1, Y0, or Y1 and / or the results of the operations for X0, X1, Y0, and / or Y1 are shaded to indicate that they are clock-gated to be turned off. The remaining unshaded logic components 700 are unshaded to indicate that they are not clock-gated to be turned off.
[0089] 8 illustrates a method 800 for AI QoS determination, according to one embodiment. Referring to FIGS. 1-8, method 800 may be implemented in a computing device (e.g., 100), in general-purpose hardware, in dedicated hardware (e.g., 210), in software executing in a processor (e.g., processor 104, AI processor 124, AI QoS manager 210, AI processing subsystem 300, AI processors 124a-124f, I / O interface 302, memory controller / physical layer components 304a-304f), or in a combination of software-configured processors and dedicated hardware, such as processors executing software in a dynamic neural network quantization system, including other individual components, various memory / cache controllers, etc. To encompass the alternative reconfigurations possible in various embodiments, the hardware that implements method 800 is referred to herein as an "AI QoS device."
[0090] In block 802, an AI QoS device may receive AI QoS factors. The AI QoS device may be communicatively coupled to any number and combination of sensors, such as temperature sensors, voltage sensors, current sensors, etc., and a processor. The AI QoS device may receive data signals representing the AI QoS factors from these communicatively coupled sensors and / or processor. The AI QoS factors may be operating conditions, and the dynamic neural network quantization logic reconfiguration may be based on the operating conditions to modify quantization, masking, and / or neural network pruning. These operating conditions may include temperature, power consumption, processing unit utilization, performance, etc., of the AI processor, the SoC having the AI processor (e.g., 102), memory accessed by the AI processor (e.g., 106, 114), and / or other peripheral devices of the AI processor (e.g., 122). For example, the temperature may be a temperature sensor value representing the temperature at a location on the AI processor. In a further example, the power may be a value representing the peak of a power supply rail compared to the power supply, and / or the capability of a power management integrated circuit, and / or the state of charge of a battery. As a further example, performance may be a value representing utilization, completely idle time, frames per second, and / or end-to-end latency of an AI processor. In some embodiments, an AI QoS manager may be configured to receive the AI QoS factors at block 802. In some embodiments, an I / O interface and / or memory controller / physical layer component may be configured to receive the AI QoS factors at block 802.
[0091] In decision block 804, the AI QoS device may determine whether to dynamically configure neural network quantization. In some embodiments, in decision block 804, an AI QoS manager may be configured to determine whether to dynamically configure neural network quantization. In some embodiments, in decision block 804, an I / O interface and / or memory controller / physical layer component may be configured to determine whether to dynamically configure neural network quantization. The AI QoS device may determine whether to implement dynamic neural network quantization reconfiguration from operating conditions. The AI QoS device may determine to dynamically configure neural network quantization based on a level of operating conditions that increase constraints on the AI processor's processing power. The AI QoS device may determine to dynamically configure neural network quantization based on a level of operating conditions that reduce constraints on the AI processor's processing power. Constraints on the AI processor's processing power may be caused by operating condition levels, such as levels of heat buildup, power consumption, and processing unit utilization, which affect the AI processor's ability to maintain a level of processing power.
[0092] In some embodiments, the AI QoS device may be configured with any number of algorithms, thresholds, lookup tables, etc., and combinations thereof, for determining whether to perform dynamic neural network quantization reconfiguration from the operating conditions. For example, the AI QoS device may compare the received operating conditions to an operating condition threshold. In response to an unfavorable result of the comparison of the operating conditions to the operating condition threshold, such as exceeding the threshold, the AI QoS device may determine to perform dynamic neural network quantization reconfiguration at decision block 804. Such an unfavorable comparison may indicate to the AI QoS device that the operating conditions have increased the processing power constraints of the AI processor. In response to a favorable result of the comparison of the operating conditions to the operating condition threshold, such as being below the threshold, the AI QoS device may determine to perform dynamic neural network quantization reconfiguration at decision block 804. Such a favorable comparison may indicate to the AI QoS device that the operating conditions have decreased the processing power constraints of the AI processor.
[0093] In some embodiments, the AI QoS device may compare a plurality of received operating conditions to a plurality of threshold operating condition values and determine to perform dynamic neural network quantization reconfiguration based on a combination of unfavorable and / or favorable comparison results. In some embodiments, the AI device may be configured with an algorithm for combining the plurality of received operating conditions and compare the results of the algorithm to a threshold value. In some embodiments, the plurality of received operating conditions may be of the same type and / or different types. In some embodiments, the plurality of received operating conditions may be for a particular time and / or over a period of time.
[0094] In response to determining to dynamically configure neural network quantization (i.e., decision block 804="Yes"), the AI QoS device may determine an AI QoS value in block 805. For dynamic neural network quantization reconfiguration, the AI QoS device may determine an AI QoS value to be achieved by the AI processor, taking into account the throughput of the AI processor and the accuracy of the AI processor's results to be achieved as a result of the dynamic neural network quantization reconfiguration, and / or the AI processor operating frequency of the AI processor under certain operating conditions. The AI QoS value may represent a level of latency, quality, accuracy, etc. for the AI processor that is perceptible to a user and / or acceptable for mission-critical applications.
[0095] In some embodiments, the AI QoS device may be configured with any number of algorithms, thresholds, lookup tables, etc., and combinations thereof, for determining an AI QoS value from an operating condition. For example, the AI QoS device may determine an AI QoS value that considers the throughput of an AI processor and the accuracy of the AI processor's results as a goal to be achieved by an AI processor that exhibits a temperature above a temperature threshold. As a further example, the AI QoS device may determine an AI QoS value that considers the throughput of an AI processor and the accuracy of the AI processor's results as a goal to be achieved by an AI processor that exhibits a current (power consumption) above a current threshold. As a further example, the AI QoS device may determine an AI QoS value that considers the throughput of an AI processor and the accuracy of the AI processor's results as a goal to be achieved by an AI processor that exhibits a throughput value and / or utilization value above a throughput threshold and / or utilization threshold. The foregoing examples described with respect to operating conditions exceeding thresholds are not intended to limit the scope of the claims or the specification and are equally applicable to embodiments in which operating conditions are below thresholds. In some embodiments, the AI QoS manager may be configured to determine the AI QoS value at block 805. In some embodiments, at block 805, an I / O interface and / or memory controller / physical layer component may be configured to determine an AI QoS value.
[0096] In operation block 806, the AI QoS device may determine whether to reduce the AI processor operating frequency. The AI QoS device may also determine whether to implement a conventional reduction of the AI processor operating frequency, alone or in combination with dynamic neural network quantization reconfiguration. For example, some of the thresholds for the operating conditions may be associated with a conventional reduction of the AI processor operating frequency and / or dynamic neural network quantization reconfiguration. An unfavorable result of comparing any number of received operating conditions or combinations thereof against thresholds associated with an AI processor operating frequency reduction and / or dynamic neural network quantization reconfiguration may cause the AI QoS device to determine to implement an AI processor operating frequency reduction and / or dynamic neural network quantization reconfiguration. In some embodiments, in optional decision block 806, the AI QoS manager may be configured to determine whether to reduce the AI processor operating frequency. In some embodiments, an I / O interface and / or a memory controller / physical layer component may be configured to determine whether to reduce the AI processor operating frequency in optional decision block 806.
[0097] Following determining the AI QoS value in block 805, or in response to determining not to reduce the AI processor operating frequency (i.e., optional decision block 806="No"), in block 808, the AI QoS device may determine an AI quantization level to achieve the AI QoS value. The AI QoS device may determine the AI quantization level that takes into account the throughput of the AI processor and the accuracy of the AI processor's results to be achieved as a result of the dynamic neural network quantization reconfiguration under certain operating conditions. For example, the AI QoS device may determine the AI quantization level that takes into account the throughput of the AI processor and the accuracy of the AI processor's results as goals to be achieved by an AI processor that exhibits a temperature above a temperature threshold. In some embodiments, the AI QoS device may be configured to execute an algorithm that calculates the AI quantization level from any number of values representing the accuracy of the AI processor and the throughput of the AI processor, or a combination thereof, such as the AI QoS value. For example, the algorithm may be an additive and / or minimum function of the accuracy of the AI processor and the throughput of the AI processor. As a further example, a value representing the accuracy of an AI processor may include an error value of the output of a neural network executed by the AI processor, and a value representing the throughput of an AI processor may include a value of the inferences per time period produced by the AI processor. Algorithms may be weighted to prioritize either the accuracy of the AI processor or the throughput of the AI processor. In some embodiments, the weights may be associated with any number and combination of operating conditions of the AI processor, the SoC having the AI processor, memory accessed by the AI processor, and / or other peripheral devices of the AI processor. The AI quantization level may vary relative to a previously calculated AI quantization level based on the impact of the operating conditions on the processing power of the AI processor. For example, operating conditions that indicate to the AI QoS device increased constraints on the processing power of the AI processor may result in an increase in the AI quantization level.As another example, operating conditions that indicate a reduced constraint on the processing power of the AI processor to the AI QoS device may result in a reduction in the AI quantization level. In some embodiments, the AI QoS manager may be configured to determine the AI quantization level at block 808. In some embodiments, an I / O interface and / or memory controller / physical layer component may be configured to determine the AI quantization level at block 808.
[0098] At block 810, the AI QoS device may generate and transmit an AI quantization level signal. The AI QoS device may generate and transmit the AI quantization level signal having the AI quantization levels. In some embodiments, the AI QoS device may transmit the AI quantization level signal to a dynamic quantization controller (e.g., 208). In some embodiments, the AI QoS device may transmit the AI quantization level signal to an I / O interface and / or a memory controller / physical layer component. The AI quantization level signal may cause a recipient to determine parameters for performing dynamic neural network quantization reconfiguration and provide the AI quantization levels as input for the parameter determination. In some embodiments, the AI quantization level signal may also include operating conditions that caused the AI QoS device to decide to perform dynamic neural network quantization reconfiguration. The operating conditions may also be input for determining parameters for performing dynamic neural network quantization reconfiguration. In some embodiments, the operating conditions may be represented by values representing the values of the operating conditions and / or the results of an algorithm using the operating conditions, a comparison of the operating conditions to a threshold value, a value from a lookup table for the operating conditions, etc. For example, the value representing the result of the comparison may include the difference between the value of the operating condition and the value of the threshold value. In some embodiments, an AI QoS manager may be configured to generate and transmit the AI quantization level signal in block 810. In some embodiments, an I / O interface and / or memory controller / physical layer component may be configured to generate and transmit the AI quantization level signal in block 810. In block 802, an AI QoS device may receive AI QoS factors repeatedly, periodically, and / or continuously.
[0099] In response to determining to lower the AI processor operating frequency (i.e., optional decision block 806="Yes"), in optional block 812, the AI QoS device may determine an AI quantization level and an AI processor operating frequency value. The AI QoS device may determine the AI quantization level as in block 808. The AI QoS device may similarly determine the AI processor operating frequency value through the use of any number of algorithms, thresholds, lookup tables, etc., and combinations thereof. The AI processor operating frequency value may indicate the operating frequency value to which the AI processor operating frequency should be lowered. The AI processor operating frequency may be based on the AI QoS value determined in block 805. In some embodiments, the AI quantization level may be calculated along with the AI processor operating frequency to achieve the AI QoS value. In some embodiments, the AI QoS manager may be configured to determine the AI quantization level and the AI processor operating frequency value in optional block 812. In some embodiments, an I / O interface and / or memory controller / physical layer component may be configured to determine the AI quantization level and the AI processor operating frequency value in optional block 812.
[0100] In optional block 814, the AI QoS device may generate and transmit an AI quantization level signal and an AI frequency signal. The AI QoS device may generate and transmit the AI quantization level signal as in block 810. The AI QoS device may also generate and transmit an AI frequency signal to the MAC array (e.g., 200). The AI frequency signal may include an AI processor operating frequency value. The AI frequency signal may, for example, use the AI processor operating frequency value to cause the MAC array to implement a reduction in the AI processor operating frequency. In some embodiments, in optional block 814, the AI QoS manager may be configured to generate and transmit the AI quantization level signal and the AI frequency signal. In some embodiments, in optional block 814, the I / O interface and / or memory controller / physical layer component may be configured to generate and transmit the AI quantization level signal and the AI frequency signal. In block 802, the AI QoS device may receive the AI QoS factors repeatedly, periodically, and / or continuously.
[0101] In response to determining not to dynamically configure neural network quantization (i.e., decision block 804="No"), the AI QoS device may determine whether to reduce the AI processor operating frequency in optional decision block 816. The AI QoS manager may determine whether to reduce the AI processor operating frequency, as in optional decision block 806. In some embodiments, the AI QoS manager may be configured to determine whether to reduce the AI processor operating frequency in optional decision block 806. In some embodiments, the I / O interface and / or memory controller / physical layer component may be configured to determine whether to reduce the AI processor operating frequency in optional decision block 806.
[0102] In response to determining to lower the AI processor operating frequency (i.e., optional decision block 816="Yes"), in optional block 818, the AI QoS device may determine the AI processor operating frequency value. The AI QoS device may determine the AI processor operating frequency as in optional decision block 812. In some embodiments, in optional block 818, the AI QoS manager may be configured to determine the AI processor operating frequency value. In some embodiments, in optional block 818, an I / O interface and / or memory controller / physical layer component may be configured to determine the AI processor operating frequency value.
[0103] In optional block 820, the AI QoS device may generate and transmit the AI frequency signal. The AI QoS device may generate and transmit the AI frequency signal, as in optional block 814. In some embodiments, in optional block 820, the AI QoS manager may be configured to generate and transmit the AI frequency signal. In some embodiments, in optional block 820, an I / O interface and / or memory controller / physical layer component may be configured to generate and transmit the AI frequency signal. In block 802, the AI QoS device may receive the AI QoS factors repeatedly, periodically, or continuously.
[0104] In response to determining not to reduce the AI processor operating frequency (ie, optional decision block 816="No"), at block 802, the AI QoS device may receive an AI QoS factor.
[0105] 9 illustrates a method 900 for dynamic neural network quantization architecture configuration control, according to one embodiment. Referring to FIGS. 1-9, method 900 may be implemented in a computing device (e.g., 100), in general-purpose hardware, in dedicated hardware (e.g., dynamic quantization controller 208), in software executing in a processor (e.g., processor 104, AI processor 124, dynamic quantization controller 208, AI processing subsystem 300, AI processors 124a-124f, I / O interface 302, memory controller / physical layer components 304a-304f), or in a combination of software-configured processors and dedicated hardware, such as processors executing software in a dynamic neural network quantization system, including other individual components, various memory / cache controllers, etc. To encompass alternative configurations possible in various embodiments, the hardware that performs method 900 is referred to herein as a “dynamic quantization device.” In some embodiments, method 900 may be performed following block 810 and / or optional block 814 of method 800 (FIG. 8).
[0106] In block 902, a dynamic quantization device may receive an AI quantization level signal. The dynamic quantization device may receive the AI quantization level signal from an AI QoS device (e.g., AI QoS manager 210, I / O interface 302, memory controller / physical layer component 304a-304f). In some embodiments, a dynamic quantization controller may be configured to receive the AI quantization level signal in block 902. In some embodiments, an I / O interface and / or a memory controller / physical layer component may be configured to receive the AI quantization level signal in block 902.
[0107] At block 904, the dynamic quantization device may determine a number of dynamic bits for dynamic quantization. The dynamic quantization device may use the AI quantization levels received along with the AI quantization level signal to determine parameters for the dynamic neural network quantization reconfiguration. In some embodiments, the dynamic quantization device may also use the operating conditions received along with the AI quantization level signal to determine parameters for the dynamic neural network quantization reconfiguration. In some embodiments, the dynamic quantization device may be configured with an algorithm, thresholds, lookup tables, etc. to determine which parameters and / or parameter values to use for the dynamic neural network quantization reconfiguration based on the AI quantization levels and / or operating conditions. For example, the dynamic quantization device may use the AI quantization levels and / or operating conditions as input to an algorithm that may output a number of dynamic bits to use for quantizing activation values and weight values. In some embodiments, at block 904, a dynamic quantization controller may be configured to determine a number of dynamic bits for dynamic quantization. In some embodiments, at block 904, an I / O interface and / or memory controller / physical layer component may be configured to determine a number of dynamic bits for dynamic quantization.
[0108] In optional block 906, a dynamic quantization device may determine a number of dynamic bits for masking activation and weight values and bypassing portions of the MAC (e.g., 202a-202i). The dynamic quantization device may use the AI quantization levels received along with the AI quantization level signal to determine parameters for the dynamic neural network quantization reconfiguration. In some embodiments, the dynamic quantization device may also use the operating conditions received along with the AI quantization level signal to determine parameters for the dynamic neural network quantization reconfiguration. In some embodiments, the dynamic quantization device may be configured with algorithms, thresholds, lookup tables, etc. to determine which parameters and / or parameter values to use for the dynamic neural network quantization reconfiguration based on the AI quantization levels and / or operating conditions. For example, the dynamic quantization device may use the AI quantization levels and / or operating conditions as input to an algorithm that may output a number of dynamic bits for masking activation and weight values and bypassing portions of the MAC. In some embodiments, a dynamic quantization controller may be configured to determine a number of dynamic bits for masking activation and weight values and bypassing portions of the MAC in optional block 906. In some embodiments, an I / O interface and / or memory controller / physical layer component may be configured to determine a number of dynamic bits for masking activation and weight values and bypassing portions of the MAC in optional block 906.
[0109] At optional block 908, the dynamic quantization device may determine threshold weight values for dynamic network pruning. The dynamic quantization device may use the AI quantization levels received along with the AI quantization level signal to determine parameters for the dynamic neural network quantization reconfiguration. In some embodiments, the dynamic quantization device may also use the operating conditions received along with the AI quantization level signal to determine parameters for the dynamic neural network quantization reconfiguration. In some embodiments, the dynamic quantization device may be configured with algorithms, thresholds, lookup tables, etc. to determine which parameters and / or parameter values to use for the dynamic neural network quantization reconfiguration based on the AI quantization levels and / or operating conditions. For example, the dynamic quantization device may use the AI quantization levels and / or operating conditions as inputs to an algorithm that may output threshold weight values for masking weight values and bypassing the entire MAC (e.g., 202a-202i). In some embodiments, at optional block 908, a dynamic quantization controller may be configured to determine threshold weight values for dynamic network pruning. In some embodiments, in optional block 908, the I / O interface and / or memory controller / physical layer component may be configured to determine a threshold weight value for dynamic network pruning.
[0110] The AI quantization level used in block 904, optional block 906, and / or optional block 906 may be different from a previously calculated AI quantization level, which may result in differences in the determined parameters for performing the dynamic neural network quantization reconfiguration. For example, increasing the AI quantization level may cause the dynamic quantization device to determine a larger number of dynamic bits and / or a smaller threshold weight value for performing the dynamic neural network quantization reconfiguration. Increasing the number of dynamic bits and / or decreasing the threshold weight value may result in fewer bits and / or fewer MACs being used to perform the neural network calculations, which may reduce the accuracy of the neural network inference results. As another example, decreasing the AI quantization level may cause the dynamic quantization device to determine a smaller number of dynamic bits and / or a larger threshold weight value for performing the dynamic neural network quantization reconfiguration. Reducing the number of dynamic bits and / or increasing the threshold weight value may result in a larger number of bits and / or a larger number of MACs being used to perform the neural network calculations, which may increase the accuracy of the neural network inference results.
[0111] At block 910, the dynamic quantization device may generate and transmit a dynamic quantization signal. The dynamic quantization signal may include parameters for dynamic neural network quantization reconfiguration. The dynamic quantization device may transmit the dynamic quantization signal to dynamic neural network quantization logic (e.g., 212, 214). In some embodiments, the dynamic quantization device may transmit the dynamic quantization signal to an I / O interface and / or a memory controller / physical layer component. The dynamic quantization signal may cause a recipient to perform dynamic neural network quantization reconfiguration and provide parameters for performing the dynamic neural network quantization reconfiguration. In some embodiments, the dynamic quantization device may also transmit the dynamic quantization signal to a MAC array. The dynamic quantization signal may cause the MAC array to perform dynamic neural network quantization reconfiguration and provide parameters for performing the dynamic neural network quantization reconfiguration. In some embodiments, the dynamic quantization signal may include an indicator of the type of dynamic neural network quantization reconfiguration to implement. In some embodiments, the indicator of the type of dynamic neural network quantization reconfiguration may be parameters for the dynamic neural network quantization reconfiguration. In some embodiments, types of dynamic neural network quantization reconfiguration may include configuring a receiver for quantization of activation and weight values, configuring a receiver for masking of activation and weight values and configuring a MAC array and / or MAC for bypassing a portion of the MAC, and configuring a receiver for masking of weight values and configuring a MAC array and / or MAC for bypassing the entire MAC. In some embodiments, at block 910, a dynamic quantization controller may be configured to generate and transmit a dynamic quantization signal. In some embodiments, at block 910, an I / O interface and / or memory controller / physical layer component may be configured to generate and transmit the dynamic quantization signal.
[0112] FIG. 10 illustrates a method 1000 for dynamic neural network quantization architecture reconfiguration, according to one embodiment. 1-10, method 1000 may be implemented in a computing device (e.g., 100), in general-purpose hardware, in dedicated hardware (e.g., dynamic neural network quantization logic 212, 214, MAC array 200, MACs 202a-202i), in software running on a processor (e.g., processor 104, AI processor 124, AI processing subsystem 300, AI processors 124a-124f, I / O interface 302, memory controller / physical layer components 304a-304f), or in a combination of software-configured processors and dedicated hardware, such as processors executing software within a dynamic neural network quantization system including other individual components, various memory / cache controllers (e.g., AI processor 124, AI processing subsystem 300, AI processors 124a-124f, I / O interface 302, memory controller / physical layer components 304a-304f). To encompass alternative configurations possible in various embodiments, the hardware that performs method 1000 is referred to herein as a “dynamic quantization configuration device.” In some embodiments, method 1000 may be performed following block 910 of method 900 (FIG. 9).
[0113] At block 1002, a dynamic quantization configuration device may receive a dynamic quantization signal. The dynamic quantization configuration device may receive the dynamic quantization signal from a dynamic quantization controller (e.g., dynamic quantization controller 208, I / O interface 302, memory controller / physical layer component 304a-304f). In some embodiments, at block 1002, dynamic neural network quantization logic may be configured to receive the dynamic quantization signal. In some embodiments, at block 1002, the I / O interface and / or the memory controller / physical layer component may be configured to receive the dynamic quantization signal. In some embodiments, at block 1002, the MAC array may be configured to receive the dynamic quantization signal.
[0114] In block 1004, the dynamic quantization configuration device may determine a number of dynamic bits for dynamic quantization. The dynamic quantization configuration device may determine parameters for dynamic neural network quantization reconfiguration. The dynamic quantization signal may include parameters for the number of dynamic bits for configuring dynamic neural network quantization logic (e.g., dynamic neural network quantization logic 212, 214, I / O interface 302, memory controller / physical layer components 304a-304f) for quantization of activation values and weight values. In some embodiments, in block 1004, the dynamic neural network quantization logic may be configured to determine the number of dynamic bits for dynamic quantization. In some embodiments, in block 1004, the I / O interface and / or memory controller / physical layer components may be configured to determine the number of dynamic bits for dynamic quantization.
[0115] In block 1006, the dynamic quantization configuration device may configure the dynamic neural network quantization logic to quantize the activation values and weight values to the number of dynamic bits. The dynamic neural network quantization logic may be configured to quantize the activation values and weight values by rounding the bits of the activation values and weight values to the number of dynamic bits indicated by the dynamic quantization signal. The dynamic neural network quantization logic may include configurable logic gates and / or software that may be configured to round the bits of the activation values and weight values to the number of dynamic bits. In some embodiments, the logic gates and / or software may be configured to output values of zeros for the least significant bits of the activation values and weight values up to and / or including the number of dynamic bits. In some embodiments, the logic gates and / or software may be configured to output values for the most significant bits of the activation values and weight values up to and / or following the number of dynamic bits. For example, each bit of the activation values or weight values may be input to the logic gates and / or software sequentially, e.g., from the least significant bit to the most significant bit. The logic gates and / or software may output values of 0 for the lower-order bits of the activation and weight values up to and / or including the number of dynamic bits indicated by the parameter. The logic gates and / or software may output values for the upper-order bits of the activation and weight values up to and / or following the number of dynamic bits indicated by the parameter. The number of dynamic bits may be different from a default number of dynamic bits or a previous number of dynamic bits to round to for a default or previous configuration of the dynamic neural network quantization logic. Thus, the configuration of the logic gates may also be different from a default or previous configuration of the logic gates and / or software. In some embodiments, at block 1006, the dynamic neural network quantization logic may be configured to configure the dynamic neural network quantization logic to quantize the activation and weight values to the number of dynamic bits.In some embodiments, at block 1006, the I / O interface and / or memory controller / physical layer component may be configured to configure the dynamic neural network quantization logic to quantize the activation and weight values to that number of dynamic bits.
[0116] In optional decision block 1008, the dynamic quantization configuration device may determine whether to configure the quantization logic for masking and bypassing. The dynamic quantization signal may include a number of dynamic bits of parameters for configuring the dynamic neural network quantization logic for masking activation values and weight values and bypassing a portion of the MAC. The dynamic quantization configuration device may make the determination from the presence of values of the parameters for configuring the quantization logic for masking and bypassing. In some embodiments, in optional decision block 1008, the dynamic neural network quantization logic may be configured to determine whether to configure the quantization logic for masking and bypassing. In some embodiments, in optional decision block 1008, an I / O interface and / or memory controller / physical layer component may be configured to determine whether to configure the quantization logic for masking and bypassing. In some embodiments, in optional decision block 1008, the MAC array may be configured to determine whether to configure the quantization logic for masking and bypassing.
[0117] In response to determining to configure the quantization logic for masking and bypassing (i.e., optional decision block 1008="Yes"), in optional block 1010, the dynamic quantization configuration device may determine a number of dynamic bits for masking and bypassing. As described above, the dynamic quantization signal may include a number of dynamic bits parameter for configuring the dynamic neural network quantization logic (e.g., dynamic neural network quantization logic 212, 214, MAC array 200, I / O interface 302, memory controller / physical layer components 304a-304f) for masking activation values and weight values and bypassing portions of the MAC. The dynamic quantization configuration device may extract the number of dynamic bits for masking and bypassing from the dynamic quantization signal. In some embodiments, in optional block 1010, the dynamic neural network quantization logic may be configured to determine a number of dynamic bits for masking and bypassing. In some embodiments, an I / O interface and / or memory controller / physical layer component may be configured to determine a number of dynamic bits for masking and bypassing in optional block 1010. In some embodiments, a MAC array may be configured to determine a number of dynamic bits for masking and bypassing in optional decision block 1010.
[0118] In optional block 1012, the dynamic quantization configuration device may configure the dynamic quantization logic to mask a number of dynamic bits of the activation and weight values. The dynamic neural network quantization logic may be configured to quantize the activation and weight values by masking the number of dynamic bits of the activation and weight values indicated by the dynamic quantization signal.
[0119] The dynamic neural network quantization logic may include configurable logic gates and / or software that may be configured to mask the number of dynamic bits of the activation values and weight values. In some embodiments, the logic gates and / or software may be configured to output zero values for the lower-order bits of the activation values and weight values up to and / or including the number of dynamic bits. In some embodiments, the logic gates and / or software may be configured to output values for the upper-order bits of the activation values and weight values up to and / or including the number of dynamic bits. For example, each bit of the activation values and weight values may be input to the logic gates and / or software sequentially, e.g., from least significant bit to most significant bit. The logic gates and / or software may output zero values for the lower-order bits of the activation values and weight values up to and / or including the number of dynamic bits indicated by the parameter. The logic gates and / or software may output values for the upper-order bits of the activation values and weight values up to and / or including the number of dynamic bits indicated by the parameter. The number of dynamic bits may be different from the default number of dynamic bits or the previous number of dynamic bits to be masked for the default or previous configuration of the dynamic neural network quantization logic. Thus, the configuration of the logic gates and / or software may also be different from the default or previous configuration of the logic gates.
[0120] In some embodiments, the logic gates may be clock-gated to not receive and / or output lower-order bits of the activation and weight values up to and / or including the number of dynamic bits. Clock-gating the logic gates may effectively replace those lower-order bits of the activation and weight values with values of 0, since the MAC array may not receive values for those lower-order bits of the activation and weight values. In some embodiments, in optional block 1012, the dynamic neural network quantization logic may be configured to configure the dynamic quantization logic to mask a certain number of dynamic bits of the activation and weight values. In some embodiments, in optional block 1012, the I / O interface and / or memory controller / physical layer component may be configured to configure the dynamic quantization logic to mask a certain number of dynamic bits of the activation and weight values.
[0121] In optional block 1014, a dynamic quantization configuration device may configure the AI processor to clock gate and / or power down the MAC for bypass. In some embodiments, the dynamic neural network quantization logic may signal to the AI processor's MAC array a parameter for the number of dynamic bits for bypassing a portion of the MAC. In some embodiments, the dynamic neural network quantization logic may signal to the MAC array which bits of the activation values and weight values are to be masked. In some embodiments, the absence of a signal for a bit of the activation values and weight values may be a signal from the dynamic neural network quantization logic to the MAC array. The MAC array may receive a dynamic quantization signal including a parameter for the number of dynamic bits for configuring the dynamic neural network quantization logic for masking the activation values and weight values and bypassing a portion of the MAC. In some embodiments, the MAC array 200 may receive a signal from the dynamic neural network quantization logic of the parameter for the number of dynamic bits and / or which dynamic bits for bypassing a portion of the MAC. The MAC array may be configured to bypass a portion of the MAC for dynamic bits of the activation values and weight values indicated by the dynamic quantization signal and / or the signal from the dynamic neural network quantization logic. These dynamic bits may correspond to bits of the activation values and weight values that are masked by the dynamic neural network quantization logic. The MAC may include logic gates configured to implement multiplication and accumulation functions.
[0122] In some embodiments, the MAC array may clock gate logic gates in the MAC configured to multiply and accumulate activation and weight value bits corresponding to the number of dynamic bits indicated by the parameters of the dynamic quantization signal. In some embodiments, the MAC array may clock gate logic gates in the MAC configured to multiply and accumulate activation and weight value bits corresponding to the number of dynamic bits and / or dynamic most significant bits indicated by the signal from the dynamic neural network quantization logic.
[0123] In some embodiments, the MAC array may power down logic gates in the MAC configured to multiply and accumulate activation and weight value bits corresponding to the number of dynamic bits indicated by the parameters of the dynamic quantization signal. In some embodiments, the MAC array may power down logic gates in the MAC configured to multiply and accumulate activation and weight value bits corresponding to the number of dynamic bits and / or particular dynamic bits indicated by the signal from the dynamic neural network quantization logic.
[0124] The MAC may not receive activation and weight value bits corresponding to the number or particular dynamic bits, effectively masking these bits, by clock gating and / or powering down logic gates of the MAC in optional block 1014. In some embodiments, in optional block 1014, the MAC array may be configured to configure the AI processor to clock gate and / or power down the MAC for bypass.
[0125] In some embodiments, following configuring the dynamic neural network quantization logic to quantize activation and weight values to that number of dynamic bits in block 1006, the dynamic quantization configuration device may determine, in optional decision block 1016, whether to configure the quantization logic for dynamic network pruning. In some embodiments, in response to determining not to configure the quantization logic for masking and bypassing (i.e., optional decision block 1018="No"), or following configuring the AI processor to clock gate and / or power down the MAC for bypassing in optional block 1014, the dynamic quantization configuration device may determine, in optional decision block 1016, whether to configure the quantization logic for dynamic network pruning. The dynamic quantization signal may include threshold weight value parameters for configuring the dynamic neural network quantization logic for masking weight values and bypassing the entire MAC. The dynamic quantization configuration device may make the determination from the presence of values of the parameters for configuring the quantization logic for dynamic network pruning. In some embodiments, at optional decision block 1016, the dynamic neural network quantization logic may be configured to determine whether to configure the quantization logic for dynamic network pruning. In some embodiments, at optional decision block 1016, an I / O interface and / or memory controller / physical layer component may be configured to determine whether to configure the quantization logic for dynamic network pruning. In some embodiments, at optional decision block 1016, the MAC array may be configured to determine whether to configure the quantization logic for dynamic network pruning.
[0126] In response to determining to configure the quantization logic for dynamic network pruning (i.e., optional decision block 1016="Yes"), in optional block 1018, the dynamic quantization configuration device may determine threshold weight values for dynamic network pruning. As described above, the dynamic quantization signal may include threshold weight value parameters for configuring the dynamic neural network quantization logic (e.g., dynamic neural network quantization logic 212, 214, MAC array 200, I / O interface 302, memory controller / physical layer components 304a-304f) for overall weight value masking and overall MAC bypass. The dynamic quantization configuration device may retrieve the threshold weight values for masking and bypassing from the dynamic quantization signal. In some embodiments, in optional block 1018, the dynamic neural network quantization logic may be configured to determine threshold weight values for dynamic network pruning. In some embodiments, the I / O interface and / or memory controller / physical layer component may be configured to determine a threshold weight value for dynamic network pruning in optional block 1018. In some embodiments, in optional block 1018, the MAC array may be configured to determine a threshold weight value for dynamic network pruning.
[0127] In optional block 1020, the dynamic quantization configuration device may configure the dynamic quantization logic to mask the entire weight value. The dynamic neural network quantization logic may be configured to quantize the weight value by masking all of the bits of the weight value based on a comparison of the weight value with a threshold weight value indicated by the dynamic quantization signal. The dynamic neural network quantization logic may include configurable logic gates and / or software that may be configured to compare weight values received from a data source (e.g., weight buffer 204) with threshold weight values and mask weight values that result in an unfavorable comparison, such as being less than or equal to the threshold weight value. In some embodiments, the comparison may be a comparison of the absolute value of the weight value against the threshold weight value. In some embodiments, the logic gates and / or software may be configured to output a value of 0 for all of the bits of the weight value that result in an unfavorable comparison with the threshold weight value. The number of bits may be a number different from a default number of bits or a previous number of bits to be masked for a default or previous configuration of the dynamic neural network quantization logic. Thus, the configuration of the logic gates and / or software may also differ from the default or previous configuration of the logic gates. In some embodiments, the logic gates may be clock gated to prevent them from receiving and / or outputting bits of weight values that result in an unfavorable comparison with the threshold weight values. Clock gating the logic gates may effectively replace bits of weight values with a value of 0, since the MAC array may not receive values for the bits of weight values. In some embodiments, in optional block 1020, the dynamic neural network quantization logic may be configured to configure the dynamic quantization logic to mask entire weight values. In some embodiments, in optional block 1020, an I / O interface and / or memory controller / physical layer component may be configured to configure the dynamic quantization logic for entire weight values.
[0128] In optional block 1022, the dynamic quantization configuration device may configure the AI processor to clock gate and / or power down the entire MAC for dynamic network pruning. In some embodiments, the dynamic neural network quantization logic may signal to the MAC array of the AI processor which bits of the weight value are masked. In some embodiments, the absence of a signal for a bit of the weight value may be a signal from the dynamic neural network quantization logic to the MAC array. In some embodiments, the MAC array may receive a signal from the dynamic neural network quantization logic for which bits of the weight value are masked. The MAC array may interpret the entire masked weight value as a signal to bypass the entire MAC. The MAC array may be configured to bypass the MAC for weight values indicated by the signal from the dynamic neural network quantization logic. These weight values may correspond to the weight values masked by the dynamic neural network quantization logic. The MAC may include logic gates configured to implement multiplication and accumulation functions. In some embodiments, the MAC array may clock gate logic gates of the MAC configured to multiply and accumulate bits of the weight value corresponding to the masked weight values. In some embodiments, the MAC array may power down logic gates of the MAC configured to multiply and accumulate weight value bits corresponding to the masked weight values. By clock gating and / or powering down the logic gates of the MAC, the MAC does not receive activation values and weight value bits corresponding to the masked weight values. In some embodiments, in optional block 1022, the MAC array may be configured to configure the AI processor to clock gate and / or power down the MAC for dynamic network pruning.
[0129] Masking weight values by the dynamic neural network quantization logic in optional block 1020 and / or clock gating and / or powering down the MAC in optional block 1022 may prune the neural network executed by the MAC array. Removing weight values and MAC operations from the neural network may effectively remove synapses and nodes from the neural network. The weight threshold may be determined based on the fact that removing weight values that compare unfavorably to the weight threshold from the neural network execution may cause only an acceptable loss in accuracy of the AI processor's results.
[0130] In some embodiments, following configuring the dynamic neural network quantization logic to quantize activation and weight values to the number of dynamic bits in block 1006, a dynamic quantization configuration device may receive and process the activation and weight values in block 1024. In some embodiments, in response to determining not to configure the quantization logic for masking and bypassing (i.e., optional decision block 1018="No") or following configuring the AI processor to clock gate and / or power down the MAC for bypassing in optional block 1014, a dynamic quantization configuration device may receive and process the activation and weight values in block 1024. In some embodiments, in response to determining not to configure the quantization logic for dynamic network pruning (i.e., optional decision block 1016="No") or following configuring the AI processor to clock gate and / or power down the MAC for dynamic network pruning in optional block 1022, a dynamic quantization configuration device may receive and process the activation and weight values in block 1024. The dynamic quantization configuration device may receive activation values and weight values from a data source (e.g., processor 104, communication component 112, memory 106, 114, peripheral device 122, weight buffer 204, activation buffer 206, memory 106). The quantization configuration device may quantize and / or mask the activation values and / or weight values. The quantization device may bypass, clock gate, and / or power down portions of the MAC and / or the entire MAC. In some embodiments, at block 1024, dynamic neural network quantization logic may be configured to receive and process the activation values and weight values. In some embodiments, at block 1024, an I / O interface and / or a memory controller / physical layer component may be configured to receive and process the activation values and weight values.In some embodiments, at block 1024, a MAC array may be configured to receive and process the activation values and weight values.
[0131] AI processors according to various embodiments (including, but not limited to, those described above with reference to FIGS. 1-10 ) may be implemented in a wide variety of computing systems, including mobile computing devices, an example of which is suitable for use with various embodiments is shown in FIG. 11 . Mobile computing device 1100 may include a processor 1102 coupled to a touchscreen controller 1104 and internal memory 1106. Processor 1102 may be one or more multi-core integrated circuits, either general-purpose or designated for specific processing tasks. Internal memory 1106 may be volatile or non-volatile memory, and may be secure and / or encrypted, or non-secure and / or non-encrypted, or any combination thereof. Examples of memory types that may be utilized include, but are not limited to, DDR, LPDDR, GDDR, WIDEIO, RAM, SRAM, DRAM, P-RAM, R-RAM, M-RAM, STT-RAM, and embedded DRAM. The touchscreen controller 1104 and processor 1102 may be coupled to a touchscreen panel 1112, such as a resistive-sensing touchscreen, a capacitive-sensing touchscreen, an infrared-sensing touchscreen, etc. Additionally, the display of the mobile computing device 1100 need not have touchscreen capabilities.
[0132] The mobile computing device 1100 may have one or more wireless signal transceivers 1108 (e.g., Peanut, Bluetooth, ZigBee, Wi-Fi, RF radio) and an antenna 1110 coupled to each other and / or to the processor 1102 for transmitting and receiving communications. The transceiver 1108 and antenna 1110 may be used with the circuitry described above to implement various wireless transmission protocol stacks and interfaces. The mobile computing device 1100 may include a cellular network wireless modem chip 1116 that enables communication over a cellular network and is coupled to the processor.
[0133] The mobile computing device 1100 may include a peripheral device connection interface 1118 coupled to the processor 1102. The peripheral device connection interface 1118 may be configured solely to accept one type of connection or may be configured to accept various types of physical and communication connections, either common or proprietary, such as Universal Serial Bus (USB), FireWire, Thunderbolt, or PCIe. The peripheral device connection interface 1118 may also be coupled to a similarly configured peripheral device connection port (not shown).
[0134] The mobile computing device 1100 may also include a speaker 1114 for providing audio output. The mobile computing device 1100 may also include a housing 1120 constructed of plastic, metal, or a combination of materials for enclosing all or a portion of the components described herein. The mobile computing device 1100 may also include a power source 1122, such as a disposable or rechargeable battery, coupled to the processor 1102. The rechargeable battery may also be coupled to a peripheral device connection port to receive charging current from a power source external to the mobile computing device 1100. The mobile computing device 1100 may also include a physical button 1124 for receiving user input. The mobile computing device 1100 may also include a power button 1126 for turning the mobile computing device 1100 on and off.
[0135] AI processors according to various embodiments (including, but not limited to, those described above with reference to FIGS. 1-10 ) may be implemented in a wide variety of computing systems, including a laptop computer 1200, an example of which is shown in FIG. 12 . Many laptop computers include a touchpad touch surface 1217 that acts as the computer's pointing device and may therefore receive drag, scroll, and flick gestures similar to those implemented on the computing devices described above with touchscreen displays. Laptop computers 1200 typically include a processor 1202 coupled to volatile memory 1212 and a large-capacity nonvolatile memory, such as a flash memory disk drive 1213. Additionally, computer 1200 may have one or more antennas 1215 for transmitting and receiving electromagnetic radiation, which may be connected to a wireless data link and / or a cellular telephone transceiver 1216 coupled to processor 1202. The computer 1200 may also include a floppy disk drive 1214 and a compact disk (CD) drive 1215, coupled to the processor 1202. In a notebook configuration, the computer housing includes a touchpad 1217, a keyboard 1218, and a display 1219, all coupled to the processor 1202. Other configurations of computing devices may include a computer mouse or trackball coupled to the processor (e.g., via a USB input), as is well known, and may also be used with various embodiments.
[0136] An AI processor according to various embodiments (including, but not limited to, the embodiments described above with reference to FIGS. 1-10 ) may be implemented in a fixed computing system, such as any of a variety of commercially available servers. An exemplary server 1300 is shown in FIG. 13 . Such a server 1300 typically includes one or more multi-core processor assemblies 1301 coupled to volatile memory 1302 and large-capacity non-volatile memory, such as a disk drive 1304. As shown in FIG. 13 , multi-core processor assemblies 1301 may be added to the server 1300 by inserting them into a rack of assemblies. The server 1300 may also include a floppy disk drive, compact disk (CD), or digital versatile disk (DVD) disk drive 1306 coupled to the processor 1301. The server 1300 may also include a network access port 1303 coupled to the multi-core processor assembly 1301 for establishing a network interface connection with a network 1305, such as a local area network, the Internet, a public switched telephone network, and / or a cellular data network (e.g., CDMA, TDMA, GSM, PCS, 3G, 4G, LTE, or any other type of cellular data network) coupled to other broadcast system computers and servers.
[0137] Example implementations are described in the following paragraphs. While some of the following example implementations are described with reference to example methods, further example implementations may include the example methods discussed in the following paragraphs implemented by an AI processor comprising a dynamic quantization controller and a MAC array configured to perform the operations of the example methods, a computing device comprising an AI processor comprising a dynamic quantization controller and a MAC array configured to perform the operations of the example methods, and the example methods discussed in the following paragraphs implemented by an AI processor including means for performing the functions of the example methods.
[0138] Example 1. A method for processing a neural network by an artificial intelligence (AI) processor, the method comprising: receiving AI processor operating condition information; dynamically adjusting AI quantization levels for a segment of the neural network in response to the operating condition information; and processing the segment of the neural network using the adjusted AI quantization levels.
[0139] Example 2. The method of Example 1, wherein dynamically adjusting the AI quantization level for the segment of the neural network includes increasing the AI quantization level in response to operating condition information indicating a level of operating conditions that increased a processing power constraint of the AI processor, and decreasing the AI quantization level in response to operating condition information indicating a level of operating conditions that decreased a processing power constraint of the AI processor.
[0140] Example 3. The method of any of Examples 1 or 2, wherein the operating condition information is at least one of the following group: temperature, power consumption, operating frequency, or processing unit utilization.
[0141] Example 4. The method of any of Examples 1 to 3, wherein dynamically adjusting an AI quantization level for a segment of the neural network includes adjusting an AI quantization level for quantizing weight values to be processed by the segment of the neural network.
[0142] Example 5. The method of any of Examples 1 to 3, wherein dynamically adjusting an AI quantization level for a segment of the neural network includes adjusting an AI quantization level for quantizing activation values to be processed by the segment of the neural network.
[0143] Example 6. The method of any of Examples 1 to 3, wherein dynamically adjusting AI quantization levels for segments of the neural network includes adjusting AI quantization levels for quantizing weight values and activation values to be processed by the segments of the neural network.
[0144] Example 7. The method of any of Examples 1 to 6, wherein the AI quantization levels are configured to indicate dynamic bits of a value to be quantized and processed by the neural network, and wherein processing the segment of the neural network using the adjusted AI quantization levels includes bypassing a portion of a multiply-accumulate unit (MAC) associated with the dynamic bits of the value.
[0145] Example 8. The method of any of Examples 1 to 7, further comprising determining an AI Quality of Service (QoS) value using an AI QoS factor, and determining an AI quantization level to achieve the AI QoS value.
[0146] Example 9. The method of example 8, wherein the AI QoS values represent goals for the accuracy of results produced by the AI processor and the throughput of the AI processor.
[0147] Computer program code or "program code" for execution on a programmable processor to perform operations of various embodiments may be written in a high-level programming language such as C, C++, C#, Smalltalk, Java, JavaScript, Visual Basic, Structured Query Language (e.g., Transact-SQL), Perl, or a variety of other programming languages. Program code or programs stored on a computer-readable storage medium as used in this application may refer to machine code (such as object code) whose format is understandable by a processor.
[0148] The above method descriptions and process flow diagrams are provided merely as illustrative examples and do not require or imply that the operations of the various embodiments must be performed in the order presented. As will be understood by one of ordinary skill in the art, the order of operations in the above-described embodiments may be performed in any order. Words such as "then," "then," and "next" do not limit the order of operations; these words are merely used to guide the reader through the method descriptions. Furthermore, any reference to a claim element in the singular, for example, using the article "a," "an," or "the," should not be construed as limiting the element to the singular.
[0149] The various illustrative logical blocks, modules, circuits, and algorithmic operations described in connection with the various embodiments may be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and operations have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the claims.
[0150] The hardware used to implement the various exemplary logic, logic blocks, modules, and circuits described in connection with the embodiments disclosed herein may be implemented or performed using general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Alternatively, some operations or methods may be performed by circuitry specific to a given function.
[0151] In one or more embodiments, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored as one or more instructions or code on a non-transitory computer-readable or processor-readable medium. The operations of a method or algorithm disclosed herein may be embodied in a processor-executable software module, which may reside on a non-transitory computer-readable or processor-readable storage medium. A non-transitory computer-readable or processor-readable storage medium may be any storage medium that can be accessed by a computer or processor. By way of example and not limitation, such non-transitory computer-readable or processor-readable medium may include RAM, ROM, EEPROM, FLASH memory, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. As used herein, disk and disc include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically and discs reproduce data optically using lasers. Combinations of the above are also included within the scope of non-transitory computer-readable medium and non-transitory processor-readable medium. Additionally, the operations of a method or algorithm may reside as one, or any combination or set of code and / or instructions on a non-transitory processor-readable medium and / or a non-transitory computer-readable medium, which may be embodied in a computer program product.
[0152] The foregoing description of the disclosed embodiments is provided to enable any person skilled in the art to make or use the claims. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments and implementations without departing from the scope of the claims. Thus, the present disclosure is not intended to be limited to the embodiments and implementations described herein, but is to be accorded the widest scope consistent with the following claims, the principles and novel features disclosed herein. [Explanation of symbols]
[0153] 100 computing devices 102 SoC 104 processors 106 memory 108 Communication Interface 110 Memory Interface 112 Communication Components 114 memory 116 Antenna 120 Peripheral Device Interface 122 Peripheral Devices 124 AI processors 200 MAC array 202 MAC 202a-202i MAC 204 Weight Buffer 206 Activation Buffer 208 Dynamic Quantization Controller 210 AI QoS Manager 212 Dynamic Neural Network Quantization Logic 214 Dynamic Neural Network Quantization Logic 300 AI Processing Subsystem 302 I / O interface 304a-304f Memory Controller / Physical Layer 700 Logical Components 702 Logical Components 1100 Mobile Computing Devices 1102 processor 1104 Touchscreen Controller 1106 Internal Memory 1108 Radio Signal Transceiver 1110 Antenna 1112 Touch Screen Panel 1114 Speaker 1116 Cellular Network Wireless Modem Chip 1118 Peripheral device connection interface 1120 Housing 1122 Power supply 1124 Physical Buttons 1126 Power button 1200 laptop computer 1202 processor 1212 Volatile Memory 1213 disk drive 1214 Floppy disk drive 1215 Antenna 1216 Cellular Telephone Transceiver 1217 Touchpad Touch Surface 1218 keyboard 1219 Display 1300 Server 1301 Multi-core processor assembly, processor 1302 Volatile Memory 1303 Network Access Port 1304 Disk Drive 1305 Network 1306 Compact Disc (CD) or Digital Versatile (DVD) disk drive
Claims
1. 1. A method for processing a neural network, implemented by an artificial intelligence (AI) processor, comprising: receiving AI processor operating condition information; dynamically adjusting AI quantization levels for segments of the neural network in response to the operating condition information; processing the segment of the neural network using the adjusted AI quantization levels; Equipped with dynamically adjusting the AI quantization levels for the segments of the neural network; increasing the AI quantization level in response to the operating condition information indicating a level of operating conditions that increases constraints on the processing power of the AI processor; reducing the AI quantization level in response to operating condition information indicating a level of the operating condition that reduces the constraint on the processing capability of the AI processor; Equipped with the processing power constraint of the AI processor relates to the ability of the AI processor to maintain a level of processing power; The operating condition information is at least one of the group consisting of temperature, power consumption, operating frequency, or utilization rate of a processing unit; The AI quantization level is configured to indicate the dynamic bits of a value to be processed by the neural network that should be quantized; processing the segment of the neural network using the adjusted AI quantization level comprises bypassing a portion of a multiply-accumulate (MAC) unit associated with the dynamic bit of the value by clock gating and / or powering down logic components for multiplication with multiplication according to the dynamic bit. method.
2. 2. The method of claim 1 , wherein dynamically adjusting the AI quantization level for the segment of the neural network comprises adjusting the AI quantization level for quantizing weight values to be processed by the segment of the neural network.
3. 2. The method of claim 1 , wherein dynamically adjusting the AI quantization level for the segment of the neural network comprises adjusting the AI quantization level for quantizing activation values to be processed by the segment of the neural network.
4. 2. The method of claim 1 , wherein dynamically adjusting the AI quantization level for the segment of the neural network comprises adjusting the AI quantization level for quantizing weights and activation values to be processed by the segment of the neural network.
5. determining one or more AI Quality of Service (QoS) values using AI QoS factors, the QoS values being at least one of latency, quality, accuracy, result accuracy, and responsiveness of the AI processor, and the QoS factors being operating conditions of the AI processor; determining the AI quantization level to achieve the AI QoS value; The method of claim 1 further comprising:
6. The method of claim 5 , wherein the one or more AI QoS values are a plurality of QoS values representing goals for accuracy of results produced by the AI processor and throughput of the AI processor.
7. An artificial intelligence (AI) processor, Receives AI processor operating condition information; Dynamically adjusting AI quantization levels for segments of the neural network in response to the operating condition information. a dynamic quantization controller configured to a multiply-accumulate (MAC) array configured to process the segment of the neural network using the adjusted AI quantization levels; and Equipped with the dynamic quantization controller dynamically adjusting the AI quantization levels for the segments of the neural network; increasing the AI quantization level in response to the operating condition information indicating a level of operating conditions that increases constraints on the processing power of the AI processor; reducing the AI quantization level in response to operating condition information indicating a level of the operating condition that reduces the constraint on the processing capability of the AI processor; configured to include the processing power constraint of the AI processor relates to the ability of the AI processor to maintain a level of processing power; the dynamic quantization controller is configured such that the operating condition information is at least one of the group of temperature, power consumption, operating frequency, or utilization of a processing unit; The AI quantization level is configured to indicate the dynamic bits of a value to be processed by the neural network that should be quantized; an AI processor, wherein the MAC array is configured such that processing the segment of the neural network using the adjusted AI quantization level comprises bypassing a MAC of the MAC array associated with the dynamic bit of the value by clock gating and / or powering down logic components for multiplication according to the dynamic bit, and wherein the MAC is provided by the AI processor.
8. 8. The AI processor of claim 7, wherein the dynamic quantization controller is configured such that dynamically adjusting the AI quantization level for the segment of the neural network comprises adjusting the AI quantization level for quantizing weight values to be processed by the segment of the neural network.
9. 8. The AI processor of claim 7, wherein the dynamic quantization controller is configured such that dynamically adjusting the AI quantization level for the segment of the neural network comprises adjusting the AI quantization level for quantizing activation values to be processed by the segment of the neural network.
10. 8. The AI processor of claim 7, wherein the dynamic quantization controller is configured such that dynamically adjusting the AI quantization level for the segment of the neural network comprises adjusting the AI quantization level for quantizing weights and activation values to be processed by the segment of the neural network.
11. determining an AI quality of service (QoS) value using an AI QoS factor in response to determining to dynamically configure the neural network quantization; determining the AI quantization level to achieve the AI QoS value; The AI processor of claim 7, further comprising an AI QoS device configured to:
12. The AI processor of claim 11 , wherein the AI QoS device is configured such that the AI QoS value represents a goal for accuracy of results produced by the AI processor and a goal for throughput of the AI processor.
13. A computing device comprising an artificial intelligence (AI) processor according to any one of claims 7 to 12.
Citation Information
Patent Citations
Reduced computational complexity for fixed point neural network
US20160328645A1
Layer-level quantization in neural networks
US20190171927A1
Optimizing execution of a neural network based on operational performance parameters
US20210081789A1