Accelerating artificial neural network using hardware-implemented lookup table
Through hardware system design, the hardware-implemented lookup table and cross-switch array structure are used to solve the problem of low computational efficiency of matrix-vector multiplication and mathematical function in traditional computer architecture, and realize efficient, low power consumption and flexible mathematical function execution of neural network calculations.
Patent Information
- Application Number
- CN202280102172.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-27
- Publication Date
- 2025-07-08
AI Technical Summary
The existing computer architectures have problems of low efficiency and high power consumption when performing matrix-vector multiplication and mathematical function calculations in artificial neural networks. In particular, the execution of activation functions is cumbersome and expensive on traditional hardware platforms.
The hardware system design is adopted, including neural processing equipment, lookup table circuits and processing units. The search table (LUT) is used to store parameter values through hardware implementation, and the mathematical functions of the neuron output are directly calculated, avoiding the separation of memory and computing units in traditional architectures. The cross-switch array structure is used to fuse arithmetic and memory units to realize near-memory processing.
It significantly accelerates the calculation of mathematical functions, reduces storage requirements, improves computing efficiency, and supports the dynamic configuration and integration of multiple mathematical functions, suitable for various neural network architectures.
Smart Images

Figure CN120283242A_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to the field of in-memory and near-memory processing techniques (i.e., methods, devices, and systems), as well as related acceleration techniques for performing artificial neural networks (ANNs). Specifically, the present invention relates to hardware systems, including neural processing devices (e.g., having a crossbar array structure), implementing neurons, processing units, and hardware-implemented look-up tables (LUTs) that store parameter values that can be quickly accessed by the processing units to more effectively apply mathematical functions (such as activation functions) to neuron outputs. Background Art
[0002] ANNs (such as deep neural networks (DNNs)) have revolutionized the field of machine learning by providing unprecedented performance in solving cognitive data analysis tasks. ANN operations typically involve matrix-vector multiplication (MVM). MVM operations pose multiple challenges due to their repetitive nature, generality, computational, and storage requirements. Traditional computer architectures are based on the von Neumann computing concept, according to which processing power and data storage are split into separate physical units. Such architectures suffer from congestion and high power consumption because data must constantly be transferred from the storage unit through a physically limited and expensive interface to the control and arithmetic units.
[0003] One possibility for accelerating MVM is to use dedicated hardware acceleration devices, such as dedicated circuits having a crossbar array structure. This type of circuit includes input lines and output lines that are interconnected at the intersection points defining the cells. These cells contain corresponding storage means (or groups of storage means) for storing the corresponding matrix coefficients. The vector is encoded as a signal and applied to the input lines of the crossbar array to perform MVM through multiply-accumulate (MAC) operations. Such an architecture can simply and effectively map MVM. The weights can be updated as needed by reprogramming the storage elements to perform successive matrix-vector multiplications. This solution breaks the "memory wall" because it fuses the arithmetic and storage units into a single in-memory-computing (IMC) unit, thus performing processing more effectively in or near the memory (i.e., the crossbar array).
[0004] Although the main computational load of ANNs such as DNNs revolves around MAC operations, the execution of ANNs typically involves additional mathematical functions, such as activation functions. Even in quantized neural networks, activation functions are required, which are inherently more difficult to compress and generally need to be executed with floating-point precision.
[0005] The inventors of the present case have concluded that executing these functions on a hardware platform designed for the efficient execution and low power consumption of DNNs can become cumbersome and expensive in terms of computing resources. One possible solution is to offload the execution of these functions to a digital signal processor (DSP). However, doing so can be very demanding in terms of latency, area, and energy.
[0006] Therefore, the inventors of the present case have taken up the challenge and are committed to implementing a new computing architecture involving non - traditional processing means to accelerate the computation of these functions. Summary of the Invention
[0007] According to a first aspect, the present invention is implemented as a hardware system designed to implement an artificial neural network (ANN). The hardware system mainly includes a neural processing device, one or more lookup table circuits, and one or more processing units. The neural processing device is configured to implement M artificial neurons, where M≥1. One or more lookup table circuits are configured to implement a lookup table (LUT). The system further includes M' processing units, where M≥M'≥1. Each of the M' processing units is connected through at least one of the M artificial neurons so as to be able to access the value (referred to as the "first value") output by each of the at least one neuron during operation. Additionally, each processing unit is connected to the LUT circuit in one or more LUT circuits so as to be able to access the parameter values of a parameter set from the LUT during operation. Finally, each processing unit is configured to output the value (second value) of a mathematical function, where the mathematical function takes the first value as an independent variable. The mathematical function is also determined by the parameter set. During operation, each of the processing units accesses the parameter values of the parameter set from the LUT circuit.
[0008] The architecture of this hardware system is different from traditional computer architectures, where the same digital processor (or the same group of digital processors) is typically used to compute neuron output values and apply subsequent mathematical functions (e.g., activation functions). Instead, here, the processing hardware used to compute neuron output values is different from the processing unit used to apply the mathematical functions, although the processing unit in the system can well be configured as a near-memory processing device. This architecture is adopted for reasons of computational efficiency. In particular, since the hardware circuits are different for each of the neural processing device (for computing neuron output) and the processing unit (for applying mathematical functions), the LUT is implemented in hardware. Thanks to the hardware-implemented LUT, a significant acceleration is achieved. That is, the mathematical function is defined (and thus determined) by a set of parameters, the values of which can be efficiently retrieved from the hardware-implemented LUT. This results in a significant acceleration of the computation of the function output, exceeding the acceleration that might have been achieved within the neural processing device and the processing unit. Thus, the neuron output can be processed more efficiently before being passed to the next neuron layer.
[0009] In addition, since the LUT stores parameter values instead of mapping input values to output values, little memory is required; the LUT is not used to directly look up the function output, contrary to what is typically done when using a lookup table.
[0010] Finally, the present solution is compatible with integration. Specifically, the LUT circuit, the processing unit, and the neural processing device can be advantageously co-integrated in the same device, e.g., on the same chip.
[0011] In an embodiment, each processing unit is configured to output a second value by: (i) selecting the set of parameters according to a first value; (ii) performing an operation based on the first value and the parameter values of the selected set of parameters so as to output the second value. In effect, this makes it possible to reduce the number of parameters required, since a small set of parameters is sufficient to locally accurately estimate the function over an interval containing each potential input value.
[0012] Preferably, each processing unit is further configured to select the set of parameters by comparing the first value with bin boundaries to identify the relevant bin (i.e., the bin containing the first value). For this purpose, each processing unit is further configured to access the bin boundaries from the lookup table circuit. Subsequently, the set of parameters is selected according to the identified bin during operation. Thus, the bin boundaries can be accessed efficiently to enable fast comparison. Thus, the binning problem can be solved efficiently, and thus the relevant set of parameters can be quickly identified.
[0013] In a preferred embodiment, a dedicated comparator circuit is used to efficiently identify relevant bins. That is, each processing unit includes at least one comparator circuit. This circuit is designed to compare a first value with bin boundaries and transmit a selection signal encoding a selected parameter set. The processing unit can then access the corresponding parameter value based on the transmitted signal.
[0014] One can rely solely on a binary tree comparison circuit. However, more complex comparison schemes and comparator circuit layouts can be envisioned. The comparator circuit can in particular be designed to be capable of performing multi-level comparisons to accelerate binning. Specifically, the comparator circuit can advantageously be configured as a multi-level q-ary tree comparison circuit, which is designed to be capable of performing multi-level comparisons, where for one or more of the multi-levels, q is greater than or equal to three.
[0015] In an embodiment, each LUT circuit is a circuit that hard-codes parameter values. Additionally, each processing unit includes at least one multiplexer, which is connected on the one hand to the corresponding comparator circuit to receive a selection signal and on the other hand to the LUT circuit to retrieve the corresponding parameter value according to the selection signal. Such a design makes parameter retrieval extremely efficient. The drawback is that the hard-coded data cannot be changed after the LUT circuit is hard-wired.
[0016] Therefore, in a variant, a reconfigurable memory can preferably be used. Thus, if needed, the mathematical function can be dynamically reconfigured or updated as the computation progresses. For example, each LUT circuit can include addressable storage units, which are connected to the comparator circuit to receive a selection signal. In this way, the addressable storage units can retrieve the parameter values of the selected parameter set according to the received selection signal.
[0017] In a preferred embodiment, the mathematical function is a piecewise-defined polynomial function, which is a polynomial on each of its sub-domains. The sub-domains respectively correspond to the bins. In this case, the selected parameter set corresponds to the polynomial parameters of the piecewise-defined polynomial function. That is, the selected parameter set corresponds to the parameters of a locally relevant polynomial. Such a construction is very suitable for the fast computation of arithmetic units because simple arithmetic operations are required to achieve the desired result. Therefore, each processing unit can advantageously include an arithmetic unit, which is connected to the output of the LUT circuit, so that the arithmetic unit performs the operations required to compute a second value as arithmetic operations.
[0018] Interestingly, such an operation can be simply performed using a multiply-and-add circuit, that is, a circuit specifically designed to efficiently perform multiply-accumulate operations. Therefore, the arithmetic unit preferably includes a multiply-and-add circuit, which enables the output value of the mathematical function to be obtained more quickly.
[0019] In a preferred embodiment, the neural processing device includes a crossbar array structure, which includes N input lines and M output lines arranged in rows and columns, where N > 1 and M > 1, whereby the neural processing device can implement a layer of M neurons. The input lines and output lines are interconnected via storage elements. Each of the M output lines is connected to at least one of M' processing units. The crossbar array structure fuses arithmetic and storage units into a single in-memory computing unit, so that neuron outputs can be obtained efficiently.
[0020] Neural processing devices are typically designed to implement multiple neurons (M > 1) at a time. For example, the number of neurons can be greater than or equal to 256 or 512 (M ≥ 256 or M ≥ 512). Additionally, the processing units can advantageously be vector processing units, where each of the M' processing units is a vector processing unit including b processing elements, to be able to operate on a one-dimensional array of dimension b. The number of processing units M' is preferably equal to 1 or 2.
[0021] Various architectures can be considered. For example, multiple processing units (i.e., M' > 1) can be relied upon, although their number can generally be less than or equal to the number of neurons implemented at a time (i.e., M ≥ M' > 1). In this case, the LUT circuit can include M' different circuits, which are respectively mapped to the M' processing units.
[0022] According to another aspect, the invention is implemented as an operating method of the hardware system described above. That is, the provided system includes a neural processing device, which is configured to implement M artificial neurons, where M ≥ 1, and M' processing units, each of which is connected by at least one of the M artificial neurons. The hardware system also includes one or more LUT circuits implementing a LUT. The method includes operating the neural processing device to obtain M first values respectively generated by the M artificial neurons. Additionally, the method relies on the M' processing units to apply a mathematical function to the neuron outputs. That is, for each of the M first values, an output value of the mathematical function is obtained (via the M' processing units). The mathematical function is additionally determined by a parameter set. Therefore, the output value of the mathematical function is obtained based on an operand including the first value and the parameter set of parameter values, where the parameter values are retrieved from one or more LUT circuits.
[0023] Preferably, for each of the first values, the output value is obtained by selecting a parameter set according to the first value and performing an operation based on the first value and the parameter values retrieved according to the selected parameter set.
[0024] In a preferred embodiment, the parameter set is selected by comparing the first value with bin boundaries (retrieved from one or more LUT circuits) to identify the relevant bin containing the first value. Then the parameter set is selected according to the identified bin.
[0025] As described above, the mathematical function applied is preferably a piecewise-defined polynomial function. In this case, each parameter set includes two or more polynomial coefficients. The operations performed to calculate the second value can be only arithmetic operations. In a preferred embodiment, the mathematical function involves a set of linear polynomials, each polynomial corresponding to a respective bin. In this case, the parameter set corresponding to each of the linear polynomials consists of a scaling coefficient and an offset coefficient array. Also, due to the multiply-and-add circuitry, the arithmetic operations can be advantageously performed.
[0026] In an embodiment, the method further includes programming one or more LUT circuits implementing the LUT to enable one or more types of mathematical functions, such as activation functions, normalization functions, reduction functions, state update functions, classification functions, and / or prediction functions.
[0027] The method may also include upstream steps (i.e., steps performed at build time before operating the neural processing device) to determine one or more appropriate sets of bin boundaries respectively according to one or more reference functions (i.e., mathematical functions of potential interest in ANN execution). In an embodiment, bin boundaries are determined for each reference function to minimize the number of bins or the maximum error, where the error is measured as the difference between approximations of each reference function, the approximations being calculated based on the parameter values and theoretical values of the reference function. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] These and other objects, features, and advantages of the present invention will become apparent from the following detailed description of the embodiments thereof, which description should be read in conjunction with the Figure 1 accompanying drawings. The drawings are for clarity purposes to facilitate understanding of the present invention by those skilled in the art in conjunction with the detailed description. In the drawings:
[0029] Figure 1 Schematically represents a computer network involving multiple hardware systems according to an embodiment of the present invention. The network allows users to interact with the server to accelerate machine learning computing tasks, which are offloaded to the hardware systems as implemented;
[0030] Figure 2 Schematically represents selected elements of a hardware system according to an embodiment, which particularly includes a neural processing device having a crossbar switch array structure, a processing unit, and a hardware-implemented look-up table (LUT);
[0031] Figure 3 Is a diagram illustrating a possible architecture of a hardware system according to a preferred embodiment, which illustrates how neurons of a neural processing device are connected to a vector processing unit, and how the vector processing unit is connected to the LUT circuits;
[0032] Figure 4 is a circuit diagram depicting a given processing element, such as in the embodiments, connected to a corresponding LUT circuit (e.g., of a vector processing unit as shown). In this example, the processing element includes a comparator and a multiplexer, and the look-up table is implemented by a circuit that hard-codes the parameter values required to apply a mathematical function to the neuron output; Figure 3 is a variant of, where the LUT circuit is now implemented as an addressable memory (the multiplexer is not required in this example);
[0033] Figure 5 is Figure 4 is a flowchart showing the high-level steps of a method of operating a hardware system according to an embodiment as shown in
[0034] Figure 6 or Figure 2 or Figure 3 ;
[0035] Figure 7A , Figure 7B and Figure 7C are diagrams showing how a piecewise-defined polynomial function can be used to approximate a non-linear function, thanks to optimized bin boundaries, as in the embodiments; and
[0036] Figure 8 is a table showing the optimization of the number of comparators involved in each layer of a multi-layer q-ary tree comparison circuit used in the embodiments.
[0037] The accompanying drawings show a simplified representation of the apparatus or parts thereof involved in the embodiments. Unless otherwise stated, like or functionally similar elements in the drawings have been assigned the same reference numerals.
[0038] A hardware system and method for implementing the present invention will now be described by way of non-limiting examples. Detailed Description
[0039] The structure described below is as follows. Part 1 describes general embodiments and higher-order variants. Part 2 details the preferred embodiments and technical implementation details. Part 3 summarizes the final comments. Note that the method and its variants are collectively referred to as "the method". All reference numerals Sn refer to the method steps of the flowchart of FIG. 7, while the reference numerals refer to the apparatus, elements, and concepts involved in the embodiments of the present invention.
[0040] 1. General Embodiments and Higher-Order Variants
[0041] Now referring to Figures 1 to 5Describe in detail the first aspect of the present invention. This aspect relates to a hardware system 1, which is also referred to herein as the "system" for simplicity. The system 1 is designed to execute an artificial neural network (ANN) by efficiently evaluating mathematical functions (such as activation functions) applied to neuron outputs.
[0042] Figure 2 An example of such a hardware system 1 is shown. The system 1 essentially includes a neural processing device 15, a hardware-implemented look-up table (LUT) 17, and one or more processing units 18.
[0043] The neural processing device 15 is configured to implement M artificial neurons, where M≥1. However, in practice, M is typically strictly greater than 1. For example, the device 15 can enable up to 256 or 512 neurons, possibly more. However, as illustrated later, there may be cases where the neural processing device 15 may implement a single neuron at a time. The neural processing device 15 can advantageously have a crossbar array structure 15, as Figure 2 assumed.
[0044] As Figure 3 shown, the LUT is implemented in the form of one or more LUT circuits 175. Several types of LUT circuits 175a, 175b can be envisaged, as discussed in detail later.
[0045] The system further relies on M' processing units 18 to evaluate the mathematical function, where M≥M'≥1. As Figure 3 shown, each processing unit 18 can include several processing elements 185 and enable multiple effective processors.
[0046] As Figure 2 and Figure 3 shown, the neurons are connected to the processing units 18, and the processing units 18 themselves are connected to the LUT circuits. Several configurations can be envisaged. Figure 3 A preferred architecture is shown. At a minimum, each processing unit 18 is connected by at least one of the M neurons implemented by the device 15. In this way, each processing unit 18 can access the neuron outputs, that is, through the values of at least one neuron output, possibly more. Additionally, each processing unit 18 is connected to one or more of the LUT circuits 175, 175a, 175b to allow for fast calculation of a second value. For example, each processing unit can be connected to a corresponding LUT circuit 175, as Figure 3 assumed.
[0047] In the following text, the neuron output is referred to as the "first value", as opposed to the value output by processing unit 18, which is referred to as the "second value". The "first value" corresponds to one of the M values output by the neuron in each algorithmic cycle, while the "second value" corresponds to the value of a mathematical function applied to that first value, as evaluated (i.e., computed) by the processing unit. Note that the algorithmic cycle is a computational cycle triggered by the neural processing unit 15. Each algorithmic cycle begins with a computation performed by this unit 15 (see Figure 6 step S40 in
[0048] ). Further, each processing unit 18 is configured to access at least one first value (from the connected neurons) and output a second value in each algorithmic cycle. Thus, during each algorithmic cycle, the processing unit outputs M second values. However, the number of available processing elements may require several computational sub-cycles in order for the processing unit to output M second values within each algorithmic cycle.
[0049] The above specifications of system 1 define the minimum constraints regarding the architecture of processing unit 18, LUT circuit 17, and neural processing device 15. Various different embodiments can be considered. Additionally, many concepts are also relied upon, which are defined as follows.
[0050] Hardware architecture. The hardware system 1 includes several devices (i.e., one or more processing units 18, one or more LUT circuits 17, and neural processing device 15), which are connected to each other to form system 1. However, system 1 itself can be manufactured as a single device, or even as a single apparatus. Specifically, the LUT circuit 17, processing unit 18, and neural processing device 15 can all be co-integrated on the same chip, as Figure 2 is assumed. Additional components may be involved, as discussed later with reference to Figure 2 .
[0051] In principle, the neural processing device 15 can be any information processing device 15 or information processing apparatus capable of implementing an artificial neuron of an ANN. The device 15 performs the basic functions inherent to an ANN neuron. That is, the ANN neuron generates signals that are sent to other neurons, such as neurons in the next layer in a feedforward or recurrent neural network configuration. However, such signal encoding typically requires the value of post-processing (such as applying an activation function), and thus has the benefit of having a processing unit 18 connected to the neuron.
[0052] The neural processing device 15 can be a general-purpose or a special-purpose computer. However, preferably, the processing device 15 has a crossbar array structure 15 (also referred to herein as a "crossbar array", or simply as a "crossbar"). The crossbar array structure is a non-conventional processing device that is designed to efficiently process analog or digital signals to perform matrix-vector multiplication, as described in the background section. Relying on the crossbar array structure 15 can already significantly accelerate matrix-vector multiplication, such as that involved in the training and inference phases of an ANN.
[0053] The crossbar array structure 15 enables M neurons at a time (where M > 1), and can be used to implement a single neural layer (or a part thereof) at a time. The neurons are represented as ν1…ν Figure 3 in M . In principle, M can be any number permitted by the technology used to manufacture the device 15. Typically, the number M of neurons enabled by the crossbar is equal to 256, 512, or 1024. However, the problem to be solved may involve ANN layers that are disproportionate, i.e., layers that involve a different number of neurons than what the crossbar 15 can effectively permit (at a time). Therefore, a distinction should be made between the number M of neurons that the crossbar 15 actually enables at a time (which can be referred to as the physical neural layer) and the size of the abstract neural layer involved in the problem to be solved. However, in practice, this potential difference is not a problem. In fact, an ANN layer with fewer than M neurons can be directly processed by the device 15, while an ANN layer with more than M neurons can be mapped onto the crossbar 15 by repeating the operation of the crossbar 15. Thus, disproportionate ANN layers can be effectively processed in practice. For completeness, the crossbar array structure 15 involved herein can generally be used to map neurons in various ANN architectures, such as feedforward architectures (including convolutional neural networks), recurrent networks, or transformer networks.
[0054] The crossbar array structure 15 can operate periodically in a closed loop to enable this structure 15 to implement multiple consecutive neural layers of an ANN. In a variant, multiple crossbar array structures 15 can be cascaded to achieve the same effect. The neural layer implemented by the crossbar array structure 15 can be any layer (or a part thereof) of an ANN, including the final layer that may consist of a single neuron. Thus, in some cases (e.g., during the final algorithm loop), the number of neurons effectively enabled by the crossbar array structure can be equal to 1.
[0055] The architecture of the hardware system 1 is different from traditional computer architectures, where a single digital processor (or a single set of digital processors) is typically used to compute neuron output values and apply subsequent mathematical functions (e.g., activation functions). Instead, in the current context, the processing hardware 15 for computing neuron outputs is different from the hardware means 18 for applying subsequent mathematical functions. That is, the processing unit 18 will be more preferably "close" to the neural processing hardware 15. That is, in system 1, the processing unit 18 is preferably configured as a near-memory processing device, as Figure 2 hypothesized. Note that here, "near-memory" is equivalent to considering device 15 as the memory storing neuron outputs. For example, via a dedicated readout circuit 16 known per se, the neuron outputs are effectively transferred to the processing unit 18. In contrast, in traditional computerized systems, neuron outputs typically have to be transferred via a traditional computer bus, stored in the main memory (or cache) of the computerized system, and read from that memory to apply mathematical functions. That is, as Figure 2 shown, the near-memory configuration is different from a normal cache in a CPU chip.
[0056] Furthermore, the processing unit 18 preferably also involves non-traditional computing devices, such as vector processing units (as Figure 3 shown), which can further accelerate the computation.
[0057] More importantly, in this context, the LUT is implemented in hardware, thanks to different hardware circuits 17, i.e., circuits different from each of the neural processing device 15 (for computing neuron outputs) and the processing unit 18 (for applying mathematical functions). Thus, not only may the processing hardware 15, 18 be different from traditional hardware, but also the hardware-implemented LUT (implemented by different circuits 17) is relied upon to quickly retrieve parameter values and thus more efficiently apply mathematical functions.
[0058] Hardware-implemented look-up table. The LUT is implemented in hardware by means of one or more dedicated circuits 17, which can be regarded as memory circuits. Each of these circuits can implement the same numerical table, or at least partially different numerical tables. These circuits can also implement completely different tables. However, in this case, each different table can still be regarded as part of a superset forming the LUT. Overall, this table may enable multiple types of mathematical functions.
[0059] In Figure 3 , the reference numeral 17 generally refers to a group of one or more LUT circuits 175, each LUT circuit implementing its own table, the values of which may be different. Each circuit 175 can be, for example, a circuit 175a that hard-codes parameter values or an addressable memory circuit 175b, as Figure 4 andFigure 5 Assumed separately.
[0060] In Figure 5 , it is assumed that the LUT is implemented by at least one addressable memory circuit 175b that stores parameter values. It should be understood that multiple LUT circuits 175b may be involved, for example as Figure 3 shown, replacing circuit 175. Note that when the LUT circuit is implemented as an addressable memory circuit, system 1 typically includes a programming device (not shown) connected to the memory circuit 175b to rewrite (and thus update) the corresponding parameter values when necessary.
[0061] In Figure 4 , it is assumed that the LUT is implemented at least in part by a hard-coded circuit 175a, which is designed to provide the necessary parameter values, similar to a read-only memory (ROM) circuit. Again, multiple LUT circuits 175a may be involved (as in Figure 3 , replacing circuit 175). However, in practice, hard-wired circuits typically enable only a small number of functions. For example, Figure 4 the circuit 175a shown enables a single mathematical function. Therefore, it may be preferable to rely on a rewritable memory circuit, such as a random access memory (RAM), to implement a reprogrammable LUT, as Figure 5 shown.
[0062] Figure 4 And Figure 5 the examples shown assume that one LUT circuit 175a, 175b is connected to the corresponding processing element 185a, 185b. However, system 1 may actually involve multiple processing elements 185a, 185b connected to one or more LUT circuits 175a, 175b. Multiple architectures can be envisioned. At a minimum, such an architecture allows at least one LUT circuit to be mapped to the corresponding processing unit 18 (or its processing element), at least one at a time. That is, each processing unit 18 can be dynamically connected to different LUT circuits, which are switched at any time. However, when the processing unit 18 is active, at least one LUT circuit should be connected to the processing unit 18.
[0063] Nonetheless, since multiple LUT circuits can be easily provided in practice (possibly in the same device), a convenient architecture choice is to provide multiple LUT circuits 17, where at least one LUT circuit 175 is permanently connected to the corresponding processing unit 18, as Figure 3 shown. Thus, the LUT circuit 17 can include at least M' different circuits 175, which are respectively mapped to M' processing units 18. Nonetheless, the LUT circuit may include more than M' circuits. For example, in Figure 3In the example, the number T of LUT circuits 175 exceeds the number M' of processing units 18 (i.e., T > M') to allow for redundancy and / or to be able to preload (i.e., prefetch) table values during the computation, if needed, for computational speed.
[0064] Parameter types and parameter values. There is a difference between the parameter type (e.g., a given polynomial coefficient) and the actual value (e.g., 2.173) of each type of parameter retrieved from the LUT. The LUT stores parameter values for one or more parameter types. If the mathematical function is defined as a polynomial, the type of parameter can, for example, correspond to polynomial coefficients. In fact, if the mathematical function is defined as a piecewise polynomial, multiple sets of parameters can be associated with each type of function, as in the embodiments discussed later. However, more generally, several types of mathematical functions may be involved, which may require different types of parameters.
[0065] Values and signals. The values generated by components 15, 18 and the values retrieved from the LUT circuit 17 are encoded in the corresponding signals. In fact, the signal encoding the neuron output is passed from the neural processing device 15 to the processing unit 18 (usually via the readout circuit 16, see Figure 2 ). The computation performed by the processing unit 18 further requires signals to be transmitted from the LUT circuit (to retrieve the necessary parameter values). Next, further signals (encoding a second value) are passed from the processing unit 18 to the neural processing device (which can be the same device 15 or another processing device) to trigger the execution of another neural layer, and so on. This may require an input / output (I / O) unit 19, as Figure 2 assumed. The I / O unit 19 can further be used to connect the system 1 to other machines.
[0066] Processing units, vector processing units, processing elements, and efficient processors. The processing unit 18 is a processing circuit, typically in the form of an integrated circuit. As described, such a circuit is preferably configured as a near-memory processing means in the output of the neural processing device 15, see Figure 2 to efficiently process the neuron output.
[0067] In this context, the processing unit 18 can include processing elements that only require an arithmetic logic unit (to perform basic arithmetic and logical operations), or even only multiplication-and-addition processing elements. In a variant, a more complex type of processing unit is used, which can also particularly perform control and I / O operations if needed. Such operations can also be performed by other components of the system 1 (such as the I / O unit 19).
[0068] Thanks to at least one processing element 185, 185a, 185b, each processing unit 18 enables at least one effective processor. The processing unit 18 can be a standard microprocessor. However, preferably, the present processing unit 18 is a vector processing unit (as Figure 3 assumed), which allows for some parallelization when applying mathematical functions. In Figure 3 the example, each processing unit 18 includes a constant number b of processing elements 185, which enables efficient operation on a one-dimensional array of dimension b, which can be advantageously exploited in the output of the neurons. Overall, in Figure 3 , the M' vector processing units 18 include bM' processing elements 185, where b represents the degree of parallelism implemented by each vector processing unit. However, in principle, the vector processing unit 18 can have a different number of processing elements 185.
[0069] Whether in a variant or in addition to vector processing, traditional degrees of parallelism can also be further involved. That is, each unit 18 can include multiple cores. In fact, each processing element 185 can be a multi-core processor. Thus, generally speaking, one or more processing units 18 can involve one or more processor cores, where each core can enable one or more effective processors, for example, by threading that divides a physical core into multiple virtual cores. In other words, the M' processing units as a whole can produce M'' effective processors, where M'' is at least equal to M' and can be strictly greater than M'.
[0070] The number of processing units (or effective processors) relative to the number of neurons. In principle, each of M' and M'' can be greater than the number of neurons M, as described above. However, considering that each algorithmic loop executed by the neural processing device 15 typically requires at most M functions, this may be useless in a configuration such as Figure 2 shown. Therefore, a preferred setting is that the system 1 includes M' processing units (which may potentially involve M'' effective processors), where M ≥ M'' > M' ≥ 1. In a variant, the system 1 has additional processing power such that M' (or M'') can be strictly greater than M. Such a setting is particularly useful for implementing certain types of activation functions, such as the cascaded rectified linear unit (CReLU), which preserves positive and negative phase information while enforcing non-saturating non-linearity. For example, calculating CReLU(x) = [ReLU(x), ReLU(-x)], where [.,.] represents concatenation, can be done in multiple steps. However, thanks to 2M' processing units (or 2M'' processors), the performance can be improved by performing this operation in a single execution. In this case, the number of processing units (or effective processors) can advantageously be strictly greater than the number M of neurons enabled by the device 15 during each algorithmic loop.
[0071] In summary, the number M of neurons enabled in each device 15 may be less than M' (or M"). It may also be equal to M', so that each neuron output can be processed in parallel. If one or more of the M' processing units (or M" effective processors) are shared by at least some of the M neurons, other configurations may involve fewer processing units (or effective processors) than neurons, i.e., M' < M and / or M" < M (assuming M > 1). The latter case reduces the number of processing units (or effective processors), resulting in M artificial neurons taking turns using the M' processing units (or M" effective processors).
[0072] In other words, each neuron is connected to one of the M' processing units 18, but each processing unit 18 may be connected by more than one neuron. The system 1 includes at least one processing unit, which involves at least a single processor (possibly a single core). However, relying on a single processor (core) may seriously affect the throughput, so there are benefits to involving multiple processing units or at least multiple processing cores. On the contrary, vector processing is costly, so a trade-off is needed to optimize the number of effective processors according to the number of neurons.
[0073] For simplicity, the following description assumes that each processing unit 18 includes a constant number b of processing elements 185, as Figure 3 shown. In addition, it is assumed that each processing element 185 produces a single effective processor (virtual processing is not allowed for processing elements in this case), as in the examples of Figure 4 and Figure 5 Thus, the total number of processing elements is equal to bM', and the number M" of effective processors (as effectively enabled by the vector processing unit 18) is equal to the total number of processing elements (i.e., M" = bM').
[0074] As described, it may be necessary to optimize the number of effective processors relative to the number of neurons. In this regard, the ratio of M" to M is preferably between 1 / 8 and 1. A preferred architecture involves one or two vector processing units (i.e., M' = 1 or 2) for each neural device 15, where each vector processing unit 18 involves 64 processing elements, where M" is equal to 64 or 128, and the number M of neurons in each device 15 is equal to 512. For example, two vector processing units (M' = 2) can be used, each vector processing unit involving 64 processing elements (b = 64), such that M" is equal to 128 and in this case the ratio of M" to M is equal to 1 / 4.
[0075] All vector processing units may be directly connected to the neuron outputs (via the readout circuit 16), as Figure 2As assumed. In a variant, some of the vector processing units may be indirectly connected to the neurons. More precisely, some or all of the neurons may first be connected to intermediate processing units (not shown), which in turn are connected to the vector processing units. For example, in the case of using two vector processing units (M' = 2), the first vector processing unit may be directly connected to M neurons in the output of the crossbar array 15 (actually in the output of the readout circuit 16), while the second vector processing unit may be connected to a so-called depthwise processing unit (DWPU), which is itself connected to the output of the crossbar array 15. Inserting the DWPU allows for depthwise convolution operations.
[0076] The number of LUT circuits and the number of processing units. In principle, it is sufficient for the LUT to be implemented by a single circuit (e.g., a single addressable memory circuit) serving each processing unit 18. However, if a large number of processing units are relied upon, this may require a large number of interfaces or data communication channels. However, it should be noted that when a single LUT circuit is mapped onto a single vector processing unit of b processing elements (as Figure 3 shown), then the LUT circuit requires a single port, the output signal of which can be multiplexed to the b processing elements.
[0077] Sometimes it may be desirable to be able to apply multiple mathematical functions. To this end, it may be desirable to connect a processing unit 18 to multiple LUT circuits 17. However, if a single LUT circuit can already enable a large number of parameter values, such that multiple functions can be applied to the output of a neuron, then this is usually unnecessary. Thus, it may be sufficient to connect a single LUT circuit to the corresponding processing unit.
[0078] For completeness, the LUT circuits may be shared by the processing units 18, rather than by the processing elements 185 of each processing unit 18 (as Figure 3 assumed). That is, the LUT circuits may consist of J different circuits, where J < M', resulting in a configuration of M ≥ M" ≥ M' > J ≥ 1.
[0079] Processing units and mathematical functions. The number of functions L available to each processing unit 18 is greater than or equal to 1 (L ≥ 1). In the case where multiple mathematical functions are available, any one of the L functions may be selected and subsequently formed due to the corresponding parameter values accessed from the LUT. That is, each of the M' processing units can potentially apply any one of the available mathematical functions.
[0080] A convenient approach is to rely on the same general function construction (e.g., piecewise-defined polynomial functions), which are appropriately parameterized such that ultimately the same construction can be used to evaluate various functions and apply them to each neuron output, as in the embodiments discussed below.
[0081] Multiple configurations can be considered again here. In a simple scenario, the M' processing units 18 apply the same function (i.e., a single function) to the output of each neuron from the neural layer implemented by the device 15, i.e., the output at each algorithm cycle. Nevertheless, different functions may have to be applied to the output of successive neural layers. Conversely, in a more complex scenario, the M' processing units can implement up to M different functions (possibly selected from L > M potential functions) in each algorithm cycle. In this case, different functions are applied to the neuron outputs. In other words, different functions can be used from one neural layer to another, and if needed, different functions can also be applied to the neuron outputs from the same layer.
[0082] Arguments and parameters of a mathematical function. In principle, an argument is a variable passed to a mathematical function to compute its output value. The parameters of a function can also be regarded as variables. However, a parameter is a variable that determines (i.e., helps to fully specify) the function, similar to the parameters specified in a function declaration in a programming language. For example, the polynomial function f(x) = αx + β has one argument x, but involves two parameters a and b, and the values of these two parameters help to fully specify the function.
[0083] Similarly, in this context, any mathematical function involved takes the value x as an argument, i.e., the value encoded in the signal of the neuron output. Therefore, the processing unit 18 (or processing element) computes any output value of the mathematical function based on the value encoded in the signal obtained from the neurons connected to this unit 18 (or processing element). In Figure 4 and Figure 5 the argument of the function is written as x IN , while y OUT represents the output value of the function. Nevertheless, in order to be able to compute the output value, the parameter values must first be retrieved from the LUT according to the current scheme.
[0084] Advantages of the proposed solution. According to the proposed solution, the neuron output is computed by the dedicated neural processing hardware 15, and a mathematical function is applied to the output of the neuron using the dedicated processing device 18. The hardware implementation of the LUT accelerates the retrieval of the parameter values required for computing the mathematical function.
[0085] This architecture allows for accelerated computations, not only because components 15, 17, 18 can be optimized individually, but also because the intercommunication can be enhanced. For example, processing unit 18 can be configured as a near-memory processing unit 18, "close" to device 15 to accelerate the transmission of neuron outputs beyond the acceleration already achieved by the neural processing device 15 (e.g., crossbar array) implementing the neurons. Most importantly, processing unit 18 can efficiently access parameter values from a hardware-implemented look-up table, thus significantly accelerating the calculation of the function output. As a result, neuron outputs can be processed faster before being passed to the next neuron layer.
[0086] In the current context, it is important to understand that the LUT is not used to directly look up the function output (as is typically done when using a look-up table) but rather to more efficiently access the parameter values required to evaluate the function. In this way, even a moderately sized LUT already allows for the implementation of various functions (e.g., non-linear activation functions, normalization functions). Given that the LUT stores parameter values rather than mapping input values to output values, very little memory is required. Nevertheless, the LUT may be designed as a reconfigurable table, whereby, if new types of functions are needed over time, the mathematical functions can be dynamically reconfigured or updated as the computation progresses (with respect to training or inference).
[0087] Furthermore, as previously mentioned, the present solution is compatible with integration. That is, LUT circuits 17, processing units 18, and neural processing devices 15 can be advantageously co-integrated in the same device. In particular, the LUT circuits 17 can be co-integrated near their respective processing units 18. Thus, the hardware system 1 can consist of a single device (e.g., a single chip) that co-integrates all the required components. Therefore, the present system 1 can be conveniently used in a dedicated infrastructure or network to serve multiple concurrent client requests, as Figure 1 is assumed.
[0088] All of these will now be described in detail with reference to specific embodiments of the present invention. First, each processing unit 18 can be advantageously configured to obtain a mathematical function value (i.e., a second value) by first selecting each required parameter according to the value of the neuron output (a first value), and then retrieving the corresponding parameter value accordingly. That is, the operations performed to obtain the second value are, on the one hand, based on the first value and, on the other hand, based on the parameter values of the relevant parameter group selected according to the first value, where these parameter values are efficiently retrieved from the LUT. The group of suitable parameter values can be initially determined at build time. As will be understood, this makes it possible to reduce the number of parameters required for each evaluation in practice, since a small group of parameters is sufficient to accurately locally estimate the function over an interval containing the input value (the first value). For example, a linear polynomial (each polynomial only requires two coefficients) can accurately locally fit a curve.
[0089] In this regard, with more specific reference to Figure 4 and Figure 5 , each processing unit 18 may be further configured to select a relevant parameter set by comparing a first value (neuron output value) with bin boundaries. This makes it possible to identify the relevant bin containing the first value. Next, the relevant parameter set is selected according to the identified bin and then retrieved from the LUT. If necessary, further parameters may be relied upon to select the desired function type (e.g., ReLU, softmax, binary, etc.). The bin boundaries are stored in the LUT together with the parameter values. Note that the bin boundaries can also be regarded as parameters for calculating the function. However, the function of these parameters is essentially different from the parameter values used for calculating the function (e.g., the relevant polynomial coefficients). The bin boundaries can also be retrieved efficiently to allow for quick comparison. Thus, the binning problem can be effectively solved, and the relevant parameter set can be quickly identified.
[0090] For this purpose, each processing unit 18 preferably includes at least one comparator circuit 182 (see Figure 4 and 5 ), which is designed to compare the first value with the bin boundaries. That is, a dedicated comparator circuit is used to effectively identify the relevant bin. As shown in Figure 4 and Figure 5 , the comparator circuit is further designed to transmit a selection signal encoding the selected parameter set. Finally, due to the transmitted selection signal, the corresponding parameter values are retrieved. Note that the processing unit 18 may actually include more than one comparator circuit 182. For example, the processing unit 18 may include several processing elements 185 (as shown in Figure 3 ), and each of such processing elements may be designed according to Figure 4 or Figure 5 , where each element 185a, 185b includes a comparator circuit. In a variant, the comparator circuits may be partially shared in the processing unit.
[0091] The selection signal may in particular be transmitted to a multiplexer 186, which forms part of the processing element 185a, as shown in Figure 4 . In a variant, the selection signal is passed to the LUT circuit, such as an addressable memory 175b, as shown in Figure 5 .
[0092] More precisely, in Figure 4In the example, the LUT circuit 175a is a circuit for hard-coding parameter values. The multiplexer 186 is connected to its corresponding comparator circuit 182 so as to be able to receive a selection signal during operation. The multiplexer 186 is further connected to the LUT circuit 175a, which allows the multiplexer 186 to select relevant parameter values according to the received selection signal during operation. Such a design makes parameter retrieval extremely efficient. The disadvantage is that the hard-coded data cannot be changed after the circuit 175a is hard-wired. Therefore, it may be preferable to use a reconfigurable memory.
[0093] In this regard, Figure 5 the example shown depicts a LUT circuit 175b including addressable storage units 175b, which are connected to the comparator circuit 182 so as to receive a selection signal during operation. In this case, the selection signal is directly transmitted to the memory 175b (a multiplexer is not required in this case). The LUT circuit 175b is further configured to retrieve the parameter values of the relevant selected parameter group according to the received selection signal. For the rest, the comparator circuit 182 can be equivalent to Figure 4 the circuit used in Figure 4 and Figure 5 is further described in Part 2.
[0094] One can rely only on the binary tree comparison circuit. However, more complex comparison schemes and comparator circuit layouts can be envisioned, which implement multi-level comparisons to accelerate binning. Specifically, the comparator circuit 182 can advantageously be configured as a multi-level q-ary tree comparison circuit, where as Figure 8 shown in the table, for one or more of the multi-level comparisons implemented by the circuit 182, q is greater than or equal to three.
[0095] Specifically, in the q-ary tree comparison circuit, the number of comparator layers is equal to log q (K), where K represents the total number of bins used, assuming q comparators in each layer. Now, the number of layers, the total number of comparators, and the number of comparators in each layer can be jointly optimized, as Figure 8 shown. The table shows the optimal number of comparators that can be used in each layer (the second row), the total number of comparators involved (the third row), and the associated computational cost (the fourth row). Figure 8 The number of layers considered in Figure 8 varies between 1 and 6, while the number of comparators in each layer varies between 1 and 63 (the same is true for the total number of comparators). The number of layers is related to the latency: the more layers, the longer the latency. The total number of comparators also incurs a cost. Therefore, the total cost can be equal to the number of layers multiplied by the total number of comparators, as shown below
[0096] The q - ary tree comparison circuit enables q - 1 comparators, such that the number of comparators used in each layer corresponds to q - 1. In principle, the optimal value of q corresponds to the lower bound of the number q* that minimizes the cost function (q - 1)Log K (q) 2 taking into account the trade - off between the number of layers and the total number of comparators. Minimizing this function yields a value corresponding to 3 comparators.
[0097] However, the optimal number of comparators depends on the cost function chosen and the number of layers. For example, the present inventors have performed extensive optimizations based on more complex cost functions, which have led to Figure 8 the optimal values shown. According to this optimization, it is preferable to rely on three comparison layers and a quaternary tree (enabling 3 comparators), where each layer has the same number of comparators (i.e., 3).
[0098] As previously mentioned, the applied mathematical function will advantageously be constructed as a function piece - wise defined by polynomials, where each polynomial is applied to a different interval in the function domain. Thus, such a function (also called a spline), which is a polynomial on each of its sub - domains, can be mapped to binning. The polynomial coefficients can be adjusted to fit a given reference function (i.e., the theoretical function). It should be noted that the polynomials do not need to be continuous across bin boundaries in the current context (although they may be, depending on the reference function). Thus, for any given first value (the argument of the function), a relevant set of parameters can be identified, which corresponds to the polynomial parameters of the locally relevant polynomial. Then the corresponding parameter values are retrieved from the LUT to estimate the output value of the function.
[0099] As is understood, using such a construction is well - suited for the fast calculation of arithmetic units. That is, only arithmetic operations are required to achieve the desired result. Now, each processing unit 18 (or in fact each processing element 185, 185a, 185b) can include an arithmetic unit 188, which is connected to the LUT circuit 17 to perform the required arithmetic operations.
[0100] For example, an addressable storage unit can be used to store at least L×(K×(2 + l)–1) parameter values, where L represents the number of different functions to be implemented by the processing unit (L≥l), l represents the interpolation order of each interpolation polynomial (l≥1). There are K bins and K - 1 bin boundaries. Thus, if the processing unit 18 can implement a total of L functions, each function having K bins, and achieve interpolation of order l, the minimum number of parameters to be stored in the memory is equal to L×(l + 1)×K+L×(K–1)=L×(K×(2 + l)–1).
[0101] In principle, other constructs (in addition to the usual splines) can be relied upon. For example, the functions applied can also be B-spline curves or involve Bezier curves. However, using splines (especially linear polynomials) allows for very efficient computations. Additionally, in such cases, multiplication-and-addition circuits can be used to simply perform such computations, as Figure 4 and Figure 5 assumed. That is, the arithmetic unit 188 can simply consist of a multiplication-and-addition circuit. The circuit 188 is specifically designed to perform multiply-accumulate operations, which efficiently obtain the output values of mathematical functions. It should be noted that up to l operations may need to be performed in this case, where l is the polynomial order. The advantage of relying on the multiplication-and-addition circuit 188 is also that similar (or the same) circuit technologies can be used in the neural processing device 15.
[0102] In fact, the neural processing device 15 preferably includes a crossbar array structure 15, that is, a structure involving N input lines 151 and M output lines 152, where N > 1 and M > 1, as Figure 2 shown. The input lines and output lines are configured in rows and columns, which are interconnected at the intersections (i.e., nodes) via storage elements 156. Each column corresponds to a neuron, whereby the device 15 can implement M layers of neurons. Each output line 152 is connected to at least one of the M' processing units 18. The output lines are typically connected to the processing units via a readout circuit 16, as Figure 2 shown. It should be noted that each output line can be connected to a corresponding processing unit or a corresponding processing element. However, preferably, M neurons share the processing unit 18, as Figure 3 shown.
[0103] The crossbar array structure 15 can be regarded as defining N×M cells 154, that is, repeating cells corresponding to the intersections of rows and columns. As is well known, multiple conductors may actually be required for each row and each column. In a bit-serial implementation, each cell can be connected by a single physical line, which serially feeds an input signal carrying an input word. However, in a parallel data acquisition method, parallel conductors can be used to connect to each cell. That is, bits are injected in parallel into each cell via parallel conductors.
[0104] Each cell 154 includes a corresponding storage system 156, which consists of at least one storage element 156, see Figure 2 . Thus, the N×M cells contain N×M storage systems 156, which are individually referred to as a Figure 2 in 11 to a 44. The storage system 156 stores weights corresponding to matrix elements for performing matrix-vector multiplication (MVM). Each storage system 156 may include, for example, storage elements connected in series that store respective bits of the weights, where the weights are stored in corresponding cells; in this case, the multiply-accumulate (MAC) operation is performed in a bit-serial manner. The storage elements may be, for example, static random access memory (SRAM) devices, although in principle a crossbar structure can be equipped with various types of electronic storage devices (e.g., SRAM devices, flash cells, memristive devices, etc.). Any type of memristive device may be considered, such as phase change memory cells (PCM), resistive random access memory (RRAM), and electrochemically random access memory (ECRAM) devices.
[0105] The storage elements may form part of a multiply-and-add circuit ( Figure 2 not shown in the figure), whereby each column includes N multiply-and-add circuits to efficiently perform the MAC operation. The vector is encoded as a signal applied to the input lines of the crossbar array structure 15, which causes the latter to perform MVM by way of the MAC operation. Although physically limited to M neurons, the structure 15 can still be used to map larger matrix-vector multiplications, as described above. If necessary, the weights may be prefetched and stored in the corresponding cells (in an active manner, e.g., since each cell has multiple storage elements) to accelerate MVM. Generally, MVM can be performed in the digital or analog domain. Compared with all-digital IMC, the implementation in the analog domain can exhibit better performance in terms of area and energy efficiency. However, this typically comes at the cost of limited computational precision.
[0106] Now refer to Figure 6 the flowchart of to describe another aspect of the present invention. This aspect relates to a method of operating the hardware system 1 as described above. The basic features of this method have been implicitly described with reference to the first aspect of the present invention. These features are only briefly described below.
[0107] According to this method, operating the hardware system 1 requires operating the neural processing device 15, for example as Figure 6 completed in steps S20 to S50 of the flow. Generally, operating the neural processing device 15 results in obtaining M first values in each algorithm loop, see step S40. These values are respectively generated by M artificial neurons enabled by the device 15. When the device 15 is a crossbar array structure, an array of input values (i.e., vectors) is encoded as an input signal, and the input signal is applied to the input lines of the device 15 to cause it to generate an output signal at the output of M lines. According to the terms introduced above, such an output signal corresponds to the first signal.
[0108] Next, one or more mathematical functions are applied to the first value to obtain a second value. That is, for each of the M first values generated by the M neurons, the output value of the mathematical function is obtained via the M' processing units 18 (steps S60 - S110). As previously mentioned, the mathematical function takes the first value as the independent variable. Nevertheless, this function is determined by a parameter set, the values of which are retrieved from a hardware-implemented LUT. Accordingly, each mathematical function is computed based on an operand that includes the first value and the parameter set, where the S100 parameter values are retrieved efficiently from one or more LUT circuits 17. Thus, each first value gives rise to a second value, i.e., the output value of the mathematical function. Accordingly, M second signals are obtained (in each algorithm loop), which encode the M output values corresponding to the evaluation of the mathematical function.
[0109] As previously discussed, the mathematical function is preferably evaluated by first selecting the relevant parameter set according to the neuron output (the first value) S70 - S80. Subsequently, the operation S110 is performed based on the first value and the parameter values, which are retrieved S100 according to the selected parameter set. As Figure 6 Further, it can be seen that the parameter set is preferably selected by comparing the first value with the bin boundaries S70 to identify the relevant bin, whereby the relevant parameter set can subsequently be selected according to the identified bin S80. Given that the bin boundaries are retrieved from the LUT, this operation can be performed efficiently. Typically, multiple computational algorithm loops are performed, as Figure 6 shown. Each algorithm loop begins with a computation (i.e., MVM) performed by the neural processing unit 15.
[0110] A typical process is as follows. This process relies on the hardware system 1 as Figure 2 shown to perform ANN-based inference. For simplicity, it is assumed here to have a feed-forward configuration. Additionally, it is assumed that the device 15 is capable of enabling a sufficiently large number of M neurons to map any layer of the ANN. At step S10, the system 1 is provided; it particularly includes a crossbar array structure 15, LUT circuits 17 (assumed to be programmable storage circuits), and near-memory processing units 18. The parameter values required to evaluate the mathematical function are initially determined at step S5 (build time). The LUT is initialized accordingly at step S20; this is equivalent to programming the LUT circuits S20 so that they can store sufficient parameter values. If necessary, the LUT circuits can be reprogrammed later to update the function. In addition to the LUT circuits 17, the matrix coefficients (i.e., weights) must be initialized in the crossbar array 15 to configure it S30 as the first neural layer of the ANN to be executed.
[0111] Next, the input unit 11 of system 1 applies the currently selected input vector to the S40 crossbar array 15 to cause it to perform MVM. Signals are obtained at the outputs of the M columns of the crossbar 15. The obtained signals encode the neuron output values (or first values). The corresponding values are read out by the dedicated circuit 16 and passed to the processing unit 18 at S60. The comparator circuit of unit 18 compares the neuron output values with the bin boundaries at S70 to identify the relevant bin. Then, the comparator circuit forwards the corresponding selection signal at S80 to the LUT circuit to retrieve the relevant parameter values from the LUT at S100. This enables the efficient calculation of the output values of one or more mathematical functions (such as activation functions), which is preferably performed by the multiply-and-add circuit 188.
[0112] This process is repeated for each successive layer. If the current neural layer is the last layer (S120: yes), the algorithm can pass back the final function value at S130 to form the inference result. Otherwise (S120: no), if required, the function value obtained in step S110 can be passed at S140 to a further processing unit (e.g., a digital processing unit). That is, the result of step S110 may have to be passed to the digital processing unit, which performs operations at S140 that cannot be implemented by the crossbar array 15 or the processing unit 18. For example, a digital processing unit may be involved to perform a max pooling operation. Then, the value obtained at the output of step S110 (or S140) is sent at S150 to the input unit of the same crossbar array 15 or another cascaded crossbar array 15. That is, a new input vector is formed, and another algorithm loop S40 - S150 is started. Note that, in parallel with steps S60 - S140, new matrix coefficients can be stored at S50 in the (next) crossbar array to configure it for the next neural layer.
[0113] For the reasons mentioned above, the mathematical functions involved are preferably constructed as piecewise-defined polynomial functions. In this case, each parameter set consists of two or more polynomial coefficients, and the required operations can be performed as arithmetic operations only at S110. Due to the multiply-and-add circuit 188, this can be done efficiently, which enables the reuse of the same technology as that used for the neural processing device 15.
[0114] For example, a set of linear polynomials can be used to evaluate each mathematical function within the range of interest, with each linear polynomial mapped to a corresponding bin. In the case of using linear polynomials, the respective parameter sets can each consist of a scale coefficient and an offset coefficient (although the polynomials can be defined in different ways).
[0115] The degree of the polynomial to be used can be selected in step S5. This preliminary step S5 may also include determining a suitable set of bin boundaries and corresponding parameter values for each reference function (i.e., the functions potentially of interest to the ANN). Various methods can be envisioned. Generally, the determination of suitable bin boundaries can be regarded as an optimization problem. Given a predetermined number of bins to be used, the appropriate bin boundaries for S5 (for each reference function) are typically determined by minimizing the number of bins (given the maximum tolerable error at any point) or the maximum error between the approximation of the reference function (calculated based on the parameter values) and the theoretical value of the reference function. Joint optimization can also be performed to optimize both the number of bins and the maximum error. Detailed explanations and examples are provided in Part 2.
[0116] Various types of reference mathematical functions can be considered, regardless of the construction used to estimate them. The LUT circuit 17 can enable various mathematical functions conventionally required in ANN calculations, such as activation functions, normalization functions, reduction functions, state update functions, and functions for performing analytical classification, prediction, or other similar inferences.
[0117] Activation functions are an important class of functions because most such functions need to be applied to the neuron outputs. In particular, non - linear activation functions can be used, such as the so - called binary step function, Sigmoid, Tanh, ReLU, Leaky ReLU, and Softmax (i.e., normalized exponential) functions. Specific normalization functions (e.g., batch normalization, layer normalization) may also be required. Additionally, as mentioned previously, the applied mathematical functions can be any analytical functions used to perform inferences (classification, prediction). However, in some cases, the mathematical functions can be bypassed or configured as identity functions.
[0118] The parameter values retrieved from the LUT are generally sufficient for the processing unit 18 to calculate the full function output. However, in other cases, additional (external) processing may be required, for example, to calculate the sum of the exponents required in the softmax function. Additionally, other types of operations may sometimes have to be performed, such as reduction operations on a set of values, or arithmetic operations between these values. For completeness, the LUT can also store values that can be used by the system 1 to perform other tasks (such as tasks for support vector machine algorithms).
[0119] The above - described embodiments have been briefly described with reference to the accompanying drawings and can include various variations. Various combinations of the above - described features can be envisioned. Examples are given in the next part.
[0120] 2. Specific Embodiments - Technical Implementation Details
[0121] 2.1 Bin Boundary Determination and Interpolation
[0122] The LUT preferably implements a piecewise polynomial function such that for a given input value x, bin boundary vector b b and coefficient vector c, the output o of the function is typically calculated as o = f(x, c[i]), where i is an integer satisfying x > b b [i - 1], x ≤ b b [i]. Ideally, a suitable method for establishing LUT values should allow for the simple determination of the parameters used to compute the desired function. Additionally, the method should advantageously support different types of approximation of the function.
[0123] Several methods that can be used to approximate a function as a piecewise polynomial (spline) are described below. In the simplest case, the interpolation is linear and the function is defined by two parameters: a scale and an offset coefficient.
[0124] Figures 7A to 7C A method of binning using linear interpolation with the GELU function is illustrated. The normal (continuous) line represents the reference function, while the thick (striped) curve represents the interpolated version, approximated using pre - determined parameter values. The thick dots represent the optimal boundaries of the bins.
[0125] The differences in the following method examples mainly lie in the way of obtaining the bins. What they have in common is that they all attempt to minimize the error between the function and its interpolated version. However, they differ in the way of performing the minimization. Some of the proposed schemes attempt to minimize the number of bins given the maximum error at an arbitrary point, while other schemes attempt to minimize the error given the number of bins to be used.
[0126] For example, the following method minimizes the number of bins given the maximum error at an arbitrary point. The underlying algorithm is as follows:
[0127] (i) The interval of interest of the mathematical function is defined by the user. A pointer is assigned to point to this interval;
[0128] (ii) Subsequently, interpolation is performed on the points that delimit the interval;
[0129] (iii) Then, the error between the original function value and the value of the interpolated curve at the center of the interval (pointed to by the pointer) is calculated;
[0130] (iv) If the error is greater than or equal to the tolerance defined by the user, the pointed - to interval is split in half and the pointer is moved to the left - hand side of the split interval. Otherwise, the pointer moves to the next interval on the right. Once there are no more intervals on the right, the algorithm stops; otherwise, it returns to step (ii).
[0131] The comments proceed in order. For symmetric and antisymmetric functions, this algorithm is applied to the part of the function either to the left or to the right of the axis of symmetry or axis of antisymmetry, and the value of the bin is calculated by mirroring the bins according to the symmetric / antisymmetric pattern. Heuristics can be relied upon to automatically determine the limits of the interval of interest. The interval of interest can also be initially divided into a fixed number of equally spaced bins; in this case, the above algorithm can be applied to each initial bin boundary.
[0132] Assuming linear interpolation is required, each linear part requires a slope coefficient (slp) and an offset coefficient (off). Then, the specific output o of the function on a specific subdomain corresponding to a given input x can be estimated as o = slp[i] × x + off[i], where i represents the integer corresponding to the input x, e.g., determined as i|x>b b [i - 1]; x ≤ b b [i], where b b Also represents the bin boundary vector.
[0133] Let us illustrate the above binning method with an example, where the Gaussian error linear unit (GELU) function will be approximated, see Figure 7A . Initially, the number n i of points is determined (or inferred), such as Figure 7A , where n i ≥ 2. Such points delimit n i + 1 intervals. For simplicity, assume that only two points are initially determined. Binning is only required between the non - linear parts of the function. Therefore, the algorithm can first determine the intervals where the function is non - linear. This results in a single interval in the example of Figure 7A . The next step consists of binning only within the non - linear region. To this end, the algorithm can measure the error at the center of the bin and then divide the interval into two equal sub - intervals if the error exceeds the tolerance. This is shown in Figure 7B , which shows the additional points. Then the same operation can be repeated until a suitable number of intervals is reached, which results in an acceptable interpolation error. Finally, each interval is assigned a triple of optimal parameter values. For example, the first interval corresponds to b b [0], slp[0] and off[0], the second interval corresponds to b b [1], slp[1] and off[1], and so on. Each set of parameter values is stored in the LUT for later retrieval during operation.
[0134] Another scenario is to minimize the number of bins used given a maximum error (E max ) at an arbitrary point (optimization limit). The algorithm proceeds as follows:
[0135] (i) The interval of interest of the function is also defined by the user;
[0136] (ii) The algorithm starts from one boundary (left or right) of the interval of interest and considers it as one of the boundaries of the first bin. Assume the left boundary as the starting point below. In this case, the left boundary of the interval of interest is also the left boundary of the first bin. The initial values of the other bin boundaries (the right bin boundary in this case) are imposed by the user (defined via delta);
[0137] (iii) Select linear interpolation of the bin midpoints such that the maximum error (E bin_max ) between the interpolation function and the function to be interpolated in the given bin lies at its left and right boundaries and in the middle of the bin (with opposite signs). This can be achieved by interpolating the points that define the bin boundaries. Then the line is shifted by half of the mid-bin error, equal to E bin_max ;
[0138] (iv) If E bin_max < E max , the bin is increased (by moving its right boundary to the right) by the initial value of the bin (delta), and step (iii) is repeated to determine the new E bin_max . Then step (iv) is repeated until E bin_max ≥ E max when the next step (step (v)) is taken;
[0139] (v) The algorithm checks whether |E max – E bin_max | < epsilon, where epsilon is defined by the user. If so, the algorithm proceeds to step (vi). Otherwise, it moves the right boundary to the left by delta / 2 and repeats step (iii) to determine the new E bin_max . Then step (v) is repeated, halving the interval by which the boundary is moved each time and along the direction determined by the sign of E max - E bin_max (positive if to the right, otherwise to the left), until the condition |E max - E bin_max | < epsilon is true. Then, the algorithm proceeds to step (vi);
[0140] (vi) Then, the algorithm starts from step (ii) and repeats the same process for the next bin; and
[0141] (vii) The algorithm continues until all bins are determined. Finally, the function can be estimated by interpolation based on the parameter values obtained for each bin.
[0142] Note that other methods can be used to find the size of the bins, verifying |E max-E bin_max | <epsilon>. In addition, heuristic methods can be used to automatically calculate the limits of the intervals of interest.
[0143] Various other algorithms can be envisioned. For example, a non-iterative variant of the above algorithm can be designed that determines the binning in a single step. This variant minimizes the error for a given number of bins to be used and can involve any interpolation polynomial. Given the desired number of bins and the degree d of the interpolation polynomial, this algorithm approximates the d-th derivative of the function. The optimal size of each bin is inversely proportional to the rate of change of the highest-order coefficient and, accordingly, to the rate of change of the d-th derivative of the function, which is actually the (d + 1)-th derivative of the function. The cumulative sum of the (d + 1)-th derivative of the function is calculated at several sampling points (far greater than the number of bins), and the optimal binning is identified accordingly.
[0144] Further schemes can be based on neural networks, trained to minimize the number of bins or the maximum error in each bin.
[0145] 2.2 Preferred Hardware Implementations
[0146] A hardware implementation assuming linear interpolation is described below. Two types of embodiments can be envisioned in this case. The first type involves a fixed (pre-defined) function implementation, where K - 1 bin boundaries, along with K scaling coefficients and K offset coefficients, are used. These quantities remain constant during runtime. Typically, a small number of required coefficients do not require an addressable memory and can alternatively be hard-coded ( Figure 4 ) in the LUT circuit 175a, similar to a ROM circuit. The priority network of the comparator 182 provides a selection signal that is fed to the multiplexer 186. The multiplexer 186 can accordingly select the optimal scale and offset parameter values. The multiply-and-add unit 188 can be implemented, for example, as two separate units (for multiplication and addition) or a fused multiply-add unit. The value bin.b i (where 1 ≤ i ≤ K - 1) refers to the optimal bin boundary value (corresponding to the optimal vector component b b [i], see the previous subsection), which is hard-coded in the circuit 174a. For completeness, scl i and offs i refer to the scale and offset coefficients, which are also hard-coded in the circuit 174a.
[0147] Conversely, in the case of a required arbitrary function implementation, an addressable memory storing all the required bin boundaries, as well as the scaling and offset coefficients, can be used, as Figure 5 assumed in. This memory can be reprogrammed. Such an embodiment also involves the priority network 182 of the comparator and the multiply-and-add unit 188, as Figure 4As shown. However, the addressable memory 175b is used to store the bin boundaries for each desired function, as well as the scale and offset coefficients. In this case, the memory also provides a selection of the optimal scale and offset parameter values through its output decoder. bin.b i 、scl i and offs i values refer to the bin boundary values, scale coefficients, and offset coefficients, as Figure 4 shown. However, in this case, an additional value ("bin.b.bits") is required, which corresponds to (K - 1) times the number of bits used to store the bin boundaries. Additionally, the addresses of the bin boundaries (bin.b.addresses) must be passed to the memory so that the memory can retrieve the corresponding values.
[0148] Figure 4 and 5 the circuit 175a and the memory cell 175b shown can be mapped to the processing element 185 as Figure 3 shown. The neural processing device 15 is preferably implemented as a crossbar array 15 ( Figure 2 ). All components and devices required in the system 1 are preferably co-integrated on the same chip, as Figure 2 assumed. Thus, the system 1 can be assembled in a single device, including the crossbar array 15, the LUT circuit 17, and the processing unit 18, where the processing unit 18 is preferably configured as a near-memory processing unit. Additionally, the device 1 can include an input unit 11 to apply an input signal encoding the input vector components to the crossbar array 15. Further, the device 1 generally includes a readout circuit 16 and an I / O unit 19 to connect the system 1 to an external computer ( Figure 2 not shown in
[0149] Figure 1 illustrates a network 5 involving multiple systems 1 (e.g., an integrated device as Figure 2 shown). That is, the system 1 forms part of a larger computer system 5, including a server 2 that interacts with a client 4, which can be a natural person (interacting through a personal computer 3), a process, or a machine. In this example, each hardware system 1 is configured to read data from and write data to the storage unit of the server computer 2. Client requests are managed by the unit 2, which can be particularly configured to map a given computational task to vectors and weights and then pass them to the system 1. The entire computer system 5 can, for example, be configured as a composable disaggregated infrastructure, which can also include other hardware acceleration devices such as application-specific integrated circuits (ASICs) and / or field-programmable gate arrays (FPGAs).
[0150] Of course, many other architectures can be envisioned. For example, the present system 1 can be configured as a stand-alone system or a computerized system connected to one or more general-purpose computers. The system 1 can be particularly useful in distributed computing systems, such as edge computing systems.
[0151] 3. Final Remarks
[0152] The computerized device and system 1 can be designed to implement the embodiments of the present invention described herein, including the methods. In this regard, it can be understood that the methods described herein are substantially non-interactive, that is, automated. The automated part of such methods can be implemented only in hardware or in a combination of hardware and software. In an exemplary embodiment, the automated part of the methods described herein is implemented in software, as a service or an executable program (e.g., an application program), which is executed by a suitable digital processing device. However, all embodiments described herein relate to computational steps performed by means of non-conventional hardware, such as hardware-implemented LUTs, neural processing devices (such as crossbar array structures), and stand-alone processing units, preferably configured as near-memory processing units with respect to the neural processing device.
[0153] In addition, the methods described herein can also involve executable programs, scripts, or more generally, any form of executable instructions. The required computer-readable program instructions can be downloaded, for example, via a network (e.g., the Internet) from a computer-readable storage medium to a processing element.
[0154] Aspects of the present invention are particularly described herein with reference to flowcharts and block diagrams. It should be understood that each block or combination of blocks in the flowcharts and block diagrams can be implemented by computer-readable program instructions. The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of the system 1 according to various embodiments of the present invention, as well as the methods of operating them.
[0155] Although the present invention has been described with reference to a limited number of embodiments, variations, and drawings, those skilled in the art will understand that various changes can be made without departing from the scope, and equivalents can be substituted. In particular, features (similar devices or similar methods) shown in a given embodiment, variation, or drawing can be combined with or replaced by another feature in another embodiment, variation, or drawing without departing from the scope of the present invention. Therefore, various combinations of the features described with respect to any of the above embodiments or variations are contemplated, which are still within the scope of the appended claims. Additionally, many minor modifications can be made to adapt a particular situation or material to the teachings of the present invention without departing from the scope of the present invention. Accordingly, the present invention is not limited to the specific embodiments disclosed, but the present invention will include all embodiments falling within the scope of the appended claims. Additionally, many other variations are contemplated in addition to those specifically addressed above with respect to variations. For example, other types of storage element devices, selection circuits, LUT circuits, and processing units are contemplated.
Claims
1. A hardware system (1) designed to implement an artificial neural network, the system (1) comprising: A neural processing device (15) configured to implement M artificial neurons, where M ≥ 1; One or more lookup table circuits (17) configured to implement lookup tables; And M' processing units (18), where M ≥ M' ≥ 1, and each of the M' processing units: Is connected through at least one of the M artificial neurons to access a first value output by each of the at least one neuron; Is connected to a lookup table circuit (175, 175a, 175b) among the one or more lookup table circuits (17) to access parameter values of a parameter group in the lookup table; and Is configured to output a second value corresponding to a value of a mathematical function with the first value as an independent variable, the mathematical function being determined by the parameter group, and the parameter values of the parameter group being accessed by each processing unit from the lookup table circuit (175, 175a, 175b) during operation.
2. The hardware system (1) according to claim 1, wherein each processing unit (18) is configured to output the second value by: Selecting the parameter group according to the first value; and Performing an operation based on the first value and the parameter values of the selected parameter group to output the second value.
3. The hardware system (1) according to claim 2, wherein each processing unit (18) is further configured to: Access bin boundaries from the lookup table circuit (175, 175a, 175b); and Select the parameter group by comparing the first value with the accessed bin boundaries to identify a given bin in the bin containing the first value, whereby the parameter group is selected according to the given bin during operation.
4. The hardware system (1) according to claim 3, wherein Each processing unit (18) includes at least one comparator circuit (182), and the at least one comparator circuit (182) is designed to: Compare the first value with the bin boundaries; and Send a selection signal encoding the selected parameter group for the processing unit (18) to access the corresponding parameter values based on the sent signal.
5. The hardware system (1) according to claim 4, wherein The comparator circuit (182) is configured as a multi-layer q-ary tree comparison circuit designed to enable multi-layer comparison, where for one or more layers in the multi-layer, q is greater than or equal to three.
6. The hardware system (1) according to claim 4 or 5, wherein The lookup table circuit (175a) is a circuit with hard-coded parameter values, and Each processing unit (18) includes at least one multiplexer (186), and the at least one multiplexer (186) is connected to: A corresponding one of the at least one comparator circuit (182) to receive the selection signal; and The lookup table circuit (175a) to retrieve the corresponding parameter values according to the selection signal.
7. The hardware system (1) according to claim 4 or 5, wherein the lookup table circuit (175b) includes addressable storage units, the addressable storage units: being connected to the comparator circuit (182) to receive the selection signal; and being configured to retrieve parameter values of a selected parameter set according to the received selection signal.
8. The hardware system (1) according to any one of claims 4 to 7, wherein the mathematical function is a piecewise-defined polynomial function which is a polynomial on each of its sub-domains, the sub-domains corresponding to bins respectively, the selected parameter set corresponds to polynomial parameters of the piecewise-defined polynomial function, and each of the processing units includes an arithmetic unit (188), the arithmetic unit (188) being connected to the output of the lookup table circuit and being designed to perform the operation to calculate the second value as an arithmetic operation.
9. The hardware system (1) according to claim 8, wherein the arithmetic unit (188) includes a multiply-and-add circuit.
10. The hardware system (1) according to any one of claims 1 to 9, wherein the one or more lookup table circuits (17) include M' different circuits (175) which are respectively mapped to the M' processing units (18), where M ≥ M' > 1.
11. The hardware system (1) according to any one of claims 1 to 10, wherein the neural processing device (15) includes a crossbar array structure (15) which includes N input lines (151) and M output lines (152) arranged in rows and columns, where N > 1 and M > 1, such that the neural processing device (15) implements a layer of M neurons, the input lines and the output lines are interconnected via storage elements (156), and each of the M output lines (152) is connected to at least one of the M' processing units (18).
12. The hardware system (1) according to any one of claims 1 to 11, wherein the one or more lookup table circuits (17), the processing circuit (18) and the neural processing device (15) are co-integrated in the same chip (10).
13. The hardware system (1) according to any one of claims 1 to 12, wherein M > 1, preferably M ≥ 256, more preferably M ≥ 512, M' = 1 or 2, and each of the M' processing units is a vector processing unit (18) including b processing elements to operate on a one-dimensional array of dimension b.
14. A method of operating a hardware system (1), the method comprising: providing (S10) the hardware system (1), the hardware system (1) including: a neural processing device (15) configured to implement M artificial neurons, where M ≥ 1, M' processing units (18), each processing unit being connected by at least one of the M artificial neurons, and one or more lookup table circuits (17) implementing a lookup table; Operate (S20 - S50) the neural processing device (15) to obtain (S40) M first values respectively generated by the M artificial neurons; and Via the M' processing units (18), for each of the M first values, obtain (S60 - S110) an output value of a mathematical function that takes the first value as an independent variable and is based on an operand of parameter values including the first value and a parameter set, and the mathematical function is also determined by the parameter set, wherein the parameter values are retrieved (S100) from the one or more look-up table circuits (17).
15. The method according to claim 14, wherein for each of the first values, the output value is obtained by the following steps: Select (S70 - S80) the parameter set according to the first value, and Perform (S110) an operation based on the first value and the parameter values retrieved (S100) according to the selected parameter set.
16. The method according to claim 15, wherein the parameter set is selected by the following steps: Compare (S70) the first value with bin boundaries of bins to identify a given bin in the bin that contains the first value, wherein the bin boundaries are retrieved (S70) from the one or more look-up table circuits (17); and Select (S80) the parameter set according to the identified given bin.
17. The method according to claim 16, wherein before operating (S20 - S50) the neural processing device (15), the method further includes: Determine (S5) one or more groups of bin boundaries respectively according to one or more reference functions.
18. The method according to claim 17, wherein For each of the one or more reference functions, determine (S5) the bin boundaries of each group so as to minimize the number of bins or the maximum error between the approximate values of each reference function calculated based on the parameter values and theoretical values of the reference function.
19. The method according to any one of claims 16 to 18, wherein The mathematical function is a piecewise-defined polynomial function, the parameter set includes two or more polynomial coefficients, and the operations performed consist of arithmetic operations.
20. The method according to claim 19, wherein The mathematical function includes a set of linear polynomials, each linear polynomial corresponding to a respective bin, and the parameter set corresponding to each linear polynomial consists of a scaling coefficient and an offset coefficient.
21. The method according to claim 19 or 20, wherein The arithmetic operations (S110) are performed by a multiply-and-add circuit (188).
22. The method according to any one of claims 14 to 21, wherein the method further includes: Program (S20) the one or more look-up table circuits implementing the look-up table (17) such that the mathematical function can be one of the following: Activation function, Normalization function, Reduction function, State update function, Classification function, and Prediction function.
Citation Information
Cited By
Multi-core heterogeneous systems, methods, and media for machine learning models
CN121352075A