Binary neural network on programmable integrated circuit
By designing the hardware neuron layer on a programmable integrated circuit, using synchronous or (XNOR) operations and bit counting operations, an efficient binary neural network hardware implementation is achieved, solving the problem of software dependence in the prior art and improving processing speed and energy efficiency.
Patent Information
- Application Number
- CN202510122551.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2016-08-05
- Filing Date
- 2017-07-24
- Publication Date
- 2025-06-03
AI Technical Summary
In the prior art, the implementation of binary neural networks mainly relies on software and is difficult to efficiently implement on programmable integrated circuits.
A hardware neuron layer is designed, including logic circuits, counting circuits and comparison circuits, and the hardware implementation of binary neural networks is realized through synonymous or (XNOR) operations and bit counting operations.
A binary neural network that operates efficiently on programmable integrated circuits is realized, improving the processing speed and energy efficiency of input data sets.
Smart Images

Figure CN120087429A_ABST
Abstract
Description
[0001] This application is a divisional application of the patent application with the application date of July 24, 2017, application number 201780062276.4, and invention title "Binary Neural Network on a Programmable Integrated Circuit". Technical Field
[0002] Examples of the present invention generally relate to electronic circuits, and more particularly, to binary neural networks on a programmable integrated circuit (IC). Background Art
[0003] The use of programmable integrated circuits (ICs) (such as field programmable gate arrays (FPGAs)) to configure neural networks has regained attention. Currently, FPGA implementations of neural networks mainly focus on floating-point or fixed-point multiply-accumulate typically based on systolic array architectures. Recently, it has been demonstrated that even large modern machine learning problems can be solved by using neural networks with binary-represented weights and activations, while achieving high accuracy. However, the implementation of binary neural networks has hitherto been limited to software. Summary of the Invention
[0004] This application describes techniques for implementing a binary neural network on a programmable integrated circuit (IC). In an embodiment, a neural network circuit implemented in an integrated circuit (IC) includes a hardware neuron layer, the layer including a plurality of inputs, a plurality of outputs, a plurality of weights, and a plurality of thresholds. Each hardware neuron includes: a logic circuit having an input that receives a first logic signal from at least a portion of the plurality of inputs and an output that provides a second logic signal, where the second logic signal corresponds to an exclusive-NOR (XNOR) operation of the first logic signal and at least a portion of the plurality of weights; a counting circuit having an input that receives the second logic signal and an output that provides a count signal, the count signal representing the number of the second logic signals having a predetermined logic state; and a comparison circuit having an input that receives the count signal and an output that provides a logic signal, the logic signal having a logic state that represents a comparison between the count signal and one of the plurality of thresholds; wherein the logic signal output by the comparison circuit of each hardware neuron is provided as a corresponding one of the plurality of outputs.
[0005] In another embodiment, a method of implementing a neural network in an integrated circuit (IC) includes implementing a hardware neuron layer in the IC, the layer including a plurality of inputs, a plurality of outputs, a plurality of weights, and a plurality of thresholds; and in each of a plurality of neurons: receiving a first logic signal from at least a portion of the plurality of inputs and providing a second logic signal, the second logic signal corresponding to an exclusive-NOR (XNOR) operation of the first logic signal and at least a portion of the plurality of weights; receiving the second logic signal and providing a count signal, the count signal representing the number of the second logic signals having a predefined logic state; and receiving the count signal and providing a logic signal having a logic state representing a comparison between the count signal and one of the plurality of thresholds.
[0006] In another embodiment, a programmable integrated circuit (IC) includes a programmable structure configured to implement: a hardware neuron layer including a plurality of inputs, a plurality of outputs, a plurality of weights, and a plurality of thresholds. Each hardware neuron includes: a logic circuit having an input for receiving a first logic signal from at least a portion of the plurality of inputs and an output for providing a second logic signal, where the second logic signal corresponds to an exclusive-NOR (XNOR) operation of the first logic signal and at least a portion of the plurality of weights; a counting circuit having an input for receiving the second logic signal and an output for providing a count signal, the count signal representing the number of the second logic signals having a predetermined logic state; and a comparison circuit having an input for receiving the count signal and an output for providing a logic signal having a logic state representing a comparison between the count signal and one of the plurality of thresholds. The logic signal output by the comparison circuit of each of the hardware neurons is provided as a corresponding one of the plurality of outputs.
[0007] Reference will be made to the following detailed description to understand the above and other aspects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In order to understand the manner in which the above-recited features can be obtained, a more particular description, briefly summarized above, may be had by reference to the specific embodiments, some of which are illustrated in the accompanying drawings. It should be noted, however, that the drawings illustrate only typical exemplary embodiments and are not to be considered limiting of its scope.
[0009] Figure 1 is a block diagram showing the mapping of a binary neural network onto a hardware implementation in a programmable integrated circuit (IC) according to one embodiment.
[0010] Figure 2 is a block diagram showing a layer circuit of a neural network circuit according to one embodiment.
[0011] Figure 3 is a block diagram showing a hardware neuron according to an embodiment.
[0012] Figure 4 is a block diagram showing a layer circuit of a neural network circuit according to an embodiment.
[0013] Figure 5 is a block diagram showing the specialization of a hardware neuron according to an embodiment.
[0014] Figure 6 is a block diagram showing an embodiment of folding a layer of a binary neural network at a macro level (referred to as "macro-level folding").
[0015] Figure 7 is a flowchart showing a method of macro-level folding according to an embodiment.
[0016] Figure 8 is a block diagram showing an embodiment of folding a neuron of a binary neural network at a micro level (referred to as "micro-level folding").
[0017] Figure 9 shows a field programmable gate array (FPGA) architecture having a neural network implementation hardware layer according to an embodiment.
[0018] For ease of understanding, wherever possible, the same reference numerals are used to denote the same elements common to the drawings. It is understood that elements of one embodiment may be beneficially included in other embodiments. Detailed Description
[0019] The various features are described below with reference to the drawings. It should be noted that the drawings may or may not be drawn to scale, and elements of similar structure or function in all the drawings are denoted by the same reference numerals. It should be noted that the drawings are only for facilitating the description of the features. They are not an exhaustive description of the claimed invention or a limitation on the scope of the claimed invention. Additionally, the illustrated embodiments need not include all aspects or advantages shown. Aspects or advantages described in connection with a particular embodiment need not be limited to that embodiment and may be implemented in any other embodiment even if not so shown or explicitly stated.
[0020] The present invention introduces an efficient hardware implementation of a binary neural network. In one embodiment, the hardware implementation is mapped to the architecture of a programmable integrated circuit (IC), such as a field programmable gate array (FPGA). This implementation maps a large number of neurons that process input data with high computational intensity. For example, assuming that the complete neural network can be fully unfolded inside the FPGA, the neural network can classify an input data set (e.g., an image) at a clock frequency. Assuming a conservative clock frequency of 100 MHz, the neural network implementation can classify the input data set at 100 million data sets per second (e.g., images per second). In other embodiments, a certain amount of folding is performed to implement the neural network. The hardware implementation of the neural network introduced in this application is fundamentally different from the previous implementation in an FPGA that maps a floating-point neural network to a systolic array on a general-purpose processing engine. This floating-point implementation can currently process 100 - 300 data sets per second. In addition, the neural network implementation introduced in this application consumes less power compared to the previous floating-point implementation in an FPGA and the implementation using a graphics processing unit (GPU).
[0021] As described below, the benefit of a binary neural network is that standard floating-point multiply-accumulate operations become exclusive NOR operations and bit counting operations. The infrastructure includes multiple layers that can be fully connected or partially connected (e.g., convolution, pooling, etc.). Each layer includes many hardware neurons. Each hardware neuron calculates the exclusive NOR (XNOR) of all data inputs and the corresponding weights, counts the number of logical "1" bits in the result, and compares the count with a threshold. When the bit count is greater than the threshold, the hardware neuron returns "true" (logical "1"), otherwise it returns "false" (logical "0"). The fully connected and partially connected layers differ in how the hardware neurons receive input data. For a fully connected layer, the input is broadcast to all hardware neurons. For a partially connected layer, each hardware neuron operates on a part of the input data. For example, in a convolutional layer, each hardware neuron operates on a sequence of parts of the input data to produce a corresponding activation sequence. According to this implementation, a pooling layer (e.g., a max pooling layer) can be used to down-sample the output of the previous layer.
[0022] In some embodiments, the weights and thresholds may not be updated frequently (e.g., only after retraining the network). Therefore, the weights and / or thresholds can be "hardened" to generate dedicated hardware neurons. The specialization of the network should remove all resource requirements related to the XNOR operation and / or comparison operation. In addition, when implemented in an FPGA, the specialized network can consume significantly less routing resources than the non-specialized network implementation.
[0023] In some embodiments, due to resource constraints, large networks may not be fully expandable for implementation. For example, routing limitations in FPGAs may prevent the efficient implementation of large fully-connected layers with a large number of synapses. Thus, in some examples, the network can be folded onto the hardware architecture. The folding can be achieved by varying the granularity. At a macroscopic level, an entire layer can be iteratively folded onto the architecture. Alternatively, some neurons can be folded onto the same hardware neuron, and some synapses can be folded onto the same connection (reducing routing requirements). The above and other aspects of the implementation of the binary neural network are described below with reference to the accompanying drawings.
[0024] Figure 1 is a block diagram showing the mapping of a binary neural network 101 onto a hardware implementation in a programmable integrated circuit (IC). The binary neural network includes one or more layers 102. Each layer 102 includes neurons 104 having synapses 106 (e.g., inputs) and activations 108 (e.g., outputs). The inputs to layer 102 include binary inputs 110, binary weights 112, and a threshold 114. The output of layer 102 includes a binary output 116. Each binary input 110 is a 1-bit value (e.g., logic "0" or logic "1"). Similarly, each binary weight 112 is a 1-bit value. Each threshold 114 is an integer. Each binary output 116 is a 1-bit value. Each neuron 104 receives one or more binary inputs 110 through a corresponding one or more synapses 106. Each neuron 104 also receives one or more binary weights 112 and a threshold 114 for the synapses 106. Each neuron 104 provides a binary output 116 through an activation 108. A neuron 104 is "activated" when the binary inputs, binary weights, and threshold of the neuron 104 cause its binary output to be asserted (e.g., logic "1").
[0025] Some layers 102 can be "fully connected", such as Figure 2As shown. In a fully connected layer, a given set of binary inputs 110 or a given set of activations 108 from the previous layer are broadcast to all neurons 104. That is, each neuron 104 operates on the same set of inputs. Some layers 102 can be "partially connected". In a partially connected layer, each neuron 104 operates on a part of a given set of binary inputs 110 or a part of a given set of activations 108 from the previous layer. Different neurons 104 can operate on the same or different parts of a set of inputs. The binary inputs 110 can represent one or more data sets, such as one or more images. In an image embodiment, the input to the fully connected layer is the entire image, and the input to the partially connected layer is a part of the image. Some layers 102 can be "convolutional" layers. In a convolutional layer, each neuron 104 operates on a sequence of parts of a given set of binary inputs 110 or a sequence of parts of a given set of activations 108 from the previous layer.
[0026] In this embodiment, the binary neural network 101 is mapped to the hardware of the programmable IC 118. One embodiment of the programmable IC is a field programmable gate array (FPGA). An example architecture of the FPGA is described below. Although the binary neural network 101 is described as being mapped to the hardware in the programmable IC, it should be understood that the binary neural network 101 can also be mapped to the hardware of any type of IC (e.g., application specific integrated circuit (ASIC)).
[0027] The hardware implementation of the binary neural network 101 includes layer circuits 120. Each layer circuit 120 includes input connections 122, hardware neurons 124, and output connections 126. The layers 102 are mapped to the layer circuits 120. The synapses 106 are mapped to the input connections 122. The neurons 104 are mapped to the hardware neurons 124. The activations 108 are mapped to the output connections 126. The layer circuits 120 can be generated using circuit design tools based on the description of the layers 102, where the circuit design tools run on a computer system. The circuit design tools can generate configuration data for configuring the layer circuits 120 in the programmable IC 118. Alternatively, for a non-programmable IC (e.g., ASIC), the circuit design tools can generate mask data for manufacturing an IC with the layer circuits 120.
[0028] For example, as Figure 3As shown, the programmable IC 118 includes circuitry to assist in the operation of the layer circuitry 120, including a memory circuit 128 and a control circuit 130. The memory circuit 128 can store binary inputs, binary weights, thresholds, and binary outputs for the layer circuitry 120. The memory circuit 128 can include registers, first-in-first-out (FIFO) circuits, and the like. The control circuit 130 can control writes to and reads from the memory circuit 128. In other embodiments, the binary inputs can come directly from an external circuit, and the outputs can be transmitted directly to an external circuit. When, as Figure 5 shown, the layer is specialized, the weights and thresholds can be hard-coded as part of the logic. The memory circuit 128 can be on-chip or off-chip.
[0029] Figure 2 is a block diagram showing a layer circuitry 120A of a neural network circuit according to one embodiment. The layer circuitry 120A is an example of a fully connected layer and can be used as Figure 1 any of the layer circuitries 120 shown in. The layer circuitry 120A includes hardware neurons 202-1 to 202-Y (commonly referred to as hardware neurons 202), where Y is a positive integer. The data inputs (e.g., synapses) of each hardware neuron 202 receive logic signals providing binary inputs. The binary inputs can be a set of binary inputs 110 or a set of activations from a previous layer. In this embodiment, each hardware neuron 202 includes X data inputs (e.g., X synapses) for receiving X logic signals, where X is a positive integer. Each of the X logic signals provides a single binary input.
[0030] Each hardware neuron 202 includes a weight input that receives a logic signal providing a binary weight. Each neuron 202 can receive the same or different sets of binary weights. Each hardware neuron 202 receives X binary weights corresponding to the X data inputs. Each hardware neuron 202 also receives a threshold, which can be the same as or different from other hardware neurons 202. Each hardware neuron 202 generates a logic signal as an output (e.g., a 1-bit output). The hardware neurons 202-1 to 202-Y together output Y logic signals providing Y binary outputs (e.g., Y activations).
[0031] Figure 3FIG. 0 is a block diagram showing a hardware neuron 202 according to an embodiment. The hardware neuron 202 includes an exclusive-NOR (XNOR) circuit 302, a counting circuit 304, and a comparison circuit 306. The XNOR circuit 302 receives X logic signals providing binary inputs and X logic signals providing binary weights. The XNOR circuit 302 outputs X logic signals to provide corresponding XNOR results of the binary inputs and the binary weights. The counting circuit 304 is configured to count the number of signals having a predefined logic state among the X logic signals output by the XNOR circuit 302 (e.g., the number of logical "1"s output by the XNOR circuit 302). The counting circuit 304 outputs a counting signal providing a count value. The number of bits of the counting signal is log 2 (X)+1. For example, if X is 256, the counting signal includes 9 bits. The comparison circuit 306 receives the counting signal and a threshold signal providing a threshold. The threshold signal has the same width as the counting signal (e.g., log 2 (X)+1). The comparison circuit 306 outputs a logic signal providing a 1-bit output value. When the count value is greater than the threshold, the comparison circuit 306 enables the output logic signal (e.g., sets it to logical "1"), otherwise de-asserts the output logic signal (e.g., sets it to logical "0").
[0032] Regardless of the type of layer, the operations of each hardware neuron in the layer are the same. As Figure 3 shown, each hardware neuron uses the XNOR circuit 302 to calculate the XNOR between the binary input and the binary weight. The counting circuit 304 counts the number of predefined logic states in the XNOR operation result (e.g., the number of logical "1"s). The comparison circuit 306 compares the count value with the threshold and enables or de-asserts its output accordingly.
[0033] Figure 4 FIG. 13 is a block diagram showing a layer circuit 120B of a neural network circuit according to an embodiment. The layer circuit 120B is an embodiment of a partially connected layer or a convolutional layer, and can be used as Figure 1 any layer 120 as shown. Figure 4 The embodiment is different from the Figure 3 embodiment in that each hardware neuron 202 receives a separate set of binary inputs (i.e., the binary inputs are not broadcast to all hardware neurons 202). The sets of binary inputs can be a part of a set of binary inputs 110 or a part of a set of activations from the previous layer. For clarity, Figure 4 the threshold inputs of each hardware neuron 202 are omitted in Figure 4 The hardware neuron 202 shown in Figure 3 has the same structure as the neuron shown in
[0034] For example, layer circuit 120B can be used as a convolution layer. In this embodiment, the input data set (e.g., an image) can be divided into a set of N input feature maps, each having a height H and a width W. The corresponding binary weights can be divided into N*M groups, each having K*K binary weight values. In this embodiment, the value X can be set to K*K and the value Y can be set to M*W*H. A sequence of binary input groups is provided to each hardware neuron 202 to implement a convolution sequence. The sequence can be controlled by control circuit 130 by controlling memory circuit 128.
[0035] Figure 5 is a block diagram illustrating the specialization of a hardware neuron according to one embodiment. In this embodiment, the XNOR circuit 302 includes an XNOR gate for each of the X data inputs. The i-th XNOR gate outputs result [i] = input [i] XNOR weight [i]. In another example, the XNOR circuit 302 can be specialized by hardening the weights. The weights can be fixed instead of variable, and depending on the fixed state of the corresponding weight value, the XNOR gate can be reduced to an identity gate (e.g., a pass-through gate) or a NOT gate (e.g., an inverter). As Figure 5 As shown, the XNOR circuit 302 may include circuits 502-1 to 502-X, each receiving a logic signal providing a data input value. Each circuit 502 may be an identity gate or a NOT gate according to a corresponding weight value. For example, if the i-th weight is a logic "0", the i-th circuit 502-i is a NOT gate (e.g., when the data input [i] is a logic "1", the result [i] is a logic "0", and vice versa, thereby implementing the XNOR function when the weight [i] = logic "0"). If the i-th weight is a logic "1", the i-th circuit 502-i is an identity gate (e.g., when the data input [i] is a logic "1", the result [i] is a logic "1", and vice versa, thereby implementing the XNOR function when the weight [i] = logic "1"). This embodiment of a hardware neuron can be implemented when the weight value is known a priori.
[0036] In an embodiment, the binary neural network 101 is implemented without folding (e.g., "unfolded") using layer circuits 120. In the case of folding, each layer 102 of the binary neural network 101 is implemented using a corresponding layer circuit 120 in the programmable IC 118. In other embodiments, the binary neural network 101 is implemented using some type of folding.
[0037] Figure 6FIG. 0 is a block diagram illustrating an embodiment of folding layer 102 of binary neural network 101 at a macro level (referred to as "macro-level folding"). In the example shown, programmable IC 118 includes a single layer circuit 120, and binary neural network 101 includes n layers 102-1 to 102-n, where n is an integer greater than 1. Layers 102-1 to 102-n are time-multiplexed onto layer circuit 120. That is, layer circuit 120 is configured to implement layer 102-1, then implement layer 102-2, and so on until layer 102-N is implemented. The results generated by layer circuit 120 are saved for use as the input for the next implementation of layer 102. In other embodiments, programmable IC 118 may include more than one layer circuit 120, and layer 102 may be folded across multiple layer circuits 120.
[0038] Figure 7 FIG. 4 is a flowchart illustrating method 700 for macro-level folding according to one embodiment. Method 700 may be implemented using programmable IC 118 as shown in Figure 1 FIG. 0. Method 700 begins at step 702, where control circuit 130 loads a first set of binary weights and a first set of thresholds corresponding to the first layer of binary neural network 101 into memory circuit 128 for use by layer circuit 120. At step 704, layer circuit 120 processes the binary input. At step 706, control circuit 130 stores the output of layer circuit 120 in memory circuit 128. At step 708, control circuit 130 loads the next set of binary weights and the next set of thresholds corresponding to the next layer of binary neural network 101 into memory circuit 120 for use by layer circuit 120. At step 710, layer circuit 120 processes the binary output generated by layer circuit 120 for the previous layer. At step 712, control circuit 130 determines whether there are more layers 102 of binary neural network 101 to be implemented. If so, method 700 returns to step 706 and repeats. Otherwise, method 700 proceeds to step 714, where the binary output of layer circuit 120 is provided as the final output.
[0039] Figure 8 FIG. 10 is a block diagram illustrating an embodiment of folding neurons 104 of binary neural network 101 at a micro level (referred to as "micro-level folding"). In the embodiment shown, multiple sets of binary inputs, binary weights, and thresholds are provided as inputs to hardware neuron 124 (e.g., n sets of inputs). That is, n neurons 104 are folded onto hardware neuron 124. Each of the n neurons is represented by a given set of binary inputs, binary weights, and thresholds. Hardware neuron 124 generates a set of binary outputs for each set of inputs (e.g., n sets of binary outputs for n input sets).
[0040] Figure 9 shows the architecture of the FPGA 900, which includes a large number of different programmable tiles. The programmable tiles include multi-gigabit transceivers ("MGT") 1, configurable logic blocks ("CLB") 2, block random access memories ("BRAM") 3, input / output modules ("IOB") 4, configuration and clock logic ("CONFIG / CLOCKS") 5, digital signal processing modules ("DSP") 6, dedicated input / output modules ("I / O") 7 (e.g., configuration ports and clock ports), and other programmable logic 8, such as digital clock managers, analog-to-digital converters, system monitoring logic, etc. Some FPGAs also include dedicated processor blocks ("PROC") 10. The FPGA 900 can be used as Figure 1 the programmable IC 118 shown in. In this case, the layer circuit 120 is implemented using the programmable structure of the FPGA 900.
[0041] In some FPGAs, each programmable tile may include at least one programmable interconnect element ("INT") 11, and the INT 11 has connections to the input and output terminals 20 of the programmable logic elements within the same tile, as Figure 9 shown in the top embodiment. Each programmable interconnect element 11 may also include connections to the interconnect segments 22 of adjacent programmable interconnect elements within the same tile or other tiles. Each programmable interconnect element 11 may also include connections to the interconnect segments 24 of the general routing resources between the logic blocks (not shown). The general routing resources may include routing channels between the logic blocks (not shown), and the routing channels include paths of interconnect segments (e.g., interconnect segment 24) and switch blocks (not shown) for connecting the interconnect segments. The interconnect segments (e.g., interconnect segment 24) of the general routing resources may span one or more logic blocks. The programmable interconnect element 11, together with the general routing resources, implements the programmable interconnect structure ("programmable interconnect") for the shown FPGA.
[0042] In this embodiment, the CLB 2 may include configurable logic elements ("CLE") 12 that can be programmed to implement user logic and a single programmable interconnect element ("INT") 11. In addition to one or more programmable interconnect elements, the BRAM 3 may include BRAM logic elements ("BRL") 13. Generally, the number of interconnect elements included in a tile depends on the height of the tile. In the illustrated embodiment, the BRAM tile has a height of five CLBs, but other numbers (e.g., four) may also be used. In addition to an appropriate number of programmable interconnect elements, the DSP tile 6 may also include DSP logic elements ("DSPL") 14. In addition to one instance of the programmable interconnect element 11, the IOB 4 may also include two instances of, for example, input / output logic elements ("IOL") 15. As will be apparent to those skilled in the art, the I / O pads actually connected to, for example, the I / O logic element 15 are generally not limited to the area of the input / output logic element 15.
[0043] In the illustrated embodiment, a horizontal region near the center of the die (as Figure 9 shown) is used for configuration, clock, and other control logic. Vertical columns 9 extending from this horizontal region or column are used to distribute clock and configuration signals across the width of the FPGA.
[0044] Utilizing Figure 9 Some FPGAs with the illustrated architecture include additional logic blocks that disrupt the regular columnar structure that makes up most of the FPGA. The additional logic blocks can be programmable blocks and / or dedicated logic. For example, the processor block 10 spans several columns of CLBs and BRAMs. The processor block 10 can include a variety of components, ranging from a single microprocessor to a complete programmable processing system including a microprocessor, memory controller, peripherals, etc.
[0045] Note that Figure 9 The aim is to show an exemplary FPGA architecture. For example, the number of logic blocks in a row, the relative width of the rows, the number and order of the rows, the type of logic blocks included in the rows, the relative size of the logic blocks, and Figure 9 the interconnect / logic implementation at the top are all purely exemplary. For example, in an actual FPGA, there are typically more than one adjacent row of CLBs included anywhere a CLB appears to facilitate the efficient implementation of user logic, but the number of adjacent CLB rows varies with the overall size of the FPGA.
[0046] In one embodiment, a neural network circuit implemented in an integrated circuit (IC) is provided. Such a circuit may include: a hardware neuron layer, the layer including a plurality of inputs, a plurality of outputs, a plurality of weights, and a plurality of thresholds, each hardware neuron including: a logic circuit having an input that receives a first logic signal from at least a portion of the plurality of inputs and an output that provides a second logic signal, wherein the second logic signal corresponds to an exclusive-NOR (XNOR) operation of the first logic signal and at least a portion of the plurality of weights; a counting circuit having an input that receives the second logic signal and an output that provides a count signal, the count signal representing the number of the second logic signals having a predetermined logic state; and a comparison circuit having an input that receives the count signal and an output that provides a logic signal, the logic signal having a logic state that represents a comparison between the count signal and one of the plurality of thresholds; wherein the logic signal output by the comparison circuit of each of the hardware neurons is provided as a corresponding one of the plurality of outputs.
[0047] In some such circuits, the first logic signal input to the logic circuit of each of the hardware neurons is the plurality of inputs to the layer.
[0048] In some such circuits, the layer is configured to receive a plurality of weight signals that provide the plurality of weights, wherein the logic circuit in each of the hardware neurons is an XNOR circuit.
[0049] In some such circuits, the logic circuit in at least a portion of the hardware neurons is a pass circuit or an inverter depending on a fixed value of each of the plurality of weights.
[0050] In some such circuits, the neural network includes a plurality of layers, and wherein the circuit further includes: a control circuit configured to sequentially map each of the plurality of layers to the hardware neuron layer over time.
[0051] Some such circuits may further include: a control circuit configured to sequentially provide each set of input data from a plurality of sets of input data to the plurality of inputs of the layer over time.
[0052] Some such circuits may further include: a control circuit configured to sequentially provide each set of weights from a plurality of sets of weights as the plurality of weights of the layer over time.
[0053] In some such circuits, the hardware neuron layer can be a first hardware neuron layer, where the multiple inputs, the multiple outputs, the multiple weights, and the multiple thresholds are respectively a first multiple of inputs, a first multiple of outputs, a first multiple of weights, and a first multiple of thresholds, and where the circuit further comprises: a second hardware neuron layer, the second layer comprising a second multiple of inputs, a second multiple of outputs, a second multiple of weights, and a second multiple of thresholds; wherein the second multiple of inputs of the second layer are the first multiple of outputs of the first layer.
[0054] Some such circuits can further comprise: a memory circuit disposed between the first multiple of outputs of the first layer and the second multiple of inputs of the second layer.
[0055] In another embodiment, a method of implementing a neural network in an integrated circuit (IC) can be provided. Such a method can comprise: implementing a hardware neuron layer in the IC, the layer comprising multiple inputs, multiple outputs, multiple weights, and multiple thresholds; in each of a plurality of neurons: receiving a first logic signal from at least a portion of the multiple inputs and providing a second logic signal, the second logic signal corresponding to an exclusive-NOR (XNOR) operation of the first logic signal and at least a portion of the multiple weights; receiving the second logic signal and providing a count signal, the count signal representing the number of the second logic signals having a predefined logic state; receiving the count signal and providing a logic signal, the logic signal having a logic state representing a comparison of the count signal with one of the multiple thresholds.
[0056] In some such methods, the first logic signal input to each of the hardware neurons is the multiple inputs to the layer.
[0057] In some such methods, the layer is configured to receive multiple weight signals that provide the multiple weights, where each of the hardware neurons comprises an XNOR circuit.
[0058] In some such methods, at least a portion of the hardware neurons comprise a pass circuit or an inverter depending on a fixed value of each of the multiple weights.
[0059] In some such methods, the neural network comprises multiple layers, and the method further comprises: mapping each of the multiple layers to the hardware neuron layer sequentially over time.
[0060] Some such methods can further comprise: providing each of multiple sets of input data to the multiple inputs of the layer sequentially over time.
[0061] Some such methods may further include: sequentially providing each of the multiple sets of weights over time as the multiple weights of the layer.
[0062] In some such methods, the hardware neuron layer may be a first hardware neuron layer, wherein the multiple inputs, the multiple outputs, the multiple weights, and the multiple thresholds are respectively a first multiple of inputs, a first multiple of outputs, a first multiple of weights, and a first multiple of thresholds, and wherein the method further includes:
[0063] Implementing a second hardware neuron layer, the second layer including a second multiple of inputs, a second multiple of outputs, a second multiple of weights, and a second multiple of thresholds; wherein the second multiple of inputs of the second hardware neuron layer are the first multiple of outputs of the first layer.
[0064] Some such methods may further include: a memory circuit disposed between the first multiple of outputs of the first layer and the second multiple of inputs of the second layer.
[0065] In another embodiment, a programmable integrated circuit (IC) is provided. Such an IC may include: a programmable structure configured to implement:
[0066] A hardware neuron layer, the layer including multiple inputs, multiple outputs, multiple weights, and multiple thresholds, each hardware neuron including: a logic circuit having an input that receives a first logic signal from at least a portion of the multiple inputs and an output that provides a second logic signal, wherein the second logic signal corresponds to an exclusive NOR (XNOR) operation of the first logic signal and at least a portion of the multiple weights; a counting circuit having an input that receives the second logic signal and an output that provides a count signal, the count signal representing the number of the second logic signals having a predetermined logic state; and a comparison circuit having an input that receives the count signal and an output that provides a logic signal, the logic signal having a logic state that represents a comparison of the count signal with one of the multiple thresholds; wherein the logic signal output by the comparison circuit of each hardware neuron is provided as a corresponding one of the multiple outputs.
[0067] In some such ICs, the layer may be configured to receive multiple weight signals that provide the multiple weights, and the logic circuit in each hardware neuron may be an XNOR circuit.
[0068] Although the foregoing is directed to specific embodiments, other and further examples may be designed without departing from its basic scope, and the scope of the invention is determined by the appended claims.
Claims
1. A circuit of a binary neural network implemented in an integrated circuit IC, characterized in that, the binary neural network includes one or more layers, each layer includes neurons, and the neurons have synapses and activations, and the circuit includes: A layer circuit of hardware neurons, the layer circuit includes a plurality of input connections, a plurality of output connections, a plurality of weights and a plurality of thresholds, wherein the layer circuit is configured to implement a plurality of layers in a configuration folded at a macro level or a configuration folded at a micro level, wherein implementing a plurality of layers in a configuration folded at a macro level includes folding the plurality of layers onto the layer circuit, and implementing a plurality of layers in a configuration folded at a micro level includes folding the neurons of the plurality of layers onto the same hardware neuron; wherein folding the plurality of layers onto the layer circuit includes time-division multiplexing the plurality of layers onto the layer circuit, and folding the neurons of the plurality of layers onto the same hardware neuron includes time-division multiplexing the neurons of the plurality of layers onto the same hardware neuron; wherein the synapses are mapped to the plurality of input connections, the neurons are mapped to the hardware neurons, and the activations are mapped to the output connections; wherein, for a fully connected layer implemented in the layer circuit, input data is broadcast to each of the hardware neurons, and for a partially connected layer implemented in the layer circuit, each of the hardware neurons operates on a part of a given set of input data; each hardware neuron includes: a logic circuit, the logic circuit has an input for receiving a first logic signal from at least a part of the plurality of input connections and an output for providing a second logic signal, wherein the second logic signal corresponds to an exclusive-NOR (XNOR) operation of the first logic signal and at least a part of the plurality of weights; a counting circuit, the counting circuit has an input for receiving the second logic signal and an output for providing a counting signal, the counting signal representing the number of the second logic signals having a predetermined logic state; and a comparison circuit, the comparison circuit has an input for receiving the counting signal and an output for providing a logic signal, the logic signal having a logic state representing a comparison between the counting signal and one of the plurality of thresholds; wherein the logic signal output by the comparison circuit of each of the hardware neurons is provided as a corresponding one of the plurality of output connections; wherein when the layer circuit is configured to implement a plurality of layers in a configuration folded at a macro level, the circuit further includes a control circuit, the control circuit is configured to sequentially load binary weights and thresholds into a memory circuit for use by the hardware neuron layer to process binary inputs over time, thereby time-division multiplexing the plurality of layers onto the layer circuit; and wherein when the layer circuit is configured to implement a plurality of layers in a configuration folded at a micro level, the circuit further includes a control circuit, the control circuit is configured to sequentially input binary inputs, binary weights, and the threshold of the neurons into the same hardware neuron over time, thereby time-division multiplexing the plurality of layers onto the layer circuit.
2. The circuit according to claim 1, wherein, the first logic signal of the logic circuit input to each of the hardware neurons is the plurality of inputs to the layer.
3. The circuit according to any one of claims 1 or 2, wherein, the layer is configured to receive a plurality of weight signals, the plurality of weight signals providing the plurality of weights, wherein the logic circuit in each of the hardware neurons is an XNOR circuit.
4. The circuit according to any one of claims 1-2, wherein, the logic circuit in at least a portion of the hardware neurons is a pass circuit or an inverter depending on a fixed value of each of the plurality of weights.
5. A method for implementing a binary neural network in an integrated circuit IC, wherein, the binary neural network includes one or more layers, each layer including neurons having synapses and activations, and the method includes: implementing a layer circuit of hardware neurons in the IC, the layer circuit including a plurality of input connections, a plurality of output connections, a plurality of weights, and a plurality of thresholds, wherein the layer circuit implements the plurality of layers in a macro-level folded configuration or a micro-level folded configuration, wherein implementing the plurality of layers in a macro-level folded configuration includes folding the plurality of layers onto the layer circuit, and implementing the plurality of layers in a micro-level folded configuration includes folding the neurons of the layers of the plurality of layers onto the same hardware neuron; wherein folding the plurality of layers onto the layer circuit includes time-division multiplexing the plurality of layers onto the layer circuit, and folding the neurons of the layers of the plurality of layers onto the same hardware neuron includes time-division multiplexing the neurons of the layers of the plurality of layers onto the same hardware neuron; wherein the synapses are mapped to the plurality of input connections, the neurons are mapped to the hardware neurons, and the activations are mapped to the output connections; wherein, for a fully-connected layer implemented in the layer circuit, input data is broadcast to each of the hardware neurons, and for a partially-connected layer implemented in the layer circuit, each of the hardware neurons operates on a portion of a given set of input data in each of the plurality of hardware neurons: receiving a first logic signal from at least a portion of the plurality of input connections through an input of a logic circuit and providing a second logic signal through an output of the logic circuit, the second logic signal corresponding to an exclusive-NOR XNOR operation of the first logic signal and at least a portion of the plurality of weights; receiving the second logic signal through an input of a counting circuit and providing a count signal through an output of the counting circuit, the count signal representing the number of the second logic signals having a predefined logic state; and receiving the count signal through an input of a comparison circuit and providing a logic signal through an output of the comparison circuit, the logic signal having a logic state representing a comparison between the count signal and one of the plurality of thresholds. When implementing multiple layers in a configuration where the layer circuit is folded at a macroscopic level: the binary weights and thresholds are loaded into the memory circuit sequentially over time for use by the hardware neuron layer to process binary inputs, thus time-division multiplexing the multiple layers onto the layer circuit; and When implementing multiple layers in a configuration where the layer circuit is folded at a microscopic level: the binary inputs, binary weights, and the threshold of the neuron are input into the same hardware neuron sequentially over time, thus time-division multiplexing the multiple layers onto the layer circuit.
6. The method according to claim 5, wherein, the first logic signal input to each of the hardware neurons is the multiple inputs to the layer circuit.
7. The method according to claim 5, wherein, the layer circuit is configured to receive a plurality of weight signals, the plurality of weight signals providing the plurality of weights, wherein each of the hardware neurons includes an XNOR circuit.
8. The method according to claim 5, wherein, at least a portion of the hardware neurons includes a pass circuit or an inverter depending on a fixed value of each of the plurality of weights.