Probability signal generation element, probabilistic neuron, and neural network thereof

By using probabilistic signal generation elements and correlated signals within probabilistic neurons, the high computational costs of CNN processing are mitigated, achieving improved performance and energy efficiency in CNN implementations.

JP7689748B2Active Publication Date: 2025-06-09UNIV DE LAS ISLAS BALEARES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022549079
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-02-13
Filing Date
2021-02-11
Publication Date
2025-06-09
Estimated Expiration
2041-02-11

AI Technical Summary

Technical Problem

Existing CNN processing methods face high computational costs and power consumption due to the large number of operations required, especially when handling large amounts of data such as high-speed video processing.

Method used

The implementation of a probabilistic signal generation element and probabilistic neurons that utilize correlation between signals to efficiently implement convolutional neural networks, reducing the need for complex hardware and random number generators.

Benefits of technology

This approach significantly reduces hardware resource usage, simplifies the implementation of max pooling and convolution operations, and allows for the addition of deeper layers without increasing the number of random number generators, resulting in improved performance and energy efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007689748000002
    Figure 0007689748000002
  • Figure 0007689748000003
    Figure 0007689748000003
  • Figure 0007689748000004
    Figure 0007689748000004
Patent Text Reader

Abstract

The present invention provides a stochastic signal generating element comprising a first binary-to-stochastic converter having a first input for receiving a binary signal and a second input for receiving a random signal, in turn, and converting the binary signal into a first stochastic signal using the random signal, characterized in that the stochastic signal generating element comprises a processing unit having a first input for receiving the first stochastic signal and a second input for receiving a reference stochastic signal generated from a constant-valued signal using the random signal, the processing unit processing the reference stochastic signal according to at least one arithmetic function, the result of which is a stochastic output signal representative of the processing, the processing unit further comprising computational neurons implementing the stochastic signal generating element, and the processing unit further comprising a neural network implementing the computational neurons.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention will be included in the fields of electronic technology and computing intelligence. A first aspect of the present invention is composed of probabilistic signal generating elements, which are suitable for a computer system based on digital hardware and probabilistic computing techniques, for example, for use in a convolutional network of probabilistic neurons. Another aspect of the present invention is composed of probabilistic neurons, which include a neural network formed from probabilistic neurons (for example, a convolutional type known as CNN) and a probabilistic signal generating element, whereby high parallelism can be obtained, and as a result, pattern recognition can be executed very quickly.

Background Art

[0002] In recent years, the use of a wide range of neural networks (deep neural networks, DNNs) has gained great validity from their excellent ability to extract useful information from large amounts of data. Its hardware implementation enables an increase in execution speed when performing calculations in parallel, compared to methods based on the use of microprocessors using a sequential Von-Neuman type computing architecture (software solution). CNN is a type of DNN with feed-forward (no feedback) connections and is particularly useful for image shape recognition as seen in the following literature: - LeCun, Y., Bottou, L., Bengio, Y., Haffner, P. Gradient-based learning applied to document recognition (1998) Proceedings of the IEEE, 86 (11), pp. 2278-2323. - Lawrence, S., Giles, C.L., Tsoi, A.C., Back, A.D. Face recognition: A convolutional neural-network approach (1997) IEEE Transactions on Neural Networks, 8 (1), pp. 98-113.

[0003] Specifically, a CNN is a type of artificial neural network in which neurons are connected to receptive fields in a way very similar to neurons in the primary visual cortex of the biological brain, and the application is executed on a two-dimensional matrix. Therefore, it is very effective for artificial vision tasks such as image classification and segmentation among other applications. The application consists of multiple layers of artificial neurons that act as convolutional filters of one or more dimensions, and usually, after each layer, a function is added to perform a non-linear causal mapping.

[0004] Here, from the processed image (raw data obtained from the observed object), a series of calculations are performed using different neurons arranged in layers. Each neuron performs relatively simple processing on the input (which may be from the processed image or the input from other neurons), and the given output is composed of a non-linear function (called the activation function) of the weighted sum of the inputs, as shown by the following equation.

Equation

[0005] The above equation shows the dependence of the output (y j ) of the i-th neuron in relation to the sum of N inputs defined as x i . Similarly, it shows the selection of each of these inputs, the weight (w ij parameter) dedicated to each neuron, and the bias value (T i ) specific to each neuron. The activation function f is a non-linear function such as the Heaviside function, the sigmoid function, or a normalization function known as the Rectified Linear Unit (ReLU), and the input x j of the above equationEvaluate the maximum value between the linear combination of [[ID=]] and a reference value that may be zero.

[0006] In a CNN-type network, neurons are arranged in processing layers. Each neuron in the first layer receives information directly from the processed image, and the inner layers receive the responses of the neurons arranged in the preceding layer. The functionality of each layer is usually of two types: convolution and reduction. The convolutional layer is composed of convolving the preceding image using a kernel (often also called the impulse response). Mathematically, the convolutional layer can be considered as an expression of the projection of the input information onto a specific subspace defined by the kernel, and the information reduction is relatively small. The second functionality of the layer is composed of a radical information reduction process, where different regions of the processed image are selected, and the most dominant value, which may be the highest value of the signal (the type of the layer is called max pooling), or the average value of the signal (average pooling), is selected. The final result is the information reduced in the input signal from the preceding convolutional layer.

[0007] By connecting the two types of neural layers (convolution and pooling), the direction of the information of the input image changes, specific features of the same image are selected and the information is reduced, while the degree of abstraction increases. In this way, the input information bits can represent the intensity values of the processed image, and the category to be recognized in the output will be obtained.

[0008] Therefore, the main applications of CNNs would be image recognition and image processing for inference. In any case, due to its cross-cutting nature, the application fields of CNNs are huge, and currently there are thousands of scientific studies based on the use of neural processing architectures, including processes such as face recognition in the study by Lawrence et al. mentioned above, or text recognition as disclosed in the patent document EP 1598770 entitled "Low resolution optical character recognition for camera acquired documents".

[0009] One of the drawbacks of CNN processing is the high computational cost due to a large number of operations (successive convolutional and pooling processes) covering the entire generated image from the initial image. The computational cost is defined by the image processing time, as well as the power consumption associated with the process, and can become very high when applied to large amounts of data (such as in the case of high-speed video processing). In the case of conventional data processing systems (software solutions), general-purpose architectures based on the use of processors are not optimal for the implementation of CNNs, and as a result, very high consumption and response times are associated with certain applications. For this reason, hardware systems have been developed and attempts have been made to implement these processes and optimize the processing in parallel. Examples of attempts to optimize CNNs using hardware include the following publications. -Zhang, C., Li, P., Sun, G., Guan, Y., Xiao, B., Cong, J. "Optimizing FPGA-based accelerator design for deep convolutional neural networks" (2015) FPGA 2015 - 2015 ACM / SIGDA International Symposium on Field-Programmable Gate Arrays, pp. 161-170.

[0010] In this study, Zhang and co-authors presented an optimization methodology for CNNs and implemented it on a reconfigurable logic device (FPGA). Their method is based on the use of classical digital logic, and while parallelism increases in relation to implementations based on the use of general-purpose microprocessors, the inherent parallelism of CNNs cannot be fully exploited.

[0011] A plan has been developed with the aim of increasing parallelism and adapting the processing hardware to the most general DNN architectures possible, using non-conventional digital logic such as probabilistic computing. Probabilistic computing is digital logic in which the operations between signals follow probabilistic rules. This is the case for the research of the next publication where a type of deep network such as a deep belief network is applied, which is usually the type of network used for shape recognition and is inspired by statistical physics. -Sanni, K., Garreau, G., Molin, J.L., Andreou, A.G. "FPGA implementation of a Deep Belief Network architecture for character recognition using stochastic computation", (2015) 2015 49th Annual Conference on Information Sciences and Systems, CISS 2015, art. no. 7086904.

[0012] This research uses a linear approach for the non-linear sigmoid function, but when the neurons need to be arranged in the form of a CNN-type network, the implementation method of certain layers such as max pooling is not clear. On the other hand, only the use of uncorrelated signals is assumed in the implementation of the network, leaving the possibility of exploiting the correlation between probabilistic signals, which greatly helps in reducing the hardware.

[0013] Related to the implementation of non-conventional methodologies in hardware optimization, the following publications are available for convolutional neural networks. -Alawad, M., Lin, M. "Stochastic-Based Deep Convolutional Networks with Reconfigurable Logic Fabric" (2016) IEEE Transactions on Multi-Scale Computing Systems, 2 (4), art. no. 7547913, pp. 242-256. This study uses certain characteristics of probability theory, which are relationships such as the probability density function of the sum of two independent random variables and the probability densities of these variables individually (related to the convolution of both). The basis for accelerating CNN processing lies in implementing these probabilistic characteristics instead of using individual neural elements.

[0014] Other non-conventional methodologies based on the use of probabilistic characteristics are found in the following publications. -Ren, A., Li, Z., Ding, C., Qiu, Q., Wang, Y., Li, J., Qian, X., Yuan, B. "SC-DCNN: Highly-scalable deep convolutional neural network using stochastic computing" (2017) International Conference on Architectural Support for Programming Languages and Operating Systems - ASPLOS, Part F127193, pp. 405-418. In this study, probabilistic logic is used in the implementation of both convolution processing and max-pooling processing. On the other hand, a state machine is used for the implementation of the tangent bipolar function (function f in the aforementioned equation), which can considerably complicate the configuration (as seen in Figure 6). Also, the use of correlation signals is not exploited, nor is the implementation of the activation function simplified.

[0015] On the other hand, proposals in probabilistic neural networks are also found in the following recent publications. -Li, Z., Li, J., Ren, A., Cai, R., Ding, C., Qian, X., Draper, J., Yuan, B., Tang, J., Qiu, Q., Wang, Y. "HEIF: Highly Efficient Stochastic Computing-Based Inference Framework for Deep Neural Networks" (2019) IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, 38 (8), art. no. 8403283, pp. 1543-1556. In this study, as shown in Fig. 4c, a binary block known as an Approximate Parallel Counter (APC) is used. This is used to perform the weighted sum of neural inputs, but complex circuits are executed in the implementation of the ReLU-type activation function or the max-pooling block (in Fig. 6 of the said reference). In fact, the resulting activation function is not the classical ReLU but a "clipped ReLU", that is, a saturated ReLU.

[0016] Probabilistic logic (or probabilistic computing) replaces the quantity that is conventional binary logic with a logic in which the toggling frequency of bits encodes the sign, and forces such toggling to have a probabilistic nature. Probabilistic computing (SC) is an approximate computer methodology that represents signals using the toggling frequencies of time-dependent bit sequences. Each SC signal is composed of pulses in a bit sequence that represent the probability of finding a high value (logical '1'). For example, the numerical value 0.75 is represented by a bit sequence with a 75% probability of finding a logical '1' in the bit sequence, which is (1, 1, 0, 1) in 4 bits and (0, 1, 1, 0, 1, 1, 1, 1) in an 8-bit frame. This encoding is referred to as a unipolar representation, and each value is between 0 and 1. To include negative values, a different encoding is required, and similar to the bipolar encoding case, the 0 value is subtracted from the 1 value and finally divided by the total number of bits. This encoding is equivalent to the implementation of the transformation of the variable p* = 2p - 1 where 'p' is the unipolar representation of the number. Bipolar encoding gives a range of [-1, 1] as possible values.

[0017] One of the main advantages of using probabilistic logic is to implement complex functions with low-cost hardware resources and also depend on the correlation between signals. For example, an XNOR gate implements the multiplication in a bipolar code (f = x·y) when both signals are temporarily uncorrelated, and implements the absolute value of the difference between the two signals minus one (f = 1 - |x - y|) when they are correlated.

[0018] In the literature, different probabilistic neuron configurations have been proposed, but none that utilize the correlation of signals that can considerably simplify the CNN. In the present invention, the correlation between the output signals of neurons is used for an efficient implementation of a convolutional neural network. Summary of the Invention Means for Solving the Problems

[0019] The present invention mainly consists of a probabilistic signal generation element, a computing neuron comprising this element, particularly a probabilistic neuron, and similarly, a computing neural network comprising a plurality of such neurons. The proposed neuron is suitable for the implementation of deep artificial neural networks, particularly convolutional neural networks.

[0020] A first aspect of the present invention consists of a probabilistic signal generation element, which comprises a first Binary to Stochastic Converter (BSC). Usually, the BSC is composed of a binary two's complement comparator. This BSC sequentially comprises a first input for receiving a binary signal and a second input for receiving a random signal (random is understood to include either pure randomness or pseudo-randomness). It is configured to convert the binary signal into a first probabilistic signal from the random signal at its output.

[0021] This probabilistic signal generation element is characterized by comprising a processing unit. The processing unit sequentially has a first input for receiving the first probabilistic signal from the BSC and a second input for receiving a second probabilistic signal generated from a constant value signal (for example, a 0 value) and used as a reference probabilistic signal using the same random signal as that used by the BSC to generate the second probabilistic signal. This processing unit is configured to process the first probabilistic signal and the second probabilistic signal (reference probabilistic signal) from the BSC according to at least one arithmetic function and generate a probabilistic output signal representing the processing as a result.

[0022] A correlation between signals from different probabilistic signal generation elements can be obtained on the condition that the reference probabilistic signal received by the processing unit is generated using the same random signal as that used by the BSC.

[0023] As a result, it is possible to apply a simple arithmetic function such that the processing unit applies an OR-type or AND-type logic gate, thereby implementing a non-saturating activation function and obtaining a probabilistic output signal that can be correlated with other possible probabilistic signal generation elements.

[0024] Thus, as an exemplary embodiment of the probabilistic signal generation element of the first aspect of the present invention, the processing unit is an OR logic gate. Therefore, the arithmetic function applied by the same logic gate is composed of a maximum value function (max(a, b)), and the maximum value function is the basis for embodying the rectified linear unit (ReLU). As another exemplary embodiment, the processing unit is an AND logic gate. Therefore, the arithmetic function applied by the same logic gate is composed of an activation function of the min type (a, b).

[0025] For example, this probabilistic signal generation element may be part of a probabilistic neuron for a computational neural network, which is another aspect, specifically the second aspect of the present invention, where the probabilistic neuron includes a probabilistic signal generation element according to the first aspect.

[0026] The probabilistic neuron of the present invention preferably includes an approximate parallel counter (APC). The APC is arranged to receive a plurality of probabilistic input signals, add the plurality of input signals, and convert them into an output signal encoded in two's complement notation in binary as its output. This output signal becomes the binary input signal of the above-mentioned BSC of the probabilistic signal generation element.

[0027] Similarly, as an exemplary embodiment, the probabilistic neuron of the second aspect of the present invention includes a plurality of processing subunits, each arranged to receive an external probabilistic signal (such as a signal from another neuron) and a signal coupling weight signal, and to process these signals by applying an arithmetic function to generate an output signal. Since both signals received by the processing subunits are uncorrelated, in an exemplary embodiment, these subunits are composed of XNOR logic gates, and the function applied is the multiplication of both probabilistic signals in bipolar notation. As another possible embodiment, these subunits are composed of AND logic gates, and the function applied is the multiplication of both probabilistic signals in unipolar notation. Then, the output signals generated by these processing subunits are operatively connected to the APC so as to correspond to the probabilistic APC input signals.

[0028] Another aspect of the present invention, specifically the third aspect, is composed of a computational neural network implemented by a plurality of probabilistic neurons defined according to the second aspect of the above-described invention, and some are operatively interconnected with others.

[0029] Preferably, the computational neural network includes a second binary-probability converter, i.e., a BSC converter. The second BSC converter is designed to generate a probabilistic reference signal from a constant-value signal input and a random signal input and transmit it to the processing units of different neurons all at once. Similarly, as an exemplary embodiment, the neural network of the third aspect of the present invention includes a random number generator. The random number generator is configured to generate a random signal and transmit it to different binary-probability converters all at once, so that all the probabilistic signals generated by the APC of the neurons are correlated.

[0030] In most CNN implementations, the common operations used are multiplication, addition, and the max function or ReLU. These operations can be easily implemented in a probabilistic circuit when correlation is used accurately. The main advantages obtained by using correlation signals for the implementation of convolutional neural networks are: a) Saving the hardware resources used without degrading the accuracy of the results by not requiring the implementation of different random number generators for each neuron (in a probabilistic implementation, the largest proportion of resources is used for random / random number generation) b) Simplifying the implementation of the max pooling function and convolution that only includes simple computational units such as logic gates (various efforts and designs have been introduced in the literature and this function has become implementable in neural networks, but when using uncorrelated signals, the proposed designs require a considerable amount of hardware space) c) Enabling the addition of deeper layers in the neural network without the need to use each random number generator because the number of generators is constant regardless of the number of layers in the network (different from other implementations where the number of random number generators increases as the number of layers in the network increases) It is composed of

[0031] In an embodiment of a suitable computational neural network, the random number generator is of the linear feedback shift register (LFSR) type, which is a low-cost and simple implementation in terms of hardware resources.

[0032] As an exemplary embodiment, the computational neural network according to the third aspect of the present invention includes an OR group gate, which is configured to receive a probabilistic output signal from a group of probabilistic neurons and give its maximum value at its output, or is configured to implement a minimum pooling type function by replacing the OR group gate with an AND group gate.

[0033] As another exemplary embodiment, the computational neural network includes an array of binary-probability converters, which convert an initial signal, where each initial signal is received at a first input, and is configured to convert this signal into an initial probability signal using a random signal received at a second input for each output.

[0034] Preferably, the neural network of the present invention includes a second random number generator and an array of binary-probability converters. The converters of this array are configured to convert the combined weight signal received at the first input, convert this signal into a probabilistic weight signal as an output, and use the aforementioned random signal from the second generator for this purpose. In this way, uncorrelated signals can be used to generate the product of the input and the weight, and correlated signals can be used to generate the output transfer function of the neuron.

[0035] In the first embodiment, the neural network includes at least one max-pooling type layer, that is, it is implemented by neurons having probability signal generation elements, and the processing unit of the probability signal generation element is an OR logic gate. In the second embodiment, the neural network includes a minimum pooling type layer, that is, it is implemented by neurons having probability signal generation elements, and the processing unit of the probability signal generation element is an AND logic gate. And in a more preferred embodiment, the neural network is implemented as a convolutional neural network, and the convolutional neural network includes a plurality of max-pooling type layers and a plurality of minimum pooling type layers.

[0036] In 2D or 3D image processing, or in a convolutional neural network optimized for a general N-dimensional structure, the entire network is arranged in different layers of neurons. After specific features are extracted in the convolutional neural layer, a subsampling operation is usually applied to reduce the dimension of data processing in the next layer. As described, one of the most commonly used operations is the max-pooling block, in which subsampling is performed by extracting the maximum value of the output of each neuron, and each of these neurons is provided within the window of the convolutional neuron layer. In state-of-the-art implementations, a set of blocks (accumulator, comparator, and counter) is used to implement the max-pooling function, which requires a significant amount of resource usage. However, in the present invention, by using the correlation performed in all neurons of the network (by using the same random signal in all BSC converters of each neuron), it is possible to implement the maximum value classification function using a single OR logic gate. That is, it is possible to implement the max-pooling function (widely used in the implementation of deep neural networks) on the output signals of the neurons of the network layer through an OR gate, or to implement the minimum pooling type function through an AND gate.

Brief Description of the Drawings

[0037]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Best Mode for Carrying Out the Invention

[0038] Preferred embodiments of the different claimed inventions are described below with reference to FIGS. 1 to 6.

[0039] A schematic diagram of the probabilistic signal generation element that is the subject of the first aspect of the present invention is shown in FIG. 1. As can be seen, the probabilistic signal generation element (1) includes a first binary-probability converter (BSC), and this BSC sequentially includes a first input that receives a binary signal (A) and a second input that receives a random signal (R). At its output, the binary signal (A) is converted into a probabilistic signal (A*). Subsequently, the probabilistic signal generation element (1) includes a processing unit (11) that receives the probabilistic signal (A*) and a reference probabilistic signal (C*). The reference probabilistic signal (C*) is generated from a constant value signal (C) using the random signal (R). The output of this processing unit (11) is a probabilistic output signal (S*) that represents a processed arithmetic function.

[0040] Regarding the second aspect of the present invention, FIG. 2 shows a schematic diagram of a preferred embodiment of a probabilistic neuron (10), and the probabilistic neuron (10) includes a preferred embodiment of a probabilistic signal generation element (1), and the probabilistic signal generation element (1) includes an OR logic gate and is provided as a processing unit that implements a ReLU-type activation function at the output (S*). This embodiment of the probabilistic neuron (10) includes a plurality of XNOR logic gates, and each of them is arranged to receive an external probabilistic signal (X1* - Xn*) and a probabilistic weight signal (w1* - wn*). Considering that these two signals are uncorrelated, the function applied by the XNOR gate is composed of the multiplication of both probabilistic signals in bipolar codes.

[0041] Subsequently, the output of the XNOR gate of this embodiment of the probabilistic neuron (10) is computationally connected to the input of an approximate parallel counter (APC), and the approximate parallel counter (APC) sums the input signals (Y1* - Yn*) and converts their sum into a binary value encoded in two's complement at its output. In other words, the approximate parallel counter (APC) evaluates how many '1' signals are at the output of the XNOR logic gate and gives a binary signal of their sum at its output. Subsequently, this sum is transmitted to the binary-probabilistic converter (BSC) of the probabilistic signal generation element (1). As described above, this embodiment of the probabilistic signal generation element (1) includes an OR logic gate that implements a ReLU-type activation function and uses the same random signal (R) as that received by the binary-probabilistic converter (BSC) to output (S * )

[0042] Figure 3 schematically shows the main part of the third aspect of the present invention, which is an exemplary embodiment of the connection of two neurons of a neural network. This embodiment has two OR logic gates, is shown in Figure 2, and includes the two probabilistic neurons (10) described above. Similarly, it includes a second binary-probability converter (BSC2), and BSC2 is configured to generate a reference probability signal (C*) and transmit it to the OR logic gates of both probabilistic neurons (10). As can be seen, the reference probability signal (C*) is generated from a constant value signal (C) and a random signal (R) in the second binary-probability converter (BSC2).

[0043] As can be understood by those skilled in the art, the implementation of generating the correlation neural signal of the present invention is in contrast to the more complex state-of-the-art implementation of generating the ReLU function, and it also has the drawback of saturation, so it is not the standard ReLU used in typical machine learning processes.

[0044] As can be seen from Figure 4, a further effect of the present invention is that since all the outputs of the network neurons are correlated, an efficient implementation of the max-pooling function (widely used in the implementation of deep neural networks) from the group of probabilistic neurons (n 0 - n 3 ) of the network layer to the probabilistic output signals (S 0 * - S 3 *) is possible, and it is possible to give their maximum values as the output (S max *), or it is possible to implement a minimum-pooling type function by replacing the OR group gate (3) with an AND group gate.

[0045] Figure 5 is a block diagram showing a neural network in a more general aspect, and two random number generators (2, 2') are used. As can be seen, the first random number generator (2) is for the neurons (n 0 - n n , n’ 0 - n’ n) is used for the conversion of the input signal (x), the reference signal (0), and the output signal of the approximate parallel counter. The second random number generator (2’) is only used for the conversion to the probability of the weights (w) of the network.

[0046] As can be seen from the above, the output of each layer of the neural network of the present invention matches without the risk that the input signal of the neurons in the next layer is inaccurate. When generated by the first random number generator (2), this output can be multiplied by the probability weight signal (w * ) because this probability weight signal has no correlation with the input signal and is generated by an array of binary-probability converters (BSC array’) using random numbers from the second random number generator (2’).

[0047] As can be seen, this implementation of the probabilistic neuron enables the entire neural network, which was not conventionally conceivable, to be executed with only two random number generators (2, 2’) for the entire network, and simplifies the cost in terms of the digital gates implemented in the entire network.

[0048] FIG. 6 shows a comparison table of the performance between this implementation and another implementation of a deep neural network known as LeNet-5 (shown in the publication Y. LeCun, L. Bottou, Y. Bengio and P. Haffner: Gradient-Based Learning Applied to Document Recognition, Proceedings of the IEEE, 86(11):2278-2324, November 1998) in the field of machine learning. The hardware implementations compared in the proposed model are as follows. -FPGA16. S. I. Venieris and C. Bouganis, “fpgaconvnet: A framework for mapping convolutional neural networks on fpgas,” in 2016 IEEE 24th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM) , May 2016, pp. 40-47. -FPGA17a. Z. Liu, Y. Dou, J. Jiang, J. Xu, S. Li, Y. Zhou, and Y. Xu, “Throughput-optimized fpga accelerator for deep convolutional neural networks,”TRETS, vol. 10, pp. 17:1-17:23, 2017. -FPGA17b. Z. Li, L. Wang, S. Guo, Y. Deng, Q. Dou, H. Zhou, and W. Lu, “Laius: An 8-bit fixed-point cnn hardware inference engine,” in 2017 IEEE International Symposium on Parallel and Distributed Processing with Applications and 2017 IEEE International Conference on Ubiquitous Computing and Communications (ISPA / IUCC), Dec 2017, pp. 143-150. -FPGA18.S.-S. Park, K.-B. Park, and K. Chung, “Implementation of a cnn accelerator on an embedded soc platform using sdsoc,” 02 2018, pp.161-165.

[0049] For comparison, the LeNet-5 network is implemented on an FPGA. This CNN design aims to process a highly standardized database in machine learning of handwritten digit recognition (MNIST), which consists of 60,000 training images and 10,000 test images. The LeNet-5 type of CNN architecture is composed of two convolutional layers and three fully connected layers. The success rate of MNIST in this CNN is approximately 98.5%, and the hardware implementation used in the proposed neural model is 97.6% (slightly lower than the software case due to binarizing the signals and being a probabilistic calculation methodology).

[0050] The implementation is executed using an Arria 10 10AX115H2F34E1SG FPGA, operating at a clock frequency of 150 MHz, with an 8-bit accuracy for binary signals.

[0051] Thus, the proposed method is superior in performance and energy efficiency compared to other architectures. As a result, the implementation of the proposed probabilistic CNN achieves 27 times better performance (measured by inferences per second and per megahertz) compared to FPGA17a and 6.3 times better energy efficiency (measured by inferences per joule) compared to FPGA17b, showing promise for applications in embedded systems.

[0052] The proposed new neuron design reduces the overall consumption area of the system, and the proposed architecture fully and parallelly tunes a complete CNN on a single FPGA, which is in contrast to the other three implementations that only implement a part of the network and iterate based on loops (sequential implementation).

[0053] This collective parallelism is the main reason for the lower latency of the proposed design, as it does not require sequential iteration of the network.

[0054] The table in FIG. 6 also emphasizes the hardware resources required per task. The above implementation requires a large area by using specific hardware blocks such as RAM and Digital Signal Processing (DSP) blocks. In the present proposal, the DSP block is not used because a non-conventional computing technique (probabilistic computing) is used instead of classical binary logic. At the same time, since the computation is not executed recursively, a memory block is not required, reducing the main cause of power consumption, which is the data conversion with off-chip memory or the operation of memory access.

[0055] Characterizing the first aspect of the present proposal is a combination of a processing unit (11) using correlated probabilistic signals and a Binary-Stochastic Converter (BSC). Another aspect of the invention is a subsequent combination of an array of processing subunits such as XNOR gates and an Approximate Parallel Counter (APC) connected to a Binary-Stochastic Converter (BSC) implementing a probabilistic neuron (10). And the third aspect is to implement in a network equipped with a random number generator (2) common to all neurons. Considering that the random number generator (2) generates random signals (R) of both negative and positive numbers (the probability that the generated sign bit is equal to '0' or '1' is 50%), the probability signal at level 0 is composed of bits that randomly vary between 0 and 1, and each level has a probability of 50%. The gist of this invention is to correctly combine a probabilistic reference value and a signal from an Approximate Parallel Counter (APC), which is converted to a probability by a Binary-Stochastic Converter (BSC), and the same random signal (R) is used for the combination. Therefore, the probability signal at level 0 and the probability signal (A*) given by the Binary-Stochastic Converter (BSC) are completely correlated, and their combination, for example, in an OR gate, may give a maximum value signal.

Claims

1. A first input for receiving a binary signal (A), and a second input for receiving a random signal (R), in that order, and a first binary-probability converter (BSC) that converts the binary signal (A) into a first probability signal (A*) using the random signal (R). Comprising: A processing unit (11) comprising a first input for receiving the first probability signal (A*), and a second input for receiving a reference probability signal (C*) which is a second probability signal, wherein the reference probability signal (C*) is generated from a constant value signal (C) using the random signal (R), and which processes the first probability signal (A*) and the reference probability signal (C*) according to at least one arithmetic function to generate a probability output signal (S*) representing the processing. Comprising: The processing unit (11) is an OR-type logic gate, and the applied arithmetic function is composed of an activation function of the rectified linear unit (ReLU) type. Probability signal generation element (1).

2. A first input for receiving a binary signal (A), and a second input for receiving a random signal (R), in that order, and a first binary-probability converter (BSC) that converts the binary signal (A) into a first probability signal (A*) using the random signal (R). Comprising: A processing unit (11) comprising a first input for receiving the first probability signal (A*), and a second input for receiving a reference probability signal (C*) which is a second probability signal, wherein the reference probability signal (C*) is generated from a constant value signal (C) using the random signal (R), and which processes the first probability signal (A*) and the reference probability signal (C*) according to at least one arithmetic function to generate a probability output signal (S*) representing the processing. Comprising: The processing unit (11) is an AND-type logic gate, and the applied arithmetic function is composed of an activation function of the minimum value type (A*, C*). Probability signal generation element (1).

3. A probability neuron (10) for a computational neural network, comprising the probability signal generation element (1) according to claim 1 or claim 2.

4. Comprising an approximate parallel counter (APC) for receiving a plurality of probability input signals (Y1* - Yn*), summing the plurality of input signals (Y1* - Yn*), and converting the sum into a binary output signal encoded with two's complement at its output. ​ The output signal is the binary signal (A) input to the first binary - probability converter (BSC). The probabilistic neuron (10) for a computational neural network according to claim 3.

5. Comprising a plurality of processing sub - units, each processing sub - unit receiving an external probability signal (X1* - Xn*) and a probability weight signal (w1* - wn*), processing them by applying an arithmetic function, and generating an output signal that constitutes the input signal (Y1* - Yn*) of the approximate parallel counter (APC). The probabilistic neuron (10) for a computational neural network according to claim 4.

6. The processing sub - unit is composed of XNOR logic gates, and each XNOR logic gate bipolar - multiplies the external probability signal (X1* - Xn*) and the corresponding probability weight signal (w1* - wn*). The probabilistic neuron (10) for a computational neural network according to claim 5.

7. The processing sub - unit is composed of AND logic gates, and each AND logic gate unipolar - multiplies the external probability signal (X1* - Xn*) and the corresponding probability weight signal (w1* - wn*). The probabilistic neuron (10) for a computational neural network according to claim 5.

8. Comprising a plurality of probabilistic neurons (10) according to any one of claims 3 to 7, Some of the plurality of probabilistic neurons (10) are computationally interconnected with others. Computational neural network.

9. Comprising a second binary - probability converter (BSC2), the second binary - probability converter (BSC2) generating the reference probability signal (C*) from the constant value signal (C) and the random signal (R), and transmitting it to different processing units (11) of the plurality of probabilistic neurons (10) all at once. The computational neural network according to claim 8.

10. Comprising a random number generator (2), the random number generator (2) generating the random signal (R) and transmitting the random signal (R) to different first and second binary - probability converters (BSC, BSC2) of the plurality of probabilistic neurons (10) all at once. The computational neural network according to claim 9.

11. The random number generator (2) is of the linear feedback shift register type. The computational neural network according to claim 10.

12. Comprising an OR group gate (3), the OR group gate (3) receives a probabilistic output signal (S 0 * - S 3 *) from a group of probabilistic neurons (n 0 - n 3 ), and obtains the maximum value in its output (S max *). The computational neural network according to any one of claims 8 to 11.

13. Comprising an array of binary-probability converters (BSC array), wherein the array of binary-probability converters converts an initial signal (x) received at each first input and uses the random signal (R) received at each second input to convert them as their outputs into respective initial probability signals (x*). The computational neural network according to any one of claims 8 to 12.

14. Comprising a second random number generator (2') and an array of binary-probability weight converters (BSC array'), wherein the array of binary-probability weight converters converts a weight signal (w) received at a first input and uses the second random number signal received from the second random number generator (2') to convert it as its output into the probability weight signal (w*). The computational neural network according to any one of claims 8 to 13, citing claim 5.

15. Comprising a plurality of probability neurons (10) according to claim 1, Comprising a max-pooling type layer having a probability neuron (10) provided with a probability signal generation element (1) in which the processing unit (11) is an OR type logic gate. The computational neural network according to any one of claims 8 to 14.

16. Comprising a plurality of probability neurons (10) according to claim 2, Comprising a minimum-pooling type layer having a probability neuron (10) provided with a probability signal generation element (1) in which the processing unit (11) is an AND type logic gate. The computational neural network according to any one of claims 8 to 14.