An FPGA-based neural network system
By designing a neural network system based on FPGA, using counters and lookup tables to realize the search and calculation of weight values, the problem of hardware resource limitations in the existing technology is solved, and efficient big data processing needs are achieved.
Patent Information
- Application Number
- CN202311044555.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-18
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-08-18
AI Technical Summary
The implementation of existing fully connected neural networks is limited by hardware resources, which makes it impossible to effectively meet the needs of efficient big data processing.
A neural network system based on FPGA is designed, including an input layer, a multi-layer hidden layer and an output layer. The weight value search and calculation are realized through counters and lookup tables, and the calculation process is optimized using the data cache layer and activation function module.
It realizes efficient operation of fully connected neural networks in FPGAs, avoids hardware resource limitations, and meets the needs of efficient big data processing.
Smart Images

Figure CN116933848B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of neural network system design, and particularly to a neural network system based on FPGA. Background Art
[0002] Currently, the popular neural networks generally set the last layer in the form of a fully connected layer, that is, a fully connected neural network is formed. This is because the fully connected neural network has a relatively dense node connection, has a strong simulation ability for both classification and regression algorithms, and can be applied to many current application scenarios.
[0003] For the implementation of the fully connected neural network, the currently mainly used method is to set corresponding logic units in the Central Processing Unit (CPU) or the graphics processing unit (GPU) to implement the calculation logic of the fully connected neural network, thereby realizing the fully connected neural network.
[0004] However, the fully connected network model has problems of more parameters and longer calculation cycles, resulting in the existing methods being prone to making hardware resources become the bottleneck of its application-side deployment, so it can only run at a lower efficiency or run a smaller-scale model, and thus cannot well meet the current high-efficiency big data processing requirements. Summary of the Invention
[0005] Based on the above deficiencies of the prior art, this application provides a neural network system based on FPGA to solve the problem that the prior art is restricted by hardware resources and cannot well meet the high-efficiency big data processing requirements.
[0006] To achieve the above object, this application provides the following technical solutions:
[0007] The first aspect of this application provides a neural network system based on FPGA, including:
[0008] An input layer, multiple hidden layers connected in sequence, an output layer, and a clock for triggering the operation of each layer;
[0009] The input layer includes an input calculation layer and a data cache layer;
[0010] The input calculation layer includes a counter, a calculation module, and a lookup table corresponding to the input layer; wherein, the counter is used to sequentially output count values, and based on the count values, look up the weight values corresponding to the count values from the lookup tables corresponding to each layer; the calculation module calculates the calculation results of the input layer nodes by using the product of each input weight value found from the lookup table corresponding to the input layer and the corresponding input value;
[0011] The data cache layer is used to cache the calculation results of the input layer nodes and output the calculation results of the input layer nodes to the next layer;
[0012] Each hidden layer includes a hidden calculation layer and a data selection layer;
[0013] The hidden calculation layer includes a lookup table corresponding to the hidden layer and a product calculation module. When receiving the calculation results output by the previous layer each time, according to the count value received currently, the hidden weight values corresponding to the count value received currently are respectively looked up from the lookup table corresponding to the hidden layer, and the products of each hidden weight value and the received calculation results are respectively calculated through the product calculation module to obtain the calculation sub-results of each hidden layer node;
[0014] The data selection layer is used to cache the calculation sub-results of each hidden layer node, and when the count value is the maximum value, for each hidden layer node respectively, the activation function is used to calculate the sum of the calculation sub-results of the hidden layer node to obtain the calculation results of each hidden layer node, and the calculation results of each hidden layer node are output to the next layer in sequence;
[0015] The output layer includes an output calculation layer and an output latching layer; among them, the output calculation layer has the same structure as the hidden calculation layer; the output latching layer is used to cache the calculation sub-results of each output layer node, and when the count value is the maximum value, for each output layer node respectively, the calculation sub-results of the output layer node are summed to obtain the calculation results of each output layer node, and the calculation results of each output layer node are output as the system calculation results.
[0016] Optionally, in the above FPGA-based neural network system, there are multiple lookup tables corresponding to the input layer, and the number of lookup tables corresponding to the input layer is the sum of the total number of output layer nodes and 1; among them, the nth lookup table corresponding to the input layer stores the input weight values corresponding to the nth count value of each output layer node.
[0017] Optionally, in the above FPGA-based neural network system, the calculation module includes:
[0018] A summation module and an activation function module; among them, the summation module accumulatively adds the products of each found input weight value and the corresponding input value, and the activation function module uses the activation function to calculate the accumulation result to obtain the calculation result of the input layer node.
[0019] Optionally, in the above FPGA-based neural network system, the data cache layer includes:
[0020] a data selection module, a BRAM module, and a data latch module;
[0021] Among them, when the data selection module receives the calculation result of any one of the input layer nodes, if the currently received count value is not greater than n, it transmits the received calculation result of the input layer node to the BRAM module for caching. If the count value output by the currently received counter is greater than n, it transmits 1 to the BRAM module for caching, and transmits the calculation result in the BRAM module to the next layer through the data latch module.
[0022] Optionally, in the above FPGA-based neural network system, there are multiple lookup tables corresponding to the hidden layer, and the number of lookup tables corresponding to the hidden layer is the total number of hidden layer nodes; one lookup table corresponding to the hidden layer stores the hidden weight values corresponding to each count value corresponding to one hidden layer node.
[0023] Optionally, in the above FPGA-based neural network system, the data selection layer includes:
[0024] multiple cache activation modules, first latches corresponding to each cache activation module, and a selector connecting each first latch;
[0025] Each cache activation module respectively receives the count value and the calculation sub-results of the corresponding hidden layer node, accumulates the respective calculation sub-results of the corresponding hidden layer node based on the count value, and when the count value is the maximum value, calculates the current accumulated result using an activation function to obtain the calculation result of the corresponding hidden layer node;
[0026] The first latches corresponding to each cache activation module cache the calculation results of the hidden layer nodes output by the corresponding cache activation module, and transmit the calculation results of the hidden layer nodes to the selector;
[0027] The selector sequentially transmits the calculation results of the hidden layer nodes in each first latch to the next layer according to the current count value.
[0028] Optionally, in the above FPGA-based neural network system, the cache activation module includes:
[0029] a control output module, a BRAM module, and an activation function module;
[0030] Among them, the control output module is used to receive the count value and the calculation sub-result of the corresponding hidden layer node. And every time a calculation sub-result of a hidden layer node is received, the currently received calculation sub-result of the hidden layer node is accumulated with the current accumulated result stored in the BRAM module to obtain a new accumulated result, and the new accumulated result is updated to the BRAM module. And when the received count value is the maximum value, the activation function module is controlled to calculate the latest accumulated result in the BRAM module by using an activation function to obtain the calculation result of the hidden layer node, and the calculation result of the hidden layer node is output to the corresponding first latch.
[0031] Optionally, in the above FPGA-based neural network system, the output latch layer includes:
[0032] A plurality of cache control modules and second latches corresponding to each of the cache control modules;
[0033] Among them, each cache control module includes a BRAM module and a control output module; the BRAM module is used to cache the calculation sub-results of the corresponding output layer nodes; the control output module receives the count value, and when the count value is the maximum value, sums up the calculation sub-results of the output layer nodes in the BRAM module to obtain the calculation result of the corresponding output layer node, and outputs the calculation result of the corresponding output layer node to the corresponding latch for output.
[0034] An embodiment of the present application provides a neural network system based on an FPGA, including an input layer, multiple hidden layers connected in sequence, an output layer, and a clock for triggering the operation of each layer. The input layer includes an input calculation layer and a data cache layer. The input calculation layer includes a counter, a calculation module, and a lookup table corresponding to the input layer. Among them, the counter is used to sequentially output count values, and based on the count values, weight values corresponding to the count values are retrieved from the lookup tables corresponding to each layer. The calculation module calculates the calculation results of the input layer nodes by using the products of the respective input weight values retrieved from the lookup table corresponding to the input layer and the corresponding input values. The data cache layer is used to cache the calculation results of the input layer nodes and output the calculation results of the input layer nodes to the next layer. Each hidden layer includes a hidden calculation layer and a data selection layer. The hidden calculation layer includes a lookup table corresponding to the hidden layer and a product calculation module, which is used to, when receiving the calculation results output by the previous layer each time, retrieve the hidden weight values corresponding to the currently received count values from the lookup table corresponding to the hidden layer according to the currently received count values, and calculate the products of the respective hidden weight values and the received calculation results through the product calculation module to obtain the calculation sub-results of the respective hidden layer nodes. The data selection layer is used to cache the calculation sub-results of the respective hidden layer nodes, and when the count value is the maximum value, for each hidden layer node respectively, calculate the sum of the respective calculation sub-results of the hidden layer node by using an activation function to obtain the calculation results of the respective hidden layer nodes, and sequentially output the calculation results of the respective hidden layer nodes to the next layer. The output layer includes an output calculation layer and an output latch layer. Among them, the output calculation layer has the same structure as the hidden calculation layer. The output latch layer is used to cache the calculation sub-results of the respective output layer nodes, and when the count value is the maximum value, for each output layer node respectively, sum up the respective calculation sub-results of the output layer node to obtain the calculation results of the respective output layer nodes, and output the calculation results of the respective output layer nodes as the system calculation results, thereby realizing the setting of the hardware resources of the FPGA corresponding to the composition of the functional modules of the fully connected neural network in the FPGA, and these hardware resources can implement the corresponding functions of the functional modules of the fully connected neural network, and further realize the fully connected neural network in the FPGA, avoiding being restricted by hardware resources, and thus can well meet the current high-efficiency big data processing requirements. Description of the Drawings
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.
[0036] Figure 1Schematic diagram of an FPGA-based neural network system provided by an embodiment of the present application;
[0037] Figure 2 Schematic diagram of an input calculation layer provided by an embodiment of the present application;
[0038] Figure 3 Schematic diagram of another input calculation layer provided by another embodiment of the present application;
[0039] Figure 4 Schematic diagram of a data cache layer provided by this embodiment;
[0040] Figure 5 Schematic diagram of a hidden calculation layer provided by an embodiment of the present application;
[0041] Figure 6 Schematic diagram of another hidden calculation layer provided by another embodiment of the present application;
[0042] Figure 7 Schematic diagram of a hidden layer node calculation provided by an embodiment of the present application;
[0043] Figure 8 Schematic diagram of a data selection layer provided by an embodiment of the present application;
[0044] Figure 9 Schematic diagram of a cache activation module provided by an embodiment of the present application;
[0045] Figure 10 Schematic diagram of an output latch layer provided by an embodiment of the present application. Detailed implementation manners
[0046] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0047] In this application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0048] An embodiment of this application provides a neural network system based on an FPGA, such as Figure 1 , including:
[0049] An input layer, multiple hidden layers connected in sequence, an output layer, and a clock that triggers the operation of each layer.
[0050] It should be noted that the input layer is the first layer of the system, each hidden layer is the middle layer of the system, and the output layer is the last layer of the system. And the entire system is triggered and executed by a clock signal.
[0051] Also refer to Figure 1 , the input layer includes an input calculation layer and a data cache layer. Each hidden layer includes a hidden calculation layer and a data selection layer. The output layer includes an output calculation layer and an output latch layer.
[0052] For the input layer, in the embodiment of this application, specifically as Figure 2 shown, the input calculation layer includes a counter, a calculation module, and a lookup table corresponding to the input layer.
[0053] Among them, the counter is used to sequentially output count values (Count) to look up the weight values corresponding to the count values from the lookup tables corresponding to each layer based on the count values. That is, although the counter is in the input layer, the hidden layer and the output layer share this counter with the input layer, so as to realize the unified scheduling of each sub-module through the count values output by the counter.
[0054] Moreover, since the calculation of the system is continuously carried out, the counter continuously generates count values in a loop, that is, when the count value reaches the maximum value, it will return to generate the minimum value again.
[0055] It should be noted that each layer has its own corresponding lookup table, and the lookup table in the embodiment of this application needs to be mapped to the specific structure LUT of the FPGA, rather than all being occupied as registers, which will cause an explosion of register resources and ultimately make it impossible to deploy the network.
[0056] Specifically, the calculation module mainly calculates the calculation result of a node in the input layer by using the product of each input weight value found from the lookup table corresponding to the input layer and the corresponding input value, that is, each count value in the input layer corresponds to a calculation result.
[0057] The data cache layer is used to cache the calculation results of the nodes in the input layer and output the calculation results of the nodes in the input layer to the next layer.
[0058] It should be noted that in the embodiments of the present application, multiple input values are usually input in the input layer, and these input values are input to each node in the input layer. Therefore, at this time, each node in the input layer will process these input values, and correspondingly, each input value will correspond to a weight value of the node in the input layer.
[0059] Optionally, as Figure 3 shown, in another embodiment of the present application, there are multiple lookup tables corresponding to the input layer, and the number of lookup tables corresponding to the input layer is the sum of the total number of nodes in the output layer and 1.
[0060] Among them, the nth lookup table corresponding to the input layer stores the input weight values corresponding to the nth count value of each node in the output layer.
[0061] Therefore, when the calculation module calculates, it needs to find the input weight values corresponding to the current count value from each lookup table corresponding to the input layer, and then perform the calculation.
[0062] Optionally, in another embodiment of the present application, it can also be seen from Figure 3 that the calculation module includes:
[0063] A summation module and an activation function module.
[0064] Among them, the summation module accumulates the products of the found input weight values and the corresponding input values respectively, and then the activation function module uses the activation function to calculate the accumulated result, so as to obtain the calculation result of the node in the input layer corresponding to the current count value.
[0065] Optionally, as Figure 4 shown, the data cache layer provided in another embodiment of the present application specifically includes:
[0066] A data selection module, a BRAM module, and a data latch module.
[0067] Among them, the data selection module is used to implement data screening. Specifically, when receiving the calculation result of any input layer node, if the currently received count value is not greater than n, the calculation result of the received input layer node is transmitted to the BRAM module for caching. If the count value output by the currently received counter is greater than n, 1 is transmitted to the BRAM module for caching, so as to implement the augmentation operation of the input.
[0068] The data latch module is used to transmit the calculation result in the BRAM module to the next layer. Specifically, every time the BRAM module caches a calculation result, the calculation result is transmitted to the next layer, that is, transmitted to the hidden layer connected to it.
[0069] For the hidden layer, as Figure 5 shown, in the embodiment of the present application, the hidden calculation layer in the hidden layer includes a lookup table corresponding to the hidden layer and a product calculation module.
[0070] Optionally, as Figure 6 shown, in order to be adapted to the hidden layer nodes, in another embodiment of the present application, there are multiple lookup tables corresponding to one hidden layer, and the number of lookup tables corresponding to the hidden layer is the total number of hidden layer nodes in the hidden layer.
[0071] Moreover, a lookup table corresponding to one hidden layer stores the hidden weight values corresponding to each count value corresponding to one hidden layer node. That is, different from a lookup table corresponding to the input layer that stores the weight values of each node, a lookup table corresponding to the hidden layer only stores the weight values of one node and stores all its weight values.
[0072] Moreover, as Figure 6 shown, in the embodiment of the present application, there are also multiple product calculation modules. Each lookup table corresponding to the hidden layer corresponds to one product calculation module, that is, each hidden layer node corresponds to one product calculation module, which is used to calculate the output value of the previous layer, that is, the product of the current input value and the weight value corresponding to the current count value found in the lookup table corresponding to the hidden layer node.
[0073] It should be noted that, through Figure 6 it can be seen that in the embodiment of the present application, the structure of the hidden calculation layer is different from that of the input calculation layer. The main reason is that the input values of the input layer arrive simultaneously. That is, as Figure 3 shown, all input values can reach an input layer node simultaneously, so it can calculate the calculation result of the input layer node at one time. However, for the hidden calculation layer, its input is the output of the previous layer, and the output of the previous layer is output sequentially. Therefore, the input of the hidden calculation layer arrives sequentially. Therefore, as Figure 7As shown, each hidden layer node receives only one input each time, so only part of the calculation can be completed.
[0074] Specifically, the hidden calculation layer is used to, when receiving the calculation result output by the previous layer each time, according to the count value received currently, respectively look up the hidden weight values corresponding to the currently received count value from the lookup table corresponding to the hidden layer, and calculate the product of each hidden weight value and the received calculation result through the product calculation module, to obtain the calculation sub-results of each hidden layer node.
[0075] It should be noted that the structures of different hidden layers are the same, and each includes a hidden calculation layer and a data selection layer. Moreover, the structures of the hidden calculation layers of different hidden layers are also the same, that is, each includes a corresponding lookup table and a product calculation module. However, the specific information stored in the lookup tables of different hidden calculation layers is different.
[0076] Specifically, for the input calculation layer, what it looks up are the weight values corresponding to each of an input layer node and the current count value, so that the calculation result of this input layer node can be directly calculated. Therefore, the calculation layer respectively looks up a weight value corresponding to each hidden layer node and the current calculation value, which is used to calculate with the currently received calculation result, so as to complete the partial calculation of each hidden layer node.
[0077] The data selection layer is used to cache the calculation sub-results of each hidden layer node, and when the count value is the maximum value, for each hidden layer node respectively, use the activation function to calculate the sum of the calculation sub-results of the hidden layer node, to obtain the calculation results of each hidden layer node, and sequentially output the calculation results of each hidden layer node to the next layer.
[0078] It should be noted that since only part of the calculation of each hidden layer node is completed each time, it is necessary to cache the calculation sub-results of each hidden layer node. After receiving all the input values of a round of calculation and calculating to obtain all the calculation sub-results, then the summary calculation of all the calculation sub-results can be performed, so as to obtain the final calculation results of each hidden layer node. And a round of calculation is controlled by the count value of the counter. Therefore, when the count value is the maximum value, the input value received at this time is the last input value. Thus, after calculating the last calculation sub-result at this time, for each hidden layer node respectively, use the activation function to calculate the sum of the calculation sub-results of the hidden layer node, to obtain the calculation results of each hidden layer node, and sequentially output the calculation results of each hidden layer node to the next layer.
[0079] Optionally, for the sum of the respective calculation sub-results of the hidden layer nodes, it can be that after obtaining the last calculation sub-result, all the calculation sub-results are summed up, or it can be in an accumulative manner, where each received calculation sub-result is accumulated to obtain the final sum.
[0080] It should also be noted that since the next layer of a hidden layer is also a hidden layer or an output layer, and the structures of the hidden layers are the same, and the calculation layers of the output layer and the hidden layer calculation layers are also the same, the inputs need to come one by one. And in order to make the inputs correspond to the count values, although all the hidden layer nodes can be calculated simultaneously in the end, they need to be output one by one.
[0081] Optionally, in another embodiment of the present application, as Figure 8 shown, the data selection layer in the hidden layer may specifically include:
[0082] Multiple cache activation modules, first latches corresponding to each cache activation module, and a selector connecting each first latch.
[0083] Among them, the number of cache activation modules is the same as the number of hidden layer nodes, and one cache activation module corresponds to one hidden layer node, which is used to calculate the calculation result of the corresponding hidden layer node. And the selector, as the module that finally transmits the data of this layer to the next layer one by one, needs to interface with each latch to obtain the calculation results of each hidden layer node and uniformly send them down.
[0084] Specifically, each cache activation module respectively receives the count value and the calculation sub-results of the corresponding hidden layer node, accumulates the respective calculation sub-results of the corresponding hidden layer node based on the count value, and when the count value reaches the maximum value, uses the activation function to calculate the current accumulated result to obtain the calculation result of the corresponding hidden layer node.
[0085] The first latches corresponding to each cache activation module cache the calculation results of the hidden layer nodes output by the corresponding cache activation modules and transmit the calculation results of the hidden layer nodes to the selector.
[0086] The selector sequentially transmits the calculation results of the hidden layer nodes in each first latch to the next layer according to the current count value, that is, each count value corresponds to the calculation result of one hidden layer node. After the count value reaches the maximum value, it will return and start cycling from the minimum value again. At this time, the calculation results of the hidden layer nodes corresponding to the smallest count value will be transmitted to the next layer accordingly.
[0087] Specifically, in the first n cycles, the selector selects and outputs the calculation results of different nodes according to the count value. In the (n + 1)-th cycle, the constant 1 is derived, and in the (n + 2)-th cycle, the constant 0 is derived, which is to augment the input data of the next layer. Here, n is the number of nodes in the hidden layer.
[0088] Optionally, in another embodiment of the present application, as Figure 9 shown, the cache activation module includes:
[0089] a control output module, a BRAM module, and an activation function module.
[0090] Among them, the control output module is used to receive the count value and the corresponding calculation sub-results of the hidden layer nodes. And every time a calculation sub-result of a hidden layer node is received, the currently received calculation sub-result of the hidden layer node is accumulated with the current accumulated result stored in the BRAM module to obtain a new accumulated result, and the new accumulated result is updated to the BRAM module. Then, when the received count value is the maximum value, the control activation function module uses the activation function to calculate the latest accumulated result in the BRAM module to obtain the calculation result of the hidden layer node, and outputs the calculation result of the hidden layer node to the corresponding first latch.
[0091] For the output layer, it includes an output calculation layer and an output latch layer.
[0092] Among them, the output calculation layer has the same structure as the hidden calculation layer. It should be noted that the output calculation layer and the hidden calculation layer are only the same in structure, that is, they include the same modules and the functions of the modules are also the same. However, due to the different numbers of nodes included and belonging to different layers, the corresponding lookup tables included are different.
[0093] The output latch layer is used to cache the calculation sub-results of each output layer node, and when the count value is the maximum value, for each output layer node, sum up the respective calculation sub-results of the output layer node to obtain the calculation results of each output layer node, and output the calculation results of each output layer node as the system calculation result.
[0094] It should be noted that the output latch layer is similar to the calculation selection layer. However, since it is already the last layer, it no longer requires an activation function module and a selector, so it can directly sum up the calculation sub-results and output them.
[0095] Optionally, in another embodiment of the present application, as Figure 10 shown, the output latch layer includes:
[0096] a plurality of cache control modules and second latches corresponding to each cache control module.
[0097] Specifically, also refer to Figure 10 , each cache control module includes a BRAM module and a control output module.
[0098] Among them, the BRAM module is used to cache the calculation sub-results of the corresponding output layer nodes. The control output module receives the count value, and when the count value is the maximum value, sums up the respective calculation sub-results of the output layer nodes in the BRAM module to obtain the calculation result of the corresponding output layer node, and outputs the calculation result of the corresponding output layer node to the corresponding latch for output.
[0099] An embodiment of the present application provides a neural network system based on FPGA, including an input layer, multiple hidden layers connected in sequence, an output layer, and a clock for triggering the operation of each layer. The input layer includes an input calculation layer and a data cache layer. The input calculation layer includes a counter, a calculation module, and a lookup table corresponding to the input layer. Among them, the counter is used to sequentially output count values, and based on the count values, look up the weight values corresponding to the count values from the lookup tables corresponding to each layer. The calculation module calculates the calculation results of the input layer nodes by using the products of the respective input weight values found from the lookup table corresponding to the input layer and the corresponding input values. The data cache layer is used to cache the calculation results of the input layer nodes and output the calculation results of the input layer nodes to the next layer. Each hidden layer includes a hidden calculation layer and a data selection layer. The hidden calculation layer includes a lookup table corresponding to the hidden layer and a product calculation module, which is used to, when receiving the calculation results output by the previous layer each time, look up the hidden weight values corresponding to the currently received count values from the lookup table corresponding to the hidden layer according to the currently received count values, and calculate the products of the respective hidden weight values and the received calculation results through the product calculation module to obtain the calculation sub-results of the respective hidden layer nodes. The data selection layer is used to cache the calculation sub-results of the respective hidden layer nodes, and when the count value is the maximum value, for each hidden layer node respectively, use an activation function to calculate the sum of the respective calculation sub-results of the hidden layer node to obtain the calculation results of the respective hidden layer nodes, and sequentially output the calculation results of the respective hidden layer nodes to the next layer. The output layer includes an output calculation layer and an output latch layer. Among them, the output calculation layer has the same structure as the hidden calculation layer. The output latch layer is used to cache the calculation sub-results of the respective output layer nodes, and when the count value is the maximum value, for each output layer node respectively, sum up the respective calculation sub-results of the output layer node to obtain the calculation results of the respective output layer nodes, and output the calculation results of the respective output layer nodes as the system calculation results, thereby realizing the setting of the hardware resources of the FPGA corresponding to the composition of the functional modules of the fully connected neural network in the FPGA, and these hardware resources can implement the corresponding functions of the functional modules of the fully connected neural network, and further realize the fully connected neural network in the FPGA, avoiding being restricted by the hardware resources, so as to well meet the current high-efficiency big data processing requirements.
[0100] Those skilled in the art may further realize that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0101] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A neural network system based on FPGA, characterized in that, Including: An input layer, multiple sequentially connected hidden layers, an output layer, and a clock for triggering the operation of each layer; The input layer includes an input calculation layer and a data cache layer; The input calculation layer includes a counter, a calculation module, and a lookup table corresponding to the input layer; wherein, the counter is used to sequentially output count values, and based on the count values, look up the weight values corresponding to the count values from the lookup tables corresponding to each layer; the calculation module calculates the calculation results of the input layer nodes by using the products of the respective input weight values looked up from the lookup table corresponding to the input layer and the corresponding input values; The data cache layer is used to cache the calculation results of the input layer nodes and output the calculation results of the input layer nodes to the next layer; Each hidden layer includes a hidden calculation layer and a data selection layer; The hidden calculation layer includes a lookup table corresponding to the hidden layer and a product calculation module, which is used to, when receiving the calculation results output by the previous layer each time, look up the hidden weight values corresponding to the currently received count values from the lookup table corresponding to the hidden layer according to the currently received count values, and calculate the products of the respective hidden weight values and the received calculation results through the product calculation module to obtain the calculation sub-results of each hidden layer node; The data selection layer is used to cache the calculation sub-results of each hidden layer node, and when the count value is the maximum value, for each hidden layer node respectively, calculate the sum of the calculation sub-results of the hidden layer node by using an activation function to obtain the calculation results of each hidden layer node, and sequentially output the calculation results of each hidden layer node to the next layer; The output layer includes an output calculation layer and an output latch layer; wherein, the output calculation layer has the same structure as the hidden calculation layer; the output latch layer is used to cache the calculation sub-results of each output layer node, and when the count value is the maximum value, for each output layer node respectively, sum the calculation sub-results of the output layer node to obtain the calculation results of each output layer node, and output the calculation results of each output layer node as the system calculation results; 2. The system according to claim 1, characterized in that, There are multiple lookup tables corresponding to the input layer, and the number of the lookup tables corresponding to the input layer is the sum of the total number of output layer nodes and 1; wherein, the nth lookup table corresponding to the input layer stores the input weight values corresponding to the nth count value of each output layer node; 3. The system according to claim 1, wherein The calculation module includes: A summation module and an activation function module; wherein, the summation module accumulates the products of the respective input weight values looked up and the corresponding input values respectively, and the activation function module calculates the accumulated result by using an activation function to obtain the calculation results of the input layer nodes; 4. The system according to claim 1, characterized in that, The data cache layer includes: A data selection module, a BRAM module, and a data latch module; Among them, when the data selection module receives the calculation result of any one of the input layer nodes, if the currently received count value is not greater than n, it transmits the received calculation result of the input layer node to the BRAM module for caching. If the count value output by the currently received counter is greater than n, it transmits 1 to the BRAM module for caching, and transmits the calculation result in the BRAM module to the next layer through the data latch module.
5. The system according to claim 1, characterized in that, There are multiple lookup tables corresponding to the hidden layer, and the number of lookup tables corresponding to the hidden layer is the total number of hidden layer nodes; one lookup table corresponding to the hidden layer stores the hidden weight values corresponding to each count value corresponding to one hidden layer node.
6. The system according to claim 1, characterized in that, The data selection layer includes: Multiple cache activation modules, first latches corresponding to each cache activation module, and a selector connecting each first latch; Each cache activation module respectively receives the count value and the calculation sub-result of the corresponding hidden layer node, accumulates each calculation sub-result of the corresponding hidden layer node based on the count value, and when the count value is the maximum value, uses an activation function to calculate the current accumulated result to obtain the calculation result of the corresponding hidden layer node; The first latches corresponding to each cache activation module cache the calculation result of the hidden layer node output by the corresponding cache activation module, and transmit the calculation result of the hidden layer node to the selector; The selector sequentially transmits the calculation results of the hidden layer nodes in each first latch to the next layer according to the current count value.
7. The system according to claim 6, wherein The cache activation module includes: A control output module, a BRAM module, and an activation function module; Among them, the control output module is used to receive the count value and the calculation sub-result of the corresponding hidden layer node. Each time it receives a calculation sub-result of a hidden layer node, it accumulates the currently received calculation sub-result of the hidden layer node with the current accumulated result stored in the BRAM module to obtain a new accumulated result, updates the new accumulated result to the BRAM module, and when it receives that the count value is the maximum value, controls the activation function module to use the activation function to calculate the latest accumulated result in the BRAM module to obtain the calculation result of the hidden layer node, and outputs the calculation result of the hidden layer node to the corresponding first latch.
8. The system according to claim 1, wherein The output latch layer includes: Multiple cache control modules and second latches corresponding to each cache control module; Among them, each of the cache control modules includes a BRAM module and a control output module; the BRAM module is used to cache the calculation sub-results of the corresponding output layer nodes; the control output module receives the count value, and when the count value is the maximum value, sums up the respective calculation sub-results of the output layer nodes in the BRAM module to obtain the calculation result of the corresponding output layer node, and outputs the calculation result of the corresponding output layer node to the corresponding latch for output.
Citation Information
Patent Citations
A convolutional neural network module based on an FPGA
CN109711533A
Method and apparatus for performing neural network operations
CN116601643A