Long short-term memory (LSTM) network computing device and computing device
By completing LSTM network calculation in memory, storing and calculating target input vectors and weight vectors, the memory wall problem in LSTM network calculation is solved, and the computing efficiency is improved.
Patent Information
- Application Number
- CN201911193407.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-11-28
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2039-11-28
AI Technical Summary
Frequent data transmission leads to memory wall problems when performing long-term short-term memory (LSTM) network computing, especially when the processor operates at a high frequency, the data bus bandwidth is not enough to support the processor's data needs.
By completing neural network calculations in memory, the specific implementation method is to store the target input vector and corresponding target weight vectors of the hidden layer node of the LSTM network in a dynamic random access memory (DRAM) unit with multiple rows and multiple columns, and calculate using the point multiplication summing module and the activation operation module to reduce the transmission of data between memory and processor.
Implementing LSTM network computing in memory reduces the amount of data transmission, avoids memory wall problems, and improves computing efficiency.
Smart Images

Figure CN112862059B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of neural networks, and particularly relates to a long short-term memory (LSTM) network computing device and a computing device. Background Art
[0002] With the development of long short-term memory (LSTM) networks, text recognition, gesture recognition, speech recognition, etc. can be performed through LSTM network computing. Since LSTM network computing is a complex iterative computation with a large amount of data, when the processor of a computing device uses the von Neumann computing architecture for LSTM network computing, the processor first stores the weights of the LSTM network in memory. Usually, the memory is dynamic random access memory (DRAM). Then, the data bus transfers the data stored in the DRAM to the neural network computing node, and the LSTM network computation is completed in the neural network computing node. When the computation is completed, the data bus outputs the computation result to the DRAM. Since the weights need to be repeatedly read from the DRAM and the intermediate computation results need to be read and written during the LSTM network computation, the data transfer volume is huge. When the processor operating frequency requirement is high, the data bus bandwidth connecting the processor and the DRAM will not be sufficient to transfer the data required by the processor, and the memory wall problem will easily occur. Summary of the Invention
[0003] This application provides a long short-term memory (LSTM) network computing device and a computing device, which can complete neural network computing in memory and solve the memory wall problem. The technical solution is as follows:
[0004] In a first aspect, a long short-term memory (LSTM) network computing device is provided. The LSTM network computing device includes:
[0005] A storage module, the storage module includes a first storage array and a second storage array. Any one of the first storage array and the second storage array includes multiple rows and multiple columns of dynamic random access memory (DRAM) cells. The first row of DRAM cells in any one of the storage arrays is used to store the target input vector of the hidden layer nodes of the LSTM network, and the second row of DRAM cells in any one of the storage arrays is used to store the target weight vector corresponding to the target input vector. The target input vector includes a first input vector or a second input vector. The first input vector includes the input data of the LSTM network at the current moment, and the second input vector includes the output data of the hidden layer nodes at the previous moment;
[0006] The dot product summation module, connected to the storage module, is configured to perform dot product calculations on the target input vector and the target weight vector corresponding to the target input vector, obtain the dot product result corresponding to any storage array, and perform summation calculations on the dot product results corresponding to each storage array to obtain a sum value;
[0007] The activation operation module, connected to the dot product summation module, is configured to perform an activation operation on the sum value to obtain the output vector of the hidden layer node at the current moment.
[0008] By storing the target input vector and the corresponding target weight vector in any storage array, the dot product summation module can read the target input vector and the corresponding target weight vector from any storage array, and obtain the sum value of the dot product results corresponding to any storage array. The activation operation module performs an activation operation on the sum value to obtain the output vector at the current moment, thereby implementing the LSTM network calculation in the memory, reducing the transmission of LSTM network data between the memory and the processor, and solving the memory wall problem.
[0009] In a possible implementation, the first row of DRAM cells in any storage array sequentially stores a plurality of the target input vectors. The target weight vector includes a plurality of weight vectors. The second row of DRAM cells in any storage array sequentially stores a plurality of weight vectors, and each weight vector corresponds to one target input vector stored in the first row of DRAM cells in any storage array;
[0010] One DRAM cell in the first row of DRAM cells in any storage array is used to store one bit value of one input data in one target input vector, and one DRAM cell in the second row of DRAM cells in any storage array is used to store one bit value of one weight in one weight vector.
[0011] In a possible implementation, the activation operation module includes:
[0012] A plurality of activation units, each activation unit is connected to the dot product summation module, and is configured to perform an activation operation on a sum value output by the dot product summation module. The sum value is the dot product calculation result of one target input vector and one weight vector corresponding to the one target input vector;
[0013] A register, configured to store the state vector of the hidden node at the previous moment, and the state vector at the previous moment is used to indicate the state of the hidden node at the previous moment;
[0014] An output unit, connected to the multiple activation units and the register, is configured to obtain the output vector at the current moment and the state vector of the hidden node at the current moment based on the activation results of the multiple activation units and the state vector at the previous moment stored in the register. The state vector at the current moment is used to indicate the state of the hidden node at the current moment.
[0015] In a possible implementation, the multiple activation units include an input activation circuit, a forget activation circuit, an update activation circuit, and an output activation circuit. The input activation circuit is configured to perform activation calculation for the input gate in the hidden layer node. The forget activation circuit is configured to perform activation calculation for the forget gate in the hidden layer node. The update activation circuit is configured to perform activation calculation for the update gate in the hidden layer node. The output activation circuit is configured to perform activation calculation for the output gate in the hidden layer node.
[0016] In a possible implementation, the output unit includes:
[0017] A multiply-accumulate circuit, connected to the input activation circuit, the forget activation circuit, and the update activation circuit, is configured to perform multiplication calculation and addition calculation based on the activation results of the input activation circuit, the forget activation circuit, the update activation circuit, and the state vector of the hidden layer node at the previous moment to obtain the state vector at the current moment;
[0018] A target activation circuit, connected to the multiply-accumulate circuit, is configured to perform an activation operation on the state vector at the current moment output by the multiply-accumulate circuit;
[0019] A first multiplier, connected to the target activation circuit and the output activation circuit, performs a multiplication operation on the activation result of the target activation circuit and the activation result output by the output activation circuit to obtain the output vector at the current moment.
[0020] In a possible implementation, the multiply-accumulate circuit includes:
[0021] A second multiplier, connected to the output ends of the input activation circuit and the update activation circuit, is configured to perform multiplication calculation on the activation result of the input activation circuit and the activation result of the update activation circuit;
[0022] A third multiplier, connected to the forget activation circuit and the register, is configured to perform multiplication calculation on the activation result of the forget activation circuit and the state vector at the previous moment stored in the register;
[0023] An adder, connected to the second multiplier and the third multiplier, is configured to perform an addition operation on the output result of the first multiplication and the output result of the third multiplier to obtain the state vector of the previous moment of the hidden layer node.
[0024] In a possible implementation, the LSTM network computing device further includes:
[0025] A controller, connected to the driver, is configured to control the driver to select a column of DRAM cells of any one of the storage arrays every preset time duration;
[0026] The driver, connected to any one of the storage arrays, is configured to write the input data in the target input vector into the first row of DRAM cells in each column of any one of the storage arrays and write the weights in the weight vector corresponding to the target input vector into the second row of DRAM cells in each column of any one of the storage arrays based on a target current every preset time duration.
[0027] In a possible implementation, the dot product summation module includes:
[0028] A plurality of dot product units, each dot product unit is connected to multiple columns of DRAM cells in any one of the storage arrays, and is configured to perform a dot product calculation on a target input vector stored in the DRAM cells in the first row of the multiple columns of DRAM cells and a weight vector stored in the DRAM cells in the second row of the multiple columns of DRAM cells;
[0029] A summation unit, connected to the dot product unit, is configured to perform a summation calculation on the dot product results output by each dot product unit.
[0030] In a possible implementation, each dot product unit includes:
[0031] A plurality of latches, the input end of each latch is connected to a column of DRAM cells in the multiple columns of DRAM cells, the output end of each latch is connected to one of the multipliers in the plurality of multipliers, and each latch is configured to cache the value in one of the DRAM cells in the second row of any one of the storage arrays;
[0032] A plurality of multipliers, the first input end of each multiplier is respectively connected to a column of DRAM cells in the multiple columns of DRAM cells, the second input end of each multiplier is respectively connected to the output end of a latch, and the plurality of multipliers are respectively configured to perform a dot product calculation on the one target input vector and the one weight vector stored in the multiple columns of DRAM cells.
[0033] In a possible implementation, each multiplier includes:
[0034] At least one exclusive OR circuit, wherein a first input terminal of each exclusive OR circuit is connected to a latch, a second input terminal of each exclusive OR circuit is connected to one column of DRAM cells among the multiple columns of DRAM cells, and each exclusive OR circuit is configured to perform an exclusive OR calculation on one bit value of one weight in the one target weight vector cached by one latch and one bit value of one input data in one input vector stored in one column of DRAM cells among the multiple columns of DRAM cells;
[0035] A combiner, wherein a first input terminal of the combiner is connected to an output terminal of the at least one exclusive OR circuit, a second input terminal of the combiner is connected to at least one latch, and the combiner is configured to combine the data output by the at least one exclusive OR circuit and the data output by at least one latch to obtain the product of each input data in the one target input vector and one weight in the one weight vector.
[0036] A summation circuit, connected to the combiner, and configured to perform a summation calculation on the products output by the combiner multiple times.
[0037] In a possible implementation manner, the summation unit includes:
[0038] A target adder, connected to the summation circuit, and configured to perform a summation calculation on the summation results output by the summation circuit multiple times;
[0039] A buffer, connected to the target adder, and configured to store the sum value calculated by the target adder each time.
[0040] In a second aspect, a computing device is provided, and the computing device includes:
[0041] A processor: configured to receive a first input vector at each moment of the LSTM network, and input the first input vector at each moment into a memory connected to the processor, wherein the first input vector at each moment includes the input data of the LSTM network at each moment;
[0042] The memory includes:
[0043] A storage module, wherein the storage module includes a first storage array and a second storage array, and any one of the first storage array and the second storage array includes multiple rows and multiple columns of dynamic random access memory (DRAM) cells. The first row of DRAM cells of any one of the storage arrays is configured to store a target input vector of the hidden layer nodes of the LSTM network, and the second row of DRAM cells of any one of the storage arrays is configured to store a target weight vector corresponding to the target input vector. The target input vector includes the first input vector or the second input vector at the current moment, and the second input vector includes the output data of the previous moment of the hidden layer node;
[0044] The dot product summation module, connected to the storage module, is configured to perform dot product calculations on the target input vector and the target weight vector corresponding to the target input vector to obtain the dot product result corresponding to any one of the storage arrays, and perform summation calculations on the dot product results corresponding to each storage array to obtain a sum value;
[0045] The activation operation module, connected to the dot product summation module, is configured to perform an activation operation on the sum value to obtain the output vector of the hidden layer node at the current moment.
[0046] In a possible implementation, the first row of DRAM cells of any one of the storage arrays sequentially stores a plurality of the target input vectors, the target weight vector includes a plurality of weight vectors, the second row of DRAM cells of any one of the storage arrays sequentially stores a plurality of weight vectors, and each weight vector corresponds to one target input vector stored in the first row of DRAM cells of any one of the storage arrays;
[0047] One DRAM cell in the first row of DRAM cells of any one of the storage arrays is used to store one bit value of one input data in one target input vector, and one DRAM cell in the second row of DRAM cells of any one of the storage arrays is used to store one bit value of one weight in one weight vector.
[0048] In a possible implementation, the activation operation module includes:
[0049] A plurality of activation units, each activation unit is connected to the dot product summation module, and is configured to perform an activation operation on a sum value output by the dot product summation module, and the sum value is the dot product calculation result of one target input vector and one weight vector corresponding to the one target input vector;
[0050] A register, configured to store the state vector of the hidden node at the previous moment, and the state vector at the previous moment is used to indicate the state of the hidden node at the previous moment;
[0051] An output unit, connected to the plurality of activation units and the register, is configured to obtain the output vector at the current moment and the state vector of the hidden node at the current moment based on the activation results of the plurality of activation units and the state vector at the previous moment stored in the register, and the state vector at the current moment is used to indicate the state of the hidden node at the current moment.
[0052] In a possible implementation, the multiple activation units include an input activation circuit, a forgetting activation circuit, an update activation circuit, and an output activation circuit. The input activation circuit is used to implement the activation calculation of the input gate in the hidden layer node. The forgetting activation circuit is used to implement the activation calculation of the forgetting gate in the hidden layer node. The update activation circuit is used to implement the activation calculation of the update gate in the hidden layer node. The output activation circuit is used to implement the activation calculation of the output gate in the hidden layer node.
[0053] In a possible implementation, the output unit includes:
[0054] A multiply-accumulate circuit, connected to the input activation circuit, the forgetting activation circuit, and the update activation circuit, is used to perform multiplication and addition calculations based on the activation results of the input activation circuit, the forgetting activation circuit, the update activation circuit, and the state vector of the hidden layer node at the previous moment, to obtain the state vector at the current moment;
[0055] A target activation circuit, connected to the multiply-accumulate circuit, is used to perform an activation operation on the state vector at the current moment output by the multiply-accumulate circuit;
[0056] A first multiplier, connected to the target activation circuit and the output activation circuit, performs a multiplication operation on the activation result of the target activation circuit and the activation result output by the output activation circuit to obtain the output vector at the current moment.
[0057] In a possible implementation, the multiply-accumulate circuit includes:
[0058] A second multiplier, connected to the output ends of the input activation circuit and the update activation circuit, is used to perform a multiplication calculation on the activation result of the input activation circuit and the activation result of the update activation circuit;
[0059] A third multiplier, connected to the forgetting activation circuit and a register, is used to perform a multiplication calculation on the activation result of the forgetting activation circuit and the state vector of the previous moment stored in the register;
[0060] An adder, connected to the second multiplier and the third multiplier, is used to perform an addition operation on the output result of the first multiplication and the output result of the third multiplier to obtain the state vector of the hidden layer node at the previous moment.
[0061] In a possible implementation, the LSTM network computing device further includes:
[0062] A controller, connected to a driver, is used to control the driver to select a column of DRAM cells of any one of the storage arrays every preset time duration;
[0063] The driver, connected to any one of the storage arrays, is configured to write the input data in the target input vector into the first-row DRAM cells of each column of any one of the storage arrays, and write the weights in the weight vector corresponding to the target input vector into the second-row DRAM cells of each column of any one of the storage arrays every preset duration based on the target current.
[0064] In a possible implementation, the dot product summation module includes:
[0065] A plurality of dot product units, each dot product unit is connected to multiple columns of DRAM cells in any one of the storage arrays, and is configured to perform dot product calculations on a target input vector stored in the DRAM cells of the first row of the multiple columns of DRAM cells and a weight vector stored in the DRAM cells of the second row of the multiple columns of DRAM cells;
[0066] And a summation unit, connected to the dot product units, is configured to perform summation calculations on the dot product results output by each dot product unit.
[0067] In a possible implementation, each dot product unit includes:
[0068] A plurality of latches, the input end of each latch is connected to one column of DRAM cells in the multiple columns of DRAM cells, the output end of each latch is connected to one of a plurality of multipliers, and each latch is configured to cache the value in one DRAM cell in the second-row DRAM cells of any one of the storage arrays;
[0069] A plurality of multipliers, the first input end of each multiplier is respectively connected to one column of DRAM cells in the multiple columns of DRAM cells, the second input end of each multiplier is respectively connected to the output end of a latch, and the plurality of multipliers are respectively configured to perform dot product calculations on the one target input vector and the one weight vector stored in the multiple columns of DRAM cells.
[0070] In a possible implementation, each multiplier includes:
[0071] At least one exclusive OR circuit, the first input end of each exclusive OR circuit is connected to a latch, the second input end of each exclusive OR circuit is connected to one column of DRAM cells in the multiple columns of DRAM cells, and each exclusive OR circuit is configured to perform exclusive OR calculations on one bit value of one weight in the one target weight vector cached by a latch and one bit value of one input data in one column of DRAM cells stored in the multiple columns of DRAM cells;
[0072] A combiner, the first input end of the combiner is connected to the output end of the at least one exclusive - OR circuit, the second input end of the combiner is connected to at least one latch, and the combiner is used to combine the data output by the at least one exclusive - OR circuit and the data output by at least one latch to obtain the product of each input data in the one target input vector and one weight in the one weight vector.
[0073] A summing circuit, connected to the combiner, for performing a summation calculation on the products output by the combiner multiple times.
[0074] In a possible implementation manner, the summing unit includes:
[0075] A target adder, connected to the summing circuit, for performing a summation calculation on the summation results output by the summing circuit multiple times;
[0076] A buffer, connected to the target adder, for storing the sum value calculated by the target adder each time.
[0077] In a third aspect, a method for calculating an LSTM network is provided. The method is executed by an LSTM network calculation device and includes:
[0078] Receiving a first input vector of the LSTM network, a first weight vector corresponding to the first input vector, a second input vector, and a second weight vector corresponding to the second input vector. The first input vector includes the input data of the LSTM network at the current moment, and the second input vector includes the output data of the hidden - layer nodes at the previous moment;
[0079] Storing the first input vector in the first - row DRAM cells of the first storage array in the storage module, and storing the first weight vector in the second - row DRAM of the first storage array, where the first storage array includes multiple rows and multiple columns of DRAM cells;
[0080] Storing the second input vector in the first - row DRAM cells of the second storage array in the storage module, and storing the second weight vector in the second - row DRAM of the second storage array, where the second storage array includes multiple rows and multiple columns of DRAM cells;
[0081] Based on the dot - product summing module connected to the storage module, performing a dot - product calculation on the target input vector and the weight vector corresponding to the target input vector to obtain the dot - product result corresponding to any storage array, and performing a summation calculation on the dot - product results corresponding to each storage array to obtain a sum value. The target input vector includes the first input vector or the second input vector;
[0082] An activation operation module based on a connection dot product summation module performs an activation operation on the sum value to obtain the output vector of the hidden layer node at the current moment. Description of the Drawings
[0083] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0084] Figure 1 It is a schematic structural diagram of a bidirectional LSTM network provided by an embodiment of the present application;
[0085] Figure 2 It is a schematic structural diagram of a hidden node provided by an embodiment of the present application;
[0086] Figure 3 It is a schematic diagram of the calculation process of a hidden layer node provided by an embodiment of the present application;
[0087] Figure 4 It is a schematic diagram of an LSTM network calculation device provided by an embodiment of the present application;
[0088] Figure 5 It is a schematic diagram of the data deployment method of an LSTM network provided by an embodiment of the present application;
[0089] Figure 6 It is a schematic diagram of an exclusive - OR circuit provided by an embodiment of the present application;
[0090] Figure 7 It is a schematic diagram of a summation circuit provided by an embodiment of the present application;
[0091] Figure 8 It is a schematic diagram of a summation unit provided by an embodiment of the present application;
[0092] Figure 9 It is a schematic diagram of an activation operation module provided by an embodiment of the present application;
[0093] Figure 10 It is a schematic diagram of writing data provided by an embodiment of the present application;
[0094] Figure 11 It is a calculation process of an LSTM network provided by an embodiment of the present application;
[0095] Figure 12 It is a calculation process of an LSTM network calculation device provided by an embodiment of the present application;
[0096] Figure 13 It is a schematic diagram showing the decomposition of the calculation process of an implicit node provided by an embodiment of the present application;
[0097] Figure 14 It is a schematic structural diagram of a computing device provided by an embodiment of the present application;
[0098] Figure 15 It is a schematic diagram of an implementation environment provided by an embodiment of the present invention;
[0099] Figure 16 It is a schematic structural diagram of an LSTM network computing device provided by an embodiment of the present invention;
[0100] Figure 17 It is a flowchart of an LSTM network computing method provided by an embodiment of the present application. Detailed implementation manners
[0101] To make the objectives, technical solutions and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0102] To facilitate the understanding of the technical process of the present application, the basic principle of the LSTM network will be elaborated first:
[0103] The overall structure of the LSTM network: The LSTM network includes an input layer, a hidden layer, and an output layer. Among them, the input layer consists of at least one input node; when the LSTM network is a unidirectional network, the hidden layer only includes a forward hidden layer, and when the LSTM network is a bidirectional network, the hidden layer includes a forward hidden layer and a backward hidden layer. For example Figure 1 as shown in the schematic structural diagram of a bidirectional LSTM network provided by an embodiment of the present application, each hidden layer includes at least one hidden node; the output layer includes at least one output node; for each input node, it is respectively connected to the forward hidden layer node and the backward hidden layer node, and is used to output the input data to the forward hidden layer node and the backward hidden layer node respectively. The hidden nodes in each hidden layer are respectively connected to the output nodes, and are used to output their own calculation results to the output nodes. The output nodes calculate according to the output nodes of the hidden layer and output data.
[0104] Still taking Figure 1 as an example, the calculation process in the bidirectional LSTM network: The hidden nodes in the forward hidden layer calculate forward from time 1 to time t once, and obtain and save the output at each time; the hidden nodes in the backward hidden layer calculate backward along time t to time 1 once, and obtain and save the output at each time; at each time, the output layer nodes perform a fully connected calculation on the results output at the corresponding times of the forward (forward) hidden layer and the backward (backward) hidden layer to obtain the final output, which can be specifically expressed as the following formula (1):
[0105] l t = W l1 y t + W L2 y' t (1)
[0106] where t is the current time, t > 0, and W l1 and W l2 are both weight vectors, y t is the output of the hidden node in the forward hidden layer at time t, and y' t is the output of the hidden node in the backward hidden layer at time t, and l t is the output of the output layer of the LSTM network at time t.
[0107] In the LSTM network, there is a special parameter called the node state. Each hidden layer node has a node state at each time. For example, the state vector c t at time t, which is used to indicate the state of the hidden node at time t and can also be regarded as the information of the hidden node at time t, enabling the hidden node to maintain memory from time 1 to time t to achieve long-term and short-term memory. Taking the hidden node in the forward hidden layer as an example, refer to Figure 2 the schematic structural diagram of a hidden node provided by the embodiment of the present application shown in, the hidden layer node includes 4 gates, namely the input gate, the update gate, the forget gate, and the output gate. The input of this hidden node can include the input data of the LSTM network at time t, the output data at time t - 1, and the node state at time t - 1. Among them, the input data, output data, and node state can all be represented in the form of vectors. Denote the input vector at time t as x t and the output vector at time t - 1 as y t-1 , and denote the state vector at time t as c t and the state vector at time t - 1 as c t-1 . In addition, considering that both x t and y t-1 are the input vectors of the hidden node at time t, therefore, x t can be denoted as the first input vector of the hidden node at time t, and y t-1 can be denoted as the second input vector at time t. For the convenience of subsequent description, both the first input vector and the second input vector can be regarded as the target input vector, that is, the target input vector includes the first input vector or the second input vector.
[0108] Combined with Figure 2 , refer to Figure 3Schematic diagram of the calculation process of a hidden layer node provided by the embodiment of the present application. The hidden node can control the output of the hidden layer node through these 4 gates. Among them, the input gate is used to control the storage of new information into the node state, and its calculation process can be seen in the following formula (2); the update gate is used to convert new information into a form that can be added to the node state. The new information can be obtained from the input data of the LSTM network at the current moment, and the calculation process of the update gate can be the following formula (3); the forget gate is used to control how much of the node state at the previous moment can be retained in the node state at the current moment, and its calculation process can be seen in the following formula (4); the input gate, the update gate, and the forget gate can realize the control of the node state of the hidden node, and the specific calculation process can be seen in the following formula (5); the output gate is used to control how much of the node state at the current moment can be used as the output of the hidden layer node at the current moment, and its calculation process can be seen in the following formulas (6)-(7).
[0109] i t = sigmoid(w i x t + r i y t-1 + b i ) (2)
[0110] u t = tanh(w u x t + r u y t-1 + b u ) (3)
[0111] f t = sigmoid(w f x t + r f y t-1 + b f ) (4)
[0112] c t = i t u t + f t c t-1 (5)
[0113] o t = sigmoid(w o x t + r o y t-1 + b o ) (6)
[0114] y t = o t tanh(c t ) (7)
[0115] Among them, i t , u t , f t and o t are the output results of the input gate, update gate, forget gate, and output gate respectively. w i , w u , w f and w o are the weight vectors corresponding to x t in the input gate, update gate, forget gate, and output gate respectively. Each weight vector can include at least one weight, and it can be understood that the weight vector W corresponding to x t is W = [w i w u w f w o ; r i , r u , r f and r o are the weight vectors corresponding to y t-1 in the input gate, update gate, forget gate, and output gate respectively. Each weight vector can include at least one weight, and it can be understood that the weight vector R corresponding to y t-1 is R = [r i r u r f r o ; b i , b u , b f and b o are the bias parameters in the input gate, update gate, forget gate, and output gate respectively. It should be noted that for the convenience of description, the weight vector W corresponding to x t and the weight vector corresponding to y t-1 can both be regarded as target weight vectors, that is, the target weight vector includes the weight vectors W and R, where the weight vector W corresponds to x t , and the weight vector R corresponds to y t-1 .
[0116] It should be noted that the data format of a data can be represented by {sign bit, integer bit, precision bit}, for example Figure 3 the 8b{1,0,7} in it represents an 8-bit data, where this data includes 1-bit sign bit, the integer bit is 0 which means there is no integer bit, and 7-bit decimal bit. The output result of each gate is 8b{1,0,7}, and the sum value of two 8-bit data is 16 bits.
[0117] In some embodiments, the above calculation process can be implemented within a long short-term memory (LSTM) network computing device. To further illustrate the hardware structure of the LSTM network computing device, refer to Figure 4 , Figure 4 which is a schematic diagram of an LSTM network computing device provided by an embodiment of the present application. The LSTM network computing device includes: a storage module, the storage module includes a first storage array and a second storage array. Any one of the first storage array and the second storage array includes multiple rows and multiple columns of dynamic random access memory (DRAM) cells. The first row of DRAM cells in any one of the storage arrays is used to store the target input vector of the hidden layer nodes of the LSTM network. The second row of DRAM cells in any one of the storage arrays is used to store the target weight vector corresponding to the target input vector. The target input vector includes a first input vector or a second input vector. The first input vector includes the input data of the hidden layer node at the current moment. The second input vector includes the output data of the hidden layer node at the previous moment; a dot product summation module, connected to the storage module, for performing a dot product calculation on the target input vector and the target weight vector corresponding to the target input vector to obtain the dot product result corresponding to any one of the storage arrays, and performing a summation calculation on the dot product results corresponding to each storage array to obtain a sum value; an activation operation module, connected to the dot product summation module, for performing an activation operation on the sum value to obtain the output vector of the hidden layer node at the current moment.
[0118] The following description is made regarding any one of the storage arrays:
[0119] Considering that the hidden node device has four gates, and for each gate there is a weight vector corresponding to the first input vector and a weight vector corresponding to the second input vector, the LSTM network computing device can store the first input vector and the weight vector corresponding to the first input vector in a first storage array, and store the second input vector and the weight vector corresponding to the second input vector in a second storage array. And for the convenience of reading data, the LSTM network computing device can store the input vector and the corresponding weight vector in different rows of the storage array. The above-mentioned first row of DRAM cells is used to refer to any row of DRAM cells in any one of the storage arrays, and the above-mentioned second row of DRAM cells is used to refer to any row of DRAM cells in any one of the storage arrays other than the first row of DRAM.
[0120] The target input vector can be the weight vector corresponding to the first input vector. For example, w i , w u , w f and w oAny one of them, the target input vector can also be a weight vector corresponding to the second input vector. For example, r i , r u , r f and r o Any one of them. For a hidden node, the first output vector can correspond to multiple weight vectors, and the second input vector can also correspond to multiple weight vectors. Therefore, the LSTM network computing device can associate and store the target input vector and all weight vectors corresponding to the target input vector.
[0121] In a possible implementation, the first row of DRAM cells in any storage array stores multiple target input vectors in sequence. The target weight vector includes multiple weight vectors. The second row of DRAM cells in any storage array stores multiple weight vectors in sequence. Each weight vector corresponds to one target input vector stored in the first row of DRAM cells in any storage array; one DRAM cell in the first row of DRAM cells in any storage array is used to store one bit value of an input data in one target input vector, and one DRAM cell in the second row of DRAM cells in any storage array is used to store one bit value of a weight in one weight vector.
[0122] Any storage array can be a sub-array (subarry) in DRAM, and a DRAM cell can be a cell in the sub-array. Multiple DRAM cells can be arranged in a row of DRAM cells, and each row of DRAM cells is connected to a word line. Any word line is used to select a row of DRAM cells connected to the word line. Multiple DRAM cells can also be arranged in a column of DRAM cells, and each column of DRAM cells is connected to a bit line. Any bit line is used to select a column of DRAM cells connected to the bit line. When the LSTM network computing device selects any bit line and any word line, the any bit line and the any word line intersect at a DRAM cell. Furthermore, the LSTM network computing device can read the data stored in the DRAM cell or write data into the DRAM cell.
[0123] Each DRAM cell may include a capacitor and a transistor. The capacitor can store 1 bit of data volume. The amount of charge (potential level) after charging and discharging corresponds to binary data 0 and 1 respectively. A DRAM cell can be used to store 0 or 1. Since a data can be represented by the value of at least one binary bit, therefore, a data can be stored by at least one DRAM cell, that is, a 1-bit binary value after quantization of a data can be stored by a DRAM cell. Among them, one bit value of a data can be regarded as a 1-bit binary value. This data can be an input data in a target input vector or a weight in a weight vector.
[0124] In order to store data in any storage array or read data from any storage array, the LSTM network computing device may further include a row decoder and a column decoder. The row decoder is connected to each word line in any storage array, so that each row of DRAM cells in the first storage array can be connected through the word line; the column decoder is connected to each bit line in any storage array, so that each column of DRAM cells in any storage array can be connected through the bit line. The LSTM network computing device can activate the DRAM cells connected to any word line and the DRAM cells connected to any bit line through the column decoder and the row decoder, so that read operations and write operations can be performed on the intersecting DRAM cells. The read operation is used to read the data stored in the intersecting DRAM cells, and the write operation is used to store data in the intersecting DRAM cells, and then data can be read and stored in any storage array. Specifically, the row decoder can select the corresponding row address (one row address corresponds to one word line) through the row address in the row address buffer to activate a specific row address, and then realize the addressing of the row address. After the specific row address is activated, the column decoder selects the corresponding column address (one column address corresponds to one bit line) through the column address in the column address buffer to activate a specific column address, and then realizes the addressing of the column address. After the specific column address is also activated, the activated row address and the activated column address intersect at the position where a DRAM cell is located, so that the addressing of the DRAM cell can be realized, data can be stored in the DRAM cell through the write operation, or the data stored in the DRAM cell can be read through the read operation.
[0125] Since the data stored in any storage array can be quantized binary data, a binary data can be represented by data of multiple bits. For example, the binary bit 11 after quantization of the weight value 3, where the first 1 in 11 is the most significant bit of the weight value 3, and the first 1 is the second most significant bit of the weight value 3. Thus, two DRAM cells in the second row can be used to store the weight 3. Correspondingly, multiple DRAM cells in the first row can also be used to store an input data. The embodiments of the present application do not make specific limitations on the number of bits of the quantized weight and the quantized input data.
[0126] In some embodiments, any storage array may further include at least one primary sense amplifier (PSA) and at least one secondary sense amplifier (SSA). The input end of a PSA is connected to a column of DRAM cells to amplify the electric charge stored in the storage element in the DRAM. The amplified electric charge is used to indicate whether the data stored in the DRAM cell is 0 or 1. Specifically, when the amplified electric charge is greater than a preset value, it indicates that the data stored in the storage element is 1; otherwise, it indicates that the data stored in the storage element is 0. The embodiments of the present application do not make specific limitations on this preset value.
[0127] The input end of an SSA is connected to multiple any storage arrays. The SSA is used to select and read out the output result of the PSA, and then write the output result to the read-write bus of the LSTM network computing device.
[0128] When the data volume of the target input vector is small, for each hidden node, the LSTM network computing device can use one any storage array to store the target input vector and the weight vector corresponding to the target input vector. When the data volume of the target input vector is too large, considering the actual area of any storage array, one any storage array may not be able to store a large amount of data. Therefore, multiple any storage arrays can be used to jointly store the target input vector and the weight vector corresponding to the target input vector. Considering that the first input vectors at each moment of the LSTM network do not affect each other, multiple first input vectors at multiple moments can be stored in each first storage array. Since the second input vector at each moment is actually the output vector of the previous moment, a second storage array can store only the second input vector at the current moment. Additionally, considering that the hidden layer of the LSTM network can include multiple hidden nodes, the LSTM network computing device can include multiple any storage arrays. One any storage array is used to store the target input vector and the corresponding weight vector at the current moment of one hidden node. For example, Figure 5The schematic diagram of a data deployment method of an LSTM network provided by an embodiment of the present application is shown, and the storage module includes 17 first storage arrays and 17 second storage arrays, each of which corresponds to an implicit node, and each of which corresponds to an implicit node, which are nodes 0-16 respectively. For example, the first storage array of node 0 stores the first input vectors at 4 moments and the weight vectors corresponding to the first input vectors at each moment, wherein the first input vectors at the 4 moments are x t 、x t+1 、x t+2 and x t+3 , the first input vector at each moment corresponds to the weight vector W = [w i w u w f w o ], the second storage array of node 1 stores the output vector y of the hidden node at the previous moment t-1 And the output vector y at the previous moment t-1 The corresponding weight vector R = [r i r u r f r o ].
[0129] The dot product summation module is described as follows:
[0130] From the calculation formulas of each gate of the implicit node, it can be seen that for any gate, it is necessary to first perform a dot product calculation on the first input vector and the first input vector, then perform a dot product calculation on the second input vector and the second input vector, and finally perform an activation operation on the sum of the two dot product results and the bias parameter to obtain the output of any gate. In order to obtain the output result of any gate, in a possible implementation method, the dot product summation module includes: multiple dot product units, each dot product unit is connected to multiple columns of DRAM cells in any storage array, and is used to perform a dot product calculation on a target input vector stored in a DRAM cell of a first row of the multiple columns of DRAM cells and a weight vector stored in a DRAM cell of a second row of the multiple columns of DRAM cells; and a summation unit, connected to the dot product unit, and is used to sum the dot product results output by each dot product unit.
[0131] It should be noted that each dot product summation circuit can be used to obtain the dot product result of an input vector and a weight vector. Therefore, for the first input vector x t , since the four gates of the hidden node all need to be the first output vector x t Calculate, therefore, the first input vector x in the first memory array t It can correspond to multiple point multiplication circuits, for example, Figure 5 The first input vector x in tCorresponding to four dot product circuits.
[0132] In some embodiments, each dot product summation module can be connected to multiple columns of DRAM cells in any storage array through the PSA. For example, multiple columns of DRAM cells are all connected to the output terminal of the PSA, and each dot product summation module is connected to the output terminal of the PSA. Thus, each column of DRAM can output the input data and weights stored in each column of DRAM to a dot product unit through the unit PSA.
[0133] In some embodiments, multiple dot product units can be connected to the summation unit through the SSA. For example, the output terminals of multiple dot product units are connected to the output terminal of the summation unit. Each dot product unit can output the dot product result to the summation circuit through the SSA, so that the summation circuit can perform multiple summations on the dot product results output by each dot product unit to obtain the sum value between the dot product result of the first input vector and the corresponding weight vector and the dot product result of the second input vector and the corresponding weight vector.
[0134] The dot product unit can first calculate the product of each input data in an input vector and the weight corresponding to the input data in the weight vector, and then sum up the products corresponding to all input data, so as to obtain the dot product result of an input vector and a weight vector. In a possible implementation manner, each dot product unit includes: multiple latches, the input terminal of each latch is connected to one column of DRAM cells in multiple columns of DRAM cells, the output terminal of each latch is connected to one multiplier among multiple multipliers, and each latch is used to cache the value in one DRAM cell in the second row of DRAM cells in any storage array; multiple multipliers, the first input terminal of each multiplier is respectively connected to one column of DRAM cells in multiple columns of DRAM cells, the second input terminal of each multiplier is respectively connected to the output terminal of a latch, and multiple multipliers are respectively used to perform dot product calculations on a target input vector and a weight vector stored in multiple columns of DRAM cells.
[0135] Multiple multipliers can first obtain the dot product result of a weight vector and an input vector based on the weights cached in the latches and the input data stored in any storage array. In a possible implementation, each multiplier includes: at least one XOR circuit, where the first input terminal of each XOR circuit is connected to a latch, the second input terminal of each XOR circuit is connected to one column of DRAM cells among multiple columns of DRAM cells, and each XOR circuit is used to perform an XOR calculation on one bit value of a weight in a target weight vector cached in a latch and one bit value of an input data in an input vector stored in one column of DRAM cells among multiple columns of DRAM cells; a combiner, where the first input terminal of the combiner is connected to the output terminals of at least one XOR circuit, the second input terminal of the combiner is connected to at least one latch, and the combiner is used to combine the data output by at least one XOR circuit and the data output by at least one latch to obtain the product of each input data in a target input vector and a weight in a weight vector. A summing circuit is connected to the combiner and is used to perform a summation calculation on the products output by the combiner multiple times.
[0136] For the convenience of description, the output data of the combiner is called the first data, and the first data is also the product of a weight in a weight vector and an input data in an input vector. An XOR circuit can output the second highest bit of the first data, and then the combiner can combine a weight with the second highest bit of the first data to obtain the first data. To illustrate the reason for using the XOR circuit and the combiner, take a 1-bit weight K and a 2-bit input data D as an example. D can be represented by 1-bit D0 and 1-bit D1. Specifically, D0 = D0D1, and the product of D and K can be expressed by formula (8):
[0137]
[0138] As can be seen from the above formula (8), when K = 0, the highest bit of the product of D and K is K, and the second highest bit of the product of D and K is equal to the XOR of D and K. When K = 1, the highest bit 1 of 1D1D0 is K, and the second highest bit D1D0 of 1D1D0 is the XOR of D and K. Therefore, the first data can be combined based on the output result of the XOR circuit and the value of K.
[0139] When quantifying weights, if any weight is less than or equal to a first preset value, the first weight is quantified to 0. When any weight is less than or equal to the first preset value, the first weight is quantified to 1. Then, in a possible implementation, for any weight stored in the at least one latch, when the any weight is less than or equal to the first preset value, the any weight is used as the most significant bit of the first data, and the exclusive OR of the any weight and the corresponding input data is used as the second most significant bit of the first data, thereby obtaining a first data. When the any weight is greater than the first preset value, the any weight is used as the most significant bit of the second data, the first data corresponding to the any weight is used as the second most significant bit of the second data, and the second data is added to a second preset value to obtain the first data. The embodiments of the present application do not specifically limit the first preset value and the second preset value.
[0140] To further illustrate the input and output of the exclusive OR circuit, refer to Figure 6 A schematic diagram of an exclusive OR circuit provided by an embodiment of the present application is shown. A latch is connected to a switch, and this switch is connected to a PSA. When performing calculations, first, the LSTM network computing device closes the switch and selects the row decoder through a read operation to open the DRAM cells in the row where the weight K is located. The weight K stored in the opened row DRAM cells is written into the latch through the PSA. Then, the switch is disconnected, the connection between the data bus and the latch is disconnected, and the row of DRAM cells where the input data D is located is selected through a read operation. After being amplified by the PSA, the input data D1 is input to one input terminal of the exclusive OR circuit 1, and the weight K stored in the latch is input to the other input terminal of the exclusive OR circuit 1. Thus, the exclusive OR circuit 1 can output the exclusive OR of K and D1. Similarly, the exclusive OR circuit 2 can output the exclusive OR of K and D0. Then, K in the latch, the exclusive OR of K and D1, and the exclusive OR of K and D0 are input to the input terminals of the combiner so that the combiner combines them into the product of D and K. The above process is performed for each input data in an input vector, and the product of each input data and the corresponding weight can be obtained. Furthermore, the products can be added by multiple summing circuits to obtain the product of a weight vector and an input vector.
[0141] Since the input vectors and weight vectors stored in any storage array are all quantized data, each data in each input vector or weight vector may be 1-bit data or multi-bit data after quantization. Therefore, the quantity stored in any storage array is relatively large, and multiple summing circuits can be used to achieve step-by-step summation to obtain the product of a weight vector and an input vector.
[0142] To further illustrate the principle of step-by-step summation, refer to Figure 7 , Figure 7It is a schematic diagram of a summation circuit provided by an embodiment of the present application. Figure 7 One summation circuit in t can process at most 32 3-bit data. Thus, this summation circuit (which can also be called an adder tree) can add 32 3-bit data to obtain an 8-bit sum. Then, when the dimension of a weight vector is greater than 32, the dot product calculation result of a weight vector and an input vector is greater than 32 3-bit data. Then, the 8-bit sum obtained by one adder tree cannot be used as the dot product calculation result. At least the partial summation results output by multiple adder trees need to be summed again to obtain the dot product calculation result. Taking the input gate as an example, among them, the first input vector x t-1 , the second input vector y i , the weight vector w i , and the weight vector r t all have a dimension of 128, that is, x t = x t-1 [0:127], y t-1 = y i [0:127], w i = w i [0:127], r i [0:127]. For w i x t + r i y t-1 , the following decomposition is performed:
[0143] w i x t + r i y t-1 = w i [0:31]x t [0:31] + w i [32:63]x t [32:63] + w i [64:95]x t [64:95]
[0144] w i [96:127]x t [96:127] + r i [0:31]y t-1 [0:31] + r i [32:63]y t-1 [32:63]
[0145] + r i [64:95]y t-1 [64:95] + r i [96:127]y t-1[96:127]
[0146] As can be seen from the above disassembly process, if x t and w i are stored in the first storage array, at least 4 32-bit summation circuits are required to obtain w i x t ; if the summation circuits connected to the first storage array are not sufficient to calculate w i x t , the results of all summation circuits can be output to the dot product unit, and the dot product unit adds the summation results of each time of the summation circuit to obtain w i x t ; similarly, the summation results of each time of the summation circuits in the second storage array can also be output to the dot product unit, so that the output results of the summation circuits connected to the first storage array and the summation circuits connected to the second storage array are both output to the same dot product unit for summation, so as to obtain w i x t +r i y t-1 results.
[0147] In a possible implementation, the summation unit includes: a target adder, connected to the summation circuit, for performing a summation calculation on the summation results output by the summation circuit multiple times; a buffer, connected to the target adder, for storing the sum value calculated by the target adder each time. For a set of corresponding weight vectors and input vectors at the current moment, the buffer can store a bias parameter, which can be regarded as the latest sum value of the 0th time of the target adder. When each value output by the summation circuit is obtained, the target adder obtains the latest sum value calculated last time from the buffer, adds the latest sum value calculated last time to the newly obtained value to obtain the latest sum value of this time, and stores the latest sum value of this time in the buffer until the buffer obtains the sum value between the dot product results corresponding to each storage array.
[0148] Since the target adder sums the products of each input data in the input vector at the current moment and the corresponding weights, when the target adder performs each accumulation, it also needs to determine whether the latest value obtained this time is the product of the last input data in an input vector and the corresponding weight. If so, it means that the latest sum value of this time is the sum value between the dot product results corresponding to each storage array, and the target adder ends the accumulation. Otherwise, it means that the latest value of this time is not the sum value between the dot product results corresponding to each storage array, and the target adder will continue to receive new values and perform accumulation again. For example, Figure 8 the schematic diagram of a summation unit provided by the embodiment of the present application shown, Figure 8The target adder in it obtains an 8-bit data from the SSA. This 8-bit data is the data output by a summation circuit to the target adder through the SSA. The target adder obtains the result of the previous summation from the buffer, sums the newly received data and the result of the previous summation to obtain the 16-bit summation result of this time, and outputs the summation result of this time to the buffer, so that when another 8-bit data is received at input port 1 again, another summation calculation is performed. When the buffer obtains the final 16-bit summation result, the final 16-bit summation result is also w i x t +r i y t-1 , and the buffer inputs the final summation result to the activation operation module.
[0149] The activation operation module is described as follows:
[0150] The activation operation module performs an activation operation on the sum value output by the dot product summation module to obtain the node state and output data of the hidden layer node at the current moment. Among them, the activation operation refers to using an activation function to calculate the sum value, so as to perform a non-linear transformation on the sum value. The activation function can be the sigmoid function or the hyperbolic tangent function (tanh) function. In a possible implementation manner, the activation operation module includes: a plurality of first activation units, each first activation unit is connected to the dot product summation module and is used to perform an activation operation on a sum value output by the dot product summation module. A sum value is the dot product calculation result of a target input vector and a weight vector corresponding to a target input vector; a second activation unit, connected to the plurality of first activation units, and is used to perform an activation operation based on the activation results of the plurality of first activation units.
[0151] Considering that the hidden node includes an input gate, an update gate, a forget gate, and an output gate, each gate can be implemented by an activation circuit. In a possible implementation manner, the plurality of activation units include an input activation circuit, a forget activation circuit, an update activation circuit, and an output activation circuit. The input activation circuit is used to implement the activation calculation of the input gate in the hidden layer node, the forget activation circuit is used to implement the activation calculation of the forget gate in the hidden layer node, the update activation circuit is used to implement the activation calculation of the update gate in the hidden layer node, and the output activation circuit is used to implement the activation calculation of the output gate in the hidden layer node. The LSTM network computing device inputs the sum value output by the dot product summation module each time to the corresponding activation circuit, so that the activation circuit can perform an activation operation on the corresponding sum value to obtain the activation result of the activation circuit. For example, w i x t +r i y t-1+b i The sum value is input into the input activation circuit, and the input activation circuit performs an activation operation on w i x t +r i y t-1 +b i to obtain an activation result i t .
[0152] Among them, the input activation circuit realizes the activation calculation of the input gate in the hidden layer node, that is, the input activation circuit performs a first activation operation on the sum value (w i x t +r i y t-1 +b i ) corresponding to the output gate. The activation function used in this first activation operation can be the sigmoid function; the forgetting activation circuit realizes the activation calculation of the forgetting gate in the hidden layer node, that is, the forgetting activation circuit performs a second activation operation on the sum value (w f x t +r f y t-1 +b f ) corresponding to the forgetting gate. The activation function used in this second activation operation can be the sigmoid function; the update activation circuit realizes the activation calculation of the update gate in the hidden layer node, that is, the update activation circuit performs a third activation operation on the sum value (w u x t +r u y t-1 +b u ) corresponding to the update gate. The activation function used in this third activation operation can be the tanh function; the output activation circuit realizes the activation calculation of the output gate in the hidden layer node, that is, the output activation circuit performs a fourth activation operation on the sum value (w o x t +r o y t-1 +b o ) corresponding to the output gate. The activation function used in this fourth activation operation can be the sigmoid function.
[0153] As can be seen from the above formula (7), the output vector of the hidden node at the current moment (time t) is the product of the output result of the output gate and the state vector at the current moment after activation. Therefore, the output unit can first obtain the state vector of the hidden node at the current moment, and then obtain the output result of the hidden node at the current moment based on the obtained state vector at the current moment. In a possible implementation, the output unit includes: a multiply-accumulate circuit connected to the input activation circuit, the forget activation circuit, and the update activation circuit, configured to perform multiplication and addition operations based on the activation results of the input activation circuit, the forget activation circuit, the update activation circuit, and the state vector of the hidden layer node at the previous moment to obtain the state vector at the current moment; a target activation circuit connected to the multiply-accumulate circuit, configured to perform an activation operation on the state vector at the current moment output by the multiply-accumulate circuit; a first multiplier connected to the target activation circuit and the output activation circuit, configured to perform a multiplication operation on the activation result of the target activation circuit and the activation result output by the output activation circuit to obtain the output vector at the current moment.
[0154] From the above formula (5), the state vector of the hidden node at the current moment is the sum value between the product of the output results of the input gate and the update gate and the product of the state vector at the previous moment and the output result of the forget gate. Therefore, the multiply-accumulate circuit can first obtain these two products and then sum these two products. In a possible implementation, the multiply-accumulate circuit includes: a second multiplier connected to the output ends of the input activation circuit and the update activation circuit, configured to perform a multiplication operation on the activation result of the input activation circuit and the activation result of the update activation circuit; a third multiplier connected to the forget activation circuit and the register, configured to perform a multiplication operation on the activation result of the forget activation circuit and the state vector of the previous moment stored in the register; an adder connected to the second multiplier and the third multiplier, configured to perform an addition operation on the output result of the first multiplication and the output result of the third multiplier to obtain the state vector of the hidden layer node at the previous moment.
[0155] For example, Figure 9Schematic diagram of an activation operation module provided by an embodiment of the present application. The input activation circuit and the update activation circuit are connected to an 8-bit second multiplier, so that the second multiplier can obtain the output results of the input activation circuit and the update activation circuit, and obtain the product of these two activation results; the forgetting activation circuit is connected to a 16-bit third multiplier, and the third multiplier is also connected to a 16-bit register, so that the third multiplier can obtain the output result of the forgetting activation circuit and the state vector of the previous moment stored in the register, and obtain the product of the output result of the forgetting activation circuit and the state vector of the previous moment to obtain the state vector of the current moment, and store the state vector of the current moment in the register for use in the calculation of the next moment; a 16-bit adder is connected to the third multiplier and the second multiplier, so that the adder can obtain the products output by the third multiplier and the second multiplier, and sum the obtained products to obtain the state vector of the current moment; the target activation circuit is connected to the adder, and can obtain the state vector of the current moment output by the adder, and perform a fifth activation operation on the obtained state vector of the current moment. The activation function used in the fifth activation operation can be the tanh function; an 8-bit first multiplier is connected to the target activation circuit and the output activation circuit, so that the first multiplier can obtain the activation results of the target activation circuit and the output activation circuit, and perform a multiplication calculation on the two obtained activation results to obtain the output vector of the hidden node at the current moment.
[0156] After the activation operation module obtains the output vector of the current moment, the activation operation module can write the output vector back to the first row of the second storage array for the calculation of the next moment. Since the hidden node includes multiple gates, the output vector can be copied multiple times, and each copy corresponds to the weight vector of a gate and is stored in the DRAM cells of the same column. Since the data deployment is row-based, in order to improve the transmission rate and efficiency, the LSTM network computing device can keep the data bus MBL unchanged and quickly switch the column select signal line (CSL), so that multiple address writes can be achieved at one time and quickly copied to different address segments. In some embodiments, the LSTM network computing device further includes: a controller, connected to the driver, for controlling the driver to select a column of DRAM cells of any storage array every preset time; a driver, connected to any storage array, based on the target current, writing the input data in the target input vector into the DRAM cells of a certain column of any storage array, and then writing the weights in the weight vector corresponding to the target input vector into the DRAM cells of the adjacent column of any storage array.
[0157] The controller and the driver can be located in the column decoder. The preset duration can be 2 - 3 nanoseconds (ns). The target current can be a relatively large current to improve the driving performance of the driver. The controller can control the driver to select a column of DRAM cells every 2 - 3 ns. The driver can be connected to any one of the memory arrays through the address bus. The driver can select a column in any one of the memory arrays by outputting a column address strobe (CAS) signal to the address bus. For example, Figure 10 A schematic diagram of writing data provided by an embodiment of the present application as shown, Figure 10 The left figure in [the diagram] is a schematic diagram of the column decoder writing data to any one of the memory arrays, Figure 10 In the right figure [of the diagram], it is a timing diagram of the controller writing data to any one of the memory arrays. Any row of DRAM cells in any one of the memory arrays can be selected through the row decoder to form a row access. Taking the first row of DRAM cells as an example, any column in the first DRAM can be accessed through the controller in the column decoder for 2 - 3 ns to form a column access. For example, the controller can output a control signal to the driver in the column decoder every 2 - 3 ns. The control signal can be the target current. After the driver receives a control signal each time, the driver outputs a CAS to the storage address of any column of DRAM cells in any one of the memory arrays through the column address bus, so that the driver selects any column of DRAM cells in any one of the memory arrays and stores a single-bit data in the first row of DRAM in any column of DRAM cells, so that valid data is written into the first row of DRAM in any column of DRAM cells. Taking the second input vector as an example, since the output vector at the current moment is the second input vector at the next moment, after the activation operation module obtains the output vector at the current moment, the controller can control the driver to write data to each DRAM cell in the first row of the second memory array every 2 - 3 ns for the calculation at the next moment. Since a column of DRAM cells is switched every 2 - 3 ns, the data writing rate is thus improved. This method requires the data bus to remain unchanged, and when the current value of the target current is relatively large, the driving ability of the driver can be increased to ensure the efficiency of writing back data. Since the amount of data written back between multiple layers of the LSTM network computing device is relatively large and a large amount of data replication and transmission are required, adopting this method can improve the data transmission rate between the activation operation module and any one of the memory arrays and reduce the transmission delay.
[0158] It should be noted that before calculating the output vector for each moment, the LSTM network computing device can write the first input vector of this moment into the first row of the first storage array. In some embodiments, multiple first input vectors of multiple moments can also be written into the first row of the first storage array at one time. Thus, in the calculation process of each moment, the LSTM network computing device can first calculate the value of the first input vector of each moment and the corresponding weight vector, and then calculate the value of the second input vector of the current moment and the corresponding weight vector. When the LSTM network computing device calculates the output vector of the current moment of the hidden node, it then calculates the value of the second input vector of the next moment and the corresponding weight vector, and calculates the output vector of the hidden node at the next moment based on the value of the second input vector of the next moment and the corresponding weight vector and the value of the first input vector of the next moment and the corresponding weight vector calculated in advance. For example, Figure 11 As shown in the LSTM network calculation process provided by the embodiment of the present application, the LSTM network computing device first activates the first storage array and calculates the first input vector x t 、x t+1 、x t+2 、x t+3 and the product of the corresponding weight vectors to obtain 16 calculation results. Each calculation result corresponds to a gate at one moment. The LSTM network computing device stores the 16 calculation results in the summation unit; the LSTM network computing device then activates the second storage array and uses the dot product summation module and the product of the first input x t stored in the summation unit, the second input vector y t-1 stored in the second storage array, and the weight vector corresponding to the second input vector y t-1 to calculate the output vector y t at time t, and writes y t to the position of y t-1 in the second storage array; then activates the second storage array and uses the dot product summation module and the product of the first input vector x t+1 stored in the summation unit, the second input vector y t stored in the second storage array, and the weight vector corresponding to the second input vector y t to calculate the output vector y t+1 at time t + 1, and writes y t+1 to the position of y t in the second storage array. After looping 4 times, the LSTM network computing device writes the first input vectors at times t + 4 to t + 7 into the first row DRAM unit of the first storage array, and then activates the first storage array to calculate the product of the first input vectors at times t + 4 to t + 7 and the corresponding weight vectors.
[0159] It should be noted that, as can be seen from the above formula (1), the output vector of the output layer at the current moment is l t = W l1 y t + W L2 y' t , in some embodiments, the LSTM network computing device can also directly write the output vector y t of the forward hidden layer at the current moment into the first row of the third storage array, and write the output vector y' t of the backward hidden layer at the current moment into the second row of the third storage array. The third row of the third storage array is used to store the weight vector W t corresponding to y l1 , and the fourth row of the third storage array is used to store the weight vector W t corresponding to y' l2 . The third storage array is connected to the dot product summation module, so that the LSTM network computing device can reuse the dot product summation module and use the dot product summation module to obtain W l1 y t + W l2 y' t , so as to obtain the output vector of the output layer at the current moment. It should be noted that the way the third storage array stores the weight vector and the output vector is the same as the way any storage array stores a weight vector and an input vector, and the embodiments of the present application will not elaborate on this here. Moreover, the process by which the dot product summation module obtains W l1 y t + W l2 y' t is the same as the process of obtaining w i x t + r i y t-1 , and the embodiments of the present application will not elaborate on this.
[0160] To further illustrate the process of the LSTM network computing device completing the calculation process of the LSTM network at the current moment, refer to Figure 12 the calculation process of an LSTM network computing device provided by the embodiments of the present application shown in t . Taking the example that the entire LSTM network includes a forward hidden layer and a backward hidden layer, and the dimension of each hidden layer is 128, and the calculation process of the entire LSTM network for M moments is taken as an example, where M > 0, and each moment corresponds to a first input vector x. Then, the LSTM network has a total of 0 to M - 1 first input vectors x; for each hidden node, M x's are required. For the t-th moment, where <t<M, the first input vector x t-1 of each hidden node, the second input vector y t-1, the weight vector W and the weight vector R can be stored in multiple storage arrays. For any hidden node, when x is stored in the first storage array t and W = [w i w u w f w o , and y is stored in the second storage array t-1 and R = [r i r u r f r o , for each gate, the dot product unit summation module calculates the product of x t and the corresponding weight vector in the weight vector W = [w i w u w f w o and the sum of the products of the second input vector y t-1 and the corresponding weight vector in R = [r i r u r f r o , that is, the sum value corresponding to any gate. For example Figure 13 the sum value corresponding to the input gate in the decomposition schematic diagram of a hidden node calculation process provided by the embodiment of the present application shown. This sum value is calculated by the dot product circuit, and the dot product circuit inputs the sum value corresponding to each gate into the activation averaging module. The activation operation module obtains the output vector y of any hidden node at the current moment based on all the sum values t , and the LSTM network computing device executes this process for each hidden node. For example Figure 13 nodes 0 and j in, where 0 < j < 128, y t,0 is the output of node 0 at time t, y t,j is the output of node j at time t, W0 and R0 are the weight vectors of node 0, W j and R j are the weight vectors of node j; when the output vectors of each hidden node at the current moment are calculated, the output vectors are written to the third storage array so that the LSTM network computing device can implement formula (1): l t = W l1 y t + W L2 y' t , so as to obtain the output of the output node at time t; and write each input vector back to the first row of each second storage array for the next moment's calculation.
[0161] When the scale of any storage array is 1024x1024, there can be 16 adder trees (target adders) in any one storage array, and 16 dot product calculations can be supported simultaneously. The LSTM network computing device may include multiple partial banks, each bank may include multiple blocks, and each block includes 16 of any storage arrays.
[0162] It should be noted that the LSTM network computing device also includes a data input / output interface for inputting / outputting data. In the DRAM 65nm process, layout and wiring are performed on each part of the LSTM network computing device shown in this embodiment. From the final layout, PSA uses simulation tools such as cadence spectre and sadence virtuoso for layout verification, and SSA uses the simulation tool DRAMSpec for estimation. In the 65nm process, the additional area overhead is 16.05%, and the typical power consumption is 0.44W. The network operation time in the embodiment is 842ns, and the processing capacity is 4.85x106Xt / s. The analysis results show that the LSTM network computing device provided in this application achieves a very high energy efficiency ratio.
[0163] In the LSTM network computing device provided in the embodiment of this application, by storing the target input vector and the corresponding target weight vector in any storage array, the dot product summation module can read the target input vector and the corresponding target weight vector from any storage array, and obtain the sum value of the dot product results corresponding to any storage array. The activation operation module performs an activation operation on this sum value to obtain the output vector at the current moment, thereby implementing LSTM network calculation in memory, reducing the transmission of LSTM network data between memory and the processor, and solving the memory wall problem.
[0164] The LSTM network computing device can be installed on a computer device, and the computer device controls the LSTM network computing device, so that the computer device can complete the entire calculation process of the LSTM network on the LSTM network computing device. As Figure 14 shown, Figure 14 is a schematic structural diagram of a computing device provided in an embodiment of the present invention. The computing device 1400 includes:
[0165] A processor 1401, configured to receive the first input vector of the LSTM network at each moment, and input the first input vector at each moment into the memory 1402 connected to the processor. The first input vector at each moment includes the input data of the LSTM network at each moment;
[0166] The memory 1402 includes:
[0167] A storage module, which includes a first storage array and a second storage array. Any one of the first storage array and the second storage array includes multiple rows and multiple columns of dynamic random access memory (DRAM) cells. The first row of DRAM cells in any one of the storage arrays is used to store the target input vector of the hidden layer nodes of the LSTM network. The second row of DRAM cells in any one of the storage arrays is used to store the target weight vector corresponding to the target input vector. The target input vector includes the first input vector or the second input vector at the current moment. The second input vector includes the output data of the hidden layer node at the previous moment.
[0168] A dot product summation module, connected to the storage module, is used to perform dot product calculations on the target input vector and the target weight vector corresponding to the target input vector to obtain the dot product result corresponding to any one of the storage arrays, and perform summation calculations on the dot product results corresponding to each storage array to obtain a sum value.
[0169] An activation operation module, connected to the dot product summation module, is used to perform an activation operation on the sum value to obtain the output vector of the hidden layer node at the current moment.
[0170] Optionally, the first row of DRAM cells in any one of the storage arrays sequentially stores multiple target input vectors. The target weight vector includes multiple weight vectors. The second row of DRAM cells in any one of the storage arrays sequentially stores multiple weight vectors, and each weight vector corresponds to one target input vector stored in the first row of DRAM cells in any one of the storage arrays.
[0171] One DRAM cell in the first row of DRAM cells in any one of the storage arrays is used to store one bit value of one input data in one target input vector. One DRAM cell in the second row of DRAM cells in any one of the storage arrays is used to store one bit value of one weight in one weight vector.
[0172] Optionally, the activation operation module includes:
[0173] Multiple activation units, each activation unit is connected to the dot product summation module and is used to perform an activation operation on a sum value output by the dot product summation module. The sum value is the dot product calculation result of one target input vector and one weight vector corresponding to the one target input vector.
[0174] A register, which is used to store the state vector of the hidden node at the previous moment. The state vector at the previous moment is used to indicate the state of the hidden node at the previous moment.
[0175] An output unit, connected to the multiple activation units and the register, is configured to obtain the output vector at the current moment and the state vector of the hidden node at the current moment based on the activation results of the multiple activation units and the state vector of the previous moment stored in the register, where the state vector at the current moment is used to indicate the state of the hidden node at the current moment.
[0176] Optionally, the multiple activation units include an input activation circuit, a forgetting activation circuit, an update activation circuit, and an output activation circuit. The input activation circuit is configured to perform activation calculation for the input gate in the hidden layer node. The forgetting activation circuit is configured to perform activation calculation for the forgetting gate in the hidden layer node. The update activation circuit is configured to perform activation calculation for the update gate in the hidden layer node. The output activation circuit is configured to perform activation calculation for the output gate in the hidden layer node.
[0177] Optionally, the output unit includes:
[0178] A multiply-accumulate circuit, connected to the input activation circuit, the forgetting activation circuit, and the update activation circuit, is configured to perform multiplication calculation and addition calculation based on the activation results of the input activation circuit, the forgetting activation circuit, the update activation circuit, and the state vector of the hidden layer node at the previous moment to obtain the state vector at the current moment;
[0179] A target activation circuit, connected to the multiply-accumulate circuit, is configured to perform an activation operation on the state vector at the current moment output by the multiply-accumulate circuit;
[0180] A first multiplier, connected to the target activation circuit and the output activation circuit, performs a multiplication operation on the activation result of the target activation circuit and the activation result output by the output activation circuit to obtain the output vector at the current moment.
[0181] Optionally, the multiply-accumulate circuit includes:
[0182] A second multiplier, connected to the output ends of the input activation circuit and the update activation circuit, is configured to perform multiplication calculation on the activation result of the input activation circuit and the activation result of the update activation circuit;
[0183] A third multiplier, connected to the forgetting activation circuit and the register, is configured to perform multiplication calculation on the activation result of the forgetting activation circuit and the state vector of the previous moment stored in the register;
[0184] An adder, connected to the second multiplier and the third multiplier, is configured to perform an addition operation on the output result of the first multiplication and the output result of the third multiplier to obtain the state vector of the hidden layer node at the previous moment.
[0185] Optionally, the LSTM network computing device further includes:
[0186] A controller, connected to the driver, for controlling the driver to select one column of DRAM cells of any one of the storage arrays every preset time period;
[0187] The driver, connected to any one of the storage arrays, for writing the input data in the target input vector into the first row DRAM cells of each column of any one of the storage arrays and writing the weights in the weight vector corresponding to the target input vector into the second row DRAM cells of each column of any one of the storage arrays based on a target current every preset time period.
[0188] Optionally, the dot product summation module includes:
[0189] A plurality of dot product units, each dot product unit connected to multiple columns of DRAM cells in any one of the storage arrays, for performing dot product calculations on a target input vector stored in the DRAM cells of the first row of the multiple columns of DRAM cells and a weight vector stored in the DRAM cells of the second row of the multiple columns of DRAM cells;
[0190] And a summing unit, connected to the dot product unit, for performing summation calculations on the dot product results output by each dot product unit.
[0191] Optionally, each dot product unit includes:
[0192] A plurality of latches, the input end of each latch is connected to one column of DRAM cells in the multiple columns of DRAM cells, the output end of each latch is connected to one of a plurality of multipliers, and each latch is used to cache the value in one DRAM cell in the second row DRAM cells of any one of the storage arrays;
[0193] A plurality of multipliers, the first input end of each multiplier is respectively connected to one column of DRAM cells in the multiple columns of DRAM cells, the second input end of each multiplier is respectively connected to the output end of a latch, and the plurality of multipliers are respectively used to perform dot product calculations on the one target input vector and the one weight vector stored in the multiple columns of DRAM cells.
[0194] Optionally, each multiplier includes:
[0195] At least one exclusive-OR circuit, wherein a first input terminal of each exclusive-OR circuit is connected to a latch, a second input terminal of each exclusive-OR circuit is connected to one column of DRAM cells among the multiple columns of DRAM cells, and each exclusive-OR circuit is configured to perform an exclusive-OR calculation on a bit value of a weight in the one target weight vector cached by a latch and a bit value of an input data in the one input vector stored in one column of DRAM cells among the multiple columns of DRAM cells;
[0196] A combiner, a first input terminal of the combiner is connected to an output terminal of the at least one exclusive-OR circuit, a second input terminal of the combiner is connected to at least one latch, and the combiner is configured to combine the data output by the at least one exclusive-OR circuit and the data output by at least one latch to obtain a product of each input data in the one target input vector and a weight in the one weight vector.
[0197] A summing circuit, connected to the combiner, and configured to perform a summing calculation on the products output by the combiner multiple times.
[0198] Optionally, the summing unit includes:
[0199] A target adder, connected to the summing circuit, and configured to perform a summing calculation on the summing results output by the summing circuit multiple times;
[0200] A buffer, connected to the target adder, and configured to store the sum value calculated by the target adder each time.
[0201] The computing device provided by the embodiment of the present invention stores the target input vector and the corresponding target weight vector in any storage array. The dot product summing module can read the target input vector and the corresponding target weight vector from any storage array, and obtain the sum value of the dot product results corresponding to any storage array. The activation operation module performs an activation operation on the sum value to obtain the output vector at the current moment, thereby implementing the LSTM network calculation in the memory, reducing the transmission of LSTM network data between the memory and the processor, and solving the memory wall problem.
[0202] Figure 15 It is a schematic diagram of an implementation environment provided by the embodiment of the present invention. Figure 15The computer device therein includes a bus interface (peripheral component interconnect-express, PCI-e), a controller, and dual inline memory modules (DIMMs). Users can install the LSTM network computing device in the DIMM. Users can train and quantize the LSTM network according to the requirements of the application scenario to obtain a network structure and weights that meet the accuracy requirements. Then, the computer device uses a compiler to process the original data (such as pictures and audio), network structure, and weights of the application scenario. After processing, they are allocated to the memory address space, and an instruction set is generated. Then, the computer device writes the original data, weights, and instruction set into the LSTM network computing device. Finally, the LSTM network computing device automatically performs calculations. After the calculations are completed, it sends a message to the computer device's central processing units (CPU) through an interrupt. The CPU reads back the result data to complete the entire LSTM network computing operation. When compiling the original data, the compiler needs to call functions in the function library to perform the compilation. The compiler can compile the original data, etc. based on the programming language and send the compilation result to the driver of the computing device. Then, the driver controls the runtime to run the compilation result within the computer device. It should be noted that the application programming interface (API) in the user interface can also call functions in the function library to implement the functions of the API.
[0203] Figure 16 FIG. 4 is a schematic structural diagram of a computing device provided by an embodiment of the present invention. The computing device 1600 may vary greatly due to different configurations or performances, and may include one or more CPUs 1601 and one or more memories 1602. Among them, at least one instruction is stored in the memory 1602, and the at least one instruction is loaded and executed by the processor 1601 to implement the methods provided by the following various method embodiments. Of course, the LSTM network computing device 1600 may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input / output. The LSTM network computing device 1600 may also include other components for implementing the functions of the device, which will not be elaborated here.
[0204] In an exemplary embodiment, a computer-readable storage medium is further provided, such as a memory including instructions, and the above instructions can be executed by a processor in a terminal to complete the neural network calculation method in the following embodiments. For example, the computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0205] To further illustrate the process of a computing device implementing the calculations within an LSTM network, refer to Figure 17 the flowchart of an LSTM network calculation method provided by an embodiment of the present invention as shown in
[0206] 1701. The computing device receives a first input vector of the LSTM network, a first weight vector corresponding to the first input vector, a second input vector, and a second weight vector corresponding to the second input vector. The first input vector includes the input data of the LSTM network at the current moment, and the second input vector includes the output data of the hidden layer nodes at the previous moment.
[0207] This step 1701 can be executed by a processor in the computer device. The hidden layer is any hidden layer of the LSTM network to be calculated. It should be noted that at the initial moment, this step can be executed by the processor. When the time t is greater than 1, the processor can only receive the first input vector at time t, and does not receive the second input vector at time t. The second input vector at time t is calculated by the LSTM network computing device. Since the weights of the LSTM network remain unchanged at each moment, the computing device only needs to obtain the first weight vector and the second weight vector once, and there is no need to obtain them multiple times.
[0208] 1702. The LSTM network computing device stores the first input vector in the first row DRAM cells of the first storage array in the storage module, and stores the first weight vector in the second row DRAM of the first storage array, where the first storage array includes multiple rows and multiple columns of DRAM cells.
[0209] The first weight vector is also the weight vector corresponding to the first input vector. The process of storing data in any storage array has been introduced above, and this application embodiment does not elaborate on this step 1702 here. It should be noted that since the weights of the LSTM network remain unchanged at each moment, the LSTM network computing device only needs to store the first weight vector once, and there is no need to store it multiple times.
[0210] 1703. The LSTM network computing device stores the second input vector in the DRAM cells of the first row of the second storage array in the storage module, and stores the second weight vector in the DRAM of the second row of the second storage array, where the second storage array includes multiple rows and multiple columns of DRAM cells.
[0211] The second weight vector is also the weight vector corresponding to the second input vector. The process of storing data in any storage array has been introduced above. In this embodiment of the present application, step 1703 will not be elaborated here. It should be noted that at the initial moment, the LSTM network computing device stores the second input vector at the initial moment in the processor into the second storage array. At a non-initial moment, the LSTM network computing device may store the output vector at the current moment calculated in step 1705 into the second storage array, and use the output vector at the current moment as the second input vector at the next moment. It should be noted that since the weights of the LSTM network remain unchanged at each moment, the LSTM network computing device only needs to store the second weight vector once and does not need to store it multiple times.
[0212] 1704. The LSTM network computing device performs a dot product calculation on the target input vector and the weight vector corresponding to the target input vector based on the dot product summation module connected to the storage module, to obtain the dot product result corresponding to any storage array, and performs a summation calculation on the dot product results corresponding to each storage array to obtain a sum value. The target input vector includes the first input vector or the second input vector.
[0213] The process of the dot product summation module obtaining the sum value corresponding to the storage array has been introduced above. In this embodiment of the present application, step 1704 will not be elaborated here.
[0214] 1705. The LSTM network computing device performs an activation operation on the sum value based on the activation operation module connected to the dot product summation module, to obtain the output vector of the hidden layer node at the current moment.
[0215] The process of the dot activation operation module obtaining the output vector at the current moment has been introduced above. In this embodiment of the present application, step 1705 will not be elaborated here. It should be noted that if the current moment is not the last moment of the LSTM network calculation, the neural network computing device writes back the output vector of the current moment calculated by each hidden node to the second storage array, and writes the first input vector at the next moment into the first storage array for the calculation of the next moment; until the last moment, the LSTM network computing device obtains the output vectors of the forward hidden layer node and the backward hidden layer node at the last moment, and calculates the final output of the LSTM network based on formula (1) to implement the operation of the output layer. The process of the LSTM network computing device implementing the calculation of the output layer has been introduced above and will not be elaborated here.
[0216] In the method provided by the embodiment of the present application, by storing the target input vector and the corresponding target weight vector in any storage array, the dot product summation module can read the target input vector and the corresponding target weight vector from any storage array, and obtain the sum value of the dot product results corresponding to any storage array. The activation operation module performs an activation operation on the sum value to obtain the output vector at the current moment, thereby implementing the LSTM network calculation in the memory, reducing the transmission of LSTM network data between the memory and the processor, and solving the problem of the memory wall.
[0217] All of the above optional technical solutions can be combined arbitrarily to form optional embodiments of the present disclosure, which will not be elaborated herein one by one.
[0218] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing related hardware. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk, an optical disk, etc.
[0219] The above are only optional embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A long short-term memory (LSTM) network computing device, characterized in that, The LSTM network computing device includes: A storage module, the storage module includes a first storage array and a second storage array. The first storage array and the second storage array each include multiple rows and multiple columns of dynamic random access memory (DRAM) cells. The first row of DRAM cells in the first storage array is used to store the first input vector of the hidden layer nodes of the LSTM network. The second row of DRAM cells in the first storage array is used to store the target weight vector corresponding to the first input vector. The first row of DRAM cells in the second storage array is used for the second input vector of the hidden layer nodes. The second row of DRAM cells in the second storage array is used to store the target weight vector corresponding to the second input vector. The first input vector includes the input data of the LSTM network at the current moment. The second input vector includes the output data of the hidden layer nodes at the previous moment; A dot product summation module, connected to the storage module, for performing dot product calculations on the first input vector and the target weight vector corresponding to the first input vector to obtain the dot product result corresponding to the first storage array; performing dot product calculations on the second input vector and the target weight vector corresponding to the second input vector to obtain the dot product result corresponding to the second storage array; performing summation calculations on the dot product results corresponding to each storage array to obtain a sum value; An activation operation module, connected to the dot product summation module, for performing an activation operation on the sum value to obtain the output vector of the hidden layer nodes at the current moment.
2. The LSTM network computing device according to claim 1, characterized in that, There are multiple first input vectors, and the target weight vectors corresponding to the first input vectors include multiple weight vectors; The first row of DRAM cells in the first storage array sequentially stores multiple first input vectors, and the second row of DRAM cells in the first storage array sequentially stores the multiple weight vectors. Each weight vector corresponds to one of the first input vectors stored in the first row of DRAM cells in the first storage array; One DRAM cell in the first row of DRAM cells in the first storage array is used to store one bit value of one input data in one first input vector, and one DRAM cell in the second row of DRAM cells in the first storage array is used to store one bit value of one weight in one weight vector.
3. The LSTM network computing device according to claim 1, characterized in that, There are multiple second input vectors, and the target weight vectors corresponding to the second input vectors include multiple weight vectors; The first row of DRAM cells in the second storage array sequentially stores multiple second input vectors, and the second row of DRAM cells in the second storage array sequentially stores the multiple weight vectors. Each weight vector corresponds to one of the second input vectors stored in the first row of DRAM cells in the second storage array; One DRAM cell in the first row of DRAM cells in the second storage array is used to store one bit value of one input data in one second input vector, and one DRAM cell in the second row of DRAM cells in the second storage array is used to store one bit value of one weight in one weight vector.
4. The LSTM network computing device according to claim 2 or 3, characterized in that, The activation operation module includes: Multiple activation units, each activation unit is connected to the dot product summation module, and is used to perform an activation operation on a sum value output by the dot product summation module, where the sum value is the dot product calculation result of an input vector and a weight vector corresponding to the input vector; A register for storing the state vector of the hidden node at the previous moment, where the state vector at the previous moment is used to indicate the state of the hidden node at the previous moment; An output unit, connected to the multiple activation units and the register, and is used to obtain the output vector at the current moment and the state vector of the hidden node at the current moment based on the activation results of the multiple activation units and the state vector at the previous moment stored in the register, where the state vector at the current moment is used to indicate the state of the hidden node at the current moment.
5. The LSTM network computing device according to claim 4, characterized in that, The multiple activation units include an input activation circuit, a forgetting activation circuit, an update activation circuit, and an output activation circuit. The input activation circuit is used to implement the activation calculation of the input gate in the hidden layer node, the forgetting activation circuit is used to implement the activation calculation of the forgetting gate in the hidden layer node, the update activation circuit is used to implement the activation calculation of the update gate in the hidden layer node, and the output activation circuit is used to implement the activation calculation of the output gate in the hidden layer node.
6. The LSTM network computing device according to claim 5, characterized in that, The output unit includes: A multiply-accumulate circuit, connected to the input activation circuit, the forgetting activation circuit, and the update activation circuit, and is used to perform multiplication and addition calculations based on the activation results of the input activation circuit, the forgetting activation circuit, the update activation circuit, and the state vector of the hidden layer node at the previous moment to obtain the state vector at the current moment; A target activation circuit, connected to the multiply-accumulate circuit, and is used to perform an activation operation on the state vector at the current moment output by the multiply-accumulate circuit; A first multiplier, connected to the target activation circuit and the output activation circuit, and performs a multiplication operation on the activation result of the target activation circuit and the activation result output by the output activation circuit to obtain the output vector at the current moment.
7. The LSTM network computing device according to claim 6, characterized in that, The multiply-accumulate circuit includes: A second multiplier, connected to the output ends of the input activation circuit and the update activation circuit, and is used to perform a multiplication calculation on the activation result of the input activation circuit and the activation result of the update activation circuit; A third multiplier, connected to the forgetting activation circuit and the register, and is used to perform a multiplication calculation on the activation result of the forgetting activation circuit and the state vector at the previous moment stored in the register; An adder, connected to the second multiplier and the third multiplier, and is used to perform an addition operation on the output result of the first multiplication and the output result of the third multiplier to obtain the state vector of the hidden layer node at the previous moment.
8. The LSTM network computing device according to claim 2, characterized in that, The LSTM network computing device further includes: A controller, connected to the driver, and is used to control the driver to select a column of DRAM cells of the first storage array every preset time period; The driver, connected to the first storage array, is configured to write the input data in the first input vector into the first row DRAM cells of each column of the first storage array based on a target current every time the preset duration elapses, and write the weights in the target weight vector corresponding to the first input vector into the second row DRAM cells of each column of the first storage array.
9. The LSTM network computing device according to claim 3, characterized in that, The LSTM network computing device further includes: A controller, connected to the driver, is configured to control the driver to select a column of DRAM cells of the second storage array every time the preset duration elapses; The driver, connected to the second storage array, is configured to write the input data in the second input vector into the first row DRAM cells of each column of the second storage array based on a target current every time the preset duration elapses, and write the weights in the target weight vector corresponding to the second input vector into the second row DRAM cells of each column of the second storage array.
10. A computing device, characterized in that, The computing device includes: A processor, configured to receive the first input vector of the LSTM network at each moment and input the first input vector at each moment into the memory connected to the processor. The first input vector at each moment includes the input data of the LSTM network at each moment; The memory includes: A storage module, the storage module includes a first storage array and a second storage array. The first storage array and the second storage array each include multiple rows and multiple columns of dynamic random access memory (DRAM) cells. The first row DRAM cells in the first storage array are used to store the first input vector of the hidden layer nodes of the LSTM network at the current moment. The second row DRAM cells in the first storage array are used to store the target weight vector corresponding to the first input vector. The first row DRAM cells in the second storage array are used for the second input vector of the hidden layer nodes at the current moment. The second row DRAM cells in the second storage array are used to store the target weight vector corresponding to the second input vector. The second input vector includes the output data of the previous moment of the hidden layer nodes; A dot product summation module, connected to the storage module, is configured to perform dot product calculations on the first input vector and the target weight vector corresponding to the first input vector to obtain the dot product result corresponding to the first storage array; perform dot product calculations on the second input vector and the target weight vector corresponding to the second input vector to obtain the dot product result corresponding to the second storage array; perform summation calculations on the dot product results corresponding to each storage array to obtain a sum value; An activation operation module, connected to the dot product summation module, is configured to perform an activation operation on the sum value to obtain the output vector of the hidden layer nodes at the current moment.
11. The computing device according to claim 10, characterized in that, There are multiple first input vectors, and the target weight vectors corresponding to the first input vectors include multiple weight vectors; The first row of DRAM cells in the first memory array sequentially stores a plurality of first input vectors, and the second row of DRAM cells in the first memory array sequentially stores the plurality of weight vectors, with each weight vector corresponding to one of the first input vectors stored in the first row of DRAM cells in the first memory array; One DRAM cell in the first row of DRAM cells in the first memory array is used to store one bit value of an input data in one first input vector, and one DRAM cell in the second row of DRAM cells in the first memory array is used to store one bit value of a weight in one weight vector.
12. The computing device according to claim 10, characterized in that, There are a plurality of second input vectors, and the target weight vectors corresponding to the second input vectors include a plurality of weight vectors; The first row of DRAM cells in the second memory array sequentially stores a plurality of second input vectors, and the second row of DRAM cells in the second memory array sequentially stores the plurality of weight vectors, with each weight vector corresponding to one of the second input vectors stored in the first row of DRAM cells in the second memory array; One DRAM cell in the first row of DRAM cells in the second memory array is used to store one bit value of an input data in one second input vector, and one DRAM cell in the second row of DRAM cells in the second memory array is used to store one bit value of a weight in one weight vector.
13. The computing device according to claim 11 or 12, characterized in that, The activation operation module includes: A plurality of activation units, each activation unit is connected to the dot product summation module and is used to perform an activation operation on a sum value output by the dot product summation module, where the sum value is the dot product calculation result of an input vector and a weight vector corresponding to the input vector; A register for storing the state vector of the hidden node at the previous moment, and the state vector at the previous moment is used to indicate the state of the hidden node at the previous moment; An output unit, connected to the plurality of activation units and the register, is used to obtain the output vector at the current moment and the state vector of the hidden node at the current moment based on the activation results of the plurality of activation units and the state vector at the previous moment stored in the register, and the state vector at the current moment is used to indicate the state of the hidden node at the current moment.
Citation Information
Patent Citations
Apparatus and method for performing LSTM operations
CN109284825A
Runoff prediction method based on LSTM (Long Short Term Memory) multi-state vector sequence-to-sequence model
CN112949902A
Heterogeneous hybrid memory data processing method and system based on DRAM and NVM and storage medium
CN113867633A
In-memory calculation circuit and device based on transposed DRAM (Dynamic Random Access Memory) unit
CN116483773A
Method, system and equipment for generating text abstract and storage medium
CN117520535A