Storage and computing integrated operation method, memristor neural network chip and storage medium
By introducing analog storage and computing macro units and hybrid storage and computing macro units into the memristor neural network chip and combining them with analog circuits to transmit data, the problem of high analog-to-digital conversion overhead in the memristor storage and computing integrated chip is solved, and energy efficiency is improved.
Patent Information
- Application Number
- CN202210108371.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-28
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2042-01-28
AI Technical Summary
Existing memristor storage and computing integrated chips have large analog-to-digital conversion overhead during data movement, which limits the chip's energy efficiency.
A combination of analog storage and computing macro cells and hybrid storage and computing macro cells is adopted. By applying analog voltage on the memristor array and converting it into analog current, combined with clamping, subtraction and analog-to-digital conversion, efficient data transmission is achieved.
The peripheral circuits in the chip are reduced, energy consumption is reduced, and energy efficiency is improved.
Smart Images

Figure CN114418080B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of integrated storage and computing chips, and in particular to an integrated storage and computing method, a memristor neural network chip, and a storage medium. Background Art
[0002] In recent years, the field of artificial intelligence has made great progress through the use of deep learning. At the AI chip architecture level, data exchange between storage units and computing units is unavoidable.
[0003] Currently, memristor-based integrated computing and storage architectures are often used to eliminate the data movement between computing and storage units in traditional von Neumann architectures. However, the data movement in memristor-based integrated computing and storage chips incurs significant analog-to-digital conversion overhead, significantly limiting the chip's energy efficiency. Summary of the Invention
[0004] The embodiments of the present application provide a storage-computing integrated operation method, a memristor neural network chip, and a storage medium.
[0005] The technical solution of the embodiment of the present application is implemented as follows:
[0006] An embodiment of the present application provides a memristor neural network chip, comprising: at least one analog storage and computing macro unit and at least one hybrid storage and computing macro unit, wherein at least one of the analog storage and computing macro units is connected to at least one of the hybrid storage and computing macro units;
[0007] The at least one analog storage and computing macro unit is used to apply an input analog voltage to the memristor array in the unit, and convert the generated analog current into an analog voltage within a preset range and then output it;
[0008] The at least one hybrid storage and computing macro unit is used to apply the analog voltage output by the at least one analog storage and computing macro unit to the memristor array in the unit, and to clamp, subtract, and convert the generated analog current in sequence before outputting it.
[0009] The present application provides a storage-computation-integrated computing method, which is applied to the above-mentioned memristor neural network chip. The method includes:
[0010] Using at least one analog storage and computing macrocell, an analog voltage is applied to the memristor array in the cell, and the generated analog current is converted into an analog voltage within a preset range and then output;
[0011] At least one hybrid storage and computing macro unit is used to apply the analog voltage output by the at least one analog storage and computing macro unit to the memristor array in the unit, and the generated analog current is clamped, subtracted, and analog-to-digital converted in sequence before output.
[0012] An embodiment of the present application provides a computer-readable storage medium on which a computer program is stored, and is characterized in that when the computer program is executed, the above-mentioned storage-computation-in-one operation method is implemented.
[0013] The embodiment of the present application provides a storage-computation integrated operation method, a memristor neural network chip and a storage medium, wherein the memristor neural network chip includes: at least one analog storage-computation macro unit and at least one hybrid storage-computation macro unit, at least one of the analog storage-computation macro unit is connected to at least one hybrid storage-computation macro unit; the at least one analog storage-computation macro unit applies an input analog voltage to the memristor array in the unit, and converts the generated analog current into an analog voltage within a preset range and outputs it; the at least one hybrid storage-computation macro unit is used to apply the analog voltage output by the at least one analog storage-computation macro unit to the memristor array in the unit, and clamps, subtracts, and performs analog-to-digital conversion on the generated analog current in sequence and outputs it. The memristor neural network chip provided by the embodiment of the present application transmits data based on analog circuits, reduces the peripheral circuits in the chip, reduces the energy consumption of the chip, and improves the energy efficiency of the chip. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 A schematic structural diagram of a memristor neural network chip provided in an embodiment of the present application;
[0015] Figure 2 A schematic diagram of the structure of an analog storage and computing macro unit and a hybrid storage and computing macro unit provided in an embodiment of the present application;
[0016] Figure 3 A schematic structural diagram of an analog direct transmission module provided in an embodiment of the present application;
[0017] Figure 4 A schematic diagram of a convolutional layer network structure provided in an embodiment of the present application;
[0018] Figure 5 A schematic diagram of a fully connected layer network structure provided in an embodiment of the present application;
[0019] Figure 6 An exemplary convolutional layer data flow diagram provided in an embodiment of the present application;
[0020] Figure 7 A schematic diagram of the structure of a functional unit provided in an embodiment of the present application;
[0021] Figure 8 A flowchart of a storage-computing-in-one operation method provided in an embodiment of the present application. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0023] The following will specifically describe the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems through embodiments and in conjunction with the accompanying drawings. The following embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0024] In addition, the technical solutions described in the embodiments of the present application can be arbitrarily combined without conflict.
[0025] An embodiment of the present application provides a memristor neural network chip. Figure 1 This is a schematic diagram of the structure of a memristor neural network chip provided in an embodiment of the present application. Figure 1 As shown, in an embodiment of the present application, the memristor neural network chip includes: at least one analog storage and computing macro unit 10 and at least one hybrid storage and computing macro unit 11, and at least one of the analog storage and computing macro units 10 is connected to at least one of the hybrid storage and computing macro units 11;
[0026] At least one analog storage and computing macro unit 10, used to apply an input analog voltage to the memristor array 20 in the unit, and convert the generated analog current into an analog voltage within a preset range and output it;
[0027] At least one hybrid storage and computing macro unit 11 is used to apply the analog voltage output by at least one analog storage and computing macro unit 10 to the memristor array in the unit, and clamp, subtract, and convert the generated analog current in sequence before outputting it.
[0028] It should be noted that, in the embodiments of the present application, Figure 1 As shown, the memristor neural network chip may further include: a multiplexer group 12 and a controller 13. Among at least one analog storage and computing macro unit 10 and at least one hybrid storage and computing macro unit 11, the storage and computing macro units that are connected can be connected through the multiplexer group 12. The specific connection method can be set according to actual needs and application scenarios, and is not limited in the embodiments of the present application. The controller 13 is connected to at least one analog storage and computing macro unit 10 and at least one hybrid storage and computing macro unit 11, and is used to control at least one analog storage and computing macro unit 10 and at least one hybrid storage and computing macro unit 11.
[0029] Specifically, in an embodiment of the present application, at least one analog storage and computing macrocell 10 includes multiple cascaded analog storage and computing macrocells 10, wherein the analog storage and computing macrocell 10 of the last stage is connected to at least one hybrid storage and computing macrocell 11. The analog voltage input to the storage and computing macrocell of the next stage is the analog voltage output by the storage and computing macrocell of the previous stage to which it is connected.
[0030] It should be noted that, in the embodiments of the present application, the difference in the functions of the analog storage and computing macro unit 10 and the hybrid storage and computing macro unit 11 is due to the different structures of the two. The hybrid storage and computing macro unit 11 includes some modules of the analog storage and computing macro unit 10 as well as modules such as analog-to-digital conversion. The hybrid storage and computing macro unit 11 actually has two uses. One is for the last layer output of the convolutional layer network part. Because the output from the convolutional layer to the fully connected layer requires waiting until all the output data is ready, it is necessary to store the calculated data and wait for the subsequent data to be calculated. The other is for implementing the last layer of the network layer of the fully connected classifier. Because the last layer needs to be output to the outside of the chip and stored, modules such as analog-to-digital conversion are required. The following is combined with Figure 2 The structures of the analog storage and computing macro unit 10 and the hybrid storage and computing macro unit 11 are described in detail.
[0031] Specifically, in the embodiments of the present application, Figure 2 As shown, at least one analog memory-computing macrocell 10 and at least one hybrid memory-computing macrocell 11 each include a memristor array 20 capable of performing vector-matrix multiplication operations. Applying an analog voltage representing a vector to the memristor array 20 is equivalent to performing the vector-matrix multiplication operation for the vector. The memristor array 20 includes word lines (WL), bit lines (BL), and source lines (SL). The word lines WL are used to control array activation, the bit lines BL are used for analog voltage input, and the source lines SL are used to output the generated analog current.
[0032] It should be noted that in the embodiments of the present application, for at least one analog storage and computing macro unit 10, the type and specific number of the memristor array 20 included therein can be set according to actual needs. For at least one hybrid storage and computing macro unit 11, the memristor array 20 included therein can be composed of two single-transistor single-memristor arrays. Of course, it can also be set according to actual conditions. The size of the memristor array 20 included in at least one analog storage and computing macro unit 10 and at least one hybrid storage and computing macro unit 11 can be set according to the network layer in the neural network implemented by the storage and computing macro unit.
[0033] Specifically, in the embodiments of the present application, Figure 2 As shown, at least one analog storage and computing macro unit 10 includes: an analog direct transmission module 101 connected to the memristor array 20.
[0034] Specifically, such as Figure 3 As shown, in the embodiment of the present application, the analog direct transmission module 101 includes: an integrator 1011, a switched capacitor unity gain buffer 1012, a sense amplifier 1013, and a two-way multiplexer 1014;
[0035] The input end of the integrator 1011 can be connected to the analog current generated by the memristor array 20 , and the output end of the integrator 1011 can be connected to the input end of the switched capacitor unity gain buffer 1012 ;
[0036] A sense amplifier 1013 can be connected between the output of the switched capacitor unity gain buffer 1012 and the two-way multiplexer 1014;
[0037] The two-way multiplexer 1014 is connected to the output end of the switched capacitor unity gain buffer 1012 .
[0038] It should be noted that if Figure 3 As shown, in an embodiment of the present application, four switches S1, S2, S3 and S4 are deployed in the analog direct transmission module 101. By opening and closing the switches, the connection or disconnection between the devices connected at both ends can be achieved, that is, when the switch is closed, the connection between the devices connected at both ends can be achieved, and when the switch is open, the disconnection between the devices connected at both ends can be achieved.
[0039] Specifically, in the embodiment of the present application, the integrator 1011 is used to integrate the analog current generated by the memristor array 20 to obtain an integrated voltage;
[0040] A switched capacitor unity gain buffer 1012, for maintaining and outputting an integrated voltage;
[0041] A sense amplifier 1013 is used to compare the integrated voltage with a threshold voltage;
[0042] The two-way multiplexer 1014 is used to select the maximum voltage between the integrated voltage and the threshold voltage for output.
[0043] It should be noted that, in the embodiments of this application, reference Figure 3 The analog direct transmission module 101 structure shown in the figure works as follows: when the switch S1 and the switch S2 are closed, the integrator 1011 is set, then the switch S1 is closed, the switch S2 is opened, and the integrator 1011 starts to integrate, that is, specifically integrates the input analog current on the integration capacitor, then the switch S1 and the switch S2 are opened, the switch S3 and the switch S4 are closed, and the integrated voltage completes the linear rectification and is output.
[0044] It should be noted that, in the embodiments of the present application, Figure 3 As shown, V ref is the reference voltage, V th is the threshold voltage.
[0045] It should be noted that, in the embodiments of the present application, Figure 2 As shown, at least one analog storage and computing macro unit 10 further includes a unity gain buffer 30 , which is used to load the input analog voltage onto the memristor array 20 .
[0046] It can be understood that in an embodiment of the present application, in at least one analog storage and computing macro unit 10, the unit gain buffer 30 applies the input analog voltage to the internally connected memristor array 20, and then the generated analog current is converted through the analog direct transmission module 101 to output an analog voltage within a preset range.
[0047] Specifically, in the embodiments of the present application, Figure 2 As shown, at least one hybrid storage and computing macro unit 11 further includes: a unit gain buffer 30 connected to the memristor array 20, a current subtractor 111 connected to the unit gain buffer 30, and an analog-to-digital converter 112 connected to the current subtractor 111;
[0048] a unity gain buffer 30 for clamping the analog current generated by the memristor array 20 to obtain a clamped current;
[0049] a current subtractor 111 for performing a subtraction operation on the clamped current to obtain an output current;
[0050] The analog-to-digital converter 112 is used to convert the output current analog-to-digital into a digital signal for output.
[0051] It should be noted that in an embodiment of the present application, at least one hybrid storage and computing macro unit 11 also includes a unit gain buffer 30, a current subtractor 111 and an analog-to-digital converter 112, wherein the number of current subtractors 111 and analog-to-digital converters 112, as well as the unit gain buffer 30, needs to be set according to the network layer of the neural network implemented by the hybrid storage and computing macro unit 11, and is not limited in the embodiment of the present application.
[0052] It can be understood that in an embodiment of the present application, the memristor array 20 in at least one hybrid storage and computing macro unit 11 can be two single-transistor single-memristor arrays, which can only realize positive weights. Among them, for a hybrid storage and computing macro unit 11, the analog voltage output by the upper-level storage and computing macro unit connected to it will be input into each single-transistor single-memristor array, and the analog current output by each single-transistor single-memristor array needs to be fed back using a unit gain buffer 30 to clamp the current, and then the two currents are copied to the current subtractor 111 using a current mirror, and the voltage obtained by subtraction is output to the analog-to-digital converter 112 using the current subtractor 111 to be converted into a digital signal, so that it can be cached on the chip or outside the chip.
[0053] The following is a detailed description of the structure of the memristor neural network chip in conjunction with the network layer of the neural network.
[0054] Figure 4 A schematic diagram of a convolutional layer network structure provided in an embodiment of the present application. Figure 4 As shown, in an embodiment of the present application, at least one analog storage and computing macro unit 10 includes K2×K2 analog storage and computing macro units 10, the K2×K2 analog storage and computing macro units include a memristor array 20 of size (C×K1×K1, N), and the K2×K2 analog storage and computing macro units 10 are divided into K2 groups of analog storage and computing macro units 10, each group including K2 analog storage and computing macro units 10;
[0055] At least one hybrid memory-computing macro unit 11 includes a first hybrid memory-computing macro unit, and the first hybrid memory-computing macro unit includes a memristor array 20 of a size of (N×K2×K2, M);
[0056] Each group of analog storage and computing macro units 10 in the K2 groups of analog storage and computing macro units 10 is connected to the first hybrid storage and computing macro unit respectively; K1, C, N, K2 and M are natural numbers greater than or equal to 1.
[0057] It should be noted that the K2 group of analog storage and computing macro units 10 have the computing function of realizing a convolution layer with a convolution kernel size of K1×K1, the number of input channels of C, and the number of output channels of N; the first hybrid storage and computing macro unit has the computing function of realizing a convolution layer with a convolution kernel size of K2×K2, the number of input channels of N, and the number of output channels of M.
[0058] It should be noted that, in the embodiment of the present application, the first hybrid storage and computing macro unit is a specific hybrid storage and computing macro unit 11 including a memristor array 20 of a size of (N×K2×K2, M).
[0059] It should be noted that, in the embodiments of the present application, Figure 4As shown in the figure, for the hardware architecture of the first convolutional layer, to reduce data latency, the first convolutional layer is replicated K2×K2 times, divided into K2 groups, each containing K2 analog memory / computation macrocells 10 of size (C×K1×K1,N). The array word lines (WL) are transistor gate terminals used to control array activation. The array bit lines (BL) are transistor drain terminals connected to one end of the memristor for analog voltage input. The array source lines (SL) are transistor sources for analog current output. The analog input voltage is 1 / 2 VDD ± Vread, and the input value is applied to the array BL via a unity-gain buffer 30. During inference, all WLs are enabled, and the vector-matrix multiplication of the input vector and the weight matrix is performed via the memristor array 20. The output analog current is obtained at SL and converted into an analog voltage within the range of 1 / 2 VDD ± Vread by the analog direct transfer module 101. The analog direct transfer module 101 operates in two phases: first, the sampling phase, in which only the integrator 1011 is switched on. The output analog current is integrated on the integrating capacitor, with a designed integration time of 10ns. During the subsequent hold phase, the switch switches to only the switched capacitor unity-gain buffer 1012, maintaining the integrated voltage on integrator 1011 and outputting it to the next layer. Simultaneously, the sense amplifier 1013 compares the integrated voltage with a threshold voltage. If the integrated voltage exceeds the threshold voltage, the multiplexer selects the output of the integrated voltage; otherwise, it selects the output of the threshold voltage, thus fulfilling the activation function of the convolutional layer. To fully utilize the already solved results, the first layer's K2 convolution groups operate sequentially, meaning only one convolution array operates at a time, while the other convolution arrays retain the results. Therefore, a K2-way multiplexer is required to integrate and output the results of each of the K2 convolution groups.
[0060] It should be noted that, in the embodiments of the present application, Figure 4 As shown, for the hardware architecture of the second convolutional layer, i.e., the first hybrid memory-computing macrocell, its input is the analog voltage output by K2 groups of analog memory-computing macrocells 10. Specifically, the analog voltage output by K2 groups of analog memory-computing macrocells 10 is respectively input to the BL of two 1T1R memristor arrays 20 in the first hybrid memory-computing macrocell. During operation, both arrays have their WLs turned on. Since 1T1R can only achieve positive weights, it is necessary to clamp the output current before subtraction. Specifically, a unity gain buffer 30 is used as feedback to clamp the current. The two currents are then copied to the current subtractor 111 using a current mirror. The subtracted voltage is then converted into a digital signal using the analog-to-digital converter 112.
[0061] Figure 5 This is a schematic diagram of a fully connected layer network structure provided in an embodiment of the present application. Figure 5As shown, in an embodiment of the present application, at least one analog storage and computing macro unit 10 includes a first analog storage and computing macro unit and a second analog storage and computing macro unit, the first analog storage and computing macro unit includes a memristor array 20 of size (S, L), and the second analog storage and computing macro unit includes a memristor array 20 of size (L, P);
[0062] At least one hybrid memory-computing macro unit 11 includes a second hybrid memory-computing macro unit, and the second hybrid memory-computing macro unit includes a memristor array 20 of size (P, Q);
[0063] The second analog storage and computing macro unit is connected between the first analog storage and computing macro unit and the second hybrid storage and computing macro unit; S, L, P and Q are natural numbers greater than or equal to 1.
[0064] It should be noted that, in the embodiments of the present application, the first analog storage and computing macro unit has the computing function of realizing a fully connected layer with S input neurons and L output neurons; the second analog storage and computing macro unit has the computing function of realizing a fully connected layer with L input neurons and P output neurons; the second hybrid storage and computing macro unit has the computing function of realizing a fully connected layer with P input neurons and Q output neurons.
[0065] It should be noted that in the embodiments of the present application, the hardware architecture of the first fully connected layer is similar to the hardware architecture of the first convolutional layer of a network of two consecutive convolutional layers, except that there is no data waiting in the fully connected layer, so no copying is required.
[0066] It should be noted that in the embodiment of the present application, for the second-layer fully connected layer hardware architecture, if there are only two fully connected layers, its structure is the same as the second-layer convolutional layer hardware architecture of two consecutive convolutional layer networks. If there is a third fully connected layer, the simulated direct transmission module 101 is connected to the corresponding hybrid storage and computing macro unit 11 to continue processing and output.
[0067] It should be noted that in the embodiments of the present application, the analog storage and computing macro unit 10 and the hybrid storage and computing macro unit 11 can both include a current mirror. In addition, the analog storage and computing macro unit 10 can also include a current subtractor 111. The specific setting can be based on actual needs and is not limited in the embodiments of the present application.
[0068] The above-mentioned memristor neural network chip is further described in detail below with reference to a specific example.
[0069] In the embodiment of the present application, the structure of the preset neural network is shown in Table 1.
[0070]
[0071] As shown in Table 1, K is the convolution kernel size; A is the number of input channels of the convolution layer; B is the number of output channels of the convolution layer; IN and OUT are the number of input neurons and output neurons of the fully connected layer, respectively; Row and Col are the number of rows and columns of the mapped memristor array 20, respectively, replic is the number of arrays, which defaults to 1; C, H, and W are the number of channels, height, and width of the convolution layer input feature map, respectively.
[0072] For the preset neural network shown in Table 1, Figure 6 As shown, for the first convolutional layer in the hardware architecture of a network with two consecutive convolutional layers, to reduce data latency, the first convolutional layer is replicated 25 times, divided into five groups, each with five analog memory and computation macrocells 10 of size (75, 6). For the second convolutional layer in the hardware architecture of a network with two consecutive convolutional layers, a hybrid memory and computation macrocell 11 of size (150, 16) is used. The results of the first convolutional layer are input to the BLs of the two 1T1R memristor arrays 20 in the hybrid memory and computation macrocell 11. For an image in the dataset, the input feature map size is 32x32x3. When input to the hardware of the first convolutional layer, the input feature map is expanded and sequentially input to the array.
[0073] Specifically, in the embodiments of the present application, Figure 6 The 13x13x3 block marked by the black box in the input feature map is the first input to the first convolutional layer array. The small blocks in the large block are 5x5x3, each large block has 25 small blocks, and the stride between small blocks is 2. The gray part in the black box is 5 small blocks [3,5,5]. After being expanded into a vector of 5 dimensions
[75] , it will be input to the first group of 5 (75,6) analog storage macro units 10. Similarly, the 5 3x5x5 blocks in the other four columns will be input to the second to fifth groups of analog storage macro units 10 respectively. Then, five groups of multiplexers select the outputs of the first, second, third, fourth, and fifth groups of arrays respectively to the second convolutional layer.
[0074] During the second input, the black box shifts two steps to the right, so only the rightmost column (the fifth column) of the large block needs to be calculated. The calculation results for the other columns have already been calculated in the first beat and stored on the integrating capacitors of the analog direct transmission modules 102 of the second to fifth array groups. The five [3,5,5] small blocks in the fifth column are expanded into a five-dimensional
[75] vector and input to the first group of five (75,6) analog storage macro units 10. The other second to fifth array groups retain the previous results. Then, the five multiplexers select the outputs of the second, third, fourth, fifth, and first array groups, respectively, to be sent to the second convolutional layer.
[0075] During the third input, the black box shifts two steps to the right. The calculations for the rightmost column (the fifth column) of the large block now require computation. The results for the other columns have already been calculated during the first and second beats and stored on the integrating capacitors of the analog direct transmission modules 102 for the third to fifth and first array groups. The five [3,5,5] blocks in the fifth column are expanded into a five-dimensional
[75] vector and then input into the second group of five (75,6) analog storage macrocells 10. The other three to fifth and first array groups retain their previous results. Then, the five multiplexers select the outputs of the third, fourth, fifth, first, and second array groups, respectively, and send them to the second convolutional layer. The inputs thereafter follow this pattern.
[0076] It should be noted that, in the embodiment of the present application, the memristor neural network chip is suitable for recurrent neural networks. In this case, Figure 1 and Figure 7 As shown, the memristor neural network chip also includes a functional unit 14.
[0077] Specifically, such as Figure 7 As shown, in the embodiment of the present application, at least one analog storage and computing macro unit 10 includes: an analog storage and computing macro unit 10 for input and an analog storage and computing macro unit 10 for loop, at least one hybrid storage and computing macro unit includes: a hybrid storage and computing unit for output, and the memristor neural network also includes:
[0078] A functional unit 14 connected to the analog storage and computing macro unit 10 for input, the analog storage and computing macro unit 10 for loop, and the hybrid storage and computing macro unit 11 for output;
[0079] The functional unit 14 is used to obtain the total analog current output by the analog storage and computing macro unit 10 for input and the analog storage and computing macro unit 10 for circulation, and use the total analog current to generate an output voltage, and output the output voltage to the hybrid storage and computing macro unit 11 for output.
[0080] It should be noted that in the embodiment of the present application, the functional unit 14 can not only output the output voltage to the hybrid storage and computing unit 11 for output so that it can continue to complete subsequent processing, but also use the output voltage to drive the analog storage and computing macro unit 10 for circulation, so that the analog storage and computing macro unit 10 for circulation can work in a cycle.
[0081] Specifically, such as Figure 7 As shown, in the embodiment of the present application, the functional unit 14 includes: two analog direct transmission modules 101 and an activation analog circuit 141;
[0082] The activation analog circuit 141 is connected between the two analog direct transmission modules 101;
[0083] Of the two analog direct transmission modules 101 , one is an input direct transmission module and the other is an output direct transmission module;
[0084] An input direct transmission module is connected to the analog storage and calculation macro unit 10 for input and the analog storage and calculation macro unit 10 for circulation, and is used to obtain the total analog current, convert the total analog current into an analog voltage within a preset range, and output it;
[0085] The activation analog circuit 141 is used to activate the analog voltage output by the input direct transmission module and generate an activation current output;
[0086] The output direct transmission module is used to convert the activation current into an analog voltage within a preset range as the output voltage.
[0087] The embodiment of the present application provides a memristor neural network chip, comprising: at least one analog storage and computing macro unit and at least one hybrid storage and computing macro unit, at least one of the analog storage and computing macro unit is connected to at least one of the hybrid storage and computing macro unit; the at least one analog storage and computing macro unit applies an input analog voltage to the memristor array in the unit, and converts the generated analog current into an analog voltage within a preset range and outputs it; the at least one hybrid storage and computing macro unit is used to apply the analog voltage output by the at least one analog storage and computing macro unit to the memristor array in the unit, and sequentially clamps, subtracts, and performs analog-to-digital conversion on the generated analog current and outputs it. The memristor neural network chip provided by the embodiment of the present application transmits data based on analog circuits, reduces peripheral circuits in the chip, reduces the energy consumption of the chip, and improves the energy efficiency of the chip.
[0088] The embodiment of the present application also provides a storage and computing integrated method, which is applied to the above-mentioned memristor neural network chip. Figure 8 This is a flow chart of a storage and computing method provided in an embodiment of the present application. Figure 8 As shown, in an embodiment of the present application, the storage-computation-integrated computing method includes:
[0089] S801. Using at least one analog storage and computing macro unit, apply an input analog voltage to the memristor array in the unit, and convert the generated analog current into an analog voltage within a preset range and then output it.
[0090] It should be noted that in the embodiments of the present application, as described in the structure of the above-mentioned memristor neural network chip, at least one analog storage and computing macro unit 10 is capable of applying the input analog voltage to the memristor array 20 to realize vector-matrix multiplication operations, and converting the generated analog current into an analog voltage within a preset range and then outputting it, wherein the calculation steps implemented by the specific modules and devices in each analog storage and computing macro unit 10 are not repeated here.
[0091] S802. Using at least one hybrid storage-computing macro unit, apply the analog voltage output by at least one analog storage-computing macro unit to the memristor array in the unit, and sequentially clamp, subtract, and perform analog-to-digital conversion on the generated analog current before outputting it.
[0092] It should be noted that in the embodiments of the present application, as described in the structure of the above-mentioned memristor neural network chip, at least one hybrid storage and computing macro unit 11 is capable of applying the analog voltage output by at least one analog storage and computing macro unit 10 to the memristor array 20 to realize vector-matrix multiplication operations, and the generated analog current is clamped, subtracted, and analog-to-digital converted in sequence and then output, wherein the calculation steps implemented by the specific modules and devices in each hybrid storage and computing macro unit 11 are not repeated here.
[0093] It should be noted that in the embodiment of the present application, the memristor neural network chip uses a controller 13 to control the operation of different analog storage and computing macro units 10 and different hybrid storage and computing macro units 11, which is not limited in the embodiment of the present application.
[0094] It can be understood that in the embodiments of the present application, the memristor neural network chip uses at least one analog storage and computing macro unit 10 and at least one hybrid storage and computing macro unit 11 to transmit data based on analog circuits, which reduces the peripheral circuits in the chip, reduces the energy consumption of the chip, and improves the energy efficiency of the chip.
[0095] The embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the above-mentioned storage-computation-integrated operation method when executed by a processor. The computer-readable storage medium can be a volatile memory (volatile memory), such as a random-access memory (Random-Access Memory, RAM); or a non-volatile memory (non-volatile memory), such as a read-only memory (Read-Only Memory, ROM), a flash memory (flash memory), a hard disk (Hard Disk Drive, HDD) or a solid-state drive (SSD); or it can be a respective device including one or any combination of the above memories, such as a mobile phone, a computer, a tablet device, a personal digital assistant
[0096] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of hardware embodiments, software embodiments, or embodiments combining software and hardware. Furthermore, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage and optical storage, etc.) containing computer-usable program code.
[0097] The present application is described with reference to the implementation flow charts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flow charts and / or block diagrams, as well as the combination of processes and / or boxes in the flow charts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the implementation flow charts. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0098] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which is implemented in the implementation flow diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0099] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process described in the flowchart. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0100] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this utility model should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A memristor neural network chip, characterized in that: include: At least one analog storage and computing macro unit and at least one hybrid storage and computing macro unit, at least one of the analog storage and computing macro unit is connected to at least one of the hybrid storage and computing macro unit; The at least one analog storage and computing macro unit is used to apply an input analog voltage to the memristor array in the unit, and convert the generated analog current into an analog voltage within a preset range and then output it; The at least one hybrid storage and computing macro unit is configured to apply the analog voltage output by the at least one analog storage and computing macro unit to the memristor array within the unit, and sequentially clamp, subtract, and perform analog-to-digital conversion on the generated analog current before outputting the resultant analog current; The at least one analog storage and computing macro unit includes: an analog direct transmission module connected to the memristor array; The analog direct transmission module is used to integrate the analog current generated by the memristor array to obtain an integrated voltage, compare the integrated voltage with a threshold voltage, and select the largest voltage between the integrated voltage and the threshold voltage for output.
2. The memristor neural network chip according to claim 1, characterized in that: The at least one analog storage and computing macro unit includes a plurality of cascaded analog storage and computing macro units, wherein the analog storage and computing macro unit at the last level is connected to the at least one hybrid storage and computing macro unit.
3. The memristor neural network chip according to claim 1, characterized in that: The memristor array included in the at least one analog storage-computing macro unit and the at least one hybrid storage-computing macro unit has the function of realizing vector-matrix multiplication operations.
4. The memristor neural network chip according to claim 1, characterized in that: The analog direct transmission module includes: an integrator, a switched capacitor unity gain buffer, a sensitive amplifier and a two-way multiplexer; The input end of the integrator can be connected to the analog current generated by the memristor array, and the output end of the integrator can be connected to the input end of the switched capacitor unity gain buffer; The sense amplifier can be connected between the output end of the switched capacitor unity gain buffer and the two-way multiplexer; The two-way multiplexer is connected to the output end of the switched capacitor unity gain buffer.
5. The memristor neural network chip according to claim 4, characterized in that: The integrator is used to integrate the analog current generated by the memristor array to obtain an integrated voltage; The switched capacitor unity gain buffer is used to maintain and output the integrated voltage; The sensitive amplifier is used to compare the integrated voltage with a threshold voltage; The two-way multiplexer is used to select the maximum voltage between the integrated voltage and the threshold voltage for output.
6. The memristor neural network chip according to claim 1, characterized in that: The at least one hybrid storage-computation macro unit further includes: a unit gain buffer connected to the memristor array, a current subtractor connected to the unit gain buffer, and an analog-to-digital converter connected to the current subtractor; The unity gain buffer is used to clamp the analog current generated by the memristor array to obtain a clamping current; The current subtractor is used to perform a subtraction operation on the clamped current to obtain an output current; The analog-to-digital converter is used to convert the output current into a digital signal and then output it.
7. The memristor neural network chip according to claim 1, characterized in that: include: The at least one analog storage and computing macro unit includes K2×K2 analog storage and computing macro units, the K2×K2 analog storage and computing macro units include a memristor array of size (C×K1×K1, N), and the K2×K2 analog storage and computing macro units are divided into K2 groups of analog storage and computing macro units, each group including K2 analog storage and computing macro units; The at least one hybrid storage-computing macro unit includes a first hybrid storage-computing macro unit, wherein the first hybrid storage-computing macro unit includes a memristor array of size (N×K2×K2, M); Each group of analog storage and computing macro units in the K2 groups of analog storage and computing macro units is connected to the first hybrid storage and computing macro unit respectively; K1, C, N, K2, and M are natural numbers greater than or equal to 1.
8. The memristor neural network chip according to claim 7, characterized in that: The K2 group of analog storage and computing macro units has the computing function of realizing a convolution layer with a convolution kernel size of K1×K1, a number of input channels of C, and a number of output channels of N; The first hybrid storage and computing macro unit has the computing function of realizing a convolution layer with a convolution kernel size of K2×K2, an input channel number of N, and an output channel number of M.
9. The memristor neural network chip according to claim 1, characterized in that: The at least one analog storage and computing macro unit includes a first analog storage and computing macro unit and a second analog storage and computing macro unit, the first analog storage and computing macro unit includes a memristor array of size (S, L), and the second analog storage and computing macro unit includes a memristor array of size (L, P); The at least one hybrid storage-computing macro unit includes a second hybrid storage-computing macro unit, and the second hybrid storage-computing macro unit includes a memristor array of size (P, Q); The second analog storage and computing macro unit is connected between the first analog storage and computing macro unit and the second hybrid storage and computing macro unit; S, L, P, and Q are natural numbers greater than or equal to 1.
10. The memristor neural network chip according to claim 9, characterized in that: The first analog storage and calculation macro unit has the computing function of realizing a fully connected layer with S input neurons and L output neurons; The second analog storage and calculation macro unit has the computing function of realizing a fully connected layer with L input neurons and P output neurons; The second hybrid storage and computing macro unit has the computing function of realizing a fully connected layer with P input neurons and Q output neurons.
11. The memristor neural network chip according to claim 1, characterized in that: The at least one analog storage and computing macro unit includes: an analog storage and computing macro unit for input and an analog storage and computing macro unit for circulation; the at least one hybrid storage and computing macro unit includes: a hybrid storage and computing unit for output; and the memristor neural network chip further includes: A functional unit connected to the analog storage and computing macro unit for input, the analog storage and computing macro unit for circulation, and the hybrid storage and computing macro unit for output respectively; The functional unit is used to obtain the total analog current output by the analog storage and computing macro unit for output and the analog storage and computing macro unit for circulation, and use the total analog current to generate an output voltage, and output the output voltage to the hybrid storage and computing macro unit for output.
12. The memristor neural network chip according to claim 11, characterized in that: The functional unit includes: two analog direct transmission modules and an activation analog circuit; The activation analog circuit is connected between the two analog direct transmission modules; Of the two analog direct transmission modules, one is an input direct transmission module and the other is an output direct transmission module; The input direct transmission module is connected to the input analog storage and calculation macro unit and the cyclic analog storage and calculation macro unit, and is used to obtain the total analog current, convert the total analog current into an analog voltage within a preset range, and output the analog voltage; The activation analog circuit is used to activate the analog voltage output by the input direct transmission module and generate an activation current output; The output direct transmission module is used to convert the activation current into an analog voltage within a preset range as the output voltage.
13. The memristor neural network chip according to claim 1, characterized in that: Also includes: A controller is connected to the at least one analog storage and computing macro unit and the at least one hybrid storage and computing macro unit, and is used to control the at least one analog storage and computing macro unit and the at least one hybrid storage and computing macro unit.
14. A storage-computation integrated operation method, applied to the memristor neural network chip according to any one of claims 1 to 13, characterized in that: The method comprises: Using at least one analog storage and computing macrocell, an analog voltage is applied to the memristor array in the cell, and the generated analog current is converted into an analog voltage within a preset range and then output; Using at least one hybrid storage and computing macrocell, the analog voltage output by the at least one analog storage and computing macrocell is applied to the memristor array in the cell, and the generated analog current is clamped, subtracted, and analog-to-digital converted before output; The at least one analog storage and computing macro unit includes: an analog direct transmission module connected to the memristor array, and the method further includes: The analog direct transmission module is used to integrate the analog current generated by the memristor array to obtain an integrated voltage, and the integrated voltage is compared with a threshold voltage, and the maximum voltage between the integrated voltage and the threshold voltage is selected for output.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the storage-computation-integrated operation method as claimed in claim 14 is implemented.
Citation Information
Patent Citations
Equation set solver based on memristor linear neural network and operation method of equation set solver
CN111460365A
Equation solver based on memristor arrays, and operation method thereof
CN111507464A