A Convolution Acceleration Operation Array Based on Resistive Memory and Its Control Method

By optimizing the word line and bit line connection method of resistive memory array, the problem of low computing efficiency caused by device redundancy in traditional Crossbar arrays is solved, and efficient convolutional operations are achieved.

CN115910136BActive Publication Date: 2025-07-25HUAZHONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211398594.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-09
Publication Date
2025-07-25
Estimated Expiration
2042-11-09

AI Technical Summary

Technical Problem

The traditional resistor-based Crossbar array has a redundancy in the convolution operation due to the redundancy of the same input of a large number of units, resulting in low computing efficiency and cannot achieve full parallel computing.

Method used

A convolutional acceleration computing array based on resistive memory is designed. By optimizing the connection method between word lines and bit lines, the resistive memory cells on each column are connected to the same bit line, and the same resistive memory cells are inputted on each row of the array are connected to the same word line, ensuring that the weight value corresponds one by one to the corresponding relationship of the overlay element, and implementing parallel point multiplication operations.

Benefits of technology

It improves the efficiency of convolutional operations, solves the problem of device redundancy in traditional Crossbar arrays, and improves the computing power and computing efficiency of the array.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115910136B_ABST
    Figure CN115910136B_ABST
Patent Text Reader

Abstract

The present invention discloses a convolutional acceleration operation array based on a resistive memory and its control method, belonging to the technical field of memories. The number of rows is the size of the convolutional kernel, the number of columns is the total number of sliding times S required when sliding the convolutional kernel on the data D to be convolved to traverse the data D to be convolved, the total number of word lines is the total number of elements in the data D to be convolved, and the total number of bit lines is S; the resistive memory cells on each column are all connected to the same bit line; by storing all the weight values in the convolutional kernel on each column and calculating according to the principle that the dot product operation result under one sliding is output once for each bit line and the bit lines of the array can synchronously output all the dot product operation results, the inputs of the resistive memory cells in the array are determined, and the resistive memory cells with the same input on each row are connected to the same word line to determine the bit line distribution; under this bit line distribution, the convolutional acceleration array has more input ports and higher operation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of memories, and more specifically, relates to a convolutional acceleration operation array based on a resistive memory and a control method thereof. Background Art

[0002] With the development of brain science and cognitive science in recent years, in order to be closer to the human brain, biologically inspired computing methods, namely artificial neural networks, have become a research hotspot in the field of artificial intelligence. However, during the development process of the entire research on artificial neural networks, the computing power of hardware devices has always been one of the problems restricting its development. For this reason, with the significant improvement in the parallel computing power of graphics processing units (GPUs) in recent years, the training bottleneck of deep learning algorithms has been further reduced, and the research on artificial intelligence has been booming. However, based on von Neumann's central processing unit (CPU) and GPU, they not only consume kilowatt-level energy in computing, but also the computing efficiency is reduced due to the transmission of data between the computing unit and the storage unit.

[0003] To solve the above problems, researchers have proposed many kinds of neuromorphic devices, including CMOS and non-von Neumann phase change memory (PCM), resistive random access memory (RRAM), conductive bridging random access memory (CBRAM), etc. Among these technologies, memory devices based on the resistance principle have become typical devices for accelerating matrix multiplication operations because they can represent weight values by using the stored conductance values, and the current output on the bit line can represent the sum of the products of the conductance of multiple units and their respective input voltages. Since matrix multiplication operations are the basis of convolution operations in convolutional neural networks, arrays based on resistive memories can be used to accelerate convolution operations.

[0004] In the case of a traditional array connected in a Crossbar manner, although it is possible to select cells in the array through word lines and bit lines, and the arrangement of word lines and bit lines is the simplest, with the fewest input numbers of word lines and bit lines required, it is the best wiring method when the array is used as a memory. However, when performing convolution acceleration operations, since a large number of devices are connected to a single word line, all devices under this word line share the same input. Although it can accelerate the convolution operations of convolution kernels with different channels, in a neural network, the number of channels of convolution kernels is usually limited. Except for a few devices representing the weights of convolution kernels with different channels, a large number of devices will generate redundancy because they cannot utilize this input. Therefore, when a large Crossbar array performs convolution operations, the operation efficiency of the array will be greatly reduced due to the large amount of redundancy generated by the same input of a large number of cells, and parallel operations cannot be fully realized, so the advantages of resistive memory convolution acceleration cannot be exerted in the array structure designed for memory. Summary of the Invention

[0005] In view of the above defects or improvement requirements of the prior art, the present invention provides a convolution acceleration operation array based on resistive memory and its control method, aiming to solve the technical problem of low operation efficiency caused by a large amount of redundancy generated by the same input of a large number of cells in the existing resistive Crossbar array for memory.

[0006] To achieve the above object, in the first aspect, the present invention provides a convolution acceleration operation array based on resistive memory, which is used to perform convolution operations on the to-be-convolved data D with the same dimension and the convolution kernel; wherein, the size of the convolution kernel is K×K×n; the total number of sliding times S required to slide the convolution kernel over the to-be-convolved data D to traverse the to-be-convolved data D; in each sliding, the convolution kernel covers a group of elements in the to-be-convolved data D, and there is a one-to-one correspondence between the weight values in the convolution kernel and the elements it covers;

[0007] The above convolution acceleration operation array includes: resistive memory cells distributed in an array, with the number of rows being K×K×n and the number of columns being S;

[0008] The resistive memory cells in each column are all connected to the same bit line and are used to store all the weight values in the convolution kernel;

[0009] The K×K×n groups of row inputs of the array are obtained by searching from the to-be-convolved data D; wherein, the acquisition method of each group of row inputs is: starting from any un-searched element d in the to-be-convolved data D, searching for the element separated from it by m(K - 1) elements, and mapping all the searched elements and the element d to corresponding voltage values to form a group of row inputs; the number of elements in each row input is the total number of times it is covered by the convolution kernel during the process of traversing the to-be-convolved data D; wherein, m is any positive integer;

[0010] The input of each resistive memory cell on each column is obtained from the row input of the row where it is located, and it is ensured that the correspondence between the weight values stored in the resistive memory cells on column S and their inputs is in one-to-one correspondence with the correspondence between the convolutional kernel under S sliding operations and the elements it covers;

[0011] Word lines are provided on each row of the array; the total number of word lines in the array is equal to the total number of elements in the data D to be convolved; the resistive memory cells with the same input on each row are connected to the same word line.

[0012] Further preferably, the above convolutional operation is a one-dimensional convolutional operation, a two-dimensional convolutional operation, or a three-dimensional convolutional operation.

[0013] Further preferably, when the convolutional operation is a one-dimensional convolutional operation, the data D to be convolved is a one-dimensional matrix; the acquisition method of each group of row inputs is: starting from any un-searched element d in the data D to be convolved, searching for elements that are m(K - 1) elements apart from it along the row direction where it is located, and after mapping all the searched elements and the element d to corresponding voltage values, forming a group of row inputs.

[0014] Further preferably, when the convolutional operation is a two-dimensional convolutional operation, the data D to be convolved is a two-dimensional matrix; the acquisition method of each group of row inputs is: starting from any un-searched element d in the data D to be convolved, searching for elements that are m(K - 1) elements apart from it along the row direction, column direction, and diagonal direction where it is located, and after mapping all the searched elements and the element d to corresponding voltage values, forming a group of row inputs.

[0015] Further preferably, when the convolutional operation is a three-dimensional convolutional operation, the data D to be convolved is a three-dimensional matrix; the acquisition method of each group of row inputs is: starting from any un-searched element d in the data D to be convolved, searching for elements that are m(K - 1) elements apart from it along the row direction, column direction, channel direction, and diagonal direction where it is located, and after mapping all the searched elements and the element d to corresponding voltage values, forming a group of row inputs.

[0016] Further preferably, the above resistive memory is a device that stores information based on the resistance of the device, including: a phase change memory that stores information based on the change in device resistance caused by the transformation between the crystalline and amorphous states of the material, or a memristor that stores information based on the generation and breakage of conductive filaments in the material.

[0017] In a second aspect, the present invention provides a control method for the above convolutional acceleration operation array, including:

[0018] Storing all the weight values in the convolutional kernel on each column respectively;

[0019] On each line, the elements in the corresponding data D to be convolved are input into each resistive memory cell in the form of voltage through the word lines, and the dot product operation results of the convolution kernel and a group of elements in the data D to be convolved covered by it under S sliding operations are obtained in parallel, and are output through the corresponding word lines, so as to obtain the convolution operation result of the data D to be convolved and the convolution kernel;

[0020] Among them, the method for determining the input of the resistive memory cell includes:

[0021] The K×K×n groups of row inputs of the array are searched from the data D to be convolved; among them, the acquisition method of each group of row inputs is: starting from any un-searched element d in the data D to be convolved, searching for the element separated from it by m(K - 1) elements, and mapping all the searched elements and the element d into corresponding voltage values to form a group of row inputs; the number of each element in the row input is the total number of times it is covered by the convolution kernel during the process of traversing the data D to be convolved; m is any positive integer;

[0022] The input of each resistive memory cell in each column is obtained from the row input of its corresponding row, and it is ensured that the corresponding relationship between the weight values stored in the resistive memory cells in S columns and their inputs is in one-to-one correspondence with the corresponding relationship between the convolution kernel and the elements it covers under S sliding operations.

[0023] Generally speaking, through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:

[0024] 1. The present invention provides a convolution acceleration operation array based on resistive memory, the number of rows of which is the size of the convolution kernel, the number of columns is the total number of sliding operations S required when the convolution kernel slides on the data D to be convolved to traverse the data D to be convolved, the total number of word lines is the total number of elements in the data D to be convolved, and the total number of bit lines is S; the resistive memory cells in each column are all connected to the same bit line; by storing all the weight values in the convolution kernel in each column, and calculating according to the principle that the dot product operation result of each sliding operation is output through each bit line and the bit lines of the array can synchronously output all the dot product operation results, the input of each resistive memory cell in the array is determined, and the resistive memory cells with the same input in each row are connected to the same word line, so as to determine the distribution of the word lines; under this word line distribution, the convolution acceleration array has more input ports for inputting the data D to be convolved, and can obtain the dot product operation results of the convolution kernel and a group of elements in the data D to be convolved covered by it under S sliding operations in parallel, and obtain the convolution operation result of the data D to be convolved and the convolution kernel at one time, thus effectively solving the problem of device redundancy in the traditional Crossbar array when accelerating the convolution operation, greatly improving the convolution operation computing power of the array, and having a higher operation efficiency.

[0025] 2. The convolution acceleration operation array based on the resistive memory provided by the present invention controls the number of units connected to the word line through the word-bit line connection design for the convolution operation, and solves the problem that the number of units on the word line in a large-scale Crossbar array is too large, resulting in too long word lines, so that the word line electrodes have a high resistance and share a large amount of voltage in the array, affecting the operation of the units. Description of the Drawings

[0026] Figure 1 It is a schematic diagram of the data to be convolved and the convolution operation in the two-dimensional convolution operation in Embodiment 1;

[0027] Figure 2 It is a schematic diagram of the convolution acceleration operation array in Embodiment 1;

[0028] Figure 3 It is a schematic diagram of the number of resistive memory units connected to each bit line in Embodiment 1;

[0029] Figure 4 It is a schematic diagram of the number of resistive memory units connected to the word line in Embodiment 1;

[0030] Figure 5 It is a schematic diagram of the method for determining the word line distribution of the convolution operation in Embodiment 1;

[0031] Figure 6 It is a schematic diagram of the wiring distribution of the resistive memory array in Embodiment 1;

[0032] Figure 7 It is a schematic diagram of the word line and bit line mask plates of the resistive memory in a process in Embodiment 1; among them, (a) is the word line mask plate in Embodiment 1; (b) is the bit line mask plate in Embodiment 1;

[0033] Figure 8 It is a schematic diagram of the traditional cross-bar structure array in Embodiment 1;

[0034] Figure 9 It is a schematic diagram of the data to be convolved and the convolution operation in the three-dimensional convolution operation in Embodiment 2;

[0035] Figure 10 It is a schematic diagram of the convolution acceleration operation array in Embodiment 2;

[0036] Figure 11 It is a schematic diagram of the number of resistive memory units connected to each bit line in Embodiment 2;

[0037] Figure 12 It is a schematic diagram of the number of resistive memory units connected to the word line in Embodiment 2;

[0038] Figure 13It is a schematic diagram of the method for determining the word line distribution in the convolution operation in Embodiment 2;

[0039] Figure 14 It is a schematic diagram of the wiring distribution of the resistive memory array in Embodiment 2;

[0040] Figure 15 It is a schematic diagram of the word line and bit line mask plates of the resistive memory in a certain process in Embodiment 2; among them, (a) is the word line mask plate in Embodiment 2; (b) is the bit line mask plate in Embodiment 2. Detailed implementation manners

[0041] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0042] In order to solve the problem that a large amount of redundancy is generated during convolution acceleration in a resistive Crossbar array for memory, and full parallelism cannot be achieved, the present invention provides a convolution acceleration operation array based on a resistive memory, which is used to perform convolution operations on the to-be-convolved data D with the same dimension and the convolution kernel (with a size of K×K×n); among them, the total number of resistive memory cells, the number of bit lines and word lines and the number of resistive memory cells connected thereto, and the distribution of the resistive memory cells in the array are determined according to the following method according to the situation of a specific convolution operation. The resistive memory can be a phase change memory that stores information based on the resistance change of a device caused by the transformation between the crystalline state and the amorphous state of a material, a memristor that stores information based on the generation and breakage of a conductive filament in a material, and a series of devices that store information based on the resistance of a device.

[0043] Specifically, based on the definition of the convolution operation, the process of performing the convolution operation is to slide the convolution kernel on the to-be-convolved data D and traverse the entire to-be-convolved data D. At this time, the total number of sliding times required is S. Among them, each time it slides, the convolution kernel will cover a group of elements in the to-be-convolved data D (there is a one-to-one correspondence between the weight values in the convolution kernel and the elements it covers), and perform a dot product operation between the convolution kernel and this group of elements. The dot product operations under each sliding together constitute the convolution operation result of the to-be-convolved data D and the convolution kernel.

[0044] In order to implement the dot product operation under each sliding in parallel, the present invention designs the convolution acceleration operation array in the following manner. Specifically, the convolution acceleration operation array includes: resistive memory cells distributed in an array, with the number of rows being K×K×n and the number of columns being S, so that each bit line represents the output of the dot product operation result under one sliding.

[0045] Specifically, the resistive memory cells on each column are all connected to the same bit line and are used to store all the weight values in the convolution kernel;

[0046] The K×K×n groups of row inputs of the array are obtained by searching from the data D to be convolved; among them, the obtaining method of each group of row inputs is: starting from any un-searched element d in the data D to be convolved, search for the element separated from it by m(K - 1) elements, and after mapping all the searched elements and the element d to corresponding voltage values, form a group of row inputs; the number of each element in the row input is the total number of times it is covered by the convolution kernel during the process of traversing the data D to be convolved; where m is any positive integer;

[0047] The inputs of the resistive memory cells on each column are obtained from the row inputs of their corresponding rows, and it is ensured that the corresponding relationship between the weight values stored in the resistive memory cells on the S column and their inputs is in one-to-one correspondence with the corresponding relationship between the convolution kernel and the elements it covers under S sliding operations;

[0048] Word lines are provided on each row of the array; the total number of word lines in the array is equal to the total number of elements in the data D to be convolved; the resistive memory cells with the same input on each row are connected to the same word line.

[0049] It should be noted that the above convolution operation is a one-dimensional convolution operation, a two-dimensional convolution operation, or a three-dimensional convolution operation.

[0050] When the convolution operation is a one-dimensional convolution operation, the data D to be convolved is a one-dimensional matrix; the obtaining method of each group of row inputs is: starting from any un-searched element d in the data D to be convolved, search for the element separated from it by m(K - 1) elements along the direction of its row, and after mapping all the searched elements and the element d to corresponding voltage values, form a group of row inputs.

[0051] When the convolution operation is a two-dimensional convolution operation, the data D to be convolved is a two-dimensional matrix; the obtaining method of each group of row inputs is: starting from any un-searched element d in the data D to be convolved, search for the element separated from it by m(K - 1) elements along the direction of its row, column, and diagonal respectively, and after mapping all the searched elements and the element d to corresponding voltage values, form a group of row inputs.

[0052] When the convolution operation is a three-dimensional convolution operation, the data D to be convolved is a three-dimensional matrix; the obtaining method of each group of row inputs is: starting from any un-searched element d in the data D to be convolved, search for the element separated from it by m(K - 1) elements along the direction of its row, column, channel, and diagonal respectively, and after mapping all the searched elements and the element d to corresponding voltage values, form a group of row inputs.

[0053] Correspondingly, the control method of the above convolution acceleration operation array includes:

[0054] Store all the weight values in the convolution kernel on each column respectively;

[0055] On each row respectively, input the elements in the corresponding data D to be convolved into each resistive memory cell in the form of voltage through word lines, and obtain the dot product operation results of a set of elements in the convolution kernel and the data D to be convolved covered by it under S sliding operations in parallel, and output them through the corresponding word lines, so as to obtain the convolution operation result of the data D to be convolved and the convolution kernel;

[0056] Among them, the method for determining the input of the resistive memory cell includes:

[0057] The K×K×n groups of row inputs of the array are obtained by searching from the data D to be convolved; among them, the acquisition method of each group of row inputs is: starting from any un-searched element d in the data D to be convolved, search for the element separated from it by m(K - 1) elements, and map all the searched elements and the element d to the corresponding voltage values to form a group of row inputs; the number of elements in each row input is the total number of times it is covered by the convolution kernel during the process of traversing the data D to be convolved; m is any positive integer;

[0058] The input of each resistive memory cell on each column is obtained from the row input of its corresponding row, and it is ensured that the corresponding relationship between the weight values stored in the resistive memory cells on S columns and their inputs is in one-to-one correspondence with the corresponding relationship between the convolution kernel and the covered elements under S sliding operations.

[0059] Correspondingly, the preparation method of the above convolution acceleration operation array includes:

[0060] When depositing the array structure on the substrate, the patterns of each layer used for patterning are determined according to the number and distribution modes of the above resistive memory cells, word lines, and bit lines, and then combined with the deposition materials and etching processes to finally obtain the above convolution acceleration operation array.

[0061] It should be noted that the structures and materials of the devices in the array can be arbitrary.

[0062] In summary, the convolution acceleration operation array based on resistive memory in the present invention optimizes the distribution of word lines, bit lines, and resistive memory cells by combining the convolution operation and the situation of the data D to be convolved, and can effectively solve the device redundancy problem of the traditional Crossbar array in accelerating the convolution operation and improve the convolution operation computing power of the array.

[0063] In order to further illustrate the convolution acceleration operation array provided by the present invention, it will be described in detail below with reference to embodiments:

[0064] Example 1

[0065] In this embodiment, a convolution acceleration operation array for two-dimensional convolution operations is provided. The situation of the data to be convolved and the convolution operation is as follows Figure 1 shown.

[0066] The two-dimensional convolution operation is as follows Figure 1 shown, where W represents the weight value of the convolution kernel of the convolution operation, and D represents the data to be convolved. The size of the convolution kernel and the size of the data to be convolved in the figure are as shown, which are 3×3 and 5×5 respectively. The sliding step sizes in the horizontal and vertical directions of the convolution are both 1. It can be seen that the number of sliding times in the horizontal and vertical directions is 3, so the total number of sliding times is 9 times. This data will perform 9 dot product operations in total.

[0067] The method for determining the total number of devices in the resistive memory cell array for accelerating this convolution operation in the present invention is as follows Figure 2 shown. Figure 2 The W in it corresponds to the weight value of the convolution kernel in the above Figure 1 , and S represents the number of sliding times of the convolution. S1 - S9 represent the above 9 convolution operations. As shown in Figure 2 , the total number of devices required for the array is the product of the number of weight values in the convolution kernel and the total number of sliding times required for the convolution.

[0068] The method for determining the number of bit lines and the number of connected devices in the resistive memory cell array corresponding to this convolution operation is as follows Figure 3 shown. Figure 3 The W in it corresponds to the weight value in the above convolution kernel. As shown in Figure 3 , the number of devices connected to a single bit line in the array is the same as the number of weight values in the convolution kernel. According to the total number of devices determined in Figure 2 , the number of bit lines is the same as the number of sliding times of the convolution.

[0069] The method for determining the number of word lines and the number of connected devices in the above resistive memory cell array is as follows Figure 4 shown. Figure 4 The data D to be convolved in it corresponds to the input value of the above convolution operation. WL represents the word line. The number of squares on each WL in the figure represents the total number of times C that this input participates in the operation during the entire convolution operation. As shown in Figure 4 , each element in the data to be convolved corresponds to a word line, and the total number of word lines is the same as the total number of elements in the data D to be convolved. According to the total number of times C that the data input for the convolution operation participates in the operation in the 9 convolution operations corresponding to S1 - S9, the number of devices connected to each word line is determined.

[0070] The method for determining the distribution of the resistive memory array is as follows Figure 5As shown, among them, the value of the element in the rightmost matrix is the total number of times the corresponding convolution element is covered by the convolution kernel during the traversal of the data D to be convolved. After determining the total number of devices, word lines, bit lines, and the number of connected devices according to the above method, calculate according to the principle that each bit line outputs one convolution operation step by step, and the bit lines of the array can synchronously output all 9 convolution operations. The connection method of the units on each word line is determined by Figure 5 the method shown. Starting from one position of the input image, take the next value after skipping the image points with a distance equal to the side length of the convolution kernel minus one. The value-taking direction is the horizontal, vertical, and the cross-shaped direction with a 45° interval from the horizontal and vertical directions centered on the input value. On the premise of no repetition, each value-taking combination represents an arrangement method on the word line. On this basis, further determine the connection method of the word line according to the position of the bit line where the unit is located and the total number of times C that the input image value participates in the convolution operation. Among them, the bit line position represents the sliding position of the convolution kernel, that is, the action range of a specific convolution, and C represents the number of repetitions of this image value in the word line, that is, the number of connected units. The final obtained word line arrangement and input method are as shown in Figure 6 .

[0071] The word line and bit line mask plates used in the preparation process of this memory array are as shown in Figure 7 . After determining the wiring situation of the memory array according to the above method, for the device array using the via hole structure, it can be prepared by the steps of lithographing the bit line electrode pattern, depositing the bit line electrode material, depositing the insulating layer, lithographing and etching the hole array, lithographing the word line electrode pattern, depositing the functional layer and the electrode material. In all lithography processes of this preparation process, except for the lithography and etching of the holes, the mask plates used in the lithography processes of the word line and bit line all meet the requirements of the above acceleration design. Using the mask plate shown in Figure 7 can prepare an array that conforms to the wiring method shown in Figure 6 , and can realize the convolution acceleration operation for the 5×5 input and 3×3 convolution kernel described in Figure 6 ; among them, Figure 7 Figure (a) in is the word line mask plate, and figure (b) is the bit line mask plate.

[0072] It should be noted that the schematic diagram of the traditional cross-bar array is as shown in Figure 8 . The word line and bit line are respectively connected to all the units in the same row and the same column. Compared with the array designed in the present invention, in the case of the same number of devices, it has only 9 word lines for input; while the array designed in the present invention has 25 word lines for input, and the array designed in the present invention comprehensively considers the number of operations performed by the data during the convolution process and optimizes the convolution operation. The performance comparison table of the above traditional cross-bar array and the convolution acceleration array designed in the present invention is shown in Table 1.

[0073] Table 1

[0074] Array type Number of units Input port Output port Convolution computing power per clock cycle Traditional cross-bar 81 9 9 9×1MACC Convolution acceleration array 81 25 9 9×9MACC

[0075] As can be seen from Table 1, in the case of the same number of devices, since the traditional cross-bar array has only 9 input ports and all the cells in each row have the same input, it cannot meet the requirement of the convolution process for the sliding of the data to be convolved. Therefore, in one clock cycle, only the convolution operation of 9 input data can be completed, that is, the effective operations of 9 multiplications and 8 additions. So its computing power is 9 MACC. And because the convolution acceleration array provided by the present invention has more input ports, when performing the convolution operation on the data, it can match the sliding of the data to be convolved during the convolution operation. Therefore, it can complete all the convolution values of 9 slides of the 3×3 convolution kernel on the 5×5 input. So its computing power is 9×9 MACC. It can be seen that under the same number of devices, for the above convolution operation, the array designed by the present invention can achieve 9 times the computing power compared with the traditional cross-bar structure.

[0076] Example 2

[0077] In this embodiment, it is a convolution acceleration operation array for three-dimensional convolution operations. The situation of the data to be convolved and the convolution operation is as Figure 9 shown.

[0078] The three-dimensional convolution operation is as Figure 9 shown, where W represents the weight value of the convolution kernel of the convolution operation, and D represents the data to be convolved. The sizes of the convolution kernel and the data to be convolved in the figure are as shown, which are 2×2×2 and 3×3×3 respectively. Among them, the sliding step sizes in the horizontal, vertical, and channel directions of the convolution are all 1. It can be known that the sliding times in the horizontal, vertical, and channel directions are all 2. Therefore, the total sliding times are 8 times, and this data will perform 8 convolution operations in total.

[0079] The method for determining the total number of devices of the resistive memory cell array for accelerating this convolution operation in the present invention is as Figure 10 shown. Figure 9 The W in Figure 9 corresponds to the weight value of the convolution kernel in the above Figure 10 shown. The total number of devices required for the array is the product of the number of weight values in the convolution kernel and the total number of sliding times required for the convolution.

[0080] The method for determining the number of bit lines and the number of connected devices of the resistive memory cell array corresponding to this convolution operation is as Figure 11 shown. Figure 11 The W in Figure 11As shown, the number of devices connected to a single bit line in the array is the same as the number of weight values in the convolution kernel. According to Figure 10 the total number of devices determined therein, the number of bit lines is the same as the number of sliding times of the convolution.

[0081] The method for determining the number of word lines and the number of devices connected thereto in the above resistive memory cell array is as Figure 12 shown. Figure 12 In Figure 12 , D corresponds to the input value of the above convolution operation, WL represents the word line, and the number of squares on each WL in the figure represents the total number of times C that the input participates in the operation during the entire convolution operation. As Figure 12 shown, each element in the data to be convolved corresponds to a word line, and the total number of word lines is the same as the total number of elements in the data to be convolved. According to the total number of times C that the data input for the convolution operation participates in the operation in the 8 convolution operations corresponding to S1 - S8, the number of devices connected to each word line is determined.

[0082] Furthermore, the method for determining the word line distribution of the resistive memory array is as Figure 13 shown. After determining the total number of devices, word lines and bit lines and the number of devices connected thereto according to the above method, taking one step of convolution operation for each bit line output, the principle that the bit lines of the array can synchronously output all 8 convolution operations is used to obtain the array layout shown in the figure. The connection method of the units on each word line is determined by the Figure 13 shown method. Starting from one position of the input image, taking the next value after an image point with an interval of one less than the side length of the convolution kernel, the value-taking direction is the horizontal, vertical, and the direction perpendicular to the paper surface centered on the input value, as well as the three-dimensional meter-shaped with a 45° interval between them. Without repetition, each value-taking combination represents an arrangement method on the word line. On this basis, the connection method of the word line is further determined according to the position of the bit line where the unit is located and the total number of times C that the input image value participates in the convolution operation, where the bit line position represents the sliding position of the convolution kernel, that is, the action range of a specific convolution, and C represents the number of repetitions of the image value in the word line, that is, the number of connected units. The final word line arrangement and input method are as Figure 14 shown.

[0083] The word line and bit line mask plates used in the preparation process of this memory array are as Figure 15 shown. After determining the wiring situation of the memory array according to the above method, for the device array using the via structure, it can be prepared by steps such as lithographing the bit line electrode pattern, depositing the bit line electrode material, depositing the insulating layer, lithographing and etching the hole array, lithographing the word line electrode pattern, depositing the functional layer and the electrode material. In all lithography processes of this preparation process, except for the lithography and etching of the holes, the mask plates used in the lithography processes of the word lines and bit lines all meet the requirements of the above convolution acceleration operation array. Using Figure 15The mask shown can be used to fabricate an array that conforms to the Figure 14 wiring method shown, and can implement the Figure 9 convolution acceleration operation for a 3×3×3 input and a 2×2×2 convolution kernel; where Figure 15 in (a) of is the word line mask, and (b) is the bit line mask.

[0084] It should be noted that the performance comparison table between the traditional cross-bar array and the convolution acceleration array in this embodiment is shown in Table 2.

[0085] Table 2

[0086] Array type Number of units Input port Output port Convolution computing power per clock cycle Traditional cross-bar 64 8 8 8×1MACC Convolution acceleration array 64 27 8 8×8MACC

[0087] As can be seen from Table 2, in the case of the same number of devices, since the traditional cross-bar array has only 8 input ports and all units in each row have the same input, it cannot meet the requirements of the convolution process for the sliding of the data to be convolved. Therefore, only the convolution operation of 8 input data can be completed within one clock cycle, that is, 8 multiplications and 7 additions of effective operations, so its computing power is 8 MACC. And because the convolution acceleration array provided by the present invention has more input ports, it can match the sliding of the data to be convolved during the convolution operation, so it can complete all the convolution values of 8 slides of the 2×2×2 convolution kernel on the 3×3×3 input. Therefore, its computing power is 8×8 MACC. It can be seen that under the same number of devices, for the above convolution operation, the array designed by the present invention can achieve 8 times the computing power compared to the traditional cross-bar structure.

[0088] Those skilled in the art can easily understand that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A convolutional acceleration operation array based on a resistive memory, characterized in that, For performing a convolution operation on the data D to be convolved with the same dimension and the convolution kernel; the size of the convolution kernel is K×K×n; the total number of sliding times required to slide the convolution kernel on the data D to be convolved to traverse the data D to be convolved is S; Under each sliding, the convolution kernel covers a group of elements in the data D to be convolved; The convolution acceleration operation array includes: resistive memory cells distributed in an array, with the number of rows being K×K×n and the number of columns being S; The resistive memory cells on each column are all connected to the same bit line and are used to store all the weight values in the convolution kernel; The total number of word lines in the array is equal to the total number of elements in the data D to be convolved; 1 or more word lines are arranged on each row of the array; the input resistive memory cells with the same value on each row are connected to the same word line; The K×K×n groups of row inputs of the array are obtained by searching in the data D to be convolved; among them, the acquisition method of each group of row inputs is: starting from any un-searched element d in the data D to be convolved, searching for the element with an interval of m(K - 1) elements from it, and after mapping all the searched elements and the element d to corresponding voltage values, forming a group of row inputs; the number of each element in the row input is the total number of times it is covered by the convolution kernel during the process of traversing the data D to be convolved; m is any positive integer; The input of each resistive memory cell on each column is obtained from the row input of its corresponding row, and it is ensured that the corresponding relationship between the weight values stored in the resistive memory cells on the S columns and their inputs is in one-to-one correspondence with the corresponding relationship between the convolution kernel and the elements it covers under S times of sliding.

2. The convolution acceleration operation array according to claim 1, wherein The convolution operation is a one-dimensional convolution operation, a two-dimensional convolution operation, or a three-dimensional convolution operation.

3. The convolution acceleration operation array according to claim 2, wherein When the convolution operation is a one-dimensional convolution operation, the data D to be convolved is a one-dimensional matrix; the acquisition method of each group of row inputs is: starting from any un-searched element d in the data D to be convolved, searching for the element with an interval of m(K - 1) elements along the row direction where it is located, and after mapping all the searched elements and the element d to corresponding voltage values, forming a group of row inputs.

4. The convolution acceleration operation array according to claim 2, wherein When the convolution operation is a two-dimensional convolution operation, the data D to be convolved is a two-dimensional matrix; the acquisition method of each group of row inputs is: starting from any un-searched element d in the data D to be convolved, searching for the element with an interval of m(K - 1) elements along the row direction, column direction, and diagonal direction where it is located, and after mapping all the searched elements and the element d to corresponding voltage values, forming a group of row inputs.

5. The convolution acceleration operation array according to claim 2, wherein When the convolution operation is a three-dimensional convolution operation, the data D to be convolved is a three-dimensional matrix; the acquisition method of each group of row inputs is: starting from any un-searched element d in the data D to be convolved, searching for the element with an interval of m(K - 1) elements along the row direction, column direction, channel direction, and diagonal direction where it is located, and after mapping all the searched elements and the element d to corresponding voltage values, forming a group of row inputs.

6. The convolution acceleration operation array according to any one of claims 1-5, characterized in that The resistive memory is a device that stores information based on the resistance of the device, including: a phase change memory that stores information based on the change in the resistance of the device caused by the transformation between the crystalline and amorphous states of the material, or a memristor that stores information based on the generation and breakage of conductive filaments in the material, resulting in a change in the resistance of the device.

7. The control method of the convolutional acceleration operation array according to any one of claims 1-6, characterized in that including: Storing all the weight values in the convolutional kernel on each column respectively; On each row respectively, the elements in the corresponding data D to be convolved are input into each resistive memory cell in the form of voltage through the word lines, and the dot product operation results of the convolutional kernel and a set of elements in the data D to be convolved covered by it under S sliding operations are obtained in parallel, and are output via the corresponding word lines, so as to obtain the convolution operation result of the data D to be convolved and the convolutional kernel; Among them, the method for determining the input of the resistive memory cell includes: The K×K×n groups of row inputs of the array are obtained by searching from the data D to be convolved; among them, the acquisition method of each group of row inputs is: starting from any un-searched element d in the data D to be convolved, searching for the element that is m(K - 1) elements away from it, and mapping all the searched elements and the element d to the corresponding voltage values to form a group of row inputs; the number of elements in each row input is the total number of times it is covered by the convolutional kernel during the process of traversing the data D to be convolved; m is any positive integer; The input of each resistive memory cell on each column is obtained from the row input of its corresponding row, and it is ensured that the corresponding relationship between the weight values stored in the resistive memory cells on S columns and their inputs is in one-to-one correspondence with the corresponding relationship between the convolutional kernel and the elements it covers under S sliding operations.

Citation Information

Patent Citations

  • Convolution calculation and storage integrated equipment and method based on resistive random access memory array

    CN106847335A

  • Reconstructing MAC operations

    CN114365078A