Convolutional Neural Network Processing Method and Device Based on In-Memory Computing Array
By setting the three-dimensional dimensions of the sub-input feature map and the sub-convolution kernel in the memory integrated array, and deploying the sub-convolution kernel to multiple columns, the problems of storage walls and power consumption walls in the prior art are solved, and the computing efficiency and area utilization are improved.
Patent Information
- Application Number
- CN202111653640.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-30
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-12-30
AI Technical Summary
After the computing power demand of existing neural network processors have increased, they face problems such as "storage wall" and "power consumption wall". The integrated memory-based memory arrays have problems of resource waste and power consumption overhead when processing large-scale input feature maps.
By setting the three-dimensional size of the sub-input feature map and the sub-convolution kernel in the memory integrated array, and deploying the sub-convolution kernels to multiple columns of the memory integrated array in turn, the multi-row and multi-column operation unit of the memory integrated array is used for multiplication and accumulation operations, reducing data transfer operations and improving computing efficiency.
It improves the area utilization rate and computing efficiency of the integrated storage and computing array, reduces resource waste and power consumption overhead, and is suitable for processing large-scale neural network data.
Smart Images

Figure CN114298296B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to a convolutional neural network processing method and a convolutional neural network processing device based on a memory - in - computing array. Background Art
[0002] With the rapid development of information science, more and more artificial intelligence algorithms and technologies are applied in various fields of social life. Among them, the artificial neural network algorithm has been widely studied in recent years due to its diverse structures, adjustable network size and depth, learning and processing capabilities, and a certain tolerance for errors. It is the core of today's artificial intelligence technology. As a type of artificial neural network, the convolutional neural network is widely used in the field of computer vision, including many scenarios such as image classification, object detection, and semantic segmentation, due to its features such as weight sharing and multiple sampling.
[0003] Currently, neural networks are mainly deployed on central processing units and graphics processing units. However, due to the continuous increase in computing power requirements, the problems of the von Neumann architecture are prominent, facing problems such as the "memory wall" and the "power consumption wall". The memory - in - computing array based on memristors can perform multiply - accumulate operations in - situ using Ohm's law and Kirchhoff's circuit current law, eliminating the frequent data transfer operations between the storage unit and the arithmetic unit in the traditional von Neumann architecture. It is particularly suitable for completing a large number of vector - matrix multiplication operations in neural networks, greatly improving the energy efficiency of computing. Summary of the Invention
[0004] At least one embodiment of the present disclosure provides a convolutional neural network processing method based on a memory - in - computing array. The memory - in - computing array includes multiple rows and multiple columns of arithmetic units. The method includes: based on the width x_in, height y_in, and number of channels C_in of the input feature map, and the width x_k, height y_k, and number of channels C_in of the convolutional kernel, setting the width of the sub - input feature map of the input feature map to x, the height to y, and the number of channels to z, and setting the width of the sub - convolutional kernel of the convolutional kernel to x_k, the height to y_k, and the number of channels to z, where x_k is less than or equal to x, y_k is less than or equal to y, x is less than or equal to x_in, y is less than or equal to y_in, z is less than or equal to C_in, and the result of x×y×z is less than or equal to the number of rows of the memory - in - computing array; based on the situation where the sub - convolutional kernel slides in the sub - input feature map with a stride S, sequentially deploying the sub - convolutional kernel into N columns of at least one memory - in - computing array; inputting the sub - input feature map as a row input signal into at least one memory - in - computing array, and obtaining a calculation result from the column output signals of N columns of at least one memory - in - computing array.
[0005] For example, in the convolutional neural network processing method provided by at least one embodiment of the present disclosure, for the case where z is less than or equal to C_in, the number of at least one memory - in - computing array is
[0006] For example, in the convolutional neural network processing method provided by at least one embodiment of the present disclosure, based on the case where the sub-convolution kernel slides in the sub-input feature map with a stride S, deploying the sub-convolution kernel into N columns of at least one memory-computation integrated array in sequence includes: dividing the position of the sub-convolution kernel relative to the sub-input feature map into a first convolution position, a second convolution position, and a third convolution position according to the case where the sub-convolution kernel slides in the sub-input feature map with a stride S; and deploying the sub-convolution kernel into N columns of at least one memory-computation integrated array in sequence according to the first convolution position, the second convolution position, and the third convolution position, where the first convolution position corresponds to the case where the sub-convolution kernel is completely within the sub-input feature map during the sliding process, the second convolution position corresponds to the case where the sub-convolution kernel is between two sub-input feature maps during the sliding process, and the third convolution position corresponds to the case where the sub-convolution kernel is between four sub-input feature maps during the sliding process.
[0007] For example, in the convolutional neural network processing method provided by at least one embodiment of the present disclosure, deploying the sub-convolution kernel into N columns of at least one memory-computation integrated array in sequence according to the first convolution position, the second convolution position, and the third convolution position includes: deploying the sub-convolution kernel into n1 columns of at least one memory-computation integrated array according to the first convolution position, deploying the sub-convolution kernel into n2 columns of at least one memory-computation integrated array according to the second convolution position, and deploying the sub-convolution kernel into n3 columns of at least one memory-computation integrated array according to the third convolution position, where each column of the memory-computation integrated array deploys 1 sub-convolution kernel, and n1 + n2 + n3 = N.
[0008] For example, in the convolutional neural network processing method provided by at least one embodiment of the present disclosure, in the case where x_k and y_k of the sub-convolution kernel are the same, n3 = C_out × 4 × (x_k - 1) 2 ; or in the case where x_k and y_k of the sub-convolution kernel are different, n3 = Cout × 4 × (x_k - 1) × (y_k - 1).
[0009] For example, in the convolutional neural network processing method provided by at least one embodiment of the present disclosure, in the case where \(x_k\) and \(y_k\) of the sub-convolution kernel are the same, the number of weight elements corresponding to the sub-convolution kernel at the first convolution position is \(x_k\times y_k\times z\), the number of weight elements corresponding to the sub-convolution kernel at the second convolution position is \(x_k\times z\), and the number of weight elements corresponding to the sub-convolution kernel at the third convolution position is \(z\); or in the case where \(x_k\) and \(y_k\) of the sub-convolution kernel are different, the number of weight elements corresponding to the sub-convolution kernel at the first convolution position is \(x_k\times y_k\times z\), the number of weight elements corresponding to the sub-convolution kernel at the third convolution position is \(z\), and on half of the \(n2\) columns, the number of weight elements corresponding to the sub-convolution kernel at the second convolution position is \(x_k\times z\), and on the other half of the \(n2\) columns, the number of weight elements corresponding to the sub-convolution kernel at the second convolution position is \(y_k\times z\).
[0010] For example, in the convolutional neural network processing method provided by at least one embodiment of the present disclosure, the sub-input feature map is input as a row input signal into at least one memory-computation integrated array, and the calculation result is obtained from the column output signals of the N columns of at least one memory-computation integrated array, including: gradually sliding the input feature map to obtain the sub-input feature map in a sliding manner with a stride of \(x\) in the width dimension and a stride of \(y\) in the height dimension, and in each step, inputting the sub-input feature map as a row input signal into at least one memory-computation integrated array, and obtaining the calculation result corresponding to this step from the column output signals of the N columns of at least one memory-computation integrated array.
[0011] For example, in the convolutional neural network processing method provided by at least one embodiment of the present disclosure, the convolutional neural network involves \(K\) sub-convolution kernels, and \(K\) is equal to the number of output channels \(C_{out}\); based on the case where the sub-convolution kernels slide in the sub-input feature map with a stride of \(S\), deploying the \(K\) sub-convolution kernels into the N columns of at least one memory-computation integrated array in sequence, including: based on the case where the sub-convolution kernels slide in the sub-input feature map with a stride of \(S\), deploying each of the \(K\) sub-convolution kernels into the N columns of at least one memory-computation integrated array in sequence, thereby deploying the \(K\) sub-convolution kernels into the N columns of at least one memory-computation integrated array in sequence.
[0012] At least one embodiment of the present disclosure provides a convolutional neural network processing device based on a memory-computation integrated array, including: at least one memory-computation integrated array, including multiple rows and multiple columns of operation units; a setting unit configured to set the width of a sub-input feature map to x, the height to y, and the number of channels to z, and set the width of a sub-convolution kernel to x_k, the height to y_k, and the number of channels to z based on the width x_in, height y_in, and number of channels C_in of an input feature map and the width x_k, height y_k, and number of channels C_in of a convolution kernel, where x_k is less than or equal to x, y_k is less than or equal to y, x is less than or equal to x_in, y is less than or equal to y_in, z is less than or equal to C_in, and the result of x×y×z is less than or equal to the number of rows of the memory-computation integrated array; a deployment unit configured to deploy the sub-convolution kernel into N columns of at least one memory-computation integrated array in sequence based on the situation where the sub-convolution kernel slides in the sub-input feature map with a stride S; and a control unit configured to input the sub-input feature map into at least one memory-computation integrated array as a row input signal and obtain a calculation result from the column output signals of N columns of at least one memory-computation integrated array.
[0013] For example, in the convolutional neural network processing device provided by at least one embodiment of the present disclosure, the deployment unit is further configured to: divide the position of the sub-convolution kernel relative to the sub-input feature map into a first convolution position, a second convolution position, and a third convolution position according to the situation where the sub-convolution kernel slides in the sub-input feature map with a stride S; and deploy the sub-convolution kernel into N columns of at least one memory-computation integrated array in sequence according to the first convolution position, the second convolution position, and the third convolution position, where the first convolution position corresponds to the situation where the sub-convolution kernel is completely within the sub-input feature map during the sliding process, the second convolution position corresponds to the situation where the sub-convolution kernel is between two sub-input feature maps during the sliding process, and the third convolution position corresponds to the situation where the sub-convolution kernel is between four sub-input feature maps during the sliding process.
[0014] For example, in the convolutional neural network processing device provided by at least one embodiment of the present disclosure, the deployment unit is further configured to: deploy the sub-convolution kernel into n1 columns of at least one memory-computation integrated array according to the first convolution position, deploy the sub-convolution kernel into n2 columns of at least one memory-computation integrated array according to the second convolution position, and deploy the sub-convolution kernel into n3 columns of at least one memory-computation integrated array according to the third convolution position, where each column of the memory-computation integrated array deploys 1 sub-convolution kernel, and n1 + n2 + n3 = N.
[0015] For example, in the convolutional neural network processing apparatus provided in at least one embodiment of the present disclosure, the control unit is further configured to: gradually slide from the input feature map in a sliding manner with a sliding step of x in the width dimension and a sliding step of y in the height dimension to obtain a sub-input feature map, and in each step, input the sub-input feature map as a row input signal into at least one memory-computing integrated array, and obtain the calculation result corresponding to this step from the column output signals of N columns of at least one memory-computing integrated array.
[0016] For example, in the convolutional neural network processing apparatus provided in at least one embodiment of the present disclosure, the convolutional neural network involves K sub-convolution kernels, and K is equal to the number of output channels C_out; the deployment unit is further configured to: based on the situation where the K sub-convolution kernels slide in the sub-input feature map with a step of S, deploy each of the K sub-convolution kernels into N columns of at least one memory-computing integrated array in sequence, thereby deploying the K sub-convolution kernels into N columns of at least one memory-computing integrated array in sequence. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings described below only relate to some embodiments of the present disclosure and do not limit the present disclosure.
[0018] Figure 1 FIG. shows a schematic diagram of a mapping method of a convolutional neural network;
[0019] Figure 2A FIG. shows a schematic structure of a memristor array;
[0020] Figure 2B FIG. is a schematic diagram of a 1T1R structure memristor unit;
[0021] Figure 2C FIG. is a schematic diagram of a 2T2R structure memristor unit;
[0022] Figure 3 FIG. shows a schematic flowchart of a convolutional neural network processing method provided in at least one embodiment of the present disclosure;
[0023] Figure 4 FIG. shows a schematic diagram of dynamically adjusting the three-dimensional sizes of a sub-input feature map and a sub-convolution kernel;
[0024] Figure 5 FIG. shows three situations when a sub-convolution kernel performs a convolution operation with a sliding step of S within a sub-input feature map;
[0025] Figure 6A FIG. shows an example of a convolutional neural network processing method provided in at least one embodiment of the present disclosure;
[0026] Figure 6B shows the schematic diagram of setting the three-dimensional sizes of the sub-input feature map and the sub-convolution kernel in the example of Figure 6A ;
[0027] Figure 7 shows the schematic block diagram of a convolutional neural network processing device based on a processing-in-memory array provided by at least one embodiment of the present disclosure;
[0028] Figure 8A is the schematic structural diagram of a processing-in-memory array provided by at least one embodiment of the present disclosure;
[0029] Figure 8B is the schematic diagram of another processing-in-memory array provided by at least one embodiment of the present disclosure;
[0030] Figure 8C shows a processing-in-memory array constructed by memristor units adopting a 2T2R structure;
[0031] Figure 8D shows another processing-in-memory array constructed by memristor units adopting a 2T2R structure;
[0032] Figure 9 shows the schematic diagram of the convolutional neural network processing device of at least one embodiment of the present disclosure. Detailed implementation manners
[0033] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.
[0034] Unless otherwise defined, the technical terms or scientific terms used in this disclosure shall have the ordinary meanings as understood by those of ordinary skill in the art to which this disclosure pertains. The terms "first", "second" and similar terms used in this disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Similarly, the terms such as "a", "an" or "the" do not denote a quantity limitation, but mean that there is at least one. The terms such as "comprising" or "including" mean that the elements or objects appearing before this term cover the elements or objects listed after this term and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left" and "right" are only used to represent relative position relationships. When the absolute position of the object being described changes, the relative position relationship may also change accordingly.
[0035] The basic operation of a convolutional network is multiply-accumulate operation, rather than direct vector-matrix multiplication. Therefore, when mapping software weights to the in-memory computing array of hardware, some changes in element positions and orders are also required.
[0036] Figure 1 The figure shows a schematic diagram of a mapping method of a convolutional neural network. For this mapping method, the basic idea of the process of mapping software weights to the in-memory computing array based on memristors is to convert the entire convolutional operation into vector-matrix multiplication.
[0037] Memristors (such as resistive random access memories, phase change memories, conductive bridge memories, etc.) are a new type of micro-nano electronic device, and their conductance states can be adjusted by applying external stimuli. As a two-terminal device, a memristor has the characteristics of adjustable resistance and non-volatility. According to Kirchhoff's current law and Ohm's law, the in-memory computing array composed of such devices can perform analog multiply-add calculations in parallel, so as to directly process the input analog signals, and both storage and calculation occur in the memristors of the array. The in-memory computing array includes multiple rows and multiple columns of computing units, and each computing unit is a memristor unit. Figure 1 The intersection position of the middle row and column is the computing unit.
[0038] Such as Figure 1As shown, the three-dimensional size of the sub-input feature map (Sub-IFM) in the input feature map (Input Feature Map, IFM) is the same as the three-dimensional size of the convolution kernel (Kernel). The sub-input feature map in the input feature map is input into the storage-computing integrated array in the form of a row input voltage signal. A convolution kernel is deployed in each column of the storage-computing integrated array (that is, the conductance value of the memristor in the storage-computing integrated array is used to characterize the weight element (Weight) in the convolution kernel, and the number of output channels (C_out) of the convolution is equal to the number of columns used in the storage-computing integrated array. The storage-computing integrated array performs calculations through the row input voltage and the conductance value of the memristor, and obtains the calculation results from the column output current of the storage-computing integrated array.
[0039] However, due to the physical properties of the memristor itself, its high-resistance resistance is only in the megaohm range, so the current and line resistance limit the scale of the memristor-based storage and computing array. Therefore, when processing input feature maps with a large scale, the above mapping method will have the following problems:
[0040] Since the scale of the storage-computation-in-one array is limited, the number of rows of the storage-computation-in-one array is limited, which results in one column of the array being unable to fully store all the weight elements of a convolution kernel, and multiple storage-computation-in-one arrays need to be spliced. In addition, since the number of columns used in the storage-computation-in-one array is equal to the number of output channels of the convolution, if the number of output channels is small, only some columns in the storage-computation-in-one array are used, and the remaining columns cannot be used, resulting in a huge waste of resources. In addition, when deploying a convolutional neural network on a storage-computation-in-one array, it is often necessary to copy multiple copies of the weights in order to speed up the calculation and reduce latency, so the number of storage-computation-in-one arrays also needs to be copied exponentially, greatly increasing the number of memristor arrays and power consumption overhead.
[0041] At least one embodiment of the present disclosure provides a convolutional neural network processing method based on a storage-computation integrated array, the storage-computation integrated array includes a plurality of rows and columns of operation units, the method comprising: based on the width x_in, height y_in and number of channels C_in of the input feature map and the width x_k, height y_k and number of channels C_in of the convolution kernel, setting the width of the sub-input feature map of the input feature map to x, the height to y and the number of channels to z, and setting the width of the sub-convolution kernel of the convolution kernel to x_k, the height to y_k and the number of channels to z, wherein x_ k is less than or equal to x, y_k is less than or equal to y, x is less than or equal to x_in, y is less than or equal to y_in, z is less than or equal to C_in, and the result of x×y×z is less than or equal to the number of rows of the storage-computation-in-one array; based on the situation that the sub-convolution kernel slides in the sub-input feature map with a step size S, the sub-convolution kernel is deployed in sequence to the N columns of at least one storage-computation-in-one array; the sub-input feature map is input into the at least one storage-computation-in-one array as a row input signal, and the calculation result is obtained from the column output signal of the N columns of the at least one storage-computation-in-one array.
[0042] For different neural network layer parameters and different in-memory computing array scales, the convolutional neural network processing method can flexibly adjust the three-dimensional size of the sub-kernel, improving the versatility of the method, as well as the area utilization rate and computing efficiency of the in-memory computing array. And in at least one embodiment, the method can be optimized from a software perspective without changing the hardware, and can be deployed on the original in-memory computing array, with good compatibility and applicability.
[0043] At least one embodiment of the present disclosure also provides a convolutional neural network processing device corresponding to the above convolutional neural network processing method.
[0044] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings, but the present disclosure is not limited to these specific embodiments.
[0045] Figure 2A FIG. shows a schematic structure of an in-memory computing array, which is composed of a plurality of computing units (such as memristor units). The plurality of computing units form an array of M rows and N columns, where both M and N are positive integers. Each computing unit includes a switching element and one or more memristors. In Figure 2A WL<1>, WL<2>... WL <m>Respectively represent the word lines of the first row, the second row... the Mth row, and the control electrodes (such as the gates of transistors) of the switching elements in the arithmetic unit circuit of each row are connected to the word line corresponding to that row; BL<1>, BL<2>... BL <n>Bit lines respectively representing the first column, the second column... the Nth column, where the memristors in the arithmetic unit circuits of each column are connected to the bit lines corresponding to that column; SL<1>, SL<2>... SL <m>respectively represent the source lines of the first row, the second row, …, the Mth row, and the source electrodes of the transistors in the arithmetic unit circuit of each row are connected to the corresponding source line of that row. According to Kirchhoff's law, by setting the state (such as resistance value) of the arithmetic unit and applying corresponding word line signals and bit line signals to the word line and the bit line, the above-mentioned memory-computation integrated array can perform multiply-accumulate calculations in parallel.
[0046] Figure 2A The arithmetic units in the memory-computation integrated array can, for example, have a 1T1R structure or a 2T2R structure. Among them, the arithmetic unit with a 1T1R structure includes one transistor and one memristor, and the arithmetic unit with a 2T2R structure includes two transistors and two memristors. For example, memristors include, but are not limited to, RRAM, PCRAM, ECRAM, Flash, etc. It should be noted that the present disclosure does not limit the structure of the memristor unit, and other structural forms of memristor units that can implement multiply-accumulate operations can also be used, such as 1S1F, 0T1R, etc.
[0047] It should be noted that the transistors used in the embodiments of the present disclosure can all be thin film transistors or field effect transistors (such as MOS field effect transistors) or other switching devices with the same characteristics. The source electrode and the drain electrode of the transistors used here can be symmetric in structure, so there can be no difference between them in structure. In the embodiments of the present disclosure, in order to distinguish the two poles of the transistor other than the gate, one of the poles is directly described as the first pole, and the other pole is the second pole.
[0048] Figure 2B is a schematic diagram of an arithmetic unit with a 1T1R structure. As Figure 2B shown, the arithmetic unit with a 1T1R structure includes one transistor M1 and one memristor R1.
[0049] Embodiments of the present disclosure do not limit the type of transistor used. For example, when the transistor M1 is an N-type transistor, its gate is connected to the word line WL. For example, when the word line WL inputs a high level, the transistor M1 conducts; the first pole of the transistor M1 can be the source and is configured to be connected to the source line SL. For example, the transistor M1 can receive a reset voltage through the source line SL; the second pole of the transistor M1 can be the drain and is configured to be connected to the second pole (e.g., the negative pole) of the memristor R1. The first pole (e.g., the positive pole) of the memristor R1 is connected to the bit line BL. For example, the memristor R1 can receive a set voltage through the bit line BL. For example, when the transistor M1 is a P-type transistor, its gate is connected to the word line WL. For example, when the word line WL inputs a low level, the transistor M1 conducts; the first pole of the transistor M1 can be the drain and is configured to be connected to the source line SL. For example, the transistor M1 can receive a reset voltage through the source line SL; the second pole of the transistor M1 can be the source and is configured to be connected to the second pole (e.g., the negative pole) of the memristor R1. The first pole (e.g., the positive pole) of the memristor R1 is connected to the bit line BL. For example, the memristor R1 can receive a set voltage through the bit line BL. It should be noted that the memristor structure can also be implemented as other structures, such as a structure in which the second pole of the memristor R1 is connected to the source line SL. Embodiments of the present disclosure do not limit this.
[0050] In the following embodiments, the transistor M1 being an N-type transistor is taken as an example for description.
[0051] The function of the word line terminal WL is to apply a corresponding voltage to the gate of the transistor M1, thereby controlling the conduction or cutoff of the transistor M1. When operating on the memristor R1, for example, during a set operation or a reset operation, the transistor M1 needs to be turned on first, that is, a conduction voltage needs to be applied to the gate of the transistor M1 through the word line terminal WL. After the transistor M1 conducts, for example, voltages can be applied to the memristor R1 through the source line terminal SL and the bit line terminal BL to change the resistance state of the memristor R1. For example, a set voltage can be applied through the bit line terminal BL to make the memristor R1 in a low-resistance state; or for another example, a reset voltage can be applied through the source line terminal SL to make the memristor R1 in a high-resistance state. For example, the resistance value of the high-resistance state is more than 100 times, for example, more than 1000 times, that of the low-resistance state.
[0052] It should be noted that in the embodiments of the present disclosure, by applying voltages to the word line terminal WL and the bit line terminal BL simultaneously, the resistance value of the memristor R1 can be made smaller and smaller, that is, the memristor R1 changes from a high-resistance state to a low-resistance state. The operation of making the memristor R1 change from a high-resistance state to a low-resistance state is called a set operation; by applying voltages to the word line terminal WL and the source line terminal SL simultaneously, the resistance value of the memristor R1 can be made larger and larger, that is, the memristor R1 changes from a low-resistance state to a high-resistance state. The operation of making the memristor R1 change from a low-resistance state to a high-resistance state is called a reset operation. For example, the memristor R1 has a threshold voltage. When the amplitude of the input voltage is less than the threshold voltage of the memristor R1, the resistance value (or conductance value) of the memristor R1 will not change. In this case, by inputting a voltage less than the threshold voltage, calculations can be performed using the resistance value (or conductance value) of the memristor R1; by inputting a voltage greater than the threshold voltage, the resistance value (or conductance value) of the memristor R1 can be changed.
[0053] Figure 2C Schematic diagram of an arithmetic unit with a 2T2R structure. As Figure 2C shown, the arithmetic unit with a 2T2R structure includes two transistors M1 and M2 and two memristors R1 and R2. Hereinafter, an example will be given in which both the transistors M1 and M2 are N-type transistors.
[0054] The gate of the transistor M1 is connected to the word line terminal WL1. For example, when a high level is input to the word line terminal WL1 of M1, the transistor M1 is turned on. The gate of the transistor M2 is connected to the word line terminal WL2. For example, when a high level is input to the word line terminal WL2 of M2, the transistor M2 is turned on; the first pole of the transistor M1 can be the source and is configured to be connected to the source line terminal SL. For example, the transistor M1 can receive a reset voltage through the source line terminal SL. The first pole of the transistor M2 can be the source and is configured to be connected to the source line terminal SL. For example, the transistor M2 can receive a reset voltage through the source line terminal SL. The first pole of the transistor M1 is connected to the first pole of the transistor M2 and they are together connected to the source line terminal SL. The second pole of the transistor M1 can be the drain and is configured to be connected to the second pole (for example, the negative pole) of the memristor R1. The first pole (for example, the positive pole) of the memristor R1 is connected to the bit line terminal BL1. For example, the memristor R1 can receive a set voltage through the bit line terminal BL1; the second pole of the transistor M2 can be the drain and is configured to be connected to the second pole (for example, the negative pole) of the memristor R2. The first pole (for example, the positive pole) of the memristor R2 is connected to the bit line terminal BL2. For example, the memristor R2 can receive a set voltage through the bit line terminal BL2.
[0055] It should be noted that the transistors M1 and M2 in the arithmetic unit with a 2T2R structure can also both be P-type transistors, which will not be elaborated here.
[0056] Figure 3 FIG. 0 shows a schematic flowchart of a convolutional neural network processing method provided by at least one embodiment of the present disclosure. For example, the in-memory computing array is Figure 2A the in-memory computing array shown. The in-memory computing array includes multiple rows and multiple columns of computing units. For example, the structure of the computing unit is as Figure 2B or shown in 2C.
[0057] As Figure 3 shown, the convolutional neural network processing method includes the following steps S301 to S303.
[0058] Step S301: Based on the width x_in, height y_in, and number of channels C_in of the input feature map, and the width x_k, height y_k, and number of channels C_in of the convolutional kernel, set the width of the sub-input feature map of the input feature map to x, the height to y, the number of channels to z, and set the width of the sub-convolutional kernel of the convolutional kernel to x_k, the height to y_k, and the number of channels to z, where x_k is less than or equal to x, y_k is less than or equal to y, x is less than or equal to x_in, y is less than or equal to y_in, z is less than or equal to C_in, and the result of x × y × z is less than or equal to the number of rows of the in-memory computing array.
[0059] For example, in some embodiments of the present disclosure, the width, height, and number of channels of the input feature map, and the width, height, and number of channels of the convolutional kernel are used to obtain the three-dimensional dimensions of the sub-input feature map and the sub-convolutional kernel according to the convolutional neural network adopted for training or inference, and then the three-dimensional dimensions of the sub-input feature map and the sub-convolutional kernel can be adjusted within a certain range based on the width, height, and number of channels of the input feature map, and the width, height, and number of channels of the convolutional kernel. During the process of adjusting the three-dimensional dimensions of the sub-input feature map and the sub-convolutional kernel, for example, the width of the sub-input feature map is less than or equal to the width of the input feature map, or the height of the sub-input feature map is less than or equal to the height of the input feature map, or the number of channels of the sub-input feature map is less than or equal to the number of channels of the input feature map, or the width of the sub-convolutional kernel is equal to the width of the convolutional kernel and less than the width of the sub-input feature map, or the height of the sub-convolutional kernel is equal to the height of the convolutional kernel and less than the height of the sub-input feature map, or the number of channels of the sub-convolutional kernel is equal to the number of channels of the sub-input feature map, or the result of multiplying the width, height, and number of channels of the sub-input feature map is less than or equal to the number of rows of the in-memory computing array.
[0060] In at least one embodiment, for the dynamic adjustment of the three-dimensional dimensions of the sub-input feature map and the sub-convolutional kernel, reference can be made to Figure 4 . Figure 4 FIG. shows a schematic diagram of the dynamic adjustment of the three-dimensional dimensions of the sub-input feature map and the sub-convolutional kernel.
[0061] As Figure 4 As shown, the width of the input feature map is \(x_{in}\), the height is \(y_{in}\), and the number of channels is \(C_{in}\). The width of the convolutional kernel is \(x_{k}\), the height is \(y_{k}\), and the number of channels is \(C_{in}\). In this embodiment, the three-dimensional size of the sub-input feature map can be adjusted within a certain range and is no longer restricted by the three-dimensional size of the convolutional kernel (having the same three-dimensional size as the convolutional kernel). The three-dimensional size of the sub-input feature map satisfies the following conditions: the width \(x\) is less than or equal to \(x_{in}\), the height \(y\) is less than or equal to \(y_{in}\), the number of channels \(z\) is less than or equal to \(C_{in}\), and the result of \(x\times y\times z\) is less than or equal to the number of rows of the memory and computing integrated array. After setting the three-dimensional size of the sub-input feature map based on the above conditions, the three-dimensional size of the sub-convolutional kernel can be determined. The width of the sub-convolutional kernel is equal to the width of the convolutional kernel, both being \(x_{k}\), the height of the sub-convolutional kernel is equal to the height of the convolutional kernel, both being \(y_{k}\), and the number of channels of the sub-convolutional kernel is consistent with the number of channels of the sub-input feature map, both being \(z\).
[0062] In some embodiments of the present disclosure, for the case where \(z\leq C_{in}\), the number of at least one memory and computing integrated array is Here is a positive integer.
[0063] Since \(Z\leq C_{in}\), therefore, there is a need for sub-input feature maps to process all channel data. If one memory and computing integrated array can process the data calculation of 1 sub-input feature map, then without considering replicating the memory and computing integrated array, arrays are required.
[0064] Now back to Figure 3 , step S302: Based on the case where the sub-convolutional kernel slides in the sub-input feature map with a stride of \(S\), deploy the sub-convolutional kernel into \(N\) columns of at least one memory and computing integrated array in sequence. For example, the \(N\) columns can be part of or all of the columns in the memory and computing integrated array. Deploying the sub-convolutional kernel into at least one memory and computing integrated array in sequence means converting (mapping) the weight values of each element of the sub-convolutional kernel into the conductance values of the corresponding arithmetic units of the memory and computing integrated array in the required manner of the memory and computing integrated array.
[0065] In the convolution operation, the convolutional kernel slides step by step in the input feature map with a stride of \(S\) (a positive integer) and performs the convolution operation to obtain the output items in the output feature map. Similarly, the sub-convolutional kernel slides step by step in the sub-input feature map with a stride of \(S\) and performs the convolution operation.
[0066] In some embodiments of the present disclosure, step S302 may include: dividing the position of the sub-convolution kernel relative to the sub-input feature map into a first convolution position, a second convolution position, and a third convolution position according to the case where the sub-convolution kernel slides in the sub-input feature map with a stride S; and deploying the sub-convolution kernel to N columns of at least one memory-compute integrated array in sequence according to the first convolution position, the second convolution position, and the third convolution position. Here, the first convolution position corresponds to the case where the sub-convolution kernel is completely within the sub-input feature map during the sliding process, the second convolution position corresponds to the case where the sub-convolution kernel is between two sub-input feature maps during the sliding process, and the third convolution position corresponds to the case where the sub-convolution kernel is between four sub-input feature maps during the sliding process.
[0067] Figure 5 Illustrates three cases when the sub-convolution kernel performs a convolution operation with a stride of S within the sub-input feature map, namely the cases of the first convolution position, the second convolution position, and the third convolution position.
[0068] As Figure 5 shown, when the sub-convolution kernel slides within the sub-input feature map, its position relative to the sub-input feature map is divided into three cases: the first convolution position, the second convolution position, and the third convolution position. The first convolution position is represented by the letter a, where the sub-convolution kernel is completely within the sub-input feature map, the second convolution position is represented by the letter b, where the sub-convolution kernel is between two sub-input feature maps, and the third convolution position is represented by the letter c, where the sub-convolution kernel is between four sub-input feature maps.
[0069] Returning again to Figure 3 , in some embodiments of the present disclosure, deploying the sub-convolution kernel to N columns of at least one memory-compute integrated array in sequence according to the first convolution position, the second convolution position, and the third convolution position may include: deploying the sub-convolution kernel to n1 columns of at least one memory-compute integrated array according to the first convolution position, deploying the sub-convolution kernel to n2 columns of at least one memory-compute integrated array according to the second convolution position, and deploying the sub-convolution kernel to n3 columns of at least one memory-compute integrated array according to the third convolution position, where one sub-convolution kernel is deployed in each column of the memory-compute integrated array, and n1 + n2 + n3 = N.
[0070] When the sub-convolution kernel slides within the sub-input feature map, the total number of convolutions at the first convolution position is n1, the total number of convolutions at the second convolution position is n2, and the total number of convolutions at the third convolution position is n3.
[0071] The width and height of the sub-convolution kernel set in step S301 may be the same or different, and different formulas are used to calculate n1, n2, and n3 for the cases where the width and height are the same and the cases where the width and height are different.
[0072] For example, in some embodiments of the present disclosure, when \(x_k\) and \(y_k\) of the sub-convolution kernel are the same, the calculation of \(n1\) is shown in formula (1), the calculation of \(n2\) is shown in formula (2), and the calculation of \(n3\) is shown in formula (3).
[0073]
[0074]
[0075] n3 = C_out × 4 × (x_k - 1) 2 (3)
[0076] Alternatively, when \(x_k\) and \(y_k\) of the sub-convolution kernel are different, the calculation of \(n1\) is shown in formula (4), the calculation of \(n2\) is shown in formula (5), and the calculation of \(n3\) is shown in formula (6).
[0077]
[0078]
[0079] n3 = Cout × 4 × (x_k - 1) × (y_k - 1) (6)
[0080] Since each convolution position (corresponding to 1 sub-convolution kernel) needs to be deployed on one column of the memory-compute integrated array, the total number of columns N of the memory-compute integrated array required can be calculated through formulas (1) to (6), and N = n1 + n2 + n3.
[0081] The number of weight elements corresponding to the sub-convolution kernel at different convolution positions is different, and the calculation of the specific number of weight elements is as follows.
[0082] For example, in some embodiments of the present disclosure, when \(x_k\) and \(y_k\) of the sub-convolution kernel are the same, the number of weight elements corresponding to the sub-convolution kernel at the first convolution position is \(x_k×y_k×z\), the number of weight elements corresponding to the sub-convolution kernel at the second convolution position is \(x_k×z\), and the number of weight elements corresponding to the sub-convolution kernel at the third convolution position is \(z\). Alternatively, when \(x_k\) and \(y_k\) of the sub-convolution kernel are different, the number of weight elements corresponding to the sub-convolution kernel at the first convolution position is \(x_k×y_k×z\), the number of weight elements corresponding to the sub-convolution kernel at the third convolution position is \(z\), and on half of the \(n2\) columns, the number of weight elements corresponding to the sub-convolution kernel at the second convolution position is \(x_k×z\), and on the other half of the \(n2\) columns, the number of weight elements corresponding to the sub-convolution kernel at the second convolution position is \(y_k×z\).
[0083] Since the number of weight elements of the sub-convolution kernel is less than the number of elements of the sub-input feature map, the weight elements deployed on one column of the in-memory computing array are sparse, and the equivalent conductances at the remaining row positions of this column can be set to 0.
[0084] Since in this embodiment, the weight elements deployed on each column of the in-memory computing array are sparse, the current magnitude on the same column can be reduced, and the serious IR drop problem in the in-memory computing array can be suppressed, further improving the computing accuracy and reliability of the in-memory computing array.
[0085] Step S303: Input the sub-input feature map as a row input signal into at least one in-memory computing array, and obtain the calculation result from the column output signals of N columns of at least one in-memory computing array.
[0086] For example, in some embodiments of the present disclosure, step S303 may include: gradually sliding from the input feature map to obtain the sub-input feature map in a sliding manner with a sliding step of x in the width dimension and a sliding step of y in the height dimension. In each step, input the sub-input feature map as a row input signal into at least one in-memory computing array, and obtain the calculation result corresponding to this step from the column output signals of N columns of at least one in-memory computing array.
[0087] In this embodiment, sliding calculations are performed on the entire input feature map with the sub-input feature map as a unit. Each sliding step is the size of the sub-input feature map in both the width and height dimensions. Each time a slide is made, a sub-input feature map is obtained, and the elements of this sub-input feature map are input as a row input signal into an in-memory computing array.
[0088] For example, the row input signal is a voltage signal, and the values of the convolution kernel are mapped to the conductance values of memristors. Therefore, according to the voltage signal input to the in-memory computing array and the conductance values of the memristors in the in-memory computing array, the column output signals (column output currents) of N columns of the in-memory computing array can be obtained, and these column output signals are the calculation results of a certain sliding calculation.
[0089] Performing sliding calculations on the entire input feature map with the sub-input feature map as a unit ensures that there is no overlapping part between adjacent sub-input feature maps, avoids the problem of possible partial overlap of different sub-input feature maps, maximizes the utilization of input data, reduces the number of slides, and greatly improves the array computing efficiency.
[0090] For example, in some embodiments of the present disclosure, a convolutional neural network involves K sub-convolution kernels, and K is equal to the number of output channels C_out; based on the case where the sub-convolution kernels slide in the sub-input feature map with a stride S, the K sub-convolution kernels are sequentially deployed into N columns of at least one memory-compute integrated array, including: based on the case where the sub-convolution kernels slide in the sub-input feature map with a stride S, each of the K sub-convolution kernels is sequentially deployed into N columns of at least one memory-compute integrated array, whereby the K sub-convolution kernels are sequentially deployed into N columns of at least one memory-compute integrated array.
[0091] In a convolutional neural network, the number of output channels of a convolution is equal to the number of convolution kernels. Therefore, if a convolutional neural network involves K sub-convolution kernels, the number of output channels C_out is equal to K. Assuming that for the case where 1 sub-convolution kernel slides in the sub-input feature map with a stride S, N columns of the memory-compute integrated array are required, then for the case where K sub-convolution kernels slide in the sub-input feature map with a stride S, K×N columns of the memory-compute integrated array are required.
[0092] Taking the deployment of a convolutional layer in a convolutional neural network onto a memory-compute integrated array with a size of 400 rows × 150 columns as an example, the convolutional neural network processing method provided by at least one embodiment of the present disclosure is described below.
[0093] Figure 6A An example of the convolutional neural network processing method provided by at least one embodiment of the present disclosure is shown. The relevant parameters of the convolutional layer in this example are shown in Table 1 below.
[0094] Table 1
[0095]
[0096] First, the scale of the convolutional neural network and the memory-compute integrated array is analyzed and parameters are selected. If the existing mapping method is adopted, 1 convolution kernel has 3×3×64 = 576 elements, while the memory-compute integrated array has only 400 rows, which is less than 576. Therefore, all elements of the convolution kernel cannot be deployed using one memory-compute integrated array. In addition, since the number of output channels C_out = 3, only 3 columns of the memory-compute can be utilized, and the remaining 147 columns are wasted.
[0097] Figure 6B A schematic diagram showing the three-dimensional size settings of the sub-input feature map and the sub-convolution kernel in this example is shown, as Figure 6B As shown, based on the width 3840 (x_in), height 2160 (y_in), and number of channels 64 (C_in) of the input feature map, as well as the width 3 (x_k), height 3 (y_k), and number of channels 64 (C_in) of the convolutional kernel, the width of the sub-input feature map of the input feature map is set to 5 (x), height to 5 (y), and number of channels to 16 (z), and the width of the sub-convolutional kernel of the convolutional kernel is set to 3 (x_k), height to 3 (y_k), and number of channels to 16 (z). Since there are C_in = 64 input channels for the input feature map and the convolutional kernel, compute-in-memory arrays are required to complete the calculation without considering replication.
[0098] Next, based on the case where the sub-convolutional kernel slides in the sub-input feature map with a stride S (here S = 1), the sub-convolutional kernel is sequentially deployed into N columns of at least one compute-in-memory array. According to the calculation of formula (1), the number of convolutions at the first convolution position is n1 = 27, and 27 sub-convolutional kernels can be placed in 27 columns of the compute-in-memory array. According to the calculation of formula (7), the number of weight elements deployed in each column of these 27 columns is 3×3×16 = 144. So the weight elements deployed in each column are sparse, and only need to deploy the weight elements to the corresponding rows according to the positions of different convolution calculations; similarly, according to the calculation of formula (2), the number of convolutions at the second convolution position is n2 = 72, and 72 sub-convolutional kernels can be placed in 72 columns of the compute-in-memory array. According to the calculation of formula (8), the number of weight elements deployed in each column is 3×16 = 48, and the weight elements deployed in each column are also sparse; according to the calculation of formula (3), the number of convolutions at the third convolution position is n3 = 48, and 48 sub-convolutional kernels can be placed in 48 columns of the compute-in-memory array. According to the calculation of formula (9), the number of weight elements deployed in each column is 16, and the weight elements deployed in each column are also sparse. In summary, the compute-in-memory array uses a total of N = n1 + n2 + n3 = 147 columns (less than 150). It should be noted that since there are C_out = 3 convolutions with different output channels at one convolution position, there will be three columns with the same deployment position of weight elements, but different weight values. For example, as Figure 6A shown, for convolution position 1, there are three columns with the same deployment position of weight elements, but the values of the weight elements in these three columns are different; of course, for the different columns of weight elements corresponding to the convolutions of different output channels, they may not be adjacent to each other.
[0099] Next, the sub-input feature map is input as a row input signal into at least one processing-in-memory array, and the calculation result is obtained from the column output signals of N columns of at least one processing-in-memory array. The sliding calculation is performed on the entire input feature map with the sub-input feature map as a unit, and a total of 4 sub-input feature maps are obtained. These 4 sub-input feature maps are respectively input as row input signals into 4 processing-in-memory arrays. There are 5×5×16 = 400 elements in one sub-input feature map. These elements are linearly mapped to the values within the input voltage range according to the magnitude of the weight values, converted into input voltage signals through a digital-to-analog converter, and respectively input to 400 rows of one processing-in-memory array. Through the input voltage signals and the conductance values representing the weight elements, the calculation result is obtained from the column output signals of N columns of at least one processing-in-memory array.
[0100] Next, the performance of the convolutional neural network processing method provided by at least one embodiment of the present disclosure is compared with the performance of other mapping methods in conjunction with Table 2. These compared mapping methods include Mapping Method 1, New Mapping Method 2, Mapping Method 3, etc. Among them, the performance parameters of different methods are shown in Table 2 below.
[0101] Table 2
[0102]
[0103] For Mapping Method 1, see: W. Zhang et al., "Design Guidelines of RRAM based Neural-Processing-Unit: A Joint Device-Circuit-Algorithm Analysis,” in 2019 56th ACM / IEEE Design Automation Conference (DAC), Jun. 2019, pp. 1–6.
[0104] For Mapping Method 2, see: X. Peng, R. Liu, and S. Yu, "Optimizing Weight Mapping and Data Flow for Convolutional Neural Networks on Processing-in-Memory Architectures,” IEEE Transactions on Circuits and Systems I: Regular Papers, vol. 67, no. 4, pp. 1333–1343, Apr. 2020.
[0105] Mapping method 3 can be found in: W. Li, P. Xu, Y. Zhao, H. Li, Y. Xie, and Y. Lin, "Timely: Pushing Data Movements And Interfaces In Pim Accelerators Towards Local And In Time Domain,” in 2020 ACM / IEEE 47th Annual International Symposium on Computer Architecture (ISCA), May 2020, pp. 832–845.
[0106] As shown in Table 2, the convolutional neural network processing method provided by at least one embodiment of the present disclosure can greatly improve the area utilization rate of the array, reduce the number of sliding times required to complete the convolution, and improve the computing efficiency.
[0107] Figure 7 The schematic block diagram of a convolutional neural network processing device 700 provided by at least one embodiment of the present disclosure is shown, and this convolutional neural network processing device can be used to execute Figure 3 the convolutional neural network processing method shown.
[0108] As Figure 7 shown, the convolutional neural network processing device 700 includes a memory-computation integrated array 701, a setting unit 702, a deployment unit 703, and a control unit 704.
[0109] The memory-computation integrated array 701 includes multiple rows and multiple columns of computing units. For example, the structure of the memory-computation integrated array 701 is as Figure 2A shown, and the structure of the computing unit is as Figure 2B or 2C shown.
[0110] The setting unit 702 is configured to set the width of the sub-input feature map to x, the height to y, and the number of channels to z based on the width x_in, height y_in, and number of channels C_in of the input feature map and the width x_k, height y_k, and number of channels C_in of the convolutional kernel, and set the width of the sub-convolutional kernel to x_k, the height to y_k, and the number of channels to z, where x_k is less than or equal to x, y_k is less than or equal to y, x is less than or equal to x_in, y is less than or equal to y_in, z is less than or equal to C_in, and the result of x × y × z is less than or equal to the number of rows of the memory-computation integrated array.
[0111] The deployment unit 703 is configured to deploy the sub-convolutional kernel into N columns of at least one memory-computation integrated array in sequence based on the situation where the sub-convolutional kernel slides in the sub-input feature map with a stride S.
[0112] The control unit 704 is configured to input the sub-input feature map as a row input signal into at least one in-memory computing array, and obtain a calculation result from the column output signals of N columns of the at least one in-memory computing array.
[0113] For example, in at least one embodiment, the deployment unit 703 is further configured to: divide the position of the sub-convolution kernel relative to the sub-input feature map into a first convolution position, a second convolution position, and a third convolution position according to the situation where the sub-convolution kernel slides in the sub-input feature map with a stride S; and deploy the sub-convolution kernel into N columns of the at least one in-memory computing array in sequence according to the first convolution position, the second convolution position, and the third convolution position, where the first convolution position corresponds to the situation where the sub-convolution kernel is completely within the sub-input feature map during the sliding process, the second convolution position corresponds to the situation where the sub-convolution kernel is between two sub-input feature maps during the sliding process, and the third convolution position corresponds to the situation where the sub-convolution kernel is between four sub-input feature maps during the sliding process.
[0114] For example, in at least one embodiment, the deployment unit 703 is further configured to: deploy the sub-convolution kernel into n1 columns of the at least one in-memory computing array according to the first convolution position, deploy the sub-convolution kernel into n2 columns of the at least one in-memory computing array according to the second convolution position, and deploy the sub-convolution kernel into n3 columns of the at least one in-memory computing array according to the third convolution position, where each column of the in-memory computing array deploys 1 sub-convolution kernel, and n1 + n2 + n3 = N.
[0115] For example, in at least one embodiment, the control unit 704 is further configured to: gradually slide from the input feature map to obtain the sub-input feature map in a sliding manner with a stride of x in the width dimension and a stride of y in the height dimension, input the sub-input feature map as a row input signal into the at least one in-memory computing array in each step, and obtain the calculation result corresponding to this step from the column output signals of N columns of the at least one in-memory computing array.
[0116] For example, in at least one embodiment, the convolutional neural network involves K sub-convolution kernels, and K is equal to the number of output channels C_out; the deployment unit 703 is further configured to: deploy each of the K sub-convolution kernels into N columns of the at least one in-memory computing array in sequence based on the situation where the sub-convolution kernel slides in the sub-input feature map with a stride S, thereby deploying the K sub-convolution kernels into N columns of the at least one in-memory computing array in sequence.
[0117] In some embodiments of the present disclosure, during the deployment process, when mapping the weight elements to the in-memory computing array, since the values of the weight elements have positive and negative values, two memristors forming a memristor pair can be used to correspond to one weight element.
[0118] For example, a weight element can be implemented by two memristors. For example, two memristor arrays can be used to form multiple memristor pairs, and each memristor pair includes two memristors. For example, the two memristors are arranged directly adjacent to each other in the memristor array; for another example, one memristor in each memristor pair is used to receive an input voltage signal, and the other memristor in the memristor pair is used to receive an inverted input voltage signal corresponding to the above input voltage signal.
[0119] Next, through Figure 8A , Figure 8B specific examples will be given to illustrate an example of a memory - in - computing array that can implement negative - value elements.
[0120] Figure 8A FIG. is a schematic structural diagram of a memory - in - computing array provided by at least one embodiment of the present disclosure.
[0121] As Figure 8A shown, the memristor 801 and the memristor 802 can form a memristor pair. The conductance value of the memristor 801 is denoted as G 11 , and the conductance value of the memristor 802 is denoted as G 12 . Since the memristor 802 is connected to an inverter, when the memristor 801 receives a positive - polarity input voltage signal, the inverter can invert the polarity of the input voltage signal, so that the memristor 802 receives a negative - polarity input voltage signal. For example, the input voltage signal received by the memristor 801 is denoted as v(t), and the input voltage signal received by the memristor 802 is denoted as - v(t). The memristor 801 and the memristor 802 are connected to two different SLs, and the input voltage signal generates an output current through the memristor. The output current passing through the memristor 801 and the output current passing through the memristor 802 are superimposed at the end of the SL. Therefore, the result of the multiply - accumulate calculation of the memristor 801 and the memristor 802 is v(t)G 11 +( - v(t))G 12 , that is, v(t)(G 11 - G 12 ). Thus, the memristor pair composed of the memristor 801 and the memristor 802 can correspond to a weight element, and this weight element is G 11 - G 12 . By configuring the numerical relationship of G 11 - G 12 , negative - value elements can be implemented. For example, if a weight element of 0 needs to be characterized, the memristors representing positive and negative weights can be set to the same conductance state.
[0122] Figure 8B FIG. is a schematic diagram of another memory - in - computing array provided by at least one embodiment of the present disclosure.
[0123] As Figure 8B As shown, for example, memristor 801 and memristor 802 can form a memristor pair, and the conductance value of memristor 801 is denoted as G 11 , and the conductance value of memristor 802 is denoted as G 12 . Different from Figure 8A that, memristor 802 is not connected to an inverter. Therefore, when memristor 801 receives a positive-polarity input voltage signal, memristor 802 also receives a positive-polarity input voltage signal. For example, the input voltage signal received by memristor 801 is represented by v(t), and the input voltage signal received by memristor 802 is also represented by v(t). Memristor 801 and memristor 802 are connected to two different SLs, and the output current passing through memristor 801 and the output current passing through memristor 802 are subtracted at the end of the SL. Therefore, the result of the multiply-accumulate calculation of memristor 801 and memristor 802 is v(t)G 11 - v(t)G 12 , that is, v 0 (t)(G 11 - G 12 ). Therefore, the memristor pair composed of memristor 801 and memristor 802 can be a weight element, and this weight element is G 11 - G 12 . By configuring the numerical relationship of G 11 - G 12 , negative elements can be realized.
[0124] In addition, a memristor unit with a 2T2R structure as shown in Figure 2C can also be used to correspond to a weight element. The following uses Figure 8C , Figure 8D to illustrate an example of constructing a memory-computation integrated array using a memristor unit with a 2T2R structure.
[0125] Figure 8C shows a memory-computation integrated array constructed using a memristor unit with a 2T2R structure.
[0126] As Figure 8C shown, for example, a memristor unit with a 2T2R structure includes two memristors, namely memristor 801 and memristor 802. The conductance value of memristor 801 is denoted as G 11 , and the conductance value of memristor 802 is denoted as G 12 . Memristor 801 can be R1 in Figure 2C , and memristor 802 can be Figure 2C R2 in it. For example, since the memristor 802 is connected to an inverter, when the memristor 801 receives a positive-polarity input voltage signal, the inverter can invert the polarity of the input voltage signal, so that the memristor 802 receives a negative-polarity input voltage signal. For example, the input voltage signal received by the memristor 801 is represented by v(t), and the memristor 802 receives the inverted input voltage signal of v(t), that is, -v(t). The memristor 801 and the memristor 802 are connected to the same SL, and the output current passing through the memristor 801 and the output current passing through the memristor 802 are superimposed at the end of the SL. Therefore, the result of the multiply-accumulate calculation of the memristor 801 and the memristor 802 is v(t)G 11 + (-v(t))G 12 , that is, v(t)(G 11 -G 12 ). Therefore, the memristor unit with a 2T2R structure including the memristor 801 and the memristor 802 can correspond to a weight element, and this weight element is G 11 -G 12 , and negative elements can be realized by configuring the numerical relationship of G 11 -G 12 .
[0127] Figure 8D FIG. shows another memory-computation integrated array constructed by a memristor unit with a 2T2R structure.
[0128] As Figure 8D shown, for example, a memristor unit with a 2T2R structure includes two memristors, namely the memristor 801 and the memristor 802, and the conductance value of the memristor 801 is represented as G 11 , and the conductance value of the memristor 802 is represented as G 12 . Different from Figure 8C , the memristor 802 is not connected to an inverter, so when the memristor 801 receives a positive-polarity input voltage signal, the memristor 802 also receives a positive-polarity input voltage signal. For example, the input voltage signal received by the memristor 801 is represented by v(t), and the input voltage signal received by the memristor 802 is also represented by v(t). The memristor 801 and the memristor 802 are connected to different SLs, and the output current passing through the memristor 801 and the output current passing through the memristor 802 are subtracted at the end of the SL. Therefore, the result of the multiply-accumulate calculation of the memristor 801 and the memristor 802 is v(t)G 11 -v(t)G 12 , that is, v(t)(G 11 -G 12 ). Therefore, the memristor unit with a 2T2R structure including the memristor 801 and the memristor 802 can correspond to a weight element, and this weight element is G 11 -G 12 , negative elements can be achieved by configuring the numerical relationship of G 11 -G 12 .
[0129] For example, the in-memory computing array in the convolutional neural network processing device provided by at least one embodiment of the present disclosure may adopt any one of the structures provided in FIGS. 8A to 8D to achieve negative weight elements, and the present disclosure does not limit this.
[0130] For example, the above-mentioned convolutional neural network processing device 700 may be implemented by using hardware, software, firmware, and any feasible combination thereof, and the present disclosure does not limit this.
[0131] The technical effect of the convolutional neural network processing device 700 is the same as that of Figure 3 the convolutional neural network processing method shown, and will not be elaborated here.
[0132] Figure 9 FIG. shows a schematic diagram of a convolutional neural network processing device according to at least one embodiment of the present disclosure.
[0133] As Figure 9 shown, the convolutional neural network processing device is used to implement the convolutional neural network processing method according to at least one embodiment of the present disclosure. The convolutional neural network processing device includes an in-memory computing array and peripheral circuit devices, and the peripheral circuit devices include a digital-to-analog converter (DAC) and a multiplexer (MUX).
[0134] The in-memory computing array includes memristors with a 2T2R structure arranged in M rows and N columns. The memristor with a 2T2R structure includes two transistors and two memristors G P and G N . The gate of the P-type transistor is connected to the word line terminal WL P , and the gate of the N-type transistor is connected to the word line terminal WL N ; the first pole of the P-type transistor may be the source and is configured to be connected to the source line terminal SL, the first pole of the N-type transistor may be the source and is configured to be connected to the source line terminal SL, the first pole of the P-type transistor is connected to the first pole of the N-type transistor and is connected to the source line terminal SL together. The second pole of the P-type transistor may be the drain and is configured to be connected to the second pole (e.g., the negative pole) of the memristor G P , the first pole (e.g., the positive pole) of the memristor G P is connected to the bit line terminal BL P , for example, the memristor G P can receive a set voltage through the bit line terminal G P ; the second pole of the N-type transistor may be the drain and is configured to be connected to the second pole (e.g., the negative pole) of the memristor G N , the second pole of the memristor G N The first pole (e.g., the positive pole) and the bit line terminal BL N are connected. For example, the memristor R2 can receive the set voltage through the bit line terminal BL2.
[0135] The memory - in - computing array further includes 2M word lines, 2M source lines, and 2N bit lines. Each word line, each source line, and each bit line are respectively connected to a corresponding multiplexer. The multiplexers for the 2M word lines and the 2N bit lines are controlled by the signals WL_sw[1:2M] and BL_sw[1:2N], respectively. The input terminals of the 2M multiplexers connected to the 2M word lines are connected to the input voltage signal V_WL[1:2M]. For example, when the voltage signal V_WL[1:2M] is applied to the word lines through the multiplexers, the memristors corresponding to the word lines to which the voltage signal is applied are turned on. One input terminal of the 2N multiplexers connected to the 2N bit lines is connected to the ground wire (GND), and the other input terminal is connected to a DAC. The DAC converts the digital signal into a bit line voltage, and the bit line voltage is applied to the bit lines through the multiplexers. The input terminals of the 2M multiplexers connected to the 2M source lines are connected to a DAC. The DAC converts the digital signal into a source line voltage, and the source line voltage is applied to the source lines through the multiplexers. For the present disclosure, the following points need to be noted:
[0136] (1) The drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure. Other structures can refer to the general design.
[0137] (2) Without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.
[0138] The above - mentioned is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure should be subject to the protection scope of the claims.< / m> < / n> < / m>
Claims
1. A convolutional neural network processing method based on a processing-in-memory array, the processing-in-memory array including multiple rows and multiple columns of operation units, the method comprises: Based on the width x_in, height y_in, and number of input channels C_in of the input feature map, as well as the width x_k, height y_k, and number of channels C_in of the convolutional kernel, set the width of the sub-input feature map of the input feature map to be x, the height to be y, and the number of channels to be z, and set the width of the sub-convolutional kernel of the convolutional kernel to be x_k, the height to be y_k, and the number of channels to be z, where x_k is less than or equal to x, y_k is less than or equal to y, x is less than or equal to x_in, y is less than or equal to y_in, z is less than or equal to C_in, and the result of x×y×z is less than or equal to the number of rows of the processing-in-memory array; Based on the situation where the sub-convolutional kernel slides in the sub-input feature map with a stride S, deploy the sub-convolutional kernel into N columns of at least one processing-in-memory array in sequence; Input the sub-input feature map into the at least one processing-in-memory array as a row input signal, and obtain a calculation result from the column output signals of the N columns of the at least one processing-in-memory array.
2. The convolutional neural network processing method according to claim 1, wherein, For the case where z is less than or equal to C_in, the number of the at least one memory - in - computing array is 3. The convolutional neural network processing method according to claim 1, wherein, The step of deploying the sub-convolutional kernel into N columns of at least one processing-in-memory array in sequence based on the situation where the sub-convolutional kernel slides in the sub-input feature map with a stride S includes: According to the situation where the sub-convolutional kernel slides in the sub-input feature map with the stride S, divide the position of the sub-convolutional kernel relative to the sub-input feature map into a first convolution position, a second convolution position, and a third convolution position; and According to the first convolution position, the second convolution position, and the third convolution position, deploy the sub-convolutional kernel into the N columns of the at least one processing-in-memory array in sequence, where the first convolution position corresponds to the situation where the sub-convolutional kernel is completely within the sub-input feature map during the sliding process, the second convolution position corresponds to the situation where the sub-convolutional kernel is between two sub-input feature maps during the sliding process, the third convolution position corresponds to the situation where the sub-convolutional kernel is between four sub-input feature maps during the sliding process.
4. The convolutional neural network processing method according to claim 3, wherein, The step of deploying the sub-convolutional kernel into the N columns of the at least one processing-in-memory array in sequence according to the first convolution position, the second convolution position, and the third convolution position includes: Deploy the sub-convolutional kernel into n1 columns of the at least one processing-in-memory array according to the first convolution position, Deploy the sub-convolutional kernel into n2 columns of the at least one processing-in-memory array according to the second convolution position, Deploy the sub-convolutional kernel into n3 columns of the at least one processing-in-memory array according to the third convolution position, where one sub-convolutional kernel is deployed in each column of the processing-in-memory array, and n1 + n2 + n3 = N.
5. The convolutional neural network processing method according to claim 4, wherein, in the case where x_k and y_k of the sub-convolution kernel are the same, n3 = C_out × 4 × (x_k - 1) 2 ; or in the case where x_k and y_k of the sub-convolution kernel are different, n3 = Cout × 4 × (x_k - 1) × (y_k - 1).
6. The convolutional neural network processing method according to claim 5, wherein, in the case where x_k and y_k of the sub-convolution kernel are the same, the number of weight elements corresponding to the sub-convolution kernel at the first convolution position is x_k × y_k × z, the number of weight elements corresponding to the sub-convolution kernel at the second convolution position is x_k × z, the number of weight elements corresponding to the sub-convolution kernel at the third convolution position is z; or in the case where x_k and y_k of the sub-convolution kernel are different, the number of weight elements corresponding to the sub-convolution kernel at the first convolution position is x_k × y_k × z, the number of weight elements corresponding to the sub-convolution kernel at the third convolution position is z, on half of the n2 columns, the number of weight elements corresponding to the sub-convolution kernel at the second convolution position is x_k × z, on the other half of the n2 columns, the number of weight elements corresponding to the sub-convolution kernel at the second convolution position is y_k × z.
7. The convolutional neural network processing method according to claim 1, wherein, the step of inputting the sub-input feature map as a row input signal into the at least one memory-compute integrated array and obtaining a calculation result from the column output signals of the N columns of the at least one memory-compute integrated array includes: gradually sliding from the input feature map in a sliding manner with a sliding step of x in the width dimension and a sliding step of y in the height dimension to obtain the sub-input feature map, and in each step, inputting the sub-input feature map as a row input signal into the at least one memory-compute integrated array and obtaining the calculation result corresponding to this step from the column output signals of the N columns of the at least one memory-compute integrated array.
8. The convolutional neural network processing method according to any one of claims 1 - 7, wherein, the convolutional neural network involves K such convolution kernels, and K is equal to the number of output channels C_out; based on the case where the sub-convolution kernel slides in the sub-input feature map with a step size of S, the step of sequentially deploying the sub-convolution kernel into the N columns of the at least one memory-compute integrated array includes: based on the case where the sub-convolution kernel slides in the sub-input feature map with the step size of S, sequentially deploying each of the K sub-convolution kernels into the N columns of at least one memory-compute integrated array, thereby sequentially deploying the K sub-convolution kernels into the N columns of the at least one memory-compute integrated array.
9. A convolutional neural network processing device based on a memory-compute integrated array, comprising: at least one memory-compute integrated array, including multiple rows and multiple columns of operation units; A setting unit, configured to set the width of the sub-input feature map to x, the height to y, and the number of channels to z based on the width x_in, height y_in, and number of channels C_in of the input feature map, and the width x_k, height y_k, and number of channels C_in of the convolutional kernel, and set the width of the sub-convolutional kernel to x_k, the height to y_k, and the number of channels to z, where x_k is less than or equal to x, y_k is less than or equal to y, x is less than or equal to x_in, y is less than or equal to y_in, z is less than or equal to C_in, and the result of x×y×z is less than or equal to the number of rows of the memory-compute integrated array; A deployment unit, configured to sequentially deploy the sub-convolutional kernel to N columns of the at least one memory-compute integrated array based on the situation where the sub-convolutional kernel slides in the sub-input feature map with a stride S; and A control unit, configured to input the sub-input feature map into the at least one memory-compute integrated array as a row input signal, and obtain a calculation result from the column output signals of N columns of the at least one memory-compute integrated array.
10. The convolutional neural network processing device according to claim 9, wherein, the deployment unit is further configured to: divide the position of the sub-convolutional kernel relative to the sub-input feature map into a first convolution position, a second convolution position, and a third convolution position according to the situation where the sub-convolutional kernel slides in the sub-input feature map with a stride S; and sequentially deploy the sub-convolutional kernel to the N columns of the at least one memory-compute integrated array according to the first convolution position, the second convolution position, and the third convolution position, wherein the first convolution position corresponds to the situation where the sub-convolutional kernel is completely within the sub-input feature map during the sliding process, the second convolution position corresponds to the situation where the sub-convolutional kernel is between two sub-input feature maps during the sliding process, the third convolution position corresponds to the situation where the sub-convolutional kernel is between four sub-input feature maps during the sliding process.
11. The convolutional neural network processing device according to claim 10, wherein, the deployment unit is further configured to: deploy the sub-convolutional kernel to n1 columns of the at least one memory-compute integrated array according to the first convolution position, deploy the sub-convolutional kernel to n2 columns of the at least one memory-compute integrated array according to the second convolution position, deploy the sub-convolutional kernel to n3 columns of the at least one memory-compute integrated array according to the third convolution position, where one sub-convolutional kernel is deployed in each column of the memory-compute integrated array, and n1 + n2 + n3 = N.
12. The convolutional neural network processing device according to claim 9, wherein, the control unit is further configured to: gradually slide from the input feature map to obtain the sub-input feature map in a sliding manner with a sliding stride of x in the width dimension and a sliding stride of y in the height dimension, input the sub-input feature map into the at least one memory-compute integrated array as a row input signal in each step, and obtain the calculation result corresponding to the step from the column output signals of N columns of the at least one memory-compute integrated array.
13. The convolutional neural network processing device according to claim 9, wherein, the convolutional neural network involves K convolutional kernels, and K is equal to the number of output channels C_out; the deployment unit is further configured to: based on the situation that each of the K sub-convolutional kernels slides in the sub-input feature map with a stride S, deploy each of the K sub-convolutional kernels into N columns of at least one memory-computation integrated array in sequence, thereby deploying the K sub-convolutional kernels into N columns of at least one memory-computation integrated array in sequence.
Citation Information
Patent Citations
Data processing method based on memristor array and electronic device
CN113077829A
Pulse convolutional neural network algorithm, integrated circuit, computing apparatus, and storage medium
WO2021115262A1