Three-dimensional memristor array circuit, control system and multi-chip collaborative neural network calculation method
By building a three-dimensional vertical memristor array circuit and multi-chip collaborative neural network calculation method on the chip, the calculation requirements of integrated scale and reliability limitations in the existing technology are solved, and efficient neural network computing is achieved.
Patent Information
- Application Number
- CN202411702542.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-26
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, due to the integration scale and reliability of the memristor array, it is difficult to integrate a large amount of computing power on a chip, and it is impossible to effectively solve the problem of a large number of computing needs.
A three-dimensional vertical memristor array circuit is adopted to form a three-dimensional array structure through stacking of multi-layer two-dimensional memristor arrays, supporting the calculation of multi-channel convolutional neural network feature information, and collaborating with multiple chips for neural network calculation.
It achieves higher storage density and parallelism, improves the efficiency of neural network computing, can effectively handle large-scale neural network computing needs, and reduces the energy consumption and delay of data transmission.
Smart Images

Figure CN119993233A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of neural network accelerators, and specifically relates to a three-dimensional memristor array circuit, a control system and a multi-chip collaborative neural network calculation method. Background Art
[0002] With the development of applications such as deep learning and big data analysis, the "memory wall" problem has become one of the bottlenecks in computer systems. Traditional computing methods based on the von Neumann architecture require frequent data transfer between the central processor and the memory, resulting in increased computing delays and energy consumption. A memristor is a resistor with memory function that can change its resistance value according to changes in voltage or current.
[0003] Memristor-based on-chip neural network computing methods have been verified and applied in a variety of application scenarios, such as target detection, semantic segmentation, image classification, etc. The operations in these applications often have certain characteristics, such as sparsity and nonlinearity, which can match the characteristics of memristors to improve the accuracy and robustness of the calculations. Memristor-based on-chip computing methods can also be co-designed and optimized with other hardware devices, such as FPGAs and ASICs, to further improve the performance of the system. Existing memristor in-memory computing mainly relies on a two-dimensional memristor cross array to map data to the conductance value of the memristor, and then based on Ohm's law and Kirchhoff's law, the function of completing neural network operations in one step is realized. This method can make full use of the parallelism of the memristor array, greatly improve computing performance, and reduce the overhead of data transmission.
[0004] Compared with the two-dimensional memristor cross array, the three-dimensional memristor array saves on-chip space, has a higher storage density and richer device connection relationships, and can provide a more flexible matching method for the construction of state logic gates. However, in the face of a large number of computing needs, limited by the integration scale and reliability of the memristor array, the existing integrated circuit manufacturing technology is often difficult to integrate a large amount of computing power on a chip. In order to further improve computing efficiency, it is necessary to reasonably schedule computing resources and coordinate multiple chips for neural network computing. Summary of the invention
[0005] The technical problem to be solved by the present invention is to provide a three-dimensional memristor array circuit, a control system and a multi-chip collaborative neural network calculation method, which solves the problem that a large number of computing requirements in the prior art are limited by the integration scale and reliability of the memristor array.
[0006] The present invention adopts the following technical solutions to solve the above technical problems:
[0007] A three-dimensional vertical memristor array circuit includes a three-dimensional array structure composed of multiple layers of two-dimensional memristor arrays stacked together, the three-dimensional array structure having column input electrodes of multi-channel convolutional neural network characteristic information and corresponding current output electrodes; a memristor is formed at the intersection between the column input electrode and the output electrode.
[0008] The conductivity between the column input and output electrodes is used as the weight of the high-density kernel in the convolutional layer.
[0009] The two-dimensional computing layers are physically isolated from each other, and the input electrodes and output electrodes are arranged non-orthogonally.
[0010] The output current of the mth column of current output electrodes is calculated according to the following formula:
[0011]
[0012] Among them, n is the index layer, N is the total number of layers, V in is the voltage vector applied to the input terminal, and G is the conductivity matrix of the memristor array.
[0013] It includes a PC host computer, a PE module, a Tile module, and a PCIE controller; wherein the PE module includes a three-dimensional vertical memristor array circuit and its peripheral circuits, which are used to perform multi-channel convolution calculations; the Tile module is used to integrate computing resources and storage resources, and control the computing flow based on instruction information; the PCIE controller is connected to the PC host computer through an APB bus to achieve data transmission between the Tile module and the PC host computer; the PC host computer is used to combine computing resources to make corresponding arrangements for computing tasks, and send instructions to the PCIE controller.
[0014] The PE module receives the calculation results of the corresponding channels, performs post-processing operations on the received results including channel alignment, addition, and activation, and outputs the final calculation results to the cache end.
[0015] The Tile module supports DRAM-based data storage and memristor-based arithmetic logic operations, and processes network model parameters by accelerating the calculation process through multi-stage pipeline parallelism.
[0016] The PCIE controller provides DMA function for batch transmission of asynchronous data. After the PCIE receives the information, the Tile module starts the corresponding calculation function; the upstream device directly reads and writes the data space specified by the PCIE controller in the form of packets.
[0017] The multi-chip collaborative neural network computing method based on three-dimensional memristor includes the following steps:
[0018] Step 1: The column input electrode receives the input signal sent by the host computer and transmits it to the memristor array;
[0019] Step 2: The memristor array performs a weighted accumulation operation on the input signal through an adjustable weight to obtain a weighted current;
[0020] Step 3, the row output electrode receives the weighted current flowing from the memristor array;
[0021] Step 4: Input the current signal of the row output electrode into the neuron excitation function module, process the weighted sum through the nonlinear excitation function, and obtain the excitation signal of the next layer;
[0022] Step 5: The processed output signal is output to the host computer through the current output electrode.
[0023] The input data is an encoded neural network excitation signal.
[0024] Compared with the prior art, the present invention has the following beneficial effects:
[0025] 1. The present invention can realize the collaborative work of multiple memristors. Multiple three-dimensional memristor arrays can perform collaborative computing through an interconnected structure. Each memristor array is responsible for different layers or parts of computing tasks in the neural network, and the computing efficiency is improved through parallel and serial collaboration.
[0026] 2. The host computer distributes input signals to different memristor array slices through a special data distribution and scheduling mechanism. Each array slice independently completes local matrix operations according to its task assignment, and transmits the calculation results to other array slices through the internal interconnection structure to achieve large-scale neural network calculations; each memristor array can process input data asynchronously, and after the calculation is completed, the local calculation results are aggregated to specific output nodes through current output, thereby generating the final neural network inference results.
[0027] 3. During the training process, the conductance value of the memristor can be adaptively adjusted through the learning algorithm to update the weights of the neural network. Multiple array slices can update the weights of their internal memristors in parallel to speed up the training.
[0028] 4. By utilizing the integrated storage and computing characteristics of memristors, the energy consumption and delay of data transmission between storage and computing units are reduced, and the computing density is improved, which is particularly suitable for the computing needs of large-scale deep neural networks. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 Schematic diagram of a two-dimensional cross array of memristors according to the present invention.
[0030] Figure 2 Schematic diagram of a three-dimensional cross array of memristors according to the present invention.
[0031] Figure 3This is a diagram of the processing unit architecture based on a three-dimensional memristor according to the present invention.
[0032] Figure 4 It is a schematic block diagram of the monolithic structure of the present invention.
[0033] Figure 5 It is a schematic block diagram of the overall structure of the present invention. DETAILED DESCRIPTION
[0034] The structure and working process of the present invention will be further described below in conjunction with the accompanying drawings.
[0035] The purpose of this scheme is to propose a configuration method of on-chip three-dimensional memristor and related circuits in view of the limitations and deficiencies in the prior art. First, a three-dimensional memristor array is built as the computing power basis for convolution calculation. The PE module supports multi-channel convolution calculation, while the Tile module supports DRAM-based data storage and memristor-based arithmetic and logic operations, and accelerates the calculation process through multi-stage pipeline parallelization. Through PCIE connection with the PC host computer, information transmission, multi-chip calculation and calculation task allocation are realized, so that the overall system can support high-speed reasoning calculation on neural network chip.
[0036] The innovation of the present invention is
[0037] (I) By building a three-dimensional memristor cross array, it supports the input of multi-channel convolutional neural network feature information. This structure not only improves the storage density, but also increases the degree of parallelism, thereby effectively accelerating the calculation process of the neural network.
[0038] (ii) Within each 3D computing group, precise control of the programming and reading of the memristors is achieved through the optimized design of peripheral circuits, such as selectors, amplifiers, and analog-to-digital converters, thereby enhancing the computing accuracy and stability of the system.
[0039] (III) By building system monomer modules, including PE modules and Tile modules, a modular design of multi-channel convolution computing is realized. This design makes the system easier to maintain and expand, and can be integrated with multiple modules to realize multi-chip collaborative computing, improving the flexibility and parallelism of the overall system.
[0040] A three-dimensional vertical memristor array circuit includes a three-dimensional array structure composed of multiple layers of two-dimensional memristor arrays stacked together, the three-dimensional array structure having column input electrodes of multi-channel convolutional neural network characteristic information and corresponding current output electrodes; a memristor is formed at the intersection between the column input electrode and the output electrode.
[0041] Specific embodiment 1, as Figures 1 to 5 As shown,
[0042] A three-dimensional vertical memristor array circuit, firstly, build a two-dimensional cross array of memristors.
[0043] Figure 1 Shown is a two-dimensional crossbar array of memristors, a nonvolatile device with a changeable resistance whose conductance represents the elements of the matrix.
[0044] Memristor crossbar array is a method to implement matrix multiplication using the analog properties of memristors and Ohm's law, where the synaptic weight is represented by the conductance of the memristor at each intersection connecting the word line and the orthogonal bit line. By mapping the rows and columns of the two matrices to the input and output of the memristor array, respectively, the relationship between voltage and current can be used to calculate the result of matrix multiplication. Specifically, assuming there are two matrices A and B, whose product is C, then the transposed matrix of A can be mapped to the input of the memristor array, and B can be mapped to the output of the memristor array. Then a voltage vector V is applied to the input, and a current vector I will be generated at the output, each element of which is the corresponding element of C. This process can be expressed by the following formula:
[0045]
[0046] Where G is the conductance matrix of the memristor array, G A and G B It is the conductivity matrix after A and B are mapped to the memristor array. Through the electrical characteristics of the memristor, the array can easily perform large-scale matrix multiplication, which can be used to accelerate the operation of neural networks. CMOS memristor hybrid technology can achieve fully parallel programming operations that traditional vertical and horizontal arrays cannot achieve. The memristor calculation process accumulates the calculation results from the crossbar switch and generates the output using an analog-to-digital converter (ADC).
[0047] Then, build a three-dimensional memristor cross array. Figure 2 The figure shows the construction of a three-dimensional memristor cross array, which is composed of multiple layers of two-dimensional memristor arrays stacked together. The main view is also shown in the figure. It supports the input of multi-channel convolutional neural network feature information.
[0048] Most systems are now built on regular two-dimensional systems. Considering the needs of multi-channel computing of convolutional neural networks, a three-dimensional memristor cross array is built to support the input of multi-channel convolutional neural network feature information. In order to save on-chip area, the parallel computing kernel is compiled into three dimensions. By stacking multiple cross arrays, the storage density and parallelism are increased, thereby accelerating the calculation process of the neural network. Specifically, the output current of the mth column is the dot product between the output voltage from the mth column to the m+N-1th column and the memristor conductance, and the formula is expressed as follows,
[0049]
[0050] Among them, n is the index layer, N is the total number of layers, V in is the voltage vector applied to the input terminal, and G is the conductivity matrix of the memristor array.
[0051] HfO is used in the array x -Al 2 O 3 Memristors improve the reliability and uniform switching characteristics of the system; filters, self-rectifier devices and other components are added to the circuit to reduce parasitic effects, maintain normal MAC calculations, and improve the calculation accuracy and stability of the memristor array. Within each 3D computing group, memristors are formed at the intersection between the column input electrode and the 3D output electrode, and their conductivity is used as the weight of the high-density kernel in the convolution layer. The 2D computing layers are physically isolated from each other, and the input electrodes and output electrodes are arranged non-orthogonally. Build corresponding memristor peripheral circuits such as selectors, amplifiers, analog-to-digital converters, etc., which are used to control the programming and reading of memristors, as well as the conversion and amplification of signals. Bidirectional data communication is achieved between 2D computing layers through peripheral circuits.
[0052] A three-dimensional vertical memristor array control system includes a PC host computer, a PE module, a Tile module, and a PCIE controller; wherein the PE module includes a three-dimensional vertical memristor array circuit and its peripheral circuits, and is used for performing multi-channel convolution calculations; the Tile module is used for integrating computing resources and storage resources, and controlling the computing flow based on instruction information; the PCIE controller is connected to the PC host computer through an APB bus to achieve data transmission with the Tile module; the PC host computer is used for correspondingly allocating computing tasks in combination with computing resources, and sending instructions to the PCIE controller.
[0053] Specific embodiment 2:
[0054] The three-dimensional vertical memristor array control system is divided into system single module construction and overall module construction, among which,
[0055] 1. System single module construction
[0056] Figure 3 The architecture diagram of the processing unit based on three-dimensional memristors is shown. The PE module that builds the calculation foundation is used to perform multi-channel convolution calculations. After the device is initialized, the weights exist in the memristor array, and the corresponding feature map data is input. After being split and cached by the pre-processing module, it is sent to the three-dimensional memristor array for multi-channel matrix multiplication calculations. The remaining PE processing modules input the corresponding channel calculation results, and after receiving the results, they perform post-processing operations such as channel alignment, addition, activation, and output the final calculation results to the cache end.
[0057] Figure 4The schematic block diagram of the monolithic structure of the present invention is shown. The overall process consists of operations such as instruction fetching, decoding, calculation, and storage. In order to improve hardware utilization and achieve more flexible instruction scheduling, a multi-level pipeline is built for parallel acceleration, and structures such as data bypass and delay slots are constructed to improve the execution efficiency of the pipeline and reduce the occurrence of data hazards and control hazards. On the basis of the PE module, a Tile module is built to integrate computing resources and storage resources, and the computing flow is controlled based on instruction information. The PC host compiles the network model to obtain relevant parameter information, encodes it, and sends the corresponding instruction information to the instruction storage unit. The accelerator reads the corresponding instruction information, obtains the relevant network operation parameters after the compiled network model after the decoding operation, and inputs them to the control unit. The control unit collaboratively controls the DRAM-based data storage unit to perform storage operations and the arithmetic logic PE unit based on the three-dimensional memristor to perform calculation operations.
[0058] 2. System overall module construction
[0059] Figure 5 The schematic block diagram of the overall architecture of the present invention is shown. An overall system of multi-chip coordinated computing is built on the basis of a single module. The system is connected to the PC host computer through PCIE, receives corresponding instructions, and the PC host computer system performs corresponding deployment of computing tasks in combination with computing resources. The chip realizes information transmission through AXI_Stream, and configures related PCIE information through the APB bus. The host computer sends data to the base address register Bar of the PCIE data space, and the upstream device can directly read and write the specified data space in the form of packet sending. The PCIE controller also provides DMA function for batch asynchronous data transmission. After the PCIE receives the information, the Tile module starts the corresponding computing function.
[0060] The multi-chip collaborative neural network computing method based on three-dimensional memristor includes the following steps:
[0061] Step 1: The column input electrode receives the input signal sent by the host computer and transmits it to the memristor array;
[0062] Step 2: The memristor array performs a weighted accumulation operation on the input signal through an adjustable weight to obtain a weighted current;
[0063] Step 3, the row output electrode receives the weighted current flowing from the memristor array;
[0064] Step 4: Input the current signal of the row output electrode into the neuron excitation function module, process the weighted sum through the nonlinear excitation function, and obtain the excitation signal of the next layer;
[0065] Step 5: The processed output signal is output to the host computer through the current output electrode.
[0066] The input data is an encoded neural network excitation signal.
[0067] Specific embodiment three,
[0068] The multi-chip collaborative neural network computing method based on three-dimensional memristor includes the following steps:
[0069] Step 1: The column input electrode receives the input signal sent by the host computer;
[0070] The host computer transmits the pre-processed neural network input data to the column input electrodes of the three-dimensional memristor array. The input data is the encoded neural network excitation signal, which will be weighted through the conduction characteristics of the memristor.
[0071] Step 2: The memristor array performs matrix multiplication operations;
[0072] The input signal is transmitted to the memristor array and weighted accumulation is performed through the adjustable weights of the memristors. The conductance value of the memristor represents the weight of the neural network, so when the input signal passes through the memristor array, a simulated matrix-vector multiplication operation is completed.
[0073] Step 3: The row output electrodes obtain weighted current;
[0074] After the weighted accumulation operation is completed, the row output electrodes receive the weighted current flowing out of the memristor array. These current values represent the matrix multiplication results, that is, the weighted sum of neurons.
[0075] Step 4: neuron activation function processing;
[0076] The current signal of the row output electrode will be input into the neuron excitation function module, and the weighted sum will be processed by a nonlinear excitation function (such as ReLU, Sigmoid, etc.) to obtain the excitation signal of the next layer.
[0077] Step 5: Output the calculation results;
[0078] The final processed output signal represents the calculation result of the neural network and is transmitted back to the host computer for further processing or decision-making.
[0079] In summary, this scheme supports high-speed reasoning calculation on neural network chip. First, a three-dimensional memristor array is built to support the calculation of multi-channel convolutional neural network feature information. These arrays use the analog characteristics of memristors and Ohm's law to perform matrix multiplication, map the calculated matrix to the input and output ends of the memristor array, and calculate the product through the relationship between voltage and current. Secondly, a system monomer module is constructed, including a PE module and a Tile module, for multi-channel convolution calculation and network model parameter processing. The PE module can perform multi-channel convolution calculation, while the Tile module supports DRAM-based data storage and memristor-based arithmetic and logic operations, and accelerates the calculation process through multi-level pipeline parallelization. Finally, the overall system module is built to realize information transmission, multi-chip calculation and calculation task allocation. This method not only improves the speed of neural network calculation, but also has a flexible system integration method. It can be widely used in the field of artificial intelligence and provides a new solution for high-speed neural network calculation.
[0080] This solution has the following advantages:
[0081] 1. Multiple memristors work together:
[0082] Multiple three-dimensional memristor arrays perform collaborative computing through an interconnected structure. Each memristor array is responsible for different layers or parts of computing tasks in the neural network, improving computing efficiency through parallel and serial collaboration.
[0083] 2. Data distribution and scheduling mechanism:
[0084] The host computer distributes the input signals to different memristor array slices through a special data distribution and scheduling mechanism. Each array slice completes local matrix operations independently according to its task allocation, and transmits the calculation results to other array slices through the internal interconnection structure to realize large-scale neural network calculation.
[0085] 3. Asynchronous processing and result aggregation:
[0086] Each memristor array can process input data asynchronously, and after the calculation is completed, the local calculation results are aggregated to specific output nodes through current output, thereby generating the final neural network inference results.
[0087] 4. Weight update and adaptive adjustment:
[0088] During the training process, the conductance of the memristor can be adaptively adjusted through the learning algorithm to update the weights of the neural network. Multiple array slices can update the weights of their internal memristors in parallel to speed up the training.
[0089] 5. Combination of high-density storage and computing:
[0090] By utilizing the integrated storage and computing characteristics of memristors, the energy consumption and delay of data transmission between storage and computing units are reduced, and the computing density is improved, which is particularly suitable for the computing needs of large-scale deep neural networks.
[0091] It should be understood that the present solution is not limited to the above-mentioned specific implementation methods, and the devices and structures not described in detail should be understood to be implemented in a common manner in the art; any technician familiar with the art can use the above-disclosed methods and technical contents to make many possible changes and modifications to the technical solution of the present solution without departing from the scope of the technical solution of the present solution, or modify it into an equivalent embodiment with equivalent changes, which does not affect the essential content of the present solution. Therefore, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present solution without departing from the content of the technical solution of the present solution still falls within the scope of protection of the technical solution of the present solution.
Claims
1. A three-dimensional vertical memristor array circuit, characterized in that: It includes a three-dimensional array structure composed of a stack of multiple layers of two-dimensional memristor arrays, the three-dimensional array structure has a column input electrode with multi-channel convolutional neural network feature information, and a corresponding current output electrode; a memristor is formed at the intersection between the column input electrode and the output electrode.
2. The three-dimensional vertical memristor array circuit according to claim 1, characterized in that: The conductivity between the column input and output electrodes is used as the weight of the high-density kernel in the convolutional layer.
3. The three-dimensional vertical memristor array circuit according to claim 2, characterized in that: The two-dimensional computing layers are physically isolated from each other, and the input electrodes and output electrodes are arranged non-orthogonally.
4. The three-dimensional vertical memristor array circuit according to claim 1, characterized in that: The current output electrode The output current of the column is calculated according to the following formula: in, is the index layer, is the total number of layers, is the voltage vector applied to the input terminal, is the conductance matrix of the memristor array.
5. A three-dimensional vertical memristor array control system, characterized in that: It includes a PC host computer, a PE module, a Tile module, and a PCIE controller; wherein the PE module includes a three-dimensional vertical memristor array circuit and its peripheral circuits, which are used to perform multi-channel convolution calculations; the Tile module is used to integrate computing resources and storage resources, and control the computing flow based on instruction information; the PCIE controller is connected to the PC host computer through an APB bus to achieve data transmission between the Tile module and the PC host computer; the PC host computer is used to combine computing resources to make corresponding arrangements for computing tasks, and send instructions to the PCIE controller.
6. The three-dimensional vertical memristor array control system according to claim 5, characterized in that: The PE module receives the calculation results of the corresponding channels, performs post-processing operations on the received results including channel alignment, addition, and activation, and outputs the final calculation results to the cache end.
7. The three-dimensional vertical memristor array control system according to claim 5, characterized in that: The Tile module supports DRAM-based data storage and memristor-based arithmetic logic operations, and processes network model parameters by accelerating the calculation process through multi-stage pipeline parallelism.
8. The three-dimensional vertical memristor array control system according to claim 5, characterized in that: The PCIE controller provides DMA function for batch transmission of asynchronous data. After the PCIE receives the information, the Tile module starts the corresponding calculation function; the upstream device directly reads and writes the data space specified by the PCIE controller in the form of packets.
9. A multi-chip collaborative neural network computing method based on three-dimensional memristors, characterized in that: The steps include: Step 1: The column input electrode receives the input signal sent by the host computer and transmits it to the memristor array; Step 2: The memristor array performs a weighted accumulation operation on the input signal through an adjustable weight to obtain a weighted current; Step 3, the row output electrode receives the weighted current flowing from the memristor array; Step 4: Input the current signal of the row output electrode into the neuron excitation function module, process the weighted sum through the nonlinear excitation function, and obtain the excitation signal of the next layer; Step 5: The processed output signal is output to the host computer through the current output electrode.
10. The multi-chip collaborative neural network computing method based on three-dimensional memristors according to claim 9, characterized in that: The input data is an encoded neural network excitation signal.