RRAM-based neural network acceleration calculation method and system
By obtaining the hardware parameters and configuration parameters of RRAM, splitting the neural network computing tasks and dynamically adjusting the processing path, the calculation accuracy and efficiency problems caused by RRAM hardware instability are solved, and the efficiency and accuracy of neural network computing are achieved.
Patent Information
- Application Number
- CN202510821671.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-18
AI Technical Summary
Different RRAMs have hardware instability in neural network accelerated computing, resulting in reduced calculation accuracy and efficiency, especially in the conductance drift caused by differences in multi-board systems and storage structures that affect the calculation results.
By obtaining the hardware parameters and configuration parameters of RRAM, including the number of parallelizable and effective conductance ranges, the neural network computing tasks are split into linear and nonlinear computing tasks, and dynamically adjusting the processing path based on these parameters, using the high parallelism of RRAM and the high memory bandwidth of the processor for calculations, and dynamically adjusting the processing path to balance accuracy and power consumption.
It improves the accuracy and efficiency of neural network computing, reduces errors caused by RRAM hardware instability, and improves the reliability and processing speed of calculation results.
Smart Images

Figure CN120336036A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of neural network computing technology, and in particular, to a method and system for accelerating neural network computing based on RRAM. Background Art
[0002] With the surge in the computing power demand for neural network computing, traditional hardware such as GPUs has bottlenecks in energy efficiency and parallelism. Currently, RRAM is often used for neural network computing due to its advantages of high speed, low power consumption, and high parallelism.
[0003] In practical applications, when using multiple RRAMs for accelerating neural network computing, the changes in the physical characteristics of different RRAMs will reduce the accuracy of neural network computing, and the hardware configurations of different RRAMs will complicate the deployment of neural networks, thus reducing the computing efficiency of neural networks.
[0004] In view of this, how to ensure the accuracy and computing efficiency of different RRAMs in neural network acceleration computing is an urgent problem to be solved. Summary of the Invention
[0005] This application provides a method and system for accelerating neural network computing based on RRAM, which can ensure the accuracy and computing efficiency of different RRAMs in neural network acceleration computing.
[0006] In a first aspect of this application, a method for accelerating neural network computing based on RRAM is provided. The method includes: obtaining hardware parameters of one or more RRAMs, where the hardware parameters at least include the number of parallelisms, and determining configuration parameters of each of the RRAMs according to a preset tolerance, where the configuration parameters at least include an effective conductance range and a plurality of effective rows and columns; splitting the computing task of any network layer of the neural network, and generating a first computing task and a second computing task based on the splitting result; determining a first processing path of the first computing task based on the number of parallelisms and the effective conductance range, where the first processing path is connected to a target RRAM, and processing the first computing task according to the effective rows and columns of the target RRAM to obtain a first processing result; determining a second processing path of the second computing task, and processing the second computing task according to the second processing path and the first processing result to obtain a second processing result, and using the second processing result as the output result of the current network layer.
[0007] In a possible implementation manner, after obtaining the second processing result, calculating the first computing task and the second computing task according to the first processing path to obtain a third processing result; evaluating the performance parameters of the second processing result and the third processing result, and adjusting the target RRAM or the first processing path based on the evaluation result.
[0008] In a possible implementation, the performance parameters at least include the calculation accuracy and the calculation power consumption; evaluating the performance parameters of the second processing result and the third processing result, and adjusting the target RRAM or the first processing path based on the evaluation result includes: if the calculation accuracy of the second processing result is lower than that of the third processing result, adjusting the hardware parameters of the target RRAM; if the calculation power consumption of the second processing result is greater than that of the third processing result, determining the second processing path as the first processing path.
[0009] In a possible implementation, determining the configuration parameters of each of the RRAMs according to a preset tolerance includes: the RRAM includes a plurality of storage rows and columns, taking the storage rows and columns that meet the preset tolerance as valid rows and columns; obtaining the effective conductance values of the respective valid rows and columns, and determining the effective conductance range of the RRAM according to the effective conductance values.
[0010] In a possible implementation, splitting the calculation tasks of the neural network and generating a first calculation task and a second calculation task based on the splitting result includes: according to the type of the neural network, determining each calculation task of the neural network, and for any one of the calculation tasks, determining whether the calculation task belongs to a linear calculation task; if the calculation task belongs to a linear calculation task, determining the task as the first calculation task, and if the calculation task does not belong to a linear calculation task, determining the task as the second calculation task.
[0011] In a possible implementation, determining a first processing path of the first calculation task based on the parallelizable number and the effective conductance range, the first processing path being connected to a target RRAM includes: judging the RRAM that matches the first calculation task according to the parallelizable number and the effective conductance range; determining the matching RRAM as the target RRAM, and determining the path connecting the target RRAM and the neural network as the first processing path.
[0012] In a possible implementation, processing the first calculation task according to the effective rows and columns of the target RRAM to obtain a first processing result includes: determining a positive weight array and a negative weight array of the target RRAM based on the effective rows and columns; calculating the first calculation task according to the positive weight array and the negative weight array to obtain a first processing result.
[0013] In a possible implementation manner, determining the positive weight array and the negative weight array of the target RRAM based on the effective rows and columns includes: obtaining the effective conductance values of the respective effective rows and columns, determining a reference row and column in the effective rows and columns of the target RRAM, and determining the effective conductance value of the reference row and column as the reference value; determining the multiple effective rows and columns with the effective conductance value greater than or equal to the reference value as the positive weight array, and determining the multiple effective rows and columns with the effective conductance value less than the reference value as the negative weight array.
[0014] In a possible implementation manner, calculating the first calculation task according to the positive weight array and the negative weight array to obtain a first processing result includes: obtaining the positive weight matrix and the negative weight matrix of the neural network, storing the positive weight matrix in the positive weight array, storing the negative weight matrix in the negative weight array, obtaining the first output current of the positive weight array and the second output current of the negative weight array, and taking the difference between the first output current and the second output current as the first processing result.
[0015] In a possible implementation manner, the second processing path is connected to the processor; determining the second processing path of the second calculation task, and processing the second calculation task according to the second processing path and the first processing result to obtain a second processing result includes: determining the target processor connected to the neural network, determining the path connecting the target processor and the neural network as the second processing path, sending the second calculation task and the first processing result to the target processor through the first processing path, and the target processor executing the second calculation task according to the first processing result to obtain the second processing result.
[0016] The second aspect of the present application provides a system for accelerating neural network calculations based on RRAM. The system includes: a parameter acquisition unit for acquiring hardware parameters of one or more RRAMs, where the hardware parameters at least include the parallelism number, and determining configuration parameters of each RRAM according to a preset tolerance. The configuration parameters at least include an effective conductance range and a plurality of effective rows and columns; a task splitting unit for splitting the calculation task of any network layer of the neural network and generating a first calculation task and a second calculation task based on the splitting result; a task calculation unit for determining a first processing path of the first calculation task based on the parallelism number and the effective conductance range. The first processing path is connected to a target RRAM, processing the first calculation task according to the effective rows and columns of the target RRAM to obtain a first processing result, determining a second processing path of the second calculation task, and processing the second calculation task according to the second processing path and the first processing result to obtain a second processing result, and using the second processing result as the output result of the current network layer.
[0017] The third aspect of the present application provides a computer device, including a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor realizes a method for accelerating neural network calculations based on RRAM described in the first aspect by executing the computer instructions.
[0018] The technical solutions provided by one or more embodiments of the present application jointly process the neural network by automatically adapting different RRAMs and processors to improve the calculation efficiency and accuracy of the neural network. Specifically, according to the hardware parameters and configuration parameters of different RRAMs, automatically adapt the first processing path of different RRAMs, split and classify the calculation tasks of the neural network, and process each type of calculation task through the first processing path and the second processing path respectively to obtain the output result of the neural network. At the same time, dynamically adjust the adapted first processing path to balance the accuracy and power consumption of the calculation task processing, and improve the efficiency and accuracy of the neural network acceleration calculation.
[0019] It can be seen that the technical solutions provided by the present application ensure the calculation efficiency and accuracy of the neural network in the multi-configuration RRAM state. At the same time, automatically adapt different RRAMs and neural networks to solve the influence of the hardware configuration of RRAM on the neural network acceleration calculation. Description of the Drawings
[0020] To more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0021] Figure 1 Schematic diagram of steps of a method for accelerating neural network computing based on RRAM provided by an embodiment of the present application; Figure 2 Schematic diagram of steps of splitting a neural network provided by an embodiment of the present application; Figure 3 Schematic diagram of steps of determining positive and negative weight arrays provided by an embodiment of the present application; Figure 4 Schematic diagram of steps of obtaining a first processing result provided by an embodiment of the present application; Figure 5 Schematic diagram of the structure of a system for accelerating neural network computing based on RRAM provided by an embodiment of the present application; Figure 6 Schematic diagram of the structure of a computer device provided by an embodiment of the present application. Specific Embodiments
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present application.
[0023] In addition, the descriptions involving "first", "second", etc. in the present application are only for descriptive purposes and cannot be construed as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the embodiments of the present application, unless otherwise stated, the meaning of "a plurality" is two or more. Additionally, the use of "based on" or "according to" implies openness and inclusiveness, because a process, step, calculation, or other action "based on" or "according to" one or more of the stated conditions or values can in practice be based on additional conditions or values beyond those stated.
[0024] As the application of neural networks becomes more and more widespread, the requirement for computing power is also getting higher and higher. Generally speaking, traditional processors such as GPUs (Graphics Processing Unit) have limitations in energy efficiency and parallelism, making it difficult to meet the needs of neural networks. Currently, due to its advantages of high speed, low power consumption, and high parallelism, RRAM is often used instead of GPUs as a tool for accelerating neural network calculations. As a new type of non-volatile memory, RRAM can store data by utilizing the resistance change of materials and perform operations directly in the storage unit, thereby reducing data overhead.
[0025] However, in practical applications, there are hardware differences between different RRAMs. In particular, the differences in multi-board systems and storage structures complicate the deployment of neural networks, thus reducing the computing efficiency of neural networks. At the same time, affected by factors such as temperature, voltage fluctuations, and usage duration, the conductance value of RRAM will inevitably drift. Hardware instabilities such as conductance value drift will directly interfere with the accuracy of stored data, resulting in deviations in calculation results and reducing the accuracy of neural network calculations. For example, in complex image recognition tasks, subtle conductance value drift may lead to incorrect recognition results. Therefore, it is urgent to eliminate the impact of RRAM hardware instabilities on neural network accelerated computing and ensure the accuracy and computing efficiency of different RRAMs in neural network accelerated computing. In view of this, one or more embodiments of this application provide a method and system for accelerating neural network computing based on RRAM, which can solve the above problems and improve the accuracy and computing efficiency of different RRAMs in neural network accelerated computing.
[0026] Please refer to Figure 1 , one embodiment of this application provides a method for accelerating neural network computing based on RRAM, and this method may include the following multiple steps: S1: Obtain the hardware parameters of one or more RRAMs, where the hardware parameters at least include the number of parallelisms, and determine the configuration parameters of each RRAM according to a preset tolerance. The configuration parameters at least include the effective conductance range and multiple effective rows and columns.
[0027] The above-mentioned hardware parameters include the number of parallelisms, and may also include read / write voltages and storage structures. The above-mentioned number of parallelisms can be understood as the number of computing tasks that RRAM can process simultaneously, and its high parallelism can significantly accelerate neural network inference. For example, if the number of parallelisms is 128, the system can process 128 matrix multiplication operations simultaneously. The above-mentioned read / write voltages are the voltage levels required for RRAM to write and read data. Different voltages may affect the speed and accuracy of data conversion. The above-mentioned storage structures may be various memory structures such as shared memory, local memory, and on-chip memory. Different memory hierarchical structures may affect data access speed and parallel computing capabilities.
[0028] The above preset tolerance can be understood as the allowable range of hardware parameter deviations set by the user to ensure the accuracy and reliability of calculations. Preferably, a ±10% fluctuation in the conductance value is used as the preset tolerance. The above configuration parameters include the effective conductance range and the effective rows and columns. The above effective conductance range can be understood as the available conductance range of the RRAM cells in the stable state, and the above effective rows and columns can be understood as the actually available rows and columns in the RRAM array. Usually, the conductance values of the RRAM cells in each row or each column are the same. By setting the preset tolerance, the above effective conductance range and the effective rows and columns in the RRAM can be determined. For example, the rows and columns that meet the preset tolerance are used as the effective array. By obtaining the hardware configuration parameters of different RRAMs, cells with unstable conductance values can be filtered, the error impact can be reduced, and at the same time, the computing throughput can be improved by using the parallelizable number and the storage structure, providing key data support for the subsequent division of labor and acceleration of neural network tasks.
[0029] S3: Split the computing task of any network layer of the neural network, and generate a first computing task and a second computing task based on the splitting result.
[0030] The above neural network can be selected as CNN (Convolutional Neural Networks), RNN (Recurrent Neural Network), GAN (Generative Adversarial Networks), GNN (Graph Neural Network), or other neural networks. For the working types of different neural networks, the computing tasks of the neural network are split to form multiple computing tasks. Further, the above multiple computing tasks are divided into a first computing task and a second computing task.
[0031] The above first computing task can be a linear computing operation including linear combinations such as matrix multiplication and convolution operations. Since linear computing usually has high parallelism and regular data access patterns, it is suitable to use RRAM to execute the above first computing task. The above second computing task is preferably a non-linear computing operation that cannot be simply completed by linear combination. Since non-linear computing usually involves complex mathematical functions and logical judgments, it is suitable to use a processor such as a GPU to execute the above second computing task. It should be noted that the above second computing task can also be a linear computing operation.
[0032] Exemplarily, as a convolutional neural network (CNN), it mainly consists of convolutional layers, including a large number of matrix multiplications and pooling operations. The computational tasks of the CNN are split. The linear operations such as the matrix multiplication of the convolutional kernel are regarded as the first computational task, and the processing of non-linear operations such as activation and pooling is regarded as the second computational task. Among them, as a recurrent neural network (RNN), it is mainly for time series dependence, including recurrent calculations and gating mechanisms. The accelerated recurrent weight matrix multiplication can be regarded as the first computational task, and the processing of activation functions or gating functions is regarded as the second computational task.
[0033] S5: Determine the first processing path of the first computational task based on the number of parallelizable units and the effective conductance range, connect the first processing path to the target RRAM, and process the first computational task according to the effective rows and columns of the target RRAM to obtain a first processing result.
[0034] In this embodiment, the number of parallelizable units and the effective conductance range of the current RRAM are obtained to filter out RRAM cells with serious conductance value drift, and the RRAM cells that meet the tolerance range are used as the target RRAM. The first computational task is processed through the first processing path connected to the target RRAM, so as to improve the processing efficiency of some computational tasks and avoid the influence of hardware instability on the neural network acceleration calculation at the same time.
[0035] In this embodiment, if there are multiple selectable RRAMs, the first computational task is allocated to multiple first processing paths connected to multiple target RRAMs according to the number of parallelizable units and the effective conductance range of each RRAM to achieve load balancing. When the target RRAM of a certain first processing path fails due to conductance drift, the first computational task under the current first processing path will be automatically allocated to other first processing paths to obtain a first processing result.
[0036] S7: Determine the second processing path of the second computational task, process the second computational task according to the second processing path and the first processing result to obtain a second processing result, and use the second processing result as the output result of the current network layer.
[0037] In this embodiment, the output of the first computing task of the current network layer is used as the input of the second computing task. The second processing path is the task processing path for processing the second computing task. In a neural network, the first computing tasks and the second computing tasks of each network layer have an order relationship, and at the same time, there is an order relationship between each network layer. The above first processing result and the second processing result need to be spliced and transmitted in the order of the neural network layers to ensure the continuity of data and the accuracy of calculation. Exemplarily, a neural network may include a convolutional layer, an activation layer, and a fully connected layer. Taking the convolutional layer as an example, the RRAM processes all the matrix multiplications of the convolutional layer and outputs an intermediate feature map as the first processing result. The intermediate feature map is transmitted to the second processing path for activation operation to obtain an activated feature map as the second processing result. The above second processing result is used as the output result of the convolutional layer.
[0038] In view of this, the technical solution provided in the above embodiment of the present application automatically adapts different RRAMs and processors to jointly process a neural network to improve the computing efficiency and accuracy of the neural network. Specifically, according to the hardware parameters and configuration parameters of different RRAMs, the first processing path of different RRAMs is automatically adapted, and the computing tasks of the neural network are split and classified. The first processing path and the second processing path are used to process each type of computing task respectively to obtain the output result of the neural network. At the same time, the adapted first processing path is dynamically adjusted to balance the accuracy and power consumption of the computing task processing, and improve the efficiency and accuracy of the neural network accelerated computing.
[0039] In a possible embodiment, based on the above step S1, since the RRAM includes multiple storage rows and columns, determining the effective rows and columns of each RRAM according to a preset tolerance may be taking the storage rows and columns that meet the preset tolerance as the effective rows and columns, and obtaining the effective conductance values of each effective row and column, and determining the effective conductance range of the RRAM according to the effective conductance values.
[0040] The above RRAM is composed of multiple storage units, and these storage units are arranged in a matrix form to form a storage row and column structure of rows and columns. Among them, the storage units in each row or each column share a word line or a bit line. Therefore, each row or each column can be set as a storage row and column, and the conductance value of each storage row and column is the same. By selecting a specific storage row and column, all the storage units in that row and column can be accessed.
[0041] In this embodiment, effective rows and columns are screened and an effective conductance range is determined by setting a preset tolerance. Specifically, by recording the conductance values of each storage row and column and comparing them with the preset tolerance, the rows and columns with conductance values within the allowable range are screened out as effective rows and columns. Further, the conductance values of each effective row and column are recorded, and statistical analysis is performed on the conductance values of all effective rows and columns to find the minimum value and the maximum value to determine the effective conductance range of the RRAM. For example, the effective conductance range is defined as the interval from the minimum value to the maximum value of the conductance value.
[0042] Exemplarily, if there is an RRAM array including 100 rows and 100 columns, and the preset tolerance is that the allowable range of conductance value change does not exceed ±10%, through dynamic data testing, it is found that the conductance values of 80 rows or 80 columns are within the allowable range. These storage rows and columns are defined as effective rows and columns, and the conductance values of the effective rows and columns are recorded to determine the effective conductance range.
[0043] In one embodiment, please refer to Figure 2 , based on the above step S3, the calculation tasks of the neural network are split, and the first calculation task and the second calculation task can be generated based on the split result in the following manner: S301: According to the type of the neural network, determine each calculation task of the neural network, and for any one of the calculation tasks, judge whether the calculation task belongs to a linear calculation task.
[0044] S303: If the calculation task belongs to a linear calculation task, determine the task as the first calculation task; if the calculation task does not belong to a linear calculation task, determine the task as the second calculation task.
[0045] The above linear calculation task can be understood as a linear calculation operation with a linear relationship that can be completed through linear combination, such as matrix multiplication, convolution operation, loop operation, linear transformation, or other linear operations. Generally speaking, linear operations usually involve a large number of independent calculation tasks, have high parallelism, low precision, and regular data access patterns, which are suitable for the characteristics of RRAM. Assigning the above linear calculation tasks as the first calculation tasks to the target RRAM of the first calculation path can improve the calculation efficiency of neural network acceleration calculation.
[0046] The above non-linear computing tasks can be understood as non-linear computing operations where there is no linear relationship between the output and the input, such as activation operations, pooling operations, normalization operations, or other linear operations. The above activation operations can be selected as ReLU (Rectified Linear Unit), Sigmoid (Sigmoid function, an activation function), or other commonly used activation functions, and the values of the activation functions can be quickly calculated by processing multiple data points in parallel or by utilizing high-precision floating-point computing capabilities. The above pooling operations can be max pooling or average pooling, and the max pooling operation or average pooling operation can be quickly completed by processing multiple pooling windows in parallel.
[0047] In this embodiment, the computing tasks of different neural networks are split by determining whether they belong to linear tasks. Exemplarily, in the computing tasks of a CNN, the first computing task may include convolutional operations and the computing of fully connected layers, and the second computing task may include activation operations and pooling operations. In the computing tasks of an RNN, the first computing task may include matrix multiplication and recurrent operations, and the second computing task may include activation operations and gating operations. In the computing tasks of a GNN, the first computing task may include the computing of fully connected layers and convolutional operations, and the second computing task may include activation operations and normalization operations. In the computing tasks of a GNN, the first computing task may include graph convolutional operations and linear transformations, and the second computing task may include activation operations and pooling operations.
[0048] In one embodiment, based on the above step S5, it is necessary to determine the first processing path of the first computing task according to the number of parallelizable units and the effective conductance range, specifically including judging the RRAM that matches the first computing task according to the number of parallelizable units and the effective conductance range; determining the matching RRAM as the target RRAM, and determining the path connecting the target RRAM and the neural network as the first processing path.
[0049] In this embodiment, the target RRAM matching the first computing task is dynamically adjusted according to the number of parallelizable units and the effective conductance range to determine the first processing path. Specifically, if the first computing task requires high-precision computing, an RRAM with a relatively narrow effective conductance range can be selected as the target RRAM. If the first computing task requires high-parallelism processing, an RRAM with a relatively high number of parallelizable units can be selected as the target RRAM, and the path connecting the target RRAM and the neural network is determined as the first processing path.
[0050] In this embodiment, only the effective rows and columns with conductance values within the effective conductance range in the RRAM are used to process the first computing task. The target RRAM may be each RRAM region of different RRAMs. Each RRAM region includes one or more effective rows and columns. The data of the first computing task is block-parallel processed according to the number of parallelizable data, and each block of data is allocated to different RRAM regions for processing. At the same time, the conductance value drift of the target RRAM is dynamically detected. If the conductance value of a certain effective row or column exceeds the tolerance, the effective row or column is automatically skipped to improve the accuracy of the RRAM for neural network acceleration computing. Exemplarily, if the 5th column of the second RRAM fails, the task can be migrated to the spare effective rows and columns of the third RRAM.
[0051] In one embodiment, on the basis of step S5, the first computing task may be processed according to the effective array of the target RRAM to obtain a first processing result. Specifically, a positive weight array and a negative weight array of the target RRAM are determined based on the effective rows and columns, and the first computing task is calculated according to the positive weight array and the negative weight array to obtain a first processing result.
[0052] The positive weight array is used to store positive weight values, and the negative weight array is used to store negative weight values. In the effective array of the target RRAM, the positive weight array and the negative weight array are divided. By measuring the output current difference between the positive weight array and the negative weight array, and based on the result calculated from the above difference, the first processing result of the first computing task can be obtained, reducing the error caused by conductance value drift or other hardware noises in a single array, thereby improving the accuracy of the RRAM for neural network acceleration computing.
[0053] In this embodiment, please refer to Figure 3 , the positive weight array and the negative weight array are determined in the following manner: S501: Obtain the effective conductance values of the respective effective rows and columns, determine a reference row or column in the effective rows and columns of the target RRAM, and determine the effective conductance value of the reference row or column as the reference value.
[0054] S503: Determine the multiple effective rows and columns with effective conductance values greater than or equal to the reference value as the positive weight array, and determine the multiple effective rows and columns with effective conductance values less than the reference value as the negative weight array.
[0055] The above-mentioned reference row and column can be selected from any column in the effective rows and columns of the target RRAM, and the conductance value of this column is used as the reference value for comparison with the conductance values of other rows and columns. Optionally, a row and column group at a certain fixed position can be selected as the reference row and column, or the row and column with the most stable conductance value can be selected as the reference row and column, or the row and column with the conductance value in the middle range can be selected as the reference row and column. Further, the conductance values of each effective row and column are compared with the reference value, and the rows and columns with conductance values greater than or equal to the reference value are divided into positive weight arrays, and the rows and columns with conductance values less than the reference value are divided into negative weight arrays.
[0056] In this embodiment, by dividing the weights into positive and negative parts, the influence of hardware noise and conductance value drift can be offset by difference calculation. By determining the positive weight array and the negative weight array based on the effective array of the target RRAM, the first calculation task is processed to obtain a more accurate calculation result.
[0057] In this embodiment, please refer to Figure 4 , the calculation of the first calculation task according to the positive weight array and the negative weight array to obtain the first processing result can be carried out according to the following steps: S511: Obtain the positive weight matrix and the negative weight matrix of the neural network, store the positive weight matrix in the positive weight array, and store the negative weight matrix in the negative weight array.
[0058] S513: Obtain the first output current of the positive weight array and the second output current of the negative weight array, and use the difference between the first output current and the second output current as the first processing result.
[0059] The positive weight matrix and the negative weight matrix of the above-mentioned neural network represent the weight values of the current network layer of the neural network, corresponding to the positive weight and the negative weight respectively. The above-mentioned weight matrix is used to convert the input data into the output data of the current network layer. The above-mentioned positive weight matrix can be understood as the positive connection strength between the input data and the output data, and the above-mentioned negative weight matrix can be understood as the negative connection strength between the input data and the output data. Further, the above-mentioned positive weight matrix is written into the positive weight array, and the above-mentioned negative weight matrix is written into the negative weight array.
[0060] In this embodiment, the input data is applied to the positive weight array and the negative weight array to obtain the first output current of the positive weight array and the second output current of the negative weight array, and the difference between the first output current and the second output current is used as the first processing result. At this time, if the influence directions of temperature change or conductance drift on the positive and negative weight arrays are the same, their interpolation will cancel the common-mode error, making the calculation result of the RRAM for neural network acceleration calculation more accurate.
[0061] Exemplarily, obtain the weight matrix of a certain layer from the neural network , and decompose the weight matrix into a positive weight matrix and a negative weight matrix . The input data is . Store in the positive weight matrix, and store in the negative weight matrix. Further, apply the input data to the positive weight array in the form of voltage to obtain a first output current . Apply the input data to the negative weight array in the form of voltage to obtain a first output current , and calculate the difference to use the difference as the first processing result of this network layer .
[0062] In one embodiment, based on the above step S7, the above second processing path is connected to a processor. Based on the first processing result, the processor connected through the second processing path processes the above second computing task to obtain a second processing result. Specifically, determine the target processor connected to the neural network, determine the path connecting the target processor and the neural network as the second processing path, send the second computing task and the first processing result to the target processor through the first processing path, and the target processor executes the second computing task according to the first processing result to obtain the second processing result.
[0063] The second processing path is a task processing path for processing the second computing task through the target processor. The target processor can be selected as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an NPU (Neural Network Processing Unit), a TPU (Tensor Processing Unit), a DPU (Deep Learning Processing Unit), a microprocessor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or a combination of at least two of these processor forms.
[0064] Exemplarily, the first processing result The second computing task is sent to the target processor through the second processing path. The second computing task takes the activation function as an example, and the target processor takes the GPU as an example. After receiving the first processing result Y, the GPU executes the activation function to obtain the second processing result as the output result of the current network layer.
[0065] In this embodiment, by ensuring that the advance calculation task and the nonlinear calculation task are efficiently executed on different hardware, the high parallelism of RRAM and the high memory bandwidth of the processor are utilized to improve the efficiency of neural network calculation.
[0066] In one embodiment, based on the above step S7, after obtaining the second processing result, the performance parameters of the second processing result are evaluated, and dynamically adjusted according to the evaluation results. Specifically, the first computing task and the second computing task are calculated according to the second processing path to obtain a third processing result, the performance parameters of the second processing result and the third processing result are evaluated, and the target RRAM or the first processing path is adjusted based on the evaluation results.
[0067] In this embodiment, by verifying whether the result of accelerating the neural network calculation using RRAM meets the performance requirements, it can be used to optimize RRAM subsequently. Exemplarily, the above performance parameters may include one or more of accuracy, calculation speed, calculation power consumption, hardware stability, and system scalability. The above accuracy can be understood as the calculation error between the RRAM output value and the theoretical value, usually the numerical deviation of matrix multiplication caused by hardware instabilities such as conductance values. The above calculation speed can be the single-inference latency or throughput, the above calculation power consumption can be understood as the total energy consumption of the system, the above hardware stability can be the failure ratio of RRAM cells or the drift amplitude of conductance values, and the above system scalability can be the maximum neural network scale supported or the collaborative efficiency of multiple computing tasks.
[0068] In this embodiment, both the first processing task and the second processing task are processed by a processor to obtain a third processing result for comparison with the above second processing result. The above third processing result serves as the standard output result of the neural network acceleration calculation according to the processor. By quantifying the difference between the second processing result and the third processing result, the processing result of RRAM can be evaluated.
[0069] In this embodiment, the above performance parameters may include accuracy and calculation power consumption. If the calculation accuracy of the second processing result is lower than that of the third processing result, adjust the hardware parameters of the target RRAM. If the calculation power consumption of the second processing result is greater than that of the third processing result, determine the second processing path as the first processing path.
[0070] The evaluation of the above accuracy can be carried out by quantifying the difference between the second processing result and the third processing result. Exemplarily, if the accuracy of the third processing result obtained by GPU calculation is 98%, and the accuracy of the second processing result obtained by RRAM calculation is 97.2%, then the error is 0.8%. It is necessary to analyze whether the error is caused by conductance drift, and then adjust RRAM and the first processing path. Specifically, when the accuracy of the second processing result is lower than that of the third processing result, adjust the hardware parameters of the target RRAM in the first processing path, or replace it with another optional first processing path of another target RRAM.
[0071] The above evaluation of the computing power consumption can be processed by an external instrument to record the real-time power consumption of RRAM processing and GPU processing. Exemplarily, it takes 5 ms for RRAM to complete one inference with a power consumption of 0.1 J, and it takes 3 ms for GPU with a power consumption of 0.5 J. Under the same time consumption, the power consumption of RRAM is lower. Exemplarily, when the computing power consumption of the second calculation result is greater than that of the third processing result, the second processing path is determined as the first processing path, that is, both the first calculation task and the second calculation task are processed by the processor of the first processing path.
[0072] This application provides one or more embodiments of applying the above method for accelerating neural network computing based on RRAM. The embodiment is carried out according to the following steps: Embodiment 1: Deploy a CNN network for image classification.
[0073] In this embodiment, if there are four RRAM boards, each RRAM contains a 128×128 array and supports parallel computing. First, determine the conductance range and effective number of rows and columns of each RRAM board. If it is found that the conductance value of the fourth column of the second RRAM exceeds the preset tolerance, then skip the storage array of the fourth column of the second RRAM, and use other storage arrays as the effective number of rows and columns. Further, parallelly calculate the weight matrix of the CNN network in the effective number of rows and columns through RRAM, and process non-linear operations such as the ReLU activation function in the GPU to obtain the final processing result of the CNN. At the same time, optimize the data block division and transmission path of the weight matrix according to the number of parallelisms and the stored results to improve the accuracy and processing efficiency of the CNN network.
[0074] Embodiment 2: Deploy a ResNet-50 (Residual Nerwork 50 layers, a deep residual network) for image classification.
[0075] In this embodiment, if there are four RRAM boards and it is found that the conductance value of the second column of the first RRAM exceeds the preset tolerance, then skip the second column of the first board, and allocate the convolution calculation to other effective number of rows and columns. For example, store the positive and negative weight arrays in the effective number of rows and columns of the first board and the effective number of rows and columns of the third board respectively. Through difference calculation, part of the hardware error can be offset. Process all matrix multiplications of the convolutional layers, such as 3×3 convolutional kernels, through the effective number of rows and columns of RRAM to obtain the feature map, and perform normalization and residual connection processing on the feature map in blocks through the GPU.
[0076] Please refer to Figure 5 , this application also provides a system for accelerating neural network computing based on RRAM. The system includes: A parameter acquisition unit 100 is configured to acquire hardware parameters of one or more RRAMs. The hardware parameters at least include the number of parallelizable units, and determine configuration parameters of each of the RRAMs according to a preset tolerance. The configuration parameters at least include an effective conductance range and a plurality of effective rows and columns. A task splitting unit 200 is configured to split a computing task of any network layer of a neural network, and generate a first computing task and a second computing task based on the splitting result. A task computing unit 300 is configured to determine a first processing path of the first computing task based on the number of parallelizable units and the effective conductance range. The first processing path is connected to a target RRAM, process the first computing task according to the effective rows and columns of the target RRAM to obtain a first processing result, determine a second processing path of the second computing task, and process the second computing task according to the second processing path and the first processing result to obtain a second processing result, and use the second processing result as the output result of the current network layer.
[0077] In one embodiment, the parameter acquisition unit 100 is specifically configured to use the storage rows and columns that meet the preset tolerance as the effective rows and columns, acquire the effective conductance values of the respective effective rows and columns, and determine the effective conductance range of the RRAM according to the effective conductance values.
[0078] In one embodiment, the task splitting unit 200 is specifically configured to determine each computing task of the neural network according to the type of the neural network, and for any one of the computing tasks, determine whether the computing task belongs to a linear computing task. If the computing task belongs to a linear computing task, determine the task as the first computing task; if the computing task does not belong to a linear computing task, determine the task as the second computing task.
[0079] In one embodiment, the task computing unit 300 is specifically configured to determine a positive weight array and a negative weight array of the target RRAM based on the effective rows and columns, calculate the first computing task according to the positive weight array and the negative weight array to obtain a first processing result, determine a target processor connected to the neural network, determine the path connecting the target processor and the neural network as the second processing path, send the second computing task and the first processing result to the target processor through the first processing path, and the target processor executes the second computing task according to the first processing result to obtain the second processing result.
[0080] In one embodiment, the system further includes a verification unit, which is configured to calculate the first computing task and the second computing task according to the first processing path to obtain a third processing result, evaluate the performance parameters of the second processing result and the third processing result, and adjust the target RRAM or the first processing path based on the evaluation result, where the performance parameters at least include computing accuracy and computing power consumption.
[0081] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding embodiments above, and will not be elaborated here.
[0082] The system for accelerating neural network computing based on RRAM in the embodiments of the present application is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, or other devices that can provide the above functions.
[0083] In one embodiment, the above verification unit may include one or more processors (or processing circuits), a memory, and a computer program stored in the memory, which are configured to implement the above method. Among them, the processor can be selected from: CPU (Central Processing Unit), GPU (Graphics Processing Unit), NPU (Neural Network Processing Unit), TPU (Tensor Processing Unit), DPU (Deep learning Processing Unit), microprocessor, DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), FPGA (Field Programmable Gate Array), or a combination of at least two of these processor forms. The memory stores instructions executable by at least one processor, and the types of the at least one processor can be different, such as including CPU and FPGA, CPU and artificial intelligence processor, CPU and GPU, etc. The computer program can be written in the python language and can also be written in other programming languages, including but not limited to python, perl, and tcl.
[0084] In one embodiment, a processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and execution capabilities, such as a CPU, microprocessor, GPU, or DSP, etc.; in another implementation, the processor can achieve certain functions through the logical relationship of hardware circuits, and the logical relationship of the hardware circuits is fixed or can be reconfigured. For example, the processor is a hardware circuit implemented by an ASIC or PLD, such as an FPGA, etc. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the configuration of the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as a type of ASIC, such as an NPU, TPU, DPU, etc.
[0085] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of a computer device provided by an embodiment of the present application. As Figure 6 shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Figure 6 In
[0086] FIG. is taken one processor 10 as an example.
[0087] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiments.
[0088] The memory 20 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely disposed relative to the processor 10, and these remote memories may be connected to the computer device through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0089] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, a hard disk, or a solid-state drive; the memory 20 may further include a combination of the above types of memory.
[0090] The computer device further includes a communication interface 30 for communicating the computer device with other devices or a communication network.
[0091] The systems, modules, or units illustrated in the above embodiments may be specifically implemented by a computer chip or an entity, or by a product having a certain function. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0092] For convenience of description, when describing the above device, various units are described separately according to their functions. Of course, when implementing the present application, the functions of the various units may be implemented in one or more software and / or hardware.
[0093] The present application is described with reference to the flowcharts and / or block diagrams of methods and systems according to embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0094] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the functions specified in one or more of the flows Figure 1 one or more of the flows and / or boxes Figure 1 specified in one or more of the boxes or boxes.
[0095] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or boxes Figure 1 specified in one or more of the boxes or boxes.
[0096] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.
[0097] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.
[0098] The above are only the embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
[0099] Although the embodiments of the present application have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations fall within the scope defined by the appended claims.
Claims
1. A method for accelerating neural network computing based on RRAM, characterized in that, The method includes: Obtaining hardware parameters of one or more RRAMs, where the hardware parameters at least include the number of parallelizable units, and determining configuration parameters of each of the RRAMs according to a preset tolerance, where the configuration parameters at least include an effective conductance range and a plurality of effective rows and columns; Splitting a computing task of any network layer of a neural network, and generating a first computing task and a second computing task based on the splitting result; Determining a first processing path of the first computing task based on the number of parallelizable units and the effective conductance range, where the first processing path is connected to a target RRAM, and processing the first computing task according to the effective rows and columns of the target RRAM to obtain a first processing result; Determining a second processing path of the second computing task, and processing the second computing task according to the second processing path and the first processing result to obtain a second processing result, and using the second processing result as the output result of the current network layer.
2. The method according to claim 1, wherein After obtaining the second processing result, the method further includes: Calculating the first computing task and the second computing task according to the second processing path to obtain a third processing result; Evaluating performance parameters of the second processing result and the third processing result, and adjusting the target RRAM or the first processing path based on the evaluation result.
3. The method according to claim 2, wherein The performance parameters at least include computing accuracy and computing power consumption; evaluating the performance parameters of the second processing result and the third processing result and adjusting the target RRAM or the first processing path based on the evaluation result includes: If the computing accuracy of the second processing result is lower than that of the third processing result, adjusting the hardware parameters of the target RRAM; If the computing power consumption of the second processing result is greater than that of the third processing result, determining the second processing path as the first processing path.
4. The method according to claim 1, wherein The RRAM includes a plurality of storage rows and columns; determining the configuration parameters of each of the RRAMs according to a preset tolerance includes: Regarding the storage rows and columns that meet the preset tolerance as effective rows and columns; Obtaining effective conductance values of the respective effective rows and columns, and determining an effective conductance range of the RRAM according to the effective conductance values.
5. The method according to claim 1, wherein Splitting a computing task of the neural network, and generating a first computing task and a second computing task based on the splitting result includes: According to the type of the neural network, determining each computing task of the neural network, and for any one of the computing tasks, determining whether the computing task belongs to a linear computing task; If the computing task belongs to a linear computing task, determining the task as the first computing task, and if the computing task does not belong to a linear computing task, determining the task as the second computing task.
6. The method according to claim 1, wherein Determining a first processing path of the first computing task based on the number of parallelizable units and the effective conductance range, where the first processing path is connected to a target RRAM includes: Judging an RRAM that matches the first computing task according to the number of parallelizable units and the effective conductance range; Determine the matched RRAM as the target RRAM, and determine the path connecting the target RRAM and the neural network as the first processing path.
7. The method according to claim 1, wherein Processing the first computing task according to the effective rows and columns of the target RRAM to obtain a first processing result includes: Determine the positive weight array and the negative weight array of the target RRAM based on the effective rows and columns; Calculate the first computing task according to the positive weight array and the negative weight array to obtain a first processing result.
8. The method according to claim 7, wherein Determining the positive weight array and the negative weight array of the target RRAM based on the effective rows and columns includes: Obtain the effective conductance values of the respective effective rows and columns, determine a reference row and column in the effective rows and columns of the target RRAM, and determine the effective conductance value of the reference row and column as the reference value; Determine the multiple effective rows and columns with the effective conductance value greater than or equal to the reference value as the positive weight array, and determine the multiple effective rows and columns with the effective conductance value less than the reference value as the negative weight array.
9. The method according to claim 7, wherein Calculating the first computing task according to the positive weight array and the negative weight array to obtain a first processing result includes: Obtain the positive weight matrix and the negative weight matrix of the neural network, store the positive weight matrix in the positive weight array, and store the negative weight matrix in the negative weight array; Obtain the first output current of the positive weight array and the second output current of the negative weight array, and use the difference between the first output current and the second output current as the first processing result.
10. The method according to claim 1, wherein The second processing path is connected to the processor; determine the second processing path of the second computing task, and process the second computing task according to the second processing path and the first processing result to obtain a second processing result includes: Determine the target processor connected to the neural network, and determine the path connecting the target processor and the neural network as the second processing path; Send the second computing task and the first processing result to the target processor through the first processing path, and the target processor executes the second computing task according to the first processing result to obtain the second processing result.
11. A system for accelerating neural network computations based on RRAM, characterized in that, The system includes: A parameter acquisition unit for acquiring hardware parameters of one or more RRAMs, the hardware parameters at least including the parallelism number, and determining configuration parameters of each RRAM according to a preset tolerance, the configuration parameters at least including an effective conductance range and multiple effective rows and columns; A task splitting unit for splitting the computing task of any network layer of the neural network, and generating a first computing task and a second computing task based on the splitting result; A task calculation unit, configured to determine a first processing path of the first calculation task based on the parallelizable number and the effective conductance range, where the first processing path is connected to a target RRAM, process the first calculation task according to the effective row and column of the target RRAM to obtain a first processing result, determine a second processing path of the second calculation task, process the second calculation task according to the second processing path and the first processing result to obtain a second processing result, and use the second processing result as the output result of the current network layer.
12. The system according to claim 11, wherein The system further includes: A verification unit, configured to calculate the first calculation task and the second calculation task according to the first processing path to obtain a third processing result, evaluate the performance parameters of the second processing result and the third processing result, and adjust the target RRAM or the first processing path based on the evaluation result, where the performance parameters at least include calculation accuracy and calculation power consumption.
13. A computer device, characterized in that, including: A memory and a processor, where the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the method for accelerating neural network calculation based on RRAM according to any one of claims 1 to 10.
Citation Information
Patent Citations
Neural network circuits having non-volatile synapse arrays
CN111656371A
Neural network circuits having non-volatile synapse arrays
US20190164046A1