Deep neural network accelerator based on in-memory calculation and acceleration method thereof
By designing a deep neural network accelerator based on in-memory computing, and using a general processor and in-memory computing module to dynamically adapt different computing graph structures, the problem of poor universality of existing accelerators is solved, and efficient deep neural network training and inference are achieved.
Patent Information
- Application Number
- CN202510348479.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-01
AI Technical Summary
Most existing deep neural network accelerators accelerate specific network model structures, which are poor in versatility and are difficult to dynamically adapt to different computing graph structures.
Design a deep neural network accelerator based on in-memory computing, including a general-purpose processor and in-memory computing module. The general processor deploys a deep neural network based on the problem to be solved, and determines the gradient parameters of the weight parameters based on the output data and error information through the in-memory computing module, and performs iterative updates to realize pre-training and backpropagation training of deep neural networks.
It realizes acceleration of the network model type that does not depend on deep neural networks, improves the training efficiency of deep neural networks, and has a certain degree of versatility, and is suitable for solving a variety of scientific computing problems.
Smart Images

Figure CN120235191A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence technology, and more specifically, relates to a deep neural network accelerator based on in-memory computing and its acceleration method. Background Art
[0002] In today's information age, deep neural networks (DNNs) have been widely applied in fields such as image recognition, speech processing, and natural language understanding. Deep neural networks include various types. For example, convolutional neural networks (CNNs) are widely used in the field of image recognition and processing; residual networks (ResNets) solve the problem of difficult training of deep networks by introducing skip connections, improving network performance; deep operator networks (DeepONets) are suitable for function approximation and dynamic system modeling. These network architectures have their own characteristics and provide strong support for different application scenarios.
[0003] With the rapid development of deep neural network models in terms of structural diversity and data scale, higher requirements are put forward for the adaptability of the computing architecture and energy efficiency optimization. In current mainstream DNN acceleration solutions, the architecture based on the central processing unit (CPU) is limited by the sequential execution mode and is difficult to dynamically adapt to different computational graph structures such as convolution and loop; although the graphics processing unit (GPU) improves throughput through a large number of parallel computing units, its fixed stream processor architecture is mainly optimized for specific computational modes (such as matrix multiplication), resulting in a sharp drop in the support efficiency for emerging network operators (such as dynamic sparse computing); the field-programmable gate array (FPGA) solution has the characteristic of hardware reconfigurability, but it needs to redesign the data path and storage hierarchy when facing different network structures, bringing significant development cycles and verification costs. These technical status quo indicate that most existing deep neural network accelerators are designed to accelerate specific network model structures and have poor generality. Summary of the Invention
[0004] Aiming at the defects of the prior art, the purpose of this application is to provide a deep neural network accelerator based on in-memory computing and its acceleration method, aiming to solve the problem that most existing deep neural network accelerators are designed to accelerate specific network model structures and have poor generality.
[0005] To achieve the above object, in a first aspect, this application provides a deep neural network accelerator based on in-memory computing, including: A general-purpose processor, configured to deploy a deep neural network according to the problem to be solved and determine the output data during the forward propagation process of the deep neural network for the problem to be solved; The in-memory computing module is used to determine the first gradient parameters of the first weight parameters of each network layer in the deep neural network according to the first error information between the output data and the expected output during the forward propagation process, so that the general-purpose processor iteratively updates the first weight parameters according to the first gradient parameters until the deep neural network completes pre-training.
[0006] In some embodiments, the general-purpose processor is further used to determine the second gradient parameters of the updated first weight parameters of each network layer according to the second error information between the output data and the expected output during the forward propagation process of the pre-trained deep neural network for the problem to be solved, and iteratively update the updated first weight parameters according to the second gradient parameters until the deep neural network completes backpropagation training.
[0007] In some embodiments, the general-purpose processor includes: A cache for storing the output data during the forward propagation process of the pre-trained deep neural network for the problem to be solved; The near-memory processor is used to determine the second gradient parameters according to the second error information between the output data and the expected output during the forward propagation process of the pre-trained deep neural network for the problem to be solved, and update the updated first weight parameters according to the second gradient parameters until the deep neural network completes backpropagation training.
[0008] In some embodiments, the in-memory computing module is further used to determine the output data during the inference process according to the output data of each network layer during the inference process of the deep neural network for the problem to be solved after backpropagation training.
[0009] In some embodiments, the in-memory computing module includes: A pre-alignment unit, an input register, at least one in-memory computing array, and a shift accumulator; The pre-alignment unit is used to perform format conversion from floating-point numbers to fixed-point numbers on the weight parameters of each network layer in the deep neural network after backpropagation training, and obtain the fixed-point number information corresponding to the weight parameters of each network layer in the deep neural network after backpropagation training; The input register is used to split the mantissa in the fixed-point number information according to different precision requirements for each network layer; Each in-memory computing array is respectively used to store the fixed-point number information after mantissa splitting according to different precision requirements for each network layer, and determine the output data of each network layer during the inference process according to the input data during the inference process of the deep neural network for the problem to be solved after backpropagation training; The shift accumulator is used to determine the output data during the inference process according to the output data of each network layer during the inference process.
[0010] In some embodiments, each in-memory computing array consists of a digital-to-analog converter, at least one memory cell, and an analog-to-digital converter. Among them, the digital-to-analog converter is used to convert input data into a voltage vector; each memory cell is used to store the fixed-point number information after the mantissa splitting of any network layer in the deep neural network after backpropagation training according to any precision requirement, and convert the voltage vector into an output current; the analog-to-digital converter is used to convert the output current of each memory cell into a numerical quantity, and determine the output data of any network layer in the inference process of the deep neural network after backpropagation training for the problem to be solved according to the numerical quantity.
[0011] In some embodiments, the memory cell is composed of any one of the following non-volatile memories: Resistive random access memory, phase change memory, spin transfer torque magnetic memory, ferroelectric field effect transistor, and non-volatile flash memory.
[0012] In a second aspect, the present application provides an acceleration method for a deep neural network accelerator based on in-memory computing, including: Based on a general-purpose processor, deploy a deep neural network according to the problem to be solved, and determine the output data in the forward propagation process of the deep neural network for the problem to be solved; Based on the in-memory computing module, determine the first gradient parameter of the first weight parameter of each network layer in the deep neural network according to the first error information between the output data in the forward propagation process and the expected output, so that the general-purpose processor iteratively updates the first weight parameter according to the first gradient parameter until the deep neural network completes pre-training.
[0013] In some embodiments, the method further includes: Based on a general-purpose processor, determine the second gradient parameter of the updated first weight parameter of each network layer according to the second error information between the output data in the forward propagation process of the pre-trained deep neural network for the problem to be solved and the expected output, and iteratively update the updated first weight parameter according to the second gradient parameter until the deep neural network completes backpropagation training.
[0014] In a third aspect, the present application provides an electronic device, including: at least one memory for storing a program; at least one processor for executing the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method described in the second aspect or any of the embodiments of the second aspect.
[0015] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program. When the computer program runs on a processor, the processor is caused to execute the method described in the second aspect or any of the embodiments of the second aspect.
[0016] In a fifth aspect, the present application provides a computer program product. When the computer program product runs on a processor, it causes the processor to execute the method described in the second aspect or any of the embodiments of the second aspect.
[0017] Generally speaking, compared with the prior art, the above technical solution conceived by the present application has the following beneficial effects: The in-memory computing-based deep neural network accelerator and its acceleration method provided by the embodiments of the present application use a general-purpose processor to determine the output data in the forward propagation process of the deep neural network for the problem to be solved, and use the in-memory computing module to determine the first gradient parameter of the first weight parameter of each network layer according to the first error information between the output data and the expected output, so that the general-purpose processor updates the first weight parameter according to the first gradient parameter, thereby realizing the acceleration of the pre-training process of the deep neural network without depending on the type of the deep neural network model and improving the training efficiency of the deep neural network. At the same time, since the deep neural network can be widely applied to the solution of scientific computing problems, such as the solution of linear / nonlinear equations, curve least squares fitting problems, steady-state partial differential equation solutions, etc., the deep neural network accelerator of the present application has a certain generality. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is one of the schematic structural diagrams of the in-memory computing-based deep neural network accelerator provided by the embodiments of the present application; Figure 2 is another schematic structural diagram of the in-memory computing-based deep neural network accelerator provided by the embodiments of the present application; Figure 3 is the schematic structural diagram of the in-memory computing array provided by the embodiments of the present application; Figure 4 is the schematic diagram of the data processing method in the inference process provided by the embodiments of the present application; Figure 5 is the schematic structural diagram of the storage unit provided by the embodiments of the present application; Figure 6 is the schematic hardware structure diagram of the deep operator network accelerator in the training stage provided by the embodiments of the present application; Figure 7 is the schematic hardware structure diagram of the deep operator network accelerator in the inference stage provided by the embodiments of the present application; Figure 8 is the schematic flowchart of solving the steady-state partial differential equation provided by the embodiments of the present application; Figure 9 is the schematic diagram of the training process of the deep operator network accelerator provided by the embodiments of the present application; Figure 10It is a schematic diagram of the inference process of the deep operator network accelerator provided by an embodiment of the present application; Figure 11 It is a schematic flowchart of the acceleration method of the deep neural network accelerator based on in-memory computing provided by an embodiment of the present application; Figure 12 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0019] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, but not to limit the present application.
[0020] The term "and / or" in this document is an association relationship describing associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The symbol " / " in this document represents an "or" relationship between associated objects. For example, A / B represents A or B.
[0021] The terms "first" and "second" in the description and claims of this document are used to distinguish different objects, rather than to describe a specific order of objects. For example, the first error information and the second error information are used to distinguish different error information, rather than to describe the specific order of error information.
[0022] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0023] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" refers to two or more.
[0024] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application.
[0025] Please refer to Figure 1 , an embodiment of the present application provides a deep neural network accelerator based on in-memory computing, including: a general-purpose processor 110 and an in-memory computing module 120.
[0026] The general-purpose processor 110 is configured to deploy a deep neural network according to the problem to be solved, and determine the output data during the forward propagation process of the deep neural network for the problem to be solved; The in-memory computing module 120 is configured to determine, according to the first error information between the output data and the expected output during the forward propagation process, the first gradient parameters of the first weight parameters of each network layer in the deep neural network, so that the general-purpose processor updates the first weight parameters according to the first gradient parameters until the deep neural network completes pre-training.
[0027] In the embodiment of the present application, the deep neural network accelerator includes a general-purpose processor 110 and an in-memory computing module 120. Among them, the in-memory computing module 120 uses a non-volatile memory array to perform large-scale vector matrix multiplication to achieve acceleration of deep neural network training and inference.
[0028] The general-purpose processor 110 can be used to select a suitable deep neural network for initialization and deployment according to the problem to be solved, and initialize and fix the in-memory computing module 120 as a training acceleration matrix, which is set in advance.
[0029] The input data for the problem to be solved is transmitted to the deep neural network deployed on the general-purpose processor 110. The general-purpose processor obtains the output data during the forward propagation process of the deep neural network for the problem to be solved and transmits it to the in-memory computing module.
[0030] It should be noted that the problem to be solved can be various scientific computing problems such as solving linear / nonlinear equations, solving steady-state partial differential equations, and linear regression. Any scientific computing problem that can be solved using a deep neural network can be solved using this deep neural network.
[0031] The in-memory computing module 120 is configured to obtain, according to the error information (i.e., the first error information) between the output data and the expected output during the forward propagation process of the deep neural network for the problem to be solved, the gradient parameters (i.e., the first gradient parameters) corresponding to the weight parameters (i.e., the first weight parameters) of each network layer in the deep neural network, and transmit the first gradient parameters to the general-purpose processor 110. Specifically, the first gradient parameters corresponding to the first weight parameters of each network layer of the deep neural network can be obtained by performing linear / nonlinear processing on the first error information. In the embodiment of the present application, the first error information, the first weight parameters, and the first gradient parameters can all be represented in the form of matrices.
[0032] The general-purpose processor 110 iteratively updates the first weight parameters of each network layer according to the above first gradient parameters until the deep neural network completes pre-training.
[0033] The in-memory computing-based deep neural network accelerator provided by the embodiments of the present application uses a general-purpose processor to determine the output data during the forward propagation of the deep neural network for the problem to be solved, and uses the in-memory computing module to determine the first gradient parameter of the first weight parameter of each network layer according to the first error information between the output data and the expected output, so that the general-purpose processor updates the first weight parameter according to the first gradient parameter, thereby realizing the acceleration of the pre-training process of the deep neural network without depending on the type of the deep neural network model and improving the training efficiency of the deep neural network. At the same time, since the deep neural network can be widely applied to the solution of scientific computing problems, such as solving linear / nonlinear equations, curve least squares fitting problems, solving steady-state partial differential equations, etc., the deep neural network accelerator of the present application has a certain generality.
[0034] Further, in some embodiments, the general-purpose processor 110 is further configured to determine the second gradient parameter of the updated first weight parameter of each network layer according to the second error information between the output data and the expected output during the forward propagation of the pre-trained deep neural network for the problem to be solved, and update the updated first weight parameter according to the second gradient parameter until the deep neural network completes the backpropagation training.
[0035] In the embodiments of the present application, after the above deep neural network completes pre-training, the input data of the deep neural network for the problem to be solved is transmitted to the general-purpose processor 110. The general-purpose processor 110 can determine the output data during the forward propagation of the pre-trained deep neural network for the problem to be solved according to the input data, and according to the error information (i.e., the second error information) between the output data and the expected output of the deep neural network, and the in-memory computing module 120 performs a fixed calculation expansion on the second error information to obtain the error information matrix of each network layer. The error information matrix is linearly or non-linearly transformed to obtain the gradient parameter (i.e., the second gradient parameter) of the updated first weight parameter of each network layer, and the updated first weight parameter is updated again according to the second gradient parameter until the deep neural network completes the backpropagation (BP) training.
[0036] Further, in some embodiments, the general-purpose processor 110 includes: A cache for storing the output data during the forward propagation of the pre-trained deep neural network for the problem to be solved; A near-memory processor for determining the second gradient parameter according to the second error information between the output data and the expected output during the forward propagation of the pre-trained deep neural network for the problem to be solved, and updating the updated first weight parameter according to the second gradient parameter until the deep neural network completes the backpropagation training.
[0037] Please further refer to Figure 2 The general-purpose processor 110 includes a central processing unit, a cache, and a near-memory processor. Among them, the central processing unit is responsible for executing program instructions, processing data, and controlling other hardware components. The cache is used to improve data access speed and the overall system performance. The near-memory processor is responsible for processing the output of the in-memory computing module.
[0038] In a specific implementation, until the deep neural network completes pre-training, the forward propagation process of the pre-trained deep neural network for the above-mentioned problem to be solved is repeated. The output data of the forward propagation of the pre-trained deep neural network for the problem to be solved is stored in the cache and transmitted by the cache to the near-memory processor. The near-memory processor calculates the gradient based on the error information (i.e., the second error information) between the output data of the forward propagation of the pre-trained deep neural network for the problem to be solved and the expected output, and obtains the gradient parameters (i.e., the second gradient parameters) corresponding to the updated first weight parameters of each network layer in the pre-trained deep neural network, so as to update the weights of the updated first weight parameters of each network layer according to the second gradient parameters until the pre-trained deep neural network completes the backpropagation training (i.e., the accuracy of the pre-trained deep neural network reaches the expectation), thereby achieving the purpose of optimizing the data flow, reducing the memory access latency, and accelerating the backpropagation training process of the entire network.
[0039] Furthermore, in some embodiments, the in-memory computing module 120 is further configured to determine the output data during the inference process according to the output data of each network layer during the inference process of the deep neural network after backpropagation training for the problem to be solved.
[0040] In the embodiments of the present application, the in-memory computing module 120 can also be used to obtain the output data during the inference process of the deep neural network after backpropagation training for the problem to be solved according to the output data of each network layer during the inference process of the deep neural network after backpropagation training for the above-mentioned problem to be solved.
[0041] Specifically, after the deep neural network completes backpropagation training, the deep neural network after backpropagation training is deployed on the in-memory computing module 120, and the input data of the inference process of the deep neural network after backpropagation training for the problem to be solved is transmitted to the in-memory computing module 120. The in-memory computing module 120 can obtain the output data of each network layer of the inference process of the deep neural network after backpropagation training deployed on it for the problem to be solved, and based on the output data of each network layer, obtain the output data during the inference process of the deep neural network after backpropagation training for the problem to be solved.
[0042] Further, in some embodiments, the in-memory computing module 120 includes: a pre-alignment unit, an input register, at least one in-memory computing array, and a shift accumulator; The pre-alignment unit is configured to perform a format conversion from floating-point numbers to fixed-point numbers on the weight parameters of each network layer in the depth neural network after backpropagation training, and obtain the fixed-point number information corresponding to the weight parameters of each network layer in the depth neural network after backpropagation training; The input register is configured to split the mantissa in the fixed-point number information according to different precision requirements for each network layer; Each in-memory computing array is respectively configured to store the fixed-point number information after mantissa splitting according to different precision requirements for each network layer, and determine the output data of each network layer during the inference process according to the input data during the inference process of the depth neural network after backpropagation training for the problem to be solved; The shift accumulator is configured to determine the output data during the inference process according to the output data of each network layer during the inference process.
[0043] Please continue to refer to Figure 2 , the in-memory computing module 120 includes: an in-memory control unit, a data interface, a pre-alignment unit, an input register, an input driver, at least one in-memory computing array, an output driver, a shift accumulator, and an output register.
[0044] The in-memory control unit is responsible for controlling the data transmission of the depth neural network of the in-memory computing module 120. The data transmitted by the Cache is received by the data interface, processed by the pre-alignment unit, the input register, and the input driver, and then a large-scale vector matrix multiplication operation is performed on each in-memory computing array. After being processed by the output driver, the shift accumulator, and the output register, finally, the data interface returns the calculation result to the Cache. Thus, the depth neural network accelerator can implement the in-memory computing process of the depth neural network and accelerate large-scale scientific computing problems with high computational complexity. In the embodiments of the present application, the data interface uses a field-programmable gate array (FPGA).
[0045] In a specific implementation, after the deep neural network completes backpropagation training, the weight parameters of each network layer of the deep neural network after backpropagation training are passed from the Cache to the in-memory computing module 120. The in-memory control unit controls the pre-alignment unit to perform a format conversion from floating-point to fixed-point on the weight parameters of each network layer of the deep neural network after backpropagation training (the weight parameter can be a floating-point matrix, and the floating-point matrix can be in single-precision, double-precision, or extended-precision format). Specifically, the pre-alignment unit shifts all elements in the floating-point matrix corresponding to the weight parameter to the highest exponent for alignment and performs necessary mantissa padding to realize the conversion of the floating-point matrix to a fixed-point matrix during the deployment process. The mantissa part of the obtained fixed-point information (i.e., the converted fixed-point matrix) is transmitted to the input register; the in-memory control unit controls the input register to split the mantissa part according to different precision requirements for different network layers (e.g., different network layers can independently select the fixed-point bit width (such as 4bit, 8bit, 16bit, etc.)) and deploy it in blocks to each in-memory computing array to realize the mantissa deployment of the weight parameters of each network layer.
[0046] In the embodiments of the present application, each in-memory computing array stores the fixed-point information obtained by splitting the mantissa according to different precision requirements for the above-mentioned each network layer.
[0047] When the deep neural network after backpropagation training performs inference on the above-mentioned problem to be solved, after the input data in the inference process of the deep neural network after backpropagation training for this problem to be solved is transmitted to the in-memory computing module through the data interface, each in-memory computing array performs a step of large-scale vector matrix multiplication calculation. Each in-memory computing array can output the output data of the corresponding network layer in the inference process of the deep neural network after backpropagation training for the problem to be solved and transmit it to the shift accumulator.
[0048] After the shift accumulator performs shift accumulation processing on the output data of each in-memory computing array, it obtains the output data in the inference process of the deep neural network after backpropagation training for this problem to be solved and transmits it to the output register for storage.
[0049] In the embodiments of the present application, the accelerator uses in-memory computing arrays to perform large-scale vector matrix multiplication operations, realizes the calculation acceleration of the weight parameters of each network layer of the deep neural network, greatly reduces the number of data transmissions during the inference process, and enables the deep neural network accelerator to have the characteristics of high system energy efficiency and low operation time complexity.
[0050] Furthermore, this deep neural network process can be used in any deep neural network involving large-scale vector matrix multiplication operations, not limited to solving scientific computing problems. In addition to scientific computing tasks such as solving linear / nonlinear equations, curve least squares fitting problems, and solving steady-state partial differential equations, network features of other task types such as classification and recognition can also be written into the general-purpose processor.
[0051] More preferably, to improve the computing efficiency of the deep neural network accelerator, the weight parameters during the network inference process should not change, so that the accelerator only needs to perform a limited number of writing processes to be used to simulate large-scale vector matrix multiplication operations in the deep neural network. For each set of input data, no additional writing operation is performed on the in-memory computing array, and only one initialization writing operation is performed on the in-memory computing array, thereby increasing the service life of the in-memory computing array.
[0052] The deep neural network accelerator provided by the embodiments of the present application combines the efficient computing ability of the deep neural network with the computing-in-memory characteristics of emerging non-volatile memories. The accelerator includes a general-purpose processor and an in-memory computing module. The in-memory computing module is responsible for accelerating the vector matrix multiplication process of the deep neural network to achieve acceleration of the training process and the inference process; the general-purpose processor is responsible for the forward propagation, backpropagation training, etc. in the training stage of the deep neural network. Since the deep neural network can be widely used in solving scientific computing problems, such as solving linear / nonlinear equations, curve least squares fitting problems, and solving steady-state partial differential equations, the deep neural network accelerator provided by the embodiments of the present application has wide applicability.
[0053] Furthermore, in some embodiments, each in-memory computing array consists of a digital-to-analog converter, at least one memory cell, and an analog-to-digital converter. The digital-to-analog converter is used to convert input data into a voltage vector; each memory cell is used to store the fixed-point number information after the mantissa splitting of any network layer in the deep neural network after backpropagation training according to any precision requirement, and convert the voltage vector into an output current; the analog-to-digital converter is used to convert the output current of each memory cell into a numerical quantity, and determine the output data of any network layer in the inference process of the deep neural network after backpropagation training for the problem to be solved according to the numerical quantity.
[0054] Please further refer to Figure 3 , this in-memory computing array consists of a digital-to-analog converter, at least one memory cell, and an analog-to-digital converter. As Figure 3 shown, each circle represents a memory cell.
[0055] Each storage unit can be used to store the fixed-point number information after the mantissa splitting of any network layer in the deep neural network after backpropagation training according to any precision requirement (such as 4-bit, 8-bit, 16-bit, etc. mentioned above).
[0056] It should be noted that the fixed-point number information after the above mantissa splitting is mapped to the conductance value of the storage unit for storage.
[0057] During the inference process of the deep neural network for the above-mentioned problem to be solved, the input data is converted into a voltage vector by a digital-to-analog converter (DAC) (in the embodiments of the present application, the voltage vector is represented in the form of an array). The elements in the weight parameters of each network layer of the deep neural network after backpropagation training (the weight parameters are also represented in the form of a matrix) are mapped to the conductance values of the storage units and stored in the storage units. The storage units convert the voltage vector into an output current and output the output current to an analog-to-digital converter (ADC). The ADC converts the output current into a numerical quantity, and the output current is the output data of any network layer during the inference process of the deep neural network after backpropagation training for the problem to be solved.
[0058] Further, when the deep neural network after backpropagation training performs inference on the above-mentioned problem to be solved, it is necessary to process the weight parameters of each network layer in the deep neural network after backpropagation training to be written, and the processing method is as Figure 4 shown. When the weight parameters of each network layer in the deep neural network after backpropagation training in floating-point format (such as FP32, including 1-bit sign bit, 8-bit exponent, and 23-bit mantissa bits) are passed from the Cache to the in-memory computing module, the weight matrix elements will be processed by the pre-alignment unit to uniformly align the exponent parts of all elements to the highest value. The specific method is to compare the exponent bits of each floating-point element, identify and extract the maximum exponent value (denoted as Emax). Subsequently, all elements are shifted to the right until the exponent bit becomes Emax. For the insufficient mantissa part after shifting, zero padding is performed, and the mantissa part exceeding the bit width limit is truncated to complete the conversion from floating-point to fixed-point. After the alignment operation is completed, the mantissa part of the weight matrix will be split into several RRAM blocks for deployment, and the deployment bit width of each block can be different. For example, the fixed-point number is split into l 1-bit blocks and n 2-bit blocks, which are respectively mapped to the corresponding storage units of the RRAM array. During subsequent weight calculations, the input data will be multiplied by these RRAM arrays respectively, and the obtained results will be integrated by the shift accumulator and then multiplied by the pre-aligned exponent bit Emax to obtain the calculation result in floating-point format. For different layers of the network, different fixed-point conversion formats can be adopted and different mantissa bit widths can be deployed. For example, the W1 layer is deployed with a bit width of K1 bit, the W2 layer is K2 bit, and the W3 layer is K3 bit, so as to realize the mixed-precision inference of the network.
[0059] When performing the inference process, the input data of the problem to be solved by the backpropagation-trained deep neural network is input into each in-memory computing array through input driving. The calculation results of each in-memory computing array are multiplied by the pre-aligned highest exponent to obtain the output data of each network layer. Finally, after the output data of each network layer is processed by the shift accumulator, the FP32 format output of each network layer is obtained.
[0060] Furthermore, embodiments of the present application can deploy different fixed-point bit widths on the in-memory computing array according to the precision requirements of different network layers and store the corresponding alignment exponents, reducing the in-layer calculation cost and realizing mixed-precision deployment.
[0061] Further, in some embodiments, the storage unit is composed of any one of the following non-volatile memories: Resistive random access memory, phase change memory, spin transfer torque magnetic memory, ferroelectric field effect transistor, and non-volatile flash memory.
[0062] In embodiments of the present application, the storage cell array used in the in-memory computing array is a non-volatile memory array, which is used to perform analog vector matrix multiplication operations. This vector matrix multiplication operation based on the non-volatile memory array has a time complexity of O(1) for matrix operations of any scale, and thus can accelerate the inference process of the deep neural network.
[0063] More specifically, the non-volatile memory can be any one of resistive random access memory (RRAM), phase change memory (PCM), spin transfer torque magnetic memory (STT-MRAM), ferroelectric field effect transistor (FeFET), and non-volatile flash memory (NOR-FLASH).
[0064] In embodiments of the present application, the in-memory computing array utilizes the advantages of high speed, low power consumption, easy integration, and compatibility with complementary metal oxide semiconductor (CMOS) processes of non-volatile memories, stores the weight matrix for conductance operations, and realizes large-scale vector matrix multiplication. In addition, through the combination of the in-memory processing unit and the external control unit, the deep neural network accelerator provided by embodiments of the present application not only has high operation energy efficiency but also maintains high calculation accuracy.
[0065] In order to further improve the operation energy efficiency, the present application further optimizes the deep neural network model to reduce the calculation cost brought by the network scale. Through a limited number of writing processes, the circuit complexity is reduced, data transmission is minimized, and thus the circuit power consumption is reduced. Compared with the traditional method of using a computer to solve scientific computing problems, the deep neural network accelerator of the present application can effectively reduce the time complexity, realize the integration of storage and calculation, greatly save operation energy consumption and time, and ensure high reliability at the same time.
[0066] In addition, the design of the deep neural network accelerator can be compatible with various non-volatile memories such as resistive random access memory (RRAM), phase change memory (PCM), non-volatile flash memory, spin-transfer torque magnetic memory (STT-MRAM), ferroelectric field-effect transistor (FeFET), etc., expanding the application of these memories in the field of numerical computing. This feature enables the deep neural network accelerator of the embodiments of the present application to adapt to different application requirements and technological developments, providing a new solution for future scientific computing.
[0067] The in-memory computing array design in the embodiments of the present application allows for low-power operation while maintaining high computing density, which is particularly important for portable devices and Internet of Things (IoT) applications. Its high reconfigurability means that the same hardware platform can be quickly adjusted for different application scenarios, thereby improving the flexibility and adaptability of the system. By supporting multiple non-volatile memories, not only is the computing efficiency improved, but new possibilities are also provided for the diversification and integration of memory technologies, pushing the development frontier of storage computing technology.
[0068] Exemplarily, this embodiment demonstrates a deep operator network accelerator based on in-memory computing, and the non-volatile memory array used is a resistive random access memory (RRAM) array. The deep operator network computing accelerator includes: a central processor mainly composed of a CPU / GPU, an FPGA module, and an in-memory computing module. Among them, the in-memory computing module is composed of an in-memory control unit, a data interface, a pre-alignment unit, an input register, an input driver, an in-memory computing array, an output driver, a shift accumulator, and an output register.
[0069] In this embodiment, the core of the in-memory computing module is multiple non-volatile memory arrays, and this array calculates large-scale vector matrix multiplication in one step based on Ohm's law and Kirchhoff's law. Specifically, the process of the in-memory computing array performing vector matrix multiplication is as Figure 5 shown. Denote the matrix corresponding to the weight parameters stored in the non-volatile memory array in the initial state as M, and the storage unit in the non-volatile memory array is M ij where 1 and 1 , , are the number of rows and columns of matrix M, then the storage method of the weight parameters in the non-volatile memory array is , G is the conductance value of the non-volatile memory. After the data vector transmitted by the in-memory control unit to the in-memory computing array is converted into a voltage vector by a digital-to-analog converter, it is input into the non-volatile memory array M. On one in-memory computing array, define I as the output current of the in-memory computing array, Vis the voltage vector input to the digital-to-analog converter. According to Ohm's law I = V·G it can be determined that for each memory cell in the non-volatile memory array M M ij , an electric current will be output. Then, according to Kirchhoff's current law, it can be determined that the output current of each row of non-volatile memories in the non-volatile memory array is the sum of the currents of each memory cell ( M i1 , M i2 , … , M in ) on that row, that is, the output current of the th memory cell in each row : . Thus, a series of output currents can be obtained on the row lines of the in-memory computing array, forming an m-dimensional current vector. The current vector is then converted into a data vector (i.e., a numerical quantity) via an analog-to-digital converter, and this data vector is the result of the vector multiplication calculation of this layer of the network layer. By adding the input driver and the output driver, i.e., the digital-to-analog converter and the analog-to-digital converter, to the input side and the output side of the in-memory computing array, the basic structure of the in-memory computing array shown in Figure 3 is obtained.
[0070] The hardware structure of the deep neural network accelerator provided in the embodiments of this application during the training phase is as shown in Figure 6 . During this phase, the deep operator network is deployed on the CPU / GPU. The deep operator network includes a backbone network and a branch network. Among them, both the backbone network and the branch network adopt a fully connected network structure (for example, FC-1 is the backbone network and FC-2 is the branch network), and the in-memory computing module deploys a fixed weight matrix for accelerating training (i.e., the training acceleration matrix). During each iteration of training, first, the CPU / GPU calculates the output results of each network layer in the deep operator network forward, and the FPGA (as a data interface) performs linear / non-linear transformation on the output results and calculates the error information. The error information is preprocessed by the FPGA and then input to the in-memory computing module. The in-memory computing module performs one-step calculation to obtain the error information propagated to each network layer, which is transmitted to the CPU / GPU through the FPGA to replace the gradient calculation in the backpropagation training for the backpropagation process. Finally, the CPU / GPU updates the parameter information of the internal deep operator network to complete the training iteration process of this deep operator network.
[0071] The hardware structure of the deep operator network accelerator provided in the embodiments of this application during the training phase is as shown in Figure 7As shown in the figure, the trained deep operator network is deployed on the in-memory computing module. When performing an inference, the FPGA receives the input data. The main preprocessing and branch preprocessing respectively match the input data of the backbone network and the branch network of the deep operator network, and transmit the output after the above-mentioned shift alignment processing to the in-memory computing module. After the in-memory computing module completes the vector matrix multiplication in parallel, it transmits the calculation result to the FPGA for linear transformation. Repeat the above process until the calculations of the backbone network and each branch network in the deep operator network are completed. Then, the FPGA performs a non-linear activation dot product operation on the output vectors of the backbone network and the branch network, and the operation result is the output of the current batch of inference process.
[0072] In this embodiment, the deployed deep operator network structure is as Figure 8 shown. The deep operator network (i.e., the physical model in Figure 8 ) is mainly used to solve scientific computing problems such as ordinary differential equations (ODEs) and partial differential equations (PDEs). The main body of the network has a backbone-branch structure. The backbone fully-connected network and the single-branch fully-connected network structure are shown in this embodiment. For the partial differential equation to be solved, the backbone network describes the main structure of the equation, and the branch network reads the sampling parameters and other parameter conditions of the equation. After the backbone network and the branch network respectively infer the vector results, a dot product operation is performed on the two vectors to obtain the calculation result of the partial differential equation under the specified sampling conditions.
[0073] The operation method of the deep operator network accelerator includes the following steps: S1. Determine the scientific computing problem to be solved, select a suitable training structure and scale of the deep operator network, initialize and deploy the backbone network part and the branch network part of the deep operator network by the CPU / GPU, and initialize the in-memory computing module to be fixed as the training acceleration matrix.
[0074] S2. Transmit the input data for the forward propagation of the deep operator network for the scientific computing problem to be solved to the input layer of the backbone network in the CPU / GPU. After forward propagation through the hidden layer, transmit the output layer data to the FPGA.
[0075] S3. The FPGA calculates the error information between the output data and the expected output during the forward propagation of the deep operator network for the scientific computing problem to be solved, and inputs the error information into the in-memory computing module. The in-memory computing module performs fixed calculation expansion on the error information to obtain the error information matrix of each network layer, and transmits the expanded error information matrix to the FPGA for linear / non-linear transformation.
[0076] S4. The CPU / GPU receives the error information matrix processed by the FPGA, and uses the error information matrix as the gradient of the deep operator network for weight iteration update.
[0077] S5. Repeat steps S2 - S4 until the training error of the backbone network is lower than the target value, and the in - memory computing module completes the pre - training of the accelerated deep operator network. Fix the weight part of the backbone network, and perform backpropagation training on the branch network by the CPU / GPU and FPGA until the training accuracy of the deep operator network reaches the expectation. Thus, the training part of the deep operator network is completed, and the trained deep operator network is deployed on the in - memory computing module.
[0078] S6. When performing the network inference process, transfer the input data of the backbone network and the branch network of the current layer to the FPGA for pre - processing. After pre - processing, input it into the in - memory computing module for one - step calculation of large - scale vector - matrix multiplication to obtain the output data of the current layer.
[0079] S7. The FPGA performs linear transformation and non - linear activation on the output data of the current layer to obtain the output result of the current layer.
[0080] S8. Repeat steps S6 - S7 until the inference of both the backbone network and the branch network of the deep operator network is completed, obtaining the output vectors of the backbone network and each branch network. The FPGA performs dot - product operations on the output vectors of the backbone network and the branch network to obtain the output result of the current batch. Thus, the inference process of the deep operator network for the current batch is completed.
[0081] Please continue to refer to Figure 8 , when using this solver to solve the steady - state partial differential equation, assume that the deep operator network has only one backbone network and one branch network, the network type is a fully - connected neural network, and the network input is data pairs, , , where for each group of data pairs , , represents the independent variable of the partial differential equation, is the operator at a vector composed of the parameter variables and sampling results of when performing this inference, is the dependent variable of the partial differential equation, and
[0082] Furthermore, the specific training process of the deep operator network accelerator is as Figure 9 shown: Among them, (a) represents performing pre - training on the backbone network using the input - output pairs of the training dataset. In this process, the part where the deep operator network calculates the gradient matrix according to the error matrix is accelerated by the in - memory computing module; (b) after the backbone network training is completed, fix the pre - trained backbone network part, and use the sensing parameters Perform further iterative backpropagation (BP) training on the branch network, where represents the dataset sampling parameter, , . Perform a dot product operation on the output vector of the branch network and the output vector of the backbone network to achieve a more accurate solution.
[0083] Furthermore, the specific inference process of the deep operator network accelerator is as Figure 10 shown: (1) Determine that the core operator of the problem to be solved is the stationary partial differential equation ; (2) Set the network type. Here, different branches of the deep neural network are all set to fully connected networks (FNN-DeepONet); (3) Set the input of the current layer, including network layer calculation conditions such as the input format of the current layer and hyperparameters; (4) Simulate and calculate the large-scale vector matrix multiplication involved in the current layer, which is implemented by the in-memory computing module; (5) Determine whether the inference of the current backbone or branch network is completed. If not, perform non-linear activation processing on the output of the current layer using a linear activation function and pass it as the input of the next layer network; (6) Repeat the above process (4)-(5) until both the backbone / branch network have completed the calculation. The FPGA calculates the dot product result of the output of the backbone network and the output of the branch network, which is the output solution of the deep neural network.
[0084] Furthermore, when the network type is a convolutional neural network (CNN), a recurrent neural network (RNN), or other deep neural networks, the above algorithm flow only generates additional data transformation operations during inter-layer data transmission and does not generate additional operations in the in-memory computing array, so it does not affect the high energy efficiency characteristics of the accelerator.
[0085] In this application, the accelerator achieves parallel training and inference of deep neural networks by using the memory array to perform large-scale vector matrix operations, in cooperation with a general-purpose processor and an in-memory computing module, so as to achieve the purpose of efficiently solving equations. And it can accelerate the execution of various scientific computing tasks without relying on specific network features under the condition of unchanged computing architecture and standard sampling format, realizing the general training and inference acceleration of deep neural networks, and can reduce the computational complexity when the deep neural network fits the operator and improve the computing energy efficiency.
[0086] Next, the acceleration method of the deep neural network accelerator based on in-memory computing provided by this application will be described. The acceleration method of the deep neural network accelerator based on in-memory computing described below can be executed on the deep neural network accelerator based on in-memory computing described above.
[0087] Please refer further to Figure 11, an embodiment of the present application provides an acceleration method for a deep neural network accelerator based on in-memory computing, including: step 210 and step 220.
[0088] Step 210 is based on a general-purpose processor. According to the problem to be solved, a deep neural network is deployed, and the output data during the forward propagation process of the deep neural network for the problem to be solved is determined. Step 220 is based on the in-memory computing module. According to the first error information between the output data during the forward propagation process and the expected output, the first gradient parameter of the first weight parameter of each network layer in the deep neural network is determined, so that the general-purpose processor iteratively updates the first weight parameter according to the first gradient parameter until the deep neural network completes pre-training.
[0089] Further, in some embodiments, the method further includes: Based on the general-purpose processor, according to the second error information between the output data during the forward propagation process of the pre-trained deep neural network for the problem to be solved and the expected output, the second gradient parameter of the updated first weight parameter of each network layer is determined, and the updated first weight parameter is iteratively updated according to the second gradient parameter until the deep neural network completes backpropagation training.
[0090] The acceleration method for the deep neural network accelerator based on in-memory computing provided by the embodiment of the present application uses a general-purpose processor to determine the output data during the forward propagation process of the deep neural network for the problem to be solved, and uses the in-memory computing module to determine the first gradient parameter of the first weight parameter of each network layer according to the first error information between the output data and the expected output, so that the general-purpose processor updates the first weight parameter according to the first gradient parameter, thereby realizing the acceleration of the pre-training process of the deep neural network without depending on the type of the deep neural network model and improving the training efficiency of the deep neural network. At the same time, since the deep neural network can be widely applied to the solution of scientific computing problems, such as solving linear / nonlinear equations, curve least squares fitting problems, solving steady-state partial differential equations, etc., the deep neural network accelerator of the present application has a certain generality.
[0091] It can be understood that the detailed function implementation of the above-mentioned various units / modules can be referred to the introduction in the foregoing method embodiments, and will not be elaborated here.
[0092] It should be understood that the above-mentioned device is used to execute the method in the above-mentioned embodiment. For the corresponding program module in the device, its implementation principle and technical effect are similar to the description in the above-mentioned method. The working process of the device can refer to the corresponding process in the above-mentioned method, and will not be elaborated here.
[0093] Based on the method in the above-mentioned embodiment, an embodiment of the present application provides an electronic device, please refer toFigure 12 , the electronic device may include: a processor 1210, a communications interface 1220, a memory 1230, and a communication bus 1240. Among them, the processor 1210, the communications interface 1220, and the memory 1230 communicate with each other through the communication bus 1240. The processor 1210 can call the logical instructions in the memory 1230 to execute the methods in the above embodiments.
[0094] In addition, when the logical instructions in the above-mentioned memory 1230 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application.
[0095] Based on the method in the above embodiments, an embodiment of this application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program runs on a processor, it causes the processor to execute the method in the above embodiments.
[0096] Based on the method in the above embodiments, an embodiment of this application provides a computer program product. When the computer program product runs on a processor, it causes the processor to execute the method in the above embodiments.
[0097] It can be understood that the processor in the embodiments of this application may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0098] The method steps in the embodiments of this application can be implemented in a hardware manner or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in a random access memory (RAM), flash memory, read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), registers, hard disks, removable hard disks, CD-ROMs, or any other form of storage medium well-known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.
[0099] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0100] It can be understood that the various numerical numbers involved in the embodiments of this application are only for the convenience of description and are not used to limit the scope of the embodiments of this application.
[0101] Those skilled in the art can easily understand that the above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A deep neural network accelerator based on in-memory computing, characterized in that: include: A general purpose processor, configured to deploy a deep neural network according to a problem to be solved, and determine output data of the deep neural network during forward propagation of the problem to be solved; The in-memory computing module is used to determine the first gradient parameter of the first weight parameter of each network layer in the deep neural network according to the first error information between the output data and the expected output in the forward propagation process, so that the general processor iteratively updates the first weight parameter according to the first gradient parameter until the deep neural network completes pre-training.
2. The deep neural network accelerator based on in-memory computing according to claim 1, characterized in that: The general processor is further used to determine the second gradient parameters of the updated first weight parameters of each network layer according to the second error information between the output data and the expected output in the forward propagation process of the pre-trained deep neural network for the problem to be solved, and iteratively update the updated first weight parameters according to the second gradient parameters until the deep neural network completes the back-propagation training.
3. The deep neural network accelerator based on in-memory computing according to claim 2, characterized in that: The general purpose processor comprises: A cache for storing output data of a pre-trained deep neural network during the forward propagation process for the problem to be solved; The near-memory processor is used to determine a second gradient parameter based on second error information between output data and expected output during forward propagation of the pre-trained deep neural network for the problem to be solved, and update the updated first weight parameter based on the second gradient parameter until the deep neural network completes back-propagation training.
4. The deep neural network accelerator based on in-memory computing according to claim 2 or 3, characterized in that: The in-memory computing module is also used to determine the output data of the reasoning process according to the output data of each network layer in the reasoning process of the deep neural network after back-propagation training for the problem to be solved.
5. The deep neural network accelerator based on in-memory computing according to claim 4, characterized in that: The in-memory computing module comprises: a pre-alignment unit, an input register, at least one in-memory computation array, and a shift accumulator; The pre-alignment unit is used to convert the weight parameters of each network layer in the deep neural network after back-propagation training from floating point to fixed point format, and obtain the fixed-point information corresponding to the weight parameters of each network layer in the deep neural network after back-propagation training; The input register is used to split the mantissa in the fixed-point number information according to different precision requirements for each network layer; Each in-memory computing array is used to store fixed-point number information after decomposition of the mantissa of each network layer according to different precision requirements, and input data of the deep neural network after back-propagation training in the process of reasoning for the problem to be solved, and determine the output data of each network layer in the process of reasoning; The shift accumulator is used to determine the output data of the reasoning process according to the output data of each network layer in the reasoning process.
6. The deep neural network accelerator based on in-memory computing according to claim 5, characterized in that: Each of the in-memory computing arrays is composed of a digital-to-analog converter, at least one storage unit and an analog-to-digital converter, wherein the digital-to-analog converter is used to convert the input data into a voltage vector; each storage unit is used to store the fixed-point number information after the mantissa is split according to any precision requirement of any network layer in the deep neural network after back-propagation training, and convert the voltage vector into an output current; the analog-to-digital converter is used to convert the output current of each storage unit into a numerical value, and determine the output data of any network layer in the process of reasoning of the deep neural network after back-propagation training for the problem to be solved based on the numerical value.
7. The deep neural network accelerator based on in-memory computing according to claim 6, characterized in that: The storage unit is composed of any of the following non-volatile memories: Resistive memory, phase change memory, spin transfer torque magnetic memory, ferroelectric field effect transistor and non-volatile flash memory.
8. An acceleration method for a deep neural network accelerator based on in-memory computing according to any one of claims 1 to 7, characterized in that: include: Based on a general-purpose processor, a deep neural network is deployed according to the problem to be solved, and output data of the deep neural network during forward propagation of the problem to be solved is determined; Based on the in-memory computing module, according to the first error information between the output data and the expected output in the forward propagation process, the first gradient parameter of the first weight parameter of each network layer in the deep neural network is determined, so that the general processor iteratively updates the first weight parameter according to the first gradient parameter until the deep neural network completes pre-training.
9. The acceleration method of the deep neural network accelerator based on in-memory computing according to claim 8, characterized in that: The method further comprises: Based on a general-purpose processor, according to the second error information between the output data and the expected output in the forward propagation process of the pre-trained deep neural network for the problem to be solved, the second gradient parameters of the updated first weight parameters of each network layer are determined, and the updated first weight parameters are iteratively updated according to the second gradient parameters until the deep neural network completes the back-propagation training.
10. An electronic device, characterized in that: include: at least one memory for storing a computer program; At least one processor is used to execute the program stored in the memory, when the program stored in the memory is executed, the processor is used to execute the method according to claim 8 or 9.
Citation Information
Cited By
Neural network acceleration system and method, equipment, storage medium and program
CN122311317A