Neural network accelerator-oriented design method of calculation and storage fusion novel architecture
By designing a new computing and storage fusion architecture in a neural network accelerator, using weight quantization and a single input dedicated multiplier, the energy efficiency problems caused by data movement are solved, and efficient computing performance and energy efficiency balance are achieved.
Patent Information
- Application Number
- CN202510217769.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-26
AI Technical Summary
The prior art is difficult to ensure the computing performance of neural network accelerators while reducing data movement and improving computing energy efficiency.
A new architecture for computing and storage fusion for neural network accelerators is proposed. Through weight and activation quantization, designing single-input dedicated multiplier, weight analysis and data path hardware design, the storage logic calculation unit of the weight matrix is realized to reduce data movement.
It effectively reduces data movement in neural network accelerator, improves computing energy efficiency, and avoids weight-related memory access operations while ensuring computing performance.
Smart Images

Figure CN120146122A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of neural network processors, and particularly relates to a design method for a novel architecture of computing-in-memory integration for a neural network accelerator. Background Art
[0002] Neural Network (NN) technology has demonstrated superior capabilities in tasks such as natural language processing, machine vision, image and video generation. It can not only bring great convenience to human life but also assist humans in completing complex tasks and accelerating the development of various industries. However, under the current state of the art, the large parameter scale, substantial computational resources, and data access requirements pose significant challenges in terms of energy consumption and computing power, hindering the further development and application of neural network technology. Therefore, the industry urgently needs to seek new computing architectures with high energy efficiency to significantly improve computing energy efficiency while ensuring the performance of intelligent tasks as much as possible.
[0003] To improve the energy efficiency of intelligent computing systems, the field has chosen to use dedicated hardware accelerators instead of general-purpose Central Processing Units (CPUs) and Graphics Processing Units (GPUs) to undertake intelligent computing tasks. Hardware accelerators for intelligent computing generally adopt optimization means of Application Specific Integrated Circuits (ASICs) such as multi-core processing, pipelining of the computing process, and approximate circuit design, cooperate with technologies such as pruning and quantization to compress the model, and combine optimization of storage strategies and software-hardware co-design methods to improve energy efficiency, which has significant performance and power consumption optimizations compared to the general computing platforms of CPUs and GPUs.
[0004] However, these optimization means cannot effectively reduce the energy consumption overhead caused by a large amount of data movement. In contrast, a large amount of work has been done on in-memory computing and near-memory computing technologies, aiming to perform computing operations inside or near the memory or storage unit to reduce the energy consumption cost and performance loss caused by data movement. However, in-memory computing is limited by the process and scale of the memory, and its computing performance and accuracy are very limited; near-memory computing only shortens the distance of movement and still cannot avoid data movement. Summary of the Invention
[0005] The problem to be solved by the present invention is to reduce data movement and improve computing energy efficiency while ensuring the computing performance of the NN accelerator, and a design method for a novel architecture of computing-in-memory integration for a neural network accelerator is proposed.
[0006] To achieve the above object, the present invention is realized through the following technical solutions:
[0007] A design method for a new type of computing-in-memory fusion architecture for neural network accelerators, comprising the following steps:
[0008] S1. Weight and activation quantization: Using an open-source weight quantization method, on a software platform built on an open-source framework based on PyTorch, quantize the weights of the neural network and the input of each layer to obtain a quantized neural network;
[0009] S2. Design a single-input dedicated multiplier: Based on the quantized neural network obtained in step S1, design a single-input dedicated multiplier for the quantized neural network accelerator. When the quantized NN accelerator performs the multiplication operation of weights and inputs, both inputs used are multipliers with fixed-point integers. Fix one input of the multiplier to a specific multiplier to obtain a single-input dedicated multiplier;
[0010] S3. Weight analysis: For the quantized NN obtained in step S1, further analyze the weights on the software platform used in step S1, and count the number of different elements in the weight matrix to determine the number of different single-input dedicated multipliers used in a single calculation operation;
[0011] S4. Hardware design of the data path: First, complete the declaration of the module name, input, and output ports in the hardware description file on the software platform used in step S1; then use the software platform used in step S1 to traverse the weight matrix again. According to the arrangement of the weights in the weight matrix, generate corresponding connection statements according to the corresponding relationship during matrix multiplication and write them into the hardware description file; after the weights are completely traversed, supplement the statements for establishing the end module in the hardware description file to complete the hardware design work of the data path;
[0012] S5. Matrix multiplication operation of the NN accelerator with a computing-in-memory fusion architecture: The input data of each layer in the NN is transmitted to the specified single-input dedicated multiplier through the data path constructed in step S4, and then the results of the multiplication operation are accumulated by the adder according to the operation rules of matrix multiplication. Finally, the results of the matrix multiplication operation of the NN accelerator with a computing-in-memory fusion architecture are obtained.
[0013] Further, the matrix multiplication operation inside the quantized neural network in step S1 contains multiplication operations between low-bit fixed-point numbers.
[0014] Further, the range of the multiplier of the single-input dedicated multiplier in step S2 depends on the number of bits of weight quantization in step S1.
[0015] Further, the method for counting the number of different elements in the weight matrix in step S3 is to traverse all elements in the weight matrix and increment the corresponding counting variable according to the value of the element.
[0016] Further, in step S3, when the weights of the NN are the weights of the fully connected layer and the entire layer uses a large weight matrix, the weight matrix is analyzed column by column at this time; when the weights of the NN are in the form of a convolutional layer in step S3, each input channel corresponds to a convolutional kernel, and the weights of the convolutional layer are analyzed channel by channel at this time.
[0017] Further, in step S4, the hardware design of the data path realizes using wires to represent the correspondence from the input data to the single-input dedicated multiplier.
[0018] Advantages of the present invention:
[0019] The design method of a novel architecture for arithmetic-memory fusion for a neural network accelerator described in the present invention is a novel computing architecture that stores neural network weights in the form of single-input specific multipliers and data paths in the logic computing unit. The weights are placed in the structure of the logic computing circuit, which belongs to a hardware design method. It avoids loading the weight matrix from an external memory, eliminates the memory access operations related to the weights, and greatly reduces the data movement in the neural network accelerator.
[0020] The design method of a novel architecture for arithmetic-memory fusion for a neural network accelerator described in the present invention proposes to store the weight matrix in the computing unit. While avoiding a large amount of data movement, it ensures that the computing efficiency of the arithmetic-memory fusion computing component is close to the level of the currently commonly used logic operation units. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a flowchart of the design method of a novel architecture for arithmetic-memory fusion for a neural network accelerator described in the present invention;
[0022] Figure 2 is a schematic structural diagram of the single-input dedicated multiplier of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0023] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention, that is, the specific embodiments described are only a part of the embodiments of the present invention, rather than all of the specific embodiments. Usually, the components of the specific embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations, and the present invention can also have other embodiments.
[0024] Accordingly, the following detailed description of the specific embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected specific embodiments of the present invention. All other specific embodiments obtained by those skilled in the art based on the specific embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0025] To further understand the content, features and effects of the present invention, the following specific embodiments are exemplified and described in detail in conjunction with the attached Figure 1 and the attached Figure 2 as follows:
[0026] Example 1:
[0027] A design method for a novel architecture of arithmetic and memory fusion for a neural network accelerator, comprising the following steps:
[0028] S1. Weight and activation quantization: Using an open-source weight quantization method, on a software platform built on an open-source framework based on PyTorch, quantize the weights of the neural network and the input of each layer to obtain a quantized neural network;
[0029] Further, the matrix multiplication operation inside the quantized neural network in step S1 involves multiplication operations between low-bit fixed-point numbers;
[0030] Further, this embodiment needs to design a computing architecture for the quantized NN;
[0031] S2. Design a single-input dedicated multiplier: Based on the quantized neural network obtained in step S1, design a single-input dedicated multiplier for the quantized neural network accelerator. When the quantized NN accelerator performs the multiplication operation of weights and inputs, both inputs of the multiplier are fixed-point integers. Fix one input of the multiplier to a specific multiplier to obtain a single-input dedicated multiplier;
[0032] Further, the range of the multiplier of the single-input dedicated multiplier in step S2 depends on the number of bits of weight quantization in step S1;
[0033] Further, the single-input dedicated multiplier specifically processes operations such as multiplying by 2, multiplying by 3, etc. Thus, single-input dedicated multipliers with different fixed multipliers can be obtained. The range of the multiplier depends on the number of bits of weight quantization in S1. For example, for weights quantized with 3 bits, the numerical range is limited to integers within the interval (-8, 8]. Then, the integers between -7 and 8 are respectively fixed as one input of the multiplier, where 0 and 1 do not need to be fixed into the multiplier. Thus, 14 multipliers will be obtained. Similarly, for weights quantized with n bits, the integers from -(2 n - 1) to 2 nAn integer other than 0 and 1 is fixed as one input of the multiplier, resulting in (2 n+1 -2) multipliers.
[0034] S3. Weight analysis: For the quantized NN obtained in step S1, further analyze the weights on the software platform used in step S1, and count the number of different elements in the weight matrix to determine the number of different single-input dedicated multipliers used in a single calculation operation;
[0035] Furthermore, the method for counting the number of different elements in the weight matrix in step S3 is to traverse all elements in the weight matrix and increment the corresponding counting variable according to the value of the element;
[0036] Furthermore, taking quantization by 3 as an example, a total of 14 counting variables n -7 to n 8 (without n 0 and n 1 ) are set. During the process of traversing the weight matrix, if the element in the matrix is 2, execute (n 2 +1), and so on. Since the range of the number of weights of the quantized NN is limited, the scale of this statistics is also limited.
[0037] Furthermore, in step S3, when the weights of the NN are the weights of the fully connected layer and a large weight matrix is used for the entire layer, the weight matrix is analyzed column by column at this time; in step S3, when the weights of the NN are in the form of a convolutional layer and each input channel corresponds to a convolutional kernel, the weights of the convolutional layer are analyzed channel by channel at this time;
[0038] S4. Hardware design of the data path: On the software platform used in step S1, first complete the declaration of the module name, input and output ports in the hardware description file; then use the software platform used in step S1 to traverse the weight matrix again, and generate the corresponding connection statements according to the arrangement of the weights in the weight matrix and the corresponding relationship during matrix multiplication, and write them into the hardware description file; after the weights are completely traversed, supplement the statements for establishing the end module in the hardware description file to complete the hardware design work of the data path;
[0039] Furthermore, the hardware design of the data path in step S4 realizes using connections to represent the corresponding relationship from the input data to the single-input dedicated multiplier;
[0040] S5. Matrix multiplication operation of the NN accelerator with the compute-storage fusion architecture: The input data of each layer in the NN is transmitted to the specified single-input dedicated multiplier through the data path constructed in step S4, and then the results of the multiplication operations are accumulated by the adder according to the operation rules of matrix multiplication, and finally the results of the matrix multiplication operation of the NN accelerator with the compute-storage fusion architecture are obtained.
[0041] Furthermore, compared with the method of loading weights and input data from the memory and then passing these two parts of data into the matrix multiplication operation unit adopted by traditional NN accelerators, the novel computing architecture proposed in this embodiment only needs to pass in the input data and does not need to pass in the weight matrix data, avoiding large-scale data movement.
[0042] A design method of a novel architecture for arithmetic-memory fusion oriented to neural network accelerators described in this embodiment designs an arithmetic-memory fusion NN accelerator in a 28nm process library and uses Cadence Genus 15.0 software for evaluation. For the arithmetic-memory fusion NN accelerator, compared with a general NN accelerator that loads weights from external memory, the energy consumption of data movement is reduced by up to 43.52%-89.20%. Compared with the same type of methods (in-memory computing, near-memory computing), the comparison results are shown in Table 1. It can completely eliminate a part of data movement while avoiding the computing performance and accuracy problems brought by using novel storage devices.
[0043] Table 1 Comparison table of reduced data movement
[0044]
[0045] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the said element.
[0046] Although the present application has been described above with reference to specific embodiments, various improvements can be made to it and components therein can be replaced with equivalents without departing from the scope of the present application. In particular, as long as there is no structural conflict, the various features in the specific embodiments disclosed in the present application can be combined with each other in any way, and the exhaustive description of these combinations is omitted in this specification only for the sake of saving space and resources. Therefore, the present application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
Claims
1. A design method for a new computing and storage fusion architecture for a neural network accelerator, characterized in that: The steps include: S1. Weight and activation quantization: Using an open source weight quantization method, the weights of the neural network and the input of each layer are quantized on a software platform built on the open source framework of PyTorch to obtain a quantized neural network. S2. Design a single-input dedicated multiplier: Based on the quantized neural network obtained in step S1, a single-input dedicated multiplier is designed for the quantized neural network accelerator. When the quantized NN accelerator performs the multiplication operation of weights and inputs, a multiplier whose two inputs are both fixed-point integers is used. One input of the multiplier is fixed to a specific multiplier to obtain a single-input dedicated multiplier. S3. Weight analysis: For the quantized NN obtained in step S1, further analyze the weights on the software platform used in step S1, and count the number of different elements in the weight matrix to determine the number of different single-input dedicated multipliers used in a single calculation operation; S4. Hardware design of data path: First, on the software platform used in step S1, the module name, input and output port declarations are completed in the hardware description file; then, the software platform used in step S1 is used to re-traverse the weight matrix, and according to the arrangement of weights in the weight matrix and the corresponding relationship during matrix multiplication, the corresponding connection statements are generated and written into the hardware description file; After the weight traversal is complete, add the statement to end the module establishment in the hardware description file to complete the hardware design of the data path; S5. Matrix multiplication operation of the NN accelerator with a computing-storage fusion architecture: The input data of each layer in the NN is transmitted to the designated single-input dedicated multiplier through the data path constructed in step S4, and then the result of the multiplication operation is accumulated by the adder according to the operation rules of matrix multiplication, and finally the result of the matrix multiplication operation of the NN accelerator with a computing-storage fusion architecture is obtained.
2. The design method of a novel computing and storage fusion architecture for a neural network accelerator according to claim 1 is characterized in that: The matrix multiplication operation of the weights and inputs in the quantized neural network in step S1 internally includes multiplication operations between low-bit specific points.
3. The design method of a novel computing and storage fusion architecture for a neural network accelerator according to claim 2 is characterized in that: The range of the multiplier of the single-input dedicated multiplier in step S2 depends on the number of bits used to quantize the weight in step S1.
4. The design method of a novel computing and storage fusion architecture for a neural network accelerator according to claim 3 is characterized in that: The method for counting the number of different elements in the weight matrix in step S3 is to traverse all elements in the weight matrix and add 1 to the corresponding counting variable according to the value of the element.
5. The design method of a novel computing and storage fusion architecture for a neural network accelerator according to claim 4 is characterized in that: In step S3, when the weights of the NN are the weights of the fully connected layer, the whole layer uses a large weight matrix, and the weight matrix is analyzed column by column at this time; in step S3, when the weights of the NN are in the form of a convolutional layer, each input channel corresponds to a convolution kernel, and the weights of the convolutional layer are analyzed channel by channel at this time.
6. The design method of a novel computing and storage fusion architecture for a neural network accelerator according to claim 5 is characterized in that: The hardware design of the data path in step S4 implements the correspondence from input data to the single-input dedicated multiplier using wires.
Citation Information
Patent Citations
Neuron acceleration processing method and device, equipment and readable storage medium
CN112906863A
Neural network acceleration system based on logarithmic block floating point quantization
CN114626516A
Sharing-based single-input multi-weight multiplier for reduced approximation
CN116888575A
Convolutional neural network hierarchical pipeline accelerator generation method and device
CN117952164A
Neural network accelerator and neural network acceleration method based on structured pruning and low-bit quantization
US20220012593A1