Bit-slicing accelerator for efficient convolutional neural network inference
By designing a bit slice accelerator based on lookup table LUT in the DNN inference accelerator, the problem that the fixed accuracy calculation unit in the prior art is not able to adapt to variable accuracy scenarios, efficient resource and energy consumption utilization is achieved, and inference efficiency and adaptability are improved.
Patent Information
- Application Number
- CN202510173871.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2045-02-18
AI Technical Summary
Existing DNN inference accelerators are inefficient in dealing with variable accuracy scenarios because they generally use fixed-precision computing units and cannot flexibly adapt to the needs of different bit widths and sparse operations, resulting in limited resource and energy consumption efficiency.
A bit slice accelerator based on the lookup table LUT is designed, including a bit operation analysis unit BMAU and a bit-level multiplier unit. The bit-level multiplication operation is managed through the LUT mechanism, supports operations of different bit widths, and optimizes the calculation overhead through parallel calculation and pipeline operations.
It significantly improves the efficiency and energy efficiency of DNN inference in resource and energy consumption constrained environments, enhances the adaptability and overall performance of hardware accelerators, and can flexibly adapt to different levels of computing needs.
Smart Images

Figure CN119647538B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of neural network inference hardware, and particularly to a bit-slice accelerator for efficient convolutional neural network inference. Background Art
[0002] Deep Neural Networks (DNNs) exhibit excellent computational efficiency in current various artificial intelligence applications, and can significantly reduce computational requirements. Currently, the state-of-the-art DNNs still need to perform billions of operations and require a large amount of storage space to store activation values and weights. To address these challenges, existing technologies have gradually turned their attention to customized bit-slice accelerators to improve computational performance and energy efficiency. Bit-slice processing classifies multiplication operations into four categories: invalid operations, redundant operations, basic operations, and repetitive operations to achieve more efficient computing. Collectively, these innovations aim to optimize the deployment of DNNs so that they can be more widely applied in various restricted environments. Using variable precision for different convolutional layers in DNNs can effectively reduce memory and computational costs. Therefore, a DNN accelerator that can adapt to such precision changes is needed. Bit-slice accelerators improve processing efficiency by leveraging data sparsity at the bit level while maintaining computational accuracy. On the one hand, when an input of zero is detected, power consumption is reduced by gating the computing units, but this method does not effectively improve latency or throughput; on the other hand, some architectures improve latency and throughput by exploiting sparsity in compressed inputs, but usually only support sparsity of a single operand and cannot handle sparsity of two operands simultaneously. Previous DNN inference accelerators have shown insufficient efficiency in dealing with variable precision scenarios because they generally adopt fixed-precision computing units. Variable precision requirements between different levels need an accelerator with sufficient flexibility to reduce memory and computational overhead. Therefore, solving these challenges is crucial for improving the efficiency and adaptability of DNN inference accelerators under heterogeneous precision conditions.
[0003] To address this issue, a potential solution is to more effectively combine value-level sparsity and bit-level sparsity, thereby reducing unproductive bits in deep learning models. However, bit-level sparsity varies among computing units, leading to a large amount of ineffective computations during parallel processing. Therefore, precision-scalable accelerators developed through algorithm-hardware co-design represent a promising approach to address this irregularity. Quantizing deep learning models with over 6.7 billion parameters poses challenges in maintaining accuracy during inference. This makes it particularly important to ensure computational efficiency while maintaining model accuracy during inference. Algorithm design should manage the irregular distribution of zero bits in operands, while the hardware simplifies control logic and develops an efficient sparse indexing mechanism. Algorithm-hardware co-design ensures broader applicability and minimizes computational overhead. It can be seen that the prior art has the following problems in DNN inference: 1) Resource and energy consumption limitations. On hardware platforms with limited resources and energy consumption, it is difficult to effectively utilize data sparsity in DNNs, thus restricting the improvement of inference efficiency and energy efficiency; 2) Lack of flexibility. Most traditional hardware accelerators use fixed-precision computing units and lack the ability to adapt to different bit widths and sparse operations, and cannot flexibly configure computing resources to meet the requirements of different network layers. Summary of the Invention
[0004] An object of an embodiment of the present invention is to provide a bit-slice accelerator for efficient convolutional neural network inference, which can solve the above problems existing in DNN inference.
[0005] To achieve the above object, an embodiment of the present invention provides a bit-slice accelerator for efficient convolutional neural network inference. The bit-slice accelerator includes a lookup table (LUT)-based multiplier, and the LUT-based multiplier includes a bit operation analysis unit (BMAU) and a bit-level multiplier unit. Among them, the BMAU uses the LUT mechanism to manage bit-level multiplication operations and divides the input data into bit segments, and the bit-level multiplier unit processes the grouping of consecutive multiplier bits after the bit segment division to achieve bit-level data processing.
[0006] Optionally, the LUT-based multiplier configures the operation data as follows: counts the non-zero bits in the most significant bit (MSB) and the least significant bit (LSB) of the input data and stores the counting result in a flip-flop; interfaces with an LUT address generator to generate partial products, and calculates through the LUT and stores the calculation result in the flip-flop; determines the bit path according to the non-zero bits in the MSB and the LSB.
[0007] Optionally, determining the bit path according to the non-zero bits in the most significant bit (MSB) and the least significant bit (LSB) includes: when both the MSB and the LSB are zero, skipping the calculation process and setting the output to zero; when the LSB is non-zero, performing a specified shift operation to ensure alignment of the input data; when the MSB is non-zero, processing the input data to ensure effective participation in the multiplication operation; when both the MSB and the LSB are non-zero, calculating and summing multiple partial products to generate the final product.
[0008] Optionally, the multiplier bit operation analysis unit (BMAU) based on the look-up table (LUT) is further configured to optimize the calculation overhead process when processing sparse and irregular matrices. Through parallel computing and pipelining operations, the PE bit slice calculation unit processes non-zero elements and merges partial products through a reduction operation. The addition operation is initiated in each clock cycle and assigned to the digital signal processing unit (DSP) to improve the calculation speed and resource utilization.
[0009] Optionally, through parallel computing and pipelining operations, the PE bit slice calculation unit processes non-zero elements and merges partial products through a reduction operation, including: determining the convolution kernel size, identifying non-zero elements, and grouping them into multiple data blocks according to the positions of the non-zero elements; marking the multiple data blocks and loading each data block to its respective specified position; after integrating the multiple data blocks into a complete convolution kernel, preparing for calculation processing. Among them, each PE bit slice calculation unit performs multiplication calculations on the loaded data according to a predetermined operation, and during this process, the generated partial products are subsequently merged through a reduction operation.
[0010] Optionally, the addition operation is initiated in each clock cycle and assigned to the digital signal processing unit (DSP), including: the core addition unit performs an addition operation on two operands and sets a completion flag; in a parallel manner, enabling the addition operation to be initiated in each clock cycle, and the addition operations in the parallel stage are assigned to the DSP.
[0011] Optionally, the PE bit slice calculation unit is further configured to: utilize the PE bit slice calculation unit column template function to process a column of addition operations by initializing the input, managing intermediate results, and ensuring full pipelining; and utilize the top-level function to coordinate the addition process, including managing the input data stream, calling the column addition pipeline, and storing the final result.
[0012] Optionally, the PE bit slice calculation unit is further configured to: adopt data reuse and streamlined configuration to ensure that matrix data is adapted to the PE bit slice calculation unit array, and realize data flow and parallel processing through two calculation stages.
[0013] Optionally, the data multiplexing and streamlining configuration is adopted to ensure that the matrix data is adapted to the PE bit-slice computing unit array, and data flow and parallel processing are achieved through two computing stages, including: dividing the matrix data into preset slice units to utilize data locality and optimize memory bandwidth; adapting the matrix data to the PE bit-slice processing unit array so that the data is transmitted through the PE bit-slice computing unit in two computing stages. Among them, each PE bit-slice computing unit calculates partial sums and passes the data to the next PE bit-slice computing unit.
[0014] Optionally, the PE bit-slice computing unit is further configured to: use loop operations to process the iterative process of traversing the slice units within multiple cycles, excluding the stride factor, where a new operation is started in each cycle; control logic to ensure that the calculation results are stored in the output matrix; utilize the analysis of the dimensions of the matrix data to perform adaptive partitioning for irregular matrices; the PE bit-slice computing unit array selects an appropriate slice size to support parallel computing.
[0015] Through the above technical solutions, the embodiments of the present invention provide a bit-slice accelerator for efficient convolutional neural network inference, and the bit-slice accelerator includes a multiplier based on a lookup table LUT. The multiplier of the lookup table LUT includes an integrated bit-level bit operation analysis unit BMAU and a bit-level multiplier unit to support operations with different bit widths. The embodiments of the present invention can effectively utilize data sparsity in deep neural network inference, significantly improving the inference efficiency and energy efficiency in resource-constrained and energy-constrained environments. At the same time, the adaptive sparse operation and flexible bit-width support enable the embodiments of the present invention to flexibly adapt to the computing requirements of different layers, greatly enhancing the adaptability and overall performance of the hardware accelerator.
[0016] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. They are used together with the following specific implementation to explain the embodiments of the present invention, but do not constitute a limitation to the embodiments of the present invention. In the drawings:
[0018] Figure 1 is a schematic structural diagram of a bit-slice accelerator for efficient convolutional neural network inference provided by the embodiments of the present invention;
[0019] Figure 2 is a schematic structural diagram of an example bit-slice accelerator for efficient convolutional neural network inference;
[0020] Figure 3It is a schematic diagram of an example bit-slice calculation unit;
[0021] Figure 4 It is a schematic diagram of an example dynamic calculation array. Detailed implementation manners
[0022] The following will describe in detail the specific implementation manners of the embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific implementation manners described herein are only used to illustrate and explain the embodiments of the present invention, and are not used to limit the embodiments of the present invention.
[0023] As described above, data sparsity is introduced in Deep Neural Networks (DNN), thereby significantly improving the efficiency of resource-constrained and energy-constrained hardware platforms. However, in order to fully utilize data sparsity and further improve energy efficiency, these DNN models require the support of dedicated hardware accelerators. The embodiments of the invention provide a bit-slice accelerator for efficient convolutional neural network inference, which can support adaptive sparse operations between various layers of DNN.
[0024] Figure 1 It is a schematic structural diagram of a bit-slice accelerator for efficient convolutional neural network inference provided by the embodiments of the present invention. Please refer to Figure 1 , the bit-slice accelerator may include a multiplier based on a Look-Up Table (LUT). The multiplier based on the Look-Up Table (LUT) may include a Bit Manipulation Analysis Unit (BMAU) and a bit-level multiplier unit. Among them, the Bit Manipulation Analysis Unit (BMAU) uses the LUT mechanism to manage bit-level multiplication operations, divides the input data into bit segments, and the bit-level multiplier unit processes the grouping of consecutive multiplier bits after the bit segment division to achieve bit-level data processing.
[0025] The embodiments of the present invention provide a bit-slice accelerator for efficient convolutional neural network inference. The bit-slice accelerator includes a multiplier based on a Look-Up Table (LUT). The multiplier of the Look-Up Table (LUT) includes an integrated bit-level Bit Manipulation Analysis Unit (BMAU) and a bit-level multiplier unit to support operations with different bit widths, and provides technical support for improving the energy efficiency and adaptability of DNN inference in a resource-constrained environment.
[0026] The preferred multiplier based on the Look-Up Table (LUT) in the embodiments of the present invention can configure the operation data as follows: count the non-zero bits in the Most Significant Bit (MSB) and the Least Significant Bit (LSB) of the input data, and store the counting result in a flip-flop; interface with the LUT address generator to generate partial products, and calculate through the LUT, and store the calculation result in a flip-flop; determine the bit path according to the non-zero bits in the MSB and the LSB.
[0027] Please refer to Figure 2 For example, the multiplier based on the look-up table (LUT) provided by the embodiments of the present invention can support logical operations such as 4-bit, 8-bit, 16-bit, and 32-bit. By combining the BMAU and the bit-level multiplier unit, the computing efficiency of the multiplier can be improved. Taking an example to illustrate, the BMAU utilizes the LUT mechanism to manage the bit-level multiplication operation, divides the data into 2N-bit segments, and adopts a simplified addressing scheme with the address range from 0 to 2N - 1. The bit-level multiplier unit can process the grouping of 2N consecutive multiplier bits (b[2N - 1:0]), thereby achieving efficient data processing. Taking an 8-bit multiplier (b[7:0]) as an example, the BMAU counts the non-zero bits in the most significant bit (MSB) and the least significant bit (LSB). These count values are then stored in the flip-flop (FF) with the index [2N / 2 - 1:0]. Next, by interfacing the LUT address generator with the BMAU, partial products are generated. The shared bit-slice calculation unit can be integrated into the 4-bit, 8-bit, 16-bit, and 32-bit multipliers based on the look-up table (LUT).
[0028] During the implementation process, the BMAU can count the non-zero bits in the most significant bit (MSB) and the least significant bit (LSB) of the input data and store the counting results in the flip-flop. The BMAU interfaces with the LUT address generator to generate partial products and calculate through the LUT. According to the non-zero bits in the MSB and the LSB, the BMAU optimizes the calculation path to improve the efficiency. For example, if both the MSB and the LSB are zero, the calculation will be skipped and the output will be set to zero; if the LSB is non-zero, a shift operation will be performed to ensure data alignment; if the MSB is non-zero, it will ensure that the data participates in the calculation effectively.
[0029] In the preferred embodiment of the present invention, determining the bit path according to the non-zero bits in the most significant bit (MSB) and the least significant bit (LSB) (i.e., determining the operation path of the non-zero bits) may include: when both the MSB and the LSB are zero, skipping the calculation process and setting the output to zero; when the LSB is non-zero, performing a specified shift operation to ensure the alignment of the input data; when the MSB is non-zero, processing the input data to ensure effective participation in the multiplication operation; when both the MSB and the LSB are non-zero, calculating and summing multiple partial products to generate the final product.
[0030] Continuing with the above example, the BMAU analyzes the MSB and LSB of the input data to perform the appropriate multiplication function: when both the MSB and LSB are zero, the calculation process can be skipped and the output is directly set to zero, thereby improving the calculation efficiency; when the LSB is non-zero, a specified shift operation is performed to ensure the accurate alignment of the input data, which guarantees the accuracy and consistency during the data conversion process; when the MSB is non-zero, the input is processed to ensure its effective participation in the multiplication operation, and at the same time, unnecessary calculation overhead is avoided; when both the MSB and LSB are non-zero, a multiplier based on the look-up table LUT calculates and sums multiple partial products to generate the final product, ensuring the accuracy of the multiplication.
[0031] In a preferred embodiment of the present invention, the bit operation analysis unit BMAU can also be used to optimize the calculation overhead process when dealing with sparse and irregular matrices. Among them, through parallel computing and pipelining operations, the bit slice calculation unit processes non-zero elements and merges partial products through a reduction operation. The addition operation is started in each clock cycle and assigned to the digital signal processing unit DSP to improve the calculation speed and resource utilization rate.
[0032] The multiplier based on the look-up table LUT provided by the embodiments of the present invention includes an integrated bit-level BMAU and a bit-level bitwise multiplier unit to support operations with different bit widths. At the same time, the shared bit slice calculation unit (PE) optimizes the critical path of the calculation, ensuring consistency under different bit width configurations.
[0033] Please refer to Figure 3 For example, when dealing with sparse matrices, the embodiments of the present invention adopt a bit slice calculation unit to optimize the calculation overhead. Through parallel computing and pipelining operations, the bit slice calculation unit can efficiently process non-zero elements and merge partial products through a reduction operation. Among them, the addition operation is started in each clock cycle and assigned to the DSP to improve the calculation speed and resource utilization rate. In addition, stream optimization and inlining techniques help reduce latency and improve performance, especially suitable for real-time digital signal processing applications.
[0034] In a preferred embodiment of the present invention, the process of the bit slice calculation unit processing non-zero elements and merging partial products through parallel computing and pipelining operations may include: determining the size of the convolution kernel, identifying non-zero elements, and grouping them into multiple data blocks according to the positions of the non-zero elements; marking multiple data blocks and loading each data block to its respective designated position; after integrating multiple data blocks into a complete convolution kernel, preparing for calculation processing. Among them, each bit slice calculation unit performs multiplication calculations on the loaded data according to a predetermined operation, and during this process, the generated partial products are subsequently merged through a reduction operation.
[0035] Please refer to Figure 3, Continuing with the above example, determine the convolution kernel size and identify the non - zero (NNZ) elements, which are grouped into multiple data blocks according to their positions. Convolution kernel data blocks, such as labeled K1, K2, and K3, are loaded into their respective designated positions to ensure that each bit - sliced computing unit receives the data required for subsequent operations. After integrating the data blocks into a complete convolution kernel, the bit - sliced computing unit is ready for efficient computational processing. Each bit - sliced computing unit performs a multiplication calculation on the loaded data according to a predetermined operation. During this process, the generated partial products are then merged through a reduction operation. The bit - sliced computing unit sums the intermediate results and finally generates an output result in the bit - sliced computing stage.
[0036] In a preferred embodiment of the present invention, the addition operation is initiated in each clock cycle and assigned to a digital signal processing unit (DSP). It may include: the core addition unit adds two operands and sets a completion flag; in a parallel manner, the addition operation can be initiated in each clock cycle, and the addition operations in the parallel stage are assigned to the DSP.
[0037] Please refer to Figure 3 , Continuing with the above example, the core addition unit adds two operands and sets a completion flag to ensure pipeline parallelism and efficient utilization of hardware resources. In a parallel manner, the addition operation can be initiated in each clock cycle, thus achieving high throughput. The addition operations in the parallel stage are assigned to the DSP, improving the calculation speed and resource utilization rate.
[0038] In a preferred embodiment of the present invention, the bit - sliced computing unit can also be configured to: utilize a column template function to process a column of addition operations by initializing the input, managing intermediate results, and ensuring full pipelining; and utilize a top - level function to coordinate the addition process, which may include managing the input data stream, calling the column addition pipeline, and storing the final result.
[0039] Please refer to Figure 3 , Continuing with the above example, the column template function of the bit - sliced computing unit processes a column of addition operations by initializing the input, managing intermediate results, and ensuring full pipelining, thereby maximizing the throughput. On this basis, the top - level function coordinates the entire addition process, which may include managing the input data stream, calling the column addition pipeline, and storing the final result.
[0040] Further preferably, the bit - sliced computing unit can also be configured to: adopt data reuse and streamlined configuration to ensure that matrix data fits into the bit - sliced computing unit array and achieve data flow and parallel processing through two computing stages.
[0041] Please refer to Figure 4, by way of example, to further improve the computational efficiency, embodiments of the present invention adopt data reuse and streamlined configuration to ensure that matrix data can be efficiently adapted to the bit-slice computing unit array (i.e., the PE array), and realize data flow and parallel processing through two computing phases (for example, cycle 1 and cycle 2). This can ensure the optimal reuse of data, while improving memory bandwidth utilization and computational efficiency, meeting the requirements of high throughput and large-scale data processing. To further enhance the performance, embodiments of the present invention adopt flow optimization and inlining methods to achieve data reuse, so as to reduce latency. For example, flow optimization instructions can direct the first-in-first-out queue (FIFO) to facilitate efficient data flow. These optimizations significantly improve the performance, making this implementation particularly suitable for high-throughput, real-time digital signal processing applications.
[0042] In a preferred embodiment of the present invention, the adoption of data reuse and streamlined configuration to ensure that matrix data is adapted to the bit-slice computing unit array and realize data flow and parallel processing through two computing phases may include: dividing the matrix data into preset slice units to utilize data locality and optimize memory bandwidth; adapting the matrix data to the bit-slice computing unit array so that the data is transmitted through the bit-slice computing unit in two computing phases. Wherein, each bit-slice computing unit calculates a partial sum and passes the data to the next bit-slice computing unit.
[0043] Please refer to Figure 4 , continuing with the above example, embodiments of the present invention utilize data reuse to implement a sparse matrix design with pipeline parallelism. The data matrix can be divided into preset slice (T) units to utilize data locality and optimize memory bandwidth. This method can efficiently adapt matrix blocks to the bit-slice computing unit array. In the streamlined array configuration, the data is transmitted through the bit-slice computing unit in two phases (for example, cycle 1 and cycle 2), representing different stages of computing; each bit-slice computing unit calculates a partial sum and passes the data to the next unit, thus realizing efficient data flow and parallelism.
[0044] In a preferred embodiment of the present invention, the bit-slice computing unit can also be configured to: use loop operations to handle the iterative process of traversing slice units in multiple cycles, excluding the stride factor, wherein, in each cycle, a new operation is started; control logic to ensure that the calculation result is stored in the output matrix; utilize the analysis of the dimensions of the matrix data to perform adaptive division on irregular matrices; the bit-slice computing unit array selects an appropriate slice size to support parallel computing.
[0045] Please refer to Figure 4, Continuing with the above example, embodiments of the present invention manage the iterative process of traversing slice units over multiple cycles using loop operations, excluding the stride factor. For example, the first loop retrieves elements A - O, while subsequent loops access subsets such as D, J, E, G, H, O, and F. This strategy ensures optimal data reuse, significantly improving computational efficiency and performance when processing large-scale datasets. A new operation can be initiated in each clock cycle, thus increasing throughput. In addition, the control logic of the embodiments of the present invention ensures that the calculation results are stored in the output matrix at the appropriate time. The adaptive partitioning of irregular matrices in the embodiments of the present invention involves analyzing the matrix dimensions to determine the optimal structure and processing requirements. Based on these dimensions, the PE array (e.g., PE0,0, PE1,0, …, PEN,0) receives inputs from a dedicated memory buffer (memory buffer, MK) and interacts through an interconnected data path. Data flows vertically into the bit-slice calculation units, allowing for efficient data distribution, while communication between the bit-slice calculation units facilitates accumulation operations. The array of bit-slice calculation units selects an appropriate slice size to support parallel computing, thereby optimizing the entire matrix processing process.
[0046] In the embodiments of the present invention, the multiplication (MUL) operation is performed by feeding input variables (I0, I1, IG, IN) into the corresponding bit-slice calculation units and sharing intermediate results through communication between the bit-slice calculation units. The clock signal regulates different stages of the operation, ensuring synchronous data processing across all bit-slice calculation units. The addition (ADD) stage is represented by NNZ[Num], across all bit-slice calculation units, indicating the accumulation of non-zero data elements. The data flow between the bit-slice calculation units forms a chained sequence, where each bit-slice calculation unit contributes its partial result, ultimately aggregating the intermediate results. Once the calculation is complete, the output result is written back to memory to ensure the correct storage of the processed data.
[0047] Accordingly, embodiments of the present invention provide a bit-slice accelerator for efficient convolutional neural network inference. The bit-slice accelerator includes a multiplier based on a look-up table LUT. The multiplier of the look-up table LUT includes an integrated bit-level bit operation analysis unit BMAU and a bit-level multiplier unit to support operations with different bit widths. Embodiments of the present invention can effectively utilize data sparsity in deep neural network inference, significantly improving inference efficiency and energy efficiency in resource-constrained and energy-constrained environments. At the same time, adaptive sparse operations and flexible bit-width support enable embodiments of the present invention to flexibly adapt to the computational requirements of different layers, greatly enhancing the adaptability and overall performance of the hardware accelerator.
[0048] Furthermore, for the bit-slice computing unit, the embodiments of the present invention utilize core partitioning, efficient bit-slice computing, and a fully pipelined addition mechanism to optimize the computing efficiency. By combining with a digital signal processing unit (DSP), significant performance improvement has been achieved when processing sparse and irregular matrices. For the dynamic computing array, the embodiments of the present invention utilize the preset bit-slice computing unit and the two-stage systolic array data stream to maximize data reuse. In addition, the communication and adaptive partitioning between bit-slice computing units further improve the efficiency of DNN inference computing.
[0049] Accordingly, the embodiments of the present invention can offer the following significant advantages:
[0050] 1) Significantly improved utilization efficiency of sparsity. Existing technologies often fail to fully utilize the data sparsity in DNNs, resulting in waste of computing and storage resources. The embodiments of the present invention perform refined processing of data sparsity at the bit level through an adaptive sparse lookup table and a bit-slice computing unit, greatly improving the utilization efficiency of sparse data, thereby significantly reducing the computing energy consumption.
[0051] 2) Flexible bit-width support. Most traditional hardware accelerators adopt bit-slice computing units with fixed precision and lack flexible support for different bit widths. The embodiments of the present invention integrate a multiplier based on a lookup table (LUT) and a shared bit-slice computing unit, which can support operations with different bit widths and flexibly adapt to the requirements of different precisions in each layer of the neural network, thereby achieving the best balance between precision and energy efficiency at different levels.
[0052] 3) Comprehensively improved computing performance. Through the combination of a bit-slice computing unit and a fully pipelined addition mechanism, as well as the cooperation of a digital signal processing unit, the embodiments of the present invention have achieved significant performance improvement when processing sparse and irregular matrices. It not only improves the data processing speed but also reduces the latency and increases the throughput of the overall system.
[0053] 4) Dynamically configurable computing array. The dynamic computing array design of the embodiments of the present invention can adaptively partition the bit-slice computing unit according to the actual computing requirements and adopt the data stream of a two-stage systolic array to maximize data reuse. Different from the existing fixed-structure computing arrays, the dynamic configuration of the embodiments of the present invention can better adapt to the computing density requirements of different network layers, thereby improving the inference efficiency, especially having significant advantages in scenarios with limited resources and energy consumption.
[0054] 5) Significantly reduced static power consumption. Most of the bit-slice computing units in existing technologies always remain active, resulting in high static power consumption. The embodiments of the present invention perform clock gating control based on the bit-slice computing unit, reducing the power consumption of unnecessary bit-slice computing units and significantly reducing the static power consumption of the overall system.
[0055] In summary, the present invention has significant advantages in terms of sparsity utilization, bit-width flexibility, computing performance, dynamic adaptability, and energy efficiency, and is particularly suitable for DNN inference tasks with limited resources and energy consumption.
[0056] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0057] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0058] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0059] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0060] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0061] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0062] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0063] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0064] The above are only embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A bit-slicing accelerator for efficient convolutional neural network inference, characterized in that: The bit slice accelerator includes a multiplier based on a lookup table LUT, and the multiplier based on the lookup table LUT includes a bit operation analysis unit BMAU and a bit-level multiplier unit, wherein the bit operation analysis unit BMAU manages the bit-level multiplication operation by using the LUT mechanism, and performs bit segment division on the input data, and the bit-level multiplier unit processes the grouping of continuous multiplier bits after the bit segment division to realize bit-level data processing, and the multiplier based on the lookup table LUT configures the operation data as follows: Counting non-zero bits in the most significant bit MSB and the least significant bit LSB of the input data, and storing the counting result in a flip-flop; Interface with the LUT address generator, generate partial products, and calculate through the LUT, and store the calculation results in the flip-flop; Determining a bit path according to the non-zero bits in the MSB and the LSB includes: When both the MSB and the LSB are zero, the calculation process is skipped and the output is set to zero; When the LSB is non-zero and the MSB is zero, performing a specified shift operation to ensure that the input data is aligned; When the MSB is non-zero and the LSB is zero, processing the input data to ensure valid participation in a multiplication operation; When both the MSB and the LSB are non-zero, multiple partial products are calculated and summed to generate a final product.
2. The bit slice accelerator according to claim 1, characterized in that: The bit operation analysis unit BMAU is also used to optimize the process of computing overhead when processing sparse and irregular matrices. Through parallel computing and pipeline operations, the bit slice calculation unit processes non-zero elements and merges partial products through reduction operations. The addition operation is started in each clock cycle and allocated to the digital signal processing unit DSP to improve the computing speed and resource utilization.
3. The bit slice accelerator according to claim 2, characterized in that: The bit slice calculation unit processes non-zero elements through parallel calculation and pipeline operation, and merges partial products through reduction operation, including: Determine the convolution kernel size, identify the non-zero elements, and group them into multiple data blocks according to their positions; Marking the plurality of data blocks and loading each data block into a respective designated location; After the multiple data blocks are integrated into a complete convolution kernel, calculation processing is prepared, wherein each bit slice calculation unit performs multiplication calculation on the loaded data according to a predetermined operation, during which the partial products generated are subsequently merged through a reduction operation.
4. The bit slice accelerator according to claim 2, characterized in that: The addition operation is started in each clock cycle and is assigned to the digital signal processing unit DSP, including: The core addition unit adds the two operands and sets the completion flag; In a parallel manner, the addition operation can be started in each clock cycle, and the addition operation in the parallel stage is distributed to the DSP.
5. The bit slice accelerator according to claim 2, characterized in that: The bit slice calculation unit is further configured to: Use column template functions to process a column addition operation by initializing inputs, managing intermediate results, and ensuring full pipelining; And using the top-level function, coordinate the addition process, including managing the input data flow, calling the column addition pipeline, and storing the final result.
6. The bit slice accelerator according to claim 2, characterized in that: The bit slice calculation unit is further configured to: Data multiplexing and streamlined configuration are adopted to ensure that matrix data fits into the bit slice computing unit array, and data flow and parallel processing are achieved through two computing stages.
7. The bit slice accelerator according to claim 6, characterized in that: The data reuse and streamlined configuration are adopted to ensure that the matrix data fits into the bit slice computing unit array, and data flow and parallel processing are realized through two computing stages, including: Dividing the matrix data into preset slice units to utilize data locality and optimize memory bandwidth; The matrix data is adapted into the array of bit slice computation units so that the data is transmitted through the bit slice computation units in two computation phases, wherein each bit slice computation unit computes a partial sum and passes the data to the next bit slice computation unit.
8. The bit slice accelerator according to claim 7, characterized in that: The bit slice calculation unit is further configured to: Using a loop operation to handle the iterative process of traversing the slice unit in multiple cycles, excluding the stride factor, wherein in each cycle, a new operation is started; Control logic to ensure that the calculation results are stored in the output matrix; By analyzing the dimension of the matrix data, the irregular matrix is adaptively divided; The bit slice calculation unit array selects an appropriate slice size to support parallel calculation.
Citation Information
Patent Citations
Multiplier array for matrix operation and multiplier array for convolution operation
CN111652359A
System and method for accelerating training of deep learning networks
CN115885249A