Depthwise volume accumulation internal calculation device and method based on dynamic table look-up accumulation
By using dynamic table lookup and accumulation technology in Depthwise convolution calculation, the decoupling multiplication and addition operation is used to store the search table value and dynamic table lookup and accumulation, the problem of limited efficiency and speed of Depthwise convolution calculation in the existing technology is solved, and the optimization of hardware resources and energy efficiency is achieved.
Patent Information
- Application Number
- CN202510165635.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-24
AI Technical Summary
In the prior art, the efficiency and speed of Depthwise convolutional calculations are limited by a large number of matrix multiplication operations and the lack of targeted algorithm optimization and computing unit design, resulting in higher chip area and power consumption.
The Depthwise convolution in-memory computing device based on dynamic table lookup accumulation is adopted. Through the lookup table memory and the multi-access concurrent lookup table cache module, the dynamic table lookup accumulation module and the data write back module, the in-memory computing of Depthwise convolution is realized. This device decouples the traditional multiplication and addition operation into pre-computed lookup table storage and dynamic lookup table accumulation, avoiding the use of high-complexity multipliers.
Through dynamic table lookup and accumulation technology, hardware resource consumption and power consumption are significantly reduced, computing efficiency is improved, and parallel access to multiple lookup table values in a single cycle is supported, realizing in-depth optimization of hardware resources and energy efficiency.
Smart Images

Figure CN120196586A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of neural network acceleration, and relates to a Depthwise convolution in-memory computing device and method based on dynamic look-up table accumulation. Background Art
[0002] There are multiple processing chips for processing data through neural networks in a hardware chip. During the inference process of the neural network, a large number of matrix operations will be generated, consuming a large amount of computing resources and generating a considerable computing delay. To implement the deployment and application of neural networks on Internet of Things and wearable devices, researchers have not only proposed model compression technologies such as quantization, pruning, and knowledge distillation for existing network models, but also explored lightweight neural network structures, aiming to further reduce the number of model parameters while maintaining model accuracy. As a lightweight neural network operator, Depthwise convolution separates the convolution process of each channel, avoiding information interaction between channels during convolution to reduce the amount of computation and the number of parameters, thereby achieving network lightweighting.
[0003] To implement the model deployment of neural networks, dedicated neural network computing devices greatly improve the computing speed and energy consumption efficiency of neural networks through customized hardware design and algorithm optimization. Therefore, the design of neural network accelerators using high-parallel heterogeneous computing platforms such as FPGAs, GPUs, and ASICs has become a new research hotspot. Among them, FPGAs have good market prospects due to their high customization, high energy efficiency ratio, and low latency. However, a large number of matrix multiplication operations are required during Depthwise convolution calculation, and the chip area and power consumption consumed by large-scale multiplier and adder arrays are also very considerable. In addition, Depthwise convolution is usually regarded as a normal convolution with a channel number of 1 during calculation, lacking targeted algorithm optimization and computing unit design, which limits the efficiency and speed of Depthwise convolution calculation. Summary of the Invention
[0004] To solve the above problems existing in the prior art, the present invention provides a Depthwise convolution in-memory computing device and method based on dynamic look-up table accumulation. The technical problems to be solved by the present invention are realized through the following technical solutions:
[0005] A Depthwise convolution in-memory computing device based on dynamic look-up table accumulation, comprising: a look-up table memory, a multi-address concurrent look-up table cache module, a dynamic look-up table accumulation module, a data write-back module, a feature map memory, a data acquisition module, and a cache switching control module;
[0006] The lookup table memory is connected to the multi-address concurrent lookup table cache module, the multi-address concurrent lookup table cache module is connected to the dynamic lookup table accumulation module, the dynamic lookup table accumulation module is connected to the data write-back module, the data write-back module is connected to the feature map memory, the feature map memory is connected to the data acquisition module, the data acquisition module is connected to the dynamic lookup table accumulation module, and the cache switching control module is connected to the dynamic lookup table accumulation module and the multi-address concurrent lookup table cache module;
[0007] The feature map memory is used to store the input feature map data of the Depthwise convolution layer and the output feature map data after the Depthwise convolution;
[0008] The lookup table memory is used to store the table values generated for the weight data of the Depthwise convolution layer in a receptive field decoupled manner;
[0009] The multi-address concurrent lookup table cache module is used to write the lookup table data in the lookup table memory into the corresponding lookup table cache according to the flag signal of the cache switching control module;
[0010] The data acquisition module is used to acquire feature map data and generate a corresponding edge padding mode signal according to the height and width coordinates where the feature map data is located;
[0011] The dynamic lookup table accumulation module is used to obtain the calculation result of the Depthwise convolution calculation;
[0012] The data write-back module is used to generate the address and enable signal for writing into the feature map memory, and write the Depthwise convolution calculation result back to the corresponding address in the feature map memory;
[0013] The cache switching control module is used to generate corresponding cache switching flag signals in different processing stages.
[0014] In an embodiment of the present invention, the feature map memory is implemented using BlockRAM in the FPGA, with 4-bit quantization. The shapes of the input feature map data matrix and the output feature map data matrix are both [C, H, W], where C is the number of channels of the feature map, H is the height of the feature map, and W is the width of the feature map; in the channel dimension, every Tc channels are divided into a group.
[0015] In an embodiment of the present invention, the multi-address concurrent lookup table cache module is composed of two identical multi-address concurrent lookup table cache units;
[0016] Each of the multi-address concurrent lookup table cache units includes Tc multi-address concurrent lookup tables, and each of the multi-address concurrent lookup tables is divided into n according to the n×n receptive field2 Each row stores 16 4-bit values which are the indexes for the receptive field positions.
[0017] In one embodiment of the present invention, the dynamic look-up table accumulation module converts the feature map data into indexes for the multi-address concurrent look-up table cache module, performs row shift caching on the indexes, and dynamically aligns the spatial data layout required for Depthwise convolution; dynamically determines whether to perform normal look-up through the indexes or directly assign zeros according to the current edge padding mode; if performing normal look-up, dynamically selects the corresponding look-up table cache for look-up according to the cache switching flag; accumulates the data obtained from different receptive field positions, and re-constrains it to 4 bits, thus obtaining the final result of the Depthwise convolution calculation.
[0018] A Depthwise convolution in-memory computing method based on dynamic look-up table accumulation, comprising:
[0019] S1. Obtain the weight data of the Depthwise convolution layer, generate a look-up table according to the weight data in a receptive field decoupled manner, and write the generated table values into the look-up table memory;
[0020] S2. According to the cache switching flag signal of the cache switching control module, write the look-up table data in the look-up table memory into the multi-address concurrent look-up table cache module;
[0021] S3. The data acquisition module fetches Tc 4-bit feature map data from the feature map memory and generates a corresponding edge padding mode signal;
[0022] S4. The dynamic look-up table accumulation module converts the Tc 4-bit feature map data into indexes for the multi-address concurrent look-up table cache module. After row shift caching, it dynamically obtains Tc×9 4-bit indexes that align with the spatial data layout required for Depthwise convolution calculation; according to the edge padding mode signal, dynamically determines whether to perform normal look-up through the indexes or directly assign zeros; if performing normal look-up, dynamically selects the corresponding look-up table cache for look-up according to the cache switching flag, and obtains Tc×9 4-bit data;
[0023] S5. The dynamic look-up table accumulation module accumulates the Tc×9 4-bit data obtained in step S4 in the receptive field dimension to obtain Tc n-bit data, and parallelly re-constrains each n-bit data into 4 bits, which is the calculation result of the Depthwise convolution calculation;
[0024] S6. The data write-back module writes Tc 4-bit data into the corresponding addresses of the feature map memory in each round, returns to step S3, and repeats H×W rounds, writing a total of Tc×H×W 4-bit data into the feature map memory;
[0025] S7. In each round, the data write-back module writes Tc×H×W 4-bit data into the feature map memory, returns to step S2, and repeats C / Tc rounds, and a total of C×H×W 4-bit data are written into the feature map memory.
[0026] In one embodiment of the present invention, the step S1 includes:
[0027] S11. Obtain the int4-quantized weights of the Depthwise convolution layer in the convolutional neural network. The shape of the weight data matrix Wei is [C, 3, 3], where C is the number of channels for Depthwise convolution calculation, and [3, 3] represents that the size of the convolution kernel is 3×3;
[0028] S12. Generate an int4-quantized feature map index matrix Idx. Traverse from -8 to 7 to generate 16 different index values, and the index values form a feature map index matrix Idx with the shape of [16, C, 3, 3];
[0029] S13. Multiply the weight data matrix Wei and the feature map index matrix Idx to obtain a table value matrix Lut with the shape of [C, 3, 3, 16];
[0030] S14. Constrain each value in the table value matrix Lut to the range of [-8, 7], then splice them pairwise and write them into the lookup table memory in sequence.
[0031] In one embodiment of the present invention, the formula for multiplying the weight data matrix Wei and the feature map index matrix Idx is:
[0032]
[0034] Lut[c, kh, kw, i] = Wei[c, kh, kw] × Idx[i, c, kh, kw]
[0035] In one embodiment of the present invention, the step S2 includes that the multi-address concurrent lookup table cache module is composed of two identical multi-address concurrent lookup table cache units. Each multi-address concurrent lookup table cache unit includes Tc multi-address concurrent lookup tables. Each multi-address concurrent lookup table is divided into n 2 rows according to the n×n receptive field. 16 4-bit numerical values of the index for the receptive field position are stored in each row, which exactly corresponds to the shape of the table value matrix Lut. To complete a full Depthwise convolution calculation with the number of channels C, it is necessary to flush the lookup table cache in the lookup table memory C / Tc times.
[0036] In one embodiment of the present invention, the step S3 includes that the storage mode of the feature map data in the feature map memory is an interleaved layout. The data bit width of the feature map memory is Tc×4 bits. At a certain address in the BlockRAM, Tc 4-bit data of adjacent channels at the same spatial position in the feature map are stored. After obtaining the feature map data for H×W clock cycles in the spatial dimension direction of the feature map, then obtaining the feature map data for the next H×W clock cycles in the channel dimension direction. The cache switching control module switches the read / write cache at the first clock cycle of each H×W clock cycle segment to generate a corresponding edge padding mode signal.
[0037] In one embodiment of the present invention, the step S4 includes that the dynamic look-up table accumulation module has Tc parallel running sub-modules. The sub-modules parallelly convert Tc 4-bit feature map data into indexes of the multi-address concurrent look-up table cache module and then send the indexes to the row shift cache for shifting to dynamically obtain Tc×9 4-bit indexes that align with the spatial data arrangement required for the Depthwise convolution in a pipelined manner. The sub-modules parallelly take out the table values or directly assign zeros from the corresponding look-up table caches for the indexes at different receptive field spatial positions according to the edge padding mode signal and the cache switching flag signal to obtain Tc×9 4-bit data.
[0038] Advantages of the present invention:
[0039] 1. The in-memory computing device for Depthwise convolution based on dynamic look-up table accumulation of the present invention realizes the storage reconstruction of the computing logic, decouples the multiplication and addition operations in the traditional Depthwise convolution into the storage of pre-computed look-up table values and dynamic look-up table accumulation. Through the "storage is computing" method, the complex multiplication operation in the Depthwise convolution is transformed into the dynamic parallel address mapping and data reading of the storage module, completely avoiding high-complexity multipliers, thereby saving hardware resource consumption and significantly reducing power consumption. And the use of the multi-address concurrent look-up table cache can reduce the frequent migration of data between the external storage and the computing unit, realizing the deep optimization of hardware resources and energy efficiency;
[0040] 2. The Depthwise convolution in-memory computing method based on dynamic look-up table accumulation of the present invention aims at the compression and optimization of look-up tables for in-memory computing. It uses a look-up table pre-generation method with decoupled receptive fields. By separating the spatial dimension mapping relationship of the convolution kernel, the storage overhead of the look-up table is significantly reduced, making it possible to deploy the Depthwise convolution in-memory computing method on-chip. A novel multi-address concurrent look-up table cache structure is adopted to store the pre-generated look-up table values, which supports parallel access to multiple look-up table values in a single cycle. Based on this cache structure, a dynamic look-up table accumulation module is designed to equivalently implement the multiply-accumulate operation in Depthwise convolution, enabling dynamic spatial alignment of the input feature map and look-up table values, parallel look-up of multiple channels, and parallel accumulation of multiple channels, thus significantly improving the computing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 FIG. is a schematic structural diagram of a Depthwise convolution in-memory computing device based on dynamic look-up table accumulation provided by an embodiment of the present invention
[0042] Figure 2 FIG. is a schematic flow diagram of a Depthwise convolution in-memory computing method based on dynamic look-up table accumulation provided by an embodiment of the present invention;
[0043] Figure 3 FIG. is a schematic structural diagram of a multi-address concurrent look-up table cache provided by an embodiment of the present invention;
[0044] Figure 4 FIG. is a pipeline timing diagram of a Depthwise convolution in-memory computing device under the control of a cache switching control module provided by an embodiment of the present invention;
[0045] Figure 5 FIG. is a schematic order diagram of reading input feature map data from a feature map memory or writing output feature map data to a feature map memory provided by an embodiment of the present invention;
[0046] Figure 6 FIG. is a schematic diagram of edge padding modes at different feature map coordinate positions provided by an embodiment of the present invention;
[0047] Figure 7 FIG. is a schematic structural diagram of a dynamic look-up table accumulation module provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0049] Embodiment 1:
[0050] Embodiment 1 of the present invention provides a Depthwise convolution in-memory computing device based on dynamic look-up table accumulation. Referring to the attached Figure 1, the in-memory computing device for Depthwise convolution based on dynamic look-up table accumulation includes: a look-up table memory, a multi-address concurrent look-up table cache module, a dynamic look-up table accumulation module, a data write-back module, a feature map memory, a data acquisition module, and a cache switching control module.
[0051] Among them, the look-up table memory is connected to the multi-address concurrent look-up table cache module, the multi-address concurrent look-up table cache module is connected to the dynamic look-up table accumulation module, the dynamic look-up table accumulation module is connected to the data write-back module, the data write-back module is connected to the feature map memory, the feature map memory is connected to the data acquisition module, the data acquisition module is connected to the dynamic look-up table accumulation module, and the cache switching control module is connected to the dynamic look-up table accumulation module and the multi-address concurrent look-up table cache module.
[0052] In an embodiment of the present invention, the feature map memory is used to store the input feature map data of the Depthwise convolution layer and the output feature map data after the Depthwise convolution is completed, and is implemented using BlockRAM in an FPGA (Field Programmable Gate Array). In this embodiment, 4-bit quantization is adopted, and the shapes of the input feature map data matrix and the output feature map data matrix are both [C, H, W], where C is the number of channels of the feature map, H is the height of the feature map, and W is the width of the feature map; in the channel dimension, every Tc channels are divided into a group.
[0053] The look-up table memory is used to store the table values generated for the weight data of the Depthwise convolution layer in a receptive field decoupled manner, and is implemented using BlockRAM in an FPGA.
[0054] The multi-address concurrent look-up table cache module is used to write the look-up table data in the look-up table memory into the corresponding look-up table cache according to the flag signal of the cache switching control module. The multi-address concurrent look-up table cache module consists of two identical multi-address concurrent look-up table cache units; each multi-address concurrent look-up table cache unit includes Tc multi-address concurrent look-up tables, and each multi-address concurrent look-up table is divided into n 2 rows according to the n×n receptive field, and 16 4-bit values of the index numbers for the receptive field positions are stored in each row. In the embodiment provided by the present invention, Tc = 32, n = 3, and the multi-address concurrent look-up table cache is implemented using FFs (flip-flops) in an FPGA. Due to the physical characteristics of the flip-flops, they support parallel access to multiple look-up table values within a single cycle.
[0055] The data acquisition module is used to generate the address and enable signal for reading the feature map memory, and retrieve the feature map data from the corresponding position of the feature map memory; in addition, it is used to generate the corresponding edge padding mode signal according to the height and width coordinates of the feature map data retrieved at the current moment.
[0056] The dynamic look-up table accumulation module is used to obtain the calculation result of the Depthwise convolution calculation. The dynamic look-up table accumulation module converts the feature map data into the index of the multi-address concurrent look-up table cache module, performs row shift caching on the index, and dynamically aligns the spatial data layout required for the Depthwise convolution; dynamically judges whether to perform normal look-up table access or directly assign zero according to the current edge padding mode; if performing normal look-up table access, dynamically selects the corresponding look-up table cache for look-up table access according to the cache switching flag; accumulates the data obtained from different receptive field positions and re-constrains it to 4 bits, that is, obtains the final result of the Depthwise convolution calculation.
[0057] The data write-back module is used to generate the address and enable signal for writing to the feature map memory, and write back the Depthwise convolution calculation result to the corresponding address in the feature map memory.
[0058] The cache switching control module is used to generate the corresponding cache switching flag signal in different processing stages to implement the ping-pong control of two multi-address concurrent look-up table caches and improve the processing efficiency.
[0059] The Depthwise convolution in-memory computing device of this embodiment is designed based on the multi-address concurrent look-up table cache module, supports parallel access to multiple pre-stored look-up table values within a single cycle, and cooperates with the dynamic look-up table accumulation module to achieve the dynamic spatial alignment of the input feature map and the look-up table values, multi-channel parallel look-up table and accumulation operations, realizes the storage reconstruction of the computing logic, decouples the multiplication and addition operations in the traditional Depthwise convolution into the storage of pre-computed look-up table values and dynamic look-up table accumulation. Through the "storage is computing" method, the complex multiplication operation in the Depthwise convolution is transformed into the dynamic parallel address mapping and data reading of the storage module, completely avoiding high-complexity multipliers, thereby saving hardware resource consumption and significantly reducing power consumption. Moreover, the use of multi-address concurrent look-up tables can reduce the frequent migration of data between external storage and computing units, achieving in-depth optimization of hardware resources and energy efficiency.
[0060] Embodiment 2:
[0061] Embodiment 2 of the present invention provides a Depthwise convolution in-memory computing method based on dynamic look-up table accumulation. Referring to the appendix Figure 2 , the Depthwise convolution in-memory computing method based on dynamic look-up table accumulation includes:
[0062] S1. Obtain the weight data of the Depthwise convolution layer, use the Matlab tool to generate a look-up table according to the weight data in the way of decoupling the receptive field, and write the generated table values into the look-up table memory;
[0063] S2. Write the lookup table data in the lookup table memory into the multi-address concurrent lookup table cache module according to the cache switching flag signal of the cache switching control module;
[0064] S3. The data acquisition module fetches Tc 4-bit feature map data from the feature map memory and generates a corresponding edge padding mode signal;
[0065] S4. The dynamic lookup table accumulation module converts the Tc 4-bit feature map data into indexes of the multi-address concurrent lookup table cache module. After row shift caching, Tc×9 4-bit indexes that are aligned with the spatial data layout required for Depthwise convolution calculation are dynamically obtained; according to the edge padding mode signal, dynamically determine whether to perform normal lookup through the index or directly assign zero; if performing normal lookup, dynamically select the corresponding lookup table cache for lookup according to the cache switching flag, and obtain Tc×9 4-bit data;
[0066] S5. The dynamic lookup table accumulation module accumulates the Tc×9 4-bit data obtained in step S4 in the receptive field (spatial) dimension to obtain Tc n-bit data, and parallelly re-constrains each n-bit data into 4 bits, which is the calculation result of the Depthwise convolution calculation;
[0067] S6. The data write-back module writes Tc 4-bit data into the corresponding address of the feature map memory in each round, returns to step S3, and repeats H×W rounds, writing a total of Tc×H×W 4-bit data into the feature map memory;
[0068] S7. The data write-back module writes Tc×H×W 4-bit data into the feature map memory in each round, returns to step S2, and repeats C / Tc rounds, writing a total of C×H×W 4-bit data into the feature map memory.
[0069] To reduce the storage overhead of the lookup table, in this embodiment, lookup table compression and optimization for in-memory computing are performed, and a lookup table pre-generation method with decoupled receptive fields is used, that is, instead of directly generating the final result of the Depthwise convolution operation as the table value, the multiplication result in the Depthwise convolution calculation process is generated as the table value. In this way, when performing the Depthwise convolution calculation, after looking up the table value from the lookup table, the table value is accumulated in the receptive field (spatial) dimension to obtain the final result of the Depthwise convolution calculation. Since the power consumption of the multiplier is much higher than that of the adder, using the lookup table to replace a large number of multiplication operations in the Depthwise convolution calculation can significantly reduce the power consumption generated during the calculation.
[0070] Refer to Table 1 for the comparison of whether to use the lookup table with decoupled receptive fields in terms of storage capacity.
[0071] Table 1
[0072]
[0073] Table 1 shows the comparison of the storage capacity (using 4-bit quantization) of the lookup tables generated by the method without decoupling the receptive field and with decoupling the receptive field, where n is the height or width of the receptive field. In the embodiments of the present invention, the size of the receptive field is n 2 = 3×3 = 9.
[0074] In one embodiment of the present invention, step S1 specifically includes:
[0075] S11. Obtain the int4-quantized weights of the Depthwise convolutional layer in the convolutional neural network. The shape of the weight data matrix Wei is [C, 3, 3], where C is the number of channels for Depthwise convolution calculation, and [3, 3] represents that the size of the convolutional kernel is 3×3;
[0076] S12. Generate the int4-quantized feature map index matrix Idx. Traverse from -8 to 7 to generate 16 different index values, and the index values form the feature map index matrix Idx with the shape of [16, C, 3, 3];
[0077] S13. Multiply the weight data matrix Wei and the feature map index matrix Idx to obtain the table value matrix Lut with the shape of [C, 3, 3, 16]. The multiplication formula is:
[0078]
[0079] Lut[c, kh, kw, i] = Wei[c, kh, kw] × Idx[i, c, kh, kw]
[0080] S14. Constrain each value in the table value matrix Lut to the range of [-8, 7], then splice them pairwise and write them into the lookup table memory in sequence.
[0081] According to the cache switching flag signal of the cache switching control module, write the lookup table data from the lookup table memory into the corresponding lookup table cache. The ping-pong read-write control method is adopted to schedule the flow of the lookup table data among the lookup table memory, two lookup table caches, and the dynamic lookup table accumulation module in a pipelined manner, effectively improving the processing efficiency of the system.
[0082] Appendix Figure 3 is a schematic diagram of a multi-address concurrent lookup table cache structure provided by the embodiments of the present invention. In order to support various channel transformations and feature map sizes in Depthwise convolution calculation, a minimum multi-address concurrent lookup table cache structure that can be repeatedly called is designed.
[0083] In this embodiment, the multi - access concurrent lookup table cache module consists of two identical multi - access concurrent lookup table cache units. Each multi - access concurrent lookup table cache unit includes Tc multi - access concurrent lookup tables. Each multi - access concurrent lookup table is divided into n rows according to the n×n receptive field. 2 Each row stores 16 4 - bit values of the index numbers for the receptive field positions, which exactly corresponds to the shape of the table - value matrix Lut. To complete a full - fledged Depthwise convolution calculation with the number of channels C, the lookup table cache needs to be flushed C / Tc times from the lookup table memory.
[0084] Appendix Figure 4 FIG. is the pipeline timing diagram of the Depthwise convolution in - memory computing device controlled by the cache switching control module provided in the embodiment of the present invention. The cache switching control module outputs corresponding read - write cache switching flag signals in each state, and writes the lookup table data in the lookup table memory into the corresponding lookup table cache according to the cache switching flag signals. The number of input and output feature map channels of the Depthwise convolution calculation is C. In order to enable the Depthwise convolution computing device to support various input - output channel transformations, in the C dimension, every Tc channels are divided into a group in the present invention. Since the data bit - width of the lookup table memory is 1024 bit, in the embodiment provided by the present invention, Tc = 32, n = 3, and n is the height or width of the receptive field. 1024 bit data can be written into the multi - access concurrent lookup table cache per clock cycle, and the storage of the multi - access concurrent lookup table cache is Tc×n2×16×4 bit. Therefore, it can be calculated that it takes 18 clock cycles to fill the multi - access concurrent lookup table cache according to the following formula.
[0085]
[0086] Since the shape of the generated lookup table matrix Lut exactly corresponds to the structure of the designed multi - access concurrent lookup table cache, 1024 bit data can be written into the lookup table cache in sequence per clock cycle.
[0087] Appendix Figure 5 FIG. is the sequence diagram of reading input feature map data from the feature map memory or writing output feature map data to the feature map memory provided in the embodiment of the present invention. The storage mode of the feature map data in the feature map memory is interleaved layout. The data bit - width of the feature map memory is Tc×4 bit. At a certain address in the BlockRAM, the Tc 4 - bit data of adjacent channels at the same spatial position in the feature map are stored.
[0088] Since a lookup table cache can only store the lookup tables generated for Tc channels, in order to avoid frequent overwriting of the lookup table cache, after obtaining the feature map data for H×W clock cycles in the spatial dimension direction of the feature map, and then obtaining the feature map data for the next H×W clock cycles in the channel dimension direction, the cache switching control module switches the read / write cache at the first clock cycle of each segment of H×W clock cycles to generate the corresponding edge padding mode signal.
[0089] Appendix Figure 6 is a schematic diagram of the edge padding modes at different feature map coordinate positions provided by an embodiment of the present invention. As can be seen from the appendix, there are a total of 9 different edge padding mode signals.
[0090] The dynamic lookup table accumulation module converts Tc 4-bit feature map data into indexes for reading the lookup table cache. After row shift caching, it dynamically obtains Tc×9 4-bit indexes that are aligned with the spatial data layout required for Depthwise convolution calculation; according to the edge padding mode signal, it dynamically determines whether to perform normal table lookup through the index or directly assign zero; if performing normal table lookup, it dynamically selects the corresponding lookup table cache for table lookup according to the cache switching flag, and finally obtains Tc×9 4-bit data.
[0091] Appendix Figure 7 is a schematic diagram of the structure of the dynamic lookup table accumulation module provided by an embodiment of the present invention. In this embodiment, the dynamic lookup table accumulation module has Tc parallel running sub-modules. The sub-modules parallelly convert Tc 4-bit feature map data into indexes for multi-address concurrent lookup table caches, and then send the indexes to the row shift cache for shifting, and dynamically obtain Tc×9 4-bit indexes that are aligned with the spatial data layout required for Depthwise convolution in a pipelined manner. The sub-modules parallelly, according to the edge padding mode signal and the cache switching flag signal, dynamically fetch the table values from the corresponding lookup table cache or directly assign zero for the indexes at different receptive field spatial positions, and obtain Tc×9 4-bit data.
[0092] When performing Depthwise convolution calculation with the input feature map data matrix in the shape of [C, H, W] and the weight data matrix in the shape of [C, 3, 3] provided by the embodiments of the present invention, it takes H×W×C / Tc clock cycles to complete the calculation.
[0093] The effects of the present invention can be further illustrated by the following experimental data.
[0094] The present invention is compared with TinyNPU-F in reference to Table 2.
[0095] Table 2
[0096]
[0097] The operating frequency of TinyNPU-F is 10 - 100 MHz. When performing DW instruction operations at the maximum frequency, NOU-dw is in the working state, completing 16×9 + 16×8 = 272 operations per cycle. That is, the peak computing performance is: 272×0.1 = 27.2 GOP / s. The embodiment of the present invention can operate at a maximum working frequency of 200 MHz, which is equivalent to completing 32×9 + 32×8 = 544 operations per cycle. That is, the peak computing power is 544×0.2 = 108.8 GOP / s.
[0098] As can be seen from Table 2, the energy efficiency ratio of the present invention is 4.6 times that of the NOU-dw operator in the high-energy efficiency neural network accelerator TinyNPU-F, showing excellent energy efficiency ratio performance.
[0099] Traditional Depthwise convolution relies on high-complexity multiply-add operations. However, the present invention performs a storage reconstruction of the calculation logic. Through a pre-computed lookup table and a dynamic lookup table accumulation mechanism, the multiply-add operations in Depthwise convolution are decoupled into the storage of lookup table values and dynamic lookup table accumulation operations. This design completely avoids the use of multipliers, saves hardware resource consumption, and significantly reduces the power consumption generated during the calculation. Additionally, the present invention performs an efficient adaptation of the in-memory computing architecture, designs a multi-address concurrent lookup table cache structure, and maps it to the flip-flop (FF) resources of the FPGA, making full use of the high-parallel access characteristics of the flip-flops, so that the dynamic lookup table accumulation module supports concurrent reading and accumulation of multiple lookup table values within a single cycle. This design significantly improves the calculation efficiency. Moreover, the multi-address concurrent lookup table cache can reduce the frequent migration of data between external storage and the computing unit, further reducing the power consumption.
[0100] The in-memory computing method for Depthwise convolution of the present invention is oriented to the compression and optimization of the lookup table for in-memory computing. It uses a lookup table pre-generation method with decoupled receptive fields. By separating the spatial dimension mapping relationship of the convolution kernel, the storage overhead of the lookup table is significantly reduced, achieving an exponential compression of the lookup table storage overhead, making it possible to deploy the in-memory computing method for Depthwise convolution on-chip. A novel multi-address concurrent lookup table cache structure is adopted to store the pre-generated lookup table values, which supports parallel access to multiple lookup table values within a single cycle. Based on this cache structure, a dynamic lookup table accumulation module is designed to equivalently implement the multiply-add operations in Depthwise convolution, enabling dynamic spatial alignment of the input feature map and the lookup table values, multi-channel parallel lookup, and multi-channel parallel accumulation, significantly improving the calculation efficiency.
[0101] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be covered by the protection scope of the present invention.
Claims
1. A depthwise convolution in-memory computing device based on dynamic table lookup accumulation, characterized in that: include: Lookup table memory, multi-address concurrent lookup table cache module, dynamic table lookup accumulation module, data write-back module, feature map memory, data acquisition module and cache switching control module; The lookup table memory is connected to the multi-address concurrent lookup table cache module, the multi-address concurrent lookup table cache module is connected to the dynamic table lookup accumulation module, the dynamic table lookup accumulation module is connected to the data write back module, the data write back module is connected to the feature map memory, the feature map memory is connected to the data acquisition module, the data acquisition module is connected to the dynamic table lookup accumulation module, the cache switching control module is connected to the dynamic table lookup accumulation module and the multi-address concurrent lookup table cache module; The feature map memory is used to store the input feature map data of the Depthwise convolution layer and the output feature map data after the Depthwise convolution is performed; The lookup table memory is used to store table values generated for the weight data of the Depthwise convolution layer in a receptive field decoupling manner; The multi-address concurrent lookup table cache module is used to write the lookup table data in the lookup table memory into the corresponding lookup table cache according to the flag signal of the cache switching control module; The data acquisition module is used to acquire feature map data and generate a corresponding edge filling mode signal according to the height and width coordinates of the feature map data; The dynamic table lookup accumulation module is used to obtain the calculation result of the Depthwise convolution calculation; The data write-back module is used to generate an address and an enable signal for writing into the feature map memory, and write the depthwise convolution calculation result back to the corresponding address of the feature map memory; The cache switching control module is used to generate corresponding cache switching flag signals in different processing stages.
2. The depthwise convolution in-memory computing device based on dynamic table lookup accumulation according to claim 1, characterized in that: The feature map memory is implemented using BlockRAM in FPGA, using 4-bit quantization, and the shapes of the input feature map data matrix and the output feature map data matrix are both [C, H, W], where C is the number of channels of the feature map, H is the height of the feature map, and W is the width of the feature map; in the channel dimension, every Tc channels are divided into a group.
3. The depthwise convolution in-memory computing device based on dynamic table lookup accumulation according to claim 2, characterized in that: The multi-address concurrent lookup table cache module is composed of two identical multi-address concurrent lookup table cache units; The multi-address concurrent lookup table cache unit includes Tc multi-address concurrent lookup tables, each of which is divided into n according to the n×n receptive field. 2 Each row stores 16 4-bit values of the index number of the receptive field position.
4. The depthwise convolution in-memory computing device based on dynamic table lookup accumulation according to claim 3, characterized in that: The dynamic table lookup accumulation module converts the feature map data into the index of the multi-address concurrent lookup table cache module, performs row shift cache on the index, and dynamically aligns the spatial data arrangement required by the Depthwise convolution; dynamically determines whether to look up the table normally through the index or directly assign zeros according to the current edge filling mode; If the table lookup is normal, the corresponding lookup table cache is dynamically selected for table lookup according to the cache switching flag; The data obtained from different receptive field positions are accumulated and reconstrained to 4 bits to obtain the final result of the Depthwise convolution calculation.
5. A depthwise convolution in-memory calculation method based on dynamic table lookup accumulation, characterized in that: include: S1. Obtain the weight data of the Depthwise convolutional layer, generate a lookup table according to the weight data in a receptive field decoupling manner, and write the generated table value into the lookup table memory; S2, writing the lookup table data in the lookup table memory into a multi-address concurrent lookup table cache module according to a cache switching flag signal of a cache switching control module; S3, the data acquisition module takes out Tc 4-bit feature map data from the feature map memory and generates a corresponding edge filling mode signal; S4, the dynamic table lookup accumulation module converts Tc 4-bit feature map data into the index of the multi-address concurrent lookup table cache module, and after row shift cache, dynamically obtains Tc×9 4-bit indexes of the spatial data arrangement required for the alignment Depthwise convolution calculation; according to the edge filling mode signal, dynamically determines whether to look up the table normally through the index or directly assign zeros; if the table is looked up normally, dynamically select the corresponding lookup table cache for lookup according to the cache switching flag to obtain Tc×9 4-bit data; S5, the dynamic table lookup accumulation module accumulates the Tc×9 4-bit data obtained in step S4 in the receptive field dimension to obtain Tc n-bit data, and in parallel reconstrains each n-bit data to 4 bits, which is the calculation result of the Depthwise convolution calculation; S6, the data write-back module writes Tc 4-bit data into the corresponding address of the feature map memory in each round, returns to step S3, repeats H×W rounds, and writes a total of Tc×H×W 4-bit data into the feature map memory; S7. The data write-back module writes Tc×H×W 4-bit data into the feature map memory in each round, returns to step S2, repeats C / Tc rounds, and writes a total of C×H×W 4-bit data into the feature map memory.
6. The depthwise convolution in-memory calculation method based on dynamic table lookup accumulation according to claim 5, characterized in that: The step S1 comprises: S11. Get the int4 quantized weights of the Depthwise convolution layer in the convolutional neural network. The shape of the weight data matrix Wei is [C, 3, 3], where C is the number of channels calculated by the Depthwise convolution, and [3, 3] represents that the size of the convolution kernel is 3×3. S12, generate a feature map index matrix Idx after int4 quantization, traverse -8 to 7, generate 16 different index values, and the index values constitute a feature map index matrix Idx with a shape of [16, C, 3, 3]; S13, multiplying the weight data matrix Wei and the feature map index matrix Idx to obtain a table value matrix Lut with a shape of [C, 3, 3, 16]; S14, constraining each value in the table value matrix Lut to be within the range of [-8, 7], concatenating them in pairs, and writing them into the lookup table memory in sequence.
7. The depthwise convolution in-memory calculation method based on dynamic table lookup accumulation according to claim 6, characterized in that: The formula for multiplying the weight data matrix Wei and the feature map index matrix Idx is:
8. The depthwise convolution in-memory calculation method based on dynamic table lookup accumulation according to claim 7, characterized in that: The step S2 includes: the multi-address concurrent lookup table cache module is composed of two identical multi-address concurrent lookup table cache units, each of which includes Tc multi-address concurrent lookup tables, each of which is divided into n according to the n×n receptive field. 2 Each row stores 16 4-bit values of the index number for the receptive field position, which completely corresponds to the shape of the table value matrix Lut. To complete a complete Depthwise convolution calculation with C channels, a total of C / Tc lookup table caches need to be flushed from the lookup table memory.
9. The depthwise convolution in-memory calculation method based on dynamic table lookup accumulation according to claim 8, characterized in that: The step S3 includes: the feature map data is stored in the feature map memory in an interleaved layout, the data bit width of the feature map memory is Tc×4bit, and at a certain address in the BlockRAM, Tc 4-bit data of adjacent channels at the same spatial position in the feature map are stored; after acquiring the feature map data of H×W clock cycles in the direction of the feature map space dimension, the feature map data of the next H×W clock cycles is acquired in the direction of the channel dimension, and the cache switching control module switches the read-write cache in the first clock cycle of each H×W clock cycle to generate a corresponding edge fill mode signal.
10. The depthwise convolution in-memory calculation method based on dynamic table lookup accumulation according to claim 9, characterized in that: The step S4 includes: the dynamic table lookup accumulation module has Tc sub-modules running in parallel, the sub-modules convert Tc 4-bit feature map data into indexes of the multi-address concurrent lookup table cache module in parallel and then send the indexes to the row shift cache for shifting, and dynamically obtain Tc×9 4-bit indexes of the spatial data arrangement required for alignment Depthwise convolution in a pipeline manner, and the sub-modules dynamically take out table values from the corresponding lookup table cache or directly assign zeros to indexes of different receptive field spatial positions in parallel according to the edge fill mode signal and the cache switching flag signal, and obtain Tc×9 4-bit data.
Citation Information
Cited By
Multi-layer perceptron high-energy-efficiency in-memory calculation method and device based on in-memory index
CN122154795A