Pointwise convolution calculation device and method based on in-memory lookup table
By using in-memory lookup tables and table lookup accumulation pipeline units in Pointwise convolutional calculation, the problem of high computing resources and energy consumption is solved, and efficient computing and resource conservation is achieved.
Patent Information
- Application Number
- CN202510150630.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-06
AI Technical Summary
In the prior art, Pointwise convolutional calculation consumes a large amount of computing resources and energy in matrix multiplication operations, and lacks computing unit optimization and data control strategies for different channel transformation forms, which limits the computing efficiency and speed.
A Pointwise convolutional computing device based on in-memory lookup table is designed, and the search table value is generated using input channel decoupling and 4-bit quantization. The search table data is replaced by the search table accumulation pipeline unit, and the search table data is scheduled using ping pong read and write control.
It significantly saves hardware resources and reduces computing power consumption, supports multiple channel transformations and feature map sizes, improves the processing efficiency of the system, and realizes efficient Pointwise convolutional calculations.
Smart Images

Figure CN120106156A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of neural network acceleration, and in particular relates to a Pointwise convolution calculation device and method based on an in-memory lookup table. Background Art
[0002] At present, deep learning technologies represented by convolutional neural networks have been widely used in image classification, image segmentation, target detection and other aspects, and have achieved remarkable results. However, convolutional neural networks will undergo a large number of matrix operations during the inference process, consuming a lot of computing resources and energy. In order to realize the deployment and application of neural networks in the Internet of Things and wearable devices, researchers have not only proposed model compression technologies such as quantization, pruning, and knowledge distillation for existing network models, but also explored lightweight neural network structures, aiming to further reduce the number of model parameters while maintaining model accuracy. Pointwise convolution is a lightweight convolution method that uses a 1×1 convolution kernel to realize information interaction between feature map channels. Compared with large-size convolution kernels, it reduces the amount of calculation and parameters in the convolution process, thereby achieving network lightweighting.
[0003] In order to realize the model deployment of neural network, the dedicated neural network computing device greatly improves the computing speed and energy efficiency of neural network through customized hardware design and algorithm optimization. However, a large number of matrix multiplication operations occur during Pointwise convolution calculation, so large-scale multipliers and adders are required, which consumes a lot of computing resources and energy. In addition, due to the presence of multiple channel transformation forms in Pointwise convolution calculation, multiplication operations of different scales are generated during the calculation process. The current computing acceleration platform lacks computing unit optimization and data control strategies for Pointwise convolution, which limits the efficiency and speed of Pointwise convolution calculation. Summary of the invention
[0004] In order to solve the above problems existing in the prior art, the present invention provides a Pointwise convolution calculation device and method based on an in-memory lookup table. The technical problem to be solved by the present invention is achieved through the following technical solutions:
[0005] In a first aspect, the present invention provides a Pointwise convolution computing device based on an in-memory lookup table, comprising:
[0006] Feature map memory, used to store the input feature map data of the Pointwise convolution layer and the output feature map data after the Pointwise convolution;
[0007] A lookup table memory, used to store table values of a lookup table generated for weight values of a Pointwise convolution layer in a manner of input channel decoupling;
[0008] A lookup table array, for providing two lookup table arrays and a method for storing table values in each lookup array table;
[0009] The FSM control unit is used to perform ping-pong reading and writing control on the two lookup table arrays in a finite state machine manner, thereby generating a control signal;
[0010] A multiplexer MUX, configured to select, according to a control signal output by the FSM control unit, into which lookup table array the data in the lookup table memory is to be written, or to select from which lookup table array the lookup table data is to be read out to the data acquisition unit;
[0011] A data acquisition unit, used to generate an address and an enable signal for reading a feature map memory, and to retrieve feature map data from a corresponding address of the feature map memory;
[0012] The table lookup accumulation pipeline unit processes the extracted feature map data into an index for reading the lookup table array, takes out the corresponding table value from the designed lookup table array, and accumulates it in the input channel dimension in a pipeline manner to obtain an accumulation result;
[0013] A data post-processing unit is used to post-process the accumulated result to obtain a final Pointwise convolution calculation result, and send the convolution calculation result as output feature map data to the feature map memory.
[0014] In a second aspect, the present invention provides a Pointwise convolution calculation method based on an in-memory lookup table, comprising:
[0015] Step 1, obtain the weight data of the Pointwise convolution layer, and use Python tools to generate a lookup table for the weight data in a way of decoupling the input channels, and write the generated lookup table into the lookup table memory;
[0016] Step 2, writing the lookup table data in the lookup table memory into the corresponding lookup table in the corresponding lookup table array according to the control signal generated by the FSM control unit;
[0017] Step 3, using the data acquisition unit to fetch 64 clock cycles Tcin of 4-bit feature map data from the corresponding address of the feature memory each time;
[0018] Step 4, converting the feature map data into an index of a lookup table array, and taking out corresponding Tcin*Tcout 4-bit table values from the corresponding lookup table array through the index according to the control signal of the FSM control unit;
[0019] Step 5, using the table lookup accumulation pipeline unit to accumulate the table value taken out in the Tcin dimension, so as to obtain Tcout nbit results;
[0020] Step 6, using the data post-processing unit to store the calculation results of each round of Tcout nbits, and return to step 3;
[0021] Step 7, add the calculation result of Tcout nbits of the current round and the calculation result of Tcout nbits of the previous round accordingly and store them;
[0022] Step 8: Repeat steps 3 to 6 for a total of Cin / Tcin times, and reconstrain each n-bit data to 4 bits in the last step to obtain the final result of the Pointwise convolution calculation;
[0023] Step 9, write the final result into the corresponding address of the feature map memory, and repeat steps 2 to 8 a total of Cout / Tcout times to write Cout 4-bit calculation result data back to the feature map memory.
[0024] Beneficial effects:
[0025] 1. The present invention designs a novel in-memory lookup table array structure, which can efficiently store and read LUT values generated in an input channel decoupling manner. Then, based on this structure, a lookup table accumulation pipeline unit is designed, which replaces the multiplication-accumulation array responsible for matrix multiplication in Pointwise convolution calculation, thereby eliminating a large number of multipliers, significantly saving hardware resources and reducing computing power consumption. The present invention uses input channel decoupling and 4-bit quantization to generate the lookup table values of the Pointwise convolution layer, which greatly reduces the storage capacity of the lookup table, so that it can be deployed on a platform with limited hardware resources.
[0026] 2. The present invention repeatedly calls the table lookup accumulation pipeline unit multiple times to enable the Pointwise convolution calculation device to support multiple channel transformations and multiple feature map sizes; in addition, a ping-pong read and write control method is used to schedule the flow of lookup table data between the lookup table memory, the lookup table array and the table lookup accumulation pipeline unit, thereby effectively improving the processing efficiency of the system.
[0027] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 It is a schematic diagram of the structure of a Pointwise convolution computing device based on an in-memory lookup table provided by an embodiment of the present invention;
[0029] Figure 2It is a flowchart of a calculation method of a Pointwise convolution calculation device based on an in-memory lookup table provided by an embodiment of the present invention;
[0030] Figure 3 It is a schematic diagram of a specific calculation process of a calculation method of a Pointwise convolution calculation device based on an in-memory lookup table provided by an embodiment of the present invention;
[0031] Figure 4 is a state transition diagram of an FSM control unit provided by an embodiment of the present invention;
[0032] Figure 5 It is a schematic diagram of an in-memory lookup table array structure provided by an embodiment of the present invention;
[0033] Figure 6 is a schematic diagram of a sequence for reading input feature map data from a feature map memory provided by an embodiment of the present invention;
[0034] Figure 7 It is a structural schematic diagram of a table lookup accumulation pipeline unit provided by an embodiment of the present invention;
[0035] Figure 8 It is a schematic diagram of the sequence of writing output feature map data into a feature map memory provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0036] The present invention is further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.
[0037] like Figure 1 As shown, the present invention provides a Pointwise convolution computing device based on an in-memory lookup table, comprising:
[0038] Feature map memory, used to store the input feature map data of the Pointwise convolution layer and the output feature map data after the Pointwise convolution;
[0039] The feature map memory can be implemented using BlockRAM in FPGA. The present invention adopts 4-bit quantization, the shape of the input feature map data matrix is [Cin, H, W], and the shape of the output feature map data matrix is [Cout, H, W], where Cin and Cout are the number of input channels and output channels of the feature map, and H and W are the length and width of the feature map; in the present invention, in the Cin dimension, each Tcin channel is divided into a group, and in the Cout dimension, each Tcout channel is divided into a group;
[0040] A lookup table memory, used to store table values of a lookup table generated for weight values of a Pointwise convolution layer in a manner of input channel decoupling;
[0041] The lookup table memory can be implemented using BlockRAM in FPGA.
[0042] A lookup table array is used to provide two lookup table arrays and a method for storing table values in each lookup table array; wherein the lookup table array is used to provide two lookup true tables, each lookup table array includes Tcin lookup tables, each lookup table stores 16 indexes, and each index corresponds to Tcout 4-bit table values.
[0043] The lookup table array includes two identical lookup table arrays 1 and 2, each of which includes Tcin lookup tables, each of which has 16 indexes, and each index corresponds to Tcout 4-bit table values, so that the Pointwise convolution calculation device can achieve a processing speed of Tcin×Tcout parallelism. In the embodiment provided by the present invention, Tcin=Tcout=32, and the lookup table array is implemented using Distributed RAM in FPGA;
[0044] The FSM control unit is used to perform ping-pong reading and writing control on the two lookup table arrays in a finite state machine manner to effectively improve the processing speed and finally generate a control signal;
[0045] A multiplexer MUX, configured to select, according to a control signal output by the FSM control unit, into which lookup table array the data in the lookup table memory is to be written, or to select from which lookup table array the lookup table data is to be read out to the data acquisition unit;
[0046] A data acquisition unit, used to generate an address and an enable signal for reading a feature map memory, and to retrieve feature map data from a corresponding address of the feature map memory;
[0047] The table lookup accumulation pipeline unit processes the extracted feature map data into an index for reading the lookup table array, takes out the corresponding table value from the designed lookup table array, and accumulates it in the input channel dimension in a pipeline manner to obtain an accumulation result;
[0048] A data post-processing unit is used to post-process the accumulated result to obtain a final Pointwise convolution calculation result, and send the convolution calculation result as output feature map data to the feature map memory.
[0049] Among them, the data post-processing unit is used to post-process the accumulated result to obtain the final Pointwise convolution calculation result, and generate an address and an enable signal for writing into a feature map memory; and use the enable signal to enable the feature map memory so that it stores the convolution calculation result as output feature map data according to the address.
[0050] Combination Figure 2 and Figure 3 The present invention provides a Pointwise convolution calculation method based on an in-memory lookup table, including:
[0051] Step 1, obtain the weight data of the Pointwise convolution layer, and use Python tools to generate a lookup table for the weight data in a way of decoupling the input channels, and write the generated lookup table into the lookup table memory;
[0052] In order to reduce the storage capacity of the lookup table, the lookup table is generated in a way that decouples the input channels, that is, the final result of the Pointwise convolution operation is not directly generated as the table value, but the multiplication result in the Pointwise convolution calculation process is generated as the table value. In this way, when performing the Pointwise convolution calculation, after finding the table value from the lookup table, the table value is accumulated in the input channel dimension to obtain the final result of the Pointwise convolution calculation. Since the power consumption of the multiplier is much higher than that of the adder, using a lookup table to replace a large number of multiplication operations in the Pointwise convolution calculation can greatly reduce the power consumption generated in the calculation.
[0053] Table 1 Comparison of storage capacity of input channel decoupling lookup table and whether to use it
[0054] Input channel decoupling is not used Use input channel decoupling Lookup table storage <![CDATA[16 Cin ×Cout×4bit]]> 16×Cin×Cout×4bit
[0055] Table 1 shows the comparison of the storage capacity (using 4-bit quantization) of the lookup table generated by the method without using input channel decoupling and using input channel decoupling.
[0056] In a specific embodiment of the present invention, step 1 comprises:
[0057] Step 1.1, obtain the int4 quantized weight data of the Pointwise convolution layer in the convolutional neural network and form a weight data matrix; the shape of the weight data matrix is [Cout, Cin 1,1], Cout is the number of output channels of the Pointwise convolution, Cin is the number of input channels of the Pointwise convolution, and [1,1] represents the size of the convolution kernel is 1*1;
[0058] Step 1.2, according to the number of input channels of the Pointwise convolution layer, generate the int4 quantized feature map index matrix; this step traverses -8 to 7 and generates 16 different index values, which constitute the feature map index matrix Idx with a shape of [16, Cin, 1, 1];
[0059] Step 1.3, multiply the weight data matrix and the feature map index matrix to obtain a table value matrix; the shape of the table value matrix is [coutoff, cinoff, 16, TCin, Tcout]. The table value matrix is expressed as:
[0060] Lut[c ot ,c it ,i,tc i ,tc o ]=Idx[i,tc i +c it ×Tcin]×Wei[tc o +c ot ×Tcout,tc i +c it ×Tcin]
[0061] In the formula, c it ∈[0,cinoff-1],i∈[0,15],tc i ∈[0,Tcin-1],tc o ∈[0,Tcout-1],
[0062] Step 1.4, flatten the generated table value matrix and then write it into the lookup table memory in sequence.
[0063] Step 2, writing the lookup table data in the lookup table memory into the corresponding lookup table in the corresponding lookup table array according to the control signal generated by the FSM control unit;
[0064] In a specific embodiment of the present invention, step 2 comprises:
[0065] Step 2.1, designing a minimum lookup table array structure, wherein the minimum lookup table array structure includes two lookup table arrays, each lookup table array includes Tcin lookup tables, each lookup table stores 16 indexes, and each index corresponds to Tcout 4-bit table values;
[0066] Step 2.2, writing the lookup table data in the lookup table memory into the corresponding lookup table array according to the control signal output by the FSM control unit in each state.
[0067] This step uses a ping-pong read-write control method to schedule the flow of lookup table data between the lookup table memory, the lookup table array and the lookup table accumulation pipeline unit in a pipeline manner, effectively improving the processing efficiency of the system. Figure 4 A state transition diagram of an FSM control unit provided in an embodiment of the present invention.
[0068] Figure 5A schematic diagram of an in-memory lookup table array structure provided by an embodiment of the present invention. In order to support multiple channel transformations and feature map sizes in Pointwise convolution calculations, a minimum lookup table array structure that can be repeatedly called is designed: a lookup table array includes Tcin lookup tables, each lookup table has 16 indexes, and each index corresponds to Tcout 4-bit data. This structure corresponds to the shape of the table value matrix Lut generated by the input channel decoupling method in step 1. Therefore, to complete a complete Pointwise convolution calculation, a total of cinoff×coutoff lookup table arrays need to be written from the lookup table memory.
[0069] The FSM control unit will output a corresponding control signal in each state, and write the lookup table data in the lookup table memory into the lookup table array according to the control signal. Since the data bit width of the lookup table memory is 1024 bits, Tcout=Tcin=32 in the embodiment provided by the present invention, it can be seen from the following formula that it takes 64 clock cycles to fill a lookup table array.
[0070]
[0071] 1024 bits are written in each clock cycle. According to the shape of the generated lookup table matrix Lut, the first clock cycle writes Tcout*4 bits of data corresponding to the first index of the eight lookup tables from lookup table 0 to lookup table 7 in the lookup table array, and then each clock cycle writes the lookup table data into the lookup table array in this order.
[0072] Step 3, using the data acquisition unit to fetch 64 clock cycles Tcin of 4-bit feature map data from the corresponding address of the feature memory each time;
[0073] Since it takes 64 clock cycles to fully write a lookup table array, in order to perform ping-pong control, the data acquisition unit acquires 64 clock cycles of feature map data in each ping-pong cycle.
[0074] The feature map data is stored in the feature map memory in an interleaved layout. The data bit width of the feature map memory is Tcin*4 bits. At a certain address in the BlockRAM, Tcin 4-bit data of adjacent channels at the same spatial position in the feature map are stored.
[0075] Figure 6A sequential schematic diagram of reading input feature map data from a feature map memory is provided for an embodiment of the present invention. Since a lookup table array can only store lookup tables generated for Tcin input channels, in order to avoid frequently refreshing the lookup table, feature map data of 64 clock cycles are obtained in the direction of the feature map spatial position, and then the feature map data of the next 64 clock cycles are obtained in the direction of the channel position.
[0076] Step 4, converting the feature map data into an index of a lookup table array, and taking out corresponding Tcin*Tcout 4-bit table values from the corresponding lookup table array through the index according to the control signal of the FSM control unit;
[0077] In a specific embodiment of the present invention, step 4 comprises:
[0078] Step 4.1, using a table lookup accumulation pipeline unit to convert feature map data into an index of a lookup table array;
[0079] Step 4.2, according to the control signal generated by the FSM control unit, the corresponding Tcin*Tcout 4-bit table values are retrieved in parallel from the corresponding Tcin lookup tables in the corresponding lookup table array through Tcin 4-bit indexes.
[0080] This step uses the table lookup accumulation pipeline unit to convert the feature map data into the index of the lookup table array. According to the control signal of the FSM control unit, the corresponding Tcin*Tcout 4-bit table values are taken out in parallel from the corresponding Tcin lookup tables in the corresponding lookup table array through Tcin 4-bit indexes. Figure 7 It is a structural schematic diagram of a table lookup accumulation pipeline unit provided in an embodiment of the present invention.
[0081] Step 5, using the table lookup accumulation pipeline unit to accumulate the table value taken out in the Tcin dimension, so as to obtain Tcout nbit results;
[0082] Step 6, using the data post-processing unit to store the calculation results of each round of Tcout nbits, and return to step 3;
[0083] In this step, the data post-processing unit is used to store Tcout n-bit calculation results in each round; return to step 3, repeat 64 rounds, and store 64*Tcout n-bit data in total.
[0084] Step 7, add the calculation result of Tcout nbits of the current round and the calculation result of Tcout nbits of the previous round accordingly and store them;
[0085] Step 8: Repeat steps 3 to 6 for a total of Cin / Tcin times, and reconstrain each n-bit data to 4 bits in the last step to obtain the final result of the Pointwise convolution calculation;
[0086] In this step, the data post-processing unit is used to add and store the calculation results of 64*Tcout nbits in each round with the calculation results of 64*Tcout nbits in the previous round, and then returns to step 3 and repeats cinoff rounds. In the last round, each nbit data is reconstrained to 4 bits, which is the final result of the Pointwise convolution calculation.
[0087] Step 9, write the final result into the corresponding address of the feature map memory, and repeat steps 2 to 8 a total of Cout / Tcout times to write Cout 4-bit calculation result data back to the feature map memory.
[0088] If cout=Tcin=32, repeat H*W / 64 rounds and write a total of H*W*Tcout 4-bit data. Figure 8 This is a schematic diagram of the sequence of writing the output feature map data into the feature map memory provided by an embodiment of the present invention. The data post-processing unit writes H*W*Tcout 4-bit data to the feature map memory in each round; returns to step 2, repeats coutoff rounds, and finally writes H*W*Cout 4-bit data to the feature map memory.
[0089] When performing Pointwise convolution calculation with an input feature map data matrix having a shape of [Cin, H, W] and a weight data matrix having a shape of [Cout, Cin, 1, 1], the embodiment provided by the present invention requires cinoff*H*W*coutoff clock cycles to complete the calculation.
[0090] The effect of the present invention can be further illustrated by the following experimental data.
[0091] Table 2 Energy efficiency comparison between the present invention and TinyNPU-F
[0092]
[0093]
[0094] TinyNPU-F comes from the doctoral dissertation: Xu Kunran. Design and research of high-energy-efficiency deep neural network acceleration chip [D]. Xidian University, 2022. DOI: 10.27389 / d.cnki.gxadu.2022.000115. The operating frequency of TinyNPU-F is 10-100MHz. When TinyNPU-F performs Pointwise instruction calculations at the maximum frequency, NOU-conv is in working state, and each cycle completes (16×16)×2=512 operations, that is, the peak computing performance is: 512×0.1=51.2GOP / s. The method of the present invention can work at a maximum operating frequency of 200MHZ, which can be equivalent to completing (32×32)×2=2048 operations per cycle, that is, the peak computing power is 2048×0.2=409.6GOP / s.
[0095] It can be seen from Table 2 that the energy efficiency ratio of the present invention is 4.3 times that of the NOU-conv operator in the high-efficiency neural network accelerator TinyNPU-F, and has excellent energy efficiency performance.
[0096] Compared with the Pointwise convolution operator using a multiply-accumulate array, the present invention uses a novel table lookup accumulation pipeline unit to replace the multiply-accumulate unit responsible for matrix multiplication in the Pointwise convolution calculation, thereby eliminating a large number of multipliers, saving hardware resource consumption and significantly reducing the power consumption generated in the calculation. In addition, the lookup table array structure designed in the present invention facilitates parallel table lookup operations, and is implemented using the DistributedRAM on the FPGA, which has higher energy efficiency than FF and BlockRAM with extremely low storage utilization (each lookup table in the lookup table array requires a BlockRAM).
[0097] It is worth noting that the terms "first" and "second" in the present invention are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0098] Although the present application is described herein in conjunction with various embodiments, in the process of implementing the claimed application, those skilled in the art may understand and implement other variations of the disclosed embodiments by reviewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality of components or steps.
[0099] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.
Claims
1. A Pointwise convolution computing device based on an in-memory lookup table, characterized in that: include: Feature map memory, used to store the input feature map data of the Pointwise convolution layer and the output feature map data after the Pointwise convolution; A lookup table memory, used to store table values of a lookup table generated for weight values of a Pointwise convolution layer in a manner of input channel decoupling; A lookup table array, for providing two lookup table arrays and a method for storing table values in each lookup array table; The FSM control unit is used to perform ping-pong reading and writing control on the two lookup table arrays in a finite state machine manner, thereby generating a control signal; A multiplexer MUX, configured to select, according to a control signal output by the FSM control unit, into which lookup table array the data in the lookup table memory is to be written, or to select from which lookup table array the lookup table data is to be read out to the data acquisition unit; A data acquisition unit, used to generate an address and an enable signal for reading a feature map memory, and to retrieve feature map data from a corresponding address of the feature map memory; The table lookup accumulation pipeline unit processes the extracted feature map data into an index for reading the lookup table array, takes out the corresponding table value from the designed lookup table array, and accumulates it in the input channel dimension in a pipeline manner to obtain an accumulation result; A data post-processing unit is used to post-process the accumulated result to obtain a final Pointwise convolution calculation result, and send the convolution calculation result as output feature map data to the feature map memory.
2. The Pointwise convolution computing device based on an in-memory lookup table according to claim 1, characterized in that: The lookup table array is used to provide two lookup true lists. Each lookup table array includes Tcin lookup tables. Each lookup table stores 16 indexes, and each index corresponds to Tcout 4-bit table values.
3. The Pointwise convolution computing device based on an in-memory lookup table according to claim 1, characterized in that: The data post-processing unit is used to post-process the accumulated result to obtain the final Pointwise convolution calculation result, and generate an address and an enable signal written into the feature map memory; The enable signal is used to enable the feature map memory so that it stores the convolution calculation result as output feature map data according to the address.
4. A Pointwise convolution calculation method based on an in-memory lookup table, characterized in that: include: Step 1, obtain the weight data of the Pointwise convolution layer, and use Python tools to generate a lookup table for the weight data in a way of decoupling the input channels, and write the generated lookup table into the lookup table memory; Step 2, writing the lookup table data in the lookup table memory into the corresponding lookup table in the corresponding lookup table array according to the control signal generated by the FSM control unit; Step 3, using the data acquisition unit to fetch 64 clock cycles Tcin of 4-bit feature map data from the corresponding address of the feature memory each time; Step 4, converting the feature map data into an index of a lookup table array, and taking out corresponding Tcin*Tcout 4-bit table values from the corresponding lookup table array through the index according to the control signal of the FSM control unit; Step 5, using the table lookup accumulation pipeline unit to accumulate the table value taken out in the Tcin dimension, so as to obtain Tcout nbit results; Step 6, using the data post-processing unit to store the calculation results of each round of Tcout nbits, and return to step 3; Step 7, add the calculation result of Tcout nbits of the current round and the calculation result of Tcout nbits of the previous round accordingly and store them; Step 8: Repeat steps 3 to 6 for a total of Cin / Tcin times, and reconstrain each n-bit data to 4 bits in the last step to obtain the final result of the Pointwise convolution calculation; Step 9, write the final result into the corresponding address of the feature map memory, and repeat steps 2 to 8 a total of Cout / Tcout times to write Cout 4-bit calculation result data back to the feature map memory.
5. The Pointwise convolution calculation method based on an in-memory lookup table according to claim 4, characterized in that: Step 1 includes: Step 1.1, obtain the int4 quantized weight data of the Pointwise convolution layer in the convolutional neural network and form a weight data matrix; Step 1.2, generate the int4 quantized feature map index matrix according to the number of input channels of the Pointwise convolution layer; Step 1.3, multiplying the weight data matrix and the feature map index matrix to obtain a table value matrix; Step 1.4, flatten the generated table value matrix and then write it into the lookup table memory in sequence.
6. The Pointwise convolution calculation method based on an in-memory lookup table according to claim 5, characterized in that: The table value matrix in step 1.3 is expressed as: Lut[c ot ,c it ,i,tc i ,tc o ]=Idx[i,tc i +c it ×Tcin]×Wei[tc o +c ot ×Tcout,tc i +c it ×Tcin] wherein, c it ∈[0, cinoff - 1], i ∈ [0, 15], tc i ∈[0, Tcin - 1], tc o ∈[0, Tcout - 1], 7. The Pointwise convolution calculation method based on an in-memory lookup table according to claim 4, characterized in that: Step 2 includes: Step 2.1, designing a minimum lookup table array structure, wherein the minimum lookup table array structure includes two lookup table arrays, each lookup table array includes Tcin lookup tables, each lookup table stores 16 indexes, and each index corresponds to Tcout 4-bit table values; Step 2.2, writing the lookup table data in the lookup table memory into the corresponding lookup table array according to the control signal output by the FSM control unit in each state.
8. The Pointwise convolution calculation method based on an in-memory lookup table according to claim 4, characterized in that: Step 4 includes: Step 4.1, using a table lookup accumulation pipeline unit to convert feature map data into an index of a lookup table array; Step 4.2, according to the control signal generated by the FSM control unit, the corresponding Tcin*Tcout 4-bit table values are retrieved in parallel from the corresponding Tcin lookup tables in the corresponding lookup table array through Tcin 4-bit indexes.