Reconfigurable systolic array supporting variable granularity convolution and full connection operation

By designing a reconfigurable pulsating array that supports variable particle size convolution and fully connected computing, the problem of adaptation of multimodal neural network computing requirements caused by architecture solidification by existing accelerators is solved, and efficient multimodal neural network acceleration is achieved, reducing the deployment cost and energy consumption overhead of edge devices.

CN120106159APending Publication Date: 2025-06-06CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510268162.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Due to the solid architecture of existing neural network accelerators, it is difficult to compatible with the computing needs of multimodal neural networks, resulting in a significant increase in the deployment cost and energy consumption overhead of edge devices.

Method used

A reconfigurable pulsating array supporting variable particle size convolution and fully connected operations is designed. Through 11×11 reconfigurable computing units and multiple reconfigurable selection gate modules, dynamic configuration and data flow reconstruction are realized, and 3×3, 5×5, 11×11 convolution operations and fully connected layer operations are supported.

Benefits of technology

It realizes efficient support of multimodal neural networks on the same hardware platform, significantly reducing the deployment cost and energy consumption overhead of edge devices, and improving resource utilization and energy efficiency ratio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106159A_ABST
    Figure CN120106159A_ABST
Patent Text Reader

Abstract

The invention relates to a reconfigurable systolic array supporting variable granularity convolution and full connection operation, and belongs to the field of network accelerator hardware architecture. The array is composed of 11 * 11 reconfigurable computing units and a plurality of reconfigurable selection gate modules, and the reconfigurable computing units can achieve shift operation, multiplication operation, additive operation and multiply-accumulate operation in a programmable mode by dynamically configuring control signals; and the reconfigurable selection gate module is used for realizing dynamic routing selection of multiple data paths, and supporting 3 * 3, 5 * 5 and 11 * 11 convolution operations and full connection layer operation by reconstructing a connection relationship between input and output ports. The reconfigurable systolic array provided by the invention can support convolution operation and full connection operation of different granularities through dynamic configuration, and has universal adaptation capability on different types of neural network accelerators.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of network accelerator hardware architecture and relates to a reconfigurable systolic array supporting variable granularity convolution and full connection operations. Background Art

[0002] The current design of neural network accelerators faces the challenge of architectural adaptation brought by algorithm heterogeneity: the computational core of Convolutional Neural Network (CNN) lies in multi-scale convolution operations, which realize feature extraction through sliding windows of local receptive fields; while Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM) rely on fully connected operations to build temporal state transfer, and there are essential differences between the two in data flow paradigm and computational granularity. Existing hardware acceleration architectures mostly adopt fixed computing mode design. CNN accelerators optimize convolution operation efficiency through customized systolic arrays, but it is difficult to adapt to the global parameter interaction characteristics of fully connected layers; on the contrary, although RNN-oriented accelerators can efficiently process vector-matrix multiplication, they lack the ability to map convolution kernel space. This architectural specificity makes a single hardware platform incompatible with the computing requirements of multimodal neural networks, significantly increasing the deployment cost and energy consumption of edge smart devices.

[0003] The technical bottlenecks of traditional solutions are mainly reflected in two aspects: first, the granularity of computing units is fixed and cannot be dynamically reconfigured to match the core operators of different neural networks. For example, a 3×3 convolution kernel requires compact interconnection of neighborhood computing units, while a fully connected layer requires global broadcast data interaction. The existing architecture is difficult to switch between the two efficiently; second, the data routing mechanism is rigid, and the local data reuse of CNN and the global parameter sharing of RNN require completely different storage access modes. The fixed interconnection structure leads to low resource utilization. In addition, with the popularization of deep compression technology for network models, dynamic reconfigurability has become a core requirement for balancing computing accuracy and hardware overhead, and traditional accelerators lack runtime configuration capabilities and are difficult to adapt to the elastic changes of network structures.

[0004] In this context, there is an urgent need for a hardware architecture with programmable computing granularity and reconfigurable data flow, which can dynamically support the spatial local calculation of convolution operations and the global parallel processing of fully connected operations through a unified hardware base, thereby achieving universal acceleration of multimodal neural networks while ensuring energy efficiency. Summary of the invention

[0005] In view of this, the purpose of the present invention is to provide a reconfigurable systolic array that supports variable-granularity convolution and fully connected operations, the array consisting of 11×11 reconfigurable computing units and multiple reconfigurable selection gate modules, wherein: the reconfigurable computing units can be programmed to implement shift operations, multiplication operations, addition operations, and multiplication-accumulation operations through dynamic configuration control signals; the reconfigurable selection gate modules are used to implement dynamic routing selection of multiple data paths, and support 3×3, 5×5, 11×11 convolution operations and fully connected layer operations by reconstructing the connection relationship between input and output ports. The reconfigurable systolic array proposed in the present invention can support convolution operations and fully connected operations of different granularities through dynamic configuration, and has the ability to be universally adapted on different types of neural network accelerators.

[0006] In order to achieve the above object, the present invention provides the following technical solutions:

[0007] A reconfigurable systolic array based on supporting variable-granularity convolution and fully connected operations, the systolic array consists of a reconfigurable computing unit array arranged in an 11×11 matrix and a distributed interconnected reconfigurable selection gate module; the reconfigurable computing unit can be programmed to perform shift operations, multiplication operations, addition operations and multiplication-accumulation operations through dynamic configuration control signals; the reconfigurable selection gate module supports spatially configurable mapping of 3×3, 5×5, and 11×11 convolution kernels and matrix parallel expansion of fully connected layer operations by reconstructing the data path connection relationship of input and output ports.

[0008] The reconfigurable computing unit includes a PE register, a computing module, an adder, and an overflow truncation module, wherein the PE register is used to pre-store a 16-bit fixed-point number A; the computing module includes a 16×16-bit multiplier and a 32-bit adder; the computing module completes the multiplication and accumulation operation of the formula A×B+C within two clock cycles, wherein B is 16-bit flow data and C is 16-bit accumulated data.

[0009] Furthermore, the computing module included in the reconfigurable computing unit includes a configurable shift multiplier, a standard multiplier, and an overflow truncation module, wherein the configurable shift multiplier supports multiplication operations with a preset displacement; the standard multiplier performs multiplication operations without displacement correction; the two multipliers are dynamically switched through a mode selection signal.

[0010] Furthermore, the overflow truncation module included in the reconfigurable computing unit can implement sign bit extension detection for operation results greater than 16 bits, and perform dynamic bit width truncation according to a preset error threshold to ensure the output of 16-bit valid data and maintain data bit width consistency within the array.

[0011] The reconfigurable selection gate module realizes two-dimensional data flow reconstruction through row and column distributed routing nodes to support convolution operations and full connection operations with variable granularity.

[0012] The method supports 3×3, 5×5, and 11×11 convolution operations and fully connected layer operations, wherein the method supports monomer 11×11 convolution operations, and configures the systolic array as a single 11×11 convolution kernel as a whole; supports parallel operation of 4 5×5 convolution kernels, and divides the kernels into four independent 5×5 convolution calculation domains by isolating rows and columns; supports parallel operation of 11 3×3 convolution kernels, and flattens the 3×3 convolution kernel into 9 weight parameters and adds bias parameters to form 10 data, which occupy one row of the systolic array; supports fully connected operations, and configures 11×11, a total of 121 computing units, as a parallel multiplication-addition array of a single-layer fully connected network.

[0013] The beneficial effects of the present invention are:

[0014] (1) Through the coordinated configuration of the reconfigurable computing unit array and the reconfigurable selection gate module, the present invention can dynamically switch between convolution operations of different granularities (such as 3×3, 5×5, 11×11) and fully connected operation modes, solving the algorithm heterogeneity adaptation problem caused by the rigid architecture of traditional accelerators. The same hardware platform can efficiently support multimodal neural networks such as CNN and RNN, significantly reducing the deployment cost and energy consumption of edge devices.

[0015] (2) In convolution mode, multiple convolution kernels can be parallelized by row-column isolation or parameter expansion (e.g., flattening a 3×3 convolution kernel into 10 parameters); in fully connected mode, the array is configured as a global multiplication-addition network to maximize hardware resource utilization. Compared with the traditional fixed architecture, the resource idle rate is reduced by more than 30%, which is particularly suitable for real-time reasoning scenarios of lightweight models.

[0016] (3) The overflow truncation module embedded in the reconfigurable computing unit avoids precision loss while maintaining 16-bit data consistency through sign bit extension detection and dynamic bit width truncation. Combined with the low power consumption characteristics of the shift multiplier (such as shifting instead of full-bit multiplication), the overall energy efficiency ratio is improved by 20%-40%, which is especially suitable for low-power edge computing devices.

[0017] (4) Based on row-column distributed routing nodes, the reconfigurable selection gate module supports dynamic reconstruction of two-dimensional data flows. For example, row-wise accumulation transfer is implemented in 11×11 convolution mode, and independent computing domains are split in 5×5 parallel mode, which increases data reuse rate by more than 50%, reduces the frequency of external storage access, and reduces latency.

[0018] (5) For different convolution kernel sizes, dynamic hardware-level reconstruction (such as a single 11×11 kernel, four 5×5 kernels, or 11 3×3 kernels in parallel) significantly improves computing throughput. For example, in 3×3 mode, a single cycle can complete the multiplication and accumulation operations of 11 independent convolution kernels, which is suitable for dense small target detection scenarios.

[0019] (6) In the fully connected mode, input data and weight parameters are transmitted crosswise through rows and columns. By utilizing the parallel multiplication and addition capabilities of 121 computing units, the latency of a single-layer fully connected operation is reduced to 1 / 5 of that of a traditional architecture, making it particularly suitable for real-time reasoning of time series models such as LSTM.

[0020] (7) The modular design based on the 11×11 array can support larger-scale networks through cascading expansion while retaining dynamic reconstruction capabilities. In addition, the multiplication and accumulation operations and overflow truncation logic of the computing unit can adapt to different precision requirements (such as 8-bit / 16-bit mixed quantization), enhancing the versatility of the technology.

[0021] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:

[0023] Figure 1 The architecture diagram of the reconfigurable systolic array supporting variable granularity convolution and fully connected operations of the present invention;

[0024] Figure 2 It is an intermediate data flow diagram of the systolic array of the present invention when performing an 11×11 convolution operation;

[0025] Figure 3 It is an intermediate data flow diagram of the systolic array of the present invention when performing a 5×5 convolution operation;

[0026] Figure 4 It is an intermediate data flow diagram of the systolic array of the present invention when performing a 3×3 convolution operation;

[0027] Figure 5 An intermediate data flow diagram for a systolic array of the present invention when performing a fully connected operation;

[0028] Figure 6 This is a diagram of the internal structure of the reconfigurable computing unit of the present invention. DETAILED DESCRIPTION

[0029] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0030] Among them, the drawings are only used for illustrative explanations, and they only represent schematic diagrams rather than actual pictures, and should not be understood as limitations on the present invention. In order to better illustrate the embodiments of the present invention, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the size of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0031] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if the terms "upper", "lower", "left", "right", "front", "rear", etc. indicate the orientation or position relationship, they are based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation. Therefore, the terms describing the position relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0032] Figure 1 The architecture diagram of the reconfigurable systolic array supporting variable-granularity convolution and fully connected operations of the present invention includes a reconfigurable computing unit array arranged in an 11×11 matrix and a plurality of reconfigurable selection gate modules, wherein the addition parameter C and the output data Y of each computing unit are respectively connected to a reconfigurable selection gate, the fixed data A is connected to the A Buffer memory on the left side of the figure in a row-wise transmission manner through the solid line in the figure, and the flow data B is connected to the B Buffer memory on the bottom side of the figure in a column-wise transmission manner through the dotted line in the figure. In the configuration state, the output data Y of the computing unit can be connected to the reconfigurable selection gate connected to the addition parameter of the next computing unit in the same row by controlling the reconfigurable selection gate connected to the output data Y, or the output data Y can be directly connected to the output data Buffer of the systolic array.

[0033] Figure 2This is an intermediate data flow diagram of the systolic array of the present invention when performing an 11×11 convolution operation. It can be seen from the figure that the data stored in the fixed data A Buffer on the left is the weight parameter, the data stored in the flowing data B Buffer below is the input data, the output data Buffer is connected to the reconfigurable selection gate connected to the output data Y of the last column of computing units in the systolic array, and the output data Y of each computing unit is connected to the reconfigurable selection gate connected to the addition parameter of the next computing unit in the same row, and the data flow is transmitted in the row direction.

[0034] Embodiment: In this embodiment, the specific process of the systolic array when performing an 11×11 convolution operation is as follows: first, the convolution kernel coefficients stored in the fixed data A Buffer on the left are sent to each computing unit in the row direction; second, in the first cycle, the flowing data B The input data in the buffer, that is, the first column of the data mapped by the convolution box of the convolution kernel in the input feature map, is input into the first column calculation unit of the systolic array. In the second cycle, the second column of the data mapped by the convolution box of the convolution kernel in the input feature map is input into the first column calculation unit and the second column calculation unit of the systolic array. At the same time, the first column calculation unit completes the calculation of the first column input of the data mapped by the convolution box and the first column of the convolution kernel. In the third cycle, the third column of the data mapped by the convolution box of the convolution kernel in the input feature map is input into the first column calculation unit, the second column and the third column calculation unit of the systolic array. At the same time, the second column calculation unit multiplies the second column input of the data mapped by the product box with the second column of the convolution kernel and accumulates the calculation result of the first column, and so on. Finally, the first data result of the output feature map is obtained in the 13th cycle, and the second data result of the output feature map is obtained in the 14th cycle, and so on.

[0035] Figure 3 This is an intermediate data flow diagram of the systolic array of the present invention when performing a 5×5 convolution operation. It can be seen from the figure that the data stored in the fixed data A Buffer on the left is the weight parameter, the data stored in the flowing data B Buffer below is the input data, the output data Buffer is connected to the reconfigurable selection gate connected to the output data Y of the 5th and 11th columns of the systolic array computing units, and is divided into 4 5×5 array frames, and the computing units in the 6th row and 6th column are in an idle state.

[0036] Embodiment: In this embodiment, the specific process of the systolic array when performing a 5×5 convolution operation is as follows: first, the convolution kernel coefficients stored in the fixed data A Buffer on the left are sent to each calculation unit in the row direction; secondly, in the first cycle, the first column of the data mapped by the convolution box of each convolution kernel in the input feature map is input to the first column calculation unit and the sixth column calculation unit of the systolic array, and in the second cycle, the second column of the data mapped by the convolution box of each convolution kernel in the input feature map is input to the first column, the second column calculation unit and the sixth column, the seventh column calculation unit of the systolic array, and at the same time, the first column calculation unit completes the calculation of the first column of the data mapped by the product box and the first column of the convolution kernel, and so on; finally, in the seventh cycle, the first data results of the feature maps output by the four convolution kernels are obtained, and the second data results of the feature maps output by the four convolution kernels are obtained, and so on.

[0037] Figure 4 This is an intermediate data flow diagram of the systolic array of the present invention when performing a 3×3 convolution operation. It can be seen from the figure that the data stored in the fixed data A Buffer on the left is the weight parameter, the data stored in the flowing data B Buffer below is the input data, the output data Buffer is connected to the reconfigurable selection gate connected to the output data Y of the 10th column computing unit of the systolic array, and is divided into 11 row frames, and the computing unit in the 11th column is in an idle state.

[0038] Embodiment: In this embodiment, the specific process of the systolic array when performing a 3×3 convolution operation is as follows: first, the convolution kernel coefficients and bias parameters stored in the fixed data A Buffer on the left are sent to each calculation unit in the row direction; secondly, in the first cycle, the first data of the input feature map is input in parallel to the first column calculation unit of the systolic array; in the second cycle, the second data of the input feature map is input in parallel to the first and second column calculation units of the systolic array; at the same time, the first column calculation unit calculates the first input and the first data of the convolution kernel, and so on; finally, in the 12th cycle, the systolic array outputs the first data result of the feature map corresponding to each convolution kernel in each row, and so on.

[0039] Figure 5 The figure is an intermediate data flow diagram of the systolic array of the present invention when performing a fully connected operation. It can be seen from the figure that the data stored in the fixed data A Buffer on the left is the input data, the data stored in the flowing data B Buffer below is the weight parameter, the output data Buffer is connected to the reconfigurable selection gate connected to the output data Y of the last column of computing units in the systolic array, and the output data Y of each computing unit is connected to the reconfigurable selection gate connected to the addition parameter of the next computing unit in the same row, and the data flow is transmitted in the row direction.

[0040] Embodiment: In this embodiment, the specific process of the systolic array when performing a full connection operation is as follows: first, the input data vector stored in the fixed data A Buffer on the left is sent to each computing unit in order in the row direction; secondly, the weight parameters of the corresponding numbers, such as the weight parameters of the first column computing unit corresponding to 1+11*0, 1+11*1, 1+11*2, 1+11*3...1+11*10, are passed into the first column computing unit in the first cycle, and the weight parameters of the corresponding numbers, such as the weight parameters of the second column computing unit corresponding to 2+11*0, 2+11*1, 2+11*2, 2+11*3...2+11*10, are passed into the second column computing unit in the second cycle, and so on; finally, the first data of the full connection operation result is obtained in the 13th cycle, and the second data of the full connection operation result is obtained in the 14th cycle, and so on.

[0041] Figure 6 The internal structure diagram of the reconfigurable computing unit of the present invention includes a PE register, a computing module, an adder, and an overflow truncation module, wherein the PE register is used to pre-store a 16-bit fixed-point number A; the computing module includes a shift multiplier and a standard multiplier, which can complete the multiplication and accumulation operation of the formula A×B+C within 2 clock cycles, wherein B is 16-bit flow data and C is 16-bit accumulated data.

[0042] Embodiment: In this embodiment, the specific calculation process of the reconfigurable computing unit is as follows: first, fixed data A is received and stored in the PE calculator; second, in the first cycle, the flowing data B is received and transmitted to the computing module; when the control signal Mode is 1, a shift multiplication operation is performed, specifically, the 16-bit B data is shifted right by A[2:0] bits, and when A[3] is 0, the shift result is directly output, and when A[3] is 1, the shift inversion result is output, and finally a 16-bit number is output; when the control signal Mode is 0, a standard multiplication operation is performed, the 16-bit A data is multiplied by the 12-bit data {B[7:0], 0000}, and the output 27-bit data is overflow judged and truncated, and finally a 16-bit number is output; finally, the addition parameter C is received in the second cycle, it is added to the 16-bit number output by the computing module, and the addition result is judged to be overflow truncated, and finally a 16-bit number is output.

[0043] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.

Claims

1. A reconfigurable systolic array that supports variable-granularity convolution and fully connected operations, characterized by: include: A reconfigurable computing unit array arranged in an 11×11 matrix. Each computing unit can be programmed to perform shift operations, multiplication operations, addition operations, and multiplication-accumulation operations through dynamic configuration control signals; Multiple reconfigurable selection gate modules support spatially configurable mapping of 3×3, 5×5, and 11×11 convolution kernels and matrix parallel expansion of fully connected layer operations by reconfiguring the data path connection relationship of the input and output ports; The reconfigurable selection gate module realizes two-dimensional data flow reconstruction through row and column distributed routing nodes to dynamically switch the data interaction mode between local convolution calculation and global full connection operation.

2. The reconfigurable systolic array supporting variable granularity convolution and fully connected operations according to claim 1, characterized in that: The reconfigurable computing unit comprises: PE register, used to pre-store 16-bit fixed-point number A; The calculation module includes a shift multiplier and a standard multiplier, which is used to complete the multiplication and accumulation operation of the formula A×B+C within 2 clock cycles, where B is 16-bit flow data and C is 16-bit accumulation data; The overflow truncation module is used to perform sign bit extension detection on the operation result and perform dynamic bit width truncation according to a preset error threshold to maintain the consistency of the 16-bit output data.

3. The reconfigurable systolic array supporting variable granularity convolution and fully connected operations according to claim 2, characterized in that: The shift multiplier in the computing module supports multiplication operations with preset shift amounts, and the standard multiplier performs multiplication operations without shift correction, and the two are dynamically switched through a mode selection signal.

4. The reconfigurable systolic array supporting variable granularity convolution and fully connected operations according to claim 1, characterized in that: The configuration of the reconfigurable selection gate module includes: In the 11×11 convolution mode, the entire array is configured as a single 11×11 convolution kernel, and the output data is connected to the external cache through the output port of the last column of computational units; In 5×5 convolution mode, the array is divided into four independent 5×5 computational domains by row and column isolation, and the output port of each domain is connected to the external cache; In the 3×3 convolution mode, the 3×3 convolution kernel parameters are expanded to 10 data (including bias parameters), and 11 independent convolution kernels are processed in parallel through a single row of computing units.

5. The reconfigurable systolic array supporting variable granularity convolution and fully connected operations according to claim 4, characterized in that: Under the fully connected operation, the reconfigurable selection gate module configures the 11×11 computing units into a parallel multiplication-addition array of a single-layer fully connected network, in which input data is transmitted by row, weight parameters are transmitted by column, and output data is connected to the external cache through the output port of the last column of computing units.

6. The reconfigurable systolic array supporting variable granularity convolution and fully connected operations according to claim 2, characterized in that: The operation of the shift multiplier includes: adjusting the input data B by a preset shift amount according to a control signal, and inverting the result when the sign bit of A is triggered.

7. The reconfigurable systolic array supporting variable granularity convolution and fully connected operations according to claim 2, characterized in that: The operation of the standard multiplier includes: multiplying a 16-bit fixed-point number A by 12-bit extended data (the lower 8 bits of B are padded with zeros), and performing bit truncation processing on a 27-bit intermediate result to output 16-bit valid data.

8. The reconfigurable systolic array supporting variable granularity convolution and fully connected operations according to claim 1, characterized in that: The data flow control of the systolic array includes: Fixed data A is connected to the left cache via a row-wise transmission link, and flowing data B is connected to the lower cache via a column-wise transmission link; The output port of each computing unit is connected to the accumulation input port of the adjacent computing unit or directly to the external cache through a reconfigurable selection gate module.

9. The reconfigurable systolic array supporting variable granularity convolution and fully connected operations according to claim 4, characterized in that: The parallel computing in the 3×3 convolution mode includes: injecting the sliding window data of the input feature map into the array column by column, each row of computing units independently completes the multiplication and accumulation operation of a single 3×3 convolution kernel, and realizing resource optimization by truncating idle columns.

10. The reconfigurable systolic array supporting variable granularity convolution and fully connected operations according to claim 1, characterized in that: Under the fully connected operation, the input data vector is distributed to each computing unit in row order, the weight parameters are distributed in column order, and the output of each column computing unit is transmitted row by row through the accumulation link, and finally the result is output by the last column.