Channel-by-channel convolution device and operation method thereof

By designing a channel-by-channel convolution device, using a channel-by-channel convolution control module and multiple floating-point operation units for multi-channel multiplication and addition operations, the problem of large amount of channel-by-channel convolution calculation in depth separation convolution is solved, and efficient calculation is achieved and the model inference speed is improved.

CN120010922AActive Publication Date: 2025-05-16SHANGHAI BIREN TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510489160.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-05-16
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently realize channel-by-channel convolution in depth separation convolution operations, resulting in large amounts of calculation and parameters, affecting the model's inference speed and operation efficiency.

Method used

A channel-by-channel convolution device is designed, including a channel-by-channel convolution control module, a convolution kernel input control module, a feature map input control module and multiple floating-point operation units. Through the instruction scheduling of the channel-by-channel convolution control module, broadcasting and calling the convolution kernel and feature maps is realized, and multi-channel multiplication and addition operations are performed to generate channel-by-channel convolution results.

Benefits of technology

Through the design of this device, the calculation amount and parameter amount can be significantly reduced, the inference speed and operation efficiency of deep separable convolution can be improved, and it is suitable for deep learning models in artificial intelligence chips.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010922A_ABST
    Figure CN120010922A_ABST
Patent Text Reader

Abstract

The invention provides a channel-by-channel convolution device and an operation method thereof. The channel-by-channel convolution device comprises a channel-by-channel convolution control module, a convolution kernel input control module, a feature map input control module and a floating point arithmetic unit. And the convolution kernel input control module broadcasts the convolution kernel to the floating point arithmetic unit. The feature map input control module calls a part of matrix in a receptive field from the feature map matrix to the floating point arithmetic unit. The floating point arithmetic unit performs multiply-add calculations using the plurality of matrix elements of the receptive field and the plurality of weight elements of the convolution kernel in a first cycle of the multi-channel multiply-add operation to generate a phase sum value. The floating point arithmetic unit performs an accumulation calculation using the phase sum value and the previous accumulated value in a second cycle of the multi-channel multiply-add operation to generate an accumulation result of the multi-channel multiply-add operation. The floating point arithmetic unit calculates result elements in the channel-by-channel convolution result matrix based on an accumulation result of the multi-channel multiply-add operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence (AI) chip technology, and in particular to an electronic device for performing depthwise separable convolution, and particularly to a channel-by-channel convolution device for performing depthwise convolution in depthwise separable convolution and an operation method thereof. Background Art

[0002] Computing devices such as artificial intelligence chips can provide huge computing power. The huge computing power of artificial intelligence chips comes from a large number of internal hardware cores. An AI chip usually contains multiple programmable processors, such as a stream processor cluster (SPC). Each programmable processor usually contains multiple compute units (CU, or compute cores), and each compute unit usually contains multiple execution units (EU, or execution cores), such as integer (Integer, INT) core, floating point (Floating Point, FP) core, tensor core (Tcore) and / or vector core (Vcore) at least one. By organizing various types of computing units through programming, programmable multiprocessors can support general computing, scientific computing, and neural network computing.

[0003] Depthwise separable convolution is a lightweight convolution operation. The operation process of depthwise separable convolution includes two stages: channel-by-channel convolution and pointwise convolution. Depthwise separable convolution can effectively reduce the amount of calculation and parameters, and improve the reasoning speed and operation efficiency of the model. Assuming that the size of the input feature map is H×W, the number of input channels is N, the size of the convolution kernel is K×K, and the number of output channels is M, the amount of calculation of the standard convolution is H×W×N×K 2 ×M, the number of parameters is N×K 2 ×M. Under the same assumption, for depth-wise separable convolution, the number of parameters in the first stage (channel-by-channel convolution) is K 2 ×N, the number of parameters in the second stage (point-by-point convolution) is N×M (assuming the size of the convolution kernel is 1×1). Therefore, the total number of parameters for depthwise separable convolution is K 2 ×N + N×M = N×(K 2 + M). In general, N×(K 2 + M) than N×K2 ×M is much smaller.

[0004] Like standard convolution, depthwise separable convolution can be calculated using General Matrix Multiplication (GEMM) hardware, or by using Central Processing Unit (CPU) hardware, or by using Vector Unit (Vector Unit) hardware of Graphics Processing Unit (GPU). How to implement hardware specifically for depthwise separable convolution is one of the many technical issues in this field. Summary of the invention

[0005] The present invention is directed to a channel-by-channel convolution device and an operation method thereof, which are used to perform the first stage "channel-by-channel convolution" in a depth-wise separable convolution operation.

[0006] In an embodiment of the present invention, the channel-by-channel convolution device includes a channel-by-channel convolution control module, a convolution kernel input control module, a feature map input control module, and a plurality of floating-point operation units. The channel-by-channel convolution control module is used to execute the channel-by-channel convolution instruction emitted by the instruction scheduling module. The convolution kernel input control module is coupled to the channel-by-channel convolution control module. The feature map input control module is coupled to the channel-by-channel convolution control module. A plurality of floating-point operation units are coupled to the convolution kernel input control module and the feature map input control module. The convolution kernel input control module broadcasts the convolution kernel to the plurality of floating-point operation units based on the control of the channel-by-channel convolution control module. The feature map input control module calls the first partial matrix in the first receptive field from the feature map matrix to the first floating-point operation unit among the plurality of floating-point operation units based on the control of the channel-by-channel convolution control module. The first floating-point operation unit performs a first multi-channel multiplication-addition operation based on the control of the channel-by-channel convolution control module. The first floating-point operation unit performs multiplication and addition calculations using a plurality of first matrix elements of the first receptive field and a plurality of first weight elements of the convolution kernel in a first cycle of the first multi-channel multiplication and addition operation to generate a first stage sum value. The first floating-point operation unit performs accumulation calculations using the first stage sum value and the previous accumulation value in a second cycle of the first multi-channel multiplication and addition operation to generate a first accumulation result of the first multi-channel multiplication and addition operation. The first floating-point operation unit calculates a first result element in a channel-by-channel convolution result matrix based on the first accumulation result of the first multi-channel multiplication and addition operation.

[0007] In an embodiment of the present invention, the channel-by-channel convolution device includes: a channel-by-channel convolution control module executes a channel-by-channel convolution instruction issued by an instruction scheduling module; a convolution kernel input control module broadcasts the convolution kernel to the multiple floating-point operation units based on the control of the channel-by-channel convolution control module; a feature map input control module calls a first part of the matrix in the first receptive field from the feature map matrix to the first floating-point operation unit among the multiple floating-point operation units based on the control of the channel-by-channel convolution control module; the first floating-point operation unit performs a first multi-channel multiplication and addition operation based on the control of the channel-by-channel convolution control module; the first floating-point operation unit uses a plurality of first matrix elements of the first receptive field and a plurality of first weight elements of the convolution kernel to perform multiplication and addition calculations to generate a first stage sum value in a first cycle of the first multi-channel multiplication and addition operation; the first floating-point operation unit uses the first stage sum value and the previous accumulated value to perform an accumulation calculation in a second cycle of the first multi-channel multiplication and addition operation to generate a first accumulated result of the first multi-channel multiplication and addition operation; and the first floating-point operation unit calculates a first result element in the channel-by-channel convolution result matrix based on the first accumulated result of the first multi-channel multiplication and addition operation.

[0008] Based on the above, the channel-by-channel convolution control module of the channel-by-channel convolution device controls the convolution kernel input control module, the feature map input control module and the plurality of floating-point operation units to perform one or more multi-channel multiplication and addition operations on the partial matrix in the receptive field. After completing the one or more multi-channel multiplication and addition operations, the floating-point operation unit can calculate a result element in the channel-by-channel convolution result matrix. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 is a schematic diagram of a circuit block of a channel-by-channel convolution device according to an embodiment of the present invention; Figure 2 is a flow chart of an operation method of a channel-by-channel convolution device according to an embodiment of the present invention; Figure 3 is a schematic diagram of channel-by-channel convolution according to an embodiment of the present invention; Figure 4 is a schematic diagram of data layout in which a feature map matrix is ​​stored in a register group in a one-dimensional data layout manner according to an embodiment of the present invention; Figure 5 is a schematic diagram of a floating point operation unit performing multiple multi-channel multiplication and addition operations according to an embodiment of the present invention; Figure 6 is a timing diagram of a plurality of multi-channel multiplication and addition operations according to an embodiment of the present invention; Figure 7 is a schematic diagram of a circuit module of a floating-point operation unit according to an embodiment of the present invention; Figure 8 is a circuit module schematic diagram of a first multiplication circuit, a second multiplication circuit, an addition circuit and an accumulation circuit according to an embodiment of the present invention; Fig. 9 It is a circuit module diagram of a channel-by-channel convolution control module, a feature map input control module and a convolution kernel input control module according to an embodiment of the present invention. DETAILED DESCRIPTION

[0010] Reference will now be made in detail to exemplary embodiments of the present invention, examples of which are illustrated in the accompanying drawings. Whenever possible, the same reference numerals are used in the drawings and the description to refer to the same or like parts.

[0011] The term "coupled (or connected)" used in the entire specification of this case (including the claims) may refer to any direct or indirect means of connection. For example, if the text describes a first device coupled (or connected) to a second device, it should be interpreted that the first device can be directly connected to the second device, or the first device can be indirectly connected to the second device through other devices or some connection means. The terms "first", "second", etc. mentioned in the entire specification of this case (including the claims) are used to name the components (element) or to distinguish different embodiments or ranges, and are not used to limit the upper or lower limit of the number of components, nor to limit the order of components.

[0012] In addition, it should be noted that the channel-by-channel convolution device and its operating method provided in at least one embodiment of the present disclosure can be applied to artificial intelligence chips or other chips. The parameter type of the channel-by-channel convolution device and its operating method is a floating point number, and the specific physical meaning of the parameter type may be different depending on the application scenario. For example, the channel-by-channel convolution device and its operating method provided in at least one embodiment of the present disclosure can be applied in the fields of speech processing, image processing, text processing, video processing, etc.

[0013] For example, in the field of speech processing, parameters can be any parameters used, input, or generated in tasks such as feature extraction, speech enhancement, and speech recognition, such as speech feature vectors, filtering parameters, and the like.

[0014] For example, in the field of image processing, parameters can be any parameters used, input, or generated in tasks such as image preprocessing, feature extraction, image segmentation, and target detection, such as image feature vectors, various edge detection operators (such as Sobel operator, Canny operator, Prewitt operator, etc.), image filtering operators (such as Gaussian filtering, median filtering, bilateral filtering, etc.), morphological operators (such as corrosion, expansion, opening operation, closing operation, etc.), etc.

[0015] For example, in the field of text processing, parameters can be any parameters used, input, or generated in tasks such as text classification, sentiment analysis, and text generation, such as the semantic feature vector of the text.

[0016] For example, in the field of video processing, the parameters may be parameters in the field of image processing as described above, or parameters used, inputted, or generated in the field of video processing, such as an optical flow operator (used to estimate motion between video frames) or a target tracking operator (used to track a specific target in a video).

[0017] Of course, the present disclosure is not limited to this, and for other application scenarios or fields, as long as channel-by-channel convolution is required, the channel-by-channel convolution device and its operation method described in at least one embodiment of the present disclosure can be applied, and they will not be described one by one here.

[0018] Figure 1 1 is a circuit module diagram of a channel-by-channel convolution device 100 according to an embodiment of the present invention. The operation process of depth-wise separable convolution includes two stages: channel-by-channel convolution and point-by-point convolution. The channel-by-channel convolution device 100 can be used as a special hardware for "accelerating the channel-by-channel convolution stage in the depth-wise separable convolution operation". The directions of channel-by-channel convolution that the channel-by-channel convolution device 100 can support include forward convolution and backward convolution. The sizes of channel-by-channel convolution kernels that the channel-by-channel convolution device 100 can support include 1×1, 2×2, 3×3, 4×4, 5×5 or other sizes. The x-direction padding that the channel-by-channel convolution device 100 can support includes 0, +1, -1, +2, -2 or other padding. The y-direction padding that the channel-by-channel convolution device 100 can support includes 0, +1, -1, +2, -2 or other padding. The stride that the channel-by-channel convolution device 100 can support includes 1, 2 or other strides. The dilation that the channel-by-channel convolution device 100 can support includes 1, 2 or other dilations. The data type that the channel-by-channel convolution device 100 can support includes fp32, fp16, fp8, int16, int8, int4 or other data types.

[0019] When accelerating the channel-by-channel convolution application, the instruction scheduling module 11 transmits the channel-by-channel convolution instruction to the vector computing unit. Based on the channel-by-channel convolution instruction transmitted by the instruction scheduling module 11, the channel-by-channel convolution device 100 can perform the channel-by-channel convolution in the depthwise separable convolution, and then store the channel-by-channel convolution result matrix in a memory or register (e.g. Figure 1 The channel-by-channel convolution result matrix in the memory (or register group) can be used by the second stage "point-by-point convolution" of the depth-wise separable convolution operation.

[0020] exist Figure 1 In the illustrated embodiment, the channel-by-channel convolution device 100 includes a channel-by-channel convolution control module 110, a feature map input control module 120, a convolution kernel input control module 130, a register file 140, a constant cache 150, and a plurality of floating point operation units (e.g. Figure 1 The first floating-point operation unit 160_1, the second floating-point operation unit 160_2, ..., the Nth floating-point operation unit 160_N are shown, wherein the first floating-point operation unit may refer to any one of the first floating-point operation unit 160_1, the second floating-point operation unit 160_2, ..., the Nth floating-point operation unit 160_N). According to different designs, in some embodiments, at least one of the instruction scheduling module 11, the channel-by-channel convolution control module 110, the feature map input control module 120, and the convolution kernel input control module 130 may be implemented as a hardware circuit. In other embodiments, at least one of the instruction scheduling module 11, the channel-by-channel convolution control module 110, the feature map input control module 120, and the convolution kernel input control module 130 may be implemented as a combination of hardware, firmware, and software (i.e., program).

[0021] In terms of hardware, at least one of the above-mentioned instruction scheduling module 11, channel-by-channel convolution control module 110, feature map input control module 120, and convolution kernel input control module 130 can be implemented in a logic circuit on an integrated circuit. For example, the relevant functions of at least one of the instruction scheduling module 11, channel-by-channel convolution control module 110, feature map input control module 120, and convolution kernel input control module 130 can be implemented in various logic blocks, modules, and circuits in one or more controllers, hardware controllers, microcontrollers, hardware processors, microprocessors, application-specific integrated circuits (ASIC), digital signal processors (DSP), field programmable gate arrays (FPGA), central processing units, or other processing units. The relevant functions of at least one of the instruction scheduling module 11, the channel-by-channel convolution control module 110, the feature map input control module 120 and the convolution kernel input control module 130 can be implemented as hardware circuits, such as various logic blocks, modules and circuits in an integrated circuit, using hardware description languages ​​(such as Verilog HDL or VHDL) or other suitable programming languages.

[0022] In the form of "software or firmware running on hardware", the relevant functions of at least one of the above-mentioned instruction scheduling module 11, channel-by-channel convolution control module 110, feature map input control module 120 and convolution kernel input control module 130 can be implemented as programming codes. For example, at least one of the instruction scheduling module 11, channel-by-channel convolution control module 110, feature map input control module 120 and convolution kernel input control module 130 is implemented using general programming languages ​​(programming languages, such as C, C++ or assembly language) or other suitable programming languages. The programming code can be recorded or stored in a "non-transitory machine-readable storage medium". In some embodiments, the non-transitory machine-readable storage medium includes, for example, a semiconductor memory and (or) a storage device. An electronic device (such as a CPU, a hardware controller, a microcontroller, a hardware processor or a microprocessor) can read and execute programming code from a non-temporary machine-readable storage medium to implement the relevant functions of at least one of the instruction scheduling module 11, the channel-by-channel convolution control module 110, the feature map input control module 120 and the convolution kernel input control module 130.

[0023] Figure 1 The constant cache 150 and the register group 140 shown can be used as internal components of the channel-by-channel convolution device 100, but in other embodiments, at least one of the constant cache 150 and the register group 140 can be used as an external component of the channel-by-channel convolution device 100. This embodiment does not limit the implementation of the constant cache 150 and the register group 140. For example, at least one of the constant cache 150 and the register group 140 can be a register, a cache, a main memory, or other memory. The constant cache 150 is used to store the convolution kernel. The constant cache 150 is coupled to the convolution kernel input control module 130 to provide the convolution kernel. The register group 140 is used to store the feature map matrix. The register group 140 is coupled to the feature map input control module 120 to provide the feature map matrix.

[0024] The number N of the first floating-point operation unit 160_1 to the Nth floating-point operation unit 160_N can be determined according to the actual design. For example, assuming that the feature map matrix is ​​a u×v matrix, the number N of the first floating-point operation unit 160_1 to the Nth floating-point operation unit 160_N can be any integer in the range of 1 to u×v in one embodiment. If the number N of the first floating-point operation unit 160_1 to the Nth floating-point operation unit 160_N is u×v and the stride is 1, the channel-by-channel convolution device 100 can complete the channel-by-channel convolution of a feature map matrix by performing a single iteration. If the number N of the first floating-point operation unit 160_1 to the Nth floating-point operation unit 160_N is 1 and the stride is 1, the channel-by-channel convolution device 100 needs to perform u×v iterations to complete the channel-by-channel convolution of a feature map matrix.

[0025] Figure 2 FIG. 1 is a flow chart of an operation method of a channel-by-channel convolution device according to an embodiment of the present invention. Figure 1 and Figure 2 , in step S210, the channel-by-channel convolution control module 110 executes the channel-by-channel convolution instruction issued by the instruction scheduling module 11. The feature map input control module 120 and the convolution kernel input control module 130 are coupled to the channel-by-channel convolution control module 110. The channel-by-channel convolution control module 110 controls the feature map input control module 120 and the convolution kernel input control module 130 based on the channel-by-channel convolution instruction. The first floating point operation unit 160_1 to the Nth floating point operation unit 160_N are coupled to the convolution kernel input control module 130 and the feature map input control module 120. In step S220, the convolution kernel input control module 130 calls the convolution kernel from the constant cache 150 based on the control of the channel-by-channel convolution control module 110, and broadcasts the convolution kernel to each of the first floating point operation unit 160_1 to the Nth floating point operation unit 160_N. In addition, in step S220, the feature map input control module 120 calls a partial matrix in the receptive field (which can be generally referred to as the first partial matrix in the first receptive field) from the feature map matrix of the register group 140 based on the control of the channel-by-channel convolution control module 110 to the corresponding ones of the 1st floating-point operation unit 160_1 to the Nth floating-point operation unit 160_N (which can be generally referred to as the first floating-point operation unit).

[0026] Figure 3 FIG. 1 is a schematic diagram of channel-by-channel convolution according to an embodiment of the present invention. Figure 3 In the example shown, the feature map matrix 310 is assumed to be a 6×16 matrix, and the convolution kernel 320 is assumed to be a 3×3 matrix. Figure 1 and Figure 3Based on the control of the channel-by-channel convolution control module 110, the convolution kernel input control module 130 broadcasts the convolution kernel 320 to each of the first floating-point operation unit 160_1 to the Nth floating-point operation unit 160_N. Taking the first floating-point operation unit as the first floating-point operation unit 160_1 as an example, the feature map input control module 120 calls the corresponding partial matrix (which can be generally referred to as the first partial matrix in the first receptive field, and the matrix elements in the corresponding partial matrix include "Pad", "Pad", "Pad", "Pad", "1", "2", "Pad", "17" and "18") from the feature map matrix 310 to the corresponding first floating-point operation unit 160_1. Among them, "Pad" means zero padding. The first floating-point operation unit 160_1 uses the corresponding partial matrix elements of the feature map matrix 310 and the convolution kernel 320 to calculate the corresponding result elements (Pad·W1 + Pad·W2 + Pad·W3 + Pad·W4 +1·W5 + 2·W6 + Pad·W7 + 17·W8 + 18·W9) in the channel-by-channel convolution result matrix (which can be generally referred to as the first result elements), where "W1" to "W9" are weight elements in the convolution kernel 320.

[0027] The channel-by-channel convolution device 100 moves the multiplication-addition window corresponding to the convolution kernel 320 to different positions of the feature map matrix 310 according to the step size, thereby forming different receptive fields. Then, the channel-by-channel convolution device 100 performs multiplication-addition calculations on the data elements (data elements, such as feature points) of the partial matrix in the receptive field after each movement of the multiplication-addition window to obtain a result element in the channel-by-channel convolution result matrix. After the multiplication-addition window completely scans the entire feature map matrix 310 according to the stride, the channel-by-channel convolution device 100 can store the channel-by-channel convolution result matrix in the register group 140. For example (taking the first floating-point operation unit as the second floating-point operation unit 160_2 as an example), assuming that the step size is 1, the feature map input control module 120 calls the corresponding partial matrix (which can be generally referred to as the first partial matrix in the first receptive field, and the matrix elements in the corresponding partial matrix include "Pad", "Pad", "Pad", "1", "2", "3", "17", "18" and "19") from the feature map matrix 310 to the corresponding second floating-point operation unit 160_2. The second floating-point operation unit 160_2 uses the corresponding partial matrix elements of the feature map matrix 310 and the convolution kernel 320 to calculate the corresponding result element (Pad·W1 +Pad·W2 + Pad·W3 + 1·W4 + 2·W5 + 3·W6 + 17·W7 + 18·W8 + 19·W9) in the channel-by-channel convolution result matrix (relative to the first result element which can be generally referred to as the first result element of the above-mentioned first floating-point operation unit 160_1, this result element can be generally referred to as the second result element). The remaining floating point operation units (eg, the Nth floating point operation unit 160_N) can refer to the related description of the first floating point operation unit 160_1 and the second floating point operation unit 160_2 and be deduced by analogy. The first floating point operation unit 160_1 to the Nth floating point operation unit 160_N each calculate at least one result element in the channel-by-channel convolution result matrix.

[0028] There are multiple layout modes for storing the feature map matrix 310 in the register group 140, such as one-dimensional data layout and two-dimensional data layout. Figure 4 3 is a schematic diagram of data layout of a feature map matrix 310 stored in a register set 140 in a one-dimensional data layout manner according to an embodiment of the present invention. Figure 4 The upper part again shows the feature map matrix 310 . Figure 4 The lower part shows a schematic diagram of the data layout of the first register 1, the second register 2, the third register 3, the fourth register 4, the fifth register 5, the sixth register 6, the seventh register 7, the eighth register 8 and the ninth register 9 in the register group 140. Figure 4In the illustrated example, the one-dimensional data layout is assumed to be in row priority mode, the stride is assumed to be 1, and the zero padding width is assumed to be 1, where “Pad” indicates padding with zeros at the boundary of the feature map matrix 310 .

[0029] A partial matrix of the feature map matrix 310 in the first receptive field (generally referred to as a first partial matrix) is stored in the first address of each of the first register 1 to the ninth register 9 (i.e., a plurality of registers) in the register group 140, for example Figure 4 The first column on the left of the lower part. The partial matrix of the feature map matrix 310 in the second receptive field (generally referred to as the second partial matrix) is stored in the second address of each of the first register 1 to the ninth register 9 (i.e., multiple registers) in the register group 140, for example Figure 4 The second column from the left of the lower part. Based on the control of the channel-by-channel convolution control module 110, the feature map input control module 120 calls the first part of the feature map matrix 310 from the first address of each of the first register 1 to the ninth register 9 (i.e., multiple registers) to the first floating-point operation unit 160_1 (which can be generally referred to as the first floating-point operation unit). Based on the control of the channel-by-channel convolution control module 110, the feature map input control module 120 calls the second part of the feature map matrix 310 from the second address of each of the first register 1 to the ninth register 9 (i.e., multiple registers) to the second floating-point operation unit 160_2 (which can be generally referred to as the second floating-point operation unit). The remaining floating-point operation units (e.g., the Nth floating-point operation unit 160_N) can refer to the relevant description of the first floating-point operation unit 160_1 and the second floating-point operation unit 160_2 and make analogies.

[0030] Please refer to Figure 1 and Figure 2In step S230, each of the 1st floating-point operation unit 160_1 to the Nth floating-point operation unit 160_N performs one or more multi-channel multiplication and addition operations based on the control of the channel-by-channel convolution control module 110. Each multi-channel multiplication and addition operation includes a first cycle and a second cycle. Each of the 1st floating-point operation unit 160_1 to the Nth floating-point operation unit 160_N (generally referred to as the first floating-point operation unit) uses a plurality of matrix elements (generally referred to as a plurality of first matrix elements) of a corresponding receptive field (generally referred to as the first receptive field) and a plurality of weight elements (generally referred to as a plurality of first weight elements) of a convolution kernel to perform multiplication and addition calculations in the first cycle of the multi-channel multiplication and addition operation to generate a stage sum value (generally referred to as a first stage sum value) (step S240). Each of the 1st floating-point operation unit 160_1 to the Nth floating-point operation unit 160_N (generally referred to as the first floating-point operation unit) uses the stage sum value (generally referred to as the first stage sum value) and the previous accumulated value to perform accumulation calculation in the second cycle of the multi-channel multiplication and addition operation (generally referred to as the first multi-channel multiplication and addition operation) to generate an accumulation result (generally referred to as the first accumulation result) of the multi-channel multiplication and addition operation (generally referred to as the first multi-channel multiplication and addition operation) (step S250). In the absence of a previous accumulated value, the floating-point operation unit (generally referred to as the first floating-point operation unit) uses the stage sum value and the initial value (or bias value) to perform accumulation calculation in the second cycle. The initial value (or bias value) can be determined based on actual design and application. For example, the initial value (or bias value) can be 0 or other values. In step S260, each of the 1st floating-point operation unit 160_1 to the Nth floating-point operation unit 160_N (generally referred to as the first floating-point operation unit) calculates the corresponding result element (generally referred to as the first result element) in the channel-by-channel convolution result matrix based on the accumulated result (generally referred to as the first accumulated result) of the multi-channel multiplication-addition operation (generally referred to as the first multi-channel multiplication-addition operation).

[0031] Figure 5 1 is a schematic diagram of a floating point operation unit (e.g., the first floating point operation unit 160_1) performing multiple multi-channel multiplication and addition operations according to an embodiment of the present invention. The remaining second floating point operation units 160_2 to the Nth floating point operation unit 160_N can refer to the relevant description of the first floating point operation unit 160_1 and can be deduced by analogy. Figure 5In the illustrated embodiment, the first floating-point arithmetic unit 160_1 performs five multi-channel multiplication-addition operations, namely, a first multi-channel multiplication-addition operation 510, a second multi-channel multiplication-addition operation 520, a third multi-channel multiplication-addition operation 530, a fourth multi-channel multiplication-addition operation 540, and a fifth multi-channel multiplication-addition operation 550. The multiplication-addition calculations performed in the first cycle of each of the first multi-channel multiplication-addition operation 510 to the fifth multi-channel multiplication-addition operation 550 (which can be generally referred to as the first multi-channel multiplication-addition operation; the first multi-channel multiplication-addition operation of two consecutive multi-channel multiplication-addition operations can also be generally referred to as the first multi-channel multiplication-addition operation, and the second multi-channel multiplication-addition operation of two consecutive multi-channel multiplication-addition operations can be generally referred to as the second multi-channel multiplication-addition operation) include half-precision multiplication and half-precision addition, and the stage sum value of the first cycle (when the first cycle is in the first multi-channel multiplication-addition operation, it can be generally referred to as the second multi-channel multiplication-addition operation) includes half-precision multiplication and half-precision addition, and the stage sum value of the first cycle (when the first cycle is in the first multi-channel multiplication-addition operation, it can be generally referred to as the second multi-channel multiplication-addition operation) includes half-precision multiplication and half-precision addition. The first cycle is generally referred to as the first stage sum value, and when the first cycle is in the second multi-channel multiplication-addition operation, it can be generally referred to as the second stage sum value) is a half-precision floating point number, the accumulation calculation performed in the second cycle of each of the first multi-channel multiplication-addition operation 510 to the fifth multi-channel multiplication-addition operation 550 includes a full-precision addition, and the accumulation result of the second cycle (when the second cycle is in the first multi-channel multiplication-addition operation, it can be generally referred to as the first accumulation result, and when the second cycle is in the second multi-channel multiplication-addition operation, it can be generally referred to as the second accumulation result) is a full-precision floating point number.

[0032] In an exemplary embodiment, the plurality of first matrix elements of the first receptive field include a first element and a second element, and the plurality of first weight elements of the convolution kernel include a first weight and a second weight.

[0033] exist Figure 5In the illustrated embodiment, the feature map input control module 120 calls the corresponding partial matrix (which can be generally referred to as the first partial matrix in the first receptive field, and the matrix elements in the corresponding partial matrix include "Pad", "Pad", "Pad", "Pad", "1", "2", "Pad", "17" and "18") from the feature map matrix 310 in the register group 140 to the first floating-point operation unit 160_1 (which can be generally referred to as the first floating-point operation unit), and the convolution kernel input control module 130 broadcasts the weight elements "W1", "W2", "W3", "W4", "W5", "W6", "W7", "W8" and "W9" of the convolution kernel 320 to the first floating-point operation unit 160_1 (which can be generally referred to as the first floating-point operation unit). In the first cycle of the first multi-channel multiply-add operation 510 (generally referred to as the first multi-channel multiply-add operation), the first floating-point operation unit 160_1 (generally referred to as the first floating-point operation unit) uses the first element "Pad" (generally referred to as the first element) in the matrix element of the receptive field and the first weight "W1" (generally referred to as the first weight) in the weight element to perform half-precision multiplication to generate a dot product result Pad·W1 (generally referred to as the first dot product result). At the same time, the first floating-point operation unit 160_1 (generally referred to as the first floating-point operation unit) uses the second element "Pad" (generally referred to as the second element) in the matrix element of the receptive field and the second weight "W2" (generally referred to as the second weight) in the weight element to perform half-precision multiplication to generate another dot product result Pad·W2 (generally referred to as the second dot product result). The first floating-point operation unit 160_1 (generally referred to as the first floating-point operation unit) uses the dot product result Pad·W1 (generally referred to as the first dot product result) and the dot product result Pad·W2 (generally referred to as the second dot product result) to perform half-precision addition to generate the stage sum value of the first cycle of the first multi-channel multiplication-addition operation 510 (i.e., the stage sum value S1 = Pad·W1 + Pad·W2) (generally referred to as the first stage sum value). The first floating-point operation unit 160_1 (generally referred to as the first floating-point operation unit) performs accumulation calculation using the stage sum value S1 of the first cycle (generally referred to as the first stage sum value) and the initial value (or offset value, such as 0) in the second cycle of the first multi-channel multiplication-addition operation 510 (generally referred to as the first multi-channel multiplication-addition operation) to generate an accumulation result SUM1 (generally referred to as the first accumulation result) of the first multi-channel multiplication-addition operation 510 (generally referred to as the first multi-channel multiplication-addition operation), that is, [(SUM1 = Pad·W1 + Pad·W2) +0].The accumulated result SUM1 of the first multi-channel multiplication-addition operation 510 (generally referred to as the first accumulated result) can be used as the “previous accumulated value” (equivalent to the first accumulated result) used by the next second multi-channel multiplication-addition operation 520 (generally referred to as the second multi-channel multiplication-addition operation).

[0034] The first floating-point operation unit 160_1 (generally referred to as the first floating-point operation unit) performs a second multi-channel multiplication-addition operation 520 (generally referred to as the second multi-channel multiplication-addition operation) based on the control of the channel-by-channel convolution control module 110. The first floating-point operation unit 160_1 (generally referred to as the first floating-point operation unit) performs multiplication-addition calculations using matrix elements of the receptive field (generally referred to as the plurality of second matrix elements of the first receptive field) and a plurality of weight elements of the convolution kernel 320 (generally referred to as the second weight elements) in the first cycle of the second multi-channel multiplication-addition operation 520 (generally referred to as the second multi-channel multiplication-addition operation) to generate a stage sum value S2 (generally referred to as the second stage sum value) of the second multi-channel multiplication-addition operation 520. In detail, the first floating-point operation unit 160_1 (generally referred to as the first floating-point operation unit) uses the matrix element "Pad" of the receptive field and the weight "W3" in the weight element to perform half-precision multiplication in the first cycle of the second multi-channel multiplication-addition operation 520 (if the first multi-channel multiplication-addition operation 510 is generally referred to as the first multi-channel multiplication-addition operation, the second multi-channel multiplication-addition operation 520 can be generally referred to as the second multi-channel multiplication-addition operation) to generate a dot product result Pad·W3. At the same time, the first floating-point operation unit 160_1 uses the matrix element "Pad" of the receptive field and the weight "W4" in the weight element to perform half-precision multiplication to generate another dot product result Pad·W4. The first floating-point operation unit 160_1 uses the dot product result Pad·W3 and the dot product result Pad·W4 to perform half-precision addition to generate the stage sum value (i.e., stage sum value S2 = Pad·W3 + Pad·W4) of the first cycle of the second multi-channel multiplication-addition operation 520 (which can be generally referred to as the second multi-channel multiplication-addition operation) (which can be generally referred to as the second stage sum value). The first floating-point operation unit 160_1 (generally referred to as the first floating-point operation unit) performs accumulation calculation in the second cycle of the second multi-channel multiplication-addition operation 520 (generally referred to as the second multi-channel multiplication-addition operation) using the stage sum value S2 of the first cycle (generally referred to as the second stage sum value) and the previous accumulated value (i.e., the accumulated result SUM1 of the first multi-channel multiplication-addition operation 510) (generally referred to as the first accumulated result) to generate the accumulated result SUM2 (generally referred to as the second accumulated result) of the second multi-channel multiplication-addition operation 520 (generally referred to as the second multi-channel multiplication-addition operation), i.e., [(SUM2 = Pad·W3 + Pad·W4) + (Pad·W1 + Pad·W2)].The accumulated result SUM2 of the second multi-channel multiplication-addition operation 520 can be used as the “previous accumulated value” (if the second multi-channel multiplication-addition operation 520 is generally referred to as the first multi-channel multiplication-addition operation, the third multi-channel multiplication-addition operation 530 can be generally referred to as the second multi-channel multiplication-addition operation) for the next third multi-channel multiplication-addition operation 530 (if the second multi-channel multiplication-addition operation 520 is generally referred to as the first multi-channel multiplication-addition operation, the “previous accumulated value” can be generally referred to as the first accumulated result).

[0035] The first floating-point operation unit 160_1 (generally referred to as the first floating-point operation unit) uses the matrix element "1" of the receptive field and the weight "W5" in the weight element to perform half-precision multiplication in the first cycle of the third multi-channel multiplication-addition operation 530 (if the second multi-channel multiplication-addition operation 520 is generally referred to as the first multi-channel multiplication-addition operation, the third multi-channel multiplication-addition operation 530 can be generally referred to as the second multi-channel multiplication-addition operation) to generate a dot product result 1·W5. At the same time, the first floating-point operation unit 160_1 uses the matrix element "2" of the receptive field and the weight "W6" in the weight element to perform half-precision multiplication to generate another dot product result 2·W6. The first floating-point operation unit 160_1 uses the dot product result 1·W5 and the dot product result 2·W6 to perform half-precision addition to generate the stage sum value (i.e., stage sum value S3 = 1·W5 + 2·W6) of the first cycle of the third multi-channel multiplication-addition operation 530 (which can be generally referred to as the second multi-channel multiplication-addition operation) (which can be generally referred to as the second stage sum value). The first floating-point operation unit 160_1 (generally referred to as the first floating-point operation unit) uses the stage sum value S3 of the first cycle (generally referred to as the second stage sum value) and the previous accumulated value (generally referred to as the first accumulated result) to perform accumulation calculation in the second cycle of the third multi-channel multiplication-addition operation 530 (generally referred to as the second multi-channel multiplication-addition operation) to generate an accumulated result SUM3 (generally referred to as the second accumulated result) of the third multi-channel multiplication-addition operation 530 (generally referred to as the second multi-channel multiplication-addition operation), that is, [(SUM3 = 1·W5 + 2·W6) + (Pad·W1 + Pad·W2) + (Pad·W3 + Pad·W4)].

[0036] The first floating-point operation unit 160_1 (generally referred to as the first floating-point operation unit) uses the matrix element "Pad" of the receptive field and the weight "W7" in the weight element to perform half-precision multiplication in the first cycle of the fourth multi-channel multiplication-addition operation 540 (if the third multi-channel multiplication-addition operation 530 is generally referred to as the first multi-channel multiplication-addition operation, the fourth multi-channel multiplication-addition operation 540 can be generally referred to as the second multi-channel multiplication-addition operation) to generate a dot product result Pad·W7. At the same time, the first floating-point operation unit 160_1 uses the matrix element "17" of the receptive field and the weight "W8" in the weight element to perform half-precision multiplication to generate another dot product result 17·W8. The first floating-point operation unit 160_1 uses the dot product result Pad·W7 and the dot product result 17·W8 to perform half-precision addition to generate the stage sum value (i.e., stage sum value S4 = Pad·W7 + 17·W8) of the first cycle of the fourth multi-channel multiplication-addition operation 540 (which can be generally referred to as the second multi-channel multiplication-addition operation) (which can be generally referred to as the second stage sum value). The first floating-point operation unit 160_1 (which may be generally referred to as the first floating-point operation unit) performs accumulation calculation using the stage sum value S4 of the first cycle (which may be generally referred to as the second stage sum value) and the previous accumulated value (which may be generally referred to as the first accumulated result) in the second cycle of the fourth multi-channel multiplication-addition operation 540 (which may be generally referred to as the second multi-channel multiplication-addition operation) to generate an accumulated result SUM4 (which may be generally referred to as the second accumulated result) of the fourth multi-channel multiplication-addition operation 540 (which may be generally referred to as the second multi-channel multiplication-addition operation), i.e., [(SUM4 = Pad·W7 + 17·W8) + (Pad·W1 + Pad·W2) + (Pad·W3 + Pad·W4) + (1·W5 + 2·W6)].

[0037] The first floating-point operation unit 160_1 (which can be generally referred to as the first floating-point operation unit) uses the matrix element "18" of the receptive field and the weight "W9" in the weight element to perform half-precision multiplication in the first cycle of the fifth multi-channel multiplication-addition operation 550 (if the fourth multi-channel multiplication-addition operation 540 is generally referred to as the first multi-channel multiplication-addition operation, then the fifth multi-channel multiplication-addition operation 550 can be generally referred to as the second multi-channel multiplication-addition operation, or the fifth multi-channel multiplication-addition operation 550 can also be generally referred to as the first multi-channel multiplication-addition operation) to generate a dot product result 18·W9 (that is, the stage sum value S5 = 18·W9 of the fifth multi-channel multiplication-addition operation 550) (corresponding to the fifth multi-channel multiplication-addition operation 550 can be generally referred to as the first multi-channel multiplication-addition operation or the second multi-channel multiplication-addition operation, and the stage sum value S5 can be generally referred to as the first stage sum value or the second stage sum value). The first floating-point operation unit 160_1 (which may be generally referred to as the first floating-point operation unit) performs accumulation calculation using the stage sum value S5 (which may be generally referred to as the first stage sum value or the second stage sum value) and the previous accumulated value (which may also be generally referred to as the first accumulated result) in the second cycle of the fifth multi-channel multiplication-addition operation 550 (which may be generally referred to as the first multi-channel multiplication-addition operation or the second multi-channel multiplication-addition operation) to generate an accumulated result SUM5 (which may be generally referred to as the first accumulated result or the second accumulated result) of the fifth multi-channel multiplication-addition operation 550 (which may be generally referred to as the first multi-channel multiplication-addition operation or the second multi-channel multiplication-addition operation), i.e., [SUM5 = (18·W9) +(Pad·W1 + Pad·W2) + (Pad·W3 + Pad·W4) + (1·W5 + 2·W6) + (Pad·W7 + 17·W8)]. The accumulated result SUM5 (which may be generally referred to as the first accumulated result or the second accumulated result) of the fifth multi-channel multiplication-addition operation 550 (which may be generally referred to as the first multi-channel multiplication-addition operation or the second multi-channel multiplication-addition operation) may be used as a corresponding result element in the channel-by-channel convolution result matrix (if the fifth multi-channel multiplication-addition operation 550 is generally referred to as the first multi-channel multiplication-addition operation, the corresponding result element may be generally referred to as the first result element, and if the fifth multi-channel multiplication-addition operation 550 is generally referred to as the second multi-channel multiplication-addition operation, the corresponding result element may be generally referred to as the second result element). Therefore, the floating-point operation unit base 160_1 calculates a corresponding result element (which may be generally referred to as the first result element or the second result element) in the channel-by-channel convolution result matrix based on the accumulated results of the first multi-channel multiplication-addition operation 510 to the fifth multi-channel multiplication-addition operation 550.

[0038] In one application example, the first cycle of the next multi-channel multiplication-addition operation (e.g., the second multi-channel multiplication-addition operation 520) (generally referred to as the second multi-channel multiplication-addition operation) is located in time after the second cycle of the current multi-channel multiplication-addition operation (e.g., the first multi-channel multiplication-addition operation 510) (generally referred to as the first multi-channel multiplication-addition operation) ends. In another application example, the first cycle of the next multi-channel multiplication-addition operation (e.g., the second multi-channel multiplication-addition operation 520) (generally referred to as the second multi-channel multiplication-addition operation) overlaps in time with the second cycle of the current multi-channel multiplication-addition operation (e.g., the first multi-channel multiplication-addition operation 510) (generally referred to as the first multi-channel multiplication-addition operation). For example, Figure 6 FIG. 4 is a timing diagram of a plurality of multi-channel multiplication and addition operations according to an embodiment of the present invention. Figure 6 The horizontal axis represents time.

[0039] Figure 6 The first multi-channel multi-channel addition operation 510 to the fifth multi-channel multi-channel addition operation 550 can be used as Figure 5 The illustrated embodiment is one of many implementation examples of the first multi-channel multi-channel addition operation 510 to the fifth multi-channel multi-channel addition operation 550 . Figure 6 R1, R2, R3, R4, R5, R6, R7, R8 and R9 respectively represent different registers in register group 140 (for example Figure 4 and Figure 5 The first floating-point arithmetic unit 160_1 uses the dot product result R1·W1 and the dot product result R2·W2 to perform half-precision addition to generate the phase sum value S1 of the first cycle of the first multi-channel multiplication and addition operation 510. The first floating-point arithmetic unit 160_1 uses the phase sum value S1 of the first cycle and the initial value SUM0 to perform accumulation calculation in the second cycle of the first multi-channel multiplication and addition operation 510 to generate the accumulation result SUM1 of the first multi-channel multiplication and addition operation 510. Figure 6 In the illustrated embodiment, the first cycle of the second multi-channel multiplication-addition operation 520 temporally overlaps with the second cycle of the first multi-channel multiplication-addition operation 510, the first cycle of the third multi-channel multiplication-addition operation 530 temporally overlaps with the second cycle of the second multi-channel multiplication-addition operation 520, the first cycle of the fourth multi-channel multiplication-addition operation 540 temporally overlaps with the second cycle of the third multi-channel multiplication-addition operation 530, and the first cycle of the fifth multi-channel multiplication-addition operation 550 temporally overlaps with the second cycle of the fourth multi-channel multiplication-addition operation 540.

[0040] The first floating-point arithmetic unit 160_1 uses the dot product result R3·W3 and the dot product result R4·W4 to perform half-precision addition to generate the phase sum value S2 of the first cycle of the second multi-channel multiplication and addition operation 520. The first floating-point arithmetic unit 160_1 uses the phase sum value S2 of the first cycle and the previous accumulated value (i.e., the accumulated result SUM1 of the first multi-channel multiplication and addition operation 510) to perform accumulation calculation in the second cycle of the second multi-channel multiplication and addition operation 520 to generate the accumulated result SUM2 of the second multi-channel multiplication and addition operation 520. The first floating-point arithmetic unit 160_1 uses the dot product result R5·W5 and the dot product result R6·W6 to perform half-precision addition to generate the phase sum value S3 of the first cycle of the third multi-channel multiplication and addition operation 530. The first floating-point arithmetic unit 160_1 performs accumulation calculation using the phase sum value S3 of the first cycle and the previous accumulated value (i.e., the accumulated result SUM2 of the second multi-channel ... The first floating-point operation unit 160_1 uses the matrix element "R9" provided by the ninth register 9 and the weight "W9" in the weight element to perform half-precision multiplication in the first cycle of the fifth multi-channel multiplication-addition operation 550 to generate a dot product result R9·W9 (i.e., the stage sum value S5 of the fifth multi-channel multiplication-addition operation 550). The first floating-point operation unit 160_1 uses the stage sum value S5 and the previous accumulated value (i.e., the accumulated result SUM4 of the fourth multi-channel multiplication-addition operation 540) to perform accumulation calculation in the second cycle of the fifth multi-channel multiplication-addition operation 550 to generate an accumulated result SUM5 of the fifth multi-channel multiplication-addition operation 550. The accumulated result SUM5 of the fifth multi-channel multiplication-addition operation 550 can be used as a corresponding result element (generally referred to as the first result element or the second result element) in the channel-by-channel convolution result matrix.

[0041] Figure 7 FIG. 4 is a schematic diagram of a circuit module of a floating-point operation unit 700 according to an embodiment of the present invention. Figure 7 The floating point unit 700 shown can refer to Figure 1 The relevant description of any one of the 1st floating-point operation unit 160_1 to the Nth floating-point operation unit 160_N is shown. Figure 7 The floating point unit 700 shown can be used as Figure 1One of many implementation examples of any one of the first floating point operation unit 160_1 to the Nth floating point operation unit 160_N. Figure 7 In the illustrated embodiment, the floating-point operation unit 700 (generally referred to as the first floating-point operation unit) includes a first multiplication circuit 710, a second multiplication circuit 720, an addition circuit 730, and an accumulation circuit 740. A first input terminal of the first multiplication circuit 710 is coupled to the convolution kernel input control module 130 to receive the convolution kernel. A second input terminal of the first multiplication circuit 710 is coupled to the feature map input control module 120 to receive a plurality of elements of a partial matrix of the receptive field (generally referred to as the first partial matrix of the first receptive field). A first input terminal of the second multiplication circuit 720 is coupled to the convolution kernel input control module 130 to receive the convolution kernel. A second input terminal of the second multiplication circuit 720 is coupled to the feature map input control module 120 to receive a plurality of elements of a partial matrix of the receptive field (generally referred to as the first partial matrix of the first receptive field). For example, in the first multi-channel multiplication and addition operation 510, the first multiplication circuit 710 receives the corresponding weight element "W1" of the convolution kernel 320 and the corresponding partial matrix element "Pad" of the feature map matrix 310, and the second multiplication circuit 720 receives the corresponding weight element "W2" of the convolution kernel 320 and the corresponding partial matrix element "Pad" of the feature map matrix 310.

[0042] The first input terminal of the adding circuit 730 is coupled to the output terminal of the first multiplying circuit 710. The second input terminal of the adding circuit 730 is coupled to the output terminal of the second multiplying circuit 720. The input terminal of the accumulating circuit 740 is coupled to the output terminal of the adding circuit 730. For example, in the first cycle of the first multi-channel multiplication-addition operation 510, the first multiplying circuit 710 performs half-precision multiplication to generate a dot product result Pad·W1. At the same time, the second multiplying circuit 720 performs half-precision multiplication to generate another dot product result Pad·W2. The adding circuit 730 performs half-precision addition to generate a phase sum value S1 of the first cycle of the first multi-channel multiplication-addition operation 510. The accumulating circuit 740 performs an accumulation calculation in the second cycle of the first multi-channel multiplication-addition operation 510 using the phase sum value S1 of the first cycle and an initial value (or offset value, such as 0) to generate an accumulation result SUM1 of the first multi-channel multiplication-addition operation 510.

[0043] In the second multi-channel multiplication and addition operation 520, the first multiplication circuit 710 receives the corresponding weight element "W3" of the convolution kernel 320 and the corresponding partial matrix element "Pad" of the feature map matrix 310, and the second multiplication circuit 720 receives the corresponding weight element "W4" of the convolution kernel 320 and the corresponding partial matrix element "Pad" of the feature map matrix 310. In the first cycle of the second multi-channel multiplication and addition operation 520, the first multiplication circuit 710 performs half-precision multiplication to generate a dot product result Pad·W3, and the second multiplication circuit 720 performs half-precision multiplication to generate another dot product result Pad·W4. The addition circuit 730 performs half-precision addition to generate the stage sum value S2 of the first cycle of the second multi-channel multiplication and addition operation 520. The accumulation circuit 740 performs accumulation calculation using the phase sum value S2 of the first cycle and the accumulation result SUM1 of the first multi-channel ...

[0044] Figure 8 It is a circuit module diagram of a first multiplication circuit 710 , a second multiplication circuit 720 , an addition circuit 730 and an accumulation circuit 740 according to an embodiment of the present invention. Figure 8 The first multiplication circuit 710, the second multiplication circuit 720, the addition circuit 730 and the accumulation circuit 740 shown in FIG. Figure 7 Related instructions. Figure 8 The first multiplication circuit 710, the second multiplication circuit 720, the addition circuit 730 and the accumulation circuit 740 can be used as Figure 7 The first multiplication circuit 710, the second multiplication circuit 720, the addition circuit 730 and the accumulation circuit 740 are one of many implementation examples.

[0045] exist Figure 8In the illustrated embodiment, the first multiplication circuit 710 includes a first multiplexer 711, a second multiplexer 712, and a first half-precision multiplier 713. In this embodiment, the first half-precision multiplier 713 can be a component of a standard floating-point arithmetic unit, and the first multiplexer 711 and the second multiplexer 712 are added to the standard floating-point arithmetic unit to implement hardware specifically for channel-by-channel convolution operations. The first input end of the first multiplexer 711 is coupled to the feature map input control module 120 to receive multiple elements of the partial matrix of the receptive field. The first input end of the second multiplexer 712 is coupled to the convolution kernel input control module 130 to receive the convolution kernel. The first input end of the first half-precision multiplier 713 is coupled to the output end of the first multiplexer 711. The second input end of the first half-precision multiplier 713 is coupled to the output end of the second multiplexer 712. The output end of the first half-precision multiplier 713 is coupled to the first input end of the addition circuit 730.

[0046] In the case where the first multiplication circuit 710 is used as hardware dedicated to channel-by-channel convolution operations, the first multiplexer 711 couples the feature map input control module 120 to the first input terminal of the first half-precision multiplier 713, and the second multiplexer 712 couples the convolution kernel input control module 130 to the second input terminal of the first half-precision multiplier 713. In the case where the first multiplication circuit 710 is used as a multiplication circuit of a standard floating-point arithmetic unit, and in the case where this standard floating-point arithmetic unit performs floating-point multiplication "a1·b1", the second input terminal of the first multiplexer 711 receives the floating-point number a1, and the second input terminal of the second multiplexer 712 receives the floating-point number b1. At this time, the first multiplexer 711 transmits the floating-point number a1 to the first input terminal of the first half-precision multiplier 713, and the second multiplexer 712 transmits the floating-point number b1 to the second input terminal of the first half-precision multiplier 713.

[0047] exist Figure 8In the illustrated embodiment, the second multiplication circuit 720 includes a third multiplexer 721, a fourth multiplexer 722, and a second half-precision multiplier 723. In this embodiment, the second half-precision multiplier 723 can be a component of a standard floating-point arithmetic unit, and the third multiplexer 721 and the fourth multiplexer 722 are added to the standard floating-point arithmetic unit to implement hardware specifically for channel-by-channel convolution operations. The first input end of the third multiplexer 721 is coupled to the feature map input control module 120 to receive multiple elements of the partial matrix of the receptive field. The first input end of the fourth multiplexer 722 is coupled to the convolution kernel input control module 130 to receive the convolution kernel. The first input end of the second half-precision multiplier 723 is coupled to the output end of the third multiplexer 721. The second input end of the second half-precision multiplier 723 is coupled to the output end of the fourth multiplexer 722. An output terminal of the second half-precision multiplier 723 is coupled to a second input terminal of the adding circuit 730 .

[0048] In the case where the second multiplication circuit 720 is used as hardware dedicated to channel-by-channel convolution operation, the third multiplexer 721 couples the feature map input control module 120 to the first input terminal of the second half-precision multiplier 723, and the fourth multiplexer 722 couples the convolution kernel input control module 130 to the second input terminal of the second half-precision multiplier 723. In the case where the second multiplication circuit 720 is used as a multiplication circuit of a standard floating-point operation unit, and in the case where this standard floating-point operation unit performs floating-point multiplication "a2·b2", the second input terminal of the third multiplexer 721 receives the floating-point number a2, and the second input terminal of the fourth multiplexer 722 receives the floating-point number b2. At this time, the third multiplexer 721 transmits the floating-point number a2 to the first input terminal of the second half-precision multiplier 723, and the fourth multiplexer 722 transmits the floating-point number b2 to the second input terminal of the second half-precision multiplier 723.

[0049] exist Figure 8 In the illustrated embodiment, the adding circuit 730 includes a half-precision fast adder 731. A first input terminal of the half-precision fast adder 731 is coupled to the output terminal of the first multiplication circuit 710. A second input terminal of the half-precision fast adder 731 is coupled to the output terminal of the second multiplication circuit 720. The half-precision fast adder 731 performs fast addition on the half-precision outputs of the two channels and outputs a fast addition result (a half-precision floating point number, i.e., a phase sum value of the first cycle). The output terminal of the half-precision fast adder 731 is coupled to the input terminal of the accumulation circuit 740 to provide a phase sum value.

[0050] exist Figure 8In the illustrated embodiment, the accumulation circuit 740 includes a format conversion module 746, a fifth multiplexer 741, a sixth multiplexer 742, a full-precision adder 743, a floating-point number normalization module 744, and a loop accumulation result access module 745. In this embodiment, the full-precision adder 743 and the floating-point number normalization module 744 can be components of a standard floating-point arithmetic unit, and the format conversion module 746, the fifth multiplexer 741, the sixth multiplexer 742, and the loop accumulation result access module 745 are added to the standard floating-point arithmetic unit to implement hardware specifically for channel-by-channel convolution operations. The input end of the format conversion module 746 is coupled to the output end of the addition circuit 730. The format conversion module 746 converts the half-precision floating point number into a full-precision floating point number to the fifth multiplexer 741. The first input end of the fifth multiplexer 741 is coupled to the output end of the addition circuit 730. A first input terminal of the full-precision adder 743 is coupled to the output terminal of the fifth multiplexer 741 . A second input terminal of the full-precision adder 743 is coupled to the output terminal of the sixth multiplexer 742 .

[0051] In the case where the accumulation circuit 740 is used as hardware dedicated to channel-by-channel convolution operation, the fifth multiplexer 741 couples the output of the addition circuit 730 to the first input of the full-precision adder 743, and the sixth multiplexer 742 couples the output of the loop accumulation result access module 745 to the second input of the full-precision adder 743. Alternatively, the sixth multiplexer 742 transmits the externally stored accumulation value (or bias value) to the second input of the full-precision adder 743. In the case where the accumulation circuit 740 is used as an accumulation circuit of a standard floating-point arithmetic unit, and in the case where this standard floating-point arithmetic unit performs addition "A + B", the second input of the fifth multiplexer 741 receives the floating-point number A, and the second input of the sixth multiplexer 742 receives the floating-point number B. At this time, the fifth multiplexer 741 transmits the floating point number A to the first input terminal of the full-precision adder 743 , and the sixth multiplexer 742 transmits the floating point number B to the second input terminal of the full-precision adder 743 .

[0052] The input end of the floating point number standardization module 744 is coupled to the output end of the full-precision adder 743. The input end of the loop accumulation result access module 745 is coupled to the output end of the floating point number standardization module 744. The floating point number standardization module 744 standardizes the output value of the full-precision adder 743 based on the floating point number standard format to generate a standardized sum value for the loop accumulation result access module 745. The output end of the loop accumulation result access module 745 is coupled to the first input end of the sixth multiplexer 742. The loop accumulation result access module 745 saves the result of the current loop (current multi-channel multiplication and addition operation) according to the current loop number information sent by the multi-channel multiplication and addition instruction (DP2A) decoding module 112, and sends it to the next loop accumulation.

[0053] For example, in the first multi-channel multiplication-addition operation 510, the loop accumulation result memory access module 745 provides an initial value (or a bias value, such as 0) to the sixth multiplexer 742. In the second multi-channel multiplication-addition operation 520, the loop accumulation result memory access module 745 provides the accumulation result SUM1 of the first multi-channel multiplication-addition operation 510 to the sixth multiplexer 742. In the third multi-channel multiplication-addition operation 530, the loop accumulation result memory access module 745 provides the accumulation result SUM2 of the second multi-channel multiplication-addition operation 520 to the sixth multiplexer 742. In the fourth multi-channel multiplication-addition operation 540, the loop accumulation result memory access module 745 provides the accumulation result SUM3 of the third multi-channel multiplication-addition operation 530 to the sixth multiplexer 742. In the fifth multi-channel multiply-add operation 550 , the loop accumulation result memory access module 745 provides the accumulation result SUM4 of the fourth multi-channel multiply-add operation 540 to the sixth multiplexer 742 .

[0054] According to application requirements, the feature map and weight data format usually uses half precision (16 bits), so each double word (DWORD, i.e. 32 bits) can store data from two different channels. The final result usually requires accuracy, so the accumulation circuit 740 outputs a full-precision format floating point number to improve the accuracy of feature extraction.

[0055] Fig. 9 It is a circuit module diagram of a channel-by-channel convolution control module 110, a feature map input control module 120, and a convolution kernel input control module 130 according to an embodiment of the present invention. Fig. 9 The channel-by-channel convolution control module 110, the feature map input control module 120, and the convolution kernel input control module 130 can refer to Figure 1 Related instructions. Fig. 9 The channel-by-channel convolution control module 110, the feature map input control module 120, and the convolution kernel input control module 130 can be used as Figure 1The channel-by-channel convolution control module 110, the feature map input control module 120 and the convolution kernel input control module 130 are one of many implementation examples.

[0056] exist Fig. 9 In the illustrated embodiment, the channel-by-channel convolution control module 110 includes an instruction decoding module 111 and a multi-channel multiplication-addition instruction (DP2A) decoding module 112. The instruction decoding module 111 is used to decode the instruction issued by the instruction scheduling module. The instruction decoding module 111 executes the channel-by-channel convolution instruction and issues the multi-channel multiplication-addition instruction DP2A to the multi-channel multiplication-addition instruction decoding module 112 for further decoding of the multi-channel multiplication-addition instruction DP2A.

[0057] The multi-channel multiplication-addition instruction decoding module 112 is coupled to the instruction decoding module 111, the convolution kernel input control module 130 and the feature map input control module 120. The multi-channel multiplication-addition instruction decoding module 112 executes the multi-channel multiplication-addition instruction DP2A to control the convolution kernel input control module 130, the feature map input control module 120 and the first floating point operation unit 160_1 to the Nth floating point operation unit 160_N to perform a multi-channel multiplication-addition operation (for example, any one of the first multi-channel multiplication-addition operation 510 to the fifth multi-channel multiplication-addition operation 550). The multi-channel multiplication-addition instruction decoding module 112 broadcasts the floating point operation control information to each thread according to the instruction to guide the calculation process until the final result is output. The convolution kernel input control module 130 calls the convolution kernel to the first floating point operation unit 160_1 to the Nth floating point operation unit 160_N based on the control of the multi-channel multiplication-addition instruction decoding module 112. The feature map input control module 120 calls a partial matrix in the receptive field from the feature map matrix of the register group 140 to the first floating point operation unit 160_1 to the Nth floating point operation unit 160_N based on the control of the multi-channel multiplication and addition instruction decoding module 112 .

[0058] exist Fig. 9 In the illustrated embodiment, the feature map input control module 120 includes an operand acquisition module 121 and a feature map buffer module 122. The operand acquisition module 121 is coupled to the multi-channel multiplication-addition instruction decoding module 112. The multi-channel multiplication-addition instruction decoding module 112 acquires the address of the corresponding feature map and convolution kernel and passes it to the operand acquisition module 121. The register group 140 is coupled to the operand acquisition module 121 to provide a feature map matrix.

[0059] exist Fig. 9In the illustrated embodiment, the convolution kernel input control module 130 includes a convolution kernel acquisition module 131 and a convolution kernel buffer module 132. The convolution kernel acquisition module 131 is coupled to the multi-channel multiplication-addition instruction decoding module 112. The multi-channel multiplication-addition instruction decoding module 112 obtains the address of the corresponding feature map and the convolution kernel and passes them to the convolution kernel acquisition module 131. The constant cache 150 is coupled to the convolution kernel acquisition module 131 to provide the convolution kernel.

[0060] The feature map buffer module 122 and the convolution kernel buffer module 132 are controlled by the multi-channel multiplication and addition instruction decoding module 112, and one of the channels in the dual-channel data storage format is allocated to the 1st floating-point operation unit 160_1 to the Nth floating-point operation unit 160_N of each thread for calculation, and the content of an unused channel is saved for the operation of the next instruction cycle. Therefore, even if each multi-channel multiplication and addition instruction DP2A requires two instruction cycles, this embodiment only accesses the data containing the dual channels in the first instruction cycle. For a one-dimensional data layout, the data of the feature map is independent for each thread, while the weight data of the convolution kernel is shared by all threads.

[0061] In summary, this embodiment utilizes the hardware resources of a standard vector computing unit, adds a small amount of control and computing logic, and implements a hardware specifically used to process one-dimensional data layout and accelerate its channel-by-channel convolution application. In order to call this hardware, this embodiment adds a multi-channel multiplication-addition instruction DP2A that supports mixed-precision floating-point multiplication-addition operations. Taking a 3×3 convolution kernel as an example, the conventional method requires 9 half-precision multiplications and 18 mixed-precision additions, a total of 27 instruction cycles; this embodiment only requires 5 multi-channel multiplication-addition operations (a single multi-channel multiplication-addition operation requires 2 instruction cycles). Therefore, this embodiment can complete channel-by-channel convolution operations with fewer instructions, and at a faster speed. Different parameter configurations of channel-by-channel convolution, including feature map size, receptive field size, zero padding width, step size, neural network operation direction (forward or reverse), and various data formats (including fp16, bf16, etc.), can all be simply mapped to the multi-channel multiplication and addition instruction DP2A under a one-dimensional data layout. Therefore, this embodiment can adapt to different convolutional neural network (CNN) application requirements and is very flexible. For dual-channel operations, this embodiment automatically caches temporarily unused half-precision data in dual-channel registers within the instruction, and automatically transfers full-precision accumulation results between instructions. Therefore, this embodiment can perform fewer memory accesses, thereby better releasing computing power.

[0062] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A channel-by-channel convolution device, characterized in that: The channel-by-channel convolution device comprises: A channel-by-channel convolution control module, used for executing the channel-by-channel convolution instructions issued by the instruction scheduling module; A convolution kernel input control module, coupled to the channel-by-channel convolution control module; a feature map input control module, coupled to the channel-by-channel convolution control module; and A plurality of floating point operation units are coupled to the convolution kernel input control module and the feature map input control module, wherein: The convolution kernel input control module broadcasts the convolution kernel to the plurality of floating point operation units based on the control of the channel-by-channel convolution control module; The feature map input control module calls a first part of the matrix in the first receptive field from the feature map matrix to a first floating point operation unit among the plurality of floating point operation units based on the control of the channel-by-channel convolution control module; The first floating-point operation unit performs a first multi-channel multiplication-addition operation based on the control of the channel-by-channel convolution control module; The first floating-point operation unit performs multiplication and addition calculations using a plurality of first matrix elements of the first receptive field and a plurality of first weight elements of the convolution kernel in a first cycle of the first multi-channel multiplication and addition operation to generate a first stage sum value; The first floating point operation unit performs an accumulation calculation using the first stage sum value and a previous accumulation value in a second cycle of the first multi-channel multiplication-addition operation to generate a first accumulation result of the first multi-channel multiplication-addition operation; and The first floating-point operation unit calculates a first result element in a channel-by-channel convolution result matrix based on the first accumulated result of the first multi-channel multiplication-addition operation.

2. The channel-by-channel convolution device according to claim 1, characterized in that: The channel-by-channel convolution device also includes: A constant cache is used to store the convolution kernel, wherein the constant cache is coupled to the convolution kernel input control module to provide the convolution kernel.

3. The channel-by-channel convolution device according to claim 1, characterized in that: The channel-by-channel convolution device also includes: A register group is used to store the feature map matrix, wherein the register group is coupled to the feature map input control module to provide the feature map matrix.

4. The channel-by-channel convolution device according to claim 3, characterized in that: The first partial matrix of the feature map matrix in the first receptive field is stored in a first address of each of the multiple registers of the register group, and the second partial matrix of the feature map matrix in the second receptive field is stored in a second address of each of the multiple registers of the register group. The feature map input control module calls the first partial matrix from the first address of each of the multiple registers to the first floating-point operation unit based on the control of the channel-by-channel convolution control module, and the feature map input control module calls the second partial matrix from the second address of each of the multiple registers to the second floating-point operation unit among the multiple floating-point operation units based on the control of the channel-by-channel convolution control module.

5. The channel-by-channel convolution device according to claim 1, characterized in that: The multiplication and addition calculations performed in the first cycle of the first multi-channel multiplication and addition operation include half-precision multiplication and half-precision addition, the first stage sum is a half-precision floating point number, the accumulation calculations performed in the second cycle of the first multi-channel multiplication and addition operation include full-precision addition, and the first accumulation result is a full-precision floating point number.

6. The channel-by-channel convolution device according to claim 5, characterized in that: The multiple first matrix elements of the first receptive field include a first element and a second element, and the multiple first weight elements of the convolution kernel include a first weight and a second weight, In the first cycle of the first multi-channel multiply-add operation, the first floating-point operation unit uses the first element and the first weight to perform the half-precision multiplication to generate a first dot product result, the first floating-point operation unit uses the second element and the second weight to perform the half-precision multiplication to generate a second dot product result, and the first floating-point operation unit uses the first dot product result and the second dot product result to perform the half-precision addition to generate the first stage sum value.

7. The channel-by-channel convolution device according to claim 1, characterized in that: The first floating-point operation unit performs a second multi-channel multiplication-addition operation based on the control of the channel-by-channel convolution control module; The first floating-point operation unit performs multiplication and addition calculations using a plurality of second matrix elements of the first receptive field and a plurality of second weight elements of the convolution kernel in a first cycle of the second multi-channel multiplication and addition operation to generate a second stage sum value; The first floating-point operation unit performs accumulation calculation using the second stage sum and the first accumulation result in a second cycle of the second multi-channel multiplication-addition operation to generate a second accumulation result of the second multi-channel multiplication-addition operation; as well as The first floating-point operation unit calculates a second result element in a channel-by-channel convolution result matrix based on the second accumulated result of the second multi-channel multiplication-addition operation.

8. The channel-by-channel convolution device according to claim 7, characterized in that: The first period of the second multi-channel multiply-add operation temporally overlaps with the second period of the first multi-channel multiply-add operation.

9. The channel-by-channel convolution device according to claim 7, characterized in that: The first cycle of the second multi-channel multiply-add operation is located after the second cycle of the first multi-channel multiply-add operation ends in time.

10. The channel-by-channel convolution device according to claim 1, characterized in that: The first floating-point operation unit comprises: a first multiplication circuit, wherein a first input terminal of the first multiplication circuit is coupled to the convolution kernel input control module to receive the convolution kernel, and a second input terminal of the first multiplication circuit is coupled to the feature map input control module to receive a plurality of elements of the first partial matrix of the first receptive field; a second multiplication circuit, wherein a first input terminal of the second multiplication circuit is coupled to the convolution kernel input control module to receive the convolution kernel, and a second input terminal of the second multiplication circuit is coupled to the feature map input control module to receive a plurality of elements of the first partial matrix of the first receptive field; an adding circuit, wherein a first input terminal of the adding circuit is coupled to an output terminal of the first multiplying circuit, and a second input terminal of the adding circuit is coupled to an output terminal of the second multiplying circuit; and An accumulation circuit, wherein an input terminal of the accumulation circuit is coupled to an output terminal of the adding circuit.

11. The channel-by-channel convolution device according to claim 10, characterized in that: The first multiplication circuit comprises: a first multiplexer, wherein a first input terminal of the first multiplexer is coupled to the convolution kernel input control module to receive the convolution kernel; a second multiplexer, wherein a first input terminal of the second multiplexer is coupled to the feature map input control module to receive a plurality of elements of the first partial matrix of the first receptive field; and A first half-precision multiplier, wherein a first input terminal of the first half-precision multiplier is coupled to an output terminal of the first multiplexer, a second input terminal of the first half-precision multiplier is coupled to an output terminal of the second multiplexer, and an output terminal of the first half-precision multiplier is coupled to a first input terminal of the adding circuit.

12. The channel-by-channel convolution device according to claim 10, characterized in that: The second multiplication circuit comprises: a third multiplexer, wherein a first input terminal of the third multiplexer is coupled to the convolution kernel input control module; a fourth multiplexer, wherein a first input terminal of the fourth multiplexer is coupled to the feature map input control module; and A second half-precision multiplier, wherein a first input terminal of the second half-precision multiplier is coupled to the output terminal of the third multiplexer, a second input terminal of the second half-precision multiplier is coupled to the output terminal of the fourth multiplexer, and an output terminal of the second half-precision multiplier is coupled to the second input terminal of the adding circuit.

13. The channel-by-channel convolution device according to claim 10, characterized in that: The adding circuit comprises: A half-precision fast adder, wherein a first input terminal of the half-precision fast adder is coupled to an output terminal of the first multiplication circuit, a second input terminal of the half-precision fast adder is coupled to an output terminal of the second multiplication circuit, and an output terminal of the half-precision fast adder is coupled to an input terminal of the accumulation circuit.

14. The channel-by-channel convolution device according to claim 10, characterized in that: The accumulating circuit comprises: a format conversion module, wherein an input terminal of the format conversion module is coupled to an output terminal of the adding circuit, and the format conversion module converts the half-precision floating point number into a full-precision floating point number; a fifth multiplexer, wherein a first input terminal of the fifth multiplexer is coupled to an output terminal of the format conversion module; a sixth multiplexer; a full-precision adder, wherein a first input terminal of the full-precision adder is coupled to an output terminal of the fifth multiplexer, and a second input terminal of the full-precision adder is coupled to an output terminal of the sixth multiplexer; a floating point number normalization module, wherein an input terminal of the floating point number normalization module is coupled to an output terminal of the full precision adder; and A loop accumulation result memory access module, wherein the input end of the loop accumulation result memory access module is coupled to the output end of the floating point number standardization module, and the output end of the loop accumulation result memory access module is coupled to the first input end of the sixth multiplexer.

15. The channel-by-channel convolution device according to claim 1, characterized in that: The channel-by-channel convolution control module comprises: an instruction decoding module, configured to decode the instruction emitted by the instruction scheduling module, wherein the instruction decoding module executes the channel-by-channel convolution instruction and emits a multi-channel multiplication-addition instruction; and A multi-channel multiply-add instruction decoding module is coupled to the instruction decoding module, the convolution kernel input control module and the feature map input control module, wherein the multi-channel multiply-add instruction decoding module executes the multi-channel multiply-add instruction to control the convolution kernel input control module and the feature map input control module, the convolution kernel input control module calls the convolution kernel to the multiple floating-point operation units based on the control of the multi-channel multiply-add instruction decoding module, and the feature map input control module calls the first part of the matrix in the first receptive field from the feature map matrix to the first floating-point operation unit based on the control of the multi-channel multiply-add instruction decoding module.

16. A method for operating a channel-by-channel convolution device, characterized in that: The operation method comprises: The channel-by-channel convolution control module of the channel-by-channel convolution device executes the channel-by-channel convolution instruction issued by the instruction scheduling module, wherein the convolution kernel input control module of the channel-by-channel convolution device is coupled to the channel-by-channel convolution control module, the feature map input control module of the channel-by-channel convolution device is coupled to the channel-by-channel convolution control module, and the plurality of floating-point operation units of the channel-by-channel convolution device are coupled to the convolution kernel input control module and the feature map input control module; The convolution kernel input control module broadcasts the convolution kernel to the multiple floating-point operation units based on the control of the channel-by-channel convolution control module; The feature map input control module calls a first partial matrix in a first receptive field from a feature map matrix to a first floating point operation unit among the plurality of floating point operation units based on the control of the channel-by-channel convolution control module; The first floating-point operation unit performs a first multi-channel multiplication-addition operation based on the control of the channel-by-channel convolution control module; The first floating-point operation unit performs multiplication and addition calculations using a plurality of first matrix elements of the first receptive field and a plurality of first weight elements of the convolution kernel in a first cycle of the first multi-channel multiplication and addition operation to generate a first stage sum value; performing, by the first floating-point arithmetic unit, an accumulation calculation using the first stage sum value and a previous accumulation value in a second cycle of the first multi-channel multiplication-addition operation to generate a first accumulation result of the first multi-channel multiplication-addition operation; and The first floating-point operation unit calculates a first result element in a channel-by-channel convolution result matrix based on the first accumulated result of the first multi-channel multiplication-addition operation.

17. The operating method according to claim 16, characterized in that: The operation method further includes: storing the first partial matrix of the feature map matrix in the first receptive field at a first address of each of a plurality of registers of a register group; storing a second partial matrix of the feature map matrix in a second receptive field at a second address of each of the plurality of registers of the register group; The feature map input control module calls the first partial matrix from the first address of each of the plurality of registers to the first floating point operation unit based on the control of the channel-by-channel convolution control module; and The feature map input control module calls the second partial matrix from the second address of each of the plurality of registers to a second floating-point operation unit among the plurality of floating-point operation units based on the control of the channel-by-channel convolution control module.

18. The operating method according to claim 16, characterized in that: The multiplication and addition calculations performed in the first cycle of the first multi-channel multiplication and addition operation include half-precision multiplication and half-precision addition, the first stage sum is a half-precision floating point number, the accumulation calculations performed in the second cycle of the first multi-channel multiplication and addition operation include full-precision addition, and the first accumulation result is a full-precision floating point number.

19. The operating method according to claim 18, characterized in that: The plurality of first matrix elements of the first receptive field include a first element and a second element, the plurality of first weight elements of the convolution kernel include a first weight and a second weight, and the operation method further includes: In the first cycle of the first multi-channel multiply-add operation, the first floating-point operation unit uses the first element and the first weight to perform the half-precision multiplication to generate a first dot product result, the first floating-point operation unit uses the second element and the second weight to perform the half-precision multiplication to generate a second dot product result, and the first floating-point operation unit uses the first dot product result and the second dot product result to perform the half-precision addition to generate the first stage sum value.

20. The operating method according to claim 16, characterized in that: The operation method further includes: The first floating-point operation unit performs a second multi-channel multiplication-addition operation based on the control of the channel-by-channel convolution control module; The first floating-point operation unit performs multiplication and addition calculations using a plurality of second matrix elements of the first receptive field and a plurality of second weight elements of the convolution kernel in a first cycle of the second multi-channel multiplication and addition operation to generate a second stage sum value; performing accumulation calculation by the first floating-point operation unit in a second cycle of the second multi-channel multiplication-addition operation using the second stage sum and the first accumulation result to generate a second accumulation result of the second multi-channel multiplication-addition operation; and The first floating-point operation unit calculates a second result element in a channel-by-channel convolution result matrix based on the second accumulated result of the second multi-channel multiplication-addition operation.

21. The operating method according to claim 20, characterized in that: The first period of the second multi-channel multiply-add operation temporally overlaps with the second period of the first multi-channel multiply-add operation.

22. The operating method according to claim 20, characterized in that: The first cycle of the second multi-channel multiply-add operation is located after the second cycle of the first multi-channel multiply-add operation ends in time.

Citation Information

Patent Citations

  • Floating point data inverse quantization and quantization method and equipment

    CN111240746A

  • Implementation method of channel-by-channel convolution

    CN116957018A

  • Low precision convolution operations

    US20190354568A1

  • Vector packed matrix multiplication and accumulation processors, methods, systems, and instructions

    US20250004768A1