Per-channel Convolution Device and Its Operation Method
By designing a channel-by-channel convolution device, the channel-by-channel convolution operation in convolution can be optimized to separate the depth, solve the problem of large amount of calculation and parameter, improve the model's inference speed and operation efficiency, and is suitable for convolutional neural network applications of various data types and convolution kernel sizes.
Patent Information
- Application Number
- CN202510489160.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The prior art is difficult to efficiently implement channel-by-channel convolution operations in deep separation convolution, resulting in large amounts of calculations and parameters, affecting the model inference speed and operation efficiency.
A channel-by-channel convolution device is designed, including a channel-by-channel convolution control module, a convolution kernel input control module, a feature map input control module and multiple floating-point operation units. The control module coordinates the operation unit to perform multi-channel multiplication and addition operations to optimize the convolution calculation process.
By optimizing the channel-by-channel convolution operation, the calculation amount and parameter amount are reduced, the inference speed and operation efficiency of the model are improved, and it is suitable for a variety of data types and convolution kernel sizes, and supports forward and reverse convolution.
Smart Images

Figure CN120010922B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence (AI) chips, and more particularly to an electronic device for performing depthwise separable convolution, and particularly to a depthwise convolution device for performing depthwise convolution in depthwise separable convolution and an operation method thereof. Background Art
[0002] Computing devices such as artificial intelligence chips can provide huge computing power. The huge computing power of artificial intelligence chips stems from a large number of internal hardware cores. An AI chip usually includes multiple programmable processors, such as a stream processor cluster (SPC). Each programmable processor usually includes multiple compute units (CUs, or computing cores), and each compute unit usually includes multiple execution units (EUs, or execution cores), such as at least one of an integer (INT) core, a floating point (FP) core, a tensor core (Tcore), and / or a vector core (Vcore). By programming to organize various types of compute units, the programmable multi-processor can support general computing, scientific computing, and neural network computing.
[0003] Depthwise separable convolution is a lightweight convolution operation. The operation process of depthwise separable convolution includes two stages: depthwise convolution and pointwise convolution. Depthwise separable convolution can effectively reduce the amount of computation and the number of parameters, and improve the inference speed and running efficiency of the model. Assume that the size of the input feature map is H×W, the number of input channels is N, the size of the convolution kernel is K×K, and the number of output channels is M. Then the amount of computation of standard convolution is H×W×N×K 2 ×M, and the number of parameters is N×K 2 ×M. Under the same assumption conditions, for depthwise separable convolution, the number of parameters in the first stage (depthwise convolution) is K 2 ×N, and the number of parameters in the second stage (pointwise convolution) is N×M (assuming a convolution kernel size of 1×1). Therefore, the total number of parameters of depthwise separable convolution is K 2 ×N + N×M = N×(K 2 + M). Generally, N×(K 2 + M) is smaller than N×K2 ×M is much smaller.
[0004] Like standard convolutions, depthwise separable convolutions can be computed using General Matrix Multiplication (GEMM) hardware, or by Central Processing Unit (CPU) hardware, or by the Vector Unit hardware of a Graphics Processing Unit (GPU). How to implement hardware specifically for depthwise separable convolutions is one of many technical issues in this field. SUMMARY OF THE INVENTION
[0005] The present invention is directed to a per-channel convolution device and an operation method thereof for performing the first stage, "per-channel convolution", in a depthwise separable convolution operation.
[0006] In an embodiment according to the present invention, the per-channel convolution device includes a per-channel convolution control module, a convolution kernel input control module, a feature map input control module, and a plurality of floating-point arithmetic units. The per-channel convolution control module is configured to execute per-channel convolution instructions issued by an instruction scheduling module. The convolution kernel input control module is coupled to the per-channel convolution control module. The feature map input control module is coupled to the per-channel convolution control module. The plurality of floating-point arithmetic units are coupled to the convolution kernel input control module and the feature map input control module. The convolution kernel input control module broadcasts the convolution kernel to the plurality of floating-point arithmetic units based on the control of the per-channel convolution control module. The feature map input control module calls a first partial matrix in a first receptive field from a feature map matrix to a first floating-point arithmetic unit among the plurality of floating-point arithmetic units based on the control of the per-channel convolution control module. The first floating-point arithmetic unit performs a first multi-channel multiply-accumulate operation based on the control of the per-channel convolution control module. The first floating-point arithmetic unit uses a plurality of first matrix elements of the first receptive field and a plurality of first weight elements of the convolution kernel to perform multiply-accumulate calculations in a first cycle of the first multi-channel multiply-accumulate operation to generate a first-stage sum value. The first floating-point arithmetic unit uses the first-stage sum value and a previous accumulated value to perform an accumulation calculation in a second cycle of the first multi-channel multiply-accumulate operation to generate a first accumulated result of the first multi-channel multiply-accumulate operation. The first floating-point arithmetic unit calculates a first result element in a per-channel convolution result matrix based on the first accumulated result of the first multi-channel multiply-accumulate operation.
[0007] In an embodiment according to the present invention, the per-channel convolution device includes: executing, by a per-channel convolution control module, the per-channel convolution instructions emitted by an instruction scheduling module; broadcasting, by a convolution kernel input control module based on the control of the per-channel convolution control module, the convolution kernel to the plurality of floating-point operation units; invoking, by a feature map input control module based on the control of the per-channel convolution control module, a first partial matrix in a first receptive field from a feature map matrix to a first floating-point operation unit among the plurality of floating-point operation units; performing, by the first floating-point operation unit based on the control of the per-channel convolution control module, a first multi-channel multiply-accumulate operation; using, by the first floating-point operation unit in a first cycle of the first multi-channel multiply-accumulate operation, a plurality of first matrix elements of the first receptive field and a plurality of first weight elements of the convolution kernel to perform multiply-accumulate calculation to generate a first-stage sum value; using, by the first floating-point operation unit in a second cycle of the first multi-channel multiply-accumulate operation, the first-stage sum value and a previous accumulated value to perform accumulation calculation to generate a first accumulated result of the first multi-channel multiply-accumulate operation; and calculating, by the first floating-point operation unit based on the first accumulated result of the first multi-channel multiply-accumulate operation, a first result element in a per-channel convolution result matrix.
[0008] Based on the above, the per-channel convolution control module of the per-channel convolution device controls the convolution kernel input control module, the feature map input control module, and the plurality of floating-point operation units to perform one or more multi-channel multiply-accumulate operations on a partial matrix in the receptive field. After completing the one or more multi-channel multiply-accumulate operations, the floating-point operation unit can calculate a result element in the per-channel convolution result matrix. Description of the Drawings
[0009] Figure 1 is a schematic diagram of a circuit block of a per-channel convolution device according to an embodiment of the present invention;
[0010] Figure 2 is a schematic flowchart of an operation method of a per-channel convolution device according to an embodiment of the present invention;
[0011] Figure 3 is a schematic diagram of per-channel convolution illustrated according to an embodiment of the present invention;
[0012] Figure 4 is a schematic diagram of a data layout in which a feature map matrix is stored in a register bank in a one-dimensional data layout manner according to an embodiment of the present invention;
[0013] Figure 5 is a schematic diagram of a plurality of multi-channel multiply-accumulate operations performed by a floating-point operation unit according to an embodiment of the present invention;
[0014] Figure 6It is a timing schematic diagram of multiple multi-channel multiply-accumulate operations illustrated according to an embodiment of the present invention;
[0015] Figure 7 It is a schematic diagram of circuit modules of a floating-point arithmetic unit illustrated according to an embodiment of the present invention;
[0016] Figure 8 It is a schematic diagram of circuit modules of a first multiplication circuit, a second multiplication circuit, an addition circuit, and an accumulation circuit illustrated according to an embodiment of the present invention;
[0017] Figure 9 It is a schematic diagram of circuit modules of a per-channel convolution control module, a feature map input control module, and a convolution kernel input control module illustrated according to an embodiment of the present invention. Detailed Embodiments
[0018] Reference will now be made in detail to the exemplary embodiments of the present invention, examples of which are illustrated in the accompanying drawings. Wherever possible, the same reference numerals are used in the drawings and the description to refer to the same or like parts.
[0019] As used throughout the specification of this application (including the claims), the term "coupled (or connected)" can refer to any direct or indirect connection means. For example, if it is described in the text that a first device is coupled (or connected) to a second device, it should be interpreted that the first device can be directly connected to the second device, or the first device can be indirectly connected to the second device through other devices or some connection means. The terms "first", "second", etc. mentioned throughout the specification of this application (including the claims) are used to name components (elements), or to distinguish different embodiments or scopes, rather than to limit the upper or lower limits of the number of components, nor to limit the order of components.
[0020] In addition, it should be noted that the per-channel convolution device and its operation method provided by at least one embodiment of the present disclosure can be applied to an artificial intelligence chip or other chips. The parameter type of the per-channel convolution device and its operation method is a floating-point number, and the specific physical meaning of the parameter type can be different according to different application scenarios. For example, the per-channel convolution device and its operation method provided by at least one embodiment of the present disclosure can be applied to fields such as speech processing, image processing, text processing, and video processing.
[0021] For example, in the field of speech processing, the parameters can be any parameters used, input, or generated in tasks such as feature extraction, speech enhancement, and speech recognition, such as speech feature vectors, filtering parameters, etc.
[0022] For example, in the field of image processing, the parameter can be any parameter used, input, or generated in tasks such as image preprocessing, feature extraction, image segmentation, object detection, etc., such as an image feature vector, various edge detection operators (such as Sobel operator, Canny operator, Prewitt operator, etc.), image filtering operators (such as Gaussian filtering, median filtering, bilateral filtering, etc.), morphological operators (such as erosion, dilation, opening operation, closing operation, etc.), and other parameters used, input, or generated.
[0023] For example, in the field of text processing, the parameter can be any parameter used, input, or generated in tasks such as text classification, sentiment analysis, text generation, etc., such as a semantic feature vector of the text.
[0024] For example, in the field of video processing, the parameter can be a parameter in the field of image processing as described above, or a parameter used, input, or generated specifically in the field of video processing, such as an optical flow operator (used to estimate the motion between video frames), an object tracking operator (used to track a specific object in a video), etc.
[0025] Of course, the present disclosure is not limited to this. For other application scenarios or fields, as long as per-channel convolution is required, the per-channel convolution device and its operation method described in at least one embodiment of the present disclosure can be applied, and they will not be elaborated one by one here.
[0026] Figure 1 It is a schematic diagram of a circuit module of a per-channel convolution device 100 according to an embodiment of the present invention. The operation process of depthwise separable convolution includes two stages: per-channel convolution and pointwise convolution. The per-channel convolution device 100 can be used as a dedicated hardware for "accelerating the per-channel convolution stage in the depthwise separable convolution operation". The directions of per-channel convolution that the per-channel convolution device 100 can support include forward convolution and backward convolution. The sizes of per-channel convolution kernels that the per-channel convolution device 100 can support include 1×1, 2×2, 3×3, 4×4, 5×5, or other sizes. The padding amounts in the x direction that the per-channel convolution device 100 can support include 0, +1, -1, +2, -2, or other padding amounts. The padding amounts in the y direction that the per-channel convolution device 100 can support include 0, +1, -1, +2, -2, or other padding amounts. The stride amounts that the per-channel convolution device 100 can support include 1, 2, or other stride amounts. The dilation amounts that the per-channel convolution device 100 can support include 1, 2, or other dilation amounts. The data types that the per-channel convolution device 100 can support include fp32, fp16, fp8, int16, int8, int4, or other data types.
[0027] When accelerating the per-channel convolution application, the instruction scheduling module 11 issues per-channel convolution instructions to the vector calculation unit. Based on the per-channel convolution instructions issued by the instruction scheduling module 11, the per-channel convolution device 100 can perform the per-channel convolution in the depthwise separable convolution, and then store the per-channel convolution result matrix in memory or registers (such as Figure 1 the register bank 140 shown). The per-channel convolution result matrix in memory (or register bank) can be used for the second stage of the depthwise separable convolution operation, "pointwise convolution".
[0028] In Figure 1 the embodiment shown, the per-channel convolution device 100 includes a per-channel convolution control module 110, a feature map input control module 120, a convolution kernel input control module 130, a register file 140, a constant cache 150, and a plurality of floating-point arithmetic units (such as Figure 1 the first floating-point arithmetic unit 160_1, the second floating-point arithmetic unit 160_2,..., the Nth floating-point arithmetic unit 160_N shown, where the first floating-point arithmetic unit can refer to any one of the first floating-point arithmetic unit 160_1, the second floating-point arithmetic unit 160_2,..., the Nth floating-point arithmetic unit 160_N). According to different designs, in some embodiments, the implementation of at least one of the instruction scheduling module 11, the per-channel convolution control module 110, the feature map input control module 120, and the convolution kernel input control module 130 can be a hardware circuit. In other embodiments, the implementation of at least one of the instruction scheduling module 11, the per-channel convolution control module 110, the feature map input control module 120, and the convolution kernel input control module 130 can be a combination of multiple forms of hardware, firmware, and software (i.e., programs).
[0029] In terms of hardware, at least one of the above instruction scheduling module 11, per-channel convolution control module 110, feature map input control module 120, and convolution kernel input control module 130 can be implemented as logic circuits on an integrated circuit. For example, the related functions of at least one of the instruction scheduling module 11, per-channel convolution control module 110, feature map input control module 120, and convolution kernel input control module 130 can be implemented in one or more controllers, hardware controllers, microcontrollers, hardware processors, microprocessors, application-specific integrated circuits (ASICs), digital signal processors (DSPs), field programmable gate arrays (FPGAs), central processing units, or various logic blocks, modules, and circuits in other processing units. The related functions of at least one of the instruction scheduling module 11, per-channel convolution control module 110, feature map input control module 120, and convolution kernel input control module 130 can be implemented as hardware circuits, such as various logic blocks, modules, and circuits in an integrated circuit, using hardware description languages (such as Verilog HDL or VHDL) or other suitable programming languages.
[0030] In the form of "software or firmware running on hardware", the related functions of at least one of the above instruction scheduling module 11, per-channel convolution control module 110, feature map input control module 120, and convolution kernel input control module 130 can be implemented as programming codes. For example, use general programming languages (such as C, C++, or assembly language) or other suitable programming languages to implement at least one of the instruction scheduling module 11, per-channel convolution control module 110, feature map input control module 120, and convolution kernel input control module 130. The programming codes can be recorded or stored in a "non-transitory machine-readable storage medium". In some embodiments, the non-transitory machine-readable storage medium includes, for example, semiconductor memory and / or storage devices. An electronic device (such as a CPU, hardware controller, microcontroller, hardware processor, or microprocessor) can read and execute the programming codes from the non-transitory machine-readable storage medium to implement the related functions of at least one of the instruction scheduling module 11, per-channel convolution control module 110, feature map input control module 120, and convolution kernel input control module 130.
[0031] Figure 1 The shown constant cache 150 and register bank 140 can be internal components of the per-channel convolution device 100. However, in other embodiments, at least one of the constant cache 150 and register bank 140 can be an external component of the per-channel convolution device 100. The implementation manners of the constant cache 150 and register bank 140 are not limited in this embodiment. For example, at least one of the constant cache 150 and register bank 140 can be a register, cache, main memory, or other memory. The constant cache 150 is used to store convolution kernels. The constant cache 150 is coupled to the convolution kernel input control module 130 to provide the convolution kernels. The register bank 140 is used to store the feature map matrix. The register bank 140 is coupled to the feature map input control module 120 to provide the feature map matrix.
[0032] The number N of the first floating-point operation unit 160_1 to the Nth floating-point operation unit 160_N can be determined according to the actual design. For example, assuming that the feature map matrix is a u×v matrix, the number N of the first floating-point operation unit 160_1 to the Nth floating-point operation unit 160_N can be any integer in the range of 1 to u×v in one embodiment. If the number N of the first floating-point operation unit 160_1 to the Nth floating-point operation unit 160_N is u×v and assuming a stride of 1, the per-channel convolution device 100 can complete the per-channel convolution of a feature map matrix by performing a single iteration. If the number N of the first floating-point operation unit 160_1 to the Nth floating-point operation unit 160_N is 1 and assuming a stride of 1, the per-channel convolution device 100 needs to perform u×v iterations to complete the per-channel convolution of a feature map matrix.
[0033] Figure 2 is a schematic flowchart of an operation method of a per-channel convolution device according to an embodiment of the present invention. Please refer to Figure 1 and Figure 2 , in step S210, the per-channel convolution control module 110 executes the per-channel convolution instruction issued by the instruction scheduling module 11. The feature map input control module 120 and the convolution kernel input control module 130 are coupled to the per-channel convolution control module 110. The per-channel convolution control module 110 controls the feature map input control module 120 and the convolution kernel input control module 130 based on the per-channel convolution instruction. The first floating-point operation unit 160_1 to the Nth floating-point operation unit 160_N are coupled to the convolution kernel input control module 130 and the feature map input control module 120. In step S220, the convolution kernel input control module 130 calls the convolution kernel from the constant cache 150 based on the control of the per-channel convolution control module 110, and broadcasts the convolution kernel to each of the first floating-point operation unit 160_1 to the Nth floating-point operation unit 160_N. In addition, in step S220, the feature map input control module 120 calls a partial matrix (which can be generally referred to as the first partial matrix in the first receptive field) in the receptive field from the feature map matrix of the register bank 140 to the corresponding ones (which can be generally referred to as the first floating-point operation unit) of the first floating-point operation unit 160_1 to the Nth floating-point operation unit 160_N based on the control of the per-channel convolution control module 110.
[0034] Figure 3 is a schematic diagram of per-channel convolution shown according to an embodiment of the present invention. In Figure 3 the example shown, the feature map matrix 310 is assumed to be a 6×16 matrix, and the convolution kernel 320 is assumed to be a 3×3 matrix. Please refer to Figure 1 and Figure 3, based on the control of the per-channel convolution control module 110, the convolution kernel input control module 130 broadcasts the convolution kernel 320 to each of the first floating-point operation unit 160_1 to the Nth floating-point operation unit 160_N. Taking the first floating-point operation unit as the first floating-point operation unit 160_1 as an example, the feature map input control module 120 calls the corresponding partial matrix from the feature map matrix 310 (which can be generally referred to as the first partial matrix in the first receptive field, and the matrix elements in the corresponding partial matrix include "Pad", "Pad", "Pad", "Pad", "1", "2", "Pad", "17", and "18") to the corresponding first floating-point operation unit 160_1. Among them, "Pad" represents zero padding. The first floating-point operation unit 160_1 uses the corresponding partial matrix elements of the feature map matrix 310 and the convolution kernel 320 to calculate the corresponding result element in the per-channel convolution result matrix (Pad·W1 + Pad·W2 + Pad·W3 + Pad·W4 +1·W5 + 2·W6 + Pad·W7 + 17·W8 + 18·W9) (which can be generally referred to as the first result element), where "W1" to "W9" are the weight elements in the convolution kernel 320.
[0035] The per-channel convolution device 100 moves the multiplication and addition window corresponding to the convolution kernel 320 to different positions of the feature map matrix 310 according to the stride, thereby forming different receptive fields. Then, after each movement of the multiplication and addition window, the per-channel convolution device 100 performs multiplication and addition calculations on the data elements (such as feature points) of the partial matrix within the receptive field to obtain a result element in the per-channel convolution result matrix. After the multiplication and addition window completely scans the entire feature map matrix 310 according to the stride, the per-channel convolution device 100 can store the per-channel convolution result matrix in the register bank 140. For example (taking the first floating-point operation unit as the second floating-point operation unit 160_2), assuming the stride is 1, the feature map input control module 120 calls the corresponding partial matrix from the feature map matrix 310 (which can be generally referred to as the first partial matrix within the first receptive field, and the matrix elements in this corresponding partial matrix include "Pad", "Pad", "Pad", "1", "2", "3", "17", "18", and "19") to the corresponding second floating-point operation unit 160_2. The second floating-point operation unit 160_2 uses the corresponding partial matrix elements of the feature map matrix 310 and the convolution kernel 320 to calculate the corresponding result element in the per-channel convolution result matrix (Pad·W1 + Pad·W2 + Pad·W3 + 1·W4 + 2·W5 + 3·W6 + 17·W7 + 18·W8 + 19·W9) (which can be generally referred to as the second result element relative to the first result element generally referred to for the above-mentioned first floating-point operation unit 160_1). The remaining floating-point operation units (such as the Nth floating-point operation unit 160_N) can refer to the relevant descriptions of the first floating-point operation unit 160_1 and the second floating-point operation unit 160_2 and make analogies. The first floating-point operation unit 160_1 to the Nth floating-point operation unit 160_N each calculate at least one result element in the per-channel convolution result matrix.
[0036] There are various layout ways for the feature map matrix 310 stored in the register bank 140, such as one-dimensional data layout and two-dimensional data layout. Figure 4 FIG. is a schematic diagram of the data layout in which the feature map matrix 310 is stored in the register bank 140 in a one-dimensional data layout manner according to an embodiment of the present invention. Figure 4 The feature map matrix 310 is shown again at the upper part. Figure 4 The lower part shows a schematic diagram of the data layout of the first register 1, the second register 2, the third register 3, the fourth register 4, the fifth register 5, the sixth register 6, the seventh register 7, the eighth register 8, and the ninth register 9 in the register bank 140. In Figure 4In the example shown, the one-dimensional data layout is assumed to be in row priority mode, the stride is assumed to be 1, and the zero-padding width is assumed to be 1, where "Pad" means padding zeros at the boundary of the feature map matrix 310.
[0037] The partial matrix of the feature map matrix 310 in the first receptive field (which can be generally referred to as the first partial matrix) is stored at the first address of each of the first register 1 to the ninth register 9 (i.e., multiple registers) in the register bank 140, for example Figure 4 The leftmost first column of the lower part. The partial matrix of the feature map matrix 310 in the second receptive field (which can be generally referred to as the second partial matrix) is stored at the second address of each of the first register 1 to the ninth register 9 (i.e., multiple registers) in the register bank 140, for example Figure 4 The leftmost second column of the lower part. The feature map input control module 120 calls the first partial matrix of the feature map matrix 310 from the first address of each of the first register 1 to the ninth register 9 (i.e., multiple registers) to the first floating-point operation unit 160_1 (which can be generally referred to as the first floating-point operation unit) based on the control of the per-channel convolution control module 110. The feature map input control module 120 calls the second partial matrix of the feature map matrix 310 from the second address of each of the first register 1 to the ninth register 9 (i.e., multiple registers) to the second floating-point operation unit 160_2 (which can be generally referred to as the second floating-point operation unit) based on the control of the per-channel convolution control module 110. The remaining floating-point operation units (such as the Nth floating-point operation unit 160_N) can refer to the relevant descriptions of the first floating-point operation unit 160_1 and the second floating-point operation unit 160_2 and be analogized accordingly.
[0038] Please refer to Figure 1 and Figure 2, in step S230, each of the first floating-point arithmetic unit 160_1 to the Nth floating-point arithmetic unit 160_N performs one or more multi-channel multiply-accumulate operations based on the control of the per-channel convolution control module 110. Each multi-channel multiply-accumulate operation includes a first cycle and a second cycle. Each of the first floating-point arithmetic unit 160_1 to the Nth floating-point arithmetic unit 160_N (collectively referred to as the first floating-point arithmetic unit) performs multiply-accumulate calculations using multiple matrix elements (collectively referred to as multiple first matrix elements) of the corresponding receptive field (collectively referred to as the first receptive field) and multiple weight elements of the convolution kernel (collectively referred to as multiple first weight elements) in the first cycle of the multi-channel multiply-accumulate operation to generate a stage sum value (collectively referred to as the first stage sum value) (step S240). Each of the first floating-point arithmetic unit 160_1 to the Nth floating-point arithmetic unit 160_N (collectively referred to as the first floating-point arithmetic unit) performs accumulation calculations using the stage sum value (collectively referred to as the first stage sum value) and the previous accumulated value in the second cycle of the multi-channel multiply-accumulate operation (collectively referred to as the first multi-channel multiply-accumulate operation) to generate an accumulation result of the multi-channel multiply-accumulate operation (collectively referred to as the first accumulation result) (step S250). In the absence of a previous accumulated value, the floating-point arithmetic unit (collectively referred to as the first floating-point arithmetic unit) performs accumulation calculations using the stage sum value and an initial value (or bias value) in the second cycle. The initial value (or bias value) can be determined based on the actual design and application. For example, the initial value (or bias value) can be 0 or other values. In step S260, each of the first floating-point arithmetic unit 160_1 to the Nth floating-point arithmetic unit 160_N (collectively referred to as the first floating-point arithmetic unit) calculates the corresponding result element (collectively referred to as the first result element) in the per-channel convolution result matrix based on the accumulation result of the multi-channel multiply-accumulate operation (collectively referred to as the first multi-channel multiply-accumulate operation).
[0039] Figure 5 is a schematic diagram of a floating-point arithmetic unit (such as the first floating-point arithmetic unit 160_1) performing multiple multi-channel multiply-accumulate operations according to an embodiment of the present invention. The remaining second floating-point arithmetic unit 160_2 to the Nth floating-point arithmetic unit 160_N can refer to the relevant description of the first floating-point arithmetic unit 160_1 and be analogized accordingly. In Figure 5In the illustrated embodiment, the first floating-point arithmetic unit 160_1 performs five multi-channel multiply-accumulate operations, namely, the first multi-channel multiply-accumulate operation 510, the second multi-channel multiply-accumulate operation 520, the third multi-channel multiply-accumulate operation 530, the fourth multi-channel multiply-accumulate operation 540, and the fifth multi-channel multiply-accumulate operation 550. Each of the first multi-channel multiply-accumulate operation 510 to the fifth multi-channel multiply-accumulate operation 550 (which can be generally referred to as the first multi-channel multiply-accumulate operation; the previous multi-channel multiply-accumulate operation of two consecutive multi-channel multiply-accumulate operations can also be generally referred to as the first multi-channel multiply-accumulate operation, and the subsequent multi-channel multiply-accumulate operation of two consecutive multi-channel multiply-accumulate operations can be generally referred to as the second multi-channel multiply-accumulate operation) performs multiply-accumulate calculations in the first cycle, including half-precision multiplication and half-precision addition. The stage sum value in the first cycle (which can be generally referred to as the first stage sum value when in the first multi-channel multiply-accumulate operation in the first cycle and can be generally referred to as the second stage sum value when in the second multi-channel multiply-accumulate operation in the first cycle) is a half-precision floating-point number. The accumulation calculation performed in the second cycle of each of the first multi-channel multiply-accumulate operation 510 to the fifth multi-channel multiply-accumulate operation 550 includes full-precision addition, and the accumulation result in the second cycle (which can be generally referred to as the first accumulation result when in the first multi-channel multiply-accumulate operation in the second cycle and can be generally referred to as the second accumulation result when in the second multi-channel multiply-accumulate operation in the second cycle) is a full-precision floating-point number.
[0040] In the illustrative embodiment, the plurality of first matrix elements of the first receptive field include a first element and a second element, and the plurality of first weight elements of the convolution kernel include a first weight and a second weight.
[0041] In Figure 5In the illustrated embodiment, the feature map input control module 120 calls a corresponding partial matrix (which can be generally referred to as the first partial matrix in the first receptive field, and the matrix elements in the corresponding partial matrix include "Pad", "Pad", "Pad", "Pad", "1", "2", "Pad", "17", and "18") from the feature map matrix 310 in the register bank 140 to the first floating-point arithmetic unit 160_1 (which can be generally referred to as the first floating-point arithmetic unit), and the convolution kernel input control module 130 broadcasts the weight elements "W1", "W2", "W3", "W4", "W5", "W6", "W7", "W8", and "W9" of the convolution kernel 320 to the first floating-point arithmetic unit 160_1 (which can be generally referred to as the first floating-point arithmetic unit). In the first cycle of the first multi-channel multiply-accumulate operation 510 (which can be generally referred to as the first multi-channel multiply-accumulate operation), the first floating-point arithmetic unit 160_1 (which can be generally referred to as the first floating-point arithmetic unit) uses the first element "Pad" (which can be generally referred to as the first element) in the matrix elements of the receptive field and the first weight "W1" (which can be generally referred to as the first weight) in the weight elements to perform a half-precision multiplication to generate a dot product result Pad·W1 (which can be generally referred to as the first dot product result). At the same time, the first floating-point arithmetic unit 160_1 (which can be generally referred to as the first floating-point arithmetic unit) uses the second element "Pad" (which can be generally referred to as the second element) in the matrix elements of the receptive field and the second weight "W2" (which can be generally referred to as the second weight) in the weight elements to perform a half-precision multiplication to generate another dot product result Pad·W2 (which can be generally referred to as the second dot product result). The first floating-point arithmetic unit 160_1 (which can be generally referred to as the first floating-point arithmetic unit) uses the dot product result Pad·W1 (which can be generally referred to as the first dot product result) and the dot product result Pad·W2 (which can be generally referred to as the second dot product result) to perform a half-precision addition to generate the stage sum value of the first cycle of the first multi-channel multiply-accumulate operation 510 (i.e., the stage sum value S1 = Pad·W1 + Pad·W2) (which can be generally referred to as the first stage sum value). The first floating-point arithmetic unit 160_1 (which can be generally referred to as the first floating-point arithmetic unit) uses the stage sum value S1 of the first cycle (which can be generally referred to as the first stage sum value) and the initial value (or bias value, such as 0) in the second cycle of the first multi-channel multiply-accumulate operation 510 (which can be generally referred to as the first multi-channel multiply-accumulate operation) to perform an accumulation calculation to generate the accumulation result SUM1 of the first multi-channel multiply-accumulate operation 510 (which can be generally referred to as the first accumulation result), that is, [(SUM1 = Pad·W1 + Pad·W2) + 0].The accumulation result SUM1 of the first multi-channel multiply-accumulate operation 510 (which can be generally referred to as the first accumulation result) can be used as the "previous accumulated value" (equivalent to the first accumulation result) for the next second multi-channel multiply-accumulate operation 520 (which can be generally referred to as the second multi-channel multiply-accumulate operation).
[0042] The first floating-point operation unit 160_1 (which can be generally referred to as the first floating-point operation unit) performs the second multi-channel multiply-accumulate operation 520 (which can be generally referred to as the second multi-channel multiply-accumulate operation) under the control of the per-channel convolution control module 110. The first floating-point operation unit 160_1 (which can be generally referred to as the first floating-point operation unit) uses the matrix elements of the receptive field (which can be generally referred to as multiple second matrix elements of the first receptive field) and multiple weight elements of the convolution kernel 320 (which can be generally referred to as second weight elements) in the first cycle of the second multi-channel multiply-accumulate operation 520 (which can be generally referred to as the second multi-channel multiply-accumulate operation) to perform multiply-accumulate calculations to generate the stage sum value S2 of the second multi-channel multiply-accumulate operation 520 (which can be generally referred to as the second stage sum value). Specifically, the first floating-point operation unit 160_1 (which can be generally referred to as the first floating-point operation unit) uses the matrix element "Pad" of the receptive field and the weight "W3" in the weight elements in the first cycle of the second multi-channel multiply-accumulate operation 520 (if the first multi-channel multiply-accumulate operation 510 is generally referred to as the first multi-channel multiply-accumulate operation, then the second multi-channel multiply-accumulate operation 520 can be generally referred to as the second multi-channel multiply-accumulate operation) to perform half-precision multiplication to generate the dot product result Pad·W3. At the same time, the first floating-point operation unit 160_1 uses the matrix element "Pad" of the receptive field and the weight "W4" in the weight elements to perform half-precision multiplication to generate another dot product result Pad·W4. The first floating-point operation unit 160_1 uses the dot product result Pad·W3 and the dot product result Pad·W4 to perform half-precision addition to generate the stage sum value in the first cycle of the second multi-channel multiply-accumulate operation 520 (which can be generally referred to as the second multi-channel multiply-accumulate operation) (i.e., the stage sum value S2 = Pad·W3 + Pad·W4) (which can be generally referred to as the second stage sum value). The first floating-point operation unit 160_1 (which can be generally referred to as the first floating-point operation unit) uses the stage sum value S2 in the first cycle (which can be generally referred to as the second stage sum value) and the previous accumulated value (i.e., the accumulated result SUM1 of the first multi-channel multiply-accumulate operation 510) (which can be generally referred to as the first accumulated result) in the second cycle of the second multi-channel multiply-accumulate operation 520 (which can be generally referred to as the second multi-channel multiply-accumulate operation) to perform accumulation calculations to generate the accumulated result SUM2 of the second multi-channel multiply-accumulate operation 520 (which can be generally referred to as the second accumulated result), that is, [(SUM2 = Pad·W3 + Pad·W4) + (Pad·W1 + Pad·W2)].The accumulated result SUM2 of the second multi-channel multiply-accumulate operation 520 can be used as the "previous accumulated value" (which can be generally referred to as the first accumulated result if the second multi-channel multiply-accumulate operation 520 is generally referred to as the first multi-channel multiply-accumulate operation) for the next third multi-channel multiply-accumulate operation 530 (which can be generally referred to as the second multi-channel multiply-accumulate operation if the second multi-channel multiply-accumulate operation 520 is generally referred to as the first multi-channel multiply-accumulate operation).
[0043] The first floating-point operation unit 160_1 (which can be generally referred to as the first floating-point operation unit) uses the matrix element "1" of the receptive field and the weight "W5" in the weight element to perform half-precision multiplication to generate the dot product result 1·W5 in the first cycle of the third multi-channel multiply-accumulate operation 530 (which can be generally referred to as the second multi-channel multiply-accumulate operation if the second multi-channel multiply-accumulate operation 520 is generally referred to as the first multi-channel multiply-accumulate operation). At the same time, the first floating-point operation unit 160_1 uses the matrix element "2" of the receptive field and the weight "W6" in the weight element to perform half-precision multiplication to generate another dot product result 2·W6. The first floating-point operation unit 160_1 uses the dot product result 1·W5 and the dot product result 2·W6 to perform half-precision addition to generate the stage sum value in the first cycle of the third multi-channel multiply-accumulate operation 530 (which can be generally referred to as the second multi-channel multiply-accumulate operation) (i.e., the stage sum value S3 = 1·W5 + 2·W6) (which can be generally referred to as the second stage sum value). The first floating-point operation unit 160_1 (which can be generally referred to as the first floating-point operation unit) uses the stage sum value S3 in the first cycle (which can be generally referred to as the second stage sum value) and the previous accumulated value (which can be generally referred to as the first accumulated result) to perform an accumulation calculation in the second cycle of the third multi-channel multiply-accumulate operation 530 (which can be generally referred to as the second multi-channel multiply-accumulate operation) to generate the accumulated result SUM3 (which can be generally referred to as the second accumulated result) of the third multi-channel multiply-accumulate operation 530 (which can be generally referred to as the second multi-channel multiply-accumulate operation), that is, [(SUM3 = 1·W5 + 2·W6) + (Pad·W1 + Pad·W2) + (Pad·W3 + Pad·W4)].
[0044] The first floating-point arithmetic unit 160_1 (which can be generally referred to as the first floating-point arithmetic unit) uses the matrix element "Pad" of the receptive field and the weight "W7" in the weight element to perform half-precision multiplication to generate the dot product result Pad·W7 in the first cycle of the fourth multi-channel multiply-accumulate operation 540 (if the third multi-channel multiply-accumulate operation 530 is generally referred to as the first multi-channel multiply-accumulate operation, then the fourth multi-channel multiply-accumulate operation 540 can be generally referred to as the second multi-channel multiply-accumulate operation). At the same time, the first floating-point arithmetic unit 160_1 uses the matrix element "17" of the receptive field and the weight "W8" in the weight element to perform half-precision multiplication to generate another dot product result 17·W8. The first floating-point arithmetic unit 160_1 uses the dot product result Pad·W7 and the dot product result 17·W8 to perform half-precision addition to generate the stage sum value in the first cycle of the fourth multi-channel multiply-accumulate operation 540 (which can be generally referred to as the second multi-channel multiply-accumulate operation) (i.e., the stage sum value S4 = Pad·W7 + 17·W8) (which can be generally referred to as the second stage sum value). The first floating-point arithmetic unit 160_1 (which can be generally referred to as the first floating-point arithmetic unit) uses the stage sum value S4 in the first cycle (which can be generally referred to as the second stage sum value) and the previous accumulated value (which can be generally referred to as the first accumulated result) to perform an accumulation calculation in the second cycle of the fourth multi-channel multiply-accumulate operation 540 (which can be generally referred to as the second multi-channel multiply-accumulate operation) to generate the accumulated result SUM4 (which can be generally referred to as the second accumulated result) of the fourth multi-channel multiply-accumulate operation 540, that is, [(SUM4 = Pad·W7 + 17·W8) + (Pad·W1 + Pad·W2) + (Pad·W3 + Pad·W4) + (1·W5 + 2·W6)].
[0045] The first floating-point arithmetic unit 160_1 (which can be generally referred to as the first floating-point arithmetic unit) uses the matrix element "18" of the receptive field and the weight "W9" in the weight elements in the first cycle of the fifth multi-channel multiply-accumulate operation 550 (if the fourth multi-channel multiply-accumulate operation 540 is generally referred to as the first multi-channel multiply-accumulate operation, the fifth multi-channel multiply-accumulate operation 550 can be generally referred to as the second multi-channel multiply-accumulate operation, or the fifth multi-channel multiply-accumulate operation 550 can also be generally referred to as the first multi-channel multiply-accumulate operation) to perform half-precision multiplication to generate a dot product result 18·W9 (that is, the stage sum value S5 of the fifth multi-channel multiply-accumulate operation 550 = 18·W9) (corresponding to the fifth multi-channel multiply-accumulate operation 550 being generally referred to as the first multi-channel multiply-accumulate operation or the second multi-channel multiply-accumulate operation, the stage sum value S5 can be generally referred to as the first stage sum value or the second stage sum value). The first floating-point arithmetic unit 160_1 (which can be generally referred to as the first floating-point arithmetic unit) uses the stage sum value S5 (which can be generally referred to as the first stage sum value or the second stage sum value) and the previous accumulated value (which can also be generally referred to as the first accumulation result) in the second cycle of the fifth multi-channel multiply-accumulate operation 550 (which can be generally referred to as the first multi-channel multiply-accumulate operation or the second multi-channel multiply-accumulate operation) to perform an accumulation calculation to generate the accumulation result SUM5 (which can be generally referred to as the first accumulation result or the second accumulation result) of the fifth multi-channel multiply-accumulate operation 550 (which can be generally referred to as the first multi-channel multiply-accumulate operation or the second multi-channel multiply-accumulate operation), that is, [SUM5 = (18·W9) +(Pad·W1 + Pad·W2) + (Pad·W3 + Pad·W4) + (1·W5 + 2·W6) + (Pad·W7 + 17·W8)]. The accumulation result SUM5 (which can be generally referred to as the first accumulation result or the second accumulation result) of the fifth multi-channel multiply-accumulate operation 550 (which can be generally referred to as the first multi-channel multiply-accumulate operation or the second multi-channel multiply-accumulate operation) can be used as a corresponding result element in the per-channel convolution result matrix (if the fifth multi-channel multiply-accumulate operation 550 is generally referred to as the first multi-channel multiply-accumulate operation, this corresponding result element can be generally referred to as the first result element, if the fifth multi-channel multiply-accumulate operation 550 is generally referred to as the second multi-channel multiply-accumulate operation, this corresponding result element can be generally referred to as the second result element). Therefore, the floating-point arithmetic unit base 160_1 calculates a corresponding result element (which can be generally referred to as the first result element or the second result element) in the per-channel convolution result matrix based on the accumulation results of the first to fifth multi-channel multiply-accumulate operations 510 to 550.
[0046] In one application example, the first cycle of the next multi-channel multiply-add operation (e.g., the second multi-channel multiply-add operation 520) (which can be generally referred to as the second multi-channel multiply-add operation) is located in time after the end of the second cycle of the current multi-channel multiply-add operation (e.g., the first multi-channel multiply-add operation 510) (which can be generally referred to as the first multi-channel multiply-add operation). In another application example, the first cycle of the next multi-channel multiply-add operation (e.g., the second multi-channel multiply-add operation 520) (which can be generally referred to as the second multi-channel multiply-add operation) overlaps in time with the second cycle of the current multi-channel multiply-add operation (e.g., the first multi-channel multiply-add operation 510) (which can be generally referred to as the first multi-channel multiply-add operation). For example, Figure 6 is a timing schematic diagram of multiple multi-channel multiply-add operations shown according to an embodiment of the present invention. Figure 6 The horizontal axis of
[0047] Figure 6 shows that the first multi-channel multiply-add operation 510 to the fifth multi-channel multiply-add operation 550 can be used as Figure 5 one of many implementation examples of the first multi-channel multiply-add operation 510 to the fifth multi-channel multiply-add operation 550 shown. Figure 6 Shown, R1, R2, R3, R4, R5, R6, R7, R8, and R9 respectively represent matrix elements provided by different registers in the register bank 140 (e.g., Figure 4 and Figure 5 the first register 1 to the ninth register 9 shown), and SUM0 represents an initial value (or bias value, e.g., 0). The first floating-point operation unit 160_1 uses the dot product result R1·W1 and the dot product result R2·W2 to perform a half-precision addition to generate the stage sum value S1 in the first cycle of the first multi-channel multiply-add operation 510. The first floating-point operation unit 160_1 uses the stage sum value S1 in the first cycle and the initial value SUM0 in the second cycle of the first multi-channel multiply-add operation 510 to perform an accumulation calculation to generate the accumulation result SUM1 of the first multi-channel multiply-add operation 510. In Figure 6 the shown embodiment, the first cycle of the second multi-channel multiply-add operation 520 overlaps in time with the second cycle of the first multi-channel multiply-add operation 510, the first cycle of the third multi-channel multiply-add operation 530 overlaps in time with the second cycle of the second multi-channel multiply-add operation 520, the first cycle of the fourth multi-channel multiply-add operation 540 overlaps in time with the second cycle of the third multi-channel multiply-add operation 530, and the first cycle of the fifth multi-channel multiply-add operation 550 overlaps in time with the second cycle of the fourth multi-channel multiply-add operation 540.
[0048] The first floating-point arithmetic unit 160_1 uses the dot product results R3·W3 and R4·W4 to perform half-precision addition to generate the stage sum value S2 in the first cycle of the second multi-channel multiply-accumulate operation 520. The first floating-point arithmetic unit 160_1 performs an accumulation calculation in the second cycle of the second multi-channel multiply-accumulate operation 520 using the stage sum value S2 in the first cycle and the previous accumulated value (i.e., the accumulated result SUM1 of the first multi-channel multiply-accumulate operation 510) to generate the accumulated result SUM2 of the second multi-channel multiply-accumulate operation 520. The first floating-point arithmetic unit 160_1 uses the dot product results R5·W5 and R6·W6 to perform half-precision addition to generate the stage sum value S3 in the first cycle of the third multi-channel multiply-accumulate operation 530. The first floating-point arithmetic unit 160_1 performs an accumulation calculation in the second cycle of the third multi-channel multiply-accumulate operation 530 using the stage sum value S3 in the first cycle and the previous accumulated value (i.e., the accumulated result SUM2 of the second multi-channel multiply-accumulate operation 520) to generate the accumulated result SUM3 of the third multi-channel multiply-accumulate operation 530. The first floating-point arithmetic unit 160_1 uses the dot product results R7·W7 and R8·W8 to perform half-precision addition to generate the stage sum value S4 in the first cycle of the fourth multi-channel multiply-accumulate operation 540. The first floating-point arithmetic unit 160_1 performs an accumulation calculation in the second cycle of the fourth multi-channel multiply-accumulate operation 540 using the stage sum value S4 in the first cycle and the previous accumulated value (i.e., the accumulated result SUM3 of the third multi-channel multiply-accumulate operation 530) to generate the accumulated result SUM4 of the fourth multi-channel multiply-accumulate operation 540. In the first cycle of the fifth multi-channel multiply-accumulate operation 550, the first floating-point arithmetic unit 160_1 uses the matrix element "R9" provided by the ninth register 9 and the weight "W9" in the weight elements to perform half-precision multiplication to generate the dot product result R9·W9 (i.e., the stage sum value S5 of the fifth multi-channel multiply-accumulate operation 550). The first floating-point arithmetic unit 160_1 performs an accumulation calculation in the second cycle of the fifth multi-channel multiply-accumulate operation 550 using the stage sum value S5 and the previous accumulated value (i.e., the accumulated result SUM4 of the fourth multi-channel multiply-accumulate operation 540) to generate the accumulated result SUM5 of the fifth multi-channel multiply-accumulate operation 550. The accumulated result SUM5 of the fifth multi-channel multiply-accumulate operation 550 can be used as a corresponding result element in the per-channel convolution result matrix (which can be generally referred to as the first result element or the second result element).
[0049] Figure 7 is a schematic diagram of a circuit module of the floating-point arithmetic unit 700 illustrated according to an embodiment of the present invention. Figure 7 The illustrated floating-point arithmetic unit 700 can be referred to Figure 1 the relevant description of any one of the first floating-point arithmetic unit 160_1 to the Nth floating-point arithmetic unit 160_N shown. Figure 7 The illustrated floating-point arithmetic unit 700 can be used as Figure 1One of many implementation examples of any one of the first floating-point operation units 160_1 to the Nth floating-point operation units 160_N shown. In Figure 7 In the illustrated embodiment, the floating-point operation unit 700 (which can be generally referred to as the first floating-point operation unit) includes a first multiplication circuit 710, a second multiplication circuit 720, an addition circuit 730, and an accumulation circuit 740. The first input terminal of the first multiplication circuit 710 is coupled to the convolution kernel input control module 130 to receive the convolution kernel. The second input terminal of the first multiplication circuit 710 is coupled to the feature map input control module 120 to receive a plurality of elements of a partial matrix of the receptive field (which can be generally referred to as the first partial matrix of the first receptive field). The first input terminal of the second multiplication circuit 720 is coupled to the convolution kernel input control module 130 to receive the convolution kernel. The second input terminal of the second multiplication circuit 720 is coupled to the feature map input control module 120 to receive a plurality of elements of a partial matrix of the receptive field (which can be generally referred to as the first partial matrix of the first receptive field). For example, in the first multi-channel multiply-accumulate operation 510, the first multiplication circuit 710 receives the corresponding weight element "W1" of the convolution kernel 320 and the corresponding partial matrix element "Pad" of the feature map matrix 310, while the second multiplication circuit 720 receives the corresponding weight element "W2" of the convolution kernel 320 and the corresponding partial matrix element "Pad" of the feature map matrix 310.
[0050] The first input terminal of the addition circuit 730 is coupled to the output terminal of the first multiplication circuit 710. The second input terminal of the addition circuit 730 is coupled to the output terminal of the second multiplication circuit 720. The input terminal of the accumulation circuit 740 is coupled to the output terminal of the addition circuit 730. For example, in the first cycle of the first multi-channel multiply-accumulate operation 510, the first multiplication circuit 710 performs a half-precision multiplication to generate a dot product result Pad·W1. At the same time, the second multiplication circuit 720 performs a half-precision multiplication to generate another dot product result Pad·W2. The addition circuit 730 performs a half-precision addition to generate the stage sum value S1 of the first cycle of the first multi-channel multiply-accumulate operation 510. The accumulation circuit 740 uses the stage sum value S1 of the first cycle and an initial value (or bias value, such as 0) to perform an accumulation calculation in the second cycle of the first multi-channel multiply-accumulate operation 510 to generate the accumulation result SUM1 of the first multi-channel multiply-accumulate operation 510.
[0051] In the second multi-channel multiply-accumulate operation 520, the first multiplication circuit 710 receives the corresponding weight element "W3" of the convolution kernel 320 and the corresponding partial matrix element "Pad" of the feature map matrix 310, while the second multiplication circuit 720 receives the corresponding weight element "W4" of the convolution kernel 320 and the corresponding partial matrix element "Pad" of the feature map matrix 310. In the first cycle of the second multi-channel multiply-accumulate operation 520, the first multiplication circuit 710 performs half-precision multiplication to generate the dot product result Pad·W3, and the second multiplication circuit 720 performs half-precision multiplication to generate another dot product result Pad·W4. The addition circuit 730 performs half-precision addition to generate the stage sum value S2 in the first cycle of the second multi-channel multiply-accumulate operation 520. The accumulation circuit 740 uses the stage sum value S2 in the first cycle and the accumulation result SUM1 of the first multi-channel multiply-accumulate operation 510 to perform an accumulation calculation in the second cycle of the second multi-channel multiply-accumulate operation 520 to generate the accumulation result SUM2 of the second multi-channel multiply-accumulate operation 520. And so on, the accumulation circuit 740 uses the stage sum value S5 and the previous accumulation value to perform an accumulation calculation in the second cycle of the fifth multi-channel multiply-accumulate operation 550 to generate the accumulation result SUM5 of the fifth multi-channel multiply-accumulate operation 550. Therefore, the floating-point operation unit 700 calculates a corresponding result element (which can be generally referred to as the first result element or the second result element) in the per-channel convolution result matrix based on the accumulation results of the first multi-channel multiply-accumulate operation 510 to the fifth multi-channel multiply-accumulate operation 550.
[0052] Figure 8 FIG. is a schematic diagram of circuit modules of the first multiplication circuit 710, the second multiplication circuit 720, the addition circuit 730, and the accumulation circuit 740 according to an embodiment of the present invention. Figure 8 The illustrated first multiplication circuit 710, second multiplication circuit 720, addition circuit 730, and accumulation circuit 740 can be referred to Figure 7 for related descriptions. Figure 8 The illustrated first multiplication circuit 710, second multiplication circuit 720, addition circuit 730, and accumulation circuit 740 can be used as Figure 7 one of many implementation examples of the illustrated first multiplication circuit 710, second multiplication circuit 720, addition circuit 730, and accumulation circuit 740.
[0053] In Figure 8In the illustrated embodiment, the first multiplication circuit 710 includes a first multiplexer 711, a second multiplexer 712, and a first half-precision multiplier 713. In this embodiment, the first half-precision multiplier 713 may be a component of a standard floating-point arithmetic unit, and the first multiplexer 711 and the second multiplexer 712 are added to the standard floating-point arithmetic unit to implement a hardware specifically for per-channel convolution operations. The first input terminal of the first multiplexer 711 is coupled to the feature map input control module 120 to receive a plurality of elements of a partial matrix of the receptive field. The first input terminal of the second multiplexer 712 is coupled to the convolution kernel input control module 130 to receive a convolution kernel. The first input terminal of the first half-precision multiplier 713 is coupled to the output terminal of the first multiplexer 711. The second input terminal of the first half-precision multiplier 713 is coupled to the output terminal of the second multiplexer 712. The output terminal of the first half-precision multiplier 713 is coupled to the first input terminal of the addition circuit 730.
[0054] When the first multiplication circuit 710 is used as a hardware specifically for per-channel convolution operations, the first multiplexer 711 couples the feature map input control module 120 to the first input terminal of the first half-precision multiplier 713, and the second multiplexer 712 couples the convolution kernel input control module 130 to the second input terminal of the first half-precision multiplier 713. When the first multiplication circuit 710 is used as a multiplication circuit of a standard floating-point arithmetic unit, and when the standard floating-point arithmetic unit performs a floating-point multiplication "a1·b1", the second input terminal of the first multiplexer 711 receives the floating-point number a1, and the second input terminal of the second multiplexer 712 receives the floating-point number b1. At this time, the first multiplexer 711 transmits the floating-point number a1 to the first input terminal of the first half-precision multiplier 713, and the second multiplexer 712 transmits the floating-point number b1 to the second input terminal of the first half-precision multiplier 713.
[0055] In Figure 8In the illustrated embodiment, the second multiplication circuit 720 includes a third multiplexer 721, a fourth multiplexer 722, and a second half-precision multiplier 723. In this embodiment, the second half-precision multiplier 723 may be a component of a standard floating-point arithmetic unit, and the third multiplexer 721 and the fourth multiplexer 722 are added to the standard floating-point arithmetic unit to implement a hardware specifically for per-channel convolution operations. The first input terminal of the third multiplexer 721 is coupled to the feature map input control module 120 to receive a plurality of elements of a partial matrix of the receptive field. The first input terminal of the fourth multiplexer 722 is coupled to the convolution kernel input control module 130 to receive the convolution kernel. The first input terminal of the second half-precision multiplier 723 is coupled to the output terminal of the third multiplexer 721. The second input terminal of the second half-precision multiplier 723 is coupled to the output terminal of the fourth multiplexer 722. The output terminal of the second half-precision multiplier 723 is coupled to the second input terminal of the addition circuit 730.
[0056] When the second multiplication circuit 720 is used as a hardware specifically for per-channel convolution operations, the third multiplexer 721 couples the feature map input control module 120 to the first input terminal of the second half-precision multiplier 723, and the fourth multiplexer 722 couples the convolution kernel input control module 130 to the second input terminal of the second half-precision multiplier 723. When the second multiplication circuit 720 is used as a multiplication circuit of a standard floating-point arithmetic unit, and when the standard floating-point arithmetic unit performs a floating-point multiplication "a2·b2", the second input terminal of the third multiplexer 721 receives the floating-point number a2, and the second input terminal of the fourth multiplexer 722 receives the floating-point number b2. At this time, the third multiplexer 721 transmits the floating-point number a2 to the first input terminal of the second half-precision multiplier 723, and the fourth multiplexer 722 transmits the floating-point number b2 to the second input terminal of the second half-precision multiplier 723.
[0057] In Figure 8 In the illustrated embodiment, the addition circuit 730 includes a half-precision fast adder 731. The first input terminal of the half-precision fast adder 731 is coupled to the output terminal of the first multiplication circuit 710. The second input terminal of the half-precision fast adder 731 is coupled to the output terminal of the second multiplication circuit 720. The half-precision fast adder 731 performs a fast addition on the half-precision outputs of two channels and outputs a fast addition result (a half-precision floating-point number, that is, the stage sum value of the first cycle). The output terminal of the half-precision fast adder 731 is coupled to the input terminal of the accumulation circuit 740 to provide the stage sum value.
[0058] In Figure 8In the illustrated embodiment, the accumulation circuit 740 includes a format conversion module 746, a fifth multiplexer 741, a sixth multiplexer 742, a full-precision adder 743, a floating-point normalization module 744, and a loop accumulation result memory access module 745. In this embodiment, the full-precision adder 743 and the floating-point normalization module 744 can be components of a standard floating-point arithmetic unit, while the format conversion module 746, the fifth multiplexer 741, the sixth multiplexer 742, and the loop accumulation result memory access module 745 are added to the standard floating-point arithmetic unit to implement a hardware specifically for per-channel convolution operations. The input end of the format conversion module 746 is coupled to the output end of the addition circuit 730. The format conversion module 746 converts a half-precision floating-point number into a full-precision floating-point number and gives it to the fifth multiplexer 741. The first input end of the fifth multiplexer 741 is coupled to the output end of the addition circuit 730. The first input end of the full-precision adder 743 is coupled to the output end of the fifth multiplexer 741. The second input end of the full-precision adder 743 is coupled to the output end of the sixth multiplexer 742.
[0059] When the accumulation circuit 740 is used as a hardware specifically for per-channel convolution operations, the fifth multiplexer 741 couples the output end of the addition circuit 730 to the first input end of the full-precision adder 743, and the sixth multiplexer 742 couples the output end of the loop accumulation result memory access module 745 to the second input end of the full-precision adder 743. Alternatively, the sixth multiplexer 742 transmits an externally stored accumulated value (or bias value) to the second input end of the full-precision adder 743. When the accumulation circuit 740 is used as the accumulation circuit of a standard floating-point arithmetic unit, and when the standard floating-point arithmetic unit performs an addition "A + B", the second input end of the fifth multiplexer 741 receives the floating-point number A, and the second input end of the sixth multiplexer 742 receives the floating-point number B. At this time, the fifth multiplexer 741 transmits the floating-point number A to the first input end of the full-precision adder 743, and the sixth multiplexer 742 transmits the floating-point number B to the second input end of the full-precision adder 743.
[0060] The input end of the floating-point normalization module 744 is coupled to the output end of the full-precision adder 743. The input end of the loop accumulation result memory access module 745 is coupled to the output end of the floating-point normalization module 744. The floating-point normalization module 744 normalizes the output value of the full-precision adder 743 based on the floating-point standard format to generate a normalized sum value for the loop accumulation result memory access module 745. The output end of the loop accumulation result memory access module 745 is coupled to the first input end of the sixth multiplexer 742. The loop accumulation result memory access module 745 saves the result of the current loop (the current multi-channel multiply-accumulate operation) according to the current loop count information sent by the multi-channel multiply-accumulate instruction (DP2A) decoding module 112 and sends it to the next loop accumulation.
[0061] For example, in the first multi-channel multiply-accumulate operation 510, the loop accumulation result memory access module 745 provides an initial value (or a bias value, such as 0) to the sixth multiplexer 742. In the second multi-channel multiply-accumulate operation 520, the loop accumulation result memory access module 745 provides the accumulation result SUM1 of the first multi-channel multiply-accumulate operation 510 to the sixth multiplexer 742. In the third multi-channel multiply-accumulate operation 530, the loop accumulation result memory access module 745 provides the accumulation result SUM2 of the second multi-channel multiply-accumulate operation 520 to the sixth multiplexer 742. In the fourth multi-channel multiply-accumulate operation 540, the loop accumulation result memory access module 745 provides the accumulation result SUM3 of the third multi-channel multiply-accumulate operation 530 to the sixth multiplexer 742. In the fifth multi-channel multiply-accumulate operation 550, the loop accumulation result memory access module 745 provides the accumulation result SUM4 of the fourth multi-channel multiply-accumulate operation 540 to the sixth multiplexer 742.
[0062] According to application requirements, the feature map and weight data formats are usually in half-precision (16 bits), so each double word (DWORD, i.e., 32 bits) can store data of two different channels. The final result usually requires accuracy, so the accumulation circuit 740 outputs a floating-point number in full-precision format to improve the accuracy of feature extraction.
[0063] Figure 9 It is a schematic diagram of the circuit modules of the per-channel convolution control module 110, the feature map input control module 120, and the convolution kernel input control module 130 shown according to an embodiment of the present invention. Figure 9 The per-channel convolution control module 110, the feature map input control module 120, and the convolution kernel input control module 130 shown can be referred to Figure 1 for the relevant description. Figure 9 The per-channel convolution control module 110, the feature map input control module 120, and the convolution kernel input control module 130 shown can be used as Figure 1One of many implementation examples of the per-channel convolution control module 110, the feature map input control module 120, and the convolution kernel input control module 130 shown.
[0064] In Figure 9 In the illustrated embodiment, the per-channel convolution control module 110 includes an instruction decoding module 111 and a multi-channel multiply-accumulate instruction (DP2A) decoding module 112. The instruction decoding module 111 is used to decode the instructions issued by the instruction scheduling module. The instruction decoding module 111 executes the per-channel convolution instruction and issues the multi-channel multiply-accumulate instruction DP2A to the multi-channel multiply-accumulate instruction decoding module 112 for further decoding of the multi-channel multiply-accumulate instruction DP2A.
[0065] The multi-channel multiply-accumulate instruction decoding module 112 is coupled to the instruction decoding module 111, the convolution kernel input control module 130, and the feature map input control module 120. The multi-channel multiply-accumulate instruction decoding module 112 executes the multi-channel multiply-accumulate instruction DP2A to control the convolution kernel input control module 130, the feature map input control module 120, and the first floating-point operation unit 160_1 to the Nth floating-point operation unit 160_N to perform a multi-channel multiply-accumulate operation (such as any one of the first multi-channel multiply-accumulate operation 510 to the fifth multi-channel multiply-accumulate operation 550). The multi-channel multiply-accumulate instruction decoding module 112 broadcasts the floating-point operation control information to each thread according to the instruction to guide the calculation process until the final result is output. The convolution kernel input control module 130 calls the convolution kernel for the first floating-point operation unit 160_1 to the Nth floating-point operation unit 160_N based on the control of the multi-channel multiply-accumulate instruction decoding module 112. The feature map input control module 120 calls a partial matrix in the receptive field from the feature map matrix of the register bank 140 for the first floating-point operation unit 160_1 to the Nth floating-point operation unit 160_N based on the control of the multi-channel multiply-accumulate instruction decoding module 112.
[0066] In Figure 9 In the illustrated embodiment, the feature map input control module 120 includes an operand acquisition module 121 and a feature map buffer module 122. The operand acquisition module 121 is coupled to the multi-channel multiply-accumulate instruction decoding module 112. The multi-channel multiply-accumulate instruction decoding module 112 obtains the addresses of the corresponding feature map and convolution kernel and gives them to the operand acquisition module 121. The register bank 140 is coupled to the operand acquisition module 121 to provide the feature map matrix.
[0067] In Figure 9In the illustrated embodiment, the convolution kernel input control module 130 includes a convolution kernel acquisition module 131 and a convolution kernel buffer module 132. The convolution kernel acquisition module 131 is coupled to the multi-channel multiply-accumulate instruction decoding module 112. The multi-channel multiply-accumulate instruction decoding module 112 acquires the addresses of the corresponding feature map and the convolution kernel and delivers them to the convolution kernel acquisition module 131. The constant cache 150 is coupled to the convolution kernel acquisition module 131 to provide the convolution kernel.
[0068] The feature map buffer module 122 and the convolution kernel buffer module 132 are controlled by the multi-channel multiply-accumulate instruction decoding module 112 to allocate one of the channels in the dual-channel data storage format to the first floating-point arithmetic units 160_1 to the Nth floating-point arithmetic units 160_N of each thread for calculation, and save the content of the unused channel for the operation in the next instruction cycle. Therefore, even though each multi-channel multiply-accumulate instruction DP2A requires two instruction cycles, in this embodiment, only the data containing dual channels is accessed from memory in the first instruction cycle. For the one-dimensional data layout, the data of the feature map is independent for each thread, while the weight data of the convolution kernel is shared by all threads.
[0069] In summary, this embodiment utilizes the hardware resources of a standard vector arithmetic unit, adds a small amount of control and calculation logic, and implements a hardware specifically for processing one-dimensional data layout and accelerating the execution of per-channel convolution applications. To call this hardware, this embodiment adds a multi-channel multiply-accumulate instruction DP2A that supports mixed-precision floating-point multiply-accumulate operations. Taking a 3×3 convolution kernel as an example, the conventional method requires 9 half-precision multiplications and 18 mixed-precision additions, a total of 27 instruction cycles; this embodiment only requires 5 multi-channel multiply-accumulate operations (each single multi-channel multiply-accumulate operation requires 2 instruction cycles). Therefore, this embodiment can complete the per-channel convolution operation with fewer instructions and faster speed. Different parameter configurations of per-channel convolution, including the size of the feature map, the size of the receptive field, the zero-padding width, the stride, the neural network operation direction (forward or backward), and various data formats (including fp16, bf16, etc.), can be simply mapped to the multi-channel multiply-accumulate instruction DP2A in the one-dimensional data layout. Therefore, this embodiment can adapt to the application requirements of different Convolutional Neural Networks (CNNs) and is very flexible. For dual-channel operations, this embodiment automatically caches the temporarily unused half-precision data in the dual-channel register within the instruction and automatically transfers the full-precision accumulation result between instructions. Therefore, this embodiment can perform less memory access and thus better release computing power.
[0070] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A per-channel convolution device, characterized in that The per-channel convolution device includes: A per-channel convolution control module for executing the per-channel convolution instructions issued by the instruction scheduling module; A convolution kernel input control module coupled to the per-channel convolution control module; A feature map input control module coupled to the per-channel convolution control module; and Multiple floating-point operation units coupled to the convolution kernel input control module and the feature map input control module, where The convolution kernel input control module broadcasts the convolution kernel to the multiple floating-point operation units based on the control of the per-channel convolution control module; The feature map input control module, based on the control of the per-channel convolution control module, calls a first partial matrix in the first receptive field from the feature map matrix to the first floating-point operation unit among the multiple floating-point operation units; The first floating-point operation unit performs a first multi-channel multiply-accumulate operation based on the control of the per-channel convolution control module; The first floating-point operation unit uses multiple first matrix elements of the first receptive field and multiple first weight elements of the convolution kernel to perform multiply-accumulate calculations in the first cycle of the first multi-channel multiply-accumulate operation to generate a first-stage sum value; The first floating-point operation unit uses the first-stage sum value and a previous accumulated value to perform an accumulation calculation in the second cycle of the first multi-channel multiply-accumulate operation to generate a first accumulation result of the first multi-channel multiply-accumulate operation; and The first floating-point operation unit calculates a first result element in the per-channel convolution result matrix based on the first accumulation result of the first multi-channel multiply-accumulate operation.
2. The per-channel convolution device according to claim 1, wherein The per-channel convolution device further includes: A constant cache for storing the convolution kernel, where the constant cache is coupled to the convolution kernel input control module to provide the convolution kernel.
3. The per-channel convolution device according to claim 1, wherein The per-channel convolution device further includes: A register bank for storing the feature map matrix, where the register bank is coupled to the feature map input control module to provide the feature map matrix.
4. The per-channel convolution device according to claim 3, wherein The first partial matrix of the feature map matrix in the first receptive field is stored at a first address of each of the multiple registers of the register bank, the second partial matrix of the feature map matrix in the second receptive field is stored at a second address of each of the multiple registers of the register bank, the feature map input control module calls the first partial matrix from the first address of each of the multiple registers to the first floating-point operation unit based on the control of the per-channel convolution control module, and the feature map input control module calls the second partial matrix from the second address of each of the multiple registers to the second floating-point operation unit among the multiple floating-point operation units based on the control of the per-channel convolution control module.
5. The per-channel convolution device according to claim 1, wherein, The multiply-accumulate calculation performed in the first cycle of the first multi-channel multiply-accumulate operation includes half-precision multiplication and half-precision addition, the first-stage sum value is a half-precision floating-point number, the accumulation calculation performed in the second cycle of the first multi-channel multiply-accumulate operation includes full-precision addition, and the first accumulation result is a full-precision floating-point number.
6. The channel-by-channel convolution device according to claim 5, wherein The multiple first matrix elements of the first receptive field include a first element and a second element, and the multiple first weight elements of the convolution kernel include a first weight and a second weight. In the first cycle of the first multi-channel multiply-accumulate operation, the first floating-point operation unit uses the first element and the first weight to perform the half-precision multiplication to generate a first dot product result, the first floating-point operation unit uses the second element and the second weight to perform the half-precision multiplication to generate a second dot product result, and the first floating-point operation unit uses the first dot product result and the second dot product result to perform the half-precision addition to generate the first stage sum value.
7. The per-channel convolution device according to claim 1, wherein: The first floating-point operation unit performs a second multi-channel multiply-accumulate operation based on the control of the per-channel convolution control module; The first floating-point operation unit performs multiply-accumulate calculations using multiple second matrix elements of the first receptive field and multiple second weight elements of the convolution kernel in the first cycle of the second multi-channel multiply-accumulate operation to generate a second stage sum value; The first floating-point operation unit performs an accumulation calculation using the second stage sum value and the first accumulation result in the second cycle of the second multi-channel multiply-accumulate operation to generate a second accumulation result of the second multi-channel multiply-accumulate operation; And The first floating-point operation unit calculates a second result element in the per-channel convolution result matrix based on the second accumulation result of the second multi-channel multiply-accumulate operation.
8. The per-channel convolution device according to claim 7, wherein The first cycle of the second multi-channel multiply-accumulate operation overlaps in time with the second cycle of the first multi-channel multiply-accumulate operation.
9. The channel-by-channel convolution device according to claim 7, wherein The first cycle of the second multi-channel multiply-accumulate operation is located after the end of the second cycle of the first multi-channel multiply-accumulate operation in time.
10. The per-channel convolution device according to claim 1, characterized in that, The first floating-point operation unit includes: A first multiplication circuit, wherein a first input terminal of the first multiplication circuit is coupled to the convolution kernel input control module to receive the convolution kernel, and a second input terminal of the first multiplication circuit is coupled to the feature map input control module to receive multiple elements of the first partial matrix of the first receptive field; A second multiplication circuit, wherein a first input terminal of the second multiplication circuit is coupled to the convolution kernel input control module to receive the convolution kernel, and a second input terminal of the second multiplication circuit is coupled to the feature map input control module to receive multiple elements of the first partial matrix of the first receptive field; An addition circuit, wherein a first input terminal of the addition circuit is coupled to the output terminal of the first multiplication circuit, and a second input terminal of the addition circuit is coupled to the output terminal of the second multiplication circuit; and An accumulation circuit, wherein an input terminal of the accumulation circuit is coupled to the output terminal of the addition circuit.
11. The per-channel convolutional device according to claim 10, wherein The first multiplication circuit includes: A first multiplexer, wherein a first input terminal of the first multiplexer is coupled to the convolution kernel input control module to receive the convolution kernel; A second multiplexer, wherein a first input terminal of the second multiplexer is coupled to the feature map input control module to receive a plurality of elements of the first partial matrix of the first receptive field; and A first half-precision multiplier, wherein a first input terminal of the first half-precision multiplier is coupled to an output terminal of the first multiplexer, a second input terminal of the first half-precision multiplier is coupled to an output terminal of the second multiplexer, and an output terminal of the first half-precision multiplier is coupled to a first input terminal of the addition circuit.
12. The per-channel convolution device according to claim 10, wherein, The second multiplication circuit includes: A third multiplexer, wherein a first input terminal of the third multiplexer is coupled to the convolution kernel input control module; A fourth multiplexer, wherein a first input terminal of the fourth multiplexer is coupled to the feature map input control module; and A second half-precision multiplier, wherein a first input terminal of the second half-precision multiplier is coupled to an output terminal of the third multiplexer, a second input terminal of the second half-precision multiplier is coupled to an output terminal of the fourth multiplexer, and an output terminal of the second half-precision multiplier is coupled to a second input terminal of the addition circuit.
13. The per-channel convolutional device according to claim 10, wherein The addition circuit includes: A half-precision fast adder, wherein a first input terminal of the half-precision fast adder is coupled to an output terminal of the first multiplication circuit, a second input terminal of the half-precision fast adder is coupled to an output terminal of the second multiplication circuit, and an output terminal of the half-precision fast adder is coupled to an input terminal of the accumulation circuit.
14. The per-channel convolution device according to claim 10, characterized in that, The accumulation circuit includes: A format conversion module, wherein an input terminal of the format conversion module is coupled to an output terminal of the addition circuit, and the format conversion module converts a half-precision floating-point number into a full-precision floating-point number; A fifth multiplexer, wherein a first input terminal of the fifth multiplexer is coupled to an output terminal of the format conversion module; A sixth multiplexer; A full-precision adder, wherein a first input terminal of the full-precision adder is coupled to an output terminal of the fifth multiplexer, and a second input terminal of the full-precision adder is coupled to an output terminal of the sixth multiplexer; A floating-point number normalization module, wherein an input terminal of the floating-point number normalization module is coupled to an output terminal of the full-precision adder; and A cyclic accumulation result memory access module, wherein an input terminal of the cyclic accumulation result memory access module is coupled to an output terminal of the floating-point number normalization module, and an output terminal of the cyclic accumulation result memory access module is coupled to a first input terminal of the sixth multiplexer.
15. The per-channel convolutional device according to claim 1, characterized in that, The per-channel convolution control module includes: An instruction decoding module for decoding the instructions issued by the instruction scheduling module, wherein the instruction decoding module executes the per-channel convolution instruction and issues multi-channel multiply-accumulate instructions; and The multi-channel multiply-accumulate instruction decoding module is coupled to the instruction decoding module, the convolution kernel input control module, and the feature map input control module, wherein the multi-channel multiply-accumulate instruction decoding module executes the multi-channel multiply-accumulate instruction to control the convolution kernel input control module and the feature map input control module. The convolution kernel input control module invokes a convolution kernel to the multiple floating-point arithmetic units based on the control of the multi-channel multiply-accumulate instruction decoding module, and the feature map input control module invokes the first partial matrix in the first receptive field from the feature map matrix to the first floating-point arithmetic unit based on the control of the multi-channel multiply-accumulate instruction decoding module.
16. A method for operating a per-channel convolution device, characterized in that, The operation method includes: Executing, by the per-channel convolution control module of the per-channel convolution device, the per-channel convolution instruction transmitted by the instruction scheduling module, wherein the convolution kernel input control module of the per-channel convolution device is coupled to the per-channel convolution control module, the feature map input control module of the per-channel convolution device is coupled to the per-channel convolution control module, and the multiple floating-point arithmetic units of the per-channel convolution device are coupled to the convolution kernel input control module and the feature map input control module; Broadcasting, by the convolution kernel input control module based on the control of the per-channel convolution control module, the convolution kernel to the multiple floating-point arithmetic units; Invoking, by the feature map input control module based on the control of the per-channel convolution control module, the first partial matrix in the first receptive field from the feature map matrix to the first floating-point arithmetic unit among the multiple floating-point arithmetic units; Performing, by the first floating-point arithmetic unit based on the control of the per-channel convolution control module, a first multi-channel multiply-accumulate operation; Performing, by the first floating-point arithmetic unit in the first cycle of the first multi-channel multiply-accumulate operation, a multiply-accumulate calculation using multiple first matrix elements of the first receptive field and multiple first weight elements of the convolution kernel to generate a first-stage sum value; Performing, by the first floating-point arithmetic unit in the second cycle of the first multi-channel multiply-accumulate operation, an accumulation calculation using the first-stage sum value and a previous accumulated value to generate a first accumulated result of the first multi-channel multiply-accumulate operation; and Calculating, by the first floating-point arithmetic unit based on the first accumulated result of the first multi-channel multiply-accumulate operation, a first result element in the per-channel convolution result matrix.
17. The operating method according to claim 16, wherein, The operation method further includes: Storing the first partial matrix of the feature map matrix in the first receptive field at a first address of each of the multiple registers in the register bank; Storing the second partial matrix of the feature map matrix in the second receptive field at a second address of each of the multiple registers in the register bank; Invoking, by the feature map input control module based on the control of the per-channel convolution control module, the first partial matrix from the first address of each of the multiple registers to the first floating-point arithmetic unit; The second part of the matrix is called from the second address of each of the multiple registers by the feature map input control module under the control of the per-channel convolution control module and given to the second floating-point arithmetic unit among the multiple floating-point arithmetic units.
18. The operating method according to claim 16, wherein The multiply-accumulate calculation performed in the first cycle of the first multi-channel multiply-accumulate operation includes half-precision multiplication and half-precision addition. The first-stage sum is a half-precision floating-point number. The accumulation calculation performed in the second cycle of the first multi-channel multiply-accumulate operation includes full-precision addition, and the first accumulation result is a full-precision floating-point number.
19. The operating method according to claim 18, characterized in that, The multiple first matrix elements of the first receptive field include a first element and a second element. The multiple first weight elements of the convolution kernel include a first weight and a second weight. And the operation method further includes: In the first cycle of the first multi-channel multiply-accumulate operation, the first floating-point arithmetic unit uses the first element and the first weight to perform the half-precision multiplication to generate a first dot product result, uses the second element and the second weight to perform the half-precision multiplication to generate a second dot product result, and uses the first dot product result and the second dot product result to perform the half-precision addition to generate the first-stage sum.
20. The operating method according to claim 16, characterized in that, The operation method further includes: The first floating-point arithmetic unit performs a second multi-channel multiply-accumulate operation based on the control of the per-channel convolution control module; The first floating-point arithmetic unit uses the multiple second matrix elements of the first receptive field and the multiple second weight elements of the convolution kernel to perform a multiply-accumulate calculation in the first cycle of the second multi-channel multiply-accumulate operation to generate a second-stage sum; The first floating-point arithmetic unit uses the second-stage sum and the first accumulation result to perform an accumulation calculation in the second cycle of the second multi-channel multiply-accumulate operation to generate the second accumulation result of the second multi-channel multiply-accumulate operation; and The first floating-point arithmetic unit calculates the second result element in the per-channel convolution result matrix based on the second accumulation result of the second multi-channel multiply-accumulate operation.
21. The operating method according to claim 20, characterized in that, The first cycle of the second multi-channel multiply-accumulate operation overlaps in time with the second cycle of the first multi-channel multiply-accumulate operation.
22. The operating method according to claim 20, wherein The first cycle of the second multi-channel multiply-accumulate operation is located after the end of the second cycle of the first multi-channel multiply-accumulate operation in time.
Citation Information
Patent Citations
Floating point data inverse quantization and quantization method and equipment
CN111240746A
Implementation method of channel-by-channel convolution
CN116957018A