Artificial intelligence chip and its pooling method
By introducing pooling control circuits and floating-point operation units into the artificial intelligence chip, the instruction execution of pooling operations is optimized, and the problem of long calculation time of pooling operations in the existing technology is solved, and a more efficient pooling process is achieved.
Patent Information
- Application Number
- CN202510363402.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-26
AI Technical Summary
When performing pooling operations in the prior art, a large number of floating-point operation instructions are required, resulting in a long calculation time and the pooling operation cannot be efficiently completed.
An artificial intelligence chip is designed, including an instruction scheduling module, an instruction decoding module, a register group, a pooling control circuit and a floating-point operation unit. The number of execution cycles and the value of the read feature map is determined through the pooling control circuit, and the floating-point operation unit is used to perform pooling operations to optimize the pooling process.
By reducing the number of instructions and improving the computing efficiency, faster pooling operations are achieved, especially mean pooling and maximum pooling, which significantly reduces the calculation time.
Smart Images

Figure CN119886242B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of integrated circuit technology, and particularly to an artificial intelligence chip and its pooling method. Background Art
[0002] Artificial Intelligence (AI) chips can provide huge computing power. The huge computing power of AI chips stems from a large number of internal hardware cores. An AI chip usually contains multiple programmable processors, such as a Stream Processor Cluster (SPC). Each programmable processor usually contains multiple Compute Units (CUs, or computing cores), and each computing unit usually contains multiple Execution Units (EUs, or execution cores), such as at least one of an Integer (INT) core, a Floating Point (FP) core, a Tensor core (Tcore), and a Vector core (Vcore). By programming to organize various types of computing units, the AI chip can support general computing, scientific computing, and neural network computing.
[0003] Neural network computing performs operations such as convolution, activation, and pooling. Pooling operations are very common in convolutional neural networks. Pooling is a non-linear downsampling operation. "Pooling" originated in computer vision, referring to the merging and integration of resources, and imitating the human visual system to reduce the dimension of data. The pooling operation moves the pooling window to different positions of the data matrix (such as a tensor or a feature map) according to the stride, thus forming different receptive fields. Then, after each movement of the pooling window, pooling calculations (such as taking the maximum value or averaging) are performed on multiple data elements (such as feature points) in the receptive field to obtain a data element of the pooling matrix. After the pooling window completely scans the entire data matrix according to the stride, the execution unit can store the pooling matrix in the memory. The pooling operation reduces the dimensional size of the data matrix, thereby reducing the number of parameters and the amount of calculation in neural network computing. The pooling operation can control overfitting to a certain extent.
[0004] Common pooling operations include max pooling, average pooling, or other pooling calculation methods. If the size of the receptive field is K×K, and the data elements in the receptive field are P_1, P_2, P_3, …, P_k×k. The result of average pooling is (P_1 + P_2 + P_3 + … + P_k×k) / (K×K). It is generally believed that the background information of the feature map can be better retained through average pooling. The result of max pooling is MAX(P_1, P_2, P_3, …, P_k×k), that is, the maximum value among the K×K feature points. It is generally believed that the texture information of the feature map can be better retained through max pooling.
[0005] Conventional techniques use the conventional floating-point operation hardware of a central processing unit (CPU) or a graphics processing unit (GPU) to complete the pooling operation. The disadvantage of this method is that multiple general-purpose floating-point instructions are required to meet the application requirements. If the receptive field is 3×3 (assuming the data elements in the receptive field are P1, P2, P3, P4, P5, P6, P7, P8, and P9), then average pooling requires "tmp1 = P1 + P2", "tmp2 = P3 + tmp1", "tmp3 = P4 + tmp2", "tmp4 = P5 + tmp3", "tmp5 = P6 + tmp4", "tmp6 = P7 + tmp5", "tmp7 = P8 + tmp6", "tmp8 = P9 + tmp7", and "result = (1 / 9) × tem8", a total of 9 conventional floating-point operation instructions; max pooling requires "tmp1 = MAX(P1, P2)", "tmp2 = MAX(P3, tmp1)", "tmp3 = MAX(P4, tmp2)", "tmp4 = MAX(P5, tmp3)", "tmp5 = MAX(P6, tmp4)", "tmp6 = MAX(P7, tmp5)", "tmp7 = MAX(P8, tmp6)", and "result = MAX(P9, tmp7)", a total of 8 instructions. It can be seen that the shortcoming of the conventional technique is that a large number of instructions lead to a long calculation time. How to perform the pooling operation is one of many technical issues in this field. Summary of the Invention
[0006] The present invention is directed to an artificial intelligence chip and its pooling method for performing a pooling operation on multiple feature points of the feature map values.
[0007] In an embodiment according to the present invention, the artificial intelligence chip includes an instruction scheduling module, an instruction decoding module, a register file, a pooling control circuit, and a first floating point operation unit. The instruction decoding module is coupled to the instruction scheduling module to receive a pooling instruction. The instruction decoding module decodes the pooling instruction to generate a decoding result. The pooling control circuit is coupled to the instruction decoding module and the register file. The first floating point operation unit is coupled to the pooling control circuit. The pooling control circuit determines the number of execution cycles of the pooling instruction based on the decoding result, and provides the number of execution cycles to the first floating point operation unit. The pooling control circuit reads the feature map values from the register file based on the decoding result, and provides multiple feature points of the first receptive field of the feature map values to the first floating point operation unit. The first floating point operation unit performs a pooling operation on the multiple feature points of the first receptive field based on the number of execution cycles to generate a pooling result of the first receptive field.
[0008] In an embodiment according to the present invention, the pooling method includes: decoding, by an instruction decoding module of an artificial intelligence chip, a pooling instruction to generate a decoding result; determining, by a pooling control circuit of the artificial intelligence chip, the number of execution cycles of the pooling instruction based on the decoding result; providing, by the pooling control circuit, the number of execution cycles to a first floating point operation unit; reading, by the pooling control circuit, feature map values from a register file based on the decoding result; providing, by the pooling control circuit, multiple feature points of a first receptive field of the feature map values to the first floating point operation unit; and performing, by the first floating point operation unit, a pooling operation on the multiple feature points of the first receptive field based on the number of execution cycles to generate a pooling result of the first receptive field.
[0009] Based on the above, the pooling control circuits of the embodiments of the present invention determine the number of execution cycles of the pooling instruction and read the feature map values based on the decoding result of the pooling instruction, and the pooling control circuit determines the first receptive field by moving the pooling window based on the stride. The pooling control circuit provides multiple feature points of the first receptive field of the feature map values to the first floating-point operation unit, and the first floating-point operation unit performs a pooling operation (such as average pooling or maximum pooling) on the feature points of the first receptive field based on the number of execution cycles. In some embodiments, the artificial intelligence chip is arranged with multiple floating-point operation units, and the pooling control circuit determines different receptive fields by moving the pooling window based on the stride. The pooling control circuit provides the feature points of each receptive field among the multiple receptive fields to the corresponding floating-point operation unit, so that different floating-point operation units can simultaneously perform pooling operations on the feature points of different receptive fields. In some embodiments, a small amount of control and computing logic is added to the standard floating-point operation unit to implement a hardware resource dedicated to pooling operations (including average pooling and maximum pooling), thereby accelerating the pooling process. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 is a schematic diagram of a circuit block of an artificial intelligence chip according to an embodiment of the present invention;
[0011] Figure 2 is a schematic flowchart of a pooling method of an artificial intelligence chip according to an embodiment of the present invention;
[0012] Figure 3 is a schematic diagram of a circuit block of a pooling control circuit illustrated according to an embodiment of the present invention;
[0013] Figure 4 is a schematic diagram of a circuit block of a floating-point operation unit illustrated according to an embodiment of the present invention;
[0014] Figure 5 is a schematic diagram of a circuit block of a floating-point operation unit illustrated according to another embodiment of the present invention;
[0015] Figure 6 is a schematic diagram of a circuit block of a floating-point operation unit illustrated according to still another embodiment of the present invention;
[0016] Figure 7 is a schematic diagram of a circuit block of an average pooling circuit and a maximum pooling circuit illustrated according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] Reference will now be made in detail to the exemplary embodiments of the present invention, examples of which are illustrated in the accompanying drawings. Whenever possible, the same reference numerals will be used in the drawings and the description to refer to the same or like parts.
[0018] As used throughout the specification (including the claims) of this case, the term "coupled (or connected)" can refer to any direct or indirect connection means. For example, if it is described in the text that a first device is coupled (or connected) to a second device, it should be interpreted that the first device can be directly connected to the second device, or the first device can be indirectly connected to the second device through other devices or certain connection means. The terms "first", "second", etc. mentioned throughout the specification (including the claims) of this case are used to name components (elements), or to distinguish different embodiments or scopes, rather than to limit the upper or lower limits of the number of components, nor to limit the order of components.
[0019] In addition, it should be noted that the parameter type of the artificial intelligence chip and its pooling method provided in at least one embodiment of this case is a floating-point number. The parameter type can have different specific physical meanings according to different application scenarios. For example, the AI chip and its pooling method provided in at least one embodiment of this case can be applied in fields such as speech processing, image processing, text processing, video processing, etc.
[0020] For example, in the field of speech processing, the parameter can be any parameter used, input, or generated in tasks such as feature extraction, speech enhancement, speech recognition, etc., such as speech feature vectors, filtering parameters, etc.
[0021] For example, in the field of image processing, the parameter can be any parameter used, input, or generated in tasks such as image preprocessing, feature extraction, image segmentation, target detection, etc., such as image feature vectors, various edge detection operators (such as Sobel operator, Canny operator, Prewitt operator, etc.), image filtering operators (such as Gaussian filtering, median filtering, bilateral filtering, etc.), morphological operators (such as erosion, dilation, opening operation, closing operation, etc.), etc.
[0022] For example, in the field of text processing, the parameter can be any parameter used, input, or generated in tasks such as text classification, sentiment analysis, text generation, etc., such as semantic feature vectors of text, etc.
[0023] For example, in the field of video processing, the parameter can be the parameters in the image processing field as described above, or parameters used, input, or generated specifically in the video processing field, such as optical flow operators (used to estimate the motion between video frames), target tracking operators (used to track specific targets in the video), etc.
[0024] Of course, this case is not limited to this. For other application scenarios or fields, as long as pooling is required, the pooling method described in at least one embodiment of this case can be applied, and details are not elaborated here one by one.
[0025] Figure 1It is a schematic circuit diagram of an artificial intelligence chip 100 according to an embodiment of the present invention. In different application cases, the artificial intelligence chip 100 can be used as a graphics processing unit, a general-purpose graphics processing unit (General-Purpose computing on GPU, GPGPU), or other processing circuits. The artificial intelligence chip 100 can provide huge computing power. The huge computing power of the artificial intelligence chip 100 comes from a large number of internal hardware cores.
[0026] In Figure 1 the illustrated embodiment, the artificial intelligence chip 100 includes an instruction scheduling module 110, an instruction decoding module 120, a pooling control circuit 130, a register bank 140, and a plurality of floating-point arithmetic units, such as Figure 1 the illustrated floating-point arithmetic unit 150_1, floating-point arithmetic unit 150_2,..., floating-point arithmetic unit 150_N. The number N of the floating-point arithmetic units 150_1 to 150_N can be any integer determined based on actual design and application. The floating-point arithmetic units 150_1 to 150_N are coupled to the pooling control circuit 130 and the register bank 140. The pooling control circuit 130 is coupled to the instruction decoding module 120 and the register bank 140. The instruction decoding module 120 is coupled to the instruction scheduling module 110 to receive pooling instructions. In this embodiment, the pooling instructions include two instructions, POOL.MEAN and POOL.MAX. The POOL.MEAN instruction is used to perform an average pooling operation, and the POOL.MAX instruction is used to perform a maximum pooling operation.
[0027] Figure 2 It is a schematic flowchart of a pooling method for an artificial intelligence chip according to an embodiment of the present invention. Please refer to Figure 1 and Figure 2 . In step S210, the instruction decoding module 120 decodes the pooling instructions from the instruction scheduling module 110 to generate a decoding result. The pooling instructions support different parameter configurations, including pooling type, receptive field size, padding bit width, stride, neural network operation direction (forward or backward), and various data formats (including fp16, fp8, int8, int4), and can adapt to the application requirements of different convolutional neural networks (CNNs). In step S220, the pooling control circuit 130 determines relevant parameters such as the number of execution cycles of the pooling instruction, the pooling type (enabling average or maximum calculation resources), etc., based on the decoding result of the pooling instruction, and reads the feature map values (usually a data matrix) from the register bank 140.
[0028] Assume that the size of the receptive field is K×K, and the number of feature points received or processed by any one of the floating-point arithmetic units 150_1 to 150_N in each loop of the iterative operation is L. Then the number of execution cycles is (K×K) / L. For example, assume that the size of the receptive field is 3×3 feature points, and the number of feature points received or processed by any one of the floating-point arithmetic units 150_1 to 150_N in each loop of the iterative operation is 3. Then the number of execution cycles is (3×3) / 3 = 3, that is, 1 iterative operation performs 3 loops.
[0029] In step S230, the pooling control circuit 130 provides relevant parameters such as the number of execution cycles and the pooling type (enabling the mean or maximum calculation resources) to the floating-point arithmetic units 150_1 to 150_N. In addition, the pooling control circuit 130 determines different receptive fields in the feature map by moving the pooling window based on the stride parameter of the pooling instruction. The pooling control circuit 130 also provides multiple feature points of the corresponding receptive field of the feature map values to each of the floating-point arithmetic units 150_1 to 150_N in step S230. For example, the pooling control circuit 130 provides multiple feature points of the first receptive field of the feature map values to the floating-point arithmetic unit 150_1, and provides multiple feature points of the second receptive field of the feature map values to the floating-point arithmetic unit 150_2.
[0030] In step S240, the floating-point arithmetic units 150_1 to 150_N perform a pooling operation on the feature points of different receptive fields based on the number of execution cycles to generate pooling results of different receptive fields. For example, the floating-point arithmetic unit 150_1 performs a pooling operation on multiple feature points of the first receptive field based on the number of execution cycles to generate a pooling result of the first receptive field. Similarly, the floating-point arithmetic unit 150_2 performs a pooling operation on multiple feature points of the second receptive field based on the number of execution cycles to generate a pooling result of the second receptive field.
[0031] In summary, the pooling control circuit 130 determines the number of execution cycles of the pooling instruction and reads the feature map values based on the decoding result of the pooling instruction, and the pooling control circuit 130 determines the first receptive field by moving the pooling window based on the stride parameter. The pooling control circuit 130 provides multiple feature points of one receptive field of the feature map values to the floating-point operation unit 150_1, and the floating-point operation unit 150_1 performs a pooling operation (such as average pooling or maximum pooling) on the feature points of this one receptive field based on the number of execution cycles to generate a pooling result. Similarly, the pooling control circuit 130 provides multiple feature points of another receptive field of the feature map values to the floating-point operation unit 150_2, and the floating-point operation unit 150_2 performs a pooling operation on the multiple feature points of this another receptive field based on the number of execution cycles to generate a pooling result. The pooling control circuit 130 provides the feature points of multiple receptive fields to the floating-point operation units 150_1 to 150_N. Therefore, different floating-point operation units 150_1 to 150_N can simultaneously perform pooling operations on the feature points of different receptive fields, thereby accelerating the pooling process.
[0032] Figure 3 It is a schematic block diagram of the pooling control circuit 130 shown according to an embodiment of the present invention. Figure 3 The instruction scheduling module 110, instruction decoding module 120, pooling control circuit 130, register bank 140, and floating-point operation units 150_1 to 150_N shown can be referred to Figure 1 for the relevant descriptions, so they will not be elaborated here. Figure 3 The pooling control circuit 130 shown can be used as Figure 1 one of many implementation examples of the pooling control circuit 130 shown. In Figure 3 the embodiment shown, the pooling control circuit 130 includes a pooling operation control module 131, an operand acquisition module 132, and a feature map input control module 133. According to different designs, in some embodiments, the implementation manner of at least one of the instruction scheduling module 110, instruction decoding module 120, pooling control circuit 130, pooling operation control module 131, operand acquisition module 132, and feature map input control module 133 can be a hardware circuit, firmware, software (i.e., a program), or a combination of the foregoing.
[0033] In terms of hardware, at least one of the above instruction scheduling module 110, instruction decoding module 120, pooling control circuit 130, pooling operation control module 131, operand acquisition module 132, and feature map input control module 133 can be implemented as logic circuits on an integrated circuit. For example, the relevant functions of at least one of the instruction scheduling module 110, instruction decoding module 120, pooling control circuit 130, pooling operation control module 131, operand acquisition module 132, and feature map input control module 133 can be implemented in one or more hardware controllers, microcontrollers, hardware processors, microprocessors, application-specific integrated circuits (ASICs), digital signal processors (DSPs), field-programmable gate arrays (FPGAs), central processing units, or various logic blocks, modules, and circuits in other processing units. The relevant functions of at least one of the instruction scheduling module 110, instruction decoding module 120, pooling control circuit 130, pooling operation control module 131, operand acquisition module 132, and feature map input control module 133 can be implemented as hardware circuits, such as various logic blocks, modules, and circuits in an integrated circuit, using hardware description languages (such as Verilog HDL or VHDL) or other suitable programming languages.
[0034] In terms of software or firmware, the related functions of at least one of the above instruction scheduling module 110, instruction decoding module 120, pooling control circuit 130, pooling operation control module 131, operand acquisition module 132, and feature map input control module 133 can be implemented as programming codes. For example, the instruction scheduling module 110, instruction decoding module 120, pooling control circuit 130, pooling operation control module 131, operand acquisition module 132, and feature map input control module 133 are implemented using a general programming language (such as C, C++, or assembly language) or other suitable programming languages. The programming codes can be recorded or stored in a "non-transitory machine-readable storage medium". In some embodiments, the non-transitory machine-readable storage medium includes, for example, semiconductor memory and / or storage devices. An electronic device (such as a computer, CPU, hardware controller, microcontroller, hardware processor, or microprocessor) can read and execute the programming codes from the non-transitory machine-readable storage medium to implement the related functions of at least one of the instruction scheduling module 110, instruction decoding module 120, pooling control circuit 130, pooling operation control module 131, operand acquisition module 132, and feature map input control module 133.
[0035] The pooling operation control module 131 is coupled to the instruction decoding module 120 to receive the decoding result. The pooling instruction passes through the instruction scheduling module 110 and the instruction decoding module 120 and is sent to the pooling operation control module 131 for further decoding of the pooling parameters. The pooling operation control module 131 determines the number of execution cycles of the pooling instruction based on the decoding result of the pooling instruction, and then provides the number of execution cycles to floating-point operation units 150_1 to 150_N. The pooling operation control module 131 also directly broadcasts the pooling calculation resource control information to all floating-point operation units 150_1 to 150_N. For example, the pooling operation control module 131 determines the number of execution cycles of the current instruction according to parameters such as the receptive field size and stride in the instruction, and sends this instruction in a loop until the current per-channel convolution is completed.
[0036] The operand acquisition module 132 is coupled to the pooling operation control module 131 and the register bank 140. The operand acquisition module 132 reads the feature map values from the register bank 140 based on the control of the pooling operation control module 131. For example, the pooling operation control module 131 calculates the address of the feature map according to the parameters and then sends the address to the operand acquisition module 132, so that the operand acquisition module 132 reads the feature map values from the register bank 140. The feature map input control module 133 is coupled to the operand acquisition module 132 to receive the feature map values. The feature map input control module 133 provides the feature points of different receptive fields of the feature map values to the floating-point operation units 150_1 to 150_N. For example, the feature map input control module 133 provides multiple feature points of the first receptive field of the feature map values to the floating-point operation unit 150_1. Similarly, the feature map input control module 133 provides multiple feature points of the second receptive field of the feature map values to the floating-point operation unit 150_2.
[0037] After receiving the relevant feature map, the feature map input control module 133 selects the corresponding feature point inputs for the floating-point operation units 150_1 to 150_N according to the current loop count. For example, assuming that the size of the receptive field is 3×3 feature points, and the data elements (feature points) in a certain receptive field of the feature map values are P1, P2, P3, P4, P5, P6, P7, P8, and P9, and the number of feature points received or processed by the floating-point operation unit 150_1 in each loop of the iterative operation is 3. In the first loop of the iterative operation, the feature map input control module 133 provides the feature points P1, P2, and P3 to the floating-point operation unit 150_1. In the second loop of the iterative operation, the feature map input control module 133 provides the feature points P4, P5, and P6 to the floating-point operation unit 150_1. In the third loop of the iterative operation, the feature map input control module 133 provides the feature points P7, P8, and P9 to the floating-point operation unit 150_1. The operations of the feature map input control module 133 on the other floating-point operation units 150_2 to 150_N can be referred to the operation description of the feature map input control module 133 on the floating-point operation unit 150_1 and extrapolated, so they will not be elaborated here. The floating-point operation units 150_1 to 150_N can receive 3 feature points from the feature map input control module 133 in each loop period and perform the pooling operation. After 3 loop periods, the floating-point operation units 150_1 to 150_N can complete the pooling operations of different receptive fields and output the pooling results of different receptive fields to the register bank 140.
[0038] In some embodiments, a small amount of control and computing logic can be added to a standard floating-point arithmetic unit to implement floating-point arithmetic units 150_1 to 150_N as hardware dedicated to pooling operations (including mean pooling and max pooling), thereby accelerating the pooling process. The pooling control circuit 130 supports different parameter configurations, including pooling type, receptive field size, zero-padding width, stride, neural network operation direction (forward or backward), and various data formats (including fp16, fp8, int8, int4). The pooling control circuit 130 can adapt to the application requirements of different convolutional neural networks and is very flexible.
[0039] Figure 4 FIG. 4 is a schematic circuit diagram of a floating-point arithmetic unit 400 according to an embodiment of the present invention. Figure 4 The illustrated floating-point arithmetic unit 400 can be referred to Figure 1 or Figure 3 and analogized according to the relevant description of any one of the floating-point arithmetic units 150_1 to 150_N shown. Figure 4 The illustrated floating-point arithmetic unit 400 can be used as Figure 1 or Figure 3 one of many implementation examples of any one of the floating-point arithmetic units 150_1 to 150_N shown. In Figure 4 the illustrated embodiment, the floating-point arithmetic unit 400 includes a carry save adder (CSA) 410, an adder 420, a multiplexer 430, a multiplexer 440, an adder 450, a floating-point normalization circuit 460, a loop intermediate result memory access circuit 470, a multiplier 480, and a multiplexer 490. In this embodiment, the adder 450, the floating-point normalization circuit 460, the multiplier 480, and the multiplexer 490 can be components of a standard floating-point arithmetic unit, while the carry save adder 410, the adder 420, the multiplexer 430, the multiplexer 440, and the loop intermediate result memory access circuit 470 are added to the standard floating-point arithmetic unit to implement hardware dedicated to mean pooling operations.
[0040] The instruction decoding module 120 decodes the pooling instruction POOL.MEAN from the instruction scheduling module 110 to generate a decoding result. The pooling control circuit 130 determines the number of execution cycles of the pooling instruction and the feature points for reading the feature map values from the register bank 140 based on the decoding result of the pooling instruction POOL.MEAN, and supplies them to the carry-save adder 410. The carry-save adder 410 is coupled to the pooling control circuit 130 to receive multiple feature points of a corresponding receptive field, such as feature point IN41, feature point IN42, and feature point IN43. The carry-save adder 410 generates a sum value tmp41 and a carry value tmp42 based on the multiple feature points of the corresponding receptive field, that is, IN41 + IN42 + IN43 = tmp41 + tmp42. The adder 420 is coupled to the carry-save adder 410 to receive the sum value tmp41 and the carry value tmp42. The adder 420 generates a sum value tmp43 based on the sum value tmp41 and the carry value tmp42, that is, tmp41 + tmp42 = tmp43. The first input terminal of the multiplexer 430 is coupled to the adder 420 to receive the sum value tmp43. The output terminal of the multiplexer 430 is coupled to the first input terminal of the adder 450. The first input terminal of the multiplexer 440 is coupled to the output terminal of the loop intermediate result memory circuit 470. The loop intermediate result memory circuit 470 provides the sum value tmp44 of the previous execution cycle in multiple loops (execution cycles) of the average pooling operation. In the case of the first loop of the average pooling operation in the floating-point arithmetic unit 400, the sum value tmp44 is zero. The output terminal of the multiplexer 440 is coupled to the second input terminal of the adder 450.
[0041] When the floating-point arithmetic unit 400 is used as the hardware dedicated to the average pooling operation, the multiplexer 430 transfers the sum value tmp43 to the first input terminal of the adder 450, and the multiplexer 440 transfers the sum value tmp44 of the previous execution cycle to the second input terminal of the adder 450. When the floating-point arithmetic unit 400 is used as a standard floating-point arithmetic unit and the standard floating-point arithmetic unit performs the floating-point addition "a + b", the second input terminal of the multiplexer 430 receives the floating-point number a, and the second input terminal of the multiplexer 440 receives the floating-point number b. At this time, the multiplexer 430 transfers the floating-point number a to the first input terminal of the adder 450, and the multiplexer 440 transfers the floating-point number b to the second input terminal of the adder 450.
[0042] The input terminal of the floating-point normalization circuit 460 is coupled to the output terminal of the adder 450 to receive the sum value tmp45. The output terminal of the floating-point normalization circuit 460 is coupled to the input terminal of the loop intermediate result memory access circuit 470 and the first input terminal of the multiplexer 490. The floating-point normalization circuit 460 normalizes the sum value tmp45 based on the floating-point standard format to generate the normalized sum value tmp45' for the loop intermediate result memory access circuit 470 and the multiplexer 490.
[0043] The instruction decoding module 120 decodes the floating-point multiplication instruction FMUL regarding the average pooling operation and from the instruction scheduling module 110 to generate a decoding result. The pooling control circuit 130 reads the feature point total value SUM41 for the multiplier 480 from the register bank 140 based on the decoding result of the floating-point multiplication instruction FMUL. The first input terminal of the multiplier 480 receives the feature point total value SUM41 regarding the receptive field corresponding to the floating-point arithmetic unit 400 from the pooling control circuit 130. The second input terminal of the multiplier 480 receives the average coefficient AC41 regarding the corresponding receptive field from the pooling control circuit 130. The second input terminal of the multiplexer 490 is coupled to the output terminal of the multiplier 480.
[0044] For example, assume that the size of the corresponding receptive field is 3×3 feature points, and assume that the data elements in the corresponding receptive field are P1, P2, P3, P4, P5, P6, P7, P8, and P9. Then, the mean pooling operation for this corresponding receptive field includes summation iteration and multiplication. The summation iteration includes 3 summation loops, and the average coefficient AC41 of the corresponding receptive field is 1 / (3×3) = 1 / 9. In the first summation loop, the feature points IN41, IN42, and IN43 are P1, P2, and P3 respectively. The loop intermediate result memory circuit 470 provides the sum value tmp44 with a value of zero to the multiplexer 440, and the floating-point normalization circuit 460 writes the sum value tmp45’ (i.e., P1 + P2 + P3) of the first summation loop into the loop intermediate result memory circuit 470. In the second summation loop, the feature points IN41, IN42, and IN43 are P4, P5, and P6 respectively. The loop intermediate result memory circuit 470 provides the sum value tmp44 (i.e., P1 + P2 + P3) of the first summation loop to the multiplexer 440, and the floating-point normalization circuit 460 writes the sum value tmp45’ (i.e., P4 + P5 + P6 + tmp44 = P1 + P2 + P3 + P4 + P5 + P6) of the second summation loop into the loop intermediate result memory circuit 470. In the third summation loop, the feature points IN41, IN42, and IN43 are P7, P8, and P9 respectively. The loop intermediate result memory circuit 470 provides the sum value tmp44 (i.e., P1 + P2 + P3 + P4 + P5 + P6) of the second summation loop to the multiplexer 440, and the floating-point normalization circuit 460 writes the sum value tmp45’ (i.e., SUM41 = P7 + P8 + P9 + tmp44 = P1 + P2 + P3 + P4 + P5 + P6 + P7 + P8 + P9) of the third summation loop into the register bank 140 through the multiplexer 490. In the multiplication of the mean pooling operation, the first input terminal of the multiplier 480 receives the sum value of the feature points SUM41 (i.e., P1 + P2 + P3 + P4 + P5 + P6 + P7 + P8 + P9) from the pooling control circuit 130. The second input terminal of the multiplier 480 receives the average coefficient AC41 (e.g., 1 / 9) from the pooling control circuit 130. The multiplier 480 writes the mean pooling result (i.e., (P1 + P2 + P3 + P4 + P5 + P6 + P7 + P8 + P9) / 9) into the register bank 140 through the multiplexer 490.
[0045] In summary, in the mean pooling operation with a 3×3 receptive field, the conventional technique requires 9 instructions and 9 instruction cycles, while the above embodiments only require 2 instructions, namely POOL.MEAN and FMUL, to complete the mean pooling operation, taking a total of 4 instruction cycles. That is, the above embodiments can complete the mean pooling operation with fewer instructions and at a faster speed.
[0046] Figure 5 FIG. 5 is a schematic block diagram of a floating-point arithmetic unit 500 according to another embodiment of the present invention. Figure 5 The illustrated floating-point arithmetic unit 500 can be referred to Figure 1 or Figure 3 the relevant description of any one of the floating-point arithmetic units 150_1 to 150_N shown in FIG. and analogized therefrom. Figure 5 The illustrated floating-point arithmetic unit 500 can be used as Figure 1 or Figure 3 one of many implementation examples of any one of the floating-point arithmetic units 150_1 to 150_N shown in FIG. In Figure 5 the illustrated embodiment, the floating-point arithmetic unit 500 includes a comparator 510, a comparator 520, a multiplexer 530, a multiplexer 540, a comparator 550, and a loop intermediate result memory access circuit 560. In this embodiment, the comparator 550 can be a component of a standard floating-point arithmetic unit, and the comparator 510, the comparator 520, the multiplexer 530, the multiplexer 540, and the loop intermediate result memory access circuit 560 are added to the standard floating-point arithmetic unit to implement hardware specifically for the maximum pooling operation.
[0047] The instruction decoding module 120 decodes the pooling instruction POOL.MAX from the instruction scheduling module 110 to generate a decoding result. The pooling control circuit 130 determines the number of execution cycles of the pooling instruction based on the decoding result of the pooling instruction POOL.MAX, and reads the feature points of the feature map values from the register bank 140 to the comparator 510 and the comparator 520. The comparator 510 is coupled to the pooling control circuit 130 to receive the first part of the multiple feature points corresponding to the receptive field, such as the feature points IN51 and IN52. The comparator 510 selects the maximum feature point tmp52 from the first part. The comparator 520 is coupled to the pooling control circuit 130 to receive the second part of the multiple feature points corresponding to the receptive field, such as the feature point IN53. The comparator 520 is further coupled to the output terminal of the loop intermediate result memory access circuit 560 to receive the maximum feature point tmp51 of the previous execution cycle in multiple cycles (execution cycles) of the maximum pooling operation from the loop intermediate result memory access circuit 560. In the case of the first cycle of the maximum pooling operation performed by the floating point arithmetic unit 500, the initial value of the maximum feature point tmp51 is negative infinity to ensure that the final result does not come from the initial value of the maximum feature point tmp51. The comparator 520 selects the maximum feature point tmp53 from the second part (such as the feature point IN53) of the current execution cycle and the maximum feature point tmp51 of the previous execution cycle.
[0048] The first input terminal of the multiplexer 530 is coupled to the output terminal of the comparator 510 to receive the maximum feature point tmp52. The output terminal of the multiplexer 530 is coupled to the first input terminal of the comparator 550. The first input terminal of the multiplexer 540 is coupled to the output terminal of the comparator 520 to receive the maximum feature point tmp53. The output terminal of the multiplexer 540 is coupled to the second input terminal of the comparator 550. When the floating-point arithmetic unit 500 is used as the hardware dedicated to the max pooling operation, the multiplexer 530 transmits the maximum feature point tmp52 to the first input terminal of the comparator 550, and the multiplexer 540 transmits the maximum feature point tmp53 to the second input terminal of the comparator 550. When the floating-point arithmetic unit 500 is used as a standard floating-point arithmetic unit and the standard floating-point arithmetic unit performs a comparison operation, the second input terminal of the multiplexer 530 receives the floating-point number a, and the second input terminal of the multiplexer 540 receives the floating-point number b. At this time, the multiplexer 530 transmits the floating-point number a to the first input terminal of the comparator 550, and the multiplexer 540 transmits the floating-point number b to the second input terminal of the comparator 550. The comparator 550 selects the maximum feature point tmp54 from the maximum feature point tmp52 output by the multiplexer 530 and the maximum feature point tmp53 output by the multiplexer 540. The input terminal of the loop intermediate result memory access circuit 560 is coupled to the output terminal of the comparator 550 to receive the maximum feature point tmp54 of the current execution cycle of the max pooling operation.
[0049] For example, assume that the size of the corresponding receptive field is 3×3 feature points, and assume that the data elements in the corresponding receptive field are P1, P2, P3, P4, P5, P6, P7, P8, and P9. Then, the max pooling operation (comparison iteration) for this corresponding receptive field includes 3 comparison loops. In the first comparison loop, the feature points IN51, IN52, and IN53 are P1, P2, and P3 respectively. The loop intermediate result memory circuit 560 provides the maximum feature point tmp51 with a value of negative infinity to the comparator 520, and the comparator 550 writes the maximum feature point tmp54 of the first comparison loop (i.e., the largest of P1, P2, and P3) to the loop intermediate result memory circuit 560. In the second comparison loop, the feature points IN51, IN52, and IN53 are P4, P5, and P6 respectively. The loop intermediate result memory circuit 560 provides the maximum feature point tmp51 of the first comparison loop (i.e., the largest of P1, P2, and P3) to the comparator 520, and the comparator 550 writes the maximum feature point tmp54 of the second comparison loop (i.e., the largest of P1, P2, P3, P4, P5, and P6) to the loop intermediate result memory circuit 560. In the third comparison loop, the feature points IN51, IN52, and IN53 are P7, P8, and P9 respectively. The loop intermediate result memory circuit 560 provides the maximum feature point tmp51 of the second comparison loop (i.e., the largest of P1, P2, P3, P4, P5, and P6) to the comparator 520, and the comparator 550 writes the maximum feature point tmp54 of the third comparison loop (i.e., the largest of P1, P2, P3, P4, P5, P6, P7, P8, and P9) to the register bank 140.
[0050] In summary, in performing the max pooling operation for a 3×3 receptive field, the conventional technique requires 8 instructions and 8 instruction cycles, while the above embodiment only requires 1 instruction, namely POOL.MAX, to complete the max pooling operation, which takes a total of 3 instruction cycles. That is, the above embodiment can complete the max pooling operation with fewer instructions and at a faster speed.
[0051] Figure 6 FIG. 6 is a circuit block diagram of a floating-point arithmetic unit 600 according to another embodiment of the present invention. Figure 6 The floating-point arithmetic unit 600 shown can be referred to Figure 1 or Figure 3 the relevant description of any one of the floating-point arithmetic units 150_1 to 150_N shown in FIG. 15 and analogized accordingly. Figure 6 The floating-point arithmetic unit 600 shown can be used as Figure 1 or Figure 3 one of many implementation examples of any one of the floating-point arithmetic units 150_1 to 150_N shown in FIG. 15. InFigure 6 In the illustrated embodiment, the floating-point arithmetic unit 600 includes an average pooling circuit 610, a maximum pooling circuit 620, a loop intermediate result memory access circuit 630, and a multiplexer 640.
[0052] The average pooling circuit 610 is coupled to the pooling control circuit 130 to receive a plurality of feature points in a corresponding receptive field, such as feature point IN61, feature point IN62, and feature point IN63. The average pooling circuit 610 selectively performs an average pooling operation on the plurality of feature points in the corresponding receptive field based on the control of the pooling control circuit 130. The average pooling operation of the average pooling circuit 610 can refer to Figure 4 the relevant description of the average pooling operation in the illustrated embodiment and can be analogized. The maximum pooling circuit 620 is coupled to the pooling control circuit 130 to receive a plurality of feature points in a corresponding receptive field. The maximum pooling circuit 620 selectively performs a maximum pooling operation on the plurality of feature points in the corresponding receptive field based on the control of the pooling control circuit 130. The maximum pooling operation of the maximum pooling circuit 620 can refer to Figure 5 the relevant description of the maximum pooling operation in the illustrated embodiment and can be analogized.
[0053] The loop intermediate result memory access circuit 630 is coupled to the average pooling circuit 610 to receive the loop intermediate result tmp61 of the current execution cycle in multiple cycles (execution periods) of the average pooling operation. The loop intermediate result memory access circuit 630 is coupled to the average pooling circuit 610 to provide the loop intermediate result tmp62 of the previous execution cycle in the average pooling operation. The loop intermediate result memory access circuit 630 is coupled to the maximum pooling circuit 620 to receive the loop intermediate result tmp63 of the current execution cycle in multiple cycles (execution periods) of the maximum pooling operation. The loop intermediate result memory access circuit 630 is coupled to the maximum pooling circuit 620 to provide the loop intermediate result tmp64 of the previous execution cycle in the maximum pooling operation.
[0054] The first input terminal of the multiplexer 640 is coupled to the output terminal of the average pooling circuit 610. The second input terminal of the multiplexer 640 is coupled to the output terminal of the maximum pooling circuit 620. If the pooling operation in response to the pooling instruction is an average pooling operation, the maximum pooling circuit 620 is disabled, and the multiplexer 640 writes the average pooling result of the average pooling circuit 610 into the register bank 140. If the pooling operation in response to the pooling instruction is a maximum pooling operation, the average pooling circuit 610 is disabled, and the multiplexer 640 writes the maximum pooling result of the maximum pooling circuit 620 into the register bank 140.
[0055] Figure 7It is a schematic block diagram of the average pooling circuit 610 and the maximum pooling circuit 620 shown according to an embodiment of the present invention. Figure 7 The average pooling circuit 610, the maximum pooling circuit 620, the loop intermediate result memory access circuit 630, and the multiplexer 640 shown can be referred to Figure 6 for the relevant descriptions, so they will not be elaborated here. Figure 7 The average pooling circuit 610 and the maximum pooling circuit 620 shown can be used as Figure 6 one of many implementation examples of the average pooling circuit 610 and the maximum pooling circuit 620 shown.
[0056] In Figure 7 the embodiment shown, the average pooling circuit 610 includes a carry-save adder 611, an adder 612, a multiplexer 613, a multiplexer 614, an adder 615, a floating-point normalization circuit 616, a multiplier 617, and a multiplexer 618. The carry-save adder 611 is coupled to the pooling control circuit 130 to receive multiple feature points corresponding to the receptive field, such as feature point IN61, feature point IN62, and feature point IN63. The carry-save adder 611 generates a sum value tmp71 and a carry value tmp72 based on the feature point IN61, the feature point IN62, and the feature point IN63. The adder 612 is coupled to the carry-save adder 611 to receive the sum value tmp71 and the carry value tmp72. The adder 612 generates a sum value tmp73 based on the sum value tmp71 and the carry value tmp72. The first input terminal of the multiplexer 613 is coupled to the adder 612 to receive the sum value tmp73. The first input terminal of the adder 615 is coupled to the output terminal of the multiplexer 613. The second input terminal of the adder 615 is coupled to the output terminal of the multiplexer 614. The input terminal of the floating-point normalization circuit 616 is coupled to the output terminal of the adder 615 to receive the sum value tmp75. The first input terminal of the multiplier 617 receives the total sum SUM41 of the feature points of the corresponding receptive field from the pooling control circuit 130. The second input terminal of the multiplier 617 receives the average coefficient AC41 of the corresponding receptive field from the pooling control circuit 130. The first input terminal of the multiplexer 618 is coupled to the output terminal of the floating-point normalization circuit 616. The second input terminal of the multiplexer 618 is coupled to the output terminal of the multiplier 617. The output terminal of the multiplexer 618 is coupled to the first input terminal of the multiplexer 640.
[0057] The loop intermediate result memory access circuit 630 is coupled to the output terminal of the floating-point normalization circuit 616 to receive the loop intermediate result tmp61 of the current execution cycle in multiple cycles (execution periods) of the average pooling operation. The loop intermediate result memory access circuit 630 is coupled to the first input terminal of the multiplexer 614 to provide the loop intermediate result tmp62 of the previous execution cycle of the average pooling operation.Figure 7 The carry-save adder 611, adder 612, multiplexer 613, multiplexer 614, adder 615, floating-point normalization circuit 616, multiplier 617, multiplexer 618, and loop intermediate result memory access circuit 630 shown can be referred to Figure 4 the relevant descriptions of the carry-save adder 410, adder 420, multiplexer 430, multiplexer 440, adder 450, floating-point normalization circuit 460, multiplier 480, multiplexer 490, and loop intermediate result memory access circuit 470 shown and analogized, so they will not be elaborated here.
[0058] In Figure 7 the embodiment shown, the maximum pooling circuit 620 includes a comparator 621, a comparator 622, a multiplexer 623, a multiplexer 624, and a comparator 625. The comparator 621 is coupled to the pooling control circuit 130 to receive a first portion of a plurality of feature points in the corresponding receptive field, such as the feature points IN61 and IN62. The comparator 621 selects the maximum feature point tmp76 from the first portion. The comparator 622 is coupled to the pooling control circuit 130 to receive a second portion of a plurality of feature points in the corresponding receptive field, such as the feature point IN63. The comparator 622 is further coupled to the loop intermediate result memory access circuit 630 to receive the maximum feature point (loop intermediate result tmp64) of the previous execution cycle in multiple cycles (execution periods) of the maximum pooling operation. In the case of the first cycle of the maximum pooling operation performed by the floating-point arithmetic unit 600, the loop intermediate result tmp64 is zero. The comparator 622 selects the maximum feature point tmp77 from the second portion (such as the feature point IN63) of the current execution cycle and the maximum feature point (loop intermediate result tmp64) of the previous execution cycle.
[0059] The first input terminal of the multiplexer 623 is coupled to the output terminal of the comparator 621 to receive the maximum feature point tmp76. The first input terminal of the multiplexer 624 is coupled to the output terminal of the comparator 622 to receive the maximum feature point tmp77. The first input terminal of the comparator 625 is coupled to the output terminal of the multiplexer 623. The second input terminal of the comparator 625 is coupled to the output terminal of the multiplexer 624. The comparator 625 selects the maximum feature point (the intermediate result tmp63 of the loop) from the output of the multiplexer 623 and the output of the multiplexer 624. The intermediate result memory circuit 630 of the loop is coupled to the output terminal of the comparator 625 to receive the intermediate result tmp63 of the current execution cycle in multiple execution cycles (execution periods) of the maximum pooling operation. The intermediate result memory circuit 630 of the loop is coupled to the comparator 622 to provide the intermediate result tmp64 of the previous execution cycle of the maximum pooling operation. The comparator 622 selects the maximum feature point tmp77 from the second part of the current execution cycle (such as the feature point IN63) and the intermediate result tmp64 of the previous execution cycle of the maximum pooling operation. Figure 7 The comparators 621, 622, multiplexers 623, 624, comparator 625, and the intermediate result memory circuit 630 of the loop shown can be referred to Figure 5 the relevant descriptions of the comparators 510, 520, multiplexers 530, 540, comparator 550, and the intermediate result memory circuit 560 shown and analogized, so details are not repeated here.
[0060] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An artificial intelligence chip, characterized in that, The artificial intelligence chip includes: An instruction scheduling module; An instruction decoding module, coupled to the instruction scheduling module to receive pooling instructions, wherein the instruction decoding module decodes the pooling instructions to generate a decoding result; A register bank; A pooling control circuit, coupled to the instruction decoding module and the register bank; and A first floating-point arithmetic unit, coupled to the pooling control circuit, wherein The pooling control circuit determines the number of execution cycles of the pooling instruction based on the decoding result, and provides the number of execution cycles to the first floating-point arithmetic unit; The pooling control circuit reads the feature map values from the register bank based on the decoding result, and provides multiple feature points of a first receptive field of the feature map values to the first floating-point arithmetic unit; and The first floating-point arithmetic unit performs a pooling operation on the multiple feature points of the first receptive field based on the number of execution cycles to generate a pooling result of the first receptive field.
2. The artificial intelligence chip according to claim 1, characterized in that, The artificial intelligence chip further includes: A second floating-point arithmetic unit, coupled to the pooling control circuit, wherein The pooling control circuit provides the number of execution cycles to the second floating-point arithmetic unit; The pooling control circuit provides multiple feature points of a second receptive field of the feature map values to the second floating-point arithmetic unit based on the decoding result; and The second floating-point arithmetic unit performs a pooling operation on the multiple feature points of the second receptive field based on the number of execution cycles to generate a pooling result of the second receptive field.
3. The artificial intelligence chip according to claim 1, wherein, The pooling control circuit includes: A pooling operation control module, coupled to the instruction decoding module to receive the decoding result, wherein the pooling operation control module determines the number of execution cycles of the pooling instruction based on the decoding result, and provides the number of execution cycles to the first floating-point arithmetic unit; An operand acquisition module, coupled to the pooling operation control module and the register bank, wherein the operand acquisition module reads the feature map values from the register bank based on the control of the pooling operation control module; and A feature map input control module, coupled to the operand acquisition module to receive the feature map values, wherein the feature map input control module provides multiple feature points of the first receptive field of the feature map values to the first floating-point arithmetic unit.
4. The artificial intelligence chip according to claim 1, wherein The first floating-point arithmetic unit includes: A carry-save adder, coupled to the pooling control circuit to receive the multiple feature points of the first receptive field, wherein the carry-save adder generates a first sum value and a carry value based on the multiple feature points of the first receptive field; A first adder, coupled to the carry-save adder to receive the first sum value and the carry value, wherein the first adder generates a second sum value based on the first sum value and the carry value; A first multiplexer, wherein a first input terminal of the first multiplexer is coupled to the first adder to receive the second sum value; A second multiplexer; A second adder, wherein a first input terminal of the second adder is coupled to an output terminal of the first multiplexer, and a second input terminal of the second adder is coupled to an output terminal of the second multiplexer; A floating-point normalization circuit, wherein an input terminal of the floating-point normalization circuit is coupled to an output terminal of the second adder; A loop intermediate result memory access circuit, wherein an input terminal of the loop intermediate result memory access circuit is coupled to an output terminal of the floating-point normalization circuit, and an output terminal of the loop intermediate result memory access circuit is coupled to a first input terminal of the second multiplexer; A multiplier, wherein a first input terminal of the multiplier receives a sum value of feature points of the first receptive field from the pooling control circuit, and a second input terminal of the multiplier receives an average coefficient of the first receptive field from the pooling control circuit; and A third multiplexer, wherein a first input terminal of the third multiplexer is coupled to an output terminal of the floating-point normalization circuit, and a second input terminal of the third multiplexer is coupled to an output terminal of the multiplier.
5. The artificial intelligence chip according to claim 1, wherein The first floating-point arithmetic unit includes: A first comparator, coupled to the pooling control circuit to receive a first part of a plurality of feature points of the first receptive field, wherein the first comparator selects a first maximum feature point from the first part; A second comparator, coupled to the pooling control circuit to receive a second part of a plurality of feature points of the first receptive field, wherein the second comparator outputs a second maximum feature point; A first multiplexer, wherein a first input terminal of the first multiplexer is coupled to an output terminal of the first comparator to receive the first maximum feature point; A second multiplexer, wherein a first input terminal of the second multiplexer is coupled to an output terminal of the second comparator; A third comparator, wherein a first input terminal of the third comparator is coupled to an output terminal of the first multiplexer, a second input terminal of the third comparator is coupled to an output terminal of the second multiplexer, and the third comparator selects a third maximum feature point from the output of the first multiplexer and the output of the second multiplexer; and A loop intermediate result memory access circuit, wherein an input terminal of the loop intermediate result memory access circuit is coupled to an output terminal of the third comparator, and an output terminal of the loop intermediate result memory access circuit is coupled to the second comparator to provide the third maximum feature point of the previous execution cycle; Wherein the second comparator selects the second maximum feature point from the second part of the current execution cycle and the third maximum feature point of the previous execution cycle.
6. The artificial intelligence chip according to claim 1, wherein The first floating-point arithmetic unit includes: An average pooling circuit, coupled to the pooling control circuit to receive a plurality of feature points of the first receptive field, wherein the average pooling circuit selectively performs an average pooling operation on the plurality of feature points of the first receptive field based on the control of the pooling control circuit; A maximum pooling circuit, coupled to the pooling control circuit to receive multiple feature points of the first receptive field, wherein the maximum pooling circuit selectively performs a maximum pooling operation on the multiple feature points of the first receptive field based on the control of the pooling control circuit; A loop intermediate result memory access circuit, wherein the loop intermediate result memory access circuit is coupled to the average pooling circuit to receive the loop intermediate result of the average pooling operation in the current execution cycle, the loop intermediate result memory access circuit is coupled to the average pooling circuit to provide the loop intermediate result of the average pooling operation in the previous execution cycle, the loop intermediate result memory access circuit is coupled to the maximum pooling circuit to receive the loop intermediate result of the maximum pooling operation in the current execution cycle, and the loop intermediate result memory access circuit is coupled to the maximum pooling circuit to provide the loop intermediate result of the maximum pooling operation in the previous execution cycle; and A first multiplexer, wherein a first input terminal of the first multiplexer is coupled to an output terminal of the average pooling circuit, and a second input terminal of the first multiplexer is coupled to an output terminal of the maximum pooling circuit.
7. The artificial intelligence chip according to claim 6, wherein The average pooling circuit includes: A carry-save adder, coupled to the pooling control circuit to receive multiple feature points of the first receptive field, wherein the carry-save adder generates a first sum value and a carry value based on the multiple feature points of the first receptive field; A first adder, coupled to the carry-save adder to receive the first sum value and the carry value, wherein the first adder generates a second sum value based on the first sum value and the carry value; A second multiplexer, wherein a first input terminal of the second multiplexer is coupled to the first adder to receive the second sum value; A third multiplexer; A second adder, wherein a first input terminal of the second adder is coupled to an output terminal of the second multiplexer, and a second input terminal of the second adder is coupled to an output terminal of the third multiplexer; A floating-point normalization circuit, wherein an input terminal of the floating-point normalization circuit is coupled to an output terminal of the second adder; A multiplier, wherein a first input terminal of the multiplier receives the total sum of the feature points of the first receptive field from the pooling control circuit, and a second input terminal of the multiplier receives the average coefficient of the first receptive field from the pooling control circuit; and A fourth multiplexer, wherein a first input terminal of the fourth multiplexer is coupled to an output terminal of the floating-point normalization circuit, a second input terminal of the fourth multiplexer is coupled to an output terminal of the multiplier, and an output terminal of the fourth multiplexer is coupled to the first input terminal of the first multiplexer; The loop intermediate result memory access circuit is coupled to the output of the floating-point normalization circuit to receive the loop intermediate result of the mean pooling operation in the current execution cycle, and the loop intermediate result memory access circuit is coupled to the first input terminal of the third multiplexer to provide the loop intermediate result of the mean pooling operation in the previous execution cycle.
8. The artificial intelligence chip according to claim 6, wherein The maximum pooling circuit includes: A first comparator, coupled to the pooling control circuit to receive a first part of a plurality of feature points in the first receptive field, wherein the first comparator selects a first maximum feature point from the first part; A second comparator, coupled to the pooling control circuit to receive a second part of a plurality of feature points in the first receptive field, wherein the second comparator outputs a second maximum feature point; A second multiplexer, wherein a first input terminal of the second multiplexer is coupled to the output of the first comparator to receive the first maximum feature point; A third multiplexer, wherein a first input terminal of the third multiplexer is coupled to the output of the second comparator; and A third comparator, wherein a first input terminal of the third comparator is coupled to the output of the second multiplexer, a second input terminal of the third comparator is coupled to the output of the third multiplexer, and the third comparator selects a third maximum feature point from the output of the second multiplexer and the output of the third multiplexer; The loop intermediate result memory access circuit is coupled to the output of the third comparator to receive the loop intermediate result of the maximum pooling operation in the current execution cycle, and the loop intermediate result memory access circuit is coupled to the second comparator to provide the loop intermediate result of the maximum pooling operation in the previous execution cycle; The second comparator selects the second maximum feature point from the second part in the current execution cycle and the loop intermediate result of the maximum pooling operation in the previous execution cycle.
9. A pooling method for an artificial intelligence chip, characterized in that, The pooling method includes: Decoding a pooling instruction by an instruction decoding module of the artificial intelligence chip to generate a decoding result, wherein the instruction decoding module is coupled to an instruction scheduling module of the artificial intelligence chip to receive the pooling instruction; Determining, by a pooling control circuit of the artificial intelligence chip, the number of execution cycles of the pooling instruction based on the decoding result, wherein the pooling control circuit is coupled to the instruction decoding module, a register bank of the artificial intelligence chip, and a first floating-point arithmetic unit of the artificial intelligence chip; Providing, by the pooling control circuit, the number of execution cycles to the first floating-point arithmetic unit; Reading, by the pooling control circuit, feature map values from the register bank based on the decoding result; Providing, by the pooling control circuit, a plurality of feature points in the first receptive field of the feature map values to the first floating-point arithmetic unit; and Performing, by the first floating-point arithmetic unit, a pooling operation on the plurality of feature points in the first receptive field based on the number of execution cycles to generate a pooling result of the first receptive field.
10. The pooling method according to claim 9, wherein The pooling method further includes: The pooling control circuit provides the number of execution cycles to a second floating-point arithmetic unit of the artificial intelligence chip, wherein the second floating-point arithmetic unit is coupled to the pooling control circuit; The pooling control circuit provides multiple feature points of a second receptive field of the feature map values to the second floating-point arithmetic unit based on the decoding result; and The second floating-point arithmetic unit performs a pooling operation on the multiple feature points of the second receptive field based on the number of execution cycles to generate a pooling result of the second receptive field.
11. The pooling method according to claim 9, wherein The pooling method further includes: A pooling operation control module of the pooling control circuit determines the number of execution cycles of the pooling instruction based on the decoding result, wherein the pooling operation control module is coupled to the instruction decoding module to receive the decoding result; The pooling operation control module provides the number of execution cycles to the first floating-point arithmetic unit; An operand acquisition module of the pooling control circuit reads the feature map values from the register bank based on the control of the pooling operation control module, wherein the operand acquisition module is coupled to the pooling operation control module and the register bank; and A feature map input control module of the pooling control circuit provides multiple feature points of the first receptive field of the feature map values to the first floating-point arithmetic unit, wherein the feature map input control module is coupled to the operand acquisition module to receive the feature map values.
Citation Information
Patent Citations
On-chip architecture, pooling computing accelerator array, unit and control method
CN112905530A
Memristor-based programmable neural network accelerator
CN113869504A