Matrix operation array system and operation method for deep learning acceleration kernel
Through the combined design of PE array and pipeline control unit, the problems of large area and high power consumption in deep learning acceleration core hardware are solved, flexible multi-data type computing is realized, and system performance is improved.
Patent Information
- Application Number
- CN202510394576.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
AI Technical Summary
The existing deep learning accelerated core hardware design has problems such as large area, high power consumption and poor flexibility, and cannot support flexible computing of multiple data types.
The combined design of PE array, input control unit, pipeline control unit and output control unit is adopted to support fixed-point and floating-point operations, and the working status of the PE module in each operation cycle is controlled through pipeline processing technology. It supports INT8, UINT8, INT16, UINT16, FP16, BF16 data types, and designs the FP19 data format to multiplex floating-point adder.
It effectively reduces power consumption, reduces circuit area, improves system performance and flexibility, and supports flexible computing of various data types.
Smart Images

Figure CN120336688A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a deep learning acceleration core, and specifically to a matrix operation array system and operation method for a deep learning acceleration core. Background Art
[0002] A deep learning acceleration core is a dedicated high-performance hardware architecture in the field of artificial intelligence, which supports large-scale matrix parallel operations. Compared with general-purpose processors, it can complete the same tasks with lower power consumption.
[0003] A deep learning acceleration core generally consists of a control unit, a matrix operation array (MAC), on-chip memory close to the matrix operation array, and an on-chip bus. Among them, the matrix operation array, as the operation core of the deep learning acceleration core, is composed of numerous PE modules, and supports calculation modes such as neural network inference operations, matrix multiplication operations, and two-dimensional convolution operations for fixed-point and floating-point data types. The matrix operation array receives the input feature map in the row direction and the convolution kernel weights in the column direction, multiplies them, and accumulates the results to the accumulation register.
[0004] When it is necessary to meet the operations of different data types, the existing hardware designs for each data type are usually implemented separately, which have the disadvantages of large occupied area and high power consumption, and cannot support multiple data types. In addition, each sub-module unit cannot be arbitrarily parallel, with poor flexibility and unable to be dynamically configured. Summary of the Invention
[0005] (1) Technical Problems to be Solved
[0006] In view of the above-mentioned disadvantages existing in the prior art, the present invention provides a matrix operation array system and operation method for a deep learning acceleration core, which can effectively overcome the defects of large occupied area, high power consumption, and poor flexibility existing in the prior art.
[0007] (2) Technical Solutions
[0008] To achieve the above object, the present invention is realized through the following technical solutions:
[0009] A matrix operation array system for a deep learning acceleration core includes a PE array, an input control unit, a pipeline control unit, and an output control unit;
[0010] The PE array, as the operation core of the matrix operation array, is composed of multiple PE modules. Each PE module can perform fixed-point and floating-point multiplication and accumulation operations, and process the result overflow generated during the addition process;
[0011] The input control unit distributes the input data from the memory as operands to the PE modules, and controls the work of the PE modules by sending work enable signals to the PE modules;
[0012] The pipeline control unit uses pipeline processing technology to control the working state of the PE module in each operation cycle;
[0013] The output control unit receives the operation results of the PE module, and after receiving the output enable, processes the operation results of the PE module according to the output mode and then outputs them to the external memory.
[0014] Preferably, the PE module includes a fixed-point arithmetic unit and a floating-point arithmetic unit;
[0015] The fixed-point arithmetic unit performs fixed-point multiplication and accumulation operations, and processes the result overflow generated during the addition process;
[0016] The floating-point arithmetic unit performs floating-point multiplication and accumulation operations.
[0017] Preferably, the fixed-point arithmetic unit includes a fixed-point multiplier, an accumulator, and a saturation processing unit;
[0018] The fixed-point multiplier performs fixed-point multiplication operations;
[0019] The accumulator accumulates the multiplication operation results of the fixed-point multiplier to the local fixed-point register;
[0020] The saturation processing unit processes the result overflow generated during the addition process;
[0021] The floating-point arithmetic unit includes a floating-point multiplier and a floating-point adder;
[0022] The floating-point multiplier performs floating-point multiplication operations;
[0023] The floating-point adder accumulates the multiplication operation results of the floating-point multiplier to the local floating-point register.
[0024] Preferably, the input control unit includes a data distribution unit and an enable control unit;
[0025] The data distribution unit distributes the input data from the memory as operands to the PE module;
[0026] The enable control unit controls the operation of the PE module by sending a work enable to the PE module;
[0027] Among them, multiple PE modules form a 16*16 PE array. The data distribution unit distributes 4096-bit input data from the memory to 256 PE modules. The lower 2048-bit input data is distributed to 16 rows of PE modules at 128 bits per row, and the higher 2048-bit input data is distributed to 16 columns of PE modules at 128 bits per column.
[0028] Preferably, the pipeline control unit uses pipeline processing technology to control the working state of the PE module in each operation cycle, including:
[0029] Using pipeline processing technology to control the PE module to perform the following operations in 11 operation cycles:
[0030] cycle0: The PE module preprocesses the received input data and sends it to the port of the multiplier;
[0031] cycle1: Perform multiplication operations to obtain 8-bit and 16-bit multiplication results;
[0032] cycle2: For fixed-point operations, sum the fixed-point multiplication results through a fixed-point adder tree and accumulate them with the values in the local fixed-point register to obtain the fixed-point operation result; for floating-point operations, perform post-processing according to the floating-point multiplication result to obtain 8 floating-point multiplication results;
[0033] cycle3: The first beat operation of the first-level floating-point adder tree;
[0034] cycle4: The second beat operation of the first-level floating-point adder tree to obtain the operation result of the first-level floating-point adder tree;
[0035] cycle5: The first beat operation of the second-level floating-point adder tree;
[0036] cycle6: The second beat operation of the second-level floating-point adder tree to obtain the operation result of the second-level floating-point adder tree;
[0037] cycle7: The first beat operation of the third-level floating-point adder tree;
[0038] cycle8: The second beat operation of the third-level floating-point adder tree to obtain the operation result of the third-level floating-point adder tree;
[0039] cycle9: Accumulate the sum of the 8 floating-point multiplication results with the values in the local floating-point register;
[0040] cycle10: The second beat operation of the floating-point accumulation operation to obtain the floating-point operation result.
[0041] Preferably, the PE module supports multiplication and accumulation operations of data types INT8, UINT8, INT16, UINT16, FP16, and BF16. To support the reuse of the floating-point adder, a new data format of FP19 is designed. When performing floating-point addition operations, first convert the data formats of FP16 and BF16 into the FP19 data format, and then perform the floating-point addition algorithm;
[0042] Among them, the data format of FP16 is 1 sign bit, 5 exponent bits, and 10 mantissa bits; the data format of BF16 is 1 sign bit, 8 exponent bits, and 7 mantissa bits; the data format of FP19 is 1 sign bit, 8 exponent bits, and 10 mantissa bits.
[0043] An operation method for a matrix operation array of a deep learning acceleration core includes the following steps:
[0044] S1. After receiving an instruction from the upper layer, the input control unit distributes the input data from the memory as operands to the PE modules and controls the PE modules to work by sending a work enable to the PE modules.
[0045] S2. The PE modules perform fixed-point and floating-point multiplication and accumulation operations and handle the result overflow generated during the addition process.
[0046] S3. The output control unit receives the operation results of the PE modules and, after receiving the output enable, processes the operation results of the PE modules according to the output mode and then outputs them to the external memory.
[0047] Among them, during the process of the PE modules performing fixed-point and floating-point multiplication and accumulation operations, the pipeline control unit uses pipeline processing technology to control the working state of the PE modules in each operation cycle.
[0048] Preferably, the pipeline control unit uses pipeline processing technology to control the working state of the PE modules in each operation cycle, including:
[0049] Using pipeline processing technology to control the PE modules to perform the following operations in 11 operation cycles:
[0050] cycle0: The PE module preprocesses the received input data and sends it to the port of the multiplier.
[0051] cycle1: Perform a multiplication operation to obtain 8-bit and 16-bit multiplication operation results.
[0052] cycle2: For fixed-point operations, sum the fixed-point multiplication operation results through a fixed-point adder tree and accumulate them with the values in the local fixed-point register to obtain fixed-point operation results; for floating-point operations, perform post-processing based on the floating-point multiplication operation results to obtain 8 floating-point multiplication operation results.
[0053] cycle3: The first beat operation of the first-level floating-point adder tree.
[0054] cycle4: The second beat operation of the first-level floating-point adder tree to obtain the operation result of the first-level floating-point adder tree.
[0055] cycle5: The first - beat operation of the second - level floating - point adder tree;
[0056] cycle6: The second - beat operation of the second - level floating - point adder tree to obtain the operation result of the second - level floating - point adder tree;
[0057] cycle7: The first - beat operation of the third - level floating - point adder tree;
[0058] cycle8: The second - beat operation of the third - level floating - point adder tree to obtain the operation result of the third - level floating - point adder tree;
[0059] cycle9: Accumulate the sum of the results of 8 floating - point multiplication operations and the value in the local floating - point register;
[0060] cycle10: The second - beat operation of the floating - point accumulation operation to obtain the floating - point operation result.
[0061] (III) Advantageous Effects
[0062] Compared with the prior art, the matrix operation array system and operation method for deep - learning acceleration kernels provided by the present invention have the following advantageous effects:
[0063] 1) The input control unit controls the PE module to work by sending a work enable signal to the PE module, so that each PE module is under enable control, effectively reducing power consumption while the computing power is scalable;
[0064] 2) Through hardware reuse, power consumption can be effectively reduced and circuit area can be reduced;
[0065] 3) Adopting a modular design enables flexible expansion and maintenance;
[0066] 4) The pipeline control unit uses pipeline - type processing technology to control the working state of the PE module in each operation cycle. Through pipeline - type processing, the critical path of the system can be shortened, effectively improving the system performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0068] Figure 1 It is a schematic diagram of the system of the present invention;
[0069] Figure 2Schematic diagram showing that the pipeline control unit in the present invention uses pipeline processing technology to control the working state of the PE module in each operation cycle;
[0070] Figure 3 Hardware structure diagram of the fixed-point arithmetic unit in the present invention;
[0071] Figure 4 Schematic diagram of the data formats of FP16, BF16, and FP19 in the present invention;
[0072] Figure 5 Schematic flow diagram of the present invention. Detailed implementation manners
[0073] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0074] A matrix operation array system for a deep learning acceleration core, as Figure 1 shown, includes a PE array, an input control unit, a pipeline control unit, and an output control unit;
[0075] The PE array, which is the operation core of the matrix operation array, is composed of multiple PE modules. Each PE module can perform fixed-point and floating-point multiplication and accumulation operations and handle the result overflow generated during the addition process;
[0076] The input control unit distributes the input data from the memory as operands to the PE modules and controls the operation of the PE modules by sending a work enable signal to the PE modules;
[0077] The pipeline control unit uses pipeline processing technology to control the working state of the PE modules in each operation cycle;
[0078] The output control unit receives the operation results of the PE modules and processes the operation results of the PE modules according to the output mode and then outputs them to the external memory after receiving the output enable.
[0079] ① The PE module includes a fixed-point arithmetic unit and a floating-point arithmetic unit;
[0080] The fixed-point arithmetic unit performs fixed-point multiplication and accumulation operations and handles the result overflow generated during the addition process;
[0081] The floating-point arithmetic unit performs floating-point multiplication and accumulation operations.
[0082] Specifically, 1) the fixed-point arithmetic unit includes a fixed-point multiplier, an accumulator, and a saturation processing unit;
[0083] The fixed-point multiplier performs fixed-point multiplication operations;
[0084] The accumulator accumulates the multiplication operation results of the fixed-point multiplier into the local fixed-point register;
[0085] The saturation processing unit processes the result overflow generated during the addition process;
[0086] 2) The floating-point arithmetic unit includes a floating-point multiplier and a floating-point adder;
[0087] The floating-point multiplier performs floating-point multiplication operations;
[0088] The floating-point adder accumulates the multiplication operation results of the floating-point multiplier into the local floating-point register.
[0089] Figure 3 is the hardware structure diagram of the fixed-point arithmetic unit. The input data is allocated to different multiplier ports according to the data type, enters the adder tree stage after completing the multiplication operation, and finally accumulates into the local ACC register.
[0090] ② The input control unit includes a data distribution unit and an enable control unit;
[0091] The data distribution unit distributes the input data from the memory as operands to the PE modules;
[0092] The enable control unit controls the operation of the PE modules by sending a work enable to the PE modules;
[0093] Among them, multiple PE modules form a 16*16 PE array. The data distribution unit distributes 4096-bit input data from the memory to 256 PE modules. The lower 2048-bit input data is distributed to 16 rows of PE modules at 128 bits per row, and the upper 2048-bit input data is distributed to 16 columns of PE modules at 128 bits per column.
[0094] ③ The pipeline control unit uses pipeline processing technology to control the working state of the PE modules in each operation cycle, including:
[0095] Such as Figure 2 shown, pipeline processing technology is used to control the PE modules to perform the following 11 operation cycles of operations:
[0096] cycle0: The PE module preprocesses the received input data and sends it to the port of the multiplier;
[0097] cycle1: Perform multiplication operations to obtain 8-bit and 16-bit multiplication results;
[0098] cycle2: For fixed-point operations, sum the fixed-point multiplication results through a fixed-point adder tree and accumulate them with the values in the local fixed-point register to obtain the fixed-point operation result; for floating-point operations, perform post-processing based on the floating-point multiplication results to obtain 8 floating-point multiplication results;
[0099] cycle3: The first beat operation of the first-level floating-point adder tree;
[0100] cycle4: The second beat operation of the first-level floating-point adder tree to obtain the operation result of the first-level floating-point adder tree;
[0101] cycle5: The first beat operation of the second-level floating-point adder tree;
[0102] cycle6: The second beat operation of the second-level floating-point adder tree to obtain the operation result of the second-level floating-point adder tree;
[0103] cycle7: The first beat operation of the third-level floating-point adder tree;
[0104] cycle8: The second beat operation of the third-level floating-point adder tree to obtain the operation result of the third-level floating-point adder tree;
[0105] cycle9: Accumulate the sum of the 8 floating-point multiplication results with the values in the local floating-point register;
[0106] cycle10: The second beat operation of the floating-point accumulation operation to obtain the floating-point operation result.
[0107] In the technical solution of this application, the PE module supports multiplication and accumulation operations for data types INT8, UINT8, INT16, UINT16, FP16, and BF16. To support the reuse of the floating-point adder, a new data format with the data type FP19 is designed. When performing floating-point addition operations, first convert the data formats of FP16 and BF16 into the FP19 data format, and then perform the floating-point addition algorithm, which can effectively reduce the circuit area. At the same time, since the floating-point data bit width is widened, the precision of the floating-point operation can also be improved;
[0108] Among them, as Figure 4 shown, the data format of FP16 is 1 sign bit, 5 exponent bits, and 10 mantissa bits; the data format of BF16 is 1 sign bit, 8 exponent bits, and 7 mantissa bits; the data format of FP19 is 1 sign bit, 8 exponent bits, and 10 mantissa bits.
[0109] Based on the matrix operation array system for deep learning acceleration cores disclosed above in the technical solution of this application, a method for operating a matrix operation array for deep learning acceleration cores is also disclosed. As Figure 5 shown, it includes the following steps:
[0110] S1. After receiving an instruction from the upper layer, the input control unit distributes the input data from the memory as operands to the PE modules, and controls the PE modules to work by sending a work enable to the PE modules.
[0111] S2. The PE modules perform fixed-point and floating-point multiplication and accumulation operations, and handle the result overflow generated during the addition process.
[0112] S3. The output control unit receives the operation results of the PE modules, and after receiving the output enable, processes the operation results of the PE modules according to the output method and then outputs them to the external memory.
[0113] Among them, during the process of the PE modules performing fixed-point and floating-point multiplication and accumulation operations, the pipeline control unit uses pipeline processing technology to control the working state of the PE modules in each operation cycle.
[0114] Specifically, the pipeline control unit uses pipeline processing technology to control the working state of the PE modules in each operation cycle, including:
[0115] Using pipeline processing technology to control the PE modules to perform the following operations in 11 operation cycles:
[0116] cycle0: The PE modules preprocess the received input data and send it to the ports of the multipliers.
[0117] cycle1: Perform multiplication operations to obtain 8-bit and 16-bit multiplication operation results.
[0118] cycle2: For fixed-point operations, sum the fixed-point multiplication operation results through a fixed-point adder tree and accumulate them with the values in the local fixed-point registers to obtain fixed-point operation results; for floating-point operations, perform post-processing based on the floating-point multiplication operation results to obtain 8 floating-point multiplication operation results.
[0119] cycle3: The first beat operation of the first-level floating-point adder tree.
[0120] cycle4: The second beat operation of the first-level floating-point adder tree to obtain the operation result of the first-level floating-point adder tree.
[0121] cycle5: The first beat operation of the second-level floating-point adder tree.
[0122] cycle6: The second beat operation of the second-level floating-point adder tree to obtain the operation result of the second-level floating-point adder tree;
[0123] cycle7: The first beat operation of the third-level floating-point adder tree;
[0124] cycle8: The second beat operation of the third-level floating-point adder tree to obtain the operation result of the third-level floating-point adder tree;
[0125] cycle9: Accumulate the sum of the results of 8 floating-point multiplication operations and the value in the local floating-point register;
[0126] cycle10: The second beat operation of the floating-point accumulation operation to obtain the floating-point operation result.
[0127] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. Matrix operation array system for deep learning acceleration kernel, characterized in that: It includes a PE array, an input control unit, a pipeline control unit, and an output control unit; The PE array is the operation core of the matrix operation array, which is composed of multiple PE modules. Each PE module can perform fixed-point and floating-point multiplication and accumulation operations, and handle the result overflow generated during the addition process; The input control unit distributes the input data from the memory as operands to the PE modules, and controls the operation of the PE modules by sending a work enable to the PE modules; The pipeline control unit uses pipeline processing technology to control the working state of the PE modules in each operation cycle; The output control unit receives the operation results of the PE modules, and after receiving the output enable, processes the operation results of the PE modules according to the output mode and then outputs them to the external memory.
2. The matrix operation array system for deep learning acceleration cores according to claim 1, wherein: The PE module includes a fixed-point operation unit and a floating-point operation unit; The fixed-point operation unit performs fixed-point multiplication and accumulation operations, and handles the result overflow generated during the addition process; The floating-point operation unit performs floating-point multiplication and accumulation operations.
3. The matrix operation array system for deep learning acceleration kernel according to claim 2, characterized in that: The fixed-point operation unit includes a fixed-point multiplier, an accumulator, and a saturation processing unit; The fixed-point multiplier performs fixed-point multiplication operations; The accumulator accumulates the multiplication operation results of the fixed-point multiplier to the local fixed-point register; The saturation processing unit handles the result overflow generated during the addition process; The floating-point operation unit includes a floating-point multiplier and a floating-point adder; The floating-point multiplier performs floating-point multiplication operations; The floating-point adder accumulates the multiplication operation results of the floating-point multiplier to the local floating-point register.
4. The matrix operation array system for deep learning acceleration cores according to claim 2, wherein: The input control unit includes a data distribution unit and an enable control unit; The data distribution unit distributes the input data from the memory as operands to the PE modules; The enable control unit controls the operation of the PE modules by sending a work enable to the PE modules; Among them, multiple PE modules form a 16*16 PE array. The data distribution unit distributes 4096-bit input data from the memory to 256 PE modules. The lower 2048-bit input data is distributed to 16 rows of PE modules at 128 bits per row, and the upper 2048-bit input data is distributed to 16 columns of PE modules at 128 bits per column.
5. The matrix operation array system for deep learning acceleration kernel according to claim 4, characterized in that: The pipeline control unit uses pipeline processing technology to control the working state of the PE modules in each operation cycle, including: Using pipeline processing technology to control the PE modules to perform the following 11 operation cycle operations: cycle0: The PE module preprocesses the received input data and sends it to the port of the multiplier; cycle1: Perform multiplication operations to obtain 8-bit and 16-bit multiplication operation results; cycle2: For fixed-point operations, sum the fixed-point multiplication operation results through a fixed-point adder tree and accumulate them with the values in the local fixed-point register to obtain the fixed-point operation results; for floating-point operations, perform post-processing based on the floating-point multiplication operation results to obtain 8 floating-point multiplication operation results; cycle3: The first beat operation of the first-level floating-point adder tree; cycle4: The second beat operation of the first-level floating-point adder tree to obtain the operation results of the first-level floating-point adder tree; cycle5: The first - beat operation of the second - level floating - point adder tree; cycle6: The second - beat operation of the second - level floating - point adder tree, obtaining the operation result of the second - level floating - point adder tree; cycle7: The first - beat operation of the third - level floating - point adder tree; cycle8: The second - beat operation of the third - level floating - point adder tree, obtaining the operation result of the third - level floating - point adder tree; cycle9: Accumulating the sum of the results of 8 floating - point multiplication operations and the value in the local floating - point register; cycle10: The second - beat operation of the floating - point accumulation operation, obtaining the floating - point operation result.
6. The matrix operation array system for deep learning acceleration cores according to any one of claims 1 to 5, characterized in that: The PE module supports multiplication and accumulation operations with data types of INT8, UINT8, INT16, UINT16, FP16, and BF16. To support the reuse of the floating - point adder, a new data format with the data type of FP19 is designed. When performing floating - point addition operations, the data formats of FP16 and BF16 are first converted to the FP19 data format, and then the floating - point addition algorithm is performed; Among them, the data format of FP16 is 1 sign bit, 5 exponent bits, and 10 mantissa bits; the data format of BF16 is 1 sign bit, 8 exponent bits, and 7 mantissa bits; the data format of FP19 is 1 sign bit, 8 exponent bits, and 10 mantissa bits.
7. An operation method for a matrix operation array of a deep learning acceleration core, applied to the matrix operation array system for a deep learning acceleration core described in claim 1, characterized in that: It includes the following steps: S1. After receiving the instructions from the upper layer, the input control unit distributes the input data from the memory as operands to the PE module and controls the PE module to work by sending a work enable to the PE module; S2. The PE module performs fixed - point and floating - point multiplication and accumulation operations and processes the result overflow generated during the addition process; S3. The output control unit receives the operation result of the PE module and, after receiving the output enable, processes the operation result of the PE module according to the output method and then outputs it to the external memory; Among them, during the process of the PE module performing fixed - point and floating - point multiplication and accumulation operations, the pipeline control unit uses pipeline - type processing technology to control the working state of the PE module in each operation cycle.
8. The operation method of the matrix operation array for deep learning acceleration cores according to claim 7, characterized in that: The pipeline control unit uses pipeline - type processing technology to control the working state of the PE module in each operation cycle, including: Using pipeline - type processing technology to control the PE module to perform the following operations in 11 operation cycles: cycle0: The PE module pre - processes the received input data and sends it to the port of the multiplier; cycle1: Performing multiplication operations to obtain 8 - bit and 16 - bit multiplication operation results; cycle2: For fixed - point operations, summing the fixed - point multiplication operation results through a fixed - point adder tree and accumulating them with the value in the local fixed - point register to obtain the fixed - point operation result; for floating - point operations, post - processing the floating - point multiplication operation results to obtain 8 floating - point multiplication operation results; cycle3: The first - beat operation of the first - level floating - point adder tree; cycle4: The second - beat operation of the first - level floating - point adder tree, obtaining the operation result of the first - level floating - point adder tree; cycle5: The first - beat operation of the second - level floating - point adder tree; Cycle 6: The second - beat operation of the second - level floating - point adder tree to obtain the operation result of the second - level floating - point adder tree; Cycle 7: The first - beat operation of the third - level floating - point adder tree; Cycle 8: The second - beat operation of the third - level floating - point adder tree to obtain the operation result of the third - level floating - point adder tree; Cycle 9: Accumulate the sum of the results of 8 floating - point multiplication operations with the value in the local floating - point register; Cycle 10: The second - beat operation of the floating - point accumulation operation to obtain the floating - point operation result.
Citation Information
Cited By
Automatic driving data processing method and device based on systolic array, equipment and storage medium
CN120892010A