Coarse-grained reconfigurable array system for deep learning and calculation method

By designing a coarse-grained reconfigurable array system for deep learning, using the combination of controller and processing units, the compatibility between floating-point operations and integer operations is achieved, solving the problem that the existing CGRA architecture cannot process floating-point data, and improving computing efficiency and accuracy.

CN119987861AActive Publication Date: 2025-05-13UNIV OF SCI & TECH OF CHINA
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510160346.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-13
Estimated Expiration
2045-02-13

AI Technical Summary

Technical Problem

The existing coarse-grained reconfigurable array (CGRA) architecture can only perform integer data calculations and cannot meet the accuracy requirements for floating-point data in deep learning model training. Direct embedding of floating-point processing units will lead to increased hardware overhead and reduce energy efficiency advantages.

Method used

A coarse-grained reconfigurable array system for deep learning is designed, and the controller inputs mode control instructions and operation instructions to enable the processing unit to perform floating-point operations or integer operations. The processing unit includes an operand selector and an arithmetic logic subunit. By multiplexing the integer computing resources in traditional CGRA, floating-point calculations in BF16 format are realized in a single cycle.

Benefits of technology

It realizes the effect of calculating both integer and floating-point data, avoids the problem of increasing hardware overhead, improves computing efficiency, and supports high-precision training of deep learning models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987861A_ABST
    Figure CN119987861A_ABST
Patent Text Reader

Abstract

The invention provides a coarse-grained reconfigurable array system for deep learning and a computing method, which can be applied to the technical field of reconfigurable arrays. The system comprises a controller used for determining input information input to at least one processing unit, and the input information comprises to-be-calculated data, an operation instruction and a mode control instruction; and the processing unit group comprises a plurality of processing units, the plurality of processing units form a reconfigurable array, and each processing unit is used for performing floating-point operation or integer operation on the data to be calculated based on the mode control instruction and the operation instruction to obtain an operation result. The processing unit comprises an operand selector used for determining first operation data and second operation data from the neighbor operation result, the weight data and the input data based on the operation instruction; and the arithmetic logic subunit is used for performing floating-point operation or integer operation on the first operation data and the second operation data based on the operation subinstruction corresponding to the mode control instruction to obtain an operation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of reconfigurable array technology, and in particular to a coarse-grained reconfigurable array system and a computing method for deep learning. Background Art

[0002] A coarse-grained reconfigurable array (CGRA) is an architecture that performs computing tasks based on instructions and data information configured by data streams. In existing coarse-grained reconfigurable arrays, most architectures only support integer data computing, considering the high energy efficiency of computing and the accuracy of supporting the target network. However, in scenarios with higher accuracy requirements, such as the training of deep learning models, floating-point data is needed to improve training accuracy.

[0003] The following methods are usually used in related technologies to implement CGRA calculations on floating-point data, such as directly embedding a dedicated floating-point processing unit (FPU) in the processing unit of CGRA. However, this method will greatly increase the hardware overhead and reduce the energy efficiency advantage of the coarse-grained reconfigurable array, thereby limiting its application. Summary of the invention

[0004] In view of the above problems, the present disclosure provides a coarse-grained reconfigurable array system and a computing method for deep learning.

[0005] According to one aspect of the present disclosure, a coarse-grained reconfigurable array system for deep learning is provided, the system comprising:

[0006] A controller is used to determine input information input to at least one processing unit, wherein the input information includes data to be calculated, operation instructions and mode control instructions; a processing unit group includes multiple processing units, and the multiple processing units form a reconfigurable array, each processing unit is used to perform floating-point operations or integer operations on the data to be calculated based on the mode control instructions and the operation instructions to obtain operation results; wherein the data to be calculated includes input data and weight data, and the processing unit includes: an operand selector and an arithmetic logic subunit; the operand selector is used to receive the neighbor operation results output by the neighbor processing unit, and based on the operation instructions, determine the first operation data and the second operation data from the neighbor operation results, weight data and input data, wherein the neighbor processing unit is a processing unit adjacent to the processing unit; the arithmetic logic subunit is used to determine the operation sub-instruction corresponding to the mode control instruction from the operation instruction, so as to perform floating-point operations or integer operations on the first operation data and the second operation data based on the operation sub-instruction to obtain the operation result.

[0007] Another aspect of the present disclosure provides a coarse-grained reconfigurable array computing method for deep learning, comprising: determining input information to at least one processing unit through a controller, wherein the input information includes data to be calculated, operation instructions and mode control instructions; inputting data to be calculated to at least one processing unit through an input bus; inputting operation instructions to at least one processing unit through a configuration bus; performing floating-point operations or integer operations on the data to be calculated through each processing unit according to the mode control instructions and the operation instructions to obtain operation results; and outputting the operation results through an output bus.

[0008] According to the coarse-grained reconfigurable array system for deep learning disclosed in the present invention, mode control instructions and operation instructions are input through a controller so that a processing unit group can perform floating-point operations or integer operations on the calculation data based on the mode control instructions and the operation instructions. Specifically, the processing unit includes an operand selector and an arithmetic logic subunit. When performing calculations, the operand selector can be used to select the first operation data and the second operation data, and the arithmetic logic subunit can be used to determine the operation subinstruction corresponding to the mode control instruction from the operation subinstruction, so as to perform floating-point or integer operations on the first operation data and the second operation data based on the operation subinstruction, thereby solving the technical problem that the CGRA architecture in the related art can only perform integer calculations and the hardware overhead is large if the floating-point processing unit is directly embedded, that is, the computing effect of being able to calculate both integer data and floating-point data is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The above contents and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:

[0010] Figure 1 A schematic diagram of a coarse-grained reconfigurable array system for deep learning according to an embodiment of the present disclosure is schematically shown;

[0011] Figure 2 A schematic diagram schematically shows a floating point operation data BF16 format according to an embodiment of the present disclosure;

[0012] Figure 3 A schematic diagram of a processing unit of a coarse-grained reconfigurable array system for deep learning according to an embodiment of the present disclosure is schematically shown;

[0013] Figure 4 A schematic diagram of an arithmetic logic subunit of a coarse-grained reconfigurable array system for deep learning according to an embodiment of the present disclosure is schematically shown;

[0014] Figure 5A schematic diagram schematically illustrates the reconstruction of an arithmetic logic subunit of a coarse-grained reconfigurable array system for deep learning when performing floating-point multiplication according to an embodiment of the present disclosure;

[0015] FIG6( a ) schematically shows a first reconstruction diagram of an arithmetic logic subunit of a coarse-grained reconfigurable array system for deep learning when performing floating-point addition according to an embodiment of the present disclosure;

[0016] FIG6( b ) schematically shows a second reconstruction diagram of the arithmetic logic subunit of the coarse-grained reconfigurable array system for deep learning when performing floating-point addition according to an embodiment of the present disclosure;

[0017] FIG6( c ) schematically shows a third reconstruction diagram of the arithmetic logic subunit of the coarse-grained reconfigurable array system for deep learning when performing floating-point addition according to an embodiment of the present disclosure;

[0018] FIG6( d ) schematically shows a fourth reconstruction diagram of the arithmetic logic subunit of the coarse-grained reconfigurable array system for deep learning when performing floating-point addition according to an embodiment of the present disclosure;

[0019] Figure 7 A schematic diagram of a normalized circuit of a coarse-grained reconfigurable array system for deep learning according to the present disclosure is schematically shown;

[0020] Figure 8 A schematic diagram of an arithmetic logic subunit of a coarse-grained reconfigurable array system for deep learning according to another embodiment of the present disclosure is schematically shown;

[0021] Fig. 9 A flowchart of a coarse-grained reconfigurable array computing method for deep learning according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0022] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present disclosure. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.

[0023] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise", "include", etc. used herein indicate the existence of the features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components.

[0024] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0025] When using expressions such as "at least one of A, B, and C, etc.", they should generally be interpreted according to the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0026] It should be noted that the coarse-grained reconfigurable array system and computing method for deep learning disclosed in the present invention can be used in the field of reconfigurable array technology, and can also be used in any field outside the field of reconfigurable array technology, such as the field of artificial intelligence technology. There is no limitation on the application field of the coarse-grained reconfigurable array system and computing method for deep learning disclosed in the present invention.

[0027] During the research process, it was found that in the existing coarse-grained reconfigurable arrays, considering the high energy efficiency of calculations and the accuracy of supporting target networks, most architectures only support integer data calculations. When higher accuracy is required or other networks are intended to be supported, the calculation results of this architecture are not enough to support real applications. The BF16 (Brain Floating Point 16-bit) format is a floating-point data representation mainly for deep learning. Its exponent bits are the same as FP32, but the mantissa bits are only 7 bits. Many modern hardware architectures have provided native support for BF16, including Google TPU and NVIDIA A100 GPU, which can greatly improve AI performance while ensuring the accuracy of neural networks. If a dedicated floating-point processing unit (FPU) is directly embedded in the processing unit of the array, the hardware overhead will be greatly increased (about 25%), which will reduce the energy efficiency advantage of the coarse-grained reconfigurable array, thereby limiting the application. If the instruction reconstruction method is used to complete floating-point calculations, although there is no hardware overhead, it takes more than 100 clock cycles to complete each floating-point operation on average, which cannot support the computing power requirements of neural networks.

[0028] In view of this, an embodiment of the present disclosure provides a coarse-grained reconfigurable array system for deep learning, comprising: a controller, used to determine input information input to at least one processing unit, wherein the input information includes data to be calculated, operation instructions and mode control instructions; a processing unit group, including multiple processing units, the multiple processing units form a reconfigurable array, each processing unit is used to perform floating-point operations or integer operations on the data to be calculated based on the mode control instructions and the operation instructions to obtain operation results; wherein the data to be calculated includes input data and weight data, and the processing unit comprises: an operand selector and an arithmetic logic subunit; the operand selector is used to receive the neighbor operation results output by the neighbor processing unit, and based on the operation instructions, determine the first operation data and the second operation data from the neighbor operation results, weight data and input data, wherein the neighbor processing unit is a processing unit adjacent to the processing unit; the arithmetic logic subunit is used to determine the operation sub-instruction corresponding to the mode control instruction from the operation instruction, so as to perform floating-point operations or integer operations on the first operation data and the second operation data based on the operation sub-instruction to obtain the operation result.

[0029] Figure 1 A schematic diagram of a coarse-grained reconfigurable array system for deep learning according to an embodiment of the present disclosure is schematically shown.

[0030] like Figure 1 As shown, the system includes a controller 11 , a processing unit group 12 , an input storage unit 13 , an output storage unit 14 , an input bus 15 , a configuration bus 16 and an output bus 17 .

[0031] A controller 11, used to determine input information input to at least one processing unit, wherein the input information includes data to be calculated, operation instructions and mode control instructions;

[0032] The processing unit group 12 includes a plurality of processing units, and the plurality of processing units form a reconfigurable array, wherein each processing unit is used to perform floating-point operations or integer operations on the data to be calculated based on the mode control instruction and the operation instruction to obtain the operation result; wherein the data to be calculated includes input data and weight data, and the processing unit includes: an operand selector and an arithmetic logic subunit; the operand selector is used to receive the neighbor operation result output by the neighbor processing unit, and based on the operation instruction, determine the first operation data and the second operation data from the neighbor operation result, the weight data and the input data, wherein the neighbor processing unit is a processing unit adjacent to the processing unit; the arithmetic logic subunit is used to determine the operation sub-instruction corresponding to the mode control instruction from the operation instruction, so as to perform floating-point operations or integer operations on the first operation data and the second operation data based on the operation sub-instruction to obtain the operation result.

[0033] According to an embodiment of the present disclosure, the controller 11 is used to control the start and end of operations, data loading, and array computing mode according to configuration information transmitted by the data stream.

[0034] According to an embodiment of the present disclosure, the configuration information includes instruction information such as state instructions and mode control instructions. The state instructions are used to control the related operations of the operation instructions, and may include: the number of operation instructions, the execution order of operation instructions, the number of state instructions, the number of weights, the number of small outer loops, the number of large outer loops, the parameters of the small outer loop weight base address increment, and the large outer loop weight base address increment, etc. The mode control instruction is used to control the operation mode of the current operation.

[0035] According to an embodiment of the present disclosure, the operation instruction is an instruction for controlling the processing unit to perform operations.

[0036] According to the embodiments of the present disclosure, in some embodiments, the controller 11 configures the number of operation instructions, the number of state instructions, the number of weights, the number of small outer loops, the number of large outer loops, the parameters of the small outer loop weight base address increment, the large outer loop weight base address increment, and the mode control instructions to the processing unit group 12 according to the instruction order in the configuration information.

[0037] According to an embodiment of the present disclosure, the processing unit group 12 forms an array, the length and width of the array can be configured according to parameters, and each processing unit can be interconnected with the processing units above, below, left and right. Each processing unit initially receives configuration information from the controller 11, such as the number of cycles, floating point control, and load operation instructions, load weight data, output result data, etc. The processing unit array shares the output bus 17 by row and the input bus 15 by column to reduce wiring pressure.

[0038] According to an embodiment of the present disclosure, the processing unit group 12 includes multiple processing units, each processing unit operates on specified data according to the operation instructions stored therein, and each processing unit supports two sets of instruction sets for integer and BF16 format data.

[0039] According to an embodiment of the present disclosure, the input storage unit 13 and the output storage unit 14 are composed of a group of input and output FIFOs. The input storage unit 13 is used to cache input data and use the bus and the top-level controller 11 to configure the working mode, instruction data and weight data of the processing unit and output the final calculation result.

[0040] According to an embodiment of the present disclosure, the input bus 15 is used to input weight data and input data to at least one processing unit. The configuration bus 16 is used to input configuration information and operation instructions to at least one processing unit. The output bus 17 is used for at least one processing unit to output operation results.

[0041] According to an embodiment of the present disclosure, the array uses distributed storage to reduce dynamic power consumption, and each processing unit has a local register to store instructions and weight data to support deep learning-oriented calculations.

[0042] According to an embodiment of the present disclosure, the array needs to prepare the operation instructions and weight data of each processing unit before starting to execute the computing task. The controller 11 controls the data input FIFO and the instruction FIFO to preload the operation instructions and weight data and configuration information to the processing unit. The configuration information will include whether each processing unit performs floating-point operations and thus performs calculations according to different instruction sets.

[0043] According to an embodiment of the present disclosure, each processing unit has two operands, namely, first operation data and second operation data, when calculating in a single cycle. It can be determined whether the instruction calculation adopted for the first operation data and the second operation data is an integer operation or a floating-point operation according to the mode control instruction. The floating-point calculation in the BF16 format is implemented in a single cycle by reusing the integer computing resources in the conventional coarse-grained reconfigurable array.

[0044] According to an embodiment of the present disclosure, at the beginning of calculation, the data to be calculated, the operation instructions and the mode control instructions are preloaded to the processing unit group 12 through the input FIFO. The arrangement of the configuration information is shown in Table 1. In the case where there are 24 processing units in the processing unit group 12, fp_ctrl[23:0] is a mode control instruction for controlling each processing unit respectively. When fp_ctrl is high, the corresponding processing unit is converted from integer data type calculation to BF16 data type calculation.

[0045] Table 1

[0046]

[0047] According to the embodiments of the present disclosure, preloading can be achieved by first loading instruction information and weight data information from the FIFO into the corresponding processing unit according to the configuration information of the controller 11. When the preloading process is completed, the controller 11 controls the processing unit group 12 to start calculation. At this time, each processing unit in the processing unit group 12 starts calculation according to the operation instructions stored in its own processing unit. When the processing unit reads an instruction indicating an output request, it will output the result of this calculation to the output storage unit 14. After all processing units have completed the execution of the control logic configured according to the configuration information, the coarse-grained reconfigurable array system will send a calculation end signal to wait for the next round of calculation.

[0048] Figure 2 The diagram schematically shows a floating-point operation data BF16 format according to an embodiment of the present disclosure.

[0049] like Figure 2As shown, the BF16 format is a floating-point data representation mainly for deep learning. The data identifier of the BF16 format is shown in the following formula (1).

[0050] ; (1)

[0051] in, is the representation of the sign bit, is the representation of the exponent bits, and Mantissa_true is the representation of the mantissa bits.

[0052] According to an embodiment of the present disclosure, the binary representation of Mantissa_true is .like Figure 2 In the Value_BF16, the value is 4.5, the sign bit is 0, the exponent bit is 8'b10000001, which is 129 in decimal, and the mantissa is 7'b0010000. In binary, {1'b1,7'b0010000} is equal to 1.125 in decimal, so the value represented is shown in the following formula (2).

[0053] ; (2)

[0054] Figure 3 A schematic diagram of a processing unit of a coarse-grained reconfigurable array system for deep learning according to an embodiment of the present disclosure is schematically shown.

[0055] like Figure 3 As shown, the processing unit includes: an operand selector 31, an arithmetic logic subunit 32, an instruction register subunit 33, a weight register subunit 34, a local register subunit 35 and an output register subunit 36.

[0056] According to an embodiment of the present disclosure, the operand selector 31 can be represented by two selectors, namely MUX, the arithmetic logic subunit 32 can be represented by ALU, the instruction register subunit 33 can be represented by CTRL, the weight register subunit 34 can be represented by WRF, the local register subunit 35 can be represented by LRF, and the output register subunit 36 ​​can be represented by OUT REG.

[0057] According to an embodiment of the present disclosure, the instruction register subunit 33 is used to store operation instructions, state instructions and mode control instructions. The weight register subunit 34 is used to store weight data or input data; the local register subunit 35 is used to store intermediate data generated during the operation; and the output register subunit 36 ​​is used to store the final operation result of the processing unit.

[0058] According to an embodiment of the present disclosure, each processing unit will decode and obtain operation data according to the operation instruction in the current clock cycle. The source of the operation data is the output of the processing units above, below, left and right, the data bus, the weight register output, the local register output, etc.

[0059] According to an embodiment of the present disclosure, the internal design of the arithmetic logic subunit 32 in the processing unit is a design that optimizes the critical path as much as possible, so that two BF16 multiplications or one BF16 addition can be implemented in each cycle, thereby improving the operation efficiency.

[0060] According to an embodiment of the present disclosure, the array uses distributed storage to reduce dynamic power consumption, and each processing unit has a local register to store instructions and weight data to support deep learning-oriented calculations.

[0061] According to an embodiment of the present disclosure, in each processing unit, two different instruction formats are implemented according to the high and low levels of the mode control instruction. When the mode control instruction is high, the operation is a floating point calculation; when the mode control instruction is low, an integer operation is performed on the same instruction.

[0062] According to an embodiment of the present disclosure, the processing unit receives a 32-bit input from the data bus and a control instruction from the controller 11, and the controller 11 controls whether the current data bus information is used for preloading. When not used for preloading, the information of the data bus will be used as one of the sources for selecting the operation data of the arithmetic logic unit. When the controller 11 controls the processing unit to preload, the operation instruction of the corresponding processing unit will be first stored in the local storage of the processing unit. When the operation instruction width is 20 bits, 32 instructions can be stored.

[0063] According to an embodiment of the present disclosure, when the controller 11 controls the processing unit to preload the weight, the information of the data bus is stored in the weight register, a total of 32 32-bit data. When the calculation task starts, the arithmetic logic subunit receives information from the local instruction storage, and the operand selector 31 selects the corresponding operation data according to the operation instruction. The ALU calculation result is stored in the 32-bit output register subunit 36 ​​and provided to the adjacent processing unit. It is selected whether to write to the local register subunit 35 according to the operation instruction, and is written to the data output FIFO when there is an output demand. The ALU receives a mode control instruction to select the execution of different operation instructions.

[0064] According to an embodiment of the present disclosure, Table 2 shows the format of the operation instruction, mux_ctrl1 and mux_ctrl2 respectively control the operand selection of two selectors in the operand selector 31 in the processing unit. The weight register subunit 34 outputs weight data as one of the operand sources of the data selector according to the address provided by WRF_addr.

[0065] According to the embodiment of the present disclosure, there is no mode control instruction in the instruction format in Table 2, because the mode control instruction is already given by the initial configuration information, and adding floating point control to the operation instruction will increase the instruction storage cost. The first line represents the control information, and the second line represents the address information.

[0066] Table 2

[0067]

[0068] According to an embodiment of the present disclosure, op_code is an operation sub-instruction for controlling the arithmetic logic sub-unit 32, wherein different op_codes can be selected through different mode control instructions, wherein Table 3 is an operation sub-instruction when the mode control instruction represents integer operations, and Table 4 is an operation sub-instruction when the mode control instruction represents floating-point operations.

[0069] Table 3

[0070]

[0071] Table 4

[0072]

[0073] Figure 4 A schematic diagram of an arithmetic logic subunit of a coarse-grained reconfigurable array system for deep learning according to an embodiment of the present disclosure is schematically shown.

[0074] like Figure 4 As shown, the arithmetic logic subunit 32 includes: a multiplier group 41, a first adder circuit 42, a second adder circuit 43, a normalization circuit 44, a third adder circuit 45, a fourth adder circuit 46, a result selector 47, a logic operator group 48, a consistency operator 49 and a sign bit selector 50. The multiplier group 41 includes: a first multiplier 411, a second multiplier 412, a third multiplier 413, and a fourth multiplier 414.

[0075] The first adder circuit 42, the first input end of the first adder circuit 42 is connected to the first multiplier 411, the second input end of the first adder circuit 42 is connected to the output end of the second adder circuit 43, the third input end of the first adder circuit 42 is connected to the second multiplier 412, and the output end of the first adder circuit 42 is respectively connected to the input end of the normalization circuit 44, the first input end of the third adder circuit 45 and the input end of the result selector 47.

[0076] The normalization circuit 44 has an output terminal connected to a second input terminal of the third adder circuit 45 .

[0077] The second adder circuit 43, the first input end of the second adder circuit 43 is connected to the third multiplier 413, the second input end of the second adder circuit 43 is connected to the fourth multiplier 414, and the output end of the second adder circuit 43 is respectively connected to the second input end of the first adder circuit 42, the third input end of the third adder circuit 45, the first input end of the fourth adder circuit 46 and the input end of the result selector 47.

[0078] The third adder circuit 45 , the output end of the third adder circuit 45 is connected to the result selector 47 .

[0079] A fourth adder circuit 46 , an output end of the fourth adder circuit 46 is connected to a result selector 47 .

[0080] The first adder circuit 42, the second adder circuit 43, the third adder circuit 45 and the fourth adder circuit 46 respectively include at least one selector and an adder, and the input of the adder is selected by at least one selector based on the operation sub-instruction.

[0081] According to the embodiment of the present disclosure, each processing unit includes a basic structure of four 8-bit multipliers, one 8-bit adder, two 16-bit adders, one 17-bit adder, one 35-bit adder, and right shift, left shift, OR, AND, XOR, etc. By adding a reasonable data selector (MUX) to the input of each basic structure to realize module reuse, more functions can be realized while reducing hardware overhead.

[0082] According to an embodiment of the present disclosure, the normalization circuit 44 is used to provide a result of the required number of shift bits for floating point operations.

[0083] According to the embodiments of the present disclosure, a traditional CGRA may also include four 8-bit multipliers, one 8-bit adder, two 16-bit adders, one 17-bit adder, and one 35-bit adder. The present application adds a low-cost normalization circuit 44 and a selector on this basis to achieve multiplexing of the above-mentioned operators to reduce hardware overhead to implement floating-point operations, thereby avoiding the overhead caused by adding an additional floating-point processor FPU.

[0084] According to an embodiment of the present disclosure, the arithmetic logic subunit 32 further includes: a consistency operator 49 and a logic operator group 48. The logic operator group 48 may include a plurality of logic operators such as a comparator, a shifter, an XOR, and an inverter. The input end of the consistency operator 49 may be connected to the output end of the logic operator group 48. The consistency operator 49 is used to compare the input data with the preset data bit by bit, and output 1 if the two data are exactly the same, otherwise output 0.

[0085] According to an embodiment of the present disclosure, the arithmetic logic subunit 32 also includes: a consistency operator 49; the first adder circuit 42 includes: a first selector 421, a second selector 422, a first adder 423, a first XOR operator 424, a second XOR operator 425, a first shift operator 426, a bit number expander 427, a third selector 428 and a third XOR operator 429.

[0086] A first input terminal of the first selector 421 is connected to an output terminal of the first multiplier, and an output terminal of the first selector 421 is connected to a first input terminal of the first adder 423 .

[0087] A first input terminal of the second selector 422 is connected to an output terminal of the first XOR operator 424 , a second input terminal of the second selector 422 is connected to a second multiplier, and an output terminal of the second selector 422 is connected to a second input terminal of the first adder 423 .

[0088] A first input terminal of the first XOR operator 424 is connected to an output terminal of the first shift operator 426 , and a second input terminal of the first XOR operator 424 is connected to an output terminal of the bit number expander 427 .

[0089] An output terminal of the second XOR operator 425 is connected to a third input terminal of the first adder 423 .

[0090] An output terminal of the first adder 423 is connected to an input terminal of the normalization circuit 44 and a first input terminal of the third adder circuit 45 , respectively.

[0091] A first input terminal of the first shift operator 426 is connected to the third selector 428 , and a second output terminal of the first shift operator 426 is connected to an output terminal of the second adder circuit 43 .

[0092] A first input terminal of the third selector 428 is connected to an output terminal of the consistency operator 49 .

[0093] An input terminal of the bit number expander 427 is connected to a second input terminal of the third XOR operator 429 .

[0094] According to an embodiment of the present disclosure, the arithmetic logic subunit 32 also includes: a consistency operator 49; the second adder circuit 43 includes: a fourth selector 431, a second adder 432, a first negation operator 433, a fifth selector 434, a sixth selector 435, a seventh selector 436 and an eighth selector 437.

[0095] A first input terminal of the fourth selector 431 is connected to the output terminal of the third multiplier, a second input terminal of the fourth selector 431 is connected to the output terminal of the first negation operator 433 , and an output terminal of the fourth selector 431 is connected to a first input terminal of the second adder 432 .

[0096] An input terminal of the first negation operator 433 is connected to an output terminal of the fifth selector 434 .

[0097] An input terminal of the fifth selector 434 is connected to an output terminal of the consistency operator 49 .

[0098] A first input terminal of the sixth selector 435 is connected to the fourth multiplier, and an output terminal of the sixth selector 435 is connected to a second input terminal of the second adder 432 .

[0099] An input end of the seventh selector 436 is connected to an output end of the consistency operator 49 , and an output end of the seventh selector 436 is connected to an input end of the eighth selector 437 .

[0100] An output terminal of the eighth selector 437 is connected to the third input terminal of the second adder 432 .

[0101] According to an embodiment of the present disclosure, the arithmetic logic subunit 32 further includes: a sign bit selector 50. The third adder circuit 45 includes: a second negation operator 451, a ninth selector 452, an AND logic operator 453, a third adder 454, a tenth selector 455, an eleventh selector 456, a twelfth selector 457, a thirteenth selector 458 and a second shift operator 459.

[0102] The input end of the second negation operator 451 is connected to the output end of the normalization circuit 44 and the output end of the second XOR operator 425 in the first adder circuit 42, respectively. The output end of the second negation operator 451 is connected to the first input end of the ninth selector 452 and the first input end of the logic operator 453, respectively.

[0103] The second input terminal of the ninth selector 452 is connected to the output terminal of the first adder circuit 42, and the output terminal of the ninth selector 452 is connected to the first input terminal of the third adder 454;

[0104] The second input terminal of the third adder 454 is connected to the output terminal of the tenth selector 455, the third input terminal of the third adder 454 is connected to the output terminal of the eleventh selector 456, and the output terminal of the third adder 454 is connected to the first input terminal of the twelfth selector 457;

[0105] The second input terminal of the AND logic operator 453 is connected to the output terminal of the first adder circuit 42, and the output terminal of the AND logic operator 453 is connected to the first input terminal of the thirteenth selector 458 and the second input terminal of the twelfth selector 457 respectively;

[0106] A second input terminal of the thirteenth selector 458 is connected to an output terminal of the second shift operator 459;

[0107] A first input terminal of the second shift operator 459 is connected to the output terminal of the normalization circuit 44 , and a second input terminal of the second shift operator 459 is connected to the output terminal of the first adder circuit 42 .

[0108] According to an embodiment of the present disclosure, the selectors in the arithmetic logic subunit 32 can be used to determine the input sub-data at the preset data address from the first operation data and the second operation data respectively based on the instruction signal corresponding to the selector in the operation sub-instruction, and determine the output data from the input sub-data.

[0109] Figure 5 A schematic diagram schematically illustrates the reconstruction of the arithmetic logic subunit of the coarse-grained reconfigurable array system for deep learning when performing floating-point multiplication according to an embodiment of the present disclosure.

[0110] like Figure 5 As shown, when the floating-point operation is a floating-point multiplication, based on the operator instruction, the calculation process of each of the following structures in the first adder circuit 42 is as follows.

[0111] The first multiplier is used to perform multiplication operation on the first mantissa bit sub-data in the first operation data and the second mantissa bit sub-data in the second operation data to obtain a multiplication calculation result, wherein the first mantissa bit sub-data and the second mantissa bit sub-data are output data of two selectors connected to the input end of the first multiplier.

[0112] The first selector 421 is used for outputting the first exponent bit sub-data in the first data.

[0113] The second selector 422 is used to output the second exponent bit sub-data in the second data.

[0114] The first adder 423 is used for performing addition calculation on the first exponent bit sub-data and the second exponent bit sub-data to obtain a first exponent bit result.

[0115] The second XOR operator 425 is used for performing bitwise XOR on the sign bit sub-data of the first operation data and the sign bit sub-data of the second operation data to obtain an output sign bit result.

[0116] According to an embodiment of the present disclosure, in the case where the floating-point operation is a floating-point multiplication, based on the operator instruction, the calculation process of each of the following structures in the third adder circuit 45 is as follows.

[0117] a ninth selector 452, for receiving the first exponent bit result output by the first adder circuit 42, and for outputting the first exponent bit result;

[0118] a tenth selector 455, configured to output a second preset value;

[0119] an eleventh selector 456, configured to receive the multiplication result output by the first multiplier, and to output a sub-result of the multiplication result whose data address is in the first range;

[0120] A third adder 454 is used to perform addition calculation on the first exponent bit result, the second preset value and the sub-result to obtain an output exponent bit result;

[0121] According to an embodiment of the present disclosure, the second preset value is 8'd129, and the first range may be the 15th bit.

[0122] According to an embodiment of the present disclosure, Figure 5 The inputs of the first multiplier, the first adder 423 and the third adder 454 are all selected by their respective selectors.

[0123] According to an embodiment of the present disclosure, data1 and data2 are respectively operation data selected by data selectors MUX1 and MUX2 in the processing unit structure, data1[15:0] and data2[15:0] respectively represent two BF16 format data to be calculated, data1[14:7] is the first exponent bit sub-data in the first operation data, data1[6:0] is the first mantissa bit sub-data, data1

[15] is the first sign bit sub-data, and the second operand is similar, the numbers in [] represent the range of the data address of the data, and the data address can specifically be the number of data bits.

[0124] According to an embodiment of the present disclosure, when performing floating-point multiplication, two 8-bit adders are used instead of one 8-bit adder. However, since the traditional CGRA includes two 17-bit adders, in order to realize the multiplexing of integer operations and floating-point operations, and the 17-bit adder can meet the calculation requirements of floating-point multiplication, two 17-bit adders can still be used to replace the two 8-bit adders.

[0125] According to the embodiments of the present disclosure, Figure 5As shown, data1[14:7] and data2[14:7] are added as two inputs of an 8-bit adder, which represents the addition of the exponent bits of the two operation data in the BF16 format. The mantissa bits {1'b1, data1[6:0]} and {1'b1, data2[6:0]} are multiplexed with an 8-bit multiplier, outputting the 16-bit multiplication result product0.

[0126] According to the embodiments of the present disclosure, for the multiplication of floating-point numbers in BF16 format, the normalized result can only be a carry or remain unchanged. In order to obtain the exponent of the final multiplication result, 127 should be subtracted from the result of adding the exponents of two BF16 format data. This is reflected in the 8-bit binary adder as a negation plus 1, i.e. 8'd129. Therefore, in order to perform normalization processing. Another adder is reused to add the exponent addition result and 8'd129. The input carry is product0

[15] , i.e., the exponent sub-data in the multiplication result, which is used to determine whether a carry is required to indicate the normalized result, thereby obtaining the final output exponent result.

[0127] According to an embodiment of the present disclosure, the multiplexing result selector 47 may select sub-data whose data addresses are within a preset address range from the multiplication calculation result as the final output bit number result.

[0128] According to an embodiment of the present disclosure, the second XOR operator 425 may be reused and connected to the result selector 47 to obtain a final output sign bit result.

[0129] According to an embodiment of the present disclosure, the result selector 47 can be used to splice and output the output exponent bit result, the output bit result and the output sign bit result, so as to obtain the final floating-point multiplication operation result.

[0130] According to an embodiment of the present disclosure, in some embodiments, multiplication calculation of two groups of BF16 format data is implemented by multiplexing other multipliers and adders of the integer calculation module, such as: the two groups of data are data1[15:0]*data2[15:0] and data1[31:16] multiplied by *data2[31:16].

[0131] FIG6( a ) schematically shows a first reconstruction diagram of an arithmetic logic subunit of a coarse-grained reconfigurable array system for deep learning when performing floating-point addition according to an embodiment of the present disclosure.

[0132] As shown in FIG6 (a), the comparator CMP may be a device in the logic operator group 48, and the absolute values ​​of the first operation data and the second operation data are compared by the comparator to obtain a comparison result, data1_cmp_data2. The comparison result is sent to the consistency operator. The consistency operator 49 compares the larger operation data in the comparison result with the first operation data or the second operation data bit by bit, thereby obtaining an input signal representing the comparison result between the first operation data and the second operation data, namely op1_biggger in the figure.

[0133] According to an embodiment of the present disclosure, the sign bit selector 50 may output a sign bit result of the largest data among the first operation data and the second operation data based on the above input signal.

[0134] FIG6( b ) schematically shows a second reconstruction diagram of the arithmetic logic subunit of the coarse-grained reconfigurable array system for deep learning when performing floating-point addition according to an embodiment of the present disclosure.

[0135] According to an embodiment of the present disclosure, in the case where the floating-point operation is a floating-point addition, based on the operator instruction, the calculation process of each of the following structures in the second adder circuit 43 is as follows.

[0136] The fifth selector 434 is used to select the exponent bit sub-data of the minimum data from the first exponent bit sub-data of the first operation data and the second mantissa bit sub-data of the second operation data based on the input signal output by the consistency operator 49 for representing the comparison result between the first operation data and the second operation data.

[0137] The first negation operator 433 is used for negating the exponent bit sub-data of the minimum data to obtain first negated sub-data.

[0138] The fifth selector 434 is used for outputting the first inverted sub-data to the second adder 432 .

[0139] The sixth selector 435 is used to output null data.

[0140] The seventh selector 436 is used to select the exponent bit sub-data of the maximum data from the first exponent bit sub-data and the second mantissa bit sub-data based on the input information.

[0141] The eighth selector 437 is used for outputting the exponential bit sub-data of the maximum data and the preset sub-data to the second adder 432 .

[0142] The second adder 432 is used for performing addition calculation on the exponent bit sub-data of the maximum data, the preset sub-data and the first inverted sub-data to obtain a target order difference between the first exponent bit sub-data and the second exponent bit sub-data.

[0143] According to an embodiment of the present disclosure, in the case where the floating-point operation is a floating-point addition, based on the operator instruction, the calculation process of each of the following structures in the first adder circuit 42 is as follows.

[0144] The third selector 428 is used to select the mantissa sub-data of the maximum data from the first mantissa sub-data and the second mantissa sub-data based on the input signal output by the consistency operator 49 and used to characterize the comparison result between the first operation data and the second operation data, wherein the maximum data is the largest of the first operation data and the second operation data.

[0145] The first shift operator 426 is used to shift the mantissa bit sub-data of the minimum data by a target order difference to obtain the first shifted sub-data, wherein the target order difference is used to characterize the order difference between the first exponent bit sub-data and the second exponent bit sub-data, and the target order difference is the output data of the second adder circuit 43.

[0146] The third XOR operator 429 is used for performing bitwise XOR on the sign bit sub-data of the first operation data and the sign bit sub-data of the second operation data to obtain first XOR sub-data.

[0147] The bit number expander 427 is used to expand the data bit number of the first XOR sub-data to a target bit to obtain extended sub-data.

[0148] The first XOR operator 424 is used to perform bitwise XOR on the first shifted sub-data and the extended sub-data to obtain target input sub-data.

[0149] The second selector 422 is used for outputting the target input sub-data to the first adder 423 .

[0150] The second XOR operator 425 is used for performing bitwise XOR on the sign bit sub-data of the first operation data and the sign bit sub-data of the second operation data to obtain second XOR sub-data.

[0151] The first adder 423 is used to add the second XOR sub-data, the target input sub-data and the mantissa sub-data of the minimum data to obtain intermediate mantissa sub-data, wherein the minimum data is the smallest of the first operation data and the second operation data, and the mantissa sub-data of the minimum data is the output data of the third selector 428.

[0152] According to an embodiment of the present disclosure, the maximum data is the largest of the two operation data, and the minimum data is the smallest of the two operation data.

[0153] According to an embodiment of the present disclosure, the preset sub-data may be 1'b1.

[0154] According to an embodiment of the present disclosure, based on the input signal of the comparison result between the first operation data and the second operation data, the exponent bit sub-data of the minimum data can be selected and inverted, and the exponent bit sub-data of the maximum data can be selected, and the inverted exponent bit sub-data of the minimum data and the exponent bit sub-data of the maximum data can be added to obtain the target order difference between the first operation data and the second operation data.

[0155] According to an embodiment of the present disclosure, the minimum data is shifted by the target order difference, so as to align the order of the minimum data and the maximum data. For example, when calculating the addition of 8.0 and 4.0, the mantissas of both are 8'b10000000. During the addition process, the mantissa of 4.0 needs to be shifted. After the mantissa part representing 4.0 is shifted to 8'b01000000, it is added to the mantissa part representing 8.0, 8'b100000000.

[0156] According to the embodiments of the present disclosure, whether to subtract the mantissa of the data with smaller absolute value can be determined according to whether the sign bits of the two operation data are the same. If the sign bits are different, when calculating the mantissa of the addition result, the shift result of the mantissa of the data with smaller absolute value is subtracted from the data with larger absolute value; if the sign bits are the same, the addition result of the two after shifting can be calculated.

[0157] According to the embodiments of the present disclosure, specifically, the third XOR operator 429 can perform XOR processing on the sign bits of the two operands, that is, data1

[15] xor data2

[15] ; and the XOR result is expanded by 8 bits, and then XORed with the shift result of the mantissa of the smallest data, so that if the sign bits of the operands of the two operation data are opposite, that is, subtraction processing is required, the mantissa of the smaller operand is XORed to all 1 to complete the inversion operation. If the sign bits are the same, addition processing is required, and the mantissa of the smaller operand is XORed to all 0, that is, it remains unchanged. The output result of the first XOR operator 424 is added to the mantissa bit sub-data of the largest data, and the carry input of the adder is the XOR result of the third XOR operator 429. Thus, if the sign bits of the two are opposite, subtraction processing is performed, that is, the mantissa shift result of the smaller operand is inverted and then 1 is added, and this 1 is reflected in the carry input of the adder; if the sign bits are the same, addition processing is performed, that is, the carry input remains 0.

[0158] According to the embodiments of the present disclosure, the above-mentioned processing procedure can be used to avoid performing more data selection processing based on the sign bit result, thereby eliminating multiple selectors. For example, there is no need to add an additional selector in front of the adder after the mantissa shift processing, thereby saving area and timing overhead and shortening the critical path.

[0159] FIG6( c ) schematically shows a third reconstruction diagram of the arithmetic logic subunit of the coarse-grained reconfigurable array system for deep learning when performing floating-point addition according to an embodiment of the present disclosure.

[0160] According to the embodiment of the present disclosure, in the case where the floating-point operation is floating-point addition, the calculation process of each of the following structures in the third adder circuit 45 is as follows.

[0161] The second negation operator 451 is used to invert the shift data output by the normalization circuit 44 to obtain second negated sub-data, and is also used to invert the second XOR sub-data output by the second XOR operator 425 to obtain third negated sub-data, wherein the shift data represents the number of shift bits of the middle mantissa sub-data output by the first adder circuit 42.

[0162] The tenth selector 455 is used to output the exponential bit sub-data of the maximum data.

[0163] The eleventh selector 456 is used for outputting a third preset value.

[0164] The third adder 454 is used to add the second negated sub-data, the exponent bit sub-data of the maximum data and the third preset value to obtain a third exponent bit result.

[0165] The AND logic operator 453 is used to perform AND logic processing on the intermediate mantissa bit sub-data output by the first adder circuit 42 and the third inverted sub-data to obtain processed sub-data.

[0166] The twelfth selector 457 is used to determine the output exponent bit result from the third exponent bit result and the preset input based on the operation sub-data and the processing sub-data, wherein the preset input is the sum of the exponent bit sub-data of the maximum data and the fourth preset value.

[0167] The second shift operator 459 is used to shift the sub-data with data addresses in the second range in the middle mantissa sub-data based on the processed sub-data to obtain second shifted sub-data.

[0168] The thirteenth selector 458 is used to select and output the mantissa bit result from the sub-data with data addresses in the third range in the second shifted sub-data and the intermediate mantissa bit sub-data based on the processed sub-data.

[0169] According to an embodiment of the present disclosure, the second negation operator 451 may be reused to achieve the negation of the shifted data and the second XOR sub-data respectively.

[0170] According to an embodiment of the present disclosure, data1_cmp_data2[14:7] is the exponential bit sub-data of the maximum data, the third preset value may be binary 1, the preset input may be data1_cmp_data2[14:7]+1b1', which represents adding 1 to the exponential bit sub-data of the maximum data, and the fourth preset value may be binary 1.

[0171] According to an embodiment of the present disclosure, the normalization circuit 44 is used to ensure that the highest bit in the final output mantissa bit sub-data is 1, that is, to maintain The format of the intermediate mantissa bit data output by the first adder 423 is determined by the normalization circuit 44 to determine how many bits need to be shifted to keep it For example, if 8'b01000000 is input, the normalization circuit 44 outputs 1, and 8'b01000000 is shifted left by 1 bit. The shifted data output by the normalization circuit 44 can be represented by norm_code.

[0172] According to the embodiment of the present disclosure, when determining the output exponent bit sub-data of floating-point multiplication, if the shift data is to shift the intermediate mantissa bit sub-data, i.e., psum18, left by n bits, then the exponent bit data of the maximum data should be subtracted by n; and if the shift data is to carry the intermediate mantissa bit data, then the exponent bit sub-data of the maximum data should be added by 1. Thus, the exponent bit sub-data of the above two schemes are calculated respectively, and whether there is a carry is determined based on the processed sub-data output by the AND logic operator 453, and the output exponent bit result is determined from the exponent bit sub-data outputted by the above two schemes respectively.

[0173] According to an embodiment of the present disclosure, a mantissa result is selected from the shift result of the intermediate mantissa sub-data and from the sub-data whose data address is in the third range in the intermediate mantissa sub-data by determining whether there is a carry.

[0174] According to an embodiment of the present disclosure, the third range may be the first 7 bits of the middle mantissa sub-data.

[0175] According to the embodiments of the present disclosure, the arithmetic logic subunit 32 of the present disclosure can realize parallel processing for different sign bit situations, and obtain the final output result through selector selection, thereby reducing the critical path delay.

[0176] FIG6( d ) schematically shows a fourth reconstruction diagram of the arithmetic logic subunit of the coarse-grained reconfigurable array system for deep learning when performing floating-point addition according to an embodiment of the present disclosure.

[0177] As shown in FIG6(d), during floating-point addition, the final output sign bit result is output by the sign bit selector 50, the output exponent bit result and the output mantissa bit result are output by the twelfth selector 457 and the thirteenth selector 458 respectively, and the above three output results are all received and integrated by the result selector 47.

[0178] According to an embodiment of the present disclosure, the outputs of the adders in FIG. 6 (a) to FIG. 6 (c) are all obtained by selection by selectors. For the convenience of drawing, some selectors are not drawn.

[0179] Figure 7 A schematic diagram of a normalized circuit of a coarse-grained reconfigurable array system for deep learning according to the present disclosure is schematically shown.

[0180] like Figure 7 As shown, the normalization circuit 44 includes: a fourteenth selector 441 , a fifteenth selector 442 , a first OR logic operator 223 , a second OR logic operator 444 , a third OR logic operator 445 and a non-logic operator group 446 .

[0181] The first OR logic operator 223 is used to perform logic OR processing on the sub-data with data addresses in the fourth range in the middle mantissa sub-data to obtain a first selection signal, wherein the middle mantissa sub-data is the output data of the first adder circuit 42.

[0182] The fourteenth selector 441 is used to determine the first output data from the sub-data with data addresses in the fifth range and the sub-data with data addresses in the sixth range in the middle mantissa bit sub-data based on the first selection signal.

[0183] The second OR logic operator 444 is used to perform a logic OR process on the first output data to obtain a second selection signal.

[0184] The fifteenth selector 442 is used to determine the second output data from the sub-data with data addresses in the seventh range and the sub-data with data addresses in the eighth range in the first output data based on the second selection signal.

[0185] The third OR logic operator 445 is used to perform a logic OR process on the second output data to obtain a third selection signal.

[0186] The NOT logic operator group 446 is used to convert the first selection signal, the second selection signal and the third selection signal respectively to obtain shift data.

[0187] According to an embodiment of the present disclosure, the normalization circuit 44 can process the intermediate mantissa bit data into a mantissa form in a standard BF16 format. Figure 7As shown, when the middle mantissa sub-data Data[7:0] is normalized, Mux0[3:0] can be selected from the middle mantissa sub-data Data[7:4] and Data[3:0] according to the instruction signal corresponding to the fourteenth selector 441 in the operation sub-instruction and the first selection signal, and Mux0[1:0] can be determined from Mux0[3:2] and Mux0[1:0].

[0188] According to an embodiment of the present disclosure, the NOT logic operator group 446 may include a plurality of NOT logic operators, and the shift data is obtained by inputting the first selection signal, the second selection signal and the third selection signal into corresponding NOT logic operators respectively.

[0189] According to an embodiment of the present disclosure, for example, when 8'b01000000 is input, the shift data Result=3'b001 is obtained, which means that only 1 bit shift is required to convert it into a standard mantissa form of 8'b10000000, and the final output mantissa is 7'b0000000.

[0190] According to the embodiments of the present disclosure, the coarse-grained reconfigurable array system for deep learning disclosed in the present disclosure can realize floating-point data multiplication and addition calculations and integer data multiplication and accumulation calculations, and reuse the hardware resources for integer calculations as much as possible to complete the multiplication and addition operations of formatted floating-point data, thereby reducing hardware overhead, and all calculation results are obtained within one cycle. Flexible configuration of data flow is used to enable each processing unit to support instructions suitable for calculations of different data types, and the number of instructions executable by the processing unit is increased to perform multi-precision operations without increasing the instruction length and the instruction logic control cost.

[0191] According to the embodiments of the present disclosure, the coarse-grained reconfigurable array system for deep learning of the present disclosure, in BF16 multiplication and addition, by reusing integer multipliers and existing integer adders, and adding a normalization module, completes BF16 format floating-point multiplication and addition for deep learning under the general CGRA architecture, and the accuracy is consistent with the C language compiler. And the implementation efficiency is greatly improved compared with the instruction reconstruction method alone.

[0192] According to the embodiments of the present disclosure, the coarse-grained reconfigurable array system for deep learning of the present disclosure realizes the above-mentioned multi-precision operations with only 3.9% additional hardware overhead. The dynamic configuration of the instruction type executed by each processing unit is satisfied without increasing the instruction width, that is, without increasing the instruction data storage cost.

[0193] Figure 8 A schematic diagram of an arithmetic logic subunit of a coarse-grained reconfigurable array system for deep learning according to another embodiment of the present disclosure is schematically shown.

[0194] like Figure 8 As shown, the arithmetic logic subunit 32 includes: 4 8-bit multipliers, two 16-bit adders, a 17-bit adder and a 35-bit adder, a logic operator group 48, multiple selectors and multiple logic operators. The basic_logic module is a logic operator group 48 including modules such as shift, XOR, and inverter.

[0195] According to an embodiment of the present disclosure, Figure 8 Schematically shows the circuits or devices included in the arithmetic logic subunit 32, as well as the related devices included in the circuit and the connection relationship between the related devices. The arithmetic logic subunit 32 includes a multiplier group 41, a first adder circuit 42, a second adder circuit 43, a normalization circuit 44, a third adder circuit 45, a fourth adder circuit 46, a result selector 47, a logic operator group 48, a consistency operator 49 and a sign bit selector 50.

[0196] According to an embodiment of the present disclosure, the fourth adder circuit 46 may include a sixteenth selector 461 and a fourth adder 462 .

[0197] According to an embodiment of the present disclosure, Figure 8 The connection relationship between individual devices is not marked. If the input of some devices is the output of other devices, it can be considered that there is a connection relationship between the two devices. For example, psuml8 is the output of the first adder 423 when performing floating-point addition, so the devices with input psuml8 are all connected to the output of the first adder 423.

[0198] Fig. 9 A flowchart of a coarse-grained reconfigurable array computing method for deep learning according to an embodiment of the present disclosure is schematically shown.

[0199] like Fig. 9 As shown, the method includes: operation S910 to operation S950.

[0200] In operation S910, the controller 11 determines input information to at least one processing unit, wherein the input information includes data to be calculated, operation instructions, and mode control instructions.

[0201] In operation S920 , data to be calculated is input to at least one processing unit through the input bus 15 .

[0202] In operation S930 , an operation instruction is input to at least one processing unit through the configuration bus 16 .

[0203] In operation S940, each processing unit performs floating point operation or integer operation on the data to be calculated according to the mode control instruction and the operation instruction to obtain an operation result.

[0204] In operation S950 , the operation result is output through the output bus 17 .

[0205] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level process and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, Java, C++, python, "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on the remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect through the Internet).

[0206] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0207] It will be appreciated by those skilled in the art that the features described in the various embodiments of the present disclosure may be combined and / or combined in a variety of ways, even if such combinations or combinations are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in the various embodiments of the present disclosure may be combined and / or combined in a variety of ways. All of these combinations and / or combinations fall within the scope of the present disclosure.

[0208] The embodiments of the present disclosure are described above. However, these embodiments are only for illustrative purposes and are not intended to limit the scope of the present disclosure. Although the embodiments are described above, this does not mean that the measures in the various embodiments cannot be used in combination to advantage. Without departing from the scope of the present disclosure, those skilled in the art may make a variety of substitutions and modifications, which should all fall within the scope of the present disclosure.

Claims

1. A coarse-grained reconfigurable array system for deep learning, characterized in that: The system comprises: A controller, configured to determine input information to be input to at least one processing unit, wherein the input information includes data to be calculated, operation instructions, and mode control instructions; A processing unit group, comprising a plurality of the processing units, wherein the plurality of the processing units form a reconfigurable array, and each of the processing units is used to perform floating-point operations or integer operations on the data to be calculated based on the mode control instruction and the operation instruction to obtain an operation result; Wherein, the data to be calculated includes input data and weight data, and the processing unit includes: an operand selector and an arithmetic logic subunit; The operand selector is configured to receive a neighbor operation result output by a neighbor processing unit, and determine first operation data and second operation data from the neighbor operation result, the weight data and the input data based on the operation instruction, wherein the neighbor processing unit is a processing unit adjacent to the processing unit; The arithmetic logic subunit is used to determine the operation sub-instruction corresponding to the mode control instruction from the operation instruction, so as to perform floating-point operation or integer operation on the first operation data and the second operation data based on the operation sub-instruction to obtain an operation result.

2. The system according to claim 1, characterized in that The arithmetic logic subunit comprises: The first adder circuit, wherein a first input terminal of the first adder circuit is connected to a first multiplier, a second input terminal of the first adder circuit is connected to an output terminal of a second adder circuit, a third input terminal of the first adder circuit is connected to the second multiplier, and an output terminal of the first adder circuit is connected to an input terminal of a normalization circuit, a first input terminal of the third adder circuit, and an input terminal of a result selector, respectively; The normalization circuit, the output end of the normalization circuit is connected to the second input end of the third adder circuit; the second adder circuit, wherein a first input terminal of the second adder circuit is connected to the third multiplier, a second input terminal of the second adder circuit is connected to the fourth multiplier, and an output terminal of the second adder circuit is respectively connected to the second input terminal of the first adder circuit, the third input terminal of the third adder circuit, the first input terminal of the fourth adder circuit and an input terminal of the result selector; The third adder circuit, the output end of the third adder circuit is connected to the result selector. The fourth adder circuit, the output end of the fourth adder circuit is connected to the result selector; The first adder circuit, the second adder circuit, the third adder circuit and the fourth adder circuit respectively include at least one selector and an adder, and the input of the adder is obtained by the at least one selector based on the operation sub-instruction.

3. The system according to claim 2, characterized in that The arithmetic logic subunit further includes: a consistency operator; the first adder circuit includes: a first selector, a second selector, a first adder, a first XOR operator, a second XOR operator, a first shift operator, a bit number expander, a third selector and a third XOR operator; The first input terminal of the first selector is connected to the output terminal of the first multiplier, and the output terminal of the first selector is connected to the first input terminal of the first adder; The first input end of the second selector is connected to the output end of the first XOR operator, the second input end of the second selector is connected to the second multiplier, and the output end of the second selector is connected to the second input end of the first adder; The first input end of the first XOR operator is connected to the output end of the first shift operator, and the second input end of the first XOR operator is connected to the output end of the bit expander; The output terminal of the second XOR operator is connected to the third input terminal of the first adder; The output terminal of the first adder is connected to the input terminal of the AND normalization circuit and the first input terminal of the third adder circuit respectively; The first input end of the first shift operator is connected to the third selector, and the second output end of the first shift operator is connected to the output end of the second adder circuit; The first input terminal of the third selector is connected to the output terminal of the consistency operator; The input end of the bit number expander is connected to the second input end of the third XOR operator.

4. The system according to claim 3, characterized in that In the case where the floating-point operation is a floating-point multiplication, The first multiplier is used to perform a multiplication operation on the first mantissa sub-data in the first operation data and the second mantissa sub-data in the second operation data to obtain a multiplication result, wherein the first mantissa sub-data and the second mantissa sub-data are output data of two selectors connected to the input end of the first multiplier; The first selector is used to output the first index bit sub-data in the first data; The second selector is used to output the second index bit sub-data in the second data; The first adder is used to perform addition calculation on the first exponent bit sub-data and the second exponent bit sub-data to obtain a first exponent bit result; The second XOR operator is used to perform bitwise XOR on the sign bit sub-data of the first operation data and the sign bit sub-data of the second operation data to obtain an output sign bit result; In the case where the floating-point operation is a floating-point addition, The third selector is used to select the mantissa digit sub-data of the maximum data from the first mantissa digit sub-data and the second mantissa digit sub-data based on the input signal output by the consistency operator and used to represent the comparison result between the first operation data and the second operation data, wherein the maximum data is the largest one between the first operation data and the second operation data; The first shift operator is used to shift the mantissa bit sub-data of the minimum data by a target order difference to obtain first shifted sub-data, wherein the target order difference is used to represent the order difference between the first exponent bit sub-data and the second exponent bit sub-data, and the target order difference is the output data of the second adder circuit; The third XOR operator is used to perform bitwise XOR on the sign bit sub-data of the first operation data and the sign bit sub-data of the second operation data to obtain first XOR sub-data; The bit number expander is used to expand the data bit number of the first XOR sub-data to a target bit to obtain extended sub-data; The first XOR operator is used to perform bitwise XOR on the first shifted sub-data and the extended sub-data to obtain target input sub-data; The second selector is used to output the target input sub-data to the first adder; The second XOR operator is used to perform bitwise XOR on the sign bit sub-data of the first operation data and the sign bit sub-data of the second operation data to obtain second XOR sub-data; The first adder is used for adding the second XOR sub-data, the target input sub-data and the mantissa sub-data of the minimum data to obtain intermediate mantissa sub-data, wherein the minimum data is the smallest of the first operation data and the second operation data, and the mantissa sub-data of the minimum data is the output data of the third selector.

5. The system according to claim 2, characterized in that The arithmetic logic subunit further includes: a consistency operator; the second adder circuit includes: a fourth selector, a second adder, a first negation operator, a fifth selector, a sixth selector, a seventh selector and an eighth selector; The first input end of the fourth selector is connected to the output end of the third multiplier, the second input end of the fourth selector is connected to the output end of the first negation operator, and the output end of the fourth selector is connected to the first input end of the second adder; The input end of the first negation operator is connected to the output end of the fifth selector; The input end of the fifth selector is connected to the output end of the consistency operator; The first input terminal of the sixth selector is connected to the fourth multiplier, and the output terminal of the sixth selector is connected to the second input terminal of the second adder; The input end of the seventh selector is connected to the output end of the consistency operator, and the output end of the seventh selector is connected to the input end of the eighth selector; An output terminal of the eighth selector is connected to a third input terminal of the second adder.

6. The system according to claim 5, characterized in that In the case where the floating-point operation is a floating-point addition, the fifth selector is configured to select the exponent bit sub-data of the smallest data from the first exponent bit sub-data of the first operation data and the second mantissa bit sub-data of the second operation data based on the input signal output by the consistency operator and used to represent the comparison result between the first operation data and the second operation data; The first negation operator is used to perform negation processing on the exponent bit sub-data of the minimum data to obtain first negated sub-data; The fourth selector is used to output the first inverted sub-data to the second adder; The sixth selector is used to output empty data; The seventh selector is used to select the exponent bit sub-data of the maximum data from the first exponent bit sub-data and the second mantissa bit sub-data based on the input information; the eighth selector is used for outputting the exponential bit sub-data and the preset sub-data of the maximum data to the second adder; The second adder is used to perform addition calculation on the exponent bit sub-data of the maximum data, the preset sub-data and the first negated sub-data to obtain a target order difference between the first exponent bit sub-data and the second exponent bit sub-data.

7. The system according to claim 2, characterized in that The third adder circuit comprises: a second negation operator, a ninth selector, an AND logic operator, a third adder, a tenth selector, an eleventh selector, a twelfth selector, a thirteenth selector and a second shift operator; The input end of the second negation operator is connected to the output end of the normalization circuit and the output end of the second XOR operator in the first adder circuit respectively, and the output end of the second negation operator is connected to the first input end of the ninth selector and the first input end of the AND logic operator respectively; The second input terminal of the ninth selector is connected to the output terminal of the first adder circuit, and the output terminal of the ninth selector is connected to the first input terminal of the third adder; The second input terminal of the third adder is connected to the output terminal of the tenth selector, the third input terminal of the third adder is connected to the output terminal of the eleventh selector, and the output terminal of the third adder is connected to the first input terminal of the twelfth selector; The second input terminal of the AND logic operator is connected to the output terminal of the first adder circuit, and the output terminal of the AND logic operator is connected to the first input terminal of the thirteenth selector and the second input terminal of the twelfth selector respectively; The second input terminal of the thirteenth selector is connected to the output terminal of the second shift operator; A first input terminal of the second shift operator is connected to an output terminal of the normalization circuit, and a second input terminal of the second shift operator is connected to an output terminal of the first adder circuit.

8. The system according to claim 7, characterized in that In the case where the floating-point operation is a floating-point multiplication, The ninth selector is used to receive the first exponent bit result output by the first adder circuit, and to output the first exponent bit result; The tenth selector is used to output a second preset value; The eleventh selector is used to receive the multiplication result output by the first multiplier, and to output a sub-result of the multiplication result whose data address is in the first range; The third adder is used to perform addition calculation on the first exponent bit result, the second preset value and the sub-result to obtain an output exponent bit result; In the case where the floating-point operation is a floating-point addition, The second negation operator is used to negate the shift data output by the normalization circuit to obtain second negated sub-data, and is also used to negate the second XOR sub-data output by the second XOR operator to obtain third negated sub-data, wherein the shift data represents the number of shift bits of the intermediate mantissa bit sub-data output by the first adder circuit; The tenth selector is used to output the exponential bit sub-data of the maximum data; The eleventh selector is used to output a third preset value; The third adder is used to add the second negated sub-data, the exponent bit sub-data of the maximum data and the third preset value to obtain a third exponent bit result; The AND logic operator is used to perform AND logic processing on the intermediate mantissa bit sub-data output by the first adder circuit and the third inverted sub-data to obtain processed sub-data; The twelfth selector is used to determine an output exponent bit result from the third exponent bit result and a preset input based on the operation sub-data and the processing sub-data, wherein the preset input is a sum of the exponent bit sub-data of the maximum data and a fourth preset value; The second shift operator is used to shift the sub-data with data addresses in the second range in the intermediate mantissa sub-data based on the processed sub-data to obtain second shifted sub-data; The thirteenth selector is used to select and output the mantissa result from the sub-data with data addresses in the third range in the second shifted sub-data and the intermediate mantissa sub-data based on the processed sub-data.

9. The system according to claim 2, characterized in that The normalization circuit includes: a fourteenth selector, a fifteenth selector, a first OR logic operator, a second OR logic operator, a third OR logic operator and a non-logic operator group; The first OR logic operator is used to perform a logic OR process on the sub-data with data addresses in the fourth range in the intermediate mantissa sub-data to obtain a first selection signal, wherein the intermediate mantissa sub-data is output data of the first adder circuit; The fourteenth selector is used to determine the first output data from the sub-data with data addresses in the fifth range and the sub-data with data addresses in the sixth range in the intermediate mantissa bit sub-data based on the first selection signal; The second OR logic operator is used to perform a logic OR process on the first output data to obtain a second selection signal; The fifteenth selector is used to determine the second output data from the sub-data with data addresses in the seventh range and the sub-data with data addresses in the eighth range in the first output data based on the second selection signal; The third OR logic operator is used to perform a logic OR process on the second output data to obtain a third selection signal; The NOT logic operator group is used to convert the first selection signal, the second selection signal and the third selection signal respectively to obtain shift data.

10. A coarse-grained reconfigurable array computing method for deep learning, comprising: Determining, by a controller, input information to be input to at least one processing unit, wherein the input information includes data to be calculated, operation instructions, and mode control instructions; Inputting the data to be calculated into at least one of the processing units via an input bus; Inputting the operation instruction to at least one of the processing units via a configuration bus; By each of the processing units performing floating point operations or integer operations on the data to be calculated in accordance with the mode control instruction and the operation instruction, a calculation result is obtained; The operation result is output through an output bus.

Citation Information

Patent Citations

  • Coarse-grained reconfigurable array system for deep learning and calculation method

    CN115168284A

  • Coarse-grained reconfigurable array simulator system for deep learning and calculation method

    CN116187434A

  • Coarse-grained reconfigurable array operator design method and system for deep learning

    CN116301892A

  • Single-precision floating point arithmetic device

    CN116382618A

  • Reconfigurable method and system supporting multi-precision floating point or fixed point operation

    CN116627379A