A neural network computing terminal and a neural network computing method
By adopting the design of multiplexed addition modules in the neural network computing terminal, the problems of high computational complexity and high power consumption in the embedded system are solved, and the hardware area is reduced and the computing efficiency is improved.
Patent Information
- Application Number
- CN202210295538.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-24
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-03-24
AI Technical Summary
The existing convolutional neural network computing hardware has high computational complexity and high power consumption in embedded systems, making it difficult to optimize the hardware area.
A neural network computing terminal is designed, using N input channels and N computing arrays. Each input channel corresponds to a computing array. The computing array includes multiple data processing units. The data processing unit includes a multiplication module and an addition module. Multiplication and addition operations are performed by multiplexing the addition module to reduce hardware area consumption.
It realizes that within the power consumption limit range, it ensures both operation accuracy and meets multiple computing modes, improves the utilization rate and throughput of the computing terminal, and reduces hardware area consumption.
Smart Images

Figure CN114781627B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of convolutional neural network hardware design, and particularly relates to a neural network computing terminal and a neural network computing method. Background Art
[0002] Currently, neural network research has become a research hotspot, and convolutional neural network models are applied to various fields, such as image recognition, video coding, action recognition, and speech recognition. The superior effects demonstrated by convolutional neural networks are becoming more and more obvious, and the number of parameters or weights that can be processed is also increasing, reaching up to millions or more. As a result, the requirements for computational complexity are getting higher and higher, and the requirements for the computing power of hardware are also very high. Nowadays, some high-performance GPUs (i.e., Graphics Processing Unit, graphics processors) can already meet these performance requirements, but the power consumption is quite large, and the area optimization cannot be better, so it is difficult to be applied to embedded systems. Summary of the Invention
[0003] In view of this, the purpose of this application is to provide a neural network computing terminal and a neural network computing method, which can reduce the hardware area consumption of the computing terminal. The specific solutions are as follows:
[0004] In a first aspect, this application discloses a neural network computing terminal, including N input channels and N computing arrays. Among them, each of the input channels corresponds to one of the computing arrays, and each of the computing arrays includes multiple data processing units, and each of the data processing units includes a multiplication module and an addition module;
[0005] The multiplication module is used to perform a multiplication operation on the data input by the input channel, and output the multiplication operation result of the multiplication operation to the addition module;
[0006] The addition module is used to perform an addition operation on the multiplication operation result to obtain an addition operation result, and input multiple addition operation results to the addition module for addition operation or use the addition operation result as the target data of the computing terminal.
[0007] Optionally, the computing terminal further includes a selector module, the selector module includes multiple selection terminals, and the connection terminals of the multiple selection terminals include the output terminal of the multiplication module, the input terminal of the input channel, and the output terminal of the addition module.
[0008] Optionally, the input channel includes a first input terminal and a second input terminal, and the multiplication module is used to perform a multiplication operation on the first data of the first input terminal and the second data of the second input terminal to obtain a multiplication operation result.
[0009] Optionally, the adder modules in any of the computing arrays are connected into an adder hierarchical structure, and the adder hierarchical structure includes a first layer, an intermediate layer, and an output layer;
[0010] The adder modules in the first layer are connected to the output ends of the multiplication modules, the input ends of the adder modules in the intermediate layer are connected to the output ends of the adder modules in the first layer, and the input ends of the adder modules in the output layer are connected to the output ends of the adder modules in the intermediate layer;
[0011] The adder modules in the first layer are configured to receive the multiplication operation results, perform an addition operation on the multiplication operation results to obtain first-layer addition operation results, and output multiple first-layer addition operation results to the adder modules in the intermediate layer;
[0012] The adder modules in the intermediate layer are configured to receive the first-layer addition operation results, perform an addition operation on the first-layer addition operation results to obtain intermediate-layer addition operation results;
[0013] The adder modules in the output layer are configured to receive the intermediate-layer addition operation results, perform an accumulation operation on the intermediate-layer addition operation results, and output the final operation result as the target data of the computing terminal.
[0014] Optionally, the intermediate layer includes multiple levels of adder modules. Any adder module in any level of the multiple levels of adder modules is configured to receive the addition operation results output by two adder modules in the previous level of the current adder module, perform an addition operation on the two addition operation results to obtain corresponding addition operation results, and the addition operation result output by the last level of adder module in the multiple levels of adder modules is the intermediate-layer addition operation result.
[0015] Optionally, the computing array further includes a register configuration unit, and the register configuration unit is configured to determine different operation types according to the register values of the register configuration unit.
[0016] Optionally, the data processing unit further includes a subtractor connected to the input channel;
[0017] The subtractor is configured to perform a subtraction operation processing for zero drift on the input data of the input channel.
[0018] Optionally, the data processing unit further includes:
[0019] A first register connected to the output end of the multiplication module, configured to store the multiplication operation results output by the multiplication module;
[0020] A second register connected to the output end of the adder module, configured to store the addition operation results output by the adder module.
[0021] In a second aspect, the present application discloses a neural network computing method, which is applied to a neural network computing terminal. The neural network computing terminal includes N input channels and N computing arrays. Among them, each input channel corresponds to one computing array, and each computing array includes multiple data processing units. Each data processing unit includes an addition module. The computing method includes:
[0022] Performing a multiplication operation on the data input to the input channel, and outputting the multiplication result of the multiplication operation to the addition module;
[0023] Performing an addition operation on the multiplication result to obtain an addition result, and inputting multiple addition results to the addition module or using the addition result as the target data of the computing terminal.
[0024] Optionally, the addition modules in any one of the computing arrays are connected into an addition hierarchical structure, and the addition hierarchical structure includes a first layer, an intermediate layer, and an output layer;
[0025] The addition module of the first layer receives the multiplication result, performs an addition operation on the multiplication result to obtain a first-layer addition result, and outputs multiple first-layer addition results to the addition module of the intermediate layer;
[0026] The addition module of the intermediate layer receives the first-layer addition result, performs an addition operation on the first-layer addition result to obtain an intermediate-layer addition result;
[0027] The addition module of the output layer receives the intermediate-layer addition result, performs an accumulation operation on the intermediate-layer addition result, and outputs the final operation result as the target data of the computing terminal.
[0028] It can be seen that a neural network computing terminal disclosed in the present application includes N input channels and N computing arrays. Among them, each input channel corresponds to a computing array, and the computing array includes multiple data processing units. Each data processing unit includes a multiplication module and an addition module. The multiplication module is used to perform a multiplication operation on the data input by the input channel and output the multiplication operation result of the multiplication operation to the addition module. The addition module is used to perform an addition operation on the multiplication operation result to obtain an addition operation result, and input multiple addition operation results into the addition module for an addition operation or use the addition operation result as the target data of the computing terminal. That is, in the present application, the addition module in each data processing unit can perform an addition operation on the multiplication operation result obtained by the multiplication module to obtain the target data of the computing terminal, and can also perform an addition operation on the operation result input by the input addition module to obtain the target data of the computing terminal, realizing the reuse of the addition module, thereby reducing the hardware area consumption of the computing terminal. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0030] Figure 1 Schematic diagram of a neural network computing terminal disclosed in the present application;
[0031] Figure 2 Schematic diagram of a data processing unit disclosed in the present application;
[0032] Figure 3 Schematic diagram of a computing array in a specific neural network computing terminal disclosed in the present application;
[0033] Figure 4 Schematic diagram of a specific data processing unit disclosed in the present application;
[0034] Figure 5 Schematic diagram of a specific fully connected operation disclosed in the present application;
[0035] Figure 6 Flowchart of a neural network computing method disclosed in the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0037] See Figure 1 As shown, an embodiment of the present application discloses a neural network computing terminal, including N input channels and N computing arrays. Among them, each of the input channels corresponds to one of the computing arrays, and each computing array includes a plurality of data processing units, and each data processing unit includes a multiplication module 001 and an addition module 002;
[0038] Among them, the computing array can be a MAC (i.e., Multiply Accumulate) array.
[0039] The multiplication module 001 is configured to perform a multiplication operation on the data input by the input channel, and output the multiplication result of the multiplication operation to the addition module.
[0040] Specifically, the input channel may include one or more input ends. In some embodiments, if the input channel has only one input end, the input end data of two input channels is input to a multiplication module for multiplication calculation.
[0041] In this embodiment, the input channel is a data channel with two input ends. Specifically, the input channel includes a first input end and a second input end, and the multiplication module is configured to perform a multiplication operation on the first data of the first input end and the second data of the second input end to obtain a multiplication result.
[0042] It should be noted that the embodiments of the present application can reuse the multiplication module for multiplication operations in point-to-point multiplication operations, ordinary convolution operations, and fully connected operations. Moreover, in point-to-point operations and fully connected operations, the first data and the second data can be two different feature pictures. In ordinary convolution operations, the first data can be a feature picture, and the second data can be a convolution kernel. It can be understood that the reuse of the multiplication module further reduces the hardware area consumption of the computing terminal.
[0043] In some other embodiments, the input channel can also be a multi-input end channel with more than two input ends, which is not specifically limited herein.
[0044] The addition module 002 is configured to perform an addition operation on the multiplication operation result to obtain an addition operation result, and input multiple addition operation results to the addition module for addition operation or use the addition operation result as the target data of the computing terminal.
[0045] Further, the computing terminal further includes a selector module. The selector module includes multiple selection terminals, and the connection terminals of the multiple selection terminals include the output terminal of the multiplication module, the input terminal of the input channel, and the output terminal of the addition module.
[0046] In a specific embodiment, the selector module includes a first selector and a second selector, where
[0047] The first selector includes three selection terminals, namely a first selection terminal, a second selection terminal, and a third selection terminal. The first selection terminal is connected to the output terminal of the multiplication module of the current data processing unit and inputs the multiplication operation result of the multiplication module of the current data processing unit. The second selection terminal is connected to the first input terminal and inputs the first data input by the first input terminal. The third selection terminal is connected to the output terminal of other data processing units and inputs the multiplication operation result of the multiplication module of other data processing units or the addition operation result of the addition module of other data processing units.
[0048] The second selector includes three selection terminals, namely a fourth selection terminal, a fifth selection terminal, and a sixth selection terminal. The fourth selection terminal is connected to the output terminal of the addition module of the current data processing unit and inputs the addition operation result of the addition module of the current data processing unit. The fifth selection terminal is connected to the second input terminal and inputs the second data input by the second input terminal. The sixth selection terminal is connected to the output terminal of other data processing units and inputs the multiplication operation result of the multiplication module of other data processing units or the addition operation result of the addition module of other data processing units.
[0049] Moreover, the data processing unit further includes: a first register connected to the output terminal of the multiplication module for storing the multiplication operation result output by the multiplication module; a second register connected to the output terminal of the addition module for storing the addition operation result output by the addition module.
[0050] In a specific embodiment, the first selection terminal of the first selector is connected to the first register, and the fourth selection terminal of the second selector is connected to the second register.
[0051] Further, the computing array further includes a register configuration unit, which is configured to determine different operation types according to the register values of the register configuration unit. Where
[0052] The operation types include ordinary convolution operations. The selection result of the first selector is the multiplication operation result of the multiplication module, and the selection result of the second selector is the addition operation result of the addition module. The addition module is used to perform an addition operation based on the multiplication operation result and the addition operation result to implement ordinary convolution operation processing.
[0053] The operation types include point operations. The selection result of the first selector is the first data of the first input end, and the selection result of the second selector is the second data of the second input end. The addition module is used to perform an addition operation based on the first data and the second data to implement point addition operation processing.
[0054] Further, in any of the computing arrays, the addition modules are connected into an addition hierarchical structure. The addition hierarchical structure includes a first layer, an intermediate layer, and an output layer. The addition modules in the first layer are connected to the output ends of the multiplication modules. The input ends of the addition modules in the intermediate layer are connected to the output ends of the addition modules in the first layer. The input ends of the addition modules in the output layer are connected to the output ends of the addition modules in the intermediate layer. The addition modules in the first layer are used to receive the multiplication operation results, perform an addition operation on the multiplication operation results to obtain first-layer addition operation results, and output multiple first-layer addition operation results to the addition modules in the intermediate layer. The addition modules in the intermediate layer are used to receive the first-layer addition operation results, perform an addition operation on the first-layer addition operation results to obtain intermediate-layer addition operation results. The addition modules in the output layer are used to receive the intermediate-layer addition operation results, perform an accumulation operation on the intermediate-layer addition operation results to obtain a final operation result, and output the final operation result as the target data of the computing terminal.
[0055] Among them, the intermediate layer includes multiple levels of addition modules. Any addition module in any level of the multiple levels of addition modules is used to receive the addition operation results output by two addition modules in the previous-level addition module of this level of addition module, and perform an addition operation on the two addition operation results to obtain corresponding addition operation results. The addition operation result output by the last-level addition module in the multiple levels of addition modules is the intermediate-layer addition operation result.
[0056] Moreover, the two addition operation results received by any addition module in any level of the multiple levels of addition modules are not the addition operation results output by the same addition module as the two addition operation results received by other addition modules in the same level of addition module.
[0057] Correspondingly, the operation type includes fully connected operations. The selection result of the first selector connected to the addition module of the first layer is the multiplication operation result of the multiplication module of the current data processing unit, and the selection result of the second selector is the multiplication operation result of the multiplication module of other data processing units; the selection result of the first selector connected to the addition module of the middle layer is the addition operation result of the addition module of other data processing units, and the selection result of the second selector is the addition operation result of the addition module of other data processing units.
[0058] Moreover, the addition module of the first layer is configured to perform an addition operation based on the multiplication operation results of the current data processing unit and other data processing units to obtain a first addition operation result; input multiple first addition operation results into the addition module of the middle layer in pairs respectively, and use the multi-stage addition module of the middle layer to obtain the addition operation result of the middle layer. Then, input the addition operation result of the middle layer into the addition module of the output layer. The addition module of the output layer performs an addition and accumulation operation on the addition operation result of the middle layer to obtain a final operation result, which is output as the target data of the calculation terminal. Among them, the number of addition and accumulation operations is determined according to the input data size and the number of data processing units in the array.
[0059] That is to say, the operation type includes fully connected operations. The selection result of the first selector of the addition module of the first layer is the multiplication operation result of the multiplication module of the current data processing unit, and the selection result of the second selector is the multiplication operation result of the multiplication module of other data processing units; the addition module of the first layer is configured to perform an addition operation based on the multiplication operation results of the current data processing unit and other data processing units to obtain a first addition operation result; input multiple first addition operation results into the selection terminals of the first selector or the second selector of multiple other idle data processing units in pairs respectively, so that the addition modules of the multiple other idle data processing units perform an addition operation on the first addition operation result to obtain a second addition operation result; repeatedly use the addition modules of the multiple other idle data processing units to perform an addition operation on the second addition operation results in pairs until a final addition operation result is obtained. The N calculation arrays corresponding to the N data channels are processed in parallel, and each calculation array performs a self-accumulation operation on the final addition operation results of multiple beats to obtain a fully connected operation result.
[0060] That is, the embodiments of the present application can reuse the addition module for point-to-point addition operations, addition operations in ordinary convolution operations, and addition operations in fully connected operations. When performing point-to-point addition operations, the first data and the second data input from the input channels are subjected to point-to-point addition operations to obtain an addition operation result, which is used as the target data of the computing terminal. When performing ordinary convolution operations, the addition module is used to perform self-accumulation operations on the multiplication operation results to obtain an addition operation result, and the addition operation result is used as the target data of the computing terminal. When performing fully connected operations, addition operations are performed on the multiplication operation results output by the multiplication modules of the same data processing unit and the multiplication modules output by the multiplication modules of other data processing units, or addition operations are performed on the addition operation results of the addition modules of two other data processing units.
[0061] In addition, the calculation types of the computing terminal also include dot product operations, taking the maximum value, taking the minimum value operations, left shift and right shift operations, etc., all of which can be determined according to the configuration values of the registers.
[0062] See Figure 2 As shown, the embodiments of the present application disclose a schematic diagram of the structure of a data processing unit, including:
[0063] A multiplication module 101, configured to perform multiplication operation processing on the first data input from the first input end and the second data input from the second input end in the input channel to obtain a multiplication operation result; wherein, the multiplication operation processing may include multiplication operation processing corresponding to convolution operations, multiplication operation processing corresponding to fully connected operations, and point-to-point multiplication operation processing;
[0064] An addition module 102, configured to perform addition operations on the multiplication operation results to obtain an addition operation result, and input multiple addition operation results to the addition module for addition operations or use the addition operation result as the target data of the computing terminal. In a specific implementation manner, the multiplication operation results corresponding to convolution operations can be accumulated, corresponding addition operation processing can be performed on the multiplication operation results of multiple data processing units during fully connected operations, and point-to-point addition operation processing can be performed.
[0065] Further, the data processing unit further includes a subtractor connected to the input channel; the subtractor is configured to perform subtraction operation processing on the input data of the input channel for zero drift. It should be noted that subtracting zero drift can improve the calculation accuracy and reduce errors.
[0066] In a specific embodiment, the data processing unit further includes:
[0067] A first subtractor 103 connected to the first input end and a second subtractor 104 connected to the second input end;
[0068] The first subtractor 103 and the second subtractor 104 are respectively used to perform subtraction operation processing for zero drift on the first data and the second data.
[0069] It should be noted that two subtractors are used respectively to subtract the zero offset value, improve the accuracy and reduce the error.
[0070] Moreover, the data processing unit further includes a quantization processing unit connected to the subtractor, which is used to perform quantization processing on the output result of the subtractor when performing point-to-point addition operation, and no quantization processing is required for other types of operations. Among them, the quantization processing unit includes a left shift processing unit, a multiplier and a right shift processing unit. In this way, when performing point addition operation, the data of the two feature maps can be quantized into a unified quantization range.
[0071] In a specific embodiment, the data processing unit further includes a quantization processing unit, and the quantization processing unit specifically includes:
[0072] A first left shift processing unit 105 connected to the first subtractor 103, a first multiplier 106 connected to the first left shift processing unit, and a first right shift processing unit 107;
[0073] A second left shift processing unit 108 connected to the second subtractor 104, a second multiplier 109 connected to the second left shift processing unit, and a second right shift processing unit 110;
[0074] Among them, the first left shift processing unit 105 and the second left shift processing unit 108 are respectively used to perform left shift operations on the processing results of the first subtractor 103 and the second subtractor 104, the first multiplier 106 and the second multiplier 109 are respectively used to perform multiplication operations on the left shift operation results of the first left shift processing unit 105 and the second left shift processing unit 108 and a preset quantization coefficient, and the first right shift processing unit 107 and the second right shift processing unit 110 are respectively used to perform right shift operations on the operation results of the first multiplier 106 and the second multiplier 109.
[0075] It should be noted that after subtracting the zero offset, left shift -> multiply -> right shift is performed. This operation is only applied to specific point-to-point addition operation processing, in order to make the data of the two feature maps quantized into a unified quantization range during point addition operation. If it is not specific point-to-point addition operation processing, the left shift -> multiply -> right shift operation is not performed. Specifically, it can be controlled whether to execute the left shift -> multiply -> right shift operation through a bypass.
[0076] Correspondingly, the addition module 102 can be specifically configured to perform point-to-point addition operation processing on the right shift operation results of the first right shift processing unit and the second right shift processing unit.
[0077] In a specific embodiment, the multiplication module 101 is specifically configured to perform corresponding multiplication operation processing on the processing results of the first subtractor 103 and the second subtractor 104, including multiplication operation processing corresponding to convolution operation, multiplication operation processing corresponding to fully connected operation, and point-to-point multiplication operation processing.
[0078] Furthermore, the data processing unit further includes: a first register 111 connected to the multiplication module, configured to store the multiplication operation result.
[0079] Moreover, the data processing unit further includes a first selector 112 and a second selector 113 connected to the addition module; wherein, the first selector 112 and the second selector 113 are configured to input the selected data into the addition module for corresponding type of addition operation processing.
[0080] That is, in the embodiment of the present application, the first selector and the second selector are configured to select data corresponding to convolution operation or data corresponding to fully connected operation or data corresponding to point-to-point operation from their input data, and then input the data into the addition module to complete the addition operation of the corresponding type. Among them, the data corresponding to the fully connected operation is the multiplication operation result output by the multiplication module of other data processing units or the addition operation result output by the addition module of other data processing units.
[0081] It should be noted that the addition module in the embodiment of the present application can be reused. Therefore, the data input into the addition module includes data corresponding to convolution operation, data corresponding to fully connected operation, and data corresponding to point-to-point operation. Through the selector for selection, the selected data is input into the addition module to complete the addition operation of the corresponding type, thereby realizing the reuse of the addition module and saving circuit resources.
[0082] Moreover, the data processing unit further includes:
[0083] a second register 114 connected to the addition module, configured to store the addition operation result output by the addition module. Correspondingly, the addition module is specifically configured to perform an accumulation operation of convolution operation by using the addition operation result in the second register.
[0084] It should be noted that the second register 114 stores the addition operation result of this time. When the next addition operation is performed, the addition operation result stored in the second register 114 is used to perform addition operation processing with the current data to be added, and the number of accumulation times is related to the convolution kernel size.
[0085] Further, the data processing unit further includes:
[0086] A saturation truncation processing unit, configured to perform saturation truncation processing on the output result of the multiplication module to obtain data with a specified number of digits.
[0087] In addition, in a specific implementation manner, this embodiment may include 8 input channels, each input channel includes a first input end and a second input end, and each of the computing arrays includes 64 data processing units. For example, refer to Figure 3 as shown Figure 3 is a schematic diagram of a computing array in a specific neural network computing terminal disclosed in an embodiment of the present application, including 8 computing arrays, namely MAC_64 arrays.
[0088] Correspondingly, specifically, in the fully connected operation, the addition module is configured to respectively pass the multiplication operation results of the multiplication module 101 in the 32nd to 63rd data processing units into the addition module 102 in the 0th to 31st data processing units, and perform addition operation processing with the multiplication operation results of the multiplication module 101 in the 0th to 31st data processing units respectively to obtain corresponding addition operation results, and then sequentially pass the addition operation results of the 0th to 31st data processing units in pairs into the addition module 102 in the 32nd to 47th data processing units for addition operation processing, and then sequentially pass the addition operation results of the 32nd to 47th data processing units in pairs to the addition module 102 in the 48th to 55th data processing units for addition operation processing, and so on, until the addition module 102 of the 62nd data processing unit determines the sum of 64 points, and passes the sum into the addition module of the 63rd data processing unit for accumulation.
[0089] It should be noted that the process of reusing the addition module by global average pooling can refer to the process of reusing the fully connected operation described above.
[0090] That is to say, the neural network computing terminal provided by the embodiment of the present application can calculate operations such as ordinary convolution, fully connected operation, point-to-point operation (calculation of corresponding points of two feature maps, such as point addition and point multiplication), etc., can be customized to achieve different computing powers, realize the reuse of each operation unit of the computing terminal, and reduce the area consumption. And, it includes multiple input channels, each input channel corresponds to a computing array, and each computing array includes multiple data processing units, which can achieve a large throughput, support simultaneous operation of multiple data processing units. In this way, within the power consumption limit, it can not only ensure the operation accuracy, but also meet various computing modes, and at the same time improve the utilization rate of the computing terminal.
[0091] For example, refer to Figure 4 As shown, the embodiments of the present application disclose a schematic diagram of a specific data processing unit.
[0092] 1) The data input corresponding to A is the data input of the feature image A. When B performs a convolution operation, it is the data input of the convolution kernel. If it is a fully connected or point-to-point operation, it is the data input corresponding to the feature image B. Two subtractors are used respectively to subtract the zero-point offset value after quantization, so as to improve the accuracy and reduce the error. Among them, In_a_byte_mode, In_b_byte_mode, In_a_zp, and In_b_zp are all configuration parameters. In_a_byte_mode and In_b_byte_mode respectively represent the data formats of data A and data B, and In_a_zp and In_b_zp are the zero-point drift values to be subtracted corresponding to data A and data B respectively.
[0093] 2) After subtracting the zero-point offset, there will be a left shift -> multiplication -> right shift. This operation is only applied to specific point-to-point addition operations, aiming to quantize the data of the two feature images to the same quantization range during the point addition operation; if it is not a point-to-point addition operation, bypass is selected. Elt_a_lshft and Elt_b_lshft are the left shift bits corresponding to data A and data B respectively, Elt_a_scale and Elt_b_scale are the quantization coefficients corresponding to data A and data B respectively, Elt_a_rshft and Elt_b_rshft are the right shift bits corresponding to data A and data B respectively, and ~In_a_scale_en and ~In_b_scale_en respectively represent the bypass of the two-way data. int32 and int64 respectively represent the data types processed at the corresponding steps. And, in the figure, the << mark represents the left shift processing unit, the X mark represents the multiplier, the >> mark represents the right shift processing unit, and the REG mark represents the register for storing data. Since the data can be directly used without a register after right shift because it can be directly used at the current clock cycle and meets the comprehensive timing convergence. And, in the actual circuit, right shift is equivalent to intercepting the data, and left shift is equivalent to bit splicing, which can be implemented by a circuit through synthesis.
[0094] 3) After the right shift operation, connect to the multiplication module to satisfy the multiplication operation of the convolutional neural network, perform a saturation truncation process on the multiplied result, process the result as int32 type and output it through the register, and this multiplication module will also be reused during the point-to-point multiplication arithmetic operation.
[0095] 4) An addition module that can perform accumulation during convolution operations, output from registers, and use the output result as the input for self-accumulation. The number of accumulations is related to the size of the convolution kernel. This addition module is reused multiple times during different operations, including performing addition operation processing between corresponding points, such as Figure 4 where A’ and B’ are both data after right shift processing, and point-to-point operations are performed. During fully connected or global average pooling operations, the register outputs after the multiplication modules of the 32nd pe_core to the 63rd pe_core are respectively passed into the 0th pe_core to the 31st pe_core, and the multiplication results of the two pe_cores are added together; the outputs of the addition modules of the 0th pe_core to the 31st pe_core are pairwise passed into the 32nd pe-core to the 47th pe_core, and then the results are pairwise passed into other pe_cores not occupied by the addition module, and so on. Finally, the sum of 64 points is obtained in the 62nd pe_core; the 62nd pe_core obtains the sum of 64 points, and at this time, the sum is passed into the 63rd pe_core for self-accumulation, and the number of accumulations is related to the size of the input data.
[0096] Next, taking the feature size of the first input end as 8*8*8 (i.e., width * height * number of input channels of the feature map), and the weight size of the second input end is also 8*8*8, and they are input in parallel. That is, taking the number of input channels of the feature map as 8 as an example, the calculation process of the fully connected layer is described in detail. This example only requires one computing array. In other embodiments, multiple computing arrays can be used for processing according to needs.
[0097] 1. A data block of 8*8 is passed into 64 pe_cores point by point, with one pe_core corresponding to one pixel point, and the input data of the feature and the weight are passed in parallel.
[0098] 2. Subtract the corresponding zero offset values from the input data of the feature and the weight to improve the quantization accuracy so that the accuracy is not significantly lower than the quantization accuracy.
[0099] 3. Pass the data after subtracting the zero offset into the two input ends of the multiplication module respectively for point multiplication operations, and output the operation results to the registers; 64 pe_cores perform parallel calculations, so that the multiplication results between corresponding points of an 8*8 block are calculated simultaneously.
[0100] 4. Refer to Figure 5 as shown Figure 5 is a schematic diagram of a specific fully connected operation disclosed in the embodiments of the present application. Among them, PE is the pe_core, such as Figure 4As shown in the first row, the result of multiplying PE32 - PE63 is output through the pe_ext_out interface, and the outputs are respectively connected to the Ext_b of PE0 - PE31. The register outputs of the multiplication of PE0 - PE31 and the pe_ext_out of PE32 - PE63 are connected to the addition module of PE0 - PE31 to achieve the addition between Ext_a and mul_reg (the register after the multiplication module). The addition result is output through the pe_out register, and D represents the output result.
[0101] 5. In the previous step, 64 points have been added pairwise once, becoming 32 values output from the pe_out of PE0 - PE31 respectively. The output results are respectively passed pairwise into the unoccupied PE's Ext_a and Ext_b of the addition module to achieve the addition of Ext_a and Ext_b. As Figure 4 shown in the second to sixth rows, adding 32 values pairwise results in 16 values, and then adding them pairwise again becomes 8 values. Finally, in PE62, the values of 8 * 8 points are added into one point.
[0102] 6. The value of PE62 is passed into PE63. Only the sum of 64 points of the first feature map channel comes in this cycle. There are 8 feature map input channels, that is, it takes 8 beats to pass in the sum of 8 data blocks of 8 * 8. The addition module here needs to achieve self - accumulation, that is, the number in this cycle is added to the input of the next cycle, and so on until all the numbers are added up. Here there are 8 feature map channels, so it will self - accumulate 8 times to complete the fully - connected calculation and obtain the fully - connected calculation result pe_out = fc_out. Among them, the number of self - accumulation times in a computing array is the result obtained by dividing the width * height * number of feature map input channels by the number of data processing units in a computing array.
[0103] It should be noted that the two selectors connected to the addition module are used to select the three - way input data and perform corresponding types of addition operation processing.
[0104] Furthermore, as Figure 4 shown, the neural network computing terminal provided by the embodiment of the present application may further include processing units such as left shift (<<), data selector (OP), maximum value calculation unit (Max), minimum value calculation unit (Min), etc. In some other embodiments, at the position where the multiplication module is located, there may also be processing units such as A << B, A >> B, bitwise AND, bitwise OR, bitwise XOR, etc. Furthermore, a subtraction module is also included. The embodiment of the present application controls the specific operations to be executed by the computing mode controller based on the configuration parameters, and Conv_type & elt_type is the configuration parameter.
[0105] In addition, the data processing unit can also select and output data A and B according to data C.
[0106] See Figure 6 As shown, an embodiment of the present application discloses a neural network computing method. The method is applied to a neural network computing terminal, and the neural network computing terminal includes N input channels and N computing arrays. Among them, each input channel corresponds to one computing array, and each computing array includes a plurality of data processing units. Each data processing unit includes an addition module. The computing method includes:
[0107] Step S11: Perform a multiplication operation on the data input to the input channel, and output the multiplication result of the multiplication operation to the addition module;
[0108] Step S12: Perform an addition operation on the multiplication result to obtain an addition result, and input a plurality of the addition results to the addition module or use the addition result as the target data of the computing terminal.
[0109] It can be seen that in the embodiment of the present application, the addition module in each data processing unit can perform an addition operation on the multiplication result to obtain the target data of the computing terminal, and can also perform an addition operation on the operation result input to the addition module to obtain the target data of the computing terminal, realizing the reuse of the addition module, thereby reducing the hardware area consumption of the computing terminal.
[0110] Furthermore, in any one of the computing arrays, the addition modules are connected into an addition hierarchical structure. The addition hierarchical structure includes a first layer, an intermediate layer, and an output layer. The addition module of the first layer receives the multiplication result, performs an addition operation on the multiplication result to obtain a first-layer addition result, and outputs a plurality of the first-layer addition results to the addition module of the intermediate layer. The addition module of the intermediate layer receives the first-layer addition result, performs an addition operation on the first-layer addition result to obtain an intermediate-layer addition result. The addition module of the output layer receives the intermediate-layer addition result, performs an accumulation operation on the intermediate-layer addition result, and outputs the final operation result as the target data of the computing terminal.
[0111] For the specific content of the neural network computing terminal, reference can be made to the content disclosed in the foregoing embodiments, and details will not be described herein again.
[0112] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. The steps of the method or algorithm described in combination with the embodiments disclosed in this article can be implemented directly by hardware, software modules executed by a processor, or a combination of both. The software module can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0113] The above provides a detailed introduction to a neural network computing terminal and a neural network computing method provided by this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A neural network computing terminal, characterized in that, It includes N input channels and N computing arrays, where each of the input channels corresponds to one of the computing arrays, and each of the computing arrays includes a plurality of data processing units, and each of the data processing units includes a multiplication module and an addition module; The multiplication module is configured to perform a multiplication operation on the data input by the input channel and output the multiplication operation result of the multiplication operation to the addition module; The addition module is configured to perform an addition operation on the multiplication operation result to obtain an addition operation result, and input a plurality of the addition operation results into the addition module for an addition operation or use the addition operation result as the target data of the computing terminal; The computing terminal further includes a selector module, the selector module includes a first selector and a second selector, the addition module is connected to the first selector, and the first selector includes a first selection end, a second selection end, and a third selection end; the first selection end is connected to the output end of the multiplication module of the current data processing unit and inputs the multiplication operation result of the multiplication module of the current data processing unit; the second selection end is connected to the first input end of the input channel and inputs the first data input by the first input end; the third selection end is connected to the output end of other data processing units and inputs the multiplication operation result of the multiplication module of other data processing units or the addition operation result of the addition module of other data processing units; the second selector includes three selection ends, namely a fourth selection end, a fifth selection end, and a sixth selection end; the fourth selection end is connected to the output end of the addition module of the current data processing unit and inputs the addition operation result of the addition module of the current data processing unit; the fifth selection end is connected to the second input end of the input channel and inputs the second data input by the second input end; the sixth selection end is connected to the output end of other data processing units and inputs the multiplication operation result of the multiplication module of other data processing units or the addition operation result of the addition module of other data processing units; various operation types are executed based on the selection results of the first selector and the second selector, and the various operation types include ordinary convolution operations and point operations.
2. The computing terminal according to claim 1, characterized in that The multiplication module is configured to perform a multiplication operation on the first data of the first input end and the second data of the second input end to obtain a multiplication operation result.
3. The computing terminal according to claim 1, wherein In any one of the computing arrays, the addition modules are connected into an addition hierarchical structure, and the addition hierarchical structure includes a first layer, an intermediate layer, and an output layer; The addition modules of the first layer are connected to the output ends of the multiplication modules, the input ends of the addition modules of the intermediate layer are connected to the output ends of the addition modules of the first layer, and the input ends of the addition modules of the output layer are connected to the output ends of the addition modules of the intermediate layer; The addition modules of the first layer are configured to receive the multiplication operation result, perform an addition operation on the multiplication operation result to obtain a first layer addition operation result, and output a plurality of the first layer addition operation results to the addition modules of the intermediate layer; The addition module of the middle layer is configured to receive the addition operation result of the first layer, perform an addition operation on the addition operation result of the first layer, and obtain the addition operation result of the middle layer; The addition module of the output layer is configured to receive the addition operation result of the middle layer, perform an accumulation operation on the addition operation result of the middle layer, and output the final operation result as the target data of the computing terminal.
4. The computing terminal according to claim 3, wherein The middle layer includes multiple levels of addition modules. Any addition module in any level of the multiple levels of addition modules is configured to receive the addition operation results output by two addition modules in the previous-level addition module of this level of addition module, and perform an addition operation on the two addition operation results to obtain the corresponding addition operation result. The addition operation result output by the last-level addition module in the multiple levels of addition modules is the addition operation result of the middle layer.
5. The computing terminal according to claim 1, characterized in that, The computing array further includes a register configuration unit, and the register configuration unit is configured to determine different operation types according to the register values of the register configuration unit.
6. The computing terminal according to claim 1, wherein The data processing unit further includes a subtractor connected to the input channel; The subtractor is configured to perform a subtraction operation for zero drift on the input data of the input channel.
7. The computing terminal according to claim 1, wherein The data processing unit further includes: A first register connected to the output end of the multiplication module, configured to store the multiplication operation result output by the multiplication module; A second register connected to the output end of the addition module, configured to store the addition operation result output by the addition module.
8. A neural network computing method, characterized in that, The method is applied to a neural network computing terminal. The neural network computing terminal includes N input channels and N computing arrays. Among them, each input channel corresponds to one computing array. The computing array includes multiple data processing units, and each data processing unit includes an addition module. The computing method includes: Performing a multiplication operation on the data input by the input channel, and outputting the multiplication operation result of the multiplication operation to the addition module; Performing an addition operation on the multiplication operation result to obtain an addition operation result, and inputting multiple addition operation results to the addition module or using the addition operation result as the target data of the computing terminal; Among them, the computing terminal further includes a selector module, and the selector module includes a first selector and a second selector. The addition module is connected to the first selector. The first selector includes a first selection terminal, a second selection terminal, and a third selection terminal. The first selection terminal is connected to the output terminal of the multiplication module of the current data processing unit and inputs the multiplication operation result of the multiplication module of the current data processing unit. The second selection terminal is connected to the first input terminal of the input channel and inputs the first data input by the first input terminal. The third selection terminal is connected to the output terminal of other data processing units and inputs the multiplication operation result of the multiplication module of other data processing units or the addition operation result of the addition module of other data processing units. The second selector includes three selection terminals, namely a fourth selection terminal, a fifth selection terminal, and a sixth selection terminal. The fourth selection terminal is connected to the output terminal of the addition module of the current data processing unit and inputs the addition operation result of the addition module of the current data processing unit. The fifth selection terminal is connected to the second input terminal of the input channel and inputs the second data input by the second input terminal. The sixth selection terminal is connected to the output terminal of other data processing units and inputs the multiplication operation result of the multiplication module of other data processing units or the addition operation result of the addition module of other data processing units. Multiple operation types are performed based on the selection results of the first selector and the second selector, and the multiple operation types include ordinary convolution operations and point operations.
9. The calculation method according to claim 8, characterized in that In any one of the computing arrays, the addition modules are connected into an addition hierarchical structure, and the addition hierarchical structure includes a first layer, an intermediate layer, and an output layer. The addition modules of the first layer receive the multiplication operation results, perform addition operations on the multiplication operation results to obtain first-layer addition operation results, and output multiple first-layer addition operation results to the addition modules of the intermediate layer. The addition modules of the intermediate layer receive the first-layer addition operation results, perform addition operations on the first-layer addition operation results to obtain intermediate-layer addition operation results. The addition modules of the output layer receive the intermediate-layer addition operation results, perform accumulation operations on the intermediate-layer addition operation results to obtain a final operation result, and output it as the target data of the computing terminal.
Citation Information
Patent Citations
Pipeline-based neural network processing system and processing method
CN107862374A
Operation method and device
CN111445017A
Neural network computing module, method and communication device
CN113792868A