Data processing system and method, coding unit, processing unit and storage medium
By introducing coding units into the data processing system of the convolutional neural network, data encoding processing is performed and the logic and hardware design of the processing unit is simplified, the problems of high complexity and low processing efficiency in the prior art are solved, and more efficient data processing is achieved.
Patent Information
- Application Number
- CN201980096378.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-07-03
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2039-07-03
AI Technical Summary
The processing unit (PE) in the existing convolutional neural network undertakes complex coding, partial product operations, intra-multiplier accumulation and inter-multiplier accumulation tasks, resulting in low processing efficiency and further optimization space.
By introducing an encoding unit into the data processing system, the data to be encoded is encoded and the encoded data is input into the processing unit to simplify the internal processing logic and hardware design of the processing unit, thereby reducing hardware complexity and improving processing efficiency.
The hardware design and processing logic of the processing unit are simplified, processing efficiency is improved, and the hardware complexity and cost overhead of the system are reduced.
Smart Images

Figure CN113841158B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a data processing system and method, an encoding unit, a processing unit, and a storage medium. Background Art
[0002] Convolutional Neural Networks (CNNs) have broad application prospects in fields such as image and speech recognition.
[0003] Convolutional neural networks are generally used to perform convolution operations on multiple convolutional kernel data and at least one feature layer data. Specifically, the convolution operations are implemented by multiple Processing Engines (PEs) included in the CNN. Specifically, the input of each PE includes two aspects of data: convolutional kernel data and feature layer data. One of them is used as the multiplier, and the other is used as the multiplicand. After encoding the multiplier by multiple multipliers set inside the PE respectively, each multiplier performs partial product operation and carry addition operation on the encoded multiplier and the multiplicand to obtain the operation result of each multiplier, and the accumulation tree of the PE accumulates and sums the operation results output by each multiplier to obtain the convolution result output by the PE.
[0004] In the existing convolutional neural networks, the PE undertakes the tasks of encoding, partial product operation, accumulation inside the multiplier, and accumulation between multipliers. The internal processing logic and hardware structure of the PE are relatively complex, which to a certain extent affects the processing efficiency of the PE and there is room for further optimization. Summary of the Invention
[0005] The data processing system and method, encoding unit, processing unit, and storage medium provided by this application are used to reduce the internal processing logic and hardware complexity of the processing unit and improve the processing efficiency of the processing unit.
[0006] In a first aspect, this application provides a data processing system, including an encoding unit and a processing unit; the encoding unit is used to: obtain data to be encoded; perform encoding processing on the data to be encoded to obtain first input data; the processing unit is used to: perform partial product generation processing on the first input data and second input data to obtain a partial product result; perform accumulation processing on the partial product result to obtain a convolution result; where the data to be encoded is convolutional kernel data and the second input data is feature layer data; or the data to be encoded is feature layer data and the second input data is convolutional kernel data. Through the foregoing design, the PE no longer undertakes the encoding logic, simplifies the hardware design inside the PE, is beneficial to reducing the hardware complexity, and also is beneficial to improving the processing efficiency of the PE due to the simplified processing steps.
[0007] In another possible design, the processing unit includes a first accumulation module and at least two multipliers, where each multiplier includes a partial product module and a second accumulation module; the partial product module is configured to perform the partial product generation process; the second accumulation module is configured to perform a first accumulation process on the partial product results to obtain an accumulation result; the first accumulation module is configured to perform a second accumulation process on the accumulation results respectively output by the second accumulation modules of the at least two multipliers to obtain the convolution result. That is, in the multipliers inside the PE, an encoding logic module for executing encoding logic is no longer provided, and the partial product module directly performs the partial product generation process according to the received data.
[0008] In another possible design, the number of encoding units is less than the number of processing units. In this design, the encoding units do not need to be set one-to-one with the processing units. One encoding unit can be responsible for pre-encoding the data to be input to multiple PEs, which can, to a certain extent, reduce the impact of the number of encoding units on the complexity of the data processing system.
[0009] In another possible design, multiple processing units reuse the first input data output by one encoding unit. When the first input data is reused by multiple PEs, the encoding unit does not need to encode each data to be input separately, but only needs to perform encoding once and reuse the encoded first input data, which can conveniently implement data processing and is beneficial to further improving the processing efficiency.
[0010] In an implementation scenario of reusing the first input data, the storage media of the multiple processing units are connected pairwise in sequence to form a processing unit sequence. One processing unit located at one end of the processing unit sequence is configured to receive the first input data output by the one encoding unit, and the other processing units obtain the first input data from the storage media of the previous processing unit to which they are connected. In this design, multiple PEs that reuse the same first input data are connected by data lines. Therefore, any PE can obtain the data stored in the storage media of its adjacent PE to which it is connected. Thus, the encoding unit only needs to input the first input data into one PE, and the other PEs can obtain the reused first input data through the connection relationship between the storage media. In terms of the processing method and effect, it has a high processing efficiency and simple and easy steps.
[0011] In an implementation scenario of reusing the first input data, the one encoding unit is further configured to: input the first input data into the multiple processing units respectively. That is, each PE is connected to the encoding unit, and the encoding unit can directly input the encoded first input data into each PE without passing through the PEs.
[0012] In another possible design, the first input data is used to control the processing unit to perform the partial product generation process. In other words, the first input data can be embodied as a control signal for the partial product module, which can be used to instruct the partial product module to perform the partial product generation process. In this way, during the partial product generation process performed by the partial product module, the multiplier is no longer encoded, but the pre-generated control signal and the multiplicand are directly used for operation to obtain the convolution result.
[0013] In another possible design, the second input data is the multiplexed data among multiple processing units. In a specific implementation scenario, the data processing system can multiplex at least one of the first input data and the second input data. By means of data multiplexing, the complexity of the data flow transfer in the data processing system is reduced, which is also beneficial to reducing the system complexity and improving the processing efficiency.
[0014] In another possible design, the encoding process for the data to be encoded includes: encoding the data to be encoded according to a preset encoding rule. Among them, the encoding rule can include but is not limited to the Booth encoding rule, that is, it includes the Booth encoding rule and its various derivative rules.
[0015] In another possible design, the second accumulation module of the processing unit is composed of multiple carry-save adders (CSA). At this time, the first accumulation module of the processing unit is composed of multiple carry-save adders (CSA) and one carry-lookahead adder (CPA). With this design, when performing the accumulation process in each PE, it is not necessary to calculate the exact product of each pair of values in the step of multiplying numbers pairwise. The accumulation can be achieved by using the characteristic of the equivalent product expression for accumulation and generating the convolution result. Thus, on the premise of ensuring the mathematical equivalence of the output result inside the PE, the slow and high-logic-cost CPA accumulation logic can be removed, and the first layer of the first accumulation module can be modified to adapt to this modification, thereby improving the logic operation speed and reducing the device overhead.
[0016] Second aspect, the present application provides a data processing method, which is applied to an encoding unit. The method includes: obtaining data to be encoded; performing encoding processing on the data to be encoded to obtain first input data; inputting the first input data into a processing unit, so that the processing unit performs partial product generation processing on the first input data and second input data to obtain a partial product result, and performs accumulation processing on the partial product result to obtain a convolution result; wherein, the data to be encoded is convolution kernel data, and the second input data is feature layer data; or, the data to be encoded is feature layer data, and the second input data is convolution kernel data. Through the foregoing design, the PE no longer undertakes the encoding logic, simplifies the internal hardware design of the PE, is beneficial to reducing the hardware complexity, and since the processing steps are simplified, it is also beneficial to improving the processing efficiency of the PE.
[0017] In another possible design, the number of encoding units is less than the number of processing units. In this design, the encoding units do not need to be set one-to-one with the processing units. One encoding unit can be responsible for pre-encoding the data to be input of multiple PEs, which can, to a certain extent, reduce the impact of the number of encoding units on the complexity of the data processing system. In another possible design, the first input data is reused by multiple processing units. When the first input data is reused by multiple PEs, the encoding unit does not need to encode each data to be input separately, but only needs to perform encoding once and reuse the encoded first input data, so as to conveniently implement data processing, which is beneficial to further improving the processing efficiency.
[0018] In an implementation scenario of reusing the first input data, the storage media of the multiple processing units are sequentially connected in pairs to form a processing unit sequence, and the method further includes: inputting the first input data into one of the processing units located at one end of the processing unit sequence. In this design, multiple PEs that reuse the same first input data are connected by data lines. Therefore, any PE can obtain the data stored in the storage media of its adjacent PEs connected thereto. Thus, the encoding unit only needs to input the first input data into one of the PEs, and other PEs can obtain the reused first input data through the connection relationship between the storage media. In terms of the processing method and effect, it has high processing efficiency and simple and feasible steps.
[0019] In another implementation scenario of reusing the first input data, the method further includes: inputting the first input data into the multiple processing units respectively. That is, each PE is respectively connected to the encoding unit, and the encoding unit can directly input the encoded first input data into each PE without passing through the PEs.
[0020] In another possible design, the first input data is used to control the processing unit to perform the partial product generation process. In other words, the first input data can be embodied as a control signal for the partial product module, which can be used to instruct the partial product module to perform the partial product generation process. In this way, during the process of the partial product module performing the partial product generation process, the multiplier is no longer encoded, but the convolution result is directly obtained by operating on the pre-generated control signal and the multiplicand.
[0021] In another possible design, the second input data is the multiplexed data among multiple processing units. In a specific implementation scenario, the data processing system can multiplex at least one of the first input data and the second input data. By means of data multiplexing, the complexity of the data flow transfer in the data processing system is reduced, which is also beneficial to reducing the system complexity and improving the processing efficiency.
[0022] In another possible design, the encoding process for the data to be encoded includes: encoding the data to be encoded according to a preset encoding rule. Among them, the encoding rule can include but is not limited to the Booth encoding rule, that is, it includes the Booth encoding rule and its various derivative rules. Thirdly, the present application provides a data processing method applied to a processing unit. The method includes: performing a partial product generation process on the first input data and the second input data to obtain a partial product result; performing an accumulation process on the partial product result to obtain a convolution result; wherein, the first input data is obtained by an encoding unit acquiring the data to be encoded and encoding the data to be encoded; wherein, the data to be encoded is convolution kernel data, and the second input data is feature layer data; or, the data to be encoded is feature layer data, and the second input data is convolution kernel data. Through the foregoing design, the PE no longer bears the encoding logic, simplifies the internal hardware design of the PE, is beneficial to reducing the hardware complexity, and since the processing steps are simplified, it is also beneficial to improving the processing efficiency of the PE.
[0023] In another possible design, the processing unit includes at least two multipliers; the performing an accumulation process on the partial product result to obtain a convolution result includes: performing a first accumulation process on the partial product results of each multiplier to obtain an accumulation result; performing a second accumulation process on the accumulation results of the at least two multipliers to obtain the convolution result. That is, in the multipliers inside the PE, an encoding logic module for executing the encoding logic is no longer provided, and the partial product module directly performs the partial product generation process according to the received data.
[0024] In another possible design, the number of the encoding units is less than the number of the processing units. In this design, the encoding units do not need to be set one-to-one with the processing units. One encoding unit can be responsible for pre-encoding the data to be input for multiple PEs, which can, to a certain extent, reduce the impact of the number of encoding units on the complexity of the data processing system.
[0025] In another possible design, multiple ones of the processing units reuse the first input data output by one of the encoding units. When the first input data is reused by multiple PEs, the encoding unit does not need to encode each data to be input separately, but only needs to perform encoding once and reuse the encoded first input data, so as to conveniently implement data processing, which is beneficial to further improving the processing efficiency.
[0026] In an implementation scenario of reusing the first input data, the storage media of the multiple processing units are sequentially connected pairwise to form a processing unit sequence; when the processing unit is at one end of the processing unit sequence, the method further includes: receiving the first input data output by the one encoding unit; when the processing unit is another processing unit in the processing unit sequence, the method further includes: obtaining the first input data stored in the storage media of the previous processing unit connected thereto. In this design, multiple PEs reusing the same first input data are connected by data lines. Therefore, any PE can obtain the data stored in the storage media of its adjacent PE connected thereto. Thus, the encoding unit only needs to input the first input data into one of the PEs, and other PEs can obtain the reused first input data through the connection relationship between the storage media. In terms of the processing method and processing effect, it has a high processing efficiency and the steps are simple and feasible.
[0027] In an implementation scenario of reusing the first input data, the method further includes: receiving the first input data output by the one encoding unit. That is, each PE is respectively connected to the encoding unit. Therefore, the encoded first input data sent by the encoding unit can be directly received without being transmitted between PEs.
[0028] In another possible design, the first input data is used to control the processing unit to perform the partial product generation process. In other words, the first input data can be specifically embodied as the control signal of the partial product module and can be used to instruct the partial product module to perform the partial product generation process. Thus, during the process of the partial product module performing the partial product generation process, the multiplier is no longer encoded, but the convolution result is directly obtained by operating on the pre-generated control signal and the multiplicand.
[0029] In another possible design, the second input data is the multiplexed data among multiple processing units. In a specific implementation scenario, the data processing system can multiplex at least one of the first input data and the second input data. By means of data multiplexing, the flow complexity of the data stream in the data processing system can be reduced, which is also beneficial to reducing the system complexity and improving the processing efficiency.
[0030] In another possible design, each multiplier includes a second accumulation module, and the second accumulation module is composed of multiple carry-save adders (CSA). At this time, the processing unit further includes a first accumulation module, and the first accumulation module is composed of multiple carry-save adders (CSA) and a carry-lookahead adder (CPA). With this design, when performing accumulation processing in each PE, it is not necessary to calculate the exact product of each pair of values in the step of multiplying numbers pairwise. The accumulation can be achieved by using the equivalent product expression and generating the characteristics of the convolution result. Thus, on the premise of ensuring the mathematical equivalence of the output results inside the PE, the slow and high-logic-cost CPA accumulation logic can be removed, and the first layer of the first accumulation module can be modified to adapt to this modification, thereby improving the logic operation speed and reducing the device overhead.
[0031] In a fourth aspect, the present application provides an encoding unit, including: an acquisition module, configured to acquire data to be encoded; an encoding module, configured to perform encoding processing on the data to be encoded to obtain first input data; a sending module, configured to input the first input data into a processing unit, so that the processing unit performs partial product generation processing on the first input data and second input data to obtain a partial product result, and perform accumulation processing on the partial product result to obtain a convolution result; where the data to be encoded is convolution kernel data, and the second input data is feature layer data; or the data to be encoded is feature layer data, and the second input data is convolution kernel data. Through the foregoing design, the PE no longer undertakes the encoding logic, simplifies the internal hardware design of the PE, is beneficial to reducing the hardware complexity, and since the processing steps are simplified, it is also beneficial to improving the processing efficiency of the PE.
[0032] In another possible design, the number of encoding units is less than the number of processing units. In this design, it is not necessary to set the encoding units one-to-one with the processing units. One encoding unit can be responsible for pre-encoding the data to be input of multiple PEs, which can, to a certain extent, reduce the impact of the number of encoding units on the complexity of the data processing system.
[0033] In another possible design, the first input data is multiplexed by multiple processing units. When the first input data is multiplexed by multiple PEs, the encoding unit does not need to encode each piece of data to be input separately, but only needs to perform encoding once and multiplex the encoded first input data, so as to conveniently implement data processing, which is beneficial to further improving the processing efficiency.
[0034] In an implementation scenario of multiplexing the first input data, the storage media of the multiple processing units are sequentially connected in pairs to form a processing unit sequence. The sending module is specifically configured to: input the first input data into one of the processing units located at one end of the processing unit sequence. In this design, multiple PEs that multiplex the same first input data are connected by data lines. Therefore, any PE can obtain the data stored in the storage media of its adjacent PEs connected thereto. Thus, the encoding unit only needs to input the first input data into one of the PEs, and other PEs can obtain the multiplexed first input data through the connection relationship between the storage media. In terms of the processing method and processing effect, it has high processing efficiency and simple and feasible steps.
[0035] In another implementation scenario of multiplexing the first input data, the sending module is specifically configured to: input the first input data into the multiple processing units respectively. That is, each PE is respectively connected to the encoding unit, and the encoding unit can directly input the encoded first input data into each PE without passing through the PEs.
[0036] In another possible design, the first input data is used to control the processing unit to perform the partial product generation process. In other words, the first input data can be specifically embodied as the control signal of the partial product module, which can be used to instruct the partial product module to perform the partial product generation process. Thus, during the process of the partial product module performing the partial product generation process, the multiplier is no longer encoded, but the convolution result is directly obtained by operating on the pre-generated control signal and the multiplicand.
[0037] In another possible design, the second input data is multiplexed data among multiple processing units. In a specific implementation scenario, the data processing system can multiplex at least one of the first input data and the second input data. By means of data multiplexing, the flow complexity of the data stream in the data processing system is reduced, which is also beneficial to reducing the system complexity and improving the processing efficiency.
[0038] In another possible design, the encoding module is specifically configured to: perform encoding processing on the data to be encoded according to a preset encoding rule. Among them, the encoding rule can include but is not limited to the Booth encoding rule, that is, it includes the Booth encoding rule and its various derivative rules.
[0039] In a fifth aspect, the present application provides a processing unit, including: a partial product module for performing partial product generation processing on first input data and second input data to obtain a partial product result; an accumulation processing module for performing accumulation processing on the partial product result to obtain a convolution result; wherein, the first input data is obtained by an encoding unit acquiring data to be encoded and performing encoding processing on the data to be encoded; wherein, the data to be encoded is convolution kernel data, and the second input data is feature layer data; or, the data to be encoded is feature layer data, and the second input data is convolution kernel data. Through the foregoing design, the PE no longer undertakes the encoding logic, simplifies the internal hardware design of the PE, is beneficial to reducing the hardware complexity, and since the processing steps are simplified, it is also beneficial to improving the processing efficiency of the PE.
[0040] In another possible design, the accumulation processing module includes: a first accumulation module and at least two second accumulation modules, wherein one of the second accumulation modules belongs to one multiplier; the second accumulation module is used for performing first accumulation processing on the partial product result to obtain an accumulation result; the first accumulation module is used for performing second accumulation processing on the accumulation results respectively output by the second accumulation modules of the at least two multipliers to obtain the convolution result. That is, in the multiplier inside the PE, an encoding logic module for executing encoding logic is no longer provided, and the partial product module directly performs partial product generation processing according to the received data.
[0041] In another possible design, the number of encoding units is less than the number of processing units. In this design, the encoding units do not need to be set one-to-one with the processing units. One encoding unit can be responsible for pre-encoding the data to be input of multiple PEs, and can, to a certain extent, reduce the influence of the number of encoding units set on the complexity of the data processing system.
[0042] In another possible design, multiple processing units reuse the first input data output by one encoding unit. When the first input data is reused by multiple PEs, the encoding unit does not need to encode each data to be input separately, but only needs to perform encoding once and reuse the encoded first input data, so as to conveniently implement data processing, which is beneficial to further improving the processing efficiency.
[0043] In an implementation scenario of reusing the first input data, the storage media of the multiple processing units are sequentially connected pairwise to form a processing unit sequence; when the processing unit is at one end of the processing unit sequence, the partial product module is further configured to: receive the first input data output by the one encoding unit; when the processing unit is another processing unit in the processing unit sequence, the partial product module is further configured to: obtain the first input data stored in the storage media of the previous connected processing unit. In this design, multiple PEs that reuse the same first input data are connected through data lines. Therefore, any PE can obtain the data stored in the storage media of its adjacent connected PE. Thus, the encoding unit only needs to input the first input data into one of the PEs, and other PEs can obtain the reused first input data through the connection relationship between the storage media. In terms of the processing method and processing effect, it has high processing efficiency and simple and easy steps.
[0044] In an implementation scenario of reusing the first input data, the partial product module is further configured to: receive the first input data output by the one encoding unit. That is, each PE is respectively connected to the encoding unit. Therefore, the encoded first input data sent by the encoding unit can be directly received without being transmitted between PEs.
[0045] In another possible design, the first input data is used to control the processing unit to perform the partial product generation process. In other words, the first input data can be specified as the control signal of the partial product module, which can be used to instruct the partial product module to perform the partial product generation process. Thus, during the process of the partial product module performing the partial product generation process, the multiplier is no longer encoded, and the convolution result is directly obtained by operating on the pre-generated control signal and the multiplicand.
[0046] In another possible design, the second input data is the reused data among the multiple processing units. In a specific implementation scenario, the data processing system can reuse at least one of the first input data and the second input data. By means of data reuse, the flow complexity of the data stream in the data processing system is reduced, which is also beneficial to reducing the system complexity and improving the processing efficiency.
[0047] In another possible design, the second accumulation module of the processing unit is composed of multiple carry-save adders (CSAs). At this time, the first accumulation module of the processing unit is composed of multiple carry-save adders (CSAs) and a carry-lookahead adder (CPA). With this design, when performing accumulation processing in each PE, it is not necessary to calculate the exact product of each pair of values in the step of multiplying numbers pairwise. Instead, the accumulation can be achieved by using an equivalent product expression and generating the characteristics of the convolution result. Thus, on the premise of ensuring the mathematical equivalence of the output result inside the PE, the CPA accumulation logic with slow speed and high logic cost can be removed, and the first layer of the first accumulation module can be modified to adapt to this change, thereby improving the logical operation speed and reducing the device overhead.
[0048] In a sixth aspect, the present application provides an encoding unit, including: a memory; a processor; and a computer program; wherein, the computer program is stored in the memory and is configured to be executed by the processor to implement the method according to any one of the fourth aspect.
[0049] In a seventh aspect, the present application provides a processing unit, including: a memory; a processor; and a computer program; wherein, the computer program is stored in the memory and is configured to be executed by the processor to implement the method according to any one of the fifth aspect.
[0050] In an eighth aspect, the present application provides a computer-readable storage medium storing computer-executable instructions, which are used to implement the method according to any one of the second aspect or the third aspect when being executed by a processor.
[0051] In a ninth aspect, the present application provides a computer program product, which includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions, and when the program instructions are executed by a computer, the computer is caused to execute the method according to any one of the second aspect or the third aspect.
[0052] In summary, the data processing system and method, encoding unit, processing unit, and storage medium provided in the embodiments of the present application encode the data to be encoded through the encoding unit and input the encoded first input data into the processing unit. Thus, the processing unit can perform convolution processing based on the first input data and the second input data and obtain a convolution result. Thereby, the software and hardware settings required for executing the encoding logic inside the processing unit are saved, which is beneficial to reducing the logic overhead of the processing unit and improving the processing efficiency. Moreover, the encoder is removed inside the processing unit, reducing the device complexity of the processing unit and also reducing the cost overhead to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1Schematic diagram of the convolution processing principle of a convolutional neural network in the prior art;
[0054] Figure 2 Schematic diagram of the structure of a processing unit in the prior art;
[0055] Figure 3 For Figure 2 Structure of the Booth multiplier inside the processing unit shown;
[0056] Figure 4 Schematic diagram of the architecture of a data processing system provided by the present application;
[0057] Figure 5 Interaction schematic diagram of a data processing method provided by the present application;
[0058] Figure 6 Schematic diagram of the architecture of a processing unit provided by the present application;
[0059] Figure 7 Schematic diagram of the architecture of a multiplier in the processing unit provided by the present application;
[0060] Figure 8 Schematic diagram of the principle of reusing convolution kernel data in convolution processing in the present application;
[0061] Figure 9 Schematic diagram of the principle of reusing feature layer data in convolution processing in the present application;
[0062] Figure 10 Schematic diagram of the principle of simultaneously reusing convolution kernel data and feature layer data in convolution processing in the present application;
[0063] Figure 11 Schematic diagram of the architecture of another data processing system provided by the present application;
[0064] Figure 12 Schematic diagram of the architecture of another data processing system provided by the present application;
[0065] Figure 13 Schematic diagram of the processing principle of the radix-4 Booth multiplier provided by the present application;
[0066] Figure 14 Schematic diagram of the architecture of another multiplier in the processing unit provided by the present application;
[0067] Figure 15 Functional block diagram of the encoding unit provided by the present application;
[0068] Figure 16 Schematic diagram of the physical structure of the encoding unit provided by the present application;
[0069] Figure 17 Functional block diagram of the processing unit provided by this application;
[0070] Figure 18 Schematic diagram of the physical structure of the processing unit provided by this application. Detailed implementation manners
[0071] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. Hereinafter, this application will be specifically described with reference to the drawings.
[0072] The specific application scenarios of this application are as follows: scenarios for optimizing convolutional neural networks, or scenarios for performing convolutional operation processing using convolutional neural networks, or scenarios for constructing and testing convolutional neural networks.
[0073] Figure 1 Schematic diagram of the convolutional processing principle of the existing convolutional neural network. As Figure 1 shown, the convolutional neural network performs convolutional processing on the convolutional kernel data and the feature layer data (illustrated by taking at least one feature map as an example). Specifically, for each convolutional kernel, it starts from the first pixel of the feature map and moves pixel by pixel in the row direction. When reaching the end of this row, it moves down one pixel in the column direction, returns to the starting point in the row direction, and repeats the row direction movement process until all pixels of the feature map are traversed.
[0074] In the existing convolutional neural network, especially in a matrix processor, there are multiple identical processing units PE. In the prior art, generally, the convolutional kernel data and the feature layer data are loaded onto the PE, and the PE performs convolutional processing according to the convolutional principle to obtain their respective convolutional results. For the data that needs to be reused, the storage media (such as register) of each PE are connected through data lines so that each PE can obtain the reused data from another connected PE.
[0075] Figure 2 Shows the structure of the processing unit in the prior art. As Figure 2 shown, the inside of each PE includes an accumulation tree and multiple multipliers. Among them, each multiplier is used to perform convolutional processing on the input convolutional kernel data and the feature layer data and output the convolutional result of the multiplier, and the accumulation tree is used to accumulate the convolutional results of each multiplier to obtain the convolutional result of the PE.
[0076] In the actual application process, there can be various different choices for the multiplier inside the PE. For example, Booth multipliers, Wallace tree multipliers, etc. Among them, the Booth multiplier encodes every X bits of the multiplicand (for example, the radix-4 Booth multiplier encodes every 3 bits, and the radix-8 Booth multiplier encodes every 4 bits). By parsing the encoding, the partial products obtained by multiplying the X bits of the multiplicand with the multiplier are obtained. Then, all partial products are shifted according to the positions of their corresponding X bits and accumulated to find the final product.
[0077] Figure 3 shows Figure 2 the structure of the Booth multiplier inside the processing unit shown. As Figure 3 shown, the Booth multiplier mainly consists of three modules: an encoding logic module ( Figure 2 and Figure 3 represented by ENC), a partial product module ( Figure 2 and Figure 3 represented as Partial Product), and an accumulator module inside the multiplier ( Figure 2 and Figure 3 represented as a combination of multiple carry-save adders (CSA) and a carry-propagate adder (CPA)). As Figure 3 shown, the encoding logic unit is mainly used to encode the multiplicand to obtain the encoded multiplicand data, that is, to obtain the control signal of the partial product module. And the partial product module has two inputs: the multiplier and the control signal. The partial product module is used to process the multiplier according to the control signal to obtain the partial product of the multiplier and the encoded multiplicand. Thus, the accumulator module inside the multiplier is used to accumulate the partial products output by the partial product module to obtain the convolution result of the multiplier.
[0078] It can be seen from this that in the existing convolutional neural network as Figure 2 shown, the PE undertakes the tasks of encoding, partial product operation, accumulation inside the multiplier, and accumulation between multipliers. The internal processing logic and hardware structure of the PE are relatively complex, which to a certain extent affects the processing efficiency of the PE and there is room for further optimization.
[0079] The technical solution provided by this application aims to solve the above technical problems in the prior art and gives the following inventive concept: Before inputting the multiplicand data into the PE, the data is encoded in advance and the encoded data is input into the PE, thereby simplifying the internal processing logic and hardware structure of the PE.
[0080] The technical solutions of the present invention and how the technical solutions of the present application solve the above technical problems will be described in detail below with specific embodiments. The following several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will be described below with reference to the accompanying drawings.
[0081] Embodiment 1
[0082] An embodiment of the present invention provides a data processing system and a data processing method. Please refer to Figure 4 and Figure 5 , where Figure 4 shows a schematic architecture diagram of the data processing system, as Figure 4 shown, the data processing system 400 includes: an encoding unit 410 and a processing unit 420. Figure 5 shows an interaction schematic diagram of a data processing method:
[0083] On the one hand, the encoding unit 410 is used to perform the following steps:
[0084] S502, obtain the data to be encoded.
[0085] The data to be encoded is convolution kernel data or feature layer data. According to the design of the convolutional neural network, one of the convolution kernel data and the feature layer data can be used as the data to be encoded, and the encoding unit 410 performs encoding processing.
[0086] In the convolutional neural network, the output data of the Kth layer is the input data of the (K + 1)th layer. Therefore, the data to be encoded can have at least two sources: for the first-layer PE, the data to be encoded comes from the input data of the convolutional neural network (i.e., feature layer data) or the convolution kernel; for other layers of PE except the first layer, the data to be encoded can come from the output data of the previous layer of PE.
[0087] S504, perform encoding processing on the data to be encoded to obtain the first input data.
[0088] After the encoding processing by the encoding unit, the first input data that meets the format and other requirements for post-convolution processing by the processing unit can be obtained. There is no special limitation on the data format of the first input data in this application. It can be specifically represented as a set of numbers or characters, or can also be specifically represented as a control signal.
[0089] S506, input the first input data into the processing unit.
[0090] The first input data is used to control the processing unit to execute the partial product generation process. Specifically, the first input data is input into each multiplier in the processing unit. Further, it is input into the partial product module in the multiplier of the PE to perform the partial product generation process.
[0091] On the other hand, the processing unit 420 is used to execute the following steps:
[0092] S508, perform a partial product generation process on the first input data and the second input data to obtain a partial product result.
[0093] As described above, the two aspects of data for the processing unit to perform convolution processing include: convolution kernel data and feature layer data, and the second input data is the other aspect of data except the data to be encoded. In other words, in one implementation scenario, if the data to be encoded is convolution kernel data, then the second input data is feature layer data. Or, in another implementation scenario, if the data to be encoded is feature layer data, then the second input data is convolution kernel data.
[0094] Specifically, similar to the first input data, the second input data also has the following sources: feature layer data or convolution kernel data input by the convolutional neural network.
[0095] S510, perform an accumulation process on the partial product result to obtain a convolution result.
[0096] Specifically, the convolution result can be obtained through internal accumulation in the multiplier and external accumulation of the multiplier. Among them, the accumulation process involved in this application is to perform an accumulation summation process on each partial product result to obtain the sum of each partial product result as the convolution result. In a specific implementation scenario, when the data type of the accumulation process is binary data, the accumulation summation process can be implemented through carry addition.
[0097] Through the foregoing design, it can be seen that in the data processing system 400, the encoding of the multiplier can be implemented through an encoding unit outside the processing unit, and the encoded data is input to the PE. The PE performs convolution processing on the encoded data and the other aspect of data to obtain a convolution result. In this process, the PE no longer needs to bear the encoding logic, which is beneficial to simplifying the internal processing steps of the PE and improving the processing efficiency of the PE; moreover, the encoding logic modules of each multiplier in the PE can be correspondingly deleted, which also reduces the hardware complexity of the PE and the entire data processing system.
[0098] Hereinafter, the implementation manners of the processing unit and the encoding unit will be described separately.
[0099] On the one hand, the schematic diagram of the architecture of the PE and its internal multiplier can be referred to Figure 6 andFigure 7 Among them, as Figure 6 shown, the PE includes a first accumulation module and at least two multipliers ( Figure 6 shows 4 multipliers, namely multiplier 0 to multiplier 3), as Figure 6 and Figure 7 shown, each multiplier inside the PE includes a partial product module and a second accumulation module.
[0100] As Figure 6 or Figure 7 shown, the partial product module in the multiplier has multiple sub - modules ( Figure 6 , Figure 7 shows 5), and each sub - module has two aspects of inputs. One input is the second input data ( Figure 6 and Figure 7 is represented as Multiplicand, specifically Multiplicand0 to Multiplicand3), and the other input is the first input data encoded by the encoding unit ( Figure 6 and Figure 7 is represented as the control signal Control signal, specifically Control signal 0 to Control signal 3). Each sub - module of the partial product module is used to perform partial product generation processing on the two - aspect input data of itself to obtain a partial product result. It can be seen that based on the different numbers of sub - modules in the partial product module, the number of partial product results output by it is also different. In a specific implementation scenario, the partial product module can have at least two sub - modules, that is, it can output at least two partial product results.
[0101] In addition, it should be noted that Figure 6 , Figure 7 represents the partial product module as Partial ProductGeneration, which is only used to indicate that this module is used to perform partial product generation processing. In an actual application scenario, the partial product module can include but is not limited to: MUX SHFT, that is, use the MUX module and the Shifter module to implement partial product generation processing.
[0102] At this time, as Figure 7 shown, the multiplier is further designed with a second accumulation module, and the second accumulation module is used to perform the first accumulation processing on the at least two partial product results output above, so as to obtain an accumulation result, that is, the convolution result of the current multiplier.
[0103] And in as Figure 6In the internal architecture of the processing unit shown, the PE may include at least two multipliers. Therefore, it is necessary to further accumulate the accumulated results of multiple multipliers to obtain the convolution result of the PE. Therefore, Figure 6 In the scenario shown, a first accumulation module is further designed. The first accumulation module is used to perform a second accumulation process on the accumulation results respectively output by the second accumulation modules of the at least two multipliers to obtain the convolution result.
[0104] In addition, in a special application scenario, there may also be a situation where there is only one multiplier in the PE. At this time, the accumulated result output by the multiplier is the convolution result of the PE. In this implementation scenario, there is no need to additionally design a second accumulation module in the PE, which can further save hardware costs.
[0105] On the other hand, the present application has no special limitations on the number and form of encoding units.
[0106] In a possible design, the number of encoding units may be the same as the number of processing units. At this time, they are in one-to-one correspondence. One encoding unit is used to encode the data to be encoded in one processing unit and input it into each multiplier in the PE. Compared with Figure 3 the existing technology shown in which each multiplier and each sub-module in its encoding logic unit are encoded separately, this processing method can also improve and reduce the system complexity in the PE to a certain extent and improve the processing efficiency.
[0107] In another possible design, the number of encoding units is less than the number of processing units. At this time, at least one encoding unit is used to perform pre-encoding processing on at least two processing units to obtain their corresponding first input data. For example, in a data processing system, there may be 5 processing units and 4 encoding units. In this implementation method, there must be one encoding unit for pre-encoding 2 processing units.
[0108] In a specific process of implementing the foregoing design, the data to be encoded corresponding to any two processing units is different. For example, if the data processing system has only 3 PEs and 2 encoding units, at this time, one encoding unit is responsible for encoding one PE, and the other encoding unit is responsible for encoding the data to be encoded of two PEs. Then, in a possible implementation method, if the data to be encoded between any two PEs is different, there must be one encoding unit responsible for encoding different data to be encoded. In this implementation method, the encoding unit undertakes the encoding task. Compared with the existing Figure 2 implementation method shown, it can effectively reduce the system complexity in the PE, reduce the maintenance cost, and is also beneficial to improving the data processing efficiency in the PE.
[0109] In the process of implementing the foregoing design in another way, if the number of encoding units is less than the number of processing units, there is also a situation where multiple said processing units reuse the first input data output by one encoding unit. Still taking the foregoing example, at this time, if there are two PEs with the same data to be encoded among the three PEs, one encoding unit can perform encoding processing on the data to be encoded once and input it to these two PEs.
[0110] Specifically, the so-called reuse means reusing the same convolutional kernel data and / or feature layer data among PEs. In this way, for the data to be encoded, the encoding unit only needs to perform encoding once and reuse the encoded first input data, so as to conveniently implement data processing, which is beneficial to further improving the processing efficiency.
[0111] Figure 8 and Figure 9 respectively show the principles of reusing convolutional kernel data and feature layer data respectively in convolutional processing. As Figure 8 shown, when performing convolutional processing, the convolutional kernel parameter A1 needs to perform the same operation with the feature layer data W1, W2... Wm one by one. Therefore, when performing this part of convolutional processing, the same convolutional kernel parameter A1 is reused among the feature layer data W1, W2... Wm. And as Figure 9 shown for the feature layer data W1, it needs to perform the same operation with the convolutional kernel data A1, A2... An one by one. Therefore, when performing this part of convolutional processing, the same feature layer data W1 is reused among the convolutional kernel data A1, A2... An.
[0112] In addition to Figure 8 or Figure 9 the reuse of unilateral data shown, it is also possible to simultaneously implement the reuse of feature layer data and convolutional kernel data, which will not be elaborated here. Figure 10 shows the principle of simultaneously reusing convolutional kernel data and feature layer data. As Figure 9 shown, the convolutional kernel data A1, A2... An are respectively transmitted and reused among PEs in the row direction, while the feature layer data W1, W2... Wm are respectively transmitted and reused among PEs in the column direction.
[0113] Based on the reuse principle as Figures 8 - 10 shown, in the technical solution provided by this application, there are at least the following implementation manners:
[0114] In one implementation manner, in this data processing system, only the first input data encoded by the encoding unit is reused. At this time, reference can be made to Figure 11 the schematic diagram of another data processing system shown. It should be noted that for the convenience of understanding, Figure 11The encoding unit is not shown, and only the first input data encoded by the encoding unit is shown, that is, Figure 11 the control signal in
[0115] As Figure 11 shown, among the m processing units shown in this data processing system, the control signal (that is, the first input data) is multiplexed in each PE (PE1~PEm). At this time, for any control signal that needs to be multiplexed, the encoding module only needs to perform one encoding process, and the processing module can multiplex the control signal obtained by this processing of the encoding module and execute the subsequent partial product generation process and accumulation process.
[0116] In another implementation, the second input data is the multiplexed data between multiple processing units.
[0117] At this time, in this data processing system, both the first input data and the second input data encoded by the encoding unit can be multiplexed. In other words, the second input data is the multiplexed data between multiple processing units. It should be noted that the multiple processing units that multiplex the first input data and the multiple processing units that multiplex the second input data can be the same or different, and this application has no special limitation on this.
[0118] At this time, reference can be made to Figure 12 the schematic diagram of another data processing system shown. In Figure 12 the data processing system shown, m×n PEs are shown. Figure 12 The data processing system shown is composed of n data processing systems as shown in Figure 11 connected by data lines. As Figure 12 shown, Figure 12 in the data processing system shown, the m×n PEs are arranged in a matrix. Among them, the m PEs in each column jointly multiplex the same control signal, that is, the first input data is multiplexed in the column direction; the n PEs in each row jointly multiplex the same second input data, that is, the second input data is multiplexed in the row direction.
[0119] In the data processing system as shown in Figure 12 shown, the encoding unit only needs to encode n data to be encoded and obtain n first input data, and then input the n first input data into the corresponding data multiplexing queues respectively. In contrast, in the existing data processing system composed of m×n PEs arranged in a matrix, at least m×n encoding logic modules encode the input data (that is, the multiplier data), with a huge encoding workload and low efficiency, and moreover, it also increases the hardware complexity.
[0120] Or, only the second input data can be multiplexed, and the first input data is not multiplexed, which will not be elaborated.
[0121] It should be noted that there is no special limitation on the arrangement mode of each PE in the data processing system in this application. In the actual application scenario, it can be arranged in a matrix form as shown in Figure 12 shown, or it can be arranged in sequence as shown in Figure 11 shown. Or, it can also be arranged in any custom manner. In this case, as long as each PE meets the algorithm of the convolutional neural network, this application has no special limitation on the physical arrangement mode and connection mode of each PE.
[0122] Hereinafter, the interaction mode between the encoding unit and the processing unit for the first input data will be described. Specifically, there are at least the following interaction modes between the two:
[0123] First, the storage media of the multiple processing units are sequentially connected in pairs to form a processing unit sequence. One of the processing units located at one end of the processing unit sequence is used to receive the first input data output by the one encoding unit, and the other processing units obtain the first input data from the storage media of the previous processing unit to which they are connected.
[0124] At this time, taking the data processing system shown in Figure 11 or Figure 12 as an example, in this implementation, one first input data is multiplexed by m PEs arranged in sequence in each column. At this time, the storage media of the m PEs are sequentially connected in pairs. Therefore, each PE can obtain the stored data of an adjacent PE through this connection method. Then, the encoding unit only needs to input the encoded first input data into one of the PEs, and the other adjacent processing units can obtain the first input data of this PE through the connection relationship between the storage media and use it as their own first input data. Specifically, the PE into which the encoding unit inputs the first input data can be any one of the m PEs. In the actual application scenario, it can be the PE located at the end point in the sequentially connected PE sequence. Such as Figure 11 or Figure 12 , the encoding unit can input the first input data into the first PE in each column. Then, the first PE in each column can receive the first input data output by the encoding unit and store it in its own storage media, and the second and subsequent PEs can also obtain the first input data stored in the storage media of its previous PE. Each PE can store the first input data in its own storage media after obtaining it.
[0125] Among them, the storage method can be cache or temporary storage or permanent storage, and this application has no special limitation on this. The type of the storage media can also include but is not limited to: register.
[0126] This implementation method requires connections between the storage media of PEs that reuse the same first input data. In this way, the encoding unit does not need to repeatedly input the first input data into each PE, and can also improve the processing efficiency to a certain extent. In addition, this implementation method has no special requirements for the arrangement order of PEs.
[0127] Second, the encoding unit inputs the first input data into multiple processing units respectively. In other words, the encoding unit can be connected to multiple processing units (wired or wirelessly), and then input the encoded first input data to each PE through this connection. This implementation method has no restrictions on the physical positions and connection relationships of PEs in the data processing system, and has higher flexibility.
[0128] Based on any of the foregoing designs, the hardware complexity of the data processing system can be reduced to a certain extent, and the processing efficiency can be improved.
[0129] For the encoding unit, when encoding the data to be encoded, it can encode the data to be encoded according to a preset encoding rule to obtain the first input data.
[0130] Among them, the encoding rules involved in this application may include but are not limited to: Booth encoding rules. Among them, the Booth encoding rules may include Booth encoding rules and their derivative rules. For example, the derivative rules of Booth encoding may include but are not limited to: Radix 4 encoding rules, or Radix 8 encoding rules.
[0131] Specifically, the Booth multiplier encodes every X bits of the multiplier (for example, the radix 4 Booth multiplier encodes every 3 bits, and the radix 8 Booth multiplier encodes every 4 bits), obtains the partial product of multiplying the X bits of the multiplier by the multiplicand by parsing the encoding, and then shifts and accumulates all partial products according to their corresponding positions of X bits to find the final product.
[0132] For example, Figure 13 shows a schematic diagram of the processing principle of a radix 4-bit Booth multiplier. As Figure 13 shown, a partial product selection table can be preset in advance. The partial product selection table is actually a concretization of the right partial product formula. That is to say, substituting the multiplier value into the right formula to obtain the partial product and forming this partial product selection table. In the actual scenario, the expression of the partial product formula can be custom preset. This application only uses Figure 13 the principle shown as an example to illustrate its encoding rules.
[0133] As Figure 13As shown, a radix-4 Booth multiplier encodes every 3 bits. That is, an "inspection bit" is added to the lower bits of the multiplier. For example, if inspecting y i+1 y i y i-1 , then the y i-1 among them is the inspection bit; if inspecting y i+3 y i+2 y i+1 , then the y i+1 among them is the inspection bit.
[0134] In the specific encoding process, the situation of padding is often involved. That is, when the data to be encoded does not meet the bit requirement of encoding, padding is performed on it to make it meet the encoding bit requirement. As Figure 13 shown, the data to be encoded is: 10011010, and its number of bits does not meet the requirement. Therefore, before the multiplier performs encoding, 2 bits and 1 bit are respectively inserted at the beginning and end of the data to be encoded ( Figure 13 shown in the form of parentheses plus numbers), so that the data to be encoded is extended to 11 bits, thus meeting the bit requirement of encoding every 3 bits.
[0135] After the number of bits meets the requirement, the multiplier encodes the 1st, 3rd, 5th, 7th, 9th bit and its adjacent 2 bits (for example, 1 bit → 0 / 1 / 2 bit) in the padded data to be encoded to generate the control signal of the partial product module as the first input data. As Figure 13 shown, its specific encoding method is to determine the partial product corresponding to every 3-bit data according to the partial product selection table, and obtain the sum of partial products after numerically representing each partial product.
[0136] When the first input data is the control signal encoded by Booth encoding or its improved encoding rule, the encoding unit inputs the control signal into the partial product module, and the partial product module and the accumulation module (the first and the second) perform subsequent processing on the control signal and the multiplicand to obtain the convolution result of the PE.
[0137] For the method described above, the required parameters are converted into control signals by software methods before convolution processing. Thus, during the convolution processing, the PE only needs to process this control signal instead of the actual data to be encoded, eliminating the internal encoding logic of the hardware, reducing the device complexity and improving the device operation speed. And in the convolutional neural network, since the output data of the Kth layer is the input data of the (K + 1)th layer, therefore, for any data of the (K + 1)th layer, the output data of the Kth layer can be converted into a control signal, and convolution processing is performed on the control signal.
[0138] In addition, it should be noted that in the present application, there can be multiple different designs for the form of the encoding unit.
[0139] In one possible design, the encoding unit can be one or more dedicated hardware devices actually existing in the data processing system. For example, dedicated logic circuits. At this time, only an additional hardware device needs to be set in the data processing system, and this hardware device can transmit data to at least one PE.
[0140] In another possible design, the encoding unit can reuse the existing hardware in the data processing system, such as the software code of a CPU, GPU, or other processing devices, and implement the encoding process through software. For example, if the data processing system is a computer processor and each PE is a processing sub-module in the processor, then this encoding unit can also directly reuse the processor (or one or more sub-modules different from the processing sub-module), and the processor implements the foregoing method in software. At this time, there is no need to additionally set up hardware devices, which is more conducive to saving costs and reducing hardware complexity.
[0141] In addition, for the consideration of further improving efficiency and reducing the hardware complexity of the system, the present application further improves the hardware composition of the PE: replacing the carry-propagate adder (CPA) at the last stage in the second accumulation module of the multiplier with a carry-save adder (CSA). At this time, the second accumulation module of the processing unit is composed of multiple carry-save adders CSA. And the first accumulation module of the processing unit is composed of multiple carry-save adders CSA and a carry-propagate adder CPA.
[0142] Figure 14 shows a schematic diagram of the architecture of the multiplier in another processing unit provided by the present application. Compared with Figure 7 the multiplier in the shown PE, Figure 14 in the shown PE, the CPA of the multiplier is removed. Thus, the data output at the last stage of the multiplier is the carry and sum signals generated by the CSA (final multiplication result = carry + sum). Through this design, when the first accumulation module performs the accumulation process, it only needs to continue to perform carry addition processing on the carry and sum signals output by the second accumulation module, and finally summarize them in the CPA, and the CPA outputs the numerical convolution result.
[0143] Figure 7 With Figure 14The 4:2 in the CSA shown means 4 inputs and 2 outputs, and the 2:1 means 2 inputs and 1 output. Among them, the signals of the CSA are in pairs, and a pair of signals includes 1 carry and 1 sum signal. Therefore, when it has 2 inputs, the value of the input signals is 4; when it has 1 input, the value of the input signals is 2. The CPA, on the other hand, can output a numerical result, and its output data is 1. Therefore, after the design as described in Figure 14 Since the last stage of each multiplier is changed to a CSA, the model of the first-stage CSA of the first accumulation module also needs to be adjusted adaptively, from the 2:1 shown in Figure 7 to 4:2 as shown in
[0144] Compared with the architecture of the PE shown in Figure 7 when the PE shown in Figure 14 performs the accumulation process, it does not need to calculate the accurate product of each pair of values in the step of multiplying numbers pairwise. It can use an equivalent product expression to perform the accumulation and generate the convolution result. For the last-stage CPA of the multiplier logic, the carry and sum output by the penultimate-stage CSA are accumulated in the subsequent accumulation unit to obtain the final convolution calculation result. In this way, on the premise of ensuring the mathematical equivalence of the output result inside the PE, the slow and high-logic-cost CPA accumulation logic is removed, and the first layer of the first accumulation module is modified to adapt to this modification, thereby improving the logic operation speed and reducing the device overhead.
[0145] Based on the foregoing design, the present application also claims protection for an encoding unit and a processing unit.
[0146] First, for the encoding unit 410:
[0147] On the one hand, Figure 15 shows a functional block diagram of the encoding unit 410, and the encoding unit 410 includes:
[0148] An acquisition module 411 for acquiring data to be encoded;
[0149] An encoding module 412 for encoding the data to be encoded to obtain first input data;
[0150] A sending module 413 for inputting the first input data into the processing unit, so that the processing unit performs partial product generation processing on the first input data and second input data to obtain a partial product result, and performs accumulation processing on the partial product result to obtain a convolution result;
[0151] wherein, the data to be encoded is convolution kernel data, and the second input data is feature layer data; or, the data to be encoded is feature layer data, and the second input data is convolution kernel data.
[0152] In one possible design, the number of the encoding units is smaller than the number of the processing units.
[0153] In another possible design, the first input data is multiplexed by multiple processing units.
[0154] In another possible design, the storage media of the multiple processing units are connected in pairs in sequence to form a processing unit sequence, and the sending module 413 is specifically used to:
[0155] The first input data is input to one of the processing units located at one end of the sequence of processing units.
[0156] In another possible design, the sending module 413 is specifically used to:
[0157] The first input data is input into the plurality of processing units respectively.
[0158] In another possible design, the first input data is used to control the processing unit to perform the partial product generation process.
[0159] In another possible design, the second input data is multiplexed data between multiple processing units.
[0160] In another possible design, the encoding module 412 is specifically used to:
[0161] The data to be encoded is encoded according to a preset encoding rule.
[0162] It should be understood that the above Figure 15 The division of the modules of the coding unit 410 shown is only a division of logical functions. In actual implementation, they can be fully or partially integrated into one physical entity, or they can be physically separated. And these modules can all be implemented in the form of software calling through processing elements; they can also be all implemented in the form of hardware; some modules can also be implemented in the form of software calling through processing elements, and some modules can be implemented in the form of hardware. For example, the coding module 412 can be a separately established processing element, or it can be integrated in the coding unit 410, such as a chip of the terminal. In addition, it can also be stored in the memory of the coding unit 410 in the form of a program, and called and executed by a processing element of the coding unit 410. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together, or they can be implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each module above can be completed by an integrated logic circuit of hardware in the processor element or an instruction in the form of software.
[0163] For example, these modules above can be one or more integrated circuits configured to implement the above methods, such as: one or more Application Specific Integrated Circuits (ASICs), or, one or more digital signal processors (DSPs), or, one or more Field Programmable Gate Arrays (FPGAs), etc. Again, when a certain module above is implemented in the form of a processing element scheduler, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processors that can call programs. Again, these modules can be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0164] On the other hand, Figure 16 A schematic diagram of the physical structure of the encoding unit 410 is shown. The encoding unit 410 includes:
[0165] A memory 4101;
[0166] A processor 4102; and
[0167] A computer program;
[0168] Wherein, the computer program is stored in the memory 4101 and is configured to be executed by the processor 4102 to implement the method described in any implementation manner implemented by the encoding unit as described above.
[0169] Wherein, the number of processors 4102 in the encoding unit 410 can be one or more. The processor 4102 can also be referred to as a processing unit and can implement certain control functions. The processor 4102 can be a general-purpose processor or a dedicated processor, etc. In an optional design, the processor 4102 can also store instructions, and the instructions can be run by the processor 4102 so that the encoding unit 410 executes the method described in the above method embodiments.
[0170] In yet another possible design, the encoding unit 410 can include a circuit, and the circuit can implement the functions of sending, receiving, or communicating in the foregoing method embodiments.
[0171] Optionally, the number of memories 4101 in the encoding unit 410 may be one or more. Instructions or intermediate data are stored on the memories 4101. The instructions can be run on the processor 4102, so that the encoding unit 410 executes the method described in the foregoing method embodiments. Optionally, other relevant data may also be stored in the memories 4101. Optionally, instructions and / or data may also be stored in the processor 4102. The processor 4102 and the memories 4101 may be provided separately or integrated together.
[0172] In addition, as Figure 16 shown, a transceiver 4103 is further provided in the encoding unit 410. Herein, the transceiver 930 may be referred to as a transceiver unit, a transceiver, a transceiver circuit, or a transceiver, etc., and is used for data transmission or communication with a test device or other terminal devices, which will not be elaborated herein.
[0173] As Figure 16 shown, the memories 4101, the processor 4102, and the transceiver 4103 are connected and communicate through a bus.
[0174] When the encoding unit 410 is used to implement the method corresponding to Figure 5 , for example, the transceiver 4103 may issue the packet under test to each test terminal, and the transceiver 4103 may also be used to receive the test running data fed back by each test terminal. The processor 4102 is used to complete the corresponding determination or control operations. Optionally, corresponding instructions may also be stored in the memories 4101. The specific processing manners of each component may refer to the relevant descriptions in the foregoing embodiments.
[0175] Secondly, for the processing unit 420:
[0176] On the one hand, Figure 17 shows a functional block diagram of the processing unit 420. The processing unit 420 includes: a partial product module 4211 and an accumulation processing module 422;
[0177] The partial product module 4211 is used to perform partial product generation processing on the first input data and the second input data to obtain a partial product result;
[0178] The accumulation processing module 422 is used to perform accumulation processing on the partial product result to obtain a convolution result;
[0179] Among them, the first input data is obtained by the encoding unit 410 acquiring the data to be encoded and performing encoding processing on the data to be encoded;
[0180] Wherein, the data to be encoded is convolution kernel data, and the second input data is feature layer data; or, the data to be encoded is feature layer data, and the second input data is convolution kernel data.
[0181] In another possible design, the accumulation processing module 422 includes: a first accumulation module 4222 and at least two second accumulation modules 4212, wherein one of the second accumulation modules 4212 belongs to one multiplier 421;
[0182] The second accumulation module 4212 is configured to perform a first accumulation process on the partial product result to obtain an accumulation result;
[0183] The first accumulation module 4222 is configured to perform a second accumulation process on the accumulation results respectively output by the second accumulation modules 4212 of the at least two multipliers 421 to obtain the convolution result.
[0184] In another possible design, the number of the encoding units 410 is less than the number of the processing units 420.
[0185] In another possible design, multiple processing units 420 multiplex the first input data output by one encoding unit 410.
[0186] In another possible design, the storage media of the multiple processing units are connected in sequence pairwise to form a processing unit sequence;
[0187] When the processing unit is at one end of the processing unit sequence, the partial product module is further configured to: receive the first input data output by one encoding unit;
[0188] When the processing unit is another processing unit in the processing unit sequence, the partial product module is further configured to: obtain the first input data stored in the storage media of the previous processing unit connected thereto.
[0189] In another possible design, the partial product module 4211 is further configured to:
[0190] Receive the first input data output by one encoding unit.
[0191] In another possible design, the first input data is used to control the processing unit to perform the partial product generation process.
[0192] In another possible design, the second input data is multiplexed data among multiple processing units.
[0193] In another possible design, the second accumulation module of the processing unit is composed of multiple carry-save adders CSA.
[0194] In another possible design, the first accumulation module of the processing unit is composed of a plurality of carry-save adders (CSA) and a carry-lookahead adder (CPA).
[0195] It should be understood that the division of each module of the processing unit 420 shown above is only a division of logical functions. In actual implementation, they can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by processing elements; they can also all be implemented in the form of hardware; or some modules can be implemented in the form of software called by processing elements, and some modules can be implemented in the form of hardware. For example, the accumulation processing module 422 can be a separately established processing element, or can be integrated in the processing unit 420, such as implemented in a certain chip of the terminal. In addition, it can also be stored in the memory of the processing unit 420 in the form of a program, and called and executed by a certain processing element of the processing unit 420 to perform the functions of each above module. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together or can be independently implemented. The processing element mentioned here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each above module can be completed by the integrated logic circuit in the processor element or the instruction in the form of software. Figure 17 For example, these above modules can be one or more integrated circuits configured to implement the above method, such as: one or more application specific integrated circuits (ASIC), or, one or more digital signal processors (DSP), or, one or more field programmable gate arrays (FPGA), etc. Again, when a certain above module is implemented in the form of a processing element scheduling program, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call programs. Again, these modules can be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0196] On the other hand,
[0197] shows a schematic diagram of the physical structure of the processing unit 420, and the processing unit 420 includes: Figure 18
[0198] a memory 4201;
[0199] a processor 4202; and
[0200] Computer program;
[0201] Wherein, the computer program is stored in the memory 4201 and is configured to be executed by the processor 4202 to implement the method described in any of the implementation manners implemented by the encoding unit as described above.
[0202] Wherein, the number of the processors 4202 in the processing unit 420 may be one or more. The processor 4202 may also be referred to as a processing unit and can implement certain control functions. The processor 4202 may be a general-purpose processor or a dedicated processor, etc. In an alternative design, the processor 4202 may also store instructions, and the instructions may be run by the processor 4202, so that the processing unit 420 executes the method described in the above method embodiments.
[0203] In yet another possible design, the processing unit 420 may include a circuit, and the circuit may implement the functions of sending, receiving, or communicating in the foregoing method embodiments.
[0204] Optionally, the number of the memories 4201 in the processing unit 420 may be one or more. Instructions or intermediate data are stored on the memory 4201, and the instructions may be run on the processor 4202, so that the processing unit 420 executes the method described in the above method embodiments. Optionally, other relevant data may also be stored in the memory 4201. Optionally, instructions and / or data may also be stored in the processor 4202. The processor 4202 and the memory 4201 may be provided separately or integrated together.
[0205] In addition, as Figure 18 shown, a transceiver 4203 is further provided in the processing unit 420. Wherein, the transceiver 930 may be referred to as a transceiver unit, a transceiver, a transceiver circuit, or a transceiver, etc., and is used for data transmission or communication with a test device or other terminal devices, which will not be elaborated herein.
[0206] As Figure 18 shown, the memory 4201, the processor 4202, and the transceiver 4203 are connected and communicate through a bus.
[0207] When the processing unit 420 is used to implement the method corresponding to Figure 5 , for example, the transceiver 4203 may publish the package under test to each test terminal, and the transceiver 4203 may also be used to receive the test operation data fed back by each test terminal. The processor 4202 is used to complete corresponding determination or control operations. Optionally, corresponding instructions may also be stored in the memory 4201. The specific processing manners of each component may refer to the relevant descriptions in the foregoing embodiments.
[0208] In addition, an embodiment of the present invention provides a readable storage medium storing computer-executable instructions, which are used to implement the method described in any implementation manner of the encoding unit and / or the processing unit when being executed by a processor.
[0209] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A data processing system, characterized in that, It includes an encoding unit and a processing unit; The encoding unit is used for: Obtain data to be encoded; Perform encoding processing on the data to be encoded to obtain first input data, where the first input data is a control signal for a partial product generation module; The processing unit is used for: Perform partial product generation processing on the first input data and second input data to obtain a partial product result; Perform accumulation processing on the partial product result to obtain a convolution result; Wherein, the data to be encoded is convolution kernel data, and the second input data is feature layer data; or, the data to be encoded is feature layer data, and the second input data is convolution kernel data; The number of the encoding units is less than the number of the processing units, and multiple processing units share the first input data output by one encoding unit.
2. The system according to claim 1, characterized in that, The processing unit includes a first accumulation module and at least two multipliers, and each multiplier includes a partial product module and a second accumulation module; The partial product module is used for performing the partial product generation processing; The second accumulation module is used for performing first accumulation processing on the partial product result to obtain an accumulation result; The first accumulation module is used for performing second accumulation processing on the accumulation results respectively output by the second accumulation modules of the at least two multipliers to obtain the convolution result.
3. The system according to claim 1, characterized in that, The storage media of the multiple processing units are sequentially connected pairwise to form a processing unit sequence. One processing unit located at one end of the processing unit sequence is used to receive the first input data output by the one encoding unit, and other processing units obtain the first input data from the storage media of the previous processing unit to which they are connected.
4. The system according to claim 1, characterized in that, The one encoding unit is further used for: Input the first input data into the multiple processing units respectively.
5. The system according to any one of claims 1-4, characterized in that, The first input data is used to control the processing unit to execute the partial product generation processing.
6. The system according to any one of claims 1-4, characterized in that, The second input data is shared data among the multiple processing units.
7. The system according to any one of claims 1-4, characterized in that, The performing encoding processing on the data to be encoded includes: Perform encoding processing on the data to be encoded according to a preset encoding rule.
8. The system according to any one of claims 2-4, characterized in that, The second accumulation module of the processing unit is composed of multiple carry-save adders (CSA).
9. The system according to claim 8, characterized in that, The first accumulation module of the processing unit is composed of multiple carry-save adders (CSA) and one carry-lookahead adder (CPA).
10. A data processing method, characterized in that, Applied to the encoding unit, the method includes: Obtain data to be encoded; Perform encoding processing on the data to be encoded to obtain first input data, where the first input data is a control signal for a partial product generation module; Input the first input data into the processing unit, so that the processing unit performs partial product generation processing on the first input data and second input data to obtain a partial product result, and perform accumulation processing on the partial product result to obtain a convolution result; Wherein, the data to be encoded is convolution kernel data, and the second input data is feature layer data; or, the data to be encoded is feature layer data, and the second input data is convolution kernel data; the number of the encoding units is less than the number of the processing units, and the first input data is shared by the multiple processing units.
11. The method according to claim 10, characterized in that, The storage media of the multiple processing units are connected pairwise in sequence to form a processing unit sequence, and the method further includes: Input the first input data into one of the processing units located at one end of the processing unit sequence.
12. The method according to claim 10, characterized in that, The method further includes: Input the first input data into the multiple processing units respectively.
13. The method according to any one of claims 10-12, characterized in that, The first input data is used to control the processing unit to perform the partial product generation process.
14. The method according to any one of claims 10-12, characterized in that, The second input data is multiplexed data between multiple processing units.
15. The method according to any one of claims 10-12, characterized in that, The encoding process for the data to be encoded includes: Encoding the data to be encoded according to a preset encoding rule.
16. A data processing method, characterized in that, Applied to a processing unit, the method includes: Performing a partial product generation process on the first input data and the second input data to obtain a partial product result; Performing an accumulation process on the partial product result to obtain a convolution result; Wherein, the first input data is obtained by an encoding unit acquiring the data to be encoded and performing an encoding process on the data to be encoded, and the first input data is a control signal for the partial product generation module; Wherein, the data to be encoded is convolution kernel data, and the second input data is feature layer data; or, the data to be encoded is feature layer data, and the second input data is convolution kernel data; The number of encoding units is less than the number of processing units, and the multiple processing units multiplex the first input data output by one encoding unit.
17. The method according to claim 16, characterized in that, The processing unit includes at least two multipliers; the performing an accumulation process on the partial product result to obtain a convolution result includes: Performing a first accumulation process on the partial product results of the multipliers to obtain an accumulation result; Performing a second accumulation process on the accumulation results of the at least two multipliers to obtain the convolution result.
18. The method according to claim 16, characterized in that, The storage media of the multiple processing units are connected pairwise in sequence to form a processing unit sequence; When the processing unit is located at one end of the processing unit sequence, the method further includes: receiving the first input data output by the one encoding unit; When the processing unit is another processing unit in the processing unit sequence, the method further includes: obtaining the first input data stored in the storage media of the previous connected processing unit.
19. The method according to claim 16, characterized in that, The method further includes: Receiving the first input data output by the one encoding unit.
20. The method according to any one of claims 17-19, characterized in that, The first input data is used to control the processing unit to perform the partial product generation process.
21. The method according to any one of claims 17-19, characterized in that, The second input data is multiplexed data between multiple processing units.
22. The method according to any one of claims 17-19, characterized in that, Each multiplier includes a second accumulation module, and the second accumulation module is composed of multiple carry-save adders (CSA).
23. The method according to claim 22, characterized in that, The processing unit further includes a first accumulation module, and the first accumulation module is composed of multiple carry-save adders (CSA) and one carry-lookahead adder (CPA).
24. An encoding unit, characterized in that, Includes: An acquisition module for acquiring the data to be encoded; An encoding module for encoding the data to be encoded to obtain the first input data, and the first input data is a control signal for the partial product generation module; A sending module, configured to input the first input data into a processing unit, so that the processing unit performs partial product generation processing on the first input data and second input data to obtain a partial product result, and performs accumulation processing on the partial product result to obtain a convolution result; Wherein, the data to be encoded is convolution kernel data, and the second input data is feature layer data; or, the data to be encoded is feature layer data, and the second input data is convolution kernel data; The number of encoding units is less than the number of processing units, and the first input data is multiplexed by multiple processing units.
25. The coding unit according to claim 24, wherein The storage media of the multiple processing units are sequentially connected pairwise to form a processing unit sequence. The sending module is specifically configured to: Input the first input data into one processing unit located at one end of the processing unit sequence.
26. The coding unit according to claim 24, wherein The sending module is specifically configured to: Input the first input data into the multiple processing units respectively.
27. The coding unit according to any one of claims 24-26, wherein The first input data is used to control the processing unit to perform the partial product generation processing.
28. The coding unit according to any one of claims 24-26, wherein The second input data is multiplexed data among the multiple processing units.
29. The coding unit according to any one of claims 24-26, wherein The encoding module is specifically configured to: Perform encoding processing on the data to be encoded according to a preset encoding rule.
30. A processing unit, wherein Includes: A partial product module and an accumulation processing module; The partial product module is configured to perform partial product generation processing on the first input data and second input data to obtain a partial product result; The accumulation processing module is configured to perform accumulation processing on the partial product result to obtain a convolution result; Wherein, the first input data is obtained by an encoding unit acquiring the data to be encoded and performing encoding processing on the data to be encoded, and the first input data is a control signal for generating the partial product module; Wherein, the data to be encoded is convolution kernel data, and the second input data is feature layer data; or, the data to be encoded is feature layer data, and the second input data is convolution kernel data; the number of encoding units is less than the number of processing units, and multiple processing units multiplex the first input data output by one encoding unit.
31. The processing unit according to claim 30, wherein The accumulation processing module includes: a first accumulation module and at least two second accumulation modules, wherein one second accumulation module belongs to one multiplier; The second accumulation module is configured to perform first accumulation processing on the partial product result to obtain an accumulation result; The first accumulation module is configured to perform second accumulation processing on the accumulation results respectively output by the second accumulation modules of the at least two multipliers to obtain the convolution result.
32. The processing unit according to claim 30, wherein The storage media of the multiple processing units are sequentially connected pairwise to form a processing unit sequence; When the processing unit is located at one end of the processing unit sequence, the partial product module is further configured to: receive the first input data output by the one encoding unit; When the processing unit is another processing unit in the processing unit sequence, the partial product module is further configured to: obtain the first input data stored in the storage medium of the previous processing unit connected thereto.
33. The processing unit according to claim 30, wherein The partial product module is further configured to: Receive the first input data output by the one encoding unit.
34. The processing unit according to any one of claims 30-33, wherein The first input data is used to control the processing unit to execute the partial product generation process.
35. The processing unit according to any one of claims 30-33, wherein The second input data is multiplexed data among multiple processing units.
36. The processing unit according to any one of claims 30 - 33, wherein The second accumulation module of the processing unit is composed of multiple carry-save adders (CSA).
37. The processing unit according to claim 36, wherein The first accumulation module of the processing unit is composed of multiple carry-save adders (CSA) and a carry-lookahead adder (CPA).
38. An encoding unit, wherein Comprising: A memory; A processor; And A computer program; Wherein, the computer program is stored in the memory and is configured to be executed by the processor to implement the method according to any one of claims 10-15.
39. A processing unit, wherein Comprising: A memory; A processor; And A computer program; Wherein, the computer program is stored in the memory and is configured to be executed by the processor to implement the method according to any one of claims 16-23.
40. A computer-readable storage medium, wherein Stored with computer-executable instructions, when the computer-executable instructions are executed by the processor, they are used to implement the method according to any one of claims 10-23.
41. A computer program product, wherein Stored with computer-executable instructions, when the computer-executable instructions are executed by the processor, they are used to implement the method according to any one of claims 10-23.
Citation Information
Patent Citations
Neural network method and circuit for optimizing sparsity matrix operation
CN108647774A