A data processing method and device, chip, equipment, and storage medium
By introducing multiply-accumulators and intermediate result accumulators into the depthwise separable convolution computation unit, the problem of excessively long data transmission paths in depthwise separable convolution computation is solved, and low-power data processing is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
- Filing Date
- 2022-01-11
- Publication Date
- 2026-05-01
AI Technical Summary
In depthwise separable convolution computation, the high power consumption is caused by the long data transmission path, especially when no accumulation operation is required in the input channel direction. The excessively long data transmission path leads to increased energy consumption.
Multiple depthwise separable convolutional computation units are deployed in the computational unit array. Each unit contains multiple multiply-accumulators, multiple intermediate result accumulators, and a multiplexer. The data transmission path is shortened through internal accumulation operations, and the convolutional results are accumulated using the intermediate result accumulators, thereby reducing the data transmission outside the computational unit array.
By using internal accumulation operations, the data transmission path is shortened, data transmission power consumption is reduced, and computing efficiency is improved.
Smart Images

Figure CN116484925B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of deep learning, and more particularly to a data processing method and apparatus, chip, device, and storage medium. Background Technology
[0002] In artificial intelligence network algorithms, due to the large computational cost of traditional convolution, depthwise separable convolution is often used for computation. For depthwise separable convolution, there is no need to perform an accumulation operation in the direction of the input channel. Instead, the convolution results of each input channel are accumulated. Therefore, an accumulator (ACC) needs to be added to the computational unit array, and the convolution results of each input channel are input into the ACC to perform the accumulation operation. This results in a long data transmission path, which in turn leads to the problem of high power consumption during data transmission. Summary of the Invention
[0003] This application provides a data processing method, apparatus, chip, device, and storage medium that can reduce data transmission power consumption.
[0004] The technical solution of this application is implemented as follows:
[0005] In a first aspect, embodiments of this application propose a data processing apparatus, wherein a computing unit array is deployed in the apparatus, the computing unit array including multiple depthwise separable convolution computing units, each depthwise separable convolution computing unit including: multiple multiply-accumulate units, multiple intermediate result accumulators, and a multiplexer; wherein the output of each multiply-accumulate unit is connected to the input of an intermediate result accumulator, and the input of the multiplexer is connected to the outputs of the multiple multiply-accumulate units and the outputs of the multiple intermediate result accumulators respectively;
[0006] The multiple multiply-accumulate operators are used to perform multiply-accumulate operations on multiple convolutional window data in one channel of input feature data.
[0007] The plurality of intermediate result accumulators are used to perform accumulation operations on the intermediate results generated by the corresponding multiply-accumulators respectively.
[0008] The multiplexer is used to select the output result of each multiply-accumulate unit or to select the output result of each intermediate result accumulator.
[0009] Secondly, embodiments of this application propose a data processing method applied in the aforementioned data processing apparatus, the method comprising:
[0010] Based on the number of multiply-accumulators in a depth-separable convolutional computation unit, multiple convolutional window data are determined from the input feature data of each channel in the feature data memory.
[0011] The multiple convolutional window data are input into a corresponding depthwise separable convolutional computation unit for depthwise separable convolution computation to obtain output feature data; and the output feature data is output to the feature data memory.
[0012] Thirdly, embodiments of this application provide a chip that includes the data processing apparatus described above.
[0013] Fourthly, this application provides a computing device comprising: a processor, a memory, and a communication bus; the communication bus is used to establish a communication connection between the processor and the memory; the processor implements the above-mentioned data processing method when executing a running program stored in the memory.
[0014] Fifthly, embodiments of this application propose a storage medium storing a computer program thereon, which, when executed by a processor, implements the aforementioned data processing method.
[0015] This application provides a data processing method, apparatus, chip, device, and storage medium. The data processing apparatus deploys a computing unit array, which includes multiple depthwise separable convolutional computing units. Each depthwise separable convolutional computing unit includes: multiple multiply-accumulate units, multiple intermediate result accumulators, and a multiplexer. The output of each multiply-accumulate unit is connected to the input of an intermediate result accumulator, and the input of the multiplexer is connected to the outputs of both the multiple multiply-accumulate units and the multiple intermediate result accumulators. The multiple multiply-accumulate units perform multiply-accumulate operations on multiple convolutional window data in one channel of input feature data at a time. The multiple intermediate result accumulators accumulate the intermediate results generated by their respective multiply-accumulate units. The multiplexer selects to output either the result of each multiply-accumulate unit or the accumulated result of each intermediate result accumulator. Using the above-mentioned data processing device implementation scheme, a depth-separable convolution calculation unit is proposed. The depth-separable convolution calculation unit includes a multiply-accumulator, an intermediate result accumulator, and a multiplexer. It can perform the accumulation operation of convolution results inside the depth-separable convolution calculation unit, shorten the data transmission path, and thus reduce the data transmission power consumption. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of a depthwise separable convolution calculation process;
[0017] Figure 2 This is a schematic diagram of a hardware deployment architecture for existing depthwise separable convolutions.
[0018] Figure 3This is a schematic diagram of an existing data mapping method for convolutional window data to computational units.
[0019] Figure 4 This is a schematic diagram of an existing diagonal mapping method;
[0020] Figure 5 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application;
[0021] Figure 6 This is an exemplary structural diagram of a computing unit array with multiple depth-separable convolution computing units arranged diagonally, as provided in an embodiment of this application.
[0022] Figure 7 This is an exemplary structural diagram of a column of depth-separable convolutional computing units arranged on the far right of a computing unit array, provided for an embodiment of this application.
[0023] Figure 8 A schematic diagram of an exemplary depth-separable convolutional computation unit provided in an embodiment of this application;
[0024] Figure 9 A schematic diagram illustrating an exemplary data mapping method from convolutional window data to computational units, provided for embodiments of this application;
[0025] Figure 10 This is an exemplary schematic diagram of flattening convolution window data and feeding it into a depth-separable convolution calculation unit, provided as an embodiment of this application.
[0026] Figure 11 This is an exemplary flowchart illustrating how a 5x5 convolution window of data is mapped to a depthwise separable convolution unit in stages and then used for depthwise separable convolution calculation, as provided in this application embodiment.
[0027] Figure 12 A flowchart illustrating a data processing method provided in an embodiment of this application;
[0028] Figure 13 This is a schematic diagram of the structure of a chip provided in an embodiment of this application;
[0029] Figure 14 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application. Detailed Implementation
[0030] In order to gain a more detailed understanding of the features and technical content of the embodiments of this application, the implementation of the embodiments of this application will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for reference and illustration only and are not intended to limit the embodiments of this application.
[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0032] In the following description, references to "some embodiments" refer to a subset of all possible embodiments. It is understood that "some embodiments" may be the same or different subsets of all possible embodiments and may be combined with each other without conflict. It should also be noted that the terms "first, second, third" used in the embodiments of this application are merely for distinguishing similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0033] The computation process of depthwise separable convolution is as follows: Figure 1 As shown, k0-k3 are the convolutional kernels of 4 channels, and f0-f3 are the convolutional window data of 4 channels in the input feature data. The convolutional kernel data of each channel is subjected to DWC operation with the convolutional window data of each corresponding channel, and then the feature data is directly output without the need for accumulation in the input channel direction.
[0034] Currently, the hardware deployment architecture for depthwise separable convolutions is as follows: Figure 2 As shown, each box represents a computation unit. Vertically, the accumulated result of each computation unit is passed up to the next computation unit and accumulated with the result of the next computation unit, continuing upwards until it reaches the accumulator at the top of the computation unit array. For a single computation unit, the data mapping method from the convolution window to the computation unit can be found in [reference needed]. Figure 3 , Figure 3 This is a 3x3 convolution data mapping method. For a single computation unit, box 1 in the input image represents a search window region. Each computation unit can process data within one search window region simultaneously, which can contain up to 16 3x3 convolution computation windows. After the data from the 16 3x3 convolution windows is read out, a data flattening operation is performed, and the data is transmitted to the computation unit. Each computation unit pre-stores the weight information required for the convolution operation, and each computation unit shares a weight kernel. Figure 3 In the w0 array, after the computation unit receives the convolution window data, it can perform multiplication and accumulation operations. Based on the characteristics of convolution operations, a 4*10 convolution window can generate 2*8 output data points. For example... Figure 3 As shown in box 2, the size is 2*8=16.
[0035] Since depthwise separable convolutions do not have accumulation behavior in the input channel direction, Figure 2 The convolution acceleration engine accumulates the convolution results of each input channel vertically. Therefore, when mapping data, it needs to be done diagonally, such as... Figure 4 As shown, the input feature data of channel 0 is transmitted to computing unit 0 in the computing unit array. Computing unit 0 transmits the convolution result to ACC0 for accumulation, and ACC0 returns the accumulated result to the feature data storage unit, Static Random-Access Memory (SRAM). The input feature data of channel 1 is transmitted to computing unit 1 in the computing unit array. Computing unit 1 transmits the convolution result to ACC1 for accumulation, and ACC1 returns the accumulated result to the feature data storage unit, SRAM. The input feature data of channel 2 is transmitted to computing unit 2 in the computing unit array. Computing unit 2 transmits the convolution result to ACC2 for accumulation, and ACC2 returns the accumulated result to the feature data storage unit, SRAM. The input feature data of channel 3 is transmitted to computing unit 3 in the computing unit array. Computing unit 3 transmits the convolution result to ACC3 for accumulation, and ACC3 returns the accumulated result to the feature data storage unit, SRAM. This mirror-reflection path aims to ensure that the accumulated result of each ACC has the same path length back to SRAM, and can arrive at SRAM at the same time.
[0036] While the above approach can functionally achieve depthwise separable convolution and support both traditional and depthwise separable convolution using a single hardware architecture, for depthwise separable convolution, to maintain the same data transmission path as traditional convolution, the convolution result of each computational unit needs to be sent to the ACC (Accumulated Accumulator) before being returned. This results in a longer data transmission path for each operation, leading to higher power consumption. Although depthwise separable convolution and traditional convolution share the same computational unit array, the data transmission path of traditional convolution is excessively wasteful for depthwise separable convolution.
[0037] To address the aforementioned problems, embodiments of this application provide a data processing apparatus 1, such as... Figure 5As shown, the device 1 deploys a computing unit array 10, which includes multiple depthwise separable convolution computing units 100. Each depthwise separable convolution computing unit 100 includes: multiple multiply-accumulate units 1001, multiple intermediate result accumulators 1002, and a multiplexer 1003. The output of each multiply-accumulate unit 1001 is connected to the input of an intermediate result accumulator 1002, and the input of the multiplexer 1003 is connected to the outputs of the multiple multiply-accumulate units 1001 and the outputs of the multiple intermediate result accumulators 1002, respectively.
[0038] The multiple multiply-accumulate units 1001 are used to perform multiply-accumulate operations on multiple convolutional window data in one channel of input feature data.
[0039] The plurality of intermediate result accumulators 1002 are used to perform accumulation operations on the intermediate results generated by the corresponding multiply-accumulators respectively;
[0040] The multiplexer 1003 is used to select the output result of each multiply-accumulator or to select the output result of each intermediate result accumulator.
[0041] The data processing apparatus proposed in this application embodiment is applicable to scenarios involving depthwise separable convolution computation.
[0042] refer to Figure 5 In this embodiment of the application, the computing unit array 10 is composed of depthwise separable convolution computing units 100 and non-depthwise separable convolution computing units 101. If a depthwise separable convolution operation is performed, multiple depthwise separable convolution computing units are started and non-depthwise separable convolution computing units are disconnected. Correspondingly, if a depthwise separable convolution operation is not performed, multiple depthwise separable convolution computing units are disconnected and non-depthwise separable convolution computing units are started.
[0043] In this embodiment, the multiple depthwise separable convolutional computation units can be depthwise separable convolutional computation units on the diagonal of the computational unit array, or at least one column of depthwise separable convolutional computation units in the computational unit array; wherein, the at least one column of depthwise separable convolutional computation units can be located in the computational unit array at least one column closest to the feature data storage. The specific arrangement of the multiple depthwise separable convolutional computation units in the computational unit array can be selected according to the actual situation, and this embodiment does not impose specific limitations.
[0044] For example, multiple depthwise separable convolutional computation units (PE1) are arranged on the diagonal of the computational unit array, and non-depthwise separable convolutional computation units (PE2) are arranged in the remaining positions as follows: Figure 6As shown, however, because the depthwise separable convolutional computation unit adds some optimized depthwise separable convolutional computation logic, the area of the depthwise separable convolutional computation unit is larger than that of the non-depthwise separable convolutional computation unit. Setting multiple depthwise separable convolutional computation units on the diagonal of the computation unit array will result in irregular wiring layout and reduced area utilization, increasing the difficulty of subsequent wiring.
[0045] For example, a column of depthwise separable convolutional computational units (PE1) is arranged in the rightmost column of the computational unit array, and the remaining positions are arranged in a manner similar to that of non-depthwise separable convolutional computational units (PE2). Figure 7 As shown, this allows for a neat wiring layout, thereby reducing the difficulty of wiring layout.
[0046] In this embodiment, the depthwise separable convolution computation unit includes multiple multiply-accumulate units, multiple intermediate result accumulators, and a multiplexer, such as... Figure 8 As shown, the output of each multiply-accumulate unit is connected to the input of an intermediate result accumulator. The outputs of each multiply-accumulate unit and each intermediate result accumulator are also connected to the input of a multiplexer. Because the depthwise separable convolution unit contains an intermediate result accumulator, when the amount of data in the convolution window exceeds the number of multipliers in the multiply-accumulate unit, the convolution window data can be divided into multiple sets of input data and transmitted to the depthwise separable convolution unit in stages. The intermediate result accumulator is then used to accumulate the convolution results, thereby obtaining output feature data of different sizes. Therefore, the depthwise separable convolution unit can process convolution kernel data of different sizes.
[0047] It should be noted that each multiply-accumulator includes multiple multipliers and one adder. The adder is used to superimpose the results of the operations in multiple multipliers to obtain the result of the convolution operation.
[0048] It should be noted that if the depthwise separable convolutional computation unit only needs to perform one multiply-accumulate operation to obtain the final output feature data, then the multiply-accumulate unit does not need to output the intermediate result to the intermediate result accumulator. In this case, the multiplexer selects the operation result of the output multiply-accumulate unit. If the depthwise separable convolutional computation unit needs to perform multiple multiply-accumulate operations to obtain the output feature data, then the multiply-accumulate unit inputs the intermediate result generated by the convolutional computation into the intermediate result accumulator for accumulation. In this case, the multiplexer selects the accumulation result of the output intermediate result accumulator.
[0049] Optionally, the device further includes: a feature data memory; the feature data memory includes multi-channel input feature data; the feature data memory is connected to the input of multiple multiply-accumulate units and the output of a multiplexer in each depthwise separable convolution computation unit;
[0050] The feature data storage is used to input multiple convolution window data of each channel input feature data into multiple multiply-accumulators of a depth-separable convolution calculation unit;
[0051] The multiplexer is also used to select the output result of each multiply-accumulator to the feature data memory, or to select the output result of each intermediate result accumulator to the feature data memory.
[0052] In this embodiment, the feature data memory includes multi-channel input feature data. The feature data memory can input each channel's input feature data into a depthwise separable convolutional computation unit, as shown in the reference. Figure 9 The feature data memory transmits the input feature data of channel 0 to the depthwise separable convolution calculation unit 0 for depthwise separable convolution calculation, and returns the convolution calculation result to the feature data memory. Similarly, the feature data memory transmits the input feature data of channel 1 to the depthwise separable convolution calculation unit 1 for depthwise separable convolution calculation, and returns the convolution calculation result to the feature data memory. The feature data memory transmits the input feature data of channel 2 to the depthwise separable convolution calculation unit 2 for depthwise separable convolution calculation, and returns the convolution calculation result to the feature data memory. The feature data memory transmits the input feature data of channel 3 to the depthwise separable convolution calculation unit 3 for depthwise separable convolution calculation, and returns the convolution calculation result to the feature data memory. Furthermore, if the feature data memory also contains input feature data for channel 4, after the depthwise separable convolution calculation unit 0 processes the input feature data of channel 0, the feature data memory transmits the input feature data of channel 4 to the depthwise separable convolution calculation unit 4 for depthwise separable convolution calculation, and so on, until the depthwise separable convolution calculation of the input feature data of all channels is completed. It should be noted that the input feature data of channel 4 and the depthwise separable convolutional computation unit 4 are not in... Figure 9 As shown in the image.
[0053] Optionally, the apparatus further includes a data flattening unit located between the feature data memory and the depth-separable convolution calculation unit;
[0054] The data flattening unit is used to divide each convolution window data into at least one set of convolution window data according to the number of multipliers in the multiply-accumulator; and input the at least one set of convolution window data into the multiply-accumulator corresponding to each convolution window data in stages;
[0055] The multiply-accumulator corresponding to each convolutional window data is used to perform at least one multiply-accumulate operation on at least one set of convolutional window data.
[0056] The intermediate result accumulator corresponding to each convolution window data is used to accumulate the intermediate results generated by each multiplication and accumulation operation if each convolution window data is divided into multiple groups of convolution window data.
[0057] The multiplexer corresponding to each convolutional window data is used to select the operation result of the corresponding multiply-accumulate to the feature data memory if each convolutional window data is divided into a group of convolutional window data; and to select the accumulation result of the corresponding intermediate result accumulator to the feature data memory if each convolutional window data is divided into multiple groups of convolutional window data.
[0058] Taking a 5x5 convolution kernel as an example, the data in each convolution window needs to be flattened and fed into the depthwise separable convolutional computation unit, such as... Figure 10 As shown, each convolution window requires 5*5=25 multiplication operations. Since each multiply-accumulator in a depthwise separable convolutional computation unit contains 9 multipliers, the convolution window data needs to be divided into three segments: 9 data points, 9 data points, and 7 data points. Because a depthwise separable convolutional computation unit includes 16 multiply-accumulators, 16 sets of convolution window data are input into the depthwise separable convolutional computation unit each time, in three separate inputs to complete the depthwise separable convolution calculation. The first input is 16 sets of the first segment of convolution window data; the second input is 16 sets of the second segment of convolution window data; and the third input is 16 sets of the third segment of convolution window data. The results of these three convolutions are then accumulated in the intermediate result accumulator within the depthwise separable convolutional computation unit.
[0059] refer to Figure 11The process of mapping a 5*5 convolution window of data to a depthwise separable convolutional computation unit and performing depthwise separable convolution computation is as follows: A 5*5 convolution window contains 25 points. Since a multiply-accumulate unit contains 9 multipliers, the data of a single convolution window needs to be mapped three times. At time T0, the first 9 data points are fed into multiply-accumulate unit 0. The convolution result of these 9 data points and the weight data is P0. Then, P0 is fed into the ACC module. At time T1, the middle 9 data points are fed into multiply-accumulate unit 0. The convolution result of these 9 data points and the weight data is P1. Then, P1 is fed into the ACC module and added to the previous P0. The ACC module continues to temporarily store the result P0+P1. At time T2, the last 7 data points are fed into multiply-accumulate unit 0. The convolution result of these 7 data points and the weight data is P2. Therefore, at time T2, the result of the ACC module is P0+P1+P2. After these three clock cycles, the convolution result of the 5*5 convolution window data is calculated, and the accumulated result of ACC can be selected as the output through the multiplexer.
[0060] The above example illustrates the calculation process of a 5x5 convolution window. For a 3x3 depthwise separable convolution, since each convolution window only requires 9 multiplication operations, which is the same as the number of multipliers in a multiply-accumulate, the depthwise separable convolution operation of a convolution window can be completed in just one clock cycle. At this time, the result of the multiply-accumulate operation can be selected as the output by a multiplexer.
[0061] It is understood that this application proposes a depthwise separable convolution computation unit, which includes a multiply-accumulator, an intermediate result accumulator, and a multiplexer. This enables the accumulation operation of convolution results to be performed within the depthwise separable convolution computation unit, shortening the data transmission path and thus reducing data transmission power consumption.
[0062] Based on the above embodiments, this application also proposes a data processing method, applied in the above-mentioned data processing device 1, such as... Figure 12 As shown, the method may include:
[0063] S101. Determine multiple convolution window data from the input feature data of each channel in the feature data memory according to the number of multiply-accumulators of a depth-separable convolution computation unit.
[0064] S102. Input multiple convolution window data into a corresponding depthwise separable convolution calculation unit to perform depthwise separable convolution calculation to obtain output feature data; and output the output feature data to the feature data memory.
[0065] In this embodiment, after determining multiple convolutional window data from the feature data input from each channel in the feature data memory, each convolutional window data is divided into at least one group of convolutional window data according to the number of multipliers in each multiplier-accumulator. Then, using the multiplier-accumulator corresponding to each convolutional window data in a depthwise separable convolutional computation unit, at least one multiply-accumulate operation is performed on the at least one group of convolutional window data. Next, using the intermediate result accumulator corresponding to each convolutional window data in a depthwise separable convolutional computation unit, the intermediate results generated from each multiply-accumulate operation are accumulated. Using the multiplexer corresponding to each convolutional window data in a depthwise separable convolutional computation unit, if each convolutional window data is divided into multiple groups of convolutional window data, the operation result of the corresponding multiplier-accumulator is selected and output to the feature data memory; if each convolutional window data is divided into multiple groups of convolutional window data, the accumulation result of the corresponding intermediate result accumulator is selected and output to the feature data memory.
[0066] It should be noted that if the amount of data in each convolutional window is less than or equal to the number of multipliers in the multiply-accumulate unit, then each convolutional window data is divided into a group of convolutional window data; if the amount of data in each convolutional window data is greater than the number of multipliers in the multiply-accumulate unit, then each convolutional window data is divided into multiple groups of convolutional window data.
[0067] It should be noted that the method of dividing each convolution window data into multiple groups of convolution window data can be based on the number of multipliers. For example, for 9 multipliers, the 5*5 convolution window data can be divided into three groups of 9, 9, and 7. Alternatively, without exceeding the number of multipliers, the convolution window data can be evenly divided to obtain multiple groups of convolution window data. For example, for 9 multipliers, the 4*4 convolution window data can be divided into two groups of 8 and 8. The specific division method can be selected according to the actual situation, and this application embodiment does not impose specific limitations.
[0068] It is understood that this application proposes a depthwise separable convolution computation unit, which includes a multiply-accumulator, an intermediate result accumulator, and a multiplexer. This allows the accumulation operation of the convolution result to be performed inside the depthwise separable convolution computation unit without going through the accumulator in the computation unit array, thus shortening the data transmission path and reducing data transmission power consumption.
[0069] Based on the above embodiments, this application also proposes a chip 2, such as... Figure 13 As shown, the chip 2 includes the aforementioned data processing device 1.
[0070] Based on the above embodiments, this application also provides a computing device 3, such as... Figure 14As shown, the device 3 includes a processor 30, a memory 31, and a communication bus 32. The processor 30 can be at least one of the following: an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field-programmable gate array (FPGA), a CPU, a controller, a microcontroller, or a microprocessor. It is understood that for different devices, the electronic devices used to implement the above-mentioned processor functions can also be other types, and this embodiment does not specifically limit them.
[0071] The communication bus 32 is used to establish a communication connection between the processor 30 and the memory 31; when the processor 30 executes the running program stored in the memory 31, it implements the following data processing methods:
[0072] Based on the number of multiply-accumulators in a depthwise separable convolution computation unit, multiple convolution window data are determined from the input feature data of each channel in the feature data memory; the multiple convolution window data are input into a corresponding depthwise separable convolution computation unit for depthwise separable convolution computation to obtain output feature data; and the output feature data is output to the feature data memory.
[0073] Furthermore, the processor 30 is also configured to divide each convolutional window data into at least one group of convolutional window data according to the number of multipliers in each multiplier-accumulator; perform at least one multiply-accumulate operation on the at least one group of convolutional window data using the multiplier-accumulator corresponding to each convolutional window data in a depthwise separable convolutional computation unit; accumulate the intermediate results generated by each multiply-accumulate operation using the intermediate result accumulator corresponding to each convolutional window data in a depthwise separable convolutional computation unit; and select the operation result of the corresponding multiplier-accumulator to be output to the feature data memory if each convolutional window data is divided into one group of convolutional window data using the multiplexer corresponding to each convolutional window data in a depthwise separable convolutional computation unit; and select the accumulation result of the corresponding intermediate result accumulator to be output to the feature data memory if each convolutional window data is divided into multiple groups of convolutional window data.
[0074] This application provides a storage medium storing a computer program thereon. The computer-readable storage medium stores one or more programs, which can be executed by one or more processors and applied in a computing device. The computer program implements the data processing method described above.
[0075] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0076] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause an image display device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this disclosure.
[0077] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application.
Claims
1. A data processing apparatus, characterized in that, The device deploys a computing unit array, which includes multiple depthwise separable convolution computing units. Each depthwise separable convolution computing unit includes: multiple multiply-accumulate units, multiple intermediate result accumulators, and a multiplexer. The output of each multiply-accumulate unit is connected to the input of an intermediate result accumulator, and the input of the multiplexer is connected to the outputs of both the multiple multiply-accumulate units and the multiple intermediate result accumulators. The multiple multiply-accumulate operators are used to perform multiply-accumulate operations on multiple convolutional window data in one channel of input feature data. The plurality of intermediate result accumulators are used to perform accumulation operations on the intermediate results generated by the corresponding multiply-accumulators respectively. The multiplexer is used to select the output result of each multiply-accumulator or to select the output result of each intermediate result accumulator. The device further includes: a feature data memory and a data flattening unit located between the feature data memory and the depth-separable convolution calculation unit; The data flattening unit is used to divide each convolution window data into at least one set of convolution window data according to the number of multipliers in the multiply-accumulator. The multiplexer corresponding to each convolutional window data is used to select the operation result of the corresponding multiply-accumulate to the feature data memory if each convolutional window data is divided into a group of convolutional window data; and to select the accumulation result of the corresponding intermediate result accumulator to the feature data memory if each convolutional window data is divided into multiple groups of convolutional window data.
2. The apparatus according to claim 1, characterized in that, The plurality of depthwise separable convolutional computation units are at least one column of depthwise separable convolutional computation units in the computational unit array.
3. The apparatus according to claim 2, characterized in that, The at least one column of depth-separable convolutional computation units is located in the computation unit array, in the column closest to the feature data storage.
4. The apparatus according to claim 1, characterized in that, The device further includes: a feature data memory; the feature data memory includes multi-channel input feature data; the feature data memory is connected to the input of multiple multiply-accumulates and the output of a multiplexer in each depthwise separable convolutional computation unit; The feature data storage is used to input multiple convolution window data of each channel input feature data into multiple multiply-accumulators of a depth-separable convolution calculation unit; The multiplexer is also used to select the output result of each multiply-accumulator to the feature data memory, or to select the output result of each intermediate result accumulator to the feature data memory.
5. The apparatus according to claim 4, characterized in that, The device further includes a data flattening unit located between the feature data storage and the depth-separable convolution calculation unit; The data flattening unit is used to divide each convolution window data into at least one set of convolution window data according to the number of multipliers in the multiply-accumulator; and input the at least one set of convolution window data into the multiply-accumulator corresponding to each convolution window data in stages; The multiply-accumulator corresponding to each convolutional window data is used to perform at least one multiply-accumulate operation on at least one set of convolutional window data. The intermediate result accumulator corresponding to each convolution window data is used to accumulate the intermediate results generated by each multiplication and accumulation operation if each convolution window data is divided into multiple groups of convolution window data. The multiplexer corresponding to each convolutional window data is used to select the operation result of the corresponding multiply-accumulate to the feature data memory if each convolutional window data is divided into a group of convolutional window data; and to select the accumulation result of the corresponding intermediate result accumulator to the feature data memory if each convolutional window data is divided into multiple groups of convolutional window data.
6. The apparatus according to claim 1, characterized in that, The computing unit array further includes: a non-depth-separable convolution computing unit; the device is also configured to, when performing depth-separable convolution operations, activate the plurality of depth-separable convolution computing units and disconnect the non-depth-separable convolution computing units.
7. A data processing method, characterized in that, Applied in the data processing apparatus according to any one of claims 1-5, the method comprises: Based on the number of multiply-accumulators in a depth-separable convolutional computation unit, multiple convolutional window data are determined from the input feature data of each channel in the feature data memory. The multiple convolutional window data are input into a corresponding depthwise separable convolutional computation unit for depthwise separable convolution computation to obtain output feature data; and the output feature data is output to the feature data memory. The step of inputting the multiple convolutional window data into a corresponding depthwise separable convolutional computation unit for depthwise separable convolution computation to obtain output feature data, and then outputting the output feature data to the feature data memory, includes: Based on the number of multipliers in the multiply-accumulator, each convolutional window data is divided into at least one set of convolutional window data; Using a multiplexer corresponding to each convolution window data in a depth-separable convolution computation unit, if each convolution window data is divided into a group of convolution window data, the operation result of the corresponding multiply-accumulate is selected and sent to the feature data memory; if each convolution window data is divided into multiple groups of convolution window data, the accumulation result of the corresponding intermediate result accumulator is selected and sent to the feature data memory.
8. The method according to claim 7, characterized in that, The multiple convolutional window data are input into a corresponding depthwise separable convolutional computation unit for depthwise separable convolutional computation to obtain output feature data; And outputting the output feature data to the feature data memory, including: Divide each convolution window data into at least one set of convolution window data according to the number of multipliers in each multiply-accumulator; Using the multiply-accumulator corresponding to each convolution window data in a depth-separable convolution computation unit, perform at least one multiply-accumulate operation on the at least one set of convolution window data; By using an intermediate result accumulator corresponding to the data of each convolution window in a depthwise separable convolutional computation unit, the intermediate results generated by each multiplication-accumulation operation are accumulated. Using a multiplexer corresponding to each convolution window data in a depth-separable convolution computation unit, if each convolution window data is divided into a group of convolution window data, the operation result of the corresponding multiply-accumulate is selected and sent to the feature data memory; if each convolution window data is divided into multiple groups of convolution window data, the accumulation result of the corresponding intermediate result accumulator is selected and sent to the feature data memory.
9. A chip, characterized in that, The chip includes the data processing device as described in any one of claims 1-6.
10. A computing device, characterized in that, The device includes: a processor, a memory, and a communication bus; the communication bus is used to realize a communication connection between the processor and the memory; when the processor executes the running program stored in the memory, it implements the method as described in any one of claims 7-8.
11. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method as described in any one of claims 7-8.
Citation Information
Patent Citations
Convolution calculation module, neural network processor, chip and electronic equipment
CN111222090A
Convolution operation circuit and operation method thereof
CN113869498A