Depth separation convolution acceleration method based on depth-first processing and electronic equipment

By adopting a deep separation convolution acceleration method based on depth priority processing in neural network computing, using flexible cache architecture and cross-layer pipeline scheduling method, the problems of large area, high energy consumption and data delay in traditional neural network computing are solved, and efficient deep convolution calculation and flexible cache access are achieved.

CN120146108APending Publication Date: 2025-06-13SHENZHEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510213607.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Traditional neural network computing has the problem of layer by layer processing leading to large area and high energy consumption. The pulsating array cannot quickly utilize intermediate features during depth calculations, and the traditional line cache architecture cannot adapt to random access requirements, resulting in data delay.

Method used

The deep separation convolution acceleration method based on depth priority processing is adopted. Through a flexible cache architecture and cross-layer pipeline scheduling method, efficient calculation of deep convolution and point-by-point convolution is realized, and random access to cross-channel and cross-space coordinates is supported.

Benefits of technology

The calculation efficiency of deep separation convolution is optimized, the rapid utilization of intermediate feature data is realized, the transmission bandwidth and storage requirements of intermediate feature maps are reduced, and flexible cache access and multiple convolution modes are supported.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146108A_ABST
    Figure CN120146108A_ABST
Patent Text Reader

Abstract

The invention relates to a depth-first processing-based depth separation convolution acceleration method and electronic equipment, and the method comprises the steps: employing a flexible cache architecture in a process of executing depth separation convolution calculation; comprising the following steps: processing an input feature map by adopting a depth-first calculation mode; performing deep convolution calculation in a deep convolution stage based on a depth-first calculation mode to obtain a calculation result of the deep convolution stage; and in the point-by-point convolution stage, multiplexing the calculation result of the deep convolution stage in each period to obtain an intermediate result of the point-by-point convolution stage, and respectively accumulating the intermediate results of the corresponding output channels to obtain an output feature map. According to the method, the calculation efficiency of deep separation convolution can be optimized, a depth-first calculation mode is adopted, rapid utilization of intermediate feature data can be realized, the transmission bandwidth and the storage requirement of an intermediate feature graph are reduced, and meanwhile, a flexible cache architecture is adopted, so that a flexible access row cache structure and multiple convolution modes can be supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of neural network computing, and more specifically, to a depthwise separable convolution acceleration method and an electronic device based on depth-first processing. Background Art

[0002] Traditional neural network processing is carried out layer by layer, that is, all calculations of one layer are completed to generate a complete output feature map (FM), and then this feature map is used for the processing of the next layer. However, since the size of the feature map is generally large, it needs to be stored in a large-capacity off-chip memory or a large on-chip cache. Therefore, it will result in a large area requirement, and at the same time, frequent transmission is required between the off-chip memory and the on-chip memory, resulting in increased energy consumption. Traditionally, the implementation of a neural network based on a systolic array realizes efficient convolution acceleration by reusing the input feature map data horizontally and multiplying it with the weights of different output channels. However, for depthwise separable convolution, due to the one-to-one correspondence between its input channels and output channels, its utilization rate in the acceleration structure of the systolic array is generally low. To improve the utilization rate of depthwise convolution in the systolic array, the commonly used method is to perform oblique propagation of data and improve the utilization rate of computing units through the reuse relationship of feature data in the row direction. However, this method will also bring a high requirement for transmission bandwidth, and the control of the data stream will be more complex. Another solution is to use the result of depthwise convolution for the calculation of pointwise convolution. However, due to the need to wait for the completion of depthwise convolution, there is also a large amount of idle time in the calculation of the pointwise convolution part, resulting in low utilization rate. Further, in the traditional scheme, in order to align the convolution window data, the traditional approach is based on a row buffer architecture of FIFO or shift register. Although this architecture is simple, it only supports sequential sliding window access and cannot meet the random access requirements across channels and spatial coordinates in depth-first calculation. When the calculation needs to jump between different channels to read data, the traditional row buffer needs to load / unload data multiple times, increasing the latency.

[0003] Based on this, the traditional neural network calculation has the following problems: layer-by-layer processing leads to a large area and high energy consumption. When using the systolic array architecture for depth calculation, it cannot quickly utilize intermediate features, has high transmission bandwidth and storage requirements for intermediate feature maps, and uses a sequential access row buffer architecture, which cannot meet the random access needs and also has data latency problems. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a depthwise separable convolution acceleration method and an electronic device based on depth-first processing in view of the problems existing in the prior art.

[0005] The technical solution adopted by the present invention to solve its technical problems is: to construct a depth-separable convolution acceleration method based on depth-first processing, and adopt a flexible cache architecture during the execution of depth-separable convolution calculation; the depth-separable convolution acceleration method based on depth-first processing includes the following steps:

[0006] Process the input feature map using a depth-first calculation mode;

[0007] Based on the depth-first calculation mode, perform depth convolution calculation in the depth convolution stage to obtain the calculation result of the depth convolution stage;

[0008] In the pointwise convolution stage, reuse the intermediate result of the depth convolution stage in each cycle to obtain the intermediate results of different output channels in the pointwise convolution stage, and accumulate the calculation results corresponding to the same output channel to obtain the output feature maps of different output channels.

[0009] In the depth-separable convolution acceleration method based on depth-first processing of the present invention, the process of processing the input feature map using a depth-first calculation mode to obtain an output feature map includes:

[0010] Determine the basic processing unit;

[0011] Process the input feature map according to the basic processing unit using a cross-layer pipelining scheduling method and a channel-first processing method.

[0012] In the depth-separable convolution acceleration method based on depth-first processing of the present invention, the determination of the basic processing unit includes:

[0013] Determine the convolution window;

[0014] Determine the number of rows required for the convolution window according to the convolution window;

[0015] Determine the basic processing unit according to the number of rows required for the convolution window.

[0016] In the depth-separable convolution acceleration method based on depth-first processing of the present invention, the process of processing the input feature map according to the basic processing unit using a cross-layer pipelining scheduling method and a channel-first processing method includes:

[0017] After the calculation of the feature row in the current layer is completed, determine whether the cached result output by the current layer meets the calculation requirements of the next layer;

[0018] If it is satisfied, process the data of the next layer;

[0019] During the execution of the above calculation process, the input feature map is processed using the channel-first method for the processing of in-row data.

[0020] In the depthwise separable convolution acceleration method based on depth-first processing according to the present invention, performing depthwise convolution calculation in the depthwise convolution stage to obtain the calculation result of the depthwise convolution stage includes:

[0021] Sliding the input feature map window according to the row caching rule, and independently processing each input channel; calculating the depthwise convolution based on each column of processing units to generate a single-channel calculation result; the single-channel calculation result is the calculation result of the depthwise convolution stage; the calculation result of the depthwise convolution stage includes: the single-channel calculation result generated by the first column of processing units.

[0022] In the depthwise separable convolution acceleration method based on depth-first processing according to the present invention, in the pointwise convolution stage, reusing the calculation result of the depthwise convolution stage in each cycle to obtain the intermediate results of different output channels in the pointwise convolution stage, and correspondingly accumulating the intermediate results of the same output channel to obtain the output feature maps of different output channels includes:

[0023] In the pointwise convolution stage, after outputting the single-channel calculation result generated by the first column of processing units, starting the pointwise convolution calculation, and reusing the calculation result of the depthwise convolution stage in each cycle;

[0024] Multiplying by the weights corresponding to different output channels to obtain the intermediate results of different output channels; accumulating the intermediate results of the same output channel to obtain the output feature maps of different output channels.

[0025] In the depthwise separable convolution acceleration method based on depth-first processing according to the present invention, the method further includes:

[0026] In the pointwise convolution stage, caching the calculation results of each output channel using a shift register bank, and each output channel corresponds to an independent shift register, and the intermediate results are accumulated in the shift register.

[0027] In the depthwise separable convolution acceleration method based on depth-first processing according to the present invention, the flexible cache architecture includes:

[0028] Setting K independent static random access memories in each row and writing data in a polling manner; where K is the convolution kernel size;

[0029] The flexible cache architecture generates addresses using the stride adaptive address method.

[0030] In the depthwise separable convolution acceleration method based on depth-first processing according to the present invention, the stride adaptive address method includes:

[0031] Set the shift register stride and offset parameters;

[0032] Encode the segmented address;

[0033] Dynamically adjust the stride of the shift register by modifying the configuration of the shift register; perform a jump operation on the channel by updating the channel address.

[0034] The present invention also provides an electronic device, including a memory and a processor. A computer program is stored in the memory. The processor executes the steps of the depthwise separable convolution acceleration method based on depth-first processing as described above by calling the computer program stored in the memory.

[0035] Implementing the depthwise separable convolution acceleration method and the electronic device based on depth-first processing of the present invention has the following beneficial effects: During the execution of depthwise separable convolution calculation, a flexible cache architecture is adopted; including: processing the input feature map in a depth-first calculation mode, outputting a convolution window matching the depth-first calculation mode, and calculating the output feature map through a processing unit; based on the depth-first calculation mode, performing depthwise convolution calculation in the depthwise convolution stage to obtain the calculation result of the depthwise convolution stage; in the pointwise convolution stage, reusing the calculation result of the depthwise convolution stage in each cycle to obtain the intermediate result of the pointwise convolution stage of different output channels, and accumulating the results of the same output channel to obtain the output feature map. Through the present invention, the calculation efficiency of depthwise separable convolution can be optimized, and by adopting the depth-first calculation mode, the rapid utilization of intermediate feature data can be achieved, reducing the transmission bandwidth and storage requirements of intermediate feature maps. At the same time, by adopting a flexible cache architecture, a row cache structure supporting flexible access and multiple convolution modes can be supported. Description of the Drawings

[0036] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:

[0037] Figure 1 is the architecture diagram of the existing layer-by-layer processing;

[0038] Figure 2 is the flowchart of the depthwise separable convolution acceleration method based on depth-first processing of the present invention;

[0039] Figure 3 is the architecture diagram of the depth-first calculation mode adopted by the present invention;

[0040] Figure 4 is the architecture diagram of the processing unit of the present invention;

[0041] Figure 5 is the overall architecture diagram provided by the present invention. Detailed Embodiments

[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0043] Refer to Figure 2 , Figure 2 which is a flowchart of the depthwise separable convolution acceleration method based on depth-first processing provided by the present invention. Among them, in the process of performing depthwise separable convolution calculation, the depthwise separable convolution acceleration method based on depth-first processing adopts a flexible cache architecture.

[0044] Specifically, as Figure 2 shown, the depthwise separable convolution acceleration method based on depth-first processing includes the following steps:

[0045] Step S201: Process the input feature map using a depth-first calculation mode.

[0046] Specifically, after processing the input feature map using a depth-first calculation mode, a convolution window matching the depth-first calculation mode can be output, and after being calculated by the processing unit, an output feature map can be obtained.

[0047] Optionally, in the embodiments of the present invention, processing the input feature map using a depth-first calculation mode includes: determining a basic processing unit; processing the input feature map using a cross-layer pipeline scheduling method and a channel-first processing method according to the basic processing unit. Among them, determining the basic processing unit includes: determining a convolution window, determining the number of rows required for the convolution window according to the convolution window; determining the basic processing unit according to the number of rows required for the convolution window. Optionally, the convolution window is mainly related to the convolution mode, so the convolution window can be determined according to the convolution mode. Among them, the convolution mode can be determined according to the neural network model used.

[0048] Optionally, in the embodiments of the present invention, processing the input feature map using a cross-layer pipeline scheduling method and a channel-first processing method according to the basic processing unit includes: after the feature rows of the current layer (layer L i ayer) are calculated, determining whether the cached result output by the current layer meets the calculation requirements of the next layer (layer L i+1 ayer); if it meets, processing the data of the next layer; in the process of performing the above calculations, the data within the row is processed using a channel-first method to process the input feature map.

[0049] Specifically, as Figure 3As shown in the figure, the present invention adopts a row-based grouping processing method, that is, the number of rows required by the convolution window of the input feature map (set as N rows) is used as the basic processing unit, rather than processing the entire feature map layer by layer. As Figure 1 shown, in the traditional scheme, a layer-by-layer sequential processing method is adopted, that is, after processing the input feature map layer, then processing the intermediate feature map layer, and then after all the intermediate feature map layers are processed, finally processing the data of the output feature map layer. However, the present invention uses N rows required by the convolution window of the input feature map as the basic processing unit, rather than the layer-by-layer processing method, so that the feature map data required to be stored across convolutional layers is also compressed to N rows required by the convolution window, which can significantly improve the computational efficiency of depthwise separable convolution.

[0050] Further, in the embodiment of the present invention, in the depth-first computing mode, a cross-layer pipeline scheduling method is simultaneously adopted for data processing, that is, when the i layer's feature row calculation is completed, if the cached feature results can meet the calculation requirements of the i+1 layer, the data of the i+1 layer is preferentially processed to avoid waiting for the generation of the entire layer of data. Specifically, as Figure 3 shown, when the feature row calculation of the intermediate feature map layer is completed, if the cached feature results can meet the calculation requirements of the output feature map layer, the data of the output feature map layer is preferentially processed, without waiting for all the data of the intermediate feature map layer to be calculated before processing, that is, waiting for the generation of the entire layer of data can be avoided. Among them, the calculation requirement here corresponds to the cached data of the next layer and the convolution mode of the next layer. For example, if the convolution mode of the next layer is 3*3 convolution, then if the cached result of the feature row calculation in the intermediate feature map meets the 3*3 convolution, the calculation requirement is met.

[0051] Further, in the embodiment of the present invention, when processing the feature row data, the in-row data is mainly processed in channel-first order. After the data of all data channels corresponding to a row is calculated, the row data of the previous layer can be discarded, thereby effectively reducing the cache pressure. Specifically, the traditional calculation method is that in each layer, the calculation of the entire feature map is to calculate the in-row data first and then execute in channel order. In the present invention, as Figure 3 shown, assuming that the data of the first row of the output feature map depends on the data of the first, second, and third rows of the intermediate feature map, then after all the output channels of the first row of the output feature map are calculated, the data of the first row in the intermediate feature map can be discarded. Similarly, the same method is used for processing between the intermediate feature map and the input feature map.

[0052] Step S202: Based on the depth-first calculation mode, perform depth convolution calculation in the depth convolution stage to obtain the calculation result of the depth convolution stage. Optionally, in the embodiments of the present invention, performing depth convolution calculation in the depth convolution stage to obtain the calculation result of the depth convolution stage includes: sliding the input feature map window according to the row buffer rule, and independently processing each input channel; calculating depth convolution based on each column of processing units to generate a single-channel calculation result; the single-channel calculation result is the calculation result of the depth convolution stage; the calculation result of the depth convolution stage includes: the single-channel calculation result generated by the first column of processing units.

[0053] Specifically, in the depth convolution stage (i.e., the DW stage): the input feature map window slides according to the row buffer rule, and each input channel is independently processed. For example, assume that the input feature map has 9 rows, and each row has 9 data from 1 to 9. Then, first process the 9 data of input channel 0. After processing the 9 data of input channel 0, then process the 9 data of channel 1, and slide in sequence. At the same time, calculate the depth convolution through the first column of processing units (PE) to generate a single-channel calculation result (which can be defined as the calculation result in the first single channel). Among them, the depth convolution of each window is K 2 cycles (K is the convolution kernel size). When the depth convolution calculation of the first column of processing units is completed, continue to calculate the depth convolution of other columns of processing units in the DW stage and output the corresponding single-channel calculation results.

[0054] Step S203: In the pointwise convolution stage, reuse the calculation result of the depth convolution stage in each cycle to obtain the intermediate results of different output channels in the pointwise convolution stage, and accumulate the intermediate results corresponding to the same output channel to obtain the output feature maps of different output channels.

[0055] It should be noted that in the embodiments of the present invention, there is no strict order requirement for Step S201, Step S202, and Step S203, that is, Step S201, Step S202, and Step S203 do not require sequential execution, and specifically, it depends on the actual neural network calculation.

[0056] Optionally, in the embodiments of the present invention, in the pointwise convolution stage, the calculation results of the depth convolution stage are reused within each cycle to obtain intermediate results of different output channels in the pointwise convolution stage, and the intermediate results of the same output channel are correspondingly accumulated to obtain the output feature maps of different output channels, including: in the pointwise convolution stage, after generating the single-channel calculation result of the first column of processing units, start the pointwise convolution calculation, and reuse the calculation results of the depth convolution stage in each cycle, multiply with the corresponding weights of different output channels to obtain the intermediate results of different output channels; accumulate the intermediate results of the same output channel to obtain the output feature maps of different output channels. The accumulation of the intermediate results of the same output channel here means that the intermediate results of the same output channel need to be accumulated.

[0057] Further, in the embodiments of the present invention, in the pointwise convolution stage, the calculation results of each output channel are cached using shift registers, and each output channel corresponds to an independent shift register, where the intermediate results are accumulated in the shift register.

[0058] Specifically, in the pointwise convolution stage (i.e., the PW stage): within the K 2 cycles of DW calculation, the calculation results output by DW are reused in each cycle. Specifically, first calculate and accumulate the partial products of different output channels. At the same time, a parallel design is adopted, that is, one output channel is processed in each cycle, then K 2 cycles complete the calculation of K 2 channels; for the allocation of registers, the present invention caches the calculation results through a group of shift registers, where each output channel corresponds to an independent shift register, thereby avoiding data competition. As Figure 4 shown is the structure of the array processing unit (i.e., the PE unit). It can be seen from Figure 4 that in each processing unit, the in_a data and the in_b data first perform a multiplication operation, then accumulate with the value in the shift register, and then store the accumulated result in the shift register. Figure 4 In, Reg is the register, and Shift Reg is the shift register.

[0059] Further, in the embodiments of the present invention, since each output channel uses an independent shift register to store the accumulated result, therefore, the accumulated results of all output channels can be directly accumulated to the output feature map without an additional accumulation stage.

[0060] It should be noted that in the embodiments of the present invention, the calculations in the DW stage and the PW stage adopt a mode of hybrid calculation of depthwise separable convolution. That is, when the first column of processing units in the DW stage performs depthwise convolution and outputs the first single-channel calculation result, all output channels in the PW stage start pointwise convolution. That is to say, the startup waiting time of pointwise convolution in the PW stage is the depthwise convolution calculation time of the first column of processing units in the DW stage. When the first column of processing units in the DW stage performs depthwise convolution and outputs the first single-channel calculation result, all output channels in the PW stage start pointwise convolution, and all channels calculate and accumulate partial products by multiplexing this first single-channel intermediate result, and all channels are designed in a parallel manner. At the same time, during the process of performing pointwise convolution calculation in the PW stage, other columns of processing units in the DW stage also perform depthwise convolution calculation synchronously, and the single-depth convolution results output by each column of processing units will be multiplexed by the PW stage. That is, after the PW stage multiplexes the first single-channel calculation result, it continues to multiplex the second single-channel calculation result output by the DW stage, and after generating the first single-channel intermediate result, the DW stage and the PW stage perform calculations simultaneously. That is, after the PW stage waits for K 2 cycles, the PW stage and the DW stage calculate simultaneously.

[0061] The present invention adopts a depth-first pipeline scheduling method. The PW stage parallel processes K 2 output channels within K 2 cycles of the DW stage. When the number of output channels of the PW layer is greater than or equal to K 2 *(Col pe −1), the utilization rate is increased to 100%. (Col_pe is the number of columns of the computing array). The specific algorithm is as follows:

[0062]

[0063] In the embodiments of the present invention, the flexible cache architecture is a flexible row cache architecture, which may include: setting K independent static random access memories in each row (for example, for a 3*3 convolution, 3 SRAMs (static random access memories) are allocated), and writing data in a polling fixed-in manner to ensure conflict-free storage; where K is the convolution kernel size. The flexible cache architecture generates addresses by using a stride adaptive address method. Among them, the flexible row cache architecture outputs convolution windows corresponding to multiple different input channels and sends them to the array for calculation.

[0064] Optionally, in the embodiments of the present invention, the stride adaptive address method includes: setting a shift register stride and an offset parameter; encoding the segmented address; dynamically adjusting the stride of the shift register by modifying the configuration of the shift register; and performing a channel jump operation by updating the channel address.

[0065] Specifically, the row buffer architecture of the present invention controls the address offset through shift sequence, and combines segmented address encoding (Row_Addr, Cin_Addr, Wid_Addr) to achieve flexible access with multiple strides and multiple channels. The specific process is as follows:

[0066] 1) Shift register configuration and offset generation:

[0067] Stride = 1: The 3-bit shift register is initialized to 001 and cyclically shifted left by 1 bit per cycle.

[0068] Stride = 2: The 3-bit shift register is initialized to 011 and cyclically shifted left by 2 bits per cycle.

[0069] 2) Segmented address encoding:

[0070] Address segmentation definition:

[0071] Row Addr : Controls the vertical access (row address) of the feature map (here, the feature map refers to the feature map of the current processing layer).

[0072] Cin Addr : Controls the jump of the input channels (channel address).

[0073] Wid Addr : Controls the horizontal access (width address) of the feature map.

[0074] Address calculation formula:

[0075]

[0076] Parameter description:

[0077] S c : Row address segment width (can make L c be the number of bits of the row address).

[0078] S w : Channel address segment width (can make L w be the number of bits of the row address).

[0079] 3) Flexible dynamic configuration and channel jump:

[0080] Stride adjustment: By modifying the width and initial value of the shift register, support for windowing with any stride is achieved.

[0081] Channel jump: By updating the Cin_Addr value, data of different input channels can be directly accessed.

[0082] In the embodiments of the present invention, the input channels are defined relative to the current layer, and the output channels are defined relative to the next layer. That is, if the data being processed currently is the data of the current layer, it is an input channel; if the data being processed currently is the data output to the next layer, it is an output channel.

[0083] Specifically, the flexible row buffer architecture (IFM_Buffer) can output the data of the convolution window corresponding to multiple different input channels and send the data to the PE array for calculation. Figure 5 is the overall architecture diagram. From Figure 5 it can be seen that the overall architecture mainly includes six parts: the external interface part (External Interface), which mainly communicates with the outside based on the bus protocol. CFG_Reg_Group is the configuration register group, storing a series of convolution parameters including input and output channels, convolution stride, whether to pad, etc. IFM Buffer stores the feature map data, WGT_buffer stores the weight parameters, Global Arbiter is the arbiter, which is used to judge whether the next layer meets the calculation requirements. If it meets, it controls the IFM_Buffer and WGT_buffer to output the corresponding feature map and weight of the next layer and input them into the PE_Array for calculation. At the same time, the result calculated by the PE array will also be output to the position in the IFM_buffer for storing the data of the next layer after being judged by the arbiter. PE_Array is the calculation array, which is used to perform calculation operations. Figure 4 corresponding to Figure 5 a calculation unit (i.e., a small square box) in the

[0084] The present invention can be applied to the field of neural network computing, such as image classification, target recognition, FPGA technology, deep learning, etc., and can also be applied to industrial automation production, autonomous driving, etc.

[0085] The following effects can be achieved through the present invention:

[0086] 1. The computing efficiency is improved.

[0087] Traditional architecture: The effective utilization rate is approximately 85%.

[0088] The present invention: The utilization rate of the PW unit is increased to 100% (when the number of output channels of the PW layer is greater than K 2 *(Col_pe - 1)).

[0089] 2. Bandwidth optimization (taking a 4*4*9 systolic array as an example).

[0090] Traditional architecture: The input bandwidth requirement is 12 rows (9 rows of input + 3 rows of intermediate data).

[0091] The present invention: Through depth - first calculation and row cache reuse, only 9 rows of input bandwidth are required.

[0092] 3. Flexibility and compatibility.

[0093] It supports asynchronous stride sliding windows and flexible access across channels and spatial coordinates, and is adapted to models such as MobileNet and EfficientNet.

[0094] In addition, an electronic device of the present invention includes a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program to implement the depth - separable convolution acceleration method based on depth - first processing as described in any one of the above. Specifically, according to an embodiment of the present invention, the process described with reference to the flowchart above can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer - readable medium, and the computer program contains program codes for executing the method shown in the flowchart. In such an embodiment, when the computer program is downloaded and installed by the electronic device and executed, it executes the above - defined functions in the method of the embodiment of the present invention. The electronic device in the present invention can be a terminal such as a notebook, a desktop computer, a tablet computer, a smart phone, etc., or a server.

[0095] In addition, a storage medium of the present invention stores a computer program, and when the computer program is executed by a processor, it implements the depth-separable convolution acceleration method based on depth-first processing in any one of the above. Specifically, it should be noted that the above storage medium of the present invention may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, and the computer-readable signal medium may send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0096] The above computer-readable medium may be included in the above electronic device; or it may exist separately and not be assembled into the electronic device.

[0097] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference may be made to each other. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and reference may be made to the description of the method part for the relevant parts.

[0098] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0099] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well known in the technical field.

[0100] The above embodiments are only for illustrating the technical concept and features of the present invention, and their purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly, and cannot limit the protection scope of the present invention. All equivalent changes and modifications made to the scope of the claims of the present invention shall fall within the scope covered by the claims of the present invention.

Claims

1. A depth separation convolution acceleration method based on depth-first processing, characterized in that: In the process of performing depth separation convolution calculation, a flexible cache architecture is adopted; the depth separation convolution acceleration method based on depth priority processing includes the following steps: The input feature map is processed using a depth-first calculation mode; Based on the depth-first calculation mode, a depth convolution calculation is performed in the depth convolution stage to obtain a calculation result of the depth convolution stage; In the point-by-point convolution stage, the calculation results of the depth convolution stage are reused in each cycle to obtain intermediate results of different output channels of the point-by-point convolution stage, and the intermediate results of the same output channel are accumulated accordingly to obtain output feature maps of the different output channels.

2. The depth separation convolution acceleration method based on depth priority processing according to claim 1, characterized in that: The processing of the input feature map using the depth-first calculation mode includes: Identify basic processing units; The input feature map is processed according to the basic processing unit using a cross-layer pipeline scheduling method and a channel priority processing method.

3. The depth separation convolution acceleration method based on depth priority processing according to claim 2 is characterized in that: The determining basic processing unit comprises: Determine the convolution window; Determining the number of rows required for the convolution window according to the convolution window; The basic processing unit is determined according to the number of rows required by the convolution window.

4. The depth separation convolution acceleration method based on depth priority processing according to claim 2 is characterized in that: The processing of the input feature map by using a cross-layer pipeline scheduling method and a channel priority processing method according to the basic processing unit includes: After the feature row calculation of the current layer is completed, determine whether the cached result output by the current layer meets the calculation requirements of the next layer; If satisfied, the data of the next layer is processed; In the process of executing the above calculation, the input feature map is processed in the channel priority manner for processing the intra-row data.

5. The depth separation convolution acceleration method based on depth priority processing according to claim 1, characterized in that: The step of performing the depth convolution calculation in the depth convolution stage to obtain the calculation result of the depth convolution stage includes: The input feature map window is slid according to the row cache rule, and each input channel is processed independently; Based on each column of processing units, a deep convolution is calculated to generate a single-channel calculation result; the single-channel calculation result is the calculation result of the deep convolution stage; the calculation result of the deep convolution stage includes: the single-channel calculation result generated by the first column of processing units.

6. The depth separation convolution acceleration method based on depth priority processing according to claim 5 is characterized in that: In the point-by-point convolution stage, the calculation results of the depth convolution stage are reused in each cycle to obtain the intermediate results of different output channels in the point-by-point convolution stage, and the intermediate results of the same output channel are accumulated accordingly to obtain the output feature maps of the different output channels, including: In the point-by-point convolution stage, after outputting the single-channel calculation result generated by the first column processing unit, starting the point-by-point convolution calculation, and reusing the calculation result of the depth convolution stage in each cycle; Multiplying the weights corresponding to different output channels to obtain intermediate results of the different output channels; The intermediate results of the same output channel are accumulated to obtain the output feature maps of the different output channels.

7. The depth separation convolution acceleration method based on depth priority processing according to claim 6 is characterized in that: The method further comprises: In the point-by-point convolution stage, the calculation result of each output channel is cached using a shift register, and each output channel corresponds to an independent shift register, and the intermediate results are accumulated in the shift register.

8. The depth separation convolution acceleration method based on depth priority processing according to claim 1, characterized in that: The flexible cache architecture includes: Each row is equipped with K independent static random access memories, and data is written in a polling manner; where K is the convolution kernel size; The flexible cache architecture adopts a stride adaptive address method to generate addresses.

9. The depth separation convolution acceleration method based on depth priority processing according to claim 8, characterized in that: The method for adaptive stride address includes: Set the shift register stride and offset parameters; Encode the segment address; dynamically adjusting the stride of the shift register by modifying the configuration of the shift register; The channel jump operation is performed by updating the channel address.

10. An electronic device, characterized in that: It includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the steps of the depth separation convolution acceleration method based on depth-first processing as described in any one of claims 1 to 9 by calling the computer program stored in the memory.