Cache scheduling method for bottleneck residual block with attention mechanism

By adopting a cache scheduling method of bottleneck residual blocks with attention mechanism on microcontroller equipment, the memory leak problem is solved and the memory usage peak is significantly reduced.

CN119988030APending Publication Date: 2025-05-13UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510158882.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing neural network blocks containing bottleneck residual blocks with attention mechanisms are easily caused by memory leakage when deploying on devices of the microcontroller level.

Method used

By providing a caching scheduling method for bottleneck residual blocks with attention mechanism, it includes step-by-step reading and processing of feature map data, reasonably applying and freeing cache space, and performing multi-layer cascade operations to reduce the cache demand of the intermediate layer.

Benefits of technology

The decoupling of the SE block operation process and the entire bottleneck residual block operation process is realized, making the multi-layer cascade operation method re-available, greatly reducing the peak of memory usage and avoiding memory leakage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988030A_ABST
    Figure CN119988030A_ABST
Patent Text Reader

Abstract

The invention discloses a cache scheduling method for a bottleneck residual block with an attention mechanism, relates to the field of cache scheduling, and aims to eliminate the front-back dependency relationship between operational data and balance the operand on two branches by decoupling the process of solving the result of the SE attention mechanism block and the process of finally solving the output result of the whole Block. According to the method, the problem that the demand of deploying the neural network of the bottleneck residual block with the attention mechanism on the DRAM-free system on the on-chip RAM is too high is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of cache scheduling, and in particular to a cache scheduling method for a bottleneck residual block with an attention mechanism. Background Art

[0002] In 2017, Hu et al. proposed SENet (Squeeze-and-Excitation Network) in the paper "Squeeze-and-Excitation Networks", which is a typical implementation of the channel attention mechanism. SENet dynamically assigns weights to each channel by introducing the two stages of Squeeze and Excitation, thereby enhancing the network's ability to express features. SENet achieved remarkable results in the ImageNet competition and attracted widespread attention from academia and industry. Later, this mechanism was also introduced in Goolge's "Searching for MobilenetV3" to improve network performance. However, machine learning algorithms based on this attention mechanism often require a larger cache space to cache intermediate data, which is not friendly to some ultra-small devices on the end side, such as single-chip microcomputers.

[0003] At the same time, the deployment scheme is currently divided into the following two categories according to the size:

[0004] a. Deployment solutions represented by Thinker, Eyeriss and other DRAM-based general-purpose NPUs. Inference is performed by deploying some lightweight neural network models suitable for edge deployment on these devices. This method causes high system power consumption due to the presence of DRAM. The typical power consumption is between several hundred milliwatts and several watts, and stable power supply is required for inference.

[0005] b. Deployed on microcontroller-level devices, this method greatly reduces power consumption. The typical full-load power consumption is only tens of milliwatts, and only a button battery is needed to keep the detection equipment running for days or even months. The disadvantage of this method is that since the on-chip cache is too small and there is no off-chip dynamic memory, memory leaks are very likely to occur. Summary of the invention

[0006] In view of the above-mentioned deficiencies in the prior art, the present invention provides a cache scheduling method for a bottleneck residual block with an attention mechanism, which solves the problem that the existing neural network blocks containing bottleneck residual blocks with an attention mechanism are easily prone to memory leakage when deployed on devices at the microcontroller level.

[0007] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is:

[0008] A cache scheduling method for a bottleneck residual block with an attention mechanism is provided, which comprises the following steps:

[0009] S1. Using the input feature map of the neural network block including the bottleneck residual block with the attention mechanism as the target feature map, reading the target feature map in a channel-first order and performing a 1×1 convolution operation, outputting pixel by pixel in a feature map-first manner, and obtaining a first operation result; applying for a first cache space corresponding to the amount of data of the overlapping part in the 3×3 depth-separable convolution operation for the first operation result, and discarding the rest of the data;

[0010] S2, reading the first operation result in a feature map priority order and performing a 3×3 depth-separable convolution operation, outputting pixel by pixel in a feature map priority manner, and obtaining a second operation result;

[0011] S3, adding pixel data of the second operation result in the channel through an adder to obtain an accumulated result; dividing the accumulated result by the number of input feature map pixels through a divider to obtain a division result; releasing the first cache space, applying for a second cache space of the same size as the number of channels, and writing the division result into the second cache space;

[0012] S4, performing two full-connection operations on the division result in the second cache space to obtain a third operation result; wherein, in each full-connection operation, a third cache space equivalent to the number of full-connection output channels is applied, and when each full-connection operation is completed, the third cache space corresponding to each full-connection operation is released;

[0013] S5, cutting the neural network block containing the bottleneck residual block with the attention mechanism into a number of rectangular tiles in the length and width directions of the target feature map;

[0014] S6, performing a 1×1 convolution operation on the target feature map in the order of channel priority for input and feature map priority for output within the rectangular tile to obtain a fourth operation result; applying for a fourth cache space corresponding to the amount of data of the overlapping part in the 3×3 depth-separable convolution operation for the fourth operation result, and discarding the remaining data;

[0015] S7, performing a 3×3 depth-separable convolution operation on the fourth operation result in the order of input as feature map priority and output as feature map priority in the rectangular tile, to obtain a fifth operation result; releasing the fourth cache space, applying for a fifth cache space equivalent to the output size of the rectangular tile for each rectangular tile, and inputting the fifth operation result into the corresponding fifth cache space;

[0016] S8. Perform channel multiplication and 1×1 depth-wise separable convolution operations on the third operation result and the fifth operation result in the order of channel priority for input and feature map priority for output within the rectangular Tile to obtain a sixth operation result; perform a residual connection on the sixth operation result and the target feature map to complete the operation of the bottleneck residual block with an attention mechanism, and release the cache space for storing the target feature map and the fifth cache space.

[0017] An electronic device is provided, comprising:

[0018] a memory storing executable instructions; and

[0019] A processor is configured to execute executable instructions in the memory to implement a cache scheduling method for a bottleneck residual block with an attention mechanism.

[0020] A readable storage medium is provided, on which executable instructions are stored. When the executable instructions are executed by a processor, a cache scheduling method for a bottleneck residual block with an attention mechanism is implemented.

[0021] The beneficial effects of the present invention are as follows: in the conventional operation process, layer-by-layer operation requires caching the data of the entire intermediate feature map, and the multi-layer cascade operation method can effectively reduce the cache required for the intermediate layer. However, when this method operates the residual block with the SE mechanism, the intermediate data cannot be used for the next operation before the complete intermediate feature map operation is completed, that is, the newly generated intermediate data cannot be consumed immediately, so the multi-layer cascade operation method will be invalid, and the intermediate feature map data is 4 to 6 times the size of the input and output feature map data, and the focus of the cache overhead lies in this. This method completes the decoupling of the SE block operation process from the entire bottleneck residual block operation process, so that the multi-layer cascade operation method can be used again, thereby greatly reducing the peak memory usage. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is a schematic diagram of the process of this method;

[0023] Figure 2 A schematic diagram of cache usage and release in step S1;

[0024] Figure 3 A schematic diagram of cache usage and release in steps S2, S3, and S4;

[0025] Figure 4 A schematic diagram of the use and release of the cache in the Nth Tile in steps S5 and S6;

[0026] Figure 5 A schematic diagram of the use and release of the cache in the N+1th Tile in step S6;

[0027] Figure 6 A schematic diagram of the use and release of the cache in the Nth Tile in step S7;

[0028] Figure 7 Schematic diagram of cache usage and release in the Nth Tile in the channel multiplication and 1×1 depthwise separable convolution operation in step S8;

[0029] Figure 8 Schematic diagram of cache usage and release in the Nth Tile in the residual connection in step S8. DETAILED DESCRIPTION

[0030] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.

[0031] like Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 , Figure 6 , Figure 7 and Figure 8 As shown (red is the cache that needs to be occupied, green is the free cache, yellow is the cache that will be released after the step is completed, RAMbuffer represents the feature map cache, WGT_buffer represents the weight cache, MAC Core represents the computing acceleration hardware, Block_IFM represents the feature map cache that has been used before S5, IDLE represents the free feature map cache, Block_IFM_Tile represents the cache occupied by each Tile block in the tiled feature map after S5, the cache scheduling method of the bottleneck residual block with attention mechanism includes the following steps:

[0032] S1. The input feature map of the neural network block including the bottleneck residual block with the attention mechanism is used as the target feature map, the target feature map is read in a channel-first order and a 1×1 convolution operation is performed, and the pixel-by-pixel output is performed in a feature map-first manner to obtain a first operation result; a first cache space corresponding to the amount of data of the overlapping part in the 3×3 depth-separable convolution operation is applied for the first operation result, and the rest of the data is discarded, that is, a first cache space that can hold the data of two rows plus two pixel positions in the third row is applied, and the rest of the data is discarded and not written into the cache;

[0033] S2, reading the first operation result in a feature map priority order and performing a 3×3 depth-separable convolution operation, outputting pixel by pixel in a feature map priority manner, and obtaining a second operation result;

[0034] S3, adding pixel data of the second operation result in the channel through an adder to obtain an accumulated result; dividing the accumulated result by the number of input feature map pixels (global average pooling) through a divider to obtain a division result; releasing the first cache space, applying for a second cache space of the same size as the number of channels, and writing the division result into the second cache space;

[0035] S4, performing two full-connection operations on the division result in the second cache space to obtain a third operation result; wherein, in each full-connection operation, a third cache space equivalent to the number of full-connection output channels is applied, and when each full-connection operation is completed, the third cache space corresponding to each full-connection operation is released;

[0036] S5, cutting the neural network block containing the bottleneck residual block with the attention mechanism into a number of rectangular tiles in the length and width directions of the target feature map;

[0037] S6, performing a 1×1 convolution operation on the target feature map in the order of channel priority for input and feature map priority for output within the rectangular tile to obtain a fourth operation result; applying for a fourth cache space corresponding to the amount of data of the overlapping part in the 3×3 depth-separable convolution operation for the fourth operation result, and discarding the remaining data;

[0038] S7, performing a 3×3 depth-separable convolution operation on the fourth operation result in the order of input as feature map priority and output as feature map priority in the rectangular tile, to obtain a fifth operation result; releasing the fourth cache space, applying for a fifth cache space equivalent to the output size of the rectangular tile for each rectangular tile, and inputting the fifth operation result into the corresponding fifth cache space;

[0039] S8. Perform channel multiplication and 1×1 depth-wise separable convolution operations on the third operation result and the fifth operation result in the order of channel priority for input and feature map priority for output within the rectangular Tile to obtain a sixth operation result; perform a residual connection on the sixth operation result and the target feature map to complete the operation of the bottleneck residual block with an attention mechanism, and release the cache space for storing the target feature map and the fifth cache space.

[0040] The electronic device includes:

[0041] a memory storing executable instructions; and

[0042] A processor is configured to execute executable instructions in the memory to implement a cache scheduling method for a bottleneck residual block with an attention mechanism.

[0043] The readable storage medium stores executable instructions thereon, and when the executable instructions are executed by a processor, a cache scheduling method for a bottleneck residual block with an attention mechanism is implemented.

[0044] In the specific implementation process, the accumulation and division in step S3 are performed once for each pixel output in step S2, instead of writing the result into the cache and waiting for the convolution of step S2 to be completely completed before accumulation. This is equivalent to step S2 and step S3 forming a pipeline with a scalar granularity.

[0045] In step S8, the multiplication operation resources can be reused by controlling the data path. At the first moment, the operation result of a pixel point in step S7 is used as an input of the multiplier, and the operation result of the corresponding channel in step S4 is used as another input; at the second moment, the output result is returned to one end of the multiplier input, and the other end is input with the weight of the 1×1 depth-separable convolution in step S8. Finally, the result is residually connected with the target feature map, and the residual connection result is written back to the cache, and the corresponding part of the target feature map that has been completely calculated is released.

[0046] In this embodiment, the output data block size of the 1×1 depthwise separable convolution in step S1, step S2, step S3, step S6 and step S8 is one pixel; the data block size output by step S4 is a vector; the size of the output data block of the channel multiplication and residual connection in step S7 and step S8 is a rectangular Tile.

[0047] In one embodiment of the present invention, the original input image size of the mobilenetv3 network is 128×128, and the rectangular tile size is 8×8. The input of the third layer block of the mobilenetv3 network is 32×32×16, the output is 32×32×24, the inverted residual block dimension magnification is 6, and the SE block feature shrinkage is 4. This layer can reflect the amount of computation of each layer of block in mobilenetv3 to a certain extent. In the following formula, Cin is the number of channels of the Block input feature map, Cexp is the number of channels of the 3×3 depth-separable convolution input feature map, which is equal to Cin The inverted residual block dimension magnification, kennel_size is the convolution window size, L is the side length of the two-dimensional feature map, Tile_L is the side length of the Tile block, Cr is the number of channels of the output feature map of the first fully connected layer in the SE block, and the number of channels of the input feature map of the second fully connected layer is equal to Cexp / SE block feature contraction ratio, and Cout is the number of channels of the Block output feature map.

[0048] Existing methods:

[0049] Cin Cexp L L (input 1×1 convolution operation amount) + kennel_size Cexp L L (3×3 depth-separable convolution operation amount)

[0050] +Cout Cexp L L (output 1×1 convolution operation amount) + Cexp Cr 2(Full connection operation amount in SE) = 4821504

[0052] This method:

[0053] (Cin Cexp L L (input 1×1 convolution operation amount) + kennel_size Cexp L L (3×3 depth-wise separable convolution operation amount) 2

[0054] +Cout Cexp L L (output 1×1 convolution operation amount) + Cexp Cr 2(Full connection operation amount in SE) = 7279104

[0056] The amount of computation per SE block is increased to about 150% of the original amount. However, in an actual forward reasoning process, this block only accounts for about 50% (refer to the structure of Large and Small in the searching for mobilenetV3 paper, https: / / ieeexplore.ieee.org / document / 9008835), and the actual increase in the amount of computation in the final forward reasoning process does not exceed 130%.

[0057] Peak storage usage:

[0058] Traditional method: Cin L L(input feature map cache occupancy) + Cexp L L(intermediate feature map cache occupancy) + (Cexp + Cr) (SE AVGPooling occupancy) = 114808

[0060] + (Cin + kennel_size + Cexp) (weight cache occupancy) 681

[0062] =115489

[0063] This method: Cin L L (input feature map cache occupancy) + 2 Cexp (Tile_L) (Tile_L)(intermediate feature map cache occupancy) + (Cexp+Cr)(SE AVGPooling occupancy) = 28792

[0065] + (Cin + kennel_size + Cexp) (weight cache occupancy) = 681

[0067] =29473

[0068] It can be seen that the peak storage usage of this method is reduced to 25.5% of the original.

[0069] In summary, the present invention greatly reduces the peak memory usage and solves the problem that the existing neural network blocks containing bottleneck residual blocks with attention mechanisms are prone to memory leakage when deployed on devices at the level of single-chip microcomputers.

Claims

1. A cache scheduling method for bottleneck residual blocks with an attention mechanism, characterized in that: The following steps are involved: S1. Using the input feature map of the neural network block including the bottleneck residual block with the attention mechanism as the target feature map, reading the target feature map in a channel-first order and performing a 1×1 convolution operation, outputting pixel by pixel in a feature map-first manner, and obtaining a first operation result; applying for a first cache space corresponding to the amount of data of the overlapping part in the 3×3 depth-separable convolution operation for the first operation result, and discarding the rest of the data; S2, reading the first operation result in a feature map priority order and performing a 3×3 depth-separable convolution operation, outputting pixel by pixel in a feature map priority manner, and obtaining a second operation result; S3, accumulating the pixel data of the second operation result in the channel through an adder to obtain an accumulation result; The accumulated result is divided by the number of pixels of the input feature map through a divider to obtain a division result; Release the first cache space, apply for a second cache space of the same size as the number of channels, and write the division result into the second cache space; S4, performing two full-connection operations on the division result in the second cache space to obtain a third operation result; wherein, in each full-connection operation, a third cache space equivalent to the number of full-connection output channels is applied, and when each full-connection operation is completed, the third cache space corresponding to each full-connection operation is released; S5, cutting the neural network block containing the bottleneck residual block with the attention mechanism into a number of rectangular tiles in the length and width directions of the target feature map; S6, performing a 1×1 convolution operation on the target feature map in the order of channel priority for input and feature map priority for output within the rectangular tile to obtain a fourth operation result; applying for a fourth cache space corresponding to the amount of data of the overlapping part in the 3×3 depth-separable convolution operation for the fourth operation result, and discarding the remaining data; S7, performing a 3×3 depth-separable convolution operation on the fourth operation result in the order of input as feature map priority and output as feature map priority in the rectangular tile, to obtain a fifth operation result; releasing the fourth cache space, applying for a fifth cache space equivalent to the output size of the rectangular tile for each rectangular tile, and inputting the fifth operation result into the corresponding fifth cache space; S8. Perform channel multiplication and 1×1 depth-wise separable convolution operations on the third operation result and the fifth operation result in the order of channel priority for input and feature map priority for output within the rectangular Tile to obtain a sixth operation result; perform a residual connection on the sixth operation result and the target feature map to complete the operation of the bottleneck residual block with an attention mechanism, and release the cache space for storing the target feature map and the fifth cache space.

2. An electronic device, characterized in that: include: A memory storing executable instructions; as well as A processor is configured to execute the executable instructions in the memory to implement the method of claim 1.

3. A readable storage medium having executable instructions stored thereon, characterized in that: When the executable instructions are executed by a processor, the method according to claim 1 is implemented. The cache scheduling method of the bottleneck residual block with attention mechanism is used for the processor to perform the bottleneck residual operation with attention mechanism.