A self - selective receptive field block, an image processing method and an application
By introducing self-selective receptive field blocks into the multi-scale convolution module, using the pyramid convolution group and the adaptive weighted fusion module, the problems of receptive field fixed and large calculation amount are solved, the full understanding of scene information and the control of calculation amount are achieved, and the accuracy and performance of pedestrian target detection are improved.
Patent Information
- Application Number
- CN202210786208.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-04
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-07-04
AI Technical Summary
The existing multi-scale convolution modules have problems such as fixed receptive fields, high design complexity, and increased calculation and parameter quantities. It is difficult to effectively take into account locality and globality, and it is difficult to fully understand scene information.
A self-selective receptive field block is designed. Through the pyramid convolution group and the adaptive weighted fusion module, the receptive field size is automatically selected according to the task, and the semantic information from different levels in the image from local to global are obtained, and the calculation amount and parameter amount are controlled.
It has achieved a full understanding of the target scenario, significantly controlled the calculation amount and parameter amount, improved the accuracy and performance of pedestrian target detection, and strong adaptability.
Smart Images

Figure CN115273141B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and specifically relates to a self - selective receptive field block, an image processing method and an application. Background Art
[0002] Person Re - identification (Re - ID) technology is a technology that aims to solve the problem of matching the same pedestrian at different times and places based on the whole - body information of the pedestrian. It can make up for the visual limitations of fixed cameras and has wide application value in security fields such as shopping malls and airports.
[0003] In the network design of Re - ID, multi - scale convolution modules are often introduced to enhance the diversity of the receptive fields of pedestrian feature maps (such as RFB and ASPP modules) to obtain different levels of spatial context information of pedestrians. For example, the pedestrian target detection method and system disclosed in the patent document with the publication number CN114332908A. Among them, pedestrian re - identification is implemented by using a trained pedestrian re - identification network; the pedestrian re - identification network incorporates a multi - scale feature extraction module to expand the receptive field and improve the ability of target feature extraction. On the improved pedestrian re - identification network, a multi - scale feature extraction module is fused to further enhance the accuracy of pedestrian target detection, enabling more diverse processing of target detection in complex scenes and improving the accuracy of pedestrian target detection. However, most of the existing multi - scale convolution modules have the following two technical problems:
[0004] 1) Most multi - scale convolution modules use fixed receptive fields, which are difficult to be adjusted autonomously according to different tasks and different models, and cannot achieve a full understanding of scene information.
[0005] 2) The design complexity of these modules is relatively high, which slows down the inference speed. The most direct manifestation is a large increase in the amount of computation and the number of parameters. Summary of the Invention
[0006] The purpose of the present invention is to provide a self - selective receptive field block, an image processing method and an application. The present invention designs a pyramid convolution group by leveraging the advantages of multiple different - scale dilated convolutions to obtain different levels of semantic information from local to global in the image, and enables the model to autonomously select the required receptive field size according to the task through an adaptive weighted fusion method, which can fully understand the target scene and significantly control the amount of computation and the number of parameters, effectively solving the technical problems that the existing models cannot effectively balance locality and globality, are difficult to fully understand scene information, and have a large amount of computation and a large number of parameters.
[0007] To achieve the above - mentioned purpose, the technical solution adopted by the present invention is as follows:
[0008] A self - selective receptive field block, characterized in that: it includes an image input module, a squeezing module, a pyramid convolution group, a weighted fusion module, an excitation module and an image output module. The pyramid convolution group includes a plurality of parallel convolution branches. Among them,
[0009] The image input module is used to input the feature map to be processed;
[0010] The squeezing module is used to compress the number of channels of the feature map to be processed;
[0011] The pyramid convolution group processes the compressed feature map to be processed through convolution branches respectively. Each convolution branch first encodes the spatial information of the image to obtain an intermediate feature map with the same dimension as the input, and then fuses the spatial and channel information together in the local receptive field to obtain an enhanced feature map;
[0012] The weighted fusion module is used to adaptively weight - fuse the enhanced feature maps output by each convolution branch by combining weight factors respectively, and obtain a fusion feature map with high - fine - grained representation ability;
[0013] The excitation module is used to restore the number of channels of the fusion feature map and obtain a restored feature map;
[0014] The image output module is used to add the restored feature map and the feature map to be processed element - by - element to obtain a final output feature map that takes into account both high - level semantic information and original detail information.
[0015] The pyramid convolution group includes four parallel convolution branches. Each convolution branch uses a dilated convolution layer with a kernel size of 3×3, and the dilation rate rate of each convolution branch is different; among them, each convolution branch first uses a dilated convolution layer with a kernel of 3 and Groups = C to encode the spatial information of the image, where C represents the number of channels of the feature map to be processed. After encoding, an intermediate feature map with the same dimension as the input is obtained, and then a convolution layer with a kernel of 1 and Groups = 1 is applied to the intermediate feature map to fuse the spatial and channel information together in the local receptive field to obtain an enhanced feature map.
[0016] The squeezing module compresses the number of channels of the feature map to be processed to 1 / r of the original based on the convolution layer Conv1×1, where r is a scaling factor; the excitation module realizes the restoration of the number of channels of the fusion feature map based on the composite convolution layer of Conv1×1 - ReLU.
[0017] An image processing method, characterized by including the following steps:
[0018] Step A: Input a feature map to be processed with dimensions of H×W×C, and use a squeezing operation to compress the number of channels of the feature map to be processed;
[0019] Step B: Input the compressed feature map to be processed into a pyramid convolution group containing multiple parallel convolution branches for processing. Each convolution branch first encodes the spatial information of the image to obtain an intermediate feature map with the same dimension as the input, and then fuses the spatial and channel information together in the local receptive field to obtain an enhanced feature map;
[0020] Step C: Adaptive weighted fusion is performed on the enhanced feature maps output by multiple convolution branches by combining weight factors respectively to obtain a fusion feature map with high fine-grained representation ability;
[0021] Step D: Perform an excitation operation on the fusion feature map to restore the compressed number of channels of the fusion feature map to obtain a restored feature map;
[0022] Step E: Introduce a Shortcut operation to add the restored feature map and the feature map to be processed element by element to obtain a final output feature map that takes into account both high-level semantic information and original detailed information.
[0023] In step A, based on the convolutional layer Conv1×1, the number of channels of the feature map to be processed is compressed to 1 / r of the original, where r is the scaling factor; among them, assuming the input feature map is F and the compressed feature map to be processed is F1, then:
[0024] F1 = F sq (F), F1 ∈ R H×W×C / r
[0025] In the formula, F sq The essence of the (·) function is the convolutional layer Conv1×1.
[0026] In step B, the pyramid convolution group includes four parallel convolution branches. Each convolution branch uses a dilated convolution layer with a kernel size of 3×3, and the dilation rate rate of each convolution branch is different; among them, each convolution branch first uses a dilated convolution layer with a kernel of 3 and Groups = C to encode the spatial information of the image, where C represents the number of channels of the feature map to be processed. After encoding, an intermediate feature map with the same dimension as the input is obtained, and then a convolution layer with a kernel of 1 and Groups = 1 is applied to the intermediate feature map to fuse the spatial and channel information together in the local receptive field to obtain an enhanced feature map; assuming the rates of the four convolution branches are 1, 2, 3, and 4 respectively, the output enhanced feature maps are F 3×3 、F 5×5 、F 7×7 and F 9×9 respectively.
[0027] In step C, assuming the fusion feature map is F2, then:
[0028] F2 = w 3×3·F 3×3 +w 5×5 ·F 5×5 +w 7×7 ·F 7×7 +w 9×9 ·F 9×9 ,R H×W×C / r
[0029] In the formula, F 3×3 、F 5×5 、F 7×7 and F 9×9 respectively represent the enhanced feature maps output by four different convolutional branches with rate = 1, 2, 3, and 4; w 3×3 、w 5×5 、w 7×7 and w 9×9 correspond to the weight factors of the four convolutional branches respectively; + represents element-wise addition.
[0030] In step D, the composite convolutional layer based on Conv1×1-ReLU restores the number of channels C / r of the fused feature map to C; where, assuming the restored feature map is F3, then:
[0031] F3 = F ex (F2), F3 ∈ R H×W×C
[0032] In the formula, the essence of the F ex (·) function is the Conv1×1-ReLU composite convolutional layer.
[0033] In step E, assuming the final output feature map is F', then:
[0034] F′ = F + F3, F′ ∈ R H×W×C
[0035] In the formula, + represents element-wise addition.
[0036] An application of a self-selective receptive field block, characterized in that: the self-selective receptive field block described in any one of claims 1-3 is applied to neural network models including but not limited to VGGNet series, ResNet series, and MobileNet series.
[0037] Adopting the above technical solution, the beneficial technical effects of the present invention are:
[0038] 1. The self-selected receptive field block described in the present invention refers to the Self-selected Receptive Field block, abbreviated as the SRF block (the same hereinafter), which includes an image input module, a squeezing module, a pyramid convolution group including a plurality of parallel convolution branches, a weighted fusion module, an excitation module, and an image output module. Compared with the prior art, the present invention designs a pyramid convolution group by leveraging the advantages of multiple dilated convolutions with different scales to obtain different levels of semantic information from local to global in the image, and through an adaptive weighted fusion method, enables the model to autonomously select the required receptive field size according to the task, which can fully understand the target scene, significantly control the computational amount and the number of parameters, and effectively solve the technical problems that the existing models cannot effectively balance locality and globality, are difficult to fully understand the scene information, and have a large computational amount and a large number of parameters.
[0039] 2. The self-selected receptive field block described in the present invention is a plug-and-play and truly lightweight module, and its specific advantages are as follows:
[0040] (1) In terms of feature processing. The SRF block obtains context information with different scales, different levels, and different semantics in the pedestrian image through multiple dilated convolution layers with different kernel sizes. This information takes into account both locality and globality, and can achieve a full understanding of the pedestrian target and the image scene, and has good adaptability to problems such as pedestrian multi-scale and target occlusion. In addition, the adaptive weighting operation in the SRF block enables the learning of each convolution branch of the pyramid convolution group to supervise and complement each other, greatly promoting the performance expression of the SRF block.
[0041] (2) In terms of parameters and computational amount. On the one hand, the SRF block first maps the input feature map to a low dimension to reduce the increase in the number of parameters and the computational amount brought by subsequent pyramid convolution operations; on the other hand, each branch of the pyramid convolution group all uses 3×3 convolutions with different dilation rates. Compared with the standard 3×3 convolution, while increasing the receptive field, it still maintains a low number of parameters and a low computational amount; in particular, all branches of the pyramid convolution group adopt the operation technique of "group convolution", which further reduces the number of parameters and the computational amount.
[0042] (3) In terms of flexibility. While processing the input feature map, the SRF block does not change its spatial size and channel structure, so it can be easily transplanted to various positions of the existing convolutional neural network or used as a substitute for other modules, thereby playing a role in feature enhancement.
[0043] 3. The image processing method described in the present invention includes the following five steps, where
[0044] Step A uses a squeezing operation to compress the number of channels of the feature map to be processed. Its advantage is that it can make the operation platform of the subsequent pyramid convolution group at a low dimension, thereby reducing the number of parameters and the amount of computation.
[0045] Step B processes the compressed feature map to be processed using multiple parallel convolution branches respectively. The dilated convolution layers used all adopt the operation technique of "group convolution". This technique enables each dilated convolution layer to obtain different local information while significantly reducing the number of parameters and the amount of computation.
[0046] Step C adaptively weights and fuses the enhanced feature maps output by multiple convolution branches by combining weight factors respectively. Since the weight factors used can be autonomously learned through network training instead of setting unchanging hyperparameters, the required receptive field size can be autonomously selected according to the task, thereby achieving a full understanding of the target scene.
[0047] Step D restores the compressed number of channels of the fused feature map through an excitation operation on the fused feature map, which can perform non-linear mapping on the features processed by the pyramid convolution group, thereby improving the abstract expression ability of the model.
[0048] Step E introduces a Shortcut operation, which is beneficial to the gradient transmission of the SRF enhancement module, thereby making the training and learning of the SRF block easier.
[0049] 4. When the self-selected receptive field block of the present invention is applied to neural network models including but not limited to the VGGNet series, ResNet series, and MobileNet series, it has the following advantages:
[0050] (1) In terms of performance, the SRF block designs a pyramid convolution group by leveraging the advantages of multiple dilated convolutions at different scales to obtain different levels of semantic information from local to global in the image, and enables the model to autonomously select the required receptive field size according to the task through adaptive weighted fusion, thereby achieving a full understanding of the target scene.
[0051] (2) In terms of complexity, the SRF block first constrains the complexity of the model from a global perspective by introducing a channel scaling factor. Secondly, the SRF block specifically designs a "group convolution operation", which can significantly control the amount of computation and the number of parameters generated by the model while modeling the feature space information and channel information. Experiments show that the performance of the ResNet-50 baseline model after using the SRF block on the selected Re-ID occlusion dataset has increased by more than 10%, while the increase in its amount of computation and the number of parameters is only about 3%. Description of the Drawings
[0052] Figure 1This is the schematic diagram of the present invention;
[0053] Figure 2 This is the comparison diagram of the grouped convolution operation and the standard convolution in the SRF block;
[0054] Figure 3 This is the schematic diagram of the adaptive weighted fusion in the SRF block;
[0055] Figure 4 This is the flow chart of the present invention;
[0056] Figure 5 This is the integration scheme diagram of the SRF block and the residual unit in ResNet-50;
[0057] Figure 6 This is the schematic diagram of the SRF-ResNet-50 model. Detailed implementation manners
[0058] Example 1
[0059] The present invention discloses a self-selective receptive field block (SRF block), as Figure 1 shown, which includes an image input module, a squeezing module, a pyramid convolution group, a weighted fusion module, an excitation module and an image output module. The pyramid convolution group includes a plurality of parallel convolution branches. The functions of each module are as follows:
[0060] The image input module is used to input the feature map to be processed.
[0061] The squeezing module is used to compress the number of channels of the feature map to be processed. The function of this squeezing module is to compress the number of channels of the feature map to be processed to 1 / r of the original based on the convolutional layer Conv1×1, where r is the scaling factor. Here, Conv1×1 refers to the convolutional layer with a kernel size of 1×1 (the same below). Through the squeezing operation, the operation platform of the subsequent pyramid convolution group can be in a low dimension, thereby reducing the number of parameters and the computational amount.
[0062] The pyramid convolution group processes the compressed feature map to be processed through convolution branches respectively, encoding the object and the image context at multiple scales. To prevent the performance degradation caused by the "grid effect" when the dilation rate is too large, the present invention preferably includes four parallel convolution branches in the pyramid convolution group. Each convolution branch uses a dilated convolutional layer with a kernel size of 3×3, and the dilation rate rate of each convolution branch is different, but preferably increases in sequence. As Figure 1 shown, when the dilation rates rate of the four convolution branches are 1, 2, 3, and 4 respectively, the actual receptive field sizes of the four dilated convolutional layers are 3×3, 5×5, 7×7, and 9×9 respectively.
[0063] For each convolution branch of the pyramid convolution group in the SRF block, the "grouped convolution operation" adopted by the present invention is as follows: Figure 2 As shown in (a): For any convolution branch, first, a hole convolution layer with a kernel of 3 and Groups = C is used to encode the compressed feature map F to be processed. i The spatial information of the feature map is C, which represents the number of channels of the feature map to be processed. After the encoding is completed, the intermediate feature map with the same dimension as the input is obtained. Secondly, a convolutional layer with a kernel of 1 and Groups = 1 is applied to the intermediate feature map to fuse the spatial and channel information in the local receptive field to obtain the enhanced enhanced feature map F. o . At this point, the feature enhancement process performed by the grouped convolution operation in a convolution branch is completed. In particular, in practical applications, for other different convolution branches of the pyramid convolution group, the present invention only needs to set different rates to perform the same processing flow. It can be seen that during the operation of the entire pyramid convolution group, the present invention always keeps the feature dimensions of the input and output unified, which makes the adjustment and transplantation of the SRF block structure inherently flexible.
[0064] The weighted fusion module is used to adaptively weight the enhanced feature maps output by each convolution branch by combining the weight factors, and obtain a fused feature map with high fine-grained representation capabilities. The weight factors can be learned autonomously through network training instead of setting constant hyperparameters.
[0065] What needs to be explained about the weighted fusion module is that Figure 3 As shown in the figure, the weighted fusion module uses adaptive weighted fusion to aggregate the multi-scale context information generated by the pyramid convolution group. In particular, the weight factors used in the SRF block are coefficients rather than vectors. This is because the enhanced feature maps after the pyramid convolution group operation have semantic similarities in their respective channel directions, so there is no need to let each channel feature learn different weight factors; based on this, the feature fusion method in the SRF block is "element addition" rather than "channel merging".
[0066] Further, such as Figure 3 As shown, assuming and Represent the feature map F 3×3 、F 5×5 、F 7×7 、F 9×9 and F fus The feature plane of the i-th channel, i = 1, 2, ..., C, then the fusion feature plane The calculation formula is:
[0067]
[0068] In the formula, + represents the addition of corresponding elements; w 3×3 , w 5×5 , w 7×7 , w 9×9 respectively represent weight factors, and w 3×3 , w 5×5 , w 7×7 , w 9×9 ∈[0, 1]. The present invention does not set w 3×3 + w 5×5 + w 7×7 + w 9×9 = 1, because the feature information of the four convolutional branches is not mutually exclusive. If the outputs of the four convolutional branches all play a positive role in the target task, then they should all be emphasized.
[0069] The excitation module is used to restore the number of channels of the fused feature map and obtain the restored feature map. The function of this excitation module is to realize the restoration of the number of channels of the fused feature map based on the composite convolutional layer of Conv1×1 - ReLU. Specifically, it is to restore the number of channels C / r to C. The purpose of the ReLU function is to perform a non - linear mapping on the fused feature map to improve the abstract expression ability of the model.
[0070] The image output module is used to perform element - by - element addition of the restored feature map and the feature map to be processed, that is, corresponding element addition, and finally obtain the final output feature map that takes into account both high - level semantic information and original detail information. Among them, the Shortcut operation can be used to achieve element - by - element addition. The use of the Shortcut operation is beneficial to the gradient transfer of the SRF enhancement module, thus making the training and learning of the SRF block easier.
[0071] In addition, in order to fully illustrate the advantage of the pyramid convolution group in the SRF block in terms of model complexity, for whether to use grouped convolution operations, the present invention makes a comparison of the parameter quantity and computational complexity before and after. As Figure 2 (a) shows, the parameter quantity and computational complexity generated by the SRF block are respectively:
[0072] Params = 3 2 ·C + 1 2 ·C·C = (9 + C)·C
[0073] FLOPs = 2·3 2 ·C·(W·H)+2·1 2 ·C·C·(W·H)=(18 + 2C)·C·(W·H)
[0074] And as Figure 2(b) For the case of using standard convolution, it is a direct one-step process. It uses C standard convolution layers with a kernel size of 3×3 (Groups = 1) to process the input and obtains an output of the same dimension. The required number of parameters and the amount of computation are respectively:
[0075] Params = 3 2 ·C·C = 9·C 2
[0076] FLOPs = 2·3 2 ·C·C·(W·H) = 18·C 2 ·(W·H)
[0077] As can be seen from the comparison of the above formulas, compared with standard convolution, the number of parameters and the amount of computation generated by using the grouped convolution operation of the present invention are both R times the original, that is:
[0078]
[0079] In the formula, k is the kernel size.
[0080] Therefore, after using the grouped convolution operation of the present invention, the number of parameters and the computational cost of the convolution layer are reduced to a multiple approximately equal to the square factor of the kernel, which is very beneficial under the conditions of limited computing resources or mobile platform applications.
[0081] The SRF block described in this embodiment can be applied to neural network models including but not limited to the VGGNet series, ResNet series, and MobileNet series. During the actual training of the neural network model, considering that the neural network model can only learn real-valued parameters, in order to learn four weight factors with a value range of [0,1], the present invention cleverly introduces the Sigmoid function to constrain the four real-valued parameters to be learned by the neural network model within the interval [0,1]. The Sigmoid function is also called the Logistic function, which is a common S-shaped function in biology and is also called the S-shaped growth curve. In information science, due to its properties such as monotonic increase and monotonic increase of the inverse function, it has advantages such as smoothness and easy differentiation, so it is often used as the activation function of neural networks.
[0082] Furthermore, let w1, w2, w3, and w4 respectively correspond to the four weight factors w to be trained 3×3 , w 5×5 , w 7×7 and w 9×9 , then:
[0083]
[0084] As shown in the above formula, by means of the Sigmoid function, the neural network model's w3×3 , w 5×5 , w 7×7 and w 9×9 The learning problem of w is transformed into the learning problems of four parameter values x1, x2, x3, and x4 respectively.
[0085] The specific implementation process of this embodiment is as follows:
[0086] Assume that the dimension of the input feature map F to be processed is C×H×W, then:
[0087] (1) Use a squeezing module essentially a Conv1×1 convolutional layer to perform a squeezing operation on the feature map F to be processed, and obtain a compressed feature map F1 to be processed with a dimension of C / r×H×W;
[0088] (2) Input the compressed feature map F1 to be processed into the pyramid convolutional group to perform pyramid convolution operations, and obtain enhanced feature maps F 3×3 , F 5×5 , F 7×7 and F 9×9 ;
[0089] (3) Pass the enhanced feature maps F 3×3 , F 5×5 , F 7×7 and F 9×9 through the weighted fusion module, multiply them by different weight factors with value ranges all in [0, 1], and then perform element-wise addition to obtain a fused feature map F2 with a dimension of C / r×H×W;
[0090] (4) Use an excitation module essentially a composite convolutional layer of Conv1×1-ReLU to perform an excitation operation on the fused feature map F2 to obtain a restored feature map F3 with a dimension of C×H×W;
[0091] (5) In the image output module, perform element-wise addition on the feature map F to be processed and the restored feature map F3 to obtain a final output feature map F' with a dimension of C×H×W.
[0092] Compared with the initially input feature map F to be processed, the finally obtained final output feature map F' increases the receptive field of the feature map to be processed, enhances the scale diversity and semantic richness of the feature map to be processed, and thus can more accurately obtain relevant features in the feature map to be processed.
[0093] Embodiment 2
[0094] The present invention discloses an image processing method, which is implemented based on the SRF block described in Embodiment 1, such as Figure 1 , 4As shown, it includes the following steps:
[0095] Step A: Input the feature map to be processed with dimensions of H×W×C, and use a squeezing operation to compress the number of channels of the feature map to be processed.
[0096] Specifically, in this step, based on the convolutional layer Conv1×1, the number of channels of the feature map to be processed is compressed to 1 / r of the original, where r is the scaling factor. Let the input feature map be F, and the compressed feature map to be processed be F1, then:
[0097] F1 = F sq (F), F1 ∈ R H×W×C / r
[0098] In the formula, F sq The essence of the (·) function is the convolutional layer Conv1×1. The starting point of this step is that it can reduce the dimension of the feature map to be processed, so that there will be no large number of parameters and large computational complexity.
[0099] Step B: Input the compressed feature map to be processed into a pyramid convolutional group containing multiple parallel convolutional branches for processing, and encode the object and image context at multiple scales. Each convolutional branch first encodes the spatial information of the image to obtain an intermediate feature map with the same dimension as the input, and then fuses the spatial and channel information together in the local receptive field to obtain an enhanced feature map.
[0100] Specifically, in this step, it is preferred that the pyramid convolutional group includes four parallel convolutional branches. Each convolutional branch uses a dilated convolutional layer with a kernel size of 3×3, and the dilation rate rate of each convolutional branch is different; as Figure 1 shown, when the dilation rates rate of the four convolutional branches are 1, 2, 3, and 4 respectively, the actual receptive field sizes of the four dilated convolutional layers are 3×3, 5×5, 7×7, and 9×9 respectively. Among them, each convolutional branch first uses a dilated convolutional layer with a kernel of 3 and Groups = C to encode the spatial information of the image, where C represents the number of channels of the feature map to be processed. After encoding, an intermediate feature map with the same dimension as the input is obtained, and then a convolutional layer with a kernel of 1 and Groups = 1 is applied to the intermediate feature map to fuse the spatial and channel information together in the local receptive field to obtain an enhanced feature map; let the rates of the four convolutional branches be 1, 2, 3, and 4 respectively, then the output enhanced feature maps are F 3×3 、F 5×5 、F 7×7 and F 9×9Thus, the process of feature enhancement performed by the grouped convolution operation in a convolutional branch ends. In particular, in practical applications, for other different convolutional branches of the pyramid convolution group, the present invention only needs to set different rates to perform the same processing flow. It can be seen that during the operation of the entire pyramid convolution group, the present invention always maintains the unity of the input and output feature dimensions.
[0101] Step C: In order to make the image context information at different levels play a full role, in this step, the enhanced feature maps output by multiple convolutional branches are adaptively weighted and fused respectively with weight factors to obtain a fused feature map with high fine-grained representation ability.
[0102] Specifically, setting the fused feature map as F2, then:
[0103] F2 = w 3×3 ·F 3×3 + w 5×5 ·F 5×5 + w 7×7 ·F 7×7 + w 9×9 ·F 9×9 ,R H×W×C / r
[0104] In the formula, F 3×3 、F 5×5 、F 7×7 and F 9×9 respectively represent the enhanced feature maps output by four different convolutional branches with rate = 1, 2, 3, and 4; w 3×3 、w 5×5 、w 7×7 and w 9×9 correspond to the weight factors of the four convolutional branches respectively; + represents element-wise addition. It should be noted that the four weight factors can be autonomously learned through network training instead of setting invariant hyperparameters.
[0105] Step D: Perform an excitation operation on the fused feature map to restore the compressed number of channels of the fused feature map to obtain a restored feature map.
[0106] Specifically, in this step, based on the composite convolutional layer of Conv1×1-ReLU, the number of channels C / r of the fused feature map is restored to C; where, setting the restored feature map as F3, then:
[0107] F3 = F ex (F2), F3 ∈ R H×W×C
[0108] In the formula, the essence of the F ex (·) function is the composite convolutional layer of Conv1×1-ReLU.
[0109] Different from the F(·) function in step A, the F(·) function sq adds an activation layer ReLU after Conv1×1. The purpose is to perform a non-linear mapping on the features processed by the pyramid convolution group, improving the model's abstract expression ability. In particular, ex the BN layer is not added in the F(·) function to avoid the normalization operation from destroying the mutual dependence relationship between the original feature information. ex H×W×C
[0110] Step E: Introduce the Shortcut operation to perform element-wise addition of the restored feature map and the feature map to be processed, that is, add corresponding elements. After completion, the final output feature map that takes into account both high-level semantic information and original detailed information is obtained.
[0111] Specifically, setting the final output feature map as F', then:
[0112] F′ = F + F3, F′ ∈ R H×W×C
[0113] In the formula, + represents element-wise addition.
[0114] Generally speaking, the present invention designs a pyramid convolution group by leveraging the advantages of multiple dilated convolutions with different scales to obtain different levels of semantic information from local to global in the image, and through the method of adaptive weighted fusion, enables the model to autonomously select the required receptive field size according to the task, can fully understand the target scene and significantly control the computational amount and the number of parameters, effectively solving the technical problems that the existing models cannot effectively balance locality and globality, are difficult to fully understand the scene information, and have large computational amounts and numbers of parameters. And compared with the initially input feature map F to be processed, the finally obtained final output feature map F' increases the receptive field of the feature map to be processed, enhances the scale diversity and semantic richness of the feature map to be processed, so as to more accurately obtain the relevant features in the feature map to be processed.
[0115] Embodiment 3
[0116] The present invention discloses an application of a self-selective receptive field block. The present invention mainly applies the self-selective receptive field block described in Embodiment 1 to neural network models including but not limited to the VGGNet series, ResNet series, and MobileNet series, so as to facilitate the recognition of target features.
[0117] It should be noted that the SRF block is a plug-and-play feature enhancement module that can be flexibly inserted into all general neural network models such as the VGGNet series, ResNet series, MobileNet series, etc., to form an SRFNet instantiated model, which is then used for the Re-ID task.
[0118] For example, if the ResNet-50 network model is used for pedestrian re-identification, with the input being a pedestrian image and the output being the result of pedestrian re-identification. When an SRF block is inserted at a certain position in ResNet-50, it can be called the SRF-ResNet-50 network model. Its mechanism is as follows: with the same input of a pedestrian image, when the input information flow passes through the SRF block, its image features will be enhanced, and then the enhanced features continue to be passed backward, with the output being the result of pedestrian re-identification. Similarly, the same applies to the VGGNet network.
[0119] Next, the present invention provides an integration scheme for the SRF block and the most common ResNet-50. As Figure 5 shown, the present invention inserts the SRF block after the "non-linear operation" and before summing with the "identity mapping branch" in the residual unit of ResNet-50, and it is called SRF-ResNet-50. At this time, the scaling factor r of the SRF block in SRF-ResNet-50 is set to 16.
[0120] Specifically, for the SRF-ResNet-50 model, as Figure 6 shown. It can be seen that the present invention only inserts the SRF-ResBlock into Stage_3, while for other Stages, the original structure of the ResBlock is still retained. This is because for Stage_1 and Stage_2 at the shallow layer of the model, the receptive field of their feature maps is small. During network training, they mostly process small-scale targets and focus on low-level information such as the detailed texture of the image, and do not require the "blessing" of too much high-level semantic information. Therefore, when the SRF block is inserted into Stage_1 and Stage_2, it will seem like "using a sledgehammer to crack a nut" and it is difficult to fully play its role. In addition, for Stage_4, which is at the deep layer of the model, due to multiple samplings of the convolutional layer and pooling layer in ResNet-50, the receptive field of the feature maps at this stage is much larger than that of the shallow layer. Therefore, when performing convolutional operations, the feature information captured by adjacent kernel parameters is too remote (mapped to the original image), resulting in partial or complete loss of some local information. After using the SRF block, its increasing dilation rate will further exacerbate this phenomenon, leading to a decline in model performance due to discontinuous information. Therefore, for the SRF block, integrating it into Stage_3 with an appropriate depth can fully play its role.
[0121] The above are only specific embodiments of the present invention. Any feature disclosed in this specification, unless specifically stated, can be replaced by other equivalent or similar-purpose alternative features; all the disclosed features, or all the steps in any method or process, can be combined in any way except for mutually exclusive features and / or steps.
Claims
1. A self - selective receptive field block, which is a plug - and - play feature enhancement module, and is characterized in that: It includes an image input module, a squeezing module, a pyramid convolution group, a weighted fusion module, an excitation module, and an image output module. The pyramid convolution group includes multiple parallel convolution branches. Among them, The image input module is used to input the feature map to be processed; The squeezing module is used to compress the number of channels of the feature map to be processed; The pyramid convolution group processes the compressed feature map to be processed through convolution branches respectively. Each convolution branch first encodes the spatial information of the image to obtain an intermediate feature map with the same dimension as the input, and then fuses the spatial and channel information together in the local receptive field to obtain an enhanced feature map; The weighted fusion module is used to adaptively weight and fuse the enhanced feature maps output by each convolution branch by combining weight factors respectively, and obtain a fusion feature map with high fine-grained representation ability; The excitation module is used to restore the number of channels of the fusion feature map and obtain a restored feature map; The image output module is used to add the restored feature map and the feature map to be processed element by element to obtain a final output feature map that takes into account both high-level semantic information and original detail information; The pyramid convolution group includes four parallel convolution branches. Each convolution branch uses a dilated convolution layer with a kernel size of 3×3, and the dilation rate rate of each convolution branch is different; among them, each convolution branch first uses a dilated convolution layer with a kernel of 3 and Groups = C to encode the spatial information of the image, where C represents the number of channels of the feature map to be processed. After encoding, an intermediate feature map with the same dimension as the input is obtained, and then a convolution layer with a kernel of 1 and Groups = 1 is applied to the intermediate feature map to fuse the spatial and channel information together in the local receptive field to obtain an enhanced feature map.
2. The self - selective receptive field block according to claim 1, characterized in that: The squeezing module compresses the number of channels of the feature map to be processed to 1 / r of the original based on the convolution layer Conv1×1, where r is the scaling factor; the excitation module restores the number of channels of the fusion feature map based on the composite convolution layer of Conv1×1-ReLU.
3. An image processing method, characterized in that, It includes the following steps: Step A: Input a feature map to be processed with dimensions of H×W×C, and use a squeezing operation to compress the number of channels of the feature map to be processed; Step B: Input the compressed feature map to be processed into a pyramid convolution group containing multiple parallel convolution branches for processing. Each convolution branch first encodes the spatial information of the image to obtain an intermediate feature map with the same dimension as the input, and then fuses the spatial and channel information together in the local receptive field to obtain an enhanced feature map; Step C: Respectively combine weight factors to perform adaptive weighted fusion on the enhanced feature maps output by multiple convolution branches to obtain a fusion feature map with high fine-grained representation ability; Step D: Perform an excitation operation on the fusion feature map to restore the compressed number of channels of the fusion feature map to obtain a restored feature map; Step E: Introduce a Shortcut operation to add the restored feature map and the feature map to be processed element by element to obtain a final output feature map that takes into account both high-level semantic information and original detail information; In step B, the pyramid convolution group includes four parallel convolution branches, each of which uses a hole convolution layer with a kernel size of 3×3, and the hole rate rate of each convolution branch is different; wherein, each convolution branch first uses a hole convolution layer with a kernel of 3 and Groups=C to encode the spatial information of the image, where C represents the number of channels of the feature map to be processed. After the encoding is completed, an intermediate feature map with the same dimension as the input is obtained, and then a convolution layer with a kernel of 1 and Groups=1 is applied to the intermediate feature map, and the spatial and channel information are fused together in the local receptive field to obtain an enhanced feature map; the rates of the four convolution branches are set to 1, 2, 3, and 4, respectively, and the output enhanced feature maps are F and F, respectively. 3×3 、F 5×5 、F 7×7 and F 9×9 .
4. An image processing method according to claim 3, characterized in that: In step A, based on the convolutional layer Conv1×1, the number of channels of the feature map to be processed is compressed to 1 / r of the original, where r is the scaling factor. Assuming the input feature map is F and the processed feature map after compression is F1, then: F1 = F sq (F), F1 ∈ R H×W×C / r where F sq (·) is essentially a 1×1 convolutional layer Conv1×1.
5. An image processing method according to claim 3, characterized in that: In step C, assuming the fused feature map is F2, then: F2 = w 3×3 ·F 3×3 +w 5×5 ·F 5×5 +w 7×7 ·F 7×7 +w 9×9 ·F 9×9 ,F2 ∈ R H×W×C / r where F 3×3 , F 5×5 , F 7×7 and F 9×9 respectively represent the enhanced feature maps output by four different convolutional branches with rate = 1, 2, 3, and 4; w 3×3 , w 5×5 , w 7×7 and w 9×9 correspond to the weight factors of the four convolutional branches respectively; + represents element-wise addition.
6. An image processing method according to claim 5, wherein: In step D, based on the composite convolutional layer of Conv1×1-ReLU, the number of channels C / r of the fused feature map is restored to C. Assuming the restored feature map is F3, then: F3 = F ex (F2), F3 ∈ R H×W×C where F ex (·) The essence of the function is a Conv1×1-ReLU composite convolutional layer.
7. An image processing method according to claim 6, characterized in that: In step E, assuming the final output feature map is F', then: F′ = F + F3, F′ ∈ R H×W×C In the formula, + represents element-wise addition.
8. Application of a self - selective receptive field block, characterized in that: Apply the self-selective receptive field block according to claim 1 or 2 to neural network models including VGGNet series, ResNet series, and MobileNet series.
Citation Information
Patent Citations
Pedestrian target detection method and system
CN114332908A
Scene semantic segmentation method based on deep learning
CN112381097A
Compact multi-scale video foreground segmentation method
CN113592878A