Image feature processing method based on channel dynamic weighting self-enhancement attention mechanism

By dynamically adjusting channel weights in the image feature extraction network, the feature focus on occluded targets and noise suppression are enhanced, which solves the problem of insufficient accuracy of occluded target detection in the existing technology and achieves high-precision occluded target detection.

CN120599424APending Publication Date: 2025-09-05江西省通讯终端产业技术研究院有限公司 +3
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510709430.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-05

Smart Images

  • Figure CN120599424A_ABST
    Figure CN120599424A_ABST
Patent Text Reader

Abstract

The invention discloses an image feature processing method based on a channel dynamic weighting self-intensifying attention mechanism, which belongs to a visual detection image processing technology, and dynamically adjusts the channel weight in the image feature propagation process in the process of performing feature extraction network training by adopting a detection image so as to improve the accuracy of feature extraction. The method comprises the following steps: performing channel expansion on an input feature map of a detection image, decomposing the input feature map into three part feature maps with the same dimensions, performing parallel channel attention calculation on two part feature maps to obtain a new channel weight, and performing weighted fusion on the channel weight and a corresponding initial part feature map; the two fused partial feature maps and the third initial partial feature map are re-spliced and compressed into a single-channel output feature map, and regularization parameters are dynamically adjusted in real time, so that the detection capability of the image feature extraction network on the sheltered target in the image is improved, the high-precision detection requirement for detecting the sheltered target in the image is met, and the detection efficiency is improved. The method is especially suitable for target detection application based on deep learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention discloses an image feature processing method based on a channel dynamic weighted self-reinforcement attention mechanism, which belongs to the field of visual detection image processing technology. Background Art

[0002] Currently, deep learning-based target detection methods are widely used in visual reasoning tasks such as surface defect detection of complex industrial products and detection of specific ground targets by drones. Due to the complex structure of complex industrial products or the changing perspective of drones, defects in the acquired images and specific ground targets waiting to be detected inevitably occlude each other, resulting in missed detection of occluded targets by deep learning-based target detection methods. Introducing an attention mechanism into deep learning-based target detection methods can guide the network to focus on key target areas, helping to improve the network's detection performance. However, the existing spatial attention mechanism still has shortcomings in processing the features of occluded targets. It has weak feature processing capabilities for occluded targets and cannot effectively capture the features of occluded targets, making it difficult to meet the requirements for high-precision detection of occluded targets in detection images. Summary of the Invention

[0003] The technical problem solved by the present invention is: in response to the problem that the existing self-attention method cannot effectively capture the features of occluded targets in the detection image, an image feature processing method based on channel dynamic weighted self-reinforcement attention mechanism is provided to improve the detection accuracy of occluded targets in the image by the target detection method based on deep learning.

[0004] The present invention is implemented by the following technical solutions:

[0005] The present invention first discloses an image feature processing method based on a channel dynamic weighted self-reinforcement attention mechanism. During the feature extraction network training process using the detection image, the channel weights in the image feature propagation process are dynamically adjusted. The adjustment process is as follows:

[0006] The input feature map of the detection image is channel-expanded and decomposed into three initial partial feature maps with the same dimensions. Two of the initial partial feature maps use parallel channel attention calculation to obtain new channel weights, and the channel weights are weightedly fused with the corresponding initial partial feature maps. The two fused partial feature maps are re-spliced ​​with the third initial partial feature map and compressed into a single-channel output feature map.

[0007] In the image feature processing method based on the channel dynamic weighted self-reinforcement attention mechanism of the present invention, further, the dimensions of the input feature map include batch size, channel, height and width.

[0008] In the image feature processing method based on the channel dynamic weighted self-reinforcement attention mechanism of the present invention, further, the input feature map uses a multi-layer perceptron to expand the feature channel C to 3C.

[0009] In the image feature processing method based on the channel dynamic weighted self-reinforcement attention mechanism of the present invention, further, the channel attention calculation process adopted by the partial feature map is as follows:

[0010] The two initial partial feature maps are multiplied by the weight vector normalized by the absolute value of the weight after the batch normalization layer, and then normalized by the Sigmoid activation function and multiplied by the corresponding initial partial feature map. The formula is as follows:

[0011] 、

[0012] ,

[0013] Where, X MLP_1 、X MLP_2 are two initial partial feature maps, and BN is the batch normalization layer; 、 The absolute value normalization result of the weight vector of the batch normalization layer corresponding to the two initial partial feature maps is obtained by the n normalized weights of the batch normalization layer. composition, , S is the total weight w in the batch normalization layer weight vector i The absolute value of X' MLP_1 、X' MLP_2 It is the partial feature map after two fusions.

[0014] In the image feature processing method based on the channel dynamic weighted self-reinforcement attention mechanism of the present invention, further, the feature channel after splicing the fused partial feature map and the third initial partial feature map is 3C, and the feature channel 3C is compressed to C through a multi-layer perceptron.

[0015] The present invention also discloses an image detection method, which places the above-mentioned image feature processing method of the present invention in the fourth stage of the feature extraction network of YOLOv5, uses YOLOv5 as a benchmark model to input a detection image for model training, and uses the trained YOLOv5 model for image detection.

[0016] The present invention also discloses a computer-readable storage medium storing a computer program, wherein the computer program is called by a processor to implement the above-mentioned image detection method of the present invention.

[0017] A visual inspection device for industrial product surface defects using the present invention is disclosed. The visual inspection device uses the image inspection method of the present invention to perform visual inspection of industrial product surface defects.

[0018] A visual inspection device for drone inspection using the present invention is provided. The visual inspection device is carried by a drone and uses the image detection method of the present invention to perform visual inspection on the appearance of equipment, including but not limited to power transmission and transformation equipment erected at high altitude.

[0019] When adjusting the channel weights in the process of image feature propagation, the present invention expands the input feature map into three channels and then decomposes it into three partial feature maps of the same dimension. Each partial feature map is an independent feature space block. The same channel attention calculation is used for the first partial feature map and the second partial feature map. New channel weights are obtained through two parallel channel attention calculations to achieve dynamic adjustment of image recognition regularization parameters, strengthen the feature extraction network's attention to occlusion-related semantic information in the image and suppress noise generated by occluded targets; then, the first partial feature map after attention calculation, the second partial feature map after attention calculation and the original third partial feature map are spliced ​​on the channel. On the one hand, the attention to occlusion-related semantic information in the image and the suppression of noise are strengthened, and on the other hand, the reuse of the original features in the image is maintained.

[0020] In the image detection method applied by the present invention, the loss function and back propagation are used to dynamically adjust the regularization parameters of the network for image feature recognition in real time during the training process of the image feature extraction network. A larger regularization force is applied to the channels with lower importance caused by occlusion, and a smaller regularization force is applied to the channels of the occluded related semantic features with higher importance. The channel weights of the image features that are focused on are increased, and the channel weights of the image features that are not to be focused on are reduced. The attention of the feature map channel weights is automatically adjusted during the training process of the image feature extraction network. Through this process, the feature extraction network is guided to re-attention to the feature channels of the occluded target when they fail, and at the same time, the noise that interferes with the occluded target, such as obstructions and complex backgrounds, is feature suppressed. The weight vector of the batch normalization layer is updated in real time during the back-propagation process of training. Even if the feature extraction network loses the semantic feature channel for detecting occluded objects in the image in the previous step of training, a smaller regularization force is applied in the next back-propagation to make the model pay attention to it again, and a larger regularization force is applied to make the model not pay attention to irrelevant features such as occluded objects and complex backgrounds, so that the value of the expected loss function of the training process is minimized. The various parameters of the network are dynamically adjusted during the back-propagation process, and the image detection training model is optimized.

[0021] In summary, the image feature processing method based on the channel dynamic weighted self-reinforcement attention mechanism proposed in the present invention improves the detection ability of the image feature extraction network for occluded targets in the image by dynamically adjusting the regularization parameters in real time, and meets the high-precision detection requirements for occluded targets in the detection image.

[0022] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 This is a flow chart of the image feature processing method based on the channel dynamic weighted self-reinforcement attention mechanism of the present invention. DETAILED DESCRIPTION

[0024] Example

[0025] This embodiment is one of the visual inspection tasks applied by the present invention. It detects substation equipment from infrared images acquired by drones. The detected infrared images of substation equipment include seven types of substation equipment: lightning arrester 1, lightning arrester 2, current transformer 1, current transformer 2, voltage transformer, disconnector, and support porcelain bottle. The image feature detection training is performed using the YOLOv5 model. The image feature processing method based on the channel dynamic weighted self-reinforcement attention mechanism of the present invention is set in the fourth stage of the feature extraction network of YOLOv5. The input feature map X extracted in this stage is calculated. , where B represents the batch size of the input feature map, C represents the channel of the input feature map, H represents the height of the input feature map, and W represents the width of the input feature map.

[0026] In this embodiment, during the feature extraction network training process for the detection image, the channel weights in the image feature propagation process are dynamically adjusted, and the input feature map of the detection image is decomposed into three initial partial feature maps with the same dimensions after channel expansion. Two of the initial partial feature maps use parallel channel attention calculations to obtain new channel weights, and the channel weights are weightedly fused with the corresponding initial partial feature maps. The two fused partial feature maps are re-spliced ​​with the third initial partial feature map and compressed into a single-channel output feature map.

[0027] The specific process of image feature processing method is as follows Figure 1 As shown, it can be summarized into the following steps:

[0028] S100, input feature map Use the multi-layer perceptron MLP to expand its feature channel to 3C to obtain the expanded feature map X MLP , ;

[0029] S200, from the channel dimension Decomposed into three initial partial feature maps with the same batch size, channel, height and width, denoted as X MLP_1 , X MLP_2 and X MLP_3 ;

[0030] S300, the first initial partial feature map X MLP_1 After the batch normalization layer Batch Normalization and the weight vector normalized by the absolute value of the weight are multiplied, the formula is as follows:

[0031] .

[0032] Where, X MLP_11 is the first initial partial feature map X MLP_1 After the updated channel weights, BN is the batch normalization layer Batch Normalization, For X MLP_1 The absolute value normalization result of the weight vector of the batch normalization layer, By n normalized weights composition, , S is the batch normalization layer weight vector w BN1 All the values ​​in the i The sum of the absolute values ​​of , w BN1 By n weights w i Composition, ʘ is multiplication.

[0033] For the second initial partial feature map X MLP_2 After the batch normalization layer Batch Normalization and the weight vector normalized by the absolute value of the weight are multiplied, the formula is as follows:

[0034] .

[0035] Where, X MLP_21 is the second initial partial feature map X MLP_2 After the updated channel weights, BN is the batch normalization layer Batch Normalization, For X MLP_2 The absolute value normalization result of the weight vector of the batch normalization layer, By n normalized weights composition, , S is the batch normalization layer weight vector w BN2 All the values ​​in the i The sum of the absolute values ​​of , w BN2 By n weights w i Composition, ʘ is multiplication.

[0036] S400, X MLP_11 After the Sigmoid activation function and the first initial part feature map X MLP_1 Multiply, and the normalized channel weights are reapplied to X MLP_1 ,

[0037] .

[0038] X MLP_21 Through the Sigmoid activation function, and the second initial part feature map X MLP_2 Multiply, and the normalized channel weights are reapplied to X MLP_2 ,

[0039] .

[0040] Where, is the Sigmoid activation function, X' MLP_1 is the first part of the feature map after fusion channel weights, X' MLP_2 It is the second part of the feature map after fusing the channel weights.

[0041] S500, X' MLP_1 、X' MLP_2 and the third initial part feature map X MLP_3 Splicing, and then using the multi-layer perceptron MLP to compress the spliced ​​feature map into the output feature map Y of the single channel C,

[0042] .

[0043] Where F3C-C MLP is a multi-layer perceptron MLP that reduces the feature map channel 3C to C, and ⊕ represents the channel concatenation operation.

[0044] In the visual inspection task of detecting substation equipment from infrared images obtained by drones in this embodiment, the above-mentioned image feature processing is performed on the detected infrared images of the substation equipment, and the weight vectors of the batch normalization layer are updated in real time during the back propagation process of the feature extraction network training. Even if the semantic feature channel of the occluded object in the image is lost in the previous step of the feature extraction network training, applying a smaller regularization force in the next back propagation can make the feature extraction network model pay attention again. At the same time, the first part of the feature map after attention calculation, the second part of the feature map after attention calculation, and the original third part of the feature map channel are spliced. On the one hand, the attention to the occlusion-related semantic information and the suppression of noise in the image are strengthened, and on the other hand, the reuse of the original features of the graphics is maintained.

[0045] In some embodiments, a computer-readable storage medium based on the above-described image detection method is also provided, storing a computer program that is invoked by a processor to implement the YOLOv5 image detection method described above in this embodiment. YOLOv5 is a mature deep learning-based object detection model, and this embodiment does not describe the specific training process and data processing of YOLOv5.

[0046] The readable storage medium is a computer-readable storage medium, which can be an internal storage unit of the software and hardware device described in any of the aforementioned embodiments, such as a hard disk or memory of a controller. The readable storage medium can also be an external storage device of the controller, such as a plug-in hard disk equipped on the controller, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. Furthermore, the readable storage medium can also include both an internal storage unit of the controller and an external storage device. The readable storage medium is used to store the computer program and other programs and data required by the controller. The readable storage medium can also be used to temporarily store data that has been output or is to be output.

[0047] Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes instructions for causing a computer device (such as a personal computer, server, or network device) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned readable storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0048] To verify that this embodiment can effectively capture obscured targets in infrared images of substation equipment acquired by drones, the following comparative experiments compare the output feature maps processed by this embodiment's feature processing method with several mainstream attention modules in the prior art, including SimAM, SE, ECA, and CBAM. The experimental results are shown in Table 1. Table 1 shows that processing image features using this embodiment improves the detection performance of the conventional YOLOv5 model by +1.47%, achieving the highest improvement in YOLOv5 detection performance among all attention mechanisms. Furthermore, the number of parameters (Params) and computational complexity (FLOPs) of this embodiment are comparable to those of other attention mechanisms.

[0049] Table 1 Comparison between this embodiment and other attention methods.

[0050] .

[0051] The application of the improved YOLOv5 model based on this embodiment in visual inspection equipment with visual occlusion can also be used in visual inspection equipment for surface defects of industrial products, effectively improving the accuracy of visual inspection of surface quality image features of industrial products with occlusion.

[0052] In this document, the directions or positional relationships indicated by terms such as "up", "down", "front", "back", "left", "right", "top", "bottom", "inside", "outside", "vertical", and "horizontal" are based on the directions or positional relationships shown in the accompanying drawings and are only for the clarity of the technical solution and the convenience of description, and therefore should not be understood as limiting the present invention.

[0053] As used herein, the terms "comprises," "comprising," or any other variation thereof, are intended to cover a non-exclusive inclusion of elements other than the listed elements and may also include additional elements not specifically listed.

[0054] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. An image feature processing method based on a channel-dynamic weighted self-reinforcement attention mechanism, characterized by: During the feature extraction network training using the detection image, the channel weights in the image feature propagation process are dynamically adjusted. The adjustment process is as follows: The input feature map of the detection image is channel-expanded and decomposed into three initial partial feature maps with the same dimensions. Two of the initial partial feature maps use parallel channel attention calculation to obtain new channel weights, and the channel weights are weightedly fused with the corresponding initial partial feature maps. The two fused partial feature maps are re-spliced ​​with the third initial partial feature map and compressed into a single-channel output feature map.

2. The image feature processing method based on the channel dynamic weighted self-reinforcement attention mechanism according to claim 1 is characterized by: The dimensions of the input feature map include batch size, channels, height, and width.

3. The image feature processing method based on the channel dynamic weighted self-reinforcement attention mechanism according to claim 1 is characterized by: The input feature map uses a multi-layer perceptron to expand the feature channel C to 3C.

4. The image feature processing method based on the channel dynamic weighted self-reinforcement attention mechanism according to claim 1 is characterized by: The channel attention calculation process used in the above-mentioned feature maps is as follows: The two initial partial feature maps are multiplied by the weight vector normalized by the absolute value of the weight after the batch normalization layer, and then normalized by the Sigmoid activation function and multiplied by the corresponding initial partial feature map. The formula is as follows: 、 , Where, X MLP_1 、X MLP_2 are two initial partial feature maps, and BN is the batch normalization layer; 、 The absolute value normalization result of the weight vector of the batch normalization layer corresponding to the two initial partial feature maps is obtained by the n normalized weights of the batch normalization layer. composition, , S is the total weight w in the batch normalization layer weight vector i The absolute value of X' MLP_1 、X' MLP_2 It is the partial feature map after two fusions.

5. The image feature processing method based on the channel dynamic weighted self-reinforcement attention mechanism according to claim 1 is characterized by: The feature channel after the fusion of the partial feature map and the third initial partial feature map is 3C, and the feature channel 3C is compressed to C through the multi-layer perceptron.

6. An image detection method, characterized in that: The image feature processing method according to any one of claims 1 to 5 is set in the fourth stage of the feature extraction network of YOLOv5, YOLOv5 is used as a benchmark model to input the detection image for model training, and the trained YOLOv5 model is used for image detection.

7. A computer-readable storage medium based on the image detection method according to claim 6, characterized in that: A computer program is stored, and the computer program is called by a processor to implement the image detection method according to claim 6.

8. Visual inspection equipment for industrial product surface defects, characterized by: The visual inspection device uses the image inspection method of claim 6.

9. UAV inspection visual inspection equipment, characterized by: The visual inspection device uses the image inspection method of claim 6.

Citation Information

Cited By

  • Image multi-target detection system and method based on channel attention mechanism

    CN121053377A