A multi-scale object detection method based on feature enhancement and pixel inversion dehazing

By introducing feature enhancement modules and multi-scale object detection models in underwater target video detection, the problem of low detection accuracy of small targets in complex backgrounds is solved, and higher detection accuracy and adaptability are achieved.

CN119359696BActive Publication Date: 2025-05-13AUTOLINK INFORMATION TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411839589.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-05-13
Estimated Expiration
2044-12-13

AI Technical Summary

Technical Problem

The existing underwater target video detection method has low detection accuracy when the target object is too small under complex background conditions.

Method used

Using a multi-scale object detection method based on feature enhancement and pixel inversion defog removal, the network's adaptability to the size changes of the target object is enhanced by building feature enhancement modules and multi-scale object detection models.

Benefits of technology

It effectively improves the detection accuracy of small target objects, reduces the loss of context information of feature maps in deep networks, and enhances the accuracy of biometric identification at different scales underwater.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119359696B_ABST
    Figure CN119359696B_ABST
Patent Text Reader

Abstract

The present application provides a multi-scale target detection method based on feature enhancement and pixel inversion dehazing, which introduces the SE attention mechanism in the specified convolutional layer and combines it with the enhanced feature extraction module to construct a feature enhancement module, solves the problem of channel attention in the multi-scale feature fusion process, reduces the context information loss of the feature map in the deep network, and effectively improves the accuracy of underwater biological recognition at different scales; a backbone network of the multi-scale target detection model is constructed based on VGG16 as the basic network, and the feature enhancement module is embedded in the convolutional layer 4, the second fully connected FC layer and the convolutional layer 6 respectively, and at the same time, the convolutional layer 4, the second fully connected FC layer and the convolutional layer 6 of the backbone network are respectively subjected to multi-scale enhancement by the feature enhancement module, and then stacked and connected for feature extraction, thereby realizing the feature fusion operation of different convolutional layers, which can enable more shallow detail information to be transmitted to the deep layer, thereby extracting richer information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image recognition technology, and in particular to a multi-scale target detection method based on feature enhancement and pixel inversion defogging. Background Art

[0002] When developing intelligent vehicles that can wade through water, it is an essential research and development direction to equip the vehicle with various underwater target detection functions based on on-board image acquisition equipment. Our company is currently using a patent with application number CN202411020787.8, which discloses a lightweight underwater target video detection method. A lightweight feature extraction backbone network is built using an anti-residual hole convolution module. After deep convolution learning of features and classification and positioning regression of the convolution network, an SSH detection head is used to obtain multi-channel detection results for fusion detection, which can achieve a good balance between underwater target detection speed and detection accuracy. However, in actual applications, it was found that the existing technical solution has low detection accuracy when detecting underwater targets in complex backgrounds and when the detected target is too small. Summary of the invention

[0003] In order to solve the problem of low detection accuracy in underwater target video detection methods in the prior art when the detection target is too small under complex background conditions, the present application provides a multi-scale target detection method based on feature enhancement and pixel inversion dehazing, which can enhance the network's adaptability to changes in target size and effectively improve the detection accuracy of small targets.

[0004] The technical solution of the present application is as follows: a multi-scale target detection method based on feature enhancement and pixel inversion defogging, characterized in that it includes the following steps:

[0005] S1: build feature enhancement module;

[0006] Embedding the SE attention mechanism module into the specified convolutional layer to construct the feature enhancement module;

[0007] S2: Build a multi-scale object detection model;

[0008] The multi-scale target detection model includes: a pre-processing module, a backbone network and an output layer connected in sequence;

[0009] The backbone network uses VGG16 as the basic network and integrates the feature enhancement module SFEM, which includes: sequentially connected convolutional layers 1 to 5, the first fully connected FC layer, the second fully connected FC layer and convolutional layers 6 to 9;

[0010] A prediction module of different sizes is set after convolutional layer 4, the second fully connected FC layer, and convolutional layers 6 to 9 respectively;

[0011] Embedding one feature enhancement module in the convolutional layer 4, the second fully connected FC layer and the convolutional layer 6 respectively;

[0012] The output of convolution layer 4 is processed by the feature enhancement module and then sent to the prediction module corresponding to this layer;

[0013] The outputs of the convolution layer 4 and the second fully connected FC layer are respectively subjected to feature enhancement processing and a convolution operation, and then subjected to feature enhancement processing by the feature enhancement module, and the output result is recorded as: secondary enhanced feature; the secondary enhanced feature is sent to the prediction module corresponding to the second fully connected FC layer;

[0014] The output of the convolution layer 6 is processed by the feature enhancement module, and then subjected to a convolution operation with the secondary enhancement feature, and then processed by the feature enhancement module, and then sent to the prediction module corresponding to the convolution layer 6;

[0015] All feature maps output by all prediction modules are superimposed to obtain the final output result;

[0016] S3: constructing a training data set and a verification data set based on historical data, and training the multi-scale object detection model based on the training data set to obtain a trained multi-scale object detection model;

[0017] S4: Identify the image to be identified based on the trained multi-scale object detection model.

[0018] It is further characterized by:

[0019] The feature enhancement module includes: a SE attention mechanism module, an enhanced feature extraction module and an add operation connected in sequence;

[0020] Adding the SE attention mechanism module before the convolution operation of the specified channel of the convolutional network layer to be processed, the SE attention mechanism module extracts the attention weight of the output feature map of the specified channel;

[0021] The attention weight of the convolutional layer channel extracted by the SE attention mechanism module is multiplied by the feature map output by the channel through the Scale operation, and then sent to the enhanced feature extraction module for enhanced feature extraction operation;

[0022] Finally, the output feature map of the enhanced feature extraction module is connected to the channel input feature map of the next layer by using the add method, and the channel of the final output feature map is adjusted to the size of the channel of the next layer;

[0023] The enhanced feature extraction module includes: a convolution operation with a convolution kernel of 1*1, a convolution operation with a convolution kernel of 3*3 and a step length of 2, and a convolution operation with a convolution kernel of 1*1, which are sequentially set;

[0024] The prediction module includes a detector Detector and a classifier Classifler;

[0025] The preprocessing module includes: image inversion operation, dark channel calculation, atmospheric light intensity estimation operation, transmittance estimation, transmission optimization and image inversion operation connected in sequence;

[0026] The number of channels of the convolutional layers 1 to 5 are set to 64, 128, 256, 512, and 512 respectively;

[0027] The number of channels corresponding to the convolutional layers 6 to 9 are set to 512, 256, 256, and 256 respectively;

[0028] The corresponding scales of the prediction modules are: 38*38, 19*19, 10*10, 5*5, 3*3, and 1*1.

[0029] The present application provides a multi-scale target detection method based on feature enhancement and pixel inversion defogging, which introduces the SE attention mechanism in the specified convolution layer and combines it with the enhanced feature extraction module to construct a feature enhancement module, solves the problem of channel attention in the multi-scale feature fusion process, reduces the context information loss of the feature map in the deep network, and effectively improves the accuracy of underwater biological recognition at different scales; based on VGG16 as the basic network, a backbone network of the multi-scale target detection model is constructed, and the feature enhancement module is embedded in the convolution layer 4, the second fully connected FC layer and the convolution layer 6 respectively, and at the same time, the convolution layer 4, the second fully connected FC layer and the convolution layer 6 of the backbone network are respectively subjected to multi-scale enhancement by the feature enhancement module, and then stacked and connected, and feature extraction is performed, so that the feature fusion operation of different convolution layers is realized, so that more shallow detail information can be transmitted to the deep layer, thereby extracting richer information; the present application adopts a feature fusion strategy and integrates the feature enhancement module to fully extract features of different levels and scales, reduce information loss in the feature propagation process, and improve the performance level of the network. This method sets up 6 prediction modules and extracts feature maps of six different scales to adapt to the detection needs of objects of different sizes. This strategy greatly enhances the network's adaptability to scale changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is the structural diagram of SFEM module;

[0031] Figure 2 It is a schematic diagram of the structure of the multi-scale target detection model;

[0032] Figure 3 Output feature maps for SFEM-SSD and SSD to visualize the result graphs;

[0033] Figure 4 This is the loss function diagram of the SSD target detection algorithm;

[0034] Figure 5 It is the loss function diagram of SFEM-SSD network algorithm;

[0035] Figure 6 Comparison chart of SSD training loss and multi-scale feature fusion target detection network training loss;

[0036] Figure 7 Comparison chart of SSD verification loss and multi-scale feature fusion target detection network verification loss;

[0037] Figure 8 Comparison of underwater fish detection results using different target detection networks. DETAILED DESCRIPTION

[0038] The present application includes a multi-scale target detection method based on feature enhancement and pixel inversion defogging, which includes the following steps.

[0039] S1: Construct feature enhancement module.

[0040] Considering that the small target objects in the image to be identified are small in size, account for a small proportion of pixels, and have poor anti-interference ability, a large amount of feature information will be lost after multiple convolutions and pooling, resulting in a decrease in the accuracy of small target detection. In order to solve this problem, this application embeds the SE attention mechanism module into the network to improve the network's anti-interference ability and the accuracy of small target detection. This module is named the feature enhancement module (SE-Feature-Enhancement module, hereinafter referred to as the SFEM module). The network structure diagram of the feature enhancement module is shown below. Figure 1 shown.

[0041] The feature enhancement module includes: a SE attention mechanism module, an enhanced feature extraction module and an add operation connected in sequence; the enhanced feature extraction module includes: a convolution operation with a convolution kernel of 1*1, a convolution operation with a convolution kernel of 3*3 and a step size of 2, and a convolution operation with a convolution kernel of 1*1.

[0042] The ‌SE attention mechanism (Squeeze-and-Excitation mechanism) consists of two main steps: Squeeze and Excitation‌. In the Squeeze step, the input feature map is compressed into a vector through a global average pooling operation, which can capture the global statistical information of each channel. Then, in the Excitation step, a fully connected layer and a sigmoid function are used to generate the weights of each channel, and these weights are multiplied with the original input feature map to obtain the weighted feature map. In this way, the model can adaptively learn the importance of each channel and adjust it.

[0043] The SFEM module is a top-down structure that consists of three parts. First, the SE attention mechanism module is added before the first convolution operation of the specified channel to explicitly express the interdependence between channels and to adaptively readjust the feature response of the channel. The SE attention mechanism module assigns weights to each channel so that multiple channels can have an effect on the result. These weights represent the influence of each channel on feature extraction. When the weight is larger, the value of the feature map of the channel will be smaller, and the influence on the final output will be smaller. This means that when extracting image features, some feature maps output by the convolution layer have a greater impact on the final result, while some feature maps output by the convolution layer have a smaller impact on the final result. Therefore, by using the weights obtained by the channel itself and applying them to these feature maps, the channel weights can be adaptively given based on the features extracted by the convolution layer, so that the feature maps that have a greater impact on the final result have a greater influence.

[0044] Secondly, after the SE attention mechanism module processes the feature map of the specified channel for the first convolution operation, the feature extraction module is further enhanced. First, the output feature map of the previous layer is convolved with a convolution kernel of 1*1, and the number of channels is adjusted. Then, the convolution kernel is 3*3 and the step length is 2. The width and height of the feature map are adjusted, and the amount of calculation is simplified. It is further input into the next convolution layer with a convolution kernel of 1*1 and a convolution layer with a convolution kernel of 3*3 for enhanced feature extraction. Finally, the convolution layer with a convolution kernel of 1*1 is adjusted for the number of channels. Through the downsampling of the convolution operation, the deep features can have a larger receptive field, so as to better capture the important features in the image. Secondly, compared with the weighted sum operation of the pooling layer, the point-by-point operation of the 1*1 convolution operation is more conducive to optimization and solution. Therefore, the pooling layer after the first convolution layer with a 1*1 convolution kernel and the convolution layer with a 3*3 convolution kernel in the original sense is cancelled and replaced with a convolution layer with a 1*1 convolution kernel, thereby reducing the risk of effective information being lost, improving information fusion, and strengthening feature extraction.

[0045] Finally, the add method is used to connect the feature map output by the enhanced convolutional network layer with the input feature map of the next layer channel, and adjust the channel to the corresponding size of the next layer channel. Compared with the concat method, the add method superimposes the extracted information multiple times, highlighting the proportion of correct classification, which is beneficial to the final target classification and achieves high activation of the correct classification.

[0046] Specifically, after adding the feature enhancement module to the specified convolutional layer, the following operations are performed for every two adjacent channels of the convolutional network layer to be processed:

[0047] For example, a convolutional layer includes 3 channels, then channels 1 and 2, and channels 2 and 3 are adjacent channels. Figure 1 In the figure, FM1 (H*W*C) and FM2 (H'*W'*C') are two adjacent channels. The SE attention mechanism module is set before the first convolution of the FM1 channel. After extracting the attention weight for the channel where FM1 is located, the attention weight extracted by the SE attention mechanism module is multiplied by the output feature map of the FM1 channel through the Scale operation, and then sent to the enhanced feature extraction module for enhanced feature extraction operation; finally, the output feature map of the enhanced feature extraction module is connected with the input feature map of the FM2 channel by the add method, and the channel of the connected feature map is adjusted to the corresponding size (H'*W'*C') of the feature extraction convolutional network layer of the next layer, and the feature map New_FM_2 is obtained and sent to the channel where FM2 is located.

[0048] S2: Build a multi-scale object detection model.

[0049] like Figure 2 As shown, the multi-scale target detection model includes: a preprocessing module, a backbone network and an output layer connected in sequence.

[0050] The input part of the network is responsible for receiving the original image data and preprocessing it to improve the quality of the input data and improve the efficiency of subsequent network learning. The preprocessing module is implemented based on the double inversion defog module, which specifically includes: sequentially connected image inversion operation, dark channel calculation, atmospheric light intensity estimation operation, transmittance estimation, transmission optimization and image inversion operation. The input image is first reversed (Reverse Image) and then the dark channel is calculated (Calculated dark channels), that is, the darkest value of the RGB three channels is calculated, the atmospheric light intensity corresponding to the feature map is estimated (Estimation of atmospheric light), the transmittance corresponding to the collected image is estimated (transmittance estimation), and after calculating the transmission optimization (transmission optimization), the feature map obtained is the enhanced image to be identified, and finally the image is reversed (Reverse Image) again. The feature map obtained is the preprocessed band-identification image.

[0051] The backbone network uses VGG16 as the basic network and integrates the feature enhancement module SFEM, which includes: sequentially connected convolutional layer 1 (Conv1) ~ convolutional layer 5 (Conv5), the first fully connected FC layer (fc6), the second fully connected FC layer (fc7) and convolutional layer 6 (Conv6) ~ convolutional layer 9 (Conv9).

[0052] The number of channels of convolutional layer 1 to convolutional layer 5 is set to 64, 128, 256, 512, and 512 respectively. When extracting features, increasing the number of channels along with the convolution operation can obtain more image feature information.

[0053] The number of channels corresponding to convolutional layers 6 to 9 is set to 512, 256, 256, and 256, respectively; although increasing the number of channels can enhance the expressiveness of features, it may also lead to a significant increase in the amount of calculation and the number of parameters. In order to solve this problem, the present application reduces the number of channels in the subsequent layers of the network, that is, reduces the depth of the feature map. The main purpose of reducing the number of channels is to reduce the computational complexity and storage requirements of the network, thereby improving the efficiency of the network.

[0054] A prediction module is set after convolution layer 4, the second fully connected FC layer, and convolution layers 6 to 9. The prediction module includes a detector and a classifier; all feature maps output by all prediction modules are superimposed to obtain the final output result.

[0055] There are 6 prediction modules in the multi-scale object detection model, which extract feature maps of six different scales to meet the detection needs of objects of different sizes. The corresponding output feature map sizes are: 38*38, 19*19, 10*10, 5*5, 3*3, and 1*1.

[0056] Feature enhancement modules are embedded in convolutional layer 4, the second fully connected FC layer, and convolutional layer 6. In this method, features at different levels are used to capture multi-scale information. Convolutional layer 4 is used as a shallow feature to better capture small objects; the output of convolutional layer 6 and the output of FC7 are used as deep features to help process large objects and complex backgrounds. By fusing these features, the network can focus on small objects and large objects at the same time, avoiding missed detection or false detection.

[0057] The outputs of convolution layer 4 and the second fully connected FC layer are subjected to feature enhancement processing after a convolution operation, and then to feature enhancement processing by the feature enhancement module. The output result is recorded as: secondary enhanced feature; the secondary enhanced feature is sent to the prediction module corresponding to the second fully connected FC layer;

[0058] The output of convolution layer 6 is processed by the feature enhancement module, and then it is processed by the feature enhancement module after a convolution operation with the secondary enhancement feature, and then sent to the prediction module corresponding to convolution layer 6;

[0059] After the convolution layer 4, the second fully connected FC layer and the convolution layer 6 are respectively subjected to multi-scale enhancement by the feature enhancement module, the enhanced features are again extracted by SFEM. The high-level semantic information is integrated with the low-level detail information through the multi-layer SFEM module, so that more shallow detail information can be transmitted to the deep layer.

[0060] In the backbone network, convolution layer 4, the second fully connected FC layer, and convolution layer 6 each include a 3*3 convolution kernel; the third convolution result of convolution layer 4, the convolution result of the fc7 layer, and the second convolution result of convolution layer 6 are stacked and connected after multi-scale feature enhancement and feature extraction. This method uses features at different levels to capture multi-scale information, and fuses high-level semantic information with low-level detail information by transferring multiple layers of features from top to bottom. Rich feature information is retained through stacking connections, and abstract high-level features are further extracted through subsequent convolution layers; by fusing these features, the network can focus on small objects and large objects at the same time to avoid missed detection or false detection. Without the need to relearn redundant features, this feature fusion strategy can obtain a large amount of feature information through fewer convolutions, thereby greatly optimizing the information features in the neural network and enhancing the expression ability of the network model.

[0061] The backbone network is based on VGG16 as the basic network, which is used to extract basic feature information from the input image. However, using only the final output of VGG16 as a feature source cannot fully utilize the rich information in the image. Therefore, a key improvement is to stack and connect the multi-layer output results of VGG16. At the same time, the enhanced feature extraction network receives the output of the backbone feature extraction network as input, which not only retains the shallow detail features, but also integrates the deep abstract features, greatly enriching the feature expression ability. The design purpose of this part is to further enhance the expressiveness and diversity of features. By reprocessing and fusing features at different levels, richer and more powerful feature representations can be generated, providing more accurate information for subsequent target detection. In particular, this study extracts feature maps of six different scales to meet the detection needs of targets of different sizes. This strategy greatly enhances the network's adaptability to scale changes. Finally, the results of convolutional neural network feature extraction are localized and classified and regressed to achieve the purpose of underwater target detection.

[0062] In this application, a feature fusion strategy is added to the backbone feature extraction network and combined with a feature enhancement module so that the neural network can focus on multi-dimensional feature information at different scales during the feature extraction stage to enhance the detection of underwater targets.

[0063] S3: Build a training data set and a validation data set based on historical data, train the multi-scale object detection model based on the training data set, and obtain a trained multi-scale object detection model;

[0064] S4: Identify the image to be identified based on the trained multi-scale object detection model.

[0065] In order to confirm the performance of this method, a simulation experiment was constructed. The test platform configuration is: Windows 11 operating system CUDA10.1, CUDNN 7.6.5 Python 3.7 version, processor: AMD Ryzen 7 5800H with RadeonGraphics 3.20 GHz. The graphics card model is NVIDIA GeForce RTX 3060, and the tensorflow framework is used.

[0066] In order to ensure the fairness of the experiment, the same initial training parameters are set for each group of experiments. In this experiment, the target detection network uses pre-trained weights to pre-train the backbone network to shorten the training time. Batchsize is set to 16, Epoch is set to 50, and the learning rate is 2×10−3. During the unfreezing stage, the parameters of the entire network will be adjusted, Batchsize is set to 8, Epoch is set to 100, and the learning rate is 2×10−4. This chapter uses the SGD method to adjust the loss function, and the weight decay coefficient is set to 5×10−4. During the test, the confidence is set to 0.5, and the nms iou size used for non-maximum suppression is set to 0.2.

[0067] The original SSD (Single Shot Multibox Detector) model is constructed based on VGG-16. The multi-scale target detection model in this method is introduced: SFEM-SSD model.

[0068] Figure 3 The left column in the figure is the feature map A, C, E of the original SSD output, and the right column is the feature map B, D, F of the SFEM-SSD model. From the three figures A, C, and E on the left, it can be seen that: as the network depth increases, although the information obtained is more fine-grained, this high-grained information is mainly limited to local areas, and the ability to capture global information is limited. This limitation is particularly obvious when dealing with complex scenes or small target detection, because in these cases the network needs to be able to comprehensively consider global and local information to achieve more accurate target detection. From the three figures B, D, and F on the right, it can be seen that by introducing feature enhancement and attention mechanisms, the target detection network combined with multi-scale feature fusion can effectively improve the quality of feature maps. By fusing feature maps of different depth levels, the high semantic information of the deep network is retained, and the high-resolution detail information of the shallow network is combined, thereby achieving a more fine-grained feature expression. In particular, after the introduction of the attention mechanism, the network can adaptively focus on more important feature areas, further enhancing the model's ability to detect targets.

[0069] At the same time, it can be seen from B, D, and F that the combination of multi-scale feature fusion and attention mechanism makes the feature information of small targets more significant, thereby improving the ability of the target detection network to detect small targets. Small targets have more obvious contours and other details in the shallow feature map. By fusing with deep features, it can not only enhance the expression of these details, but also enrich the overall semantic information, so that the network can accurately identify and locate small targets even in complex scenes.

[0070] The loss function is used to measure the difference or error between the model prediction results and the actual observed values. The loss function is usually a non-negative real number. The smaller the value, the closer the model's prediction results are to the actual values. Conversely, the larger the value, the greater the difference between the prediction results and the actual values.

[0071] Figure 4 This is the loss function diagram of the SSD target detection algorithm. Figure 5 This is the loss function diagram of the SFEM-SSD network algorithm. The loss functions used in the figure are: smoothL1, train loss represents the training loss of smoothL1, valloss represents the test loss of smoothL1, smooth train loss represents the curve of train loss quadratic fitting, and smoothval loss represents the curve of test loss val loss quadratic fitting. In the figure, the horizontal axis loss represents the loss value, and the vertical axis Epoch represents the number of training rounds.

[0072] from Figure 4 and Figure 5 It can be seen that due to the addition of feature fusion strategy, SFEM-SSD can extract features at different scales. The model can take into account the details and overall information of the input image at the same time, so as to better capture the diverse features of the target and make the loss function drop faster. At the same time, through the SFEM module, the original features can be enhanced to make the feature information more distinguishable and representative, which helps the model to better learn the key features of the target and thus better fit.

[0073] from Figure 6 It is a comparison between the SSD training loss and the multi-scale feature fusion network SFEM-SSD training loss in this application. Figure 7 This is a comparison of the SSD test loss and the SFEM-SSD test loss. It can be seen that compared with the SSD target detection algorithm, the training loss and verification loss values ​​of the multi-scale feature fusion target detection network are lower, proving that the multi-scale feature fusion target detection network can better detect targets, and the improved model is more robust.

[0074] Figure 8The three images A, B, and C are real underwater images, among which A is an image of a school of fish with small underwater targets, B is an image of a school of fish with mixed underwater targets of different sizes, and C is an image of a school of fish with complex underwater backgrounds. After detection by SSD (VGG), SSD (Mobilenetv2), YOLOv3, and the SFEM-SSD target detection algorithm of this method, A1, A2, and A3 all have problems of missed detection and false detection of small targets, while in the case of mixed large and small targets, B1, B2, and B3 have serious problems of missed detection and false detection of small targets. When the background is complex, C1, C2, and C3 can see that the detection results of SSD (VGG) and YOLOv3 are similar, but they have the same missed detection problem as SSD (Mobilenetv2). By observing the SFEM-SSD detection results of A4, B4, and C4, it can be seen that SFEM-SSD has more advantages in small target detection and target detection under complex backgrounds.

[0075] The comparison of evaluation indicators is shown in Table 1 below.

[0076] Table 1 Comparison of accuracy, recall, and mAP between SFEM-SSD target detection network and different target detection networks

[0077]

[0078] After experimental comparison, it was found that the multi-scale feature fusion deep learning target detection network model using SFEM module and feature fusion strategy can more accurately identify targets in complex underwater scenes. Compared with the original SSD target detection network, the network has multiple detection errors and missed detections in complex scenes. The deep learning target detection network proposed in this paper has improved the accuracy by 2.27%, 3.13% and 4.99% respectively compared with SSD (VGG), SSD (Mobilenetv2) and YOLOv3 target detection network. By comparing the accuracy, recall and mAP index of different algorithms, it can be found that SFEM-SSD basically meets the requirements of underwater target detection.

[0079] After using the technical solution of the present application, the VGG16 network structure is used as the basic network to realize the initial feature extraction. In view of the fact that the features of the neural network at the shallow level are more sensitive to smaller objects, while the features at the deep level contain better semantic information, this method constructs the SFEM module based on the SE attention mechanism and combines it with the enhanced feature extraction module to improve the extraction of features, solves the problem of channel attention in the multi-scale feature fusion process, reduces the context information loss of the feature map in the deep network, and effectively improves the accuracy of underwater biological recognition at different scales. In order to fully retain the detailed features in the original underwater picture, this method adopts a feature fusion strategy and integrates the SFEM module to fully extract features of different levels and scales, reduces the information loss in the feature propagation process, and improves the performance level of the network. Experimental results show that compared with other underwater target detection methods, the multi-scale target detection model based on feature enhancement and pixel inversion dehazing algorithm proposed in this patent has achieved better average detection accuracy on the underwater target detection dataset.

Claims

1. A multi-scale target detection method based on feature enhancement and pixel inversion defogging, characterized in that: The following steps are involved: S1: Construct feature enhancement module; Embedding the SE attention mechanism module into the specified convolutional layer to construct the feature enhancement module; S2: Build a multi-scale object detection model; The multi-scale target detection model includes: a pre-processing module, a backbone network and an output layer connected in sequence; The backbone network uses VGG16 as the basic network and integrates the feature enhancement module SFEM, which includes: sequentially connected convolutional layers 1 to 5, the first fully connected FC layer, the second fully connected FC layer and convolutional layers 6 to 9; A prediction module of different sizes is set after convolutional layer 4, the second fully connected FC layer, and convolutional layers 6 to 9 respectively; Embedding one feature enhancement module in the convolutional layer 4, the second fully connected FC layer and the convolutional layer 6 respectively; The output of convolution layer 4 is processed by the feature enhancement module and then sent to the prediction module corresponding to this layer; The outputs of the convolution layer 4 and the second fully connected FC layer are respectively subjected to feature enhancement processing and a convolution operation, and then subjected to feature enhancement processing by the feature enhancement module, and the output result is recorded as: secondary enhanced feature; the secondary enhanced feature is sent to the prediction module corresponding to the second fully connected FC layer; The output of the convolution layer 6 is subjected to feature enhancement processing by the feature enhancement module, and then subjected to a convolution operation with the secondary enhancement feature, and then subjected to feature enhancement processing by the feature enhancement module, and then sent to the prediction module corresponding to the convolution layer 6; All feature maps output by all prediction modules are superimposed to obtain the final output result; S3: constructing a training data set and a verification data set based on historical data, and training the multi-scale object detection model based on the training data set to obtain a trained multi-scale object detection model; S4: Identify the image to be identified based on the trained multi-scale object detection model.

2. According to claim 1, a multi-scale target detection method based on feature enhancement and pixel inversion defogging is characterized by: The feature enhancement module includes: a SE attention mechanism module, an enhanced feature extraction module and an add operation connected in sequence; Adding the SE attention mechanism module before the convolution operation of the specified channel of the convolutional network layer to be processed, the SE attention mechanism module extracts the attention weight of the output feature map of the specified channel; The attention weight of the convolutional layer channel extracted by the SE attention mechanism module is multiplied by the feature map output by the channel through the Scale operation, and then sent to the enhanced feature extraction module for enhanced feature extraction operation; Finally, the output feature map of the enhanced feature extraction module is connected to the channel input feature map of the next layer using the add method, and the channel of the final output feature map is adjusted to the size of the channel of the next layer.

3. According to claim 2, a multi-scale target detection method based on feature enhancement and pixel inversion defogging is characterized in that: The enhanced feature extraction module includes: a convolution operation with a convolution kernel of 1*1, a convolution operation with a convolution kernel of 3*3 and a step size of 2, and a convolution operation with a convolution kernel of 1*1, which are set in sequence.

4. The multi-scale target detection method based on feature enhancement and pixel inversion defogging according to claim 1, characterized in that: The prediction module includes a detector and a classifier.

5. The multi-scale target detection method based on feature enhancement and pixel inversion defogging according to claim 1, characterized in that: The preprocessing module includes: an image inversion operation, a dark channel calculation, an atmospheric light intensity estimation operation, a transmittance estimation, a transmission optimization and an image inversion operation which are connected in sequence.

6. The multi-scale target detection method based on feature enhancement and pixel inversion defogging according to claim 1, characterized in that: The number of channels of the convolutional layers 1 to 5 are set to 64, 128, 256, 512, and 512, respectively.

7. The multi-scale target detection method based on feature enhancement and pixel inversion defogging according to claim 1, characterized in that: The number of channels corresponding to the convolutional layers 6 to 9 are set to 512, 256, 256, and 256, respectively.

8. The multi-scale target detection method based on feature enhancement and pixel inversion defogging according to claim 1, characterized in that: The corresponding scales of the prediction modules are: 38*38, 19*19, 10*10, 5*5, 3*3, and 1*1.

Citation Information

Patent Citations

  • Lightweight underwater target video detection method

    CN118552839A

  • Weed detection method based on multi-scale fusion module and feature enhancement

    CN113657326A

  • Multi-scale target detection method combining equilibrium features and deformable convolution

    CN114913433A