A Mobile Target Detection Method Based on Visual Perception
By improving the feature extraction layer and feature fusion network in the YOLOv5 model, the problem of low detection accuracy of vehicle or ship targets under low visual environment and complex background is solved, and higher detection accuracy and lower missed detection error rate are achieved.
Patent Information
- Application Number
- CN202310500160.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-06
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2043-05-06
AI Technical Summary
The prior art has problems of low accuracy and high missed detection error rate in vehicle or ship target detection in low visual environments and complex backgrounds, especially in multi-scale target detection and feature extraction.
By replacing the C3 module as GC module in the YOLOv5 model, the AC-SPPF module is built, the BIFPN structure is used for feature fusion, and a four-scale detection head is added to enhance the model's feature extraction and information fusion capabilities.
It improves the accuracy and generalization ability of vehicle or ship target detection, reduces the missed detection error rate, especially in low visual environments and complex backgrounds to significantly improve detection performance.
Smart Images

Figure CN116524268B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to target detection technology, and specifically to a moving target detection method based on visual perception. Background Art
[0002] In traditional lane and inland river monitoring systems, generally, specialized staff are required to subjectively judge based on the video images transmitted by fixed cameras, and relevant departments will be informed to handle the situation only after a dangerous situation or accident occurs. This not only increases the labor cost, but also leaves very little time for rescue or crime fighting at the first time. The monitoring environment and targets of lanes and inland river channels are very complex. First, the background of the lane and inland river environment is complex, making it difficult for the detection algorithm to characterize and adapt to the target in low visibility environments (foggy days, nights) and cluttered backgrounds, resulting in low inspection accuracy. Then, the sizes of vehicles driving in the lane and ships navigating in the inland river channel vary, which poses a great challenge to the multi-scale vehicle or ship feature extraction of the target detection network.
[0003] For vehicles in the lane and ships near the inland river, it is difficult to extract global features in low visibility environments. How to model the global context from a global perspective and capture long-range dependencies is the key to improving the detection accuracy of vehicle and ship targets. In addition, there is serious interference from background information such as buildings and trees in lane and inland river channel target detection. If the detection accuracy is to be improved, at this time, it is more necessary to retain rich semantic information, reduce the loss of local information, and obtain a larger receptive field to better locate and detect vehicle or ship targets. In addition, the targets in lane and inland river monitoring images generally vary in size. How to improve the multi-scale fusion method and enhance the model detection ability is the key to avoiding missed and false detections during the detection process and improving the target detection accuracy.
[0004] Due to the interference of low visibility environments (foggy days, nights) and strong scattering in inland areas. In the field of vehicle or ship target detection in low visibility environments, because convolution in the convolutional neural network can only perform context modeling in a local area, the receptive field is limited. If the way the network extracts features is only a simple stacking of convolutions, this situation is similar to repeating a certain type of function all the time, which will lead to the loss of semantic information in the features extracted by the network and is not conducive to the feature extraction of vehicle or ship targets in low visibility environments.
[0005] In the Backbone layer of YOLOv5, after the input image undergoes multiple convolutional operations, it enters the SPPF module and uses the maximum pooling layer of the same size to obtain features with different receptive fields, realizing the fusion of global features and local features. However, the operation of extracting features by convolution and then performing pooling downsampling will lead to the loss of local information, and the ability of convolution operation to extract features of complex background images is limited. As a result, the feature extraction and feature fusion capabilities decline, leading to poor detection accuracy.
[0006] In a convolutional neural network, the shallow network can effectively represent detailed information and capture more small-scale vehicle or ship information. However, it cannot achieve a strong ability to represent feature semantic information. The deep network has poor ability to represent detailed information, but has a strong ability to represent feature semantic information and can capture more medium and large-scale vehicle or ship information. Therefore, by fusing the multi-scale feature information of the deep and shallow networks and integrating the advantages of the shallow and deep networks, the accuracy of object localization and classification in vehicle or ship target detection can be improved. The Neck structure in YOLOv5 is composed of FPN and a bottom-up path PAN added on this basis to form the PANet structure. The PANet structure utilizes semantic and detailed feature information in feature map processing to achieve multi-scale feature fusion of vehicles or ships. However, this method still has certain deficiencies. The main problem is that all the input information of PAN is processed by FPN, and it is difficult to use the original information obtained from the backbone network for fusion. As the depth of the convolutional network increases in object detection, it is more difficult to retain the semantic information of small-scale vehicles or ships, which will limit network learning and affect the accuracy of the detection results to a certain extent.
[0007] The YOLOv5 network has a relatively deep hierarchy. After continuously performing convolutional downsampling, abstract semantic features can be extracted from the input image. However, this also reduces the size of the feature map, which may result in the loss of feature information for small-sized vehicles or ships and insufficient feature information retention. Combining the above analysis, it is difficult to meet the requirements using deep features for small target prediction. From the structure of YOLOv5, there are 3 detection heads in YOLOv5, which perform object detection in feature maps of 20×20, 40×40, and 80×80 respectively. Due to the complex conditions of lanes and inland waterways in China, the sizes of vehicles or ships detected in images vary. If a vehicle or ship occupies 8 pixels in terms of pixel size in the image, and the size of the input image is 640 pixels, even in the feature map with the largest scale of 80×80, the scale occupied by the target object is only 1 pixel; if the scale of the target object is less than 8 pixels, the entire target will disappear in the feature map, and the previous 80×80 size detection layer may not be able to accurately detect smaller vehicles or ships in the image. Summary of the Invention
[0008] The purpose of the present invention is to address the deficiencies in the prior art and propose a moving target detection method based on visual perception. This method has high detection accuracy and can reduce missed detections and false detections in vehicle or ship target detection in visual pictures or videos.
[0009] The technical solution to achieve the purpose of the present invention is as follows:
[0010] A moving target detection method based on visual perception, the input of which is an image, and the method includes the following steps:
[0011] S1 Replace the C3 module with the GC module in the YOLOv5 feature extraction layer;
[0012] S2 Construct an AC-SPPF module;
[0013] S2-1 Send the features into the SPPF module through layer-by-layer convolution for feature extraction and information fusion;
[0014] S2-2 The SPPF module fuses the ACmix module and constructs the AC-SPPF module by building a residual structure;
[0015] S3 Replace the feature fusion network PANet with the BIFPN structure;
[0016] S4 Add a four-scale detection head to detect the target.
[0017] The specific steps of replacing the C3 module with the GC module in the YOLOv5 feature extraction layer in step S1 are as follows:
[0018] S1-1 The image enters the feature extraction layer Bcakbone after preprocessing;
[0019] S1-2 Directly replace the C3 module with the GC module;
[0020] S1-3 Use the GC module to extract features from the feature map of the feature extraction layer Backbone from shallow to deep;
[0021] S1-4 The output of each GC module is connected to the feature fusion Neck layer.
[0022] The specific steps of constructing the AC-SPPF module in step S2 are as follows:
[0023] S2-1 Send the features into the SPPF module through layer-by-layer convolution for feature extraction and information fusion;
[0024] S2-2 The SPPF module fuses the ACmix module and constructs the AC-SPPF module by building a residual structure;
[0025] S2-3 After the input feature map undergoes a convolution operation, it is further enhanced by ACmix for information aggregation to weaken the interference of complex background information;
[0026] S2-4 Then, the feature map in step S2-3 is passed through a CBS convolution with a convolution kernel size of 1×1 to reduce the number of channels, and then different receptive field features are obtained through 3 max-pooling layers to achieve the fusion of global features and local features;
[0027] S2-5 is further passed through a CBS convolution module with a convolution kernel size of 1×1 to adjust the channels and restore to the number of channels of the original features. Finally, it is added and fused with the original feature map through a residual structure to retain the rich local information in the original features, and then the features are output.
[0028] The specific steps for replacing the feature fusion network PANet with the BIFPN structure in step S3 are as follows:
[0029] S3-1 First, replace the original feature fusion network PANet with a weighted bidirectional feature pyramid network BIFPN;
[0030] S3-2 Send the features of step S2-5 into the feature fusion layer of the BIFPN structure;
[0031] S3-3 At this time, the original information of the backbone network participates in the fusion through two skip connections and two cross-scale connections;
[0032] S3-4 Set learnable weight parameters through the BFConcat layer to achieve bidirectional fusion of shallow and deep layer features, enhance local perception and the feature transfer ability between network layers to determine the influence of features at each scale on the output features;
[0033] S3-5 Send the fused features into the detection layer.
[0034] The specific steps for adding a four-scale detection head in step S4 are as follows:
[0035] S4-1 Add a 4-fold downsampling once in the Backbone layer;
[0036] S4-2 Add an upsampling and a downsampling once on the original Neck layer;
[0037] S4-3 Fuse the output feature map of the first GC in the Backbone layer with the 3 times upsampling in the Neck layer into the Concat layer, and then obtain a detection head with a detection scale of 160×160;
[0038] S4-4 After using a detection feature map with a higher resolution of 160×160 in the detection head, the number of pixels occupied by the vehicle or ship will increase, making it easier to be detected.
[0039] In this technical solution, in the improved YOLOv5, the GC module replaces the C3 module in the Backbone layer to strengthen the extraction of image features from a global perspective, and the construction of the AC-SPPF model strengthens information aggregation and reduces local information loss. In the Neck layer, the BIFPN structure is adopted to achieve cross-scale feature fusion and realize multi-scale information fusion, enabling the model to fuse more information. In addition, cross-connection operations are added to effectively alleviate the problem that the original information obtained from the backbone network is difficult to participate in the fusion. By setting learnable weight parameters in the BFConcat layer, bidirectional fusion of shallow and deep layer features is achieved, enhancing local perception and the feature transfer ability between network layers, thereby determining the influence of features at each scale on the output features. A large-size detection head is added to the detection head to improve the detection ability of vehicle or ship targets.
[0040] Compared with the original YOLOv5 method, the improved YOLOv5 method in this technical solution effectively improves the detection accuracy of vehicle or ship targets in low visibility environments (rainy, foggy days, nights) and complex backgrounds, reduces the false detection and missed detection probabilities of vehicles or ships in the above environments, and improves the generalization ability of the method in multi-scenario target detection. Brief Description of the Drawings
[0041] Figure 1 It is a diagram of the YOLOv5 model after the GC replaces the C3 module in the embodiment;
[0042] Figure 2 It is a structural diagram of the AC-SPPF module in the embodiment;
[0043] Figure 3 It is a diagram of the model with AC-SPPF embedded in YOLOv5 in the embodiment;
[0044] Figure 4 It is a schematic diagram of different feature fusion structures in the embodiment;
[0045] Figure 5 It is a schematic diagram of the four-scale detection head in the embodiment;
[0046] Figure 6 It is a diagram of the improved YOLOv5 model in the embodiment;
[0047] Figure 7 It is a diagram of the visualization comparison results of YOLOv5 and GC-YOLOv5 in the embodiment;
[0048] Figure 8 It is a visualization comparison effect diagram of actual target detection in low visibility environments in the embodiment. Detailed Implementation Manner
[0049] The following further describes the invention in detail in conjunction with the drawings and specific embodiments, but it is not a limitation to the present invention. Embodiment
[0050] Referring to Figure 6 , a moving target detection method based on visual perception, the input of which is an image, and the method includes the following steps:
[0051] S1 Replace the C3 module with the GC module in the YOLOv5 feature extraction layer, as Figure 1 shown;
[0052] S2 Construct the AC-SPPF module, as Figure 2 , 3 shown;
[0053] S2-1 Send the features into the SPPF module through multiple layers of convolution for feature extraction and information fusion;
[0054] S2-2 The SPPF module fuses the ACmix module, and constructs the AC-SPPF module by building a residual structure, as Figure 2 shown;
[0055] S3 Replace the feature fusion network PANet with the BIFPN structure, as Figure 4 shown;
[0056] S4 Add a four-scale detection head, as Figure 5 shown, to detect the target.
[0057] The specific steps of replacing the C3 module with the GC module in the YOLOv5 feature extraction layer in step S1 are as follows:
[0058] S1-1 The image enters the feature extraction layer Backbone after preprocessing;
[0059] S1-2 Directly replace the C3 module with the GC module, as Figure 1 shown;
[0060] S1-3 Use the GC module to extract features from the feature map of the feature extraction layer Backbone from the shallow layer to the deep layer;
[0061] S1-4 The output of each GC module is connected to the feature fusion Neck layer.
[0062] The specific steps of constructing the AC-SPPF module in step S2 are as follows:
[0063] S2-1 Send the features into the SPPF module through multiple layers of convolution for feature extraction and information fusion;
[0064] S2-2 The SPPF module fuses the ACmix module, and constructs the AC-SPPF module by building a residual structure, as Figure 2 shown;
[0065] After the input feature map of S2-3 undergoes a convolution operation, it is then enhanced by ACmix for information aggregation to weaken the interference of complex background information;
[0066] S2-4 Then, the feature map from step S2-3 is passed through a CBS convolution with a kernel size of 1×1 to reduce the number of channels. Then, different receptive field features are obtained through 3 max pooling layers to achieve the fusion of global and local features;
[0067] S2-5 It is then passed through a CBS convolution module with a kernel size of 1×1 to adjust the channels and restore to the number of channels of the original feature. Finally, it is added and fused with the original feature map through a residual structure to retain the rich local information in the original feature, and then the feature is output, as Figure 3 shown.
[0068] The specific steps for replacing the feature fusion network PANet with the BIFPN structure in step S3 are as follows:
[0069] S3-1 First, replace the original feature fusion network PANet with a weighted bidirectional feature pyramid network BIFPN, as Figure 4 shown;
[0070] S3-2 Send the feature from step S2-5 into the feature fusion layer of the BIFPN structure;
[0071] S3-3 At this time, the original information of the backbone network participates in the fusion through two skip connections and two cross-scale connections;
[0072] S3-4 Set learnable weight parameters through the BFConcat layer to achieve bidirectional fusion of shallow and deep layer features, enhance local perception and the feature transfer ability between network layers to determine the influence of features at each scale on the output feature;
[0073] S3-5 Send the fused feature into the detection layer.
[0074] The specific steps for adding a four-scale detection head in step S4 are as follows:
[0075] S4-1 Add a 4-fold downsampling in the Backbone layer;
[0076] S4-2 Add an upsampling and a downsampling on the original Neck layer;
[0077] S4-3 Fuse the output feature map of the first GC in the Backbone layer with the 3 times upsampling in the Neck layer into the Concat layer, and then obtain a detection head with a detection scale of 160×160, as Figure 5 shown;
[0078] After using a detection feature map with a higher resolution of 160×160 in the detection head, the pixels occupied by the vehicle or ship will be more, making it easier to be detected.
[0079] Experimental results and analysis of GC module replacement:
[0080] Comparison of ship target detection results of different methods:
[0081] Table 1 Comparison of ship target detection results of different methods
[0082] 。
[0083] Table 1 shows the data comparison of different selected methods such as Faster R-CNN, SSD, YOLOv5 and GC-YOLOv5 on the self-made ship dataset. For the two-stage object detection method, the mAP of the GC-YOLOv5 method proposed in this example is 9.8% higher than that of Faster R-CNN. For the single-stage detection method, it is 13.5% higher than SSD and 5.5% higher than YOLOv5. Considering the comprehensive mAP and FPS indicators, the GC-YOLOv5 method in this example has the best performance, achieving real-time ship target detection in a low-visibility environment while improving the accuracy.
[0084] Comparison of performance indicators between YOLOv5 and GC-YOLOv5:
[0085] Table 2 Comparison of performance indicators between YOLOv5 and GC-YOLOv5
[0086] 。
[0087] Since the detection accuracy of the YOLOv5 method is higher than that of the selected Faster R-CNN and SSD methods, it further verifies the superiority of the GC-YOLOv5 compared with YOLOv5 in performance. It can be analyzed from Table 2 that the precision of GC-YOLOv5 has increased by 8.4 percentage points, the recall rate has increased by 6 percentage points, mAP@0.5 has increased by 5.5 percentage points, and mAP@0.5:0.95 has increased by 6.4 percentage points compared with YOLOv5 on the validation set. Due to the replacement and application of the lightweight module GCnet, the detection speed has increased by 6 frames.
[0088] The visual comparison results of image detection of YOLOv5 and GC-YOLOv5 methods are as Figure 7 shown. By observing, it is found that in the night environment, YOLOv5 has cases of missed detection and false detection. Specifically, as Figure 7 (a)In the night environment, a ship is misdetected as two ships, Figure 7 (c)The ship is missed. In foggy weatherFigure 7 There are undetected ships in both (e) and (g). However, Figure 7 All ships in (b), (d), (f), and (h) are correctly detected. Therefore, GC-YOLOv5 effectively reduces the false detection and missed detection rates of ships, and improves the detection accuracy compared to YOLOv5.
[0089] Verification of ship target detection in low visibility environment:
[0090] To further verify the usability of GC-YOLOv5 in low visibility environments such as foggy days and nights. In this example, ship navigation videos are obtained using the riverside video monitoring system in the Yongjiang River section of Nanning. The selected videos for detection respectively include foggy days and night environments. YOLOv5 and GC-YOLOv5 are used to detect ships in the same video in foggy days and night environments respectively. Finally, the following comparison diagrams of detection effects are obtained by randomly intercepting the same frame (435, 4).
[0091] Through Figure 8 It is found in (a) that when detecting ships in the night environment, YOLOv5 did not detect two ships. In Figure 8 In the (c) foggy day environment, YOLOv5 misdetected a distant billboard as a ship. While in Figure 8 In (b), GC-YOLOv5 detected both ships in the night environment. In Figure 8 In the (d) foggy day environment, GC-YOLOv5 correctly detected the ship without false detection. Based on the above analysis, the practical performance of the GC-YOLOv5 method in ship target detection in actual low visibility environments is better than the original method, and the detection accuracy of ship targets in low visibility environments is improved.
Claims
1. A moving target detection method based on visual perception, characterized in that, The input of this method is an image, and the method includes the following steps: S1 Replace the C3 module with the GC module in the YOLOv5 feature extraction layer; S2 Construct the AC-SPPF module; S2-1 Send the features into the SPPF module through layer-by-layer convolution for feature extraction and information fusion; S2-2 The SPPF module fuses the ACmix module, and constructs the AC-SPPF module by building a residual structure; S3 Replace the feature fusion network PANet with the BIFPN structure; S4 Add a four-scale detection head to detect the target.
2. The moving target detection method based on visual perception according to claim 1, characterized in that, The specific steps of replacing the C3 module with the GC module in the YOLOv5 feature extraction layer in step S1 are as follows: S1-1 The image enters the feature extraction layer Bcakbone after preprocessing; S1-2 Directly replace the C3 module with the GC module; S1-3 Use the GC module to extract features from the feature map of the feature extraction layer Backbone from shallow to deep; S1-4 The output of each GC module is connected to the feature fusion Neck layer.
3. The moving target detection method based on visual perception according to claim 1, characterized in that, The specific steps of constructing the AC-SPPF module in step S2 are as follows: S2-1 Send the features into the SPPF module through layer-by-layer convolution for feature extraction and information fusion; S2-2 The SPPF module fuses the ACmix module, and constructs the AC-SPPF module by building a residual structure; S2-3 After the input feature map undergoes a convolution operation, it is further enhanced by ACmix for information aggregation to weaken the interference of complex background information; S2-4 Then, the feature map in step S2-3 undergoes a CBS convolution with a kernel size of 1×1 to reduce the number of channels, and then different receptive field features are obtained through 3 max pooling layers to achieve the fusion of global features and local features; S2-5 Then, the channel is adjusted through a CBS convolution module with a kernel size of 1×1 to restore to the number of channels of the original feature, and finally, it is added and fused with the original feature map through a residual structure to retain the rich local information in the original feature, and then the feature is output.
4. The moving target detection method based on visual perception according to claim 3, characterized in that, The specific steps of replacing the feature fusion network PANet with the BIFPN structure in step S3 are as follows: S3-1 First, replace the original feature fusion network PANet with the weighted bidirectional feature pyramid network BIFPN; S3-2 Send the features in step S2-5 into the feature fusion layer of the BIFPN structure; S3-3 At this time, the original information of the backbone network participates in the fusion through two skip connections and two cross-scale connections; S3-4 Set learnable weight parameters through the BFConcat layer to achieve bidirectional fusion of shallow and deep layer features, enhance local perception and the feature transfer ability between network layers to determine the influence of each scale feature on the output feature; S3-5 Send the fused features into the detection layer.
5. The moving target detection method based on visual perception according to claim 1, characterized in that, The specific steps of adding a four-scale detection head in step S4 are as follows: S4-1 Add a 4-fold downsampling once in the Backbone layer; S4-2 Add an upsampling and a downsampling once on the original Neck layer; S4-3 Fuse the first GC output feature map of the Backbone layer with the 3 times upsampling in the Neck layer into the Concat layer, and then obtain a detection head with a detection scale of 160×160; S4-4 After using the detection feature map with a high resolution of 160×160 in the detection head, the pixels occupied by vehicles or ships will increase, making them easier to detect.
Citation Information
Patent Citations
Small target detection method based on multilevel residual network perception and attention mechanism
CN114821246A
Distribution line tree obstacle detection method
CN115205292A