YOLOv8 foggy day target detection method and system based on feature extraction fusion reinforcement
By introducing star modules, expansion residual modules and cross-attention deep semantic fusion modules in the YOLOv8 network, the FE-YOLOv8 network model is constructed using the anchor box-free recognition method, which solves the problem of reduced target detection accuracy and robustness in foggy days, and achieves efficient and accurate target detection in foggy days.
Patent Information
- Application Number
- CN202411977644.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, under severe weather conditions, especially in foggy days, it is difficult to accurately identify vehicles and pedestrians on the road, resulting in reduced detection accuracy and robustness, and high computational complexity, making it difficult to meet real-time requirements.
A YOLOv8 foggy day object detection method based on feature extraction and fusion enhancement is proposed. By introducing a star module, an expansion residual module and a deep semantic fusion module based on cross attention, an anchor frame recognition method is adopted to build a FE-YOLOv8 network model to enhance the feature extraction ability and feature fusion effect of foggy day images.
It improves the speed and accuracy of target detection in foggy days, reduces the missed detection of small background targets, reduces the amount of computing in the network, and meets the real-time requirements.
Smart Images

Figure CN119942298A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image recognition technology, and in particular to a YOLOv8 foggy target detection method and system based on feature extraction fusion enhancement. Background Art
[0002] Severe weather conditions increase risks in key areas such as intelligent traffic management, public safety maintenance, environmental protection supervision, and emergency rescue response. For example, in foggy weather, the light intensity is reduced and the edges of targets in the image are blurred, making it difficult for autonomous driving systems to accurately identify vehicles and pedestrians on the road, posing a major threat to driving safety.
[0003] In the field of target detection, target detection technology is currently divided into two categories: methods based on image processing and machine learning, and target detection methods based on deep learning. They usually rely on artificially designed features (such as HOG (Histogram of Oriented Gradients), SIFT (Scale-Invariant Feature Transform), etc.) combined with classifiers (such as SVM (Support Vector Machine), AdaBoost (Adaptive Boosting), etc.) for target detection. These methods perform well in scenes with good lighting conditions and simple backgrounds, but in bad weather, due to the sharp decline in image quality, it is difficult for traditional methods to extract effective features, resulting in a significant decrease in detection accuracy and robustness. In addition, existing methods face high computational complexity and are difficult to meet real-time requirements. In recent years, with the rapid development of deep learning technology, methods based on deep learning have made significant progress in the field of target detection. These methods learn and extract shallow and deep features of images through convolutional neural networks, and can better cope with target detection tasks in complex scenes. However, existing deep learning-based target detection methods still have many shortcomings under severe weather conditions: First, due to cross-domain differences (such as differences in image distribution under different weather conditions), existing methods are prone to domain shift problems, resulting in reduced performance in foggy target detection. Second, most existing methods focus on target detection in a single field, and it is difficult to handle target detection tasks in multiple severe weather scenarios at the same time. In addition, in foggy images, the pixel proportion of targets (especially small targets) in the image is small, and occlusion and blurred boundaries often occur, resulting in limited features extracted by the target detection algorithm and poor detection effect.
[0004] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the invention
[0005] The main purpose of the embodiments of the present application is to propose a YOLOv8 foggy target detection method and system based on feature extraction fusion enhancement, which can enrich the features extracted by the target detection algorithm, enable a deeper interaction between the target and environmental features, and reduce the missed detection of small background targets, thereby improving the detection speed and accuracy of foggy targets.
[0006] To achieve the above purpose, an embodiment of the present application proposes a YOLOv8 foggy target detection method based on feature extraction fusion enhancement, the method comprising:
[0007] Obtain a dataset of foggy images to be identified;
[0008] Based on the YOLOv8 network model, the star module, the dilated residual module, and the deep semantic fusion module based on cross attention are introduced, and the anchor-free box recognition method is adopted to construct the FE-YOLOv8 network model;
[0009] Based on the FE-YOLOv8 network model, image target recognition and detection processing is performed on the foggy image data set to be identified to obtain a foggy image target recognition and detection result.
[0010] In some embodiments, the FE-YOLOv8 network model includes a backbone feature extraction network module, a neck feature fusion network module and a head detection network module, the head detection network module adopts the anchor-free frame recognition method, the backbone feature extraction network module, the neck feature fusion network module and the head detection network module are connected in sequence, wherein:
[0011] The backbone feature extraction network module includes a convolution batch normalization activation module, a star module, a dilated residual module and a spatial pyramid pooling module, wherein the star module includes a first depth-separable convolution layer, a first fully connected layer, a second fully connected layer, a third fully connected layer and a second depth-separable convolution layer, and the dilated residual module includes a regional residual module, a multi-rate dilated depth convolution module, a batch normalization module and a point-by-point convolution fusion module;
[0012] The neck feature fusion network module includes a downsampling module, a feature splicing module, a spatial channel feature fusion module, a deep semantic fusion module based on cross attention, and a convolution batch normalization activation module, wherein the deep semantic fusion module based on cross attention includes a dense layer, a linear projection layer, a KQV embedding vector calculation layer, an activation layer, a splicing layer, and a convolution layer;
[0013] The head detection network module includes an anchor-free detection head module, an IoU-aware query selection module and an auxiliary detection head decoder module.
[0014] In some embodiments, the image target recognition detection processing is performed on the foggy image dataset to be identified based on the FE-YOLOv8 network model to obtain the foggy image target recognition detection result, including:
[0015] Inputting the foggy image data set to be identified into the FE-YOLOv8 network model;
[0016] Based on the backbone feature extraction network module of the FE-YOLOv8 network model, feature extraction processing is performed on the foggy image data set to be identified to obtain a foggy feature image;
[0017] Based on the neck feature fusion network module of the FE-YOLOv8 network model, feature fusion processing is performed on the foggy feature image to obtain a foggy feature fusion image;
[0018] Based on the head detection network module of the FE-YOLOv8 network model, anchor-free frame recognition and detection processing is performed on the foggy feature fusion image to obtain the foggy image target recognition and detection result.
[0019] In some embodiments, the backbone feature extraction network module based on the FE-YOLOv8 network model performs feature extraction processing on the foggy image dataset to be identified to obtain a foggy feature image, including:
[0020] Inputting the foggy image data set to be identified into the backbone feature extraction network module of the FE-YOLOv8 network model;
[0021] Based on the convolution batch normalization activation module of the backbone feature extraction network module, feature extraction processing is performed on the foggy image dataset to be identified to obtain a preliminary foggy feature image;
[0022] Based on the star module of the backbone feature extraction network module, the preliminary foggy weather feature image is subjected to feature mapping processing to obtain a mapped foggy weather feature image;
[0023] Based on the expansion residual module of the backbone feature extraction network module, the mapped foggy feature image is subjected to feature interactive fusion processing to obtain an interactively fused foggy feature image;
[0024] Based on the spatial pyramid pooling module of the backbone feature extraction network module, multi-scale pooling processing is performed on the interactively fused foggy feature image to obtain the foggy feature image.
[0025] In some embodiments, the star module based on the backbone feature extraction network module performs feature mapping processing on the preliminary foggy weather feature image to obtain a mapped foggy weather feature image, including:
[0026] Inputting the preliminary foggy weather feature image into the star module of the backbone feature extraction network module;
[0027] Based on the first depthwise separable convolutional layer of the star module, high-dimensional feature extraction processing is performed on the preliminary foggy weather feature image to obtain a high-dimensional foggy weather feature image;
[0028] Based on the first fully connected layer and the second fully connected layer of the star-shaped module, linear combination processing is performed on the high-dimensional foggy weather feature image to obtain a first high-dimensional foggy weather feature image and a second high-dimensional foggy weather feature image;
[0029] Performing matrix dot product calculation on the first high-dimensional foggy weather feature image and the second high-dimensional foggy weather feature image to obtain a calculated high-dimensional foggy weather feature image;
[0030] Based on the third fully connected layer of the star-shaped module, nonlinear mapping is performed on the calculated high-dimensional foggy weather characteristic image to obtain a foggy weather characteristic image after preliminary mapping;
[0031] Based on the second depthwise separable convolutional layer of the star-shaped module, feature extraction processing is performed on the initially mapped foggy feature image to obtain a mapped foggy feature image.
[0032] In some embodiments, the dilation residual module based on the backbone feature extraction network module performs feature interactive fusion processing on the mapped fog feature image to obtain the interactively fused fog feature image, including:
[0033] Inputting the mapped foggy feature image into the dilation residual module of the backbone feature extraction network module;
[0034] Based on the regional residual module of the dilated residual module, a preliminary target feature extraction process is performed on the mapped foggy feature image to obtain a target feature map with different foggy foregrounds;
[0035] Based on the multi-rate dilated deep convolution module of the dilated residual module, morphological filtering is performed on the feature map of targets with different foggy foregrounds to obtain a filtered feature map of targets with different foggy foregrounds;
[0036] Based on the batch normalization module of the dilated residual module, the filtered feature map with different foggy foreground targets is standardized to obtain a standardized feature map with different foggy foreground targets;
[0037] Based on the point-by-point convolution fusion module of the dilated residual module, the standardized feature map with different foggy foreground targets and the mapped foggy feature image are subjected to feature interactive fusion processing to obtain the interactively fused foggy feature image.
[0038] In some embodiments, the neck feature fusion network module based on the FE-YOLOv8 network model performs feature fusion processing on the foggy feature image to obtain a foggy feature fusion image, including:
[0039] Inputting the foggy weather feature image into the neck feature fusion network module of the FE-YOLOv8 network model;
[0040] Based on the downsampling module of the neck feature fusion network module, the foggy weather feature image is downsampled to obtain a downsampled foggy weather feature image;
[0041] Based on the feature stitching module of the neck feature fusion network module, feature stitching is performed on the downsampled fog feature image to obtain a preliminary stitched fog feature image;
[0042] Based on the spatial channel feature fusion module of the neck feature fusion network module, the preliminary spliced foggy feature image is subjected to spatial channel feature fusion processing to obtain a spliced foggy feature image;
[0043] Based on the cross-attention based deep semantic fusion module of the neck feature fusion network module, the spliced foggy feature image is subjected to deep semantic fusion processing to obtain a preliminary foggy feature fusion image;
[0044] Based on the convolution batch normalization activation module of the neck feature fusion network module, the preliminary foggy feature fusion image is subjected to normalization mapping processing to obtain the foggy feature fusion image.
[0045] In some embodiments, the cross-attention based deep semantic fusion module based on the neck feature fusion network module performs deep semantic fusion processing on the spliced foggy feature image to obtain a preliminary foggy feature fusion image, including:
[0046] Inputting the spliced foggy feature image into the cross-attention based deep semantic fusion module of the neck feature fusion network module;
[0047] Based on the dense layer of the cross-attention based deep semantic fusion module, the spliced foggy feature image is enhanced to obtain an enhanced foggy feature image;
[0048] Based on the linear projection layer of the cross-attention based deep semantic fusion module, projecting the enhanced foggy feature image to obtain a projected foggy feature image;
[0049] Based on the KQV embedding vector calculation layer of the cross-attention based deep semantic fusion module, KQV calculation is performed on the projected foggy feature image to obtain a feature vector of the foggy image;
[0050] Reconstructing the feature vector of the foggy image and the spliced foggy feature image based on the activation layer and the splicing layer of the cross-attention based deep semantic fusion module to obtain a reconstructed foggy feature fusion image;
[0051] Based on the convolution layer of the cross-attention based deep semantic fusion module, the reconstructed foggy feature fusion image is convolved to obtain the preliminary foggy feature fusion image.
[0052] In some embodiments, the head detection network module based on the FE-YOLOv8 network model performs anchor-free recognition detection processing on the foggy feature fusion image to obtain the foggy image target recognition detection result, including:
[0053] Inputting the foggy feature fusion image into the head detection network module of the FE-YOLOv8 network model;
[0054] Based on the anchor-free detection head module of the head detection network module, target information extraction processing is performed on the foggy feature fusion image to obtain predicted foggy image target feature information;
[0055] Based on the IoU perception query selection module of the head detection network module, the predicted foggy image target feature information is screened to obtain the screened predicted foggy image target feature information;
[0056] Based on the auxiliary detection head decoder module of the head detection network module, target detection processing is performed on the filtered predicted foggy image target feature information to obtain the foggy image target recognition detection result.
[0057] To achieve the above purpose, another aspect of the embodiment of the present application proposes a YOLOv8 foggy target detection system based on feature extraction fusion enhancement, the system comprising:
[0058] The first module is used to obtain a foggy image dataset to be identified;
[0059] The second module is used to build a FE-YOLOv8 network model based on the YOLOv8 network model by introducing a star module, an expansion residual module, and a deep semantic fusion module based on cross attention and adopting an anchor-free frame recognition method;
[0060] The third module is used to perform image target recognition and detection processing on the foggy image data set to be identified based on the FE-YOLOv8 network model to obtain a foggy image target recognition and detection result.
[0061] The embodiments of the present application include at least the following beneficial effects: The present application provides a YOLOv8 foggy target detection method and system based on feature extraction fusion enhancement. The scheme first obtains a foggy image dataset to be identified, and then introduces a star module, an expansion residual module and a deep semantic fusion module based on cross-attention based on the YOLOv8 network model, adopts an anchor-free frame recognition method, and constructs a FE-YOLOv8 network model. Finally, the foggy image dataset to be identified is subjected to image target recognition and detection processing. The feature extraction capability of the foggy image is enhanced by the star module and the expansion residual module, so that the model can capture more foggy image features, thereby enriching the features extracted by the target detection algorithm. The deep semantic fusion module based on cross-attention is used to fuse foggy features of different depths and shallow layers, so that the target and the environmental features interact more deeply, reducing the phenomenon of missed detection of small background targets. The anchor-free frame recognition method reduces the amount of network calculation, removes the fault tolerance of the anchor frame that may not be applicable to targets of all sizes and shapes, improves the performance of small background target detection in foggy weather, and thereby improves the detection speed and accuracy of foggy targets. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 It is a flow chart of the YOLOv8 foggy target detection method based on feature extraction fusion enhancement provided in an embodiment of the present application;
[0063] Figure 2 It is a structural diagram of a YOLOv8 foggy target detection system based on feature extraction fusion enhancement provided in an embodiment of the present application;
[0064] Figure 3 It is a structural diagram of the FE-YOLOv8 network model provided in an embodiment of the present application;
[0065] Figure 4 is a schematic diagram of the structure of a star-shaped module provided in an embodiment of the present application;
[0066] Figure 5 is a schematic diagram of the structure of the expansion residual module provided in an embodiment of the present application;
[0067] Figure 6It is a structural diagram of a deep semantic fusion module based on cross attention provided in an embodiment of the present application;
[0068] Figure 7 It is a comparative schematic diagram of the simulation experiment results provided in the embodiments of the present application. DETAILED DESCRIPTION
[0069] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are only examples of systems and methods consistent with some aspects of the embodiments of the present application as detailed in the attached claims.
[0070] It is understood that the terms "first", "second", etc. used in this application can be used to describe various concepts in embodiments of the present invention, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determination".
[0071] The terms "at least one", "multiple", "each", "any", etc. used in this application, at least one includes one, two or more, multiple includes two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.
[0072] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present invention have the same meanings as those commonly understood by those skilled in the art to which the present application belongs. The terms used in the embodiments of the present invention are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0073] Reference Figure 1 , Figure 1 A flowchart of a YOLOv8 foggy target detection method based on feature extraction fusion enhancement provided by an embodiment of the present invention, referring to Figure 1 , the method comprises the following steps:
[0074] S100, obtaining a foggy image dataset to be identified;
[0075] In some specific embodiments, the experiment selects the public dataset RTTS (Real-world Task-driven Testing Set) for foggy object detection. The RTTS dataset is a foggy dataset in a real scene, containing a total of 4320 images, including 5 types of objects, namely Bus, People, Car, Motor and Bicycle.
[0076] S200, based on the YOLOv8 network model, introduces the star module, the dilated residual module and the deep semantic fusion module based on cross attention, adopts the anchor-free frame recognition method, and constructs the FE-YOLOv8 network model;
[0077] It should be noted that, in some specific embodiments, the FE-YOLOv8 network model includes a backbone feature extraction network module, a neck feature fusion network module and a head detection network module. The head detection network module adopts an anchor-free frame recognition method, and the backbone feature extraction network module, the neck feature fusion network module and the head detection network module are connected in sequence.
[0078] It should be noted that the embodiment of the present invention optimizes and improves the Backbone (backbone feature extraction network), Neck (neck feature fusion network) and Head (head detection network) of YOLOv8, and proposes a FE-YOLOv8 (Feature Extraction Fusion Enhanced YOLOv8) foggy target detection algorithm. The overall network structure of FE-YOLOv8 is as follows: Figure 3As shown in the figure, the Backbone module introduces StarBlock, which uses element multiplication (star operation) to map the fuzzy features in the foggy image to a high-dimensional nonlinear feature space without widening the network, thereby enhancing the fuzzy features to the greatest extent. The DWR (Dilation-wise Residual) module uses a multi-branch structure to expand the feature map size of the foggy features. Each branch uses a hole convolution of different sizes to improve the features of small targets in the foggy background and avoid missed detection. The Deep Semantic Fusion Module with Cross Attention (PSFM) based on the Neck part is added to avoid the feature loss of small targets in the foggy image through the interaction of shallow and deep features. The Head part of YOLOv8 adopts an anchor-free calculation method, which brings flexibility and reduces the dependence on anchor frames. However, in a foggy environment, due to the increased uncertainty of the target shape and position and the blurred target boundary, it is difficult for the Head to accurately predict the target's bounding box, affecting its detection and positioning accuracy. Inspired by the codec structure in RTDETR, the detection head structure design in RTDETR is adopted. The multi-scale feature fusion capability of this structure is used to adapt to the scale changes of targets in fog, enhance the focus on the area of interest, reduce background noise interference, and improve the detection accuracy and robustness in foggy environments.
[0079] It should be further explained that the backbone feature extraction network module includes a convolution batch normalization activation module, a star module, a dilated residual module and a spatial pyramid pooling module, wherein the star module includes a first depth-separable convolution layer, a first fully connected layer, a second fully connected layer, a third fully connected layer and a second depth-separable convolution layer, and the dilated residual module includes a regional residual module, a multi-rate dilated depth convolution module, a batch normalization module and a point-by-point convolution fusion module;
[0080] The neck feature fusion network module includes a downsampling module, a feature splicing module, a spatial channel feature fusion module, a deep semantic fusion module based on cross attention, and a convolution batch normalization activation module. The deep semantic fusion module based on cross attention includes a dense layer, a linear projection layer, a KQV embedding vector calculation layer, an activation layer, a splicing layer, and a convolution layer.
[0081] The head detection network module includes an anchor-free detection head module, an IoU-aware query selection module and an auxiliary detection head decoder module.
[0082] S300, based on the FE-YOLOv8 network model, performing image target recognition and detection processing on the foggy image data set to be recognized, and obtaining a foggy image target recognition and detection result;
[0083] It should be noted that, in some embodiments, step S300 may include:
[0084] S310, inputting the foggy image data set to be identified into the FE-YOLOv8 network model;
[0085] S320, a backbone feature extraction network module based on the FE-YOLOv8 network model, performs feature extraction processing on the foggy image data set to be identified to obtain a foggy feature image;
[0086] Specifically, in this embodiment, step S320 may include:
[0087] S321, inputting the foggy image data set to be identified into the backbone feature extraction network module of the FE-YOLOv8 network model;
[0088] S322, based on the convolution batch normalization activation module of the backbone feature extraction network module, performing feature extraction processing on the foggy image dataset to be identified to obtain a preliminary foggy feature image;
[0089] S323, based on the star module of the backbone feature extraction network module, performing feature mapping processing on the preliminary foggy feature image to obtain a mapped foggy feature image;
[0090] Specifically, a preliminary fog feature image is input into the star module of the backbone feature extraction network module; based on the first depth-separable convolutional layer of the star module, a high-dimensional feature extraction process is performed on the preliminary fog feature image to obtain a high-dimensional fog feature image; based on the first fully connected layer and the second fully connected layer of the star module, a linear combination process is performed on the high-dimensional fog feature image to obtain a first high-dimensional fog feature image and a second high-dimensional fog feature image; a matrix dot product calculation is performed on the first high-dimensional fog feature image and the second high-dimensional fog feature image to obtain a calculated high-dimensional fog feature image; based on the third fully connected layer of the star module, a nonlinear mapping process is performed on the calculated high-dimensional fog feature image to obtain a preliminary mapped fog feature image; based on the second depth-separable convolutional layer of the star module, a feature extraction process is performed on the preliminary mapped fog feature image to obtain a mapped fog feature image.
[0091] In this embodiment, in order to balance the computational complexity and performance in the foggy object detection process, StarBlock realizes the mapping of high-dimensional and nonlinear feature spaces through star operations, and can obtain richer feature representation without increasing the computational complexity. This module can implicitly integrate high-dimensional features into low-dimensional foggy image features, which not only realizes more accurate feature extraction, but also speeds up the model calculation efficiency. The StarBlock module structure is as follows: Figure 4 As shown, the module uses batch normalization to improve efficiency and places the normalized input foggy image after the deep convolution. Then, a deep convolution is introduced at the end of each block, where the convolution kernel of the deep separable convolution is 7 and the step size is 1. The embodiment of the present invention sets the channel expansion factor to 4, so that the model can extract richer implicit features.
[0092] S324, based on the expansion residual module of the backbone feature extraction network module, performing feature interactive fusion processing on the mapped foggy feature image to obtain an interactively fused foggy feature image;
[0093] Specifically, the mapped fog feature image is input into the dilated residual module of the backbone feature extraction network module; the regional residual module based on the dilated residual module performs preliminary target feature extraction processing on the mapped fog feature image to obtain a feature map of targets with different fog foregrounds; the multi-rate dilated deep convolution module based on the dilated residual module performs morphological filtering processing on the feature map of targets with different fog foregrounds to obtain a filtered feature map of targets with different fog foregrounds; the batch normalization module based on the dilated residual module performs standardization processing on the filtered feature map of targets with different fog foregrounds to obtain a standardized feature map of targets with different fog foregrounds; the point-by-point convolution fusion module based on the dilated residual module performs feature interactive fusion processing on the standardized feature map of targets with different fog foregrounds and the mapped fog feature image to obtain an interactively fused fog feature image.
[0094] In this embodiment, in order to reduce the phenomenon of missed detection of small targets in the background of foggy images, the embodiment of the present invention embeds the DWR structure in the Backbone structure to extract more context features, thereby further enhancing its feature extraction capability. Figure 5As shown in the figure, DWR can efficiently obtain multi-scale context information through the design of residual structure. This module efficiently obtains multi-scale context information of foggy targets through a two-step method and fuses to generate multi-scale receptive field feature maps. In the first step RR, the input foggy feature map is used to extract rough foggy large target features through regional residualization. In this process, a series of feature maps with different foggy foreground targets are generated. In the second step SR, the feature map calculated in the first step is input, and the multi-rate dilated deep convolution is used to perform morphological filtering on the foggy background small target features. To avoid feature redundancy, the DWR module only applies one required receptive field to each channel feature, and refines the background small target features. After the above two steps are completed, all feature maps are spliced, and then batch normalized, and point-by-point convolution is used to fuse features to form the final residual feature. Finally, the final residual feature is added to the input feature map to complete the feature interaction fusion. The DWR module uses a two-step method to transform the interactive fusion of foggy features from complex feature acquisition to concise feature map expression, and uses morphological filtering to remove the noise interference of some foggy targets, making the FE-YOLOv8 algorithm learning process more orderly and more efficient in acquiring multi-scale contextual information.
[0095] S325. The spatial pyramid pooling module based on the backbone feature extraction network module performs multi-scale pooling processing on the interactively fused foggy feature image to obtain the foggy feature image.
[0096] S330, a neck feature fusion network module based on the FE-YOLOv8 network model, performs feature fusion processing on the foggy feature image to obtain a foggy feature fusion image;
[0097] Specifically, in this embodiment, step S330 may include:
[0098] S331, inputting the foggy feature image into the neck feature fusion network module of the FE-YOLOv8 network model;
[0099] S332, based on the downsampling module of the neck feature fusion network module, performing a downsampling operation on the foggy weather feature image to obtain a downsampled foggy weather feature image;
[0100] S333, based on the feature stitching module of the neck feature fusion network module, performing feature stitching on the downsampled foggy feature image to obtain a preliminary stitched foggy feature image;
[0101] S334, performing spatial channel feature fusion processing on the preliminary spliced foggy feature image based on the spatial channel feature fusion module of the neck feature fusion network module to obtain a spliced foggy feature image;
[0102] S335, a cross-attention based deep semantic fusion module based on the neck feature fusion network module, performing deep semantic fusion processing on the spliced foggy feature image to obtain a preliminary foggy feature fusion image;
[0103] Specifically, the spliced fog feature image is input into the cross-attention based deep semantic fusion module of the neck feature fusion network module; the dense layer based on the cross-attention deep semantic fusion module performs enhancement processing on the spliced fog feature image to obtain the enhanced fog feature image; the linear projection layer based on the cross-attention deep semantic fusion module projects the enhanced fog feature image to obtain the projected fog feature image; the KQV embedding vector calculation layer based on the cross-attention deep semantic fusion module performs KQV calculation on the projected fog feature image to obtain the feature vector of the fog image; the activation layer and the splicing layer of the cross-attention deep semantic fusion module reconstruct the feature vector of the fog image and the spliced fog feature image to obtain the reconstructed fog feature fusion image; the convolution layer based on the cross-attention deep semantic fusion module performs convolution processing on the reconstructed fog feature fusion image to obtain a preliminary fog feature fusion image.
[0104] In this embodiment, in order to improve the feature expression of small targets in the background of foggy images, the embodiment of the present invention proposes to use a deep semantic fusion module PSFM based on cross attention to integrate deep features and embed them into the Neck structure, such as Figure 6 shown.
[0105] PSFM first uses a dense layer to enhance the foggy image features extracted by Backbone and outputs the enhanced depth features. The whole process is calculated as:
[0106]
[0107] In the above formula, x∈{ir,vi} represents the mode, represents the key K, represents the value V, Conv(·) and Reshape(·) are the 3×3 kernel convolution layer and the reshaping operation of the blurred features, respectively. i , W i and C i Represents the input features The height, width, and number of channels.
[0108] Multi-scale features are incorporated into the generation of modality-invariant queries, as shown in the following formula, which fully utilizes the complementary properties of multi-modal features. The expression is:
[0109]
[0110] in, Q i represents the query matrix Q, represents the feature concatenation operation, and express different modal characteristics.
[0111] Attention map calculated for the input convolutional layer x The expression is:
[0112]
[0113] In the above formula, represents the attention map of different modalities x, Denotes the query matrix Q i and key matrix The transposed multiplication matrix of is used to calculate the attention map.
[0114] Subsequently, the calculated feature values are multiplied with the characteristic diagnosis map obtained by attention to obtain features with global context. The global features are added to the original features of the other branch and the resulting features are concatenated along the channel dimension. The concatenated features are input into the convolutional layer to obtain the fused features. This process can be expressed as:
[0115]
[0116] In the above formula, represents the fused feature map, and Represented as raw features of different modes and represents the attention maps of different modalities, and represents different modal values of V.
[0117] S336. Based on the convolution batch normalization activation module of the neck feature fusion network module, normalization mapping processing is performed on the preliminary foggy feature fusion image to obtain a foggy feature fusion image.
[0118] S340, based on the head detection network module of the FE-YOLOv8 network model, performs anchor-free recognition and detection processing on the foggy feature fusion image to obtain the foggy image target recognition and detection result.
[0119] Specifically, in this embodiment, step S340 may include:
[0120] S341, inputting the foggy feature fusion image into the head detection network module of the FE-YOLOv8 network model;
[0121] S342, the anchor-free detection head module based on the head detection network module performs target information extraction processing on the foggy feature fusion image to obtain the predicted foggy image target feature information;
[0122] S343, the IoU perception query selection module based on the head detection network module screens and processes the predicted foggy image target feature information to obtain the screened predicted foggy image target feature information;
[0123] S344, the auxiliary detection head decoder module based on the head detection network module performs target detection processing on the filtered predicted foggy image target feature information to obtain the foggy image target recognition detection result.
[0124] In some specific embodiments, the anchor-free transformer detection head is designed: In order to solve the problem of reduced positioning and detection accuracy due to decreased image contrast and blurred targets caused by haze, inspired by the detection head in RTDETR, a Transformer Decoder-based detection head is designed to replace and modify the benchmark YOLOv8 to better capture long-distance pixel dependencies, and combined with multi-scale features, it enhances the model's target recognition and positioning accuracy in complex foggy environments, and significantly improves the detection effect of foggy targets.
[0125] like Figure 3 As shown in the Head part of FE-YOLOv8, the baseline YOLOv8 uses three detection heads to detect large, medium and small targets respectively. Each detection head is divided into two paths to locate and calculate the category, position and confidence of foggy targets by calculating Bbox loss and target loss respectively. The anchor-free transformer detection head designed in the embodiment of the present invention uses a layer of detection heads to fuse the different feature maps output by YOLOv8, and uses IoU-aware query selection to filter a fixed number of image features as the initial target feature query of the decoder. Finally, the target detection is completed by the decoder with an auxiliary detection head, and the position, confidence and category information of the corresponding target are generated.
[0126] The design of the transformer detection head without anchor boxes changes the output mode of the target detection results of the YOLO series. The design method of using a layer of detection heads can exclude some useless anchor points, thereby accelerating the convergence speed of training and improving the prediction accuracy of the model. The loss function of the detection head is consistent with RT DETR.
[0127] The following is a detailed description and explanation of the solution of the embodiment of the present invention in conjunction with a specific application example:
[0128] Specifically, in the embodiment of the present application, in order to verify the effectiveness of the method of the embodiment of the present invention in the foggy weather target detection task, the experiment selects the foggy weather target detection public data set RTTS (Real-world Task-driven Testing Set). The RTTS data set is a foggy weather data set in a real scene, containing a total of 4320 images, including 5 types of targets, namely Bus, People, Car, Motor and Bicycle. The sizes of the images in the RTTS data set are different, and the density of the fog is also different, which can greatly simulate the complex situations encountered in the actual driving environment and test the performance of the target detection algorithm in the foggy weather scene. The data set is divided into a training set, a validation set and a test set in a ratio of 8:1:1. The number of images in the training set is 3456, the number of images in the test set is 432, and the number of images in the validation set is 432.
[0129] The experiment uses the deep learning framework Pytorch 2.3.1, Python 3.8.19, CUDA 12.1, the hardware is 13th Gen Intel (R) Core (TM) i5-13600KF, and the GPU is NVIDIA GeForce RTX 4070S for single card training. The input image size is 640 × 640, and the pre-trained weights are used in the training phase. The Adam optimizer is selected to optimize the model, and the batch-size is set to 12, the num workers is 4, the epoch is 200, the initial learning rate is set to 0.01, and the weight decay coefficient is 0.000005. Its graphics card memory is 12GB.
[0130] When evaluating the performance of foggy object detection algorithms, detection accuracy is one of the core considerations, and AP (average precision for a single category), mAP (average precision for all categories), and mmAP (average precision at different intersection-over-union (IOU) thresholds, from 0.5 to 0.95, increasing by 0.05) are three key indicators. The higher the values of these indicators, the more accurate and stable the algorithm is in identifying and locating targets. In addition to detection accuracy, model size and real-time detection performance are also factors that cannot be ignored. Parameters and model size MS (Model Size) are directly related to the structural complexity and computational requirements of the model. Smaller parameters and GFLOPs mean that the model is lighter, which can reduce computing resource consumption and improve computing speed without sacrificing too much accuracy. In addition, detection speed (FPS, frames per second) is a direct indicator of the real-time performance of the algorithm. It is generally believed that when FPS reaches or exceeds 30, the algorithm can meet the needs of real-time detection. The higher the FPS value, the more efficient the algorithm is on the given hardware, and the faster it can complete the detection task and reduce latency.
[0131] In order to verify the effectiveness of the FE-YOLOv8 method, an objective evaluation and visualization comparison analysis were performed with several current mainstream target detection methods (YOLOv5m, YOLOv8m and YOLOv10) under the same test set and experimental environment. The results are shown in Tables 1 and Figure 7 shown.
[0132] Table 1 Experimental comparison data of different methods
[0133] method mAP / % mmAP / % R / % Para / M MS / MB FPS frames / s YOLOv5m 72.17 49.02 64.69 22.21 25.98 167.94 YOLOv10l 75.21 50.17 67.05 24.41 99.55 95.16 YOLOv8m 76.08 51.94 70.01 25.93 50.29 132.21 FE-YOLOv8 79.36 53.39 72.17 26.04 72.74 118.37
[0134] As can be seen from Table 1, the method EC-RTDETR proposed in the embodiment of the present invention is the best in the mAP, mmAP and R3 evaluation indicators, with values of 79.36%, 53.39% and 73.17% respectively. Since StarBlock is used in the backbone, the DWR module is used to expand the receptive field to improve performance, the PSFM module is added to the Neck part, and the RTDETR detection head is used in the Head part, the number of parameters and model size increase, the detection speed decreases, but the overall real-time foggy target detection requirements are met. Compared with the benchmark YOLOv8m, the mAP of FE-YOLOv8 is increased by 3.28%, the mmAP is increased by 1.45%, and the recall rate R is increased by 2.16%, indicating that FE-YOLOv8 can complete the foggy target detection task with high accuracy, and has certain superiority and comparability in the comparison algorithm.
[0135] The visual comparison results are as follows Figure 7 As shown, three test images in the RTTS dataset are randomly selected for comparison. Figure 7 (a) shows the simulation results obtained by YOLOv5m. Figure 7 (b) shows the simulation results obtained by YOLOv10l. Figure 7 (c) shows the simulation results obtained by YOLOv8m. Figure 7(d) shows the simulation results of FE-YOLOv8 constructed by the embodiment of the present invention. In the first column, YOLOv5m, YOLOv101 and YOLOv8m all have missed detection phenomenon, and the Bus in the middle of the image is not detected. The FE-YOLOv8 algorithm of the embodiment of the present invention correctly detects the Bus category. In the second column, there are three pedestrians in the woods far away in the park in the background of the image. The third column is a partial enlarged view of the second column. It can be seen that YOLOv5m and YOLOv8m both have missed detection phenomenon, and YOLOv101 and FE-YOLOv8 algorithms both correctly detect pedestrians in the background. Moreover, comparing the detection results of YOLOv101 and FE-YOLOv8, the FE-YOLOv8 algorithm proposed in the embodiment of the present invention has higher detection accuracy.
[0136] In summary, the embodiment of the present invention proposes an improved method FE-YOLOv8 (Feature extraction fusion Enhanced YOLOv8 foggy target detection) for target detection under foggy weather conditions. By introducing StarBlock and DWR modules to optimize Backbone, the feature extraction capability of foggy images is enhanced, so that the model can capture more foggy image features, and at the same time reduce network parameters, so that the network can meet the real-time foggy target detection. In the Neck part, a PSFM module based on cross attention is proposed, which improves the feature fusion effect by fusing different depths and shallow layers of foggy features, enables the target and environmental features to interact more deeply, reduces the phenomenon of missing small background targets, and enhances the adaptability of the network to the complex environment of foggy weather. Finally, inspired by the Transformer detection head in RTDETR, the Head part is improved so that FE-YOLOv8 can more flexibly handle targets of different shapes and sizes when detecting small targets in foggy weather, and directly predict the bounding box and category of the object. This anchor-free detection method not only reduces the computational effort of the network, but also removes the fault tolerance of the anchor box that may not be applicable to targets of all sizes and shapes, thus improving the performance of small target detection in foggy backgrounds, and thus improving the detection speed and accuracy of targets in foggy backgrounds. Experiments show that compared with the baseline YOLOv8, the mAP of FE-YOLOv8 is improved by 3.28%, the mmAP is improved by 1.45%, and the recall rate R is improved by 2.16%, indicating that FE-YOLOv8 can complete the task of foggy target detection with high accuracy, meeting the needs of target detection in actual foggy scenes.
[0137] See also Figure 2 The embodiment of the present application also provides a YOLOv8 foggy weather target detection system based on feature extraction fusion enhancement, which can implement the above-mentioned YOLOv8 foggy weather target detection method based on feature extraction fusion enhancement, and the system includes:
[0138] The first module 201 is used to obtain a foggy image dataset to be identified;
[0139] The second module 202 is used to introduce a star module, an expansion residual module, and a deep semantic fusion module based on cross attention based on the YOLOv8 network model, and adopt an anchor-free frame recognition method to build a FE-YOLOv8 network model;
[0140] The third module 203 is used to perform image target recognition and detection processing on the foggy image dataset to be recognized based on the FE-YOLOv8 network model to obtain foggy image target recognition and detection results.
[0141] It can be understood that the contents of the above method embodiments are all applicable to the present system embodiments, the functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0142] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.
Claims
1. The YOLOv8 foggy target detection method based on feature extraction fusion enhancement is characterized by: The method comprises the following steps: Obtain a dataset of foggy images to be identified; Based on the YOLOv8 network model, the star module, the dilated residual module, and the deep semantic fusion module based on cross attention are introduced, and the anchor-free box recognition method is adopted to construct the FE-YOLOv8 network model; Based on the FE-YOLOv8 network model, image target recognition and detection processing is performed on the foggy image data set to be identified to obtain a foggy image target recognition and detection result.
2. The method according to claim 1, characterized in that The FE-YOLOv8 network model includes a backbone feature extraction network module, a neck feature fusion network module and a head detection network module. The head detection network module adopts the anchor-free frame recognition method. The backbone feature extraction network module, the neck feature fusion network module and the head detection network module are connected in sequence, wherein: The backbone feature extraction network module includes a convolution batch normalization activation module, a star module, a dilated residual module and a spatial pyramid pooling module, wherein the star module includes a first depth-separable convolution layer, a first fully connected layer, a second fully connected layer, a third fully connected layer and a second depth-separable convolution layer, and the dilated residual module includes a regional residual module, a multi-rate dilated depth convolution module, a batch normalization module and a point-by-point convolution fusion module; The neck feature fusion network module includes a downsampling module, a feature splicing module, a spatial channel feature fusion module, a deep semantic fusion module based on cross attention, and a convolution batch normalization activation module, wherein the deep semantic fusion module based on cross attention includes a dense layer, a linear projection layer, a KQV embedding vector calculation layer, an activation layer, a splicing layer, and a convolution layer; The head detection network module includes an anchor-free detection head module, an IoU-aware query selection module and an auxiliary detection head decoder module.
3. The method according to claim 2, characterized in that The method of performing image target recognition and detection processing on the foggy image dataset to be identified based on the FE-YOLOv8 network model to obtain a foggy image target recognition and detection result includes: Inputting the foggy image data set to be identified into the FE-YOLOv8 network model; Based on the backbone feature extraction network module of the FE-YOLOv8 network model, feature extraction processing is performed on the foggy image data set to be identified to obtain a foggy feature image; Based on the neck feature fusion network module of the FE-YOLOv8 network model, feature fusion processing is performed on the foggy feature image to obtain a foggy feature fusion image; Based on the head detection network module of the FE-YOLOv8 network model, anchor-free frame recognition and detection processing is performed on the foggy feature fusion image to obtain the foggy image target recognition and detection result.
4. The method according to claim 3, characterized in that The backbone feature extraction network module based on the FE-YOLOv8 network model performs feature extraction processing on the foggy image data set to be identified to obtain a foggy feature image, including: Inputting the foggy image data set to be identified into the backbone feature extraction network module of the FE-YOLOv8 network model; Based on the convolution batch normalization activation module of the backbone feature extraction network module, feature extraction processing is performed on the foggy image dataset to be identified to obtain a preliminary foggy feature image; Based on the star module of the backbone feature extraction network module, the preliminary foggy weather feature image is subjected to feature mapping processing to obtain a mapped foggy weather feature image; Based on the expansion residual module of the backbone feature extraction network module, the mapped foggy feature image is subjected to feature interactive fusion processing to obtain an interactively fused foggy feature image; Based on the spatial pyramid pooling module of the backbone feature extraction network module, multi-scale pooling processing is performed on the interactively fused foggy feature image to obtain the foggy feature image.
5. The method according to claim 4, characterized in that The star module based on the backbone feature extraction network module performs feature mapping processing on the preliminary foggy weather feature image to obtain a mapped foggy weather feature image, including: Inputting the preliminary foggy weather feature image into the star module of the backbone feature extraction network module; Based on the first depthwise separable convolutional layer of the star module, high-dimensional feature extraction processing is performed on the preliminary foggy weather feature image to obtain a high-dimensional foggy weather feature image; Based on the first fully connected layer and the second fully connected layer of the star-shaped module, linear combination processing is performed on the high-dimensional foggy weather feature image to obtain a first high-dimensional foggy weather feature image and a second high-dimensional foggy weather feature image; Performing matrix dot product calculation on the first high-dimensional foggy weather feature image and the second high-dimensional foggy weather feature image to obtain a calculated high-dimensional foggy weather feature image; Based on the third fully connected layer of the star-shaped module, nonlinear mapping is performed on the calculated high-dimensional foggy weather characteristic image to obtain a foggy weather characteristic image after preliminary mapping; Based on the second depthwise separable convolutional layer of the star-shaped module, feature extraction processing is performed on the initially mapped foggy feature image to obtain a mapped foggy feature image.
6. The method according to claim 4, characterized in that The dilation residual module based on the backbone feature extraction network module performs feature interactive fusion processing on the mapped fog feature image to obtain the interactively fused fog feature image, including: Inputting the mapped foggy feature image into the dilation residual module of the backbone feature extraction network module; Based on the regional residual module of the dilated residual module, a preliminary target feature extraction process is performed on the mapped foggy feature image to obtain a target feature map with different foggy foregrounds; Based on the multi-rate dilated deep convolution module of the dilated residual module, morphological filtering is performed on the feature map of targets with different foggy foregrounds to obtain a filtered feature map of targets with different foggy foregrounds; Based on the batch normalization module of the dilated residual module, the filtered feature map with different foggy foreground targets is standardized to obtain a standardized feature map with different foggy foreground targets; Based on the point-by-point convolution fusion module of the dilated residual module, the standardized feature map with different foggy foreground targets and the mapped foggy feature image are subjected to feature interactive fusion processing to obtain the interactively fused foggy feature image.
7. The method according to claim 3, characterized in that The neck feature fusion network module based on the FE-YOLOv8 network model performs feature fusion processing on the foggy feature image to obtain a foggy feature fusion image, including: Inputting the foggy weather feature image into the neck feature fusion network module of the FE-YOLOv8 network model; Based on the downsampling module of the neck feature fusion network module, the foggy weather feature image is downsampled to obtain a downsampled foggy weather feature image; Based on the feature stitching module of the neck feature fusion network module, feature stitching is performed on the downsampled fog feature image to obtain a preliminary stitched fog feature image; Based on the spatial channel feature fusion module of the neck feature fusion network module, the preliminary spliced foggy feature image is subjected to spatial channel feature fusion processing to obtain a spliced foggy feature image; Based on the cross-attention based deep semantic fusion module of the neck feature fusion network module, the spliced foggy feature image is subjected to deep semantic fusion processing to obtain a preliminary foggy feature fusion image; Based on the convolution batch normalization activation module of the neck feature fusion network module, the preliminary foggy feature fusion image is subjected to normalization mapping processing to obtain the foggy feature fusion image.
8. The method according to claim 7, characterized in that The cross-attention based deep semantic fusion module based on the neck feature fusion network module performs deep semantic fusion processing on the spliced foggy feature image to obtain a preliminary foggy feature fusion image, including: Inputting the spliced foggy feature image into the cross-attention based deep semantic fusion module of the neck feature fusion network module; Based on the dense layer of the cross-attention based deep semantic fusion module, the spliced foggy feature image is enhanced to obtain an enhanced foggy feature image; Based on the linear projection layer of the cross-attention based deep semantic fusion module, projecting the enhanced foggy feature image to obtain a projected foggy feature image; Based on the KQV embedding vector calculation layer of the cross-attention based deep semantic fusion module, KQV calculation is performed on the projected foggy feature image to obtain a feature vector of the foggy image; Reconstructing the feature vector of the foggy image and the spliced foggy feature image based on the activation layer and the splicing layer of the cross-attention based deep semantic fusion module to obtain a reconstructed foggy feature fusion image; Based on the convolution layer of the cross-attention based deep semantic fusion module, the reconstructed foggy feature fusion image is convolved to obtain the preliminary foggy feature fusion image.
9. The method according to claim 3, characterized in that: The head detection network module based on the FE-YOLOv8 network model performs anchor-free recognition detection processing on the foggy feature fusion image to obtain the foggy image target recognition detection result, including: Inputting the foggy feature fusion image into the head detection network module of the FE-YOLOv8 network model; Based on the anchor-free detection head module of the head detection network module, target information extraction processing is performed on the foggy feature fusion image to obtain predicted foggy image target feature information; Based on the IoU perception query selection module of the head detection network module, the predicted foggy image target feature information is screened to obtain the screened predicted foggy image target feature information; Based on the auxiliary detection head decoder module of the head detection network module, target detection processing is performed on the filtered predicted foggy image target feature information to obtain the foggy image target recognition detection result.
10. The YOLOv8 foggy target detection system based on feature extraction fusion enhancement is characterized by: The system comprises: The first module is used to obtain a foggy image dataset to be identified; The second module is used to build a FE-YOLOv8 network model based on the YOLOv8 network model by introducing a star module, an expansion residual module, and a deep semantic fusion module based on cross attention and adopting an anchor-free frame recognition method; The third module is used to perform image target recognition and detection processing on the foggy image data set to be identified based on the FE-YOLOv8 network model to obtain a foggy image target recognition and detection result.
Citation Information
Cited By
Method, system and equipment for detecting and counting rice blast
CN121414759A
Self-adaptive full-scale infrared target detection network based on YOLO
CN121725341A