Haze scene road multi-target detection method based on multi-stack space channel cooperation and multi-scale feature fusion
By improving the YOLO11 algorithm and combining the multi-stack spatial channel collaboration and multi-scale feature fusion methods, the problems of environmental adaptability and detection accuracy of multi-target detection on roads in haze environments are solved, and efficient and real-time multi-target detection on roads in haze scenes is achieved.
Patent Information
- Application Number
- CN202510836557.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-21
- Publication Date
- 2025-10-03
AI Technical Summary
Existing technologies for multi-target detection on roads in haze environments have limitations in environmental adaptability and detection range, insufficient dynamic target detection accuracy, and robustness in complex scenarios. Especially when haze and complex road backgrounds overlap, it is difficult to strike a balance between detection accuracy and real-time performance.
A multi-target detection method for roads in haze scenes is adopted based on multi-stacked spatial channel collaboration and multi-scale feature fusion. By improving the YOLO11 algorithm, designing the MSSCA module and SPPF-AIFI module, and combining the atmospheric scattering model to construct a hybrid dataset, the DFI-YOLO11 model is trained to improve the detection accuracy and recall rate.
In haze environments, the accuracy and recall rate of the detection model are improved, the false detection rate is reduced, the operating efficiency and real-time performance of the algorithm are maintained, and it is suitable for complex environment detection in different fog concentration scenes.
Smart Images

Figure CN120747464A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent traffic monitoring technology, and in particular to a multi-target detection method for roads in haze scenes based on multi-stacked spatial channel collaboration and multi-scale feature fusion. Background Art
[0002] In the field of intelligent traffic monitoring, achieving accurate, real-time detection of multiple road targets (vehicles, pedestrians, traffic facilities, etc.) in haze-induced scenarios is an extremely challenging task. Current mainstream detection algorithms have significant limitations in their recognition capabilities in complex environments, particularly in scenarios such as haze-induced low visibility, blurred image features, variable target scales, vehicle occlusion, and complex road backgrounds. This makes it difficult to achieve both accurate and real-time detection of multiple road targets. Therefore, there is an urgent need to develop efficient and accurate technical means to address these challenges and improve the application of road target detection technology in haze-induced environments within intelligent transportation systems.
[0003] Road multi-target detection technology is a core pillar of modern transportation infrastructure security, directly impacting traffic safety, traffic flow management, and the reliability of intelligent driving systems. Traditional detection solutions primarily rely on road radar, LiDAR, and conventional cameras combined with traditional image processing techniques.
[0004] In hazy environments, atmospheric particulate matter can severely interfere with the propagation of electromagnetic waves and laser signals, significantly shortening detection distances and increasing false alarm rates. When using conventional cameras with traditional algorithms, haze-induced issues such as reduced image contrast and blurred edge features can lead to ineffective target feature extraction and a sharp decline in detection performance.
[0005] The YOLO (You Only Look Once) family of algorithms is a mainstream target detection solution in computer vision. Versions such as YOLOv5, YOLOv8, and YOLOv10 have demonstrated excellent performance in common scenarios. While these algorithms have overcome some of the shortcomings of traditional technologies, they still face significant limitations when dealing with foggy and smoggy environments. Low-resolution images (caused by smog), blurred target boundaries, complex road scenes, and vehicle occlusions significantly reduce detection effectiveness. In particular, for targets like vehicle taillights and pedestrian silhouettes that are translucent in smog and highly obscured from the background, the detection accuracy of existing algorithms is insufficient to meet practical requirements, necessitating targeted optimization solutions.
[0006] The existing technology has the following core problems:
[0007] 1. Environmental adaptability and limited detection range: Traditional detection methods are highly sensitive to environmental factors such as haze concentration and light changes. The sensor's signal attenuation is severe under strong haze conditions, resulting in a short detection range. Furthermore, the detection stability is poor under different haze levels and is easily affected by environmental interference, resulting in false alarms.
[0008] 2. Insufficient dynamic target detection accuracy: When processing haze-blurred images, the lightweight YOLO algorithm has insufficient feature extraction capabilities for small targets (such as distant pedestrians and motorcycles) and partially occluded targets. The superposition of motion blur of dynamic targets and haze blur significantly reduces the detection recall rate.
[0009] 3. Robustness defects in complex scenarios: Both traditional solutions and the YOLO algorithm are prone to confusion between target and background features when haze and complex road backgrounds are superimposed. Summary of the Invention
[0010] To address the aforementioned technical issues, this paper proposes a multi-target road detection method in foggy and hazy scenes based on the synergy of multiple stacked spatial channels and the fusion of multi-scale features. This method achieves high recognition accuracy for target detection in scenes with varying fog concentrations. While addressing the aforementioned technical issues and challenges, it maximizes the model's accuracy and recall, and reduces the false positive rate, while maintaining a certain level of algorithmic efficiency and real-time performance. This provides a superior detection method for multi-target road detection in complex environments.
[0011] In order to achieve the above object, the present invention adopts the following technical solutions:
[0012] A multi-target detection method for roads in hazy scenes based on multi-stacked spatial channel collaboration and multi-scale feature fusion includes the following steps:
[0013] Step S1: Collect and obtain the original images of the road fog and haze scene dataset, and use the atmospheric scattering physical model to construct the Hybrid_FogData hybrid road fog and haze scene dataset;
[0014] Step S2: Design a multi-stacked spatial and channel attention mechanism called MSSCA, propose the SPPF-AIFI intra-scale feature interaction module, use SPD-Conv to improve the traditional convolutional layer, and propose a new DFI-YOLO11 detection network architecture based on the above modules;
[0015] Step S3: The model is trained using the mixed road haze scene dataset constructed in step S1 to obtain the trained DFI-YOLO11 model weights;
[0016] Step S4: Use the trained DFI-YOLO11 network model to detect the image to be detected, and output the detection results including the location, category and confidence of pedestrians and vehicles;
[0017] Furthermore, the specific construction process of step 1 is as follows:
[0018] 5000 original images are randomly selected from 8240 images in the Voc_2012 dataset as input. Through the atmospheric scattering model, the β in the transmittance t(x) is adjusted to control the fog concentration. In this study, β is controlled in the range of 0.01 to 0.1, and the output is a randomly generated haze image.
[0019] The atmospheric scattering model representation of the mixed data set is as follows:
[0020] I(x)=J(x)e -βd(x) +A(1-e -βd(x) );
[0021] Where I(x) represents the synthesized foggy image, J(x) represents the fog-free image, A represents the atmospheric light value, β represents the atmospheric scattering coefficient, and d(x) represents the scene depth.
[0022] The RTTS (Real-world Task-driven Testing Set) real-world dataset was used as a test set to verify the training results, ultimately forming a mixed image dataset with a data volume of 12,562;
[0023] Use the Labelme annotation tool to annotate the collected original images and create a YOLO format mixed road haze scene dataset and corresponding labels;
[0024] The key to the multi-target detection method for road in haze scenes based on multi-stacked spatial channel collaboration and multi-scale feature fusion described in the present invention is to select a preferred implementation method for the problem to be solved. In step S2, the original algorithm is improved to design an original MSSCA module and an SPPF-AIFI intra-scale feature interaction module. Based on the above modules, a new DFI-YOLO11 detection network architecture is proposed. The steps are as follows:
[0025] In the YOLO11 feature extraction backbone network of step S2, the newly designed MSSCA multi-stacked spatial and channel attention mechanism is used. By deeply fusing spatial attention and channel attention, a multi-stacked strategy is adopted to pass part of the channels of the input feature map into multiple consecutive modules, thereby achieving layer-by-layer refinement and enhancement of spatial attention;
[0026] The key to this module is the original structural improvement of the original SCSA module. It adopts an asymmetric processing strategy, divides the input feature map into two branches in the channel dimension, applies deep attention operations only to a part of the channels, and uses identity mapping or lightweight convolution operations to preserve the original features in the shallow branch.
[0027] SCSA consists of a shared multi-semantic spatial attention (SMSA) and a progressive channel-wise self-attention (PCSA);
[0028] Input feature map X∈R in SMSA B×C×H×W Decomposed into two one-dimensional sequences X in height and width directions H and X W , and then normalize to generate the spatial attention map. The main calculation formula is as follows:
[0029]
[0030]
[0031] in, Represents the spatial structure information of the i-th sub-feature obtained after the lightweight convolution operation, k i represents the convolution kernel applied to the i-th sub-feature;
[0032]
[0033]
[0034] SMSA(X)=X s =Attn H ×Attn W ×X
[0035] σ(·) represents the normalized Sigmoid function, and Respectively, they represent group normalization with K groups in the height (H) and width (W) dimensions;
[0036] The PCSA module uses a progressive compression strategy and a channel self-attention mechanism to optimize channel features. It pools and compresses input features, and a single-head self-attention mechanism calculates the similarity between channels. The main calculation formula is as follows:
[0037]
[0038]
[0039]
[0040]
[0041]
[0042] represents a pooling operation with a kernel size of k×k, which rescales the resolution from (H,W) to (H′,W′). proj (·) represents the linear projection used to generate query, key, and value;
[0043] Unlike the traditional attention mechanism that processes all channels uniformly, this design improvement divides the input feature map into two parts along the channel dimension. The first half of the channel is processed by n layers of continuous LSABlock, and the second half of the channel remains unchanged or undergoes a lightweight convolution transformation;
[0044] This partial deep attention mechanism can effectively control computational complexity and only perform deep processing on key channels;
[0045] By extracting spatial attention information hierarchically, it improves sensitivity to fine-grained features;
[0046] The input feature map is Divide it into and The main calculation formula is as follows:
[0047] Y1=LSA n (X1),Y2=Conv1×1(X2)
[0048] Y=Concat(Y1,Y2),Output=Conv fusion (Y)
[0049] Among them, LSA n (·) represents the LSABlock operation of stacking n layers continuously. Y processes the input feature map into two parts and then splices them together to integrate the features obtained by the two parts with different processing methods. Output outputs the final features after performing a fused convolution operation on Y;
[0050] Improve the MSSCA multi-stacked spatial and channel attention mechanism into the C2PSA feature extraction module in the backbone network to improve the model's feature expression capability in complex scenarios;
[0051] Improve the backbone network using SPD-Conv spatial pyramid pooling and depthwise separable convolution, and connect the improved structure with the above-mentioned improved output and input channels and the detection head;
[0052] SPD-Conv does not use strided convolutional layers and pooling layers, but instead uses spatial-to-depth (SPD) layers and non-strided convolutional layers. It downsamples the feature maps while retaining all information in the channel dimension, effectively eliminating information loss and having the advantage of capturing multi-scale information that is critical for small object detection.
[0053] This paper combines the Transformer Encoder in RT-DETR to improve the SPPF (Spatial Pyramid Pooling Fast) module and proposes an SPPF-AIFI intra-scale feature interaction module.
[0054] The module's global information modeling function enables the model to more effectively utilize the environmental pixels around the target to reduce false detections;
[0055] By focusing on the internal scale interactions at the S5 level (i.e., the high-level feature layer), the model can more accurately distinguish objects and calculate the global context association to encode the features;
[0056] Multiple self-attention heads process features in parallel, completing dimensionality transformation through cascaded outputs and a learnable weight matrix. This effectively compensates for the shortcomings of CNN local feature modeling without significantly increasing the computational load, providing richer global contextual information support for road object detection in haze scenes.
[0057] The key to the multi-target detection method for road surfaces in haze scenes based on multi-stacked spatial channel collaboration and multi-scale feature fusion described in the present invention is to select a preferred implementation method for the problem being solved. In step S3, based on the proposed feature extraction network structure, the optimal training weights of the DFI-YOLO11 model for road haze scene detection are obtained by training on a constructed mixed road haze scene dataset.
[0058] Use the training set and model configuration file to train the model. Adjust the training parameter settings during the training process to ultimately train and output the optimal multi-target detection model for roads in foggy and hazy scenes. Specifically, the following steps are involved:
[0059] Adjust the batch size Batch_Size to 64 to increase the model training speed;
[0060] The initial learning rate is set to 1×10 -3 , weight decay is 5×10 -4 ;
[0061] The cosine annealing decay strategy is used to adjust the learning rate, and Mosaic online data enhancement is turned off in the last 10 epochs;
[0062] Other hyperparameter settings except the above adjustments remain default;
[0063] The present invention discloses a method for detecting multiple targets on roads in haze scenes based on the synergy of multiple stacked spatial channels and the fusion of multi-scale features, characterized in that, in step S4, the specific process is as follows: Figure 2 As shown, the steps are as follows:
[0064] Based on the training results on the test set, we conducted comparative experiments with baseline and cutting-edge models in terms of parameter size, GFLOPs, model memory size, and detection metrics such as mAP50, mAP95, Precision, and Recall. The DFI-YOLO11 model for road multi-object detection in foggy and hazy scenes was trained to achieve the best results.
[0065] Specifically, based on the DFI-YOLO11 network model and trained network weights, target detection is performed on road images of foggy scenes. The network performs a forward inference, and the detection heads (including three standard detection heads and a tiny object detection head) output the bounding box coordinates and category of the detected targets.
[0066] The above technical solution proposed by the present invention can obviously solve the problems of the prior art and has the following beneficial effects:
[0067] We selected the single-stage target detection algorithm YOLO11 from the YOLO series as a benchmark and improved the original YOLO11 algorithm to obtain an improved DFI-YOLO11 network structure. We constructed a dataset to address existing problems and model training needs. We then used the dataset to train the model and conduct comparative experiments. We obtained a DFI-YOLO11 model for multi-target detection on roads in foggy and hazy scenes, which showed significant improvements in various indicators compared to existing technologies and benchmark algorithms. The specific optimization and improvement steps are as follows:
[0068] To train a model that is more robust to generalized scenarios, we constructed a mixed road haze scene dataset using an atmospheric scattering model. We used the Labelme annotation tool to annotate the collected original images and create a mixed road haze scene dataset in YOLO format and the corresponding labels.
[0069] The designed MSSCA multi-stacked spatial and channel attention mechanism module is used to improve the backbone network of YOLO11, achieving layer-by-layer refinement and enhancement of spatial attention without increasing the computational complexity and parameter count of the baseline network.
[0070] The backbone network is improved by using SPD-Conv spatial pyramid pooling and depthwise separable convolution, which downsamples the feature maps while retaining all information in the channel dimension, effectively eliminating information loss and having advantages in capturing multi-scale information that is critical for small object detection.
[0071] Combining the Transformer Encoder in RT-DETR with the improved SPPF (Spatial Pyramid Pooling Fast) module, a SPPF-AIFI intra-scale feature interaction module is proposed, which effectively makes up for the shortcomings of CNN local feature modeling and provides richer global context information support for road object detection in haze scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 : Overall flow chart of the road target detection method based on foggy scenes;
[0073] Figure 2 : Flowchart of an example implementation of the present invention;
[0074] Figure 3 : Example diagram of some detection results during the DFI-YOLO11 network training process;
[0075] Figure 4 : Internal structure diagram of the MSSCA multi-stacked spatial and channel attention mechanism module;
[0076] Figure 5 : SPD-Conv spatial pyramid pooling and depth-wise separable convolution internal structure diagram;
[0077] Figure 6 :Internal structure diagram of SPPF-AIFI scale feature interaction module
[0078] Figure 7 : The final network architecture diagram of the training model; DETAILED DESCRIPTION
[0079] In order to enable those skilled in the art to better understand the technical solutions in this application, the following will clearly and completely describe the technical solutions in the embodiments of this application in conjunction with the drawings in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0080] Example 1
[0081] like Figure 1As shown in FIG, a multi-target detection method for roads in haze scenes based on multi-stacked spatial channel collaboration and multi-scale feature fusion specifically includes the following steps:
[0082] Step S1 acquires images and constructs a hybrid fog scene dataset Hybrid_FogData:
[0083] Since there are currently few public datasets for target detection in hazy weather, in order to achieve ideal detection performance under both normal and hazy weather conditions, the present invention uses an atmospheric scattering model to generate fog of different concentrations and simulates a dataset to expand the number of training samples.
[0084] 5000 original images are randomly selected from 8240 images in the Voc_2012 dataset as input. Through the atmospheric scattering model, the β in the transmittance t(x) is adjusted to control the fog concentration. In this study, β is controlled in the range of 0.01 to 0.1, and the output is a randomly generated haze image.
[0085] The RTTS real-world dataset was used as a test set to verify the training results, ultimately forming a mixed image dataset with a data volume of 12,562;
[0086] Step S2 improves the original algorithm to obtain a new DFI-YOLO11 detection network:
[0087] In the YOLO11 feature extraction backbone network in step S2, the newly designed MSSCA multi-stacked spatial and channel attention mechanism is used. By deeply fusing spatial attention and channel attention, a multi-stacked strategy is adopted to pass a portion of the input feature map's channels into multiple consecutive modules, thereby achieving layer-by-layer refinement and enhancement of spatial attention. The key to this module is its asymmetric processing strategy, which divides the input feature map into two branches along the channel dimension. Deep attention is applied to only a portion of the channels, while the shallow branch uses identity mapping or lightweight convolution operations to preserve the original features.
[0088] like Figure 4 As shown in Figure 2, SCSA consists of a shared multi-semantic spatial attention (SMSA) and a progressive channel-wise self-attention (PCSA):
[0089] Input feature map X∈R in SMSA B×C×H×W Decomposed into two one-dimensional sequences X in height and width directions H and X W , and then normalize to generate the spatial attention map:
[0090]
[0091]
[0092] SMSA(X)=X s =Attn H ×Attn W ×X
[0093] The PCSA module uses a progressive compression strategy and a channel self-attention mechanism to optimize channel features. The input features are pooled and compressed, and the single-head self-attention mechanism calculates the similarity between channels:
[0094]
[0095]
[0096]
[0097]
[0098] This design divides the input feature map into two parts along the channel dimension. The first half of the channels are processed through n layers of continuous LSABlock, while the second half remains unchanged or undergoes a light convolution transformation. This partial deep attention mechanism effectively controls computational complexity, performing deep processing only on key channels. It also extracts spatial attention information in layers, improving sensitivity to fine-grained features.
[0099] The input feature map is Divide it into and The main calculation formula is as follows:
[0100] Y1=LSA n (X1),Y2=Conv1×1(X2)
[0101] Y=Concat(Y1,Y2),Output=Conv fusion (Y)#(13)
[0102] The input feature map is processed into two parts and then spliced together to integrate the features obtained by the two different processing methods. The final feature is output after the fusion convolution operation on Y;
[0103] In the YOLO11 network in step S2, the SPD-Conv spatial pyramid pooling and depth-separable convolution are used to improve the convolution layers in the backbone and neck networks. Figure 5 As shown in Figure 2, while downsampling the feature map, all information in the channel dimension is retained, effectively eliminating information loss and having advantages in capturing multi-scale information that is crucial for small target detection.
[0104] By combining the Transformer Encoder in RT-DETR to improve the SPPF (SpatialPyramid Pooling Fast) module, a SPPF-AIFI intra-scale feature interaction module is proposed. Figure 6 As shown in Figure 2, the module's global information modeling capability enables the model to more effectively utilize the surrounding pixels around the target to reduce false detections. Focusing on internal scale interactions at the S5 level (i.e., the high-level feature layer) helps the model distinguish targets more accurately and calculates global contextual associations to encode features.
[0105] Multiple self-attention heads process features in parallel, completing dimensionality transformation through cascaded outputs and a learnable weight matrix. This effectively compensates for the shortcomings of CNN local feature modeling without significantly increasing the computational load, providing richer global contextual information support for road object detection in haze scenes.
[0106] S3 performs model pre-training for the proposed DFI-YOLO11 network structure:
[0107] Based on the proposed feature extraction network structure, the mixed road haze scene dataset constructed in step S1 is trained. Some results of the training process are shown in the following figure. Figure 3 As shown, the optimal training weights for road haze scene detection are obtained;
[0108] Use the training set and model configuration file to train the model. Adjust the training parameter settings during the training process to ultimately train and output the optimal multi-target detection model for roads in foggy and hazy scenes. Specifically, the following steps are involved:
[0109] Adjust the batch size Batch_Size to 64 to increase the model training speed;
[0110] The initial learning rate is set to 1×10 -3 , weight decay is 5×10 -4 ;
[0111] The cosine annealing decay strategy is used to adjust the learning rate, and Mosaic online data enhancement is turned off in the last 10 epochs;
[0112] Other hyperparameter settings except the above adjustments remain default;
[0113] During the training process of the mixed road haze scene dataset, the hardware experimental environment configuration is as follows:
[0114] Table 1 Experimental environment configuration
[0115] S4 Result Analysis and Application: The final network used to train the model is as follows Figure 7 As shown in the figure, the trained DFI-YOLO11 best model weights are deployed to mobile storage devices. From the comparative experiments, it can be seen that DFI-YOLO11 has the best mAP when the IoU intersection-over-union ratio is 0.5 and 0.95. 50 、mAP 95 The average precision, accuracy and recall rate are higher than other models. At the same time, its parameter amount, computational complexity and model size are better than the existing state-of-the-art methods. It can effectively perform real-time detection, indicator analysis, category classification and monitoring processing on pictures, videos and real-time image streams collected by cameras connected to mobile devices. The specific steps of step S4 are:
[0116] The image data is input into the DFI-YOLO11 optimal model obtained in steps S2 and S3 to monitor the road haze scene and determine whether there is a target in the given image stream;
[0117] Based on comparative experiments on the training results of the test set with the parameters, GFLOPs, and model memory size of baseline and cutting-edge models, as well as detection indicators such as mAP50, mAP95, Precision, and Recall, the best multi-target detection model for road in haze scenes based on multi-stacked spatial channel collaboration and multi-scale feature fusion was trained.
[0118] The training results and comparative experiments on the validation set are as follows:
[0119] Table 2 Ablation comparison experiment results
[0120] Example 2
[0121] This embodiment provides an electronic device, including a processor and a memory, wherein the memory stores a program that can be run on the above-mentioned processor. The processor is used to perform deep learning calculations and program execution. The processor used in the embodiment runs based on Python-3.9.7 and the deep learning development environment pytorch-2.1.0. When the program is executed by the above-mentioned processor, the steps of the above-mentioned embodiment of the method for multi-target detection on roads in haze scenes based on multi-stacked spatial channel collaboration and multi-scale feature fusion are implemented.
[0122] Example 3
[0123] This embodiment provides a computer-readable storage medium storing at least one program. The at least one program is executable by at least one processor to implement the steps of the method for multi-target road detection in haze scenes based on multi-stacked spatial channel collaboration and multi-scale feature fusion described in the above embodiment. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory, a random access memory, a magnetic disk, or an optical disk.
[0124] The above is a detailed introduction to the multi-target detection method for roads in haze scenes based on multi-stacked spatial channel collaboration and multi-scale feature fusion proposed in the present invention, and the principles and implementation methods of the present invention are explained. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those skilled in the art, according to the ideas of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A multi-target detection method for roads in haze scenes based on multi-stacked spatial channel collaboration and multi-scale feature fusion, characterized by: The following steps are involved: Step S1: Collect and obtain the original images of the mixed road haze scene dataset, filter, annotate and divide the collected images according to proportion; Step S2: Design a multi-stacked spatial and channel attention mechanism called MSSCA, propose the SPPF-AIFI intra-scale feature interaction module, use SPD-Conv to improve the traditional convolutional layer, and propose a new DFI-YOLO11 detection network architecture based on the above modules; Step S3: The model is trained using the mixed road haze scene dataset constructed in step S1 to obtain the trained DFI-YOLO11 model weights; Step S4: Use the trained DFI-YOLO11 network model to detect the image to be detected, and output the detection results including the location, category and confidence of pedestrians and vehicles.
2. The mixed road haze scene dataset according to claim 1, characterized in that: Here are the steps: 5000 original images are randomly selected from the 8240 images in the Voc_2012 dataset as input. Through the atmospheric scattering model, the β in the transmittance t(x) is adjusted to control the fog concentration. In this study, β is controlled in the range of 0.01 to 0.
1. The output is a randomly generated haze image. The atmospheric scattering model of the mixed dataset is shown as follows: I(x)=J(x)e -βd(x) +A(1-e -βd(x) ) (1) Where I(x) represents the synthesized foggy image, J(x) represents the fog-free image, A represents the atmospheric light value, β represents the atmospheric scattering coefficient, and d(x) represents the scene depth. The RTTS (Real-world Task-driven Testing Set) real-world dataset was used as a test set to verify the training results, ultimately forming a mixed image dataset with a data volume of 12,562; Use the Labelme annotation tool to annotate the collected original images and create a YOLO format mixed road haze scene dataset and corresponding labels.
3. The multi-target detection method for roads in haze scenes based on multi-stacked spatial channel collaboration and multi-scale feature fusion of DFI-YOLO11 according to claim 1 is characterized in that: In the YOLO11 feature extraction backbone network in S2, the designed MSSCA multi-stacked spatial and channel attention mechanism is used. By deeply fusing spatial attention and channel attention, a multi-stacked strategy is adopted to pass part of the channels of the input feature map into multiple consecutive modules, thereby achieving layer-by-layer refinement and enhancement of spatial attention. The key to this module is the original structural improvement of the original SCSA module. It adopts an asymmetric processing strategy, divides the input feature map into two branches in the channel dimension, applies deep attention operations only to a part of the channels, and uses identity mapping or lightweight convolution operations to preserve the original features in the shallow branch. SCSA consists of a shared multi-semantic spatial attention (SMSA) and a progressive channel-wise self-attention (PCSA): Input feature map X∈R in SMSA B×C×H×W Decomposed into two one-dimensional sequences X in height and width directions H and X W , and then normalize to generate the spatial attention map. The main calculation formula is as follows: in, Represents the spatial structure information of the i-th sub-feature obtained after the lightweight convolution operation, k i represents the convolution kernel applied to the i-th sub-feature; SMSA(X)=X s =Attn H ×Attn W ×X (6) σ(·) represents the normalized Sigmoid function, and Respectively, they represent group normalization with K groups in the height (H) and width (W) dimensions; The PCSA module uses a progressive compression strategy and a channel self-attention mechanism to optimize channel features, pool and compress input features, and a single-head self-attention mechanism to calculate the similarity between channels. The main calculation formula is as follows: Represents a pooling operation with a kernel size of k×k, which rescales the resolution from (H,W) to (H ′ ,W ′ ), F proj (·) represents the linear projection used to generate query, key, and value; Unlike the traditional attention mechanism that processes all channels uniformly, this design improvement divides the input feature map into two parts along the channel dimension. The first half of the channel is processed by n layers of continuous LSABlock, and the second half of the channel remains unchanged or undergoes a lightweight convolution transformation; This partial deep attention mechanism can effectively control computational complexity, perform deep processing only on key channels, and extract spatial attention information hierarchically, improving sensitivity to fine-grained features. The input feature map is Divide it into and The main calculation formula is as follows: Y1=LSA n (X1), Y2=Conv1×1(X2) (12) Y=Concat(Y1,Y2), Output=Conv fusion (Y) (13) Among them, LSA n (·) represents the LSABlock operation of stacking n layers continuously. Y processes the input feature map into two parts and then splices them together to integrate the features obtained by the two parts in different processing methods. Output outputs the final features by performing a fused convolution operation on Y.
4. The method for multi-target detection on roads in haze scenes based on multi-stacked spatial channel collaboration and multi-scale feature fusion according to claim 1 is characterized in that: In step S2, the specific process is as follows: Improve the backbone network using SPD-Conv spatial pyramid pooling and depthwise separable convolution, and connect the improved structure with the above-mentioned improved output and input channels and the detection head; SPD-Conv does not use strided convolutional layers and pooling layers, but instead uses spatial-to-depth (SPD) layers and non-strided convolutional layers. It downsamples the feature maps while retaining all information in the channel dimension, effectively eliminating information loss and having the advantage of capturing multi-scale information that is critical for small object detection. This paper combines the Transformer Encoder in RT-DETR to improve the SPPF (SpatialPyramid Pooling Fast) module and proposes an SPPF-AIFI intra-scale feature interaction module. The module's global information modeling capability enables the model to more effectively utilize the surrounding pixels around the target to reduce false detections. It helps the model distinguish targets more accurately by focusing on internal scale interactions at the S5 level (i.e., the high-level feature layer) and calculating global context associations to encode features. Multiple self-attention heads process features in parallel, and complete dimensionality transformation through cascaded output and learnable weight matrix. Without significantly increasing the amount of computation, it effectively makes up for the shortcomings of CNN local feature modeling and provides richer global context information support for road target detection in haze scenes.
5. The method for multi-target detection on roads in haze scenes based on multi-stacked spatial channel collaboration and multi-scale feature fusion according to claim 1 is characterized in that: In step S3, based on the proposed feature extraction network structure, the mixed road haze scene dataset constructed in step S1 is used for training to obtain the optimal training weights of the model for road haze scene detection.
6. The method for multi-target detection on roads in haze scenes based on multi-stacked spatial channel collaboration and multi-scale feature fusion according to claim 1 is characterized in that: In step S4, based on the training results on the test set, a comparative experiment was conducted with the parameters, GFLOPs, and model memory size of the baseline and cutting-edge models, as well as detection indicators such as mAP50, mAP95, Precision, and Recall, to train the DFI-YOLO11 model for multi-target detection on roads in haze scenes with the best results.
7. A detection system for executing the method for multi-target detection on roads in haze scenes based on multi-stacked spatial channel collaboration and multi-scale feature fusion as described in any one of claims 1 to 6, characterized in that: The detection system includes: a construction module, an acquisition and processing module, a training module and a detection output module; The construction module integrates the MSSCA module, SPPF-AIFI module and SPD-Conv module into the original YOLO11 target detection model to construct the DFI-YOLO11 model; The acquisition and processing module collects road image data from different haze concentrations and performs preprocessing; The training module designs a loss function based on the characteristics of the target detection task in haze scenes, performs weighted processing on the detection loss of small targets, selects an optimizer, and trains the DFI-YOLO11 model; The detection output module inputs the image to be detected into the trained DFI-YOLO11 model and outputs the detection results including the location, category and confidence of targets such as pedestrians and vehicles.
8. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium for storing computer instructions, characterized in that: When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.