Foggy-day pedestrian and vehicle detection method based on improved YOLO11

By improving the YOLO11 network, the introduction of dual backbone mechanisms, dense high-level connections and iterative attention characteristics fusion have solved the problem of detection accuracy and lightweight by pedestrians and vehicles in foggy days, achieving high-precision and lightweight detection effects, and are suitable for edge devices.

CN120472429APending Publication Date: 2025-08-12南京理工大学紫金学院
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510602605.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-08-12

Smart Images

  • Figure CN120472429A_ABST
    Figure CN120472429A_ABST
Patent Text Reader

Abstract

The invention discloses a foggy-day pedestrian and vehicle detection method based on improved YOLO11, which comprises the following steps: firstly, acquiring pedestrian and vehicle images in a foggy-day scene, constructing a data set, then constructing a foggy-day pedestrian and vehicle detection model based on improved YOLO11n, training and verifying the constructed model by using the data set, and finally obtaining a foggy-day pedestrian and vehicle detection result. And finally, carrying out pedestrian and vehicle target detection by using the trained and verified foggy day pedestrian and vehicle detection model to obtain a detection result. According to the scheme, improvement is carried out based on an existing YOLO11n network, a double-trunk mechanism, an iAFF mechanism and EUCB up-sampling are introduced, targets such as pedestrians and automobiles in a foggy day scene can be accurately detected, the detected map50 value reaches 56.6% and is improved by 8.1% compared with a yo11n original model, the high precision and the lightweight effect are very remarkable, and the method has a certain application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent traffic target detection, and specifically relates to a method for detecting pedestrians and vehicles in foggy weather based on an improved YOLO11. Background Art

[0002] Fog and haze significantly impact vehicle safety. Statistics show that in low visibility conditions, the risk of rear-end collisions increases by 1.2-2 times compared to normal conditions, with chain reactions being particularly common. Therefore, improving the accuracy of pedestrian and vehicle detection in foggy conditions is crucial for ensuring safe driving.

[0003] Currently, mainstream object detection algorithms are mostly based on various deep learning network structures, such as SSD, Fast-RCNN, Faster-RCNN, and the YOLO series of models. The YOLO series of models excels in comprehensive performance, including detection accuracy and speed, and has become the leading object detection algorithm, widely used in various fields. Since Redmon et al. proposed YOLOv1 in 2016, YOLO has evolved to YOLO11 (1.REDMON J,DIVVALAS,GIRSHICK R,et al.You only look once: unified,real-time object detection[C] / / Proceedings of 2016IEEE Conference on Computer Vision and Pattern Recognition,2016:779-788.). Numerous papers have also improved upon the YOLO series of models for object detection in foggy road traffic safety, achieving considerable success.

[0004] For example, Lai Jingan et al. (2. Lai Jingan, Chen Ziqiang, Sun Zongwei, Pei Qingqi. Lightweight foggy target detection method based on YOLOv5 [J]. Computer Engineering and Applications, 2024, 60(6): 78-88.) proposed a lightweight foggy detection algorithm based on an improved YOLOv5. On the real foggy dataset RTTS, they achieved a map50 value 4.9% higher than the original YOLOv5l model, and the number of parameters was only 54.6% of YOLOv5l. However, its number of parameters is still as high as 25.4M and its GFLOPs is 92.8, which still greatly limits its deployment and application on edge devices. Summary of the Invention

[0005] In view of the above problems, the purpose of the present invention is to provide a method for detecting pedestrians and vehicles in foggy weather based on improved YOLO11.

[0006] The specific technical solutions for achieving the purpose of the present invention are as follows:

[0007] A method for detecting pedestrians and vehicles in foggy weather based on improved YOLO11 includes the following steps:

[0008] Step 1: Collect pedestrian and vehicle images in foggy scenes and build a dataset;

[0009] Step 2: Build a foggy pedestrian and vehicle detection model based on the improved YOLO11n;

[0010] Step 3: Use the dataset to train and verify the constructed model;

[0011] Step 4: Use the trained and verified foggy pedestrian and vehicle detection model to detect pedestrians and vehicles and obtain detection results.

[0012] Compared with the prior art, the present invention has the following beneficial effects:

[0013] This solution is based on the existing YOLO11n network and is improved. First, a dual-trunk mechanism is designed to reduce the problem of information loss during deep network feature extraction. The DHLC composite connection strategy is adopted when the auxiliary branch transfers data to the main branch, so that high-level features can be better transferred to low-level features, gradually expanding the receptive field, thereby enhancing feature expression capabilities and improving detection performance. In addition, the iAFF mechanism is introduced in the Botteleneck jump connection of the C3K2 module, designed as the C3K2_iAFF module. iAFF performs more effective feature fusion of global and local information at the data fusion level, taking into account the detection effects of large and small targets. Finally, the efficient EUCB upsampling method is used in the Neck network of YOLO11 to enhance the smooth integration of information between different levels and stages, and efficiently enhance the expressiveness of feature maps without significantly increasing computational overhead.

[0014] The improved YOLO11n foggy pedestrian and vehicle detection model of this solution provides a high-precision and lightweight foggy pedestrian and vehicle detection method, which can detect pedestrians, cars and other targets in foggy scenes. The detection map50 value reaches 56.6%, which is 8.1% higher than the original yolo11n model and 1.5% higher than the improved YOLOv5l model proposed in the existing technology. The parameter volume is 5.26M and the GFLOPs is 13.8, which are 20.7% and 14.9% of the existing model respectively. The high precision and lightweight effects are very significant.

[0015] The present invention will be further described below with reference to specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1The figure is a flow chart of the method for detecting pedestrians and vehicles in foggy weather based on the improved YOLO11 of the present invention.

[0017] Figure 2 Schematic diagram of the network structure of the foggy pedestrian and vehicle detection model based on the improved YOLO11n in an embodiment of the present invention.

[0018] Figure 3 Schematic diagram of a composite backbone architecture adopting the DHLC composite connection strategy in an embodiment of the present invention.

[0019] Figure 4 Schematic diagram of the process of feature fusion using the iAFF mechanism in an embodiment of the present invention, wherein Figure 4 (a) is the traditional AFF structure diagram, Figure 4 (b) is the iAFF structure diagram, Figure 4 (c) is the structural diagram of MS-CAM in feature fusion.

[0020] Figure 5 Schematic diagram of the C3K2_iAFF structure in an embodiment of the present invention, wherein Figure 5 (a) is a schematic diagram of the C3K2 structure. Figure 5 (b) Schematic diagram of the information fusion of the skip connection layer using the iAFF mechanism in the Bottleneck module.

[0021] Figure 6 Schematic diagram of the EUCB structure in an embodiment of the present invention.

[0022] Figure 7 This is a graph showing the index changes of the improved model in the embodiment of the present invention during 200 rounds of iterative training, including four indicators: P, R, map50, and map50-90. Figure 7 (a) is the traditional YOLO11n performance index, Figure 7 (b) is the performance index of the improved model in this embodiment.

[0023] Figure 8 is a visual comparison diagram of the detection result A in the embodiment of the present invention, wherein Figure 8 (a) is the original image, Figure 8 (b) is the detection situation of the original YOLO11n model. Figure 8 Figure (c) shows the detection situation of the improved model in this embodiment.

[0024] Figure 9 This is a visual comparison diagram of the test result B in the embodiment of the present invention, wherein Figure 9 (a) is the original image, Figure 9 (b) is the detection situation of the original YOLO11n model. Figure 9Figure (c) shows the detection situation of the improved model in this embodiment.

[0025] Figure 10 This is a comparison chart of the map50 values of multiple models in an embodiment of the present invention during 200 rounds of iterative training. DETAILED DESCRIPTION

[0026] Example

[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. The described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts are within the scope of protection of the present invention.

[0028] As used in this application and the claims, the words "a," "an," "an," and / or "the" are not intended to refer to the singular but may include the plural unless the context clearly indicates otherwise. Generally speaking, the terms "comprises" and "include" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.

[0029] Unless otherwise specifically stated, the relative arrangement of the parts and steps, numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present application. At the same time, it should be understood that, for ease of description, the sizes of the various parts shown in the drawings are not drawn according to actual proportional relationships. The techniques, methods and equipment known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the techniques, methods and equipment should be considered as part of the authorization specification. In all examples shown and discussed here, any specific values should be interpreted as being merely exemplary and not as limitations. Therefore, other examples of the exemplary embodiments may have different values. It should be noted that similar numbers and letters represent similar items in the following figures, and therefore, once an item is defined in one figure, it does not need to be further discussed in subsequent figures.

[0030] Combine Figure 1 , a method for detecting pedestrians and vehicles in foggy weather based on improved YOLO11, comprising the following steps:

[0031] Step 1: Collect pedestrian and vehicle images in foggy scenes to build a dataset. Specifically, the public dataset RTTS can be used. The pedestrian and vehicle images include images of cars, buses, bicycles, motorcycles, and pedestrians.

[0032] RTTS is a collection of images of pedestrians and vehicles in real-world foggy scenes. Jointly released by several universities and research institutes, RTTS is specifically designed to test the performance of image dehazing algorithms. It contains 4,322 images of real-world foggy traffic and driving scenes, labeled for five categories: cars, buses, bicycles, motorcycles, and pedestrians. RTTS serves as both a training and validation set, while a separate test set of 100 images from similar scenes is used to verify the generalization ability of the trained network.

[0033] Step 2: Combine Figure 2 , build a pedestrian and vehicle detection model in foggy weather based on the improved YOLO11n;

[0034] This solution uses the state-of-the-art model YOLO11n as the base model. The YOLO11 model makes two key adjustments to the YOLOv8 architecture: replacing the original C2F module with the C3k2 module and adding the C2PSA module after the SPPF module. The head of YOLO11 also draws on the design philosophy of YOLOv10, using depthwise separable convolution technology to reduce repetitive computations and improve recognition efficiency.

[0035] In terms of improvement, combined Figure 3 ,The described foggy pedestrian and vehicle detection model adopts a composite backbone architecture ,and a composite connection strategy with dense ,high-level connections when the auxiliary backbone transmits data to the ,main backbone to improve the detection performance;

[0036] More specifically, it is to stack multiple identical pre-trained backbone networks and use the feature fusion strategy between multiple layers to achieve multi-level reuse and enhancement of input information, so as to reduce the information loss problem in the process of deep network feature extraction;

[0037] In addition, a composite connection strategy of dense high-level connections (Dense Higher-Level Composition, DHLC) is selected between the multiple identical pre-trained backbone networks. The core idea is to deeply integrate the high-level semantic information retained in the auxiliary backbone with the low-level spatial information of the main backbone through dense cross-trunk high-level-low-level connections. Specifically, it has the advantages of multi-level feature complementarity, gradual expansion of receptive field and resistance to information attenuation. From an operational point of view, all high-level features of the auxiliary backbone are integrated into each low-level stage of the main backbone through upsampling, so that the high-level features can be better transmitted to the low-level features, gradually expanding the receptive field, thereby enhancing the feature expression ability and improving the detection performance.

[0038] Combine Figure 4 and Figure 5The foggy pedestrian and vehicle detection model introduces the iAFF mechanism (iterative attention feature fusion mechanism) in the Botteleneck jump connection of the C3K2 module in YOLO11n. This mechanism introduces an iterative mechanism based on the attention feature fusion (AFF) and optimizes the feature fusion process through a two-stage attention strategy: in the first stage, the input features are preliminarily fused with the help of the multi-scale channel attention module (MS-CAM), thereby generating attention weight coefficients containing multi-scale context associations; in the second stage, the preliminary fusion results are used as feedback and passed through the MS-CAM again with the original input features to iteratively optimize the fusion weights and alleviate the bottleneck effect of the initial integration. iAFF is used to perform more effective feature fusion of global information and local information at the data fusion level, taking into account the detection effects of large and small targets.

[0039] Specifically in operation, the iAFF mechanism first processes the original input features to generate integrated features containing preliminary attention weights, and then performs a secondary fusion based on the integrated features to form an iteratively optimized feature fusion result:

[0040]

[0041] Among them, X and Y represent the features to be fused, and M(·) represents the attention weight generated by MS-CAM.

[0042] Combine Figure 6 The foggy pedestrian and vehicle detection model uses the EUCB (Efficient Up-Convolution Block) upsampling method in the Neck network of YOLO11n, which enhances the smooth integration of information between different levels and stages, and efficiently enhances the representation of feature maps without significantly increasing computational overhead.

[0043] That is, after completing the double upsampling operation, the size of the input feature map is doubled;

[0044] Then apply the DWC convolution operation to the enlarged feature map;

[0045] Next, batch normalization and ReLU activation operations are performed, and finally a 1*1 convolution is performed to change the number of channels of the input feature map to match the number of channels of the next layer of the network.

[0046] EUCB uses deep convolution for efficient computation, reducing the overhead typically associated with standard 3×3 convolutions. This helps EUCB produce high-quality, high-resolution images while reducing computational resources by efficiently performing multi-stage feature aggregation and refinement.

[0047] like Figure 2Figure 2 shows the network structure of the improved model in this embodiment. Compared with the original YOLO11 model, the improved model incorporates the C3K2 module into the iAFF attention mechanism. In the traditional Botteleneck architecture, skip connections are key to feature fusion. This improvement introduces skip connections into the iAFF module. Its core advantage lies in its ability to fully utilize global and local attention information at the data fusion level. Global attention captures the overall context of the target, while local attention focuses on detailed features. The combination of the two enables more effective feature fusion. For large targets, global attention grasps the overall structure to avoid misjudgment; for small targets, local attention captures subtle features to prevent them from being overlooked during fusion.

[0048] The Neck network connects the upper and lower levels of object detection, responsible for feature fusion and transmission. Traditional upsampling modules suffer from information loss or discontinuity, but the introduction of the EUCB module effectively addresses this drawback. Its unique structural design enhances the smooth integration of information across different levels and stages. During upsampling, it not only considers features from the current level but also utilizes features from adjacent levels, combining high-level semantics with low-level detail information through an efficient fusion strategy to maintain information integrity and coherence. For example, when processing complex background images, EUCB smoothly integrates high-level category information with low-level edge information to avoid detection errors.

[0049] The improved model incorporates the dual-backbone concept from YOLOv9 and employs the DHLC composite connection strategy. This dual-backbone approach simultaneously extracts and combines features from two main networks, breaking the limitations of a single backbone and capturing more comprehensive image features. The auxiliary backbone uses the DHLC composite connection strategy to optimize feature transfer, effectively transferring its own high-level features to the underlying layers of the dominant backbone. In traditional single-backbone networks, the downward transfer of high-level features is prone to information loss due to depth. However, DHLC ensures the integration of high-level features into the underlying layers through skip connections and feature fusion. While preserving detailed low-level features, it also incorporates high-level semantics, gradually expanding the receptive field and enhancing understanding of the target context.

[0050] Step 3: Use the dataset to train and verify the constructed model;

[0051] First, the constructed foggy pedestrian and vehicle detection model is trained using the images in the dataset. That is, the dataset images are input and the classification results predicted by the model are output.

[0052] After model training is completed and verified, the key performance indicators of the model are verified using test data, including the mean average precision (mAP50) and mAP50-90 at confidence levels of 0.5 and 0.5-0.9, the model's floating-point computational power (GFLOPs), and the model's parameter count (Parameters).

[0053]

[0054] Where TP is the number of true positives, FP is the number of false positives, FN is the number of false negatives, AP is the average precision of a category, C is the number of detected categories, and P(r) represents the accuracy when the recall rate is r.

[0055] Step 4: Use the trained and verified foggy pedestrian and vehicle detection model to detect pedestrians and vehicles and obtain detection results.

[0056] The simulation implementation in this embodiment uses the Moment Pool cloud server, and the configuration is shown in Table 1:

[0057] Table 1 Hardware environment configuration table

[0058]

[0059] The software version configuration is shown in Table 2:

[0060] Table 2 Software environment configuration table

[0061]

[0062] The hyperparameter settings such as the number of iterations are shown in Table 3.

[0063] Table 3 Experimental parameter settings

[0064]

[0065] The key performance indicators of the verification include precision (P), recall (R), mean average precision (mAP50) and mAP50-90 at confidence levels of 0.5 and 0.5-0.9, model floating-point computational power (GFLOPs), and model parameter count (Parameters).

[0066] These performance indicators are used to measure the model's detection accuracy, complexity and other performance.

[0067] Since this solution makes multiple improvements to the YOLO model, we conducted ablation experiments to verify the effectiveness of each improvement in improving detection accuracy. The results are shown in Table 4. As can be seen, the dual-backbone mechanism has the greatest impact on improving accuracy, with map@0.5 improving by 4.2%, followed by C3k2_iAFF with a 2.6% improvement, and EUCB with a 1% improvement. The combination of any two of the three modules outperforms any single module: the dual-backbone and C3k2_iAFF combination improves by 7.2%, the dual-backbone and EUCB combination improves by 5.7%, and the C3k2_iAFF and EUCB combination improves by 3.2%. When all three modules work together, the improvement reaches 8.1%, significantly improving detection accuracy.

[0068] Table 4 Ablation experiment performance indicators

[0069]

[0070] Comparative experiment

[0071] (1) In order to verify the effectiveness of this model, comparative experiments were conducted with multiple current SOTA models under the same experimental conditions.

[0072] We selected YOLOv5n, YOLOv8n, and YOLO11n for experiments. Table 5 shows the accuracy and model complexity of each model on the RTTS dataset. As can be seen from the table, compared with models like v5n, v8n, and 11n, which have fewer parameters, our model significantly improves detection accuracy, while the increase in parameters is acceptable. Existing technologies have improved the accuracy and lightweighting of YOLOv5l models, employing the receptive field attention module RFAblock and the lightweight network Slimneck, as well as the PNMS non-maximum suppression method and SIOU as the loss function. The improved Map50 is 1.5% lower than ours, and Map50-90 is 1.7% higher. However, its computational complexity is as high as 92.8%, 6.7 times that of ours, which severely limits its real-time application.

[0073] It can be seen that our model has achieved the best comprehensive performance in detection accuracy and model complexity, exceeding the performance of the current SOTA model.

[0074] Table 5 Comparison of detection performance of different models

[0075]

[0076]

[0077] (2) Verification experiments on different models

[0078] The improved method was applied to the YOLO11s model to verify its effectiveness. The training accuracy values are shown in Table 6. It can be seen that Map50 achieved 61.4%, a 4.0% improvement over the original model's 57.4%. These results demonstrate that the improved algorithm proposed in this paper is also applicable to other models.

[0079] Table 6 Algorithm improvements on YOLO11s

[0080]

[0081] (3) Generalization experiment

[0082] This experiment aims to verify that this improved method can also improve accuracy in different datasets. The FoggyCityscapes dataset was used for this experiment. This dataset contains artificially synthesized foggy images with varying levels of haze. The FoggyCityscapes dataset is widely used to evaluate and compare the performance of different dehazing algorithms and object detection. It contains 4,428 artificially synthesized foggy scene images, uniformly covering images of light, moderate, and heavy fog, ensuring accuracy in evaluating and comparing different levels of haze and object detection performance. The dataset contains eight category labels: car, person, rider, truck, bus, train, motorcycle, and bicycle.

[0083] As shown in Table 7, the proposed algorithm still performs well on the Foggy Cityscapes dataset. Map50 is 40.0%, a 2.5% improvement over the original model. This experiment demonstrates that the proposed algorithm has good generalization performance.

[0084] Table 7 Experimental data of the algorithm in the Foggy Cityscapes dataset

[0085]

[0086] Combine Figure 7-10 , it can be seen that the detection model proposed in this paper has superiority in detection results. All performance indicators are greatly improved compared with the YOLO11n model before improvement. YOLO11n has false detection.

[0087] Combine Figure 8 , is a schematic diagram of the test result A, Figure 8 In (b), the original model made a clear mistake in identifying the left side of the image. When identifying the person wearing orange clothes, it misjudged part of the target area as a bicycle, and a bicycle that did not exist in the actual scene was mistakenly detected. This kind of misdetection can have serious consequences in application scenarios such as traffic monitoring, causing the detection results to deviate from the actual situation. In contrast, the detection results of the improved algorithm using this solution are Figure 8 In (c), the above-mentioned false detections are completely eliminated, the detection results are correct, and the confidence level is also improved.

[0088] Combine Figure 9 , is a schematic diagram of the detection result B, which clearly shows the excellent performance of the improved algorithm in processing targets at different distances. Figure 9(b) is the detection result of the original YOLO11 model. For large targets that are close to the screen, such as pedestrians and vehicles close to the camera, the original model can basically detect them, showing that it has certain capabilities in recognizing close-range and large-scale targets. However, when the target is far away, the haze environment makes the target features blurred and weak, and the limitations of the original model become apparent - pedestrians and vehicles in the distance are almost not detected, which may lead to serious safety hazards in actual scenarios such as traffic monitoring and autonomous driving. Figure 9 (c) The improved model detection results show a significant change in the image. Not only are nearby pedestrians and vehicles still accurately detected, but pedestrians and vehicles farther away that were missed by the original model are also clearly labeled under the improved algorithm.

[0089] Combine Figure 10 ,It can be seen that the improved model of this scheme achieves the best performance in ,detection accuracy and model complexity compared with multiple SOTA models.

[0090] The present invention also provides a foggy pedestrian and vehicle detection system based on improved YOLO11, which includes the following modules:

[0091] The dataset construction module is used to collect pedestrian and vehicle images in foggy scenes and construct a dataset;

[0092] Model building module: used to build foggy pedestrian and vehicle detection models based on the improved YOLO11n;

[0093] Model training module: Use the data set to train and verify the constructed model;

[0094] Detection module: used to detect pedestrians and vehicles using the trained and verified foggy pedestrian and vehicle detection models to obtain detection results.

[0095] The present invention further provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the following steps are implemented:

[0096] Step 1: Collect pedestrian and vehicle images in foggy scenes and build a dataset;

[0097] Step 2: Build a foggy pedestrian and vehicle detection model based on the improved YOLO11n;

[0098] Step 3: Use the dataset to train and verify the constructed model;

[0099] Step 4: Use the trained and verified foggy pedestrian and vehicle detection model to detect pedestrians and vehicles and obtain detection results.

[0100] The present invention further provides a computer storable medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:

[0101] Step 1: Collect pedestrian and vehicle images in foggy scenes and build a dataset;

[0102] Step 2: Build a foggy pedestrian and vehicle detection model based on the improved YOLO11n;

[0103] Step 3: Use the dataset to train and verify the constructed model;

[0104] Step 4: Use the trained and verified foggy pedestrian and vehicle detection model to detect pedestrians and vehicles and obtain detection results.

[0105] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A method for detecting pedestrians and vehicles in foggy weather based on improved YOLO11, characterized in that: The following steps are involved: Step 1: Collect pedestrian and vehicle images in foggy scenes and build a dataset; Step 2: Build a foggy pedestrian and vehicle detection model based on the improved YOLO11n; Step 3: Use the dataset to train and verify the constructed model; Step 4: Use the trained and verified foggy pedestrian and vehicle detection model to detect pedestrians and vehicles and obtain detection results.

2. The method for detecting pedestrians and vehicles in foggy weather based on improved YOLO11 according to claim 1, characterized in that: The dataset in step 1 adopts the public dataset RTTS, and the pedestrian and vehicle images include images of cars, buses, bicycles, motorcycles and pedestrians.

3. The method for detecting pedestrians and vehicles in foggy weather based on improved YOLO11 according to claim 1, characterized in that: The foggy pedestrian and vehicle detection model in step 2 uses the SOTA model YOLO11n as the basic model; The foggy pedestrian and vehicle detection model adopts a composite backbone architecture and adopts a composite connection strategy of dense high-level connections when the auxiliary backbone transmits data to the main backbone to improve detection performance; The foggy pedestrian and vehicle detection model introduces the iAFF mechanism in the Botteleneck jump connection of the C3K2 module in YOLO11n to enhance the feature fusion effect; The foggy pedestrian and vehicle detection model uses the EUCB upsampling method in the Neck network in YOLO11n to enhance the representation of feature images.

4. The method for detecting pedestrians and vehicles in foggy weather based on improved YOLO11 according to claim 3, characterized in that: The foggy pedestrian and vehicle detection model adopts a composite backbone architecture, which stacks multiple identical pre-trained backbone networks and uses a multi-layer feature fusion strategy to achieve multi-level reuse and enhancement of input information. The multiple identical pre-trained backbone networks are connected through a composite connection strategy with dense high-level connections, that is, all high-level features of the auxiliary backbone are integrated into each low-level stage of the main backbone through upsampling.

5. The method for detecting pedestrians and vehicles in foggy weather based on improved YOLO11 according to claim 3, characterized in that: The foggy pedestrian and vehicle detection model uses the iAFF mechanism for feature fusion. That is, the original input features are first processed to generate integrated features containing preliminary attention weights. Then, a secondary fusion is performed based on the integrated features to form an iteratively optimized feature fusion result: Among them, X and Y represent the features to be fused, and M(·) represents the attention weight generated by MS-CAM.

6. The method for detecting pedestrians and vehicles in foggy weather based on improved YOLO11 according to claim 3, characterized in that: The foggy pedestrian and vehicle detection model introduces the EUCB upsampling mechanism in the Neck network: That is, after completing the double upsampling operation, the size of the input feature map is doubled; Then apply the DWC convolution operation to the enlarged feature map; Next, batch normalization and ReLU activation operations are performed, and finally a 1*1 convolution is performed to change the number of channels of the input feature map to match the number of channels of the next layer of the network.

7. The method for detecting pedestrians and vehicles in foggy weather based on improved YOLO11 according to claim 1, characterized in that: The step 3 uses the dataset to train and verify the constructed model, specifically: First, the constructed foggy pedestrian and vehicle detection model is trained using the images in the dataset. That is, the dataset images are input and the classification results predicted by the model are output. After the model training is completed and passed the verification, the key performance indicators of the model are verified using the test data, including precision P, recall R and mean average precision mAP: Where TP is the number of true positives, FP is the number of false positives, FN is the number of false negatives, AP is the average precision of a category, C is the number of detected categories, and P(r) represents the accuracy when the recall rate is r.

8. A pedestrian and vehicle detection system in foggy weather based on improved YOLO11, characterized in that: Includes the following modules: The dataset construction module is used to collect pedestrian and vehicle images in foggy scenes and construct a dataset; Model building module: used to build foggy pedestrian and vehicle detection models based on the improved YOLO11n; Model training module: Use the data set to train and verify the constructed model; Detection module: used to detect pedestrians and vehicles using the trained and verified foggy pedestrian and vehicle detection models to obtain detection results.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

10. A computer storable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Intelligent positioning and grabbing method for crosshead part assembly

    CN120816498A

  • Lightweight traffic violation behavior detection method, device and equipment based on automobile data recorder and medium

    CN120833593A

  • Lightweight traffic violation detection method, device and equipment based on vehicle event data recorder and medium

    CN120833593B