Lightweight fire and smoke detection method based on improved YOLOv11n
By improving the network structure of YOLOv11n, including replacing the backbone network with RepGhostNet, introducing CCFM and GSConv modules, and using the Focaler-GIoU loss function, the problem of insufficient detection accuracy of YOLOv11n in small targets and complex backgrounds has been solved, and more efficient fire and smoke detection has been achieved.
Patent Information
- Application Number
- CN202510800096.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-11-07
AI Technical Summary
Existing fire and smoke detection methods have low accuracy, high false negative rate and insufficient real-time performance when dealing with small targets and complex backgrounds. In particular, the YOLOv11n model may cause false negatives or false positives when the fire and smoke targets are small and the background is cluttered.
By replacing the YOLOv11n backbone network with RepGhostNet, introducing the lightweight cross-scale feature fusion module CCFM and GSConv convolution, and adopting the improved regression loss function Focaler-GIoU, the network structure is optimized to improve the model's detection capability and adaptability.
It significantly improves the accuracy and real-time performance of fire and smoke detection, especially the detection performance of small targets and complex backgrounds, reduces the false negative rate and false positive rate, and improves the model's inference efficiency and the ability to detect targets at multiple scales.
Smart Images

Figure CN120912846A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer vision target detection, and particularly relates to a lightweight fire and smoke detection method based on improved YOLOv11n. BACKGROUND
[0002] Fire, as a sudden and highly hazardous disaster, seriously threatens human life and property safety and the ecological environment. Traditional fire detection methods, such as smoke detectors and temperature sensors, mainly detect temperature changes or smoke concentration to identify fire. However, these methods have certain limitations, such as high sensitivity to environmental factors, inability to respond in the early stages of fire, and limitations on device installation location. In order to overcome these problems, in recent years, fire and smoke detection technology based on computer vision and deep learning has been widely applied. Compared with traditional sensor methods, vision-based fire detection technology can detect fire at an early stage, especially when the flame or smoke in the image has not yet spread.
[0003] With the rapid development of target detection technology, fire and smoke detection methods based on convolutional neural networks (CNN) and deep learning have attracted more and more attention. Among these methods, the YOLO series of target detection networks are widely used in fire and smoke detection tasks due to their high detection accuracy and real-time performance. YOLOv11n, as the latest version, combines stronger feature extraction and detection capabilities, and performs well in handling complex scenes, but still faces challenges in small object detection accuracy and background complexity, especially in cases where fire and smoke targets are small and the background is cluttered, which may lead to missed detection or false detection.
[0004] How to solve the above technical problems is the subject faced by the present application. SUMMARY
[0005] Existing fire and smoke detection methods can detect fire signals to some extent, but often face technical problems such as low detection accuracy, high missed detection rate, and insufficient real-time performance, especially in handling small targets and complex backgrounds. In order to solve these technical problems, the purpose of the present application is to provide a lightweight fire and smoke detection method based on improved YOLOv11n, which optimizes the network structure, improves the loss function, etc. to improve the detection ability of small targets, enhance the adaptability of the model to complex scenes and improve the inference efficiency.
[0006] The application idea of the application is: in order to improve the detection accuracy, especially in the fire and smoke detection task facing multi-scale, small target and complex background, the structure of YOLOv11n is optimized, firstly, the backbone network is replaced by RepGhostNet, the structural reparameterization technology is used to improve the inference efficiency of the model, then the lightweight cross-scale feature fusion module CCFM and GSConv convolution are fused in the neck, which effectively improves the detection ability of multi-scale target. In addition, the improved regression loss function Focaler-GIoU is used, which further enhances the detection ability and regression accuracy of fire and smoke in small target and complex background, thereby improving the performance of fire and smoke detection.
[0007] In order to achieve the above application purpose, the technical scheme adopted by the application is as follows: a lightweight fire and smoke detection method based on improved YOLOv11n, comprising the following steps:
[0008] S1, constructing a fire dataset
[0009] By downloading the public dataset, collecting fire videos and using the network crawler method to obtain fire data, a fire dataset containing multiple scenes, multiple scales and multiple forms is constructed, and the dataset is labeled using the picture labeling tool LabelImg and divided into training set, validation set and test set according to the ratio of 8:1:1.
[0010] S2, replacing the YOLOv11n backbone network with the RepGhostNet network
[0011] In the target detection task, the YOLOv11n model backbone network is a C3k2 module stacking structure, although this network performs excellently in precision, but its parameter quantity and calculation quantity are large, which leads to slow inference speed, especially in edge device and mobile terminal deployment, there may be performance bottleneck, in order to solve this problem, RepGhostNet is used to replace the Yolov11n backbone network.
[0012] RepGhostNet is a lightweight neural network architecture, which adopts structural reparameterization technology and Ghost module, aiming to improve the inference efficiency and accuracy of the model; specifically, RepGhostNet generates diversified feature maps through multi-branch structure in the training stage, and fuses these branches into a standard convolution layer in the inference stage, thereby reducing the calculation amount and delay in inference. In addition, RepGhostNet generates more feature maps through the Ghost module, further improving the expression ability of the model while maintaining low computational complexity.
[0013] S3, introducing a lightweight cross-scale feature fusion module CCFM
[0014] In the fire and smoke detection task, the neck structure of the YOLOv11n model bears the bridge function of connecting the backbone network and the detection head, responsible for the fusion and transmission of multi-scale features. However, the traditional YOLOv11n neck design has problems such as large computational overhead, structural redundancy, and insufficient cross-scale feature interaction capability when facing complex and variable fire and smoke scenes, which limits the model's detection ability for small and blurred targets.
[0015] To solve the above technical problems, the original neck structure of YOLOv11n is replaced with the CCFM (Cross-scale Communication Feature Fusion) module in RT-DETR. CCFM effectively fuses feature maps from different scales through cross-layer connection and communication mechanism, enhances the model's multi-scale semantic expression ability, and significantly improves the detection performance of fire and smoke targets sensitive to scale changes.
[0016] In addition, to further optimize the computational efficiency and feature expression ability, the GSConv lightweight convolution structure is introduced in the key convolution layer of the CCFM module to replace the traditional standard convolution. GSConv combines the advantages of grouped convolution and point-wise convolution, which can reduce redundant computation and effectively extract target-related key information, improving the model's recognition ability for small targets and edge blurred areas. The combination of CCFM and GSConv realizes the dual improvement of feature fusion capability and computational efficiency, providing better structural support for fire and smoke detection.
[0017] S4, replace the bounding box regression loss function with Focaler-GIoU
[0018] In the fire and smoke detection task, the target usually has the characteristics of small size, edge blur, and low contrast with the background, and the actual deployment scene is often accompanied by light changes, occlusions, and complex backgrounds, further increasing the detection difficulty. Although the CIoU loss function used by YOLOv11n can take into account the positioning accuracy and frame overlap quality to some extent, it is prone to problems such as insensitive response to difficult samples and overfitting to background samples when facing small targets or complex scenes.
[0019] To this end, the original CIoU loss function is replaced with a more targeted Focaler-GIoU loss function to improve the model's bounding box regression performance in complex environments. Focaler-GIoU loss combines two key ideas: Focaler Loss introduces a difficult sample mining mechanism by dynamically adjusting the gradient weight of different prediction boxes, strengthening the model's attention to low-quality predictions (such as small targets, occluded areas), effectively alleviating the training bias caused by class imbalance; GIoU (Generalized IoU) extends the traditional IoU definition, even if the prediction box has no intersection with the true box, it can provide a continuous and derivable optimization target, improving the stability of box regression. The composite loss can effectively guide the model to focus on difficult-to-detect bounding boxes during the training phase, improving the positioning ability of small targets and occluded smoke areas.
[0020] S5, using the fire smoke data set to train the improved yolov11n network model, obtaining the trained model. Then input the to-be-detected image and video into the trained model for inference prediction to detect whether a fire occurs.
[0021] Compared with the prior art, the present application has the following advantages:
[0022] (1) The backbone network of the present application is replaced by RepGhostNet. RepGhostNet uses structural reparameterization technology to introduce auxiliary branches during training to enhance model expression ability, and merges these branches during inference, which simplifies the Yolov11n network structure after replacement and further improves the inference efficiency. In addition, RepGhostNet optimizes the feature reuse method by introducing the reparameterization technology, reduces redundant calculation, and improves the inference speed.
[0023] (2) The neck structure of YOLOv11 is improved in the present application, and a CCFM (cross-scale communication feature fusion) module is introduced, and a lightweight and efficient GSConv is used instead of traditional convolution in the fusion path; the CCFM module can efficiently integrate feature information of different scales, enhance the model's adaptability to scale changes of fire and smoke; and the GSConv effectively retains key information related to the target by reducing calculation redundancy and strengthening feature expression, especially improving the model's performance in small target detection. The combination of the two significantly improves the model's detection accuracy and multi-scale feature utilization efficiency.
[0024] (3) The application adopts Focaler-GIoU as a bounding box regression loss function. BRIEF DESCRIPTION OF DRAWINGS
[0025] The accompanying drawings are included to provide a further understanding of the application and are incorporated in and constitute a part of this specification, illustrate embodiments of the application and are used to explain the application, but do not limit the application.
[0026] Figure 1 The fire detection method flow chart of the improved YOLOv11n network structure of the application.
[0027] Figure 2 The bottleneck structure diagram when the RepGhostNet is trained in the application.
[0028] Figure 3 The bottleneck structure diagram when the RepGhostNet is inferred in the application.
[0029] Figure 4 The CCFM structure diagram in the application.
[0030] Figure 5 The GSConv flow chart in the application.
[0031] Figure 6 The improved Yolo11 structure chart of the application.
[0032] Figure 7 The Precision curve comparison chart before and after the model is improved in the application.
[0033] Figure 8 The mAP@0.5 curve comparison chart before and after the model is improved in the application.
[0034] Figure 9 The detection effect comparison chart before and after the improvement.
[0035] Figure 10 The heat map effect comparison chart before and after the improvement. DETAILED DESCRIPTION
[0036] In order to make the purpose, technical scheme and advantages of the application more clear, the application is further described in detail below in combination with the drawings and examples. Of course, the specific examples described here are only used to explain the application, and do not limit the application.
[0037] Embodiment 1
[0038] Reference Figure 1 With Figure 8 The embodiment provides a lightweight fire and smoke detection method based on an improved YOLOv11n, which comprises the following steps:
[0039] S1, constructing a fire dataset
[0040] First, fire and smoke pictures are collected through downloading public datasets, collecting fire videos, and python crawler technology; second, the collected pictures are sorted and unified into JPG format; finally, LabelImg tool is used for labeling processing, including two categories of fire (Fire) and smoke (Smoke), and the dataset is divided into a training set, a validation set and a test set according to a ratio of 8:1:1 through a python code.
[0041] S2, replacing the YOLOv11n backbone network with a RepGhostNet network
[0042] In the target detection task, the YOLOv11n model backbone network is a C3k2 module stacking structure, the network is excellent in precision, but the parameter quantity and the calculation quantity are large, which leads to slow inference speed, especially in the edge device and mobile terminal deployment, there may be performance bottleneck, and it is difficult to meet the real-time demand. In order to solve this problem, the RepGhostNet is used to replace the Yolo11n backbone network, the efficient feature extraction and structure reparameterization design of the RepGhostNet module are utilized, the model complexity is reduced, and the detection performance is optimized.
[0043] RepGhostNet is a lightweight neural network based on Ghost module, which optimizes the structure design of GhostNet through reparameterization technology. GhostNet module generates basic features through a small number of convolution kernels, and generates "ghost" features through linear transformation, so as to extract more features with smaller parameter amount; RepGhostNet further adopts structural reparameterization on this basis, which is equivalent to converting the multi-branch Ghost module in the training stage into a single convolution structure in the inference stage. Specifically, the training stage adopts multi-branch design: the main path compresses the channel through 1×1 convolution and extracts local features through depth separable convolution, and the identity mapping branch adjusts the dimension through 1×1 convolution with BN layer to match the output of the main path; the inference stage combines multiple branches into a single 3×3 convolution through weight superposition, eliminating the memory copy overhead caused by explicit splicing operation, while retaining the feature reuse capability. In this process, the intermediate channel number is dynamically compressed in proportion, and the non-linear expression ability is enhanced through the post ReLU activation function, while the BN layer is eliminated through parameter fusion after improving stability in the training stage, and finally the inference speed is accelerated in the hardware device.
[0044] The specific process of replacing the backbone network with RepGhostNet in this embodiment includes: first, removing the convolution architecture in the original backbone of YOLOv11n, connecting the input layer to the input of the RepGhostNet network; then, the RepGhostNet backbone performs feature extraction operations in turn. RepGhostNet is composed of multiple RepGhost bottleneck modules in series, each module extracts deep features while maintaining low computational complexity, and performs down-sampling through convolution with a stride of 2 at a specific layer, thereby generating feature maps of different scales at each level. Next, the feature maps output by the RepGhostNet at different stages are mapped to the feature layers corresponding to the original YOLOv11n backbone of the corresponding scale, and input to the original YOLO detection head for target detection. In the model structure after replacing the backbone network, the backbone part adopts RepGhostNet module integration, and the output features at each level are fused through the detection head for target positioning and classification. Through the above replacement, the model effectively reduces redundant computation and parameter amount while ensuring detection accuracy, thereby improving inference speed and deployment performance.
[0045] S3, CCFM and GSConv in neck fusion RT-DETR
[0046] In the fire and smoke detection task, the neck structure of YOLOv11n model plays a bridging role between the backbone network and the detection head, responsible for the fusion and transmission of multi-scale features. This structure is composed of a feature pyramid network composed of a spatial pyramid pooling module (SPPF) and multi-layer convolution. The multi-scale feature maps P3, P4, P5 output by the backbone network are sequentially upsampled and convolved by the neck: first, the highest level feature (such as P5) is spatially pyramid pooled to expand the receptive field, then upsampled and spliced with the middle-level feature P4; then, the low-level feature P3 is again upsampled and fused to form a top-down feature pyramid; at the same time, the high-resolution feature is fed back to the deep layer through the down-sampling path to realize bidirectional feature fusion. This PANet type structure enriches the feature representation of different scales, which is used to support the detection of large, medium and small targets. However, the above-mentioned neck design has the problems of structural redundancy and insufficient cross-scale interaction: a large number of convolution, up / down sampling operations cause computational redundancy, and different scale features are mainly fused through simple splicing or accumulation of adjacent layers, lacking direct and efficient cross-scale information exchange, which limits the efficiency of feature utilization.
[0047] In order to improve the ability and efficiency of multi-scale fusion of the model, the CCFM (Cross-scale Communication Feature Fusion) module in the RT-DETR model is introduced into Yolo11. CCFM is a cross-scale feature fusion module based on convolutional neural network, which aims to efficiently integrate feature information of different scales. Its internal contains communication fusion units for each scale feature map, which realizes efficient fusion of multi-scale features through cross-layer connection and feature interaction. Specifically, CCFM receives multi-scale feature maps S3, S4, S5 from the output of the backbone network, the deep feature S5 is upsampled after adjusting the channel number by 1x1 convolution, spliced with the middle layer feature S4 after adjusting the channel number by 1x1 convolution, and generates the intermediate feature F4 through the RepC3 module fusion, F4 is upsampled after adjusting the channel number by 1x1 convolution and spliced with the middle layer feature S3 after adjusting the channel number by 1x1 convolution, to generate high-resolution feature F3; shallow feature F3 is down-sampled and spliced with F4, and then fused through the RepC3 module to generate middle layer fusion feature F4', F4' is down-sampled and spliced with S5 after up-sampling to generate deep layer fusion feature F5.
[0048] To further optimize the efficiency and accuracy of feature expression in the fusion process, the GSConv (Ghost Shuffle Convolution) structure is introduced in the CCFM module to replace the traditional standard convolution operation. GSConv combines grouped convolution and pointwise convolution, with lower computational overhead and higher parameter utilization, effectively suppressing background interference information in shallow features while retaining key semantic response areas. This lightweight and efficient convolution form provides stronger expression ability for feature fusion and provides more abundant and refined multi-scale semantic support for the subsequent detection head.
[0049] When the Yolo11 network neck fuses the CCFM module, the original C3K2 module is retained, and the structure of the CCFM module is borrowed. In the bottom-up path of FPN, a layer of lightweight convolution operation composed of GSConv is added before upsampling and splicing, replacing the traditional standard convolution. GSConv combines the advantages of grouped convolution and pointwise convolution, balancing information expression and computational efficiency, and has significant advantages in compressing channel redundancy and retaining key semantic information. This convolution is composed of a group convolution layer (Group Conv), a pointwise convolution (1x1Conv), and a SiLU activation function, with strong information extraction, normalization stability, and nonlinear expression capabilities. Specifically, the GSConv layer can effectively compress the channel dimension, enhance local perception ability, and reduce background noise in shallow features; the BatchNorm layer further standardizes feature distribution, improving the convergence speed and stability of network training; and the SiLU activation function enhances the expression ability of features through smooth nonlinear mapping, helping the network to extract more discriminative semantic features in multi-scale scenarios and strengthen the details of small target regions.
[0050] S4, replace the bounding box regression loss function with Focaler-GIoU
[0051] In fire and smoke detection tasks, especially in monitoring systems deployed in actual environments, the detection performance of objects is often plagued by small target detection, class imbalance, complex background, and other problems. Although the traditional YOLOv11n bounding box regression loss function CIoU can handle frame positioning well, it may face the problems of insufficient accuracy and excessive optimization of background samples when facing small targets and complex backgrounds. Therefore, replacing the bounding box regression loss function of YOLOv11n with Focaler-GIoU loss, combined with Focaler Loss and GIoU loss, can more effectively address these challenges.
[0052] Focaler-GIoU loss combines Focaler-IoU and GIoU loss, aiming to improve the regression accuracy in fire detection caused by small targets, class imbalance and complex background. Specifically: Focaler-IoU introduces a dynamic loss adjustment mechanism to map the IoU value to an interval, so that the model can focus on regression samples of different difficulties. The loss formula of Focaler Loss is:
[0053]
[0054] where d is the low confidence threshold, and u is the high confidence threshold. When IoU < d, set IoU to 0 to filter noise samples with very low IoU (such as false detection of background boxes), avoiding invalid gradient interference with model convergence; when IoU > u, set IoU to 1 to suppress the optimization intensity of simple samples with high IoU (such as accurately located targets), preventing them from dominating the training process; when IoU is between the interval [d, u], it is a medium difficulty sample, and by dynamically weighting the gradient of these samples, the model pays more attention to difficult samples with IoU in this interval, such as small targets or fuzzy boundary targets. In this embodiment, in order to improve the detection effect of small targets and complex background, d is set to 0 and u is set to 0.95.
[0055] GIoU loss is an improvement on traditional IoU loss. In addition to considering the overlap of the predicted box and the real box (IoU), it also introduces the minimum closed box (i.e. the smallest rectangular box enclosing the predicted box and the real box) to further optimize the positioning accuracy of the box. The formula of GIoU loss is:
[0056]
[0057] where IoU is the intersection over union of the predicted box and the real box, C is the area of the minimum closed box, A∪B is the union area of the predicted box and the real box, and |C| is the area of the closed box. GIoU loss not only focuses on the overlapping area, but also considers the optimization of the shape and position of the box, which can effectively reduce the influence of shape mismatch and position deviation.
[0058] Focaler-GIoU loss combines the advantages of the above two, by setting a reasonable confidence threshold, dynamically weighting the gradient of samples of different difficulties, so that the model pays more attention to the detection of small targets and complex background, and suppresses the optimization intensity of simple samples, optimizing the detection effect of small targets and complex background, while optimizing the accuracy of the boundary box. The final loss formula is:
[0059] L Focaler-GIoU =L GIoU +IoU-IoU Focaler
[0060] S5, training the improved YOLOv11n network model using the fire data set to obtain a trained model.
[0061] The parameters of the YOLOv11n network structure to be improved are set, the model is selected as YOLOv11n, the learning rate is 0.01, the epoch is set to 200 rounds, the batch-size is set to 16, the data set is used for training, and the network structure parameters are continuously optimized according to the training condition until the improved model with the best training effect is obtained.
[0062] Finally, the to-be-detected image and video are input into the trained model for inference prediction to obtain the result of whether a fire occurs and output.
[0063] In order to verify the effectiveness of the embodiment and the improvement, the improved YOLOv11n model is compared with YOLOv11n, YOLOv8n and YOLOv5n in the same data set for comparative experiments and ablation experiments. The evaluation indexes are accuracy P, mAP50 and weight size as evaluation criteria, and the results are shown in Table 1:
[0064] Table 1 Comparison of precision and weight of different models
[0065]
[0066] The results in Table 1 show that the method of the embodiment has a higher mAP50 and a smaller weight size. The higher mAP50 reflects the detection performance and model generalization of the method of the embodiment, and the smaller weight size reflects the lightweight of the model of the embodiment. In addition, the ablation experiment shows that after adding each module, the mAP50 of the embodiment is improved, indicating the effectiveness of the improvement.
[0067] Embodiment 2
[0068] Figure 9 The results show that the method of the embodiment successfully detects the small flame target and the thin smoke that are missed by the original model, indicating that the embodiment has an advantage in detecting difficult samples such as small targets and thin smoke. The third row shows that the embodiment can also maintain high accuracy in night environment, reflecting the adaptability of the embodiment to complex environments such as night. These two points indicate the effectiveness of the improvement.
[0069] Embodiment 3
[0070] Figure 10The results show that in the three fire scenarios of high-rise building, urban group and night ancient building, the YOLOv11n model before improvement has the problem of missing small flames, and the improved YOLOv11n heat map not only covers the original missed area completely, but also significantly strengthens the heat response intensity of the flame target, especially for the thin smoke and small flame, the overall heat distribution is more accurate and coherent. This directly verifies the effectiveness of the algorithm improvement for small target detection, and proves that it still maintains stable recognition ability in complex lighting environment (such as night scene).
[0071] The above merely describes preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A lightweight fire and smoke detection method based on improved YOLOv11n, characterized in that, The method comprises the following steps: S1, constructing a fire data set By downloading public data sets, collecting fire videos and using web crawler methods to obtain fire data, a fire data set containing various scenes, various scales and various forms is constructed, the data set is labeled using the picture labeling tool LabelImg, and the data set is divided into a training set, a verification set and a test set according to a ratio of 8:1:1; S2, replacing the YOLOv11n backbone network with a RepGhostNet network In the target detection task, the YOLOv11n backbone network is a C3k2 module stacking structure; RepGhostNet is a lightweight neural network architecture that uses structured reparameterization technology and Ghost modules. RepGhostNet generates diverse feature maps through multi-branch structure during the training stage, and fuses these branches into a standard convolution layer during the inference stage. RepGhostNet generates multiple feature maps through the Ghost module; S3, introducing a lightweight cross-scale feature fusion module CCFM Replace the original neck structure of YOLOv11n with the CCFM module in RT-DETR. The CCFM module fuses feature maps from different scales through cross-layer connection and communication mechanism; S4, replacing the bounding box regression loss function with Focaler-GIoU; S5, training the improved yolov11n network model using the fire smoke data set, obtaining the trained model, and inputting the to-be-detected image and video into the trained model for inference prediction to detect whether a fire occurs.
2. The improved YOLOv11n-based lightweight fire and smoke detection method according to claim 1, characterized in that, In the step S1, first, by downloading public data sets, collecting fire videos and using web crawler methods to obtain fire data, a data set containing various scenes, scales and forms is constructed; then, using the picture labeling tool LabelImg to label the data set, including two categories of fire and smoke; finally, using a python script to divide the data set into a training set, a verification set and a test set according to a ratio of 8:1:
1.
3. The lightweight fire and smoke detection method based on improved YOLOv11n according to claim 1 or 2, characterized in that, In the step S2, first, remove the convolution architecture in the original backbone of YOLOv11n, and connect the input layer to the input of the RepGhostNet network; then, the RepGhostNet backbone performs feature extraction operations in turn; then, the feature maps output by the RepGhostNet at different stages are mapped to the feature layers corresponding to the original YOLOv11n backbone of corresponding scales, and input to the original YOLO detection head for target detection.
4. The lightweight fire and smoke detection method based on improved YOLOv11n according to claim 1, characterized in that, In the step S3, in the bottom-up path of the neck FPN of YOLOv11n, a convolution operation is added before upsampling and splicing, and the convolution is replaced with GSConv instead of the traditional standard convolution.
5. The lightweight fire and smoke detection method based on improved YOLOv11n according to claim 1, characterized in that, In the step S4, first, modify the loss.py module in the YOLOv11n file, and replace the original bounding box loss function CIoU with Focaler-GIoU, wherein the loss formula of Focaler Loss is: Wherein, IoU is the intersection over union of the predicted box and the real box, d is the low confidence threshold, u is the high confidence threshold, when IoU < d, set IoU as 0, filter the noise samples with extremely low IoU; when IoU > u, set IoU as 1, suppress the optimization intensity of the simple samples with high IoU, prevent them from dominating the training process, when IoU is between the interval [d, u], it is a medium difficulty sample, by dynamically weighting the gradient of these samples, make the model pay more attention to the difficult samples with IoU in this interval. The loss formula of GIoU is: Wherein, IoU is the intersection over union of the predicted box and the real box, C is the area of the minimum closed box, AUB is the union area of the predicted box and the real box, |C| is the area of the closed box, and GIoU loss not only pays attention to the overlapping area.
6. The lightweight fire and smoke detection method based on improved YOLOv11n according to claim 1, characterized in that, In the step S5, first, the parameters of the improved YOLOv11n network structure are set, the model is selected as YOLOv11n, the learning rate is 0.01, the epoch is set to 200 rounds, the batch-size is set to 16, the data set is used for training, the best weight best.pt is obtained, then the to-be-detected image and video are input into the trained model for inference prediction, and the result of whether a fire occurs is obtained and output.