Fire smoke target detection method and system based on improved YOLOv11
By introducing the CIB module into the C3K2 structure of the YOLOv11 model, the inverted residual structure and depth separation convolution is used to improve the fire smoke target detection method, solve the problems of insufficient detection accuracy and poor real-time performance in the prior art, and achieve more efficient fire smoke detection.
Patent Information
- Application Number
- CN202510292556.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-06
AI Technical Summary
The existing fire smoke detection technology has insufficient detection accuracy and poor real-time performance in complex environments, and cannot provide information for fire rescue quickly and accurately.
Based on the improved fire smoke target detection method of YOLOv11, by introducing a CIB module into the C3K2 structure of the YOLOv11 model, the inverted residual structure and depth can be separated and convolutionized, reducing the calculation amount and maintaining the feature expression ability.
Improves the accuracy and real-time performance of fire smoke detection, enhances the model's detection capabilities of small targets and low-quality examples, and reduces the calculation cost and parameter volume.
Smart Images

Figure CN120107561A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision target detection, and in particular to a fire smoke target detection method and system based on improved YOLOv11. Background Art
[0002] Fire is a type of disaster that frequently rages in reality. From forest fires caused by natural factors, to car ignition and explosions caused by traffic accidents, to circuit fires caused by aging household appliances in daily life, the hazards are widespread and serious. Fires not only threaten the safety of the natural ecology and disrupt social traffic order, but also directly endanger the lives of the people and bring inestimable economic losses to the country and the people. Therefore, being able to detect the fire in time and respond quickly when it first appears is of vital importance for protecting natural resources, protecting people’s lives and property, and maintaining social public stability.
[0003] In the early research stage of fire smoke detection, temperature sensors and smoke sensors were commonly used. However, this method has many limitations: limited application scenarios and difficult to adapt to diverse environments; relatively narrow detection range, prone to blind spots; high equipment costs, not conducive to large-scale promotion; sensors are fragile and easily damaged; slow detection speed, difficult to detect dangerous situations in time; and frequent false alarms and false triggers, resulting in waste of manpower and material resources.
[0004] Subsequently, a method based on traditional image processing came into being. This method recognizes and detects fire smoke by analyzing the characteristic information of the fireworks generated by the fire, and to a certain extent overcomes some of the drawbacks of the sensor detection method. However, once faced with complex environments, such as scenes with changing light, uneven smoke concentration, and many interference sources, this detection method will still be seriously interfered with, resulting in problems such as insufficient detection accuracy and poor real-time detection capabilities, and it is unable to accurately and quickly provide reliable information for fire rescue.
[0005] In today's fire smoke target detection tasks, how to ensure that the model has high-precision detection performance while effectively improving the detection speed has become a key problem. The YOLO series algorithm models have shown excellent performance in this regard. In view of this, the present invention is based on YOLOv11, deeply analyzes the difficulties in the flame smoke specific target detection task, and conducts comprehensive research and exploration, striving to break through the existing technical bottlenecks and bring new solutions to the field of fire smoke detection. Summary of the invention
[0006] To solve the above problems, the present invention aims to propose a fire smoke target detection method and system based on improved YOLOv11. Based on the original YOLOv11, the CIB module is introduced in the core structure C3K2, and the inverted residual structure and depthwise separable convolution are used to reduce the amount of calculation while maintaining the feature expression ability.
[0007] To achieve the above object, the technical solution of the present invention is achieved as follows:
[0008] A fire smoke target detection method based on improved YOLOv11 includes the following steps:
[0009] S1, build a fire smoke image dataset;
[0010] S2, by adding CIB module to C3K2 in YOLOv11 model, using inverted residual structure and depth-separable convolution to improve, and build a fire smoke target detection model;
[0011] S3. Select data enhancement method and set training parameters;
[0012] S4, training a fire smoke target detection model based on improved YOLOv11;
[0013] S5. Test the trained fire smoke target detection model based on improved YOLOv11 to obtain the test results.
[0014] Furthermore, step S1 is specifically as follows: the user collects and downloads a fire smoke image dataset on the Internet or uses photographic equipment to collect the fire smoke image dataset, performs data cleaning on the collected image dataset, uses Labelme to annotate the image, and names the fire smoke image file and the label file according to the standardized YOLO dataset format, and divides the fire smoke image file into a training set, a validation set, and a test set.
[0015] Further, the step S2 comprises the following steps:
[0016] Step S21: Use a 1×1 convolutional layer to expand the number of channels of the input feature map of size H×W×C to C1. The calculation formula for this process is:
[0017] X 1 =σ(W 1 *X+b 1 );
[0018] Where W 1 is the convolution kernel weight, σ is the Relu activation function;
[0019] Step S22: Use 3×3 depth-separable convolution to extract spatial features and output feature map X 2 ,The calculation formula of this process is:
[0020] X 2 =σ(Pointwise(Depthwise(X 1 )));
[0021] Where Pointwise(·) is point-wise convolution, Depthwise(·) is depth-separable convolution, the formula is
[0022]
[0023] Where W dw is a 3×3 depth convolution kernel, (x, y) is the spatial position of the output feature map, (i, j) is the spatial position of the convolution kernel, and C1 is the number of channels;
[0024] Step S23: Compress the feature map through 1×1 point-by-point convolution and output the final feature map X 3 ,The calculation formula of this process is:
[0025] X 3 =σ(W 4 *X 2 +b 4 ).
[0026] Furthermore, the step S2 also includes: performing multi-scale optimization, introducing an improved MSDSConv module to perform multi-scale feature extraction and information fusion on the feature map output by the C3K2CIB module.
[0027] Furthermore, the MSDSConv module includes two parallel depth-separable convolutions of multi-scale convolution kernels, and uses the feature map output by the C3K2CIB module as input; specifically, first, a 3×3, 5×5 depth-separable convolution layer and an AvgPooling layer are respectively passed through, and then the three intermediate feature maps are respectively reduced in channels through a 1×1 convolution layer and concatenated into a feature map, and finally the output result is obtained through a 1×1 convolution layer. The calculation formula of this process is:
[0028] X 2 =
[0029] Pointwise(Concat(Pointwise(Depthwise(X 3 ))+Pointwise(Depthwise(X 3 ))+Pointwise(AvgPool(X3));
[0030] Where Concat(·) is a channel concatenation operation.
[0031] Furthermore, it also includes: after introducing the improved MSDSConv module to perform multi-scale feature extraction and information fusion on the feature map output by the C3K2CIB module, the space-channel decoupled downsampling is realized through the sampling SCDown module, specifically: using 3×3 deep convolution to compress the space while focusing on local small target features, and adjusting the number of channels through 1×1 point-by-point convolution operation to further fuse information and adjust the number of channels to achieve lightweight and efficient downsampling.
[0032] Furthermore, the model training in step S4 uses Wise-IoU to enhance the detection capability of the model for scenes with many small targets and low-quality examples. The calculation formula of this process is:
[0033]
[0034]
[0035] in, Significantly enlarge the normal quality anchor box Significantly reduce the high quality anchor box And when the anchor box and the target box overlap well, more attention is paid to the distance between the center points; is the size of the smallest enclosing box; to prevent Produces a speed that hinders convergence, Separate from the computation graph.
[0036] Furthermore, the step S3 is specifically as follows: setting the input image resolution to 640×640, setting the BatchSize during training to 32, setting the BatchSize during testing to 8, setting the number of training epochs to 200, selecting SGD as the optimizer, setting the momentum size to 0.9, setting the weight decay coefficient to 0.0005, using the OneCycle learning rate scheduler, setting the maximum learning rate to 0.1, and using the default data enhancement method.
[0037] Furthermore, the indicator for testing the model performance in step S5 is the mean average precision (mAP), which is calculated by: Where k represents the number of categories and AP represents the average precision of each category.
[0038] In order to achieve the above object, the present invention also provides a fire smoke target detection system based on improved YOLOv11, comprising:
[0039] Data module: used to build fire smoke image dataset;
[0040] Fire smoke target detection module: By adding the CIB module to C3K2 in the YOLOv11 model, the inverted residual structure and depth-separable convolution are used to improve it, and a fire smoke target detection model is built based on this;
[0041] Setting module: used to select data enhancement method and set training parameters;
[0042] Training module: training the fire smoke target detection model based on improved YOLOv11;
[0043] Detection module: used to test the trained fire smoke target detection model based on improved YOLOv11 and obtain test results.
[0044] Beneficial effects: Based on the original YOLOv11, the present invention introduces the CIB module into the core structure C3K2, and utilizes the inverted residual structure and the depth-separable convolution to reduce the amount of calculation while maintaining the feature expression capability; performs multi-scale optimization of feature extraction, introduces the MSDSConv module, and improves the target detection accuracy and real-time performance by enhancing the multi-scale feature extraction capability and reducing the amount of calculation; adds the SCDown module to realize space-channel decoupling downsampling, and achieves a balance between computational efficiency and feature retention capability by separating the reduction of spatial resolution and the adjustment of the number of channels. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The accompanying drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the accompanying drawings:
[0046] Figure 1 This is a flow chart of a fire smoke target detection method based on improved YOLOv11 according to an embodiment of the present invention;
[0047] Figure 2 A model structure diagram of an improved YOLOv11 in a fire smoke target detection method based on an improved YOLOv11 according to an embodiment of the present invention;
[0048] Figure 3 It is a C3K2CIB module diagram in the fire smoke target detection method based on improved YOLOv11 described in an embodiment of the present invention;
[0049] Figure 4 It is a CIB module diagram in the fire smoke target detection method based on improved YOLOv11 according to an embodiment of the present invention;
[0050] Figure 5 It is a diagram of an improved MSDSConv module in the fire smoke target detection method based on improved YOLOv11 described in an embodiment of the present invention;
[0051] Figure 6 It is a diagram of the SCDown module in the fire smoke target detection method based on improved YOLOv11 described in an embodiment of the present invention;
[0052] Figure 7 This is a diagram showing the actual detection effect of the fire smoke target detection method based on the improved YOLOv11 described in an embodiment of the present invention;
[0053] Figure 8 It is a structural schematic diagram of a fire smoke target detection system based on improved YOLOv11 described in an embodiment of the present invention. DETAILED DESCRIPTION
[0054] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.
[0055] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0056] Example 1
[0057] See also Figure 1-7 :A fire smoke target detection method based on improved YOLOv11, comprising the following steps:
[0058] S1. Construct a fire smoke image dataset; the dataset includes a training set, a validation set, and a test set;
[0059] S2, by adding CIB module to C3K2 in YOLOv11 model, using inverted residual structure and depth-separable convolution to improve, and build a fire smoke target detection model;
[0060] S3. Select data enhancement method and set training parameters;
[0061] S4, training a fire smoke target detection model based on improved YOLOv11; inputting the training set and the validation set established in step S1 into the fire smoke target detection model based on improved YOLOv11 constructed in step S2 for training;
[0062] S5. Test the trained fire smoke target detection model based on improved YOLOv11 to obtain test results; input the test set established in step S1 into the fire smoke target detection model based on improved YOLOv11 trained in step S4 for testing.
[0063] This embodiment is based on the original YOLOv11, introduces the CIB module in the core structure C3K2, and uses the inverted residual structure and depth-wise separable convolution to reduce the amount of calculation while maintaining the feature expression capability.
[0064] In a specific example, step S1 is specifically as follows: a user collects and downloads a fire smoke image dataset on the Internet or uses photographic equipment to collect the fire smoke image dataset, performs data cleaning on the collected image dataset, uses Labelme to annotate the image, and names the fire smoke image file and the label file according to the standardized YOLO dataset format, and divides the fire smoke image file into a training set, a validation set, and a test set.
[0065] In the specific implementation, the dataset can be divided into 5600 images for training, 1600 images for validation, and 800 images for test according to the ratio of 7:2:1; the dataset folder is placed in the YOLOv11 project directory.
[0066] In a specific example, the step S2 comprises the following steps:
[0067] Step S21: Use a 1×1 convolutional layer to expand the number of channels of the input feature map of size H×W×C to C1. The calculation formula for this process is:
[0068] X 1 =σ(W 1 *X+b 1 );
[0069] Where W 1 is the convolution kernel weight, σ is the Relu activation function;
[0070] Step S22: Use 3×3 depth-separable convolution to extract spatial features and output feature map X 2 ,The calculation formula of this process is:
[0071] X 2 =σ(Pointwise(Depthwise(X 1 )));
[0072] Where Pointwise(·) is point-wise convolution, Depthwise(·) is depth-separable convolution, the formula is
[0073]
[0074] Where W dw is a 3×3 depth convolution kernel, (x, y) is the spatial position of the output feature map, (i, j) is the spatial position of the convolution kernel, and C1 is the number of channels;
[0075] Step S23: Compress the feature map through 1×1 point-by-point convolution and output the final feature map X 3 ,The calculation formula of this process is:
[0076] X 3=σ(W 4 *X 2 +b 4 ).
[0077] This embodiment introduces the CIB module into the C3K2 structure. Specifically, the C3K2 module in the YOLOv11 backbone network is replaced with the C3K2CIB module. The method is to use the inverted residual structure of the CIB module, use 1×1 convolution to perform dimensionality increase operation, and finally use 1×1 convolution to perform dimensionality reduction operation. In addition, 3×3 depthwise separable convolution is used to extract richer features in the high-dimensional space. This is used as a new Bottleneck to replace the Bottleneck module in the C3K2 module, so as to achieve lightweight feature extraction while maintaining the scale invariance of the feature map, thereby significantly improving the computational efficiency and feature extraction capability.
[0078] In a specific example, the step S2 further includes: performing multi-scale optimization, introducing an improved MSDSConv module to perform multi-scale feature extraction and information fusion on the feature map output by the C3K2CIB module.
[0079] This embodiment performs multi-scale optimization of feature extraction and introduces an improved MSDSConv module to improve the accuracy and real-time performance of target detection by enhancing the multi-scale feature extraction capability and reducing the amount of calculation.
[0080] In a specific example, the MSDSConv module includes two parallel depth-separable convolutions of multi-scale convolution kernels, and uses the feature map output by the C3K2CIB module as input; specifically: first, a 3×3, 5×5 depth-separable convolution layer and an AvgPooling layer are respectively passed through, and then the three intermediate feature maps are respectively reduced in channels through a 1×1 convolution layer and spliced into a feature map, and finally the output result is obtained through a 1×1 convolution layer. The calculation formula of this process is:
[0081] X 2 =
[0082] Pointwise(Concat(Pointwise(Depthwise(X 3 ))+Pointwise(Depthwise(X 3 ))+Pointwise(AvgPool(X3));
[0083] Where Concat(·) is a channel concatenation operation.
[0084] In this embodiment, an improved MSDSConv module is added to the existing YOLOv11 model to perform multi-scale optimization, thereby expanding the receptive field and enhancing the feature expression capability.
[0085] In a specific example, it also includes: after introducing an improved MSDSConv module to perform multi-scale feature extraction and information fusion on the feature map output by the C3K2CIB module, space-channel decoupled downsampling is realized through the sampling SCDown module, specifically: using 3×3 deep convolution to compress the space while focusing on local small target features, and adjusting the number of channels through 1×1 point-by-point convolution operations to further fuse information and adjust the number of channels to achieve lightweight and efficient downsampling.
[0086] This embodiment adds an SCDown module after the improved MSDSConv module processing to achieve space-channel decoupled downsampling, and achieves a balance between computational efficiency and feature retention capability by separating spatial resolution reduction and channel number adjustment.
[0087] In a specific example, the model training in step S4 uses Wise-IoU to enhance the detection capability of the model for scenes with many small objects and low-quality examples. The calculation formula of the process is:
[0088]
[0089]
[0090] in, Significantly enlarge the normal quality anchor box Significantly reduce the high quality anchor box And when the anchor box and the target box overlap well, more attention is paid to the distance between the center points; is the size of the smallest enclosing box; to prevent Produces a speed that hinders convergence, Separate from the computation graph (superscript * denotes this operation).
[0091] In a specific example, step S3 is specifically as follows: setting the input image resolution to 640×640, setting the BatchSize during training to 32, setting the BatchSize during testing to 8, setting the number of training epochs to 200, selecting SGD as the optimizer, setting the momentum size to 0.9, setting the weight decay coefficient to 0.0005, using the OneCycle learning rate scheduler, setting the maximum learning rate to 0.1, and using the default data enhancement method.
[0092] In a specific example, the indicator for testing the model performance in step S5 is the mean average precision mAP, which is calculated by the formula: Where k represents the number of categories and AP represents the average precision of each category.
[0093] The data set used in this example contains 2 categories, so k is 2. The detection effect of the model on the test set is as follows: Figure 7 As shown in the figure, after testing, the model achieved a map of 71.2% on the test set, which is 2.8% higher than the original YOLOv11.
[0094] In summary, the method of this embodiment is based on the YOLOv11 model. By adding the C3K2CIB module, it can extract and fuse features more efficiently and significantly reduce the amount of calculation; by adding the improved MSDSConv module, the receptive field is expanded to better adapt to targets of different scales, enhance the feature expression ability of the network, improve the detection accuracy, and enhance the robustness of the model; by introducing the efficient downsampling module SCDown, the computational cost and parameter amount are reduced while maintaining accuracy, and the feature information is effectively retained; by introducing Wise-IoU, the model's detection ability for low-quality and small target examples is improved; the method of this embodiment effectively improves the accuracy of fire smoke target detection tasks.
[0095] Example 2
[0096] To achieve the above purpose, see Figure 8 : This embodiment also provides a fire smoke target detection system based on improved YOLOv11, including:
[0097] Data module: used to build fire smoke image dataset;
[0098] Fire smoke target detection module: By adding the CIB module to C3K2 in the YOLOv11 model, the inverted residual structure and depth-separable convolution are used to improve it, and a fire smoke target detection model is built based on this;
[0099] Setting module: used to select data enhancement method and set training parameters;
[0100] Training module: training the fire smoke target detection model based on improved YOLOv11;
[0101] Detection module: used to test the trained fire smoke target detection model based on improved YOLOv11 and obtain test results.
[0102] The fire smoke target detection system based on improved YOLOv11 of this embodiment has the same advantages as the above-mentioned fire smoke target detection method based on improved YOLOv11 over the prior art, which will not be repeated here.
[0103] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A fire smoke target detection method based on improved YOLOv11, characterized in that: The following steps are involved: S1, build a fire smoke image dataset; S2, by adding CIB module to C3K2 in YOLOv11 model, using inverted residual structure and depth-separable convolution to improve, and build a fire smoke target detection model; S3. Select data enhancement method and set training parameters; S4, training a fire smoke target detection model based on improved YOLOv11; S5. Test the trained fire smoke target detection model based on improved YOLOv11 to obtain the test results.
2. The fire smoke target detection method based on improved YOLOv11 according to claim 1, characterized in that: The step S1 is specifically as follows: the user collects and downloads a fire smoke image dataset on the Internet or uses photographic equipment to collect the fire smoke image dataset, cleans the collected image dataset, uses Labelme to annotate the image, names the fire smoke image file and the label file according to the standardized YOLO dataset format, and divides the fire smoke image file into a training set, a validation set, and a test set.
3. The fire smoke target detection method based on improved YOLOv11 according to claim 1, characterized in that: Described step S2 comprises the following steps: Step S21, using a 1×1 convolutional layer to expand the number of channels of the input feature map of size H×W×C to C1. The calculation formula for this process is: X1=σ(W1*X+b1); Where W1 is the convolution kernel weight and σ is the Relu activation function; Step S22, use 3×3 depth separable convolution to extract spatial features and output feature map X2. The calculation formula of this process is: X2=σ(Pointwise(Depthwise(X1))); Where Pointwise(·) is point-wise convolution, Depthwise(·) is depth-separable convolution, the formula is Where W dw is a 3×3 depth convolution kernel, (x, y) is the spatial position of the output feature map, (i, j) is the spatial position of the convolution kernel, and C1 is the number of channels; Step S23, compress the feature map through 1×1 point-by-point convolution, and output the final feature map X3. The calculation formula of this process is: X3=σ(W4*X2+b4).
4. The fire smoke target detection method based on improved YOLOv11 according to claim 1, characterized in that: The step S2 also includes: performing multi-scale optimization, introducing an improved MSDSConv module to perform multi-scale feature extraction and information fusion on the feature map output by the C3K2CIB module.
5. The fire smoke target detection method based on improved YOLOv11 according to claim 4, characterized in that: The MSDSConv module contains two parallel depth-separable convolutions of multi-scale convolution kernels, and uses the feature map output by the C3K2CIB module as input; specifically: first, a 3×3, 5×5 depth-separable convolution layer and an AvgPooling layer are passed respectively, and then the three intermediate feature maps are respectively reduced in channels through a 1×1 convolution layer and concatenated into a feature map, and finally the output result is obtained through a 1×1 convolution layer. The calculation formula of this process is: X2= Pointwise(Concat(Pointwise(Depthwise(X3))+Pointwise(Depthwise(X3))+Pointwise(AvgPool(X3)); Where Concat(·) is a channel concatenation operation.
6. The fire smoke target detection method based on improved YOLOv11 according to claim 4, characterized in that: It also includes: after introducing the improved MSDSConv module to perform multi-scale feature extraction and information fusion on the feature map output by the C3K2CIB module, the spatial-channel decoupled downsampling is realized through the sampling SCDown module. Specifically, 3×3 deep convolution is used to compress the space while focusing on local small target features, and the number of channels is adjusted through 1×1 point-by-point convolution operation to further fuse information and adjust the number of channels to achieve lightweight and efficient downsampling.
7. The fire smoke target detection method based on improved YOLOv11 according to claim 1, characterized in that: The model training in step S4 uses Wise-IoU to enhance the detection capability of the model for scenes with many small targets and low-quality examples. The calculation formula of this process is: in, Significantly enlarge the normal quality anchor box Significantly reduce the high quality anchor box And when the anchor box and the target box overlap well, more attention is paid to the distance between the center points; is the size of the smallest enclosing box; to prevent Produces a speed that hinders convergence, Separate from the computation graph.
8. The fire smoke target detection method based on improved YOLOv11 according to claim 1, characterized in that: The step S3 is specifically as follows: setting the input image resolution to 640×640, setting the BatchSize during training to 32, setting the BatchSize during testing to 8, setting the number of training epochs to 200, selecting SGD as the optimizer, setting the momentum size to 0.9, the weight decay coefficient to 0.0005, using the OneCycle learning rate scheduler, setting the maximum learning rate to 0.1, and using the default data enhancement method.
9. The fire smoke target detection method based on improved YOLOv11 according to claim 1, characterized in that: The indicator for testing the model performance in step S5 is the mean average precision (mAP), which is calculated by the following formula: Where k represents the number of categories and AP represents the average precision of each category.
10. A fire smoke target detection system based on improved YOLOv11, characterized in that: include: Data module: used to build fire smoke image dataset; Fire smoke target detection module: By adding the CIB module to C3K2 in the YOLOv11 model, the inverted residual structure and depth-separable convolution are used to improve it, and a fire smoke target detection model is built based on this; Setting module: used to select data enhancement method and set training parameters; Training module: training the fire smoke target detection model based on improved YOLOv11; Detection module: used to test the trained fire smoke target detection model based on improved YOLOv11 and obtain test results.
Citation Information
Cited By
Single-plant-scale tree positioning and identifying method
CN120747757A
Forest fire identification method, system and device based on edge calculation and medium
CN121305379A