Flame smoke detection method and system based on dynamic convolution and multi-scale feature fusion

By improving the YOLOv5s model and introducing dynamic convolution and multi-scale feature fusion techniques, the problems of false detection and missed detection in flame and smoke detection under complex backgrounds were solved, the detection accuracy and robustness were improved, and the model's adaptability in different environments was enhanced.

CN120876829APending Publication Date: 2025-10-31HANGZHOU YALONG INTELLIGENT TECH CO LTD +1
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510977235.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing flame and smoke detection methods have low accuracy in complex backgrounds, low contrast, small targets, and dynamic changes, and have poor generalization ability, making it difficult to effectively distinguish between flames and smoke and background. In particular, they are prone to false detection or missed detection in complex environments.

Method used

We adopt a method based on dynamic convolution and multi-scale feature fusion. By improving the YOLOv5s model, we introduce the C3-ODConv module, the global context module GCB, and the bidirectional feature pyramid network BiFPN. Combined with the Soft-NMS algorithm, we optimize the feature extraction and fusion process and enhance the robustness and adaptability of the model.

Benefits of technology

It improves the accuracy and robustness of flame and smoke detection, enhances the model's detection stability and generalization ability in complex scenarios, and can better adapt to changes in flame and smoke under different environments, reducing false detections and missed detections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876829A_ABST
    Figure CN120876829A_ABST
Patent Text Reader

Abstract

The invention discloses a flame smoke detection method and system based on dynamic convolution and multi-scale feature fusion, and belongs to the technical field of computer vision target detection, and the method comprises the steps: 1, obtaining a fire data set, carrying out the preprocessing of the fire data set, and dividing the fire data set into a training set and a test set; 2, constructing an improved YOLOv5s model: replacing a C3 module with C3-ODConv, introducing a GCB module into a backbone network, and adopting BiFPN on a neck; 3, training the improved YOLOv5s model by using the training set, and testing the improved YOLOv5s model by using the test set; and step 4, performing fire data detection by using the improved YOLOv5s model after training and testing. According to the invention, the learning and identification capability of flame and smoke characteristics can be enhanced, and the detection precision of the method in a complex scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision target detection technology, and more specifically to a method and system for detecting flames and smoke based on dynamic convolution and multi-scale feature fusion. Background Technology

[0002] Visual object detection is an important research area that involves detecting the location and category of single or multiple target objects. This task processes images or video sequences, typically providing the location of targets in the image and their category labels. Over the past few decades, object detection algorithms based on Region Proposal Networks (RPNs) and Convolutional Neural Networks (CNNs) have led the development of object detection. In recent years, deep learning technology has become a core driving force for advancements in the field of visual detection. However, the complexity and diversity of object detection scenarios affect the generalization ability of detectors in practical applications. Therefore, it is necessary to develop more discriminative and robust feature representation methods to achieve accurate object detection.

[0003] The choice of feature extraction backbone determines the upper limit of detection performance. Existing object detectors typically use CNNs or visual Transformers as feature extraction backbones. Since traditional pre-trained network weights are not entirely suitable for object detection tasks, current feature extraction networks are initialized with pre-trained weights and fine-tuned on object detection datasets. Regarding feature fusion, detectors based on Region Proposal Networks (RPNs) employ the concept of RPNs to generate candidate regions to help predict object bounding boxes and class labels.

[0004] The goal of feature extraction is to obtain more accurate feature representations, while feature fusion aims to obtain highly discriminative target features. After obtaining the feature representations, the next step is to accurately predict the bounding box location and class label of the target using a detection head. Its main function is to perform bounding box regression and classification tasks. Object detection can be viewed as a two-stage learning task because it requires providing target information in the image during training and learning how to identify and locate these targets.

[0005] Currently, most mainstream object detection methods are based on multi-stage detection frameworks, such as Faster R-CNN and YOLO, which utilize different techniques to improve detection accuracy and speed. Meanwhile, Transformer-based object detection frameworks, such as DETR and YOLOv5, have successfully combined the advantages of CNN and Transformer, improving the overall performance of object detection.

[0006] However, existing methods for detecting flames and smoke using machine vision have the following drawbacks:

[0007] 1. Background Interference and Complex Environments: Flames and smoke often appear against complex backgrounds, especially indoors or outdoors, where there may be a lot of noise or moving objects (such as people, vehicles, etc.). Existing object detection methods, especially algorithms based on convolutional neural networks (CNNs), have poor robustness against complex backgrounds. The low contrast, color variations, and transparency or translucency of flames and smoke make them even more difficult to distinguish from the background, leading to a significant number of false positives and false negatives.

[0008] 2. Diversity and Dynamic Changes of Flames and Smoke: Flames and smoke are dynamic phenomena with high uncertainty. The shape and color of flames, as well as the diffusion pattern of smoke, constantly change with factors such as time, wind speed, and temperature. This high degree of uncertainty makes it difficult for target detection models to capture the time-varying characteristics of these features. Existing detection methods, especially those based on traditional CNNs, often struggle to adapt to the behavior of flames and smoke in different times and environments, resulting in low detection accuracy under some extreme conditions.

[0009] 3. Difficulty in detecting small-scale smoke: Smoke has a complex and variable morphology, especially in the early stages, where its area is relatively blurry and may be difficult to detect at low resolution. Existing target detection methods, especially those based on CNN networks, often cannot handle small-scale, blurry smoke targets well, easily resulting in missed detections and leading to the failure of fire early warning systems.

[0010] 4. Poor generalization ability: Due to the wide variation in detection scenarios for flames and smoke (e.g., indoors, outdoors, and under different weather conditions), existing target detection methods exhibit poor generalization ability across different environments. In certain specific scenarios, such as low light or low smoke concentration, the performance of the model often degrades significantly.

[0011] Therefore, how to provide a flame and smoke detection method and system based on dynamic convolution and multi-scale feature fusion is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0012] In view of this, the present invention provides a flame and smoke detection method and system based on dynamic convolution and multi-scale feature fusion to solve the technical problems existing in the prior art.

[0013] To achieve the above objectives, the present invention provides the following technical solution:

[0014] A flame and smoke detection method based on dynamic convolution and multi-scale feature fusion includes the following steps:

[0015] Step 1: Obtain the fire dataset and divide it into training and test sets after preprocessing;

[0016] Step 2: Construct a YOLOv5s model. Improve the C3 module in the YOLOv5s model to a C3-ODConv module that integrates dynamic convolution mechanism. Introduce a global context module GCB into the backbone network of the YOLOv5s model. Replace the neck feature fusion module of the YOLOv5s model with a bidirectional feature pyramid network BiFPN to obtain an improved YOLOv5s model.

[0017] Step 3: Train the improved YOLOv5s model using the training set, and test the improved YOLOv5s model using the test set;

[0018] Step 4: Detect fire data using the improved YOLOv5s model after training and testing.

[0019] Furthermore, step 2 also includes post-processing the detection results of the YOLOv5s model using the Soft-NMS algorithm, and using a Gaussian function to softly suppress the confidence related to the Intersection over Union (IoU).

[0020] Furthermore, the soft suppression of the confidence related to the Intersection over Union (IoU) using a Gaussian function includes:

[0021]

[0022] Among them, S i σ represents the confidence level, M is the bounding box with the highest confidence level, σ is a hyperparameter, and B... i This represents the remaining prediction boxes in the set.

[0023] Furthermore, the C3-ODConv module dynamically adjusts the convolutional kernel parameters through a multi-dimensional attention mechanism, including the weight allocation of spatial dimension, input channel dimension, output channel dimension, and convolutional kernel attention scalar, specifically as follows:

[0024] Spatial dimension: a si ∈R k×k Normalized by the Sigmoid function;

[0025] Input channel dimensions: Normalize using the Sigmoid function;

[0026] Output channel dimensions: Normalize using the Sigmoid function;

[0027] Convolution kernel attention scalar: a βi ∈R n×1 Normalization is achieved through the Softmax function;

[0028] Finally, the dynamic convolution result is obtained by weighted summation of the outputs of all convolution kernels;

[0029] Where, k×k, c in ×1、c out ×1 and n×1 represent the feature vectors output by the fully connected layers of the four branches, respectively.

[0030] Furthermore, the Global Context Module (GCB) enhances global feature representation by fusing non-local attention and channel attention mechanisms to capture long-distance dependencies and adaptively adjust channel weights. Specifically, global context modeling is achieved through the following steps:

[0031] Global average pooling is performed on the input feature map to generate a global context vector;

[0032] The context vector is compressed and nonlinearly activated through a bottleneck transformation layer;

[0033] The activated vector is weighted and fused with the original input feature map to generate an enhanced feature map.

[0034] Furthermore, the Bidirectional Feature Pyramid Network (BiFPN) optimizes multi-scale feature transfer through a bidirectional information fusion mechanism and a weighted feature fusion strategy, and introduces learnable dynamic weight allocation.

[0035] Furthermore, the bidirectional information fusion mechanism and weighted feature fusion strategy include:

[0036] Remove redundant single-input nodes and add cross-scale bidirectional connection paths;

[0037] A fast normalization fusion strategy is adopted, which dynamically weights the input features using learnable weights, as shown in the formula:

[0038]

[0039] Where n represents the total input features, I i O and W represent the input and output feature matrices, respectively. i Indicates input feature I i The corresponding weights are c, which is a hyperparameter, c = 0.0001, and E is the identity matrix.

[0040] A flame and smoke detection system based on dynamic convolution and multi-scale feature fusion includes:

[0041] Data acquisition module: Acquires fire datasets and preprocesses them into training and test sets;

[0042] Model building module: Construct the YOLOv5s model, improve the C3 module in the YOLOv5s model to the C3-ODConv module which integrates dynamic convolution mechanism, introduce the global context module GCB into the backbone network of the YOLOv5s model, and replace the neck feature fusion module of YOLOv5s with the bidirectional feature pyramid network BiFPN to obtain the improved YOLOv5s model.

[0043] Model training module: Trains the improved YOLOv5s model using the training set, and tests the improved YOLOv5s model using the test set;

[0044] Detection module: Fire data detection is performed using an improved YOLOv5s model that has been trained and tested.

[0045] As can be seen from the above technical solution, compared with the prior art, the present invention discloses a flame and smoke detection method and system based on dynamic convolution and multi-scale feature fusion, with the following beneficial effects:

[0046] (1) Improve detection accuracy. Flames and smoke appear in various and uncertain forms in images. Traditional detection algorithms often result in false detections or missed detections in complex backgrounds, low contrast, small targets, and dynamic changes. This invention can enhance the learning and recognition capabilities of flame and smoke features, thereby improving its detection accuracy in complex scenes.

[0047] (2) Enhance background noise processing capabilities. Flames and smoke often occur in complex backgrounds (such as indoor or outdoor environments where there may be dynamic objects and other interference), and existing detection methods often struggle to effectively distinguish between the target and the background. Improved algorithms can better separate flame and smoke targets from background noise through more advanced feature extraction and background modeling methods, thereby improving the stability and robustness of detection.

[0048] (3) Improve generalization ability. The occurrence of flames and smoke is affected by factors such as environment, weather, and lighting. Existing detection algorithms have poor generalization ability in different scenarios. This invention can enhance the model's adaptability to flames and smoke in different scenarios and environmental conditions, thereby improving the detection effect in diverse environments.

[0049] (4) Enhanced dynamic updating and adaptability. In practical applications, the morphology of flames and smoke changes rapidly, therefore the detection algorithm needs to have online updating and adaptability. This invention enables the model to adjust and optimize in real time through online learning and adaptation mechanisms, adapting to changes in flames and smoke at different times and in different environments. Attached Figure Description

[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0051] Figure 1 This is a schematic diagram of the method flow of the present invention;

[0052] Figure 2 This is a structural diagram of the improved YOLOv5s model (AFSD-YOLO network) of this invention;

[0053] Figure 3 The ODConv structure diagram provided by this invention;

[0054] Figure 4 The present invention provides a diagram of the four attention scalars of ODConv.

[0055] Figure 5 The structural diagrams of the NL module and SE module provided by this invention;

[0056] Figure 6 The simplified NL module and GCNet module structure diagram provided by this invention;

[0057] Figure 7 Comparison diagram of the three fusion methods of FPN, PAN and BiFPN provided for this invention;

[0058] Figure 8 This is a schematic diagram of the system structure provided by the present invention;

[0059] Figure 9 The comparison chart shows the detection performance of the improved YOLOv5s model (AFSD-YOLO network) provided by this invention with that of YOLOv5s, YOLOv8s, and YOLOv10s. Detailed Implementation

[0060] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0061] Example 1:

[0062] See Figure 1Embodiment 1 of this invention discloses a flame and smoke detection method based on dynamic convolution and multi-scale feature fusion, comprising the following steps:

[0063] Step 1: Obtain the fire dataset and divide it into training and test sets after preprocessing;

[0064] Step 2: Construct a YOLOv5s model. Improve the C3 module in the YOLOv5s model to a C3-ODConv module that integrates dynamic convolution mechanism. Introduce a global context module GCB into the backbone network of the YOLOv5s model. Replace the neck feature fusion module of the YOLOv5s model with a bidirectional feature pyramid network BiFPN to obtain an improved YOLOv5s model.

[0065] Step 3: Train the improved YOLOv5s model using the training set, and test the improved YOLOv5s model using the test set;

[0066] Step 4: Detect fire data using the improved YOLOv5s model after training and testing.

[0067] Specifically, the improved algorithm structure is as follows: Figure 2 As shown, the C3 module in the original network is replaced with the C3 omni-dimensional dynamic convolution (C3OD) module. This module dynamically adjusts the convolution kernel parameters by introducing a multi-dimensional attention mechanism, enabling the convolution to extract features more flexibly in the neural network. Furthermore, a Global Context-Block (GCB) module, which integrates non-local and channel attention (SE, squeeze and excitation) mechanisms, is introduced into the backbone network, effectively enhancing the model's ability to capture global information. To more efficiently fuse multi-scale features, the neck feature fusion module is replaced with a BiFPN (bidirectional feature union) module. Finally, the Soft-NMS (Soft non-maximum suppression) algorithm is used to effectively reduce the probability of false positives and false negatives in object detection, further improving the overall performance of the model.

[0068] Flame and smoke targets are typically characterized by blurriness and dynamic changes, and YOLOv5s has limitations in extracting detailed features of such complex targets. To improve the network's ability to represent target features in complex scenes, this invention improves the original network's C3 module into a C3-ODConv module that incorporates a dynamic convolution mechanism.

[0069] The C3-ODConv module introduces a multi-dimensional attention mechanism and employs a parallel strategy to learn the four dimensions of the convolutional kernel (including spatial size, number of input channels, number of output channels, and number of convolutional kernels). By flexibly learning attention in these dimensions, the C3-ODConv module can help the network extract features more accurately, further enhancing the model's learning ability. The calculation formula for the C3-ODConv module is shown in equation (1):

[0070]

[0071] In equation (1), x and y represent the input feature and output feature, respectively, and W i Let a represent the i-th convolutional kernel. βi Represents the convolution kernel W i attention scalar, a si ∈R k×k , These represent three attention scalars introduced along the spatial dimension, the input channel dimension, and the output channel dimension, respectively.

[0072] Specifically, the structure of the ODConv attention mechanism is as follows: Figure 3 As shown, the input feature map is compressed into a dataset of length C using global average pooling. in The feature vector is then compressed through a fully connected layer to extract and transform feature information. The feature vector is then activated by ReLU (rectified linear unit) to obtain a global context information vector containing the input feature map, which serves as the basis for generating multi-dimensional attention weights. Finally, four head branches are constructed, each containing a fully connected layer and a non-linear activation layer, to generate the four attention scalars of the ODConv module. Its function diagram is shown below. Figure 4 As shown. The four fully connected layers output k×k, c respectively. in ×1、c out The eigenvectors are k×1 and n×1, where k×k is the spatial kernel size, and c in and c out represents the number of input and output channels, respectively, and n represents the number of convolutional kernels. The first three branches are normalized using the sigmoid activation function to obtain the convolutional kernel W. i Three attention scalars a si a ci and a fi The fourth branch generates the convolution kernel W using the Softmax activation function. i The fourth attention scalar a βiODConv is obtained by multiplying each convolutional kernel by its corresponding attention scalar, and then summing all the weighted convolutional kernel results. Finally, ODConv is used to process the input data. By introducing an attention mechanism on different dimensions of the convolutional kernel, ODConv dynamically adjusts the convolutional kernel weights, thereby improving the model's feature representation ability and adaptive capture ability of target features, and further improving model performance.

[0073] Specifically, global context information can effectively enhance the model's ability to express local features, thereby improving the accuracy of image analysis. To further expand the network's receptive field, this invention introduces GCNet, a global context network that integrates non-local concepts and channel attention (SE) mechanisms, to establish dependencies between distant pixels. The structure diagram of the SE module is shown below. Figure 5 As shown in section (b), GCNet reduces computational complexity while maintaining model performance by calculating global context information and applying it uniformly to all locations in the input feature map. Furthermore, its lightweight design further improves the network's computational efficiency.

[0074] The basic NL (non-local) module structure is as follows: Figure 5 As shown in part (a), given the input feature map F∈R C×H×W Here, C, H, and W represent the number of channels, height, and width of the feature map, respectively. The core idea of ​​the NL module is to calculate the similarity between each position in the feature map and all other positions to construct a global attention feature map, and then perform weighted fusion of features at each position based on this attention feature map, thereby capturing global contextual information.

[0075] The specific calculation process of the NL module is shown in equation (2):

[0076]

[0077] In equation (2), i represents the query location index, j represents traversing all possible locations in the feature map, and x represents the position of the query location index. i and z i Let f(x) represent the feature vectors at position i in the input and output feature maps, respectively. i ,x j ) represents the similarity relationship between positions i and j, C(x) is the normalization factor, and W z and W v Let N be a linear transformation matrix. p This represents the total number of spatial locations in the feature map. Since the NL module needs to calculate each query bit...

[0078] The similarity between the input feature map and all other locations is calculated, so the computational and spatial complexity of the NL module is proportional to the square of the input feature map size.

[0079] Since the global context information obtained after training the NL module is independent of the query location, the generated global attention mechanism feature map is shared, while W in equation (2) is discarded. z And W v Moving it outside the attention pooling layer, we get a simplified NL module, such as... Figure 6 As shown in part (a). Based on this, and drawing on ideas from the SE module, GCNet is finally obtained. Its structure is as follows... Figure 6 For part (b), the calculation formula is shown in equation (3):

[0080]

[0081] In equation (3), W k and W v All are linear transformation matrices. Figure 6 The value of r is 16. Among them, These are the weights of the attention pooling layer. This indicates the bottleneck transform, and LN indicates layer normalization of the input.

[0082] Specifically, to more effectively enhance the feature representation capability of flame and smoke targets, this invention replaces the original FPN+PAN feature fusion module in the Neck part of the YOLOv5 network with a weighted bidirectional feature pyramid network (BiFPN). Traditional FPN uses a top-down approach to perform multi-scale feature fusion from layers P7 to P3, but it is easily limited by unidirectional information flow. PAN adds a bottom-up path to FPN to achieve bidirectional feature flow. The structures of FPN and PAN are as follows... Figure 7 As shown. However, both FPN and PAN structures have limitations. The unidirectional information flow of FPN limits the efficiency of utilizing shallow spatial information, while although PAN forms bidirectional information interaction, it lacks a dynamic weight adjustment mechanism, making it difficult to achieve effective feature fusion and limiting the network's ability to express flame and smoke features.

[0083] BiFPN is a bidirectional feature fusion structure based on an improvement upon FPN. Its core idea is to perform bottom-up feature fusion on top of top-down fusion to obtain richer bidirectional information interaction. Specifically, BiFPN simplifies the network structure by removing redundant nodes that only have a single input edge and do not participate in feature fusion; simultaneously, it adds additional connection paths between the original input and output nodes at the same feature scale to fuse more effective features. Furthermore, BiFPN achieves higher-level feature fusion by reusing features on each bidirectional path. Finally, it introduces a learnable weight mechanism to differentially assign weights to input features, allowing the network to dynamically adjust the importance of each feature according to specific task requirements. This further improves the performance of flame and smoke target detection. Through this series of structural optimizations and weight sharing strategies, BiFPN achieves more efficient and accurate feature fusion while reducing the number of parameters and computational cost.

[0084] Weighted feature fusion methods typically fall into three categories: infinite fusion, softmax-based fusion, and fast normalization fusion. Traditional infinite fusion methods can lead to training instability, while softmax-based fusion incurs high computational costs. Therefore, the BiFPN structure employs fast normalization fusion, using the ReLU activation function and regularization techniques to control the output weights within the range of 0 to 1, thus avoiding unreasonable results caused by differences between the input feature values ​​and the image. The calculation formula is shown below:

[0085]

[0086] In equation (4), n represents the total input features, I i O and W represent the input and output feature matrices, respectively. i Indicates input feature I i The corresponding weights. To ensure the stability of the calculation, the hyperparameter c = 0.0001 is set, and E is the identity matrix. Taking P6 as an example, this layer fuses two features and calculates the output value of P6. The calculation process is shown in equation (5).

[0087]

[0088] In equation (5), Resize represents the upsampling or downsampling operation during the resolution matching process. Output result P6 out The calculation process is as follows Figure 7 As shown, P6 td As intermediate features from top to bottom, P6 out This represents the output features from bottom to top, with intermediate features P6. td via P6 in and P7 inThe product is obtained by multiplying each component by its respective weight and then summing them. Conv represents the convolution operation performed during feature processing. The final output is P6. out We obtain it from the following equation (6):

[0089]

[0090] A BiFPN feature fusion network is employed to enrich the information representation of the feature map through a bidirectional feature fusion mechanism. This reduces redundant parameters and computational load, significantly improving the detection performance of flame and smoke targets while enhancing feature fusion efficiency.

[0091] Specifically, in flame and smoke detection tasks, targets often overlap or occlude each other. Traditional non-maximum suppression (NMS) algorithms, when removing redundant candidate boxes, may mistakenly delete some valid but overlapping detection boxes, leading to low target recall and missed detections of flames or smoke. The NMS calculation formula is as follows:

[0092]

[0093] In equation (7), S i Indicates the confidence level, M represents the box with the highest confidence level, and B represents the box with the highest confidence level. i This represents the remaining predicted bounding boxes in the set, where θ is the Intersection over Union (IoU) threshold, typically 0.5. Traditional Non-Maximum Suppression (NMS) algorithms compare the IoU of predicted bounding boxes to a set threshold θ, setting the confidence of bounding boxes with an IoU greater than or equal to the threshold to zero. If there are no occluded flames or smoke in the image, detection at a lower threshold usually does not result in missed detections. However, when flames or smoke are occluded, if the IoU of two boxes is higher than the threshold θ, the NMS algorithm will set the confidence of one box to zero, potentially leading to missed detections.

[0094] Specifically, the Soft-NMS algorithm of this invention introduces a unique penalty strategy, which uses a Gaussian function to apply a penalty to the confidence score S. i A positive correlation is established with the Intersection over Union (IoU). Based on the IoU value, the confidence of bounding boxes is gradually reduced rather than directly zeroed out, avoiding information loss caused by directly deleting potentially valuable detection boxes. In complex scenes where images contain overlapping or occluded targets, the Soft-NMS algorithm can more effectively retain candidate boxes with significant detection value, significantly improving the recall and matching performance for target detection. Its calculation formula is shown below:

[0095]

[0096] In equation (8), S iB represents the confidence level, σ represents the hyperparameter, which is typically 0.5. When the IoU is 0, Soft-NMS has no penalty mechanism; the larger the IoU, the stronger the penalty mechanism. i This represents the remaining prediction boxes in the set.

[0097] Specifically, by introducing the Soft-NMS algorithm, the problem of missed detection caused by erroneous deletion of prediction boxes in the traditional NMS algorithm is effectively alleviated, the recall rate and detection accuracy of flame and smoke targets are significantly improved, and the effectiveness of the AFSD-YOLO algorithm in detecting flame and smoke in complex scenes is enhanced.

[0098] On the other hand, see Figure 8 Embodiment 1 of the present invention also discloses a flame and smoke detection system based on dynamic convolution and multi-scale feature fusion, comprising:

[0099] Data acquisition module: Acquires fire datasets and preprocesses them into training and test sets;

[0100] Model building module: Build a YOLOv5s model, replace the C3 module in the YOLOv5s model with the C3-ODConv module that integrates dynamic convolution mechanism, introduce the global context module GCB into the backbone network of the YOLOv5s model, and replace the neck feature fusion module of YOLOv5s with the bidirectional feature pyramid network BiFPN to obtain the improved YOLOv5s model.

[0101] Model training module: Trains the improved YOLOv5s model using the training set, and tests the improved YOLOv5s model using the test set;

[0102] Detection module: Fire data detection is performed using an improved YOLOv5s model that has been trained and tested.

[0103] In summary, to extract features more flexibly and dynamically adjust convolutional kernel parameters, this invention replaces the C3 module in the original network with C3OD (C3 Omni-Dimensional Dynamic Convolution). C3OD, by introducing a dynamic convolution mechanism, allows the parameters of the convolutional kernel to automatically adjust according to the input features, thereby improving the network's adaptability to different scenes and targets. Unlike traditional convolutional layers, C3OD can perform feature fusion across multiple dimensions, enhancing the ability to identify complex backgrounds and dynamic targets, and improving overall detection accuracy and robustness. Simultaneously, to enhance the model's ability to capture global information, this invention introduces a Global Context Block (GCB) in the backbone network, which integrates nonlocal and channel attention (SE, Squeeze, and Excitation) mechanisms. GCB effectively improves feature representation by adaptively adjusting channel weights and capturing long-distance dependencies. By strengthening the modeling of the global context, GCB can more accurately extract target features in complex scenes, improving the model's ability to identify targets of different scales and dynamic changes. Furthermore, to more efficiently fuse multi-scale features, this invention replaces the neck feature fusion module with a Bidirectional Feature Pyramid Network (BiFPN) module. BiFPN optimizes feature transfer and integration across different scales through a bidirectional information fusion mechanism, effectively improving the representation capability of multi-scale features. This module performs feature fusion in a weighted manner, which not only improves the network's ability to perceive targets at different scales but also enhances the network's robustness in complex scenes, thereby improving overall detection accuracy and efficiency. Finally, this invention employs the Soft-NMS (Soft Non-Maximum Suppression) algorithm, effectively reducing the probability of false positives and false negatives in target detection. Unlike the traditional NMS algorithm, Soft-NMS softly suppresses candidate boxes, retaining more potential target boxes and reducing information loss due to improper threshold settings. This improvement enhances detection accuracy while strengthening the model's stability and robustness in complex environments, further improving overall detection performance.

[0104] Example 2:

[0105] The experimental environment in this embodiment is Windows 10 operating system, Intel(R) Core(TM) i5-10400F CPU@2.90GHz, 32GiB memory, and NVIDIA GeForce RTX 2080Ti GPU with 22GiB video memory. The experiment was trained for 200 epochs with a batch size of 16. The uniform input image resolution was 640×640. The SGD gradient descent algorithm was used with a learning rate of 0.01, a momentum of 0.937, and an optimizer decay of 0.0005. The experiment used a publicly available flame and smoke training set, which included labeled images of both flame and smoke objects, totaling 7501 images, including 5252 images for training, 1501 images for testing, and 751 images for validation.

[0106] To verify the detection performance of the proposed improved YOLOv5s model (AFSD-YOLO), experiments were conducted using accuracy (P), recall (R), and mean precision (P0.05). mA ), number of model parameters (P) arams The model is evaluated using both computational cost (FLOPs) and computational cost (FLOPs).

[0107]

[0108] In equations (9) and (10): TP represents the number of true examples; FN represents the number of false negatives; FP represents the number of false positives.

[0109]

[0110] Equations (11) and (12): P A Indicates average precision; P mA P represents the mean of the average precision. Ai This represents the average precision of the i-th category.

[0111] Referring to Table 1, YOLOv5, a classic single-stage target detection algorithm, has a lightweight version, YOLOv5n, which achieves accurate detection. However, its detection accuracy is far lower than the baseline model YOLOv5s, failing to effectively meet the high-precision requirements of flame and smoke detection tasks. In comparison, YOLOv8n performs better in flame and smoke detection tasks, achieving mAP@0.5 and mAP@0.5-0.95 of 0.836 and 0.623, respectively, representing improvements of 0.7% and 5% compared to YOLOv5n. However, it still suffers from missed detections when handling flames and smoke targets with varying shapes. Therefore, YOLOv8n is not the optimal choice for flame and smoke detection tasks. The mAP@0.5 of YOLOv8s is 0.9% lower than that of the baseline model YOLOv5s. The reason for this phenomenon may be that YOLOv8s has a reduced ability to extract target features of flames and smoke in complex scenes, which limits the performance of the model.

[0112] Compared to the latest YOLO series algorithms, AFSD-YOLO shows a 9.4% improvement in mAP@0.5-0.95 compared to YOLOv11n, and a 6.9% improvement in mAP@0.5. Furthermore, compared to YOLOv10n, YOLOv10s, and YOLOv9c, AFSD-YOLO achieves improvements of 22.3%, 19.4%, and 6.5% in mAP@0.5-0.95, respectively. YOLOv11, YOLOv10, and YOLOv9 do not perform well in mAP@0.5 and mAP@0.5-0.95 metrics, especially failing to accurately identify flames and smoke in complex scenes, thus failing to meet expected performance targets.

[0113] Table 1 Comparative experiments of YOLO series models on public flame and smoke datasets.

[0114]

[0115] Referring to Table 2, compared with the latest YOLO series algorithms, AFSD-YOLO improves the mAP@0.5-0.95 by 9.4% compared to YOLOv11n, and by 6.9% on mAP@0.5. Furthermore, compared with YOLOv10n, YOLOv10s, and YOLOv9c, AFSD-YOLO improves the mAP@0.5-0.95 by 22.3%, 19.4%, and 6.5%, respectively. YOLOv11, YOLOv10, and YOLOv9 do not perform well in the mAP@0.5 and mAP@0.5-0.95 metrics, especially in accurately identifying flames and smoke in complex scenes, failing to meet expected performance targets. Their detection speed is also slow, making real-time and effective detection of flames and smoke impossible. The Transformer-based RT-DETR algorithm achieved mAP@0.5-0.95 and mAP@0.5 of 0.656 and 0.862 respectively on the flame and smoke dataset, lagging behind AFSD-YOLO by 0.032. Furthermore, its FLOPs were 8 times higher than AFSD-YOLO, failing to demonstrate any significant advantage. Experimental results show that, compared to existing mainstream object detection algorithms, AFSD-YOLO achieves higher accuracy in flame and smoke detection with lower model parameter count and complexity.

[0116] Table 2 Comparison of AFSD-YOLO with mainstream object detection models

[0117]

[0118]

[0119] Specifically, to improve the detection accuracy of AFSD-YOLO, improvements were made to the backbone network, neck network, and non-maximum suppression part, and relevant comparative experiments were conducted. The experimental results are shown in Table 3. Analysis of the experimental data revealed that introducing the C3 fusion convolutional module ODConv into the backbone network achieved detection accuracies of 0.888 and 0.654 at mAP@0.5 and mAP@0.5-0.95, respectively, representing improvements of 1.5% and 2.4% compared to the baseline algorithm (YOLOv5s). Building upon this, further combining the GCB global attention mechanism resulted in detection accuracies of 0.89 and 0.657 at mAP@0.5 and mAP@0.5-0.95, respectively, representing improvements of 1.7% and 2.7% compared to the basic algorithm. Using other types of modules did not significantly improve model accuracy. Therefore, jointly introducing the C3 fusion convolutional module ODConv and the GCB global attention mechanism into the backbone network effectively enhances the model's feature representation ability and significantly improves object detection performance.

[0120] Table 3 compares experiments with different improvements to the backbone network.

[0121]

[0122]

[0123] Based on the introduction of the C3 fusion convolutional module ODConv and the GCB module, to enhance the multi-scale feature fusion capability of the AFSD-YOLO model in the target detection process, the feature fusion module FPN+PAN structure in the Neck part is replaced with a weighted bidirectional feature pyramid network BiFPN. The detection accuracy is 0.891 and 0.659 at mAP@0.5 and mAP@0.5-0.95, respectively, which is 1.8% and 2.9% higher than the benchmark algorithm. Experimental data show that the application of the BiFPN module in the Neck part can significantly improve the detection performance of AFSD-YOLO for flame and smoke targets in complex scenes.

[0124] To further improve the accuracy of the AFSD-YOLO model in flame and smoke detection, Soft-NMS was used to replace the traditional NMS method, and comparative experiments were designed with different non-maximum suppression strategies. Comparing the experimental results in Table 4, it was found that under the same network structure, Soft-NMS performed best in the target detection task, achieving mAP@0.5 and mAP@0.5-0.95 scores of 0.894 and 0.715, respectively. These represent improvements of 0.3% and 6.6% compared to the traditional NMS method. Experimental data show that Soft-NMS can more effectively preserve high-quality bounding boxes when handling overlapping targets and effectively reduce the number of false detections, thus significantly improving the overall performance of the model.

[0125] Table 4 Comparison of different nonmaximum suppression experiments

[0126]

[0127] To more comprehensively evaluate the effectiveness of each improved module in AFSD-YOLO, Table 5 presents ablation experiments of the improved modules to analyze their impact on detection accuracy. The baseline model uses the YOLOv5s algorithm, achieving a detection accuracy of 0.873 at mAP@0.5 and 0.63 at mAP@0.5-0.95. Building upon this, the addition of the C3-fused convolutional module ODConv, the global attention mechanism GCB, the weighted bidirectional feature pyramid network BiFPN, and non-maximum suppression Soft-NMS all improved detection accuracy. Furthermore, by introducing the BiFPN structure on top of the C3-fused convolutional module ODConv and the GCB module on the backbone network, the model's accuracy improved to 0.891 at mAP@0.5 and to 0.659 at mAP@0.5-0.95. Finally, by combining Soft-NMS with ODConv, GCB, and BiFPN, AFSD-YOLO achieved a detection accuracy of 0.894 at mAP@0.5, a 2.1% improvement over the baseline model, and a detection accuracy of 0.715 at mAP@0.5-0.95, an 8.5% improvement over the baseline model. This ablation experiment demonstrates that all improved modules effectively enhanced model performance.

[0128] Table 5 Ablation Experiments of the AFSD-YOLO Improved Module

[0129]

[0130] Specifically, Figure 9 This paper presents a visual comparative analysis of the improved YOLOv5s model (AFSD-YOLO) provided by this invention with the benchmark models YOLOv5s, YOLOv8s, and YOLOv10s in flame and smoke detection tasks. The figures show that flame and smoke detection tasks typically face the problem of high similarity between the target object and the background, especially under varying lighting conditions and interference from smoke backgrounds. The detection performance of the benchmark models YOLOv5s, YOLOv8s, and YOLOv10s is significantly affected in these situations. Because flames and smoke are visually similar to other objects, YOLOv5s, YOLOv8s, and YOLOv10s may experience false positives or false negatives in these situations, making it impossible to accurately distinguish flames from the background. In contrast, the AFSD-YOLO model exhibits stronger robustness when dealing with similar flame and smoke backgrounds, and can more effectively identify flame and smoke targets. Even in environments with uneven lighting or dense smoke, AFSD-YOLO can still maintain high detection accuracy and achieve precise flame and smoke detection.

[0131] This invention addresses the low accuracy problem in flame and smoke detection tasks in complex scenes by proposing the AFSD-YOLO algorithm and making targeted improvements. In the backbone network, the C3 convolutional module ODConv is fused, thereby improving the model's feature representation and information integration capabilities. A BiFPN module is introduced at the network neck, utilizing a bidirectional feature fusion mechanism to enhance the fusion capability of target features. A GCB module is used to expand the network's receptive field and enhance the representation capability of local features. Simultaneously, the Soft-NMS algorithm is employed to retain more effective information in flame and smoke detection. Experimental results show that AFSD-YOLO achieves a detection accuracy of 0.715 on the comprehensive evaluation index mAP@0.5-0.95, an improvement of 8.5% compared to the benchmark model, proving that AFSD-YOLO can achieve high accuracy in flame and smoke detection tasks in complex scenes.

[0132] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0133] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A flame and smoke detection method based on dynamic convolution and multi-scale feature fusion, characterized in that, Includes the following steps: Step 1: Obtain the fire dataset and divide it into training and test sets after preprocessing; Step 2: Construct a YOLOv5s model. Improve the C3 module in the YOLOv5s model to a C3-ODConv module that integrates dynamic convolution mechanism. Introduce a global context module GCB into the backbone network of the YOLOv5s model. Replace the neck feature fusion module of the YOLOv5s model with a bidirectional feature pyramid network BiFPN to obtain an improved YOLOv5s model. Step 3: Train the improved YOLOv5s model using the training set, and test the improved YOLOv5s model using the test set; Step 4: Detect fire data using the improved YOLOv5s model after training and testing.

2. The flame and smoke detection method based on dynamic convolution and multi-scale feature fusion according to claim 1, characterized in that, Step 2 further includes post-processing the detection results of the YOLOv5s model using the Soft-NMS algorithm, and using a Gaussian function to softly suppress the confidence related to the Intersection over Union (IoU).

3. The flame and smoke detection method based on dynamic convolution and multi-scale feature fusion according to claim 2, characterized in that, The confidence level related to the Intersection over Union (IoU) is softly suppressed using a Gaussian function. include: Among them, S i σ represents the confidence level, M is the bounding box with the highest confidence level, σ is a hyperparameter, and B... i This represents the remaining prediction boxes in the set.

4. The flame and smoke detection method based on dynamic convolution and multi-scale feature fusion according to claim 1, characterized in that, The C3-ODConv module dynamically adjusts the convolution kernel parameters through a multi-dimensional attention mechanism, including the weight allocation of spatial dimension, input channel dimension, output channel dimension, and convolution kernel attention scalar, specifically: Spatial dimension: a si ∈R k×k Normalized by the Sigmoid function; Input channel dimensions: Normalize using the Sigmoid function; Output channel dimensions: Normalize using the Sigmoid function; Convolution kernel attention scalar: a βi ∈R n×1 Normalization is achieved through the Softmax function; Finally, the dynamic convolution result is obtained by weighted summation of the outputs of all convolution kernels; Where, k×k, c in ×1、c out ×1 and n×1 represent the feature vectors output by the fully connected layers of the four branches, respectively.

5. The flame and smoke detection method based on dynamic convolution and multi-scale feature fusion according to claim 1, characterized in that, The Global Context Module (GCB) enhances global feature representation by fusing nonlocal attention and channel attention mechanisms to capture long-distance dependencies and adaptively adjust channel weights. Specifically, global context modeling is achieved through the following steps: Global average pooling is performed on the input feature map to generate a global context vector; The context vector is compressed and nonlinearly activated through a bottleneck transformation layer; The activated vector is weighted and fused with the original input feature map to generate an enhanced feature map.

6. The flame and smoke detection method based on dynamic convolution and multi-scale feature fusion according to claim 1, characterized in that, The Bidirectional Feature Pyramid Network (BiFPN) optimizes multi-scale feature transfer through a bidirectional information fusion mechanism and a weighted feature fusion strategy, and introduces learnable dynamic weight allocation.

7. The flame and smoke detection method based on dynamic convolution and multi-scale feature fusion according to claim 6, characterized in that, The bidirectional information fusion mechanism and weighted feature fusion strategy include: Remove redundant single-input nodes and add cross-scale bidirectional connection paths; A fast normalization fusion strategy is adopted, which dynamically weights the input features using learnable weights, as shown in the formula: Where n represents the total input features, I i O and W represent the input and output feature matrices, respectively. i Indicates input feature I i The corresponding weights are c, which is a hyperparameter, c = 0.0001, and E is the identity matrix.

8. A flame and smoke detection system based on dynamic convolution and multi-scale feature fusion, characterized in that, include: Data acquisition module: Acquires fire datasets and preprocesses them into training and test sets; Model building module: Construct the YOLOv5s model, improve the C3 module in the YOLOv5s model to the C3-ODConv module which integrates dynamic convolution mechanism, introduce the global context module GCB into the backbone network of the YOLOv5s model, and replace the neck feature fusion module of YOLOv5s with the bidirectional feature pyramid network BiFPN to obtain the improved YOLOv5s model. Model training module: Trains the improved YOLOv5s model using the training set, and tests the improved YOLOv5s model using the test set; Detection module: Fire data detection is performed using an improved YOLOv5s model that has been trained and tested.

Citation Information

Cited By

  • Flame detection method for intelligent operation and maintenance fire-fighting inspection robot

    CN121582548A

  • A flame detection method for an intelligent operation and maintenance fire patrol robot

    CN121582548B

  • Danger source and illegal behavior identification method based on deep learning

    CN121904707A

  • Flame target detection method fusing channel statistic pruning and adaptive feature pyramid

    CN122049340A