Flame and smoke detection method based on RoseConv2d

By introducing the RoseConv2d convolution module, ECA attention mechanism and CBAM attention mechanism into the fire detection model, the problems of low fire detection accuracy and large model parameters in the existing technology are solved, and flame and smoke detection with high accuracy, low parameter volume and strong generalization capabilities are achieved.

CN120071247AActive Publication Date: 2025-05-30CHONGQING UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510136069.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-30
Estimated Expiration
2045-02-07

AI Technical Summary

Technical Problem

The existing fire detection algorithms have low accuracy when detecting small flames and facing dark environments, and the large number of model parameters lead to slow operation, long training time, and insufficient generalization capability on embedded devices.

Method used

A flame and smoke detection method based on RoseConv2d is proposed. By adding ECA attention mechanism and CBAM attention mechanism to the Fire Detection Model, and using IoU loss function for training, combined with the RoseConv2d convolution module to reduce model parameters and improve deep feature extraction capabilities.

Benefits of technology

While reducing model parameters, it improves the ability to extract deep features, improves the accuracy of flame and smoke detection, enhances the generalization ability of the model, and ensures high accuracy and low error detection rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071247A_ABST
    Figure CN120071247A_ABST
Patent Text Reader

Abstract

The invention provides a real-time flame and smoke detection method based on RoseConv2d. The real-time flame and smoke detection method comprises the following steps: S1, data acquisition: acquiring flame and smoke images in real time; s2, a flame and smoke detection model is constructed, wherein RoseConv2d, an ECA attention mechanism and a CBAM attention mechanism are added to a Fire Desection Model part of the model; s3, flame and smoke detection results are obtained, wherein the preprocessed to-be-detected flame and / or smoke images are input into a flame and smoke detection model, and the detection results are obtained; the detection result comprises a smoke value and a flame value. The detection method provided by the invention can ensure high accuracy and low false detection rate in fire detection in various scenes. Compared with an existing deep learning fire detection algorithm, the method can provide a more reliable detection effect when scenes with the background similar to the colors of flames and smoke and fire behaviors of different scales are processed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of fire detection in computer vision in artificial intelligence, and particularly to a method for detecting flames and smoke based on RoseConv2d. Background Art

[0002] In the field of fire detection, traditional technologies rely on sensors to monitor the characteristic changes in the environment during a fire, such as light perception type, temperature perception type, smoke detection type, and gas identification type, etc., to judge the fire situation. With the development of artificial intelligence technology, fire detection technology based on computer vision has emerged, which identifies flames and smoke by analyzing image or video data.

[0003] However, current fire detection algorithms face severe challenges. On the one hand, although large models can provide high detection accuracy, due to their huge number of parameters, their running speed is extremely slow, and the training process also consumes a large amount of time and resources. On the other hand, small models have a fast running speed, but their accuracy is difficult to meet the actual needs. More critically, whether it is a large or small model, their performance is unsatisfactory when detecting small flames or in a dark environment. Even large models trained with a comprehensive flame dataset still have a large room for improvement in detection accuracy. At the same time, the way some algorithms process the dataset is too complex, which not only fails to improve the detection effect but instead has a negative impact, and there is also a lack of model generalization ability on embedded devices.

[0004] To solve the above problems, especially the two core problems of insufficient accuracy and too large number of parameters, it is particularly important to innovate the convolution technology and reconstruct the algorithm. The shapes of flames and smoke are not fixed and are extremely vulnerable to environmental interference, which requires the model to have a strong ability to extract deep image information, which usually leads to an increase in the number of model parameters and poses high requirements for the performance of embedded devices. If the model size is overly compressed just to achieve lightweight, the detection accuracy will decrease accordingly, and the risks of misjudgment and missed judgment will also increase significantly. Therefore, innovating convolution and reconstructing the algorithm based on the underlying logic of the image detection algorithm is the only way to build an ideal fire detection model with lightweight, high accuracy, and low number of parameters. Summary of the Invention

[0005] The present invention aims to at least solve the technical problems existing in the prior art, and particularly innovatively proposes a method for detecting flames and smoke based on RoseConv2d.

[0006] To achieve the above object of the present invention, the present invention provides a method for detecting flames and smoke based on RoseConv2d, including the following steps:

[0007] S1, collect flame images and / or smoke images in real time and perform preprocessing;

[0008] S2. Input the image into a flame and smoke detection model to obtain a detection result, where the detection result includes the value of the flame and / or the value of the smoke.

[0009] The flame and smoke detection model includes three modules. The output end of the first part module is connected to the input end of the second part module, and the output end of the second part module is connected to the input end of the third part module. Among them, the first part module is successively composed of a convolutional layer, a normalization layer, an activation function layer, and an ECA layer. The second part module is successively composed of a residual layer, a convolutional layer, a normalization layer, an activation function layer, a CBAM layer, and a residual layer. The third part module is successively composed of an adaptive average pooling layer, a flattening layer, a fully connected layer, and an activation function layer.

[0010] The input end of the first convolutional module is connected to the output end of the input image. The input end of the first normalization module is connected to the output end of the first convolutional module. The input end of the first activation function module is connected to the output end of the first normalization module. The input end of the ECA module is connected to the output end of the first activation function module. The input end of the first residual module is connected to the output end of the ECA module for residual connection. The input end of the second convolutional module is connected to the output end of the ECA module. The input end of the second normalization layer is connected to the output end of the second convolutional module. The input end of the second activation function is connected to the output end of the second normalization module. The input end of the CBAM module is connected to the output end of the second activation function module. The input end of the second residual connection is connected to the output end of the CBAM module, and the input end of the second residual connection is connected to the output end of the first residual module. The input end of the adaptive average pooling module is connected to the output end of the second residual module. The input end of the flattening module is connected to the output end of the adaptive average pooling module. The input end of the fully connected module is connected to the output end of the flattening module. The input end of the third activation function module is connected to the output end of the fully connected module. Through these connections, the model realizes the complete process from the input image to the prediction of categories and bounding boxes.

[0011] In the application model, to solve the problem of low recognition accuracy of small flames, an ECA attention mechanism is added to its Fire Detection Model part. The concept is that the process of weight learning should be directly one-to-one corresponding. The ECA attention mechanism module directly uses a 1×1 convolutional layer after the global average pooling layer, removing the fully connected layer. This module avoids dimensional reduction and effectively captures cross-channel interactions. And ECA can achieve good results with only a few parameters. In the application model, to enhance the correlation between channels, a CBAM attention mechanism is added to its Fire Detection Model part. CBAM is used to enhance the representation ability of convolutional neural networks.

[0012] Further, the convolutional layer is Roseconv2d, which is a custom two-dimensional convolutional layer, and its construction process is as follows:

[0013] First, set the original convolution kernel. The original convolution kernel is a two-dimensional matrix, and each element in the matrix represents a weight used to perform a weighted summation operation on the input image. The height and width of the original convolution kernel are equal and regarded as a square;

[0014] Then, divide the original convolution kernel into four identical triangular convolution kernels through two diagonals. Each triangular convolution kernel is regarded as a rectangular convolution kernel. The value in the triangular inner region is x, where x is a random value based on a normal distribution, and the value outside the triangle is 0;

[0015] Next, overlap and arrange the four triangular convolution kernels into a square of the same size as the original convolution kernel: place the right angles of the four triangular convolution kernels at the upper left corner, upper right corner, lower right corner, and lower left corner in sequence; each triangular matrix will be overlapped by two adjacent triangular matrices; the formed matrix moves according to the stride, and its subsequent movement mode is the same as that of the traditional convolution.

[0016] When performing Roseconv2d convolution, the input data is respectively convolved through the four triangular convolution kernels. Each convolution kernel slides on the input data and calculates the dot product with its specific shape and weight to generate the corresponding convolution feature map; the four obtained convolution feature maps are added element by element to form the final output.

[0017] In addition, in order to further enhance the expression ability and flexibility of the model, a bias term is added to the result after addition. This bias term is a learnable parameter that allows fine-tuning of the final output. Finally, after this series of processing steps, the obtained output is the result that integrates the feature information of the four triangular convolution kernels and is adjusted by the bias.

[0018] Further, the activation function used is the RELU activation function.

[0019] Further, the flame and smoke detection model is a trained flame and smoke detection model, and the IoU loss function is used during training. Since there is a large amount of extraction of deep features in the algorithm of the present invention and high requirements for model evaluation metrics, to solve this problem, IoU is used as the loss function to accelerate model convergence and more accurately evaluate the model quality.

[0020] Further, the detection result also includes judging the fire situation:

[0021] If the value of the smoke exceeds the first smoke threshold, it is determined that there is a fire situation;

[0022] If the value of the flame exceeds the first flame threshold, it is determined that there is a fire situation;

[0023] If the value of the smoke exceeds the second smoke threshold and the value of the flame exceeds the second flame threshold, it is determined that there is a fire situation;

[0024] Wherein, the first flame threshold is greater than the second flame threshold; the first smoke threshold is greater than the second smoke threshold;

[0025] When there is a fire situation, the alarm program is started and the alarm information is sent to the terminal.

[0026] When there is a fire situation, the alarm program is started and the alarm information is sent to the terminal.

[0027] Spontaneous combustion fires and forest fires usually follow the pattern of smoke first and then fire because combustible substances produce smoke before complete combustion. However, not all fires follow this rule. For example, the spontaneous combustion of electric vehicles often shows fire first and then smoke because the chemical reaction inside the battery may directly trigger a flame, and the smoke is produced subsequently. Therefore, when designing a fire detection system, it is possible to flexibly set detection parameters and trigger conditions.

[0028] In summary, due to the adoption of the above technical solutions, the flame and smoke detection model constructed by the present invention reduces the model parameters while effectively improving the ability to extract deep features, thereby enhancing the accuracy of the model. For dynamic and variable features such as flame and smoke, the model provides an ideal solution. In addition, the model also exhibits excellent generalization ability.

[0029] Moreover, the detection method proposed by the present invention can ensure high accuracy and low false detection rate in fire detection in various scenarios. Compared with the existing deep learning fire detection algorithms, in dealing with scenarios where the background is similar in color to the flame and smoke, and different scales of fire, the present invention can provide a more reliable detection effect. This not only significantly improves the accuracy of flame detection, but also reduces the risk of misidentification, and at the same time has strong robustness, indicating its great potential in practical applications.

[0030] The additional aspects and advantages of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The above and / or additional aspects and advantages of the present invention will become obvious and easy to understand from the description of the embodiments in conjunction with the following drawings, wherein:

[0032] Figure 1Flowchart of the flame and smoke detection method based on RoseConv2d of the present invention.

[0033] Figure 2 Schematic diagram of the structure of the detection model of the present invention.

[0034] Figure 3 Schematic diagram of the RoseConv2d convolution process of the present invention.

[0035] Figure 4 Training parameter diagram of the flame and smoke detection method based on RoseConv2d of the present invention.

[0036] Figure 5 IoU loss of the result of the code training diagram of this embodiment.

[0037] Figure 6 Detection accuracy mAP@50-95 of the result of the code training diagram of this embodiment. Detailed implementation manners

[0038] The embodiments of the present invention are described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals indicate the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention.

[0039] The present invention is a flame and smoke detection method based on the RoseConv2d module, and the process is as Figure 1 shown: First, the monitoring system inputs the real-time detection video into the computer, and the computer divides the video frame by frame and then inputs it into the detection module. After the detection module detects the picture, it outputs the result to the visualization interface, recombines the single-frame pictures into a video, and displays the detection result with a visualization box. Flames and smoke will be identified separately and detection labels will be added. If there is a flame or smoke, the result of detecting a fire will be output for early warning. The method of the present invention can effectively act on environments such as commercial areas, factories, and forests, and improve the accuracy, generalization ability, and detection rate of the flame detection algorithm.

[0040] The method of the present invention is implemented based on the RoseConv2d module: By adding an ECA attention mechanism and a CBAM attention mechanism in the Fire Detection Model part, and adding an IoU loss function to the evaluation index. After training, it is deployed to the monitoring device for real-time monitoring. The structure of the flame and smoke detection method based on RoseConv2d is as Figure 2As shown in the figure. The left column is the first part module of the Fire Detection Model, which consists of a convolutional layer, a normalization layer, an activation function layer, and an ECA layer. An ECA module is added to enhance the information interaction between channels. The middle column is the second part module of the Fire Detection Model, which consists of a residual module layer, a convolutional layer, a normalization layer, an activation function layer, and a CBAM layer. The right column is the third part module of the Fire Detection Model, which consists of an adaptive average pooling layer, a flattening layer, a fully connected layer, and an activation function layer. Together, they constitute a detection algorithm based on RoseConv2d.

[0041] Since a series of unpredictable situations such as humans and weather will occur in the real video environment, the algorithm model itself needs to have good generalization. The present invention innovatively proposes the RoseConv2d convolutional module, as Figure 3 shown in the figure. Its design purpose is to process the input data through four triangular convolutional kernels with different shapes, namely the upper left triangular convolutional kernel, the upper right triangular convolutional kernel, the lower right triangular convolutional kernel, and the lower left triangular convolutional kernel. The non-triangular areas of the convolutional kernels are set to zero, and the random values in the triangular areas are based on the standard normal distribution. The input data is convolved through the four triangular convolutional kernels respectively. The convolution results are added together, and a bias term is set to form the final output. The ECA and CBAM modules are embedded in the algorithm to enhance the attention mechanism, so that both small flames and large flames can be accurately and quickly recognized during real-time detection. In terms of fire detection, for features such as flames and smoke that have no fixed form and whose targets change greatly according to the environment and there are too many environmental interference factors, it is necessary to extract deep information from the image, which inevitably increases the model parameters and is very unfriendly to embedded devices. However, if a smaller model is used for lightweight purposes, the consequence is that the image detection accuracy is insufficient, and misjudgments and missed detections occur frequently. This is not what is desired in fire detection. The reason for choosing to embed the ECA and CBAM modules is that it can not only reduce the model parameters, but also effectively improve the deep feature extraction, increase the model accuracy, and is the best choice for features such as flames and smoke.

[0042] The specific embodiments are as follows:

[0043] Step 1: Collect the data set. The data set source is the mainstream data platform on github. The data set includes a total of 30,000 images in cities, factories, families, etc. The source website of the public data set: GitHub - gaiasd / DFireDataset: D - Fire: an image data set for fire and smoke detection. These data sets are divided into a training set and a test set in a ratio of 8:2 as the input of the model. For the basic data set division, the file is named datasets.

[0044] Step 2: The algorithm scales the image size of the dataset proportionally to 256×256. The long side is reduced proportionally, and the short side is padded with gray edges. Then the picture is divided into N×N cells. After the RoseConv2d algorithm, the optimizer selects Adam. An ECA attention mechanism is added to the Fire Detection Model module. In fact, ECA is an improved version of SENET. It removes the fully connected layer in the original SENET and replaces it with a 1*1 convolutional kernel for processing, making the model parameters smaller and more lightweight. Because the convolution of ECA has good cross-channel information capture ability, it is not necessary to capture the information of all channels, so the fully connected layer is cancelled and replaced with a 1*1 convolution. ECA (ECA-Net) is a new convolutional neural network structure, and its advantages are as follows: (1) It can effectively capture long-range dependencies. ECA-Net introduces a new channel attention mechanism that can weight the features of different channels, enabling the network to better capture long-range dependencies. (2) Small number of parameters and high computational efficiency. The number of parameters of ECA-Net is smaller than that of traditional convolutional neural networks, and it can also greatly improve the computational efficiency while ensuring the accuracy. (3) It can adapt to different input sizes. ECA-Net can adapt to different input sizes, so it has good adaptability when processing images of different sizes. The core idea of the ECA module is to capture the dependencies between channels through one-dimensional convolution.

[0045] The core idea of the ECA module is to capture the dependencies between channels through one-dimensional convolution. Compared with traditional attention mechanisms, the ECA module avoids complex downsampling and upsampling processes, thus achieving efficient and lightweight characteristics. Specifically, the ECA module first adaptively calculates the kernel size k of the one-dimensional convolution according to the number of channels. The calculation formula for the kernel size is as follows:

[0046]

[0047] This formula is used to calculate the kernel size k of the one-dimensional convolution. Here, C is the number of channels of the input feature, and γ and b are hyperparameters. Taking the absolute value and rounding down to the nearest odd number is to ensure that the kernel size is odd. After obtaining the kernel size k, the ECA module applies one-dimensional convolution to the input feature to learn the importance of each channel relative to other channels. This process can be expressed by the following formula:

[0048] out=Conv1D k (in)

[0049] This formula means that the input feature in is converted into the output feature out through a one-dimensional convolution operation (kernel size k). Conv1Dk Denotes a one-dimensional convolution operation with a kernel size of k.

[0050] Add the CBAM attention mechanism to the Fire Detection Model module. CBAM consists of two sub-modules: the Channel Attention Module and the Spatial Attention Module. The Channel Attention Module is used to adjust the importance between channels, while the Spatial Attention Module is used to adjust the importance of spatial positions. This attention mechanism helps the network better capture the correlations between features and improve the model performance. Assume the input feature map is: F ∈ R C×H×W ; Use CBAM to sequentially derive the one-dimensional channel attention map M c ∈ R C×1×1 and the two-dimensional spatial attention map M s ∈ R 1×H×W , and the overall attention process can be summarized as:

[0051]

[0052] where F' is generated after the feature map F passes through the channel attention module, and F'' is generated after F' passes through the spatial attention module. M c is the channel attention module, and M s is the two-dimensional spatial attention module.

[0053] M c : Channel Attention Module (Channel Attention Module)

[0054] Utilize the channel relationships between features to generate the channel attention map. Since each channel of the feature map is considered a feature detector, the attention of the channel focuses on "what" is meaningful in the given input image; to effectively calculate the channel attention, a method of compressing the spatial dimension of the input feature map is adopted; at the same time, the methods of AvgPool (average pooling) and MaxPool (maximum pooling) are used, and it is proved that this approach is more representative than using only one pooling method;

[0055]

[0056] where σ is the sigmoid function;

[0057] W 0 、W 1 both represent weights;

[0058] W 0 ∈ R C / r×C ,W1 ∈R C×C / r ; r represents the dimensionality reduction ratio, which is used to adjust the dimension of the middle layer of the MLP, thereby controlling the model complexity and computational volume, and balancing the model performance and efficiency.

[0059] MLP() represents a multi-layer perceptron;

[0060] F represents a feature map;

[0061] represents average pooling of the feature map over the channels;

[0062] represents max pooling of the feature map over the channels;

[0063] The weights W of the MLP 0 and W 1 are shared, and the ReLU activation function is applied before W 0 ;

[0064] M s is a spatial attention module (Spatial Attention Module).

[0065] Generates a spatial attention map using the spatial relationship between features. Different from the channel attention module, the spatial attention module focuses on "where" the information part is as a supplement to the channel attention module; to calculate the spatial attention, first apply average pooling and max pooling operations along the channel axis and concatenate them to generate an effective feature descriptor; use two pooling operations to aggregate the channel information of a feature map to generate two 2D maps:

[0066] and

[0067] where represents average pooling of the feature map over the spatial features;

[0068] represents max pooling of the feature map over the spatial features;

[0069] Each represents the average pooling feature and max pooling feature of the channels, and then a standard convolutional layer is used for concatenation and convolution operations to obtain a two-dimensional spatial attention map;

[0070]

[0071] In the formula, M s (F) represents the two-dimensional spatial attention map;

[0072] σ is the sigmoid function;

[0073] [;] indicates connection through a convolutional layer;

[0074] f 7×7 () represents performing a convolution operation using a 7×7 convolutional kernel;

[0075] For the convolution operation in the TriangleConv2d class, it can be expressed as the following matrix expression: Assume the size of the input tensor x is C in ×H×W, where C in is the number of input channels, and H and W are the height and width of the input respectively. The size of the output tensor is C out ×H’×W’, where C out is the number of output channels, and H’ and W’ are the height and width of the output. For each output channel C out , each element of the output can be expressed by the following formula;

[0076]

[0077] where: K is the size of the convolutional kernel; W is the convolutional kernel matrix, which is one of the four triangular matrices here; stride is the step size; padding and dilation affect the index calculation of x. The sum of the four triangular convolutional kernels is the output result as the formula:

[0078] output = left_top_output + right_top_output + right_bottom_output

[0079] + left_bottom_output

[0080] Finally, add the bias term, as the formula:

[0081] output = output + bias[C out

[0082] bias is the bias term. The above steps describe how to implement the convolution operation through matrix operations and add the convolution results of four different shapes to obtain the final output.

[0083] Step 3: Input the dataset into the model for training, perform multiple parameter tuning based on the training results, and retain the weight file with the highest detection speed and accuracy. epochs = 10 (number of training epochs), image size = 256 (input image size), batch = 128 (batch size), workers = 2 (CPU and GPU operate together). The training results are: FPS = 50, mAP@50 - 95 = 0.5787. Obtain the training weight pth file

[0084] ​Step 4: Embed the detection program based on the optimal weight file into the monitoring device to monitor the input results in real time and give an early warning in case of fire. Embedding into the monitoring device: Apply the cv2 module, import the training weight pth file, use cap.read to read the input video of the camera frame by frame for detection, and output the detection video in real time. If fire is detected, the video will be marked and the program will give an alarm.

[0085] Thus, the real-time fire and smoke detection based on RoseConv2d is completed.

[0086] To prove the effectiveness of the method of the present invention, a control experiment was conducted. Under the same basic environment and dataset, only the network structure was changed and the module described in the present invention was added. The basic environment parameters are: Python = 3.8, CPU = i5-12600KF, GPU = NVIDIA GeForce RTX4060Ti, number of training rounds = 10, input size = 256, initial learning rate = 0.01. In the control group, YOLOv8s, YOLOv8m, YOLOv8l, and YOLOv8x do not add ECA and CBAM modules, do not modify the loss function, and do not add image enhancement modules. The experimental results are shown in Table 1. Comparison table of the training results of the present invention and the YOLOv8 model; it can be clearly seen that compared with the original YOLOv8 series algorithms, the number of parameters of the algorithm in this paper has decreased from the tens of millions level to the two hundred thousand level at most, the floating-point value has decreased from the two hundred GFLOPs level to 18 GFLOPs, and mAP@50-95 has increased from 0.46 to 0.57.

[0087] Table 1 Comparison of the training results with the YOLOv8 series algorithms

[0088] Algorithm Name Number of Parameters Floating-point Value mAP@50-95 FPS YOLOv8s 11166544 28.8 GFLOPs 0.4652 50 YOLOv8m 25902624 79.3 GFLOPs 0.4764 45 YOLOv8l 43691504 165.7 GFLOPs 0.4881 45 YOLOv8x 68229632 258.5 GFLOPs 0.4922 35 RoseConv2d 212171 18.56 GFLOPs 0.5787 50

[0089] For the reconstruction algorithm, the present algorithm adds triangular convolution, ECA, and CBAM (i.e., the last row of Table 2). Comparing the impact on the algorithm before and after adding each component, the training results of the model are shown in Table 2. Compared with traditional convolution, the number of parameters has increased by 50,000, but the floating-point value has decreased by 160 GFLOPs, and mAP@50-95 has increased from 0.34 to 0.57. It shows that the innovative convolution and optimization module of this paper play an important role in improving performance.

[0090] Table 2 Comparison of the training results of the improvement effects of each component

[0091] Add Module Number of Parameters Floating-point Value mAP@50-95 Without Addition (Traditional Convolution) 154,280 181.23 GFLOPs 0.3401 Triangular Convolution 211,558 17.73 GFLOPs 0.5628 Triangular Convolution, ECA 211,561 17.73 GFLOPs 0.5673 Triangular Convolution, CBAM 212,168 18.56 GFLOPs 0.5644 Triangular Convolution, ECA, CBAM 212171 18.56 GFLOPs 0.5787

[0092] The YOLOv8 algorithm. In terms of the number of parameters, the present invention has even fewer parameters than YOLOv8s (a lightweight model), reducing from tens of millions to hundreds of thousands. In this way, a large amount of time can be saved in model training, and the pressure on the device during model deployment will be reduced significantly. The detection accuracy of this algorithm is the best among all YOLOv8 algorithms. Compared with the traditional convolutional algorithm, the floating-point value has decreased by 160 GFLOPs, and the detection accuracy has increased by 70%. The detection speed is at a relatively high level. For a fire detection algorithm with such high accuracy and low number of parameters on a public dataset, this is worthy of recognition.

[0093] Figure 4 It is a data size and parameter diagram for the flame and smoke detection model structure. The first column on the left shows the names of each part module and the corresponding number of layers; the middle column shows the data size; the right column shows the number of parameters of this layer module. The lower part of this figure shows that the total number of parameters of the model is 212,171, and the floating-point value is 18.56 GFLOPs.

[0094] Figure 5 It is the change of the IoU_loss value of this algorithm with the increase of the number of training rounds. It can be seen that the loss finally stabilizes at 0.95.

[0095] Figure 6 It is the change of the mAP@50 - 95 of this algorithm with the increase of the number of training rounds. It can be seen that the mAP stabilizes near 0.6. Considering that if the number of training rounds is increased subsequently, the mAP is expected to break through 0.6.

[0096] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.

Claims

1. The flame and smoke detection method based on RoseConv2d is characterized by: The following steps are involved: S1, real-time acquisition of flame images and / or smoke images; S2, inputting the image into a flame and smoke detection model to obtain a detection result, where the detection result includes a flame value and / or a smoke value; The flame and smoke detection model includes three modules, the output end of the first part module is connected to the input end of the second part module, and the output end of the second part module is connected to the input end of the third part module; wherein the first part module is composed of a convolution layer, a normalization layer, an activation function layer and an ECA layer in sequence, the second part module is composed of a residual layer, a convolution layer, a normalization layer, an activation function layer, a CBAM layer and a residual layer in sequence, and the third part module is composed of an adaptive average pooling layer, a flattening layer, a fully connected layer and an activation function layer in sequence.

2. The flame and smoke detection method based on RoseConv2d according to claim 1, characterized in that: The convolutional layer is Roseconv2d, which is a two-dimensional convolutional layer, and its construction process is as follows: First, set the original convolution kernel, which is a two-dimensional matrix. Each element in the matrix represents a weight, where the height and width of the original convolution kernel are equal and is regarded as a square. Then, the original convolution kernel is divided into four identical triangular convolution kernels by two diagonal lines, where each triangular convolution kernel is regarded as a rectangular convolution kernel, the value in the triangle area is x, x is a random value based on the normal distribution, and the value outside the triangle is 0; Next, overlap the four triangular convolution kernels to form a square of the same size as the original convolution kernel: place the right angles of the four triangular convolution kernels at the upper left corner, upper right corner, lower right corner, and lower left corner in sequence; When implementing Roseconv2d convolution, the input data is convolved through four triangular convolution kernels to generate corresponding convolution feature maps; the four convolution feature maps are added element by element to form the final output.

3. The flame and smoke detection method based on RoseConv2d according to claim 1, characterized in that: The activation function used is the RELU activation function.

4. The flame and smoke detection method based on RoseConv2d according to claim 1, characterized in that: The flame and smoke detection model is a trained flame and smoke detection model, and an IoU loss function is used during training.

5. The flame and smoke detection method based on RoseConv2d according to claim 1, characterized in that: The detection results also include the judgment of fire conditions: If the smoke value exceeds the first smoke threshold, it is determined that a fire situation exists; If the flame value exceeds the first flame threshold, it is determined that a fire situation exists; If the smoke value exceeds the second smoke threshold and the flame value exceeds the second flame threshold, it is determined that a fire situation exists; Wherein, the first flame threshold is greater than the second flame threshold; The first smoke threshold is greater than the second smoke threshold; When there is a fire, the alarm program is started and the alarm information is sent to the terminal.

Citation Information

Patent Citations

  • Antenna array DOA estimation method and system based on triangular convolution and application

    CN112561033A

  • YOLOv8s traffic sign detection method based on edge and shape feature fusion

    CN118609092A

  • Real-time flame and smoke detection method based on YOLOv8

    CN119274131A