Video flame detection method based on context key subblocks

By adopting a context-based key subblock method in video flame detection, the problems of high computational complexity and high false alarm rate in the prior art are solved, and the flame detection effect with high accuracy and low complexity is achieved.

CN120164145AInactive Publication Date: 2025-06-17XINYANG NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510313034.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

While improving detection accuracy, the existing video flame detection methods have high computational complexity and are not suitable for deployment on resource-constrained edge devices, and there are a large number of false alarms.

Method used

The video flame detection method based on the context key subblock is adopted. The video images are extracted frame by frame, divided into image blocks with 50×50 resolution, and the fire pixels of the same type are marked and counted. The first K key pixels are selected to extract the context key subblocks, and the trained spatial channel attention network is input for classification.

Benefits of technology

It significantly improves the accuracy and efficiency of early flame detection, effectively reduces the computational complexity, and is suitable for deployment on resource-constrained edge devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164145A_ABST
    Figure CN120164145A_ABST
Patent Text Reader

Abstract

The invention discloses a video flame detection method based on context key subblocks, and relates to the technical field of video image processing. The method comprises the following steps: acquiring to-be-detected image data; segmenting the image data into a plurality of specified resolution image blocks, marking fire-like pixels in the blocks and counting the number of the fire-like pixels; sorting the image blocks in a descending order according to the number, selecting a plurality of first image blocks, extracting low three bit planes, and combining the low three bit planes; taking the fire-like pixels as the center, extracting a fixed window and calculating a fire-like pixel bit variance value, and sorting the fire-like pixels in a descending order based on the value to obtain a fire-like pixel sequence; selecting a plurality of first key pixels, taking the key pixels as a center, extracting a central sub-block and a surrounding window, calculating and normalizing correlation surfaces of the central sub-block and neighbor sub-blocks in each channel, obtaining context features, splicing the context features with the normalized central sub-block, and obtaining context key sub-blocks; and inputting the context key sub-blocks into the trained space channel attention network model to obtain a video flame detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of video image processing, and particularly to a video flame detection method based on context key sub-blocks. Background Art

[0002] With the advancement of the global digitalization process, the Internet of Things (IoT) intelligent monitoring system provides a comprehensive intelligent management solution for network transmission, environmental monitoring, and security control, and is widely used in the real-time monitoring of various disasters, such as fires, earthquakes, and floods. Due to its suddenness and destructiveness, fire has become a major hidden danger threatening human life and property safety. If not detected and addressed in a timely manner, fires will have a profound impact on the ecological environment, social order, and economic development. Therefore, the timely detection of early fires is particularly important. Early fire detection mainly relies on contact sensors, such as gas, temperature, smoke, and particulate sensors. Since such sensors will trigger an alarm only when a preset threshold is reached, the response delay is relatively long, and timely alarms for early fires cannot be achieved. In addition, contact sensors also face many challenges in large-scale outdoor environments, being vulnerable to damage and having high procurement, installation, and maintenance costs. Given the limitations of the above methods, researchers have shifted their research focus to vision-based flame detection technology. Visual flame detection (VFD) uses vision sensors and image processing technology to identify flames by analyzing each frame of the video captured by a camera. VFD generally includes traditional manual, machine learning, and deep learning methods.

[0003] Traditional manual methods usually rely on manually extracted features, such as color, shape, texture, and motion information. For example, the literature "Comparative analysis of simple rules for flame recognition" (E. Buza, E. Turajlic and A. Akagic, 2022 30th Telecommunications Forum (TELFOR), Belgrade, Serbia, pp. 1-4, 2022.) compared the flame detection performance of 20 threshold methods based on different color rules, revealing that simple color rules can effectively identify fires in urban environments. However, this method is difficult to distinguish objects with colors similar to flames, such as yellow cars, red flags, and fire hydrants, resulting in a high false alarm rate and poor robustness in detection. The introduction of machine learning methods has improved the accuracy and efficiency of flame detection. For example, the literature "BoWFire: Detection of Fire in Still Images by Integrating Pixel Color and Texture Analysis" (D.Y.T. Chino, L.P.S. Avalhais, J.F.R. Rodrigues and A.J.M. Traina, 2015 28th SIBGRAPI Conference on Graphics, Patterns and Images, Salvador, Brazil, pp. 95-102, 2015.) combined the Naive Bayes (NB) and K-Nearest Neighbor (KNN) classifiers to achieve pixel-level color and texture classification. In addition, a new semantic segmentation dataset was created, covering the application scope of intelligent transportation systems (ITS), including urban fire and wildfire images. Machine learning-based methods effectively improve the accuracy of flame detection, but further research and verification are needed in terms of environmental adaptability and robustness. The emergence of deep learning (DL) methods effectively overcomes the limitations of traditional manual methods and machine learning methods. The rapid development of DL and IoT has promoted the application of intelligent monitoring in the field of flame detection.For example, the literature "Efficient Fire Detection for Uncertain Surveillance Environment" (K. Muhammad, S. Khan, M. Elhoseny, S. Hassan Ahmed and S. Wook Baik, IEEE Transactions on Industrial Informatics, vol. 15, no. 5, pp. 3113 - 3122, 2019.) proposes an efficient flame detection method named EMN_Fire based on convolutional neural network (CNN), which is specifically designed to handle complex and uncertain surveillance environments such as smoke, haze, and snowfall. Experimental results on benchmark datasets show that EMN_Fire performs excellently in terms of detection accuracy. To efficiently deploy deep learning methods in intelligent surveillance systems, the literature "Efficient Fire Segmentation for Internet-of-Things-Assisted Intelligent Transportation Systems" (K. Muhammad, H. Ullah, S. Khan, M. Hijji and J. Lloret, IEEE Transactions on Intelligent Transportation Systems, vol. 24, no. 11, pp. 13141 - 13150, 2023.) also proposes a CNN model based on the UNet architecture for the fire segmentation task, cleverly replacing the UNet encoder with ShuffleNetV1 units, which reduces the model complexity while maintaining high accuracy. To comprehensively evaluate the performance of the proposed model, a new fire semantic segmentation dataset is created and annotated, and the experimental results show that this method achieves a good balance between model efficiency and detection performance. However, these methods will produce false alarms when detecting fire-like visual information, such as objects with sunlight reflection and street lights, etc.

[0004] In summary, current deep learning-based methods have significantly improved detection accuracy. However, this excellent result depends on complex model structures and a large number of parameters, which are not suitable for deployment on resource-constrained edge devices, and there are still a large number of false alarms. Therefore, there is an urgent need for a detection method to significantly reduce the computational complexity while improving the detection accuracy. Summary of the Invention

[0005] Based on this, it is necessary to provide a video flame detection method based on context key sub-blocks for the above technical problems.

[0006] The present invention adopts the following technical solutions:

[0007] The present invention provides a video flame detection method based on context key sub - blocks, including:

[0008] Obtain the image data in the video to be detected; divide the image data into several image blocks with a specified resolution, mark the fire - like pixels in each image block and count the number of fire - like pixels;

[0009] Sort all the image blocks in descending order according to the number of fire - like pixels, select the specified number of image blocks with the top rankings as target image blocks, convert each target image block into a grayscale image, extract the lower three - bit planes of each target image block through a binary mask and merge them; taking the marked fire - like pixels in each target image block as the center, extract a resolution window on the merged bit plane, calculate the bit variance value of the fire - like pixels within their corresponding resolution windows, and sort the fire - like pixels in each target image block in descending order according to the bit variance value to obtain a fire - like pixel sequence;

[0010] Select the first K key pixels from the fire - like pixel sequence, and taking the first K key pixels as the center, extract a central sub - block and a surrounding window with a fixed size, calculate the correlation surface of the neighbor sub - blocks within the central sub - block and the surrounding window on each channel in the RGB space and convert it into a normalized histogram to obtain context features; splice the context features with the normalized central sub - block to obtain the context key sub - block of each target image block;

[0011] Batch - input the context key sub - blocks of each target image block into a trained spatial - channel attention network model for classification to obtain the video flame detection result.

[0012] Preferably, marking the fire - like pixels in the image block specifically includes:

[0013] Given any image block X, convert X from the RGB space to the YCbCr space;

[0014] Adopt a color filtering method to mark each pixel in X in the YCbCr space, and the formula is:

[0015]

[0016] In the formula, l i,j is a label defined based on a simple color rule, is the final label defined by color filtering, Y i,j 、Cb i,j and Cr i,j are the three - channel values of the pixel at the coordinate (i, j) in X in the YCbCr space respectively;

[0017] When the pixel values in the YCbCr space conform to the rules in Formula (1) and Formula (2), the pixel is the same as the flame color, is a fire-like pixel, and is marked as 1; otherwise, the pixel is different from the flame color, is a non-fire-like pixel, and is marked as 0.

[0018] Preferably, the formula for calculating the bit variance value of fire-like pixels within their corresponding resolution windows is:

[0019]

[0020] In Formula (3), Γ i,j is the set of pixel coordinates within the 7×7 window corresponding to the pixel at coordinates (i, j), b i',j' is the bit value at coordinates (i', j'), and (i', j') belongs to Γ i,j , μ i,j is the average bit value within this window. In Formula (4), v i,j is the bit variance value of the fire-like pixel at (i, j) within this window.

[0021] Preferably, calculating the correlation surface of the central sub-block and the neighbor sub-blocks within the surrounding window on each channel in the RGB space and converting it into a normalized histogram specifically includes:

[0022] Assume that Ω K is the set of coordinates of K key pixels in the target image block X, and the key pixel coordinates (x, y) belong to Ω K ;

[0023] Calculate the similarity weights on each channel of RGB. Among them, the calculation formula for the similarity weight on the R channel is:

[0024]

[0025] In Formula (5), is the similarity weight on the R channel, and are the values of the central sub-block P x,y and P m,n on the R channel respectively, α is the normalization factor, and P m,n is the neighbor sub-block with a resolution of p×p extracted with coordinates (m, n) as the center in the window W x,y ;

[0026] According to Formula (5), the similarity weights of the G channel and the B channel in the RGB space are calculated in the same way;

[0027] Calculate the correlation surface of the R channel through the similarity weight on the R channel. The formula is:

[0028]

[0029] In formula (6) is the central sub-block P x,y On the relevant surface of the R channel, calculate the relevant surfaces of the G channel and the B channel based on formula (6) and

[0030] Convert the relevant surfaces of the R channel, G channel and B channel into normalized histograms and perform horizontal splicing to obtain context features. The formula is:

[0031]

[0032] In formula (7), is the context feature, and are the normalized histograms of the R channel, G channel and B channel respectively

[0033] Preferably, the way to splice the context feature and the normalized central sub-block is vertical splicing. The formula is:

[0034]

[0035] In formula (8), is the context feature, is the normalized central sub-block, is the context key sub-block

[0036] Preferably, the trained spatial channel attention network model includes: a channel attention module, a spatial attention module and a feature extraction network

[0037] Preferably, the channel attention module is used to calculate the attention weights of the R channel, G channel and B channel. The calculation formula is:

[0038]

[0039] In formula (9), σ is the Sigmoid activation function, represents element-wise addition, is the obtained channel attention weight; in formula (10), represents element-wise multiplication, is the channel attention feature map

[0040] Preferably, the spatial attention module is used to calculate the attention weight of each spatial position in the RGB space. The calculation formula is:

[0041]

[0042] In formula (11), σ is the Sigmoid activation function, is the spatial attention weight; in Equation (12), represents element-wise multiplication, is the spatial attention feature map.

[0043] Preferably, the feature extraction network is used to extract features from the spatial attention feature map and perform voting classification, and output the classification result of the image data.

[0044] Preferably, for the classification result of the image data, if the proportion of 1s in the classification result is greater than 0.5, then there is a flame in the image data; when the proportion of 0s is greater than 0.5, there is no flame in the image.

[0045] The above at least one technical solution adopted by the present invention can achieve the following beneficial effects:

[0046] The method of the present invention is simple, novel and unique. Images are extracted frame by frame from the video. For a certain frame of image, it is divided into several image blocks with a resolution of 50×50. Color filtering is applied to each image block, and the number of fire-like pixels in each block is marked and counted, and the image blocks are sorted in descending order according to this. Select the first several image blocks and extract the context key sub-blocks from them, and input the extracted context key sub-blocks into the trained spatial channel attention network, and obtain the classification prediction result through the majority voting method. The present invention can effectively solve the problem of untimely and inaccurate early flame detection in the field of video image processing, significantly improve the accuracy and efficiency of early flame detection, and effectively reduce the computational complexity. Description of the Drawings

[0047] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0048] Figure 1 is a schematic flowchart of a video flame detection method based on context key sub-blocks provided by the present invention;

[0049] Figure 2 is a video flame detection framework diagram of a video flame detection method based on context key sub-blocks provided by the present invention;

[0050] Figure 3 is a process diagram of extracting the context of a sub-block of a video flame detection method based on context key sub-blocks provided by the present invention;

[0051] Figure 4 is an internal structure diagram of a spatial channel attention network of a video flame detection method based on context key sub-blocks provided by the present invention. Detailed Embodiments

[0052] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in the specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.

[0053] The following will detail the technical solutions provided by each embodiment of this application in conjunction with the drawings.

[0054] Figure 1 The following is a schematic flowchart of a video flame detection method based on context key sub-blocks in the present invention, which specifically includes the following steps:

[0055] S101: Obtain the image data in the video to be tested; divide the image data into several image blocks with a specified resolution, mark the fire-like pixels in each image block, and count the number of fire-like pixels.

[0056] Specifically, extract images frame by frame from the video to be tested. For a certain frame of image, after mirror expansion, it is divided into several non-overlapping image blocks with a resolution of 50×50. Apply a color filtering method to each image block to mark the fire-like pixels and count the number of fire-like pixels in each block. Sort the image blocks in descending order according to the number of fire-like pixels. Given an image block X, convert X from the RGB space to the YCbCr space. The three channel values of the pixel at coordinates (i, j) in X are Y i,j , Cb i,j and Cr i,j , and the calculation formula of the corresponding color filtering method is:

[0057]

[0058] In formula (1), l i,j is a label defined based on simple color rules. In formula (2), is the final label defined by color filtering.

[0059] When the pixel values in the YCbCr space conform to the rules in formula (1) and formula (2), the pixel has the same color as the flame, is a fire-like pixel, and is marked as 1; otherwise, the pixel has a different color from the flame, is a non-fire-like pixel, and is marked as 0.

[0060] S102: Sort all the image patches in descending order according to the number of fire-like pixels, select the specified number of image patches with the top rankings as target image patches, convert each target image patch into a grayscale image, extract the low three bit planes of each target image patch through a binary mask and merge them; taking the marked fire-like pixels in each target image patch as the center, extract a resolution window on the merged bit plane, calculate the bit variance value of the fire-like pixels within their corresponding resolution windows, and sort the fire-like pixels within each target image patch in descending order according to the bit variance value to obtain a fire-like pixel sequence.

[0061] Specifically, extract the low three bit planes of the image patch and merge them. Taking the fire-like pixels marked based on color filtering as the center, extract a 7×7 resolution window on the merged bit plane, calculate the bit variance of the fire-like pixels within the corresponding window, and sort the fire-like pixels in descending order according to the variance value. The corresponding calculation method is as follows:

[0062]

[0063] In formula (3), Γ i,j is the set of pixel coordinates within the 7×7 window corresponding to the pixel at coordinates (i, j), and b i',j' is the bit value at coordinates (i', j'), and (i', j') belongs to Γ i,j , and μ i,j is the average bit value within this window. In formula (4), v i,j is the bit variance value of the fire-like pixel at (i, j) within this window.

[0064] S103: Select the top K key pixels from the fire-like pixel sequence, and taking the top K key pixels as the center, extract a central sub-block and surrounding windows of a fixed size, calculate the correlation planes of the neighboring sub-blocks within the central sub-block and surrounding windows on each channel in the RGB space and normalize them to obtain context features; splice the context features with the normalized central sub-block to obtain the context key sub-blocks of each target image patch.

[0065] Calculate the correlation planes of the neighboring sub-blocks within the central sub-block and surrounding windows on each channel in the RGB space and normalize them to obtain context features, including:

[0066] Assume that Ω K is the set of coordinates of the K key pixels in the target image patch X, and the key pixel coordinates (x, y) belong to Ω K ; calculate the similarity weights on each channel of RGB. Among them, the calculation formula for the similarity weight on the R channel is:

[0067]

[0068] In formula (5), is the similarity weight on the R channel, and are the values of the central sub - block P x,y and P m,n on the R channel respectively. α is the normalization factor. P m,n is the neighbor sub - block with a resolution of p×p extracted from the window W x,y centered at the coordinates (m,n); According to formula (5), the similarity weights of the G channel and B channel in the RGB space are calculated in the same way.

[0069] Through the similarity weight on the R channel, the correlation surface of the R channel is calculated, and the formula is:

[0070]

[0071] In formula (6), is the correlation surface of the central sub - block P x,y on the R channel. Based on formula (6), the correlation surfaces of the G channel and B channel are calculated and The correlation surfaces of the R channel, G channel and B channel are converted into normalized histograms and horizontally concatenated to obtain the context feature, and the formula is:

[0072]

[0073] In formula (7), is the context feature, and are the normalized histograms of the R channel, G channel and B channel respectively.

[0074] The way to concatenate the context feature and the normalized central sub - block is vertical concatenation, and the formula is:

[0075]

[0076] In formula (8), is the context feature, is the normalized central sub - block, is the context key sub - block.

[0077] Specifically, the first K key pixels are selected from the sorted pixel sequence, and the context key sub - blocks of these K pixels are extracted from the image block. The specific calculation process of this step is as Figure 2 shown. Assume that Ω K is the coordinate set of the K key pixels in X and (x,y) belongs to Ω K , for the extraction process of the context key sub - block of the pixel at the coordinate (x,y), it includes: taking the coordinate (x,y) as the center, and extracting the central sub - block P x,yand the surrounding window W with a resolution of w×w x,y , calculate P x,y and P x,y the similarity weights on each RGB channel with the neighbor sub-blocks P m,n within W; the similarity weights of each channel form the correlation map on each channel, and calculate the correlation maps of the G channel and the B channel based on formula (6) and To obtain a more compact representation, convert the correlation maps of each channel into normalized histograms, respectively and The normalized histograms of each channel form the context features of P x,y Finally, the context features are concatenated with the normalized central sub-block to obtain the context key sub-block

[0078] S104: Batch input the context key sub-blocks of each target image block into the trained spatial channel attention network model for classification to obtain the video flame detection result.

[0079] The trained spatial channel attention network model includes: a channel attention module, a spatial attention module, and a feature extraction network.

[0080] The channel attention module is used to calculate the attention weights of the R channel, the G channel, and the B channel, and the calculation formula is:[[]]

[0081]

[0082] In formula (9), σ is the Sigmoid activation function,[[]] denotes element-wise addition,[[]] is the obtained channel attention weight; in formula (10) denotes element-wise multiplication,[[]] is the channel attention feature map.

[0083] The spatial attention module is used to calculate the attention weights of each spatial position in the RGB space, and the calculation formula is:[[]]

[0084]

[0085] In formula (11), σ is the Sigmoid activation function,[[]] is the spatial attention weight; in formula (12) denotes element-wise multiplication,[[]] is the spatial attention feature map.

[0086] ​A feature extraction network is used to extract features from the spatial attention feature map and classify them, outputting the classification result of the image data. If the proportion of 1s in the classification result is greater than 0.5, there is a flame in the image data; when the proportion of 0s is greater than 0.5, there is no flame in the image.

[0087] Specifically, the extracted context key sub-blocks are batch-input into the trained spatial channel attention network. The internal structure of the spatial channel attention network is as Figure 3 shown, including a channel attention module, a spatial attention module, and a feature extraction network. The channel attention module includes a parallel pooling layer and an MLP. The parallel pooling layer includes adaptive average pooling and max pooling. The MLP includes two fully connected layers, with a ReLU activation function in the middle, and a Sigmoid activation function is connected after the MLP layer to obtain the attention weight of each channel. The spatial attention module includes a parallel pooling layer and a convolutional layer. The parallel pooling layer includes global average pooling and max pooling, and a Sigmoid activation function is connected after the convolutional layer to obtain the attention weight of each spatial position Input into this spatial attention module to obtain the spatial attention feature map Input into the feature extraction network for feature extraction and classification. The feature extraction network includes three convolutional layers and two max pooling layers. To prevent overfitting, two ReLU layers are connected before the two pooling layers respectively, and a Batch Norm layer is connected after the last convolutional layer, followed by two fully connected layers and a Sigmoid layer. After is convolved and features are extracted, it is input into the Sigmoid classification layer to obtain a list containing 0s and 1s. When the proportion of 1s in the final output list is greater than 0.5, there is a flame in the frame image; when the proportion of 0s is greater than 0.5, there is no flame in the image.

[0088] In summary, the video flame detection framework provided by the present invention is shown in Figure 2 , images are extracted frame by frame from the video to be tested. For a certain frame of image, it is divided into several image blocks with a resolution of 50×50. Based on the color filtering method, the number of fire-like pixels in each block is marked and counted, and the image blocks are sorted in descending order based on this. Select the first several image blocks, and based on the texture information, evaluate the reliability of each fire-like pixel and sort them in descending order. Select the first several fire-like pixels, extract the central sub-blocks centered on them in the image blocks and calculate the corresponding context, and merge the normalized central sub-blocks with the context to obtain the final context key sub-blocks. Input the context key sub-blocks into the trained spatial channel attention network for detection, and apply the majority voting method to the detection results to obtain the final detection result. See Figure 3, is the process diagram for extracting the context of a sub-block provided by the present invention. Centering on the fire-like pixels, a central sub-block of a fixed size and surrounding windows are extracted. Based on the self-similarity descriptor, the similarity weights between the central sub-block and each sub-block in the surrounding windows are calculated. The similarity weights are normalized into a histogram to obtain the context feature, and the normalized central sub-block is concatenated with the context feature to obtain the final context key sub-block. See Figure 4 , is the internal structure diagram of the spatial channel attention network provided by the present invention. Its internal structure includes a channel attention module, a spatial attention module, and a feature extraction network. The channel attention module includes a parallel pooling layer, an MLP layer, and a Sigmoid layer. The spatial attention module includes a parallel pooling layer, a convolutional layer, and a Sigmoid layer. The feature extraction network includes a convolutional layer, a max pooling layer, a ReLU layer, a BatchNorm layer, a fully connected layer, and a Sigmoid layer.

[0089] 1000 image patches with a resolution of 50×50 are used to train the spatial channel attention network proposed by the present invention. These 1000 image patches include 500 flame and 500 non-fire image patches. K context key sub-blocks are extracted from each image patch, and these context key sub-blocks are used to train the spatial channel attention network. 100 image patches with a resolution of 50×50 are used to evaluate the present invention and the comparative techniques. The comparative techniques are respectively:

[0090] 1) Comparative Technique 1: The literature "Comparative analysis of simple rules for flame recognition" (E. Buza, E. Turajlic and A. Akagic, 2022 30th Telecommunications Forum (TELFOR), Belgrade, Serbia, pp. 1-4, 2022.) compares the flame detection performance of 20 different color-based threshold rules and finally concludes that a set of color threshold rules in the YCbCr color space can effectively segment the flame in the image.

[0091] 2) Comparative Technique 2: The literature "BoWFire: Detection of Fire in Still Images by Integrating Pixel Color and Texture Analysis" (D.Y.T. Chino, L.P.S. Avalhais, J.F.R. Rodrigues and A.J.M. Traina, 2015 28th SIBGRAPI Conference on Graphics, Patterns and Images, Salvador, Brazil, pp. 95-102, 2015.) combines the Naive Bayes (NB) and K-Nearest Neighbor (KNN) classifiers to achieve pixel-level color and texture classification.

[0092] 3) Comparative Technique 3: The literature "Efficient Fire Detection for Uncertain Surveillance Environment" (K. Muhammad, S. Khan, M. Elhoseny, S.H. Hassan Ahmed and S. Wook Baik, IEEE Transactions on Industrial Informatics, vol. 15, no. 5, pp. 3113-3122, 2019.) presents the EMN_Fire, an efficient CNN-based fire detection method specifically designed to handle complex and uncertain surveillance environments.

[0093] 4) Comparative Technique 4: The literature "Efficient Fire Segmentation for Internet-of-Things-Assisted Intelligent Transportation Systems" (K. Muhammad, H. Ullah, S. Khan, M. Hijji and J. Lloret, IEEE Transactions on Intelligent Transportation Systems, vol. 24, no. 11, pp. 13141-13150, 2023.) describes a CNN model based on the UNet architecture that replaces the UNet encoder with ShuffleNetV1 units, reducing the model complexity while maintaining high accuracy.

[0094] The techniques proposed in the present invention and the comparative techniques are used to detect a test data set, and the precision, recall, accuracy, and F1 value are used to verify the detection performance and compare the model sizes of the model of the present invention and the models used in the comparative techniques. Referring to Table 1, it shows the comparison of the precision, recall, accuracy, F1 value, and model size for the detection of flames in videos by different techniques, listing the precision, recall, accuracy, F1 value, and model size comparison of Comparative Technique 1, Comparative Technique 2, Comparative Technique 3, Comparative Technique 4, and the video flame detection technique based on context key sub-blocks proposed in the present invention. It can be clearly seen from the experimental data that for the detection of flames in images using the present invention, its precision, recall, accuracy, and F1 value are significantly higher than those of Comparative Technique 1, Comparative Technique 2, Comparative Technique 3, and Comparative Technique 4. The precision of the present invention is 0.2476 higher than that of Comparative Technique 1, 0.294 higher than that of Comparative Technique 2, 0.1045 higher than that of Comparative Technique 3, and 0.0369 higher than that of Comparative Technique 4; the recall of the present invention is 0.06 higher than that of Comparative Technique 1, 0.59 higher than that of Comparative Technique 2, 0.2 higher than that of Comparative Technique 3, and 0.32 higher than that of Comparative Technique 4; the accuracy of the present invention is 0.22 higher than that of Comparative Technique 1, 0.34 higher than that of Comparative Technique 2, 0.14 higher than that of Comparative Technique 3, and 0.15 higher than that of Comparative Technique 4; the F1 value of the present invention is 0.1705 higher than that of Comparative Technique 1, 0.4721 higher than that of Comparative Technique 2, 0.1532 higher than that of Comparative Technique 3, and 0.2 higher than that of Comparative Technique 4. In terms of the model size, the model size of the model proposed in the present invention is significantly smaller than that of Comparative Techniques 3 and 4 based on deep learning. The model size of the present invention is reduced by 28.74 MB compared to Comparative Technique 3 and 0.22 MB compared to Comparative Technique 4. Thus, it can be shown that the present invention has a high detection accuracy for the detection of flames in videos and significantly reduces the computational complexity.

[0095] It can be clearly seen from the above that compared with the prior art, the present invention has the following outstanding technical effects: The present invention extracts images from the video to be tested frame by frame, divides the images into non-overlapping image blocks with a resolution of 50×50, marks the fire-like pixels in the image blocks based on the color filtering method and counts the number of fire-like pixels, and sorts the image blocks in descending order based on this. Select the first several sorted image blocks, extract the texture of the fire-like pixels in the blocks and sort the pixels to obtain the key fire-like pixels and extract their context key sub-blocks, and input them into the trained spatial channel attention network to detect whether there is a flame in the extracted images. Compared with the prior art, its precision can reach 0.8654, the recall rate can reach 0.9000, the accuracy rate can reach 0.8800, and the F1 value can reach 0.8824, all of which are higher than those of Comparative Technique 1, Comparative Technique 2, Comparative Technique 3, and Comparative Technique 4. The model size of the present invention is only 0.03MB, which is significantly smaller than that of Comparative Technique 3 and Comparative Technique 4. It can be seen from this that the present invention provides an effective, simple, and efficient method for early flame detection in videos, effectively ensuring the accurate and rapid detection of early fires and protecting the lives and property of humans.

[0096] Table 1 Comparison of Precision, Recall Rate, Accuracy Rate, F1 Value, and Model Size of Different Techniques for Detecting Flames in Videos

[0097]

[0098]

[0099] *The bold numerical values represent the highest values in the test result comparison.

[0100] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as falling within the scope recorded in the present invention.

Claims

1. A video flame detection method based on context key sub-blocks, characterized in that: include: Obtain image data in the video to be tested; The image data is divided into several image blocks of specified resolution, the fire-like pixels in each image block are marked and the number of fire-like pixels is counted; Sort all image blocks in descending order according to the number of fire-like pixels, select the image blocks with the highest number of pixels as target image blocks, convert each target image block into a grayscale image, extract the lower three bit planes of each target image block through a binary mask and merge them; Taking the marked fire-like pixels in each target image block as the center, extracting the resolution window on the merged bit plane, calculating the bit variance value of the fire-like pixels in its corresponding resolution window, and sorting the fire-like pixels in each target image block in descending order according to the bit variance value to obtain a fire-like pixel sequence; Select the first K key pixels from the fire-like pixel sequence, and extract the fixed-size central sub-block and surrounding window with the first K key pixels as the center. Calculate the correlation surface of the central sub-block and the neighboring sub-blocks in the surrounding window in each channel of the RGB space and normalize them to obtain the context feature; The context features are concatenated with the normalized central sub-block to obtain the context key sub-block of each target image block; The context key sub-blocks of each target image block are batch-inputted into the trained spatial channel attention network model for classification to obtain the video flame detection results.

2. A video flame detection method based on context key sub-blocks as claimed in claim 1, characterized in that, The fire-like pixels in the marked image block specifically include: Given any image block X, convert X from RGB space to YCbCr space; Using the color filtering method, each pixel in X is marked in the YCbCr space. The formula is: In the formula, l i,j For labels defined based on simple color rules, The final label defined for color filtering, Y i,j , Cb i,j and Cr i,j They are the three-channel values ​​of the pixel at coordinate (i, j) in X in the YCbCr space; When the pixel value in the YCbCr space meets the rules in formula (1) and formula (2), the pixel has the same color as the flame, is a fire-like pixel, and is marked as 1; otherwise, the pixel has a different color from the flame, is a non-fire-like pixel, and is marked as 0.

3. A video flame detection method based on context key sub-blocks as claimed in claim 1, characterized in that, The formula for calculating the bit square difference of the fire-like pixel in its corresponding resolution window is: In formula (3), Γ i,j is the pixel coordinate set in the 7×7 window corresponding to the pixel at coordinate (i, j), b i',j' is the bit value at coordinate (i', j'), and (i', j') belongs to Γ i,j , μ i,j is the bit average value in the window. i,j is the bit variance value of the fire-like pixel at (i, j) within the window.

4. A video flame detection method based on context key sub-blocks as claimed in claim 1, characterized in that, The calculation of the correlation surface between the central sub-block and the neighboring sub-blocks in the surrounding window in each channel of the RGB space and normalization specifically includes: Assume Ω K is the coordinate set of K key pixels in the target image block X, and the key pixel coordinates (x, y) belong to Ω K ; Calculate the similarity weight on each RGB channel, where the calculation formula for the similarity weight on the R channel is: In formula (5), is the similarity weight on the R channel, and The center sub-block P x,y and P m,n The value on the R channel, α is the normalization factor, P m,n For window W x,y The neighbor sub-blocks of p×p resolution extracted with the coordinates (m,n) as the center; According to formula (5), the similarity weights of the G channel and the B channel in the RGB space are calculated in the same way; The correlation surface of the R channel is calculated by the similarity weight on the R channel, and the formula is: In formula (6) The center sub-block P x,y On the correlation surface of the R channel, the correlation surface of the G channel and the B channel is calculated based on formula (6): and The R channel, G channel and B channel correlation surfaces are converted into normalized histograms and horizontally spliced ​​to obtain context features. The formula is: In formula (7), is the context feature, and They are the normalized histograms of the R channel, G channel, and B channel respectively.

5. A video flame detection method based on context key sub-blocks as claimed in claim 1, characterized in that, The context feature and the normalized central sub-block are spliced ​​vertically, and the formula is: In formula (8), is the context feature, is the normalized central sub-block, is the context key sub-block.

6. A video flame detection method based on context key sub-blocks as claimed in claim 1, characterized in that: The trained spatial channel attention network model includes: a channel attention module, a spatial attention module and a feature extraction network.

7. A video flame detection method based on context key sub-blocks as claimed in claim 6, characterized in that: The channel attention module is used to calculate the attention weights of the R channel, the G channel, and the B channel. The calculation formula is: In formula (9), σ is the Sigmoid activation function, represents element-by-element addition, is the obtained channel attention weight; in formula (10), it represents element-by-element multiplication, is the channel attention feature map.

8. A video flame detection method based on context key sub-blocks as claimed in claim 6, characterized in that: The spatial attention module is used to calculate the attention weight of each spatial position in the RGB space, and the calculation formula is: In formula (11), σ is the Sigmoid activation function, is the spatial attention weight; in formula (12), it represents element-by-element multiplication. It is the spatial attention feature map.

9. A video flame detection method based on context key sub-blocks as claimed in claim 6, characterized in that: The feature extraction network is used to extract features from the spatial attention feature map and perform voting classification, and output the image data classification result.

10. A video flame detection method based on context key sub-blocks as claimed in claim 9, characterized in that: In the classification result of the image data, if the proportion of votes of 1 in the classification result is greater than 0.5, there is flame in the image data; when the proportion of votes of 0 is greater than 0.5, there is no flame in the image.