Network Model Optimization Method, System and Storage Medium for Fire Smoke Detection

Through rich fire smoke data sets and optimized object detection models, combined with CBAM attention mechanism and model pruning technology, the problems of slow fire smoke detection speed and low accuracy are solved, and more efficient fire smoke detection is achieved.

CN119206402BActive Publication Date: 2025-07-11SOUTHWEST JIAOTONG UNIV +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411366168.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-29
Publication Date
2025-07-11
Estimated Expiration
2044-09-29

AI Technical Summary

Technical Problem

The existing fire smoke detection methods have problems such as slow detection speed and low accuracy, especially in complex environments, the real-time and accuracy of the model are difficult to meet the needs.

Method used

By collecting rich fire smoke data sets, including positive and negative samples, feature images in fire smoke videos are extracted, and the target detection model is optimized using CBAM attention mechanism and model pruning technology, and appropriate image processing methods are selected to improve detection performance.

Benefits of technology

It improves the accuracy and speed of fire smoke detection, reduces the computational complexity and resource consumption, and enhances the model's ability to identify fire smoke characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119206402B_ABST
    Figure CN119206402B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of data processing, and specifically relates to a method, system, and storage medium for optimizing a network model for fire smoke detection. The method includes collecting a fire smoke data set, which includes positive samples and negative samples. It also extracts fire smoke feature images representing different states from the fire smoke video by calculating the similarity between different frames in the fire smoke video and adds them to the fire smoke data set. Analyze all the images in the fire smoke data set to obtain an analysis result, and based on the analysis result, use different processing methods for different images, and use the processed images as a training data set to train an object detection model for identifying fire smoke. During the process of training the object detection model, the CBAM attention mechanism is incorporated to extract image features, and the object detection model is optimized by adjusting parameters. The present invention can improve the recognition speed and accuracy of the model by optimizing the object detection model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and particularly relates to a method, a system and a storage medium for optimizing a network model for fire smoke detection. Background Art

[0002] With the acceleration of the urbanization process, the problem of fire safety has become increasingly prominent. Traditional fire smoke detection methods have problems such as slow detection speed and low accuracy. In recent years, deep learning technology has made remarkable progress in the field of image recognition. In particular, the YOLO series of models have performed excellently in target detection tasks. However, the detection of fire smoke requires higher real-time performance and accuracy of the model. Therefore, it is necessary to further optimize the model to adapt to complex environments. For a similar prior art, a Chinese patent application with the publication number CN110096942A provides a smoke detection algorithm based on video analysis. For the production of the training set, first, pictures containing smoke are obtained from existing forest smoke videos, the smoke-containing areas are manually labeled to produce label images, the training of the full convolutional network model is carried out, the trained network model is used to obtain the suspected smoke areas in a single-frame image, the dynamic growth analysis of the suspected smoke area area, the formula for calculating the sum of the areas of the suspected smoke areas corresponding to a continuous video sequence, and the formula for calculating the relative growth of the suspected smoke area area. However, this document does not consider the problem that the lack of richness in the training set will reduce the accuracy of the model. Another similar prior art is a Chinese patent application with the publication number CN112507925A, which discloses a fire detection method based on slow feature analysis, extracts frame images from the real-time video data of a building, preprocesses the frame images, uses slow feature analysis to extract the flame parameter features in the frame images, and according to the global flame parameter features, combines the fuzzy logic algorithm to quickly locate and extract the flame suspected areas in the frame images, and conducts local feature analysis of the flame parameters. The preprocessed frame images are divided into grids to obtain sub-images, each sub-image is matched with a standard flame image for feature points, and according to the local feature analysis results of the flame parameters and the feature point matching results, an accurate judgment is made on whether a fire has occurred. However, this document does not consider the problem of slow processing speed when processing images. Therefore, the present invention provides a method, a system and a storage medium for optimizing a network model for fire smoke detection. Summary of the Invention

[0003] The present invention prepares a fire smoke data set, including positive samples and negative samples, which increases the richness of the fire smoke data set. In order to further increase the richness of the fire smoke data set, fire smoke videos are also collected, and fire smoke feature images in different states are extracted from the fire smoke videos. Since the amount of data in the fire smoke data set is large, in order to quickly process the images in the fire smoke data set, different image processing methods are selected for the images according to different image features, so as to achieve the purpose of improving the performance of fire smoke detection.

[0004] To achieve the above-mentioned invention objectives, the present invention provides an optimization method for a network model for fire smoke detection, which is implemented by performing the following steps:

[0005] Collect a fire smoke dataset, which includes positive samples and negative samples. The positive samples include fire smoke images and simulated smoke images under different environments and conditions, and the negative samples include similar images similar to the fire smoke images. A fire smoke video is also collected;

[0006] For the obtained fire smoke video, by calculating the similarity between different frames in the fire smoke video, extract fire smoke feature images that can represent different states from the fire smoke video, and add the extracted fire smoke feature images to the fire smoke dataset;

[0007] Analyze all the images in the fire smoke dataset to obtain an analysis result, and based on the analysis result, use different processing methods for different images, and use the processed images as a training dataset to train an object detection model for identifying fire smoke;

[0008] During the process of training the object detection model, the CBAM attention mechanism is also incorporated to extract image features, the object detection model is optimized by adjusting parameters, and the model pruning technology is also applied to reduce the computational amount of the model.

[0009] As a preferred technical solution of the present invention, extracting fire smoke feature images that can represent different states from the fire smoke video includes the following steps:

[0010] Intercept multiple first images from the fire smoke video based on a preset time period, use the first first image as a standard image, convert the standard image into a second image, the second image is a single-channel image only containing luminance information, use a scaling algorithm to adjust the second image to a preset size, use a preset method to extract a third image from the second image, and also perform contour extraction on the third image to obtain a fourth image, and generate an attribute code based on the fourth image;

[0011] Obtain the second first image, generate the attribute code of the second first image, calculate the similarity between the second first image and the first first image based on the attribute code. If the similarity is greater than or equal to the first threshold, continue to obtain the third first image and generate the corresponding attribute code, and calculate the similarity between the third first image and the first first image. If the similarity is less than the first threshold, add the standard image as the smoke feature image to the fire smoke dataset, and use the third first image as the new standard image, and continue to compare it with the remaining first images until all the first images are compared.

[0012] As a preferred technical solution of the present invention, generating an attribute code based on the fourth image includes the following steps:

[0013] Divide the fourth image into multiple squares on average. Calculate the difference in grayscale values between the central pixel point and adjacent pixel points for each square to generate multiple differences, and also calculate the average value of the absolute values of the multiple differences. Compare the corresponding difference and the average value for each pixel point in the square. If the difference is greater than the average value, set the value of the corresponding pixel point to one, otherwise set the value of the corresponding pixel point to zero. Combine the values corresponding to each pixel point in each square to generate the first attribute code corresponding to the square. The third image corresponds to multiple first attribute codes, and combine the multiple first attribute codes to generate the attribute code.

[0014] As a preferred technical solution of the present invention, analyzing all the images in the fire smoke dataset to obtain an analysis result includes the following steps:

[0015] Perform a first analysis on the images in the fire smoke dataset to obtain a first analysis result. The first analysis refers to analyzing the basic attribute information of the images. The first analysis result includes the type data, size data, number of pixels, brightness data, and clarity of the images. Also perform a second analysis on the images in the fire smoke dataset and obtain a second analysis result. The second analysis refers to dividing the brightness data into multiple categories and analyzing the number of pixels in each brightness category in the image. Also perform image segmentation to obtain multiple segmented regions and analyze the number of pixels in each brightness category in each segmented region. The second analysis result includes the brightness distribution data and contrast data of the images. Also perform a third analysis on the images in the fire smoke dataset and obtain a third analysis result. The third analysis refers to analyzing the features and quality of the images. The third analysis result refers to the edge and texture features, color distribution, and whether the image is blurred of the images.

[0016] As a preferred technical solution of the present invention, different processing methods are used for different images based on the analysis results, including the following steps:

[0017] Pre-define multiple processing methods, where the multiple processing methods include image value normalization processing, brightness adjustment, sharpness adjustment, and histogram smoothing. Select several different images with different analysis results from the fire and smoke dataset as the first test images, and use different processing methods to perform image processing on the first test images respectively to obtain the processed images as the second test images. Input the different second test images into the target detection model to obtain detection results, obtain the influence of different processing methods on the detection results based on the detection results, determine the optimal processing method corresponding to different analysis results, create an optimal processing method matching table based on the analysis results and the corresponding optimal processing methods, and process the images in the fire and smoke dataset based on the optimal processing method matching table.

[0018] As a preferred technical solution of the present invention, training the target detection model includes the following steps:

[0019] Divide the training dataset into multiple data types, select several first data and several second data from the training datasets of each data type as training sample data, extract the feature vectors of each training sample in the training sample data, estimate the proportions of the first data and the second data in each data type based on the extracted feature vectors, calculate the first difference, second difference, and third difference corresponding to each data type based on the proportions, calculate the total difference based on the first difference, the second difference, and the third difference, and reduce the total difference by updating parameters during the calculation of the total difference. Repeat this step until the total difference is less than or equal to a preset second threshold.

[0020] As a preferred technical solution of the present invention, calculating the first difference, second difference, and third difference respectively based on the proportions includes the following steps:

[0021] Arbitrarily select three first data from the data belonging to the same data type, and obtain the feature vectors of the three first data. Use a preset method to calculate the sub-differences between the three feature vectors, obtain all combinations of the three first data and calculate the sub-differences of each combination, and sum up all the calculated sub-differences to obtain the first difference;

[0022] Arbitrarily select two of the first data and one of the second data from the same data type, and obtain the feature vectors. Use the preset method to calculate the sub-differences between the three feature vectors. Also, obtain all combinations of two first data and one second data and calculate the sub-differences for each combination. Sum all the sub-differences to obtain the first sub-difference sum C1. Calculate the second difference D2 using the first formula: D2 = C1 - r * a, where r is the proportion of the second data in the current data type, and a is the preset maximum difference between any two data in the same data type.

[0023] Select one of the first data and one of the second data from the same data type, and also obtain one of the second data from other data types. Obtain the corresponding feature vectors. Use the preset method to calculate the sub-differences between the three feature vectors. Also, obtain all combinations of one first data, one second data from the same data type, and one second data from any other data type, and calculate the sub-differences for each combination. Sum all the calculated sub-differences to obtain the second sub-difference sum C2. Calculate the third difference D3 using the second formula: D3 = C2 - (1 - r) * a.

[0024] As a preferred technical solution of the present invention, calculating the total difference based on the first difference, the second difference, and the third difference includes the following steps:

[0025] Calculate the total difference D using the first formula:

[0026]

[0027] where n is the number of data types, D1, D2, and D3 are the first difference, the second difference, and the third difference respectively, and α is the adjustment parameter.

[0028] The present invention also provides a network model optimization system for fire smoke detection, including the following modules:

[0029] A collection module for collecting a fire smoke data set, which includes positive samples and negative samples. The positive samples include fire smoke images and simulated smoke images under different environments and conditions, and the negative samples include similar images similar to the fire smoke images. Also collect fire smoke videos;

[0030] An expansion module for the obtained fire smoke video, by calculating the similarity between different frames in the fire smoke video, extracting fire smoke feature images that can represent different states from the fire smoke video, and adding the extracted fire smoke feature images to the fire smoke data set;

[0031] A processing module, configured to analyze all images in the fire smoke dataset to obtain an analysis result, use different processing methods for different images based on the analysis result, and use the processed images as a training dataset to train an object detection model for identifying fire smoke;

[0032] A training module, configured to, during the process of training the object detection model, also incorporate the CBAM attention mechanism to extract image features, optimize the object detection model by adjusting parameters, and also apply model pruning technology to reduce the computational amount of the model.

[0033] The present invention also provides a storage medium storing program instructions, wherein when the program instructions run, the device where the storage medium is located is controlled to execute the method described in any one of the above.

[0034] Compared with the prior art, the beneficial effects of the present invention are at least as follows:

[0035] In the present invention, first, a fire smoke dataset is prepared, including positive samples and negative samples, which increases the richness of the fire smoke dataset. To further increase the richness of the fire smoke dataset, fire smoke videos are also collected, and fire smoke feature images in different states are extracted from the fire smoke videos. Since the amount of data in the fire smoke dataset is large, in order to quickly process the images in the fire smoke dataset, different image processing methods are selected for the images according to different image features, so as to achieve the purpose of improving the performance of fire smoke detection. Finally, during the process of training the object detection model, an attention mechanism is incorporated when extracting image features. Through the combination of channel attention and spatial attention, the recognition ability of the model for fire smoke features is enhanced, and the deep learning efficiency is also improved by model pruning, so that the object detection model reduces the computational complexity and resource consumption while maintaining the performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a flowchart of the steps of the network model optimization method for fire smoke detection according to the present invention;

[0037] Figure 2 It is a structural composition diagram of the network model optimization system for fire smoke detection according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0039] It will be understood that the terms "first", "second", etc. used in the present application may be used herein to describe various elements, but unless otherwise specified, these elements are not limited by these terms. These terms are only used to distinguish a first element from another element. For example, without departing from the scope of the present application, the first xx script may be referred to as the second xx script, and similarly, the second xx script may be referred to as the first xx script.

[0040] The present invention provides a method for optimizing a network model for fire smoke detection as Figure 1 shown, which is implemented by performing the following steps:

[0041] Step S1: Collect a fire smoke data set, which includes positive samples and negative samples. The positive samples include fire smoke images and simulated smoke images under different environments and conditions, and the negative samples include similar images similar to the fire smoke images. Fire smoke videos are also collected.

[0042] Specifically, in order to monitor fire smoke, relevant data of fire smoke need to be collected first to make a fire smoke data set. In order to provide rich data resources for training and optimizing fire smoke detection, positive samples and negative samples are collected. The positive samples include fire smoke images and simulated smoke images under different environments and different conditions. The fire smoke images refer to the images taken when a real fire occurs, and the simulated smoke images refer to the fire smoke images artificially made and simulated under certain environmental conditions. In order to improve the accuracy of fire smoke detection, images that are not fire smoke images but are similar to fire smoke images are also collected as negative samples. In order to ensure the richness of the fire smoke data set, videos of fires are also collected, and image data at each stage during the fire occurrence process can be obtained from the video data.

[0043] Step S2: For the obtained fire smoke video, by calculating the similarity between different frames in the fire smoke video, extract fire smoke feature images that can represent different states from the fire smoke video, and add the extracted fire smoke feature images to the fire smoke data set.

[0044] Specifically, after obtaining the fire smoke video data, due to the continuity of the fire smoke video, there is a great similarity between adjacent frames of the fire smoke video data. If a large number of similar image data are obtained as training data, it will not only reduce the efficiency of fire smoke detection but also reduce the accuracy of fire smoke detection. Therefore, based on the similarity between different frames in the fire smoke video, fire smoke feature images that can represent different states are extracted from the fire smoke video, and the extracted fire smoke feature images are added to the fire smoke data set to increase the data volume of the fire smoke data set, providing data support for subsequent training of the fire smoke detection model and optimization of the model.

[0045] Step S3: Analyze all the images in the fire smoke dataset to obtain the analysis results, use different processing methods for different images based on the analysis results, and use the processed images as the training dataset to train the object detection model for identifying fire smoke.

[0046] Specifically, a large number of images are collected in the fire smoke dataset, and the images in the fire smoke dataset need to be analyzed and processed. To improve the performance of fire smoke detection, appropriate processing methods can be selected according to the characteristics of the images to solve the problems of long processing time and loss of data features that may be brought by using the same processing method. Therefore, the analysis results are obtained based on the analysis of the images in the fire smoke dataset, and different processing methods are used for different images based on the analysis results. The specific selection of the analysis and processing methods will be explained in detail later. The processed images are used as the training dataset to train the object detection model for identifying fire smoke.

[0047] Step S4: During the process of training the object detection model, the CBAM attention mechanism is also incorporated to extract image features, the object detection model is optimized by adjusting parameters, and the model pruning technology is also applied to reduce the computational amount of the model.

[0048] Specifically, the training dataset is used to train the object detection model. The specific training process will be explained in detail later. During the process of training the object detection model, due to the complexity and diversity of the fire smoke data, there are problems of inaccurate positioning and false detection in object detection. To enhance the accuracy and robustness of the model detection, the CBAM attention mechanism is incorporated when extracting image features. Limited by hardware resources, and the object detection model has high requirements for computing power. To reduce the computational complexity and resource consumption while maintaining high performance, the complexity and computational amount of the object detection model are also reduced through model pruning to improve the efficiency of deep learning.

[0049] Through the above method, first, the fire smoke dataset is prepared, including positive samples and negative samples, which increases the richness of the fire smoke dataset. To further increase the richness of the fire smoke dataset, fire smoke videos are also collected, and the fire smoke feature images in different states are extracted from the fire smoke videos. Since the amount of data in the fire smoke dataset is large, to quickly process the images in the fire smoke dataset, different image processing methods are selected for the images according to different image features, so as to achieve the purpose of improving the performance of fire smoke detection. Finally, during the process of training the object detection model, the attention mechanism is incorporated when extracting image features to enhance the accuracy and robustness of the model detection, and the efficiency of deep learning is also improved through model pruning, so that the object detection model reduces the computational complexity and resource consumption while maintaining the performance.

[0050] Furthermore, extracting fire smoke feature images that can represent different states from the fire smoke video includes the following steps:

[0051] Based on a preset time period, multiple first images are captured from a fire smoke video, the first first image is used as a standard image, the standard image is converted into a second image, the second image is a single-channel image containing only brightness information, a scaling algorithm is used to adjust the second image to a preset size, a third image is extracted from the second image using a preset method, contour extraction is performed on the third image to obtain a fourth image, and attribute coding is generated based on the fourth image.

[0052] Specifically, in order to quickly and accurately extract fire smoke feature images from fire smoke videos, since the differences between video frames at adjacent times are not much or even exactly the same, multiple first images are first captured from the fire smoke video based on a preset time period. The preset period is generally set to 500ms or 1s. Although there is a time interval between the multiple first images, there is also a problem of no difference between the multiple first images. If there are many identical images in the fire smoke dataset, the image processing speed will be reduced, thereby reducing the training speed of the target detection model and the recognition accuracy of the target detection model. Therefore, images with larger differences are selected from the multiple first images as fire smoke feature images and added to the fire smoke dataset.

[0053] Multiple first images are stored in the order in which they are captured. First, the first image is used as a standard image, and then the standard image is converted into a single-channel image containing only brightness information, namely a second image, which can also be called a grayscale image. In order to reduce computational complexity and improve processing speed, an image scaling algorithm is used to adjust the second image to a preset size, such as uniformly adjusting it to 720×480 pixels. The scaling algorithm is a prior art and has a variety of options such as bilinear difference. Then, in order to identify and separate the area of ​​the main object from the second image, a preset method such as Gaussian normal distribution is used to extract a third image from the second image. The third image refers to the foreground image of the second image. Then, edge mask detection is used to extract the contour of the third image to obtain a fourth image and generate attribute coding based on the fourth image. The attribute coding refers to the feature vector of the image, and the specific generation method will be explained in detail later.

[0054] Obtain the second first image, generate the attribute code of the second first image, calculate the similarity between the second first image and the first first image based on the attribute code. If the similarity is greater than or equal to the first threshold, continue to obtain the third first image and generate the corresponding attribute code, and calculate the similarity between the third first image and the first first image. If the similarity is less than the first threshold, use the standard image as the smoke feature image and add it to the fire smoke dataset, and use the third first image as the new standard image, and continue to compare it with the remaining first images until all the first images are compared.

[0055] Specifically, in order to detect the similarity between two adjacent first images, after generating the attribute code of the first first image, use the same method to generate the second first image and generate the corresponding attribute code, and calculate the similarity between the two based on the attribute code. If the similarity is greater than or equal to the first threshold, it means that the similarity between the second first image and the first first image is relatively high, and there is no need to add it as a smoke feature image to the fire smoke dataset. Then continue to obtain the attribute code of the third first image and calculate the similarity with the attribute code of the first first image, that is, the standard image. If the similarity between the two is less than the first threshold, it means that their similarity is not high. Add the standard image, that is, the first first image, to the fire smoke dataset, and use the third first image as the new standard image to continue comparing with the remaining first images. After all the first images are compared, all the first images with low similarity are added to the fire smoke dataset as fire smoke feature images, providing a data basis for training the object detection model for identifying fire smoke subsequently.

[0056] Further, generating the attribute code based on the fourth image includes the following steps:

[0057] Divide the fourth image into multiple squares on average. Calculate the difference in grayscale values between the central pixel point and adjacent pixel points for each square to generate multiple differences, and also calculate the average value of the absolute values of the multiple differences. Compare the corresponding difference and the average value for each pixel point in the square. If the difference is greater than the average value, set the value of the corresponding pixel point to one, otherwise set the value of the corresponding pixel point to zero. Combine the values corresponding to each pixel point in each square to generate the first attribute code corresponding to the square. The third image corresponds to multiple first attribute codes, and combine the multiple first attribute codes to generate the attribute code.

[0058] Specifically, in order to quickly obtain the attribute encoding of an image, first divide the fourth image into multiple squares. Taking the example of dividing the fourth image into multiple 3×3 pixel squares, calculate the differences between the central pixel and adjacent pixels for each square. For example, if the gray value of the central pixel is 32, and the gray values of the adjacent pixels adjacent to the central pixel are 75, 80, 90, 26, 56, 24, 18, 39 respectively, then the differences in gray values between the central pixel and adjacent pixels are 43, 48, 58, -6, 24, -8, -14, 7 respectively. The average value of the absolute values of multiple differences is 26. Compare the size of the difference of each pixel with the average value. Set the value of the pixel with a difference greater than the average value to 1, otherwise set it to 0. Then the corresponding first attribute encoding obtained is (1, 1, 1, 0, 1, 1, 0, 0, 1). Calculate the corresponding first attribute encoding for each square using the same method as above, and combine all the first attribute encodings to generate the attribute encoding corresponding to the fourth image. Subsequently, the attribute encoding can be used as a feature vector, and the cosine similarity calculation formula can be used to calculate the similarity between two attribute encodings. The calculated similarity is used as the similarity between two first images. The cosine similarity calculation formula is a prior art and will not be explained in detail here.

[0059] Further, analyze all the images in the fire smoke dataset to obtain the analysis results, including the following steps:

[0060] Perform a first analysis on the images in the fire smoke dataset to obtain the first analysis results. The first analysis refers to analyzing the basic attribute information of the images. The first analysis results include the type data, size data, number of pixels, brightness data, and clarity of the images. Also, perform a second analysis on the images in the fire smoke dataset and obtain the second analysis results. The second analysis refers to dividing the brightness data into multiple categories and analyzing the number of pixels in each brightness category in the image. Also, segment the image to obtain multiple segmented regions and analyze the number of pixels in each brightness category in each segmented region. The second analysis results include the brightness distribution data and contrast data of the image. Also, perform a third analysis on the images in the fire smoke dataset and obtain the third analysis results. The third analysis refers to analyzing the features and quality of the images. The third analysis results refer to the edge and texture features, color distribution, and whether the image is blurred.

[0061] Specifically, in order to select an appropriate image processing method for different image features to solve the problems of long processing time and loss of important image data that may be caused by using a unified image processing method, first perform similar feature analysis on the images using the above first analysis, second analysis, and third analysis to obtain the analysis results.

[0062] Further, based on the analysis results, use different processing methods for different images, including the following steps:

[0063] Pre - define multiple processing methods. The multiple processing methods include image value normalization, brightness adjustment, sharpness adjustment, and histogram smoothing. Select several different images with different analysis results from the fire and smoke dataset as the first test images. Use different processing methods to perform image processing on the first test images respectively, and obtain the processed images as the second test images. Input the different second test images into the target detection model to obtain detection results. Based on the detection results, obtain the influence of different processing methods on the detection results, determine the optimal processing method corresponding to different analysis results, create an optimal processing method matching table based on the analysis results and the corresponding optimal processing methods, and process the images in the fire and smoke dataset based on the optimal processing method matching table.

[0064] Specifically, to ensure the selection of the most suitable pre - processing algorithm under different image features, pre - define multiple processing methods, select several images from the fire and smoke dataset as the first test images, perform image processing on the first test images using the pre - defined multiple processing methods to obtain the processed second test images, input the second test images into the target detection model to obtain detection results, obtain the influence of different processing methods on the detection results, determine which processing method can make the detection results of the target detection model more accurate, and take the corresponding processing method as the optimal processing method for the analysis results corresponding to the first test images. Create an optimal processing method matching table based on the analysis results and the optimal processing methods. Subsequently, the optimal processing method can be selected for different images according to the optimal processing method matching table. Through the above method, it can ensure the selection of the most suitable pre - processing algorithm under different image features, and further improve the accuracy and efficiency of object recognition.

[0065] Furthermore, training the target detection model includes the following steps:

[0066] Divide the training dataset into multiple data types. Select several first data and several second data from the training datasets of each data type as training sample data. Extract the feature vectors of each training sample in the training sample data, estimate the proportions of the first data and the second data in each data type based on the extracted feature vectors, calculate the first difference, the second difference, and the third difference corresponding to each data type based on the proportions, calculate the total difference based on the first difference, the second difference, and the third difference, and reduce the total difference by updating the parameters during the calculation of the total difference. Repeat this step until the total difference is less than or equal to a preset second threshold.

[0067] Specifically, when training an object detection model, in order to avoid overfitting problems caused by unbalanced data types in the training data, for example, the amount of positive samples in the training dataset is much larger than that of negative samples, the parameters of the model are optimized by combining the proportions of different data types in the training dataset, so as to effectively train the object detection model even when the data proportions of different data types are different, and improve the accuracy of the object detection model.

[0068] Obtain a number of first data and second data from the training dataset. The first data refers to image data with a predefined label added, and the second data refers to image data without a predefined label added. Extract the feature vectors of each sample data in the training sample data. The first difference, second difference, and third difference refer to calculating the losses in different cases within the same data type using a loss function, and considering the proportion of the first data and the second data in the same data type. The specific calculation methods of the first difference, second difference, and third difference will be explained in detail later. When calculating the total difference, the proportion is also combined as a weight to weight the losses of each data type, so that for the data type with a larger proportion of the second data, its contribution to the calculation of the difference in the total difference is also greater. The total difference refers to the total loss, which reflects the performance of the model under the current parameters. During the process of calculating the total difference, the parameters are updated to reduce the total difference. The smaller the total difference, the better the performance of the model. Repeat the above process until the total difference is less than or equal to a preset second threshold, indicating that the performance of the model has reached the expectation, and the training of the model can be stopped.

[0069] Furthermore, calculate the first difference, second difference, and third difference respectively based on the proportion, including the following steps:

[0070] Arbitrarily select three first data from the data belonging to the same data type, and obtain the feature vectors of the three first data. Use a preset method to calculate the sub-differences between the three feature vectors. Obtain all combinations of the three first data and calculate the sub-differences of each combination. Sum up all the calculated sub-differences to obtain the first difference;

[0071] Arbitrarily select two first data and one second data from the same data type, and obtain the feature vectors. Use a preset method to calculate the sub-differences between the three feature vectors. Also obtain all combinations of two first data and one second data and calculate the sub-differences of each combination. Sum up all the sub-differences to obtain the first sub-difference C1, and calculate the second difference D2 using the first formula. The first formula is: D2 = C1 - r * a, where r refers to the proportion of the second data in the current data type, and a refers to the preset maximum difference between any two data in the same data type;

[0072] Select a first data and a second data from the same data type, and also obtain a second data from another data type. Obtain the corresponding feature vectors, calculate the sub-differences between the three feature vectors using a preset method. Also obtain combinations of a first data, a second data from the same data type, and a second data from any other data type, and calculate the sub-differences for each combination. Sum all the calculated sub-differences to obtain the second sub-difference sum C2, and calculate the third difference D3 using the second formula, where the second formula is D3 = C2 - (1 - r) * a.

[0073] Specifically, the preset method refers to obtaining three feature vectors x, y, z, calculating the difference S(x, y) between x and y and the difference S(x, z) between x and z through a similarity calculation formula, and calculating the sub-difference d = max{(S(x, y) - S(x, z) + a), 0} using the following formula. Calculate the first difference, second difference, and third difference through the above steps. The first difference refers to the difference calculated for the first data, i.e., the labeled data, in the same data type. The second difference takes into account the first data and the second data, aiming to use the second data to improve the learning process. The third difference also takes into account the first data and the second data, but it focuses on the differences between different data types. The calculations of the first difference, second difference, and third difference all depend on specific combinations of data, through which the model learns how to distinguish the differences within the same data type and the differences between different data types simultaneously during the training process, achieving the purpose of reinforcement learning.

[0074] Furthermore, calculate the total difference based on the first difference, second difference, and third difference, including the following steps:

[0075] Calculate the total difference D using the third formula, where the third formula is:

[0076]

[0077] where n is the number of data types, D1, D2, D3 are the first difference, second difference, and third difference respectively, and α is the adjustment parameter.

[0078] Specifically, calculate the total difference using the above third formula. During the calculation of the total difference, update the adjustment parameter to minimize the total difference, and gradually optimize the adjustment parameter during the iteration process to improve the recognition ability of the target detection model for fire smoke.

[0079] According to another aspect of the embodiments of the present invention, as shown in Figure 2 also provided is a network model optimization system for fire smoke detection, including a collection module, an expansion module, a processing module, and a training module, used to implement the network model optimization method for fire smoke detection described above. The specific functions of each module are as follows:

[0080] A collection module, configured to collect a fire smoke dataset, where the fire smoke dataset includes positive samples and negative samples. The positive samples include fire smoke images and simulated smoke images under different environments and conditions, and the negative samples include similar images similar to the fire smoke images. The fire smoke videos are also collected.

[0081] An extension module, configured to, for the obtained fire smoke video, extract fire smoke feature images capable of representing different states from the fire smoke video by calculating the similarity between different frames in the fire smoke video, and add the extracted fire smoke feature images to the fire smoke dataset.

[0082] A processing module, configured to analyze all the images in the fire smoke dataset to obtain an analysis result, and use different processing methods for different images based on the analysis result, and use the processed images as a training dataset to train an object detection model for recognizing fire smoke.

[0083] A training module, configured to, during the process of training the object detection model, also incorporate the CBAM attention mechanism to extract image features, optimize the object detection model by adjusting parameters, and also apply model pruning technology to reduce the computational amount of the model.

[0084] According to another aspect of the embodiments of the present invention, a storage medium is further provided. The storage medium stores program instructions, where, when the program instructions run, the device where the storage medium is located is controlled to execute the network model optimization method for fire smoke detection in any one of the above.

[0085] In summary, the network model optimization method, system and storage medium for fire smoke detection of the present invention. The method includes collecting a fire smoke dataset, where the dataset includes positive samples and negative samples. It also extracts fire smoke feature images capable of representing different states from the fire smoke video by calculating the similarity between different frames in the fire smoke video, and adds them to the fire smoke dataset. Analyze all the images in the fire smoke dataset to obtain an analysis result, and use different processing methods for different images based on the analysis result, and use the processed images as a training dataset to train an object detection model for recognizing fire smoke. During the process of training the object detection model, incorporate the CBAM attention mechanism to extract image features, and optimize the object detection model by adjusting parameters. The present invention can improve the recognition speed and recognition accuracy of the model by optimizing the object detection model.

[0086] It should be understood that although the steps in the flowcharts of the embodiments of the present invention are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in each embodiment may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0087] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The above program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0088] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0089] The above-mentioned embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent for the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent for the present invention shall be subject to the appended claims.

[0090] The above is only the preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for optimizing a network model for fire smoke detection, characterized in that, It includes the following steps: Collect a fire smoke dataset, which includes positive samples and negative samples. The positive samples include fire smoke images and simulated smoke images under different environments and conditions, and the negative samples include similar images similar to fire smoke images. Also collect fire smoke videos; For the obtained fire smoke videos, by calculating the similarity between different frames in the fire smoke videos, extract fire smoke feature images that can represent different states from the fire smoke videos, and add the extracted fire smoke feature images to the fire smoke dataset; Analyze all the images in the fire smoke dataset to obtain an analysis result, and based on the analysis result, use different processing methods for different images, and use the processed images as a training dataset to train an object detection model for identifying fire smoke; During the process of training the object detection model, also incorporate the CBAM attention mechanism to extract image features, optimize the object detection model by adjusting parameters, and also apply model pruning techniques to reduce the computational amount of the model; Among them, analyzing all the images in the fire smoke dataset to obtain an analysis result includes the following steps: Conduct a first analysis on the images in the fire smoke dataset to obtain a first analysis result. The first analysis refers to analyzing the basic attribute information of the images, and the first analysis result includes the type data, size data, number of pixels, brightness data, and clarity of the images. Also conduct a second analysis on the images in the fire smoke dataset and obtain a second analysis result. The second analysis refers to dividing the brightness data into multiple categories and analyzing the number of pixels in each brightness category in the image. Also segment the image to obtain multiple segmented regions and analyze the number of pixels in each brightness category in each segmented region. The second analysis result includes the brightness distribution data and contrast data of the image. Also conduct a third analysis on the images in the fire smoke dataset and obtain a third analysis result. The third analysis refers to analyzing the features and quality of the images, and the third analysis result refers to the edge and texture features, color distribution, and whether the image is blurred of the image.

2. The method according to claim 1, wherein Extracting fire smoke feature images that can represent different states from the fire smoke videos includes the following steps: Intercept multiple first images from the fire smoke video based on a preset time period. Take the first of the first images as a standard image, convert the standard image into a second image, where the second image is a single-channel image containing only brightness information. Use a scaling algorithm to adjust the second image to a preset size. Use a preset method to extract a third image from the second image. Also perform contour extraction on the third image to obtain a fourth image, and generate an attribute code based on the fourth image; Obtain the second first image, generate the attribute code of the second first image, calculate the similarity between the second first image and the first first image based on the attribute code. If the similarity is greater than or equal to the first threshold, continue to obtain the third first image and generate the corresponding attribute code, and calculate the similarity between the third first image and the first first image. If the similarity is less than the first threshold, add the standard image as the smoke feature image to the fire smoke dataset, and use the third first image as the new standard image, and continue to compare it with the remaining first images until all the first images are compared.

3. The method according to claim 2, characterized in that Generating an attribute code based on the fourth image includes the following steps: Divide the fourth image into multiple squares on average. Calculate the difference in gray values between the central pixel point and adjacent pixel points for each square to generate multiple differences. Also calculate the average value of the absolute values of the multiple differences. Compare the corresponding difference and the average value for each pixel point in the square. If the difference is greater than the average value, set the value of the corresponding pixel point to one, otherwise set the value of the corresponding pixel point to zero. Combine the values corresponding to each pixel point in each square to generate the first attribute code corresponding to the square. The third image corresponds to multiple first attribute codes. Combine the multiple first attribute codes to generate the attribute code.

4. The method according to claim 1, characterized in that And use different processing methods for different images based on the analysis results, including the following steps: Pre-define multiple processing methods, where the multiple processing methods include image value normalization processing, brightness adjustment, sharpness adjustment, and histogram smoothing. Select several different images with different analysis results from the fire smoke dataset as the first test images, and perform image processing on the first test images using different processing methods respectively. Obtain the processed images as the second test images. Input the different second test images into the target detection model to obtain the detection results. Based on the detection results, obtain the influence of different processing methods on the detection results, determine the optimal processing method corresponding to different analysis results, create an optimal processing method matching table based on the analysis results and the corresponding optimal processing methods, and process the images in the fire smoke dataset based on the optimal processing method matching table.

5. The method according to claim 1, wherein Training the target detection model includes the following steps: Divide the training data set into multiple data types, select a number of first data and a number of second data from the training data set of each data type as training sample data, extract the feature vectors of each training sample in the training sample data, estimate the proportions of the first data and the second data in each data type based on the extracted feature vectors, calculate the first difference, second difference, and third difference corresponding to each data type based on the proportions, calculate the total difference based on the first difference, the second difference, and the third difference, and reduce the total difference by updating parameters during the calculation of the total difference. Repeat this step until the total difference is less than or equal to a preset second threshold.

6. The method according to claim 5, characterized in that, Calculating the first difference, second difference, and third difference respectively based on the proportions includes the following steps: Arbitrarily select three first data from the data belonging to the same data type, and obtain the feature vectors of the three first data. Use a preset method to calculate the sub-differences between the three feature vectors, obtain all combinations of the three first data and calculate the sub-differences of each combination, and sum up all the calculated sub-differences to obtain the first difference; Arbitrarily select two of the first data and one of the second data from the same data type, and obtain the feature vectors. Use the preset method to calculate the sub-differences between the three feature vectors. Also obtain all combinations of two first data and one second data and calculate the sub-differences of each combination, and sum up all the sub-differences to obtain the first sub-difference and C1. Calculate the second difference D2 using the first formula, and the first formula is: D2 = C1 - r * a, where r refers to the proportion of the second data in the current data type, and a refers to the preset maximum difference between any two data in the same data type; Select one of the first data and one of the second data from the same data type, and also obtain one of the second data from other data types. Obtain the corresponding feature vectors, use the preset method to calculate the sub-differences between the three feature vectors. Also obtain all combinations of one of the first data, one of the second data in the same data type, and one of the second data from any other data type, and calculate the sub-differences of each combination, and sum up all the calculated sub-differences to obtain the second sub-difference and C2. Calculate the third difference D3 using the second formula, and the second formula is: D3 = C2 - (1 - r) * a.

7. The method according to claim 5, wherein Calculating the total difference based on the first difference, the second difference, and the third difference includes the following steps: Calculate the total difference D using the third formula, and the third formula is: , where n is the number of data types, which are respectively the first difference, the second difference, and the third difference, and refer to adjustment parameters.

8. A network model optimization system for fire smoke detection, which is used to implement the network model optimization method for fire smoke detection described in any one of claims 1-7, and is characterized in that, Includes the following modules: A collection module for collecting a fire smoke data set, where the fire smoke data set includes positive samples and negative samples. The positive samples include fire smoke images and simulated smoke images under different environments and conditions, and the negative samples include similar images similar to the fire smoke images. Also collect fire smoke videos; An expansion module, which is used for the obtained fire smoke video. By calculating the similarity between different frames in the fire smoke video, fire smoke feature images representing different states are extracted from the fire smoke video, and the extracted fire smoke feature images are added to the fire smoke dataset; A processing module, which is used to analyze all the images in the fire smoke dataset to obtain an analysis result, and based on the analysis result, use different processing methods for different images, and use the processed images as a training dataset to train an object detection model for identifying fire smoke. Among them, analyzing all the images in the fire smoke dataset to obtain an analysis result includes the following steps: performing a first analysis on the images in the fire smoke dataset to obtain a first analysis result. The first analysis refers to analyzing the basic attribute information of the images, and the first analysis result includes the type data, size data, number of pixels, brightness data, and clarity of the images. Also, performing a second analysis on the images in the fire smoke dataset and obtaining a second analysis result. The second analysis refers to dividing the brightness data into multiple categories and analyzing the number of pixels in each brightness category in the images. Also, segmenting the images to obtain multiple segmented regions and analyzing the number of pixels in each brightness category in each segmented region. The second analysis result includes the brightness distribution data and contrast data of the images. Also, performing a third analysis on the images in the fire smoke dataset and obtaining a third analysis result. The third analysis refers to analyzing the features and quality of the images, and the third analysis result refers to the edge and texture features, color distribution, and whether the image is blurred; A training module, which is used to also incorporate the CBAM attention mechanism to extract image features during the process of training the object detection model, optimize the object detection model by adjusting parameters, and also apply model pruning techniques to reduce the computational amount of the model.

9. A storage medium, characterized in that, The storage medium stores program instructions. Among them, when the program instructions run, they control the device where the storage medium is located to execute the network model optimization method for fire smoke detection according to any one of claims 1-7.

Citation Information

Patent Citations

  • Smoke detection algorithm based on video analysis

    CN110096942A

  • Fire detection method based on slow feature analysis

    CN112507925A

  • Method and device for identifying few-sample forest fire smoke based on improved prototype network

    CN112989932A