A Fire Image Recognition Method Based on the Mutual Information of Color Features

Through a method based on the mutual information amount of color features, combined with the probability density estimation of the kernel method and the 160-layer single hidden layer fully connected ANN network model, the shortcomings of the existing fire image recognition technology in terms of computing time, memory and recognition accuracy are solved, and efficient and accurate fire image recognition is achieved.

CN115620065BActive Publication Date: 2025-06-17CIVIL AVIATION FLIGHT UNIV OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211320911.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-26
Publication Date
2025-06-17
Estimated Expiration
2042-10-26

AI Technical Summary

Technical Problem

The existing fire image recognition technology has shortcomings in computing time, memory and recognition accuracy, and there are few researches on recognition methods based on machine learning.

Method used

The fire image recognition method based on the mutual information amount of color features is adopted. The color cast factor and variance are extracted as color features through three color modes: Lab, RGB, and HSV. The feature selection is performed by combining the probability density estimation of the kernel method and the mutual information amount. The 160-layer single hidden layer fully connected ANN network model is used for training.

Benefits of technology

It achieves high recognition efficiency and accuracy, with recognition accuracy as high as 93.82%, Pierman's correlation coefficient as high as 0.8747, and the training cost is only 256 seconds, reducing the redundancy of the input data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115620065B_ABST
    Figure CN115620065B_ABST
Patent Text Reader

Abstract

The invention discloses a method for fire image recognition based on the mutual information of color features. First, a fire image library and a non-fire image library are constructed based on real-scene shooting and a forest fire image library, and at the same time, the color deviation factors and variances of the images in three color models, namely Lab, RGB, and HSV, are extracted as color feature information. Secondly, the color feature combination is optimized based on mutual information, and the optimized color combination features are used as input data. Finally, an identification model is trained based on a 160-layer single-hidden-layer fully connected network, and at the same time, the parameters of the training model are optimized to complete the fire image recognition. Compared with the color features of traditional fire image recognition, the present invention selects the color deviation factors and variances in three color models, namely Lab, RGB, and HSV, as color features. At the same time, in order to better reduce the computational cost caused by multiple inputs, the present invention also performs feature selection based on mutual information, which well reduces the redundancy of the input data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fire image recognition, and particularly relates to a fire image recognition method based on the mutual information of color features. Background Art

[0002] Fire image recognition, as a difficult point in the field of fire accident prevention, is widely used in the fire recognition process in occasions such as forests, industrial and mining areas, and transportation vehicles. In recent years, the fire image recognition algorithms based on machine learning have developed rapidly, and scholars at home and abroad have conducted numerous studies on fire image recognition from multiple aspects such as fire image preprocessing, feature engineering, network structure, and recognition methods. For the fire image recognition features, currently, they are mainly manifested in two major aspects of the dynamic and static features of flames and smoke, specifically including color features, texture features, geometric shape features, gray-level co-occurrence matrix features, fire image edge features, etc. For the fire image recognition network structure, currently, there are mainly improved DenseNet architectures, SqueezeNet architectures, and CNN architectures. The above recognition features and network structures each have some deficiencies in terms of recognition accuracy, calculation time, memory, etc.

[0003] Although a large number of studies have been conducted on fire image recognition methods from four aspects: fire image preprocessing, fire image feature engineering, fire recognition network structure, and fire image recognition algorithms, in specific implementations, most of them use methods such as feature engineering and traditional recognition algorithms, and relatively few studies on recognition methods based on machine learning. Currently, mainly from improving the existing network structure Dense Net structure, the Squeeze Net architecture based on fire detection, positioning, and semantic understanding of the fire scene, and the cost-effective CNN architecture. The above network frameworks have not achieved ideal effects in terms of calculation time, memory, recognition accuracy, etc. Summary of the Invention

[0004] To solve the problems existing in the prior art, the present invention provides a fire image recognition method based on the mutual information of color features, selects color features with strong robustness as the input to obtain strong semantic information and high recognition efficiency; in terms of the recognition network structure, selects a single-hidden-layer fully connected network to reduce the time cost in the model training process, and solves the problems mentioned in the above background art.

[0005] To achieve the above object, the present invention provides the following technical solution: A fire image recognition method based on the mutual information of color features, comprising the following steps:

[0006] Step 1: Collect fire image and non-fire image samples to form a sample database, and randomly divide it into a training set and a test set according to a ratio of 3:1;

[0007] Step 2: Extract the color cast factors and variances var of fire and non-fire images as color feature data based on the three color models of Lab, RGB, and HSV, denoted as: Ka, Kb1, Var1; Kr, Kg, Kb2, Var2; Kh, Ks, Kv, Var3, a total of 11 color features, to form a fire image color feature library;

[0008] Step 3: Based on the probability density estimation of the kernel method, detect the density curves of the two types of data under different color features to obtain the density curve graphs of the two types of data under different color features;

[0009] Step 4: According to the density curve graphs, characterize the contribution degrees of different color features to the classification of conventional images and fire image data, and rank the contribution degrees of color feature recognition based on mutual information;

[0010] Step 5: Train based on a 160-layer single-hidden-layer fully connected ANN network model, and optimize the training model parameters at the same time to obtain the final fire image recognition model;

[0011] Step 6: Input the test set data into the final fire image recognition model for fire image recognition.

[0012] Furthermore, the calculation method of the color cast factor and its variance var is as follows:

[0013]

[0014]

[0015]

[0016]

[0017]

[0018]

[0019]

[0020]

[0021] Among them, dr, dg, and db are the average values of the respective component information of the RGB image, M and N are the pixel dimensions of the image, mr, mg, and mb are the average color casts of the respective component information of the RGB image, kr, kg, and kb are the color cast factors of the three components, and var is the variance of the color cast factor.

[0022] Further, the probability density estimation based on the kernel method is specifically as follows: Let the fire and non-fire image samples X be a "hypercube" centered at a point x in a D-dimensional space, and define the kernel function as:

[0023]

[0024] is used to represent whether a sample z of a fire image falls into this cube, x is the center point, i represents the i-th sample, where H is the width of the kernel function;

[0025] Given N samples of fire images The number K of samples falling into the region R is

[0026]

[0027] n represents the number of samples, then the density estimation of point x is

[0028]

[0029] where HD represents the volume of the hypercube R.

[0030] Further, the specific way of characterizing the contribution degree of different color features to the classification of conventional images and fire image data according to the density curve graph is as follows: The greater the difference between the density curves of the two types of data in the density curve graph of the same feature, the greater the contribution degree of this feature to the classification of conventional images and non-conventional images; conversely, the smaller the difference between the density curves of the two types of data in the density curve graph of the same feature, the smaller the contribution degree of this feature to the classification of conventional images and non-conventional images.

[0031] Further, in the calculation of mutual information, let a pair of random variables X and Y, and the information entropy of the random variable X is defined as:

[0032]

[0033] The joint entropy refers to the uncertainty of a pair of random variables X and Y, and the joint entropy of two random variables X and Y is defined as:

[0034]

[0035] x, y are specific samples, and the mutual information of two random variables X and Y is defined as:

[0036]

[0037] According to the definition of mutual information, extract more effective color feature information to provide a direct classification basis for fire image recognition. The mutual information of color features is defined as:

[0038]

[0039] Among them, p(x mn , x pq ) represents the joint probability distribution of X mn , X pq , and p1(x mn ) and p2(x pq ) are the marginal probability distributions of X and Y respectively.

[0040] Furthermore, in Step 5, it specifically includes:

[0041] Step 1: Train a 160-layer single-hidden-layer fully-connected ANN model based on 11 features in the color feature library, optimize the training model parameters at the same time, and use the test accuracy of the training model as the threshold T;

[0042] Step 2: For any integer k, where k ∈ [1, 10], gradually eliminate the k features with lower contribution degrees in the feature library, and calculate the recognition test accuracy T (11-k) ;

[0043] Step 3: Extract the maximum recognition test accuracy max(T (11-k) ) of the (11 - k) features, and compare it with the threshold T to determine the final fire image recognition model.

[0044] Furthermore, the comparison between max(T (11-k) ) and the threshold T specifically means:

[0045] If max(T (11-k) ) < T, then the 11-feature single-hidden-layer fully-connected neural network training model is the final fire image recognition model;

[0046] If max(T (11-k) ) ≥ T, then the (11 - k)-feature single-hidden-layer fully-connected neural network training model is the final fire image recognition model.

[0047] Furthermore, in the model training, the Adam optimizer is adopted, the ReLU activation function is adopted, the cross-entropy loss function is used as the loss function, and the maximum number of iterations is 2000.

[0048] The beneficial effects of the present invention are:

[0049] 1) Compared with the traditional fire image recognition color features, the present invention selects the color deviation factors and variances under three color models of Lab, RGB, and HSV as color features. At the same time, in order to better reduce the computational cost brought by multiple inputs, the present invention also conducts feature selection based on mutual information, which well reduces the redundancy of the input data;

[0050] 2) The present invention collects and constructs a real fire image recognition dataset. The data sources mainly come from real scene shootings (80%) and forest fire image libraries (20%). Among them, for real scene shootings, the influences of weather and lighting factors are mainly considered. Shootings are carried out based on red backgrounds, green backgrounds, and blue backgrounds under natural light on sunny days, natural light on cloudy days, and no light conditions in a dark box, fully reflecting the diversity of data sources and the balance of sample data;

[0051] 3) The present invention selects a 160-layer single-hidden-layer fully connected network for the training network structure, fully considering the great advantages of the single-hidden-layer fully connected network in recognition effects, especially in terms of computational time cost and operation memory. The recognition accuracy is as high as 93.82%, the Pearson correlation coefficient is as high as 0.8747, and the training cost is only 256 seconds. In addition, through comparative analysis, the optimal hyperparameters for fire image recognition are found. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 It is a density curve graph of two types of data under different color features;

[0053] Figure 2 It is a distribution graph of the contribution degree of color feature fire image recognition based on mutual information;

[0054] Figure 3 It is a schematic flow chart of the method steps of the present invention;

[0055] Figure 4 It is a schematic diagram of the structure of a single-hidden-layer fully connected neural network;

[0056] Figure 5 It is a schematic diagram of the performance comparison between the present invention and other methods;

[0057] Figure 6 It is a distribution graph of the recognition accuracy of fire images with different feature combinations;

[0058] Figure 7 It is a comparative analysis graph of the recognition accuracy of different feature combinations under different activation functions;

[0059] Figure 8 It is a comparison result graph of the effectiveness of the size of the single-hidden layer on fire image recognition;

[0060] Figure 9 It is a comparison result graph of the effectiveness of different optimization functions;

[0061] Figure 10 It is a visual difference graph of the recognition accuracy of feature combinations and hidden layer sizes under the ReLU activation function. DETAILED DESCRIPTION OF THE INVENTION

[0062] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0063] Regarding the problems such as complex calculation amount, high calculation cost, and unsatisfactory recognition effect in the process of fire image recognition, please refer to Figures 1 - 10 , the present invention proposes a fire image recognition method based on the mutual information of color features. First, a fire image and non-fire image library are constructed based on real-scene shooting and a forest fire image library. At the same time, the color deviation factors and variances of the images in three color models of Lab, RGB, and HSV are extracted as color feature information. Secondly, the color feature combination is optimized based on the mutual information, and the optimized color combination features are used as input data. Finally, an identification model is trained based on a 160-layer single-hidden-layer fully connected network, and the parameters of the training model are optimized at the same time. The method flow is specifically as Figure 3 shown, including the following steps:

[0064] Including the following steps:

[0065] Step 1: Collect fire image and non-fire image samples to form a sample database, and randomly divide it into a training set and a test set according to a ratio of 3:1.

[0066] A self-built database is used for the test experiment. The images in the self-built database are from the simulated real fire scenes photographed by a Canon EOS80D camera and the online forest fire image library. Table 1 shows the specific source information table of the image data.

[0067] Table 1 Specific Source Information Table of Image Data

[0068]

[0069] The total number of images in the image library is 7,775, including 3,777 fire images and 3,998 non-fire images. They are randomly divided into two parts in a ratio of 3:1, one as the training set and the other as the test set. For fire images, fire scenes are mainly photographed under three lighting conditions: sunny natural light, cloudy natural light, and dark box without light, based on red background, green background, and blue background, with 400 images for each. The forest fire image library mainly consists of 177 randomly selected images from 400 images on the Internet, totaling 3,777 fire scene images. For non-fire images, conventional scenes are mainly photographed under three lighting conditions: sunny natural light, cloudy natural light, and dark box without light, based on red background, green background, and blue background, with 400 images for each. Then, 398 non-fire scenes such as sunrise, sunset, campus exterior, and interior of teaching buildings are randomly photographed, totaling 3,998 non-fire images.

[0070] Step 2: Extract the color deviation factors and variances var of fire and non-fire images as color feature data based on three color models: Lab, RGB, and HSV, denoted as: Ka, Kb1, Var1; Kr, Kg, Kb2, Var2; Kh, Ks, Kv, Var3, a total of 11 color features, which form the fire image color feature library. The variables and their meanings are described in Table 2.

[0071] Table 2 Variables and Their Meanings

[0072]

[0073] Based on three common color models (Lab, RGB, HSV), the color deviation factors and the variance var between the color deviation factors are used to characterize the image color features in conventional scenes and fire scenes. The calculation process of the color deviation factors and their variance var is as follows (taking the RGB color model as an example):

[0074]

[0075]

[0076]

[0077]

[0078]

[0079]

[0080]

[0081]

[0082] where \(d_r\), \(d_g\), and \(d_b\) are the average values of the respective components of the RGB image, \(M\) and \(N\) are the pixel dimensions of the image, \(m_r\), \(m_g\), and \(m_b\) are the average color deviation values of the respective components of the RGB image, \(k_r\), \(k_g\), and \(k_b\) are all color deviation factors for the three components, and \(var\) is the variance of the color deviation factor.

[0083] Step 3: Based on kernel method-based probability density estimation, detect the density curves of the two types of data under different color features to obtain the density curve graphs of the two types of data under different color features.

[0084] Kernel Density Estimation is an improvement of the histogram method. Specifically, for kernel method-based probability density estimation: Let the fire and non-fire image samples \(X\) be a "hypercube" centered at a point \(x\) in a \(D\)-dimensional space, and define the kernel function as:

[0085]

[0086] to represent whether a sample \(z\) of the fire image falls into this cube, \(x\) is the center point, \(i\) represents the \(i\)-th sample, where \(H\) is the width of the kernel function;

[0087] Given \(N\) samples of fire images The number \(K\) of samples falling into the region \(R\) is

[0088]

[0089] \(n\) represents the number of samples, then the density estimate of point \(x\) is:

[0090]

[0091] where \(H\) D represents the volume of the hypercube \(R\).

[0092] For some data features, the distributions for different categories of data vary greatly, and it is considered that these data are helpful for classifying fire and non-fire images; while for some data, the distributions are very similar for different categories of data, then it can be considered that these feature data are not helpful for classifying fire and non-fire images. Based on this principle, detect the density curves of the two types of data under different features to preliminarily screen the color features.

[0093] Analysis of the density curves of fire color feature data

[0094] Such as Figure 1The figure shows the density curves of two types of data under different color features. The density curves characterize the contribution degrees of different features to the classification of conventional images and fire images. The greater the difference between the density curves of the two types of data in the density curve graph of the same feature, the greater the contribution degree of this feature to the classification of conventional images and unconventional images; conversely, the smaller the difference between the density curves of the two types of data in the density curve graph of the same feature, the smaller the contribution degree of this feature to the classification of conventional images and unconventional images. From Figure 1 it can be seen that: (1) Compared with other color features, the density distribution curves of the X 34 feature for the two types of data are basically the same in the low-order and high-order data ranges. Therefore, the X 34 feature has a small contribution to the classification of fire images and non-fire images; (2) Compared with other color features, the density distribution curves of the X 24 feature for the two types of data basically overlap in the low-order and medium-high-order data ranges. Therefore, the X 24 feature has a small contribution to the classification of fire images and non-fire images; (3) Compared with other color features, the density distribution curves of the X 22 feature for the two types of data have large differences in the distribution trends throughout the range. Therefore, the X 22 feature has a high contribution to the classification of fire images and non-fire images; (4) Compared with other color features, the density distribution curves of the X 32 feature for the two types of data have large differences in the distribution trends throughout the range. Therefore, the X 32 feature has a high contribution to the classification of both fire images and non-fire images.

[0095] Step 4: Based on the density curve graph, the contribution degrees of different color features to the classification of conventional images and fire images are characterized, and the contribution degrees of color feature recognition are sorted based on mutual information.

[0096] Specifically, based on the density curve graph, the contribution degrees of different color features to the classification of conventional images and fire images are as follows: The greater the difference between the density curves of the two types of data in the density curve graph of the same feature, the greater the contribution degree of this feature to the classification of conventional images and unconventional images; conversely, the smaller the difference between the density curves of the two types of data in the density curve graph of the same feature, the smaller the contribution degree of this feature to the classification of conventional images and unconventional images.

[0097] The sorting of the contribution degrees of color features for fire image recognition based on mutual information can effectively quantify the contribution degrees of each color feature to the fire type. As Figure 2 shown in the distribution graph of the contribution degrees of color features for fire image recognition based on mutual information.

[0098] From Figure 2It can be seen that: (1) In the Lab color model, the color feature of the b channel (X 12 ) contributes significantly more to the classification of fire images and non-fire images than the color feature of the a channel (X 11 ); (2) In the RGB color model, the color feature of the green channel (X 22 ) contributes the most to the classification of fire images and non-fire images. The color feature of the blue channel (X 23 ) contributes the second most to the classification of fire images and non-fire images, and the color feature of the red channel (X 21 ) contributes the least to the classification of fire images and non-fire images; (3) In the HSV color model, the color feature of the saturation channel (X 32 ) contributes the best to the classification of fire images and non-fire images. The color feature of the hue channel (X 31 ) contributes the second most to the classification of fire images and non-fire images, and the color feature of the value channel (X 33 ) contributes the least to the classification of fire images and non-fire images; (4) From the variance values of the color feature distributions of each channel in the three color models, the variance values of the color features of the red (R), green (G), and blue (B) channels (X 24 ) in the RGB color model contribute the most to the classification of fire images and non-fire images. The variance values of the color features of the hue (H), saturation (S), and value (V) channels (X 34 ) in the HSV color model contribute the second most to the classification of fire images and non-fire images, and the variance values of the color features of the a and b channels (X 13 ) in the Lab color model contribute the least to the classification of fire images and non-fire images; (5) Overall, the RGB color model contributes the most to the classification of fire images and non-fire images, the HSV color model contributes the second most to the classification of fire images and non-fire images, and the Lab color model contributes the least to the classification of fire images and non-fire images.

[0099] Calculation of mutual information

[0100] The importance of the color features of fire images not only reflects their own characteristics but also reflects the status and influence of the features on the classification of fire image categories, and this status and influence are generally reflected in the contribution degrees among the features. In information theory, entropy is the degree of uncertainty of a random variable or the average value of self-information. Let any feature of fire and non-fire images be a random variable X i , and the entropy of a random variable X is a measure of the information contained in the random variable X. Let the random variable X i = [x i1 , x i2 , …, x is, p(xij) (where both i and j are positive integers) is the probability distribution of the random variable X.

[0101] In the calculation of mutual information, given a pair of random variables X and Y, the information entropy of the random variable X is defined as:

[0102]

[0103] The joint entropy refers to the uncertainty of a pair of random variables X and Y. The joint entropy of two random variables X and Y is defined as:

[0104]

[0105] x and y are specific samples.

[0106] Mutual Information is another useful information measure, which refers to the correlation between two random variables. The mutual information between two random variables X and Y is defined as:

[0107]

[0108] Among them, p(x,y) represents the joint probability distribution of X and Y, and p1(x) and p2(y) are the marginal probability distributions of X and Y respectively. At the same time, I(X,Y) satisfies the following properties:

[0109] Property 1: I(X,Y) = H(X) + H(Y) - H(X,Y);

[0110] Property 2: I(X,Y) = H(X);

[0111] Property 3: I(X,X) = H(X);

[0112] Property 4: I(X,Y) ≥ 0;

[0113] Property 5: If X and Y are independent of each other, I(X,Y) = 0.

[0114] Calculation of the mutual information of color features

[0115] Let X ij be the color feature variable, where i ∈ {1, 2, 3}, i = 1 represents the Lab color model, i = 2 represents the RGB color model, and i = 3 represents the HSV color model. In the Lab color model, X 11 represents the color feature of the a channel, and X 12 represents the color feature of the b channel; in the RGB color model, X 21 represents the color feature of the R channel, X 22 represents the color feature of the G channel, and X 23Represents the color feature of channel B; in the HSV color model, X 31 Represents the color feature of the H channel, X 32 Represents the color feature of the S channel, X 33 Represents the color feature of the V channel. For specific variable information of the color feature, see Table 3.

[0116] Table 3 Color Feature Variables and Their Meanings

[0117]

[0118]

[0119] According to the definition of mutual information, study the correlation magnitude between color features, extract more effective color feature information, and provide a direct classification basis for fire image recognition. The definition of the mutual information of color features is:

[0120]

[0121] where p(x mn ,x pq ) represents the joint distribution column of X mn ,X pq , and p1(x mn ), p2(x pq ) are the marginal distribution columns of X and Y respectively.

[0122] Step Five: Train based on the 160-layer single-hidden-layer fully connected ANN network model, and at the same time optimize the training model parameters to obtain the final fire image recognition model.

[0123] Step 1: Train the 160-layer single-hidden-layer fully connected ANN model based on 11 features in the color feature library, optimize the training model parameters at the same time, and use the test accuracy of its training model as the threshold T;

[0124] Step 2: For any integer k, and k ∈ [1, 10], gradually eliminate the k features with lower contribution degrees in the feature library, and calculate the recognition test accuracy T (11-k) ;

[0125] Step 3: Extract the maximum recognition test accuracy max(T (11-k) ) of the (11 - k) features, and compare it with the threshold T to determine the final fire image recognition model.

[0126] The comparison between max(T (11-k) ) and the threshold T specifically means:

[0127] If max(T (11-k)) If <T, then the 11-feature single-hidden-layer fully-connected neural network training model is the final fire image recognition model;

[0128] If max(T (11-k) ) ≥ T, then the (11 - k)-feature single-hidden-layer fully-connected neural network training model is the final fire image recognition model.

[0129] During model training, the Adam optimizer is adopted, the ReLU activation function is used, the cross-entropy loss function is used as the loss function, and the maximum number of iterations is 2000.

[0130] The fire image recognition algorithm considering the mutual information of color features is based on a single-hidden-layer fully-connected network for model training. As Figure 4 shown in the structure diagram of the single-hidden-layer fully-connected neural network, it mainly consists of 11 input layers, 10 hidden layers, and 2 output layers. The input layer contains the channel color feature information and the variance value features of the color feature values in three color models, namely Lab, RGB, and HSV. The output layer contains two layers: fire images and non-fire images.

[0131] Activation function

[0132] The Sigmoid-type function refers to a class of S-shaped curve functions saturated at both ends. Commonly used Sigmoid-type functions include the Logistic function and the Tanh function. The expression of the Logistic function is shown below. Its characteristic is to "squeeze" an input in the real number domain to (0, 1). When the input value is near 0, the Sigmoid function is approximately a linear function. When the input value approaches both ends, the input is suppressed. The smaller the input, the closer it is to 0, and the larger the input, the closer it is to 1. The Tanh function is also a modified Sigmoid function. Its specific expression is as follows and can be regarded as an amplified and translated Logistic function. Its value range is (-1, 1).

[0133]

[0134]

[0135] The ReLU (Rectified Linear Unit) function is an activation function often used in current deep neural networks. It is actually a ramp function, and its specific expression is as follows:

[0136]

[0137] The advantages of the ReLU function are as follows: The neurons sampling ReLU only need to perform addition, multiplication, and comparison operations, which is more computationally efficient; in terms of optimization, compared with the two-end saturation of the Sigmoid function, ReLU is a left-saturated function, and its derivative is 1 when x > 1, which alleviates the gradient disappearance problem of the neural network to a certain extent and accelerates the convergence speed of gradient descent.

[0138] The Identity function is the simplest activation function, and its output is proportional to the input. The problem is that its derivative is a constant, and the gradient is also a constant, so gradient descent cannot work, and f(x) = x.

[0139] Loss function

[0140] The loss function is a non-negative real-valued function used to quantify the difference between the model prediction and the true label. The better the loss function, the better the performance of the model in general. Different models usually use different loss functions. Commonly used loss functions include: 0-1 loss function, squared loss function, cross-entropy loss function, Hinge loss function. Table 4 shows the comparison table of the expressions, advantages and disadvantages of common loss functions.

[0141] Table 4 Expressions, advantages and disadvantages of common loss functions

[0142]

[0143] Since the cross-entropy loss function has the conditional probability distribution of the class label as the output and can be well used for classification problems, the cross-entropy loss function is selected in the model training process of the present invention.

[0144] Optimizer

[0145] The parameter learning of deep neural networks mainly uses the gradient descent method to find a set of parameters that can minimize the structural risk. In specific implementation, gradient descent can be divided into three forms: batch gradient descent, stochastic gradient descent, and mini-batch gradient descent. According to different data volumes and parameter numbers, a specific implementation form can be selected. The optimization algorithm is mainly optimized from two aspects: (1) adjusting the learning rate to make the optimization more stable; (2) correcting the gradient estimation to optimize the training speed. Here, the mini-batch gradient descent (MBGD) algorithm, stochastic gradient descent (SGD) algorithm, AdaGrad algorithm, RMSprop algorithm, Adam algorithm, and L-BFGS algorithm are mainly introduced.

[0146] MBGD Algorithm: When training a neural network, the scale of the training data is relatively large. If the velocity on the entire training data needs to be calculated for each iteration during gradient descent, a relatively large amount of computing resources is required. Additionally, the data in large-scale training sets is usually highly redundant, and it is not necessary to calculate the gradient on the entire training set. Therefore, the mini-batch gradient descent method (Mini-Batch Gradient Descent) is often used to train deep neural networks.

[0147] SGD Optimizer: Stochastic Gradient Descent (SGD for short) is an implementation of the gradient descent algorithm. To obtain the minimum value of the loss function, the gradient of the loss function needs to be obtained first. The goal is to obtain the minimized loss function, that is, to update in the direction where the gradient is negative.

[0148] AdaGrad Optimizer: The AdaGrad algorithm (Adaptive Gradient Algorithm) adaptively adjusts the learning rate of each parameter during each iteration. In the AdaGrad algorithm, if the accumulated partial derivative of a certain parameter is relatively large, its learning rate is relatively small; conversely, if its accumulated partial derivative is relatively small, its learning rate is relatively large. The disadvantage of the AdaGrad algorithm is that when the optimal point has not been found after a certain number of iterations, since the learning rate is already very small at this time, it is difficult to continue to find the optimal point.

[0149] RMSprop Optimizer: The RMSprop algorithm is an algorithm with an adaptive learning rate, which can avoid the disadvantage in the AdaGrad algorithm that the learning rate continuously and monotonically decreases and decays prematurely in some cases. The difference between the RMSprop algorithm and the AdaGrad algorithm lies in that the calculation of \(G_t\) changes from an accumulative method to an exponentially decaying moving average. During the iteration process, the learning rate of each parameter does not show a decaying trend and can either become smaller or larger.

[0150] Adam Optimizer: The Adam (Adaptive momentum) algorithm is an adaptive momentum stochastic optimization algorithm, which can be regarded as a combination of the momentum method and the RMSprop algorithm. It not only uses momentum as the direction of parameter update but also can adaptively adjust the learning rate. The Adam algorithm calculates the exponentially weighted average of the squared gradient on the one hand and the exponentially weighted average of the gradient \(g\) on the other hand t The calculation process is shown in the following formula:

[0151] M t =\(\beta_1M\) t-1 +(1 - \(\beta_1\))g t

[0152] G t = β2G t-1 +(1 - β2)g t ⊙g t

[0153] Among them, β1 and β2 are the decay rates of two moving averages respectively, and usually take values of β1 = 0.9 and β2 = 0.99.

[0154] Assume M0 = 0, G0 = 0, then at the initial stage of iteration, the values of M t and G t will be smaller than the true mean and variance. Especially when both β1 and β2 are close to 1, the deviation will be large. Therefore, it is necessary to correct the deviation.

[0155]

[0156]

[0157] The parameter update difference of the Adam algorithm is

[0158]

[0159] Among them, the learning rate α is usually set to 0.001 and can be decayed, such as

[0160] L - BFGS algorithm: L - BFGS is a new inverse Newton algorithm proposed by relevant scholars using the DFP algorithm to construct an approximate matrix, abbreviated as the BFGS method. However, this method still needs to store the matrix, which increases the inversion time and reduces the efficiency. To address this issue, relevant scholars proposed the L - BFGS method on this basis. The advantage of this algorithm is to construct an approximate Hessian matrix and only utilize the curvature information of the most recent iteration.

[0161] Step 6: Input the test set data into the final fire image recognition model for fire image recognition.

[0162] Experiment and Analysis

[0163] The dataset of this embodiment is 7775 fire and non - fire images from a self - built database. During the experiment, the positive - negative sample ratio of the dataset is 3777:3998. The present invention conducted a series of experiments to verify the effectiveness of the fire image classification method proposed by the present invention from different perspectives. The statistical situation of the dataset is shown in Table 1. All experiments were carried out on a server configured with a 2.4 GHZ Iter CORE i5 processor and 1 NVIDIA IATian XP graphics card.

[0164] In the experiment, 70% and 10% of the samples were randomly selected from the fire images and non-fire images respectively for training and validation, and the remaining 20% of the images were used as test samples. The network model was trained and tested based on the deep learning framework TensorFlow. During training, the Adam optimizer was used, the ReLU activation function was adopted, the cross-entropy loss function was used as the loss function, and the maximum number of iterations was 2000.

[0165] Evaluation metrics: Both Accuracy and AUC are commonly used performance metrics for evaluating model classification. For the fire image classification task, the present invention follows the practice of previous work and selects Accuracy and Spearman to evaluate the quality of the model. The present invention first measures the accuracy (Acc) of the algorithm's judgment on the recognition effect of fire images. Given the test image set S, the accuracy metric is defined as follows:

[0166]

[0167] For the image to be inspected, y i , respectively represent the true type of the image to be inspected, and I is an indicator function. If the given Boolean expression is true, its value is 1; otherwise, its value is 0.

[0168] At the same time, the Spearman Rank Correlation Coefficient is used to measure the correlation between the classification results of the algorithm of the present invention and the true results. The higher the Spearman rank correlation coefficient, the better the classification performance of the algorithm for the test images. The Spearman rank correlation coefficient p is defined as follows:

[0169]

[0170] where r(i) and are the true classification and predicted classification of the image to be inspected respectively.

[0171] Comparative experiment

[0172] In the experiment, the method of the 160 single-hidden-layer fully connected neural network based on mutual information proposed by the present invention was compared with traditional machine learning methods, including: rough, medium, fine, optimized decision trees, linear, quadratic, cubic, rough Gaussian, medium Gaussian, fine Gaussian support vector machines, narrow, medium, wide, double-layer, and triple-layer neural networks. The performance comparison between different methods is as Figure 5 shown, from Figure 5It can be seen that, compared with the other 15 methods, the ANN algorithm for fire image recognition based on the mutual information of color features proposed in the present invention has achieved the best results in the fire image prediction task. Specifically, the method of the present invention has reached 93.82% and 0.8747 in terms of accuracy and Spearman rank correlation coefficient respectively, far exceeding the sub-optimal method, the wide neural network, and obtaining relative improvements of 0.99% and 1.51% respectively in the two indicators of accuracy and Spearman correlation rank.

[0173] Figure 6 Shown is the distribution map of the recognition accuracy of fire images with different feature combinations. From Figure 6 it can be seen that the 9-feature combination proposed in the present invention has the highest prediction accuracy and Pearson correlation coefficient among all combined features, which are 93.82% and 0.8747 respectively. At the same time, the training time is 256 seconds and the training cost is acceptable. This is sufficient to show that the 9-feature combination based on mutual information in the method of the present invention is significantly superior to other feature combinations in both the prediction accuracy and Spearman coefficient of fire image recognition, and the training cost is also small, significantly effectively improving the prediction ability of the model. In addition, there are only 11 cases of the feature combination based on mutual information. If randomly combined based on 11 features, there are a total of

[0174]

[0175] cases. The cost of the feature combination based on mutual information in the process is much lower than that of the random feature combination. In terms of computational cost, the method of the feature combination based on mutual information of the present invention is only 0.54% of the computational cost of the random feature combination.

[0176] Effectiveness analysis of different activation functions

[0177] The comparison of the recognition accuracy of color feature combinations based on mutual information can effectively find the best combination features, overcoming the problems of high computational complexity and blind operability brought by the traditional method based on multi-feature permutation and combination. As Figure 7 shown is the comparative analysis diagram of the recognition accuracy of different feature combinations under different activation functions. It can be seen from the figure that different activations have relative stability for fire image recognition under different hidden layer sizes. As the hidden layer size increases from 10 layers to 10240 layers, the recognition performances of the ReLU and Tanh activation functions become more and more stable, and their recognition accuracies can reach 95.22% and 93.93% respectively. The recognition accuracy of the ReLU activation function is higher than that of the Tanh activation function; the recognition accuracy of the Logistic activation function first increases and then decreases as the hidden layer size increases, and its stability is relatively poor; the recognition performance of the Identity activation function also becomes more and more stable as the hidden layer size increases, but its accuracy is significantly lower than that of the above three types of activation functions.

[0178] Effectiveness Analysis of the Size of a Single Hidden Layer

[0179] The size of a single hidden layer has a significant impact on the effectiveness of fire image recognition. Figure 8 As shown in the figure comparing the effectiveness of the size of a single hidden layer in fire image recognition, it can be seen from the figure that: (1) As the size of the single hidden layer increases, the test accuracy of fire images first increases and then decreases. When the size of the single hidden layer increases to 160 to 320, the highest recognition accuracy of fire images can reach 93.82%. (2) When the number of single hidden layers is 160, the Pearson correlation coefficient between the predicted results and the true results of fire images is the highest, reaching 0.8747, which is highly consistent with the highest fire recognition accuracy. (3) In terms of training cost, as the number of single hidden layers increases from 10 to 20489, the training time cost shows an exponential increase trend; at the same time, when the number of single hidden layers is 160, the training time is 256 seconds, and the relative time for the method with 20480 hidden layer numbers is 0.0067. Therefore, when the size of the single hidden layer is 160, the recognition accuracy of fire image recognition, the Pearson correlation coefficient between the predicted results and the true results is the highest, and the training time cost is the smallest.

[0180] Effectiveness Analysis of Different Optimization Functions

[0181] The parameter learning of deep neural networks mainly uses the gradient descent method to find a set of parameters that can minimize the structural risk. The gradient descent of the optimization function can be divided into three forms: batch gradient descent, stochastic gradient descent, and mini-batch gradient descent. Figure 9 Shown are the results of the effectiveness analysis of three common different optimization functions.

[0182] From Figure 9 It can be seen that in terms of prediction accuracy, the adam, lbfgs, and sgd optimization functions reach 93.82%, 94.44%, and 86.84% respectively; however, the lbfgs optimizer shows overfitting as the number of iterations increases, and the adam optimizer has good stability during training and does not show overfitting. In terms of the Spearman correlation coefficient, the adam, sgd, and lbfgs optimization functions reach 0.8844, 0.7587, and 0.8628 respectively; in terms of training time cost, the adam, sgd, and lbfgs optimization functions are 1623.15 seconds, 1397.13 seconds, and 1328.23 seconds respectively. To sum up, in terms of training cost, both the lbfgs and sgd optimizers are significantly lower than the adam optimizer. The lbfgs optimizer is 13.92% lower than the adam optimizer, and the sgd optimizer is 18.17% lower than the adam optimizer; however, in terms of prediction accuracy stability and the Spearman correlation coefficient, the adam optimizer is significantly better than the lbfgs and sgd optimizers.

[0183] Verification Analysis of Feature Combination and Hidden Layer Size under the ReLU Activation Function

[0184] like Figure 10 The following is a visualization of the difference in recognition accuracy between different feature combinations and hidden layer sizes under the ReLU activation function. Figure 10 It can be seen that for the best-performing ReLU activation function, the 9-feature combination has the highest overall recognition accuracy and good stability; the 10-feature combination has a higher overall accuracy, but its stability is slightly worse than that of the 9-feature combination; 11 features, 8 features, and 7 features are obviously inferior to 9 features in terms of overall accuracy and stability. Therefore, the 9-feature combination under the ReLU activation function is the best feature combination. For the best-performing ReLU activation function, as the hidden layer size increases, its prediction accuracy performance also increases, but due to Figure 5 The effectiveness comparison of the size of a single hidden layer shows that as the size of the hidden layer increases, its training time cost also increases exponentially. Considering the two perspectives of training time cost and prediction accuracy, the training cost of a network with a single hidden layer size of 160 is 256 seconds, and its accuracy is 93.82; compared with a network with a single hidden layer size of 640, its accuracy is increased by 0.7% and the training cost is reduced by 85%. Therefore, the network with a single hidden layer size of 160 under the ReLU activation function is the best prediction network.

[0185] Summarize

[0186] The present invention proposes to study the ANN algorithm for fire image recognition based on the color information, and constructs effective color information features by analyzing and mining the color information of the image, thereby identifying the type of fire image. The present invention proposes an image feature input based on the mutual information of color features, introduces a single hidden layer fully connected network with a size of 160, uses the ReLU activation function, and uses Adam as the optimization function for model training. The experimental results on the real data set prove that the 160 single hidden layer fully connected ANN algorithm based on the mutual information of color features proposed by the present invention has high effectiveness in the problem of fire image recognition.

[0187] The results show that: (1) There are significant differences in the contribution degrees of various color deviation factors and variance color features to the recognition of fire images under the three color models of Lab, RGB, and HSV. Among them, the 9-feature combination is the best input for the recognition of fire images; (2) Considering comprehensively from the aspects of recognition accuracy, Pearson correlation coefficient, and training cost, the 160 single-hidden-layer fully connected network is the best network for the recognition of fire images, and it has obvious advantages compared with single-hidden-layer fully connected networks such as 10, 20, 40, 80, 320, 640, 1280, 2560, 5120, 10240, and 20480. Its recognition accuracy is as high as 93.82%, the Pearson correlation coefficient is as high as 0.8747, and the training cost is only 256 seconds; (3) Through comparative analysis with algorithms such as decision tree class, support vector machine class, and different-depth neural network class algorithms, and comprehensively considering algorithm performance indicators such as recognition accuracy, Pearson correlation coefficient, and training cost, the comprehensive recognition performance of the method of the present invention is better than the above algorithms.

[0188] It is hoped that the present invention can stimulate the research interest of relevant scholars in fire image recognition algorithms. Further research can train more image data through multi-source heterogeneous features such as the texture, edge, and shape of images, providing more reliable technical support for safety managers. In future research, based on indicators such as recognition accuracy, Pearson correlation coefficient, and training cost as the evaluation basis for model performance, high-dimensional features of images are extracted by deeper networks, enabling the model to start learning from easy samples first, and then gradually progress to difficult samples and further verify the effectiveness of the method on a larger-scale dataset.

[0189] Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for fire image recognition based on the mutual information of color features, characterized in that, It includes the following steps: Step 1: Collect fire image and non-fire image samples to form a sample database, and randomly divide it into a training set and a test set according to a ratio of 3:1; Step 2: Based on three color models of Lab, RGB, and HSV, extract the color deviation factors and variances var of fire and non-fire images as color feature data, denoted as: Ka, Kb1, Var1; Kr, Kg, Kb2, Var2; Kh, Ks, Kv, Var3, a total of 11 color features, to form a fire image color feature library; among them, Ka represents the size of the color information of the a channel in the Lab color space, Kb1 represents the size of the color information of the b channel in the Lab color space, Var1 represents the variance value of the color information of the a and b channels in the Lab color space, Kr represents the size of the color information of the red channel in the RGB color space, Kg represents the size of the color information of the green channel in the RGB color space, Kb2 represents the size of the color information of the blue channel in the RGB color space, Var2 represents the variance value of the color information of the three channels in the RGB color space, Kh represents the size of the color information of the hue channel in the HSV color space, Ks represents the size of the color information of the saturation channel in the HSV color space, Kv represents the size of the color information of the value channel in the HSV color space, and Var3 represents the variance value of the color information of the three channels in the HSV color space; Step 3: Based on the probability density estimation of the kernel method, detect the density curves of the two types of data under different color features to obtain the density curve graphs of the two types of data under different color features; Step 4: According to the density curve graphs, characterize the contribution degrees of different color features to the classification of normal images and fire image data, and sort the contribution degrees of color feature recognition based on the mutual information; Step 5: Train based on a 160-layer single-hidden-layer fully connected ANN network model, and optimize the training model parameters at the same time to obtain the final fire image recognition model; specifically including: Step 1: Train a 160-layer single-hidden-layer fully connected ANN model based on 11 features of the color feature library, optimize the training model parameters at the same time, and use the test accuracy of the training model as the threshold T; Step 2. For any integer k, where k ∈ [1, 10], gradually eliminate the k features with lower contribution degrees in the feature library, and calculate the recognition test accuracy T of (11 - k) features respectively (11-k) ; Step 3: Extract the maximum accuracy max(T (11-k) ) of the recognition test for (11-k) features, and compare it with the threshold T to determine the final fire image recognition model; Step 6: Input the test set data into the final fire image recognition model for fire image recognition.

2. The method for fire image recognition based on the mutual information of color features according to claim 1, characterized in that: The calculation method of the color deviation factor and its variance var is: where dr, dg, and db are the average values of the respective component information of the RGB image, M and N are the pixel dimensions of the image, mr, mg, and mb are the color deviation averages of the respective component information of the RGB image, kr, kg, and kb are all color deviation factors of the three components, and var is the variance of the color deviation factor.

3. The method for fire image recognition based on the mutual information of color features according to claim 1, characterized in that: The probability density estimation based on the kernel method is specifically: Assume that the fire and non-fire image samples X are a "hypercube" centered at a point x in a D-dimensional space, and define the kernel function as: to represent whether a sample z of a fire image falls into this cube, where H is the width of the kernel function; Given a sample of N fire images The number of samples K that fall into region R is Then the density estimation of the point x is where H D represents the volume of the hypercube R.

4. The method for fire image recognition based on the mutual information of color features according to claim 1, characterized in that: The contribution degrees of different color features to the classification of conventional images and fire image data are characterized by the density curve graph as follows: the greater the difference between the density curves of the two types of data in the same feature density curve graph, the greater the contribution degree of this feature to the classification of conventional images and unconventional images; conversely, the smaller the difference between the density curves of the two types of data in the same feature density curve graph, the smaller the contribution degree of this feature to the classification of conventional images and unconventional images.

5. The method for fire image recognition based on the mutual information of color features according to claim 1, characterized in that: In the calculation of mutual information, let a pair of random variables X and Y, and the information entropy of random variable X is defined as: The joint entropy refers to the uncertainty of a pair of random variables X and Y, and the joint entropy of two random variables X and Y is defined as: The mutual information of two random variables X and Y is defined as: According to the definition of mutual information, more effective color feature information is extracted to provide a direct classification basis for fire image recognition. The mutual information of color features is defined as: Among them, p(x mn ,x pq ) represents the joint distribution list of X mn ,X pq , and p1(x mn ) and p2(x pq ) are the marginal distribution lists of X and Y respectively.

6. The fire image recognition method based on color feature mutual information according to claim 1, characterized in that: max(T (11-k) ) is compared with the threshold value T specifically means that: If max(T (11-k) ) < T, then the 11-feature single-hidden-layer fully-connected neural network training model is the final fire image recognition model; If max(T (11-k) ) ≥ T, then the (11 - k) feature single-hidden-layer fully-connected neural network training model is the final fire image recognition model.

7. The fire image recognition method based on color feature mutual information according to claim 1, characterized in that: In model training, the Adam optimizer is used, the ReLU activation function is used, the cross-entropy loss function is used as the loss function, and the maximum number of iterations is 2000.

Citation Information

Patent Citations

  • Straw burning fire detection method based on multi-feature fusion of image feature extraction

    CN107067007A

  • Fire image color cast quantification method based on three color modes

    CN112668426A