Lightweight Fire Detection Method Based on Double-Level Pruning and Adaptive Quantization
By using two-stage pruning and adaptive quantization technology in the fire detection model, the structure and parameters of the convolutional neural network are optimized, and the problems of large resource consumption and low operation efficiency of the existing model are solved, achieving more efficient and more accurate fire detection.
Patent Information
- Application Number
- CN202411190138.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-28
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-08-28
AI Technical Summary
Due to the large number of parameters and complex structures of the existing fire detection model, it has large resource consumption and low operating efficiency. The existing pruning and quantization methods cannot take into account high flexibility, high efficiency and high precision.
The lightweight fire detection method based on double-stage pruning and adaptive quantization is adopted to optimize the structure and parameters of the model by fusing branched convolution blocks.
It improves the flexibility and accuracy of the model, reduces the model scale and inference time, and solves the problems of high hardware costs and low operating efficiency.
Smart Images

Figure CN119169420B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of machine learning applications for fire detection, and particularly relates to a lightweight fire detection method based on double-level pruning and adaptive quantization. Background Art
[0002] In high-fire-risk areas such as new energy vehicles, photovoltaic power generation, and flammable material processing plants, fire detection devices are generally used to monitor fires. The traditional fire monitoring technical means are mainly contact fire detectors, such as carbon monoxide sensors, carbon dioxide sensors, and smoke sensors, etc. However, contact fire detectors are difficult to be applied to outdoor open spaces. Therefore, currently, pictures of the detection area are generally obtained in real time, and a convolutional neural network model is used to perform real-time detection and recognition on the video pictures to achieve the monitoring of high-fire-risk areas.
[0003] Existing fire detection convolutional neural network models can effectively detect and identify fires. However, due to the usually large number of parameters and complex structure of convolutional neural network models, they face problems such as high resource consumption and low operation efficiency in practical applications. Currently, pruning and quantization are mainly used to solve the above problems. Pruning reduces the complexity of the model by removing unimportant weights or neurons in the model, while quantization reduces the storage and calculation requirements of the model by reducing the number of bits representing weights and activation values. Each of the two methods has its unique advantages.
[0004] Currently, most pruning methods are channel-level pruning. Channel-level pruning has a low granularity and high flexibility, and can effectively reduce the number of model parameters, but it cannot reduce the number of model layers. And the number of layers has a great positive correlation with the inference time during hardware deployment. Generally speaking, the more layers there are, the more time is required. Hierarchical pruning methods can effectively reduce the number of layers and greatly reduce the inference time consumption, but they lack the flexibility of channel-level pruning. Existing pruning methods cannot take into account the characteristics of greatly reducing the inference time consumption and flexibility.
[0005] Quantization methods can effectively reduce the model size, but the accuracy of the model will decrease to a certain extent after quantization, resulting in poor model usage effects. An important reason for the decrease in model accuracy is the existence of outliers. As Figure 1 shown in the floating-point and quantization value correspondence situation of eight-bit symmetric quantization, where the floating-point maximum value deviates greatly from the overall floating-point value, that is, the quantization blank interval is too large, resulting in a more limited information representation ability of the quantization value, and finally causing a significant decrease in model accuracy. Values similar to the floating-point maximum value in the figure are usually defined as outliers. Existing quantization methods lead to the problem of low model accuracy.
[0006] Existing lightweight methods cannot balance the characteristics of high flexibility, high efficiency, and high precision, resulting in the inability to effectively solve the problems of high hardware costs and low operating efficiency of fire detection models. Summary of the Invention
[0007] The technical problem to be solved by the present invention is to provide a lightweight fire detection method based on double-level pruning and adaptive quantization in view of the deficiencies in the above-mentioned prior art. This method improves the flexibility of model lightweighting, enhances the accuracy of the lightweight model, reduces the scale of the lightweight model, and reduces the inference time consumption of the lightweight model.
[0008] To solve the above technical problems, the technical solution adopted by the present invention is:
[0009] A lightweight fire detection method based on double-level pruning and adaptive quantization, the method comprising the following steps:
[0010] Step 1: Construct a non-lightweight convolutional neural network model according to the fire detection task to be implemented. The convolutional neural network model includes an input layer, a convolutional layer, a pooling layer, and an output layer. The input layer receives known fire image data and transmits it to the convolutional layer. The convolutional layer outputs corresponding feature maps according to the input feature maps. The pooling layer downsamples the feature maps output by the convolutional layer and outputs corresponding feature maps. The output layer makes a final recognition result according to the feature maps output by the convolutional layer and the pooling layer. The convolutional neural network includes multiple convolutional layers. All convolutional layers in the convolutional neural network model adopt fusible branch convolutional blocks. The fusible branch convolutional block includes a main branch and a shortcut branch. The outputs of the main branch and the shortcut branch are fused as the final output of the fusible branch convolutional block.
[0011] In the main branch, the input feature map passes through a first convolutional layer and a batch normalization layer in sequence to output a feature map. The feature map of each channel performs a linear multiplication operation with the channel-level importance factor of the corresponding channel to obtain the feature map of each channel, and then performs a linear multiplication operation with the hierarchical importance factor to complete the output of the first feature map. In the shortcut branch, after the input feature map passes through the feature map matching operation, the feature map of each channel performs a linear multiplication operation with the channel-level importance factor of the corresponding channel, and then performs a linear multiplication operation with the adaptive information control factor to complete the output of the second feature map.
[0012] Step 2: Perform L1 sparsification training on the convolutional neural network model to make the hierarchical importance factor and the channel-level importance factor tend to zero. After training, unimportant layers and channels are trimmed. The fusible branch convolutional blocks that are not trimmed are fused into ordinary convolutional layers. The fused convolutional neural network model is called a pruning model. An unimportant layer refers to a layer whose hierarchical importance factor is less than 0.0001, and an unimportant channel refers to a channel whose channel-level importance factor is less than 0.0001.
[0013] Step 3: Quantize the pruned model, and remove the weight outliers and activation outliers for each layer of the network according to the weight outlier percentage parameter and the activation outlier percentage parameter; the quantized pruned model is called the dynamic quantization model;
[0014] Step 4: Take the minimum loss value of the feature maps output by the dynamic quantization model and the pruned model as the optimization goal, and continuously iterate the weight outlier percentage parameter and the activation outlier percentage parameter until the final weight outlier percentage parameter and activation outlier percentage parameter are obtained;
[0015] Step 5: Perform actual quantization on the pruned model according to the final weight outlier percentage parameter and activation outlier percentage parameter, and the actually quantized model is the lightweight fire detection model based on double-level pruning and adaptive quantization;
[0016] Step 6: Input the fire image to be detected into the lightweight fire detection model based on double-level pruning and adaptive quantization, and output the corresponding recognition result.
[0017] Furthermore, the feature map matching operation includes a no-operation, a 1×1 convolution operation, and an average pooling operation. According to the input feature map and the convolutional layer of the backbone branch, it is judged whether the number of channels and the size of the input feature map are consistent with those of the output first feature map. When the number of channels of the input feature map is inconsistent with that of the output first feature map, the 1×1 convolution operation is used for the feature map matching operation. When the size of the input feature map is inconsistent with that of the output first feature map, the average pooling operation is used for the feature map matching operation. When the number of channels and the size of the input feature map are both consistent with those of the output first feature map, the no-operation is used for the feature map matching operation.
[0018] Furthermore, the fusion branch convolution block in Step 2 has the characteristic of being fused into a normal convolution, and its expression is as follows:
[0019] Y = (X * W + b) + (X * W s + b s ) · h i · g i = X * W f + b f
[0020] where * is the convolution operation, · represents the linear multiplication, h i is the channel-level importance factor of the i-th layer, g i is the shortcut branch adaptive information control factor of the i-th layer, X and Y are the feature maps input and output by the corresponding fusion branch convolution block, W and b are the weights and biases of the convolution of the backbone branch convolution layer after fusing the batch normalization layer and the importance factor, W s and bs The weights and biases of the convolution transformed from the shortcut branch feature map matching operation, W f and b f are the weights and biases of the ordinary convolution after the fusion of the backbone branch and the shortcut branch.
[0021] Furthermore, when the feature map matching operation uses a 1×1 convolution operation, is equal to the bias value of the corresponding 1×1 convolution, The expression of
[0022]
[0023] When the feature map matching operation uses average pooling operation, is zero, The expression of
[0024]
[0025] When the feature map matching operation uses no operation, is zero, The expression of
[0026]
[0027] where j is the output channel, q is the input channel, v represents the convolution kernel size of the convolutional layer, and are the weights and biases of the j-th output channel and the q-th input channel of the convolution transformed from the shortcut branch feature map matching operation, is to pad 0 around to generate a convolution kernel with a size of v×v, is the weight of the j-th output channel and the q-th input channel of the 1×1 convolution on the shortcut branch, Copy(0) is to copy 0 to generate a convolution kernel with a size of v×v, is to copy to generate a convolution kernel with a size of v×v, Padding(1) is to pad 0 around the value 1 to generate a convolution kernel with a size of v×v.
[0028] Furthermore, the initial values of both the weight outlier percentage parameter and the activation value outlier percentage parameter are randomly generated by the computer.
[0029] Further, the removal process in step 3 is as follows: Obtain the absolute value of the weight to be quantized, take the product of the maximum value in the absolute value of the weight to be quantized and the weight outlier percentage parameter as the outlier threshold, and remove the weight to be quantized in the absolute value of the weight to be quantized that is greater than the outlier threshold; Obtain the absolute value of the activation to be quantized, take the product of the maximum value in the absolute value of the activation to be quantized and the activation outlier percentage parameter as the outlier threshold, and remove the activation to be quantized in the absolute value of the activation to be quantized that is greater than the outlier threshold.
[0030] The present invention has the following advantages compared with the prior art:
[0031] By integrating the advantages of pruning and quantization, the present invention makes the lightweight model smaller in size and higher in accuracy. Specifically, the method first constructs a non-lightweight convolutional neural network model according to the fire detection task, and replaces the convolutional layers in the original convolutional neural network model with fusible branch convolutional blocks to build a convolutional neural network model that supports two-level pruning and quantization. The fusible branch convolutional block includes a main branch and a shortcut branch, and pruning is performed simultaneously at the channel level and the hierarchical level through sparsification of the channel-level importance factor and the hierarchical importance factor during model training. After pruning, the two branches of the fusible branch convolutional block are fused into a normal convolutional layer to obtain the final pruned model; then, the pruned model is quantized through the optimal weight outlier percentage parameter and activation outlier percentage parameter adaptively. This method performs pruning simultaneously at the channel level and the hierarchical level, and quantizes adaptively according to the characteristics of each layer of the model, effectively realizing the lightweight of the model, making the performance of the lightweight model higher and the size smaller; by lightweighting the fire detection model, the problems of high hardware cost and low operation efficiency of the existing fire detection model are solved.
[0032] The technical solution of the present invention will be further described in detail below with reference to the drawings and embodiments. Description of the Drawings
[0033] Figure 1 Schematic diagram of the corresponding relationship between floating-point and quantization values of eight-bit symmetric quantization in the prior art;
[0034] Figure 2a Structural diagram of the fusible branch convolutional block when the feature map matching operation in the embodiment of the lightweight fire detection method based on two-level pruning and adaptive quantization of the present invention is a no-operation;
[0035] Figure 2b Schematic diagram of the structure of the fusible branch convolutional block when the feature map matching operation in the embodiment of the lightweight fire detection method based on two-level pruning and adaptive quantization of the present invention is a 1×1 convolutional kernel;
[0036] Figure 2cThe structural diagram of the fusible branch convolution block when the feature map matching operation of the lightweight fire detection method embodiment based on double-level pruning and adaptive quantization in the present invention is average pooling;
[0037] Figure 3 The schematic diagram of the non-lightweight model structure of the fire detection of the lightweight fire detection method embodiment based on double-level pruning and adaptive quantization in the present invention;
[0038] Figure 4 The schematic diagram of the fusion process of the fusible branch convolution block of the lightweight fire detection method embodiment based on double-level pruning and adaptive quantization in the present invention. Detailed implementation manners
[0039] Embodiment of the lightweight fire detection method based on double-level pruning and adaptive quantization:
[0040] The lightweight fire detection method based on double-level pruning and adaptive quantization includes the following steps:
[0041] Step 1. Construct a non-lightweight convolutional neural network model according to the fire detection task to be implemented. The convolutional neural network model includes an input layer, a convolutional layer, a pooling layer, and an output layer. The input layer receives known fire image data and passes it to the convolutional layer. The convolutional layer outputs corresponding feature maps according to the input feature maps. The pooling layer downsamples the feature maps output by the convolutional layer and outputs corresponding feature maps. The output layer makes a final recognition result according to the feature maps output by the convolutional layer and the pooling layer. All convolutional layers in this convolutional neural network model adopt fusible branch convolution blocks, that is, the original convolution in the convolutional neural network model is replaced with this fusible branch convolution block. The fusible branch convolution block includes a main branch and a shortcut branch, and the outputs of the main branch and the shortcut branch are fused as the final output of the fusible branch convolution block. This convolutional neural network model can be: R-CNN model, YOLO model, or SSD model, etc.
[0042] When the convolutional neural network model is an R-CNN model, the R-CNN model includes a region proposal network, a convolutional layer, a pooling layer, and a classifier. The region proposal network is used to generate proposal windows. The RPN network is used to receive the feature maps extracted by the convolutional layer as inputs and output feature maps. The pooling layer converts the feature maps output by the convolutional layer into feature maps of a fixed size. The classifier performs specific classification operations according to the above-mentioned feature maps and outputs the final detection result; by replacing this convolutional layer with the fusible branch convolution block of the present application, the lightweight operation of this model is realized.
[0043] When the convolutional neural network model is the YOLO model, the YOLO model includes an input end, a backbone network, a neck network, and an output end, where the backbone network includes 5 convolutional layers and 3 fully connected layers; similarly, by replacing all 5 convolutional layers with the fusible branch convolutional block of the present application, the lightweight operation of the model is realized.
[0044] When the convolutional neural network model is the SSD model, the SSD model includes a feature extraction layer, additional feature layers, a detection layer, anchor boxes, and a loss function. The additional feature layers include multiple feature maps that form a feature pyramid for object detection, and object detection is performed at multiple scales to improve the detection accuracy; the detection layer sets feature layers of different scales, enabling the model to predict objects on feature maps with different receptive fields, thereby realizing the detection of objects of different sizes. By replacing all convolutional layers with the fusible branch convolutional block of the present application, the lightweight operation of the model is realized.
[0045] As Figure 3 shown, in the above main branch, the input feature map passes through the first convolutional layer and the batch normalization layer in sequence to output the feature map. The feature map of each channel performs a linear multiplication operation with the channel-level importance factor of the corresponding channel to obtain the feature map of each channel, and then performs a linear multiplication operation with the hierarchical importance factor to complete the output of the first feature map; in the above shortcut branch, after the input feature map undergoes a feature map matching operation, the feature map of each channel performs a linear multiplication operation with the channel-level importance factor of the corresponding channel, and then performs a linear multiplication operation with the adaptive information control factor to complete the output of the second feature map. As Figure 3 shown, the fusion process is as follows: First, the product operation of the channel-level and hierarchical importance factors is a linear multiplication, which can be fused into the batch normalization layer in front of it after model training; then, the batch normalization layer and the discrete convolutional layer in front of it can also be linearly fused; therefore, the main branch can be finally fused into a convolutional layer after the model training ends; the expression of the fusible branch convolutional block after the main branch is fused is as follows:
[0046] Y = X * W + b + h i · g i · f(X)
[0047] where X and Y are the feature maps input and output by the corresponding fusible branch convolutional block, f(X) represents performing a feature map matching operation on the input X, * is the convolutional operation, W and b are the weights and biases of the convolution after the main branch convolutional layer is fused with the batch normalization layer and the importance factors, · is the linear multiplication, g i is the adaptive information control factor of the i-th layer shortcut branch, and h i is the channel-level importance factor of the i-th layer.
[0048] The shortcut branches are fused into a single convolutional layer, and this convolutional layer is fused with the convolutional layer of the main branch to form a single convolutional layer. The expression of the fusible branch convolutional block is as follows:
[0049] Y = X * W f + b f
[0050] where W f and b f are the weights and biases of the ordinary convolution after fusing the main branch and the shortcut branches.
[0051] Express f(·) using a convolution with weight W s and bias b s . Then the expressions of W f and b f are as follows:
[0052]
[0053] where t and u represent the total number of output and input channels respectively; and are the weights and biases of the j-th output channel and the q-th input channel of the convolution transformed from the feature map matching operation of the shortcut branch, and are the weights and biases of the j-th output channel and the q-th input channel of the ordinary convolution after fusing the main branch and the shortcut branches. W jq and b jq are the weights and biases of the j-th output channel and the q-th input channel of the convolution after fusing the batch normalization layer and the importance factor in the convolutional layer of the main branch.
[0054] Figures 2a - 2c In i , m ij is the hierarchical importance factor, used to represent the importance degree of the i-th layer, h i is the channel-level importance factor, used to represent the importance degree of the j-th channel of the i-th layer, g Figure 4 is the adaptive information control factor of the shortcut branch, used to adaptively control the flow of information of the shortcut branch at the i-th layer. All three are trainable adaptive parameters. As shown in Figure 2, the above feature map matching operation includes an existence operation, a 1×1 convolution operation, and an average pooling operation, where f(·) represents the feature map matching operation, and it determines whether the number of channels and the size of the input feature map are consistent with those of the output first feature map based on the input feature map and the convolutional layer of the main branch;
[0055] As Figure 2b shown, when the number of channels of the input feature map is inconsistent with that of the output first feature map, the feature map matching operation uses a 1×1 convolution operation. At this time, the above-mentioned The expression is as follows:
[0056]
[0057] As Figure 2c shown, when the size of the input feature map is inconsistent with that of the first output feature map, the feature map matching operation adopts average pooling operation. The above-mentioned The expression is as follows:
[0058]
[0059] As Figure 2a shown, when the number of channels and the size of the input feature map are both consistent with those of the first output feature map, the feature map matching operation adopts no operation. The above-mentioned The expression is as follows:
[0060]
[0061] Among them, Padding is to obtain a v×v convolution kernel by padding 0 around the operation value, and Copy is to obtain a v×v convolution kernel by copying the operation value. is the weight of the j-th output channel and the q-th input channel of the 1×1 convolution on the shortcut branch, and v represents the size of the convolution kernel of the convolutional layer. Copy(0) is to generate a v×v convolution kernel by copying 0 (the convolution kernel is a matrix, that is, the values in the matrix are all 0). is to copy to generate a v×v convolution kernel (the values in the convolution kernel matrix are all ), is to pad 0 around to generate a v×v convolution kernel (the value at the center position of the convolution kernel matrix is and the other values are 0), and Padding(1) is to pad 0 around the value 1 to generate a v×v convolution kernel (the value at the center position of the convolution kernel matrix is 1 and the other values are 0).
[0062] The above-mentioned calculation is relatively simple. Specifically, when the feature map matching operation on the shortcut branch adopts 1×1 convolution operation, is equal to the bias value of the corresponding 1×1 convolution; when the feature map matching operation is no operation or average pooling operation, are both 0.
[0063] In summary, the structure of the fusible branch convolutional block is fused into an ordinary convolutional layer through mathematical linear calculation; therefore, the fusible branch convolutional block can replace its ordinary convolutional layer in the convolutional neural network model without affecting the final structure of the model.
[0064] Step 2: Conduct L1 sparsification training on the convolutional neural network model to make the layer importance factor and channel-level importance factor tend to zero. After training, trim the unimportant layers and channels, and fuse the fusible branch convolutional blocks that are not trimmed into ordinary convolutional layers. The fused convolutional neural network model is called the pruning model; an unimportant layer refers to a layer whose layer importance factor is less than 0.0001, and an unimportant channel refers to a channel whose channel-level importance factor is less than 0.0001. L1 sparsification means making many values in the model parameters tend to zero by using the L1 regularization technique, thereby achieving the sparsification of the model.
[0065] Step 3: Quantify the pruning model, and remove the weight outliers and activation outliers for each layer of the network according to the weight outlier percentage parameter and activation outlier percentage parameter; the pruned model after quantization is called the dynamic quantization model. The initial values of the above-mentioned weight outlier percentage parameter and activation outlier percentage parameter are both randomly generated by the computer.
[0066] The removal process is as follows: Obtain the absolute value of the weight to be quantized, use the product of the maximum value in the absolute value of the weight to be quantized and the weight outlier percentage parameter as the outlier threshold, and remove the weight to be quantized in the absolute value of the weight to be quantized that is greater than the outlier threshold; obtain the absolute value of the activation to be quantized, use the product of the maximum value in the absolute value of the activation to be quantized and the activation outlier percentage parameter as the outlier threshold, and remove the activation to be quantized in the absolute value of the activation to be quantized that is greater than the outlier threshold.
[0067] Step 4: Take the minimum loss value of the feature maps output by the dynamic quantization model and the pruning model as the optimization goal. The loss value is determined by the difference between the feature maps output by the dynamic quantization model and the pruning model. The weight outlier percentage parameter and activation outlier percentage parameter are continuously iterated until the final weight outlier percentage parameter and activation outlier percentage parameter are obtained.
[0068] Step 5: Conduct actual quantization on the pruning model according to the final weight outlier percentage parameter and activation outlier percentage parameter. The actually quantized model is a lightweight fire detection model based on double-level pruning and adaptive quantization.
[0069] Step 6: Input the fire image to be detected into the lightweight fire detection model based on double-level pruning and adaptive quantization, and output the corresponding recognition result.
[0070] As Figure 3As shown, this lightweight method can be applied to the lightweighting of fire detection models. First, a convolutional neural network model is built based on the SSD object detection algorithm. Since the SSD algorithm has a high speed advantage in processing detection tasks, in order to meet the high real-time requirements of fire detection, the SSD algorithm is selected here to build a convolutional neural network model for fire detection. Of course, other object detection algorithms can also be used to build this convolutional neural network model. The image data for training uses the public dataset of the School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen). The dataset contains a total of 2,688 flame images, with 3,273 flame targets, including many complex scenarios, covering day, night, indoor, outdoor, forest, house, etc. Use the above steps 1-5 to lightweight this model to obtain a corresponding lightweight fire detection model based on double-level pruning and adaptive quantization. During the lightweighting process, after pruning, the layers and channels of the model are reduced. Before quantization, 32-bit floating-point numbers are used, and after quantization, they become 8-bit fixed-point numbers, and the scale of the corresponding model becomes smaller. The lightweighted model is compared with the non-lightweighted model and the model lightweighted using existing technologies. The main comparison is the detection accuracy and scale of the models. The specific comparison results are shown in Table 1 below.
[0071] Table 1
[0072]
[0073]
[0074] According to the experimental results in Table 1, it can be seen that double-level pruning can effectively compress the model: compared with single channel-level pruning and layer-level pruning, the lightweighted model obtained by double-level pruning has higher accuracy and smaller model scale; adaptive quantization can effectively compress the model: compared with general quantization, the lightweighted model obtained by adaptive quantization has higher accuracy; applying the lightweight fire detection method based on double-level pruning and adaptive quantization of the present invention to lightweight the fire detection model can compress the model to less than one-tenth of its original scale without a significant decrease in performance.
[0075] The above fire detection model examples show that the lightweight fire detection method based on double-level pruning and adaptive quantization proposed by the present invention can effectively reduce the model scale while ensuring the model performance, and has outstanding technical advancement and engineering practicability.
[0076] The above is only a preferred embodiment of the present invention, and does not impose any limitation on the present invention. Any simple modification, change, and equivalent structural change made to the above embodiments according to the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A lightweight fire detection method based on two-stage pruning and adaptive quantization, characterized in that: The method comprises the following steps: Step 1: construct a non-lightweight convolutional neural network model according to the fire detection task to be implemented, wherein the convolutional neural network model includes an input layer, a convolutional layer, a pooling layer and an output layer, wherein the input layer receives known fire image data and passes it to the convolutional layer, the convolutional layer outputs a corresponding feature map according to the input feature map, the pooling layer downsamples the feature map output by the convolutional layer, and outputs the corresponding feature map, and the output layer makes a final recognition result according to the feature maps output by the convolutional layer and the pooling layer; the convolutional neural network includes multiple convolutional layers; All convolutional layers in the convolutional neural network model use fusible branch convolution blocks, which include trunk branches and shortcut branches. The outputs of the trunk branches and shortcut branches are fused as the final output of the fusible branch convolution block. The input feature map in the trunk branch is sequentially passed through the first convolution layer and the batch normalization layer to output the feature map, the feature map of each channel is linearly multiplied with the channel-level importance factor of the corresponding channel to obtain the feature map of each channel, and then linearly multiplied with the level importance factor to complete the output of the first feature map; The input feature map in the shortcut branch is first subjected to a feature map matching operation, and then the feature map of each channel is linearly multiplied with the channel-level importance factor of the corresponding channel, and then linearly multiplied with the adaptive information control factor to complete the output of the second feature map; Step 2: Perform L1 sparsification training on the convolutional neural network model to make the layer-level importance factor and channel-level importance factor approach zero. After training, unimportant layers and channels are pruned, and the unpruned fusible branch convolution blocks are fused into ordinary convolution layers. The fused convolutional neural network model is called a pruned model; an unimportant layer refers to a layer whose layer-level importance factor is less than 0.0001, and an unimportant channel refers to a channel-level importance factor of a channel less than 0.0001; Step 3, quantize the pruned model, and remove weight outliers and activation value outliers from each layer of the network according to the weight outlier percentage parameter and the activation value outlier percentage parameter; the quantized pruned model is called a dynamic quantization model; Step 4: Taking the minimum loss value of the feature map output by the dynamic quantization model and the pruning model as the optimization goal, the weight outlier percentage parameter and the activation value outlier percentage parameter are continuously iterated until the final weight outlier percentage parameter and the activation value outlier percentage parameter are obtained; Step 5: The pruned model is actually quantized according to the final weight outlier percentage parameter and the activation value outlier percentage parameter. The model after actual quantization is a lightweight fire detection model based on dual-stage pruning and adaptive quantization. Step 6: Input the fire image to be detected into the lightweight fire detection model based on two-level pruning and adaptive quantization, and output the corresponding recognition result.
2. The lightweight fire detection method based on double-stage pruning and adaptive quantization according to claim 1 is characterized in that: The feature map matching operation includes a yes / no operation, a 1×1 convolution operation, and an average pooling operation. It is determined whether the number of channels and the size of the input feature map and the output first feature map are consistent based on the input feature map and the convolution layer of the trunk branch. When the number of channels of the input feature map and the output first feature map are inconsistent, the feature map matching operation adopts a 1×1 convolution operation. When the size of the input feature map and the output first feature map are inconsistent, the feature map matching operation adopts an average pooling operation. When the number of channels and the size of the input feature map and the output first feature map are consistent, the feature map matching operation adopts a no operation.
3. The lightweight fire detection method based on double-stage pruning and adaptive quantization according to claim 2 is characterized in that: The fusible branch convolution block in step 2 has the property of being fused into a normal convolution, and its expression is as follows: Y=(X*W+b)+(X*W s +b s )·h i ·g i =X*W f +b f Among them, * is the convolution operation, · represents linear multiplication, and h i is the channel-level importance factor of the i-th layer, g i is the adaptive information control factor of the shortcut branch of the i-th layer, X and Y are the feature maps of the corresponding fusion branch convolution block input and output, W and b are the weights and biases of the convolution after the main branch convolution layer is fused with the batch normalization layer and the importance factor, W s and b s is the weight and bias of the convolution transformed by the shortcut branch feature map matching operation, W f and b f The weights and biases of the ordinary convolution after the fusion of the trunk branch and the shortcut branch.
4. The lightweight fire detection method based on double-stage pruning and adaptive quantization according to claim 3 is characterized in that: When the feature map matching operation adopts a 1×1 convolution operation, is equal to the bias value corresponding to the 1×1 convolution, The expression is as follows: When the feature map matching operation adopts the average pooling operation, is zero, The expression is as follows: When the feature map matching operation adopts no operation, is zero, The expression is as follows: Among them, j is the output channel, q is the input channel, and v represents the convolution kernel size of the convolution layer. and The weights and biases of the j-th output channel and the q-th input channel of the convolution transformed by the shortcut branch feature map matching operation, For Fill the surrounding with 0 to generate a convolution kernel of size v×v. is the weight of the j-th output channel and the q-th input channel of the 1×1 convolution on the shortcut branch. Copy(0) generates a convolution kernel of size v×v by copying 0. To copy Generate a convolution kernel of size v×v. Padding(1) is to fill 0 around the value 1 to generate a convolution kernel of size v×v.
5. The lightweight fire detection method based on double-stage pruning and adaptive quantization according to claim 1 is characterized in that: The initial values of the weight outlier percentage parameter and the activation value outlier percentage parameter are randomly generated by a computer.
6. The lightweight fire detection method based on double-stage pruning and adaptive quantization according to claim 1 is characterized in that: The removal process of step 3 is as follows: obtaining the absolute value of the weight to be quantified, taking the product of the maximum value among the absolute values of the weight to be quantified and the weight outlier percentage parameter as the outlier threshold, and removing the weight to be quantified that is greater than the outlier threshold among the absolute values of the weight to be quantified; obtaining the absolute value of the activation to be quantified, taking the product of the maximum value among the absolute values of the activation to be quantified and the activation outlier percentage parameter as the outlier threshold, and removing the activation to be quantified that is greater than the outlier threshold among the absolute values of the activation to be quantified.
Citation Information
Patent Citations
Visible light forest fire detection method based on lightweight anchor-free detection model
CN115423998A
Method for quickly acquiring total energy response matrix
CN116009055A