Improved YOLOv5 Fire Detection Method and Device Incorporating Adjustable Coordinate Residual Attention
By constructing the fire dataset and integrating the improved YOLOv5 neural network with adjustable coordinate residual attention, the problem of inaccurate flame target segmentation and extraction is solved, and the rapid and accurate identification of fire detection is achieved, which is suitable for real-time detection on mobile terminals.
Patent Information
- Application Number
- CN202210981425.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-16
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-08-16
AI Technical Summary
In the existing fire detection technology, there is inconsistent noise area in the segmentation and extraction of flame targets, resulting in insufficient recognition speed and accuracy. The traditional system is greatly affected by environmental changes and lacks diversity in data sets, which affects the accuracy and speed of fire recognition.
The fire dataset is constructed, and the improved YOLOv5 neural network with adjustable coordinate residual attention is integrated. The position information is encoded horizontally and vertically through the attention mechanism, combined with the improved Bottleneck-CSP module, the YOLOv5 model is optimized, and the adaptive anchor box and focus loss function are used to improve detection accuracy.
It realizes rapid and accurate identification of flames and smoke, reduces the risk of missing the best remediation time in early fires, and is suitable for real-time mobile detection.
Smart Images

Figure CN115457428B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and image-based fire detection and recognition, in particular to an improved YOLOv5 fire detection method and device incorporating adjustable coordinate residual attention. Background Technique
[0002] Accurately identifying early fire detection is an important means of fire safety. It is valuable and necessary to study a fire monitoring and alarm system with fast response capabilities. Fire monitoring and alarm systems have been studied for decades.
[0003] Chinese Patent Application CN113869567A describes a control method, device, computer equipment, and storage medium based on fire prediction information applicable to multiple scenarios. It mainly conducts fire prediction and control for fire scenario data such as temperature, smoke flow, and fire protection data. The above implementation method requires a certain amount of computing power support and has a relatively high cost. Chinese Patent Application CN113673748A discloses a fire prediction method based on the XGBoost model, which mainly uses the XGBoost model to achieve fire prediction. However, the XGBoost model cannot model spatio-temporal positions and cannot capture images well.
[0004] Muhammad et al. [Muhammad K, Ahmad J et al (2019) Efficient deep CNN-based fire detection and localization in video surveillance applications. IEEE Trans Syst Man Cybern Syst 19(7):1419–1434] classified fire detection methods into two categories: traditional fire alarms and vision sensor-assisted fire detection. Currently, most fire detection and fire alarm systems are based on traditional fire detection or fire alarm systems. For example, Xu et al. [Xu Y, Zhang J et al (2013) The structure of automatic fire alarm system based on virtual instrument. J Tianjin Univ Technol 29(3):30–36] proposed a fire alarm system based on a fire alarm controller, temperature, and smoke detectors. Hu et al. [Hu X (2013) Research and product development of MIR flame detector system. J Zhejiang Univ 10(1):78] proposed a multi-band infrared fire detector. However, the monitoring range of systems based on these sensors is limited, and the performance of these systems is vulnerable to environmental changes.
[0005] With the popularization of video surveillance systems, the research on visual sensor-assisted fire detection has attracted much attention. The advantages of fire detection based on images / videos include rapid response, insensitivity to environmental temperature, and real-time images or videos of the accompanying fire scene. In image / video-based fire detectors, fire objects are abstracted into image features generated from color, brightness, texture, shape, and motion information. Toptas proposed a remote video surveillance system based on network cameras and image processing technology for fire monitoring and alarm technology [Toptas B, Hanbay D. A new artificial bee colonyalgorithm-based color space for fire / flame detection[J]. Soft Computing, 2019(2): 1-12.]. Wan et al. [Wan Z (2020) Fire detection from images based on singleshot multibox detector. Hohai university,Nanjing] proposed an improved SSD to detect fires in images by using data augmentation and modifying the scale and number of default boxes, but its accuracy only reached 84.75%. Shen et al. [Shen D, Chen X, Yan W (2018) Flame detection using deep learning.In: Proceedings of the 2018 4th international conference on controlautomation and robotics, pp 416–420] proposed an optimized YOLO model for detecting flame objects from video frames. However, the dataset adopted lacks data diversity because the samples are from 194 images.
[0006] The difficulty of the above-mentioned fire recognition technology based on digital image processing lies in the segmentation and extraction of flame targets. Previously, the extraction of flame and smoke targets was mainly achieved through mining methods and contour tracking techniques. However, in practical applications, the obtained images are noisy, and the noise areas vary in size, often resulting in image damage. This not only takes a lot of time but also requires the difference between the object and the actual contour. Without specifically using attention to extract more feature information from the feature extraction part of the training images, there will be cases of misrecognition or missed recognition, affecting the speed and accuracy of flame recognition. Summary of the Invention
[0007] In view of the above problems, the present invention proposes an improved YOLOv5 fire detection method and device incorporating adjustable coordinate residual attention to overcome or at least partially solve the above problems.
[0008] The present invention provides an improved YOLOv5 fire detection method incorporating adjustable coordinate residual attention, and the method includes:
[0009] Construct a fire dataset, where the fire dataset includes video data and first picture data of different fire degrees collected in a laboratory ignition experiment, extract second picture data from the video data, and add labels of flames and / or smoke to the first picture data and the second picture data;
[0010] Establish an improved YOLOv5 neural network incorporating adjustable coordinate residual attention, and use the fire dataset to train the improved YOLOv5 neural network as a fire detection model;
[0011] Deploy the fire detection model to a mobile device. After the mobile device receives real-time video data captured by a camera, the mobile device uses the fire detection model to detect and identify fire targets in the real-time video data.
[0012] Optionally, the improved YOLOv5 neural network incorporating adjustable coordinate residual attention includes: a backbone network Backbone, a neck network Neck, and a head network Head;
[0013] Among them, the backbone network Backbone is mainly used to extract key features from the input image; the neck network Neck is mainly used to create a feature pyramid; the head network Head is mainly responsible for the final detection step, and it constructs a final output vector with class probabilities, objectness scores, and bounding boxes using anchor boxes.
[0014] Optionally, the establishment of the improved YOLOv5 neural network incorporating coordinate attention includes:
[0015] Add an attention mechanism to the backbone network of YOLOv5 feature extraction, use the attention mechanism to encode long-range dependencies and position information from the input image in the horizontal and vertical spatial directions respectively, and then aggregate the features.
[0016] Optionally, the final output of the attention mechanism is expressed as follows:
[0017]
[0018] Among them, represents the input feature map, and respectively represent the attention weights in two spatial directions. The formula is as follows:
[0019]
[0020]
[0021] where and are the feature tensors of the information decomposition of feature F in two directions respectively. (·) and (·) respectively represent the convolution operations with a convolution kernel of 1×1. is a hyperparameter that can automatically adjust the feature weights in the horizontal and vertical directions.
[0022]
[0023]
[0024] where, and respectively represent the original features in two directions. Concat(·) represents the concatenation operation of two features.
[0025]
[0026]
[0027] where, H and W are the height and width of the input feature map respectively, is the output of the c-th channel with height h, is the output of the c-th channel with width w, is the input image feature of the c-th channel.
[0028] Optionally, the backbone network includes four Bottleneck-CSP-New modules to replace the Bottleneck-CSP module in the original YOLOv5 neural network;
[0029] The Bottleneck-CSP-New module includes a first module and a second module; the first module uses a 1×1 convolutional layer to halve the number of channels, then passes through the residual structure Bottleneck module, controls the number of channels of the hidden layer in the Bottleneck module through parameters, and then passes through a Conv.2Dxl module without passing through BN and activation functions; the second module performs a shortcut connection operation on the input feature without any change and the output of the first module, and finally outputs after passing through BN+Relu and a normal Conv.2Dxl convolution.
[0030] Optionally, the loss function of the fire detection model is as follows:
[0031]
[0032] This loss function is the total loss function of the model and is specifically as follows:
[0033]
[0034] Among them, represents the aspect ratio of the target detection box, represents the aspect ratio of the predicted detection box.
[0035]
[0036]
[0037] Among them, A and B respectively represent the predicted detection box and the target detection box. a and respectively represent the center points of the predicted detection box and the target detection box. represents the Euclidean distance between the two center points, and c represents the diagonal length of the smallest closed area of the predicted frame that simultaneously contains the target frame. is a hyperparameter of variable parameters.
[0038]
[0039] Among them, is set to 1, is set to 2, is the size of the prediction probability, and y is to judge whether it is a positive sample. The focal loss function replaces the cross-entropy loss function as the confidence and classification loss of the network.
[0040] The present invention also provides an improved YOLOv5 fire detection device incorporating adjustable coordinate residual attention. The device includes:
[0041] A data collection module for constructing a fire data set. The fire data set includes video data and first picture data of different fire degrees collected in a laboratory ignition experiment, extracting second picture data from the video data, and adding marks of flames and / or smoke to the first picture data and the second picture data;
[0042] A model establishment module for establishing an improved YOLOv5 neural network incorporating coordinate attention, and training the improved YOLOv5 neural network using the fire data set as a fire detection model;
[0043] A model deployment module is used to deploy the fire detection model to a mobile device. After the mobile device receives real-time video data captured by a camera, the mobile device uses the fire detection model to detect and identify fire targets in the real-time video data.
[0044] The present invention also provides a computer-readable storage medium, which is used to store program code for executing the method described in any one of the above.
[0045] The present invention also provides a computing device, which includes a processor and a memory:
[0046] The memory is used to store program code and transmit the program code to the processor;
[0047] The processor is used to execute the method described in any one of the above according to the instructions in the program code.
[0048] Aiming at the problems of low accuracy and slow speed of existing manual and sensor fire detections, based on the analysis of fire image features, the present invention provides an improved YOLOv5 fire detection method and device incorporating adjustable coordinate residual attention. The present invention uses the YOLOv5 neural network to automatically extract and learn image features. First, by adding an attention mechanism, position information is embedded into channel attention, enabling the network to obtain a wider range of information and improving the detection accuracy for small targets and blurred smoke boundaries. At the same time, the Bottleneck-CSP model in the backbone network is improved to reduce model parameters and the volume of the model, providing effective support for the deployment of the model. This method can quickly and accurately identify detection objects and can detect fires intuitively in real time. This method can not only identify and detect flame information generated by fires, but also identify and detect smoke generated in the early stage of fires, reducing the loss of missing the best remedial time in the early stage of fires. For early fires, it reduces missing the best remedial time and conducts early fire detection in a timely manner.
[0049] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present invention more obvious and understandable, the following specifically describes the embodiments of the present invention.
[0050] According to the following detailed description of the specific embodiments of the present invention in conjunction with the drawings, those skilled in the art will understand the above and other purposes, advantages, and features of the present invention more clearly. Description of the Drawings
[0051] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to denote the same components. In the drawings:
[0052] Figure 1 Schematic diagram of the improved YOLOv5 fire detection method incorporating adjustable coordinate residual attention according to an embodiment of the present invention;
[0053] Figure 2 Schematic diagram of the overall structure of the improved YOLOv5 network according to an embodiment of the present invention;
[0054] Figure 3 Schematic diagram of the overall implementation of the attention mechanism according to an embodiment of the present invention;
[0055] Figure 4 Schematic diagram of the comparison before and after improving the original BottleneckCSP module according to an embodiment of the present invention;
[0056] Figure 5 Schematic diagram of the structure of the improved YOLOv5 fire detection device incorporating adjustable coordinate residual attention according to an embodiment of the present invention. Detailed implementation manners
[0057] The following will describe the exemplary embodiments of the present invention in more detail with reference to the drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be completely conveyed to those skilled in the art.
[0058] An embodiment of the present invention provides an improved YOLOv5 fire detection method incorporating adjustable coordinate residual attention. As Figure 1 shown, the improved YOLOv5 fire detection method incorporating coordinate attention according to an embodiment of the present invention may at least include the following steps S101 to S103.
[0059] S101, construct a fire dataset, where the fire dataset includes video data of different fire intensities collected in a laboratory ignition experiment and first picture data, extract second picture data from the video data, and add marks of flames and / or smoke to the first picture data and the second picture data.
[0060] Optionally, multiple ignition tests can be conducted in the laboratory to collect videos and pictures of fires at different levels. Among them, the size of the combustion pan is simulated according to the size of the combustion pan specified in the national standard "Special Fire Detectors" for image fire detectors to collect small-target fire picture data and / or video data. The video data is intercepted to obtain fire picture data in various combustion states. Use the picture marking tool (labelImg) to mark the fire pictures, mark the area of interest in each picture, and manually mark the flame and smoke parts in each picture. Optionally, the first picture data and the second picture data can be divided into: only flame targets, only smoke targets, and targets that contain both smoke and flame targets. That is, for any image data, target marking can be aligned. For example, it can be determined that the image data contains a smoke target, a flame target, or both a smoke target and a flame target.
[0061] S102. Establish an improved YOLOv5 neural network incorporating coordinate attention, and use the fire dataset to train the improved YOLOv5 neural network as a fire detection model.
[0062] S103. Deploy the fire detection model to the mobile device. After the mobile device receives the real-time video data captured by the camera, the mobile device uses the fire detection model to detect and identify fire targets in the real-time video data.
[0063] This embodiment can improve the traditional YOLOv5 neural network by adding an attention mechanism; input the labeled fire dataset in the format required by the neural network, input it into the improved YOLOv5 neural network for training and test the results; deploy the trained model to the mobile device to perform fire inspection and identification tasks.
[0064] As Figure 2 shown, the improved YOLOv5 neural network incorporating coordinate attention includes: a backbone network Backbone, a neck network Neck, and a head network Head; among them, the backbone network Backbone is mainly used to extract key features from the input image; the neck network Neck is mainly used to create a feature pyramid; the head network Head is mainly responsible for the final detection step, and it uses anchor boxes to construct the final output vector with class probabilities, objectness scores, and bounding boxes.
[0065] The above step S102 of establishing an improved YOLOv5 neural network incorporating adjustable coordinate residual attention includes: adding an attention mechanism to the backbone network of YOLOv5 feature extraction, using the attention mechanism to encode the long-range dependence relationship and position information in the horizontal and vertical spatial directions of the input image respectively, and then aggregating the features.
[0066] When integrating the improved attention mechanism, the coordinate attention can be first added to the backbone network of YOLOv5 feature extraction. Coordinate attention is a lightweight and efficient attention mechanism that embeds position information into channel attention, allowing the mobile network to acquire knowledge in a larger range. Figure 3 FIG. shows the overall implementation schematic diagram of the attention mechanism according to an embodiment of the present invention. The attention mechanism of this embodiment is a coordinate attention mechanism. This attention encodes the long-range dependency and position information from the horizontal and vertical spatial directions respectively, and then aggregates the features. Therefore, it is necessary to decompose the features to capture the position information spatially. Specifically, it is decomposed along the horizontal and vertical directions. The input feature map, that is , using pooling kernels and encode the features in the horizontal and vertical directions respectively, and the outputs of the c-th channel with height h and width w are respectively expressed as:
[0067]
[0068]
[0069] where H and W are the height and width of the input feature map respectively, is the output of the c-th channel with height h, is the output of the c-th channel with width w, The input image features for the c channel. The above two transformations aggregate the features with two spatial directions (X and Y). They generate a pair of direction-aware feature maps, enabling the attention mechanism to capture the long-range information of the feature map along one spatial path and retain the accurate position information along the other spatial path. The attention mechanism is widely used to improve the performance of the model. The inspiration for the attention mechanism comes from the way the human eye observes things, as the human eye always focuses on the most important aspects of things. Similarly, it allows the network to focus on important features, which helps improve the accuracy of the network. By applying the attention mechanism to the network model, the classification accuracy will be further improved. The essence of the attention mechanism is to weight the feature map, enabling the model to focus on important feature information and improve the generalization ability of the model. The SE attention mechanism uses 2D global pooling to calculate the channel attention weights and weights the feature information to optimize the model. However, the SE attention weights the channel dimension of the feature map but ignores the spatial dimension, which is crucial in computer vision tasks. CBAM uses channel pooling and convolution to weight the spatial dimension. However, convolution cannot capture the correlation of long-range information, which is crucial for visual tasks. Therefore, the present invention proposes a network fusion coordinate attention mechanism, which can obtain cross information, position-sensitive and direction-aware information. It helps the model focus on useful feature information. Global average pooling (GAP) is usually used to calculate the channel attention weights and globally encode the spatial information, performing GAP on each image feature in the spatial dimension H×W. However, it calculates the channel attention weights by compressing the global spatial information, thus losing the spatial information. Therefore, the two-dimensional global pooling is decomposed into one-dimensional global pooling in the horizontal and vertical directions to effectively utilize the spatial and spectral information. Specifically, 1D horizontal global pooling and vertical global pooling are used to encode each spectral dimension in the feature maps with spatial ranges of (H, 1) and (1, W). The above two formulas allow obtaining the correlation of long-range information in one spatial direction while retaining the position information in the other spatial direction, which helps the network focus on more information useful for classification. Then the two feature maps generated in the horizontal and vertical directions are encoded into two attention weights, each weight capturing the correlation of long-range information from the input feature map in one spatial direction.
[0070] The above two transformations are concatenated in the spatial dimension, and 1×1 convolution is used to compress the channels. Then BatchNorm and non-linearity are used to encode the spatial information in the vertical and horizontal directions, segment the encoded information, and use 1×1 convolution to adjust the channels of the attention map to be equal to the number of channels of the input feature map. Then, the sigmoid function is used for normalization and weighted fusion. The final output of the attention mechanism is represented as follows:
[0071]
[0072] Among them, represents the input feature map, c represents the c-th channel, h and w respectively represent the height and width of the input feature map. i and j respectively represent the height and width of the current vector. and respectively represent the attention weights in two spatial directions. The formula is as follows:
[0073]
[0074]
[0075] Among them and are the feature tensors of the information decomposition of feature F in two directions respectively. (·) and (·) respectively represent the convolution operations with a convolution kernel of 1×1. is a hyperparameter that can automatically adjust the feature weights in the horizontal and vertical directions.
[0076] This method is for target detection of flames. Considering that the flame morphology changes continuously over time and has different change characteristics in the horizontal and vertical directions. Therefore, through the hyperparameter the impacts of the changes in the horizontal and vertical directions on recognition are adjusted respectively. At the same time, the initial features of the flame are retained through residual connections, and the initial features are combined with the coordinate attention features to achieve better recognition effects.
[0077]
[0078]
[0079] Among them, and respectively represent the original features in two directions. Concat(·) represents the splicing operation of two features.
[0080] Furthermore, the backbone network of the embodiment of the present invention includes four Bottleneck-CSP-New modules to replace the Bottleneck-CSP modules in the original YOLOv5 neural network; the Bottleneck-CSP-New module includes a first module and a second module; the first module uses a 1×1 convolutional layer to halve the number of channels, then passes through the residual structure Bottleneck module, controls the number of channels of the hidden layer in the Bottleneck module through parameters, and then passes through a Conv.2Dxl module without passing through BN and activation functions; the second module performs a shortcut connection operation on the input features without any change and the output of the first module, and finally outputs after passing through BN+Relu and a normal Conv.2Dxl convolution.
[0081] The size of the modified model should be reduced as much as possible for deployment on hardware devices. Therefore, the backbone network of the YOLOv5 model is modified. The backbone network of the YOLOv5s architecture includes four Bottleneck-CSP modules, each with many convolutional layers. Although the convolution process can extract image features, there are many parameters in the convolution kernel, resulting in many parameters in the recognition model. Therefore, the convolutional layers on different branches of the original CSP module are deleted. The depth of another branch where the input feature map of the Bottleneck-CSP module is directly connected to the output feature map greatly reduces the number of parameters in the module. The four stages of the original backbone network using the Bottleneck-CSP module are replaced by four Bottleneck-CSP-New modules, which may ultimately lead to insufficient extraction of deep features in the image due to their lightweight properties. However, combining the attention mechanism can well extract image feature information while reducing model parameters, facilitating model deployment.
[0082] Figure 4The figure shows a comparison diagram before and after improving the original BottleneckCSP module of the embodiment of the present invention. The original Bottleneck-CSP module is divided into two parts: Bottleneck and CSP. The input features pass through two different modules. First, the first module uses a 1×1 convolutional layer (conv.2Dxl+BN+Hardwish) to halve the number of channels, then passes through the residual structure Bottleneck module, and controls the number of channels of the hidden layer in the Bottleneck module through parameters. Then, it passes through a Conv.2Dxl module without passing through BN and activation functions. Secondly, it is the second module. The input features pass through a Conv2d module without passing through BN and activation functions. The outputs of the two modules are subjected to a shortcut connection operation, and finally, after BN+Relu and a normal Conv.2Dxl convolution, the output is obtained.
[0083] The modified Bottleneck-CSP-New module is also divided into two modules. The first module is the same as the original feature extraction process. The second part is to perform a shortcut connection operation on the input features without any changes and the output of the first module, and finally, after BN+Relu and a normal Conv.2Dxl convolution, the output is obtained.
[0084] After obtaining the modified YOLOv5 neural network, the training of the improved YOLOv5 neural network in step S102 can be continued. The improved YOLOv5 neural network can be trained in the following ways.
[0085] First, the labeled data set is input in the format required by the neural network. The input end uses the Mosaic data augmentation method to splice the images in a way of random scaling, random cropping, and random arrangement, which has a significant effect on the detection of small target fire images. Adaptive anchor box calculation is performed, and the optimal anchor box values for different training sets are calculated adaptively each time training is performed.
[0086] Secondly, in the network training stage, the detailed steps of each module are as follows:
[0087] (1) The backbone network Backbone is mainly used to extract key features from the input image.
[0088] The Focus layer is the initial layer of the backbone network, which is used to simplify model calculations and improve training speed. Its uses are as follows: Using the slicing technique, first, the three-channel image is divided into four slices, each slice being 3×320×320. Secondly, the four parts are deeply connected using concatenation, and the output feature map size is 12×320×320.
[0089] The Conv layer is the second layer of the backbone network. Through the convolutional layer, convolutional operations are performed on the input feature map. A convolutional layer consisting of 32 convolutional kernels is used to form an output feature map with a size of 32×320×320. Finally, the result is output to the next layer through the BN layer (batch normalization) and the Hardswish activation function.
[0090] The BottleneckCSP - New module is the third layer of the backbone network, and its design purpose is to extract the depth information of the image more effectively. The BottleneckCSP - New module is mainly composed of Bottleneck modules, which connect the convolutional layer (Conv2d + BN + ReLu activation function) with convolutional kernels of size 1×1 and size 3×3. The final output of the Bottleneck module is the sum of the output of this part and the initial input through the residual structure. The first input of the BottleneckCSP - New module is divided into two branches. The number of channels of the feature map is halved using convolutions in the two branches. Then, using concat, the output feature maps of branch 1 and branch 2 are deeply connected through the input features in the Bottleneck module and the second branch. Finally, after gradually passing through the BN layer and the Conv2d layer, the output feature map of the module is created.
[0091] The CA - Res module is the fourth layer of the backbone network. The CA - Res module adopts an attention mechanism. This attention encodes the long - range dependencies and position information of the input image from the horizontal and vertical spatial directions respectively, and then learns the horizontal and vertical direction features through an adjustable residual structure and aggregates the horizontal and vertical direction features. Specifically, for the feature map output after passing through the BottleneckCSP - New module, that is , the pooling kernels and are used to encode the horizontal and vertical direction features respectively, then learn the horizontal and vertical direction features through an adjustable residual structure and aggregate the horizontal and vertical direction features. The sigmoid function is used to normalize and weight - fuse the horizontal, vertical, and input features.
[0092] Next, the Conv module, BottleneckCSP - New module, and CA module are repeated twice, and a convolutional operation of Conv is performed on the output image features again.
[0093] The SPP module (Spatial Pyramid Pooling) is the twelfth layer of the backbone network, designed to increase the receptive field of the network by converting feature maps of arbitrary sizes into fixed-size feature vectors. After the convolutional layer loop, a feature map of size 256×20×20 is output; the convolutional kernel size is 1×1. Then, through three concurrent max pooling layers for subsampling, this feature map is deeply connected with the output feature map, and the size of the output feature map is 1024×20×20.
[0094] (2)The Neck model is mainly used to create a feature pyramid. The feature pyramid helps the model's successful generalization in object scaling and helps in identifying the same object of different sizes and proportions.
[0095] This formula is used to select the feature map, where 224 is the pre-training size of a typical graph network, and x and y are the width x and height y of the ROI (Region of Interest) respectively. is the target level to which the ROI with x*y = 224 should be mapped. The feature pyramid is very beneficial for helping the model perform well on unknown data.
[0096] (3)The Head model is mainly responsible for the final detection step. It uses anchor boxes to construct the final output vector with class probabilities, objectness scores, and bounding boxes. The detection network of the YOLOv5s structure includes three detection layers, and each detection layer inputs a feature map with a size of 80×80, 40×40, and 20×20 for detecting image objects of various sizes. Each detection layer outputs a 21-channel vector, including two classes, a class probability, four surrounding box position coordinates, and three anchor boxes. Then, the predicted bounding boxes and classes of the targets in the original image are generated and labeled, thus realizing the detection of image targets.
[0097] The calculation of the loss function adopts the loss function, and the loss function of the fire detection model is as follows:
[0098]
[0099] This loss function is the total loss function of the model.
[0100] When the predicted detection box and the target detection box do not intersect, IoU cannot reflect the distance between the two structures. At this time, the function is not differentiable and cannot be optimized. The second problem is that the IoU is the same, but the positions of the predicted detection boxes are different, so the IoU_Loss cannot distinguish the differences in their intersections. DIoU_loss solves the above problems by considering the overlapping area of the two boxes and the distance between the center points. CIoU_Loss introduces the aspect ratio of the two frames on the basis of DIoU_Loss, making the convergence faster and the regression result better when the intersection ratio is 0. The aspect ratio and The expressions of are as follows:
[0101]
[0102] Among them, represents the aspect ratio of the target detection box, represents the aspect ratio of the predicted detection box.
[0103]
[0104]
[0105] Among them, A and B represent the predicted detection box and the target detection box respectively. a and represent the center points of the predicted detection box and the target detection box respectively. represents the Euclidean distance between the two center points, and c represents the diagonal length of the smallest closed region of the prediction frame that simultaneously contains the target frame. is a hyperparameter of variable parameters.
[0106] To solve the problem of unbalanced positive and negative samples, the focal loss is used instead of the cross-entropy loss function as the confidence and classification loss of the network. It assigns a higher loss weight to the foreground image, making the model more focused on the classification of the foreground.
[0107]
[0108] Among them, is set to 1, is set to 2, is the size of the predicted probability, and y is to judge whether it is a positive sample.
[0109] Deploy the trained fire detection model to the mobile device for the detection and recognition of fire targets, including deploying the trained weights and model to the mobile device, capturing the video through the camera and extracting it to the mobile device, and then using the fire detection model to perform feature analysis on each frame image of the input video stream to determine whether there are smoke targets and / or flame targets, so as to detect the emerging fire in real time. When detecting smoke and / or flame, an alarm can be output and a detection box can be popped up to prompt relevant personnel to take fire extinguishing measures.
[0110] Input the image with fire marks into the improved network model, which passes through three modules: the backbone network, the neck network, and the head network. The backbone network includes the slicing network, the convolutional network, the improved BottleneckCSP network, the coordinate attention mechanism, and the SPP network. Using the slicing technique, first divide the three-channel image into four slices, each slice being 3×320×320. Secondly, use concatenation to deeply connect the four parts, and the output feature map size is 12×320×320. After a series of feature extraction networks extract the input features of the image, the output feature map size is 1024×20×20. The neck network retains the spatial information through upsampling and downsampling operations. Finally, sample the feature maps of different sizes and process them into the same size, perform feature fusion and convolutional operations to obtain 3 feature layers of 20*20*255, 40*40*255, and 80*80*255 respectively, calculate using the GIoU loss function, make predictions, and generate prediction results. Then use non-maximum suppression to filter the multi-object boxes.
[0111] The embodiment of the present invention also provides an improved YOLOv5 fire detection device incorporating adjustable coordinate residual attention, as Figure 5 shown. The improved YOLOv5 fire detection device incorporating coordinate attention in the embodiment of the present invention may include:
[0112] A data collection module 510, configured to construct a fire data set, where the fire data set includes video data and first picture data of different fire degrees collected in a laboratory ignition experiment, extract second picture data from the video data, and add marks of flames and / or smoke to the first picture data and the second picture data;
[0113] A model establishment module 520, configured to establish an improved YOLOv5 neural network incorporating coordinate attention, and train the improved YOLOv5 neural network using the fire data set as a fire detection model;
[0114] The model deployment module 530 is used to deploy the fire detection model to the mobile device. After the mobile device receives the real-time video data captured by the camera, the mobile device uses the fire detection model to detect and identify the fire target in the real-time video data.
[0115] An embodiment of the present invention further provides a computer-readable storage medium, which is used to store program codes, and the program codes are used to execute the method described in the above embodiments.
[0116] An embodiment of the present invention further provides a computing device, which includes a processor and a memory: the memory is used to store program codes and transmit the program codes to the processor; the processor is used to execute the method described in any one of the above according to the instructions in the program codes.
[0117] Those skilled in the art can clearly understand that the specific working processes of the systems, devices, modules, and units described above can refer to the corresponding processes in the foregoing method embodiments. For the sake of brevity, they will not be described herein again.
[0118] In addition, each functional unit in the embodiments of the present invention can be physically independent of each other, or two or more functional units can be integrated together, or all functional units can be integrated in a processing unit. The above integrated functional units can be implemented in the form of hardware, or in the form of software or firmware.
[0119] Those of ordinary skill in the art can understand that if the integrated functional unit is implemented in the form of software and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, which includes several instructions, so that a computing device (such as a personal computer, a server, or a network device, etc.) can execute all or part of the steps of the method described in the embodiments of the present invention when running the instructions. The foregoing storage media include: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks, etc., which can store program codes.
[0120] Alternatively, all or part of the steps of implementing the foregoing method embodiments can be completed by hardware related to program instructions (such as a computing device such as a personal computer, a server, or a network device), and the program instructions can be stored in a computer-readable storage medium. When the program instructions are executed by the processor of the computing device, the computing device executes all or part of the steps of the method described in the embodiments of the present invention.
[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that within the spirit and principle of the present invention, it is still possible to modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the corresponding technical solutions to deviate from the protection scope of the present invention.
Claims
1. An improved YOLOv5 fire detection method incorporating adjustable coordinate residual attention, characterized in that, The method includes: Constructing a fire dataset, which includes video data and first picture data of different fire degrees collected in a laboratory ignition experiment, extracting second picture data from the video data, and adding marks of flames and / or smoke to the first picture data and the second picture data; Establishing an improved YOLOv5 neural network integrated with adjustable coordinate residual attention, and training the improved YOLOv5 neural network with the fire dataset to be used as a fire detection model; Deploying the fire detection model to a mobile device. After the mobile device receives real-time video data captured by a camera, the mobile device uses the fire detection model to detect and identify fire targets in the real-time video data; The establishment of the improved YOLOv5 neural network integrated with adjustable coordinate residual attention includes: adding an attention mechanism to the backbone network of YOLOv5 feature extraction, using the attention mechanism to encode the long-range dependence relationship and position information of the input image in the horizontal and vertical spatial directions respectively, then learning the horizontal and vertical direction features through an adjustable residual structure, and aggregating the horizontal and vertical direction features; The final output representation of the adjustable coordinate residual attention mechanism is as follows: , where represents the input feature map, and represent the attention weights in two spatial directions respectively, and the formula is as follows: , , where and are the feature tensors of information decomposition of the feature in two directions respectively, and represent the convolution operations with convolution kernels of respectively, is a hyperparameter that can automatically adjust the feature weights in the horizontal and vertical directions; , wherein, and respectively represent the original features in two directions, represents the splicing operation of the two features; , where and input the height and width of the feature map respectively, is the output of the th channel with a height of is the output of the th channel with a width of is the input image feature of the channel.
2. The method according to claim 1, characterized in that The improved YOLOv5 neural network integrated with adjustable coordinate residual attention includes: a backbone network Backbone, a neck network Neck, and a head network Head; Among them, the backbone network Backbone is used to extract key features from the input image; the neck network Neck is used to create a feature pyramid; the head network Head is responsible for the final detection step, and it constructs a final output vector with class probabilities, objectness scores, and bounding boxes using anchor boxes.
3. The method according to claim 2, wherein The backbone network includes four Bottleneck-CSP-New modules to replace the Bottleneck-CSP module in the original YOLOv5 neural network; The Bottleneck-CSP-New module includes a first module and a second module; the first module uses a 1×1 convolutional layer to halve the number of channels, then passes through a residual structure Bottleneck module, controls the number of channels of the hidden layer in the Bottleneck module through parameters, and then passes through a Conv.2Dxl module without passing through BN and activation functions; the second module performs a shortcut connection operation on the input features without any change and the output of the first module, and finally outputs after passing through BN+Relu and a normal Conv.2Dxl convolution.
4. The method according to claim 2, wherein The loss function of the fire detection model is as follows: , the loss function is the total loss function of the model, which is specifically as follows: , where represents the aspect ratio of the target detection box, and represents the aspect ratio of the predicted detection box; , where, represent the predicted detection box and the target detection box respectively, and represent the center point of the predicted detection box and the target detection box respectively, represents the Euclidean distance between the two center points, represents the diagonal length of the smallest closed region of the predicted frame that simultaneously contains the target frame, is a hyperparameter of variable parameters; , where is set to 1, is set to 2, is the magnitude of the predicted probability, is to judge whether it is a positive sample, and the focal loss function replaces the cross-entropy loss function as the confidence and classification loss of the network.
5. An improved YOLOv5 fire detection device incorporating adjustable coordinate residual attention, characterized in that, The device includes: A data collection module, which is used to construct a fire dataset, which includes video data and first picture data of different fire degrees collected in a laboratory ignition experiment, extracting second picture data from the video data, and adding marks of flames and / or smoke to the first picture data and the second picture data; A model establishment module, configured to establish an improved YOLOv5 neural network incorporating adjustable coordinate residual attention, and train the improved YOLOv5 neural network using the fire dataset as a fire detection model; A model deployment module, configured to deploy the fire detection model to a mobile device. After the mobile device receives real-time video data captured by a camera, the mobile device uses the fire detection model to detect and identify fire targets in the real-time video data.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store program codes, and the program codes are used to execute the method according to any one of claims 1-4.
7. A computing device, characterized in that, The computing device includes a processor and a memory: The memory is used to store program codes and transmit the program codes to the processor; The processor is used to execute the method according to any one of claims 1-4 according to the instructions in the program codes.
Citation Information
Patent Citations
Fire prediction method, XGBoost model training method and related equipment
CN113673748A
Control method and device applied to multiple scenes and based on fire prediction information
CN113869567A
Training method of convolutional neural network, and image processing method and apparatus
CN108304921A
Fire detection method based on improved YOLOV5
CN114821423A